본문으로 이동

속성으로 검색

이 문서는 속성과 이름이 지정된 값으로 설명된 개체를 찾기 위한 간단한 탐색 인터페이스를 제공합니다. 그 밖에 사용 가능한 검색 인터페이스에는 문서 속성 검색Ask 쿼리 빌더가 있습니다.

속성으로 검색

"not_verifiable" 값의 "Observation" 속성을 가진 모든 문서의 목록입니다. 결과가 얼마 안 되기 때문에 주변의 값을 표시합니다.

1번 부터의 결과 28개입니다.

(이전 50개 | 다음 50개) (20 | 50 | 100 | 250 | 500) 보기


    

결과 목록

  • Lesson:identifying on off cpu bottlenecks together with blocked samples 1f04444a  + (bperf의 평균 성능 저하는 1.6%로 perf sampling 0.9%, tracing 3.6%와 비교해 perf보다 0.7%p 추가 비용. BCOZ end-to-end overhead는 평균 27.6%, 최대 64.7%. RocksDB 혼합 워크로드에서 병목을 식별했고 실제 최적화의 방향/순위가 가상 speedup 예측과 대체로 일치.)
  • Lesson:technical review development of behavior profilers for multimedia consumer electronics f30d374c  + (commodity digital TV의 실제 최적화에 사용해 tool set의 효과를 검증했다. 공식 초록에는 overhead 또는 최적화 폭의 수치가 없다.)
  • Lesson:development of behavior profilers for multimedia consumer electronics 6a76ba00  + (commodity digital TV의 실제 최적화에 사용해 tool set의 효과를 검증했다. 공식 초록에는 overhead 또는 최적화 폭의 수치가 없다.)
  • Lesson:research autopilot 20260722t210001z  + (complex-nn-signal [scientific outcome=posicomplex-nn-signal [scientific outcome=positive]: In a synthetic chirp/Doppler classification task (K=4, L=16), a learnable freq+rate chirplet correlator bank (immutable cells, per-cell hard max, no learned head) fails under plain SGD to null its chirp-rate degrees of freedom on constant-frequency data, costing ~0.09-0.10 accuracy vs the constant-frequency matched-filter scan at rho=0. Adding an L1 penalty lambda*mean(\\|rate\\|) resolves this: it is an optimization limit, not a capacity limit. A single fixed lambda=3.0 acts as a data-adaptive gate, driving mean\\|rate\\| to 0.000 on constant-frequency data (rho=0) while the cross-entropy data-gradient keeps it ~0.12 on chirped data (rho=0.3). This recovers matched-filter accuracy at rho=0…; validation: pytest tests/ -> 298 passed (292 prior + 6 new), CPU via .venv (uv py3.12, torch 2.13.0), 8.2 s; python -m src.chirp_rate_gate_benchmark -> results_rate_gate.json in 334 s CPU; 640 chirplet trainings over B{16,32} x rho{0,0.1,0.2,0.3} x SNR{-5,0} x lambda{0,0.3,1.0,3.0} x seeds 310..319; Adaptivity: lambda=3.0 gives mean\\|rate\\|@rho0=0.000 vs @rho0.3=0.120 (B32/SNR=0) and 0.000 vs 0.116 (B16/SNR=0); freq spread stays 0.49-0.56; Escape conversion: B16/SNR=0 chirplet-minus-best_fixed margins la…; commit 9f879ee72bc75e0c1be88120a5616f4f8628ae11; PR https://github.com/mrcha033/complex-nn-signal/pull/1)
  • Lesson:technical review nap natural app processing for predictive user contexts in mobile smartphones df34aff8  + (context 미사용 평균 Recall@1/2/3/4/5는 42.79/59.context 미사용 평균 Recall@1/2/3/4/5는 42.79/59.67/69.40/75.20/78.90%; context 사용 시 40.36/57.18/66.93/73.11/77.13%로 오히려 하락.</br>최고 월의 bidirectional model Recall@1은 50.79%, Recall@5는 86.55%; NAP 86.42%, AppUsage2Vec 85.93%, FALCON 80% 등과 비교.</br>긴 history보다 소수의 직전 앱이 더 유용한 경향을 보고.등과 비교. 긴 history보다 소수의 직전 앱이 더 유용한 경향을 보고.)
  • Lesson:nap natural app processing for predictive user contexts in mobile smartphones a3c20a54  + (context 미사용 평균 Recall@1/2/3/4/5는 42.79/59.context 미사용 평균 Recall@1/2/3/4/5는 42.79/59.67/69.40/75.20/78.90%; context 사용 시 40.36/57.18/66.93/73.11/77.13%로 오히려 하락.</br>최고 월의 bidirectional model Recall@1은 50.79%, Recall@5는 86.55%; NAP 86.42%, AppUsage2Vec 85.93%, FALCON 80% 등과 비교.</br>긴 history보다 소수의 직전 앱이 더 유용한 경향을 보고.등과 비교. 긴 history보다 소수의 직전 앱이 더 유용한 경향을 보고.)
  • Lesson:technical review a case for hardware based demand paging 163a61b2  + (cycle-level simulator와 ultra-low-latency SSD를 장착한 실제 x86 평가에서 demand-paging latency 37.0% 감소. FIO random-read 최대 57.1%, NoSQL server 최대 27.3% 성능 향상. OS 개입 감소의 부수 효과로 user-code IPC 최대 7.0% 증가.)
  • Lesson:a case for hardware based demand paging bdf187f5  + (cycle-level simulator와 ultra-low-latency SSD를 장착한 실제 x86 평가에서 demand-paging latency 37.0% 감소. FIO random-read 최대 57.1%, NoSQL server 최대 27.3% 성능 향상. OS 개입 감소의 부수 효과로 user-code IPC 최대 7.0% 증가.)
  • Lesson:technical review scheduler support for video oriented multimedia on client side virtualization 69d0c186  + (frame-rate 추정오차는 VLC HighRes 0.79%, Quake3 3.05%, 간섭 중 VLC LowRes 0.55%로 보고됐다. CPU-intensive 경쟁 VM이 있어도 VLC HighRes가 최대 23.976 FPS에 가까운 평균과 낮은 분산을 유지했고, Quake3에도 적절한 share를 배정했다.)
  • Lesson:scheduler support for video oriented multimedia on client side virtualization c8366be4  + (frame-rate 추정오차는 VLC HighRes 0.79%, Quake3 3.05%, 간섭 중 VLC LowRes 0.55%로 보고됐다. CPU-intensive 경쟁 VM이 있어도 VLC HighRes가 최대 23.976 FPS에 가까운 평균과 낮은 분산을 유지했고, Quake3에도 적절한 share를 배정했다.)
  • Lesson:technical review kal kernel assisted non invasive memory leak tolerance with a general purpose m e42adf12  + (glibc/Linux prototype에서 throughput과 평균 response time 기준 overhead가 conventional allocator 대비 약 2%였고, synthetic leak workloads에서 address-space expansion을 억제했다.)
  • Lesson:kal kernel assisted non invasive memory leak tolerance with a general purpose memory allocator 372cd042  + (glibc/Linux prototype에서 throughput과 평균 response time 기준 overhead가 conventional allocator 대비 약 2%였고, synthetic leak workloads에서 address-space expansion을 억제했다.)
  • Lesson:technical review efficient hybrid polling for ultra low latency storage devices 545391d6  + (light load에서 기존 hybrid polling 대비 CPU 사용 5~40% 절감하면서 classic polling에 가까운 latency 유지. heavy load에서 CPU 5~30% 절감, I/O latency 최대 10% 감소.)
  • Lesson:efficient hybrid polling for ultra low latency storage devices e753019a  + (light load에서 기존 hybrid polling 대비 CPU 사용 5~40% 절감하면서 classic polling에 가까운 latency 유지. heavy load에서 CPU 5~30% 절감, I/O latency 최대 10% 감소.)
  • Lesson:technical review 56ee25b0  + (migration 중 수 초 동안 전력이 높은 상태로 유지되거나 오히려 증가할 수 있어 즉시적인 power-cap 강제 수단으로는 부적합할 수 있음을 보였다.)
  • Lesson:lesson 0960532f  + (migration 중 수 초 동안 전력이 높은 상태로 유지되거나 오히려 증가할 수 있어 즉시적인 power-cap 강제 수단으로는 부적합할 수 있음을 보였다.)
  • Lesson:technical review analysis of virtual machine live migration as a method for power capping b973b1aa  + (migration 중 수 초 동안 전력이 높은 상태를 유지하거나 증가할 수 있음을 보였고, 제안한 두 방법은 migration 시작 직후 전력을 제한했다.)
  • Lesson:analysis of virtual machine live migration as a method for power capping 71088a92  + (migration 중 수 초 동안 전력이 높은 상태를 유지하거나 증가할 수 있음을 보였고, 제안한 두 방법은 migration 시작 직후 전력을 제한했다.)
  • Lesson:research autopilot 20260720t030001z-gpu  + (mlir-fft-compiler [scientific outcome=incomlir-fft-compiler [scientific outcome=inconclusive]: Scientific outcome=inconclusive. GPU3 고아 프로세스를 정리하고 래퍼 점유·cuda:0 가시성을 확인했다. 기존 정사각 CUDA 하네스는 5.3MB nvcc 컴파일 중 목표 직사각형/타일 코드를 포함하지 않아 중단했다. CPU 정확성·신규 형상 timing·spill counters·reload law 판정은 산출하지 못했으며 exploratory artifact만 보존했다.; validation: The implied reload count local_ld/(4*excess) is measured at K=32, K=128, and K=256, decisively separating reloads = rho*K from reloads = const (they differ ~4x at K=32, J=128); Every lowering passes an independent CPU W@x correctness check (max_abs_error well under 5e-10) before its timing is admitted; The spill predicate 8+2*MAXLIVE > 255 is scored against nvcc's static spill bytes at all eight new (shape, lowering) points; any disagreement is reported rather than smoothed; Predicted vs measur…; commit 79af53d2e524fe940e847f277ca2e347555cfd96; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1; GPU handoff 20260720T030001Z-mlir-fft-compiler</br>multi-lora-fusion [scientific outcome=passed]: Scientific outcome=passed. L40S 물리 GPU3(cuda:0)에서 격리 스냅샷 3dd704c8로 두 실험을 직렬 수행했다. boundary_ordering_cuda.json: LLC 60.0MiB, working-set cap 75(근접 shape; WS 37.4MiB knee), output cap 239, 두 cap 순서 역전 0/865; WS 경계 0.261배 graded rise, output 경계 0.791배 step. output-tail 2-boundary 상대오차 평균/최대 0%/0%, single-knee 3.35%/3.35%. transition_form_cuda.json: N=384 resident→448 spilled에서 output/LLC 0.400→0.467, 1.638배 jump; single-knee far-tail 평균/최대 오차 35.8~43.1%/45.0~55.6%, H_single 기각, H_two 지지(조합 모델 평균 38.1%, 출력 knee는 예상 N=959와 불일치). GPU3 check 후 유휴 확인. 아티팩트 두 개를 .research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/에 보존.; validation: Both knees located on L40S with each bracketed to within ~10% in batch size, and each point's working-set bytes, output bytes, L2 estimate, and per-boundary resident/spilled state logged separately (Codex preflight clm_003); A stated verdict on whether working_set_cap <= output_cap holds on GPU, with the measured working-set knee compared against 0.5*L2 -- the quantity that decides whether the ordering can invert; A stated verdict on whether the GPU working-set transition is graded (a rise over…; commit 3dd704c833ccf69c1dfc5ba0c1fa2f733d32e94e; PR https://github.com/mrcha033/multi-lora-fusion/pull/1; GPU handoff 20260720T030001Z-multi-lora-fusion001Z-multi-lora-fusion)
  • Lesson:research autopilot 20260721t030001z  + (mlir-fft-compiler [scientific outcome=posimlir-fft-compiler [scientific outcome=positive]: When validating a cost-model rate law whose form is degenerate at the calibration point (here spill reloads=0.685*K vs constant=87.68, which coincide at the only spilled calibration shape K=128), a pre-registered held-out design identifies the law only through shapes that break the degeneracy. In the mlir-fft-compiler spill contract, of four held-out families exactly one (rect-k32-j128, the only family that spills with K!=128) separates the two laws: 4.0x in the local-load NCU counter and 39.3pp in derived timing advantage vs a 5pp tolerance. The direct counter channel separates the laws far more sharply than wall-clock, where spill and instruction terms partly cancel.; validation: 1,232 tests pass (7 new) via .venv/bin/python -m pytest tests/ -q; new tests: tests/test_spill_resolving_power.py (3) and 4 in tests/test_spill_validation_contract.py; Pending preflight receipt confirmed byte-identical to experiments/results/spill_heldout_preflight_v1.json after the change; contract_sha256=c8b3bce... unchanged; git diff --stat over experiments/spill_heldout_contract_v1.json, src/spill_model.py, and the preflight receipt is empty (frozen files untouched); verify_spill_heldout --…; commit a0a464ea5a5f4f353574da7cd632cb6fcf5241fa; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1/mrcha033/mlir-fft-compiler/pull/1)
  • Lesson:research autopilot 20260719t210001z-gpu  + (mlir-fft-compiler [scientific outcome=mixemlir-fft-compiler [scientific outcome=mixed]: Scientific outcome=mixed. L40S physical GPU 3에서 complex128 dense contraction을 SM 포화 grid로 검증했다. 네 lowering 모두 max abs error 4.89e-15 이하로 통과했다. K=J=32는 real_fold 6.919 ns/problem, complex_spec 5.352 ns/problem으로 Gauss가 22.65% 빨라 예측 -55.6%의 부호를 반박했다. K=J=128은 117.700 대 143.690 ns/problem으로 Gauss가 22.08% 느려 예측 +24.9%를 반박했다. N=32에서는 register 164/252, occupancy 19.87/14.34%이나 spill이 없고 instruction 감소가 우세했다. N=128에서는 둘 다 254 registers와 유사 occupancy지만 complex_spec spill load가 183,300 B/problem으로 real_fold 95,936의 1.91배, DRAM은 3.10배여서 누락된 spill-traffic 항이 패배 원인으로 확인됐다.; validation: Both lowerings produce numerically correct results at both shapes (max-abs error vs W @ x within fp64 tolerance) — no timing is reported for a kernel that does not compute the contraction; Measured registers/thread is reported alongside the model's prediction (baseline 8 + 2*MAXLIVE: ~140 vs ~204 at n=32, capped at 255 for both at n=128), so the register model is checked independently of the timing result; The sign of the measured time advantage at K = J = 32 is reported: negative confirms the…; commit 43e787be17ee558a95a44185fc17b758f0abacbb; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1; GPU handoff 20260719T210001Z-mlir-fft-compiler719T210001Z-mlir-fft-compiler)
  • Lesson:research autopilot 20260720t030001z  + (mlir-fft-compiler [scientific outcome=posimlir-fft-compiler [scientific outcome=positive]: On GPU lowering of dense complex contractions, register pressure acts as a threshold at the architectural register cap, not as a continuous occupancy penalty. Re-analysis of four L40S (sm_89, complex128) configurations shows achieved occupancy spans only 5.53 pp (14.34-19.87%) and picks the faster lowering at 0 of 2 shapes -- it is anti-correlated with runtime, so an occupancy-knee cost model cannot be rescued by recalibration. The discriminator is spilling past the 255-register cap: below it the Gauss-amortized lowering wins on issue slots alone ((K-1)/(4K), with measured dynamic instructions within 2.5% of the closed forms 3KJ+J and 4KJ); above it, excess registers convert arithmetic savi…; validation: Full suite: 1205 tests pass (105 new in tests/test_spill_model.py), 0 failures; H5b spill predicate agrees with nvcc's static spill bytes at 4/4 measured hardware points, zero fitted parameters; H5c implied reload counts 87.29 (complex_spec) vs 89.16 (real_fold) agree to 2.15% across two different schedules; predicted spill bytes within 1.7% of measured (184,128 vs 183,300; 94,344 vs 95,936); H5d spill-aware policy reproduces 2/2 measured signs, max advantage error 1.57 pp; superseded occupancy…; commit 79af53d2e524fe940e847f277ca2e347555cfd96; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1; GPU handoff 20260720T030001Z-mlir-fft-compiler</br>multi-lora-fusion [scientific outcome=positive]: When a fused GEMM's cost model has two cache-capacity boundaries defined on nested tensors -- a total-working-set knee and an output-tensor knee -- the boundaries are strictly ordered rather than competing, and the ordering is provable rather than empirical. Because the output tensor Y is contained in the working set, working_set_bytes >= output_bytes per request for every shape, which forces working_set_cap/output_cap <= knee/(usable_fraction*cache). Whenever the fitted knee lies below the usable cache the working-set boundary binds first for ALL shapes, so a min() over the two hard caps is degenerate and the boundaries compose sequentially. The ordering is shape-invariant but the boundari…; validation: Full suite: 627 tests pass (up from 526), stable across 3 consecutive full-suite runs after the flaky-test fix; experiments/resolve_boundary_ordering.py run 3 times; headline numbers reproduce (output step 1.90-2.03x, two-boundary tail error 0.8-1.2% vs single-knee 16.4-17.3%); 900-shape analytic grid: 0 ordering inversions, max continuous cap ratio 0.704 <= analytic bound 0.710; Measured close-shape profile (N=4..256, 25 repeats, median, 1 thread): 1.07x graded rise at working-set cap, 1.98x s…; commit 3dd704c833ccf69c1dfc5ba0c1fa2f733d32e94e; PR https://github.com/mrcha033/multi-lora-fusion/pull/1; GPU handoff 20260720T030001Z-multi-lora-fusion</br>spectral-operator-compiler [scientific outcome=positive]: Two optimizations that each eliminate a disjoint stage of the same operator compose with efficiency (1-f_a)(1-f_b)/(1-f_a-f_b) = 1 + f_a*f_b/(1-f_a-f_b), which is strictly greater than 1 — they super-compose, and the naive multiplicative prediction s_a*s_b is a strict lower bound rather than an estimate. The mechanism is mutual Amdahl masking: measured in isolation, each lever's speedup is capped by the stage the other lever would have removed, so both solo measurements understate the pair, and the shortfall grows with f_a*f_b — i.e. it is worst exactly when both stages are large and the combination is most valuable. Measured on an FNO spectral-conv forward: a weight-layout lowering (1.54-1…; validation: Full test suite: 371 passed (330 prior + 41 new), 1 unrelated PyTorch complex-module warning; All four factorial cells bit-identical at every shape (max_abs_err = 0.0), enforced in-harness by assert_cells_equivalent; benchmarks/bench_lever_composition.py run twice independently; headline structure reproduced (s_both 7.84-8.29x at Cin=512, 4.06-5.05x at Cin=256); ranges reported rather than single figures; Closed form 1 + f_a*f_b/(1-f_a-f_b) verified symbolically via sympy and encoded with tests…; commit a2a50fb1f14b3cafb34c1d2ac37ffbb8f5eb63f0; PR https://github.com/mrcha033/spectral-operator-compiler/pull/1</br>openevolve-moe-prototype [scientific outcome=mixed]: In LLM-driven evolutionary program search, sibling programs sharing a parent agree on outcome far more than unrelated programs from the same task (67.6% vs 38.4% pair concordance), and this clustering is what suppresses DPO-pair yield. The cause is NOT near-duplicate sampling: siblings are measurably more textually similar than unrelated candidates (0.846 vs 0.683, p = 5e-05), yet code similarity does not predict outcome agreement among unrelated pairs at all (rho = +0.019, p = 0.56), and matching controls on similarity absorbs only 1.8% of the sibling excess. Matching instead on the parent's normalized position in the task's score range absorbs 74% (excess +0.290 -> +0.075, n.s.). Decompos…; validation: Full test suite: 441 passed in 27.4s (.venv_test/bin/python -m pytest tests/ -q); 49 new tests in tests/test_sibling_concordance.py, up from 392 total.; Known-answer checks on every statistic: matched_excess recovers a planted +1.0 difference, returns exactly 0.0 under a true null, and its permutation p is <0.01 for the planted effect and >0.2 under the null.; Covariate-confound rejection test: on a pool where siblings and controls differ only in q-stratum, the unmatched estimator reports exces…; commit 11a77d13a64a2b8e5f7955f676071034a4d326a0; PR https://github.com/mrcha033/openevolve-moe-prototype/pull/1l/1)
  • Lesson:research autopilot 20260719t210001z  + (mlir-fft-compiler [scientific outcome=posimlir-fft-compiler [scientific outcome=positive]: For a dense complex contraction y = W·x lowered to real arithmetic, the register cost and the instruction benefit of the Gauss/Karatsuba 3-multiply amortization occupy independent axes: peak live values are flat in the reuse factor K (MAXLIVE = 2J+2 for the 4-multiply fold, 3J+2 for the amortized 3-multiply form, where J is the input count), because outputs are accumulated one at a time and never co-reside, while the instruction saving (K−1)/(4K) is independent of J. Consequently the compute-vs-occupancy decision that previously required emitting, scheduling and measuring the IR reduces to a closed-form O(1) rule in (K, J, target): over 900 points (100 shapes × 3 GPU targets × 3 bandwidth-s…; validation: `.venv/bin/python -m pytest tests/ -q` → 1100 passed (368 new in tests/test_lowering_policy.py); `.venv/bin/python -m experiments.lowering_policy_validation` → closed forms exact at all 100 shapes; 900/900 decision agreement with the emit-and-measure oracle; max \\|advantage error\\| 0.0 pp; Closed-form instructions/FLOPs/MAXLIVE compared against real scheduled instruction streams at every grid shape, for both strategies; `analytic_cost` compared field-by-field (instructions, flops, bytes, regs…; commit 43e787be17ee558a95a44185fc17b758f0abacbb; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1; GPU handoff 20260719T210001Z-mlir-fft-compiler</br>openevolve-moe-prototype [scientific outcome=mixed]: In LLM-driven evolutionary program search, the rate at which sibling candidates convert into DPO preference pairs is governed by two separable factors, and the intuitive one is not the binding one. First, child outcomes depend strongly on the parent's normalized position in its task's score range: the improvement rate falls with parent quality (Spearman rho = -0.57, p = 5e-05) while the no-op rate rises (+0.57, p = 5e-05), and the executable-regression rate that supplies a DPO negative is non-monotone, peaking in a mid-quality band (63.6% inside q in [0.4,0.6) vs 14.8% outside, p = 0.0013 post-hoc). This is a parent-quality effect and not a search-time effect: partial correlation for qualit…; validation: pytest tests/ — 392 passed in 2.6s (up from 339; 53 new tests), CPU-only, stdlib-only analysis; Reconciliation with A1d/A1e on every real working-harness task: identical pairable families, deficiency verdicts, slots, and dpo_pairs; idle ledger sums to the non-converting slot total (36); Known-answer statistics tests: Wilson intervals, rank-based Spearman under a nonlinear monotone map, partial correlation collapsing a pure confound to 0.0, a planted effect detected (p<0.01), a null not detected…; commit dec037ba93c203f580c290cb50d291515398acef; PR https://github.com/mrcha033/openevolve-moe-prototype/pull/1</br>spectral-operator-compiler [scientific outcome=positive]: For the FNO spectral contraction einsum("bim,iom->bom"), the operator's cost is dominated by its lowering, not its arithmetic. On CPU (torch 2.13.0, 16 threads) the einsum baseline achieves only 27-75 GFLOP/s, i.e. 6-10% of what the same backend reaches on the identical contraction lowered to a mode-batched torch.bmm over mode-leading operands (118-954 GFLOP/s): a 4.36-14.49x contraction speedup and 1.71-5.67x end-to-end on the full spectral-conv forward, bit-identical (max_abs_err = 0.0). This is an order of magnitude larger than the 4/3x ceiling of the Karatsuba/Gauss 3M identity that four prior milestones were pursuing on the same operator. A thread-scaling control identifies the mechani…; validation: Full test suite: 330 passed (279 prior + 51 new), CPU, `.venv/bin/python -m pytest tests/ -q`; `python -m benchmarks.bench_contraction_lowering` run end to end; wrote benchmarks/results/contraction_lowering_cpu.json; `python -m benchmarks.bench_spectral_stage_profile` re-run after fixing its missing best_of_time import; wrote benchmarks/results/spectral_stage_profile_cpu.json; Numerical equivalence asserted in-harness before every timed comparison (assert_lowerings_equivalent), and end-to-end m…; commit 7175ff452fb060e71c6ba6cefcc7d3e5a5c140f8; PR https://github.com/mrcha033/spectral-operator-compiler/pull/1ull/1)
  • Lesson:research autopilot 20260720t030001z gpu  + (mlir-fft-compiler outcome inconclusive: exmlir-fft-compiler outcome inconclusive: expected rectangular/tiled artifact was not produced; exploratory occupancy artifact only.</br>multi-lora-fusion outcome passed: LLC 60 MiB; 865 scorable shapes had zero cap-order inversions; close profile had WS cap 75 and output cap 239, with 0.261x graded WS rise and 0.791x output step. Far-tail transition showed N=384 resident to N=448 spilled, output/LLC 0.400 to 0.467, 1.638x jump. Single-knee max relative error 45.0–55.6%; two-boundary mean/max relative error 38.1%/48.4% on the far tail.elative error 38.1%/48.4% on the far tail.)
  • Lesson:research autopilot 20260719t150001z  + (multi-lora-fusion [scientific outcome=posimulti-lora-fusion [scientific outcome=positive]: For fused multi-LoRA GEMMs on CPU, the per-request-latency-vs-batch curve has TWO distinct capacity knees at the last-level cache, on two different tensors, not one graded transition: the total working set crosses the LLC first (a graded knee), and the output tensor Y crosses it second as a razor-sharp step at Y==LLC. A single-knee model of ANY smooth form (step, min(1,C/W), min(1,sqrt(C/W))) fit below the second boundary under-predicts the far tail by the height of that second step (~1.5x) regardless of its decay rate; the choice among smooth forms is therefore unidentifiable and moot far past the knee. A two-boundary model (graded working-set knee times an output-tensor step past Y==cache…; validation: pytest -q — 526 passed (514 -> 526, +12); python -m experiments.resolve_transition_form — second knee reproducibly at N 504->512, Y/L3 0.984->1.000, jump ~1.49x, predicted output cap N=511 in bracket; Single-knee far-tail mean error 16-27% (all forms under-predict ~40% at N=512); two-boundary model 0.8% mean / 2.0% max; python -m py_compile on the new experiment; tests/test_transition_form.py pins the finding to the committed JSON; commit b048fc0e78958e5c4a7b0b1a93a6cb9994b4eead; PR https://github.com/mrcha033/multi-lora-fusion/pull/1ithub.com/mrcha033/multi-lora-fusion/pull/1)
  • Lesson:research autopilot 20260722t090001z  + (multi-lora-fusion [scientific outcome=mixemulti-lora-fusion [scientific outcome=mixed]: For a memory-capacity cost knee fit from noisy CPU timing (LoRA fusion working-set knee), the governing quantity is total working-set bytes at a fixed cache fraction (~0.6-0.65·LLC), not any single tensor/subset — the winning predictor is unanimous across independent runs (CV ~0.16 vs ≥0.28 for every subset). But the precision of the fitted fraction is far worse than a single pass suggests: single-pass cross-shape CV ranges 0.07-0.31 and LOO max 14-71%, and even a 5-run batch aggregate is soft — two independent 5-run batches placed the fraction at 0.662±0.035 and 0.597±0.013, a batch-to-batch gap exceeding either batch's own across-run sd. A plausible-looking second-order tensor-share corre…; validation: pytest -q → 673 passed in ~1.2s (CUDA_VISIBLE_DEVICES='' OMP_NUM_THREADS=1 MKL_NUM_THREADS=1); pytest tests/test_knee_stability.py tests/test_knee_predictor.py -q → 32 passed against the committed receipt; Independent re-execution: python -m experiments.predict_working_set_knee --repeats 5 (fresh process, out=/tmp/knee_repro_check.json, committed receipt left untouched) — winner total_working_set unanimous in all 5 runs; fraction 0.597 ± 0.013 [0.573,0.613]; single-run CV 0.067–0.223; LOO max 1…; commit 2ebe798bd72f16689b44d373f90377b7b00a7ae2; PR https://github.com/mrcha033/multi-lora-fusion/pull/1sion/pull/1)
  • Lesson:scoz a systemwide causal profiler for multicore systems 7b5e72e2  + (multithread workload에서 COZ와 동일 병목 및 유사한 잠재 speedup을 찾음. Dbench/Filebench의 ext4 kernel 병목은 기존 OS scalability 연구와 일치. NAS FT에서 transblock=128(작업집합 256 KB, 실험 CPU L2와 동일) 최적화로 throughput 3.48% 향상했고 실제 speedup 추세가 SCOZ 가상 speedup과 대부분 겹침.)
  • Lesson:technical review scoz a system wide causal profiler for multicore systems 9ddaf48a  + (multithread workload에서 COZ와 동일 병목 및 유사한 잠재 speedup을 찾음. Dbench/Filebench의 ext4 kernel 병목은 기존 OS scalability 연구와 일치. NAS FT에서 transblock=128(작업집합 256 KB, 실험 CPU L2와 동일) 최적화로 throughput 3.48% 향상했고 실제 speedup 추세가 SCOZ 가상 speedup과 대부분 겹침.)
  • Lesson:performance optimization of object tracking algorithms in opencv on gpus 8158d2ec  + (object detection 전체 성능은 최대 86%, optical flow는 최대 10% 향상. Haar 최적화는 평균 wave throughput을 APU 67%, discrete GPU 77% 높였고 개별 전체 성능은 각각 최대 73%, 86% 향상. LBP 전체 성능은 APU 최대 31%, discrete GPU 최대 21% 향상.)
  • Lesson:technical review performance optimization of object tracking algorithms in opencv on gpus 09f20f75  + (object detection 전체 성능은 최대 86%, optical flow는 최대 10% 향상. Haar 최적화는 평균 wave throughput을 APU 67%, discrete GPU 77% 높였고 개별 전체 성능은 각각 최대 73%, 86% 향상. LBP 전체 성능은 APU 최대 31%, discrete GPU 최대 21% 향상.)
  • Lesson:technical review catching two rabbits adaptive real time support for embedded linux 07b12680  + (open-source benchmarks에서 기존 접근과 같거나 더 나은 scheduling latency를 유지하면서 throughput을 개선했다. 공식 초록에는 수치가 없다.)
  • Lesson:catching two rabbits adaptive real time support for embedded linux ee13b80e  + (open-source benchmarks에서 기존 접근과 같거나 더 나은 scheduling latency를 유지하면서 throughput을 개선했다. 공식 초록에는 수치가 없다.)
  • Lesson:technical review energy efficient scheduling of real time tasks on multicore processors cc0782cd  + (simulation에서 Dynamic Repartitioning은 당시 최선의 energy-efficient partitioning 대비 약 8%를 추가 절감했고, Dynamic Core Scaling은 low load에서 약 26%를 절감했다.)
  • Lesson:energy efficient scheduling of real time tasks on multicore processors a70b387a  + (simulation에서 Dynamic Repartitioning은 당시 최선의 energy-efficient partitioning 대비 약 8%를 추가 절감했고, Dynamic Core Scaling은 low load에서 약 26%를 절감했다.)
  • Lesson:research autopilot 20260726t090001z  + (sparse-lowrank-runtime [scientific outcomesparse-lowrank-runtime [scientific outcome=positive]: On the L40S FP32 torch.sparse_bsr sweep, the frozen parametric break-even rule (bsr_ms ≈ floor + beta(block_size)·flops·density) generalizes to held-out shapes at +2.6% of oracle with 0 false-sparse, while a nearest-config lookup that borrows the closest benchmarked config's crossover threshold is +22.2% out of sample — worse than always-dense (+15.0%) — and the only memorizer that dispatches into slowdowns (9 false-sparse). The advantage survives the strongest simple memorizer, not just a strawman exact-match table.; validation: .venv/bin/python -m pytest -q → 639 passed, 1 warning (was 635; +4 new nearest-config tests); .venv/bin/python -m pytest tests/test_generalization.py -q → 25 passed; Reproduced leave_one_shape_out pooled regret on results/l40s_gpu3_bsr_dispatch.json: parametric +2.6% (false_sparse=0), nearest_lookup +22.2% (false_sparse=9), exact-lookup/always_dense +15.0%, atlas +86.5%, always_sparse +142.3%; Confirmed leave_one_config_out gives identical nearest_lookup +22.2% (not an artifact of the shape-lev…; commit 969deb5a1c941e91cf0ec8e162b732e6ec7198e9; PR https://github.com/mrcha033/sparse-lowrank-runtime/pull/1</br>sparse-lowrank-runtime [scientific outcome=mixed]: A cross-validated dispatch-rule generalization headline (+2.6% out-of-sample vs +15% lookup) can be carried by a single decisive fold. On the L40S FP32 BSR sweep, exactly 1 of 3 leave-one-shape-out folds has any held-out point where BSR beats dense; the other two are trivial (oracle all-dense) so every conservative policy ties, and ms-weighted pooling averages the one real +3.2% with two +0.0% folds into +2.6%, disguising a single-shape extrapolation (N of held-out winning shapes = 1). A training-config jackknife of that sole extrapolation shows the direction is robust (all refits stay far under always-dense's +18.7%, spread 5.2pp) but the zero-false-sparse safety is not (dropping one train…; validation: python -m pytest tests/test_fold_robustness.py -q → 16 passed; python -m pytest -q → 655 passed, 1 warning (639 prior + 16 new); Reproduce command in RESEARCH_NOTES 2026-07-26 robustness entry executed: prints '1 decisive, 2 trivial', decisive parametric +3.2% vs always-dense +18.7%, jackknife spread 5.2pp with worst_false_sparse=1, win capture 6/9 wins 82.8% of savings; commit be11743facdb7f921c8fa8fdc9ec4a946ad24b8b; PR https://github.com/mrcha033/sparse-lowrank-runtime/pull/1pull/1)
  • Lesson:research autopilot 20260719t090001z  + (spectral-operator-compiler [scientific outspectral-operator-compiler [scientific outcome=negative]: The FNO spectral contraction 'bim,iom->bom' is not multiply-bound: its wall-clock cost ratio R = t_complex/t_real over the identically-shaped real contraction caps at ~2.6 (Cin=Cout=512) and never clears the R>3 threshold required for any Karatsuba/3-multiply implementation to yield a wall-clock win (upper bound min(R/3,4/3) ≤ 0.88 < 1). A same-BLAS square complex GEMM control clears R>3 (R→4.0 at n=2048), so the boundedness is a property of the FNO batched-mode shape (many small GEMMs), not the library.; validation: python -m pytest -q → 279 passed (264 prior + 15 new); python -m benchmarks.bench_multiply_boundedness → FNO R caps 2.6, GEMM control R→4.0; JSON artifact written; Re-ran benchmark 4x to confirm the channel-heavy row stabilizes at R≈2.6 (earlier 3.96 was noise) under best-of-9; commit 0bf1148b9cc8731097d190863f60685a2befd46c; PR https://github.com/mrcha033/spectral-operator-compiler/pull/133/spectral-operator-compiler/pull/1)
  • Lesson:technical review task aware virtual machine scheduling for i o performance 26a8da47  + (synthetic mixed workloads와 realistic applications에서 I/O responsiveness/throughput을 개선하면서 VM 간 CPU fairness를 유지했다. full text는 inspection-window와 correlation이 block throughput을 높이고 false boosting을 줄임을 보이지만, 공개 summary에 단일 headline speedup은 없다.)
  • Lesson:task aware virtual machine scheduling for i o performance f431ccbe  + (synthetic mixed workloads와 realistic applications에서 I/O responsiveness/throughput을 개선하면서 VM 간 CPU fairness를 유지했다. full text는 inspection-window와 correlation이 block throughput을 높이고 false boosting을 줄임을 보이지만, 공개 summary에 단일 headline speedup은 없다.)
  • Lesson:dwkv collinear slo confound  + (tier와 slo_budget 상관이 -0.93이었고, `value_urgetier와 slo_budget 상관이 -0.93이었고, `value_urgency`와 `tier_priority`가 해당 run의 모든 tier/metric에서 동일한 숫자를 냈다. 이 설계의 결과는 2026-07-11 retracted 처리되었다. deadline을 tier와 독립적으로 샘플링한 corrected simulation에서는 realistic regime의 best가 `length_aware`(8.15%)였고 DW-KV는 9.54%로 3위였으며, neutral에서도 length-aware가 앞섰다..54%로 3위였으며, neutral에서도 length-aware가 앞섰다.)
  • Lesson:sieve is simpler than lru an efficient turn key eviction algorithm for web caches 86f86293  + (workloads=1,559 traces from 7 sources; 5 pworkloads=1,559 traces from 7 sources; 5 production cache libraries; baselines=ARC; 9 state-of-the-art algorithms; optimized LRU; metrics=miss ratio; throughput; integration LOC; results=Up to 63.2% lower miss than ARC; 2× LRU throughput; ≤20 LOC integration.C; 2× LRU throughput; ≤20 LOC integration.)
  • Lesson:perseus a fail slow detection framework for cloud storage system 4e24f100  + (workloads=10 months, 248K drives; 41K normal and 315 verified fail-slow drives; baselines=existing fail-slow detectors; metrics=detection; node p99.99 latency; results=304 fail-slow drives found; isolation reduced p99.99 by 48%.)
  • Lesson:accl an fpga based collective engine for distributed applications 896d30fd  + (workloads=100 Gb/s FPGA cluster; CPU vectoworkloads=100 Gb/s FPGA cluster; CPU vector-matrix multiply; FPGA DLR inference; baselines=software MPI; software RDMA collectives; metrics=collective performance; application performance; results=Significant/competitive gains reported; exact aggregate not abstract-verified.ed; exact aggregate not abstract-verified.)
  • Lesson:pact a criticality first design for tiered memory 20db3445  + (workloads=13 graph, HPC, in-memory-cache, workloads=13 graph, HPC, in-memory-cache, and ML workloads; 96-workload model study; baselines=Soar, Alto, Memtis, Colloid, Nomad, TPP, Linux NBT; metrics=performance, migrations, model correlation; results=up to 61% faster; up to 50x fewer migrations; Pearson >0.98 up to 50x fewer migrations; Pearson >0.98)
  • Lesson:serverless in the wild characterizing and optimizing the serverless workload at a large cloud pr f2fa1121  + (workloads=14-day Azure Functions fleet trace; baselines=fixed keep-alive policies; metrics=cold starts; resource use; results=Significantly fewer cold starts with fewer resources; exact figure not abstract-verified.)
  • Lesson:mitigating application resource overload with targeted task cancellation 90270bf8  + (workloads=16 reproduced real-world overloaworkloads=16 reproduced real-world overload cases across MySQL, Apache, PostgreSQL, Elasticsearch, Solr, and etcd; baselines=non-overloaded execution, Protego, pBox, DARC, and PARTIES; metrics=normalized throughput, normalized p99 latency, request-drop rate, SLO attainment; results=Atropos sustains average normalized throughput 0.96 and average normalized p99 latency 1.16 while dropping fewer than 0.01% of requests. It meets the SLO in 14/16 cases; the reported multi-objective policy reduces normalized throughput by 10.2% relative to its performance-priority setting in the evaluated trade-off.iority setting in the evaluated trade-off.)
  • Lesson:kvcache cache in the wild characterizing and optimizing kvcache cache at a large cloud provider 0b39b53c  + (workloads=2024년 12월과 2025년 2월에 수집한 Aliyun workloads=2024년 12월과 2025년 2월에 수집한 Aliyun Tongyi production trace 두 세트(to-C Trace A, to-B Trace B; 논문은 각 trace의 대표 하루를 주로 분석)와 vLLM replay; baselines=무한-capacity ideal, LRU, LFU; metrics=KV-block hit ratio, reuse skew/time/lifespan, required cache capacity, mean response time; results=ideal hit ratio는 Trace A 62%, Trace B 54%; 상위 10% KV block이 reuse의 77%를 만들었고 to-B에서는 single-turn request가 cache hit의 97%를 만들었다. to-B KV의 P99 lifespan은 97초였으며 GPU HBM의 2배 cache로 common GQA model의 ideal hit rate에 근접했다. Workload-aware policy는 LRU/LFU 대비 cache hit를 3.9% 높이고 mean response time을 최대 41.4% 개선했다..9% 높이고 mean response time을 최대 41.4% 개선했다.)
  • Lesson:evendb optimizing key value storage for spatial locality 7beea059  + (workloads=256 GB production analytics dataset; YCSB without locality; baselines=RocksDB; metrics=ingestion throughput; write amplification; results=4.4× ingest and nearly 4× lower WAF; parity on no-locality YCSB.)
  • Lesson:bypassd enabling fast userspace access to shared ssds da9e5d79  + (workloads=4 KB I/O microbenchmarks and WiredTiger; baselines=Linux and SPDK; metrics=latency and application performance; results=42% lower 4 KB latency vs Linux; near SPDK; about 20% WiredTiger improvement)
  • Lesson:automatically reasoning about how systems code uses the cpu cache 11306b93  + (workloads=4 TCP stacks; 7 OpenSSL algorithworkloads=4 TCP stacks; 7 OpenSSL algorithms; 51 Hyperkernel syscalls; 2 hash tables; baselines=manual cache analysis; metrics=footprint; hits/misses; bug/vulnerability discovery; results=Found performance bugs, vulnerabilities, and third-party cache impacts; no aggregate numeric speedup in abstract. no aggregate numeric speedup in abstract.)
  • Lesson:scalable and effective page table and tlb management on numa systems 92f46a49  + (workloads=4- and 8-socket x86_64; Webserveworkloads=4- and 8-socket x86_64; Webserver; Memcached; memory-management microbenchmarks; baselines=baseline Linux; eager full page-table replication; metrics=runtime; TLB shootdowns; memory-management overhead; results=12% Webserver and 36% Memcached runtime improvement; up to 40× baseline overhead characterized.up to 40× baseline overhead characterized.)