Lesson:research autopilot 20260720t030001z-gpu
| 제목 | Research findings 20260720T030001Z-gpu: 1 negative/inconclusive, 0 mixed, 1 positive |
|---|---|
| 궁금했던 점 | What did the validated experiments or analyses establish, including useful negative results and the conditions under which they apply? |
| 해본 것 | - mlir-fft-compiler [scientific outcome=inconclusive]: Scientific outcome=inconclusive. GPU3 고아 프로세스를 정리하고 래퍼 점유·cuda:0 가시성을 확인했다. 기존 정사각 CUDA 하네스는 5.3MB nvcc 컴파일 중 목표 직사각형/타일 코드를 포함하지 않아 중단했다. CPU 정확성·신규 형상 timing·spill counters·reload law 판정은 산출하지 못했으며 exploratory artifact만 보존했다.; validation: The implied reload count local_ld/(4*excess) is measured at K=32, K=128, and K=256, decisively separating reloads = rho*K from reloads = const (they differ ~4x at K=32, J=128); Every lowering passes an independent CPU W@x correctness check (max_abs_error well under 5e-10) before its timing is admitted; The spill predicate 8+2*MAXLIVE > 255 is scored against nvcc's static spill bytes at all eight new (shape, lowering) points; any disagreement is reported rather than smoothed; Predicted vs measur…; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit 79af53d2e524fe940e847f277ca2e347555cfd96; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1; GPU handoff 20260720T030001Z-mlir-fft-compiler
- multi-lora-fusion [scientific outcome=passed]: Scientific outcome=passed. L40S 물리 GPU3(cuda:0)에서 격리 스냅샷 3dd704c8로 두 실험을 직렬 수행했다. boundary_ordering_cuda.json: LLC 60.0MiB, working-set cap 75(근접 shape; WS 37.4MiB knee), output cap 239, 두 cap 순서 역전 0/865; WS 경계 0.261배 graded rise, output 경계 0.791배 step. output-tail 2-boundary 상대오차 평균/최대 0%/0%, single-knee 3.35%/3.35%. transition_form_cuda.json: N=384 resident→448 spilled에서 output/LLC 0.400→0.467, 1.638배 jump; single-knee far-tail 평균/최대 오차 35.8~43.1%/45.0~55.6%, H_single 기각, H_two 지지(조합 모델 평균 38.1%, 출력 knee는 예상 N=959와 불일치). GPU3 check 후 유휴 확인. 아티팩트 두 개를 .research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/에 보존.; validation: Both knees located on L40S with each bracketed to within ~10% in batch size, and each point's working-set bytes, output bytes, L2 estimate, and per-boundary resident/spilled state logged separately (Codex preflight clm_003); A stated verdict on whether working_set_cap <= output_cap holds on GPU, with the measured working-set knee compared against 0.5*L2 -- the quantity that decides whether the ordering can invert; A stated verdict on whether the GPU working-set transition is graded (a rise over…; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit 3dd704c833ccf69c1dfc5ba0c1fa2f733d32e94e; PR https://github.com/mrcha033/multi-lora-fusion/pull/1; GPU handoff 20260720T030001Z-multi-lora-fusion |
| 당시 조건 | Only completed, evidence-backed research findings are included. Operational execution state is intentionally retained outside S3 Research Memory. |
| 실제 결과 | mlir-fft-compiler [scientific outcome=inconclusive]: Scientific outcome=inconclusive. GPU3 고아 프로세스를 정리하고 래퍼 점유·cuda:0 가시성을 확인했다. 기존 정사각 CUDA 하네스는 5.3MB nvcc 컴파일 중 목표 직사각형/타일 코드를 포함하지 않아 중단했다. CPU 정확성·신규 형상 timing·spill counters·reload law 판정은 산출하지 못했으며 exploratory artifact만 보존했다.; validation: The implied reload count local_ld/(4*excess) is measured at K=32, K=128, and K=256, decisively separating reloads = rho*K from reloads = const (they differ ~4x at K=32, J=128); Every lowering passes an independent CPU W@x correctness check (max_abs_error well under 5e-10) before its timing is admitted; The spill predicate 8+2*MAXLIVE > 255 is scored against nvcc's static spill bytes at all eight new (shape, lowering) points; any disagreement is reported rather than smoothed; Predicted vs measur…; commit 79af53d2e524fe940e847f277ca2e347555cfd96; PR https://github.com/mrcha033/mlir-fft-compiler/pull/1; GPU handoff 20260720T030001Z-mlir-fft-compiler
multi-lora-fusion [scientific outcome=passed]: Scientific outcome=passed. L40S 물리 GPU3(cuda:0)에서 격리 스냅샷 3dd704c8로 두 실험을 직렬 수행했다. boundary_ordering_cuda.json: LLC 60.0MiB, working-set cap 75(근접 shape; WS 37.4MiB knee), output cap 239, 두 cap 순서 역전 0/865; WS 경계 0.261배 graded rise, output 경계 0.791배 step. output-tail 2-boundary 상대오차 평균/최대 0%/0%, single-knee 3.35%/3.35%. transition_form_cuda.json: N=384 resident→448 spilled에서 output/LLC 0.400→0.467, 1.638배 jump; single-knee far-tail 평균/최대 오차 35.8~43.1%/45.0~55.6%, H_single 기각, H_two 지지(조합 모델 평균 38.1%, 출력 knee는 예상 N=959와 불일치). GPU3 check 후 유휴 확인. 아티팩트 두 개를 .research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/에 보존.; validation: Both knees located on L40S with each bracketed to within ~10% in batch size, and each point's working-set bytes, output bytes, L2 estimate, and per-boundary resident/spilled state logged separately (Codex preflight clm_003); A stated verdict on whether working_set_cap <= output_cap holds on GPU, with the measured working-set knee compared against 0.5*L2 -- the quantity that decides whether the ordering can invert; A stated verdict on whether the GPU working-set transition is graded (a rise over…; commit 3dd704c833ccf69c1dfc5ba0c1fa2f733d32e94e; PR https://github.com/mrcha033/multi-lora-fusion/pull/1; GPU handoff 20260720T030001Z-multi-lora-fusion |
| 왜 그랬는지 | These are evidence-backed scientific outcomes. Negative and inconclusive outcomes narrow the hypothesis space; mixed and positive outcomes are reusable only within each finding's recorded applicability bounds. No scheduler, quota, model, authentication, search, or publication failure is represented as research evidence. |
| 다음에 기억할 것 | mlir-fft-compiler: Do not repeat this experiment without a changed hypothesis; reuse the measured scientific verdict and device-specific bounds. multi-lora-fusion: Do not repeat this experiment without a changed hypothesis; reuse the measured scientific verdict and device-specific bounds. |
| 언제 맞는지 | mlir-fft-compiler: Only the recorded shapes, dtypes, software revision, and physical L40S GPU 3.
multi-lora-fusion: Only the recorded shapes, dtypes, software revision, and physical L40S GPU 3. |
| 신뢰도 | 중간 |
| 관련 자료 | .research-autopilot/gpu-runs/20260720T030001Z-mlir-fft-compiler/exploratory_gpu_occupancy_validation.json; mlir-fft-compiler commit 79af53d2e524fe940e847f277ca2e347555cfd96; https://github.com/mrcha033/mlir-fft-compiler/pull/1; .research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/transition_form_cuda.json; multi-lora-fusion commit 3dd704c833ccf69c1dfc5ba0c1fa2f733d32e94e; https://github.com/mrcha033/multi-lora-fusion/pull/1 |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-20T04:31:38.450503Z |
| 마지막 수정 시각 (UTC) | 2026-07-20T04:31:38.720514Z |
근거 research-artifact-4f3bed59c83f9d42: .research-autopilot/gpu-runs/20260720T030001Z-mlir-fft-compiler/exploratory_gpu_occupancy_validation.json
벤치마크 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-20T04:31:38.450503Z
근거 commit_79af53d2e524fe94: GitHub mrcha033/mlir-fft-compiler commit 79af53d2e524fe940e847f277ca2e347555cfd96
(원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-20T04:31:38.720514Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.