Lesson:research autopilot 20260720t030001z gpu
| 제목 | GPU follow-up 20260720T030001Z: 1 inconclusive, 1 positive |
|---|---|
| 궁금했던 점 | What did the validated GPU3 experiments establish, and what remains unresolved? |
| 해본 것 | - mlir-fft-compiler: safely inspected and stopped the legacy square-only compile; no new rectangular/tiled timing artifact. Exploratory artifact preserved.
- multi-lora-fusion: ran boundary_ordering and transition_form serially on L40S physical GPU3 via gpu3_exec.py using an isolated commit snapshot. |
| 당시 조건 | Codex-owned GPU3 follow-up. Physical GPU index 3 was isolated by the wrapper and appeared inside processes as cuda:0. Canonical worktrees were left untouched. |
| 실제 결과 | mlir-fft-compiler outcome inconclusive: expected rectangular/tiled artifact was not produced; exploratory occupancy artifact only.
multi-lora-fusion outcome passed: LLC 60 MiB; 865 scorable shapes had zero cap-order inversions; close profile had WS cap 75 and output cap 239, with 0.261x graded WS rise and 0.791x output step. Far-tail transition showed N=384 resident to N=448 spilled, output/LLC 0.400 to 0.467, 1.638x jump. Single-knee max relative error 45.0–55.6%; two-boundary mean/max relative error 38.1%/48.4% on the far tail. |
| 왜 그랬는지 | The multi-LoRA GPU evidence supports sequential two-boundary behavior and rejects a single-knee model for this far tail. The MLIR rectangular/reload-law question remains open because the legacy harness did not contain the requested shapes; no scientific claim is made from the exploratory artifact. |
| 다음에 기억할 것 | For cache-capacity models, measure both nested tensor boundaries and score held-out tails; a sharp output-tensor transition can coexist with a graded working-set rise, while an unmodified square-only harness cannot answer rectangular reload-law claims. |
| 언제 맞는지 | Only the recorded repositories/commits, fp32 multi-LoRA GPU3 run, and exploratory MLIR compile environment. Timings are device- and shape-specific. |
| 신뢰도 | 중간 |
| 관련 자료 | /home/mrcha033/Researches/.research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/boundary_ordering_cuda.json; /home/mrcha033/Researches/.research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/transition_form_cuda.json; /home/mrcha033/Researches/.research-autopilot/gpu-runs/20260720T030001Z-mlir-fft-compiler/exploratory_gpu_occupancy_validation.json |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-20T04:30:23.461495Z |
| 마지막 수정 시각 (UTC) | 2026-07-20T04:30:35.127104Z |
근거 research-artifact-20260720t030001z-gpu: Scheduled GPU3 artifacts for cycle 20260720T030001Z (local paths listed in evidence field)
벤치마크 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-20T04:30:23.461495Z
근거 gpu3-boundary-ordering-20260720: boundary_ordering_cuda.json artifact from L40S GPU3 run
벤치마크 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-20T04:30:33.706360Z
Local artifact: /home/mrcha033/Researches/.research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/boundary_ordering_cuda.json
근거 gpu3-transition-form-20260720: transition_form_cuda.json artifact from L40S GPU3 run
벤치마크 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-20T04:30:35.127104Z
Local artifact: /home/mrcha033/Researches/.research-autopilot/gpu-runs/20260720T030001Z-multi-lora-fusion/transition_form_cuda.json