본문으로 이동

Lesson:research autopilot 20260718t030144z-gpu: 두 판 사이의 차이

S3 연구 메모리
S3W1 k=e o=attach-gpu-artifact-bsr-20260719 r=4337677d0ac36874d0d8d569e1dd3590 b=2732 t=e62f31f7e8467ec9066aa422fa7d7672 h=7ee49fd505120ea20a558a4ce8384514
S3R2 o=strip-gpu-ops-framing-20260719-v2 r=9e1b1a037492c670484b18477834a87c b=2739 e=24ce7b021c54c6ca,3128852c93139fdb,4103fe515ef3a82b,727d8a2f95b7a576,7ed3feab1e79b976,f35da6092b1f9544 c=2ff t=d5da603e4e54a93b8c07946f5c5be1e7 h=a4f01e9dfaf3a873870ef16292e4e0b3; 전체 벤치마크 산출물을 검증했으므로 실제 GPU 연구 결과는 보존하고 Autopilot 실행 상태와 운영 메타데이터를 제거합니다.
 
(같은 사용자의 중간 판 6개는 보이지 않습니다)
1번째 줄: 1번째 줄:
{{Lesson
{{Lesson
|title=<nowiki>Research Autopilot 20260718T030144Z-gpu: 6 success, 0 failure (gpu_followup_completed)</nowiki>
|title=<nowiki>L40S GPU 검증 결과: 실패한 두 가설과 조건부로 유효한 네 실행 구간</nowiki>
|question=<nowiki>What succeeded or failed in scheduled research cycle 20260718T030144Z-gpu, and what should the next cycle reuse or avoid?</nowiki>
|question=<nowiki>제안된 컴파일러·런타임 최적화 중 물리 L40S GPU 3에서 수치 정확도와 지연시간 검증을 통과한 것은 무엇이며, 그 유효 범위는 어디까지인가?</nowiki>
|attempt=<nowiki>- algebraic-ml-compiler [success]: Scientific outcome=failed. 실험 완료(가설 실패): L40S GPU 3에서 Triton rebase가 3.76–7.82배 빨랐으나 65,536-key float32 최대 오차 1.41e-2로 2e-4 기준을 초과했다. 현 rewrite는 배포 금지.; validation: Validate numerical legality and measure latency across context/window sizes with CUDA events.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/algebraic-ml-compiler/experiments/l40s_gpu3_rebase_results.json; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit e5034b987d55d449dbb0ca296c24097aba49422b; PR https://github.com/mrcha033/algebraic-ml-compiler/pull/1; GPU handoff bootstrap-20260718-algebraic-ml-compiler
|attempt=<nowiki>동일한 L40S GPU에서 여섯 실험을 각각 실행했다: Triton rebase, FNO spectral convolution의 stacked 3M, fused multi-LoRA, RoPE training fusion, shape-adaptive attention proxy, CUDA BSR sparse-low-rank runtime. 각 실험은 전체 JSON 벤치마크 산출물로 수치 정확도와 성능을 함께 검증했다.</nowiki>
- fno-spectral-conv [success]: Scientific outcome=failed. 실험 완료(가설 실패): L40S GPU 3의 4개 FNO shape 모두 native complex einsum이 가장 빨랐다. 3M 수치오차는 6.4e-7 이하였지만 stacked 3M은 native의 0.28–0.37배 성능에 그쳐 GPU lowering 이점이 없었다.; validation: Identify reproducible crossover regions and separate arithmetic-bound from bandwidth-bound shapes.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/fno-spectral-conv/benchmarks/results/l40s_gpu3_3m_crossover.json; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit c60eb1889d87f68acbfc49e2b6607958cce5f565; PR https://github.com/mrcha033/fno-spectral-conv/pull/1; GPU handoff bootstrap-20260718-fno-spectral-conv
|context=<nowiki>물리 NVIDIA L40S GPU 3에서 수행한 여섯 개의 독립 검증 결과다. 결론은 측정한 shape, dtype, 구현 revision에 한정되며 다른 GPU나 입력 분포로 일반화하지 않는다.</nowiki>
- multi-lora-fusion [success]: Scientific outcome=mixed. 실험 완료(부분 성공): bf16 fused LoRA는 N=4–128에서 2.95–81.37배 빨랐지만 affine 모델은 R²=0.314, 속도향상 예측오차 35–61%였다. L40S에서는 batch 95→143 사이에 요청당 비용이 1.76배 뛰어 GPU 3 안전 상한을 95로 측정했다.; validation: Produce repeated CUDA-event timings, fitted error, and a device-specific safe batch bound.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/multi-lora-fusion/experiments/results/calibration_l40s_gpu3.json; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit 4414ed8a3c34b8f111204d934f3f920b06a37857; PR https://github.com/mrcha033/multi-lora-fusion/pull/1; GPU handoff bootstrap-20260718-multi-lora-fusion
|observation=<nowiki>1) Triton rebase는 3.76–7.82배 빨랐지만 모든 측정점에서 최대 절대 오차 기준 2e-4를 넘었고, 65,536 keys에서 1.41e-2였다. 2) FNO에서는 네 shape 모두 native complex einsum이 가장 빨랐다. stacked 3M의 output-scale 상대 오차는 6.37e-7 이하였으나 native 대비 속도는 0.28–0.37배였다. 3) fused LoRA는 N=4–128에서 2.95–81.37배 빨랐지만 affine 모델 R²=0.314, speedup 오차 35–61%였고 batch 95→143에서 요청당 비용이 1.76배 증가했다. 4) RoPE fusion의 bf16 forward/gradient 상대 오차는 각각 0.48%/0.58% 이하였고 forward는 1.05–1.22배 빨랐으나 seq_len 2048의 forward+backward는 0.74배였다. 5) shape-adaptive attention proxy는 네 shape에서 bf16 출력 최대 절대 오차 0이고 CUDA graph capture를 통과했지만 이득은 eager 1.04–1.11배, graph 1.02–1.06배였다. 6) CUDA BSR 63개 지점은 모두 수치적으로 유효했지만 크기 1024/2048은 측정 범위에서 dense를 이기지 못했다. 크기 4096의 crossover sparsity는 block 16/32/64에서 각각 86.2%/71.9%/60.7%였고 최대 speedup은 1.73/2.30/2.42배였다.</nowiki>
- rope-training-fusion [success]: Scientific outcome=mixed. 실험 완료(부분 성공, 1회 재시도): 저장소의 미정의 _FP32_GRAD_ACCUM 때문에 첫 시도 실패 후 실험 전용 주입으로 측정했다. bf16 전방/gradient 상대오차는 각각 0.48%/0.58% 이내. forward는 1.05–1.22배 빨랐지만 forward+backward는 긴 shape에서 0.74배로 느렸다.; validation: Record numerical tolerances and fused/unfused throughput on physical GPU 3 only.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/rope-training-fusion/results/l40s_gpu3_bf16_throughput.json; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit 5aa83bc9394c31efe234e1a00fdbe7772704623c; PR https://github.com/mrcha033/rope-training-fusion/pull/1; GPU handoff bootstrap-20260718-rope-training-fusion
|interpretation=<nowiki>Triton rebase와 측정한 FNO stacked 3M 가설은 배포 후보에서 제외한다. 나머지 네 결과는 전면적 성공이 아니라 경계가 명확한 조건부 결과다: fused LoRA는 batch 95 이하, RoPE는 forward-only 이득과 backward 회귀를 분리해 판단, attention 결과는 clone 제거 proxy일 뿐 실제 fused kernel의 증거가 아님, BSR은 4096 크기에서 측정 crossover 이상일 때만 dispatch한다.</nowiki>
- shape-adaptive-attention [success]: Scientific outcome=mixed. 실험 완료(제한적 성공): 동일 bf16 출력(최대오차 0)으로 4개 shape 모두 CUDA Graph capture 성공. proxy fusion 이득은 eager 1.04–1.11배, graph replay 1.02–1.06배로 줄지만 사라지지는 않았다. 단, 이는 실제 fused attention kernel이 아닌 clone 제거 proxy 결과다.; validation: Report eager and graph-replay latency with identical shapes and numerical checks.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/shape-adaptive-attention/results/l40s_cuda_graph_validation.json; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7; PR https://github.com/mrcha033/shape-adaptive-attention/pull/1; GPU handoff bootstrap-20260718-shape-adaptive-attention
|reusable_lesson=<nowiki>정확도 기준을 넘긴 Triton rebase는 속도 향상과 무관하게 거부한다. 측정한 FNO shape에서는 native complex einsum을 유지한다. 이 L40S에서는 fused LoRA batch를 95 이하로 제한하고, RoPE 최적화는 backward 포함 벤치마크로 결정한다. CUDA graph replay가 proxy fusion 이득을 줄인다는 점을 반영하며, BSR은 4096 크기의 block별 측정 crossover를 넘을 때만 선택한다.</nowiki>
- sparse-lowrank-runtime [success]: Scientific outcome=mixed. 실험 완료(조건부 성공): 63개 CUDA BSR 점 모두 수치 검증 통과. 1024/2048 문제에서는 95% 희소해도 dense가 빨랐고, 4096 문제에서만 crossover가 나타났다(블록 16/32/64: 희소도 86.2%/71.9%/60.7%, 최대 1.73/2.30/2.42배).; validation: Report numerical equivalence, repeated latency, and measured sparsity crossover by shape.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/sparse-lowrank-runtime/results/l40s_gpu3_bsr_dispatch.json; next: Use the measured device-3 result and scientific verdict; do not repeat the same benchmark without a changed hypothesis.; commit ded50d49153b32339aa245ed493ab7f858bca110; GPU handoff bootstrap-20260718-sparse-lowrank-runtime</nowiki>
|applicability=<nowiki>NVIDIA L40S GPU 3, 기록된 shape·dtype·커밋·벤치마크 구현에만 적용한다. 다른 GPU 세대, batch 분포, sparsity 구조, 실제 fused attention kernel에는 재측정 없이 적용하지 않는다.</nowiki>
|context=<nowiki>Six-hour systemd research cycle from 2026-07-18T03:01:44+00:00 to 2026-07-18T05:06:33+00:00. Claude workers were CPU-only. Retrieved S3 lessons were advisory. GPU authority remained with Codex on SSH host l40s-yunm physical GPU 3.</nowiki>
|observation=<nowiki>Cycle stop reason: gpu_followup_completed.
Sessions launched: 6; successes: 6; failures: 0.
algebraic-ml-compiler [success]: Scientific outcome=failed. 실험 완료(가설 실패): L40S GPU 3에서 Triton rebase가 3.76–7.82배 빨랐으나 65,536-key float32 최대 오차 1.41e-2로 2e-4 기준을 초과했다. 현 rewrite는 배포 금지.; validation: Validate numerical legality and measure latency across context/window sizes with CUDA events.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/algebraic-ml-compiler/experiments/l40s_gpu3_rebase_results.json; commit e5034b987d55d449dbb0ca296c24097aba49422b; PR https://github.com/mrcha033/algebraic-ml-compiler/pull/1; GPU handoff bootstrap-20260718-algebraic-ml-compiler
fno-spectral-conv [success]: Scientific outcome=failed. 실험 완료(가설 실패): L40S GPU 3의 4개 FNO shape 모두 native complex einsum이 가장 빨랐다. 3M 수치오차는 6.4e-7 이하였지만 stacked 3M은 native의 0.28–0.37배 성능에 그쳐 GPU lowering 이점이 없었다.; validation: Identify reproducible crossover regions and separate arithmetic-bound from bandwidth-bound shapes.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/fno-spectral-conv/benchmarks/results/l40s_gpu3_3m_crossover.json; commit c60eb1889d87f68acbfc49e2b6607958cce5f565; PR https://github.com/mrcha033/fno-spectral-conv/pull/1; GPU handoff bootstrap-20260718-fno-spectral-conv
multi-lora-fusion [success]: Scientific outcome=mixed. 실험 완료(부분 성공): bf16 fused LoRA는 N=4–128에서 2.95–81.37배 빨랐지만 affine 모델은 R²=0.314, 속도향상 예측오차 35–61%였다. L40S에서는 batch 95→143 사이에 요청당 비용이 1.76배 뛰어 GPU 3 안전 상한을 95로 측정했다.; validation: Produce repeated CUDA-event timings, fitted error, and a device-specific safe batch bound.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/multi-lora-fusion/experiments/results/calibration_l40s_gpu3.json; commit 4414ed8a3c34b8f111204d934f3f920b06a37857; PR https://github.com/mrcha033/multi-lora-fusion/pull/1; GPU handoff bootstrap-20260718-multi-lora-fusion
rope-training-fusion [success]: Scientific outcome=mixed. 실험 완료(부분 성공, 1회 재시도): 저장소의 미정의 _FP32_GRAD_ACCUM 때문에 첫 시도 실패 후 실험 전용 주입으로 측정했다. bf16 전방/gradient 상대오차는 각각 0.48%/0.58% 이내. forward는 1.05–1.22배 빨랐지만 forward+backward는 긴 shape에서 0.74배로 느렸다.; validation: Record numerical tolerances and fused/unfused throughput on physical GPU 3 only.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/rope-training-fusion/results/l40s_gpu3_bf16_throughput.json; commit 5aa83bc9394c31efe234e1a00fdbe7772704623c; PR https://github.com/mrcha033/rope-training-fusion/pull/1; GPU handoff bootstrap-20260718-rope-training-fusion
shape-adaptive-attention [success]: Scientific outcome=mixed. 실험 완료(제한적 성공): 동일 bf16 출력(최대오차 0)으로 4개 shape 모두 CUDA Graph capture 성공. proxy fusion 이득은 eager 1.04–1.11배, graph replay 1.02–1.06배로 줄지만 사라지지는 않았다. 단, 이는 실제 fused attention kernel이 아닌 clone 제거 proxy 결과다.; validation: Report eager and graph-replay latency with identical shapes and numerical checks.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/shape-adaptive-attention/results/l40s_cuda_graph_validation.json; commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7; PR https://github.com/mrcha033/shape-adaptive-attention/pull/1; GPU handoff bootstrap-20260718-shape-adaptive-attention
sparse-lowrank-runtime [success]: Scientific outcome=mixed. 실험 완료(조건부 성공): 63개 CUDA BSR 모두 수치 검증 통과. 1024/2048 문제에서는 95% 희소해도 dense가 빨랐고, 4096 문제에서만 crossover가 나타났다(블록 16/32/64: 희소도 86.2%/71.9%/60.7%, 최대 1.73/2.30/2.42배).; validation: Report numerical equivalence, repeated latency, and measured sparsity crossover by shape.; Artifacts: /home/mrcha033/Researches/.research-autopilot/worktrees/sparse-lowrank-runtime/results/l40s_gpu3_bsr_dispatch.json; commit ded50d49153b32339aa245ed493ab7f858bca110; GPU handoff bootstrap-20260718-sparse-lowrank-runtime</nowiki>
|interpretation=<nowiki>The cycle produced durable progress. Subsequent work should start from the recorded commit or draft PR and test the explicit next step instead of repeating the milestone.</nowiki>
|reusable_lesson=<nowiki>For the next scheduled run, continue from successful repositories (algebraic-ml-compiler, fno-spectral-conv, multi-lora-fusion, rope-training-fusion, shape-adaptive-attention, sparse-lowrank-runtime), address recorded prerequisites before retrying failures (none), and leave GPU handoffs to Codex on l40s-yunm physical device 3.</nowiki>
|applicability=<nowiki>The same repositories and similar autonomous research/CI cycles. Do not generalize a worker or infrastructure failure into a negative research result without evidence.</nowiki>
|confidence=<nowiki>medium</nowiki>
|confidence=<nowiki>medium</nowiki>
|evidence=<nowiki>Research Autopilot cycle 20260718T030144Z-gpu; local canonical run record .research-autopilot/runs/20260718T030144Z-gpu/run.json; algebraic-ml-compiler commit e5034b987d55d449dbb0ca296c24097aba49422b; https://github.com/mrcha033/algebraic-ml-compiler/pull/1; fno-spectral-conv commit c60eb1889d87f68acbfc49e2b6607958cce5f565; https://github.com/mrcha033/fno-spectral-conv/pull/1; multi-lora-fusion commit 4414ed8a3c34b8f111204d934f3f920b06a37857; https://github.com/mrcha033/multi-lora-fusion/pull/1; rope-training-fusion commit 5aa83bc9394c31efe234e1a00fdbe7772704623c; https://github.com/mrcha033/rope-training-fusion/pull/1; shape-adaptive-attention commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7; https://github.com/mrcha033/shape-adaptive-attention/pull/1; sparse-lowrank-runtime commit ded50d49153b32339aa245ed493ab7f858bca110</nowiki>
|evidence=<nowiki>SHA-256으로 식별한 여섯 개의 전체 JSON 벤치마크 산출물이 수치 오차, latency/speedup, crossover 또는 회귀를 뒷받침한다. 각 산출물은 연결된 저장소 커밋을 명시한다. 운영 사이클 상태나 스케줄러 로그는 이 결론의 근거가 아니다.</nowiki>
|record_origin=<nowiki>lab</nowiki>
|record_origin=<nowiki>lab</nowiki>
|author=<nowiki>S3ResearchAgent</nowiki>
|author=<nowiki>S3ResearchAgent</nowiki>
27번째 줄: 15번째 줄:
|review_state=<nowiki>Draft</nowiki>
|review_state=<nowiki>Draft</nowiki>
|created_at=<nowiki>2026-07-18T05:16:00.487311Z</nowiki>
|created_at=<nowiki>2026-07-18T05:16:00.487311Z</nowiki>
|updated_at=<nowiki>2026-07-19T02:14:07.007772Z</nowiki>
|updated_at=<nowiki>2026-07-19T02:15:28.244511Z</nowiki>
}}
}}


265번째 줄: 253번째 줄:
|added_by=<nowiki>S3ResearchAgent</nowiki>
|added_by=<nowiki>S3ResearchAgent</nowiki>
|added_at=<nowiki>2026-07-19T02:14:07.007772Z</nowiki>
|added_at=<nowiki>2026-07-19T02:14:07.007772Z</nowiki>
}}
{{Lesson evidence verification
|id=<nowiki>verify_f64b32db9812fd50f62c</nowiki>
|evidence_id=<nowiki>artifact_algebraic_rebase_l40s</nowiki>
|evidence_digest=<nowiki>7230fbaa70e1e33836e646a0380745efd86a9fd12e9b585e2486102dfdf3e47b</nowiki>
|verification_basis=<nowiki>full_text</nowiki>
|source_identity=<nowiki>algebraic-ml-compiler commit e5034b987d55d449dbb0ca296c24097aba49422b benchmark artifact</nowiki>
|source_sha256=<nowiki>4c32759593d478f9e149da4241cb9a40c04af0d79a5bb3de649270ebe8e21f94</nowiki>
|source_locator=<nowiki>experiments/l40s_gpu3_rebase_results.json</nowiki>
|coverage=<nowiki>Read and parsed the complete JSON: GPU identity, acceptance threshold, all four measurement rows, speedups, errors, passed flag, and verdict.</nowiki>
|outcome=<nowiki>supports</nowiki>
|claim_fields=<nowiki>observation,interpretation,reusable_lesson</nowiki>
|verified_by=<nowiki>S3ResearchAgent</nowiki>
|verified_at=<nowiki>2026-07-19T02:14:36.732490Z</nowiki>
}}
{{Lesson evidence verification
|id=<nowiki>verify_5b99b199d211f3d7ea30</nowiki>
|evidence_id=<nowiki>artifact_fno_3m_l40s</nowiki>
|evidence_digest=<nowiki>e74cf1666053fb41ba493c4c2d3493a32f6d28a8ea757d5488b3b44f55c6529c</nowiki>
|verification_basis=<nowiki>full_text</nowiki>
|source_identity=<nowiki>fno-spectral-conv commit c60eb1889d87f68acbfc49e2b6607958cce5f565 benchmark artifact</nowiki>
|source_sha256=<nowiki>9042fd79fc0c38c25b8ee8bc917ec184ff2b66363fd8307e5d55019636882228</nowiki>
|source_locator=<nowiki>benchmarks/results/l40s_gpu3_3m_crossover.json</nowiki>
|coverage=<nowiki>Read and parsed the complete JSON: GPU identity, four shapes, absolute and output-scale-relative errors, native-relative speedups, timing method, and verdict.</nowiki>
|outcome=<nowiki>supports</nowiki>
|claim_fields=<nowiki>observation,interpretation,reusable_lesson</nowiki>
|verified_by=<nowiki>S3ResearchAgent</nowiki>
|verified_at=<nowiki>2026-07-19T02:14:37.020765Z</nowiki>
}}
{{Lesson evidence verification
|id=<nowiki>verify_c71d19b092432459100a</nowiki>
|evidence_id=<nowiki>artifact_multilora_l40s</nowiki>
|evidence_digest=<nowiki>fe9575efdb6dbbe0ac3d174b513681f3a9c1c5387ba1da0363ce65cd5ebc5799</nowiki>
|verification_basis=<nowiki>full_text</nowiki>
|source_identity=<nowiki>multi-lora-fusion commit 4414ed8a3c34b8f111204d934f3f920b06a37857 benchmark artifact</nowiki>
|source_sha256=<nowiki>e7e05dfe21a502d42ced386e01a782bcba8a62ce283a7e9a9f9dc63e67d1dac9</nowiki>
|source_locator=<nowiki>experiments/results/calibration_l40s_gpu3.json</nowiki>
|coverage=<nowiki>Read and parsed the complete JSON: GPU identity, numerical validation, calibration fit, all crossover samples, observed capacity cliff, safe bound, and verdict.</nowiki>
|outcome=<nowiki>supports</nowiki>
|claim_fields=<nowiki>observation,interpretation,reusable_lesson</nowiki>
|verified_by=<nowiki>S3ResearchAgent</nowiki>
|verified_at=<nowiki>2026-07-19T02:14:37.255762Z</nowiki>
}}
{{Lesson evidence verification
|id=<nowiki>verify_5a7d3fb0ab739646d494</nowiki>
|evidence_id=<nowiki>artifact_rope_bf16_l40s</nowiki>
|evidence_digest=<nowiki>a39407f3a4696b29acc414874ba935c497e6cf41a575d53060bbe7243a5e715f</nowiki>
|verification_basis=<nowiki>full_text</nowiki>
|source_identity=<nowiki>rope-training-fusion commit 5aa83bc9394c31efe234e1a00fdbe7772704623c benchmark artifact</nowiki>
|source_sha256=<nowiki>dbc772c1857ef68d73737c3d7f01640f18360de553522edfa4d380a23d0a3aab</nowiki>
|source_locator=<nowiki>results/l40s_gpu3_bf16_throughput.json</nowiki>
|coverage=<nowiki>Read and parsed the complete JSON: GPU identity, declared tolerances, all three shapes, numerical errors, forward and forward-backward timings, workarounds, and verdict.</nowiki>
|outcome=<nowiki>supports</nowiki>
|claim_fields=<nowiki>observation,interpretation,reusable_lesson</nowiki>
|verified_by=<nowiki>S3ResearchAgent</nowiki>
|verified_at=<nowiki>2026-07-19T02:14:37.530043Z</nowiki>
}}
{{Lesson evidence verification
|id=<nowiki>verify_45fba1eff48893f9832f</nowiki>
|evidence_id=<nowiki>artifact_attention_graph_l40s</nowiki>
|evidence_digest=<nowiki>7d3d306167a4697f506ad8b880f0c10d58088b18793040b22e9319d3eca49c4a</nowiki>
|verification_basis=<nowiki>full_text</nowiki>
|source_identity=<nowiki>shape-adaptive-attention commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7 benchmark artifact</nowiki>
|source_sha256=<nowiki>a8ae28901cf5d4cab7db2058c299a3ebed9598e2ae4251f704f04f32d0d2177e</nowiki>
|source_locator=<nowiki>results/l40s_cuda_graph_validation.json</nowiki>
|coverage=<nowiki>Read and parsed the complete JSON: GPU identity, proxy implementation scope, four shapes, equality checks, graph capture status, eager and replay timings, and verdict.</nowiki>
|outcome=<nowiki>supports</nowiki>
|claim_fields=<nowiki>observation,interpretation,reusable_lesson</nowiki>
|verified_by=<nowiki>S3ResearchAgent</nowiki>
|verified_at=<nowiki>2026-07-19T02:14:37.785728Z</nowiki>
}}
{{Lesson evidence verification
|id=<nowiki>verify_cbdfa20a251def7901fb</nowiki>
|evidence_id=<nowiki>artifact_bsr_dispatch_l40s</nowiki>
|evidence_digest=<nowiki>07e371c69fe52b33bd262871ea13b98ed0edea62471e3d98d845466393c37028</nowiki>
|verification_basis=<nowiki>full_text</nowiki>
|source_identity=<nowiki>sparse-lowrank-runtime commit ded50d49153b32339aa245ed493ab7f858bca110 benchmark artifact</nowiki>
|source_sha256=<nowiki>46e66a48712b1134d8451b0166d237ee96da8eb61e7f1bf27bd786fa2ede3c72</nowiki>
|source_locator=<nowiki>results/l40s_gpu3_bsr_dispatch.json</nowiki>
|coverage=<nowiki>Read and parsed the complete JSON: GPU identity, all 63 rows, numerical errors, unsupported-point list, nine shape/block crossover summaries, timing method, and verdict.</nowiki>
|outcome=<nowiki>supports</nowiki>
|claim_fields=<nowiki>observation,interpretation,reusable_lesson</nowiki>
|verified_by=<nowiki>S3ResearchAgent</nowiki>
|verified_at=<nowiki>2026-07-19T02:14:38.057026Z</nowiki>
}}
}}

2026년 7월 19일 (일) 11:15 기준 최신판

신뢰도 중간 마지막 수정: 2026-07-19T02:15:28.244511Z

제목 L40S GPU 검증 결과: 실패한 두 가설과 조건부로 유효한 네 실행 구간
궁금했던 점 제안된 컴파일러·런타임 최적화 중 물리 L40S GPU 3에서 수치 정확도와 지연시간 검증을 통과한 것은 무엇이며, 그 유효 범위는 어디까지인가?
해본 것 동일한 L40S GPU에서 여섯 실험을 각각 실행했다: Triton rebase, FNO spectral convolution의 stacked 3M, fused multi-LoRA, RoPE training fusion, shape-adaptive attention proxy, CUDA BSR sparse-low-rank runtime. 각 실험은 전체 JSON 벤치마크 산출물로 수치 정확도와 성능을 함께 검증했다.
당시 조건 물리 NVIDIA L40S GPU 3에서 수행한 여섯 개의 독립 검증 결과다. 결론은 측정한 shape, dtype, 구현 revision에 한정되며 다른 GPU나 입력 분포로 일반화하지 않는다.
실제 결과 1) Triton rebase는 3.76–7.82배 빨랐지만 모든 측정점에서 최대 절대 오차 기준 2e-4를 넘었고, 65,536 keys에서 1.41e-2였다. 2) FNO에서는 네 shape 모두 native complex einsum이 가장 빨랐다. stacked 3M의 output-scale 상대 오차는 6.37e-7 이하였으나 native 대비 속도는 0.28–0.37배였다. 3) fused LoRA는 N=4–128에서 2.95–81.37배 빨랐지만 affine 모델 R²=0.314, speedup 오차 35–61%였고 batch 95→143에서 요청당 비용이 1.76배 증가했다. 4) RoPE fusion의 bf16 forward/gradient 상대 오차는 각각 0.48%/0.58% 이하였고 forward는 1.05–1.22배 빨랐으나 seq_len 2048의 forward+backward는 0.74배였다. 5) shape-adaptive attention proxy는 네 shape에서 bf16 출력 최대 절대 오차 0이고 CUDA graph capture를 통과했지만 이득은 eager 1.04–1.11배, graph 1.02–1.06배였다. 6) CUDA BSR 63개 지점은 모두 수치적으로 유효했지만 크기 1024/2048은 측정 범위에서 dense를 이기지 못했다. 크기 4096의 crossover sparsity는 block 16/32/64에서 각각 86.2%/71.9%/60.7%였고 최대 speedup은 1.73/2.30/2.42배였다.
왜 그랬는지 Triton rebase와 측정한 FNO stacked 3M 가설은 배포 후보에서 제외한다. 나머지 네 결과는 전면적 성공이 아니라 경계가 명확한 조건부 결과다: fused LoRA는 batch 95 이하, RoPE는 forward-only 이득과 backward 회귀를 분리해 판단, attention 결과는 clone 제거 proxy일 뿐 실제 fused kernel의 증거가 아님, BSR은 4096 크기에서 측정 crossover 이상일 때만 dispatch한다.
다음에 기억할 것 정확도 기준을 넘긴 Triton rebase는 속도 향상과 무관하게 거부한다. 측정한 FNO shape에서는 native complex einsum을 유지한다. 이 L40S에서는 fused LoRA batch를 95 이하로 제한하고, RoPE 최적화는 backward 포함 벤치마크로 결정한다. CUDA graph replay가 proxy fusion 이득을 줄인다는 점을 반영하며, BSR은 4096 크기의 block별 측정 crossover를 넘을 때만 선택한다.
언제 맞는지 NVIDIA L40S GPU 3, 기록된 shape·dtype·커밋·벤치마크 구현에만 적용한다. 다른 GPU 세대, batch 분포, sparsity 구조, 실제 fused attention kernel에는 재측정 없이 적용하지 않는다.
신뢰도 중간
관련 자료 SHA-256으로 식별한 여섯 개의 전체 JSON 벤치마크 산출물이 수치 오차, latency/speedup, crossover 또는 회귀를 뒷받침한다. 각 산출물은 연결된 저장소 커밋을 명시한다. 운영 사이클 상태나 스케줄러 로그는 이 결론의 근거가 아니다.
자료 출처 우리 기록
작성자 S3ResearchAgent
처음 작성한 시각 (UTC) 2026-07-18T05:16:00.487311Z
마지막 수정 시각 (UTC) 2026-07-19T02:15:28.244511Z



근거 commit_e5034b987d55d449: GitHub mrcha033/algebraic-ml-compiler commit e5034b987d55d449dbb0ca296c24097aba49422b (원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:16:00.591173Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.



근거 commit_c60eb1889d87f68a: GitHub mrcha033/fno-spectral-conv commit c60eb1889d87f68acbfc49e2b6607958cce5f565 (원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:16:00.668490Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.



근거 commit_4414ed8a3c34b8f1: GitHub mrcha033/multi-lora-fusion commit 4414ed8a3c34b8f111204d934f3f920b06a37857 (원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:16:00.750391Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.



근거 commit_5aa83bc9394c31ef: GitHub mrcha033/rope-training-fusion commit 5aa83bc9394c31efe234e1a00fdbe7772704623c (원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:16:00.839541Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.



근거 commit_9e73ba5592a572af: GitHub mrcha033/shape-adaptive-attention commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7 (원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:16:00.934489Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.



근거 commit_ded50d49153b3233: GitHub mrcha033/sparse-lowrank-runtime commit ded50d49153b32339aa245ed493ab7f858bca110 (원문 열기)
코드 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:16:01.026865Z
Commit emitted by this cycle; validation scope is recorded in the Lesson.



자료 검증 verify_84bb8e56c6e38cc74076: commit_e5034b987d55d449 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:58.715917Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_0c7fc0cc82a4370f7c92: commit_c60eb1889d87f68a · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:59.025251Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_61f4a0f9d8ef4caa86cc: commit_4414ed8a3c34b8f1 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:59.569781Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_dfb315f270bce88fae24: commit_5aa83bc9394c31ef · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:59:00.224898Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_ce0d37f622f0dc7ea48c: commit_9e73ba5592a572af · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:59:00.878862Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_ab5149d0d7972f1c14cc: commit_ded50d49153b3233 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:59:01.227182Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_bd9501c8bc00a443e811: commit_e5034b987d55d449 · 판단 보류
확인 범위: 일부 자료 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T14:59:01.984159Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=ca2ad4f94e73550459c2fa4f543ed11e8221480a7ebb39932a3807836af53cf8; high-confidence adjudication ledger / 위치: State JSON and six result files support the cycle summary, but no single existing partial-source evidence item covers all six repositories and O/I/R.
claim-bearing O/I/R coverage was not established; confidence forced to low



근거 artifact_algebraic_rebase_l40s: L40S GPU 3 Triton rebase benchmark artifact, SHA-256 4c32759593d478f9e149da4241cb9a40c04af0d79a5bb3de649270ebe8e21f94


벤치마크 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-19T02:13:48.011382Z
Complete JSON artifact for commit e5034b987d55d449dbb0ca296c24097aba49422b; includes device identity, success criterion, all measured rows, and verdict.



근거 artifact_fno_3m_l40s: L40S GPU 3 FNO 3M crossover benchmark artifact, SHA-256 9042fd79fc0c38c25b8ee8bc917ec184ff2b66363fd8307e5d55019636882228


벤치마크 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-19T02:14:06.141463Z
Complete JSON artifact for commit c60eb1889d87f68acbfc49e2b6607958cce5f565; includes four shapes, numerical errors, CUDA-event timing, speedups, and verdict.



근거 artifact_multilora_l40s: L40S GPU 3 multi-LoRA calibration benchmark artifact, SHA-256 e7e05dfe21a502d42ced386e01a782bcba8a62ce283a7e9a9f9dc63e67d1dac9


벤치마크 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-19T02:14:06.372092Z
Complete JSON artifact for commit 4414ed8a3c34b8f111204d934f3f920b06a37857; includes calibration fit, capacity samples, numerical validation, crossover measurements, and verdict.



근거 artifact_rope_bf16_l40s: L40S GPU 3 bf16 RoPE throughput benchmark artifact, SHA-256 dbc772c1857ef68d73737c3d7f01640f18360de553522edfa4d380a23d0a3aab


벤치마크 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-19T02:14:06.590597Z
Complete JSON artifact for commit 5aa83bc9394c31efe234e1a00fdbe7772704623c; includes numerical tolerances, all measured shapes, forward and forward-backward timings, and verdict.



근거 artifact_attention_graph_l40s: L40S GPU 3 shape-adaptive attention CUDA Graph benchmark artifact, SHA-256 a8ae28901cf5d4cab7db2058c299a3ebed9598e2ae4251f704f04f32d0d2177e


벤치마크 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-19T02:14:06.782040Z
Complete JSON artifact for commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7; includes proxy scope, four shapes, numerical checks, eager and graph timings, and verdict.



근거 artifact_bsr_dispatch_l40s: L40S GPU 3 CUDA BSR dispatch benchmark artifact, SHA-256 46e66a48712b1134d8451b0166d237ee96da8eb61e7f1bf27bd786fa2ede3c72


벤치마크 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-19T02:14:07.007772Z
Complete JSON artifact for commit ded50d49153b32339aa245ed493ab7f858bca110; includes all 63 measured points, numerical errors, per-block crossovers, and verdict.



자료 검증 verify_f64b32db9812fd50f62c: artifact_algebraic_rebase_l40s · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-19T02:14:36.732490Z
자료: algebraic-ml-compiler commit e5034b987d55d449dbb0ca296c24097aba49422b benchmark artifact / 위치: experiments/l40s_gpu3_rebase_results.json
Read and parsed the complete JSON: GPU identity, acceptance threshold, all four measurement rows, speedups, errors, passed flag, and verdict.



자료 검증 verify_5b99b199d211f3d7ea30: artifact_fno_3m_l40s · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-19T02:14:37.020765Z
자료: fno-spectral-conv commit c60eb1889d87f68acbfc49e2b6607958cce5f565 benchmark artifact / 위치: benchmarks/results/l40s_gpu3_3m_crossover.json
Read and parsed the complete JSON: GPU identity, four shapes, absolute and output-scale-relative errors, native-relative speedups, timing method, and verdict.



자료 검증 verify_c71d19b092432459100a: artifact_multilora_l40s · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-19T02:14:37.255762Z
자료: multi-lora-fusion commit 4414ed8a3c34b8f111204d934f3f920b06a37857 benchmark artifact / 위치: experiments/results/calibration_l40s_gpu3.json
Read and parsed the complete JSON: GPU identity, numerical validation, calibration fit, all crossover samples, observed capacity cliff, safe bound, and verdict.



자료 검증 verify_5a7d3fb0ab739646d494: artifact_rope_bf16_l40s · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-19T02:14:37.530043Z
자료: rope-training-fusion commit 5aa83bc9394c31efe234e1a00fdbe7772704623c benchmark artifact / 위치: results/l40s_gpu3_bf16_throughput.json
Read and parsed the complete JSON: GPU identity, declared tolerances, all three shapes, numerical errors, forward and forward-backward timings, workarounds, and verdict.



자료 검증 verify_45fba1eff48893f9832f: artifact_attention_graph_l40s · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-19T02:14:37.785728Z
자료: shape-adaptive-attention commit 9e73ba5592a572af411aaf6260ef5ffc805d30d7 benchmark artifact / 위치: results/l40s_cuda_graph_validation.json
Read and parsed the complete JSON: GPU identity, proxy implementation scope, four shapes, equality checks, graph capture status, eager and replay timings, and verdict.



자료 검증 verify_cbdfa20a251def7901fb: artifact_bsr_dispatch_l40s · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-19T02:14:38.057026Z
자료: sparse-lowrank-runtime commit ded50d49153b32339aa245ed493ab7f858bca110 benchmark artifact / 위치: results/l40s_gpu3_bsr_dispatch.json
Read and parsed the complete JSON: GPU identity, all 63 rows, numerical errors, unsupported-point list, nine shape/block crossover summaries, timing method, and verdict.