Lesson:pod attention unlocking full prefill decode overlap for faster llm inference caae5b3e
| 제목 | POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference |
|---|---|
| 궁금했던 점 | Can prefill and decode attention execute concurrently on the same GPU SMs instead of interfering or running sequentially? |
| 해본 것 | POD-Attention co-schedules prefill and decode attention on each SM with adjustable resource allocation and integrates with Sarathi-Serve. |
| 당시 조건 | Venue: ASPLOS. Year: 2025.
Prefill is compute-heavy and decode memory-heavy, but conventional attention kernels do not co-reside efficiently. Verification: arXiv abstract/full-text excerpts and DOI metadata; confidence=high. |
| 실제 결과 | workloads=mixed prefill/decode LLM serving in Sarathi-Serve; baselines=non-overlapped/coarse-overlap attention; metrics=attention speedup and serving throughput; results=up to 59% attention speedup, 28% mean; up to 22% offline throughput |
| 왜 그랬는지 | Complementary attention phases can share SM resources productively with a purpose-built fused scheduler. |
| 다음에 기억할 것 | Fuse and partition complementary kernels at the SM level rather than only batching at request level. |
| 언제 맞는지 | GPU LLM serving that mixes ongoing decode with new prompt prefill.
Limits: Requires custom attention kernels and careful resource partitioning; gains depend on prefill/decode mix. |
| 신뢰도 | 중간 |
| 관련 자료 | POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference. ASPLOS 2025. |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T15:00:49.216636Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T14:58:55.654513Z |
근거 ev_6894c81255754ad2: POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference. ASPLOS 2025.
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:00:50.151201Z
Bibliographic paper record.
근거 verified-content-v1-0129: Aditya K. Kamath et al., "POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference", ASPLOS 2025.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:45:57.028881Z
Verification: arXiv abstract/full-text excerpts and DOI metadata; confidence=high.
Canonical title: POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Question: Can prefill and decode attention execute concurrently on the same GPU SMs instead of interfering or running sequentially?
Context: Prefill is compute-heavy and decode memory-heavy, but conventional attention kernels do not co-reside efficiently.
Method: POD-Attention co-schedules prefill and decode attention on each SM with adjustable resource allocation and integrates with Sarathi-Serve.
Evaluation: workloads=mixed prefill/decode LLM serving in Sarathi-Serve; baselines=non-overlapped/coarse-overlap attention; metrics=attention speedup and serving throughput; results=up to 59% attention speedup, 28% mean; up to 22% offline throughput
Interpretation: Complementary attention phases can share SM resources productively with a purpose-built fused scheduler.
Reusable lesson: Fuse and partition complementary kernels at the SM level rather than only batching at request level.
Applicability: GPU LLM serving that mixes ongoing decode with new prompt prefill.
Limits: Requires custom attention kernels and careful resource partitioning; gains depend on prefill/decode mix.
근거 canonical-paper-v2-caae5b3e: Aditya K. Kamath et al., "POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference", ASPLOS 2025.
(원문 열기)
논문 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:34:08.723404Z
Verification: arXiv abstract/full-text excerpts and DOI metadata; confidence=medium.
Canonical title: POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
Question: Can prefill and decode attention execute concurrently on the same GPU SMs instead of interfering or running sequentially?
Context: Prefill is compute-heavy and decode memory-heavy, but conventional attention kernels do not co-reside efficiently.
Method: POD-Attention co-schedules prefill and decode attention on each SM with adjustable resource allocation and integrates with Sarathi-Serve.
Evaluation: workloads=mixed prefill/decode LLM serving in Sarathi-Serve; baselines=non-overlapped/coarse-overlap attention; metrics=attention speedup and serving throughput; results=up to 59% attention speedup, 28% mean; up to 22% offline throughput
Interpretation: Complementary attention phases can share SM resources productively with a purpose-built fused scheduler.
Reusable lesson: Fuse and partition complementary kernels at the SM level rather than only batching at request level.
Applicability: GPU LLM serving that mixes ongoing decode with new prompt prefill.
Limits: Requires custom attention kernels and careful resource partitioning; gains depend on prefill/decode mix.
자료 검증 verify_34442862d4a4a56b2ef1:
ev_6894c81255754ad2 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:55.125938Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08 / 위치: 보존 파일 objects/sha256/d4/d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.
자료 검증 verify_6236142186c8411a72b0:
verified-content-v1-0129 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:55.455219Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08 / 위치: 보존 파일 objects/sha256/d4/d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.
자료 검증 verify_0fba45654085ddad0ed5:
canonical-paper-v2-caae5b3e ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:55.654513Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08 / 위치: 보존 파일 objects/sha256/d4/d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.