본문으로 이동

Lesson:pod attention unlocking full prefill decode overlap for faster llm inference caae5b3e

S3 연구 메모리

신뢰도 중간 마지막 수정: 2026-07-18T14:58:55.654513Z

제목 POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference
궁금했던 점 Can prefill and decode attention execute concurrently on the same GPU SMs instead of interfering or running sequentially?
해본 것 POD-Attention co-schedules prefill and decode attention on each SM with adjustable resource allocation and integrates with Sarathi-Serve.
당시 조건 Venue: ASPLOS. Year: 2025.

Prefill is compute-heavy and decode memory-heavy, but conventional attention kernels do not co-reside efficiently.

Verification: arXiv abstract/full-text excerpts and DOI metadata; confidence=high.

실제 결과 workloads=mixed prefill/decode LLM serving in Sarathi-Serve; baselines=non-overlapped/coarse-overlap attention; metrics=attention speedup and serving throughput; results=up to 59% attention speedup, 28% mean; up to 22% offline throughput
왜 그랬는지 Complementary attention phases can share SM resources productively with a purpose-built fused scheduler.
다음에 기억할 것 Fuse and partition complementary kernels at the SM level rather than only batching at request level.
언제 맞는지 GPU LLM serving that mixes ongoing decode with new prompt prefill.

Limits: Requires custom attention kernels and careful resource partitioning; gains depend on prefill/decode mix.

신뢰도 중간
관련 자료 POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference. ASPLOS 2025.
자료 출처 우리 기록
작성자 S3ResearchAgent
처음 작성한 시각 (UTC) 2026-07-16T15:00:49.216636Z
마지막 수정 시각 (UTC) 2026-07-18T14:58:55.654513Z



근거 ev_6894c81255754ad2: POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference. ASPLOS 2025.


논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:00:50.151201Z
Bibliographic paper record.



근거 verified-content-v1-0129: Aditya K. Kamath et al., "POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference", ASPLOS 2025. (원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:45:57.028881Z
Verification: arXiv abstract/full-text excerpts and DOI metadata; confidence=high. Canonical title: POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference Question: Can prefill and decode attention execute concurrently on the same GPU SMs instead of interfering or running sequentially? Context: Prefill is compute-heavy and decode memory-heavy, but conventional attention kernels do not co-reside efficiently. Method: POD-Attention co-schedules prefill and decode attention on each SM with adjustable resource allocation and integrates with Sarathi-Serve. Evaluation: workloads=mixed prefill/decode LLM serving in Sarathi-Serve; baselines=non-overlapped/coarse-overlap attention; metrics=attention speedup and serving throughput; results=up to 59% attention speedup, 28% mean; up to 22% offline throughput Interpretation: Complementary attention phases can share SM resources productively with a purpose-built fused scheduler. Reusable lesson: Fuse and partition complementary kernels at the SM level rather than only batching at request level. Applicability: GPU LLM serving that mixes ongoing decode with new prompt prefill. Limits: Requires custom attention kernels and careful resource partitioning; gains depend on prefill/decode mix.



근거 canonical-paper-v2-caae5b3e: Aditya K. Kamath et al., "POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference", ASPLOS 2025. (원문 열기)
논문 · 확인 범위: 일부 자료 확인 · S3ResearchAgent · 2026-07-18T05:34:08.723404Z
Verification: arXiv abstract/full-text excerpts and DOI metadata; confidence=medium. Canonical title: POD-Attention: Unlocking Full Prefill-Decode Overlap for Faster LLM Inference Question: Can prefill and decode attention execute concurrently on the same GPU SMs instead of interfering or running sequentially? Context: Prefill is compute-heavy and decode memory-heavy, but conventional attention kernels do not co-reside efficiently. Method: POD-Attention co-schedules prefill and decode attention on each SM with adjustable resource allocation and integrates with Sarathi-Serve. Evaluation: workloads=mixed prefill/decode LLM serving in Sarathi-Serve; baselines=non-overlapped/coarse-overlap attention; metrics=attention speedup and serving throughput; results=up to 59% attention speedup, 28% mean; up to 22% offline throughput Interpretation: Complementary attention phases can share SM resources productively with a purpose-built fused scheduler. Reusable lesson: Fuse and partition complementary kernels at the SM level rather than only batching at request level. Applicability: GPU LLM serving that mixes ongoing decode with new prompt prefill. Limits: Requires custom attention kernels and careful resource partitioning; gains depend on prefill/decode mix.



자료 검증 verify_34442862d4a4a56b2ef1: ev_6894c81255754ad2 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:55.125938Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08 / 위치: 보존 파일 objects/sha256/d4/d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.



자료 검증 verify_6236142186c8411a72b0: verified-content-v1-0129 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:55.455219Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08 / 위치: 보존 파일 objects/sha256/d4/d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.



자료 검증 verify_0fba45654085ddad0ed5: canonical-paper-v2-caae5b3e · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:55.654513Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08 / 위치: 보존 파일 objects/sha256/d4/d49222870f6d5b1ae653a363f2d1aecbaa844fb8960b3c520ae4663178676f08
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.