Lesson:nanoflow towards optimal large language model serving throughput f0633332
| 제목 | NanoFlow: Towards Optimal Large Language Model Serving Throughput |
|---|---|
| 궁금했던 점 | How close can LLM serving get to hardware-optimal throughput by overlapping compute, memory, and network work inside each GPU? |
| 해본 것 | NanoFlow forms nano-batches, duplicates selected operators, and automatically chooses count, size, order, and resource allocation for overlap. |
| 당시 조건 | Venue: OSDI. Year: 2025.
Request batching alone leaves device resources idle because Transformer operations stress different resources at different times. Verification: official USENIX page and abstract; confidence=high. |
| 실제 결과 | workloads=LLaMA-2-70B, Mixtral-8x7B, LLaMA-3-8B; baselines=state-of-the-art LLM serving systems and analytical optimum; metrics=serving throughput and percent of optimum; results=up to 1.91x; 50–72% of optimum |
| 왜 그랬는지 | Intra-device scheduling is as important as request scheduling for high-throughput serving. |
| 다음에 기억할 것 | Co-schedule operations with complementary bottlenecks at a granularity below the request batch. |
| 언제 맞는지 | GPU LLM serving with predictable operator profiles.
Limits: May duplicate work and requires model/hardware-specific profiling and scheduling. |
| 신뢰도 | 중간 |
| 관련 자료 | NanoFlow: Towards Optimal Large Language Model Serving Throughput. OSDI 2025. |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T15:00:39.766637Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T14:58:47.237637Z |
근거 ev_288b9e63b92346f9: NanoFlow: Towards Optimal Large Language Model Serving Throughput. OSDI 2025.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:00:40.735315Z
Bibliographic paper record.
근거 verified-content-v1-0125: Kan Zhu et al., "NanoFlow: Towards Optimal Large Language Model Serving Throughput", OSDI 2025.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:45:26.846011Z
Verification: official USENIX page and abstract; confidence=high.
Canonical title: NanoFlow: Towards Optimal Large Language Model Serving Throughput
Question: How close can LLM serving get to hardware-optimal throughput by overlapping compute, memory, and network work inside each GPU?
Context: Request batching alone leaves device resources idle because Transformer operations stress different resources at different times.
Method: NanoFlow forms nano-batches, duplicates selected operators, and automatically chooses count, size, order, and resource allocation for overlap.
Evaluation: workloads=LLaMA-2-70B, Mixtral-8x7B, LLaMA-3-8B; baselines=state-of-the-art LLM serving systems and analytical optimum; metrics=serving throughput and percent of optimum; results=up to 1.91x; 50–72% of optimum
Interpretation: Intra-device scheduling is as important as request scheduling for high-throughput serving.
Reusable lesson: Co-schedule operations with complementary bottlenecks at a granularity below the request batch.
Applicability: GPU LLM serving with predictable operator profiles.
Limits: May duplicate work and requires model/hardware-specific profiling and scheduling.
근거 canonical-paper-v2-f0633332: Kan Zhu et al., "NanoFlow: Towards Optimal Large Language Model Serving Throughput", OSDI 2025.
(원문 열기)
논문 · 확인 범위: 공식 초록 확인 · S3ResearchAgent · 2026-07-18T05:30:15.384786Z
Verification: official USENIX page and abstract; confidence=medium.
Canonical title: NanoFlow: Towards Optimal Large Language Model Serving Throughput
Question: How close can LLM serving get to hardware-optimal throughput by overlapping compute, memory, and network work inside each GPU?
Context: Request batching alone leaves device resources idle because Transformer operations stress different resources at different times.
Method: NanoFlow forms nano-batches, duplicates selected operators, and automatically chooses count, size, order, and resource allocation for overlap.
Evaluation: workloads=LLaMA-2-70B, Mixtral-8x7B, LLaMA-3-8B; baselines=state-of-the-art LLM serving systems and analytical optimum; metrics=serving throughput and percent of optimum; results=up to 1.91x; 50–72% of optimum
Interpretation: Intra-device scheduling is as important as request scheduling for high-throughput serving.
Reusable lesson: Co-schedule operations with complementary bottlenecks at a granularity below the request batch.
Applicability: GPU LLM serving with predictable operator profiles.
Limits: May duplicate work and requires model/hardware-specific profiling and scheduling.
자료 검증 verify_46e5dc31aaa17075f442:
ev_288b9e63b92346f9 ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:46.796726Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=7ca4384d8b6109799afa2e1b170e20ae2dbfd84d92fd89861efdef92c997701e / 위치: 보존 파일 objects/sha256/7c/7ca4384d8b6109799afa2e1b170e20ae2dbfd84d92fd89861efdef92c997701e
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.
자료 검증 verify_d1bee1e23ddb012940f3:
verified-content-v1-0125 ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:47.000902Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=7ca4384d8b6109799afa2e1b170e20ae2dbfd84d92fd89861efdef92c997701e / 위치: 보존 파일 objects/sha256/7c/7ca4384d8b6109799afa2e1b170e20ae2dbfd84d92fd89861efdef92c997701e
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.
자료 검증 verify_9204528c8cbd2e341e7b:
canonical-paper-v2-f0633332 ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:47.237637Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=7ca4384d8b6109799afa2e1b170e20ae2dbfd84d92fd89861efdef92c997701e / 위치: 보존 파일 objects/sha256/7c/7ca4384d8b6109799afa2e1b170e20ae2dbfd84d92fd89861efdef92c997701e
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.