Lesson:moe lightning high throughput moe inference on memory constrained gpus e51b6b29
| 제목 | MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs |
|---|---|
| 궁금했던 점 | How can large MoE models achieve high batch throughput on GPUs too small to hold their expert weights? |
| 해본 것 | MoE-Lightning uses CGOPipe CPU-GPU-I/O pipelining with paged weights and a Hierarchical Roofline Model to choose higher-throughput placement/schedules. |
| 당시 조건 | Venue: ASPLOS. Year: 2025.
Offloading is dominated by hierarchical CPU/GPU/I/O bottlenecks and poor overlap among expert weight transfer, compute, and KV state. Verification: full_text; confidence=high. |
| 실제 결과 | workloads=Mixtral 8x7B on one T4 16 GB; Mixtral 8x22B; DBRX; 2–4 low-cost GPUs; baselines=state-of-the-art offloading systems; FlexGen; metrics=throughput; CPU memory; resource utilization; results=up to 10.3x throughput; throughput bound with 2–3x less CPU memory |
| 왜 그랬는지 | A hierarchy-aware analytic model can choose paging and overlap plans that saturate the true bottleneck. |
| 다음에 기억할 것 | Model every tier's roofline, then pipeline transfers and compute to the limiting resource. |
| 언제 맞는지 | Batch MoE inference on low-cost, memory-constrained GPUs.
Limits: Optimized for batch throughput; results depend on model routing, batch size, I/O tiers, and T4/L4-class systems. |
| 신뢰도 | 높음 |
| 관련 자료 | MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs. ASPLOS 2025. |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T14:58:10.910062Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T14:58:45.895839Z |
근거 ev_201e1dcb4b1a4c77: MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs. ASPLOS 2025.
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T14:58:12.937185Z
Bibliographic paper record.
근거 verified-content-v1-0101: Shiyi Cao et al., "MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs", ASPLOS 2025.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:55:22.598016Z
Verification: full_text; confidence=high.
Canonical title: MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
Question: How can large MoE models achieve high batch throughput on GPUs too small to hold their expert weights?
Context: Offloading is dominated by hierarchical CPU/GPU/I/O bottlenecks and poor overlap among expert weight transfer, compute, and KV state.
Method: MoE-Lightning uses CGOPipe CPU-GPU-I/O pipelining with paged weights and a Hierarchical Roofline Model to choose higher-throughput placement/schedules.
Evaluation: workloads=Mixtral 8x7B on one T4 16 GB; Mixtral 8x22B; DBRX; 2–4 low-cost GPUs; baselines=state-of-the-art offloading systems; FlexGen; metrics=throughput; CPU memory; resource utilization; results=up to 10.3x throughput; throughput bound with 2–3x less CPU memory
Interpretation: A hierarchy-aware analytic model can choose paging and overlap plans that saturate the true bottleneck.
Reusable lesson: Model every tier's roofline, then pipeline transfers and compute to the limiting resource.
Applicability: Batch MoE inference on low-cost, memory-constrained GPUs.
Limits: Optimized for batch throughput; results depend on model routing, batch size, I/O tiers, and T4/L4-class systems.
근거 canonical-paper-v2-e51b6b29: Shiyi Cao et al., "MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs", ASPLOS 2025.
(원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:30:01.398033Z
Verification: full_text; confidence=high.
Canonical title: MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
Question: How can large MoE models achieve high batch throughput on GPUs too small to hold their expert weights?
Context: Offloading is dominated by hierarchical CPU/GPU/I/O bottlenecks and poor overlap among expert weight transfer, compute, and KV state.
Method: MoE-Lightning uses CGOPipe CPU-GPU-I/O pipelining with paged weights and a Hierarchical Roofline Model to choose higher-throughput placement/schedules.
Evaluation: workloads=Mixtral 8x7B on one T4 16 GB; Mixtral 8x22B; DBRX; 2–4 low-cost GPUs; baselines=state-of-the-art offloading systems; FlexGen; metrics=throughput; CPU memory; resource utilization; results=up to 10.3x throughput; throughput bound with 2–3x less CPU memory
Interpretation: A hierarchy-aware analytic model can choose paging and overlap plans that saturate the true bottleneck.
Reusable lesson: Model every tier's roofline, then pipeline transfers and compute to the limiting resource.
Applicability: Batch MoE inference on low-cost, memory-constrained GPUs.
Limits: Optimized for batch throughput; results depend on model routing, batch size, I/O tiers, and T4/L4-class systems.
자료 검증 verify_081937c918b6f35c0c1a:
ev_201e1dcb4b1a4c77 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:45.142288Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.
자료 검증 verify_b2a4bda7888053fae1f3:
verified-content-v1-0101 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:45.499598Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.
자료 검증 verify_5a0b1c6eaba183941a47:
canonical-paper-v2-e51b6b29 ·
지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T14:58:45.895839Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=932c469440fb76dfb3466974f55ea2ad537c539c4086bc9da25cc39108b6ca1a; independently adjudicated claim-bearing primary source / 위치: Abstract; Section 1; Section 3 HRM; Section 4 CGOPipe and HRM policy; Section 5 evaluation.
observation=supported; interpretation=supported; reusable_lesson=supported