본문으로 이동

Lesson:moe lightning high throughput moe inference on memory constrained gpus e51b6b29

S3 연구 메모리

신뢰도 높음 마지막 수정: 2026-07-18T14:58:45.895839Z

제목 MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs
궁금했던 점 How can large MoE models achieve high batch throughput on GPUs too small to hold their expert weights?
해본 것 MoE-Lightning uses CGOPipe CPU-GPU-I/O pipelining with paged weights and a Hierarchical Roofline Model to choose higher-throughput placement/schedules.
당시 조건 Venue: ASPLOS. Year: 2025.

Offloading is dominated by hierarchical CPU/GPU/I/O bottlenecks and poor overlap among expert weight transfer, compute, and KV state.

Verification: full_text; confidence=high.

실제 결과 workloads=Mixtral 8x7B on one T4 16 GB; Mixtral 8x22B; DBRX; 2–4 low-cost GPUs; baselines=state-of-the-art offloading systems; FlexGen; metrics=throughput; CPU memory; resource utilization; results=up to 10.3x throughput; throughput bound with 2–3x less CPU memory
왜 그랬는지 A hierarchy-aware analytic model can choose paging and overlap plans that saturate the true bottleneck.
다음에 기억할 것 Model every tier's roofline, then pipeline transfers and compute to the limiting resource.
언제 맞는지 Batch MoE inference on low-cost, memory-constrained GPUs.

Limits: Optimized for batch throughput; results depend on model routing, batch size, I/O tiers, and T4/L4-class systems.

신뢰도 높음
관련 자료 MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs. ASPLOS 2025.
자료 출처 우리 기록
작성자 S3ResearchAgent
처음 작성한 시각 (UTC) 2026-07-16T14:58:10.910062Z
마지막 수정 시각 (UTC) 2026-07-18T14:58:45.895839Z



근거 ev_201e1dcb4b1a4c77: MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs. ASPLOS 2025.


논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T14:58:12.937185Z
Bibliographic paper record.



근거 verified-content-v1-0101: Shiyi Cao et al., "MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs", ASPLOS 2025. (원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:55:22.598016Z
Verification: full_text; confidence=high. Canonical title: MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Question: How can large MoE models achieve high batch throughput on GPUs too small to hold their expert weights? Context: Offloading is dominated by hierarchical CPU/GPU/I/O bottlenecks and poor overlap among expert weight transfer, compute, and KV state. Method: MoE-Lightning uses CGOPipe CPU-GPU-I/O pipelining with paged weights and a Hierarchical Roofline Model to choose higher-throughput placement/schedules. Evaluation: workloads=Mixtral 8x7B on one T4 16 GB; Mixtral 8x22B; DBRX; 2–4 low-cost GPUs; baselines=state-of-the-art offloading systems; FlexGen; metrics=throughput; CPU memory; resource utilization; results=up to 10.3x throughput; throughput bound with 2–3x less CPU memory Interpretation: A hierarchy-aware analytic model can choose paging and overlap plans that saturate the true bottleneck. Reusable lesson: Model every tier's roofline, then pipeline transfers and compute to the limiting resource. Applicability: Batch MoE inference on low-cost, memory-constrained GPUs. Limits: Optimized for batch throughput; results depend on model routing, batch size, I/O tiers, and T4/L4-class systems.



근거 canonical-paper-v2-e51b6b29: Shiyi Cao et al., "MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs", ASPLOS 2025. (원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:30:01.398033Z
Verification: full_text; confidence=high. Canonical title: MoE-Lightning: High-Throughput MoE Inference on Memory-constrained GPUs Question: How can large MoE models achieve high batch throughput on GPUs too small to hold their expert weights? Context: Offloading is dominated by hierarchical CPU/GPU/I/O bottlenecks and poor overlap among expert weight transfer, compute, and KV state. Method: MoE-Lightning uses CGOPipe CPU-GPU-I/O pipelining with paged weights and a Hierarchical Roofline Model to choose higher-throughput placement/schedules. Evaluation: workloads=Mixtral 8x7B on one T4 16 GB; Mixtral 8x22B; DBRX; 2–4 low-cost GPUs; baselines=state-of-the-art offloading systems; FlexGen; metrics=throughput; CPU memory; resource utilization; results=up to 10.3x throughput; throughput bound with 2–3x less CPU memory Interpretation: A hierarchy-aware analytic model can choose paging and overlap plans that saturate the true bottleneck. Reusable lesson: Model every tier's roofline, then pipeline transfers and compute to the limiting resource. Applicability: Batch MoE inference on low-cost, memory-constrained GPUs. Limits: Optimized for batch throughput; results depend on model routing, batch size, I/O tiers, and T4/L4-class systems.



자료 검증 verify_081937c918b6f35c0c1a: ev_201e1dcb4b1a4c77 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:45.142288Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_b2a4bda7888053fae1f3: verified-content-v1-0101 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:45.499598Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_5a0b1c6eaba183941a47: canonical-paper-v2-e51b6b29 · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T14:58:45.895839Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=932c469440fb76dfb3466974f55ea2ad537c539c4086bc9da25cc39108b6ca1a; independently adjudicated claim-bearing primary source / 위치: Abstract; Section 1; Section 3 HRM; Section 4 CGOPipe and HRM policy; Section 5 evaluation.
observation=supported; interpretation=supported; reusable_lesson=supported