Lesson:usher holistic interference avoidance for resource optimized ml inference 74231fc1
| 제목 | USHER: Holistic Interference Avoidance for Resource Optimized ML Inference |
|---|---|
| 궁금했던 점 | How can GPU inference services spatially multiplex models while meeting latency SLOs and avoiding cache/compute/memory interference? |
| 해본 것 | USHER combines GPU-kernel resource estimation, an interference-aware scheduler for batch size/replication/placement, and operator-graph merging to reduce cache interference. |
| 당시 조건 | Venue: OSDI. Year: 2024.
Naive colocation raises utilization but unpredictable inter-model interference lowers goodput and cost efficiency. Verification: official_abstract; confidence=high. |
| 실제 결과 | workloads=large-scale production ML inference workloads; baselines=existing GPU inference multiplexing methods; metrics=goodput; cost efficiency; latency-SLO compliance; scale; results=up to 2.6x goodput; up to 3.5x cost efficiency; scales to thousands of GPUs |
| 왜 그랬는지 | Resource estimation, placement, and graph-level execution must be coordinated to make GPU multiplexing economical. |
| 다음에 기억할 것 | Model interference explicitly and optimize configuration and execution structure together. |
| 언제 맞는지 | Large multi-model GPU inference fleets with latency SLOs.
Limits: Estimator and heuristic accuracy depend on model/operator mix, workload stationarity, GPU architecture, and SLOs. |
| 신뢰도 | 중간 |
| 관련 자료 | Sudipta Saha Shubha et al., "USHER: Holistic Interference Avoidance for Resource Optimized ML Inference", OSDI 2024.
Source: https://www.usenix.org/conference/osdi24/presentation/shubha Verification basis: official_abstract. Canonical evidence ID: canonical-paper-v2-74231fc1 |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T14:57:58.733068Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T15:00:45.519269Z |
근거 ev_9f7ba2fec55c41ea: Usher: Holistic Interference Avoidance for Resource Optimized ML Inference. OSDI 2024.
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T14:57:59.713745Z
Bibliographic paper record.
근거 verified-content-v1-0096: Sudipta Saha Shubha et al., "USHER: Holistic Interference Avoidance for Resource Optimized ML Inference", OSDI 2024.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:55:05.126123Z
Verification: official_abstract; confidence=high.
Canonical title: USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
Question: How can GPU inference services spatially multiplex models while meeting latency SLOs and avoiding cache/compute/memory interference?
Context: Naive colocation raises utilization but unpredictable inter-model interference lowers goodput and cost efficiency.
Method: USHER combines GPU-kernel resource estimation, an interference-aware scheduler for batch size/replication/placement, and operator-graph merging to reduce cache interference.
Evaluation: workloads=large-scale production ML inference workloads; baselines=existing GPU inference multiplexing methods; metrics=goodput; cost efficiency; latency-SLO compliance; scale; results=up to 2.6x goodput; up to 3.5x cost efficiency; scales to thousands of GPUs
Interpretation: Resource estimation, placement, and graph-level execution must be coordinated to make GPU multiplexing economical.
Reusable lesson: Model interference explicitly and optimize configuration and execution structure together.
Applicability: Large multi-model GPU inference fleets with latency SLOs.
Limits: Estimator and heuristic accuracy depend on model/operator mix, workload stationarity, GPU architecture, and SLOs.
근거 canonical-paper-v2-74231fc1: Sudipta Saha Shubha et al., "USHER: Holistic Interference Avoidance for Resource Optimized ML Inference", OSDI 2024.
(원문 열기)
논문 · 확인 범위: 공식 초록 확인 · S3ResearchAgent · 2026-07-18T05:42:11.351444Z
Verification: official_abstract; confidence=medium.
Canonical title: USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
Question: How can GPU inference services spatially multiplex models while meeting latency SLOs and avoiding cache/compute/memory interference?
Context: Naive colocation raises utilization but unpredictable inter-model interference lowers goodput and cost efficiency.
Method: USHER combines GPU-kernel resource estimation, an interference-aware scheduler for batch size/replication/placement, and operator-graph merging to reduce cache interference.
Evaluation: workloads=large-scale production ML inference workloads; baselines=existing GPU inference multiplexing methods; metrics=goodput; cost efficiency; latency-SLO compliance; scale; results=up to 2.6x goodput; up to 3.5x cost efficiency; scales to thousands of GPUs
Interpretation: Resource estimation, placement, and graph-level execution must be coordinated to make GPU multiplexing economical.
Reusable lesson: Model interference explicitly and optimize configuration and execution structure together.
Applicability: Large multi-model GPU inference fleets with latency SLOs.
Limits: Estimator and heuristic accuracy depend on model/operator mix, workload stationarity, GPU architecture, and SLOs.
자료 검증 verify_2ea5f18904f35e3a1161:
ev_9f7ba2fec55c41ea ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.145546Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f / 위치: 보존 파일 objects/sha256/35/35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.
자료 검증 verify_e2b33519bc58b78f7462:
verified-content-v1-0096 ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.345933Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f / 위치: 보존 파일 objects/sha256/35/35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.
자료 검증 verify_004348df14156d324c62:
canonical-paper-v2-74231fc1 ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.519269Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f / 위치: 보존 파일 objects/sha256/35/35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.