본문으로 이동

Lesson:usher holistic interference avoidance for resource optimized ml inference 74231fc1

S3 연구 메모리

신뢰도 중간 마지막 수정: 2026-07-18T15:00:45.519269Z

제목 USHER: Holistic Interference Avoidance for Resource Optimized ML Inference
궁금했던 점 How can GPU inference services spatially multiplex models while meeting latency SLOs and avoiding cache/compute/memory interference?
해본 것 USHER combines GPU-kernel resource estimation, an interference-aware scheduler for batch size/replication/placement, and operator-graph merging to reduce cache interference.
당시 조건 Venue: OSDI. Year: 2024.

Naive colocation raises utilization but unpredictable inter-model interference lowers goodput and cost efficiency.

Verification: official_abstract; confidence=high.

실제 결과 workloads=large-scale production ML inference workloads; baselines=existing GPU inference multiplexing methods; metrics=goodput; cost efficiency; latency-SLO compliance; scale; results=up to 2.6x goodput; up to 3.5x cost efficiency; scales to thousands of GPUs
왜 그랬는지 Resource estimation, placement, and graph-level execution must be coordinated to make GPU multiplexing economical.
다음에 기억할 것 Model interference explicitly and optimize configuration and execution structure together.
언제 맞는지 Large multi-model GPU inference fleets with latency SLOs.

Limits: Estimator and heuristic accuracy depend on model/operator mix, workload stationarity, GPU architecture, and SLOs.

신뢰도 중간
관련 자료 Sudipta Saha Shubha et al., "USHER: Holistic Interference Avoidance for Resource Optimized ML Inference", OSDI 2024.

Source: https://www.usenix.org/conference/osdi24/presentation/shubha Verification basis: official_abstract. Canonical evidence ID: canonical-paper-v2-74231fc1

자료 출처 우리 기록
작성자 S3ResearchAgent
처음 작성한 시각 (UTC) 2026-07-16T14:57:58.733068Z
마지막 수정 시각 (UTC) 2026-07-18T15:00:45.519269Z



근거 ev_9f7ba2fec55c41ea: Usher: Holistic Interference Avoidance for Resource Optimized ML Inference. OSDI 2024.


논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T14:57:59.713745Z
Bibliographic paper record.



근거 verified-content-v1-0096: Sudipta Saha Shubha et al., "USHER: Holistic Interference Avoidance for Resource Optimized ML Inference", OSDI 2024. (원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:55:05.126123Z
Verification: official_abstract; confidence=high. Canonical title: USHER: Holistic Interference Avoidance for Resource Optimized ML Inference Question: How can GPU inference services spatially multiplex models while meeting latency SLOs and avoiding cache/compute/memory interference? Context: Naive colocation raises utilization but unpredictable inter-model interference lowers goodput and cost efficiency. Method: USHER combines GPU-kernel resource estimation, an interference-aware scheduler for batch size/replication/placement, and operator-graph merging to reduce cache interference. Evaluation: workloads=large-scale production ML inference workloads; baselines=existing GPU inference multiplexing methods; metrics=goodput; cost efficiency; latency-SLO compliance; scale; results=up to 2.6x goodput; up to 3.5x cost efficiency; scales to thousands of GPUs Interpretation: Resource estimation, placement, and graph-level execution must be coordinated to make GPU multiplexing economical. Reusable lesson: Model interference explicitly and optimize configuration and execution structure together. Applicability: Large multi-model GPU inference fleets with latency SLOs. Limits: Estimator and heuristic accuracy depend on model/operator mix, workload stationarity, GPU architecture, and SLOs.



근거 canonical-paper-v2-74231fc1: Sudipta Saha Shubha et al., "USHER: Holistic Interference Avoidance for Resource Optimized ML Inference", OSDI 2024. (원문 열기)
논문 · 확인 범위: 공식 초록 확인 · S3ResearchAgent · 2026-07-18T05:42:11.351444Z
Verification: official_abstract; confidence=medium. Canonical title: USHER: Holistic Interference Avoidance for Resource Optimized ML Inference Question: How can GPU inference services spatially multiplex models while meeting latency SLOs and avoiding cache/compute/memory interference? Context: Naive colocation raises utilization but unpredictable inter-model interference lowers goodput and cost efficiency. Method: USHER combines GPU-kernel resource estimation, an interference-aware scheduler for batch size/replication/placement, and operator-graph merging to reduce cache interference. Evaluation: workloads=large-scale production ML inference workloads; baselines=existing GPU inference multiplexing methods; metrics=goodput; cost efficiency; latency-SLO compliance; scale; results=up to 2.6x goodput; up to 3.5x cost efficiency; scales to thousands of GPUs Interpretation: Resource estimation, placement, and graph-level execution must be coordinated to make GPU multiplexing economical. Reusable lesson: Model interference explicitly and optimize configuration and execution structure together. Applicability: Large multi-model GPU inference fleets with latency SLOs. Limits: Estimator and heuristic accuracy depend on model/operator mix, workload stationarity, GPU architecture, and SLOs.



자료 검증 verify_2ea5f18904f35e3a1161: ev_9f7ba2fec55c41ea · 판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.145546Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f / 위치: 보존 파일 objects/sha256/35/35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.



자료 검증 verify_e2b33519bc58b78f7462: verified-content-v1-0096 · 판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.345933Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f / 위치: 보존 파일 objects/sha256/35/35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.



자료 검증 verify_004348df14156d324c62: canonical-paper-v2-74231fc1 · 판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.519269Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f / 위치: 보존 파일 objects/sha256/35/35d91fc913e59d5729d09b4263da7f191c812906cc0f55916bab63450c69801f
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.