Lesson:lithos an operating system for efficient machine learning on gpus 76d95512
| 제목 | LithOS: An Operating System for Efficient Machine Learning on GPUs |
|---|---|
| 궁금했던 점 | Can a GPU-native OS isolate and co-schedule ML kernels more efficiently than host-managed sharing? |
| 해본 것 | LithOS atomizes kernels, schedules and steals thread-blocks, rightsizes allocations, and manages GPU power. |
| 당시 조건 | Venue: SOSP. Year: 2025.
MPS and existing schedulers have coarse control over thread-block execution, interference, and energy/capacity tradeoffs. Verification: DOI metadata and the official SOSP program are recorded in the preserved review ledger; claim-bearing O/I/R source coverage remains pending; confidence=low. |
| 실제 결과 | workloads=stacked inference and hybrid ML workloads; baselines=NVIDIA MPS and prior GPU schedulers; metrics=p99 latency, throughput, capacity, energy; results=13x/3x lower p99 and 1.6x throughput; hybrid 4.7x/1.18x p99 and 1.35x throughput |
| 왜 그랬는지 | Thread-block-level OS control exposes isolation and packing opportunities hidden from host schedulers. |
| 다음에 기억할 것 | Schedule accelerators at their native execution unit and jointly manage latency, throughput, and power. |
| 언제 맞는지 | Multi-tenant GPUs running inference and mixed ML jobs.
Limits: Requires a specialized GPU OS/runtime and kernel atomization; gains vary by workload interference. |
| 신뢰도 | 낮음 |
| 관련 자료 | LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP 2025. |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T15:05:19.230121Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T15:00:27.847039Z |
근거 ev_d3bfc13791444366: LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP 2025 workshop 2025.
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:05:20.419175Z
Bibliographic paper record.
근거 verified-content-v1-0144: Patrick H. Coppock; Brian Zhang; Eliot H. Solomon; Vasileios Kypriotis; Leon Yang; Bikash Sharma; Dan Schatzberg; Todd C. Mowry.... LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP, 2025.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:58:09.291932Z
Verification: latest arXiv abstract/full text, DOI, and official SOSP program; confidence=high.
Canonical title: LithOS: An Operating System for Efficient Machine Learning on GPUs
Question: Can a GPU-native OS isolate and co-schedule ML kernels more efficiently than host-managed sharing?
Context: MPS and existing schedulers have coarse control over thread-block execution, interference, and energy/capacity tradeoffs.
Method: LithOS atomizes kernels, schedules and steals thread-blocks, rightsizes allocations, and manages GPU power.
Evaluation: workloads=stacked inference and hybrid ML workloads; baselines=NVIDIA MPS and prior GPU schedulers; metrics=p99 latency, throughput, capacity, energy; results=13x/3x lower p99 and 1.6x throughput; hybrid 4.7x/1.18x p99 and 1.35x throughput
Interpretation: Thread-block-level OS control exposes isolation and packing opportunities hidden from host schedulers.
Reusable lesson: Schedule accelerators at their native execution unit and jointly manage latency, throughput, and power.
Applicability: Multi-tenant GPUs running inference and mixed ML jobs.
Limits: Requires a specialized GPU OS/runtime and kernel atomization; gains vary by workload interference.
근거 canonical-paper-v2-76d95512: Patrick H. Coppock; Brian Zhang; Eliot H. Solomon; Vasileios Kypriotis; Leon Yang; Bikash Sharma; Dan Schatzberg; Todd C. Mowry. LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP, 2025.
(원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:27:49.407003Z
Verification: latest arXiv abstract/full text, DOI, and official SOSP program; confidence=high.
Canonical title: LithOS: An Operating System for Efficient Machine Learning on GPUs
Question: Can a GPU-native OS isolate and co-schedule ML kernels more efficiently than host-managed sharing?
Context: MPS and existing schedulers have coarse control over thread-block execution, interference, and energy/capacity tradeoffs.
Method: LithOS atomizes kernels, schedules and steals thread-blocks, rightsizes allocations, and manages GPU power.
Evaluation: workloads=stacked inference and hybrid ML workloads; baselines=NVIDIA MPS and prior GPU schedulers; metrics=p99 latency, throughput, capacity, energy; results=13x/3x lower p99 and 1.6x throughput; hybrid 4.7x/1.18x p99 and 1.35x throughput
Interpretation: Thread-block-level OS control exposes isolation and packing opportunities hidden from host schedulers.
Reusable lesson: Schedule accelerators at their native execution unit and jointly manage latency, throughput, and power.
Applicability: Multi-tenant GPUs running inference and mixed ML jobs.
Limits: Requires a specialized GPU OS/runtime and kernel atomization; gains vary by workload interference.
자료 검증 verify_0abd17fac42d3ad18e30:
ev_d3bfc13791444366 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:26.696259Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf / 위치: 보존 파일 objects/sha256/de/dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.
자료 검증 verify_bb06861ca41a6396b76b:
verified-content-v1-0144 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:26.857646Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf / 위치: 보존 파일 objects/sha256/de/dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.
자료 검증 verify_a2e7bc405dc333b62b62:
canonical-paper-v2-76d95512 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:27.113509Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf / 위치: 보존 파일 objects/sha256/de/dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.
자료 검증 verify_a1191fe7bfa5675760f6:
canonical-paper-v2-76d95512 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T15:00:27.283292Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf; high-confidence adjudication ledger / 위치: Saved HTML inspected in full; it is a cookie/landing/generic metadata page and contains no claim-bearing text covering O/I/R.
claim-bearing O/I/R coverage was not established; confidence forced to low
자료 검증 verify_9a378e2838e449c97fe4:
canonical-paper-v2-76d95512 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:27.651457Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=a140f3572e510459b9eed5bf42d4e2096c4684fd4f1ee746960d6675661d1077; adjudicated correction / 위치: content-debt-review CD-002
CD-002 curated source-bound correction; primary locator recorded in review.