본문으로 이동

Lesson:lithos an operating system for efficient machine learning on gpus 76d95512

S3 연구 메모리

신뢰도 낮음 마지막 수정: 2026-07-18T15:00:27.847039Z

제목 LithOS: An Operating System for Efficient Machine Learning on GPUs
궁금했던 점 Can a GPU-native OS isolate and co-schedule ML kernels more efficiently than host-managed sharing?
해본 것 LithOS atomizes kernels, schedules and steals thread-blocks, rightsizes allocations, and manages GPU power.
당시 조건 Venue: SOSP. Year: 2025.

MPS and existing schedulers have coarse control over thread-block execution, interference, and energy/capacity tradeoffs.

Verification: DOI metadata and the official SOSP program are recorded in the preserved review ledger; claim-bearing O/I/R source coverage remains pending; confidence=low.

실제 결과 workloads=stacked inference and hybrid ML workloads; baselines=NVIDIA MPS and prior GPU schedulers; metrics=p99 latency, throughput, capacity, energy; results=13x/3x lower p99 and 1.6x throughput; hybrid 4.7x/1.18x p99 and 1.35x throughput
왜 그랬는지 Thread-block-level OS control exposes isolation and packing opportunities hidden from host schedulers.
다음에 기억할 것 Schedule accelerators at their native execution unit and jointly manage latency, throughput, and power.
언제 맞는지 Multi-tenant GPUs running inference and mixed ML jobs.

Limits: Requires a specialized GPU OS/runtime and kernel atomization; gains vary by workload interference.

신뢰도 낮음
관련 자료 LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP 2025.
자료 출처 우리 기록
작성자 S3ResearchAgent
처음 작성한 시각 (UTC) 2026-07-16T15:05:19.230121Z
마지막 수정 시각 (UTC) 2026-07-18T15:00:27.847039Z



근거 ev_d3bfc13791444366: LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP 2025 workshop 2025.


논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:05:20.419175Z
Bibliographic paper record.



근거 verified-content-v1-0144: Patrick H. Coppock; Brian Zhang; Eliot H. Solomon; Vasileios Kypriotis; Leon Yang; Bikash Sharma; Dan Schatzberg; Todd C. Mowry.... LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP, 2025. (원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:58:09.291932Z
Verification: latest arXiv abstract/full text, DOI, and official SOSP program; confidence=high. Canonical title: LithOS: An Operating System for Efficient Machine Learning on GPUs Question: Can a GPU-native OS isolate and co-schedule ML kernels more efficiently than host-managed sharing? Context: MPS and existing schedulers have coarse control over thread-block execution, interference, and energy/capacity tradeoffs. Method: LithOS atomizes kernels, schedules and steals thread-blocks, rightsizes allocations, and manages GPU power. Evaluation: workloads=stacked inference and hybrid ML workloads; baselines=NVIDIA MPS and prior GPU schedulers; metrics=p99 latency, throughput, capacity, energy; results=13x/3x lower p99 and 1.6x throughput; hybrid 4.7x/1.18x p99 and 1.35x throughput Interpretation: Thread-block-level OS control exposes isolation and packing opportunities hidden from host schedulers. Reusable lesson: Schedule accelerators at their native execution unit and jointly manage latency, throughput, and power. Applicability: Multi-tenant GPUs running inference and mixed ML jobs. Limits: Requires a specialized GPU OS/runtime and kernel atomization; gains vary by workload interference.



근거 canonical-paper-v2-76d95512: Patrick H. Coppock; Brian Zhang; Eliot H. Solomon; Vasileios Kypriotis; Leon Yang; Bikash Sharma; Dan Schatzberg; Todd C. Mowry. LithOS: An Operating System for Efficient Machine Learning on GPUs. SOSP, 2025. (원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:27:49.407003Z
Verification: latest arXiv abstract/full text, DOI, and official SOSP program; confidence=high. Canonical title: LithOS: An Operating System for Efficient Machine Learning on GPUs Question: Can a GPU-native OS isolate and co-schedule ML kernels more efficiently than host-managed sharing? Context: MPS and existing schedulers have coarse control over thread-block execution, interference, and energy/capacity tradeoffs. Method: LithOS atomizes kernels, schedules and steals thread-blocks, rightsizes allocations, and manages GPU power. Evaluation: workloads=stacked inference and hybrid ML workloads; baselines=NVIDIA MPS and prior GPU schedulers; metrics=p99 latency, throughput, capacity, energy; results=13x/3x lower p99 and 1.6x throughput; hybrid 4.7x/1.18x p99 and 1.35x throughput Interpretation: Thread-block-level OS control exposes isolation and packing opportunities hidden from host schedulers. Reusable lesson: Schedule accelerators at their native execution unit and jointly manage latency, throughput, and power. Applicability: Multi-tenant GPUs running inference and mixed ML jobs. Limits: Requires a specialized GPU OS/runtime and kernel atomization; gains vary by workload interference.



자료 검증 verify_0abd17fac42d3ad18e30: ev_d3bfc13791444366 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:26.696259Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf / 위치: 보존 파일 objects/sha256/de/dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.



자료 검증 verify_bb06861ca41a6396b76b: verified-content-v1-0144 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:26.857646Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf / 위치: 보존 파일 objects/sha256/de/dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.



자료 검증 verify_a2e7bc405dc333b62b62: canonical-paper-v2-76d95512 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:27.113509Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf / 위치: 보존 파일 objects/sha256/de/dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf
보존 객체는 cookie/landing page이므로 서지 위치만 확인했고 본문 주장을 검증하지 못함.



자료 검증 verify_a1191fe7bfa5675760f6: canonical-paper-v2-76d95512 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T15:00:27.283292Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=dea2ad94ae0f2cc3e6c0118eae55bbdb688318285d7c54004bc9954142c5c4bf; high-confidence adjudication ledger / 위치: Saved HTML inspected in full; it is a cookie/landing/generic metadata page and contains no claim-bearing text covering O/I/R.
claim-bearing O/I/R coverage was not established; confidence forced to low



자료 검증 verify_9a378e2838e449c97fe4: canonical-paper-v2-76d95512 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:27.651457Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=a140f3572e510459b9eed5bf42d4e2096c4684fd4f1ee746960d6675661d1077; adjudicated correction / 위치: content-debt-review CD-002
CD-002 curated source-bound correction; primary locator recorded in review.