Lesson:unicom a universally high performant i o completion mechanism for modern computer systems 9172fe3e
| 제목 | UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems |
|---|---|
| 궁금했던 점 | Can one I/O completion mechanism perform well under both CPU-idle and CPU-saturated conditions? |
| 해본 것 | UnICom combines TagSched, TagPoll, and SKIP to switch/coordinate polling and interrupts within Linux. |
| 당시 조건 | Venue: FAST. Year: 2026.
Polling gives low latency but burns CPU; interrupts conserve CPU but add latency and scheduling overhead. Verification: official USENIX page and paper PDF/abstract; confidence=high. |
| 실제 결과 | workloads=Linux 6.5.1, Intel Core i9-14900K의 16 E-core, Intel Optane P5801X(주 평가)와 Kingston NV3에서 수행한 4/128 KiB direct-I/O microbenchmarks, Destor+stress-ng 혼합 workload, RocksDB YCSB(500M KV load, 10M operations, 64/200-byte values); baselines=ext4 interrupt, BypassD polling, microbenchmark의 io_uring SQ_POLL; metrics=IOPS, average/p99 latency, co-running compute throughput, application throughput; results=CPU-only 경쟁이 없을 때 4 KiB read/write IOPS가 ext4보다 평균 43.5%/34.9% 높았고 single-thread read latency는 4 KiB 42%, 128 KiB 17.4% 낮았다. CPU 경쟁 시 4 KiB read IOPS는 ext4보다 39.4%, BypassD보다 88.8% 높았지만 dedicated completion core 때문에 compute-thread 성능이 약 7.5% 낮았다. 32 C-thread에서는 BypassD 대비 82.7% 개선했다. RocksDB YCSB single-thread는 ext4 대비 64/200-byte value에서 평균 24%/28% 높았고, 32-thread에서는 BypassD 대비 34%/56% 높았다. |
| 왜 그랬는지 | Completion policy should adapt to CPU pressure rather than choosing polling or interrupts globally. |
| 다음에 기억할 것 | Unify completion modes behind load-aware scheduling and preserve per-request identity. |
| 언제 맞는지 | CPU load가 크게 변하는 Linux direct-I/O storage path와 low-latency NVMe workload에 적용합니다.
Limits: Prototype은 ext4 위 direct I/O만 지원하고 buffered I/O는 기존 POSIX path로 돌아갑니다. 전용 completion core 하나를 사용하며 단일 completion thread의 측정 상한은 약 1.82 MIOPS입니다. 대형 I/O 또는 CPU 여유가 큰 일부 조건에서는 ext4/BypassD보다 compute throughput이 낮을 수 있고, 주 평가는 Optane P5801X와 16 E-core 한 구성에 집중합니다. |
| 신뢰도 | 낮음 |
| 관련 자료 | UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems. FAST 2026. |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T15:06:02.388184Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T15:00:45.589831Z |
근거 ev_9389fb701aa74111: UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems. FAST 2026.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:06:03.833796Z
Bibliographic paper record.
근거 verified-content-v1-0150: Riwei Pan; Yu Liang; Sam H. Noh; Lei Li; Nan Guan; Tei-Wei Kuo; Chun Jason Xue. UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems. FAST, 2026.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:58:34.601338Z
Verification: official USENIX page and paper PDF/abstract; confidence=high.
Canonical title: UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems
Question: Can one I/O completion mechanism perform well under both CPU-idle and CPU-saturated conditions?
Context: Polling gives low latency but burns CPU; interrupts conserve CPU but add latency and scheduling overhead.
Method: UnICom combines TagSched, TagPoll, and SKIP to switch/coordinate polling and interrupts within Linux.
Evaluation: workloads=low- and high-CPU-load I/O workloads; baselines=ext4, BypassD, io_uring; metrics=I/O performance and CPU use; results=consistently matches or exceeds the best baseline qualitatively
Interpretation: Completion policy should adapt to CPU pressure rather than choosing polling or interrupts globally.
Reusable lesson: Unify completion modes behind load-aware scheduling and preserve per-request identity.
Applicability: Modern Linux storage across variable CPU utilization.
Limits: Benefit depends on load classification and evaluated kernel/device stack; exact per-workload figures not extracted.
근거 canonical-paper-v2-9172fe3e: Riwei Pan; Yu Liang; Sam H. Noh; Lei Li; Nan Guan; Tei-Wei Kuo; Chun Jason Xue. UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems. FAST, 2026.
(원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:41:57.227492Z
Verification: official USENIX page and paper PDF/abstract; confidence=high.
Canonical title: UnICom: A Universally High-Performant I/O Completion Mechanism for Modern Computer Systems
Question: Can one I/O completion mechanism perform well under both CPU-idle and CPU-saturated conditions?
Context: Polling gives low latency but burns CPU; interrupts conserve CPU but add latency and scheduling overhead.
Method: UnICom combines TagSched, TagPoll, and SKIP to switch/coordinate polling and interrupts within Linux.
Evaluation: workloads=low- and high-CPU-load I/O workloads; baselines=ext4, BypassD, io_uring; metrics=I/O performance and CPU use; results=consistently matches or exceeds the best baseline qualitatively
Interpretation: Completion policy should adapt to CPU pressure rather than choosing polling or interrupts globally.
Reusable lesson: Unify completion modes behind load-aware scheduling and preserve per-request identity.
Applicability: Modern Linux storage across variable CPU utilization.
Limits: Benefit depends on load classification and evaluated kernel/device stack; exact per-workload figures not extracted.
자료 검증 verify_22dde943d7a3f18bd0d0:
ev_9389fb701aa74111 ·
지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T15:00:44.752429Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=3ab592f1ba711ab2058e1f56251244727caf1a87a184f8510895bd4382ce4225; independently adjudicated claim-bearing primary source / 위치: PDF abstract and §§3–5; pp. 10–14, §6 Evaluation.
observation=supported; interpretation=supported; reusable_lesson=supported
자료 검증 verify_50542892c838e01d00b0:
verified-content-v1-0150 ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:44.966021Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=7c74c1aff9b6fc47518026e8f9fc8cef2a18aa69b8b66231e6f2cea4b5fd4238 / 위치: 보존 파일 objects/sha256/7c/7c74c1aff9b6fc47518026e8f9fc8cef2a18aa69b8b66231e6f2cea4b5fd4238
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.
자료 검증 verify_a2cc980c19606841b475:
canonical-paper-v2-9172fe3e ·
판단 보류
확인 범위: 일부 자료 확인 · 주장: context · S3ResearchAgent · 2026-07-18T15:00:45.167552Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=7c74c1aff9b6fc47518026e8f9fc8cef2a18aa69b8b66231e6f2cea4b5fd4238 / 위치: 보존 파일 objects/sha256/7c/7c74c1aff9b6fc47518026e8f9fc8cef2a18aa69b8b66231e6f2cea4b5fd4238
보존 원문 객체를 확보했으나 이 일괄 검증에서는 claim-bearing 범위를 재판정하지 않아 결론을 보류함.
자료 검증 verify_0fdd511e2888f334d7be:
ev_9389fb701aa74111 ·
반박함
확인 범위: 원문 확인 · 주장: observation,applicability · S3ResearchAgent · 2026-07-18T15:00:45.377244Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=3ab592f1ba711ab2058e1f56251244727caf1a87a184f8510895bd4382ce4225; adjudicated detailed correction / 위치: DETAILED-CD-012 PDF §§3-6 pre-revision source locator.
Current claims are contradicted or exceed the preserved source; revision is staged at low confidence.