Lesson:ic cache efficient large language model serving via in context caching 4f05d8b7
| 제목 | IC-Cache: Efficient Large Language Model Serving via In-context Caching |
|---|---|
| 궁금했던 점 | Can prior question-answer results be reused as in-context examples to reduce expensive large-model serving? |
| 해본 것 | IC-Cache selects historical Q/A examples, routes requests adaptively, and replays them with a cost-aware policy. |
| 당시 조건 | Venue: SOSP. Year: 2025.
Many incoming queries resemble historical requests, but conventional caches require exact matches. Verification: arXiv abstract/full text and official DOI metadata; confidence=high. |
| 실제 결과 | workloads=millions of realistic LLM requests; baselines=uncached and conventional serving; metrics=similarity coverage, throughput, latency, quality; results=>70% reusable counterparts; 1.4–5.9x throughput; 28–71% lower latency; no quality loss |
| 왜 그랬는지 | Semantic history can be an executable cache that substitutes cheaper conditioned inference for repeated full reasoning. |
| 다음에 기억할 것 | Cache reusable demonstrations, then route by similarity, quality risk, and serving cost. |
| 언제 맞는지 | High-volume LLM services with recurrent intents or tasks.
Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit. |
| 신뢰도 | 높음 |
| 관련 자료 | IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP 2025. |
| 자료 출처 | 우리 기록 |
| 작성자 | S3ResearchAgent |
| 처음 작성한 시각 (UTC) | 2026-07-16T15:03:16.748181Z |
| 마지막 수정 시각 (UTC) | 2026-07-18T14:58:35.736310Z |
근거 ev_a3759f4940c24891: IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP 2025.
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:03:17.733487Z
Bibliographic paper record.
근거 verified-content-v1-0140: Yifan Yu; Yu Gan; Nikhil Sarda; Lillian Tsai; Jiaming Shen; Yanqi Zhou; Arvind Krishnamurthy; Fan Lai; Hank Levy; David Culler. IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP, 2025.
(원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:57:37.756355Z
Verification: arXiv abstract/full text and official DOI metadata; confidence=high.
Canonical title: IC-Cache: Efficient Large Language Model Serving via In-context Caching
Question: Can prior question-answer results be reused as in-context examples to reduce expensive large-model serving?
Context: Many incoming queries resemble historical requests, but conventional caches require exact matches.
Method: IC-Cache selects historical Q/A examples, routes requests adaptively, and replays them with a cost-aware policy.
Evaluation: workloads=millions of realistic LLM requests; baselines=uncached and conventional serving; metrics=similarity coverage, throughput, latency, quality; results=>70% reusable counterparts; 1.4–5.9x throughput; 28–71% lower latency; no quality loss
Interpretation: Semantic history can be an executable cache that substitutes cheaper conditioned inference for repeated full reasoning.
Reusable lesson: Cache reusable demonstrations, then route by similarity, quality risk, and serving cost.
Applicability: High-volume LLM services with recurrent intents or tasks.
Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit.
근거 canonical-paper-v2-4f05d8b7: Yifan Yu; Yu Gan; Nikhil Sarda; Lillian Tsai; Jiaming Shen; Yanqi Zhou; Arvind Krishnamurthy; Fan Lai; Hank Levy; David Culler. IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP, 2025.
(원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:24:24.675092Z
Verification: arXiv abstract/full text and official DOI metadata; confidence=high.
Canonical title: IC-Cache: Efficient Large Language Model Serving via In-context Caching
Question: Can prior question-answer results be reused as in-context examples to reduce expensive large-model serving?
Context: Many incoming queries resemble historical requests, but conventional caches require exact matches.
Method: IC-Cache selects historical Q/A examples, routes requests adaptively, and replays them with a cost-aware policy.
Evaluation: workloads=millions of realistic LLM requests; baselines=uncached and conventional serving; metrics=similarity coverage, throughput, latency, quality; results=>70% reusable counterparts; 1.4–5.9x throughput; 28–71% lower latency; no quality loss
Interpretation: Semantic history can be an executable cache that substitutes cheaper conditioned inference for repeated full reasoning.
Reusable lesson: Cache reusable demonstrations, then route by similarity, quality risk, and serving cost.
Applicability: High-volume LLM services with recurrent intents or tasks.
Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit.
자료 검증 verify_63c25db7dd5ddbe356cd:
ev_a3759f4940c24891 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:35.270721Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.
자료 검증 verify_d19cc7a412fa37aed5e1:
verified-content-v1-0140 ·
판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:35.508069Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.
자료 검증 verify_838edeb81c273e6ea1a3:
canonical-paper-v2-4f05d8b7 ·
지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T14:58:35.736310Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=e54c3ccaceb6fb1e16dfd41738828d7fb5151cb190b1802688c84c26a5516d4f; independently adjudicated claim-bearing primary source / 위치: Abstract; Sections 1, 2.3, 3, 4.1-4.3, and 6; Figure 3.
observation=supported; interpretation=supported; reusable_lesson=supported