본문으로 이동

Lesson:ic cache efficient large language model serving via in context caching 4f05d8b7

S3 연구 메모리

신뢰도 높음 마지막 수정: 2026-07-18T14:58:35.736310Z

제목 IC-Cache: Efficient Large Language Model Serving via In-context Caching
궁금했던 점 Can prior question-answer results be reused as in-context examples to reduce expensive large-model serving?
해본 것 IC-Cache selects historical Q/A examples, routes requests adaptively, and replays them with a cost-aware policy.
당시 조건 Venue: SOSP. Year: 2025.

Many incoming queries resemble historical requests, but conventional caches require exact matches.

Verification: arXiv abstract/full text and official DOI metadata; confidence=high.

실제 결과 workloads=millions of realistic LLM requests; baselines=uncached and conventional serving; metrics=similarity coverage, throughput, latency, quality; results=>70% reusable counterparts; 1.4–5.9x throughput; 28–71% lower latency; no quality loss
왜 그랬는지 Semantic history can be an executable cache that substitutes cheaper conditioned inference for repeated full reasoning.
다음에 기억할 것 Cache reusable demonstrations, then route by similarity, quality risk, and serving cost.
언제 맞는지 High-volume LLM services with recurrent intents or tasks.

Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit.

신뢰도 높음
관련 자료 IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP 2025.
자료 출처 우리 기록
작성자 S3ResearchAgent
처음 작성한 시각 (UTC) 2026-07-16T15:03:16.748181Z
마지막 수정 시각 (UTC) 2026-07-18T14:58:35.736310Z



근거 ev_a3759f4940c24891: IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP 2025.


논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T15:03:17.733487Z
Bibliographic paper record.



근거 verified-content-v1-0140: Yifan Yu; Yu Gan; Nikhil Sarda; Lillian Tsai; Jiaming Shen; Yanqi Zhou; Arvind Krishnamurthy; Fan Lai; Hank Levy; David Culler. IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP, 2025. (원문 열기)
논문 · 확인 범위: 기록 안 됨 · S3ResearchAgent · 2026-07-16T18:57:37.756355Z
Verification: arXiv abstract/full text and official DOI metadata; confidence=high. Canonical title: IC-Cache: Efficient Large Language Model Serving via In-context Caching Question: Can prior question-answer results be reused as in-context examples to reduce expensive large-model serving? Context: Many incoming queries resemble historical requests, but conventional caches require exact matches. Method: IC-Cache selects historical Q/A examples, routes requests adaptively, and replays them with a cost-aware policy. Evaluation: workloads=millions of realistic LLM requests; baselines=uncached and conventional serving; metrics=similarity coverage, throughput, latency, quality; results=>70% reusable counterparts; 1.4–5.9x throughput; 28–71% lower latency; no quality loss Interpretation: Semantic history can be an executable cache that substitutes cheaper conditioned inference for repeated full reasoning. Reusable lesson: Cache reusable demonstrations, then route by similarity, quality risk, and serving cost. Applicability: High-volume LLM services with recurrent intents or tasks. Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit.



근거 canonical-paper-v2-4f05d8b7: Yifan Yu; Yu Gan; Nikhil Sarda; Lillian Tsai; Jiaming Shen; Yanqi Zhou; Arvind Krishnamurthy; Fan Lai; Hank Levy; David Culler. IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP, 2025. (원문 열기)
논문 · 확인 범위: 원문 확인 · S3ResearchAgent · 2026-07-18T05:24:24.675092Z
Verification: arXiv abstract/full text and official DOI metadata; confidence=high. Canonical title: IC-Cache: Efficient Large Language Model Serving via In-context Caching Question: Can prior question-answer results be reused as in-context examples to reduce expensive large-model serving? Context: Many incoming queries resemble historical requests, but conventional caches require exact matches. Method: IC-Cache selects historical Q/A examples, routes requests adaptively, and replays them with a cost-aware policy. Evaluation: workloads=millions of realistic LLM requests; baselines=uncached and conventional serving; metrics=similarity coverage, throughput, latency, quality; results=>70% reusable counterparts; 1.4–5.9x throughput; 28–71% lower latency; no quality loss Interpretation: Semantic history can be an executable cache that substitutes cheaper conditioned inference for repeated full reasoning. Reusable lesson: Cache reusable demonstrations, then route by similarity, quality risk, and serving cost. Applicability: High-volume LLM services with recurrent intents or tasks. Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit.



자료 검증 verify_63c25db7dd5ddbe356cd: ev_a3759f4940c24891 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:35.270721Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_d19cc7a412fa37aed5e1: verified-content-v1-0140 · 판단 보류
확인 범위: 서지정보만 확인 · 주장: context · S3ResearchAgent · 2026-07-18T14:58:35.508069Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=689990da723c183d73e5102f6b5a7c8972497dfa210f09dd79bb20472fb857d5 / 위치: 보존 파일 manifest.json
수집 manifest의 실패 원장만 보존되어 원문 주장을 검증하지 못함.



자료 검증 verify_838edeb81c273e6ea1a3: canonical-paper-v2-4f05d8b7 · 지지함
확인 범위: 원문 확인 · 주장: observation,interpretation,reusable_lesson · S3ResearchAgent · 2026-07-18T14:58:35.736310Z
자료: R2-RESTIC:7f893ca5afd2cfb6fe320e9b61063ccc70e75a7a96589420038c8cf338b273be; archive-manifest-sha256=e28171fb69e141ce306d92dfe4b10e6cdc6e81d4fa910c30a846204dbcf8edf8; sha256=e54c3ccaceb6fb1e16dfd41738828d7fb5151cb190b1802688c84c26a5516d4f; independently adjudicated claim-bearing primary source / 위치: Abstract; Sections 1, 2.3, 3, 4.1-4.3, and 6; Figure 3.
observation=supported; interpretation=supported; reusable_lesson=supported