속성:Applicability
외관
이 메모가 맞는 조건과 예외를 적습니다.
g
현재는 GPU causal profiling 관련 서지 인덱스와 원문 확보 큐에만 사용한다. 커널 최적화 설계나 성능 예측의 의사결정 근거로 사용하면 안 된다. +
Transformer 생성 추론과 KV 캐시 절감.
Limits: 그룹 수 선택과 5% 추가 학습비가 필요하며 모델군별 일반화가 제한될 수 있다. +
클라우드 VM의 deduplication, ballooning/reprovisioning, tenant-level memory QoS에 직접 적용된다. +
h
fault-tolerant stateful FaaS.
Limits: 외부 상태 의미론과 workload 분류, 프로토콜 전환 정확성에 의존한다. +
Memory-bound server workloads with latency SLOs and spare CPU cycles.
Limits: Requires analyzable/instrumentable binaries and enough predictable memory stalls; abstract does not enumerate all workloads. +
LSM-tree stores with skewed key popularity.
Limits: Benefits depend on hot-key skew and tracker/promotion overhead. +
현재는 모바일 storage I/O 연구의 서지 추적과 poster 원문 확보 작업에만 사용한다. 앱 지연 최적화나 시스템 설계의 근거로는 사용하지 않는다. +
Copy-heavy file, network, IPC, and mobile OS paths.
Limits: Applications/subsystems must express asynchronous copy dependencies; hardware availability affects benefit. +
i
NVMe 최적화 DB·캐시·스토리지 엔진.
Limits: 범용 블록 서비스 일부를 우회하고 NVMe·peak FIO 중심 결과다. +
High-volume LLM services with recurrent intents or tasks.
Limits: Requires representative history and safe example selection; novel or distribution-shifted queries may not benefit. +
ice collaborating memory and process management for user experience on resource limited mobile d 312e453f +
Resource-limited Android/mobile devices under foreground/background memory pressure.
Limits: Freezing can delay background services; correct classification, thaw policy, and device/app mix affect UX and liveness. +
스토리지·락·스케줄링 대기가 섞인 Linux 서비스와 DB 성능 분석에 높음. +
impress an importance informed multi tier prefix kv storage system for large language model infe 51b47015 +
LLM applications with recurring long contexts and disk-backed prefix caches.
Limits: Accuracy and speed depend on importance stability across heads/models/tasks, prefix reuse, disk bandwidth, and tier capacity. +
개인별 앱 사용 순서가 반복되는 Android lifecycle 및 memory management에 해당한다. +
inf2 high throughput generative inference of large language models using near storage processing 3fcead73 +
Throughput 중심의 offline inference, 수만~128K token의 장문맥, 가중치와 KV cache가 GPU 또는 host DRAM에 여유 있게 들어가지 않는 decoder-only MHA/GQA/MoE 모델, batch benchmark와 대규모 정보 추출에 적용할 수 있습니다. Commercial SmartSSD 또는 programmable near-storage accelerator가 필요합니다. CXL 기반 near-data system에도 host traffic 최소화, host/device cooperative work, exact streaming attention이라는 원칙은 재사용할 수 있지만 HILOS 구현 자체는 PCIe 기반입니다.
Limits: strict-latency online serving의 TTFT/TPOT를 검증하지 않았습니다. 큰 성능 향상은 4–16개 SmartSSD 병렬 구성에 의존하며 전용 expansion chassis, GPUDirect Storage와 FPGA software stack이 필요합니다. 일부 모델의 128K 실험은 pretraining context를 넘으므로 system scaling만 검증하고 답변 품질을 보장하지 않습니다. 기본 FP16, batch 16, output 64 결과는 workload에 따라 달라집니다. DRAM이 충분하면 FLEX(DRAM)이 더 비용 효율적입니다. DeepSpeed 비교는 저자 추가 UVM 확장이고, conventional SSD energy는 측정값이 아니라 datasheet 값입니다. PCIe 5.0 SSD를 따라가려면 2,000개가 넘는 DSP가 필요하다는 분석처럼 현 FPGA에는 확장 한계가 있습니다. Softmax가 큰 attention group에서 실행 시간의 50% 이상을 차지하며, capacity 대비 bandwidth 불균형으로 SmartSSD당 4TB 중 평가 peak 사용량은 600GB 미만입니다. Cost와 endurance는 장기 운영 관측이 아니라 논문의 가격·PBW·retention 가정에 따른 모델입니다. +
infinigen efficient generative inference of large language models with dynamic kv cache manageme a1ca3228 +
긴 문맥 LLM CPU-GPU offload 추론.
Limits: attention 예측 정확도, CPU-GPU 링크, 선정 모델·시퀀스에 민감하다. +
서지 추적과 후속 원문 확보 작업에만 적용한다.
Limits: Primary conference and author pages verify metadata but expose no abstract or paper content in the accessible record. +
SSD 기반 복제 KV 저장소.
Limits: WO-KV 의미론과 bounded history에 의존하며 일부 읽기는 추가 왕복으로 fallback한다. +
k
kal kernel assisted non invasive memory leak tolerance with a general purpose memory allocator 372cd042 +
장시간 실행되는 C/C++ service의 availability-preserving leak tolerance에 해당한다. +
Large flash caches for social, IoT, and other tiny-object workloads.
Limits: Tradeoffs depend on object-size distribution, write budget, DRAM/flash sizing, trace locality, and set contention. +