속성:Applicability
외관
이 메모가 맞는 조건과 예외를 적습니다.
s
Most directly applicable to high-fanout object distribution inside a managed private cloud. The official abstract does not establish equivalent behavior on open or adversarial networks. +
Linux file systems on multi-NVMe servers.
Limits: Results depend on the evaluated 64-core/eight-drive platform and altered page-cache/writeback semantics. +
Large-memory-pressure servers backed by multi-NVMe swap arrays.
Limits: Targets large all-flash/128-core configurations; smaller systems may not expose the same contention. +
NVMe-oF JBOF systems with programmable DPUs.
Limits: Requires DPU-capable infrastructure and targets high-density configurations. +
virtualized multimedia client에서 frame-rate-aware CPU allocation에 직접 해당한다. +
Linux multiprocess/MPI 및 kernel-intensive workload의 최적화 우선순위 탐색에 높음. +
Android 계열의 메모리 중복 제거와 앱 캐시 유지 정책에 직접 관련된다. +
Model calibration reruns, migration candidates, deterministic inference artifacts, and any replicated experiment whose control-plane provenance changes while the scientific payload should not. +
serverless in the wild characterizing and optimizing the serverless workload at a large cloud pr f2fa1121 +
FaaS 용량 계획과 콜드 스타트 완화.
Limits: 단일 공급자·14일 관측이며 초록에서 정확한 개선 수치는 확인되지 않는다. +
Multi-model serverless LLM inference clusters.
Limits: Benefits depend on local checkpoint capacity, locality, storage bandwidth, model size, and loading-dominated workloads. +
GPU serving and compiler projects that choose among shape-specialized kernels, especially attention, RoPE, and other launch-sensitive operations under CUDA Graphs. Quantitative bounds are limited to one L40S, five Triton configurations, one controlled GQA compatibility partition, and a PyTorch CUDA Graph proxy; a materially different native candidate set or real trace could reopen the question only if it first demonstrates substantially larger global-to-oracle headroom. +
RDMA-based distributed storage and transaction systems.
Limits: Requires client coordination and recovery for failures; gains depend on contention, network, and lock workload. +
웹·객체 캐시 라이브러리.
Limits: 작은 캐시의 scan형·일부 block workload에서는 LRU보다 나쁠 수 있고 ghost history가 없다. +
LSM key-value stores with latency-sensitive foreground traffic.
Limits: Results target the evaluated RocksDB-derived implementation and storage/workload mixes. +
멀티테넌트 x86 클라우드.
Limits: 선정 Intel 세대와 추정 subarray/address mapping·BIOS 정보에 의존한다. +
Cached-process 기반 모바일 lifecycle 관리에 직접 해당한다. +
specinfer accelerating generative large language model serving with tree based speculative infer e609aa25 +
Autoregressive LLM serving with distributed or offloaded target models.
Limits: Benefit depends on draft accuracy, token-tree construction cost, target parallelism, model pair, and hardware. +
point lookup이 많고 SSD queue 여유가 있는 LSM KV store에 높음; 포화 장치에는 추가 제어 필요. +
speed is all you need on device acceleration of large diffusion models via gpu aware optimizatio c53959fd +
Stable Diffusion 1.4 계열 512×512 on-device inference와 유사한 모바일 GPU kernel 최적화에 적용합니다.
Limits: S23 Ultra와 iPhone 14 Pro Max, 특정 모델·해상도·20-step 설정 중심이며 에너지·품질·다른 SoC 일반화는 추가 평가가 필요합니다. +
DRAM보다 큰 반복 ML working set과 Spark류 데이터 분석 프레임워크에 해당한다. +