속성:Applicability
외관
이 메모가 맞는 조건과 예외를 적습니다.
o
대형 생성 모델 분산 추론.
Limits: 평가 당시 모델·GPU에 맞춰졌고 최신 KV 캐시·fault handling 기능은 포함하지 않는다. +
CXL 메모리 확장과 대용량 메모리 워크로드.
Limits: 일부 결과가 추정·프로토타입 기반이며 플래시 수명과 tail에 민감하다. +
ozz identifying kernel out of order concurrency bugs with in vivo memory access reordering 0b85f639 +
Linux kernel concurrency testing, memory-barrier validation, and fuzzing.
Limits: Emulation coverage is bounded by the modeled memory behavior and explored schedules; findings focus on Linux. +
p
Multi-socket NUMA servers running buffered file I/O.
Limits: Current implementation does not support mmap; write-heavy benefits are constrained by invalidation/coherence. +
DRAM plus NUMA/persistent/CXL memory where standard performance counters are available.
Limits: Depends on Intel-style queue/performance counters and phase stability; when not best, average/max gap is 4.1%/11.8%. +
다중 테넌트 GPU 추론.
Limits: 컴파일러·런타임 통합과 하드웨어별 커널 제어가 필요하며 정확한 수치는 초록에서 검증되지 않았다. +
papi exploiting dynamic parallelism in large language model decoding with a processing in memory 5f9f53cc +
Large-model inference platforms with GPU and PIM execution support.
Limits: Benefits depend on specialized PIM hardware and the evaluated model/kernel mix. +
partial failure resilient memory management system for cxl based distributed shared memory 48948c1d +
CXL 기반 shared-memory pool.
Limits: refcount 순환 참조 가능성과 CXL·부분 장애 모델에 제약된다; 초록에 정량치가 없다. +
pegasus tolerating skewed workloads in distributed storage with in network coherence directories a4e8c8ec +
메모리 KV, 랙 단위 분산 캐시.
Limits: 스위치 ASIC 용량·토폴로지와 인메모리 KV 가정에 제약된다. +
OpenCL/OpenCV의 고정 vision kernel 최적화에 높고, 자동화 없이는 유지보수 비용이 큼. +
대규모 HDD/SSD 스토리지 플릿.
Limits: Alibaba형 플릿·라벨링 절차에 맞춰졌고 타 환경 일반화는 추가 검증이 필요하다. +
Applications able to expose or infer P-block allocation boundaries in tiered memory.
Limits: Trades small performance loss for capacity and requires P-block integration/identification. +
TLC/QLC SSD FTL read-reclaim.
Limits: SSDsim 기반이며 세 window·threshold와 trace hotness에 민감하다. +
pit optimization of dynamic sparse deep learning models via permutation invariant transformation 48509ad0 +
동적 sparse Transformer·DNN.
Limits: 연산이 PIT 규칙을 허용해야 하고 런타임 재배열·kernel 생성 비용이 있다. +
Applies to reproducible GPU inference experiments using editable installs, ignored wheel payloads, JIT caches, external symlinks, or plugin entry points. +
Applies to the B200 migration runner and any final or bridge runner using repository-local atomic lock directories plus git status --untracked-files=all. +
Applies to reproducible GPU experiments whose source tree exposes compiled modules through symlinks, editable installs, or environment-specific build locations. +
Applies to frozen final_v1 model IDs/revisions, deterministic temperature=0/top_p=1/seed=0, BF16, TP1, max_model_len=4096, output cap=256, disjoint 1,536-prompt calibration split, clean source commit, and current hardware-null target_generation inventory contract. +
Applies to experiment artifacts validated across hosts, containers, worktrees, shared filesystems, or repository renames. +
Applies to detached remote experiment launchers that drive tmux through nested local and remote shells. +