속성:Attempt
외관
직접 해본 방법이나 설정을 적습니다.
o
iteration-level scheduling과 selective batching으로 매 디코딩 반복마다 배치를 재구성한다. +
CXL 인터페이스와 SSD 내부 캐싱·prefetch를 공동 설계한다. +
ozz identifying kernel out of order concurrency bugs with in vivo memory access reordering 0b85f639 +
OEMU emulates processor memory-access reordering during live kernel execution; Ozz jointly controls reorderings and thread interleavings for systematic testing. +
p
PaCaR maintains a main page plus per-node twins, invalidates replicas, switches the main copy, and integrates replication with PFRA/LRU in Linux 6.12. +
PACT defines per-page access criticality from four hardware counters and per-tier MLP, then uses eager demotion and adaptive promotion online. +
컴파일러·클라이언트·스케줄러를 공동 설계해 소프트웨어가 커널 순서와 정책을 정한다. +
papi exploiting dynamic parallelism in large language model decoding with a processing in memory 5f9f53cc +
PAPI characterizes kernels at runtime and dynamically assigns them across a GPU and heterogeneous PIM devices. +
partial failure resilient memory management system for cxl based distributed shared memory 48948c1d +
CXL-SHM은 era 기반 비차단 reference counting으로 장애 세대와 살아 있는 참조를 구분한다. +
pegasus tolerating skewed workloads in distributed storage with in network coherence directories a4e8c8ec +
프로그래머블 스위치에 일관성 디렉터리를 두고 핫 키를 선택 복제하며 부하 인지 라우팅한다. +
프로파일링 후 일부 변수를 VGPR에서 LDS로 이동해 occupancy를 높이고 global work size/wavefront 수를 조정한다. work size가 같은 Farnebäck kernel들을 병합해 launch와 중간 데이터 이동을 줄인다. +
Perseus는 드라이브 수준 경량 회귀 모델로 지연 이상을 탐지한다. +
PET introduces application allocation units called P-blocks, with targeted selection and fast promotion in Linux 6.1.44. +
세 시간창의 working set 교집합으로 hottest/hotter/tepid를 구분해 hard threshold 전에 단계적으로 이동한다. +
pit optimization of dynamic sparse deep learning models via permutation invariant transformation 48509ad0 +
순열 불변 변환으로 sparse microtile을 dense GPU tile로 모으고 PIT 규칙을 컴파일·실행한다. +
Materialize the complete effective vLLM package before GPU work: exact-commit tracked files, every non-bytecode ignored runtime file, and stable-copied direct symlink target bytes. Freeze it read-only, launch with Python safe-path and a private PYTHONPATH, disable plugins, reject inherited VLLM overrides, and rehash before and after each model. +
Launched the source-bound migration runner at commit 01031e6e484bbd98e00bdc3320f77c45d292f6e1 on B200-2 via the approved SSH route and persistent tmux, with run_id b200-2-generation-20260717-a01 and GPU selector 0. +
The B200 migration runner enumerated repository-visible .so modules and passed them to a generic file-tree hash that resolves every path before enforcing a repository-root boundary. +
At git 8f61aa84212dc232a1392180fdbe4c0f8333d8fb, audited final_v1 protocol, run-plan inventory, generation/latency/trace validators, then ran a read-only B200-2 preflight with ssh -p 49001: GPU inventory/process query, git state, and .venv torch/CUDA/vLLM version checks. +
The calibration pipeline digest hashed resolved absolute filenames together with file bytes. +
The launcher invoked remote tmux send-keys through SSH with the full experiment command intended as one quoted argument. +