속성:Question
외관
무엇을 확인하려 했는지 적습니다.
o
생성형 Transformer 서빙의 가변 길이·반복 디코딩을 GPU에 효율적으로 배치할 수 있는가? +
SSD를 CXL 메모리 확장 장치처럼 사용하면서 지연과 수명을 감당할 수 있는가? +
ozz identifying kernel out of order concurrency bugs with in vivo memory access reordering 0b85f639 +
How can tests deterministically expose kernel bugs caused jointly by weak-memory reordering and thread interleavings? +
p
Can NUMA systems replicate page-cache data to restore locality without breaking buffered-I/O consistency? +
Which pages deserve DRAM when access frequency fails to capture their actual CPU-stall impact? +
GPU 모델 서빙에서 블랙박스 하드웨어 스케줄링을 우회해 꼬리 지연을 줄일 수 있는가? +
papi exploiting dynamic parallelism in large language model decoding with a processing in memory 5f9f53cc +
How should heterogeneous GPU and PIM resources be scheduled as LLM kernels alternate between compute- and memory-bound phases? +
partial failure resilient memory management system for cxl based distributed shared memory 48948c1d +
CXL 분산 공유 메모리에서 클라이언트 일부가 죽어도 안전하게 메모리를 회수할 수 있는가? +
pegasus tolerating skewed workloads in distributed storage with in network coherence directories a4e8c8ec +
키 편향이 심한 분산 저장소에서 핫 키 병목을 어떻게 완화할까? +
OpenCV GPU의 Haar/LBP object detection과 Farnebäck optical flow에서 register pressure, 낮은 occupancy, kernel launch/data-transfer 병목을 어떻게 줄일 것인가? +
클라우드 저장소에서 완전 고장 전의 느린 드라이브를 정확히 찾을 수 있는가? +
Can tiered memory allocate and promote data at a unit that better matches application locality than a page? +
페이지별 read counter 없이 SSD read-reclaim 지연을 줄일 수 있는가? +
pit optimization of dynamic sparse deep learning models via permutation invariant transformation 48509ad0 +
입력마다 희소 패턴이 바뀌는 DNN을 GPU에서 효율적으로 실행할 수 있는가? +
How can a B200 generation collector use ignored precompiled vLLM payloads and external native symlinks without leaving mutable runtime dependencies? +
Why did the first persistent tmux launch stop before any B200 generation work? +
How should native vLLM extension provenance be bound when repository-visible extension modules are symlinks into the active virtual environment? +
Which pre-final artifacts can be completed on B200-2 and safely migrated to the L40S/RTX PRO 5000 final experiment? +
Why can byte-identical target-generation artifacts fail validation after moving to a different host checkout? +
Why did a quoted tmux launch command arrive without spaces on the remote B200 shell? +