본문으로 이동

속성:Observation

S3 연구 메모리

Text

직접 확인한 결과를 적습니다. 원인에 대한 해석은 따로 적습니다.

( | ) (20 | 50 | 100 | 250 | 500) 보기
이 속성을 사용하는 문서 20개를 보여줍니다.
o
workloads=GPT-3 175B; baselines=NVIDIA FasterTransformer; metrics=throughput at equal latency; results=36.9× throughput at the same latency.  +
workloads=real-world memory traces; baselines=conventional SSD/memory expansion designs; metrics=access latency; estimated lifetime; results=68–91% of requests under 1 μs; estimated lifetime ≥3.1 years.  +
workloads=previously reported Linux OoO bugs; latest Linux kernel; baselines=native nondeterministic execution; metrics=bug reproduction; new confirmed bugs; results=reproduced known bugs; 11 new developer-confirmed and patched bugs  +
p
workloads=Filebench, fio, RocksDB on dual-NUMA Intel Xeon; baselines=Linux 6.12 page cache; metrics=throughput/performance and write overhead; results=up to 40% Filebench and 8% RocksDB improvement; minimal write overhead  +
workloads=13 graph, HPC, in-memory-cache, and ML workloads; 96-workload model study; baselines=Soar, Alto, Memtis, Colloid, Nomad, TPP, Linux NBT; metrics=performance, migrations, model correlation; results=up to 61% faster; up to 50x fewer migrations; Pearson >0.98  +
workloads=8-model serving mix; two extreme models; baselines=hardware-scheduled serving baseline; metrics=latency; throughput; fairness; results=Lower latency and higher pre-saturation throughput; exact aggregate not abstract-verified.  +
workloads=LLaMA-65B; GPT-3 66B and 175B; baselines=state-of-the-art heterogeneous GPU+PIM accelerator; PIM-only; metrics=inference performance; results=1.8x and 11.1x speedups, respectively  +
workloads=real CXL hardware; microbenchmarks; end-to-end applications; baselines=conventional reference-count management; metrics=safety; allocation/reclamation performance; results=Safety and low overhead verified; exact aggregate not abstract-verified.  +
workloads=skewed KV workloads; dynamic hot keys; baselines=partitioned distributed storage; metrics=throughput under latency SLO; results=>10× throughput across tested conditions.  +
object detection 전체 성능은 최대 86%, optical flow는 최대 10% 향상. Haar 최적화는 평균 wave throughput을 APU 67%, discrete GPU 77% 높였고 개별 전체 성능은 각각 최대 73%, 86% 향상. LBP 전체 성능은 APU 최대 31%, discrete GPU 최대 21% 향상.  +
workloads=10 months, 248K drives; 41K normal and 315 verified fail-slow drives; baselines=existing fail-slow detectors; metrics=detection; node p99.99 latency; results=304 fail-slow drives found; isolation reduced p99.99 by 48%.  +
workloads=tiered-memory workloads on Linux 6.1.44; baselines=page-based tiering; metrics=fast-memory footprint and performance; results=39.8% average/80.4% maximum footprint reduction at 1.7% average loss; mitigates 31% cliff  +
workloads=SSDsim; realistic disk traces; ARM Cortex-A7 controller model; baselines=RL-RR; Reallocation; existing RR scheduling; metrics=overall I/O latency; read latency; erase operations; DRAM overhead; results=Overall I/O latency -31.3%; read latency -33.6%; erases -6.1% on average.  +
workloads=diverse dynamic sparse DL models; baselines=state-of-the-art DL compilers; metrics=end-to-end speedup; results=Up to 5.9×, 2.43× average speedup.  +
A real CPU-only materialization covered 3,236 files and 553,835,742 bytes: 2,251 tracked files, 985 runtime-untracked files, 13 native extensions, and 2 direct symlinks materialized as regular files. Strict audit validation and actual vLLM package import from the private snapshot passed, and teardown left no snapshot files.  +
The runner exited with 'git source changed before lock acquisition'. The lock owner.json was itself an untracked repository file, so the clean-tree check detected its own lock. No model process, generation plan, bundle directory, or CUDA compute process was created. The lock was released and the run ID remained free.  +
The two logical FlashAttention modules vllm/vllm_flash_attn/_vllm_fa2_C.abi3.so and _vllm_fa3_C.abi3.so are symlinks into the active vLLM virtual environment. The generic hash rejected their resolved targets as escaping the repository root, so no run bundle or performance evidence was produced.  +
The inventory declares target_generation::primary and target_generation::architecture_replication with hardware=null. Promotable generation validates pinned model/tokenizer, clean git, native extension hash, backend/runtime versions, and records the producing GPU, but does not bind the artifact to L40S or RTX. Latency, regime selection, native baseline, traces, policy bindings, host preflight, and results are target hardware/config bound. B200 final cells are rejected.  +
The old digest changed when the same package tree lived at a different absolute path. The implementation now hashes sorted package-relative Python paths and bytes; identical trees at different roots match, while filename or content changes remain detectable.  +
The remote pane received a single concatenated token such as exec.venv/bin/python-mscripts..., returned command-not-found, and remained at an idle bash prompt. No global lock, run directory, or GPU compute process was created.  +