본문으로 이동

속성:Observation

S3 연구 메모리

Text

직접 확인한 결과를 적습니다. 원인에 대한 해석은 따로 적습니다.

( | ) (20 | 50 | 100 | 250 | 500) 보기
이 속성을 사용하는 문서 20개를 보여줍니다.
c
Xeon Gold 6348; 64GB DDR4 fast tier + 256GB Optane PM slow tier. Workloads: Pmbench, Graph500, Memcached, Redis, lkp-test. Baselines: Linux NUMA balancing, Auto-Tiering, Multi-Clock, TPP, Memtis. Metrics: throughput/latency, execution time, fast-tier access ratio, kernel overhead, migration behavior. Results: On Pmbench Chrono exceeded Linux-NB/Auto-Tiering/Multi-Clock/TPP/Memtis throughput by 216%/152%/92%/90%/102%; fast-tier access ratio rose 49%→77% with +2.1 percentage points kernel time vs Linux-NB. Graph500 speedups vs Linux-NB were 2.49×, 2.29×, and 2.05× across 128–256GB working sets. Memcached/Redis also improved overall throughput.  +
workloads=thousands of memory security domains; baselines=prior guard/isolation allocation; metrics=memory overhead and performance; results=7.2% average overhead, no performance loss, 4–6x less overhead than prior isolation  +
A node-local XFS bridge archive was sealed with 88 payload files and no symlinks. Exactly 27 locator-bound model/tokenizer files totaling 29,742,199,921 bytes were materialized as regular files and individually rehashed. The complete archive payload is 30,772,529,198 bytes. This protects against broken snapshot links but remains on retiring B200 infrastructure.  +
workloads=Lustre client/server workloads; baselines=original Lustre; other distributed file systems; metrics=I/O performance; results=Up to 3× vs Lustre and 13× vs other DFSs.  +
Both models completed with 1,536 records, three beliefs, promotable=true, correct single-GPU isolation, and clean teardown. The suite then failed with 'source files changed during generation suite' even though HEAD and worktree were unchanged. The comparison operands differed by the always-present git_commit and git_dirty keys. Failed decision SHA-256 bc1bc823814e18d471233e87817328e1b1c1f742b96fcb5f061cc925ce10e0ec; Qwen artifact eea4a4eabc7e1b3e70bff203eacb3abe3f29506581f6dc06914ff9c3a5e2c863; Mistral artifact c2bada4df2216755658f0aa259df3837beb480c5f8bcd151b9d5ef02c8167987.  +
Xen prototype에서 main VM에 필요한 메모리를 합리적 비용으로 제공하면서 third-party applications를 종료하지 않았다고 보고한다. 공식 초록에는 수치가 없다.  +
workloads=ten instantiated compaction strategies; baselines=representative leveling, tiering, and hybrid policies; metrics=write amplification; write throughput; point lookup; range lookup; space amplification; results=12 empirical observations; seven design takeaways; no universal winner  +
Nexus 10 Android-kernel prototype에서 fragmentation, I/O-buffer allocation time, CPU usage, energy consumption을 줄였다고 보고하지만 공개 초록에는 수치가 없다.  +
workloads=Sundial on Azure Blob; Sundial on Redis; baselines=classic 2PC; metrics=commit latency; throughput; results=Up to 1.9× latency speedup.  +
workloads=read-intensive and mixed OLTP on a 48-core server; baselines=conventional main-memory execution; metrics=transaction throughput; scalability; results=close to 2x on read-intensive workloads  +
workloads=multi-turn LLM conversations; baselines=recompute and GPU-resident KV-cache serving; metrics=TTFT; prefill throughput; end-to-end cost; results=up to 87% lower TTFT; up to 7.8x prefill throughput; up to 70% lower cost  +
이질적 함수가 한 worker를 공유할 때 AWS Lambda 대비 메모리 사용 최대 32.0% 감소. Azure trace 기반 평가에서 eviction 감소로 p99 latency 41.2% 감소.  +
workloads=microbenchmarks; macrobenchmarks; real workloads; local and remote storage; baselines=existing OS/runtime prefetchers; metrics=I/O throughput; prefetch accuracy; demand interference; results=1.22–3.7x I/O throughput  +
d
MQFQ 대비 CPU utilization을 최대 45% 줄이면서 공정성을 유지했다. LL-D2FQ는 latency를 최대 35% 낮추고 bandwidth를 최대 54% 높였다. YCSB+FIO에서는 bandwidth ratio 1.00-1.05를 보였다.  +
Nexus S와 Android apps에서 flash read I/O를 50–85% 줄였고 앱 실행성능을 8–16% 개선했으며 device reallocation은 수 ms로 제한했다.  +
workloads=fine-tuned Llama-family variants and multi-tenant request traces; baselines=vLLM and conventional model swapping; metrics=delta size, throughput, end-to-end latency, TTFT; results=up to 13x delta compression; 2–12x throughput; 1.6–16x latency/TTFT improvements  +
혼합 PARSEC workloads에서 불필요한 synchronization latency와 CPU 소비를 줄여 전체 성능을 개선했으며, 특히 communication-intensive application에서 효과가 컸다. 공식 초록에는 정량값이 없다.  +
commodity digital TV의 실제 최적화에 사용해 tool set의 효과를 검증했다. 공식 초록에는 overhead 또는 최적화 폭의 수치가 없다.  +
최대 8개 HGX A100 노드(64 GPU), LLaVA-OneVision/InternVL 계열 7B~72B 평가에서 PyTorch 및 Megatron-LM 대비 처리량 최대 3.6배. GPU idle time은 PyTorch 대비 최대 82%, Megatron-LM 대비 최대 84% 감소.  +
workloads=dRAID microbenchmarks; object-store workloads; baselines=conventional disaggregated RAID; metrics=bandwidth; object-store throughput; reconstruction efficiency; results=up to 3x bandwidth; 1.5–2.35x object-store throughput  +