본문으로 이동

속성:Context

S3 연구 메모리

Text

하드웨어, 부하, 버전, 규모처럼 결과에 영향을 줄 수 있는 조건을 적습니다.

( | ) (20 | 50 | 100 | 250 | 500) 보기
이 속성을 사용하는 문서 20개를 보여줍니다.
g
Publication scope: domestic. Lab publication metadata: 1. 권상윤,안민우,정진규, "인과관계 프로파일링의 가상 최적화를 이용한 GPU 커널 성능 향상 예측", 2022 한국소프트웨어종합학술대회 (KSC), 2022 Verification level: metadata_only. Sources checked: https://www.riss.kr/search/Search.do?colName=re_a_kor&isDetailSearch=Y&queryText=znCreator%2C%EC%95%88%EB%AF%BC%EC%9A%B0%28Minwoo+Ahn%29&searchGubun=true https://cslab.yonsei.ac.kr/publications  +
Venue: ACL. Year: 2023. MQA는 KV 캐시를 줄이지만 품질 저하가 생길 수 있고 MHA는 추론 메모리가 크다. Verification: abstract_only; confidence=high.  +
Publication scope: international. Lab publication metadata: 1. Sangwook Kim, Hwanju Kim, Joonwon Lee, and Jinkyu Jeong, "Group-based Memory Oversubscription for Virtualized Clouds," Journal of Parallel and Distributed Computing, vol. 74, issue 4, pp. 2241-2256, Apr. 2014 Verification level: official_abstract. Sources checked: https://doi.org/10.1016/j.jpdc.2014.01.001 https://www.sciencedirect.com/science/article/pii/S0743731514000033 https://yonsei.elsevierpure.com/en/publications/group-based-memory-oversubscription-for-virtualized-clouds/  +
h
Venue: SOSP. Year: 2023. 모든 상태 읽기·쓰기를 동일하게 기록하면 불필요한 로그와 지연이 발생한다. Verification: abstract_only; confidence=high.  +
Venue: OSDI. Year: 2024. SMT harvests stalls but offers coarse concurrency control and can interfere with latency-critical work. Verification: official_abstract; confidence=high.  +
Venue: USENIX ATC. Year: 2025. Coarse promotion wastes I/O and space when only a few keys in an on-disk structure are hot. Verification: official USENIX page and abstract; confidence=high.  +
Publication scope: international. Lab publication metadata: 1. Junghoon Kim, Sangwook Kim, Joonwon Lee and Jinkyu Jeong, "How Storage I/O Affects User-Perceived Latency in Mobile Apps," ACM SIGOPS/EuroSys European Conference on Computer Systems (EuroSys 2015), Poster, Bordeaux, France, April 21-24, 2015. Verification level: metadata_only. Sources checked: https://eurosys2015.labri.fr/program/posters/  +
Venue: SOSP. Year: 2025. Copies occur across files, networks, and IPC; synchronous/local handling misses overlap, SIMD/DMA, and redundant-copy elimination. Verification: official DOI/SOSP metadata, author page, and institution research report; confidence=high.  +
i
Venue: FAST. Year: 2024. 고성능 애플리케이션은 범용 블록 경로의 변환·락·큐잉 비용을 부담한다. Verification: abstract_only; confidence=high.  +
Venue: SOSP. Year: 2025. Many incoming queries resemble historical requests, but conventional caches require exact matches. Verification: arXiv abstract/full text and official DOI metadata; confidence=high.  +
Venue: EuroSys. Year: 2023. On low-memory phones, background processes can repeatedly refault pages and interfere with the foreground despite standard CPU/memory policies. Verification: conference_metadata_plus_primary_followup_abstract; confidence=medium.  +
Publication scope: international. Lab publication metadata: 1. Minwoo Ahn, Jeongmin Han, Youngjin Kwon, Jinkyu Jeong, "Identifying On-/Off-CPU Bottlenecks Together with Blocked Samples," in Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI '24), July 2024 Verification level: full_text. Sources checked: https://www.usenix.org/conference/osdi24/presentation/ahn https://www.usenix.org/system/files/osdi24-ahn.pdf  +
Venue: FAST. Year: 2025. Repeated long prefixes save prefill compute, but CPU-memory limits push KV state to disks whose I/O can erase the benefit. Verification: official_abstract; confidence=high.  +
Publication scope: international. Lab publication metadata: 1. Andre Luiz Nunes Martins, Cesar A. V. Duarte, Jinkyu Jeong, "Improving Application Launch Performance in Smartphones Using Recurrent Neural Network," in Proceedings of the 2018 International Conference on Machine Learning Technologies (ICMLT 2018), Jinan, China, May 26-28, 2018. Verification level: official_abstract. Sources checked: https://doi.org/10.1145/3231884.3231897 https://yonsei.elsevierpure.com/en/publications/improving-application-launch-performance-in-smartphones-using-rec-2/  +
Offline LLM inference는 online serving보다 지연 허용 범위가 넓어 큰 batch와 긴 sequence를 사용하기 좋으며, benchmark와 대규모 정보 추출이 대표적인 대상입니다. 기존 offloading 시스템은 가중치와 KV cache를 호스트 DRAM 또는 SSD에 두지만, KV-cache 크기는 batch와 context length에 함께 비례합니다. 저자들의 OPT-175B 동기 분석에서는 KV-cache 전송이 전체 실행 시간의 60%를 넘고 cache가 TB 규모에 이릅니다. SSD 수만 늘려도 전체 KV가 host PCIe를 통과하면 링크 병목은 남습니다. 검증 판본은 2026-02-06의 arXiv:2502.09921v2와 ASPLOS ’26 DOI 10.1145/3779212.3790119입니다. 이 revision은 v2의 새 제목, 시스템명 HILOS, 저자 7명과 v2 평가 결과만 사용합니다.  +
Venue: OSDI. Year: 2024. 모든 KV를 GPU에 두기 어렵고 CPU에서 전부 가져오면 interconnect가 병목이다. Verification: abstract_only; confidence=high.  +
Venue: ISCA. Year: 2024. Verification scope: metadata only. 공개 publisher/proceedings/author record에서 서지 존재 여부만 확인했으며 기술 본문은 확인하지 못했다.  +
Venue: FAST. Year: 2024. 모든 쓰기를 즉시 외부화·순서화하면 SSD 내부 병렬성과 follower read를 제한한다. Verification: full_text; confidence=high.  +
k
Publication scope: international. Lab publication metadata: 1. Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, and Joonwon Lee, "KAL: Kernel-assisted Non-invasive Memory Leak Tolerance with a General-purpose Memory Allocator," Software: Practice and Experience, vol. 40, no. 8, pp. 605-625, Jul. 2010 Verification level: official_abstract. Sources checked: https://doi.org/10.1002/spe.970 https://yonsei.elsevierpure.com/en/publications/kal-kernel-assisted-non-invasive-memory-leak-tolerance-with-a-gen/  +
Venue: SOSP. Year: 2021. Set-associative caches minimize DRAM but rewrite flash often; log-structured caches amortize writes but need large DRAM indexes. Verification: official_abstract; confidence=high.  +