속성:Evidence overview
외관
관련 자료를 짧게 정리합니다. 자세한 인용은 자료 항목에 따로 적습니다.
g
Publication record 1: 권상윤,안민우,정진규, "인과관계 프로파일링의 가상 최적화를 이용한 GPU 커널 성능 향상 예측", 2022 한국소프트웨어종합학술대회 (KSC), 2022 +
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints. ACL 2023. +
Publication record 1: Sangwook Kim, Hwanju Kim, Joonwon Lee, and Jinkyu Jeong, "Group-based Memory Oversubscription for Virtualized Clouds," Journal of Parallel and Distributed Computing, vol. 74, issue 4, pp. 2241-2256, Apr. 2014 +
h
Halfmoon: Log-Optimal Fault-Tolerant Stateful Serverless Computing. SOSP 2023. +
Harvesting Memory-bound CPU Stall Cycles in Software with MSH. OSDI 2024. +
HotRAP: Hot Record Retention and Promotion for LSM-trees with Tiered Storage. USENIX ATC 2025. +
Publication record 1: Junghoon Kim, Sangwook Kim, Joonwon Lee and Jinkyu Jeong, "How Storage I/O Affects User-Perceived Latency in Mobile Apps," ACM SIGOPS/EuroSys European Conference on Computer Systems (EuroSys 2015), Poster, Bordeaux, France, April 21-24, 2015. +
How to Copy Memory? Coordinated Asynchronous Copy as a First-Class OS Service. SOSP 2025. +
i
I/O Passthru: Upstreaming a Flexible and Efficient I/O Path in Linux. FAST 2024. +
IC-Cache: Efficient Large Language Model Serving via In-context Caching. SOSP 2025. +
ice collaborating memory and process management for user experience on resource limited mobile d 312e453f +
ICE: Collaborating Memory and Process Management for User Experience on Resource-limited Mobile Devices. EuroSys 2023. +
Publication record 1: Minwoo Ahn, Jeongmin Han, Youngjin Kwon, Jinkyu Jeong, "Identifying On-/Off-CPU Bottlenecks Together with Blocked Samples," in Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI '24), July 2024 +
impress an importance informed multi tier prefix kv storage system for large language model infe 51b47015 +
IMPRESS: An Importance-Informed Multi-Tier Prefix KV Storage System for Large Language Model Inference. FAST 2025. +
Publication record 1: Andre Luiz Nunes Martins, Cesar A. V. Duarte, Jinkyu Jeong, "Improving Application Launch Performance in Smartphones Using Recurrent Neural Network," in Proceedings of the 2018 International Conference on Machine Learning Technologies (ICMLT 2018), Jinan, China, May 26-28, 2018. +
inf2 high throughput generative inference of large language models using near storage processing 3fcead73 +
Hongsun Jang, Jaeyong Song, Changmin Shin, Si Ung Noh, Jaewon Jung, Jisung Park, and Jinho Lee. “A Cost-Effective Near-Storage Processing Solution for Offline Inference of Long-Context LLMs.” ASPLOS ’26. DOI: 10.1145/3779212.3790119. arXiv:2502.09921v2.
Source: https://arxiv.org/abs/2502.09921v2
Full text: https://arxiv.org/html/2502.09921v2
Verification basis: full_text.
Canonical evidence ID: canonical-paper-v2-3fcead73 +
infinigen efficient generative inference of large language models with dynamic kv cache manageme a1ca3228 +
InfiniGen: Efficient Generative Inference of Large Language Models with Dynamic KV Cache Management. OSDI 2024. +
Intel Accelerator Ecosystem: An SoC-Oriented Perspective. ISCA 2024. +
IONIA: High-Performance Replication for Modern Disk-based KV Stores. FAST 2024. +
k
kal kernel assisted non invasive memory leak tolerance with a general purpose memory allocator 372cd042 +
Publication record 1: Jinkyu Jeong, Euiseong Seo, Jeonghwan Choi, Hwanju Kim, Heeseung Jo, and Joonwon Lee, "KAL: Kernel-assisted Non-invasive Memory Leak Tolerance with a General-purpose Memory Allocator," Software: Practice and Experience, vol. 40, no. 8, pp. 605-625, Jul. 2010 +
Kangaroo: Caching Billions of Tiny Objects on Flash. SOSP 2021. +