Posts by Collection

publications

KVPR: Efficient LLM Inference with I/O-Aware Partial KV Cache Recomputation

Published in Association for Computational Linguistics (ACL) 2025, 2025

The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.

Download Paper

LEAF: Lightweight, Efficient, Adaptive and Flexible Embedding for Large-Scale Recommendation Models

Published in The ACM Conference on Recommender Systems (RecSys) 2025, 2025

The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.

Download Paper

DEL: Context-Aware Dynamic Exit Layer for Efficient Self-Speculative Decoding

Published in Conference on Language Modeling (COLM) 2025, 2025

The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.

Download Paper

HuffmanEmbed: Using Huffman Coding for Embedding Table Compression in Deep Learning Recommendation Models

Published in ACM SIGIR Conference on Research and Development in Information Retrieval (SIGIR) 2026, 2026

The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.

DELTA: Dynamic Layer-Aware Token Attention for Efficient Long-Context Reasoning

Published in Association for Computational Linguistics (ACL) 2026, 2026

The contents above will be part of a list of publications, if the user clicks the link for the publication than the contents of section will be rendered as a full page, allowing you to provide more information about the paper for the reader. When publications are displayed as a single page, the contents of the above “citation” field will automatically be included below this section in a smaller font.

Download Paper