Abstract
Large Language Models (LLMs) have become piv-otal, yet their auto-regressive inference suffers from significant memory bandwidth bottlenecks, hindering performance and energy efficiency. In this paper, we propose GATHER, a novel hardware accelerator architecture specifically designed for efficient generative AI inference. GATHER introduces two key contributions: (1) A token-stream processor that natively han-dles variable-length sequences, completely eliminating padding-related overhead. (2) A specialized Gated-Gather Engine that tackles the attention bottleneck by tightly coupling Top-K attention score selection with a dedicated address gather unit. This engine identifies the most salient tokens and issues optimized, batched memory requests to DRAM, drastically reducing off-chip traffic. Evaluation results show that our proposed architecture outperforms a single NVIDIA A100 GPU on GPT-2 and Llama-3-8B in terms of throughput and energy efficiency.
| Original language | English |
|---|---|
| Title of host publication | International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers |
| Publisher | Institute of Electrical and Electronics Engineers Inc. |
| ISBN (Electronic) | 9798331586423 |
| DOIs | |
| State | Published - 2025 |
| Event | 22nd International SoC Design Conference, ISOCC 2025 - Busan, Korea, Republic of Duration: 15 Oct 2025 → 18 Oct 2025 |
Publication series
| Name | International SoC Design Conference 2025, ISOCC 2025 - Proceedings of Technical Papers |
|---|
Conference
| Conference | 22nd International SoC Design Conference, ISOCC 2025 |
|---|---|
| Country/Territory | Korea, Republic of |
| City | Busan |
| Period | 15/10/25 → 18/10/25 |
Bibliographical note
Publisher Copyright:© 2025 IEEE.
UN SDGs
This output contributes to the following UN Sustainable Development Goals (SDGs)
-
SDG 7 Affordable and Clean Energy
Keywords
- Accelerator
- Generative AI
- Transformer
Fingerprint
Dive into the research topics of 'GATHER: A Gated-Attention Accelerator for Efficient LLM Inference'. Together they form a unique fingerprint.Cite this
- APA
- Author
- BIBTEX
- Harvard
- Standard
- RIS
- Vancouver