🎓 About Me
I am currently a third-year Ph.D. in Electrical and Computer Engineering at the University of Arizona, advised by Dr. Huanrui Yang. Prior to joining UA, I earned my M.S. from University of Chinese Academy of Sciences in 2024 and my B.E. from Beijing Forestry University in 2021.
📝 My research focuses on AI Infra, with a focus on model quantization, KV cache optimization and token pruning.
🤝 I’m on the job market and open to industry opportunities in AI Infra. Please feel free to reach out!
🔥 News
- 2026.09: 🎉🎉 One paper MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM is accepted at Neurips 2026 ! See you in Atlanta !
- 2026.07: 🎉🎉 One paper Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction is accepted at COLM 2026! We extend the idea of NPFT to ML unlearning !
- 2026.05: 🎉🎉 I am honored to be awarded the Silver Reviewer Award by ICML 2026 !
- 2026.04: 🎉🎉 One paper GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs is accepted at ICML 2026!
- 2026.02: 🎉🎉 Extend research internship at Panasonic AI, will focus on efficient Agent !
- 2025.11: 🎉🎉 Passed my PHD qualify exam. PHD candidate now!
- 2025.09:I am honored to be awarded the ICCV Broad Participation (BP) Award !
- 2025.09:I will give a live talk about our ICCV work in the AI TIME 论道 WeChat public channel.
- 2025.08: 🎉🎉 One first-authored paper FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference is accepted at EMNLP 2025 Findings!
- 2025.08: 🎉🎉 I will join Panasonic AI as a remote intern this semester!
- 2025.06: 🎉🎉 One paper MSQ: Memory-Efficient Bit Sparsification Quantization is accepted at ICCV 2025! See you in Hawaii!
- 2025.04: 🎉🎉 I am selected as a DAC Young Fellow at the 62nd Design Automation Conference. See you in SF!
- 2025.02: 🎉🎉 One first-authored paper Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization is accepted at CPAL 2025! See you in Stanford!
📝 Selected Publications (*Equal contribution)
-
Neruips 2026: MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM [PDF]
Dongwei Wang*, Jinhee Kim*, Seokho Han*, et al. -
EMNLP 2025 Findings: FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference [PDF] [CODE]
Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Tianlong Chen, Huanrui Yang. -
CPAL 2025: Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization [PDF] [CODE]
Dongwei Wang, Huanrui Yang. -
ICML 2026: GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs [PDF] [CODE]
Jianing Deng, Song Wang, Dongwei Wang, Zijie Liu, Tianlong Chen, Huanrui Yang, Jingtong Hu. -
COLM 2026: Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction [PDF]
Jialu Wang, Jianing Deng, Shuqing Luo, Yuanzhe LI, Dongwei Wang, Jingtong Hu, Huanrui Yang, Song Wang, Tianlong Chen. -
ICCV 2025: MSQ: Memory-Efficient Bit Sparsification Quantization [PDF] [CODE]
Seokho Han, Seoyeon Yoon, Jinhee Kim, Dongwei Wang, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko.
For full publications: [Google Scholar]
🎖 Honors and Awards
- 2026 ICML Silver Reviewer Award
- 2026 AI Ignite Showcase Runner-Up Award (ECE of UArizona)
- 2025 DAC Young Fellow
- 2025 ICCV Broaden Participitation Award
💻 Internships
- 2025.08 - 2025.12, Panasonic AI, US.
- 2026.02 - 2026.05, Panasonic AI, US.
🎓 Academic Services
I have served as the reviewer for :
- Neurips, CVPR, ICML, TNNLS
🌍 Visits