🎓 About Me

I am currently a third-year Ph.D. in Electrical and Computer Engineering at the University of Arizona, advised by Dr. Huanrui Yang. Prior to joining UA, I earned my M.S. from University of Chinese Academy of Sciences in 2024 and my B.E. from Beijing Forestry University in 2021.

📝 My research focuses on AI Infra, with a focus on model quantization, KV cache optimization and token pruning.

🤝 I’m on the job market and open to industry opportunities in AI Infra. Please feel free to reach out!

🔥 News

  • 2026.09:  🎉🎉 One paper MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM is accepted at Neurips 2026 ! See you in Atlanta !
  • 2026.07:  🎉🎉 One paper Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction is accepted at COLM 2026! We extend the idea of NPFT to ML unlearning !
  • 2026.05:  🎉🎉 I am honored to be awarded the Silver Reviewer Award by ICML 2026 !
  • 2026.04:  🎉🎉 One paper GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs is accepted at ICML 2026!
  • 2026.02:  🎉🎉 Extend research internship at Panasonic AI, will focus on efficient Agent !
  • 2025.11: 🎉🎉 Passed my PHD qualify exam. PHD candidate now!
  • 2025.09:I am honored to be awarded the ICCV Broad Participation (BP) Award !
  • 2025.09:I will give a live talk about our ICCV work in the AI TIME 论道 WeChat public channel.
  • 2025.08:  🎉🎉 One first-authored paper FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference is accepted at EMNLP 2025 Findings!
  • 2025.08:  🎉🎉 I will join Panasonic AI as a remote intern this semester!
  • 2025.06:  🎉🎉 One paper MSQ: Memory-Efficient Bit Sparsification Quantization is accepted at ICCV 2025! See you in Hawaii!
  • 2025.04:  🎉🎉 I am selected as a DAC Young Fellow at the 62nd Design Automation Conference. See you in SF!
  • 2025.02:  🎉🎉 One first-authored paper Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization is accepted at CPAL 2025! See you in Stanford!

📝 Selected Publications (*Equal contribution)

  • Neruips 2026: MoBiQuant: Mixture-of-Bits Quantization for Token-Adaptive Any-Precision LLM [PDF]
    Dongwei Wang*, Jinhee Kim*, Seokho Han*, et al.

  • EMNLP 2025 Findings: FIER: Fine-Grained and Efficient KV Cache Retrieval for Long-context LLM Inference [PDF] [CODE]
    Dongwei Wang, Zijie Liu, Song Wang, Yuxin Ren, Jianing Deng, Jingtong Hu, Tianlong Chen, Huanrui Yang.

  • CPAL 2025: Taming Sensitive Weights : Noise Perturbation Fine-tuning for Robust LLM Quantization [PDF] [CODE]
    Dongwei Wang, Huanrui Yang.

  • ICML 2026: GEMQ: Global Expert-Level Mixed-Precision Quantization for MoE LLMs [PDF] [CODE]
    Jianing Deng, Song Wang, Dongwei Wang, Zijie Liu, Tianlong Chen, Huanrui Yang, Jingtong Hu.

  • COLM 2026: Quantization-Robust Unlearning through the Lens of Retain-Forget Loss Landscapes Interaction [PDF]
    Jialu Wang, Jianing Deng, Shuqing Luo, Yuanzhe LI, Dongwei Wang, Jingtong Hu, Huanrui Yang, Song Wang, Tianlong Chen.

  • ICCV 2025: MSQ: Memory-Efficient Bit Sparsification Quantization [PDF] [CODE]
    Seokho Han, Seoyeon Yoon, Jinhee Kim, Dongwei Wang, Kang Eun Jeon, Huanrui Yang, Jong Hwan Ko.

For full publications: [Google Scholar]

🎖 Honors and Awards

  • 2026 ICML Silver Reviewer Award
  • 2026 AI Ignite Showcase Runner-Up Award (ECE of UArizona)
  • 2025 DAC Young Fellow
  • 2025 ICCV Broaden Participitation Award

💻 Internships

🎓 Academic Services

I have served as the reviewer for :

  • Neurips, CVPR, ICML, TNNLS

    🌍 Visits