Glaive: Cleaving Bandwidth and Latency Scheduling for Efficient Non-uniform All-to-All Collective Communication
Le Qin, Junwei Cui, Weilin Cai, Chenyu Yuan, Jiayi Huang
IEEE/ACM International Symposium on Microarchitecture (MICRO), November 2026. (Accepted)
NB-Walker: Non-blocking Page Table Walker to Enhance Address Translation in NUMA GPUs
Zihang Chen, Lieven Eeckhout, Hongyuan Liu, Tianao Ge, Xinkai Wang, Jiayi Huang
IEEE/ACM International Symposium on Microarchitecture (MICRO), November 2026. (Accepted)
Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and Inference
Weilin Cai, Le Qin, Shwai He, Junwei Cui, Ang Li, Jiayi Huang
International Conference on Machine Learning (ICML), July 2026.
Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM Inference
Junwei Cui, Le Qin, Weilin Cai, Jiayi Huang
IEEE/ACM International Symposium on Computer Architecture (ISCA), June—July 2026.
cuPTW: Leveraging Idle Compute Units for Massively Parallel GPU Page Table Walks
Zihang Chen, Tianao Ge, Lieven Eeckhout, Hongyuan Liu, Jiayi Huang
ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, June 2026.
FlexTrain: Scalable Hybrid-Parallel Training with Elastic Resource Utilization and Consistent Accuracy
Weilin Cai*, Diandian Gu*, Baoquan Zhong*, Jun Wang*, Zhuolin Zheng*, Gaohong
Liu, Kaihua Jiang, Shuguang Wang, Wencong Xiao, Jiayi Huang
Ninth Annual Conference on Machine Learning and Systems (MLSys), May 2026.
Capacity-Aware Inference: Mitigating The Straggler Effect in Mixture of Experts
Shwai He, Weilin Cai, Jiayi Huang, Ang Li
International Conference on Learning Representations (ICLR), April 2026.
XTree on EquiMesh: Topology and Algorithm Co-Design for Collective Communication
Junwei Cui, Le Qin, Weilin Cai, Jiayi Huang
Design, Automation and Test in Europe Conference (DATE), April 2026.
Harmony: A Hardware-Mapping Co-Exploration Framework for Hybrid CIM-based Vision Transformer Accelerator
Yihang Zuo, Zexin Fu, Cong Wang, Yuchao Wu, Jiayi
Huang, Yuzhe Ma
International Symposium on Quality Electronic Design (ISQED), April 2026.
Best Paper Award
Software Prefetch Multicast: Sharer-Exposed Prefetching for Bandwidth Efficiency in Manycore Processors
Yanhua Chen, Jiong Feng, Zhe Wang, Christopher J. Hughes, Jiayi Huang
IEEE/ACM International Symposium on Microarchitecture (MICRO), October 2025.
Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks
Le Qin, Junwei Cui, Weilin Cai, Meng Niu, Yan Yang, Jiayi Huang
IEEE/ACM International Symposium on Microarchitecture (MICRO), October 2025.
Optimizing Heterogeneous Compute-in-Memory with Hybrid Dataflow and In-Network Reduction for Vision Transformer
Zexin Fu, Yihang Zuo, Yuzhe Ma, Jiayi Huang
IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), August 2025.
Chimera: Communication Fusion for Hybrid Parallelism in Large Language Models
Le Qin, Junwei Cui, Weilin Cai, Jiayi Huang
ACM/IEEE International Symposium on Computer Architecture (ISCA), June 2025.
Distinguished Artifact Award
[paper]
[link]
[bibtex]
[slides]
[poster]
[artifact]
[code]
TRACI: Network Acceleration of Input-Dynamic Communication for Large-Scale Deep Learning Recommendation Model
Guyue Huang, Hao Li, Le Qin, Jiayi Huang, Yangwook Kang, Yufei Ding, Yuan Xie
ACM/IEEE International Symposium on Computer Architecture (ISCA), June 2025.
[paper]
[link]
[bibtex]
[slides]
MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training
Weilin Cai, Le Qin, Jiayi Huang
ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), March—April 2025.
[paper]
[link]
[bibtex]
[slides]
[poster]
Push Multicast: A Speculative and Coherent Interconnect for Mitigating Manycore CPU Communication Bottleneck
Jiayi Huang, Yanhua Chen, Zhe Wang, Christopher J. Hughes, Yufei Ding, Yuan Xie
IEEE International Symposium on High Performance Computer Architecture (HPCA), March 2025.
[paper]
[link]
[bibtex]
[slides]
[artifact]
[code]
A Survey on Mixture of Experts in Large Language Models
Weilin Cai*, Juyong Jiang*, Fan Wang*, Jing Tang#, Sunghun Kim#, Jiayi Huang#
(*: Equal contribution, #: Co-corresponding authors)
IEEE Tranactions on Knowledge and Data Engineering, 37(7):3896—3915, March 2025.
[paper]
[link]
[bibtex]
NoCFuzzer: Automating NoC Verification in UVM
Ruiyang Ma, Jiayi Huang, Shijian Zhang, Yuan Xie, Guojie Luo
IEEE Tranactions on Computer-Aided Design of Integrated Circuits and Systems, 44(1):371—384, Jan. 2025.
[paper]
[link]
[bibtex]
Prefender: Prefetching Defender against Cache Side Channel Attacks as A Pretender
Luyi Li, Jiayi Huang, Lang Feng, Zhongfeng Wang
IEEE Tranactions on Computers, 73(6):1457—1471, June 2024.
[paper]
[link]
[bibtex]
An Endeavor to Industrialize Hardware Fuzzing: Automating NoC Verification in UVM
Ruiyang Ma, Huatao Zhao, Jiayi Huang, Shijian Zhang,
Guojie Luo
Design, Automation and Test in Europe Conference (DATE), March 2024.
[paper]
[link]
[bibtex]
ArchExplorer: Microarchitecture Exploration via Bottleneck Analysis
Chen Bai, Jiayi Huang, Xuechao Wei, Yuzhe Ma, Sicheng
Li, Hongzhong Zheng, Bei Yu, Yuan Xie
IEEE/ACM International Symposium on Microarchitecture (MICRO), October—November 2023.
[paper]
[link]
[bibtex]
[slides]
[poster]
[artifact]
WHISTLE: CPU Abstractions for Hardware and Software Memory Safety Invariants
Sungkeun Kim, Farabi Mahmud, Jiayi Huang, Pritam Majumder, Chia-Che Tsai, Abdullah Muzahid, EJ Kim
IEEE Tranactions on Computers, 72(3):811—825, March 2023.
[paper]
[link]
[bibtex]
MPU-Sim: A Simulator for In-DRAM Near-Bank Processing Architectures
Xinfeng Xie, Peng Gu, Jiayi Huang, Yufei Ding, Yuan Xie
IEEE Computer Architecture Letters (CAL), 21(1):1—4, January—June 2022.
[paper]
[link]
[bibtex]
[code]
RVDFI: A RISC-V Architecture with Security Enforcement by High Performance Complete Data-Flow Integrity
Lang Feng*, Jiayi Huang*, Luyi Li, Haochen Zhang, Zhongfeng Wang
IEEE Tranactions on Computers, 71(10):2499—2512, October 2022.
[paper]
[link]
[bibtex]
[code]
Toward Taming the Overhead Monster for Data-Flow Integrity
Lang Feng, Jiayi Huang, Jeff Huang, Jiang Hu
ACM Transactions on Design Automation of Electronic Systems (TODAES), 27(3):1—24, May 2022.
[paper]
[link]
[bibtex]
Prefender: Prefetching Defender against Cache Side Channel Attacks as A Pretender
Luyi Li, Jiayi Huang, Lang Feng,
and Zhongfeng Wang
Design, Automation and Test in Europe Conference (DATE), March 2022.
Best Paper Award Nominee
[paper]
[link]
[bibtex]
Communication Algorithm-Architecture Co-Design for Distributed Deep Learning
Jiayi Huang, Pritam Majumder,
Sungkeun Kim, Abdullah Muzahid, Ki Hwan Yum, and Eun Jung Kim
IEEE/ACM International Symposium on Computer Architecture (ISCA), June 2021.
[paper]
[link]
[bibtex]
[slides]
[lightning]
[poster]
[example]
A Voting Approach for Adaptive Network-on-Chip Power-Gating
Jiayi Huang, Shilpa Bhosekar,
Rahul Boyapati, Ningyuan Wang, Byul Hur, Ki Hwan Yum, and Eun Jung Kim
IEEE Transactions on Computers, 70(11):1962—1975, 2021.
[paper]
[link]
[bibtex]
[code]
Remote Control: A Simple Deadlock Avoidance Scheme for Modular Systems-on-Chip
Pritam Majumder, Sungkeun Kim, Jiayi Huang,
Ki Hwan Yum, and Eun Jung Kim
IEEE Transactions on Computers, 70(11):1928—1941, 2021.
[paper]
[link]
[bibtex]
Computing En-Route for Near-Data Processing
Jiayi Huang, Pritam Majumder,
Sungkeun Kim, Troy Fulton, Ramprakash Reddy Puli, Ki Hwan Yum, and Eun Jung Kim
IEEE Transactions on Computers, 70(6):906—921, 2021.
[paper]
[link]
[bibtex]
[code]
ReViCe: Reusing Victim Cache to Prevent Speculative Cache Leakage
Sungkeun Kim, Farabi Mahmud, Jiayi Huang,
Pritam Majumder, Neophytos Christou, Abdullah Muzahid, Chia-Che Tsai, and Eun Jung Kim
IEEE Security Development Conference (SecDev), September 2020.
[paper]
[link]
[slides]
[bibtex]
Active-Routing: Compute on the Way for Near-Data Processing
Jiayi Huang, Ramprakash Reddy Puli,
Pritam Majumder, Sungkeun Kim, Rahul Boyapati, Ki Hwan Yum, and Eun Jung Kim
IEEE International Symposium on High Performance Computer Architecture (HPCA) , February 2019.
[paper]
[link]
[slides]
[lightning]
[bibtex]
[code]
Approx-NoC: A Data Approximation Framework for Network-On-Chip Architectures
Rahul Boyapati, Jiayi Huang,
Pritam Majumder, Ki Hwan Yum, and Eun Jung Kim
IEEE/ACM International Symposium on Computer Architecture (ISCA) , July 2017.
[paper]
[link]
[slides]
[lightning]
[bibtex]
[benchmarks]
Packet Coalescing Exploiting Data Redundancy in GPGPU Architectures
Kyung Hoon Kim, Rahul Boyapati,
Jiayi Huang, Yuho Jin, Ki Hwan Yum, and Eun Jung Kim
ACM International Conference on Supercomputing (ICS), June 2017.
[paper]
[link]
[bibtex]
Fly-Over: A Light-Weight Distributed Power-Gating Mechanism for Energy-Efficient Networks-on-Chip
Rahul Boyapati*, Jiayi Huang*,
Ningyuan Wang, Kyung Hoon Kim, Ki Hwan Yum, and Eun Jung Kim
IEEE International Parallel & Distributed Processing Symposium (IPDPS), May-June 2017.
[paper]
[link]
[slides]
[bibtex]
[code]