Jiayi Huang

Publications [Google Scholar] [DBLP] [ORCID]

NOTE: Names with underline are my students

Glaive: Cleaving Bandwidth and Latency Scheduling for Efficient Non-uniform All-to-All Collective Communication

Le Qin, Junwei Cui, Weilin Cai, Chenyu Yuan, Jiayi Huang

IEEE/ACM International Symposium on Microarchitecture (MICRO), November 2026. (Accepted)

NB-Walker: Non-blocking Page Table Walker to Enhance Address Translation in NUMA GPUs

Zihang Chen, Lieven Eeckhout, Hongyuan Liu, Tianao Ge, Xinkai Wang, Jiayi Huang

IEEE/ACM International Symposium on Microarchitecture (MICRO), November 2026. (Accepted)

Mining Tensor/Neuron-Level Sparsity to Maximize Mixture-of-Experts Potential in Post-Training and Inference

Weilin Cai, Le Qin, Shwai He, Junwei Cui, Ang Li, Jiayi Huang

International Conference on Machine Learning (ICML), July 2026.

Mapping and Communication Optimizations with Fault Tolerance for Wafer-Scale LLM Inference

Junwei Cui, Le Qin, Weilin Cai, Jiayi Huang

IEEE/ACM International Symposium on Computer Architecture (ISCA), June—July 2026.

cuPTW: Leveraging Idle Compute Units for Massively Parallel GPU Page Table Walks

Zihang Chen, Tianao Ge, Lieven Eeckhout, Hongyuan Liu, Jiayi Huang

ACM SIGMETRICS International Conference on Measurement and Modeling of Computer Systems, June 2026.

FlexTrain: Scalable Hybrid-Parallel Training with Elastic Resource Utilization and Consistent Accuracy

Weilin Cai*, Diandian Gu*, Baoquan Zhong*, Jun Wang*, Zhuolin Zheng*, Gaohong Liu, Kaihua Jiang, Shuguang Wang, Wencong Xiao, Jiayi Huang

Ninth Annual Conference on Machine Learning and Systems (MLSys), May 2026.

Capacity-Aware Inference: Mitigating The Straggler Effect in Mixture of Experts

Shwai He, Weilin Cai, Jiayi Huang, Ang Li

International Conference on Learning Representations (ICLR), April 2026.

XTree on EquiMesh: Topology and Algorithm Co-Design for Collective Communication

Junwei Cui, Le Qin, Weilin Cai, Jiayi Huang

Design, Automation and Test in Europe Conference (DATE), April 2026.

Harmony: A Hardware-Mapping Co-Exploration Framework for Hybrid CIM-based Vision Transformer Accelerator

Yihang Zuo, Zexin Fu, Cong Wang, Yuchao Wu, Jiayi Huang, Yuzhe Ma

International Symposium on Quality Electronic Design (ISQED), April 2026.

Best Paper Award

Software Prefetch Multicast: Sharer-Exposed Prefetching for Bandwidth Efficiency in Manycore Processors

Yanhua Chen, Jiong Feng, Zhe Wang, Christopher J. Hughes, Jiayi Huang

IEEE/ACM International Symposium on Microarchitecture (MICRO), October 2025.

Optimizing All-to-All Collective Communication with Fault Tolerance on Torus Networks

Le Qin, Junwei Cui, Weilin Cai, Meng Niu, Yan Yang, Jiayi Huang

IEEE/ACM International Symposium on Microarchitecture (MICRO), October 2025.

Optimizing Heterogeneous Compute-in-Memory with Hybrid Dataflow and In-Network Reduction for Vision Transformer

Zexin Fu, Yihang Zuo, Yuzhe Ma, Jiayi Huang

IEEE/ACM International Symposium on Low Power Electronics and Design (ISLPED), August 2025.

Chimera: Communication Fusion for Hybrid Parallelism in Large Language Models

Le Qin, Junwei Cui, Weilin Cai, Jiayi Huang

ACM/IEEE International Symposium on Computer Architecture (ISCA), June 2025.

Distinguished Artifact Award

[paper] [link] [bibtex] [slides] [poster] [artifact] [code]

TRACI: Network Acceleration of Input-Dynamic Communication for Large-Scale Deep Learning Recommendation Model

Guyue Huang, Hao Li, Le Qin, Jiayi Huang, Yangwook Kang, Yufei Ding, Yuan Xie

ACM/IEEE International Symposium on Computer Architecture (ISCA), June 2025.

[paper] [link] [bibtex] [slides]

MoC-System: Efficient Fault Tolerance for Sparse Mixture-of-Experts Model Training

Weilin Cai, Le Qin, Jiayi Huang

ACM International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), March—April 2025.

[paper] [link] [bibtex] [slides] [poster]

Push Multicast: A Speculative and Coherent Interconnect for Mitigating Manycore CPU Communication Bottleneck

Jiayi Huang, Yanhua Chen, Zhe Wang, Christopher J. Hughes, Yufei Ding, Yuan Xie

IEEE International Symposium on High Performance Computer Architecture (HPCA), March 2025.

[paper] [link] [bibtex] [slides] [artifact] [code]

A Survey on Mixture of Experts in Large Language Models

Weilin Cai*, Juyong Jiang*, Fan Wang*, Jing Tang#, Sunghun Kim#, Jiayi Huang#

(*: Equal contribution, #: Co-corresponding authors)

IEEE Tranactions on Knowledge and Data Engineering, 37(7):3896—3915, March 2025.

[paper] [link] [bibtex]

NoCFuzzer: Automating NoC Verification in UVM

Ruiyang Ma, Jiayi Huang, Shijian Zhang, Yuan Xie, Guojie Luo

IEEE Tranactions on Computer-Aided Design of Integrated Circuits and Systems, 44(1):371—384, Jan. 2025.

[paper] [link] [bibtex]

Prefender: Prefetching Defender against Cache Side Channel Attacks as A Pretender

Luyi Li, Jiayi Huang, Lang Feng, Zhongfeng Wang

IEEE Tranactions on Computers, 73(6):1457—1471, June 2024.

[paper] [link] [bibtex]

An Endeavor to Industrialize Hardware Fuzzing: Automating NoC Verification in UVM

Ruiyang Ma, Huatao Zhao, Jiayi Huang, Shijian Zhang, Guojie Luo

Design, Automation and Test in Europe Conference (DATE), March 2024.

[paper] [link] [bibtex]

ArchExplorer: Microarchitecture Exploration via Bottleneck Analysis

Chen Bai, Jiayi Huang, Xuechao Wei, Yuzhe Ma, Sicheng Li, Hongzhong Zheng, Bei Yu, Yuan Xie

IEEE/ACM International Symposium on Microarchitecture (MICRO), October—November 2023.

[paper] [link] [bibtex] [slides] [poster] [artifact]

WHISTLE: CPU Abstractions for Hardware and Software Memory Safety Invariants

Sungkeun Kim, Farabi Mahmud, Jiayi Huang, Pritam Majumder, Chia-Che Tsai, Abdullah Muzahid, EJ Kim

IEEE Tranactions on Computers, 72(3):811—825, March 2023.

[paper] [link] [bibtex]

MPU-Sim: A Simulator for In-DRAM Near-Bank Processing Architectures

Xinfeng Xie, Peng Gu, Jiayi Huang, Yufei Ding, Yuan Xie

IEEE Computer Architecture Letters (CAL), 21(1):1—4, January—June 2022.

[paper] [link] [bibtex] [code]

RVDFI: A RISC-V Architecture with Security Enforcement by High Performance Complete Data-Flow Integrity

Lang Feng*, Jiayi Huang*, Luyi Li, Haochen Zhang, Zhongfeng Wang

IEEE Tranactions on Computers, 71(10):2499—2512, October 2022.

[paper] [link] [bibtex] [code]

Toward Taming the Overhead Monster for Data-Flow Integrity

Lang Feng, Jiayi Huang, Jeff Huang, Jiang Hu

ACM Transactions on Design Automation of Electronic Systems (TODAES), 27(3):1—24, May 2022.

[paper] [link] [bibtex]

Prefender: Prefetching Defender against Cache Side Channel Attacks as A Pretender

Luyi Li, Jiayi Huang, Lang Feng, and Zhongfeng Wang

Design, Automation and Test in Europe Conference (DATE), March 2022.

Best Paper Award Nominee

[paper] [link] [bibtex]

Communication Algorithm-Architecture Co-Design for Distributed Deep Learning

Jiayi Huang, Pritam Majumder, Sungkeun Kim, Abdullah Muzahid, Ki Hwan Yum, and Eun Jung Kim

IEEE/ACM International Symposium on Computer Architecture (ISCA), June 2021.

[paper] [link] [bibtex] [slides] [lightning] [poster] [example]

A Voting Approach for Adaptive Network-on-Chip Power-Gating

Jiayi Huang, Shilpa Bhosekar, Rahul Boyapati, Ningyuan Wang, Byul Hur, Ki Hwan Yum, and Eun Jung Kim

IEEE Transactions on Computers, 70(11):1962—1975, 2021.

[paper] [link] [bibtex] [code]

Remote Control: A Simple Deadlock Avoidance Scheme for Modular Systems-on-Chip

Pritam Majumder, Sungkeun Kim, Jiayi Huang, Ki Hwan Yum, and Eun Jung Kim

IEEE Transactions on Computers, 70(11):1928—1941, 2021.

[paper] [link] [bibtex]

Computing En-Route for Near-Data Processing

Jiayi Huang, Pritam Majumder, Sungkeun Kim, Troy Fulton, Ramprakash Reddy Puli, Ki Hwan Yum, and Eun Jung Kim

IEEE Transactions on Computers, 70(6):906—921, 2021.

[paper] [link] [bibtex] [code]

ReViCe: Reusing Victim Cache to Prevent Speculative Cache Leakage

Sungkeun Kim, Farabi Mahmud, Jiayi Huang, Pritam Majumder, Neophytos Christou, Abdullah Muzahid, Chia-Che Tsai, and Eun Jung Kim

IEEE Security Development Conference (SecDev), September 2020.

[paper] [link] [slides] [bibtex]

Active-Routing: Compute on the Way for Near-Data Processing

Jiayi Huang, Ramprakash Reddy Puli, Pritam Majumder, Sungkeun Kim, Rahul Boyapati, Ki Hwan Yum, and Eun Jung Kim

IEEE International Symposium on High Performance Computer Architecture (HPCA) , February 2019.

[paper] [link] [slides] [lightning] [bibtex] [code]

Approx-NoC: A Data Approximation Framework for Network-On-Chip Architectures

Rahul Boyapati, Jiayi Huang, Pritam Majumder, Ki Hwan Yum, and Eun Jung Kim

IEEE/ACM International Symposium on Computer Architecture (ISCA) , July 2017.

[paper] [link] [slides] [lightning] [bibtex] [benchmarks]

Packet Coalescing Exploiting Data Redundancy in GPGPU Architectures

Kyung Hoon Kim, Rahul Boyapati, Jiayi Huang, Yuho Jin, Ki Hwan Yum, and Eun Jung Kim

ACM International Conference on Supercomputing (ICS), June 2017.

[paper] [link] [bibtex]

Fly-Over: A Light-Weight Distributed Power-Gating Mechanism for Energy-Efficient Networks-on-Chip

Rahul Boyapati*, Jiayi Huang*, Ningyuan Wang, Kyung Hoon Kim, Ki Hwan Yum, and Eun Jung Kim

IEEE International Parallel & Distributed Processing Symposium (IPDPS), May-June 2017.

[paper] [link] [slides] [bibtex] [code]

Poster and Preprints

PLFC: Preemptive Lossy Flow Control for Chiplet-based Systems

Jiong Feng, Yanhua Chen, Kunlin Li, Zhengrong Wang, Jiayi Huang

IEEE Chiplet Workshop: Standard, Circuits and Systems, October 2025.

Best Poster Award

[poster]

Coloring Big Graphs with AlphaGoZero

Jiayi Huang, Mostofa Patwary, Gregory Diamos

arXiv preprint arXiv:1902.10162, 2019.

[preprint] [bibtex] [media]

Fly-Over: A Light-Weight Distributed Power-Gating Mechanism for Energy-Efficient Networks-on-Chip

Rahul Boyapati*, Jiayi Huang*, Ningyuan Wang, Kyung Hoon Kim, Ki Hwan Yum, and Eun Jung Kim

International Conference on Parallel Architectures and Compilatin Techniques (PACT), September 2016.

[paper] [poster] [bibtex]