About Me

I am Chao Fu, an Associate Researcher at Shaoxin Laboratory, Fudan University. I received my Ph.D. from Fudan University, where I was advised by Prof. Jun Han.

My research focuses on computer architecture and system-level design for embodied AI deployment. To this end, my research span:

  • Chiplet-based heterogeneous architecture exploration and ESL design methodologies;
  • Memory systems and cache coherence for heterogeneous SoCs;
  • End-to-end deployment of LLMs and domain-specific architectures for embodied AI systems.

My research broadly explores how advanced computer architectures can support scalable, efficient, and deployable intelligent systems.

I also emphasize application-driven and industry-academia-oriented research, with demo systems including:

Selected Publications

Someone denotes equal contributions; Someone denotes corresponding authors.

MICRO 2026
TDMSim publication preview

SPMC: Speculative Multicast for GPU Address Translation Sharing

Zengshi Wang, Zhuoyuan Yang, Wenbin Liao, Xin Yang, Xiaoyan Xiang, Jianyi Meng, Chao Fu, Jun Han

  • This work introduces SPMC, a sharing-aware speculative multicast mechanism that reduces repeated GPU TLB misses by proactively distributing address translations across SMs.
ISCA 2026
TDMSim publication preview

TDMSim: Enabling High-Density and Energy-Efficient GPU DRAM Caches with 2D-Materials for Data-Intensive Applications

Chao Fu, Jingyang Zheng, Xinliu He, Xiangqi Dong, Zheng Cao, Yuning Zhan, Wenzhong Bao, Peng Zhou, Jun Han

  • This work bridges 2D-material devices and GPU architecture, achieving substantial performance and energy gains through retention-aware cache design.
TC 2026
Losatm publication preview

Constructing Virtual FIFOs in CPU Cache Hierarchy for Software Data Plane Orchestration

Zhiyuan Zhang, Jialin Liu, Zengshi Wang, Kanheng Jiang, Chao Fu, Jun Han

  • This paper introduces FRM, which virtualizes per-RXQ FIFOs across the CPU cache hierarchy to improve software data-plane performance and cache isolation.
TPDS 2022
Losatm publication preview

Losatm: A hardware transactional memory integrated with a low-overhead scenario-awareness conflict manager

Chao Fu, Li Wan, Jun Han

  • This work adaptively coordinates contention handling to improve multicore concurrency with minimal hardware cost.
TCAD 2025
Homonoia publication preview

Homonoia: Reinventing underutilized cross-level cache resource using reinforcement learning

Chao Fu, Li Wan, Yufan Jia, Zhiyuan Zhang, Kanheng Jiang, Jun Han

  • This work repurposes underutilized on-chip caches with reinforcement learning to expand LLC capacity for cache-sensitive workloads.
FPL 2024
Chimera publication preview

Chimera: A co-simulation framework combining with gem5 and FPGA platform for efficient verification

Chao Fu, Zengshi Wang, Jun Han

  • This work tightly integrates software simulation and hardware emulation to streamline SoC-level evaluation of RTL designs.
TCAD 2025
PCMT publication preview

PCMT: Prioritizing Coherence Message Types for NoC Protocol-level Deadlock Freedom

Yufan Jia, Zhiyuan Zhang, Chao Fu, Li Wan, Jun Han

  • This work removes virtual-network dependence by prioritizing coherence messages to resolve NoC protocol-level deadlocks efficiently.
TACO 2026
Thoth publication preview

Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling

Kanheng Jiang, Yongxin Lyu, Zhiyuan Zhang, Zengshi Wang, Chao Fu, Jun Han

  • This work uses annotation-directed load sampling to expose irregular data-dependent memory patterns and improve hardware prefetching.