About Me
I am Chao Fu, an Associate Researcher at Shaoxin Laboratory, Fudan University. I received my Ph.D. from Fudan University, where I was advised by Prof. Jun Han.
My research focuses on computer architecture and system-level design for embodied AI deployment. To this end, my research span:
- Chiplet-based heterogeneous architecture exploration and ESL design methodologies;
- Memory systems and cache coherence for heterogeneous SoCs;
- End-to-end deployment of LLMs and domain-specific architectures for embodied AI systems.
My research broadly explores how advanced computer architectures can support scalable, efficient, and deployable intelligent systems.
I also emphasize application-driven and industry-academia-oriented research, with demo systems including:
- Edge-side super-resolution for UAV scenarios [demo details];
- Interactive real-time VSLAM rendering on edge platforms [demo details];
- Hardware-software-in-the-loop systems for embodied AI [demo details].
Selected Publications
Someone† denotes equal contributions; Someone denotes corresponding authors.

SPMC: Speculative Multicast for GPU Address Translation Sharing
Zengshi Wang, Zhuoyuan Yang, Wenbin Liao, Xin Yang, Xiaoyan Xiang, Jianyi Meng, Chao Fu, Jun Han
- This work introduces SPMC, a sharing-aware speculative multicast mechanism that reduces repeated GPU TLB misses by proactively distributing address translations across SMs.

Chao Fu†, Jingyang Zheng†, Xinliu He, Xiangqi Dong, Zheng Cao, Yuning Zhan, Wenzhong Bao, Peng Zhou, Jun Han
- This work bridges 2D-material devices and GPU architecture, achieving substantial performance and energy gains through retention-aware cache design.

Constructing Virtual FIFOs in CPU Cache Hierarchy for Software Data Plane Orchestration
Zhiyuan Zhang, Jialin Liu, Zengshi Wang, Kanheng Jiang, Chao Fu, Jun Han
- This paper introduces FRM, which virtualizes per-RXQ FIFOs across the CPU cache hierarchy to improve software data-plane performance and cache isolation.

Chao Fu, Li Wan, Jun Han
- This work adaptively coordinates contention handling to improve multicore concurrency with minimal hardware cost.

Homonoia: Reinventing underutilized cross-level cache resource using reinforcement learning
Chao Fu, Li Wan, Yufan Jia, Zhiyuan Zhang, Kanheng Jiang, Jun Han
- This work repurposes underutilized on-chip caches with reinforcement learning to expand LLC capacity for cache-sensitive workloads.

Chimera: A co-simulation framework combining with gem5 and FPGA platform for efficient verification
Chao Fu, Zengshi Wang, Jun Han
- This work tightly integrates software simulation and hardware emulation to streamline SoC-level evaluation of RTL designs.

PCMT: Prioritizing Coherence Message Types for NoC Protocol-level Deadlock Freedom
Yufan Jia, Zhiyuan Zhang, Chao Fu, Li Wan, Jun Han
- This work removes virtual-network dependence by prioritizing coherence messages to resolve NoC protocol-level deadlocks efficiently.

Thoth: Uncovering Data-Dependent Memory Access Patterns via Annotation-Directed Load Sampling
Kanheng Jiang, Yongxin Lyu, Zhiyuan Zhang, Zengshi Wang, Chao Fu, Jun Han
- This work uses annotation-directed load sampling to expose irregular data-dependent memory patterns and improve hardware prefetching.