Kunming Shao

Kunming SHAO 邵堃明

PhD Candidate · HKUST ECE & AI Chip Center for Emerging Smart Systems (ACCESS)

I am a PhD candidate at The Hong Kong University of Science and Technology (HKUST), advised by Prof. Chi-Ying Tsui and Prof. Tim Kwang-Ting Cheng. I work on memory-efficient AI from circuits to systems: digital compute-in-memory (CIM) macros and their design automation, CIM accelerators for Transformers and retrieval, and LLM serving, where I study KV-cache management for agents and sparse attention.

I expect to graduate in Summer 2027 and am looking for postdoctoral positions in AI hardware, computing-in-memory, and efficient intelligence systems.

News

  • arXivTwo first-author preprints are on arXiv with code: EfficientAgent on KV-cache offloading for concurrent LLM agents, and PQ-HSA on hybrid sparse-approximate attention.
  • ICCD'26Our paper HyMoE-CIM, a hybrid-bit SRAM/ReRAM CIM architecture with hardware-aware token dispatch for MoE LLM inference, was accepted.
  • A-SSCC'26My first-authored paper PQ-SIM, a CIM-based accelerator for approximate vector search in RAG, was accepted.
  • ICCAD'26Two co-first-authored papers, Delta-MoE and DBS-CIM, were accepted.
  • GRFA GRF project that I helped apply for and participate in was funded.
Earlier news (20)
  • AwardI received the HKUST RedBird Academic Excellence Award.
  • ESSERC'26My first-authored paper SwiftCIM, a 55nm 23.2μJ/token ReRAM-coupled digital CIM accelerator for FlashAttention, was accepted.
  • TVLSIThe paper I led and co-first-authored on FP8 digital CIM with on-the-fly aligned-mantissa bitwidth prediction was accepted.
  • US PatentOur co-authored patent on a hybrid computing-in-memory device and multi-level sensing method was approved and published.
  • CN PatentOur co-authored patent on a hybrid CIM device and multi-level data-bit sensing method was approved and published.
  • DATE'26My first-authored paper DS-CIM, digital stochastic computing-in-memory with accurate OR-accumulation, was accepted.
  • A-SSCC'25Our co-authored paper Lemem, a learning-in-memory processor for edge DNN/SNN training and inference, was accepted.
  • BioCAS'25The paper I led and co-first-authored on memory-efficient retrieval for wearable medical LLM agents was accepted.
  • TCADOur co-authored paper on configurable dataflow and adaptive mapping for hybrid ReRAM/SRAM CIM was accepted.
  • AwardI received the IEEE CASS Student Travel Grant.
  • ISLPED'25My first-authored paper DIRC-RAG, edge RAG acceleration with digital in-ReRAM computation, was accepted.
  • DAC'25 WIPMy work-in-progress poster on AI accelerators based on approximate computing was accepted.
  • CICC'25Our co-first-authored paper E-NPU, an event-driven adaptive neural SoC for personalized medical wearables, was accepted.
  • ISCAS'25Our co-first-authored paper on a flexible precision-scaling DNN accelerator was accepted.
  • PQEI passed the PhD Qualifying Exam and continued as a PhD candidate.
  • DATE'25My first-authored paper SynDCIM, a performance-aware digital CIM compiler, was accepted.
  • ICCAD'24Our co-authored paper ReSCIM, variation-resilient in-memory computation with hybrid multi-level ReRAM and SRAM cells, was accepted.
  • ThesisMy BEng thesis, Digital Compute-In-Memory Automatic Design Methodology, was selected as an excellent graduation project at SCUT.
  • AwardI received the Hong Kong PhD Fellowship and the HKUST RedBird Award.
  • DAC'23Our co-authored paper AutoDCIM, an automated digital CIM compiler, was accepted.

Selected Research

All publications →

arXiv 2026LLM systems

EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?

Agent prefixes are reused only if the host tier holds the reusable context of the whole agent pool. EfficientAgent sizes the host KV tier by this reuse working set and gates writes when the tier is too small, cutting recomputed prompt tokens by 93% and end-to-end time by 39% on SWE-bench Verified coding agents.

First author arXivCode

arXiv 2026LLM systems

PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention

Sparse attention usually gives unread tokens zero weight. PQ-HSA reuses the IVF-PQ scores computed for retrieval so the unselected tokens enter the same softmax as an approximate background. At 128K context and a 1-2% budget it beats Quest and SnapKV, and its decode attention runs 1.6x faster than FlashAttention-3 inside vLLM on an H20.

First author arXivCode

A-SSCC 2026CIM for retrieval

PQ-SIM

A CIM-based accelerator for approximate vector search in RAG. It couples L-0.5 ReRAM with digital CIM macros to speed up IVF-PQ retrieval under tight edge memory and energy budgets.

First author

Background & Collaboration

Before HKUST, I received my BEng in Microelectronics from South China University of Technology (SCUT), ranked 1/120, with an excellent graduation thesis on digital CIM design automation. I hold the Hong Kong PhD Fellowship and the HKUST RedBird awards.

I coordinate a multi-institution collaboration across HKUST, SCUT, Westlake University, SYSU, WHU, and other partners on in-memory computation, approximate computing, efficient algorithms, and emerging non-volatile memories. The projects and my roles are listed in my CV.