
Kunming SHAO 邵堃明
PhD Candidate · HKUST ECE & AI Chip Center for Emerging Smart Systems (ACCESS)
I am a PhD candidate at The Hong Kong University of Science and Technology (HKUST), advised by Prof. Chi-Ying Tsui and Prof. Tim Kwang-Ting Cheng. I work on memory-efficient AI from circuits to systems: digital compute-in-memory (CIM) macros and their design automation, CIM accelerators for Transformers and retrieval, and LLM serving, where I study KV-cache management for agents and sparse attention.
I expect to graduate in Summer 2027 and am looking for postdoctoral positions in AI hardware, computing-in-memory, and efficient intelligence systems.
News
- arXivTwo first-author preprints are on arXiv with code: EfficientAgent on KV-cache offloading for concurrent LLM agents, and PQ-HSA on hybrid sparse-approximate attention.
- ICCD'26Our paper HyMoE-CIM, a hybrid-bit SRAM/ReRAM CIM architecture with hardware-aware token dispatch for MoE LLM inference, was accepted.
- A-SSCC'26My first-authored paper PQ-SIM, a CIM-based accelerator for approximate vector search in RAG, was accepted.
- ICCAD'26Two co-first-authored papers, Delta-MoE and DBS-CIM, were accepted.
- GRFA GRF project that I helped apply for and participate in was funded.
Earlier news (20)
- AwardI received the HKUST RedBird Academic Excellence Award.
- ESSERC'26My first-authored paper SwiftCIM, a 55nm 23.2μJ/token ReRAM-coupled digital CIM accelerator for FlashAttention, was accepted.
- TVLSIThe paper I led and co-first-authored on FP8 digital CIM with on-the-fly aligned-mantissa bitwidth prediction was accepted.
- US PatentOur co-authored patent on a hybrid computing-in-memory device and multi-level sensing method was approved and published.
- CN PatentOur co-authored patent on a hybrid CIM device and multi-level data-bit sensing method was approved and published.
- DATE'26My first-authored paper DS-CIM, digital stochastic computing-in-memory with accurate OR-accumulation, was accepted.
- A-SSCC'25Our co-authored paper Lemem, a learning-in-memory processor for edge DNN/SNN training and inference, was accepted.
- BioCAS'25The paper I led and co-first-authored on memory-efficient retrieval for wearable medical LLM agents was accepted.
- TCADOur co-authored paper on configurable dataflow and adaptive mapping for hybrid ReRAM/SRAM CIM was accepted.
- AwardI received the IEEE CASS Student Travel Grant.
- ISLPED'25My first-authored paper DIRC-RAG, edge RAG acceleration with digital in-ReRAM computation, was accepted.
- DAC'25 WIPMy work-in-progress poster on AI accelerators based on approximate computing was accepted.
- CICC'25Our co-first-authored paper E-NPU, an event-driven adaptive neural SoC for personalized medical wearables, was accepted.
- ISCAS'25Our co-first-authored paper on a flexible precision-scaling DNN accelerator was accepted.
- PQEI passed the PhD Qualifying Exam and continued as a PhD candidate.
- DATE'25My first-authored paper SynDCIM, a performance-aware digital CIM compiler, was accepted.
- ICCAD'24Our co-authored paper ReSCIM, variation-resilient in-memory computation with hybrid multi-level ReRAM and SRAM cells, was accepted.
- ThesisMy BEng thesis, Digital Compute-In-Memory Automatic Design Methodology, was selected as an excellent graduation project at SCUT.
- AwardI received the Hong Kong PhD Fellowship and the HKUST RedBird Award.
- DAC'23Our co-authored paper AutoDCIM, an automated digital CIM compiler, was accepted.
Selected Research
All publications →arXiv 2026LLM systems
EfficientAgent: What Makes KV Cache Offloading Work for Concurrent Agents?
Agent prefixes are reused only if the host tier holds the reusable context of the whole agent pool. EfficientAgent sizes the host KV tier by this reuse working set and gates writes when the tier is too small, cutting recomputed prompt tokens by 93% and end-to-end time by 39% on SWE-bench Verified coding agents.
arXiv 2026LLM systems
PQ-HSA: Reusing Product-Quantized Scores for Hybrid Sparse-Approximate Attention
Sparse attention usually gives unread tokens zero weight. PQ-HSA reuses the IVF-PQ scores computed for retrieval so the unselected tokens enter the same softmax as an approximate background. At 128K context and a 1-2% budget it beats Quest and SnapKV, and its decode attention runs 1.6x faster than FlashAttention-3 inside vLLM on an H20.
A-SSCC 2026CIM for retrieval
PQ-SIM
A CIM-based accelerator for approximate vector search in RAG. It couples L-0.5 ReRAM with digital CIM macros to speed up IVF-PQ retrieval under tight edge memory and energy budgets.
ESSERC 2026CIM for Transformers
SwiftCIM: a 55nm 23.2μJ/Token L-0.5 ReRAM Coupled Digital CIM Accelerator with Fully-Fused Multi-Head Attention Dataflow for FlashAttention
A 55nm ReRAM-coupled digital CIM accelerator with a fully fused multi-head attention dataflow for FlashAttention, targeting low energy per token for Transformer inference.
IEEE TVLSI 2026CIM architecture
Balancing FP8 Computation Accuracy and Efficiency on Digital CIM via Shift-Aware On-the-fly Aligned-Mantissa Bitwidth Prediction
A software-hardware co-design for FP8 digital CIM that predicts aligned-mantissa bitwidth on the fly, balancing computation accuracy against energy efficiency during inference.
DATE 2026CIM architecture
DS-CIM: Digital Stochastic Computing-In-Memory Featuring Accurate OR-Accumulation via Sample Region Remapping for Edge AI Models
Digital stochastic computing-in-memory with accurate OR-accumulation via sample-region remapping, enabling efficient edge AI inference with reduced accumulation cost.
ISLPED 2025CIM for retrieval
DIRC-RAG: Accelerating Edge RAG with Robust High-Density and High-Loading-Bandwidth Digital In-ReRAM Computation
A high-density digital In-ReRAM computation architecture for edge RAG retrieval, combining robust MLC ReRAM storage with query-stationary similarity search.
DATE 2025CIM design automation
SynDCIM: A Performance-Aware Digital Computing-in-Memory Compiler with Multi-Spec-Oriented Subcircuit Synthesis
A performance-aware DCIM compiler that synthesizes multi-spec subcircuits under PPA constraints, automating architecture search through layout generation.
Background & Collaboration
Before HKUST, I received my BEng in Microelectronics from South China University of Technology (SCUT), ranked 1/120, with an excellent graduation thesis on digital CIM design automation. I hold the Hong Kong PhD Fellowship and the HKUST RedBird awards.
I coordinate a multi-institution collaboration across HKUST, SCUT, Westlake University, SYSU, WHU, and other partners on in-memory computation, approximate computing, efficient algorithms, and emerging non-volatile memories. The projects and my roles are listed in my CV.







