Jiaming (Jamin) Cheng

AIoT · Efficient ML · On-Device LLMs · The Ohio State University

prof_pic.jpeg

After completing my B.S. in Computer and Information Science (cum laude) at The Ohio State University in 2024, I have continued doing research with the group: day to day with Subhransu Das, on projects led by Profs. Rajiv Ramnath and Brijesh Soni.

I care about making capable AI usable without a data center behind it — running on the devices people already carry, and without shipping their data elsewhere to do it.

Doing that has taught me to distrust proxies for efficiency and go measure the thing itself. On the compression side the thing itself is the hardware: structured pruning, low-bit quantization, and knowledge distillation, with segmentation models cut by up to 87% in size and parameters and deployed on a Raspberry Pi inside a few watts, and phase-wise benchmarks that show where on-device LLM latency goes.

The same distrust applies away from the device. Multi-agent LLM systems that relay KV caches credit their gains to the “latent thoughts” they exchange; swapping in a mismatched cache shows a large cache effect need not be a pairing effect — across three released systems, one relay carries example-specific content, one carries it over a subset of layers, and for one the transfer is not detected.

I am applying to PhD programs for Fall 2027.

news

Oct 01, 2025 SPICE accepted to IEEE CCNC 2026 as an oral paper.

selected publications

  1. When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
    Jiaming Cheng, Subhransu Das, and Rajiv Ramnath
    2026
    Preprint
  2. Phase-Wise Analysis of LLM Inference Acceleration on GPU, CPU, and Edge Device
    Subhransu Das, Jiaming Cheng, Swathi Vallabhajosyula, and 2 more authors
    In Practice and Experience in Advanced Research Computing (PEARC ’26), 2026
    To appear
  3. SPICE: Structured Pruning for Inference on Constrained Edge Devices
    Subhransu Das, Jiaming Cheng, Aniruddha Rakshit, and 3 more authors
    In IEEE Consumer Communications & Networking Conference (CCNC), 2026
  4. EPIC: Efficient Pruning for Inference on Constrained Devices
    Subhransu Das, Jiaming Cheng, Aniruddha Rakshit, and 2 more authors
    In Practice and Experience in Advanced Research Computing (PEARC ’25), 2025