CV
Curriculum vitae.
Contact Information
| Name | Jiaming Cheng |
| Professional Title | Efficient ML for AIoT and On-Device LLMs |
| jiaming@jiamingcheng.me | |
| Location | Hong Kong SAR, China |
Research Interests
Efficient machine learning under compute, memory, and energy constraints, with emphasis on model compression and on-device LLM inference for edge/AIoT systems: structured pruning, low-bit quantization, knowledge distillation, and hardware-aware software optimization. A complementary interest is empirically auditing whether inference-time efficiency mechanisms deliver their claimed gains.
Education
-
2020 - 2024 Columbus, Ohio, USA
Research Experience
-
2024 - present Columbus, Ohio, USA
Doing research with Subhransu Das, on projects led by Profs. Rajiv Ramnath and Brijesh Soni
The Ohio State University
Efficient Machine Learning & Model Compression for Edge / AIoT.
- Designed and led a causal audit of latent (KV-cache) communication in multi-agent LLM systems (arXiv:2608.04893), showing that benchmark gains do not by themselves establish latent-content transmission.
- Audited three released latent-communication systems; replicated the cache-matching contrast across five checkpoints from three model families; and established equivalence on standard benchmarks under a pre-registered five-seed protocol, building a reusable audit harness.
- Co-designed and implemented a DepGraph-based L2 group-magnitude pruning method in EPIC (PEARC ‘25), cutting image segmentation models by 75±10% in size and 65±15% in parameters across the model set — 77–87% for DeepLabv3 with a ResNet backbone — with accuracy recovered by fine-tuning, and the Taylor-saliency pruning method TaLK in SPICE (CCNC 2026).
- Trained and distilled pruned models at the Ohio Supercomputer Center and deployed them to Raspberry Pi under a sub-30-second wall time at roughly 4 W average power, with wall latency falling 20–30% across sparsity levels.
- Built a phase-wise on-device LLM inference benchmark spanning 10,200 runs on GPU, CPU, and Raspberry Pi, showing that 4-bit quantization cuts edge decode latency by 55–66% while prefill stays nearly flat (PEARC ‘26).
- Contributed to a survey of edge AI deployment, released as an arXiv preprint.
Publications
-
2026 Large Models for Small Devices: Recent Advances and Empirical Analysis of Edge AI Deployment
Preprint, arXiv:2608.15693
-
2026 When Does Latent Communication Pay? A Causal Audit of Relayed KV Caches in Multi-Agent LLMs
Preprint, arXiv:2608.04893
-
2026 Phase-Wise Analysis of LLM Inference Acceleration on GPU, CPU, and Edge Device
Practice and Experience in Advanced Research Computing (PEARC '26), to appear
-
2026 SPICE: Structured Pruning for Inference on Constrained Edge Devices
IEEE Consumer Communications & Networking Conference (CCNC)
-
2025 EPIC: Efficient Pruning for Inference on Constrained Devices
Practice and Experience in Advanced Research Computing (PEARC '25)
Skills
Programming: Python, Rust, TypeScript
ML & Inference: PyTorch, llama.cpp, torch-pruning, Weights & Biases
Systems & Engineering: Slurm, Docker