CV

Curriculum vitae.

Contact Information

Name Jiaming Cheng
Professional Title Efficient ML for AIoT and On-Device LLMs
Email jiaming@jiamingcheng.me
Location Hong Kong SAR, China

Research Interests

Efficient machine learning under compute, memory, and energy constraints, with emphasis on model compression and on-device LLM inference for edge/AIoT systems: structured pruning, low-bit quantization, knowledge distillation, and hardware-aware software optimization. A complementary interest is empirically auditing whether inference-time efficiency mechanisms deliver their claimed gains.

Education

  • 2020 - 2024

    Columbus, Ohio, USA

    B.S., cum laude
    The Ohio State University
    Computer and Information Science
    • GPA: 3.51/4.00

Research Experience

  • 2024 - present

    Columbus, Ohio, USA

    Doing research with Subhransu Das, on projects led by Profs. Rajiv Ramnath and Brijesh Soni
    The Ohio State University
    Efficient Machine Learning & Model Compression for Edge / AIoT.
    • Designed and led a causal audit of latent (KV-cache) communication in multi-agent LLM systems (arXiv:2608.04893), showing that benchmark gains do not by themselves establish latent-content transmission.
    • Audited three released latent-communication systems; replicated the cache-matching contrast across five checkpoints from three model families; and established equivalence on standard benchmarks under a pre-registered five-seed protocol, building a reusable audit harness.
    • Co-designed and implemented a DepGraph-based L2 group-magnitude pruning method in EPIC (PEARC ‘25), cutting image segmentation models by 75±10% in size and 65±15% in parameters across the model set — 77–87% for DeepLabv3 with a ResNet backbone — with accuracy recovered by fine-tuning, and the Taylor-saliency pruning method TaLK in SPICE (CCNC 2026).
    • Trained and distilled pruned models at the Ohio Supercomputer Center and deployed them to Raspberry Pi under a sub-30-second wall time at roughly 4 W average power, with wall latency falling 20–30% across sparsity levels.
    • Built a phase-wise on-device LLM inference benchmark spanning 10,200 runs on GPU, CPU, and Raspberry Pi, showing that 4-bit quantization cuts edge decode latency by 55–66% while prefill stays nearly flat (PEARC ‘26).
    • Contributed to a survey of edge AI deployment, released as an arXiv preprint.

Publications

Skills

Programming: Python, Rust, TypeScript
ML & Inference: PyTorch, llama.cpp, torch-pruning, Weights & Biases
Systems & Engineering: Slurm, Docker