research

Selected research projects, ongoing work, and systems contributions.

LLM inference · HPC · resource management

OPSERVE

Opportunistic LLM inference over fragmented GPU capacity in HPC systems

Batch-scheduled supercomputers can leave substantial GPU capacity temporarily unused when free resources do not match the requirements of queued jobs. OPSERVE explores how LLM inference can turn that fragmented capacity into useful serving capacity without assuming the resources will remain available.

The system combines a stable pool of persistent workers with opportunistically acquired transient workers. It adapts workers between colocated, prefill-only, and decode-only roles as both workload demand and resource availability change, while a shared KV-cache layer preserves completed prefill state when transient workers are reclaimed.

9.81%of operational node-hours observed as unallocated in a year-long Polaris trace
1.5–2.1×lower p95 end-to-end latency versus static prefill/decode allocation when both sustain load
14.3×lower observed p95 latency under overload in the evaluated scenarios
LLM servingKV cacheprefill/decodemalleable systemsHPC
Training systems · data pipelines · scheduling

BatchFlow

Benefit-aware data pipeline allocation and batch reuse for multi-job training

GitHub ↗

Modern training jobs can stall not because accelerators are slow, but because data retrieval and preprocessing cannot supply mini-batches quickly enough. The problem becomes harder when multiple jobs compete for shared CPU, storage, memory, and network resources.

BatchFlow treats data preparation as a shared cluster service. It uses online profiling to direct a finite pool of data workers toward the jobs that benefit most, while coordinating prefetching and cache management so prepared mini-batches can be reused across concurrent jobs.

3.8×higher aggregate training throughput in the evaluated multi-job workloads
2.1×higher cost efficiency relative to the evaluated prior systems
Sharedworker allocation, prefetching, and reusable mini-batch caching across training jobs
training systemsdata pipelinesbatch reuseresource allocationPyTorch
Verifiable ML · distributed systems · zero knowledge

zkInfer

A distributed system for scalable zero-knowledge proofs of machine learning inference

Zero-knowledge machine learning can make outsourced inference verifiable without exposing sensitive model or intermediate state, but proof generation can be dramatically more expensive and memory-intensive than ordinary inference.

zkInfer approaches proof generation as a distributed systems problem. It decomposes a model inference into independently executable proving jobs, binds adjacent partitions with cryptographic commitments, and schedules those jobs across prover machines using runtime and memory estimates. Compiled circuits and proving keys are reused across requests to avoid repeating expensive setup work.

69×lower end-to-end latency than monolithic proving in the evaluated distributed configuration
48×lower peak per-machine memory in the reported experiments
5.6×lower latency from decomposition alone on a single prover worker
zkMLZK-SNARKsdistributed provingresource-aware schedulingHalo2 / EZKL