research
Selected research projects, ongoing work, and systems contributions.
OPSERVE
Opportunistic LLM inference over fragmented GPU capacity in HPC systems
Batch-scheduled supercomputers can leave substantial GPU capacity temporarily unused when free resources do not match the requirements of queued jobs. OPSERVE explores how LLM inference can turn that fragmented capacity into useful serving capacity without assuming the resources will remain available.
The system combines a stable pool of persistent workers with opportunistically acquired transient workers. It adapts workers between colocated, prefill-only, and decode-only roles as both workload demand and resource availability change, while a shared KV-cache layer preserves completed prefill state when transient workers are reclaimed.
BatchFlow
Benefit-aware data pipeline allocation and batch reuse for multi-job training
Modern training jobs can stall not because accelerators are slow, but because data retrieval and preprocessing cannot supply mini-batches quickly enough. The problem becomes harder when multiple jobs compete for shared CPU, storage, memory, and network resources.
BatchFlow treats data preparation as a shared cluster service. It uses online profiling to direct a finite pool of data workers toward the jobs that benefit most, while coordinating prefetching and cache management so prepared mini-batches can be reused across concurrent jobs.
zkInfer
A distributed system for scalable zero-knowledge proofs of machine learning inference
Zero-knowledge machine learning can make outsourced inference verifiable without exposing sensitive model or intermediate state, but proof generation can be dramatically more expensive and memory-intensive than ordinary inference.
zkInfer approaches proof generation as a distributed systems problem. It decomposes a model inference into independently executable proving jobs, binds adjacent partitions with cryptographic commitments, and schedules those jobs across prover machines using runtime and memory estimates. Compiled circuits and proving keys are reused across requests to avoid repeating expensive setup work.