News
Starting a PhD at NYU Courant!
Poster accepted to PyTorch Conference 2026.
Gave a talk at NVIDIA Developer Day.
Talk on vLLM at PyTorch ATX in Austin, TX.
Presented at vLLM Office Hours on custom torch.compile passes in vLLM. Also available as a vLLM blog post.
Gave a talk at the first NYC vLLM meetup.
Joined Red Hat following the Neural Magic acquisition!
Joined Neural Magic, a fast CPU inference startup.
Graduated from MIT with M.Eng in EECS. Thesis: "Improving the Performance of Parallel Loops in OpenCilk".
vLLM
I'm a maintainer of vLLM, where I focus on efficient and portable execution of the model forward pass:torch.compile integration, fusion passes, kernels, and CUDA graphs. I am a code owner for the compilation infrastructure, as well as hardware-agnostic model definitions and related portability abstractions.
Need a PR reviewed? See instructions here. If the PR explicitly touches one of the areas mentioned above, you can DM me on vLLM Slack. Please include a short description of the PR and justification why you want me to review it. NB: I don't check GitHub mentions for the vllm repo (20+ notifications a day).
If you're looking for advice on vLLM deployment and performance, I also offer consultations. Please email me for more information.
Contact
Email is best (see top of page). For anything vLLM-related, Slack gets a faster answer.
