Hi, my name is Anup.
RSS FeedI build and debug large-scale distributed systems, distributed databases, distributed file systems, and the Linux kernel — and I work increasingly on the systems layer of AI: LLM inference, model serving, and intelligent request routing.
Currently, I’m a Member of Technical Staff on the core storage data path at Nutanix .
Recent Posts
-
I Thought ASR Needed Its Own vLLM — Here’s How Speech Models Are Actually Served
Inside the ASR inference engine: model execution, batching, streaming state, parallelism, multilingual serving, open-source runtimes, and the breakthroughs that made modern speech AI practical.
-
I Nodded Along While Inference Geeks Talked GPUs — So I Wrote the Field Guide I Wish I Had
A comprehensive, from-zero guide to GPU hardware buzzwords: HBM, memory bandwidth, FP8/FP4, NVLink, H100 vs H200 vs B200 vs GB200, NVL72 racks, AMD Instinct, TPUs, Groq, Cerebras — plus what to actually run in a homelab and why datacenter GPU infra is genuinely hard.
-
Production Inference Optimization Strategy — Full Deck (All Deliverables)
The complete slide deck covering the production inference optimization strategy: architecture, benchmarks, routing, and cost analysis.
-
Architectural Specification — Production Inference Optimization Strategy
Architectural specification for a production inference optimization strategy for the agentic inference cloud.