Tag: blog
All the articles with the tag "blog".
-
I Nodded Along While Inference Geeks Talked GPUs — So I Wrote the Field Guide I Wish I Had
A comprehensive, from-zero guide to GPU hardware buzzwords: HBM, memory bandwidth, FP8/FP4, NVLink, H100 vs H200 vs B200 vs GB200, NVL72 racks, AMD Instinct, TPUs, Groq, Cerebras — plus what to actually run in a homelab and why datacenter GPU infra is genuinely hard.
-
I Ran vLLM on a Mac Mini With No GPU — Here's Everything I Learned About Inference
A complete, beginner-friendly guide to vLLM: what inference actually is, how to build vLLM from source on an Apple Silicon Mac with no GPU, every command explained, the three errors I hit and fixed, the flags that matter, and real throughput numbers from my living room.
-
OpenCode Doesn't Have an Auto Mode. So I Built One with vLLM Semantic Router
Cursor picks the right model for you automatically. OpenCode doesn't - yet. Here's a complete, production-shaped guide to adding intelligent auto model selection to OpenCode (or any OpenAI-compatible agent) using vLLM Semantic Router and AgentGateway.
-
I Almost Built a Grafana Stack—Then AgentGateway Shipped Everything I Needed.
I Almost Built a Grafana Stack—Then AgentGateway Shipped Everything I Needed.