Writing
Case Studies
Real engagements, real numbers. How production AI infrastructure problems get solved.
FinOps for AI
Azure AKS
vLLM
KEDA
How We Cut $48K/Year in GPU Inference Costs on Azure AKS
A FinTech AI startup burning $13K/month on GPU with no attribution. 60 days later: $9K/month, full per-model cost visibility, batch workloads on spot with KEDA scale-to-zero.
May 2026
8 min read
$48K
saved/yr
LLM Inference
KServe
Coming soon
Production vLLM + KServe on EKS: Zero to 800ms p95 Latency
Series A AI startup. Mistral 7B. KServe InferenceService, DCGM monitoring, Langfuse observability. End-to-end setup from scratch.
RAG Pipelines
LangChain
Coming soon
Building a RAG Pipeline That Actually Works in Production
Why the demo worked and production didn't. LangChain + Qdrant + Langfuse — adding observability to a RAG pipeline that had no visibility into retrieval failures.