Writing

Case Studies

Real engagements, real numbers. How production AI infrastructure problems get solved.

FinOps for AI Azure AKS vLLM KEDA

How We Cut $48K/Year in GPU Inference Costs on Azure AKS

A FinTech AI startup burning $13K/month on GPU with no attribution. 60 days later: $9K/month, full per-model cost visibility, batch workloads on spot with KEDA scale-to-zero.

May 2026

8 min read

$48K

saved/yr

LLM Inference KServe Coming soon

Production vLLM + KServe on EKS: Zero to 800ms p95 Latency

Series A AI startup. Mistral 7B. KServe InferenceService, DCGM monitoring, Langfuse observability. End-to-end setup from scratch.

RAG Pipelines LangChain Coming soon

Building a RAG Pipeline That Actually Works in Production

Why the demo worked and production didn't. LangChain + Qdrant + Langfuse — adding observability to a RAG pipeline that had no visibility into retrieval failures.