AI Engineering in Production
From RAG pipelines and vector databases to MCP servers and agent security — the operational patterns for shipping LLM-backed systems that survive contact with real traffic.
Articles in this series
Building Production RAG Pipelines: Chunking, Embeddings, and Retrieval at Scale
Build RAG systems that work in production: chunking strategies, embedding selection, pgvector ops, and retrieval quality evaluation.
Vector Databases Compared: pgvector vs Pinecone vs Weaviate
Compare pgvector, Pinecone, Weaviate and Milvus (plus Qdrant and LanceDB briefly) on operational fit, scaling and pricing — with real code and each vendor's own documentation.
LLM API Integration Patterns for Backend Engineers
Production LLM API patterns: streaming, function calling, retries, token budgets, cost optimization, and observability for backend engineers.
Your LLM Provider Will Have an Outage: Circuit Breakers, Fallback Chains, and Degraded Modes
OpenAI and Anthropic both document outages, overload codes, and best-effort availability. The survival kit: LLM-aware circuit breakers, same-model-first fallback chains, and degraded modes designed before you need them.
Spring AI in Production: RAG Pipelines, Reliability, and Observability for Java Backends
Spring AI 1.1 deep-dive: production RAG pipeline with PII scrubbing, circuit breakers, Micrometer observability, and answer evaluation.
Building an MCP Server in Go with Code Mode: From 1.17M Tokens to 1,000
Cloudflare put its 2,500+ API endpoints behind one MCP server with two tools, search + execute, and cut input tokens by 99.9%. The pattern in Go.
Designing a Multi-Agent Backend: The Orchestrator Pattern
Picture one agent, one context window, one serial loop — until it hits 180K tokens and silently drops 37 of the 40 services it was asked to audit. The orchestrator pattern fans work out to isolated sub-agents in parallel, then synthesizes. Here's the backend, in compiling Go.
Securing AI Agent Infrastructure: MCP Servers, Tool Calls, and the Attack Surface You're Not Watching
AI agents calling tools via MCP create new attack surfaces: prompt injection through tool responses, credential leakage, and unauthorized execution.