Skip to content
Advanced Patterns

AI Engineering in Production

From RAG pipelines and vector databases to MCP servers and agent security — the operational patterns for shipping LLM-backed systems that survive contact with real traffic.

8 Lessons

Articles in this series

1
17 min read•Hard

Building Production RAG Pipelines: Chunking, Embeddings, and Retrieval at Scale

Build RAG systems that work in production: chunking strategies, embedding selection, pgvector ops, and retrieval quality evaluation.

2
17 min read•Hard

Vector Databases Compared: pgvector vs Pinecone vs Weaviate

Compare pgvector, Pinecone, Weaviate and Milvus (plus Qdrant and LanceDB briefly) on operational fit, scaling and pricing — with real code and each vendor's own documentation.

3
14 min read•Medium

LLM API Integration Patterns for Backend Engineers

Production LLM API patterns: streaming, function calling, retries, token budgets, cost optimization, and observability for backend engineers.

4
16 min read•Medium

Your LLM Provider Will Have an Outage: Circuit Breakers, Fallback Chains, and Degraded Modes

OpenAI and Anthropic both document outages, overload codes, and best-effort availability. The survival kit: LLM-aware circuit breakers, same-model-first fallback chains, and degraded modes designed before you need them.

5
14 min read•Hard

Spring AI in Production: RAG Pipelines, Reliability, and Observability for Java Backends

Spring AI 1.1 deep-dive: production RAG pipeline with PII scrubbing, circuit breakers, Micrometer observability, and answer evaluation.

6
14 min read•Hard

Building an MCP Server in Go with Code Mode: From 1.17M Tokens to 1,000

Cloudflare put its 2,500+ API endpoints behind one MCP server with two tools, search + execute, and cut input tokens by 99.9%. The pattern in Go.

7
24 min read•Hard

Designing a Multi-Agent Backend: The Orchestrator Pattern

Picture one agent, one context window, one serial loop — until it hits 180K tokens and silently drops 37 of the 40 services it was asked to audit. The orchestrator pattern fans work out to isolated sub-agents in parallel, then synthesizes. Here's the backend, in compiling Go.

8
13 min read•Hard

Securing AI Agent Infrastructure: MCP Servers, Tool Calls, and the Attack Surface You're Not Watching

AI agents calling tools via MCP create new attack surfaces: prompt injection through tool responses, credential leakage, and unauthorized execution.