Skip to content

Vector Databases Compared: pgvector vs Pinecone vs Weaviate

2
Part of Series: AI Engineering in Production

Lesson 2 of 8

A plain in-place REINDEX INDEX on a 12M-row pgvector table can stall a RAG ingestion pipeline until it's done. A non-concurrent rebuild locks out writes on the table and takes an AccessExclusiveLock on the HNSW index⁠[1]; because the planner locks every index of a table, nearly every query against it waits too — only cached prepared plans that skip that index get through⁠[2]. The INSERTs writing new embeddings queue up, retrieval queries queue with them, and the chatbot stops answering. The fix — REINDEX INDEX CONCURRENTLY, whose ShareUpdateExclusiveLock lets writes through, or a build-new-and-swap — is exactly the kind of thing a feature matrix never tells you.

Sharp edges like that are why this comparison is anchored to each database's own documentation — not a feature matrix.

The Answer

Choose by infrastructure fit, not feature boxes. Already run Postgres → pgvector (zero new infrastructure). Zero-ops managed → Pinecone. Hybrid BM25+vector → Weaviate. Self-hosted billion-scale → Milvus. And for a small corpus with infrequent queries, you don't need a vector database at all — exact brute-force search is fine until measured latency misses your budget.

Start with pgvector, measure, and migrate only when you hit a concrete limitation — not a theoretical one.

graph TD
    Start[Need vector search] --> PG{Already running<br/>PostgreSQL?}
    PG -->|Yes| Hybrid1{Need hybrid search<br/>BM25 plus vector?}
    PG -->|No| Hybrid2{Need hybrid search<br/>BM25 plus vector?}
    Hybrid1 -->|Yes| Wv1[Weaviate]
    Hybrid1 -->|No| Scale1{Index outgrows<br/>single-node RAM?}
    Scale1 -->|No| PGv[pgvector<br/>zero new infrastructure]
    Scale1 -->|Yes| Ops1{Want zero-ops<br/>managed service?}
    Ops1 -->|Yes| Pine[Pinecone]
    Ops1 -->|No| Throughput1{Billion-scale?}
    Throughput1 -->|Yes| Mil[Milvus<br/>distributed]
    Throughput1 -->|No| PGc[pgvector<br/>shard with Citus if needed]
    Hybrid2 -->|Yes| Wv1
    Hybrid2 -->|No| Ops1
    style PGv fill:#dfd
    style PGc fill:#dfd
    style Wv1 fill:#dfd
    style Pine fill:#ffd
    style Mil fill:#ffd

The comparison matrix

Choose based on your infrastructure and constraints, not by checking feature boxes:

DimensionpgvectorPineconeWeaviateMilvus
TypePostgreSQL extensionFully managed SaaSOpen-source DBOpen-source distributed DB
Best forAlready use PostgresZero-ops, fast prototypingHybrid BM25+vector searchSelf-hosted billion-scale
DeploymentYour existing PostgresAPI endpoint onlySelf-hosted or managedSelf-hosted or Zilliz Cloud
ScalingRead replicas / CitusNo up-front index sizingNative shardingDistributed (query, data, streaming)
Hybrid searchManual SQL joinSparse-dense in one queryNative BM25 + alpha tuningSparse + dense vectors
ACID / SQL joinsYes (full Postgres)NoNoNo
Max vectorsNo count limit; RAM-bound for latencyN/A (managed)Varies — benchmark on your dataBillion-scale+ (distributed)
Cold start latencyNone~2 s; up to ~20 s at 1BNoneNone
Operational overheadModerate (index rebalancing)NoneMedium (sharding ops)High (etcd, MinIO, multi-node)
Latest stable (Sept 2026)pgvector 0.8.6Serverless (managed)Weaviate 1.39Milvus 3.0.x

Each vendor's pricing model is broken down in its section below; re-price against the vendors' calculators before committing.

Two databases not broken out above are worth knowing in 2026: Qdrant (Rust, latest 1.19) competes directly here — native payload filtering, a free managed tier, resource-based Cloud pricing — evaluate it if your fit is "Postgres-adjacent but standalone." LanceDB is a different shape: an embedded, serverless engine on the columnar Lance format for multimodal and lakehouse-scale data rather than a long-running server.

pgvector: Start here if you have PostgreSQL

pgvector adds a vector column type and HNSW indexes to Postgres. If you already run Postgres, you get vector search with ACID transactions and SQL joins — no new infrastructure.⁠[3]

CREATE EXTENSION IF NOT EXISTS vector;
 
CREATE TABLE documents (
    id         BIGINT PRIMARY KEY,
    content    TEXT NOT NULL,
    metadata   JSONB DEFAULT '{}',
    embedding  vector(1536),  -- OpenAI text-embedding-3-small dimensions
    created_at TIMESTAMPTZ DEFAULT now()
);
 
-- HNSW index: m=16 is the default; ef_construction=200 (default 64) buys recall with build time
CREATE INDEX idx_embedding ON documents
    USING hnsw (embedding vector_cosine_ops)
    WITH (m = 16, ef_construction = 200);
 
-- Vector search with metadata filters and SQL joins
SELECT d.id, d.content, 1 - (d.embedding <=> $1::vector) AS similarity
FROM documents d
JOIN categories c ON d.category_id = c.id
WHERE c.slug = 'architecture'
  AND d.created_at > now() - INTERVAL '90 days'
ORDER BY d.embedding <=> $1::vector
LIMIT 10;

The <=> operator is cosine distance. You get full SQL expressiveness — WHERE clauses, JOINs, window functions — all with indexes. HNSW tuning matters: m=16 balances memory and recall, and raising m to 32 increases memory substantially for diminishing recall gains — the classic HNSW connectivity-vs-memory trade. Validate the exact recall delta on your own vectors, since it's dataset-dependent. Higher-dimensional vectors and larger collections push memory hard; plan for large-RAM instances at multi-million-vector scale. See ⁠[3] for the full HNSW parameter reference.

The release worth knowing about is 0.8.0 (October 2024), carried through the current 0.8.6: it added iterative index scans (SET hnsw.iterative_scan = 'relaxed_order') plus improved cost estimation for filtered queries. Before 0.8.0, a selective WHERE clause on top of an HNSW scan could "overfilter" — the index returned its ef_search candidates, the filter discarded most of them, and you got back fewer than LIMIT rows.

Iterative scans keep walking the index until the filtered result set is full, which makes the metadata-filtered queries above behave the way you'd expect. If you're on a pre-0.8.0 build (some managed Postgres lagged for months), upgrading is the single highest-leverage change for filtered RAG queries. Once on 0.8.x, run at least 0.8.4: 0.8.3 and 0.8.4 fixed HNSW-vacuuming bugs, including possible index corruption⁠[4].

The sharp edge — the one from this article's opening: a plain, in-place REINDEX INDEX locks out writes on the table and takes an AccessExclusiveLock on the index, which together block INSERTs, UPDATEs, and virtually every read of the table until it finishes (per PostgreSQL's lock-conflict table⁠[1] and REINDEX reference⁠[2]). The CONCURRENTLY variant instead takes a ShareUpdateExclusiveLock — which does not conflict with the RowExclusiveLock that writes hold, so writes keep flowing — but it runs two passes and is much slower.

Index builds on large HNSW tables are not quick either way, so a naive in-place rebuild that looks harmless can stall ingestion for a long window. Mitigation: use REINDEX INDEX CONCURRENTLY, or build a replacement with CREATE INDEX CONCURRENTLY and swap it in (DROP INDEX CONCURRENTLY the old one, then rename) instead of rebuilding in place, and time any rebuild against a measured build rate on your own data.

Pinecone: Zero-ops at massive scale

Pinecone is fully managed: no infrastructure and no up-front index sizing⁠[5]; its serverless write-up describes a customer (Gong) searching billions of vectors ⁠[5].⁠[6]

from pinecone import Pinecone
 
pc = Pinecone(api_key="your-api-key")
index = pc.Index("documents")
 
# Upsert vectors with metadata
index.upsert(vectors=[
    {"id": "doc-001", "values": embedding, "metadata": {"category": "arch"}},
    {"id": "doc-002", "values": embedding_2, "metadata": {"category": "devops"}},
])
 
# Query with metadata filter
results = index.query(
    vector=query_embedding,
    top_k=10,
    filter={"category": {"$eq": "arch"}},
    include_metadata=True
)

The tradeoff: the first query against a cold namespace pays a cold start — Pinecone's own 2024 architecture write-up puts it at a couple of seconds for most datasets and up to ~20 seconds at billion scale ⁠[5]. The serverless model bills three dimensions independently — storage ($0.33/GB-month), write units ($4–$4.50/M), and read units ($16–$18/M) at Standard-plan rates — on top of plan pricing ($0 Starter, $20/mo flat Builder, $50/mo minimum Standard, $500/mo minimum Enterprise) ⁠[7].

Each query costs 1 read unit per GB of the namespace it targets, minimum 0.25 RU⁠[8], so read spend grows with query volume × namespace size; storage adds a linear per-GB term. Model your own spend from those three rates before committing. Use namespaces for bulk deletes (instant, O(1)); avoid individual filter-based deletes.

Wins when: you want zero-ops and can accept cold-start latency; your team has no appetite for database operations. Struggles when: cost matters at massive scale; vendor lock-in concerns; you need SQL-grade filtering.

Weaviate: Native hybrid search (BM25 + vector)

Weaviate's killer feature: native BM25 + vector hybrid search in a single query with tunable alpha (0.0 = pure keyword, 1.0 = pure vector) — see ⁠[9] for the full alpha semantics. Workloads that mix factual lookup with semantic retrieval usually favor a vector-leaning blend over either extreme, but the right alpha is query-dependent — sweep it against a labelled set of your own queries rather than adopting a default blind.⁠[9]

import weaviate
client = weaviate.connect_to_local()
documents = client.collections.get("Document")
 
response = documents.query.hybrid(
    query="distributed consensus",
    alpha=0.75,  # 75% vector, 25% keyword
    limit=10,
)

Pure vector search misses exact keyword matches (searching "OAuth2 PKCE" drifts toward "authentication"). Hybrid search fixes this.

Two recent changes matter for sizing. Weaviate now ships rotational quantization (RQ; 8-bit RQ compresses vectors up to 4×), which is what keeps the per-dimension cost in check at scale — but whether new collections get it by default depends on your version and the DEFAULT_QUANTIZATION setting, so check rather than assume ⁠[10]. And Weaviate Cloud repriced on 27 October 2025: Serverless Cloud was renamed Shared Cloud, with plans starting at Flex ($45/mo), billed on vector dimensions stored — adjusted for replication, index type, and compression — plus storage and backups ⁠[11]. Budget against dimensions × objects × replication factor, not raw vector count, and re-check the calculator if you priced this before late 2025.

Wins when: users need both semantic understanding and keyword precision. Struggles when: operational overhead (shard rebalancing is not transparent).

Milvus: Distributed, billion-scale

Open-source, built for massive scale. Distributed architecture (query, data, and streaming nodes) lets each layer scale independently ⁠[12]. Milvus Distributed is designed for billion-scale or even larger deployments ⁠[13].

from pymilvus import MilvusClient, CollectionSchema, FieldSchema, DataType
 
client = MilvusClient(uri="http://localhost:19530")
 
fields = [
    FieldSchema("id", DataType.INT64, is_primary=True, auto_id=True),
    FieldSchema("embedding", DataType.FLOAT_VECTOR, dim=1536),
]
client.create_collection(collection_name="documents", schema=CollectionSchema(fields))
 
results = client.search(collection_name="documents", data=[query_embedding], limit=10)

The current major line is Milvus 3.0 (GA 29 July 2026), which adds a lake-native layer — external collections that index lakehouse files in place — and online schema add/backfill/drop. 2.6.x is still being patched, and a 3.0 deployment can roll back to 2.6 until you enable a format-changing feature such as Storage V3 ⁠[14]. At billion scale the cost lever is quantization, not node count, and 2.6 supplied it: RaBitQ 1-bit quantization compresses an index to 1/32 of its float size, with SQ8 refinement to hold recall ⁠[15]. Milvus itself is self-hosted (Apache-2.0); the managed path is Zilliz Cloud, which cut storage to $0.04/GB-month and made that the price on AWS, Azure, and GCP from January 2026 ⁠[16].

Sharp edge: etcd holds Milvus's metadata and service registration ⁠[12], so every write path depends on it. etcd's hardware guide assumes dedicated machines and warns that co-tenants cause resource contention and cluster instability ⁠[17]. Colocate etcd with query nodes and a traffic spike can starve it until it loses quorum — at which point Milvus cannot persist metadata and ingestion stalls. Fix: dedicate machines to etcd.

Wins when: billion-scale; your team can operate a distributed system. Struggles when: operational overhead exceeds value for small datasets.

Embedding model migrations: Dual-index strategy

Every team eventually swaps models — ada-002 → text-embedding-3-small, or to self-hosted nomic-embed-text to cut cost. Embeddings from different models are incompatible — you cannot mix them in one index. Re-embedding 5M documents naively (stop ingestion, re-embed, swap) leaves the system down or stale for the whole run.

The dual-index migration sequence — zero downtime, reversible until the very last step:

graph TD
    Start[Production on model v1<br/>idx_v1 hot] --> Step1[Step 1<br/>ALTER TABLE add embedding_v2 column<br/>CREATE INDEX CONCURRENTLY idx_v2]
    Step1 --> Step2[Step 2<br/>Backfill embedding_v2<br/>via batched workers + replication-lag throttling]
    Step2 --> Step3[Step 3<br/>Dual-write: writes update both v1 and v2<br/>reads still hit v1]
    Step3 --> Step4[Step 4<br/>Feature-flag cutover<br/>5 percent of reads → v2<br/>compare quality metrics]
    Step4 --> Quality{Quality<br/>regression?}
    Quality -->|Yes| Rollback[Flip flag back to v1<br/>investigate offline]
    Quality -->|No| Step5[Step 5<br/>Ramp 5 → 25 → 50 → 100 percent<br/>over 1-2 weeks]
    Step5 --> Step6[Step 6<br/>Stop dual-write<br/>drop idx_v1 + embedding_v1 column]
    Rollback --> Step3
    style Step6 fill:#dfd
    style Rollback fill:#ffd

The diagram captures the safety property: every step before Step 6 is reversible. The cutover is gradual, the rollback is one feature-flag flip, and you never have a window where queries hit a half-built index.

In pgvector, the shadow index is a second column built concurrently:

ALTER TABLE documents ADD COLUMN embedding_v3 vector(1536);
CREATE INDEX CONCURRENTLY idx_documents_embedding_v3small
    ON documents USING hnsw (embedding_v3 vector_cosine_ops)
    WITH (m = 16, ef_construction = 200);

In practice the sequence is a four-week rollout, not a weekend job. The realistic timeline for a 10M-document corpus on text-embedding-3-small:

  • Week 1 — shadow build. Add the new vector column, build the index CONCURRENTLY, enable dual-write so every fresh document hits both columns. No read traffic touches the new index yet. Smoke-test the embedding client at 100 docs/sec sustained.
  • Week 2 — backfill. Run batched workers (1k docs per batch, 4–8 workers) to embed historical rows — restartable via WHERE embedding_v3 IS NULL. Compute wall-clock time from your embedding rate limits. Throttle on replication lag — pause the workers whenever pg_stat_replication.replay_lag exceeds 30 seconds.
  • Week 3 — shadow eval. Mirror 5% of read traffic to both indexes, log the top-10 ID sets, and compute Jaccard overlap and recall@10 against a held-out judgment set. Block the cutover if recall regresses by more than 2 absolute points.
  • Week 4 — ramp and drop. Flip the feature flag from 5% → 25% → 50% → 100% over four days, watching p99 latency and answer-quality eval scores. Wait one full week with 100% traffic on the new index before dropping the old column, then reclaim disk with pg_repack ⁠[18] — not VACUUM FULL, which holds an AccessExclusiveLock for the whole table rewrite ⁠[1].

Budget the dollars up front: re-embedding 10M docs at 500 tokens each is 5B tokens; at $0.02 per 1M tokens, that is $100 of compute plus whatever your egress and worker hours cost. The expensive part is the eval judgment set, not the embeddings.

Dimension Changes

Changing dimensions (1536 → 3072) requires a new column — PostgreSQL vector is fixed-dimension per column. On Pinecone, a new index. Plan dimension changes as a separate migration from model changes.

Production checklist

Before going live with any vector DB:

  • Measure baseline performance — p50, p99, recall@10 on your actual data and query patterns. Published benchmarks run on public datasets, not yours.
  • Plan index rebuilds — an in-place HNSW REINDEX blocks writes. For pgvector: REINDEX INDEX CONCURRENTLY, or CREATE INDEX CONCURRENTLY → DROP INDEX CONCURRENTLY → rename. For Pinecone/Weaviate: automatic but plan shard rebalancing.
  • Monitor index health — Track index bloat (pgvector), shard imbalance (Weaviate/Milvus), cold-start latency (Pinecone).
  • Budget for embedding migrations — OpenAI's text-embedding-3-small costs $0.02 / 1M tokens. Re-embedding 10M docs at 500 tokens each = $100. Plan model migrations with dual-index strategy.
  • Test with your actual scale — Don't migrate prematurely. A well-tuned pgvector instance comfortably serves multi-million-vector workloads at low-latency p99 on adequate RAM, and many sub-billion-vector deployments never need to leave it. Confirm the ceiling on your own data and query mix before assuming you've outgrown it — the migration is rarely as urgent as it feels.

pgvector index health: the queries you run before users notice

Two queries that surface most pgvector production issues — bloat-driven recall regressions and the lock-contention failure mode behind the stalled-ingestion scenario this article opened with:

-- 1. Index size, bloat estimate, and last-VACUUM timestamp. Run weekly; if
--    pg_relation_size grows by >50% without matching row growth, schedule
--    REINDEX CONCURRENTLY in the next maintenance window.
SELECT
    i.relname                                                    AS index_name,
    pg_size_pretty(pg_relation_size(i.oid))                      AS index_size,
    s.n_live_tup                                                 AS table_rows,
    s.last_vacuum,
    s.last_autovacuum
FROM pg_class i
JOIN pg_index x          ON x.indexrelid = i.oid
JOIN pg_class t          ON t.oid = x.indrelid
JOIN pg_stat_user_tables s ON s.relid = t.oid
WHERE i.relkind = 'i'
  AND pg_get_indexdef(i.oid) ILIKE '%USING hnsw%'
ORDER BY pg_relation_size(i.oid) DESC;
 
-- 2. Live maintenance/blocking-lock detector. Paste during an incident:
--    mode='AccessExclusiveLock' on a vector index means a plain in-place
--    REINDEX is blocking your write path right now; mode='ShareUpdateExclusiveLock'
--    means a CONCURRENTLY reindex/VACUUM is running (slow, but writes still flow).
--    Rows with granted=false are sessions already stuck waiting on a lock.
SELECT
    a.pid,
    a.usename,
    a.query_start,
    now() - a.query_start                                        AS waiting_for,
    l.mode,
    l.relation::regclass                                         AS blocked_object,
    a.query
FROM pg_stat_activity a
JOIN pg_locks l ON l.pid = a.pid
WHERE NOT l.granted
   OR l.mode IN ('ShareUpdateExclusiveLock', 'AccessExclusiveLock')
ORDER BY a.query_start;

The first query is your "is bloat eating recall?" check; the second is your "why are inserts hanging right now?" check. Together they cover the two failure modes above. ⁠[3]

Abstract your vector store interface

Build migration flexibility from day one. Use an abstraction so swapping backends is a new class, not a pipeline rewrite:

from abc import ABC, abstractmethod
from dataclasses import dataclass
 
@dataclass
class SearchResult:
    id: str
    score: float
    metadata: dict
 
class VectorStore(ABC):
    @abstractmethod
    def upsert(self, id: str, vector: list[float], metadata: dict) -> None: ...
 
    @abstractmethod
    def search(self, vector: list[float], top_k: int = 10, filter: dict = None) -> list[SearchResult]: ...
 
    @abstractmethod
    def delete(self, ids: list[str]) -> None: ...

Implement pgvector, then Pinecone, then Weaviate. A config flag selects which backend to use. When you outgrow pgvector, you swap implementations without rewriting your application logic.

Hybrid retrieval: BM25 + vector + cross-encoder rerank

Pure vector search loses on rare tokens — product SKUs, error codes, library names, version numbers. BM25 still owns those. The production-grade pattern is a three-stage funnel: a wide BM25 + dense recall pass over the corpus, a fusion step that merges the two ranked lists, then a cross-encoder reranker that re-scores the top 50–100 candidates with a much stronger model. The reranker is too expensive to apply to the full index, but cheap enough to run on a short shortlist — measure its NDCG@10 lift against either signal alone on your own judgment set.

In Postgres you can build the hybrid stage without leaving the database. Add a tsvector column alongside the embedding, run full-text ranking with ts_rank_cd (not BM25)⁠[19], and combine the two scores with Reciprocal Rank Fusion (RRF) — which is robust to score-scale mismatch in a way that simple weighted sums are not.

-- One CTE per signal, fused by rank position (k = 60 is the standard RRF
-- constant). RRF avoids having to normalise ts_rank_cd against cosine distance.
WITH fts AS (
    SELECT id,
           ts_rank_cd(content_tsv, plainto_tsquery('english', $1)) AS score,
           row_number() OVER (ORDER BY ts_rank_cd(content_tsv,
               plainto_tsquery('english', $1)) DESC) AS rnk
    FROM documents
    WHERE content_tsv @@ plainto_tsquery('english', $1)
    LIMIT 100
),
dense AS (
    SELECT id,
           1 - (embedding <=> $2::vector) AS score,
           row_number() OVER (ORDER BY embedding <=> $2::vector) AS rnk
    FROM documents
    ORDER BY embedding <=> $2::vector
    LIMIT 100
),
fused AS (
    SELECT COALESCE(b.id, v.id) AS id,
           COALESCE(1.0 / (60 + b.rnk), 0) + COALESCE(1.0 / (60 + v.rnk), 0) AS rrf
    FROM fts b
    FULL OUTER JOIN dense v ON v.id = b.id
)
SELECT d.id, d.content, f.rrf
FROM fused f
JOIN documents d ON d.id = f.id
ORDER BY f.rrf DESC
LIMIT 50;

Feed the 50 RRF survivors into a cross-encoder — bge-reranker-v2-m3 or a Cohere Rerank call — and return the top 10. Measure the reranker's p99 on your own hardware; on CPU, shrink the shortlist until it fits your latency budget.

Capacity planning: memory and disk per million 1536-dim vectors

Three formulas cover most pgvector capacity questions. For 1536-dimension float32 vectors with HNSW at m = 16:

  • Raw vector bytes per row = 4 * dim + 8 = 4 × 1536 + 8 = 6,152 B (~6 KB)⁠[3].
  • HNSW index bytes per row ≈ dim * 4 + m * 2 * 8 + 32 = 6144 + 256 + 32 = 6,432 B — an estimate pgvector doesn't document; measure yours with pg_relation_size.
  • Heap row overhead = ~28 B (tuple header + line pointer)⁠[20] + content TEXT toasted out-of-line + JSONB metadata; e.g. budget 200 B per row.

Interactive pgvector Size Estimator

Vectors + HNSW index
10k to 10M
Index Type
Vector Type
Estimated Size
12.58 GBVectors + index
Raw vectors: 6.15 GB
HNSW index (est.): 6.43 GB

Vector bytes per row follow the pgvector README. The HNSW figure (m = 16) is an estimate pgvector doesn't document — measure yours with pg_relation_size. Heap rows, text and WAL are not included.

Choose by infrastructure fit
Vector count alone doesn't pick the database
  • Already run Postgres → pgvector
  • Zero-ops managed → Pinecone
  • Hybrid BM25 + vector → Weaviate
  • Self-hosted billion-scale → Milvus
Measure the real index size:SELECT pg_size_pretty( pg_relation_size('idx_embedding'));

Worked example for 5M documents with 1.5 KB of average text content per row:

Raw vectors:        5_000_000 * 6152   = 30.8 GB
HNSW index (est.):  5_000_000 * 6432   = 32.2 GB
Heap (text+meta):   5_000_000 * 1700   =  8.5 GB
TOAST + WAL slack:                       ~10 GB
                                        --------
Total disk:                              ~81 GB
 
Working set in RAM (keep the index cached):
  index pages           ~32 GB (est.)
  Postgres shared_buffers (typically 25% RAM)
  OS page cache

Size RAM so the index stays cached, then measure p99 before and after: indexes needn't fit in memory, but likely perform better when they do ⁠[3]. Quantizing to halfvec (float16) cuts vector and index size in half at a small, dataset-dependent recall cost; quantizing to bit (1 bit per dimension) cuts it a further 16× and is best used as a first-stage filter before a full-precision rerank ⁠[3].

Frequently Asked Questions

Should I use pgvector or a dedicated vector database?

Use pgvector if you already run PostgreSQL and need ACID, SQL joins, and want to avoid a new infrastructure dependency. Choose Pinecone for zero-ops at billions of vectors; Weaviate for hybrid BM25+vector search; Milvus for self-hosted billion-scale.

What is HNSW and why does it matter?

HNSW (Hierarchical Navigable Small World) is the dominant ANN index algorithm. In pgvector it has a better speed-recall tradeoff than IVFFlat but slower builds and more memory ⁠[3] — plan rebuild windows carefully in production.

How do you migrate embedding models without downtime?

Use a dual-index strategy: add a new embedding column, dual-write new documents to both old and new indexes, backfill existing docs in batches, validate recall parity, then swap traffic with a feature flag. Never switch models without re-embedding.

How many vectors can pgvector handle?

No fixed vector-count limit (Postgres caps a non-partitioned table at 32 TB by default); indexes needn't fit in RAM but likely perform better when they do. A 1536-dim vector is 6,152 bytes (4 × dims + 8): ~6 GB per million excluding the HNSW index ⁠[3]. When the index outgrows memory, move to Milvus or Pinecone, or shard with Citus.

Keep Reading


Coming Next

Coming Next: Building Production RAG Pipelines in Go

Selecting the right vector database is only the first step in building a resilient retrieval pipeline. In our next deep dive, we walk through constructing a production-grade RAG pipeline in Go, detailing chunking strategies, embedding requests, hybrid search execution, and retrieval evaluation metrics. Read the RAG Pipelines Guide.

Sources

  1. 1.PostgreSQL Documentation: Explicit Locking (lock modes and conflicts) — PostgreSQL Global Development Group, 2026
  2. 2.PostgreSQL Documentation: REINDEX — PostgreSQL Global Development Group, 2026
  3. 3.pgvector — Open-source vector similarity search for Postgres — pgvector project, 2026
  4. 4.pgvector — CHANGELOG — pgvector project (GitHub), 2026
  5. 5.Reimagining the vector database to enable knowledgeable AI — Pinecone, 2024
  6. 6.Pinecone — Vector database documentation — Pinecone, 2026
  7. 7.Pinecone Pricing — Pinecone, 2026
  8. 8.Understanding Pinecone cost — Pinecone, 2026
  9. 9.Weaviate — Hybrid Search (BM25 + vector) — Weaviate, 2026
  10. 10.Rotational Quantization (RQ) — Weaviate Documentation — Weaviate, 2026
  11. 11.A Simpler, More Transparent Pricing Model for Weaviate Cloud — Weaviate, 2025
  12. 12.Milvus Architecture Overview (v3.0.x docs) — Milvus / LF AI & Data (GitHub milvus-docs), 2026
  13. 13.What is Milvus (v3.0.x docs) — Milvus / LF AI & Data (GitHub milvus-docs), 2026
  14. 14.Milvus v3.0.0 release notes — milvus-io (GitHub), 2026
  15. 15.IVF_RABITQ — Milvus Documentation (v3.0.x docs) — Milvus / LF AI & Data (GitHub milvus-docs), 2026
  16. 16.New in Zilliz Cloud | Tiered Storage, BC Plan, Pricing Changes & More — Zilliz, 2025
  17. 17.etcd v3.6 — Hardware recommendations — etcd.io (CNCF), 2025
  18. 18.pg_repack — Reorganize tables in PostgreSQL with minimal locks — GitHub, 2026
  19. 19.PostgreSQL Documentation: 12.3. Controlling Text Search — PostgreSQL Global Development Group, 2026
  20. 20.PostgreSQL Documentation: Database Page Layout — PostgreSQL Global Development Group, 2026
BackendBytes Engineering Team
BackendBytes

Engineering Team

An independent engineering publication covering distributed systems, databases, and production infrastructure. Every factual claim is cited to a primary source or removed.

Read Next