Skip to content

Scaling Redis for High-Throughput Systems

6
Part of Series: Distributed Systems Mastery

Lesson 6 of 6

One trending key pins a single CPU core while five sibling shards sit green. A flash-sale or trending-content workload pushes every concurrent user at the same key — one shard, one bottleneck. The trace looks the same every time.

How a Flash Sale Bottlenecks on a Single Key

The hot-key failure mode — a recurring pattern in flash-sale and trending-content workloads:

A 6-shard Redis Cluster, each shard rated for ~100k ops/sec on a single core (Redis is single-threaded for command processing⁠[1]), runs a flash sale. The tests pass. But all the concurrent users hash to the same key — product:flash-sale:current — and therefore the same hash slot, the same primary shard, the same CPU core. The other five shards sit idle.

Within 30 seconds the hot core saturates, cache latency spikes from sub-millisecond to tens of milliseconds, the database absorbs the miss storm, and checkout fails for a sizeable fraction of users while the rest of the cluster shows green dashboards.

The diagnostic signature follows directly from the single-threaded design: cluster-level metrics look fine while the one hot shard's core sits pinned near 100% — one thread fully saturating the one core it runs on.

The Short Version

Redis scales horizontally via Cluster sharding (16,384 hash slots across nodes), but bottlenecks when access patterns concentrate on one key or node. Prevent hot keys with in-process LRU caches or key replication. Use pipelined batches instead of sequential commands. Scale connection pools carefully — concurrent Pod deploys can saturate Redis's max connections. Tune eviction and monitor CPU, not just ops/sec.

  • Scale via Redis Cluster (multi-key ops need hash tags: {user:123}:profile)
  • Mitigate hot keys with 1-second local LRU caches (about one Redis read per pod per second) or replicated reads
  • Batch commands with pipelining (up to ~10x the throughput of sequential ops) ⁠[2]
  • Pool connections correctly — 50 conns per Pod × 20 Pods = 1,000 Redis connections
  • Eviction policies: set maxmemory-policy allkeys-lru for cache workloads, tune maxmemory

The Quick Start: Sentinel vs Cluster

Redis is single-threaded for command processing — one CPU core; Redis's benchmark docs show a sample unpipelined run at ~180k SET/sec⁠[1]. Choose the right topology before scaling:

Redis SentinelRedis Cluster
PurposeHigh availability (single dataset)Horizontal scaling across nodes
ShardingNone — all data on one primary16,384 hash slots distributed
Max throughput (illustrative)One core's worth (docs' sample: ~180k SET/sec unpipelined)Roughly one core's worth × N primaries
Multi-key opsMGET/MSET always workRequire hash tags: {user:123}:*
When to usee.g. < 25GB data, < 100k ops/sec (a sizing heuristic, not a Redis-documented ceiling)Larger datasets or higher throughput

Start with Sentinel. Add Cluster when you hit a throughput or RAM limit on a single primary.

Setting Up Redis Cluster

Redis Cluster⁠[3] uses 16,384 fixed hash slots. Every key maps to one slot: slot = CRC16(key) % 16384. This distributes load across nodes — unless access patterns concentrate on a few keys.

The cluster topology in one picture — gossip-based discovery + slot ownership + client-side routing:

graph TB
    Client[Application client<br/>jedis / lettuce / go-redis] -->|MOVED redirect<br/>updates slot map cache| Slots[Slot map<br/>0-5460 P1<br/>5461-10922 P2<br/>10923-16383 P3]
    Slots --> P1[Primary 1<br/>slots 0-5460]
    Slots --> P2[Primary 2<br/>slots 5461-10922]
    Slots --> P3[Primary 3<br/>slots 10923-16383]
    P1 -.->|async replication| R1[Replica 1]
    P2 -.->|async replication| R2[Replica 2]
    P3 -.->|async replication| R3[Replica 3]
    P1 <-->|gossip protocol<br/>cluster bus on<br/>port + 10000| P2
    P2 <-->|gossip| P3
    P3 <-->|gossip| P1
    P1 -.->|failure detected<br/>by majority| Failover[Sentinel-style<br/>auto-failover<br/>R1 promotes to primary]
    style P1 fill:#dfd
    style P2 fill:#dfd
    style P3 fill:#dfd
    style R1 fill:#ffd
    style R2 fill:#ffd
    style R3 fill:#ffd
    style Failover fill:#fdd

Three production rules visible: (1) the client caches the slot map and only refreshes on MOVED redirect; (2) gossip happens on port + 10000 (e.g. 16379 for default Redis 6379); (3) failover requires majority quorum — a 3-primary cluster tolerates 1 primary loss.

Cluster setup (3 primaries, 3 replicas):

redis-cli --cluster create \
  redis-1:6379 redis-2:6379 redis-3:6379 \
  redis-4:6379 redis-5:6379 redis-6:6379 \
  --cluster-replicas 1

Multi-key operations (MGET, MSET, transactions) only work if all keys share a hash tag. The hash tag {tag} determines the slot — everything inside {} is hashed:

// All keys with {user:123} land on the same slot
keys := []string{
	fmt.Sprintf("{user:123}:profile"),
	fmt.Sprintf("{user:123}:orders"),
	fmt.Sprintf("{user:123}:prefs"),
}
results, err := client.MGet(ctx, keys...).Result() // Works: all on same slot

Without hash tags, MGET across different slots returns a CROSSSLOT error. Plan your key structure upfront — once you add hash tags, changing them requires a migration.

Cluster client setup in Go (go-redis v9):

import "github.com/redis/go-redis/v9"
 
client := redis.NewClusterClient(&redis.ClusterOptions{
	Addrs: []string{"redis-1:6379", "redis-2:6379", "redis-3:6379"},
	RouteByLatency: true,        // Read from replicas, write to primary
	PoolSize: 50,                // Connections per node
	MaxRedirects: 8,             // Retry on resharding
	ReadTimeout: 500 * time.Millisecond,
	WriteTimeout: 500 * time.Millisecond,
})

Hot Key Mitigation: In-Process Cache + Replication

A hot key is a single key receiving disproportionate requests, causing one shard to bottleneck. Symptoms: one node at 100% CPU while others idle, latency spikes for that key.

The mitigation stack is layered — each tier catches the fraction of traffic that still makes it through the layer above:

graph LR
    Req["N requests"] --> L1{"In-process<br/>LRU hit?"}
    L1 -->|hit| Serve["serve locally<br/>~0 network"]
    L1 -->|miss| L2{"Hot key<br/>replicated?"}
    L2 -->|yes| Replica["any of R replicas<br/>N/R per shard"]
    L2 -->|no| Shard["single shard<br/>full N"]
    Replica --> Redis[("Redis")]
    Shard --> Redis

A 1-second local LRU TTL cuts each pod to roughly one Redis read per second for the hot key; replica fan-out then spreads whatever still reaches Redis across several shards.

Solution 1: Local in-process LRU cache (read-heavy keys)

For a product receiving 100k reads/sec across 20 pods, a 1-second local TTL cuts Redis reads for that key to roughly 20/sec — one per pod, plus a burst at each expiry because concurrent misses are not coalesced (a small cache stampede⁠[4]):

import (
	lru "github.com/hashicorp/golang-lru/v2"
	"github.com/redis/go-redis/v9"
)
 
func GetProductCached(ctx context.Context, client *redis.ClusterClient, cache *lru.Cache[string, CacheEntry], productID string) ([]byte, error) {
	// Check local cache first (microseconds)
	if entry, ok := cache.Get(productID); ok && time.Now().Before(entry.Expires) {
		return entry.Data, nil
	}
 
	// Cache miss — fetch from Redis
	data, err := client.Get(ctx, fmt.Sprintf("product:%s", productID)).Bytes()
	if err != nil {
		return nil, err
	}
 
	// Store locally for 1 second
	cache.Add(productID, CacheEntry{
		Data:    data,
		Expires: time.Now().Add(1 * time.Second),
	})
	return data, nil
}

For product catalogs or config, a 1-second local TTL is acceptable. For inventory counts, use replicas instead.

Solution 2: Key replication (writable keys)

Store N copies on different slots, read from a random replica:

func SetProductReplicated(ctx context.Context, client *redis.ClusterClient, baseKey string, data []byte, replicas int) error {
	for i := 0; i < replicas; i++ {
		// Each copy lands on a different slot
		key := fmt.Sprintf("product:%d:%s", i, baseKey)
		if err := client.Set(ctx, key, data, 24*time.Hour).Err(); err != nil {
			return err
		}
	}
	return nil
}
 
func GetProductReplicated(ctx context.Context, client *redis.ClusterClient, baseKey string, replicas int) ([]byte, error) {
	// Random replica spread
	idx := rand.IntN(replicas)
	key := fmt.Sprintf("product:%d:%s", idx, baseKey)
	return client.Get(ctx, key).Bytes()
}

If the 10 replica keys land on 10 distinct shards, each handles ~10k RPS instead of 100k. Slots are not shards: on a 6-shard cluster some shard must own at least two replica keys, so check placement with CLUSTER KEYSLOT (or the slot() function below).

Pipelining: High-Throughput Batch Reads

At 1ms RTT, 1,000 sequential GET commands take 1 second of pure network overhead. Pipelining⁠[2] batches all commands into one round trip:

// Without pipelining: 1000 RTTs (~1 second)
for _, id := range ids {
	data, err := client.Get(ctx, fmt.Sprintf("product:%s", id)).Bytes()
	if err != nil && !errors.Is(err, redis.Nil) { // redis.Nil = key missing
		return nil, fmt.Errorf("get product %s: %w", id, err)
	}
	results = append(results, data)
}
 
// With pipelining: ~1 RTT per cluster shard, shards in parallel
cmds, err := client.Pipelined(ctx, func(pipe redis.Pipeliner) error {
	for _, id := range ids {
		pipe.Get(ctx, fmt.Sprintf("product:%s", id))
	}
	return nil
})
if err != nil && !errors.Is(err, redis.Nil) {
	return nil, fmt.Errorf("pipelined get: %w", err)
}
 
// Exec reports only the first failed command, so check each one.
for i, cmd := range cmds {
	data, err := cmd.(*redis.StringCmd).Bytes()
	if err != nil && !errors.Is(err, redis.Nil) {
		return nil, fmt.Errorf("get product %s: %w", ids[i], err)
	}
	results = append(results, data)
}

go-redis clusters automatically group commands by slot and send to the right node. Redis's docs measure 5× over loopback and up to 10× throughput with long pipelines⁠[2].

Production Checklist

  • Set maxmemory-policy allkeys-lru in redis.conf for cache workloads; use volatile-lru if mixing cached and persistent data
  • Set maxmemory to at most 50% of physical RAM if writes are heavy — Redis's own admin guidance is that RDB/AOF rewrites can use up to 2x the memory already in use⁠[5], and 2 × 50% is the full-RAM ceiling before that spike OOMs the process; read-heavy workloads with infrequent rewrites can run tighter
  • Monitor cache hit rate via redis-cli INFO stats | grep keyspace_hits and watch the trend — a falling rate is the early signal that your hot set no longer fits, before evicted_keys climbs
  • Pool size: calculate as connections_per_pod × num_pods. Start with 50 per pod; adjust if you see connection errors
  • Stagger pod restarts with readiness probes that warm the connection pool — prevents connection storms
  • Replica read routing: use RouteByLatency: true in go-redis to offload reads to replicas
  • Eviction rate: monitor evicted_keys — if rising, your hot set is larger than maxmemory or your TTLs are too aggressive
  • Replication lag: replicas should be within 10ms of primary. Monitor slave_repl_offset vs master_repl_offset

go-redis pool sizing that survives a slow Redis

Default redis.NewClient ships with PoolSize = 10 * runtime.GOMAXPROCS(0) and a multi-second read timeout (5s in go-redis v9.22, 3s in earlier v9 releases⁠[6]; PoolTimeout = ReadTimeout + 1s) — fine until Redis blocks on a single slow command and every goroutine in your service queues behind it for seconds. The configuration below is built for services pushing 200k+ ops/sec:

import (
	"time"
 
	"github.com/redis/go-redis/v9"
)
 
func NewRedisClient(addr string) *redis.Client {
	return redis.NewClient(&redis.Options{
		Addr: addr,
 
		// Pool sized for concurrency, not for parallelism.
		// PoolSize > NumCPU is fine — Redis is single-threaded but our
		// goroutines block on the network round-trip, not CPU.
		PoolSize:     200,
		MinIdleConns: 20,
 
		// The four timeouts that fail-fast instead of cascading:
		// DialTimeout — TCP connect ceiling, before any command runs.
		// ReadTimeout — per-command read budget; tighter than HTTP request budget.
		// WriteTimeout — protects against TCP backpressure from saturated link.
		// PoolTimeout — how long a goroutine waits for a free connection;
		//                exceeding this returns an error instead of hanging.
		DialTimeout:  500 * time.Millisecond,
		ReadTimeout:  200 * time.Millisecond,
		WriteTimeout: 200 * time.Millisecond,
		PoolTimeout:  100 * time.Millisecond,
 
		// Idle connection hygiene — beats most NAT/firewall idle drops.
		ConnMaxIdleTime: 5 * time.Minute,
		ConnMaxLifetime: 30 * time.Minute,
	})
}

For batch reads on a ClusterClient, one go-redis pipeline is already slot-aware: it maps each command to the node that owns its slot and sends every node its share concurrently⁠[7], so 1,000 GETs cost about one round trip per node instead of 1,000. Don't split the batch into per-slot pipelines run in turn — that costs one round trip per distinct slot, close to one per key for random keys:

type batchedReader struct {
	rdb redis.UniversalClient
}
 
// GetMany reads keys in one pipeline. Missing keys are left out of the map;
// any other per-key failure fails the whole call.
func (b *batchedReader) GetMany(ctx context.Context, keys []string) (map[string]string, error) {
	pipe := b.rdb.Pipeline()
	cmds := make([]*redis.StringCmd, len(keys))
	for i, k := range keys {
		cmds[i] = pipe.Get(ctx, k)
	}
 
	// Exec returns the first failed command's error, and a missing key
	// (redis.Nil) can mask a later real failure — so check every command.
	if _, err := pipe.Exec(ctx); err != nil && !errors.Is(err, redis.Nil) {
		return nil, fmt.Errorf("pipeline exec: %w", err)
	}
	out := make(map[string]string, len(keys))
	for i, c := range cmds {
		v, err := c.Result()
		switch {
		case errors.Is(err, redis.Nil):
			continue
		case err != nil:
			return nil, fmt.Errorf("get %s: %w", keys[i], err)
		}
		out[keys[i]] = v
	}
	return out, nil
}

Grouping by slot still matters for commands Redis itself confines to one slot — MGET, MSET, MULTI/EXEC, Lua scripts — which the server rejects with CROSSSLOT Keys in request don't hash to the same slot when keys span slots⁠[3]; hash tags fix that. The slot computation, handy for checking where hot-key replicas land, is pure CRC16 against the key (or the contents of the first {...} segment if present, the hash-tag escape hatch for forcing two keys to the same slot):

import "strings"
 
// slot returns the Redis Cluster hash slot for a key: CRC16(key) mod 16384,
// using the same CRC16/XMODEM (CCITT) variant Redis uses — polynomial 0x1021,
// initial value 0x0000. If the key contains a {tag} substring, only the tag is
// hashed, which forces co-location for keys you must batch atomically.
func slot(key string) uint16 {
    if start := strings.IndexByte(key, '{'); start >= 0 {
        if end := strings.IndexByte(key[start+1:], '}'); end > 0 {
            key = key[start+1 : start+1+end] // hash only the {...} contents
        }
    }
    var crc uint16
    for i := 0; i < len(key); i++ {
        crc ^= uint16(key[i]) << 8
        for j := 0; j < 8; j++ {
            if crc&0x8000 != 0 {
                crc = (crc << 1) ^ 0x1021
            } else {
                crc <<= 1
            }
        }
    }
    return crc % 16384
}
 
// Group keys by Cluster slot. {tag} hash-tag forces co-location for keys
// you must batch atomically (the only path to safe MULTI/EXEC across keys
// in Cluster mode).
func groupKeysBySlot(keys []string) map[uint16][]string {
    out := make(map[uint16][]string)
    for _, k := range keys {
        out[slot(k)] = append(out[slot(k)], k)
    }
    return out
}

Streams vs Pub/Sub: pick durability deliberately

Classic Redis Pub/Sub is fire-and-forget. A subscriber that disconnects for two seconds loses every message published in that window, and there is no acknowledgement, no replay, no consumer group. That works for cache-invalidation fan-out where a missed message just means a slightly stale read on one node, but it falls apart the moment you reach for it as a work queue or event log.

Redis Streams (XADD, XREADGROUP, XACK) close that gap: messages persist to RDB and AOF, consumer groups distribute work across workers with at-least-once delivery, and a pending entries list (PEL) tracks every unacked message so a crashed consumer's work can be claimed by another via XCLAIM or auto-claimed via XAUTOCLAIM after an idle threshold. The trade-off is memory pressure — a stream grows until you cap it with MAXLEN or MINID.

The decision rule that holds up in production: use Pub/Sub only for ephemeral coordination signals where loss is acceptable (cache invalidation, presence notifications, leader-election heartbeats). Use Streams for any payload representing a durable event — payment intents, audit records, async job dispatch. The Go consumer below shows the canonical pattern with bounded retries and dead-letter routing for poison messages.

// Stream consumer with consumer-group semantics, bounded retries,
// and a dead-letter stream for poison messages. Run one goroutine per
// worker; the consumer name should be unique per process (hostname+pid).
func consume(ctx context.Context, rdb *redis.Client, group, consumer string) error {
    cursor := "0-0"
    for {
        // ">" only ever delivers new entries, so a failed entry is never retried
        // unless something re-claims it. XAUTOCLAIM hands this consumer entries
        // idle > 30s (including a crashed peer's) and bumps their delivery count.
        claimed, next, err := rdb.XAutoClaim(ctx, &redis.XAutoClaimArgs{
            Stream: "orders", Group: group, Consumer: consumer,
            MinIdle: 30 * time.Second, Start: cursor, Count: 16,
        }).Result()
        if err != nil {
            return fmt.Errorf("xautoclaim: %w", err)
        }
        cursor = next
        for _, msg := range claimed {
            if err := process(ctx, rdb, group, msg); err != nil {
                return err
            }
        }
 
        res, err := rdb.XReadGroup(ctx, &redis.XReadGroupArgs{
            Group:    group,
            Consumer: consumer,
            Streams:  []string{"orders", ">"},
            Count:    16,
            Block:    5 * time.Second,
        }).Result()
        if errors.Is(err, redis.Nil) {
            continue // idle, no new messages
        }
        if err != nil {
            return fmt.Errorf("xreadgroup: %w", err)
        }
        for _, msg := range res[0].Messages {
            if err := process(ctx, rdb, group, msg); err != nil {
                return err
            }
        }
    }
}
 
// process ACKs on success. On failure it leaves the entry pending for a later
// XAUTOCLAIM until its delivery count reaches maxDeliveries, then moves it to
// the DLQ, ACKing only after the DLQ write succeeds so nothing is lost.
func process(ctx context.Context, rdb *redis.Client, group string, msg redis.XMessage) error {
    const maxDeliveries = 5
    herr := handle(ctx, msg.Values)
    if herr == nil {
        if err := rdb.XAck(ctx, "orders", group, msg.ID).Err(); err != nil {
            return fmt.Errorf("xack %s: %w", msg.ID, err)
        }
        return nil
    }
    pending, err := rdb.XPendingExt(ctx, &redis.XPendingExtArgs{
        Stream: "orders", Group: group, Start: msg.ID, End: msg.ID, Count: 1,
    }).Result()
    if err != nil {
        return fmt.Errorf("xpending %s: %w", msg.ID, err)
    }
    if len(pending) == 0 || pending[0].RetryCount < maxDeliveries {
        slog.Warn("handle failed; left pending for retry", "id", msg.ID, "err", herr)
        return nil
    }
    if err := rdb.XAdd(ctx, &redis.XAddArgs{Stream: "orders.dlq", Values: msg.Values}).Err(); err != nil {
        return fmt.Errorf("dlq xadd %s: %w", msg.ID, err)
    }
    if err := rdb.XAck(ctx, "orders", group, msg.ID).Err(); err != nil {
        return fmt.Errorf("xack dead-lettered %s: %w", msg.ID, err)
    }
    slog.Error("moved poison message to orders.dlq", "id", msg.ID,
        "deliveries", pending[0].RetryCount, "err", herr)
    return nil
}

The XAUTOCLAIM pass is what makes the retry bound real: each claim increments the entry's delivery counter unless you pass JUSTID⁠[8], and the same pass recovers work a Kubernetes pod eviction stranded in the PEL. Set MinIdle above your processing SLA. Cap stream length with XADD orders MAXLEN ~ 1000000 * so producers approximately bound memory; the ~ lets Redis trim in chunks instead of one entry at a time, which keeps XADD latency stable under load.


Client-side caching with RESP3 broadcast invalidation

Redis 6 introduced server-assisted client-side caching: clients cache reads locally, and Redis sends invalidation messages whenever a tracked key changes — as RESP3 push messages, or through a Pub/Sub redirect on RESP2⁠[9]. For hot keys, you replace a network round-trip with an in-process map lookup and remove the read entirely from Redis CPU.

Two tracking modes exist: default mode (server tracks every key each client reads, expensive for the server) and broadcast mode (clients subscribe to key-prefix patterns, server sends one invalidation per write fanned out to all matching subscribers). Broadcast mode is the only mode worth running at scale because server-side memory does not grow with the number of cached keys per client.

Three pitfalls bite teams adopting this. First, your local cache is eventually consistent — there is a window between a write landing on the primary and the invalidation reaching subscribers, so any read-your-writes guarantee must come from the application path that issued the write. Second, you must handle reconnects: if the tracking connection drops, you must invalidate the entire local cache before reconnecting, because invalidations during the disconnect were lost. Third, broadcast mode delivers invalidations even for keys you never read, so prefix selection matters for bandwidth.

// Client-side cache with RESP3 broadcast tracking. The tracking connection
// is separate from the command connection; invalidations arrive as RESP3
// push messages on the tracking conn and we apply them to the local map.
type CSCache struct {
    mu      sync.Mutex
    local   map[string][]byte
    filling map[string]uint64 // in-flight GETs; an invalidation deletes the entry
    seq     uint64
    rdb     *redis.Client
}
 
func (c *CSCache) Start(ctx context.Context, prefixes []string) error {
    // Tracking is per connection, so pin a dedicated one. Enabling it through
    // the pooled client would attach it to whichever conn Do() happened to get.
    conn := c.rdb.Conn()
    // BCAST mode: server sends invalidations for any key matching a prefix,
    // regardless of whether this client ever read it.
    args := []any{"CLIENT", "TRACKING", "ON", "BCAST"}
    for _, p := range prefixes {
        args = append(args, "PREFIX", p)
    }
    if err := conn.Do(ctx, args...).Err(); err != nil {
        return errors.Join(fmt.Errorf("enable tracking: %w", err), conn.Close())
    }
    go c.listen(ctx, conn) // consume invalidate push messages on the tracking conn
    return nil
}
 
func (c *CSCache) Get(ctx context.Context, key string) ([]byte, error) {
    c.mu.Lock()
    if v, ok := c.local[key]; ok {
        c.mu.Unlock()
        return v, nil
    }
    c.seq++
    token := c.seq
    c.filling[key] = token
    c.mu.Unlock()
 
    v, err := c.rdb.Get(ctx, key).Bytes()
 
    c.mu.Lock()
    defer c.mu.Unlock()
    // Data and invalidations travel on different connections, so an
    // invalidation can overtake this GET's reply. Cache v only if none
    // arrived for key while the GET was in flight.
    if c.filling[key] == token {
        delete(c.filling, key)
        if err == nil {
            c.local[key] = v
        }
    }
    if err != nil {
        return nil, err
    }
    return v, nil
}
 
// onInvalidate receives nil on FLUSHALL/FLUSHDB: drop everything.
func (c *CSCache) onInvalidate(keys []string) {
    c.mu.Lock()
    defer c.mu.Unlock()
    if keys == nil {
        clear(c.local)
        clear(c.filling)
        return
    }
    for _, k := range keys {
        delete(c.local, k)
        delete(c.filling, k)
    }
}

Bound the local cache size with an LRU — without one, a hostile or buggy access pattern will OOM the application before invalidations arrive. Track hit ratio per prefix as a Prometheus metric; a falling hit ratio is the early signal that your prefix is too coarse and you are paying invalidation bandwidth for keys nobody caches.


ACLs and TLS for multi-tenant production

Default Redis ships with a single password and no command-level authorization, which is fine for a development laptop and disastrous for a shared production cluster. Redis 6 added ACLs: per-user credentials with command, key-pattern, and channel-pattern restrictions. The right baseline is one user per service with the narrowest possible permission set, the legacy default user disabled, and TLS terminated at the Redis port (not at a sidecar) so that credentials never traverse the network in plaintext. Below is a users.acl implementing that baseline. An ACL file has no line continuations — each user is one line — the -command|subcommand rules need Redis 7.0+, and # comment lines load only on Redis 8.8+, so strip them for older servers⁠[10].

# Disable the default account so no client can connect without explicit creds.
user default off
 
# Read-only analytics: GET/MGET/SCAN on the analytics: prefix only,
# no write commands, no admin commands, no pub/sub.
user analytics on >REDACTED-STRONG-SECRET ~analytics:* +@read +@connection -@dangerous
 
# Application service: full data-plane access on its own keyspace,
# Streams on its own queue, no FLUSHDB/FLUSHALL/CONFIG/DEBUG/SCRIPT LOAD.
user orders-svc on >REDACTED-STRONG-SECRET ~orders ~orders:* ~orders.dlq &orders.events +@all -@dangerous -flushall -flushdb -config -debug -keys -script|load
 
# Operator account for runbooks; gated behind break-glass workflow.
# Allowed CONFIG GET but not SET, allowed CLUSTER inspection but not failover.
user oncall on >REDACTED-STRONG-SECRET ~* &* +@all -@dangerous +config|get +cluster|info -cluster|failover

Three operational rules make ACLs durable. First, version the users.acl file in a config repo (storing #<sha256> password hashes, not >plaintext) and ship it via configuration management, not ACL SETUSER over the wire — drift between nodes is the failure mode that lets an attacker keep access after a credential rotation. Second, rotate secrets through a transition user (orders-svc-v2) before deleting the old one, so the cutover does not require a deploy of every consumer at the same instant.

Third, log ACL LOG to your SIEM — every denied command is a misconfiguration or an attempted privilege escalation, and silence on that channel is what you want to hear. Combine ACLs with tls-port 6379, tls-auth-clients yes, and per-service client certificates so that even a leaked password on its own does not authenticate against the cluster.


Conclusion

Apply the fixes above to the flash-sale scenario that opened this piece — 10-way hot-key replication, a 1-second local LRU, pipelined batch reads, and connection pools sized for the spike — and the failure mode changes shape. The same flood of users hitting product:flash-sale:current no longer pins one shard while five sit idle: the local LRU absorbs most repeat reads, and replicated reads spread the rest across 10 slots. On 6 shards that is not a tenth per shard: fitting 10 replicas into 6 shards guarantees some shard owns at least two of them by pigeonhole — 2 of 10 replicas is 20% of the hot-key load at best, 3 of 10 is 30% with unlucky placement — so check with CLUSTER KEYSLOT and adjust suffixes.

Scaling Redis isn't about raw ops/sec: one instance is bounded by one core, and uneven key distribution can make a 6-shard cluster behave like one.

Start with Sentinel and a single instance. Move to Cluster only when you hit throughput limits. Prevent hot keys with in-process LRU caches. Batch reads with pipelining. Pool connections with discipline. Monitor hit rates, not just ops/sec. ⁠[2]


Frequently Asked Questions

Why is Redis single-threaded and how does that affect scaling?

Redis processes commands on a single thread, avoiding lock contention. This means a single instance's throughput is bounded by one core (Redis's benchmark docs show a sample unpipelined run at ~180k SET/sec⁠[1]), and you must use Redis Cluster for horizontal scaling across multiple CPU cores.

What is a Redis hot key and how do you fix it?

A hot key is a single key receiving disproportionate traffic, causing one shard to bottleneck while others sit idle. Fix it by replicating the key under N names that hash to different slots (e.g., product:<n>:<id> with n picked at random per read, and no {hash tag}, which would pin every copy to one slot), using read replicas, or restructuring data to distribute load.

When should you use Redis Sentinel vs Redis Cluster?

Use Redis Sentinel for high availability while one primary covers your dataset and throughput (e.g., under ~25GB and ~100k ops/sec — a rule of thumb, not a Redis limit). Use Redis Cluster when you need horizontal scaling beyond one machine's RAM or throughput, as it shards data across multiple primaries.

How do you use Redis pipelining to improve throughput?

Pipelining batches multiple Redis commands into a single network round trip instead of waiting for each response individually. This reduces network overhead dramatically — Redis's own docs measure a 5× speedup even over loopback, and throughput up to 10× the unpipelined baseline⁠[2].

Keep Reading

Sources

  1. 1.Redis benchmark (redis-benchmark utility) — redis.io, 2026
  2. 2.Redis Pipelining (Documentation) — redis.io, 2026
  3. 3.Redis Cluster Specification — redis.io, 2026
  4. 4.Optimal Probabilistic Cache Stampede Prevention (Vattani et al., 2015) — VLDB Endowment, Vol 8 No 8, 2015
  5. 5.Redis administration — redis.io, 2026
  6. 6.go-redis v9.22.0 Release Notes — GitHub (redis/go-redis), 2026
  7. 7.go-redis v9.22.0 source: osscluster.go (ClusterClient) — GitHub (redis/go-redis), 2026
  8. 8.XAUTOCLAIM (Redis command reference) — redis.io, 2026
  9. 9.Client-side caching reference (Redis documentation) — redis.io, 2026
  10. 10.ACL (Redis Access Control List documentation) — redis.io, 2026
BackendBytes Engineering Team
BackendBytes

Engineering Team

An independent engineering publication covering distributed systems, databases, and production infrastructure. Every factual claim is cited to a primary source or removed.

Read Next