Scaling Redis for High-Throughput Systems
Lesson 6 of 6
One trending key pins a single CPU core while five sibling shards sit green. A flash-sale or trending-content workload pushes every concurrent user at the same key — one shard, one bottleneck. The trace looks the same every time.
How a Flash Sale Bottlenecks on a Single Key
The hot-key failure mode — a recurring pattern in flash-sale and trending-content workloads:
A 6-shard Redis Cluster, each shard rated for ~100k ops/sec on a single core (Redis is single-threaded for command processing[1]), runs a flash sale. The tests pass. But all the concurrent users hash to the same key — product:flash-sale:current — and therefore the same hash slot, the same primary shard, the same CPU core. The other five shards sit idle.
Within 30 seconds the hot core saturates, cache latency spikes from sub-millisecond to tens of milliseconds, the database absorbs the miss storm, and checkout fails for a sizeable fraction of users while the rest of the cluster shows green dashboards.
The diagnostic signature follows directly from the single-threaded design: cluster-level metrics look fine while the one hot shard's core sits pinned near 100% — one thread fully saturating the one core it runs on.
Redis scales horizontally via Cluster sharding (16,384 hash slots across nodes), but bottlenecks when access patterns concentrate on one key or node. Prevent hot keys with in-process LRU caches or key replication. Use pipelined batches instead of sequential commands. Scale connection pools carefully — concurrent Pod deploys can saturate Redis's max connections. Tune eviction and monitor CPU, not just ops/sec.
- Scale via Redis Cluster (multi-key ops need hash tags:
{user:123}:profile) - Mitigate hot keys with 1-second local LRU caches (about one Redis read per pod per second) or replicated reads
- Batch commands with pipelining (up to ~10x the throughput of sequential ops) [2]
- Pool connections correctly — 50 conns per Pod × 20 Pods = 1,000 Redis connections
- Eviction policies: set
maxmemory-policy allkeys-lrufor cache workloads, tunemaxmemory
The Quick Start: Sentinel vs Cluster
Redis is single-threaded for command processing — one CPU core; Redis's benchmark docs show a sample unpipelined run at ~180k SET/sec[1]. Choose the right topology before scaling:
| Redis Sentinel | Redis Cluster | |
|---|---|---|
| Purpose | High availability (single dataset) | Horizontal scaling across nodes |
| Sharding | None — all data on one primary | 16,384 hash slots distributed |
| Max throughput (illustrative) | One core's worth (docs' sample: ~180k SET/sec unpipelined) | Roughly one core's worth × N primaries |
| Multi-key ops | MGET/MSET always work | Require hash tags: {user:123}:* |
| When to use | e.g. < 25GB data, < 100k ops/sec (a sizing heuristic, not a Redis-documented ceiling) | Larger datasets or higher throughput |
Start with Sentinel. Add Cluster when you hit a throughput or RAM limit on a single primary.
Setting Up Redis Cluster
Redis Cluster[3] uses 16,384 fixed hash slots. Every key maps to one slot: slot = CRC16(key) % 16384. This distributes load across nodes — unless access patterns concentrate on a few keys.
The cluster topology in one picture — gossip-based discovery + slot ownership + client-side routing:
graph TB
Client[Application client<br/>jedis / lettuce / go-redis] -->|MOVED redirect<br/>updates slot map cache| Slots[Slot map<br/>0-5460 P1<br/>5461-10922 P2<br/>10923-16383 P3]
Slots --> P1[Primary 1<br/>slots 0-5460]
Slots --> P2[Primary 2<br/>slots 5461-10922]
Slots --> P3[Primary 3<br/>slots 10923-16383]
P1 -.->|async replication| R1[Replica 1]
P2 -.->|async replication| R2[Replica 2]
P3 -.->|async replication| R3[Replica 3]
P1 <-->|gossip protocol<br/>cluster bus on<br/>port + 10000| P2
P2 <-->|gossip| P3
P3 <-->|gossip| P1
P1 -.->|failure detected<br/>by majority| Failover[Sentinel-style<br/>auto-failover<br/>R1 promotes to primary]
style P1 fill:#dfd
style P2 fill:#dfd
style P3 fill:#dfd
style R1 fill:#ffd
style R2 fill:#ffd
style R3 fill:#ffd
style Failover fill:#fdd
Three production rules visible: (1) the client caches the slot map and only refreshes on MOVED redirect; (2) gossip happens on port + 10000 (e.g. 16379 for default Redis 6379); (3) failover requires majority quorum — a 3-primary cluster tolerates 1 primary loss.
Cluster setup (3 primaries, 3 replicas):
redis-cli --cluster create \
redis-1:6379 redis-2:6379 redis-3:6379 \
redis-4:6379 redis-5:6379 redis-6:6379 \
--cluster-replicas 1Multi-key operations (MGET, MSET, transactions) only work if all keys share a hash tag. The hash tag {tag} determines the slot — everything inside {} is hashed:
// All keys with {user:123} land on the same slot
keys := []string{
fmt.Sprintf("{user:123}:profile"),
fmt.Sprintf("{user:123}:orders"),
fmt.Sprintf("{user:123}:prefs"),
}
results, err := client.MGet(ctx, keys...).Result() // Works: all on same slotWithout hash tags, MGET across different slots returns a CROSSSLOT error. Plan your key structure upfront — once you add hash tags, changing them requires a migration.
Cluster client setup in Go (go-redis v9):
import "github.com/redis/go-redis/v9"
client := redis.NewClusterClient(&redis.ClusterOptions{
Addrs: []string{"redis-1:6379", "redis-2:6379", "redis-3:6379"},
RouteByLatency: true, // Read from replicas, write to primary
PoolSize: 50, // Connections per node
MaxRedirects: 8, // Retry on resharding
ReadTimeout: 500 * time.Millisecond,
WriteTimeout: 500 * time.Millisecond,
})Hot Key Mitigation: In-Process Cache + Replication
A hot key is a single key receiving disproportionate requests, causing one shard to bottleneck. Symptoms: one node at 100% CPU while others idle, latency spikes for that key.
The mitigation stack is layered — each tier catches the fraction of traffic that still makes it through the layer above:
graph LR
Req["N requests"] --> L1{"In-process<br/>LRU hit?"}
L1 -->|hit| Serve["serve locally<br/>~0 network"]
L1 -->|miss| L2{"Hot key<br/>replicated?"}
L2 -->|yes| Replica["any of R replicas<br/>N/R per shard"]
L2 -->|no| Shard["single shard<br/>full N"]
Replica --> Redis[("Redis")]
Shard --> Redis
A 1-second local LRU TTL cuts each pod to roughly one Redis read per second for the hot key; replica fan-out then spreads whatever still reaches Redis across several shards.
Solution 1: Local in-process LRU cache (read-heavy keys)
For a product receiving 100k reads/sec across 20 pods, a 1-second local TTL cuts Redis reads for that key to roughly 20/sec — one per pod, plus a burst at each expiry because concurrent misses are not coalesced (a small cache stampede[4]):
import (
lru "github.com/hashicorp/golang-lru/v2"
"github.com/redis/go-redis/v9"
)
func GetProductCached(ctx context.Context, client *redis.ClusterClient, cache *lru.Cache[string, CacheEntry], productID string) ([]byte, error) {
// Check local cache first (microseconds)
if entry, ok := cache.Get(productID); ok && time.Now().Before(entry.Expires) {
return entry.Data, nil
}
// Cache miss — fetch from Redis
data, err := client.Get(ctx, fmt.Sprintf("product:%s", productID)).Bytes()
if err != nil {
return nil, err
}
// Store locally for 1 second
cache.Add(productID, CacheEntry{
Data: data,
Expires: time.Now().Add(1 * time.Second),
})
return data, nil
}For product catalogs or config, a 1-second local TTL is acceptable. For inventory counts, use replicas instead.
Solution 2: Key replication (writable keys)
Store N copies on different slots, read from a random replica:
func SetProductReplicated(ctx context.Context, client *redis.ClusterClient, baseKey string, data []byte, replicas int) error {
for i := 0; i < replicas; i++ {
// Each copy lands on a different slot
key := fmt.Sprintf("product:%d:%s", i, baseKey)
if err := client.Set(ctx, key, data, 24*time.Hour).Err(); err != nil {
return err
}
}
return nil
}
func GetProductReplicated(ctx context.Context, client *redis.ClusterClient, baseKey string, replicas int) ([]byte, error) {
// Random replica spread
idx := rand.IntN(replicas)
key := fmt.Sprintf("product:%d:%s", idx, baseKey)
return client.Get(ctx, key).Bytes()
}If the 10 replica keys land on 10 distinct shards, each handles ~10k RPS instead of 100k. Slots are not shards: on a 6-shard cluster some shard must own at least two replica keys, so check placement with CLUSTER KEYSLOT (or the slot() function below).
Pipelining: High-Throughput Batch Reads
At 1ms RTT, 1,000 sequential GET commands take 1 second of pure network overhead. Pipelining[2] batches all commands into one round trip:
// Without pipelining: 1000 RTTs (~1 second)
for _, id := range ids {
data, err := client.Get(ctx, fmt.Sprintf("product:%s", id)).Bytes()
if err != nil && !errors.Is(err, redis.Nil) { // redis.Nil = key missing
return nil, fmt.Errorf("get product %s: %w", id, err)
}
results = append(results, data)
}
// With pipelining: ~1 RTT per cluster shard, shards in parallel
cmds, err := client.Pipelined(ctx, func(pipe redis.Pipeliner) error {
for _, id := range ids {
pipe.Get(ctx, fmt.Sprintf("product:%s", id))
}
return nil
})
if err != nil && !errors.Is(err, redis.Nil) {
return nil, fmt.Errorf("pipelined get: %w", err)
}
// Exec reports only the first failed command, so check each one.
for i, cmd := range cmds {
data, err := cmd.(*redis.StringCmd).Bytes()
if err != nil && !errors.Is(err, redis.Nil) {
return nil, fmt.Errorf("get product %s: %w", ids[i], err)
}
results = append(results, data)
}go-redis clusters automatically group commands by slot and send to the right node. Redis's docs measure 5× over loopback and up to 10× throughput with long pipelines[2].
Production Checklist
- Set
maxmemory-policy allkeys-lruin redis.conf for cache workloads; usevolatile-lruif mixing cached and persistent data - Set
maxmemoryto at most 50% of physical RAM if writes are heavy — Redis's own admin guidance is that RDB/AOF rewrites can use up to 2x the memory already in use[5], and 2 × 50% is the full-RAM ceiling before that spike OOMs the process; read-heavy workloads with infrequent rewrites can run tighter - Monitor cache hit rate via
redis-cli INFO stats | grep keyspace_hitsand watch the trend — a falling rate is the early signal that your hot set no longer fits, beforeevicted_keysclimbs - Pool size: calculate as
connections_per_pod × num_pods. Start with 50 per pod; adjust if you see connection errors - Stagger pod restarts with readiness probes that warm the connection pool — prevents connection storms
- Replica read routing: use
RouteByLatency: truein go-redis to offload reads to replicas - Eviction rate: monitor
evicted_keys— if rising, your hot set is larger thanmaxmemoryor your TTLs are too aggressive - Replication lag: replicas should be within 10ms of primary. Monitor
slave_repl_offsetvsmaster_repl_offset
go-redis pool sizing that survives a slow Redis
Default redis.NewClient ships with PoolSize = 10 * runtime.GOMAXPROCS(0) and a multi-second read timeout (5s in go-redis v9.22, 3s in earlier v9 releases[6]; PoolTimeout = ReadTimeout + 1s) — fine until Redis blocks on a single slow command and every goroutine in your service queues behind it for seconds. The configuration below is built for services pushing 200k+ ops/sec:
import (
"time"
"github.com/redis/go-redis/v9"
)
func NewRedisClient(addr string) *redis.Client {
return redis.NewClient(&redis.Options{
Addr: addr,
// Pool sized for concurrency, not for parallelism.
// PoolSize > NumCPU is fine — Redis is single-threaded but our
// goroutines block on the network round-trip, not CPU.
PoolSize: 200,
MinIdleConns: 20,
// The four timeouts that fail-fast instead of cascading:
// DialTimeout — TCP connect ceiling, before any command runs.
// ReadTimeout — per-command read budget; tighter than HTTP request budget.
// WriteTimeout — protects against TCP backpressure from saturated link.
// PoolTimeout — how long a goroutine waits for a free connection;
// exceeding this returns an error instead of hanging.
DialTimeout: 500 * time.Millisecond,
ReadTimeout: 200 * time.Millisecond,
WriteTimeout: 200 * time.Millisecond,
PoolTimeout: 100 * time.Millisecond,
// Idle connection hygiene — beats most NAT/firewall idle drops.
ConnMaxIdleTime: 5 * time.Minute,
ConnMaxLifetime: 30 * time.Minute,
})
}For batch reads on a ClusterClient, one go-redis pipeline is already slot-aware: it maps each command to the node that owns its slot and sends every node its share concurrently[7], so 1,000 GETs cost about one round trip per node instead of 1,000. Don't split the batch into per-slot pipelines run in turn — that costs one round trip per distinct slot, close to one per key for random keys:
type batchedReader struct {
rdb redis.UniversalClient
}
// GetMany reads keys in one pipeline. Missing keys are left out of the map;
// any other per-key failure fails the whole call.
func (b *batchedReader) GetMany(ctx context.Context, keys []string) (map[string]string, error) {
pipe := b.rdb.Pipeline()
cmds := make([]*redis.StringCmd, len(keys))
for i, k := range keys {
cmds[i] = pipe.Get(ctx, k)
}
// Exec returns the first failed command's error, and a missing key
// (redis.Nil) can mask a later real failure — so check every command.
if _, err := pipe.Exec(ctx); err != nil && !errors.Is(err, redis.Nil) {
return nil, fmt.Errorf("pipeline exec: %w", err)
}
out := make(map[string]string, len(keys))
for i, c := range cmds {
v, err := c.Result()
switch {
case errors.Is(err, redis.Nil):
continue
case err != nil:
return nil, fmt.Errorf("get %s: %w", keys[i], err)
}
out[keys[i]] = v
}
return out, nil
}Grouping by slot still matters for commands Redis itself confines to one slot — MGET, MSET, MULTI/EXEC, Lua scripts — which the server rejects with CROSSSLOT Keys in request don't hash to the same slot when keys span slots[3]; hash tags fix that. The slot computation, handy for checking where hot-key replicas land, is pure CRC16 against the key (or the contents of the first {...} segment if present, the hash-tag escape hatch for forcing two keys to the same slot):
import "strings"
// slot returns the Redis Cluster hash slot for a key: CRC16(key) mod 16384,
// using the same CRC16/XMODEM (CCITT) variant Redis uses — polynomial 0x1021,
// initial value 0x0000. If the key contains a {tag} substring, only the tag is
// hashed, which forces co-location for keys you must batch atomically.
func slot(key string) uint16 {
if start := strings.IndexByte(key, '{'); start >= 0 {
if end := strings.IndexByte(key[start+1:], '}'); end > 0 {
key = key[start+1 : start+1+end] // hash only the {...} contents
}
}
var crc uint16
for i := 0; i < len(key); i++ {
crc ^= uint16(key[i]) << 8
for j := 0; j < 8; j++ {
if crc&0x8000 != 0 {
crc = (crc << 1) ^ 0x1021
} else {
crc <<= 1
}
}
}
return crc % 16384
}
// Group keys by Cluster slot. {tag} hash-tag forces co-location for keys
// you must batch atomically (the only path to safe MULTI/EXEC across keys
// in Cluster mode).
func groupKeysBySlot(keys []string) map[uint16][]string {
out := make(map[uint16][]string)
for _, k := range keys {
out[slot(k)] = append(out[slot(k)], k)
}
return out
}Streams vs Pub/Sub: pick durability deliberately
Classic Redis Pub/Sub is fire-and-forget. A subscriber that disconnects for two seconds loses every message published in that window, and there is no acknowledgement, no replay, no consumer group. That works for cache-invalidation fan-out where a missed message just means a slightly stale read on one node, but it falls apart the moment you reach for it as a work queue or event log.
Redis Streams (XADD, XREADGROUP, XACK) close that gap: messages persist to RDB and AOF, consumer groups distribute work across workers with at-least-once delivery, and a pending entries list (PEL) tracks every unacked message so a crashed consumer's work can be claimed by another via XCLAIM or auto-claimed via XAUTOCLAIM after an idle threshold. The trade-off is memory pressure — a stream grows until you cap it with MAXLEN or MINID.
The decision rule that holds up in production: use Pub/Sub only for ephemeral coordination signals where loss is acceptable (cache invalidation, presence notifications, leader-election heartbeats). Use Streams for any payload representing a durable event — payment intents, audit records, async job dispatch. The Go consumer below shows the canonical pattern with bounded retries and dead-letter routing for poison messages.
// Stream consumer with consumer-group semantics, bounded retries,
// and a dead-letter stream for poison messages. Run one goroutine per
// worker; the consumer name should be unique per process (hostname+pid).
func consume(ctx context.Context, rdb *redis.Client, group, consumer string) error {
cursor := "0-0"
for {
// ">" only ever delivers new entries, so a failed entry is never retried
// unless something re-claims it. XAUTOCLAIM hands this consumer entries
// idle > 30s (including a crashed peer's) and bumps their delivery count.
claimed, next, err := rdb.XAutoClaim(ctx, &redis.XAutoClaimArgs{
Stream: "orders", Group: group, Consumer: consumer,
MinIdle: 30 * time.Second, Start: cursor, Count: 16,
}).Result()
if err != nil {
return fmt.Errorf("xautoclaim: %w", err)
}
cursor = next
for _, msg := range claimed {
if err := process(ctx, rdb, group, msg); err != nil {
return err
}
}
res, err := rdb.XReadGroup(ctx, &redis.XReadGroupArgs{
Group: group,
Consumer: consumer,
Streams: []string{"orders", ">"},
Count: 16,
Block: 5 * time.Second,
}).Result()
if errors.Is(err, redis.Nil) {
continue // idle, no new messages
}
if err != nil {
return fmt.Errorf("xreadgroup: %w", err)
}
for _, msg := range res[0].Messages {
if err := process(ctx, rdb, group, msg); err != nil {
return err
}
}
}
}
// process ACKs on success. On failure it leaves the entry pending for a later
// XAUTOCLAIM until its delivery count reaches maxDeliveries, then moves it to
// the DLQ, ACKing only after the DLQ write succeeds so nothing is lost.
func process(ctx context.Context, rdb *redis.Client, group string, msg redis.XMessage) error {
const maxDeliveries = 5
herr := handle(ctx, msg.Values)
if herr == nil {
if err := rdb.XAck(ctx, "orders", group, msg.ID).Err(); err != nil {
return fmt.Errorf("xack %s: %w", msg.ID, err)
}
return nil
}
pending, err := rdb.XPendingExt(ctx, &redis.XPendingExtArgs{
Stream: "orders", Group: group, Start: msg.ID, End: msg.ID, Count: 1,
}).Result()
if err != nil {
return fmt.Errorf("xpending %s: %w", msg.ID, err)
}
if len(pending) == 0 || pending[0].RetryCount < maxDeliveries {
slog.Warn("handle failed; left pending for retry", "id", msg.ID, "err", herr)
return nil
}
if err := rdb.XAdd(ctx, &redis.XAddArgs{Stream: "orders.dlq", Values: msg.Values}).Err(); err != nil {
return fmt.Errorf("dlq xadd %s: %w", msg.ID, err)
}
if err := rdb.XAck(ctx, "orders", group, msg.ID).Err(); err != nil {
return fmt.Errorf("xack dead-lettered %s: %w", msg.ID, err)
}
slog.Error("moved poison message to orders.dlq", "id", msg.ID,
"deliveries", pending[0].RetryCount, "err", herr)
return nil
}The XAUTOCLAIM pass is what makes the retry bound real: each claim increments the entry's delivery counter unless you pass JUSTID[8], and the same pass recovers work a Kubernetes pod eviction stranded in the PEL. Set MinIdle above your processing SLA. Cap stream length with XADD orders MAXLEN ~ 1000000 * so producers approximately bound memory; the ~ lets Redis trim in chunks instead of one entry at a time, which keeps XADD latency stable under load.
Client-side caching with RESP3 broadcast invalidation
Redis 6 introduced server-assisted client-side caching: clients cache reads locally, and Redis sends invalidation messages whenever a tracked key changes — as RESP3 push messages, or through a Pub/Sub redirect on RESP2[9]. For hot keys, you replace a network round-trip with an in-process map lookup and remove the read entirely from Redis CPU.
Two tracking modes exist: default mode (server tracks every key each client reads, expensive for the server) and broadcast mode (clients subscribe to key-prefix patterns, server sends one invalidation per write fanned out to all matching subscribers). Broadcast mode is the only mode worth running at scale because server-side memory does not grow with the number of cached keys per client.
Three pitfalls bite teams adopting this. First, your local cache is eventually consistent — there is a window between a write landing on the primary and the invalidation reaching subscribers, so any read-your-writes guarantee must come from the application path that issued the write. Second, you must handle reconnects: if the tracking connection drops, you must invalidate the entire local cache before reconnecting, because invalidations during the disconnect were lost. Third, broadcast mode delivers invalidations even for keys you never read, so prefix selection matters for bandwidth.
// Client-side cache with RESP3 broadcast tracking. The tracking connection
// is separate from the command connection; invalidations arrive as RESP3
// push messages on the tracking conn and we apply them to the local map.
type CSCache struct {
mu sync.Mutex
local map[string][]byte
filling map[string]uint64 // in-flight GETs; an invalidation deletes the entry
seq uint64
rdb *redis.Client
}
func (c *CSCache) Start(ctx context.Context, prefixes []string) error {
// Tracking is per connection, so pin a dedicated one. Enabling it through
// the pooled client would attach it to whichever conn Do() happened to get.
conn := c.rdb.Conn()
// BCAST mode: server sends invalidations for any key matching a prefix,
// regardless of whether this client ever read it.
args := []any{"CLIENT", "TRACKING", "ON", "BCAST"}
for _, p := range prefixes {
args = append(args, "PREFIX", p)
}
if err := conn.Do(ctx, args...).Err(); err != nil {
return errors.Join(fmt.Errorf("enable tracking: %w", err), conn.Close())
}
go c.listen(ctx, conn) // consume invalidate push messages on the tracking conn
return nil
}
func (c *CSCache) Get(ctx context.Context, key string) ([]byte, error) {
c.mu.Lock()
if v, ok := c.local[key]; ok {
c.mu.Unlock()
return v, nil
}
c.seq++
token := c.seq
c.filling[key] = token
c.mu.Unlock()
v, err := c.rdb.Get(ctx, key).Bytes()
c.mu.Lock()
defer c.mu.Unlock()
// Data and invalidations travel on different connections, so an
// invalidation can overtake this GET's reply. Cache v only if none
// arrived for key while the GET was in flight.
if c.filling[key] == token {
delete(c.filling, key)
if err == nil {
c.local[key] = v
}
}
if err != nil {
return nil, err
}
return v, nil
}
// onInvalidate receives nil on FLUSHALL/FLUSHDB: drop everything.
func (c *CSCache) onInvalidate(keys []string) {
c.mu.Lock()
defer c.mu.Unlock()
if keys == nil {
clear(c.local)
clear(c.filling)
return
}
for _, k := range keys {
delete(c.local, k)
delete(c.filling, k)
}
}Bound the local cache size with an LRU — without one, a hostile or buggy access pattern will OOM the application before invalidations arrive. Track hit ratio per prefix as a Prometheus metric; a falling hit ratio is the early signal that your prefix is too coarse and you are paying invalidation bandwidth for keys nobody caches.
ACLs and TLS for multi-tenant production
Default Redis ships with a single password and no command-level authorization, which is fine for a development laptop and disastrous for a shared production cluster. Redis 6 added ACLs: per-user credentials with command, key-pattern, and channel-pattern restrictions. The right baseline is one user per service with the narrowest possible permission set, the legacy default user disabled, and TLS terminated at the Redis port (not at a sidecar) so that credentials never traverse the network in plaintext. Below is a users.acl implementing that baseline. An ACL file has no line continuations — each user is one line — the -command|subcommand rules need Redis 7.0+, and # comment lines load only on Redis 8.8+, so strip them for older servers[10].
# Disable the default account so no client can connect without explicit creds.
user default off
# Read-only analytics: GET/MGET/SCAN on the analytics: prefix only,
# no write commands, no admin commands, no pub/sub.
user analytics on >REDACTED-STRONG-SECRET ~analytics:* +@read +@connection -@dangerous
# Application service: full data-plane access on its own keyspace,
# Streams on its own queue, no FLUSHDB/FLUSHALL/CONFIG/DEBUG/SCRIPT LOAD.
user orders-svc on >REDACTED-STRONG-SECRET ~orders ~orders:* ~orders.dlq &orders.events +@all -@dangerous -flushall -flushdb -config -debug -keys -script|load
# Operator account for runbooks; gated behind break-glass workflow.
# Allowed CONFIG GET but not SET, allowed CLUSTER inspection but not failover.
user oncall on >REDACTED-STRONG-SECRET ~* &* +@all -@dangerous +config|get +cluster|info -cluster|failoverThree operational rules make ACLs durable. First, version the users.acl file in a config repo (storing #<sha256> password hashes, not >plaintext) and ship it via configuration management, not ACL SETUSER over the wire — drift between nodes is the failure mode that lets an attacker keep access after a credential rotation. Second, rotate secrets through a transition user (orders-svc-v2) before deleting the old one, so the cutover does not require a deploy of every consumer at the same instant.
Third, log ACL LOG to your SIEM — every denied command is a misconfiguration or an attempted privilege escalation, and silence on that channel is what you want to hear. Combine ACLs with tls-port 6379, tls-auth-clients yes, and per-service client certificates so that even a leaked password on its own does not authenticate against the cluster.
Conclusion
Apply the fixes above to the flash-sale scenario that opened this piece — 10-way hot-key replication, a 1-second local LRU, pipelined batch reads, and connection pools sized for the spike — and the failure mode changes shape. The same flood of users hitting product:flash-sale:current no longer pins one shard while five sit idle: the local LRU absorbs most repeat reads, and replicated reads spread the rest across 10 slots. On 6 shards that is not a tenth per shard: fitting 10 replicas into 6 shards guarantees some shard owns at least two of them by pigeonhole — 2 of 10 replicas is 20% of the hot-key load at best, 3 of 10 is 30% with unlucky placement — so check with CLUSTER KEYSLOT and adjust suffixes.
Scaling Redis isn't about raw ops/sec: one instance is bounded by one core, and uneven key distribution can make a 6-shard cluster behave like one.
Start with Sentinel and a single instance. Move to Cluster only when you hit throughput limits. Prevent hot keys with in-process LRU caches. Batch reads with pipelining. Pool connections with discipline. Monitor hit rates, not just ops/sec. [2]
Frequently Asked Questions
Why is Redis single-threaded and how does that affect scaling?
Redis processes commands on a single thread, avoiding lock contention. This means a single instance's throughput is bounded by one core (Redis's benchmark docs show a sample unpipelined run at ~180k SET/sec[1]), and you must use Redis Cluster for horizontal scaling across multiple CPU cores.
What is a Redis hot key and how do you fix it?
A hot key is a single key receiving disproportionate traffic, causing one shard to bottleneck while others sit idle. Fix it by replicating the key under N names that hash to different slots (e.g., product:<n>:<id> with n picked at random per read, and no {hash tag}, which would pin every copy to one slot), using read replicas, or restructuring data to distribute load.
When should you use Redis Sentinel vs Redis Cluster?
Use Redis Sentinel for high availability while one primary covers your dataset and throughput (e.g., under ~25GB and ~100k ops/sec — a rule of thumb, not a Redis limit). Use Redis Cluster when you need horizontal scaling beyond one machine's RAM or throughput, as it shards data across multiple primaries.
How do you use Redis pipelining to improve throughput?
Pipelining batches multiple Redis commands into a single network round trip instead of waiting for each response individually. This reduces network overhead dramatically — Redis's own docs measure a 5× speedup even over loopback, and throughput up to 10× the unpipelined baseline[2].
Keep Reading
- Caching Strategies at Scale: The Complete Guide Beyond Key-Value Stores — Cache-aside, write-through, write-behind patterns; cache stampede prevention; event-driven invalidation
- Database Indexing Strategies: B-Trees, GIN, GiST, and Production Tuning — Before reaching for Redis, index your queries properly — a 2ms indexed query doesn't need a cache layer
- Understanding Raft Consensus — How Redis Cluster's gossip topology compares to Raft consensus for distributed system reliability
- Rate Limiter Algorithms — Token bucket, sliding window log, and sliding window counter — every algorithm uses Redis for shared state across pods
- Distributed Rate Limiting (Probabilistic Drop) — When per-request Redis-Lua adds too much latency, drop_ratio gives you global enforcement with local in-memory checks
Sources
- 1.Redis benchmark (redis-benchmark utility) — redis.io, 2026
- 2.Redis Pipelining (Documentation) — redis.io, 2026
- 3.Redis Cluster Specification — redis.io, 2026
- 4.Optimal Probabilistic Cache Stampede Prevention (Vattani et al., 2015) — VLDB Endowment, Vol 8 No 8, 2015
- 5.Redis administration — redis.io, 2026
- 6.go-redis v9.22.0 Release Notes — GitHub (redis/go-redis), 2026
- 7.go-redis v9.22.0 source: osscluster.go (ClusterClient) — GitHub (redis/go-redis), 2026
- 8.XAUTOCLAIM (Redis command reference) — redis.io, 2026
- 9.Client-side caching reference (Redis documentation) — redis.io, 2026
- 10.ACL (Redis Access Control List documentation) — redis.io, 2026
Engineering Team
An independent engineering publication covering distributed systems, databases, and production infrastructure. Every factual claim is cited to a primary source or removed.
Read Next
Caching Strategies at Scale
Four caching patterns (cache-aside, write-through, write-behind, read-through), plus Go code for stampede prevention, multi-tier caching, and event-based invalidation.
Consistent Hashing: The Algorithm Behind Scalable Distributed Systems
Adding one cache server shouldn't invalidate every key. Consistent hashing with virtual nodes and bounded loads — full Go and Java implementations.
Distributed Rate Limiting at Scale: The Probabilistic Drop Architecture
Probabilistic drop rate limiting: uncoordinated enforcement that takes Redis off the request path, with no coordination per request.