How we saved $4,000/mo by replacing Redis with SQLite for our job queue

Architecture Migration: Scaling background job queues with local SQLite + Litestream instead of centralized Redis

Last year, our background worker service started hitting memory limits on AWS ElastiCache. We were spending roughly $4,500/month keeping a large Redis instance alive just to process around 15 million short-lived background jobs a day.

Instead of scaling up to a larger cluster, we experimented with running SQLite in WAL (Write-Ahead Logging) mode directly on the worker nodes with busy_timeout set to 5000ms.


The Architecture Change

*Before: 12 API nodes \rightarrow Centralized Redis Cluster \rightarrow 8 Worker nodes.

*After: 12 API nodes write incoming payloads directly to a local, append-only SQLite database on NVMe drives, synchronized using Litestream to S3 for durability.

Performance & Trade-offs

*Throughput: P99 latency dropped from 14ms (network roundtrip to Redis) to 1.2ms (local disk write).

*Cost: Cut our monthly AWS bill from $4,500 to $320 (just the EC2 NVMe instances and S3 storage).

*The Catch: Horizontal scaling requires explicit shard keying now, since each node manages its own local queue DB rather than a single shared state.

Here is a minimal benchmark script in Go showing the WAL throughput comparison: [github link]

What strategies are you using for single-node queue performance before reaching for distributed in-memory stores?

Leave a Reply

Discover more from Nationwide Private Investigator & Skip Tracing | Knoxville TN Background Checks & OSINT | Kyle's Investigation

Subscribe now to keep reading and get access to the full archive.

Continue reading