How a Missing TCP_NODELAY Flag Silently Ate 80% of Our Database Throughput

Last Tuesday at 02:14 EST, our primary checkout API suddenly started timing out.

Nothing crashed. CPU usage on the API servers sat at a relaxed 12%. CPU on our PostgreSQL primary was hovering around 18%. Yet response times for our core transaction endpoint surged from 35ms to over 4,200ms within minutes, eventually dropping incoming connections as HTTP 504s.

Here is what actually happened, how we tracked it down, and why a well-intentioned optimization in a dependency caused a quiet cascading failure.

The Incident Timeline

*02:10 UTC: Deployment of v2.14.0 completes across all 12 API nodes. The release was small: dependency updates and a minor patch to clean up database connection pool configuration.

*02:14 UTC: PagerDuty fires. P99 latency on /v1/checkout spikes to 4.2s.

*02:22 UTC: We roll back to v2.13.9. Latency immediately drops back to 35ms.

*02:45 UTC: Post-incident triage begins.

The Red Herring

Our first assumption was a slow query introduced in the release. We pulled pg_stat_activity and pg_stat_statements expecting to find a missing index or a sequential scan locking rows.

Instead, we saw something strange:

Nearly all 200 active connections were in state active, but their wait_event_type was consistently set to Client and wait_event was ClientWrite.

ClientWrite in Postgres means the database engine has processed the query result and is currently waiting for the client application socket to consume the data.

The database wasn’t struggling to compute or fetch data—it was stuck waiting for our Go API service to accept the network packets it was sending.

Digging into the Network Socket

We spun up a staging environment, mirrored a sample of production traffic, and attached strace to one of the API processes:

Looking at the syscall timings, a pattern immediately stood out:

Every time the application sent a small packet over the DB connection, the subsequent read took almost exactly 40 milliseconds to return.

40ms is not a random number in networking. It is the default timer for Nagle’s Algorithm and TCP Delayed ACK in the Linux kernel interaction.

The Root Cause: Nagle’s Algorithm Meets Small Writes

Nagle’s Algorithm (RFC 896) was designed in 1984 to prevent small TCP packets from clogging networks. It works by buffering small outgoing packets until an ACK is received for the previous packet, or until enough data accumulates to fill a full TCP Frame (MSS).

Meanwhile, Linux TCP receivers use Delayed ACKs, holding back an acknowledgement for up to 40ms to see if they can piggyback it on an outgoing response.

When you pair Nagle’s Algorithm on the sender with Delayed ACK on the receiver, they deadlock each other in a 40ms dance:

1. Sender sends a small chunk of data (e.g., query header/parameters).

2. Sender waits to send the next chunk until an ACK arrives.

3. Receiver receives the chunk, but waits 40ms to send an ACK hoping for piggyback data.

4. 40ms elapses \rightarrow Receiver sends ACK \rightarrow Sender sends second chunk.

In our v2.14.0 release, we upgraded our Go database driver. The new version removed a low-level socket initializer that previously explicitly set TCP_NODELAY (which disables Nagle’s algorithm) on all outgoing connection sockets.
Because TCP_NODELAY was no longer set, every multi-part database request (e.g., a prepared statement execution followed by parameters) was incurring a forced 40ms penalty over TCP. Under normal load, this resulted in thousands of threads waiting on network socket I/O simultaneously.

The Fix

Disabling Nagle’s algorithm on client sockets forces the network stack to send packets immediately, regardless of size.

We added an explicit socket control callback back into the driver’s connection initialization block:

Benchmarks: Before vs. After
After deploying the socket fix to staging under a simulated load of 5,000 req/sec:

*Metric: P50 Latency
-v2.14.0 (Nagle Enabled): 42.1 ms
-v2.14.1 (TCP_NODELAY Fixed): 2.4 ms

*Metric: P99 Latency
-v2.14.0 (Nagle Enabled): 4,120.0 ms
-v2.14.1 (TCP_NODELAY Fixed): 18.2 ms

*Metric: Max DB Active Conn
-v2.14.0 (Nagle Enabled): 200 (exhausted)
-v2.14.1 (TCP_NODELAY Fixed): 14

*Metric: Throughput (RPS)
-v2.14.0 (Nagle Enabled): 1,100 req/sec (max)
-v2.14.1 (TCP_NODELAY Fixed): 8,400 req/sec

What We Learned

1. ClientWrite wait events mean look at the client, not the DB. When Postgres reports ClientWrite, it has done its job. The bottleneck is in the application process or network buffer layer.

2. Never assume low-level driver defaults remain static during major version bumps. Diff the underlying connection dialers during driver upgrades, not just exported API changes.

3. Synthetic load tests need real-world latency profiles. Our staging environment tested throughput on localhost (loopback), where TCP latency is virtually zero, masking the 40ms timer effect entirely.

Leave a Reply

Discover more from Nationwide Private Investigator & Skip Tracing | Knoxville TN Background Checks & OSINT | Kyle's Investigation

Subscribe now to keep reading and get access to the full archive.

Continue reading