On a Thursday in June roughly a quarter of the outbound calls from our pricing service started timing out. Only from some pods, only on one node, and Kubernetes had no opinion about it. Every pod on that node was ready. Probes need three failures in a row, and most of them got through.
The node’s kernel log had the answer, repeated several thousand times a minute: nf_conntrack table full, dropping packet. Linux tracks every connection passing through its network stack in a table, and Kubernetes networking relies on that table for service routing. When it is full, the kernel drops the first packet of any new connection, from any pod on the node, until entries expire. Nothing was collecting kernel logs from our nodes, so nobody had seen it.
What filled it was a data export job scheduled onto the same node that morning. It called a partner API about twelve hundred times a second and opened a fresh HTTPS connection for every request, because its HTTP client was created inside the loop. A closed TCP connection stays in the table for two minutes by default. At the job’s rate that was more entries than the node’s limit of a hundred and thirty one thousand, and it stayed over the limit for as long as the job ran. The job itself was slow but succeeded, since it retried. Our pricing service, which shared nothing with it except the kernel, failed in bursts for three hours.
The fixes, in the order they mattered. The export job creates one client and reuses its connections, which took it from about twelve hundred new connections a second to a few dozen. Node metrics now include conntrack entries against the limit, with an alert at seventy percent. Kernel logs from every node go to the same place as application logs. Batch jobs that talk to the outside world run on their own node pool, tainted so nothing else lands there. And we raised the table limit on our node image, sized to the memory we can spare for it.
A pod’s isolation covers what its manifest can describe: CPU, memory, storage. The kernel keeps other shared tables that no resource request mentions, and one careless neighbour can exhaust them for everybody on the machine.
– Sergey Shinder