Which service lost its logs depended on who wrote first after midnight

작성자

카테고리:

← 피드로
DEV Community · Sergey Shinder · 2026-09-29 개발(SW)

Sergey Shinder

During an incident in August I searched our logs for errors from the fulfilment service over the previous hour and got nothing back. I took that to mean fulfilment was fine, and we spent forty minutes looking elsewhere. Fulfilment had been logging errors the whole time. None of them had been indexed.

Our logs go into a new index every day, and fields that no template describes are mapped automatically from the first document that contains them. A field called error.code had been a number in every service for years, until a library upgrade in fulfilment changed it to a string such as TIMEOUT_UPSTREAM. From then on, each day’s index took the type of whichever service happened to log an error code first after midnight. If a numeric one came first, every fulfilment error that day was rejected with a mapping exception. If fulfilment came first, the other services lost theirs.

The log shipper did not lose anything quietly on purpose. It received the rejection, wrote it to its own debug output, and moved on, because retrying a document that can never be accepted would block everything behind it. Our dashboards counted what had been indexed, and on a bad day that simply looked like a quiet service.

Going back through the shipper’s output, we found that for five weeks roughly half of all error lines in the company had not reached the index on any given day, and which half changed at midnight.

The platform team now owns an index template for the fields every service shares, with error.code mapped as a keyword so that numbers and strings both fit. Fields nobody has described are kept in the stored document but not mapped, so a new field can never make an old one fail. Rejected documents go to a dead letter index instead of a debug log, and the rejection rate is a metric with an alert above zero for more than ten minutes. And every service panel shows the lines sent by the shipper beside the lines indexed, because the gap between those two numbers is the one thing a search result can never tell you.

An empty search result reads as good news. It is only good news if you know that everything which was sent actually arrived, and until then it is just an empty result.

– Sergey Shinder

원문에서 계속 ↗