DNS is not the problem. Except this time.
Symptom: about one request in three to an internal service timed out. The other two were instant. No pattern by client, by time, or by service.
Intermittent-and-uniform is the worst shape a fault can have, because it rules out most of the things that are easy to check. A dead backend fails every time. A slow backend is slow every time. Something that fails exactly a third of the time is choosing.
It was choosing between resolvers.
The setup: a local resolver serving a split-horizon zone so internal names point at internal addresses, and — left over from an experiment in February — a second resolver still handed out by DHCP as a secondary. That second one knew nothing about the internal zone. It answered from public DNS, got the external address, and handed back a route that went out to the internet and came back to a firewall that quite correctly dropped it.
Clients pick a resolver per query, not per session. Hence one in three. Hence no pattern by anything I was looking at.
What found it was dig in a loop, twenty times, printing the answer and the server that gave it. Two different answers, alternating unpredictably. Once you can see the two answers side by side the whole thing collapses into something obvious.
I had spent the preceding four hours on the application, the reverse proxy, the container network, and MTU — because those are the things I have been burned by before, and debugging is mostly a walk through your own history of being wrong.
Fix was deleting one line from the DHCP config.
The note to self: when a fault is intermittent at a stable ratio, stop looking for something that is broken and start looking for something that is choosing. Two of anything — resolvers, routes, upstreams, replicas — where one is wrong.