The backup that had been failing for five months
Discovered on 28 August, during a restore I did not need to do. The nightly snapshot had been running since late March and backing up an empty directory.
The cause was embarrassing and will be familiar. In March I moved the data volume from /srv/data to /mnt/tank/data and updated everything that referenced it — except the backup unit, which pointed at a path that, after the move, simply no longer existed.
Here is the part worth writing down. restic backup /srv/data against a missing directory does not fail loudly. It walks nothing, finds nothing, creates a snapshot containing nothing, and exits 0. The systemd timer saw a clean exit every night for five months and had nothing to report. My monitoring checked that the unit ran. It did run. Beautifully.
The repository was fine. The credentials were fine. The schedule was fine. Every green light I had was telling me the truth about the thing it measured, and none of them measured whether any bytes had moved.
What I changed:
The unit now runs a wrapper that checks the source path exists and is non-empty before it calls restic, and exits non-zero if not.
After each run it queries the latest snapshot and compares the file count against the previous one. A drop of more than 10% fails the unit. An increase is fine; a cliff is not.
Once a month a timer restores one randomly chosen file to a scratch directory and diffs it against the original. If that diff fails, the backup is not a backup.
The lesson is not 'test your restores', which everyone already says and I already believed. It is narrower: exit code 0 means the command did what it was told, not what you wanted. I had told it to back up a directory that wasn't there, and it did that perfectly.