{
  "description" : "Notes from a small server under the stairs. Mostly what broke.",
  "feed_url" : "https://selfpubli.sh/examples/homelab-log/feed.json",
  "home_page_url" : "https://selfpubli.sh/examples/homelab-log/",
  "items" : [
    {
      "content_html" : "<p>Discovered on 28 August, during a restore I did not need to do. The nightly snapshot had been running since late March and backing up an empty directory.</p>\n<p>The cause was embarrassing and will be familiar. In March I moved the data volume from /srv/data to /mnt/tank/data and updated everything that referenced it — except the backup unit, which pointed at a path that, after the move, simply no longer existed.</p>\n<p>Here is the part worth writing down. restic backup /srv/data against a missing directory does not fail loudly. It walks nothing, finds nothing, creates a snapshot containing nothing, and exits 0. The systemd timer saw a clean exit every night for five months and had nothing to report. My monitoring checked that the unit ran. It did run. Beautifully.</p>\n<p>The repository was fine. The credentials were fine. The schedule was fine. Every green light I had was telling me the truth about the thing it measured, and none of them measured whether any bytes had moved.</p>\n<p>What I changed:</p>\n<p>The unit now runs a wrapper that checks the source path exists and is non-empty before it calls restic, and exits non-zero if not.</p>\n<p>After each run it queries the latest snapshot and compares the file count against the previous one. A drop of more than 10% fails the unit. An increase is fine; a cliff is not.</p>\n<p>Once a month a timer restores one randomly chosen file to a scratch directory and diffs it against the original. If that diff fails, the backup is not a backup.</p>\n<p>The lesson is not &#39;test your restores&#39;, which everyone already says and I already believed. It is narrower: <strong>exit code 0 means the command did what it was told, not what you wanted.</strong> I had told it to back up a directory that wasn&#39;t there, and it did that perfectly.</p>",
      "date_published" : "2026-09-01T23:00:00Z",
      "id" : "https://selfpubli.sh/examples/homelab-log/posts/the-backup-that-had-been-failing-for-five-months.html",
      "summary" : "Discovered on 28 August, during a restore I did not need to do. The nightly snapshot had been running since late March and backing up an empty directory.",
      "title" : "The backup that had been failing for five months",
      "url" : "https://selfpubli.sh/examples/homelab-log/posts/the-backup-that-had-been-failing-for-five-months.html"
    },
    {
      "content_html" : "<p>smartctl -H said PASSED right up to the day I pulled it, which is the most useful thing I can tell you about smartctl -H.</p>\n<p>The overall health assessment is a threshold test. A drive fails it when an attribute crosses the manufacturer&#39;s limit. Everything below that limit is PASSED, including a drive whose reallocated sector count has been climbing steadily for two and a half months.</p>\n<p>Mine went 0, 0, 0, then 8, then 24, then 56, then 104 over eleven weeks. The threshold was somewhere north of 1000. It would have passed for a long time yet.</p>\n<p>What made it visible was recording the numbers rather than the verdict. A cron job writes the raw values of attributes 5 (reallocated sectors), 187 (reported uncorrectable), 197 (current pending) and 198 (offline uncorrectable) to a file once a day. Nothing clever — one line per disk per day. The pattern was obvious within a week of looking at it as a series instead of a snapshot.</p>\n<p>The replacement itself was uneventful, which is the whole argument for doing it early. zpool replace, seven hours of resilver on a mirror with the array still serving files, no downtime, no restore, no stress. The failed-drive version of that same afternoon involves a degraded pool, a resilver you are watching nervously, and the non-zero chance that the second drive picks that exact window to find a bad sector of its own.</p>\n<p>A disk with 104 reallocated sectors is not dead. It may never die. But it costs about ninety pounds to stop thinking about it, and the alternative is to keep a running tally in your head of how worried to be.</p>\n<p>Old drive is in a drawer, wiped, labelled with the date and the sector count. If it turns out to run for another five years I will report back and feel foolish.</p>",
      "date_published" : "2026-08-10T23:00:00Z",
      "id" : "https://selfpubli.sh/examples/homelab-log/posts/replacing-a-drive-before-it-failed.html",
      "summary" : "smartctl -H said PASSED right up to the day I pulled it, which is the most useful thing I can tell you about smartctl -H.",
      "title" : "Replacing a drive before it failed",
      "url" : "https://selfpubli.sh/examples/homelab-log/posts/replacing-a-drive-before-it-failed.html"
    },
    {
      "content_html" : "<p>Symptom: about one request in three to an internal service timed out. The other two were instant. No pattern by client, by time, or by service.</p>\n<p>Intermittent-and-uniform is the worst shape a fault can have, because it rules out most of the things that are easy to check. A dead backend fails every time. A slow backend is slow every time. Something that fails exactly a third of the time is choosing.</p>\n<p>It was choosing between resolvers.</p>\n<p>The setup: a local resolver serving a split-horizon zone so internal names point at internal addresses, and — left over from an experiment in February — a second resolver still handed out by DHCP as a secondary. That second one knew nothing about the internal zone. It answered from public DNS, got the external address, and handed back a route that went out to the internet and came back to a firewall that quite correctly dropped it.</p>\n<p>Clients pick a resolver per query, not per session. Hence one in three. Hence no pattern by anything I was looking at.</p>\n<p>What found it was dig in a loop, twenty times, printing the answer and the server that gave it. Two different answers, alternating unpredictably. Once you can see the two answers side by side the whole thing collapses into something obvious.</p>\n<p>I had spent the preceding four hours on the application, the reverse proxy, the container network, and MTU — because those are the things I have been burned by before, and debugging is mostly a walk through your own history of being wrong.</p>\n<p>Fix was deleting one line from the DHCP config.</p>\n<p>The note to self: when a fault is intermittent at a stable ratio, stop looking for something that is broken and start looking for something that is <em>choosing</em>. Two of anything — resolvers, routes, upstreams, replicas — where one is wrong.</p>",
      "date_published" : "2026-07-22T23:00:00Z",
      "id" : "https://selfpubli.sh/examples/homelab-log/posts/dns-is-not-the-problem-except-this-time.html",
      "summary" : "Symptom: about one request in three to an internal service timed out. The other two were instant. No pattern by client, by time, or by service.",
      "title" : "DNS is not the problem. Except this time.",
      "url" : "https://selfpubli.sh/examples/homelab-log/posts/dns-is-not-the-problem-except-this-time.html"
    },
    {
      "content_html" : "<p>Drive temperatures went from a steady 38°C to touching 54°C over four days in late June, which is the sort of number that gets your attention.</p>\n<p>The cupboard is a good location for every reason except the physical one. It is central, it is near the incoming line, it is out of the way, and it means the fan noise happens somewhere nobody sits. It is also a sealed box with a door, and the inside of a sealed box ends up at whatever temperature the equipment decides.</p>\n<p>First instinct was a bigger fan. This was wrong, and it took an afternoon of moving air around to work out why: there was nowhere for the air to go. A more powerful fan in a closed cupboard is a device for stirring hot air vigorously. The intake and the exhaust were both drawing from and dumping into the same eight cubic feet.</p>\n<p>The fix was two vents, not one. A low one on the door and a high one through the side panel into the hall, so cold air is pulled in at the bottom and hot air leaves at the top on its own. Convection does most of it; the existing case fans do the rest.</p>\n<p>Steady state is now 36–39°C through the warmest week of the summer, which is better than the original figure, and the fans are quieter because they are no longer fighting.</p>\n<p>Two things I would tell anyone putting a machine in a cupboard. Measure before and after, because &#39;it feels cooler&#39; is not data and you will convince yourself of anything. And think about the path the air takes through the whole space, not about the fan — the fan was never the constraint.</p>\n<p>Total cost: two vent grilles and a hole saw.</p>",
      "date_published" : "2026-07-03T23:00:00Z",
      "id" : "https://selfpubli.sh/examples/homelab-log/posts/the-under-stairs-thermal-problem.html",
      "summary" : "Drive temperatures went from a steady 38°C to touching 54°C over four days in late June, which is the sort of number that gets your attention.",
      "title" : "The under-stairs thermal problem",
      "url" : "https://selfpubli.sh/examples/homelab-log/posts/the-under-stairs-thermal-problem.html"
    },
    {
      "content_html" : "<p>Mild irony worth admitting: the log of a self-hosted infrastructure project is a static site with no infrastructure at all.</p>\n<p>That is deliberate, and the reason is the thing this site is about.</p>\n<p>For two years the notes lived in a self-hosted wiki. It was good software. It also needed a container, a database, a reverse proxy entry, a certificate, a backup job that understood the database, and an upgrade every few weeks that occasionally wanted the database migrated. The documentation of the homelab had become a component of the homelab, with its own failure modes.</p>\n<p>Then the disk problem happened, and the notes about disk problems were on the array with the disk problem.</p>\n<p>Now it is HTML. The pages are written on a phone, generated into a folder, and pushed over SFTP to a server that does nothing except serve files. The backup is a directory of text. If the machine under the stairs is off, on fire, or halfway through a resilver, this still loads, because it is not there.</p>\n<p>There is a broader point about the homelab hobby here. The instinct is to self-host everything, and most of the time that instinct is right — you learn more, you own more, you depend on fewer people. But every service you run is a thing that can be down at the worst possible moment, and the worst possible moment for your documentation is precisely when everything else is broken.</p>\n<p>So: the array holds the data. The cupboard holds the machine. The notes about both live somewhere boring, in the dumbest possible format, on hosting I could replace in an afternoon.</p>\n<p>That is not a compromise. It is the correct architecture for the job.</p>",
      "date_published" : "2026-06-14T23:00:00Z",
      "id" : "https://selfpubli.sh/examples/homelab-log/posts/everything-here-is-a-file.html",
      "summary" : "Mild irony worth admitting: the log of a self-hosted infrastructure project is a static site with no infrastructure at all.",
      "title" : "Everything here is a file",
      "url" : "https://selfpubli.sh/examples/homelab-log/posts/everything-here-is-a-file.html"
    },
    {
      "content_html" : "<p>The machine is unremarkable on purpose: a low-power board, 32GB of ECC, four 8TB drives in two mirrored pairs, and a boot SSD that is not part of the array and holds nothing I would miss.</p>\n<p>Mirrors rather than a wider parity array, which costs half the raw capacity and buys a resilver that finishes in hours rather than days. Given the rest of this blog will mostly be about drives misbehaving, that trade has already justified itself twice.</p>\n<p>What it does: files, photographs, backups of four laptops, and a couple of small services that would be irritating rather than catastrophic to lose.</p>\n<p>What it deliberately does not do: anything anyone else in the house depends on in real time. No home automation, no DNS for the whole house, no media server that makes the television useless when I am halfway through an upgrade. That rule has survived contact with two years of temptation and is the single reason this remains a hobby rather than an obligation.</p>\n<p>The log exists because the last two machines taught me the same lesson twice. You fix something at eleven at night, it works, you go to bed, and eight months later the identical fault appears and you remember nothing except a vague sense of having been here before.</p>\n<p>So the format is deliberately boring: what changed, what broke, what fixed it, dated. Not tutorials. There are enough tutorials, most of them written by people on day one of a thing rather than month nine, which is when you find out whether the advice was any good.</p>\n<p>Starting numbers, for something to compare against later: 14TB usable, 2.1TB in use, all four drives at zero reallocated sectors, idle draw 41 watts.</p>",
      "date_published" : "2026-05-20T23:00:00Z",
      "id" : "https://selfpubli.sh/examples/homelab-log/posts/day-one.html",
      "summary" : "The machine is unremarkable on purpose: a low-power board, 32GB of ECC, four 8TB drives in two mirrored pairs, and a boot SSD that is not part of the array and holds nothing I would miss.",
      "title" : "Day one",
      "url" : "https://selfpubli.sh/examples/homelab-log/posts/day-one.html"
    }
  ],
  "title" : "homelab.log",
  "version" : "https://jsonfeed.org/version/1.1"
}