Posts / homelab
A Copy Is Not a Backup, and Other Things I Already Knew
Saw a post this week from someone on r/selfhosted who lost the lot. Ten years of a personal diary, gone. Recipes, Home Assistant config, unchecked-out code. The setup was two disks: one live, one “backup”, with a nightly dd copying everything from the first to the second. Sounds sensible until you realise what happens when the source disk is already corrupted. You don’t get a backup. You get two corrupted disks and a very bad morning.
The top comment called it “redneck RAID1.” Somebody else pointed out that’s actually worse than RAID1, because RAID1 at least has some awareness of what it’s mirroring. dd doesn’t care. It just copies bytes, good or bad, and it will do that with total confidence and zero judgement. Which is sort of the problem with a lot of automation, now that I think about it. It does exactly what you told it to, not what you meant.
I’ve got sympathy for this bloke, because I’ve built versions of this exact mistake myself. Not with dd, but with the general shape of “I’ve got a copy somewhere, that counts as safe.” It doesn’t. A copy tells you where your data was. A backup tells you where your data was at several different points in time, so that when today’s version is garbage, yesterday’s or last week’s is still sitting there, untouched, waiting to save your arse.
The 3-2-1 rule got mentioned about forty times in that thread, and one commenter made the point that actually mattered: 3-2-1 doesn’t save you here. Three copies, two media types, one offsite, sure, but if all three are being silently overwritten with the same corrupted nightly sync, you’ve just got three corrupted copies in three locations. What you actually need is versioning. Retention. The ability to reach back to last Tuesday, not just “the backup”, singular, present tense.
I run Restic on my own homelab now, after a scare a few years back that wasn’t quite this bad but was bad enough to change my behaviour. Incremental snapshots, deduplicated, going back months. It’s not glamorous. Nobody brags about their backup strategy at a barbecue, and if they did I’d quietly leave. But the peace of mind is real, in the same low-grade way that having a will or paying your car insurance is real. You don’t think about it until the one day you desperately need to.
The bit that got me, reading through that thread, was the diary. Ten years of daily writing, sitting in a Git repo, unchecked out anywhere else, gone in one bad dd. That’s not a config file you can rebuild from a template. That’s a decade of a person’s actual thoughts. I don’t know the guy, but I felt that one. I keep a similar habit myself, and it’s exactly the kind of thing you assume is safe because “it’s backed up”, right up until it isn’t.
None of this is really about dd or Restic or which acronym you should be running. It’s about the gap between what we assume our systems are doing and what they’re actually doing. I work in tech for a living and I still catch myself trusting a green tick I haven’t actually tested. Somebody in the thread put it well: a backup without a tested restore is thoughts and prayers. I’m going to go check mine this weekend, properly, not just glance at the log and nod.
Small, unglamorous lesson, but a genuine one: it’s not backed up until you’ve watched it come back from the dead.