← Back to the discography Break Stuff 02

Head-On Collision

A production incident story: one Terraform apply, one disk I attached by hand, and the Friday morning my automation and I drove straight into each other.


People talk a lot about automation.

And people talk about doing things by hand.

Nobody talks enough about the thing in between.

Half-automation.

The machine drives most of the way.

Somewhere along the road, you grab the wheel for a minute.

And you forget to tell the machine.

That’s the version that took down our production database.

On a Friday morning.


How we build VMs

Every VM I create goes through Terraform.

We run a Proxmox cluster. The VM definitions live in code. Plan, apply, the VM exists.

No clicking around in the UI.

That was the rule.

That was the rule.


Then the cluster got upgraded

At some point we upgraded the Proxmox cluster.

After that, the Terraform script stopped working.

I didn’t realise.

Not long after, the production database VM needed an extra disk.

The proper way would have been simple. Add the disk to the code. Plan. Apply.

But the Terraform side wasn’t working, and I wasn’t in the mood to find out why.

So I opened the Proxmox UI.

Add hard disk.

A few clicks.

Mounted it.

Done.

Next.

It worked.

And it kept working.

For a long, long time.

Long enough that I forgot it was the one disk Terraform knew nothing about.

So now this VM had two drivers.

Terraform, which thought it knew what the VM looked like.

And me, who knew what it actually looked like.

Only one of us had told the other.

Spoiler: it wasn’t me.


Friday morning

It was around 11 AM.

Friday.

I was on a call with my parents.

And somewhere during that call, I randomly decided:

“Today’s the day I fix the Terraform script.”

Update the version. Get it working again. A bit of housekeeping.

Very responsible.

I ran terraform plan.

It looked good.

Nothing to change.

Safe to apply.

Which, honestly, was exactly what I wanted to hear.

So I applied.

Boom.


The moment

Slack lit up.

Alerts.

Messages from devs.

Messages in the channel.

I cut the call with my parents.

SSH’d into the database VM.

No response.

Opened the console from Proxmox.

The mount was gone.

The disk I’d attached by hand wasn’t attached anymore.

The VM couldn’t start.

And this wasn’t just any VM.

It was the production database.

A shared production database.

Five or six tenants, all on the same machine.

So this wasn’t one customer having a bad morning.

It was everyone.


Panic

I’ll be honest.

I’d never faced this before.

And I panicked.

I didn’t know what to do, so I did what you do when you don’t know what to do.

I asked the datacenter for help.

And then, while I was waiting, it finally clicked.

Proxmox hadn’t deleted the disk.

It had detached it.

It was sitting right there, under Unused Disk.

The whole time.

I reattached it.

The VM came up.

The mount came back.

The database came back.

Roughly an hour, start to finish.

As far as we could tell, no data was lost.

Looking back, the fix was a few clicks away.

But panic doesn’t look for the cause.

Panic looks for the fastest way to make the alerts stop.

Every thought in my head was:

“Bring the system up.”

Not:

“Where did the disk go?”

I never went to troubleshoot in that direction.

I was so busy trying to get the patient breathing again that I never asked what actually happened to it.

The answer was sitting one tab away.


The backup I didn’t want to use

We had backups.

Every night at midnight, German time.

And I’m confident they would have restored the database just fine.

That was never the question.

The question was: was it worth it?

The incident happened around 11 AM.

Restoring the midnight backup meant throwing away several hours of data.

Across every tenant.

Every write, every change, every record since midnight.

Gone.

So yes, the backup would have brought the system up.

It just would have brought it up several hours in the past.

That’s not a fix.

That’s a price.

A backup is your last resort. Not your first answer.

I’m glad I didn’t have to pay it.

The disk was there.


What I told my lead

I messaged my team lead.

“The VM crashed.”

Technically true.

The VM did crash.

I just left out the part where it crashed because I drove into it.

I wrote the full incident report afterwards.

And now, I suppose, this post.


Two drivers, one lane

Here’s the thing.

I can’t blame Terraform for this.

The code described a VM.

The real VM was different, because I’d made it different with my own hands and never put that change back into the code.

The disk was never in the code.

And I’d just bumped the provider version, against a cluster that had been upgraded underneath it.

Exactly why that apply ended with my disk sitting under Unused Disk, I’d have to trace through the Proxmox provider and how it handled a VM that no longer matched its code.

I’m not going to pretend I know every detail of that.

But I know this much.

I’d handed a machine to a tool that only knew half of it.

That’s what half-automation really is.

Two sources of truth that don’t know about each other.

When everything’s manual, you know you’re on your own. You’re careful. You check.

When everything’s automated, the code is supposed to be the source of truth.

But half?

Half is two drivers in the same lane.

One on cruise control.

One who grabbed the wheel a few months ago and then wandered off.

Both convinced they own the road.

It’s fine for a long, long time.

Right up until the head-on collision.

Either automate it fully, or don’t automate it at all. Half is where the crash happens.

And when your automation breaks, the answer isn’t to quietly route around it by hand.

Fix it.

Or, at the very least, put what you did by hand back into the code.


After

I haven’t touched that VM with Terraform since.

Yes, I see the irony.

But the bigger problem wasn’t really the disk.

It was that one disk on one VM could take down every tenant at once.

So after that, we started planning to split the databases apart.

One bad apply shouldn’t be able to take out everyone.

And I picked up a new rule.

It’s not complicated.

It’s not clever.

Don’t touch production on a Friday.

Not for an upgrade.

Not for a quick fix.

Definitely not for “a bit of housekeeping.”

Definitely not while you’re on the phone with your parents.


The thing I actually learned

The plan said there was nothing to change.

I believed it.

That’s the problem.

A plan can only reason about what’s in the code.

It doesn’t know what you did in the UI on some random afternoon months ago.

It doesn’t know about the disk.

It doesn’t know about the shortcut.

It doesn’t know about the “I’ll fix it later.”

You do.

So if you’re going to automate something, automate all of it.

And if you’re going to do something by hand, own that too.

Write it down.

Put it in the code.

Tell the other driver.

Because the automation isn’t what hurts you.

It’s the part you did by hand and forgot to mention.

I got away with it.

An hour of downtime, a panicked call for help to the datacenter, an awkward message to my lead, and an incident report.

No data lost.


And it feels like

I’m at an all-time low.

Slightly bruised and broken.

From our head-on collision.