← Back to the discography Feeling This

Playing By Ear

Why moving to GitOps was the best infrastructure decision I’ve made: scattered Docker hosts, on-prem Kubernetes, a bash script talking to Portainer, and the day I stopped being the source of truth.


Before GitOps, our deployments were a bit like a band playing by ear.

Everyone kind of knew the song.

Nobody had the sheet music.

And if you wanted to know how it actually went, you had to ask me.


How it used to work

At first, our apps lived on Docker.

The docker-compose.yml sat in the same repo as the application code.

The pipeline built the image.

We usually had two of them: a develop image and a main image.

Then a script, triggered from CI/CD, told Portainer to redeploy.

Green pipeline.

Done.

Next.

It worked.

Until I needed to change one environment variable.

One.

And to do that, I had to run the whole pipeline again.

Build the image.

Push the image.

Trigger the script.

Redeploy.

All of that for a single line of config that had nothing to do with the code.

And that’s if everything went right.

Now imagine a developer messages me:

“Hey, can you add these two environment variables?”

I add one.

Run the pipeline.

Build. Push. Trigger. Redeploy.

Wait.

Green.

Done.

A little later:

“Hey, the second one is still missing.”

Right.

Back to the pipeline.

Build the same image again.

Push it again.

Trigger the script again.

Redeploy again.

And if I made a typo in the value?

Same thing.

Again.

The code hadn’t changed once.

The image was identical every single time.


Then came Kubernetes

We had two on-prem clusters, dev and prod.

And somewhere around them, 10 to 20 Docker hosts scattered across the infrastructure, each doing its own thing.

On the Kubernetes side, the pipeline ran helm directly.

Which meant the pipeline needed cluster credentials.

Which meant the kubeconfig lived in GitLab CI/CD variables.

So our CI system could reach straight into the cluster and do whatever it liked.


Where was the truth?

This was the real problem.

Not the scripts.

Not Portainer.

Not Helm.

The question was simple:

“What is actually supposed to be running right now?”

And the honest answer was:

It depends who you ask.

The cluster had one version.

Some wiki page had another.

The compose file in the repo had a third.

And the rest of it lived in my head.

Anyone with access could change things.

Sometimes someone hotfixed something directly.

To be fair, drift wasn’t a huge problem yet, because we were still early with Kubernetes. There wasn’t much to drift.

But the configuration was a mess.

And onboarding a new service took a lot of work: copying files, writing scripts, wiring up credentials, remembering which wiki page was the correct one.

We weren’t reading from the sheet music.

We were all playing by ear.


Fleet first

Funny thing is, GitOps didn’t start with my own projects.

It started with a message from another team, on January 29, 2024.

I’d already used Kubernetes the way I described above: kubectl, helm upgrade straight from the pipeline, credentials in GitLab.

This team was about to move a customer from Docker Swarm to Kubernetes, on a fresh on-prem setup with Rancher, RKE2 and Ceph.

And then this line:

“What’s important to him is that all deployment processes follow gitops principles and are well documented to make debugging easier.”

They also wanted a separate infra repo for all the configs not directly related to the application itself: reverse proxy, pipelines, the lot.

A separate infra repo.

I didn’t even have that concept yet.

Honestly, they wanted more than what I gave them.

I was lazy.

My knowledge was limited.

And AI wasn’t really a thing yet, so I couldn’t just ask it to explain the hard parts to me.

So I didn’t go looking for the best tool.

I went looking for the simplest one.

That was Rancher Fleet.

It came built into Rancher, and it was easy to set up without a lot of work.

At least, I thought so.

Because it didn’t work.

The repo was there.

Fleet was there.

Nothing deployed.

Meanwhile, the developer on that team just wanted to ship to dev:

“If you give me a quick introduction on how to update the deployment on dev, I can also do it.”

“Or do you think the pipelines will be ready by then?”

I didn’t think so.

So I offered the fallback.

Connect Azure DevOps to the cluster.

Run helm upgrade from the pipeline.

Yes.

The exact thing the customer had asked us not to do.

Then came the fair question:

“But what exactly is your issue with the GitOps approach again? Is it that the Azure repos don’t work but the GitLab ones do?”

Honestly?

I didn’t fully know.

So I took the same repo, the same process, and pointed another cluster at it.

It worked.

So the problem wasn’t my repo.

It wasn’t Azure.

It wasn’t GitOps.

It was the Fleet version in Rancher.

All that time, and the fix was an upgrade.

That’s where it all started.

After that, I brought it to the rest of our projects.

No big-bang migration.

One project at a time, based on priority.


Then Argo CD

I stayed on Fleet for a little over two years.

What finally moved me wasn’t a blog post or a comparison table.

It was NDC Sydney, in April 2026.

I met the Octopus Deploy team there, and they were all in on Argo.

So I tried it.

And I’ll be honest about what hooked me.

It wasn’t some deep architectural reason.

It was the deployment notifications.

That’s it. That was the hook.

But once I was in, I started to understand what I’d actually been missing.

Configuration drift. Argo shows you when the cluster doesn’t match Git, which Fleet, at least the way I was using it, never really made obvious.

Templating. The way Argo templates applications is one of the best things about it.

And then there was ApplicationSet.

More on that one in a minute.


What it looks like now

The manifests live in a separate repo.

One repo per project, a small monorepo holding the manifests and infrastructure definitions that project needs to run.

It evolved over time:

compose → Helm → infrastructure → and gradually, catalogs.

Dev and prod are just folders.

Mostly Helm. Some raw YAML. Kustomize has started showing up recently.

The flow is simple:

CI builds the image and runs the checks.

A deployment stage commits the new image tag to the manifest repo.

Argo sees the change and syncs it.

No kubeconfig in GitLab.

No script talking to Portainer.

No pipeline reaching into the cluster.

The pipeline writes to Git.

Argo reads from Git.

The cluster follows.

And no, we don’t “promote” images from dev to prod.

Dev and prod can run genuinely different code, and some frontends bake environment variables in at build time, so the same image can’t simply move forward.

I’m still looking at whether we can make that work.

But I’m not going to pretend we’re doing it just because a conference talk says we should.


Things I had to unlearn

Moving to GitOps wasn’t just new tools.

It was breaking a lot of old habits.

kubectl edit doesn’t work anymore.

One day I wanted to quickly test something.

Temporarily change an image tag.

So I did what I’d always done:

kubectl edit deployment <app>

Changed the tag.

Saved.

Felt productive.

A minute later, Argo reverted it.

WTF.

Then it clicked.

Argo wasn’t broken.

It was doing exactly what I’d asked it to do: make the cluster look like Git.

I’d just never had anything that actually held me to that before.

The code repo isn’t the only repo.

Before, all I knew was that whatever an app needed should live in the app’s repo.

Code, compose file, scripts, config. All of it.

But the code and the configuration that runs the code are not the same thing.

They change at different speeds.

They’re changed by different people.

They break for different reasons.

A separate repo for manifests sounded like extra work.

It turned out to be the whole point.

Stop answering from memory.

Before, I knew where things were.

That felt useful.

It was actually the problem.

And it’s a hard habit to break, because answering from memory is faster.

Every single time.

So now when someone asks, I have to stop myself.

Open the repo.

Find the line.

Send the link:

“The answer you want is on this line.”

Slower for me today.

Faster for everyone after.

Stop building around Portainer.

I was so obsessed with Portainer that I ended up writing Ansible playbooks for Portainer.

They created stacks.

They redeployed the stacks.

And all the credentials were in the repo itself, encrypted with Ansible Vault.

At the time, it felt like engineering.

Looking back, I had built a GitOps tool by hand.

A worse one.


The gains

I can see everything.

One glance shows the whole project.

What’s deployed.

Which version.

Whether it matches Git.

Remember that one environment variable that needed a full rebuild?

Now it’s a one-line commit.

ApplicationSet.

We run a multi-tenant setup.

Before, if something changed, I had to run the same helm command once for every tenant.

And keep the tenant list somewhere.

Usually a mix of the pipeline, my notes and my memory.

Now the tenant list is in Git.

One trigger, and every tenant gets deployed.

Want to bump the chart version?

I don’t open each tenant’s Argo file anymore.

I change one line.

Every tenant gets it.

Image Updater.

This one I never found in Fleet.

For one add-on repo, CI doesn’t commit the image tag at all.

It just builds the image and pushes it.

Argo CD Image Updater watches the registry, sees the new image and updates the deployment itself.

No tag commit.

No extra pipeline stage.

Push the image, and it shows up.

Anyone can deploy.

Developers can update ConfigMaps and check logs themselves.

They don’t have to wait for me.


The tradeoffs

It’s not free.

It’s slower.

Before, a green pipeline was the deployment. No waiting.

Now there’s roughly a 4–5 minute gap while Argo picks up the change and syncs it.

When you want something changed right now, those minutes feel long.

It hides the deployment.

This is the one I didn’t see coming.

Developers see the deployment stage go green and assume it’s deployed.

But the green tick only means the tag was committed to Git.

The sheet music changed.

The band hasn’t played it yet.

It’s not deployed until Argo says the app is healthy and the pods are actually running.

To be fair, I had to learn that one too. On Docker, if the container was running, I called it healthy.

GitOps made deploying so easy that a lot of people stopped wondering how it actually happens.

And honestly, most people want to know how it works.

Most of them just have no idea yet.

That one’s on me to fix.


Secrets: the bit I’m still fixing

Secrets didn’t move into Git.

For a long time, they were applied by hand, with kubectl apply.

And I paid for that.

When I rebuilt the dev cluster, we were still on Fleet. Re-adding the projects to Fleet was easy, and the workloads came back.

The secrets didn’t.

Finding them, rebuilding them, working out which values were current: that was the painful part.

What was in Git came back easily.

What wasn’t was a hunt.

That says more about GitOps than anything else in this post.

So now we’re slowly moving to Infisical.

Partly because developers want to see the current value without asking me.

Partly because I’m too lazy to log in and check for them.

And mostly because I don’t want to be the one person who has to be around every time a credential changes.


What’s next

Once the configuration lives in Git, the next step gets much easier.

Backstage templates.

An internal developer platform.

And, honestly, the fun part: being able to ask an AI to build a manifest based on our company standard, and getting back something that actually fits.

That only works because the standard exists.

And the standard only exists because it’s written down.


Was it worth it?

So far, it’s the best infrastructure decision I’ve made.

Not because Argo CD is magic.

Not because GitOps is fashionable.

Because for the first time, there’s one answer to:

“What is supposed to be running?”

And the answer isn’t me.

It’s in Git.

We used to be a band playing by ear.

Everyone knew roughly how the song went.

Nobody could tell you exactly.

Now there’s sheet music.

Git is the sheet music. Argo is the one who notices when someone plays the wrong note.

And I finally get to stop humming the tune for everyone.


I got no regret right now.

I’m feeling this.