Mega2580.Solutions
Blog

Notes from ops work

Short, practical write-ups since 2018. No listicle padding.

2026

August 2026 · D. Reyes

Five years of infrastructure reviews: patterns we keep seeing

After a few hundred reviews, the same handful of issues keep showing up regardless of stack or team size.

April 2026 · A. Tanaka

Monitoring that doesn't lie to you

Most alerting isn't broken because it's under-built. It's broken because nobody ever removes anything.

March 2026 · D. Reyes

Five signs your infrastructure needs a health check

The warning signs are usually visible months before something actually breaks.

January 2026 · M. Kowalski

Why we automate every deploy, even for small teams

Manual deploys feel fine right up until the one time they don't.

2025

August 2025 · M. Kowalski

Secrets management for small teams

You don't need a full vault deployment on day one. Here's what actually matters first.

March 2025 · D. Reyes

Observability vs. monitoring: a distinction worth making

They get used interchangeably, but they answer different questions.

2024

June 2024 · A. Tanaka

Moving off a legacy load balancer, without a maintenance window

How we replaced a decade-old load balancer setup while it kept serving traffic the whole time.

January 2024 · M. Kowalski

Terraform state files: the mistakes we've made

A few state-related incidents from our own history, and what we changed afterward.

2023

July 2023 · D. Reyes

What a good incident postmortem actually looks like

Blameless doesn't mean toothless. Here's the structure we actually use.

February 2023 · M. Kowalski

Cost optimization without breaking things

The order we work through a cost review in, so nothing gets cut that turns out to matter.

2022

September 2022 · A. Tanaka

On-call burnout is a system design problem

It's rarely fixed by asking people to be more resilient. It's fixed by paging them less.

March 2022 · D. Reyes

Infrastructure as code, three years in

What we'd tell ourselves at the start, now that we've maintained these codebases for a while.

2021

October 2021 · A. Tanaka

Zero-downtime database migrations, a checklist

The concrete steps we run through before any schema change that touches a live table.

April 2021 · M. Kowalski

Kubernetes: when it's worth it, and when it isn't

A framework for deciding, instead of defaulting to yes or no.

2020

August 2020 · A. Tanaka

Running incident response with a fully distributed team

What changed for us this year, and what we had to rebuild once nobody was in the same room during an outage.

February 2020 · D. Reyes

The case for boring technology

Why we default to the well-worn option, and when we don't.

2019

September 2019 · M. Kowalski

Containers didn't fix our deploy problems, process did

Containerizing a bad deploy process just gives you a bad deploy process in a container.

March 2019 · D. Reyes

What we look for in a first infrastructure review

The specific things we check in the first few days of any new engagement, before we recommend anything.

2018

November 2018 · A. Tanaka

The restore test nobody runs

A backup you haven't restored from is a hypothesis, not a backup. Here's how we actually verify them.

June 2018 · D. Reyes

Why we started Mega2580 Solutions

The problem we kept seeing at every company we worked at, and why we decided to build a small consultancy around fixing it.