Latest

Estimates Fail Because Teams Estimate the Wrong Thing

Most estimates describe how long the coding will take. Most overruns come from everything that surrounds the coding, which is exactly the part nobody put a number on.

/ 4 min read / Read article →

Latest Authority Takes

Straight-shooting analysis from the trenches

/ 4 min read

Every Dependency Is an Operational Commitment

Adding a library takes seconds and gets evaluated on features. The real cost arrives later, in upgrades, breakages, and the maintenance burden nobody scoped when the decision was made.

Read article
/ 5 min read

Your Error Messages Are Part of the Product

Teams treat error text as leftover work handled at the end of a ticket. Then support volume climbs, incident timelines stretch, and it turns out the system was communicating badly the whole time.

Read article
/ 4 min read

Configuration Is Where Systems Quietly Get Complicated

Nobody plans a configuration problem. It accumulates one environment variable at a time until deployments depend on knowledge that lives in someone's head instead of in the repository.

Read article
/ 6 min read

The Twelve Thousand Records I Did Not Delete

A junior mistake that turned out to be harmless still changed how I have written every deletion since. The instinct was right. The implementation took considerably longer to get right.

Read article
/ 3 min read

Runbooks Are Boring Until the Incident Belongs to You

Teams postpone runbooks because documentation feels secondary during calm periods. Then an incident lands on the wrong person at the wrong time and institutional memory turns out to be a very weak system.

Read article
/ 3 min read

Your API Docs Are a Reflection of Engineering Discipline

Documentation quality rarely fails because nobody had time to write. It usually fails because the team does not agree clearly enough on what the contract actually is.

Read article
/ 3 min read

AI-Generated Refactors Are Where False Confidence Gets Expensive

Refactors already carry hidden risk. AI makes it easier to perform larger ones faster, which is exactly why teams need more caution instead of less.

Read article
/ 3 min read

Cross-Functional Ownership Sounds Great Until Nobody Owns Production

Shared responsibility can improve collaboration. It can also become the sentence teams use when accountability is too blurry to survive an incident cleanly.

Read article
/ 3 min read

Your On-Call Rotation Is Telling You the Truth About Your Architecture

The incident pattern your team keeps normalizing is usually a design signal. On-call pain is one of the clearest ways a system reveals where its architecture is actually weak.

Read article
/ 3 min read

Your Test Suite Does Not Need More Tests. It Needs More Trust.

Teams keep adding tests to fix anxiety when the real problem is that nobody believes the suite is telling the truth about production risk.

Read article
/ 3 min read

The Worst Time to Design Permissions Is After You Land an Enterprise Customer

Teams love postponing access-control design until a big customer forces the issue. By then the system already has assumptions baked into it that are painful to unwind.

Read article
/ 3 min read

Retries Are Not Reliability

Repeating a failing action can be useful. It can also multiply load, duplicate side effects, and hide the fact that the system was never designed to fail safely.

Read article