Patterns from the floor

What a decade inside hyperscale infrastructure actually teaches you.

Not client case studies — I hold my engagements in confidence, and my former employers' in confidence too. These are the failure patterns I watched repeat at scale, stated generally, because they're the ones I now get hired to prevent.

Pattern 01 · The orphaned system

The thing nobody owns is the thing that takes you down.

At scale, the systems behind the worst incidents are rarely the ones under active ownership. They're the ones that quietly outlived the team that built them — still running, still load-bearing, with no name next to them on a page anywhere.

AI made this worse and faster. A team can stand up an agent in an afternoon on a corporate card. Eighteen months later it's embedded in a workflow, nobody remembers who configured it, and the person who did has left. The tool didn't go through procurement, so it isn't in the vendor register. It didn't touch production, so it isn't in the CMDB. It works, so nobody asks.

Then it stops working, or it says something it shouldn't, and the question "who owns this" gets asked for the first time under pressure.

Why it matters to you: ownership isn't an org-chart question, it's an operational one. Three different people usually own the spend, the outcome, and the failure — and they often don't know about each other. Finding that out during an incident is the expensive way.

Pattern 02 · The economics nobody re-ran

The build-vs-buy decision was right — two years ago.

Infrastructure economics move faster than procurement cycles. A decision that was correct when it was made goes expensive quietly, because nobody re-runs the math after the contract is signed. The break-even on running inference yourself, in particular, has moved meaningfully — and it will move again.

Vendors won't re-run that math for you. Their models are built to produce one answer, and each has a build bias or a rent bias baked in depending on what they sell. That isn't dishonesty; it's just what a model built by an interested party does.

The organizations that stay ahead of this treat the economics as a standing review, not a one-time decision — the same way you'd revisit a capacity plan.

Why it matters to you: workload first, math second, procurement third. Anyone who reverses that order is selling, not advising.

Pattern 03 · Governance that never touched production

A policy nobody can enforce is a liability, not a control.

The most common governance failure isn't the absence of a policy — it's a well-written policy with no technical mechanism behind it. It passes the audit. It survives the board deck. It changes nothing about what actually runs.

Worse, it creates false comfort. Leadership believes the question is handled because a document exists, so nobody funds the enforcement that would make it true. The gap only surfaces when someone asks for evidence rather than intent.

Real control looks like operational procedure: enforcement at the point of deployment, ownership recorded where the system lives, and a change process engineering will actually follow because it doesn't slow them down. Governance that fights velocity loses to velocity, every time.

Why it matters to you: if your AI governance can't tell you what's running right now, it isn't governance. It's documentation.

"The market moved from adopt AI to prove AI. The leaders who win the next 18 months won't be the ones who spent the most — they'll be the ones who can show what it returned and stand behind what it does."
— From my working framework on AI value & accountability.
The framework

These three patterns are the short version.

The longer one is a working paper on AI value and accountability — the decisions I walk clients through, published as I write them. Ask and I'll send it.