Articles
Field notes from fourteen years of running infrastructure.
A small number of pieces, kept current. Each one answers a question I actually get asked, with the tradeoffs stated and the claims checkable.
-
AWS built a deploy mode for AI agents. Build the gates first.
CloudFormation's new express mode is made for agents: up to 4x faster, rollback off by default. What that says about where IaC is going in 2026.
-
Agent governance can't wait for the standards
NIST's AI RMF, ISO 42001, and the EU AI Act were all written before tool-calling agents, and the CSA says its own standards land through 2027. Build now.
-
Your AI agents are sharing an API key
In a 900-respondent survey, 45.6% of organizations authenticate agent-to-agent traffic with shared API keys. The identity work that has to precede autonomy.
-
Your AI budget is written in tokens
AI spend doesn't behave like cloud spend: it's token-denominated, usage-driven, and volatile by design. How to forecast it before the CFO asks you to.
-
Cloud waste rose for the first time in five years
Flexera's 2026 survey has estimated cloud waste climbing to 29% of IaaS/PaaS spend after five years of decline, with AI in the blame line. What reverses it.
-
The EKS upgrade playbook
Kubernetes 1.30 gets auto-upgraded off EKS this month, and 1.33 starts billing extended support in three weeks. The playbook, with every claim cited.
-
Proving AI value is an infrastructure problem
Gartner calls 2026 the inflection year for AI spend while only 28% of finance leaders see clear value. The missing piece is metering, not a better narrative.
-
FinOps for AI: costing the workload before it costs you
AI compute is the fastest-growing line on the cloud bill and the hardest to attribute. What FinOps for AI actually takes, from someone holding ~$4M flat.
-
Five signs your cloud has outgrown your team
The recurring symptoms that send teams looking for outside cloud help, what each one really means underneath, and the ones you can fix yourself.
-
Stabilizing a Kubernetes cluster that pages
The cluster waking you up is rarely a Kubernetes problem. What actually stabilizes EKS in practice: resource discipline, node lifecycle, and GitOps.
-
SOC 2 is a sales tool that happens to improve your security
Startups don't buy SOC 2 for security. They buy it because an enterprise deal won't close without it. What the readiness work actually is, from the infra side.
-
What an AI-ready platform actually has in it
AI readiness is a platform you build in layers, not a model you buy. The data, integration, and control layers that separate a shipping program from a stalled one.
-
When the cloud stops being the cheap option
Repatriation is real, narrow, and mostly about AI compute. When owning hardware actually beats renting it, and the discipline to tell the difference.
-
Why AI pilots stall before production
Most AI pilots don't fail on model quality. They die in the gap between the demo and production, and that gap is infrastructure. What actually stalls them.
-
The attribution problem, or why "just tag everything" fails
Tags are the standard answer to AWS cost allocation and they've never once been enough on their own. What a workable attribution layer actually looks like.
-
The state of CI/CD in 2026
GitHub Actions won by default, Jenkins won't die, and most pipelines are still slower and leakier than they need to be. An assessment from someone who runs this stuff.
-
The state of infrastructure as code in 2026
Terraform after the license change, what OpenTofu did and didn't become, why CloudFormation still runs more production than anyone admits, and what I'd pick today.
-
Your AWS bill is a design document
Cost work is architecture review with a forcing function. Opening a series on cloud spend with the method behind a $500K-to-$200K bill and ~$4M/yr held flat.