Cloud, DevOps & infrastructure
How I think about cloud architecture, platform teams, and DevOps in regulated, high-growth environments.
Pillars
Multi-cloud landing zones
Account/subscription structure, SCPs, paved roads, and the boring guardrails that pass audits cheaply.
CI/CD & GitOps
Pipelines as product, signed builds, progressive delivery, and rollback you can trust at 2am.
Platform engineering
Internal platforms run like real products: roadmap, NPS, paved roads, tiered support.
DevSecOps
Shift-left controls that don't slow shipping: SAST, SBOMs, secret scanning, policy-as-code.
Reliability & SRE
SLOs that match revenue, error budgets engineers actually respect, on-call you can sustain.
FinOps & cost
Tagging that holds up, unit economics per workload, rightsizing as a recurring practice.
Reference stack
- AWS (primary)
- Azure
- EKS / Kubernetes
- ECS Fargate
- Lambda
- GitHub Actions
- ArgoCD
- Terraform
- Helm
- OpenTofu
- Datadog
- OpenTelemetry
- Grafana
- PagerDuty
- Sentry
- Wiz
- Snyk
- OPA / Gatekeeper
- Vault
- AWS Security Hub
Related writing
All cloud & DevOps postsThe Second Agent Cheap, the Fiftieth Boring
Build cost was never the constraint in a regulated shop: agent #50 hits the same security review as agent #1. Make identity and action gates reusable.
AWS Cost Levers That Moved the Needle
Cutting ~35% off a multi-region AWS footprint with no capability loss: the levers in the order they paid back, best first.
Security and DevOps Under One Roof
The case for running security and DevOps as one mandate: org-chart distance doesn't create security, and owning the pipelines changes how you protect them.
Model Selection Is Capacity Planning
Most teams pick a model like a sports team and never revisit it, but model selection is a routing, capacity, and risk decision you already know how to make.
Design AI Inference for Model Disappearance
A frontier model went dark three days after launch; here's how I make AI inference survivable on AWS when the provider is a dependency you don't control.
The Eight-Domain Azure Security Review
A tool scores your Azure posture; an assessor walks your architecture. The eight domains I review, in audit order, and the evidence each has to produce.
Cloud FinOps: Where 25-40% of Spend Hides
The press-release version of cloud savings cancels workloads and books compliance debt. The durable version is commitment management and SaaS rationalization.
An SBOM Nobody Reads Is Compliance Cosplay
Generating a software bill of materials is the easy part. Wiring it into the moment a change ships is where supply-chain security stops being theater.
Agent Memory Is a Data-Residency Problem
Give every agent a durable, MCP-connected brain and you've stood up a new data lake of PII and PCI scope nobody classified, encrypted, or can purge.
Agent Safety: Engineer the Blast Radius
Most agent "safety" is a politely worded request to a model that need not honor it. The only controls that count still hold after the model goes wrong.
An AI Agent Dropped Prod: The Change Playbook
Coding agents are committing real change to real systems. The question isn't whether to let them. It's how to give them speed without a SOC 2-fatal mistake.
Source-Map Leaks: Your Pipeline's Confession
One packaging mistake can publish hundreds of thousands of lines of internals. The leak is a confession: release controls never caught up to release velocity.
AI Found 271 Bugs in Firefox. Now Your Repos?
AI-assisted fuzzing found hundreds of bugs in hardened open-source code. The question is whether you run it before someone else runs it against you.
Dark Code Is a Control Failure, Not Tech Debt
AI is filling repos with code nobody can explain. We call it tech debt; it's a control failure, and it should fail CI like a missing approver does.
MCP Is a New Attack Surface: An IAM Playbook
Every MCP server is a new identity reaching into your cloud. Whether that's leverage or liability comes down to least-privilege IAM on every tool call.
Fine-Grained Authorization for Fintech APIs
Authorization scattered across your codebase isn't a feature. It's a liability you can't prove. The pattern multi-tenant regulated platforms actually need.
Aurora DSQL for the Ledger: Active-Active
Multi-region active-active sounds like the answer to ledger nightmares. Interrogate the consistency, recovery math, and migration before betting the books.
Zero-Downtime Database Changes Are a Process
Blue/green and serverless Aurora don't make migrations safe. The runbook does. The boring discipline that keeps schema changes from becoming incidents.
Three Token Counts, Zero You Can Attest To
Codex says one number, Claude another, your gateway a third. That isn't a metering problem. It's an attestation problem regulated industries can't afford.
Cluster Autoscaler to Karpenter: What Breaks
Karpenter is the right call for most EKS shops, but the migration breaks things unrelated to autoscaling. What to know before flipping the switch.
Shadow AI Is the New Shadow IT
Every abandoned notebook and weekend prototype is a credential-bearing asset nobody owns. The fix isn't a ban. It's discovery, demotion, and real sunsets.
Guardrails at Scale for a Three-Person Team
A lean team can govern a sprawling cloud estate without becoming a ticket queue, but only if you put the rules in the pipeline, not in your inbox.
Ransomware Recovery: A Tested-Backups Problem
Everyone has backups. Almost nobody has a restore they've actually run under fire. That gap is where ransomware turns a bad week into an existential one.
Building or rebuilding a cloud platform?
I advise fintech and regulated SaaS teams on cloud architecture, DevOps maturity, and platform engineering.