Work
The technical half
Case studies, the systems behind them in 3D, the incidents they produced, and analysis of other people's outages. All of it cross-linked — a case study links to its architecture and its incidents, and back again.
- CLOUD COST-40%Off a production bill, without quietly trading away availability.
- MANUAL TOIL-60%Tickets became a self-service portal with golden paths.
- SLA UPTIME99.9%Sustained on 24/7 on-call. Here is the alerting that made it boring.
- FIELD NOTESdiagnosedReal diagnoses, including the red herrings that cost me hours.
- ARCHITECTUREin 3DProduction architectures you can orbit, zoom and take apart.
- CI/CD + SECshift-leftScans, coverage and provenance as merge gates, not a later email.
- RESILIENCYHA failoverCheap capacity is only safe if losing it is a non-event.
- IaCGitOpsTerraform and Argo CD — peer-reviewed, audit-ready, no console edits.
- SERVICE MESHIstiomTLS, policy and traffic shifting. The upstream docs
- DELIVERYArgo CDDeclarative GitOps, and the progressive delivery docs
- POSTMORTEMSstudiedOther people's outages, studied properly. Not mine — theirs.
- COFFEE AT STAKEon meChess, badminton or basketball. Beat me and I am buying.
- Case studies10
10 pieces of work, written up
What the problem was, the order things had to happen in, the commands, and what the numbers actually moved.
- ·Cloud Cost Reduction Programme
- ·CI Quality Gate at Scale
- ·Self-Service Infrastructure Provisioning
- ·RAG Knowledge Assistant for Ops
OPEN ↗
- Architecture9
9 systems, rendered in 3D
Orbit, zoom and pan around the gate, the mesh, the self-healing loop and the platform.
- ·Production Security
- ·PR Quality Gate
- ·Zero-Trust Mesh
- ·Self-Healing Loop
- ·Self-Service Platform
- ·Edge to Pod
- ·Progressive Delivery
- ·The Build Fleet
- ·Secret Delivery Path
OPEN ↗
- Field notes16
16 incidents and their fixes
What broke, why it broke, and what actually fixed it — each one rendered failing, then recovering.
- ·connection reset by peer, immediately after enabling STRICT mTLS
- ·APM init container stuck in ImagePullBackOff across several services
- ·Dashboards green, services listed as onboarded, no data arriving
- ·Trace ingestion stopped for every tenant at once
OPEN ↗
- Postmortems2
Public outages, read closely
Analysis of published incident reports from companies operating at a scale most of us only read about.
- ·Roblox — 73 hours
- ·Cloudflare — 27 minutes
OPEN ↗