Field Notes
Incidents, diagnosed
Things that broke, why they broke, and what actually fixed them. Open one to watch the failing path stall and then clear.
- ▸connection reset by peer, immediately after enabling STRICT mTLSservice meshSEV-2
- ▸APM init container stuck in ImagePullBackOff across several servicesobservabilitySEV-3
- ▸Dashboards green, services listed as onboarded, no data arrivingobservabilitySEV-3
- ▸Trace ingestion stopped for every tenant at onceshared infraSEV-1
- ▸401 from a private package registry, but only inside CIci/cdSEV-3
- ▸Pipeline step aborts with exit 2 and no error messagepipelinesSEV-4
- ▸Spot VM reclamation taking workloads down during cost optimisationcost / availabilitySEV-2
- ▸load balancer returns a 502 error page while every pod reads Running and Readyload balancingSEV-2
- ▸pod is Ready and answering locally, but the load balancer calls it unhealthy and serves 503networkingSEV-2
- ▸canary promotion times out waiting for a pause that the rollout had already reached onceprogressive deliverySEV-1
- ▸app is permanently OutOfSync on one resource while every sync reports successgitopsSEV-2
- ▸a fresh deploy crashloops on a startup probe that the new code definitely implementsrelease engineeringSEV-2
- ▸intermittent unknown-host errors on an external name that resolves perfectly when you test it by handdnsSEV-3
- ▸builds start failing with 429 Too Many Requests from a public package registry, with no change to the codebuild infrastructureSEV-3
- ▸build agents die mid-job with lost-contact and process-spawn errors that name the wrong component entirelybuild infrastructureSEV-3
- ▸users cannot join voice sessions, while pod restart counts sit at zero and load metrics look calmrealtime mediaSEV-1