Waterfall service incident post mortem — September 29, 2026
Resolved
Sep 29, 2026 at 7:21pm UTC
Waterfall service incident — September 29, 2026
What happened
On September 29, Waterfall experienced a widespread service disruption during a production software update. A shared dependency was updated before the matching application code was deployed. This incompatibility prevented affected services from starting and processing requests.
The failure was caused by our deployment process. We apologize for the disruption to your work.
Customer impact
All API endpoints were affected. Requests could fail, and workflows depending on them could be interrupted.
Production logs show the dependency-related errors from approximately 17:36 to 17:58 UTC. Recovery was rolled out progressively, so the impact window varied by service. This interval describes the errors observed in our logs; the precise customer impact and any outstanding job recovery are still being reviewed.
Timeline
| Time on September 29 (UTC) | Event |
|---|---|
| 17:04–17:20 | The update sequence was deployed to our development environment. Manual end-to-end checks were performed after the application deployment. |
| 17:36–17:37 | The production dependency rollout began directing traffic to incompatible application versions. Service startup failures appeared. |
| During the rollout; exact time under review | We detected the outage on operational dashboards, followed shortly afterward by an alert from our external monitoring service. |
| 17:49 | A production deployment of compatible application code was initiated. |
| 17:52–17:57 | Updated application versions were progressively activated. The deployment finished at 17:57. |
| 17:58 | The last matching dependency error in the reviewed logs came from an earlier application version. No further matching errors were found through the 18:20 review boundary. |
Why testing did not catch it
Our development checks tested the final combination of application code and dependencies. They did not test the intermediate state created when the dependency update went live before the application update.
Our deployment automation also treated successful infrastructure updates as successful deployments without verifying that each affected service could start and process requests. That allowed an incompatible version to reach production traffic.
How we addressed it
We deployed the application code compatible with the updated dependency and progressively moved services to those versions. We are reviewing the remaining customer-impact details, including whether interrupted requests or jobs require further action.
Preventing recurrence
We have identified the following improvements for implementation:
- Validate the complete application and dependency combination before directing production traffic to it.
- Test service availability throughout the deployment sequence, including intermediate transitions and asynchronous processing.
- Roll out changes in smaller stages, with health checks that stop expansion when a service fails.
- Strengthen startup-failure monitoring and rehearse rapid restoration of the previous working release.
These are proposed preventive actions; their implementation and verification status will be tracked separately. Our responsibility is to make the deployment process catch this class of incompatibility before it affects customers.
Affected services