Convex reported this critical-impact incident on its official status page on Mar 4, 2026, 17:28 UTC. It was resolved after 1h.
Right now Convex is operational. Live Convex status →
Convex had downtime for some customers caused by an anomalous load spike that pushed an internal service into a load-shedding mode, which then triggered an unexpected panic in a caching library we were using. Specifically the load triggered two queue management algorithms: CoDel which proactively drops requests to keep queues small, and adaptive-LIFO which dequeues in reverse order to avoid wasting time on old requests. These are both rather subtle algorithms that large services use to avoid congestion collapse under high load or attack. The panic in the caching library was just a bug that depended on both these algorithms simultaneously. We've made some changes as a result of this incident but the key lesson is that services should try to avoid switching logical behavior during high load. When systems are stressed switching to infrequently-used codepaths can often make matters worse. We're now going to be proactively triggering CoDel and adaptive-LIFO at steady state to ensure that
Incident resolved. We're very sorry for the impact on your projects. Our team will be publishing a detailed postmortem soon.
We've identified the issue and remediated the problem. We're monitoring before we declare the all clear.
We are currently investigating an issue leading to elevated error rates on customer backends
Our 5-minute checks didn't record a change in Convex's overall status around this incident. Smaller or regional incidents often leave a provider's overall status green.
Overlapping incidents aren't necessarily related.
Convex reported 5 incidents in the last 90 days, 3 of them major or critical. A typical incident lasted 57 min. See Convex's uptime and incident history