ChatGPT and Claude both failed for some users on September 29, 2026. An IRI datacenter lesson on vendor finger-pointing, plus a 14-day fallback plan for SaaS products built on one model.

On September 29, 2026, ChatGPT and Claude both failed for some users on the same day. Your customer did not check a status page. They sat in your app and watched a spinner. I spent years in a datacenter where every vendor blamed the next one. The customer never cared whose fault it was.
Both major AI labs had service problems on the same day, and neither was a total blackout.
Forbes, citing Downdetector, said nearly 1,000 OpenAI reports landed around 2 p.m. Eastern. About 85% were ChatGPT use. OpenAI's status page said error rates were higher than usual. Some people could not log in or sign up. Some tasks did not finish. Codex and parts of the API had problems too. Reports dropped, then climbed again. It came and went.
Claude's Downdetector reports topped 11,000 around 10:15 a.m. Anthropic said errors hit the Claude website, Claude Code, and the API. It said service was back by the afternoon.
Two labs. One day. Intermittent. That last word matters. Intermittent is worse than down. Down, you can explain. Intermittent, your customer thinks your product is broken.
When nobody owns the whole stack, nobody owns the outage. I learned that at IRI.
I worked in a datacenter that ran IBM, HP, and Sun hardware at the same time. Each vendor had its own software stack on top. Each had its own tools, its own support line, its own way of failing.
On paper it looked like best of breed. In practice it was a mess. Patches fought each other. When something broke, the support call bounced from vendor to vendor. Each one pointed at another. Outages ran longer because nobody owned the full picture.
Here is what that taught me. The people who depend on the system do not care whose fault it is. They only know it is down and nobody can tell them when it comes back. Every hour of vendor finger-pointing is an hour of lost faith. Not in the vendor. In you.
The vendor's name is on the contract. Your name is on the customer's experience.
A SaaS product with one model and no fallback inherits the model's bad hour.
Your customer logs in. A task stops halfway. A report never finishes. They do not know you call OpenAI or Anthropic behind the screen. They only know your product failed them.
You did not cause the outage. It does not matter. You own the experience.
Founders whose product calls one model for the step the customer is waiting on feel it first.
Then the customer who had work in flight when the model stalled. The sales rep who was mid-demo. The analyst who needed the report before a meeting.
And then your support team, answering tickets with "it's our provider." That sentence sounds like an excuse. It is the IRI phone tree all over again.
Here is the brutal truth: vendor status is not your status.
You do not control OpenAI or Anthropic. You control three things. What the customer sees when the model stalls. What you say about it. Whether the work can finish another way.
Now a warning from the same datacenter. The fix is not to bolt on five more vendors and call it resilience. That is how IRI got its mess. Every vendor you add is another contract, another support line, another way to fail. Add a fallback only where the customer is waiting. Own the rest with clear words and a named human.
The company that can say what still works keeps the relationship. The company that shows a spinner does not.
Find the one customer step that dies when the model dies, then decide what happens instead.
That test is the step most teams skip. Do not skip it. A fallback you have never run is a hope, not a plan.
Because September 29 is still fresh for your customers. Some of them hit the spinner. Some of them asked support what happened. Right now they will notice if you fix it. In three months they will only notice the next outage.
Both had intermittent problems the same day. OpenAI reported higher error rates, login failures, and tasks that did not finish. Anthropic reported errors on the Claude website, Claude Code, and the API, and said service was back by the afternoon.
No. Both were intermittent. Some users had problems and some did not, and reports rose and fell during the day.
Only where a customer is waiting on the result. More vendors add more contracts and more ways to fail. Add a fallback where it protects the customer, not everywhere.
A plain message that says what is happening, whether the work is saved, and when it will finish. Then a named person or channel for updates.
The provider owns its outage. You own your customer's experience. Plan for both.
If your product has one model and no sentence for the bad hour, that is the gap. Start with the free audit at seriodesignfx.com/audit.
Want your fallback story told before the next outage, in your voice, on a schedule? That is what M.A.P. (Maverick Advantage Platform) does. It turns what you know into authority content, so customers trust you when the model blinks. Use the contact page to scope it.
I'm Charles K. Davis, Fractional CDO at SERIO Design FX, the team behind M.A.P. (Maverick Advantage Platform) and M.A.D. (Maverick Advantage Design).
P.S. This is for founders whose customer-facing step depends on one model. If the model is a side tool, skip this one.
M.A.D. Designs Your Brand. M.A.P. Makes You Known For It.
Stop Reading. Start Seeing.