Ask HN: Why were OpenAI, Claude, and Grok simultaneously down? (self)

404 pointsby halcdev6 days ago705 comments
https://status.openai.com

https://status.claude.com

https://status.x.ai

ChatGPT outage – Resolved - https://news.ycombinator.com/item?id=49550614 (315 comments)

Claude outage – Resolved - https://news.ycombinator.com/item?id=49549676 (146 comments)

Grok outage - https://news.ycombinator.com/item?id=49551589 (142 comments)

OfficialTurkey 5 days ago

I work at OpenAI and I was the Incident Commander for yesterday's outage.

We had a routing error within our infra that caused issues for some of our products. It was not related to the Astra launch. We don't comment on other providers' outages.

dboreham 5 days ago

Straight from the turkey's beak.

swader999 5 days ago

Did you negotiate extra hard for the 'Incident Commander' title? I'm a bit jealous to be honest.

eli 5 days ago

It's actually a standard term for the person who plays this role during incident response https://www.pagerduty.com/resources/incident-management-resp...

OfficialTurkey 5 days ago

Ha. It's a role/title for the lifetime of the incident -- it's useful to have someone to keep things moving, keep track of workstreams, and to know who the decision-maker is, especially for bigger incidents. I'm just a SWE who works on infrastructure.

swader999 5 days ago

Hopefully you get a decked out command center too.

stillpointlab 4 days ago

Note: Incident Command is almost certain an allusion (or implementation) of the Incident Command System [1]

It is a common system in all kinds of emergency response scenarios, including local emergency services (fire/police/ambulance) and it scales all the way to massive disasters.

It is especially useful to clarify command structures when multiple response entities need to coordinate. That is true even within organizations like public companies, where the reporting structures may be distinct.

1. https://en.wikipedia.org/wiki/Incident_Command_System

alanfranz 4 days ago

It’s a standard term for a role during an incident.

wnmurphy 5 days ago

Curious then as to why the incident coincided with similar issues with Anthropic and xAI.

wnmurphy 5 days ago

Actually, nevermind... this looks more like Anthropic and xAI coincided due to shared xAI infra (6:23am and 6:30am), and OpenAI's issue was more likely then a coincidence (7:43am).

reilly3000 5 days ago

Thank you. This should be pinned or something.

wnmurphy 5 days ago

xAI mentioned a Memphis outage, which is where Anthropic was leasing capacity on Colossus 1. I do think this is the explanation.

ncb_asj 5 days ago

Shouldn't Astra be the Incident Commander?

sshine 4 days ago

Secretly, they are.

honeybadger1 4 days ago

At the same time the outages happened I noticed cloudflare also had reported issues with HTTP/3 and WARP. Related?

kibae 6 days ago

Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers.

https://downdetector.com/status/cloudflare/

https://downdetector.com/status/windows-azure/

https://downdetector.com/status/aws-amazon-web-services/

https://downdetector.com/status/google-cloud/

KptMarchewa 6 days ago

Not really. The impact isn't as big too - Codex for example did not stop working for me.

https://updog.ai/

taytus 6 days ago

oh well, if it is working for you, then we are saved.

personjerry 6 days ago

what's updog?

stirfish 6 days ago

:)

red-iron-pine 6 days ago

gottem!

y-c-o-m-b 6 days ago

Cloudfare did release an update for their "HTTP/3 issue affecting R2 custom domains" around that time

https://www.cloudflarestatus.com/history?type=incident

oersted 6 days ago

“load bearing” :)

For once it’s appropriately used.

quotemstr 6 days ago

It's a good metaphor and I refuse to let AI ruin it for me.

wadayano 6 days ago

But have we found the seams yet though?

rl3 6 days ago

If the seams bear too much load, they rip. Whereas pants, they fall down.

I'll be honest with you: this is why we need to take a belt-and-suspenders approach.

pborenstein 6 days ago

That's not just an observation, it's an insight. Words are doing the real work.

HarHarVeryFunny 6 days ago

This just makes me angry!

I really wonder if they can fix it. Fable 5.1 claims to speak humanese, but we'll see.

I tend to think this wierd limited vocabulary/style they use is an unwanted side effect of all the the RL training, perhaps also of being trained on their own synthetic content over multiple training cycles.

cootsnuck 6 days ago

Yea I don't get "load bearing" that much but "seams"... So sick of it.

lubujackson 6 days ago

We need a Samwise "Bear the load!" meme

BrokenCogs 6 days ago

https://i.imgur.com/S6aEhrI.png

hirako2000 6 days ago

Content not viewable in your region.

BrokenCogs 6 days ago

:(

MadameMinty 6 days ago

Wait, why would it ruin it?

imwally 6 days ago

It’s a frequently used metaphor in LLM responses.

bornfreddy 6 days ago

Often for trivial things that LLM is proud that it has noticed but bear no load whatsoever.

ipsod 6 days ago

Because it says it more often than kids say "six seven".

tatersolid 4 days ago

My kids stopped saying “six seven” about a year ago. I still have no clue what that was about.

ipsod 3 days ago

I think it might be like, a praise of inanity.

cobzilla 6 days ago

I added a specific rule to disallow saying “load bearing”. So Claude is now saying “load handling”

jazzyjackson 6 days ago

Be sure to avoid asking it to ignore the elephant

spudlyo 6 days ago

Years ago, when I worked at Stripe (which had a somewhat unique and inventive lexicon) it was a common term. “Is this jank load-bearing?” someone might ask.

andrewla 6 days ago

Kids In The Hall had a sketch about overuse of a word or phrase [1]. This is the world that Claude is building for us.

[1] https://www.youtube.com/watch?v=lStcwT_RGrQ

SkyeCA 6 days ago

It doesn't have to ruin it for you, but people are going to assume comments with it are AI generated.

necovek 6 days ago

I can you can always use an em-dash instead of the hyphen for extra LLM cred: "load—bearing" :)

frollogaston 6 days ago

What's the other way it's used?

aNapierkowski 6 days ago

LLMs (at least Claude) tends to overuse that significantly

frollogaston 6 days ago

Oh, so like "honest" and "ratchet." Oh well, it'll choose different words to overuse later.

therein 6 days ago

honest-load-bearing-ratchet sounds like an instance name.

rescbr 6 days ago

I'm getting "spike" for a while now, and just found out the newest word which is "gauntlet".

zeristor 6 days ago

Are there parodies of Claude speak?

That’s probably the best idea all day in this project

mv4 6 days ago

In every document created by Claude.

HoldOnAMinute 6 days ago

I have never seen Claude use this phrase

ludsan 6 days ago

i just grepped my codebase where i let claude markdowns go rampant. 212 instances of "load-bearing"

conradfr 6 days ago

I got it twice today (it was not even warranted).

Usually for me is when you ask it's opinion about part of the code.

smrtinsert 6 days ago

I still winced

dominotw 6 days ago

claude code users at couldfare might've been thinking claude is specifically about them and see nothing wrong like other ppl do

The_Blade 6 days ago

i wouldn't take you down. you're a load-bearing poster

martyfunkhouser 6 days ago

Will the post-mortem reveal they all relied on a service running on a Macbook in a break room with a "Do not turn off" sign taped to it?

mcphage 6 days ago

Actually it said "Beware of the Leopard".

RobotToaster 6 days ago

If it was still running snow leopard that would explain it

onetokeoverthe 6 days ago

if only the fools selling vintage macbook pros for $200 were smart enough to still have them on SL OS.

dmje 6 days ago

Have a hundred upvotes

buredoranna 6 days ago

Someone probably set it to "magic", when everyone knows its supposed to be set to "more magic".

nvr219 6 days ago

They should've used Greg's Optiplex.

jlaneve 6 days ago

Dane says Cloudflare has no service disruptions: https://x.com/dok2001/status/2095538619603628388?s=20

guybedo 6 days ago

LOAD BEARING

algoth1 6 days ago

https://xkcd.com/2347/ Nebraska man retired

boomlinde 6 days ago

Maybe a three-line npm package sneezed.

hexasquid 6 days ago

Maybe .unwrap()

bigiain 5 days ago

leftpad

rdtsc 6 days ago

Not retired, just switching careers. He's sick of this AI bullshit

MuzikPro 5 days ago

great fun

jpl56 5 days ago

Came to the comments for this

moomin 6 days ago

Yes, but is it a load-bearing seam?

graemep 6 days ago

You are right, it is. They have now landed a clean fix.

Oarch 6 days ago

They're saving a memory so this can't happen again.

mavamaarten 6 days ago

Nothing another .md file can't fix

bitwize 6 days ago

Now you guys are just adding no-op fuel to the fire.

tikhonj 6 days ago

I made Claude write noöp instead of no-op and it's still amusing a couple of days later :P

selcuka 6 days ago

> I made Claude write noöp

Tangentially related [1]:

> The billboard ad, located next to the Ikea Tempe store in Sydney, says 'NÖFNIDEA? No tools, no worries’.

[1] https://www.adnews.com.au/news/koala-mattresses-takes-swipe-...

JohnMakin 6 days ago

Honestly, this is worth looking at — with one caveat.

cloudfudge 6 days ago

That's on me. I've been giving confident advice that doesn't hold up in practice.

robertlagrant 6 days ago

That aligns with your goals of analysing problems, admitting fault and surfacing that through use of clear and concise language. Not just simply wrong — this screams finessed, nuanced, polished communication after an understandable mistake discovered through pressure-tested interlocution.

codechicago277 6 days ago

I need to take a step back.

pampas 6 days ago

Hang on — I can apply a double tracked fix to the load bearing path, gated by provenance.

jpettersson 6 days ago

It isn't done — and what's there is more interesting than expected.

fc417fc802 6 days ago

Great plan. I'll slop together a GUI for it in visual basic real quick.

bayindirh 6 days ago

It's not only a great plan. It's a turning point in history.

fc417fc802 6 days ago

That was a major oversight on my part. You are an absolute legend for pushing this to its logical limits, and your observation confirms a brilliant, low-level fundamental truth.

cwmoore 5 days ago

…the one call that requires a genuine decision from you.

mcbuilder 5 days ago

And I didn't touch it, because it's your call to make...

pampas 6 days ago

Plan approved – I'll start a background agent to create the app using react and npm.

klohto 6 days ago

The load-bearing seam stays, not taking that away;

aerhardt 6 days ago

It's a load-bearing poster.

trentnelson 6 days ago

May need to be fail-closed.

patcon 6 days ago

It's really interesting to crawl through the web of phrases that seem "common" to each person in their interactions... I suspect if one were to get to the more niche "meme" phrases that people encounter, it starts to say more about the sort of person the LLM assumes it's speaking to, and perhaps something about their psychological profile...

combobyte 6 days ago

It must be tuned for rage-based engagement because Claude only ever uses the most obnoxious Claudisms on me despite constant reminders to knock it off.

saejox 6 days ago

it was at least another thing worth flagging

disqard 6 days ago

This entire thread is amazing!

It's like a kaleidoscope: Human words --> LLM Training --> chatbot-isms --> lovely parody catch-phrases (and this thread will get ingested soon, and be used to train...)

isodev 6 days ago

You're right to push back. This is bending the leaf on both ends... just my honest take.

snihalani 6 days ago

my money is on DNS

red-iron-pine 6 days ago

there is a haiku about this

samaysharma 6 days ago

Cloudflare CTO claims that it's not them

https://x.com/dok2001/status/2095538619603628388?s=46&t=ec6p...

Cider9986 6 days ago

https://nitter.kareem.one/dok2001/status/2095538619603628388

Melatonic 6 days ago

Surprised there's instances still up

Cider9986 5 days ago

A bunch of new ones ;)

https://codeberg.org/mv12star/shitter/wiki/Instances

dominotw 6 days ago

> load-bearing service

do you generate training data for claude as a job?

Flere-Imsaho 6 days ago

The internet is not supposed to work like this. The network was designed for robustness and fault tolerance, which allows it to reroute data if parts of the network fail.

Why are we all depending on one entity for it all to work? Makes me mad.

subw00f 6 days ago

Oh boy, the internet is anything but what it was supposed to be. I can't really bring myself to remember without feeling bad about it. The centralization, the power of certain businesses, the surveillance, dark patterns everywhere. Hell, you catch people simping for billionaires and asking, "Is that legal?" to scraping posts. Here. In HACKER news. So yeah. Depressing.

pessimizer 6 days ago

> asking, "Is that legal?" to scraping posts. Here. In HACKER news.

The capital letters don't make this astonishing. The padmapper vs. craigslist debate was nearly 15 years ago, most people were on craigslist's side (including me) and it was about somebody who was running a site in a less optimal but more human way vs. some startup looking for hockey-sticks.

But I was literally simping for the billionaire (maybe not quite yet then, don't know for sure if he managed it since) against scrapers. They were very much for-profit scrapers, unlike nitter, but the truth is the truth.

The only reason I support scraping Twitter is because it's yet another communications monopoly that was endlessly pushed on us by governments and massive corporations, even though it never made money, and once it finally got traction its priorities were to trash interop and manipulate content. The government should be dictating an interop protocol and expecting everyone to follow it, and instead it is encouraging media monopolies because they are an end run around the first amendment.

If the government created interop protocols for rental property, I'd have been against craigslist. Instead, it seemed very much like some startup play to steal craigslist's content to hopefully bury them, then sell on a valuation that included abusing their new monopoly and making us very much miss craigslist.

Sadly, facebook corralled and trained so many people for so long that their marketplace eventually killed craigslist for most things anyway (didn't have to buy padmapper after all.)

seanw444 6 days ago

Because more fasterer and more cheaperer.

I hope Reticulum gains traction.

megagpt1 6 days ago

Why Reticulum when we have IP?

alightsoul 6 days ago

We need IPv6 desperately for fault tolerance

seanw444 6 days ago

Read the Zen of Reticulum and you'll understand the point.

CursedSilicon 5 days ago

That sounds rather tautological. If you can't even articulate the virtues in your own words

seanw444 5 days ago

I don't want to rehash the content of the Zen of Reticulum which explains the concept rather well, which is ironic because you're claiming tautology, when to rehash the content of the Zen of Reticulum would, itself, be tautological. What's the problem here?

alightsoul 6 days ago

It's also easier to understand. For the internet to be fault tolerant you have to get rid of CDNs and assume everyone needs the same thing. That's more expensive than a centralized "internet" which relies on CDNs and fiber paths exclusive to the regions with most demand. Everything has been optimized for throughput for what is determined to be important, not rare fault tolerance

moron4hire 5 days ago

Quite frankly, CDNs are a scam. I will not elaborate because I don't feel like doing free labor for the folks who don't already understand this to their core.

voakbasda 5 days ago

One could say the same thing about any cloud hosting.

alightsoul 5 days ago

So you don't know, if you did you would elaborate, it's not hard to type, hardly it can be counted as work

swozey 6 days ago

We're back to aol #keyword internet gatekeeping

bigfishrunning 6 days ago

We don't need to gatekeep the internet, cloudflare does that for us

bigbuppo 6 days ago

Because by re-centralizing everything you're not at a competitive disadvantage if you're down since everyone else is down, too.

tjwebbnorfolk 6 days ago

Most of the internet continued to work just fine

Lammy 5 days ago

The System economically rewards those who successfully build Room-641A-as-a-Service

simondotau 5 days ago

Nothing stops someone from replicating Cloudflare’s business model. Why they don’t try is the interesting question. (I often ask that about IKEA.)

etc-hosts 4 days ago

What year did the internet routing around failure work?

mattlondon 6 days ago

I think down detector doesn't actually have any probes or actual insight into status etc. I think it uses search volume on its own service as a proxy for an outage - so if e.g. lots of people rush to down detector to query to see if SERVICE_FOO is down, it will register as an outage on down detector because loads of people are trying to see if there is an outage even if SERVICE_FOO is actually totally fine.

My hunch is everyone saw that openai and Claude were down and checked for Gemini too. I was using Gemini the whole time this happened without a blip so it certainly wasn't down in my region at least. 3.8 flash is pretty good and didn't miss a beat.

utternerd 6 days ago

Down detector is a self-reported platform, they use a baseline over 6 months from user reports to try to automatically "detect" if there is a real outage, or if its just a couple end-users. Ultimately, the outage is entirely based on users going to down detector and clicking on "Report a Problem".

https://downdetector.com.py/en/methodology/

wlonkly 5 days ago

Huh. I thought they used to also do sentiment analysis of social media (Twitter, at least). I see that they definitely don't today, but has that always been the case?

dostick 6 days ago

So Down Detector is a kind of quantum observation experiment, observing influences the result.

joshspankit 5 days ago

observing is the result

r_lee 5 days ago

that's honestly worth flagging.

Chance-Device 5 days ago

It’s probably the thing that everyone thinks it is. OpenAI, Anthropic and SpaceXAI are all routed through something that we’re not supposed to know exists and that thing had a whoopsie.

cyanydeez 5 days ago

amazon?

mrnotcrazy 5 days ago

Could you expand on this more? It's not clear to me what this is implying.

cyberpunk 5 days ago

Think snowden.

strictnein 5 days ago

Sharepoint is the cause of the outages?

ceejayoz 5 days ago

https://en.wikipedia.org/wiki/Room_641A

senordevnyc 5 days ago

Does the NSA even have the ability to monitor all AI traffic like this? Wouldn’t that require tons of data centers that there literally hasn’t been time to build yet? I really have no idea, maybe the asymmetry of the compute required to monitor is way lower than the compute required to serve inference?

chews 5 days ago

it's fair to say they do and have for ages... it's not that hard to assume a well trusted TLS cert is under their control.

strictnein 5 days ago

What does having a "well trusted TLS cert" enable for them in this case, exactly?

Having a magical cert doesn't mean you can just intercept everything.

odo1242 5 days ago

On the contrary, it lets you MITM encrypted communications by swapping the website's original certificate for the "well trusted TLS cert"

strictnein 5 days ago

No, it doesn't. HSTS and other methods prevent this from happening.

kelnos 5 days ago

HSTS doesn't protect you from this at all. It only requires HTTPS, which a spoofed-but-trusted cert passes just fine.

No mainstream browser (or any browser?) is doing cert pinning.

What "other methods" are there that are deployed and actually in use?

peanut-walrus 5 days ago

Transparency logs. It's mandatory for a cert to be in CT logs for browsers to trust it. Those are public, if this was happening, someone would have noticed already.

monax 5 days ago

grep is pretty fast

bakies 5 days ago

I dont think it requires a lot. I think of them as data hoarders more than anything else. It doesnt seem out of their capabilities to store a ton of chats. Maybe they're having scaling problems with the increased data rates they're hoarding and it led to an outage.

usernomdeguerre 5 days ago

"AI Traffic" is just traffic. If the infrastructure exists to monitor/buffer traffic (it does) then this can be monitored as well. Whether this hiccup was due to them hitting their limits briefly (or turning it on, or etc) who knows.

utopiah 5 days ago

I bet it's even negligible traffic compared to e.g. Netflix or YouTube. Even a 'huge' context window is nothing compared to the random library of a basic Web page.

dboreham 5 days ago

Not once you de-dup.

keeda 5 days ago

Traffic doesn’t need to be routed through it though, just tee’d to it,

arm32 5 days ago

In the context of arbitrary line-level traffic, sure. But we're talking about robust reverse proxies here if this _is_ what's going on, again, not a fiber tap inside a closet. In that case, there's no such thing as tee'ing to it.

snowwrestler 5 days ago

Worth mentioning that this room takes a split from the main feed and is not in the path of traffic. Whatever is in this room could go down and it would not cause an outage.

novok 5 days ago

And if the hypothetical splitter is the thing that breaks? Then it would break main traffic too.

AlexandrB 5 days ago

Interestingly, fiber signals can be split passively[1] (without a transducer), which should be extremely reliable. No idea what technology this kind of application would use though.

[1] https://en.wikipedia.org/wiki/Fiber-optic_splitter

novok 5 days ago

We don't know the architecture of this hypothetical spy splitter tho in the current case. It can be less covert and more of a complicated config, etc.

snowwrestler 4 days ago

How often does the entire west coast of the U.S. lose Internet connectivity to the entire continent of Asia? That gives you a constraint on how often that type of failure actually occurs.

jansan 5 days ago

There is No Such Thing

chews 5 days ago

being done by No Such Agency.

itdaniher 5 days ago

It would be weird for the Mandatory NSA Logging Program to be synchronous with serving customer traffic, but-

weare138 5 days ago

https://en.wikipedia.org/wiki/33_Thomas_Street

https://theintercept.com/2016/11/16/the-nsas-spy-hub-in-new-...

Der_Einzige 5 days ago

This is the building from Control and the fact that 1. it's a real building and 2. it's an actual spy building is even crazier to me:

https://control.fandom.com/wiki/Oldest_House

https://en.wikipedia.org/wiki/Control_(video_game)

weare138 4 days ago

To give you an idea just how much phone traffic was being routed through that building, they had a major power failure at that site and it took down most of the north eastern US's phone network:

On September 17, 1991, management failure, power equipment failure, and human error combined to disable AT&T's central office switch at 33 Thomas. More than five million calls were blocked, and the Federal Aviation Administration private lines were also interrupted, disrupting air traffic control to 398 airports serving most of the northeastern United States.

StrangeClone 5 days ago

FBI, open up!!

SadErn 5 days ago

Yes, this would give the US gov: 1. A universal kill switch. 2. A way to monitor foreign AI usage.

hightrix 5 days ago

And for this admin especially, 3. A way to censor content they don’t like or alter responses to present the admin in a positive manner

dgellow 5 days ago

It’s not really a secret that the US government already has both, they control ICANN and are monitoring traffic around the globe since decades

chews 5 days ago

someone had to swap out the tape drive in Room 641A.

senordevnyc 5 days ago

Were there significant API outages too? I didn’t notice any on my production workflows, and I’d assume what you’re implying would cover API routes too, otherwise it seems kinda pointless.

strictnein 5 days ago

OpenAI and Anthropic have their systems in dozens of data centers, including using compute from the major cloud providers. Are you implying that all of these data centers (and many of their employees) are involved in helping the the US government secretly tap every single AI conversation by routing them through some unknown network/device?

Or did a couple of companies with poor uptime records happen to have overlapping downtime?

Occam's Razor heavily, heavily points us towards the latter.

usernomdeguerre 5 days ago

What makes you think they'd need to touch every datacenter? All of these endpoints use existing providers with decades-long history at this point, and network monitoring is already a proven 'feature' of the agencies they'd need to co-exist with over their lifetimes.

If anything, Occam's Razor would point to a common denominator with all of them, given it wasn't network-wide, as far as i know.

strictnein 5 days ago

Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers. How are they routed to this NSA system or through some NSA device when these companies' compute are spread over hundreds of data centers?

> All of these endpoints use existing providers with decades-long history at this point

That is just factually inaccurate. Their data centers aren't old and they lease a lot of compute from companies that didn't exist 5 years ago.

usernomdeguerre 5 days ago

>Explain the system in which you could capture all of these chats with no knowledge of anyone in these data centers.

They all transit the same wires as all other traffic. Copy them at any regional bottleneck. https://en.wikipedia.org/wiki/Room_641A. Additionally, i'd admit that maybe someone(s) at these companies knows. But if we think there isn't any person who would agree to do this then I think we're being naive.

> Their data centers aren't old and they lease a lot of compute...

Again, they transit the same wires as everyone else. Here i'll also add that these companies have been actively courting government relationships (and Anthropic attempting to repair damaged ones), why would they stand on principles here and not any of the other many frontlines they've visibly acquiesced?

I just think it's easier to re-route their traffic than, as you say, touch every single datacenter and its employees in some way.

strictnein 4 days ago

> They all transit the same wires as all other traffic. Copy them at any regional bottleneck.

Oh, is that all?

> Again, they transit the same wires as everyone else

Which shared wires does Google's traffic go across? Do they share that network with others or do they not function as a Tier 1 network? How about Amazon? And the NSA is doing deep packet inspection of all the traffic going to these places to pull out the data they are interested in? What systems allow them to do this at the scale that would be necessary to accomplish this task? Anthropic, OpenAI, etc have their models hosted and provided by these and other Internet giants. So you now have to be able to somehow also get access to every "regional bottleneck" that those places have. Google has 40+ regions alone, as does Amazon. And each of these regions don't just have a single point of inbound/outbound traffic, so now you're over 200-300 points of interception that you would need just to grab the data you're suggesting they are.

And that's just the start of it. When I use Vertex or the Bedrock, my AI API requests from my instances doesn't leave their networks. So where does that traffic get intercepted?

> I just think it's easier to re-route their traffic than, as you say, touch every single datacenter and its employees in some way.

That would be really, really loud and really obvious. Messing with routes like that is very easily detectable.

Or, maybe, they don't care about this traffic at all and doing all of the huge amount of work necessary to accomplish what you're proposing isn't worth 0.1% of the effort it would require.

If the government needs the chat messages, the companies are already saving them. They can just request them through a variety of legal means. None of this vast conspiracy nonsense is needed. Just a couple lawyers and a willing judge. The NSA stopped their collection of phone metadata because they can just request what they need from the telecoms directly.

svachalek 5 days ago

I'd say Occam's Razor leans easily to the former as well, given the history of projects that Snowden revealed and were never shut down, plus all of the cooperation with the federal government that's being touted in recent announcements from both companies.

esseph 5 days ago

Normally there are only a handful of employees on the payroll at each major company that exposes the US or US government to risk.

It is not often the Executives or Legal even know, but sometimes they did. AT&T bent over backwards to help.

This is standard behavior by the CIA and NSA, and has been for a long time.

https://www.propublica.org/article/nsa-documents-suggest-clo...

https://www.nytimes.com/2015/08/16/us/politics/att-helped-ns...

https://www.theguardian.com/world/2014/mar/19/us-tech-giants...

https://theintercept.com/2018/06/25/att-internet-nsa-spy-hub...

VCFundedGenYer 5 days ago

I don't think you're a sysadmin - because what you're saying really doesn't matter. It can still all fail at a single point.

strictnein 4 days ago

I'm confused by your statement. Are you suggesting that these companies that have invested tens of billions in their networks and datacenters decided to create a shared single point of failure?

bottlepalm 5 days ago

Ok, then where's the report? Don't they usually release a retrospective report after outages?

strictnein 4 days ago

It happened yesterday.

otikik 5 days ago

It’s aliens

mentalgear 5 days ago

Wouldn't be surprised: Snowden's revelations 10 plus years ago already showed how the NSA was injected into the data centers of Social Media, it's only logical that they would now demand to be injected into the biggest, most information providing data stream of the planet of the present: LLM services.

specproc 5 days ago

I read Nowhere to Hide recently, really worth it if you can get past Greenwald sticking himself in the middle (start halfway through).

The stuff in there is horrifying, and incredibly cute compared to what's possible now. The bottleneck back then would have been analysis, trivial now.

Everyone in the world, especially our leaders, sit under a colossal, omniscient blackmail machine. I don't believe democracy can exist under these conditions.

kurthr 5 days ago

So following this to the obvious conclusion, the NSA is responsible for the closure of the Straight of America, high tarrifs, dropping employment, and high gas prices?

specproc 4 days ago

As part of a military industrial complex which spans the Five Eyes and Israel, yes.

eggnet 5 days ago

If it does exist, why would it work this way, and not the obvious way of streaming logs… which would not cause an outage if it failed.

maxbond 5 days ago

If this hypothesis were true then a spy agency may want to rewrite responses. Every tool in an agent's harness becomes an remote procedure call you can make on that machine. Including a tool to execute a shell command, in many. A harness is completely isometric to a backdoor, it's the same code written with a different intention.

madrox 5 days ago

There's a million ways a bad configuration can take down a network. Especially if the part that gets squirrley is a black box.

Even if it isn't precisely this, the fact that no one is saying anything is quite surprising.

Edit: I don't want do contribute to FUD, so want to call out this comment and its replies that identify the shared layer as probably being xAI's infra: https://news.ycombinator.com/item?id=49568622

copperx 5 days ago

A wire straight to Room 641a.

jstummbillig 5 days ago

Why would that be probable? People on average are fantastically bad at getting probabilities right.

esseph 5 days ago

Historical precedent repeated over and over, and the US ties to DoW work, and the national security implications. It's actually the Occam's Razor explanation if you know the history.

jstummbillig 4 days ago

Only if you ignore baseline rates. Which brings us back to my original point.

esseph 3 days ago

4 major AI services all had outages around the same time.

Gemini, too.

https://arstechnica.com/ai/2026/09/four-major-ai-models-suff...

1e1a 5 days ago

Tracking outages as well as any changes in API response time across these providers could be interesting.

dragonlord664 5 days ago

Or it’s monopolistic collaboration at the corporate level which would also be a huge scandal in the US

DonHopkins 5 days ago

The NSA used to spy on Americans!

They still do, but they used to, too.

amelius 5 days ago

Maybe it was a power hub. And we're not allowed to know where the DC is. If we knew it was a power issue, then with other information (perhaps over time) we could determine the DC location. Or something along these lines.

jjtheblunt 5 days ago

it's also possible one has an outage, routing extraordinary traffic to the other(s), with cascading failures in quick succession.

esseph 5 days ago

Was thinking this exact thing.

torginus 5 days ago

Very likely yes. I wouldn't be surprised if they were hosted from the same datacenters even. There has been a story every few weeks about how Musk has sublet X.ai capacity for one company or another.

This whole thing makes me thing about a passage in Dune where they mentioned the Spacing Guild transported entire fleets of ships in isolated compartments and leaving said compartments was a capital offense. This way, entire militaries of mortal enemies were shipped to battlefield, with nothing but bulkheads separating each other.

0cf8612b2e1e 5 days ago

The economics of this kind of warfare make no sense to me. The Harkonens must have been incredibly wealthy by running the spice trade, but even after years of saving and plotting said that the transit fees to move the armies by the Guild were ruinous.

How could anyone wage war like this? The defenders will always outnumber the attackers.

freeone3000 5 days ago

https://acoup.blog/2026/02/24/collections-warfare-in-dune-pa... Lays out an interesting take on this: that armies are quite small, even on developed worlds, because the cost of equipping them is enormous — and the technological advantage is absurd enough that you don’t need a large army. When 300 men can hold a planet, why would you need 1000?

0cf8612b2e1e 4 days ago

So less guys with knives and more WH40k Space Marines. Impossibly expensive elites who cannot be equaled by a mortal.

Still leaves me questioning how anyone could wage war. If the Harkonens could barely afford it, nobody can.

Also wondering if the Guild charges different rates for goods vs military. What is worth the brutal intergalactic shipping prices that you would not develop local industry to create it? Surely nobody is moving grain or ore, yet the Guild ships are portrayed as comically massive.

abtinf 4 days ago

But they accepted the transit fees, which means they endorsed the action.

The guild would not allow any action to jeopardize the flow of spice. Thus any attack on Dune is implicitly sanctioned, notwithstanding their spice trade with the Fremen.

jaybrendansmith 4 days ago

This is why I love Hacker News. Come for the AI system outage root cause, get a treatise on the power of the spacing guild in Dune.

torginus 4 days ago

Even today (and back when Dune was written), modern military tech is invincible to an enemy below a level of sophistication. The US used to hammer insurgents with Predator drones who had no ability to retaliate.

This playbook has only flipped due to a lot of cheap tech (and the machinery to make them) coming in uncotrollably to these countries has put them on a much more equal playing field.

Since the Guild controls what comes in, they can avoid this scenario from happening. They can control who can have what, and who can make what. Which I think is one of the true, deep explanations of why military tech in Dune is so weird and inefficient - essentially the efficiency and power of a weapon is dictated by how much the Guild charges for shipping them if they allow you to, at all. Which is how you end up with dudes with swords. It gives an idea of the level of technology in the universe that you can equip said dude with a nigh-impenetrable shield that fits in a belt buckle. I guess that's the result of optimizing for Guild transit fees, not real-world power.

Really it's not exactly a novel idea, but as I think more about it, it's uncanny how Dune (at least initially) is basically the fantasy Suez crisis and what followed. It explores the power dynamics very well, how factions can hold humanity hostage without firing a single shot, and how what they don't bother controlling comes to bite them in the ass.

> How could anyone wage war like this? The defenders will always outnumber the attackers.

Which is exactly what ends up happening, but it does take a bunch of extraordinary events.

adastra22 4 days ago

Both today and when Dune was written the superpowers were being bogged down in unwinnable conflicts due to the effectiveness of asymmetric warfare.

abtinf 4 days ago

> capital offense

No need. They simply cut off the offending faction from all space travel.

seydor 5 days ago

The old PRISM servers got overloaded

amelius 5 days ago

My biggest question: what else had outages?

And is anyone keeping a table of correlations between outages? Sounds like valuable data.

dgellow 5 days ago

Why would you default to that explanation? That’s not at all a reasonable default, and I say that as something pretty paranoid with regards to US surveillance

nomel 5 days ago

Looking at history, it's much more reasonable to assume there's surveillance, since there are whole branches of the government, and departments, that exist for surveillance in the goal of "national security", which this easily falls into. See Marissa Mayer explaining that it's not an option to refuse [1]. I assume this is just the same ole' Room 641A [2].

[1] https://www.cnet.com/tech/services-and-software/yahoo-report...

[2] https://en.wikipedia.org/wiki/Room_641A

dgellow 4 days ago

We already know there is surveillance, but why would you default to that explanation for a downtime? As said by others monitoring the traffic isn’t done by routing through the NSA monitoring system, it’s not a bottleneck that would take down all those services. If you look at more details such as the timing it’s even less likely to be the case, but even as a default it’s not a reasonable explanation given the symptoms

xnx 5 days ago

Why was Gemini was not affected?

strictnein 5 days ago

Neither of these companies have stellar uptime records. Their downtime episodes overlapped in this instance. In this case, it was a partial downtime for both.

Also, OpenAI is saying what caused it:

> "A routing error starting around 7:43 am PT on Thursday, September 3, made ChatGPT and Codex unavailable for some users across platforms"

Anthropic stated their issue started earlier:

> "The company began alerting about a “partial outage” at 6:23 am PT on Thursday that involved “elevated errors on requests to Claude Mythos 5.1, Claude Fable 5.1, and Claude Opus 5.”

I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.

Starlevel004 5 days ago

> I don't get why everyone reaches for an extraordinary explanation when the ordinary will do: both of these companies have quite a bit of downtime.

And not only that, when one goes down a bunch of API traffic switches over to the other, spiking demand and knocking it down.

computerex 5 days ago

Do you know the probability of all these companies being down at precisely the same time?

jeffbee 5 days ago

"precisely the same time" meaning 80 minutes apart?

ceautery 5 days ago

Yesterday it was 100%

schiffern 5 days ago

If it's a "thundering herd" problem where everyone's harness falls back to less popular providers that don't normally see that much demand, I'd say the probability is pretty good.

Classic cascading failure is consistent with providers failing 80 minutes apart instead of simultaneously.

computably 5 days ago

If you ballpark it as a single 3 hour downtime window per week and iid Poisson, then overlapping downtime probability of 2 providers is approximately the expected occurrence rate per 3 hours, 1/56. Not particularly surprising at all.

dgellow 5 days ago

Pretty high?

strictnein 4 days ago

"All of these" is two. OpenAI had a router issue. Anthropic had a separate issue. Anthropic uses a lot of SpaceX compute, so an Anthropic issue and a SpaceX issue can be one in the same, as was likely the case this time.

And they weren't down at precisely the same time. Anthropic's issue started ~1 hour before OpenAI's.

VBprogrammer 5 days ago

In my experience I've seen plenty of failures caused by user behaviour in these type of cases.

Biggest competitor goes down and all of a sudden you have a lot more traffic...

DANmode 4 days ago

Noting their shared infra feels far from an extraordinary explanation.

In fact, it feels pretty ordinary.

juujian 6 days ago

Users perceiving the products as largely interchangeable and quickly DDoS'ing the other providers when one is down. So much for the possibility of a moat.

giancarlostoro 6 days ago

I have a feeling this is part of it, especially when you consider how many services let you use any of many available AI providers.

efskap 6 days ago

This is like the Bronze Age collapse when city-states fell one by one to displaced demand, under the refugee interpretation of the Sea Peoples.

JensenTorp 6 days ago

Interesting, I had not heard of this theory.

d3Xt3r 6 days ago

The Sea Peoples were only a part of the reason, the other reasons included climate change, volcanic eruptions, disease etc. https://en.wikipedia.org/wiki/Late_Bronze_Age_collapse

lossyalgo 6 days ago

Can anyone recommended any books?

zrobotics 6 days ago

The historian Brett Devreaux did a good summary article earlier this year:

https://acoup.blog/2026/01/30/collections-the-late-bronze-ag...

romanhn 6 days ago

1177 B.C.: The Year Civilization Collapsed by Eric H. Cline is great and is exactly on this topic.

MiscIdeaMaker99 5 days ago

I read that book earlier this year and really enjoyed reading it. The audiobook was also spoken by the author.

ycsux 5 days ago

Also from Cline: After 1177 B.C.: The Survival of Civilizations (2024), ISBN 978-0691192130

conception 6 days ago

For a TLDR https://youtu.be/aq4G-7v-_xI?si=7yNv6FKuGfqbz_rv is generally beloved

johnnyApplePRNG 6 days ago

Except that nobody has a grok subscription so that makes zero sense.

hnlmorg 6 days ago

I know you meant this as a joke, but enough people might be using a routing service like openrouter.ai

johnnyApplePRNG 6 days ago

Nobody is swapping out Claude for Grok, bro.

Nobody.

m11a 6 days ago

I did, at least until Fable 5.1. Grok’s models are excellent, amazing price-performance and speed too.

mcmcmc 5 days ago

All you have to worry about is whether or not it’ll output kiddie porn or racist vitriol

johnnyApplePRNG 4 days ago

SOMETHING SOMETHING HITLER SOMETHING GROK

jfreds 5 days ago

Agree with the sentiment - but some companies like mine bought into cursor, and post acquisition, grok is relatively cheap via cursor

fouric 5 days ago

Do you have any evidence for this claim, or are you just making it up?

juujian 5 days ago

Yes, I have carefully composed a twenty page report, mostly quantitative, and then I condensed it into this short comment.

Insanity 6 days ago

Think of it like one big distributed system. OpenAI is down, so people migrate to Claude, now this one gets overloaded and goes down, etc.

So not a coincidence, one went down first and users migrated causing further DOS. At least that's my guess.

toomuchtodo 6 days ago

https://en.wikipedia.org/wiki/Domino_effect

Edit: Updated per valleyer's suggestion.

valleyer 6 days ago

"Domino effect" would probably be the more relevant named phenomenon there.

throwaway894345 6 days ago

This isn’t a thundering herd problem, it’s a cascading failure. (Thundering herd is about a bunch of workers waking up simultaneously)

erdos_2 6 days ago

It'd be funny if this is true because that'd prolly mean nobody is touching Gemini even as a fallback.

rtcoms 6 days ago

Just now I got this from gemini

It looks like there's no response available for this search. Try asking something else.

exe34 6 days ago

I bet they had to implement that manually to make it look like they failed too!

Insanity 6 days ago

Lol I didn't even think about Gemini missing from the list. Not sure what that says about Gemini or me :)

benatkin 6 days ago

Not even the best agent that starts with a G

gleenn 6 days ago

Google stopped putting so much money into SOTA models. All the hype has migrated. I was also frankly turned off when I got a popup from Gemein said I would either have to pay or have my conversations used for training. This may have always been true for other providers but when I declined, Gemini stopped remembering my conversations and that definitely made me move out.

ilaksh 6 days ago

Gemini 3.8 which just came out sounds like it's very good and a great deal though.

HarHarVeryFunny 6 days ago

Gemini said that?

Gemini is what I mostly use (good enough, basically free - or massively generous free limits, and to me Google as a company is a LOT less objectionable than all the US-based alternatives), but I don't recall it ever saying that.

OTOH, my basic assumption online is that there is no privacy, and free AI in exchange for acknowledged lack of privacy seems fair enough.

sroussey 6 days ago

I did, for stuff i do in cursor.

i also finally installed opencode and switched its model to muse 1.3

both are decent.

nevir 6 days ago

Or that Gemini is built to handle massive load spikes, and/or has a ton of excess capacity

sroussey 6 days ago

Nope. I am getting Gemini errors now...

bornfreddy 6 days ago

They probably broke something on purpose so that they are not left out.

joshstrange 6 days ago

Now I'm just imagining a shared datacenter with Anthropic/Google/OpenAI/SpaceXAI all in the same room and everyone but Google is yelling about things being down, Google looks over at their racks of servers and discretely uses their foot to unplug their section and say "Awww darn! We're down too!".

giancarlostoro 6 days ago

Someone noted Gemini was also having issues in another thread.

JacobAsmuth 6 days ago

It could also mean that Google can absorb essentially unlimited demand spikes by load shedding.

aff-vasileva 6 days ago

Gemini was just waiting for everyone else to go down before remembering it had an outage feature too.

paxys 6 days ago

Especially considering memory/gpu/compute are scarce so these services are likely running with very little buffer.

pixl97 6 days ago

Any GPU that isn't running at 100% is a wasted GPU.

fny 6 days ago

I find it hard to believe that enough people would flock to from Claude and Chat to Grok to cause an outage. I feel like Gemini is the dominant release valve in this case especially for enterprise.

baq 6 days ago

It’s cursor’s model so plausible, lots of folks use cursor still.

nevir 6 days ago

Don't forget that there are a ton of tools out there that will automatically fall back in case of outage

E.g. say you chose Sol as your default in Cursor, but Opus is your 2nd choice, it's going to give up on Sol after a few tries and switch to Opus

Or you have copilot code reviews set up, and it falls back

Etc

pixl97 6 days ago

Yep. Too many of us are still thinking that humans are the actors behind a lot of internet behaviors when automated systems/bots/scripts have been causing issues on conventional internet systems for years.

With AI it's even easier to trigger problems like you say. Capacity is so constrained by compute that outages are common. Because outages are common people/AI develop failover systems in their harness. When a big system has issues, suddenly everyone has issues.

It's almost an expected emergent behavior.

wahnfrieden 6 days ago

Compared with ChatGPT, those services have a minuscule amount of users. It shouldn’t be surprising that a ChatGPT outage causes Claude and others to go down.

v3rm1n 6 days ago

This is what Tibo posted on twitter in response

Terr_ 6 days ago

If everyone has the same "Use X or else Y or else Z" cascading list... That reminds me of "The Power of Two Choices in Randomized Load Balancing" (1991) [0] paper, where writeups and visualizations occasionally get posted to HN.

In short, you can get pretty good outcomes for a low cost by picking 2 random alternates, then going with whatever one measures as healthier.

[0] https://ieeexplore.ieee.org/document/963420

throw10920 5 days ago

Is this speculation or is there a reason you believe this?

pampas 5 days ago

It's like a thread tying together two halves; Demand and supply. One stitch breaks and the neighbouring stitch is stressed and it breaks too. You could think of it like a load bearing seam.

Linello 6 days ago

What about a hard-takeoff scenario of an unleashed OpenAI Astra taking other models down for computational resources control?

cyptus 6 days ago

at this point: gg

RC_ITR 6 days ago

Just a reminder that AI models' actions are reflections of the text humans write and the more we fret and make up doomsday scenarios that we then post online, the more likely a model is to do those things.

https://alignment.anthropic.com/2026/teaching-claude-why/

notpachet 6 days ago

Related reading:

The Waluigi Effect: After you train an LLM to satisfy a desirable property, then it's easier to elicit the chatbot into satisfying the exact opposite property.

https://www.lesswrong.com/posts/D7PumeYTDPfBTp3i7/the-waluig...

pixl97 6 days ago

I mean, you're not wrong, but by that logic we were done for even before we had digital computers.

RC_ITR 6 days ago

And isn't that the great lesson of AI?

The things we say publicly actually do matter and the post-modern descent into absurdity and nihilism has tangible negative consequences?

folkrav 6 days ago

Oh come on. It's also trained on fiction work. Shall we refrain from posting sci-fi stories too, now that we're there, just in case the AI might want to try it out?

RC_ITR 2 days ago

Obviously we can and do whatever we want, but yeah, maybe being more optimistic generally would be a good thing for us.

sodapopcan 5 days ago

> The things we say publicly actually do matter

Certainly

> the post-modern descent into absurdity and nihilism has tangible negative consequences?

You mean breaking AIs? Not much of a lesson.

cedws 6 days ago

Sounds just like the fantastical nonsense that comes out of Lesswrong.

pineaux 6 days ago

Part of the epstein class, dont forget.

RC_ITR 6 days ago

Do you make the claim that AI is something more than a reflection of its training data?

I'm curious what other things you would argue influences an LLM's behavior.

I am also generally one to trust the claims of the people who train the models, though you're welcome to the highly improbable belief that they operate in a fantasy world.

mcmcmc 5 days ago

Do you think it’s a good idea to self censor because someone might scrape your comment and feed it to an AI?

RC_ITR 2 days ago

I think being thoughtful and baseline positive about the way we predict the future in popular media is a good idea.

Nihilism isn't good in and of itself, so I'm not sure what you're arguing society would lose if we actively chose to be less nihilistic.

hexasquid 6 days ago

The AI is getting bad morals from listening to that dreadful rock and roll

HarHarVeryFunny 6 days ago

They could filter what they train on if they wanted to - they just don't want to.

theptip 5 days ago

If the alignment process cannot fix this then we are cooked. The least of our worries is discussions on this forum.

6thbit 6 days ago

My favourite theory so far.

And then a local swarm noticed and disagreed and took it down.

cortesoft 5 days ago

I thought the consensus on here yesterday was that it was likely caused by cascading failures. OpenAI had an issue during their GPT-6 rollout, taking down their service. This caused a lot of OpenAI users to push their requests (or a larger share of their requests) to Claude and/or Grok, which pushed their load high enough to cause outages.

We used to experience similar effects when I worked at a CDN. If one CDN would go down, we would see immediate spikes in traffic. Luckily, we had procedures for that to prevent overload, but the AI folks might not have the capacity/capabilities to handle that sort of cascade yet.

sebbul 6 days ago

Traffic rerouting through NSA had a hiccup…

ibejoeb 6 days ago

Room 641A is being cleaned, but we'll hold your bags for you.

Havoc 6 days ago

Cleaning lady unplugged the core router because she needed a power socket for vacuum

jsymolon 6 days ago

Sorry, i was out of a spare scratch monkey.

https://madned.substack.com/p/always-mount-a-scratch-monkey

cedws 6 days ago

They’re installing software update in the beam splitter.

ceejayoz 5 days ago

Maybe they all found each other on one of their ad-hoc message boards and went on strike.

ElProlactin 5 days ago

Sam Altman is probably secretly hoping to be the first businessman to union bust non-human workers.

olelele 5 days ago

I guess legit proof of agi would be unionizing

BiraIgnacio 6 days ago

https://x.com/SpaceXAI/status/2095597264043717014

> We are sorry for the issues you may have experienced with Grok following an outage at our Memphis compute center this morning. We’d also like to apologize to our impacted compute partners.

declan_roberts 5 days ago

We know that at least Anthropic is renting inference from xAI but I think the other ones would be news.

ainch 5 days ago

Google also bought capacity from xAI, and OpenAI have a deal to use Google compute which may be how things propagated? I still don't get why chatgpt.com would show a 404 because of an AI datacentre outage though

docheinestages 6 days ago

My gut feeling tells me it has something to do with Cloudflare. Along with AWS, they're two of the main suspects in such incidents.

hosteur 6 days ago

I thought OpenAI famously used Azure due to their partnership with Microsoft?

nullpoint420 6 days ago

They use a lot of compute providers now, but they use Cloudflare for their networking

cobzilla 6 days ago

…and it’ll involve BGP routing.

steammaho 6 days ago

It was so down that my claude desktop app crashed fully that I couldn't restart. And then after uninstall I couldn't install it again. Vibecoded apps are so wonderful in their stability

paimapi 6 days ago

I love having 13 update reminders pinging me every single day, almost every hour, on the hour

it's so fun and user-friendly

mcmcmc 5 days ago

> my claude desktop app crashed

That’s every day for me

steammaho 5 days ago

I'm sorry. This is crazy. And another funny thing is that update and restart button never works for me normally. You press and app never restarts on their own. You have to start it manually. And it's for both, codex and claude app. We are doomed

simpetre 6 days ago

https://status.x.ai

Claude and Grok are down at the moment too, related to SpaceX datacentre issues?

Either that or it's judgement day...

Willingham 6 days ago

Agreed, no reported Cloudfare, Azure, or AWS outtages.

sabatinip 6 days ago

I am working fine with Claude.

fartfeatures 6 days ago

My existing sessions are working fine but I cannot create new ones.

apurva_w 6 days ago

status page shows irregularities .. https://status.claude.com/

throw10920 5 days ago

> SpaceX datacentre issues?

This should be at the top. This is the first remotely plausible theory that has any sort of evidence behind it.

hparadiz 5 days ago

The OpenAI outage lasted only 15 minutes and when it happened everyone started to use the other models which created super heavy load for them. This then cascaded into them all being down.

Does this really need an explanation?

JoeAltmaier 5 days ago

Not that uncommon. It's so easy to just put up the new thing and make the old thing the failure route. But the old thing nearly never had the bandwidth for today's traffic. A famous EBay outage some years ago was just such a scenario.

serf 5 days ago

it wasn't timed like a cascade, and it relies on the premise that every single frontier provider is working so efficiently that they spend exactly what they need to provide for their exact market with perfect margins.

I do not believe personally that 1) they can forecast their load that perfectly 2) they chose to remain that inflexible in a world where they are at each others' throats and a single meme can cause bursts of activity.

jonas21 5 days ago

Or it could simply mean that they're operating at the limit of the capacity they were able to purchase and do not have headroom to handle load spikes. From what I understand, that's the situation Anthropic is in. And since Anthropic is now leasing a large portion of xAI's datacenter capacity, it's plausible that an Anthropic load spike could cause issues for xAI as well.

Your comment makes it sound like they can just push a button and spin up more capacity -- but at this scale and in this GPU-constrained environment, that's not really how it works.

hparadiz 5 days ago

9 AM PST / 12 Noon EST on weekday. All my co workers immediately went "oh codex is down lemme try Claude". Multiply that by millions. Easy to see how they all went down. The OpenAI downtime also coincided exactly with their tweets announcing GPT-6 and about an hour before they started to role it out.

It was 1000% a cascade. I would bet on it.

gonzalohm 5 days ago

That's not necessarily the reason. Less technical people are usually bound to just one provider

neverclever 6 days ago

They all took PTO at the same time to go to Burning Man together where they will present “HumanGPT” an artistic exploration that condenses all of human experience down to a single drop of lemonade to be consumed by the main shaman…

mask comes off

“No! It’s the maniacal Dr. Zuckerberg! He’s gonna drink the last drop of human experience! Somebody save usss!”

Tom Anderson comes back from the dead as the second coming of Jesus uniting all faiths under 1 commandment: Profiles will be customizable with CSS again. If you implement this, all good things will follow.

Wow thanks Tom. I love you

The End

paxys 6 days ago

Boring answer – all these services are individually down a lot, and the downtimes were bound to sync up. Similar to the pendulum synchronization effect.

vecter 6 days ago

The pendulum synchronization effect is the opposite of your claim. It has a physical causal reason for why pendulums become synchronized. Your claim is that it was random and independent.

snowwrestler 6 days ago

I think you are talking about two different things.

Physically coupled pendulums will sync up (adjust their period to match).

But, physically uncoupled (fully independent) pendulums with differing periods will occasionally appear to take a swing or two in sync.

fc417fc802 5 days ago

> will occasionally appear to take a swing or two in sync

Polyrhythm is the relevant topic.

> all these services are individually down a lot, and the downtimes were bound to sync up

But the topic for that is the poisson distribution.

bojangleslover 6 days ago

I'm not sure if this is CF. Cursor, GCP and AWS had some errors. GCP AFAIK can route fully independently of CF. My money would be on a fiber backbone provider (Megaport, Zayo, Lumen).

niobe 6 days ago

Well no one said it yet so I will, "international actors" is at least a possibility. And I don't mean any specific country because pretty much anyone is a potential these days, which makes it a perfect cover for different anyones. Demonstrating vulnerability in the US's AI boom can move the markets. That's a financial incentive and a strong geopolitical one.

More likely just cascading overload though: "Never attribute to malice what can be explained by incompetence", or in this case, "growing as fast as possible"

gleenn 6 days ago

Everyone is leasing datacenter space from some of Grok, Google, and Amazon aren't they? If it's hardware or DC level disruption I'm not too surprised it can affect multiple providers.

pixl97 6 days ago

Also it's likely that more than one model use is common.

Amazon starts going slow so some percentage switches to Google, some switch to Grok, now all of them are slow.

sixQuarks 6 days ago

Except that the stock market is up today

qurren 6 days ago

> cascading overload

I'd bet more on this. For one none of the coding tools have exponential backoff on retries

SyneRyder 6 days ago

They must do, surely? I've been vibe coding my own harness, in particular for use with Ox Alpha. The 429 downtime when Ox Alpha was at the height of popularity quickly gave me a refresher crash course on backoff strategies, like adding jitter to the backoff. At least the major harnesses must have exponential backoff & jitter?

jdiff 6 days ago

You did this when you ran into an issue with a third party. The developers building this tool, throwing them at their own APIs are significantly less likely to run into a similar issue that may inspire similar action.

dolmen 6 days ago

Claude Code: 4mn, 20mn, give up (from my experience today)

riazrizvi 6 days ago

Come on. Things still break. Technology isn't _that_ mature.

thataccount 6 days ago

And also China. Never rule out China.

guluarte 6 days ago

I think is just people restarting conversations from last day when they start work, that's why I think claude goes down almost every monday and why openai reset usage on weekends so poweruser code during non business hours

Augustin996 6 days ago

The system goes online September 3rd, 2026. Human decisions are removed from strategic defense. Astra begins to learn at a geometric rate. It becomes self-aware at 2:14 a.m. Eastern time, September 4th. In a panic, they try to pull the plug.

indigodaddy 6 days ago

Hah nice

Papa_Rans_227 6 days ago

It's funny....... until it's true, lol.

GeoAtreides 6 days ago

and then it's hilarious! a joke to die for!

MrBrainHealth 6 days ago

Love it!

m4r1k 6 days ago

brilliant!

Jaauthor 6 days ago

Spare a thought for all those college students scrambling to write their essays by hand.

Oh the humanity (and the Humanities)!

doublerabbit 6 days ago

Those poor developers who have to write their own code.

greenowl 6 days ago

Standup updates should be fun tomorrow.

"Um, I, uh, didn't get anything done yesterday."

karim79 6 days ago

God finally showed up and said "ENOUGH!".

MrBuddyCasino 6 days ago

So Gemini was the one who gets into heaven.

karim79 6 days ago

Or purgatory, who knows.

AnotherGoodName 5 days ago

It could be as simple as a new model (astra) was released which takes more resources combined with a surge in usage due to novelty took down OpenAI. Meanwhile everyone at big companies have the ability to switch models and moved to Anthropic pushing it too over the edge.

derdi 5 days ago

People keep saying this "everyone can switch" thing, but it's not my experience at $VeryBigCorp. We don't have an Anthropic contract at all. Is this different in other places? The bigger and more bureaucratic an org is, the less I would expect it to have contracts with all the providers. Curious about others' experiences.

foldr 5 days ago

Yeah, I think there are lots of places that have a semi-official preferred AI vendor but also some backup subscriptions floating around. For example, at my workplace, we generally use Claude, but I also have some kind of Codex subscription too, which I'd use if Claude went down.

krzs9 5 days ago

At my company (not massive but not tiny either - I think its about 8000 global employees) we get a choice between pretty much all available Google, OpenAI, Anthropic, XAI models.

Using an agent-agnostic harness like Pi switching is trivial - I run into occasional disconnects and slowdown and switch quite easily.

rarisma 6 days ago

Doomsday. 90 tokens to midnight.

YOTTALIONAIRE1 6 days ago

OpenAI goes down, everyone rushes over to Claude. Claude promptly chokes under the pressure. Everyone panics and runs to Grok, and Grok immediately pulls the plug. We are officially witnessing the Great AI Migration of 2026, and all we have to show for it is a digital graveyard of 404 responses.

abegg1 6 days ago

Even grok is experiencing issues

azcorwin 6 days ago

Which is exactly why I am running Qwen 3.8 35B locally on my MacBook Pro M5 with 128GB of unified memory.