Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

Ask HN: Why were OpenAI, Claude, and Grok simultaneously down?

Ask HN:为什么 OpenAI、Claude 和 Grok 会同时宕机?

Cloudflare, Azure, AWS, and Google Cloud all have a similar uptick in reported errors around 7:30. I suspect an outage on Cloudflare or another load-bearing service cascaded through all the major cloud providers. Cloudflare、Azure、AWS 和 Google Cloud 在 7:30 左右都出现了类似的错误报告激增。我怀疑是 Cloudflare 或其他核心承载服务出现了故障,并波及了所有主要的云服务提供商。

I think down detector doesn’t actually have any probes or actual insight into status etc. I think it uses search volume on its own service as a proxy for an outage - so if e.g. lots of people rush to down detector to query to see if SERVICE_FOO is down, it will register as an outage on down detector because loads of people are trying to see if there is an outage even if SERVICE_FOO is actually totally fine. 我认为 Down Detector 实际上并没有任何探测器或对服务状态的真实洞察。我认为它只是将自身服务的搜索量作为宕机的代理指标——因此,如果大量用户涌入 Down Detector 查询某个服务是否宕机,它就会在 Down Detector 上显示为宕机,即便该服务本身完全正常。

My hunch is everyone saw that openai and Claude were down and checked for Gemini too. I was using Gemini the whole time this happened without a blip so it certainly wasn’t down in my region at least. 3.8 flash is pretty good and didn’t miss a beat. 我的直觉是,大家看到 OpenAI 和 Claude 宕机后,也顺便检查了 Gemini。在整个事件期间,我一直在使用 Gemini,它没有任何异常,至少在我所在的地区它没有宕机。3.8 Flash 模型表现相当不错,完全没有受到影响。

So Down Detector is a kind of quantum observation experiment, observing influences the result. 所以 Down Detector 是一种量子观测实验,观测行为本身影响了结果。

Down detector is a self-reported platform, they use a baseline over 6 months from user reports to try to automatically “detect” if there is a real outage, or if its just a couple end-users. Ultimately, the outage is entirely based on users going to down detector and clicking on “Report a Problem”. Down Detector 是一个基于用户自报的平台,他们使用过去 6 个月的用户报告基准,试图自动“检测”是否存在真正的宕机,还是仅仅只有几个终端用户遇到问题。归根结底,宕机状态完全取决于用户是否前往 Down Detector 并点击“报告问题”。

It’s really interesting to crawl through the web of phrases that seem “common” to each person in their interactions… I suspect if one were to get to the more niche “meme” phrases that people encounter, it starts to say more about the sort of person the LLM assumes it’s speaking to, and perhaps something about their psychological profile… 深入挖掘那些在互动中对每个人似乎都“通用”的短语真的很有趣……我怀疑,如果能深入到人们遇到的更小众的“梗”短语,它就能反映出 LLM 认为它正在与什么样的人交谈,甚至可能反映出对方的心理侧写……

It must be tuned for rage-based engagement because Claude only ever uses the most obnoxious Claudisms on me despite constant reminders to knock it off. 它一定是针对基于愤怒的互动进行了调优,因为尽管我不断提醒它停止,Claude 对我时依然只会使用那些最令人讨厌的“Claude 式”用语。

This entire thread is amazing! It’s like a kaleidoscope: Human words —> LLM Training —> chatbot-isms —> lovely parody catch-phrases (and this thread will get ingested soon, and be used to train…) 整个讨论串太精彩了!这就像一个万花筒:人类语言 —> LLM 训练 —> 聊天机器人用语 —> 可爱的模仿口头禅(而且这个讨论串很快就会被抓取,并用于训练……)

Cloudflare CTO claims that it’s not them. Cloudflare 的首席技术官声称这与他们无关。