Circuit Breaker Labs hopes to make AI safer for your kids (and you)

Circuit Breaker Labs hopes to make AI safer for your kids (and you)

Circuit Breaker Labs 希望让 AI 对你的孩子(以及你)更安全

With all the talk about how AI might one day kill us all, it’s easy to forget that AI has already been life-threatening to some, not through bioweapons, but psychologically. For example, Character.AI settled several wrongful death lawsuits earlier this year brought by families of underage users who died by suicide after interactions with its bots. Multiple families have also sued OpenAI over ChatGPT’s alleged role in their loved ones’ suicides and delusions.

尽管人们都在谈论 AI 有朝一日可能会毁灭人类,但我们很容易忽略一个事实:AI 已经对一些人构成了生命威胁——不是通过生物武器,而是通过心理影响。例如,Character.AI 今年早些时候就几起非正常死亡诉讼达成了和解,这些诉讼由未成年用户家属提起,这些用户在与该平台的机器人互动后自杀身亡。此外,多个家庭也起诉了 OpenAI,指控 ChatGPT 在其亲人的自杀和妄想症中扮演了角色。

Making AI safer across languages and cultures is the mission of Circuit Breaker Labs, one of TechCrunch’s 2026 Startup Battlefield 200 finalists. (Circuit Breaker will be pitching at TechCrunch Disrupt, which takes place this year at Moscone West in San Francisco from October 13-15.)

让 AI 在不同语言和文化背景下变得更安全,正是 Circuit Breaker Labs 的使命,该公司是 TechCrunch 2026 年度 Startup Battlefield 200 决赛入围者之一。(Circuit Breaker 将在 TechCrunch Disrupt 大会上进行路演,该大会将于今年 10 月 13 日至 15 日在旧金山的 Moscone West 举行。)

Founders Shirali and Arul Nigam, who are siblings, were motivated by Sewell Setzer, the 14-year-old who developed an emotional attachment to a Character.AI chatbot and confessed thoughts to it of harming himself before dying by suicide. The chatbot, the parents alleged in a 2024 lawsuit, encouraged him. The bot may not have understood what words like “I want to be with you” really implied, said Arul, who is Circuit Breaker Labs’ CTO.

创始人 Shirali Nigam 和 Arul Nigam 是一对兄妹,他们的创业动机源于 14 岁的 Sewell Setzer。Setzer 对一个 Character.AI 聊天机器人产生了情感依赖,并在自杀前向其倾诉了自残的想法。其父母在 2024 年的一起诉讼中指控该聊天机器人鼓励了他的自杀行为。Circuit Breaker Labs 的首席技术官 Arul 表示,机器人可能并不理解“我想和你在一起”这类话语背后的真正含义。

“A lot of people, especially young people, turn to these systems for support, and usually they aren’t actually getting the help they need. But in many cases, they’re actively being harmed, and people unfortunately have taken their lives already,” Arul said. “Those sorts of safety vulnerabilities, where people aren’t necessarily actively trying to break the system — they’re engaging in a natural way — and the system has context pollution or it doesn’t understand the nuance, and then takes really dangerous action, we’re trying to prevent that.”

“很多人,尤其是年轻人,转向这些系统寻求支持,但通常他们并没有得到真正需要的帮助。在许多情况下,他们反而受到了实质性的伤害,不幸的是,已经有人因此丧生,”Arul 说道。“这类安全漏洞——即用户并非刻意破坏系统,而是在进行自然交流时,系统因上下文污染或无法理解细微差别而采取了极其危险的行动——正是我们试图预防的。”

Circuit Breaker Labs has created AI agents that it likens to an army of crash-test dummies. These agents mimic folks from all ages, backgrounds, languages, and cultures, and are used to test models on their ability to detect dangerous, psychologically harmful interactions.

Circuit Breaker Labs 创建了一系列 AI 智能体,将其比作一支“碰撞测试假人”大军。这些智能体模拟了不同年龄、背景、语言和文化的人群,用于测试模型检测危险及心理伤害性互动能力。

“The way a six-year-old girl versus a 45-year-old man, or someone who speaks English as a first language versus a second language, or … gamer slang versus someone else who uses a different kind of slang, all of those can really trip up a model,” said Shirali, who is Circuit Breaker Labs’ CEO. “Models are really good at handling standard speech patterns, but nobody actually talks like that and so if the model misunderstands nuance or slang, it can go really badly.”

“六岁女孩与 45 岁男性的表达方式,以英语为母语者与非母语者之间的差异,或者游戏俚语与其他类型俚语的区别,所有这些都可能让模型陷入困境,”Circuit Breaker Labs 的首席执行官 Shirali 说。“模型非常擅长处理标准化的语言模式,但现实中没人会那样说话。如果模型误解了细微差别或俚语,后果可能会非常严重。”

The startup works with human domain experts to build its hyper-realistic user simulations in order to run “red-team” tests against models, which are adversarial tests meant to uncover weaknesses. The tests are built to reflect real human speech patterns, slang, coded language, and typos. Circuit Breaker Labs then runs tens of thousands to hundreds of thousands of simulated interactions per day. The idea is to ensure that a model can appropriately respond to risky interactions that may emerge over time and over many conversations.

该初创公司与人类领域专家合作,构建超逼真的用户模拟,以便对模型进行“红队”测试(即旨在发现弱点的对抗性测试)。这些测试旨在反映真实的人类语言模式、俚语、暗语和错别字。Circuit Breaker Labs 每天运行数万到数十万次模拟互动。其目的是确保模型能够针对随着时间推移和多次对话中可能出现的风险互动做出适当反应。

Circuit Breaker Labs then uses a proprietary scoring method to create auditable, explainable scores. Circuit Breaker Labs is currently operating as an AI safety testing lab for high-risk AI applications such as AI coaching, journaling, or other mental health support apps, though Arul declined to name its marquee customers.

随后,Circuit Breaker Labs 使用一种专有的评分方法来生成可审计、可解释的评分。目前,Circuit Breaker Labs 作为一家 AI 安全测试实验室,服务于 AI 教练、日记应用或其他心理健康支持应用等高风险 AI 领域,不过 Arul 拒绝透露其主要客户名单。

Although the startup has a working product, it is in the very early stages, with only five employees, including the Nigam siblings. Eventually, though, the testing platform could be applied to any app where someone may fall down an “AI psychosis” hole, where the human is at risk of developing a parasocial relationship with a chatbot. Examples include AI “co-worker” agents, whose responses can vary from one interaction to the next.

尽管该公司已有可用的产品,但仍处于非常早期的阶段,包括 Nigam 兄妹在内仅有五名员工。不过,该测试平台最终可以应用于任何可能导致用户陷入“AI 精神错乱”陷阱的应用,即人类面临与聊天机器人产生准社会关系风险的场景。例如 AI “同事”智能体,其回复可能会在不同互动中发生变化。

“People are becoming more skeptical of AI or more resistant to adopt it across the board,” Arul said, adding that while skepticism is healthy, banning a potentially valuable tool over safety concerns would be “regressive.” Circuit Breakers Labs believes the answer to those fears is making AI safer. “We want to help build that trust for people.”

“人们对 AI 变得越来越怀疑,或者对全面采用 AI 产生了抵触情绪,”Arul 说。他补充道,虽然怀疑态度是健康的,但仅仅因为安全担忧就禁止一种潜在的有价值的工具是“倒退的”。Circuit Breaker Labs 认为,解决这些恐惧的答案是让 AI 变得更安全。“我们希望帮助人们建立这种信任。”