I created an interactive digital avatar of myself — and you can talk to it
I created an interactive digital avatar of myself — and you can talk to it
我为自己创建了一个交互式数字分身——你可以和它对话
When Alexandru Voica, head of corporate affairs at the video-generation startup Synthesia, sent me a link this summer to the newest addition to their PR team, I was surprised. It was an interactive virtual avatar of him, trained to answer common press questions about Synthesia, like what it does and how it works. 今年夏天,当视频生成初创公司 Synthesia 的企业事务主管 Alexandru Voica 发给我一个链接,介绍他们公关团队的最新成员时,我感到非常惊讶。那是一个交互式的虚拟分身,经过训练可以回答有关 Synthesia 的常见媒体问题,比如它是做什么的以及它是如何运作的。
The day before, I was on a panel where PR people asked me if I minded pitches that used AI-generated text. But Alexandru’s avatar was beyond that — this seemed to me the final boss of using AI in PR. 就在前一天,我参加了一个小组讨论,公关人员问我是否介意使用 AI 生成文本的推介。但 Alexandru 的分身已经超越了这一点——在我看来,这似乎是 AI 在公关领域应用的“终极形态”。
In September, Synthesia invited me to its new office space in New York. Originally based in the U.K., Synthesia is a hot digital avatar startup alongside others like D-ID, HeyGen, and Colossyan. It hit a $4 billion valuation earlier this year and said last year it had crossed $100 million in ARR. 9 月,Synthesia 邀请我参观了他们位于纽约的新办公室。Synthesia 总部最初位于英国,是目前炙手可热的数字分身初创公司之一,与 D-ID、HeyGen 和 Colossyan 等公司齐名。今年早些时候,该公司的估值达到了 40 亿美元,并表示去年其年度经常性收入(ARR)已突破 1 亿美元。
Synthesia lets enterprises build interactive training videos with AI avatars and recently launched a product called Roleplay Sessions that lets employees practice, for example, sales pitches with an interactive AI avatar that responds and scores their responses. Synthesia 允许企业利用 AI 分身制作交互式培训视频,并于近期推出了名为“角色扮演会话”(Roleplay Sessions)的产品,让员工可以与交互式 AI 分身进行练习(例如销售推介),分身会做出回应并为他们的表现打分。
When I went to the new office opening and they asked me if I would like my own AI avatar, I didn’t even hesitate to say yes. Of course I would like my own digital twin. My outfit was cute that day, and my hair was in place. 当我参加新办公室开幕式时,他们问我是否想要一个自己的 AI 分身,我毫不犹豫地答应了。我当然想要一个自己的数字孪生体。那天我的装扮很可爱,发型也很到位。
Until I met my digital twin, I had been indifferent toward avatars, but I felt they would inevitably become part of everyday online life. I heard of people on Instagram creating them in their likenesses to help them make social content. I find it all to be very interesting, and it’s perhaps why I have no qualms now presenting to you my digital twin. 在见到我的数字孪生体之前,我对分身一直持冷淡态度,但我感觉到它们终将成为日常网络生活的一部分。我听说 Instagram 上有人根据自己的形象创建分身来辅助制作社交内容。我觉得这一切都非常有趣,这或许就是我现在毫不犹豫地向大家展示我的数字孪生体的原因。
This is the first time Synethsia has made a digital avatar for a journalist (or for anyone, period, outside of Voica). It is trained on my story about why venture-backed startups commit more fraud than non-VC-backed startups and will only answer questions about that story. Below, simply press “start in a new window” to begin. You can ask it questions like: Why did I decide to write this story? What is the research paper about? What did the researchers find? 这是 Synthesia 首次为记者(或者说除了 Voica 以外的任何人)制作数字分身。它基于我关于“为什么风投支持的初创公司比非风投支持的公司更容易欺诈”的报道进行了训练,并且只会回答与该报道相关的问题。在下方,只需点击“在新窗口中开始”即可。你可以问它诸如:我为什么决定写这篇报道?研究论文讲了什么?研究人员发现了什么?等问题。
To make this, I entered a mini film studio nestled inside Synthesia’s office where they took numerous photos of me and captured a two-minute recording of my voice. I had to consent to these avatars being made and, well, digital Dom was born. 为了制作这个分身,我进入了 Synthesia 办公室里的一间小型摄影棚,他们为我拍摄了大量照片,并录制了两分钟的语音。我必须同意制作这些分身,于是,“数字版 Dom”就这样诞生了。
They created a personal avatar for me (one that just reads whatever script I give it) — with and without glasses — and they made two interactive avatars for me (ones that can talk back and listen to me), also with and without glasses. We picked an article to train the interactive avatar on, and then one of the teams built my interactive avatar, which is powered by a combination of voice-to-text, video, language, and text-to-voice models. 他们为我创建了一个个人分身(只需读取我提供的任何脚本)——包括戴眼镜和不戴眼镜的版本——并为我制作了两个交互式分身(可以与我对话并倾听我的声音),同样也有戴眼镜和不戴眼镜的版本。我们挑选了一篇文章来训练这个交互式分身,随后团队构建了我的交互式分身,它由语音转文字、视频、语言和文字转语音模型共同驱动。
My avatar’s tech stack includes Synthesia’s own video and voice models, although the company also allows customers to choose alternatives from other labs like Cartesia, ElevenLabs, Google, or OpenAI. Enterprises can also choose to host their avatars on whatever cloud they want or pay Synthesia to host them. 我的分身技术栈包括 Synthesia 自有的视频和语音模型,尽管该公司也允许客户选择其他实验室(如 Cartesia、ElevenLabs、Google 或 OpenAI)的替代方案。企业还可以选择将分身托管在他们想要的任何云端,或者付费由 Synthesia 进行托管。
The voice-to-text model turns what people say into text, the agentic language model makes sense of text and can take actions based on it, the text-to-voice model turns a response into audio, and finally a video model (built by Synthesia) animates the avatar as it talks. 语音转文字模型将人们所说的话转化为文本,代理语言模型理解文本并据此采取行动,文字转语音模型将回复转化为音频,最后由视频模型(由 Synthesia 构建)在分身说话时为其制作动画。
Overall, Synthesia builds three types of products — a video-creation and distribution platform with classic avatars, where someone types in a script and the avatar repeats it; an agentic platform called Sessions where people can interact with the avatars in surveys or roleplay; and an API platform where people can take Synthesia video and voice models and combine them with other tech services to build interactive avatars or other types of products. 总的来说,Synthesia 构建了三种类型的产品:一个带有经典分身的视频创作与分发平台,用户输入脚本,分身即可复述;一个名为 Sessions 的代理平台,人们可以在调查或角色扮演中与分身互动;以及一个 API 平台,人们可以获取 Synthesia 的视频和语音模型,并将其与其他技术服务结合,以构建交互式分身或其他类型的产品。
It took the Synthesia team a couple of days to make my avatars. I played first with the personal ones and typed a fairly generic script to see how my AI voice sounded. I had it talk about how fall has arrived in New York, my favorite time of year. The voice was fairly accurate, and I was glad it didn’t pick up any of the hoarseness I had when I recorded my audio sample. I showed it to some non-tech friends who found it both interesting and creepy. Synthesia 团队花了几天时间制作了我的分身。我先试用了个人分身,输入了一段相当通用的脚本,看看我的 AI 声音听起来如何。我让它谈论纽约秋天的到来,那是我一年中最喜欢的季节。声音相当准确,我很高兴它没有录入我录制音频样本时那种沙哑的感觉。我把它展示给一些非科技圈的朋友看,他们觉得既有趣又有些毛骨悚然。
Then I showed them my interactive one, which is deterministic, meaning it will only say what it was trained to respond to. In this case, that was my venture fraud story. I asked it questions like where I was before TechCrunch and what part of New York I lived in, but each time it directed me back to the story. 然后我向他们展示了我的交互式分身,它是确定性的,这意味着它只会说出它被训练去回答的内容。在这种情况下,就是关于我的风险投资欺诈报道。我问它诸如我在加入 TechCrunch 之前在哪里、我住在纽约的哪个区之类的问题,但它每次都把我引导回那篇报道。
My friends didn’t think the voice sounded much like mine and thought the likeness was not as good as the personal avatar, but nonetheless it was close enough to be somewhat creepy. My mom called it “amazing,” which is pretty high praise from her. She and my father kept trying to ask it questions “only they would know” about me — but the model didn’t answer, redirecting them each time to the venture fraud story. “I don’t remember giving birth to two of you,” she joked after testing the model. 我的朋友们认为声音听起来不太像我,觉得相似度不如个人分身,但尽管如此,它已经足够逼真,让人感到有些诡异。我妈妈称它为“神奇”,这对她来说是相当高的评价。她和我父亲一直试图问它一些“只有他们才知道”关于我的问题——但模型没有回答,每次都把他们引向那篇风险投资欺诈报道。“我不记得我生了两个你,”她在测试完模型后开玩笑说。
This experience has made me think about what the future of journalism could be. Would people be OK with turning on the news and it being presented by an avatar? One investor told me no immediately. Certainly there is a lot of pushback today on AI slop that has infiltrated social media and other news-sharing platforms. But others I asked weren’t so sure. Could avatars augment — or even replace — journalists? Would CEOs want to talk to an AI avatar of a journalist rather than a human one? 这次经历让我思考新闻业的未来会是什么样。如果人们打开新闻,看到的是由分身播报的,他们能接受吗?一位投资者立刻告诉我不能。当然,目前对于渗透到社交媒体和其他新闻共享平台上的“AI 垃圾内容”存在很多抵制。但我问过的其他人却不那么确定。分身能增强甚至取代记者吗?首席执行官们会愿意与记者的 AI 分身交谈,而不是与真人交谈吗?
What I like about my career is connecting with people, writing stories, and researching new topics. But the biggest part of journalism is trust. That doesn’t seem like it could ever be outsourced to an AI. Outside of journalism, I’m confident the idea of cloning yourself could be appealing. No more catching up on work after a holiday or vacation, because a version of you can always be around, answering questions. 我热爱我的职业,因为它能让我与人建立联系、撰写报道并研究新课题。但新闻业最核心的部分是信任。这似乎永远无法外包给 AI。在新闻业之外,我相信“克隆自己”这个想法会很有吸引力。假期或休假后不再需要赶进度,因为总有一个版本的你可以随时待命,回答问题。