Reverse-lookup service exposed millions of photos of people’s faces

Reverse-lookup service exposed millions of photos of people’s faces

反向搜索服务泄露了数百万张人脸照片

When someone uploads a photo to the people-search tool ClarityCheck, the website has a clear message: “Your reverse image search is private and secure.” New research, though, shows that the website left more than 9 million image files, including photographs of people’s faces, publicly exposed. And a second misconfiguration publicly exposed people’s email addresses and phone numbers. 当有人将照片上传到人物搜索工具 ClarityCheck 时,该网站会明确显示:“您的反向图像搜索是私密且安全的。”然而,最新研究显示,该网站将超过 900 万个图像文件(包括人脸照片)公之于众。此外,另一个配置错误还导致用户的电子邮件地址和电话号码被公开泄露。

Overall, according to findings from independent security researcher Jeremiah Fowler, the exposed ClarityCheck database contained roughly 450 GB of images, including what appeared to be profile images, screenshots, and other photographs of adults, teenagers, and children. All of the images were stored in an unsecured Amazon S3 bucket, with files in folders named “faces” and “profiles,” which could be accessed by anyone online through a URL included in the company’s publicly available website code. 根据独立安全研究员 Jeremiah Fowler 的调查结果,泄露的 ClarityCheck 数据库总计包含约 450 GB 的图像,其中包括看起来像是个人资料图片、截图以及成人、青少年和儿童的其他照片。所有图像都存储在一个未受保护的 Amazon S3 存储桶中,文件存放在名为“faces”(人脸)和“profiles”(个人资料)的文件夹中。任何在线用户只要通过该公司公开网站代码中包含的 URL,即可访问这些文件。

ClarityCheck is one of a number of so-called people-finder tools that have appeared online in recent years. These websites broadly claim to be able to search the web, public records, and other databases to identify individuals. ClarityCheck’s website says it can run searches on phone numbers, email addresses, vehicle identification numbers, and names. Its photo-search page says it can help “identify anyone in a photo” and find social media profiles “in seconds.” ClarityCheck 是近年来出现在网上的众多所谓“寻人工具”之一。这些网站普遍声称能够搜索网络、公共记录和其他数据库来识别个人身份。ClarityCheck 的网站称,它可以对电话号码、电子邮件地址、车辆识别码和姓名进行搜索。其照片搜索页面则表示,它可以帮助“识别照片中的任何人”,并在“几秒钟内”找到社交媒体资料。

While ClarityCheck secured the giant image database after WIRED contacted the company in July, Fowler warns that it was seemingly exposed for months, and his initial efforts to flag the problem to the company were unsuccessful. Accidental data exposures create risk for any personal information, but particularly for sensitive and unchangeable biometric data like face images. 虽然在《连线》(WIRED)杂志于 7 月联系该公司后,ClarityCheck 对这个庞大的图像数据库进行了加密,但 Fowler 警告称,该数据库似乎已经暴露了数月之久,且他最初向该公司反馈问题的尝试均未成功。意外的数据泄露会对任何个人信息造成风险,尤其是对于像人脸图像这样敏感且不可更改的生物识别数据而言,风险尤甚。

And while ClarityCheck’s website requires people to attest that they have permission to upload photos to its site, Fowler points out that in practice, people whose faces were exposed may have had no idea that ClarityCheck held their image. After all, he notes, the service is explicitly designed for identification, and people don’t typically seek to identify themselves or people they know. 尽管 ClarityCheck 的网站要求用户证明其拥有上传照片的权限,但 Fowler 指出,在实际操作中,那些被泄露人脸信息的人可能根本不知道 ClarityCheck 持有他们的照片。他指出,毕竟该服务明确旨在进行身份识别,而人们通常不会去搜索识别自己或熟人。

“If you’re trying to find out who a person is, you might not have authorization or permission, so people might not know that their image had been dumped into this database that was public,” Fowler tells WIRED. “An AI bot could crawl it, extract faces, and use them for training. And there are lots of pictures of kids in there.” “如果你试图查明某人的身份,你可能并没有获得授权或许可,因此当事人可能根本不知道他们的照片已经被丢进了这个公开的数据库中,”Fowler 告诉《连线》。“人工智能机器人可以抓取这些数据、提取人脸并将其用于训练。而且里面还有很多儿童的照片。”

In a statement sent to WIRED, a spokesperson said that ClarityCheck appreciated Fowler’s efforts to alert the company about the issues. “Once this was drawn to the attention of the appropriate teams, we acted immediately to restrict access,” the spokesperson said. The company disputed any characterization that the data was “exposed,” saying that an “ordinary member of the public” would not have come across it. 在发给《连线》的一份声明中,发言人表示 ClarityCheck 很感谢 Fowler 提醒公司注意这些问题。“一旦相关团队注意到这一点,我们立即采取行动限制了访问,”发言人说道。该公司反驳了关于数据被“泄露”的说法,称“普通公众”不会偶然发现这些数据。

“We do not accept that data in the temporary storage location was ‘publicly exposed,’ which implies large-scale public access,” the spokesperson says. “Access required knowledge of a specific, unindexed URL that was not discoverable through ordinary use of the ClarityCheck service or a general web search.” “我们不接受关于临时存储位置的数据被‘公开泄露’的说法,因为这暗示了大规模的公众访问,”发言人表示。“访问需要知道特定的、未被索引的 URL,而通过正常使用 ClarityCheck 服务或常规网络搜索是无法发现这些 URL 的。”

The security industry broadly, as well as the US federal government specifically, considers data to be exposed if it could be accessed by people who are not intended to have access—particularly if it is reachable on the open Internet without being protected by an authentication requirement, such as a username and password. 整个安全行业,特别是美国联邦政府,认为如果数据可以被非预期人员访问,即视为数据泄露——特别是如果它可以在开放的互联网上被访问,且没有受到用户名和密码等身份验证要求的保护。

“Exposure is the state in which personal or sensitive data has been left accessible, discoverable, or otherwise put at risk of unauthorized access, whether or not anyone has yet taken or misused it,” says Mark Beare, head of consumer products at the security company Malwarebytes. “A publicly reachable database backup, a misconfigured storage bucket, or credentials sitting in a system that a researcher can reach are all exposures.” “泄露是指个人或敏感数据处于可被访问、可被发现或以其他方式面临未经授权访问风险的状态,无论是否有人已经获取或滥用了这些数据,”安全公司 Malwarebytes 的消费者产品主管 Mark Beare 表示。“可公开访问的数据库备份、配置错误的存储桶,或研究人员可以触及的系统中的凭据,都属于泄露。”

“There is no suggestion of malicious access, as the researcher notes,” the ClarityCheck statement continued. “The data involved includes duplicate, cropped, and resized copies of the same files, along with non-image data, not 9 million unique images.” The company added that it has “improved” its security reporting procedures to help other researchers contact the company in the future. “正如研究人员所指出的,没有迹象表明存在恶意访问,”ClarityCheck 的声明继续说道。“所涉及的数据包括同一文件的重复、裁剪和调整大小后的副本,以及非图像数据,并非 900 万张独特的图像。”该公司补充称,它已经“改进”了安全报告流程,以帮助其他研究人员在未来与公司取得联系。

In addition to the face data, ClarityCheck had also misconfigured its APIs such that its website URLs could be manipulated to reveal data about people simply by entering names; anyone using any consumer browser could have done this. Entering a name into one of the URLs would return multiple potential email addresses, physical addresses, and phone numbers for people with that name. After WIRED contacted the company, the URLs were secured. 除了人脸数据外,ClarityCheck 还错误配置了其 API,导致其网站 URL 可以通过输入姓名被操纵,从而泄露有关人员的数据;任何使用普通浏览器的人都可以做到这一点。在 URL 中输入姓名,就会返回该姓名对应的多个潜在电子邮件地址、物理地址和电话号码。在《连线》联系该公司后,这些 URL 已被加密。

The ClarityCheck spokesperson said in the statement that the details displayed were “sourced from publicly available information and licensed third-party data providers.” ClarityCheck 的发言人在声明中表示,所显示的信息“来源于公开信息和获得许可的第三方数据提供商”。

ClarityCheck’s face-search feature allows people to upload an image and then receive a “report” about where that image may appear online and who may be shown in the photo. When a WIRED reporter tested the system using their own face image, the website said it was “scanning facial landmarks” and “mapping unique face geometry” before matching the image to others online and offering a report that could include a full name, addresses, location history, public appearances, photos, videos, social media profiles, and “hidden dating profiles” for a fee. The resulting report named the reporter, provided a biography, and linked to multiple photos of them online. ClarityCheck 的人脸搜索功能允许用户上传一张图片,然后收到一份“报告”,内容包括该图片可能出现在网上的位置以及照片中可能出现的人。当《连线》的一名记者使用自己的脸部照片测试该系统时,网站显示它正在“扫描面部特征点”并“映射独特的人脸几何结构”,随后将该图像与网上的其他图像进行匹配,并提供一份付费报告,其中可能包括全名、地址、位置历史记录、公开露面信息、照片、视频、社交媒体资料以及“隐藏的约会资料”。最终生成的报告准确说出了记者的名字,提供了个人简介,并链接了他们在网上的多张照片。

Misconfigurations and accidental exposures are unfortunately common online, but as digital platforms offer more and more automated capabilities for collecting and analyzing sensitive personal data, the stakes grow ever higher for securing information. 不幸的是,配置错误和意外泄露在网上很常见,但随着数字平台提供越来越多的自动化功能来收集和分析敏感个人数据,保护信息的风险也随之变得越来越高。