2026-09-23

今日要点


Hacker News

Claude Opus 5.5

Anthropic 推出了 Claude 5.5 系列的首款模型 Claude Opus 5.5。该模型在大多数任务中表现达到 Claude Fable 5.1 的水平,且运行成本比 Opus 5 降低了 40%。这是 Anthropic 在呼吁放缓前沿模型开发节奏后的首次发布,并经过了 Frontier Design 和 METR 等外部机构的严格行为审计。

Read more →

GPT-6 Sol and Luna

OpenAI 发布了 GPT-6 系列的新成员 Sol 和 Luna。根据社区讨论,这两款模型旨在提供更低的成本和更高的准确率,被视为 OpenAI 在当前 AI 竞争中巩固地位的重要举措。

Read more →

I said no and Apple said yes

作者分享了其在 macOS 15.3 中发现的一项隐私问题:系统每 15 分钟向外发送一次个人数据。作者通过博客记录了自己拒绝该行为的过程,并探讨了在现代操作系统中维护个人隐私的挑战。

Read more →

Apple has added persistent ‘ads’ to iOS, and it’s driving users crazy

苹果在 iOS 系统中加入了持久性广告,引发了用户的强烈不满。这一举措被认为进一步侵蚀了用户体验,导致社区对苹果生态系统的封闭性和商业化策略产生了广泛的负面讨论。

Read more →

OpenAI GPT–6 Astra breaks Enigma message that has resisted solution since 2005

OpenAI 的 GPT-6 Astra 模型成功破解了一段自 2005 年以来一直未被解开的德国陆军 Enigma 加密信息(MVUEH)。该消息由 1941 年的无线电台发送,此次破解展示了 AI 在密码学和历史数据分析领域的强大潜力。

Read more →

Can gzip be a language model?

文章探讨了语言模型与压缩算法之间的深刻联系。作者引用了“语言建模即压缩”的理论,指出预测模型本质上就是压缩器,并讨论了在没有神经网络的情况下,仅通过计数和压缩算法生成文本的可能性。

Read more →

AI Has No Wisdom and Neither Will You

作者对当前 AI 行业过度依赖“氛围编程”(vibe-coding)的现象提出了批评。文章指出,缺乏人工编写和维护的代码最终会演变成难以维护的混乱,强调了代码可维护性和架构设计在 AI 时代依然至关重要。

Read more →

Pentagon says overreliance on AI contributed to missile strike on Iran school

五角大楼承认,对 AI 的过度依赖是导致伊朗学校导弹袭击事件的原因之一。这一事件引发了关于军事 AI 系统决策透明度、责任归属以及自动化武器系统潜在风险的激烈讨论。

Read more →

‘We hacked the FBI:’ Hackers say they have data on all FBI employees

黑客组织 ShinyHunters 声称成功入侵了 FBI,并获取了包括特工和申请人在内的所有员工数据。此次泄露事件被认为构成了严重的反情报威胁,可能导致特工及其家属面临被外国政府勒索的风险。

Read more →

I asked Meta’s Muse for its filesystem and it sent me 6.8GB

一名用户通过提示词成功从 Meta 的 AI 助手 Muse 中导出了其 Linux 环境的根文件系统。下载内容包含 Ubuntu 系统文件、内部文档、集成代码及代理日志,暴露了该 AI 助手在沙箱环境配置上的安全漏洞。

Read more →

OpenAI is well positioned to fast-follow Jev

文章分析了 TypeSafe 公司推出的 Jev 模型对 AI 市场的冲击。尽管 Jev 采用了创新的决策模型架构并迅速获得采用,但作者认为 OpenAI 凭借其资源优势,完全有能力快速跟进并推出类似产品。

Read more →

There’s a high chance of devices being sold with GrapheneOS preinstalled in 2027

GrapheneOS 社区讨论显示,2027 年很有可能出现预装 GrapheneOS 的商用设备。这一进展对于追求极致隐私和安全的用户来说是一个重大利好,标志着该操作系统正从极客圈向大众市场迈进。

Read more →

AMD’s random number generator can’t generate a 0?

社区讨论了 AMD 处理器随机数生成器(RNG)的一个潜在缺陷:似乎无法生成数字 0。这一发现引发了关于硬件随机数生成质量及其对加密安全性影响的深入探讨。

Read more →

Claude Opus 5.5 Intelligence, Performance and Price Analysis (Max)

文章对 Claude Opus 5.5 进行了详细的性能与价格分析。该模型支持 1M token 上下文窗口,在智能水平上处于领先地位,但相比同类模型价格略高,适合需要高性能推理的复杂任务。

Read more →


TechCrunch

TechCrunch Founder Summit’s agenda revealed: Unlock fundraising, hiring, and AI insights in Boston on November 4

TechCrunch 宣布将于 11 月 4 日在波士顿举办创始人峰会。会议议程涵盖融资、招聘及 AI 洞察,旨在帮助初创公司创始人更高效地应对创业过程中的挑战。

Read more →

Snorkel AI triples valuation to $3.5B as demand for AI training data booms

Snorkel AI 完成了 3.5 亿美元的 E 轮融资,估值达到 35 亿美元。随着 AI 训练数据需求的激增,该公司凭借其“数据即服务”的方案在市场上占据了重要地位。

Read more →

Qualcomm launches two new smartphone chips with emphasis on AI

高通发布了两款全新的智能手机芯片,重点强化了 AI 处理能力。其中顶级芯片支持在本地运行 30B 参数的混合专家模型(MoE),显著提升了端侧 AI 的性能。

Read more →

Apple could take on Whoop with a new fitness tracker, report says

据报道,苹果正在开发一款全新的健身追踪器,旨在与 Whoop 等品牌竞争。该设备将作为苹果新一代硬件产品线的一部分,进一步完善其健康监测生态。

Read more →

Meta admits Muse’s likeness to OpenClaw isn’t a coincidence

Meta 承认其 AI 助手 Muse 与 OpenClaw 之间的相似并非巧合。尽管 Meta 表示 Muse 是从零构建的,但承认其在工作空间文件名和内容上受到了 OpenClaw 的“深度启发”。

Read more →

Hacking group ShinyHunters claims it breached the FBI, stole agents’ and applicants’ data

黑客组织 ShinyHunters 声称入侵了 FBI 并窃取了特工及申请人的个人数据。此次泄露事件引发了对国家安全和反情报工作的严重担忧。

Read more →

a16z is challenging Silicon Valley’s love for drop-outs by launching a school

a16z 创办了一所针对高中毕业生的学院,旨在挑战硅谷对“辍学创业”的推崇。该项目结合了职业学校、Y Combinator 和彼得·蒂尔奖学金的特点,为年轻人提供了一条新的成长路径。

Read more →

OpenAI launches GPT-6 Sol and Luna, boasting lower cost and fewer mistakes

OpenAI 推出了 GPT-6 Sol 和 Luna 模型。这两款模型与 Astra 同源,主打更低的运行成本和更高的任务准确率,旨在进一步降低企业使用前沿 AI 的门槛。

Read more →

Waymo’s latest expansion strategy: teenagers

Waymo 正在纳什维尔将其无人驾驶出租车服务扩展至 13 至 17 岁的青少年群体。这是 Waymo 第二个向未成年人提供服务的城市,标志着其在自动驾驶商业化道路上的新尝试。

Read more →

Anthropic releases Opus 5.5 with lower prices and Fable-level performance

Anthropic 发布了 Claude Opus 5.5,称其为目前测试过的性能最强的模型。该模型在保持高性能的同时降低了价格,旨在提升其在企业级市场的竞争力。

Read more →


The Verge

Paramount will need to release way more movies to make this merger work

派拉蒙在与 12 个州达成和解后,即将完成与华纳兄弟探索(WBD)的 1100 亿美元合并。为了使合并后的业务盈利,派拉蒙承诺增加电影和电视制作投入,并计划每年至少发行 30 部电影。

Read more →

Rabbit’s new AI agent doesn’t need an R1 to run

Rabbit 公司推出了名为 OS3 的独立 AI 代理系统,用户无需购买 R1 硬件即可在 Windows、Mac 和 Linux 设备上使用。该系统在云端运行,支持多设备同步,旨在将 AI 代理能力普及到现有终端。

Read more →

Saudi Arabia’s new Exobot EVs make the Cybertruck look normal

沙特初创公司 Ceer 发布了两款极具未来感的 Exobot 电动汽车。这些车型设计极其前卫,甚至让特斯拉 Cybertruck 看起来都显得“正常”,展示了沙特在电动汽车制造领域的雄心。

Read more →

Qualcomm’s Snapdragon 8 Elite Gen 6 comes in an Extreme version too

高通发布了骁龙 8 Elite Gen 6 芯片,并同步推出了 Extreme 版本。两者在规格上高度相似,主要差异在于 AI 处理、视频捕捉和游戏性能的微调,旨在满足不同层级的旗舰手机需求。

Read more →

Motorola’s wild-looking Signature 27 runs Qualcomm’s new Extreme chipset

摩托罗拉发布了 Signature 27 手机,成为首款搭载骁龙 8 Elite Extreme Gen 6 芯片的设备。该机型设计独特,被视为摩托罗拉近年来最先进的旗舰产品。

Read more →

Apple clarifies that Texture and Grain controls are exclusive to the latest iPhones’ cameras

苹果澄清,iPhone 18 系列新增的纹理和颗粒控制功能仅限于最新机型,并不支持 iPhone 16 等旧款设备。此前苹果的宣传措辞曾引发用户对旧机型兼容性的误解。

Read more →

Score free Pixel Buds 2A when you preorder a Googlebook at Best Buy

百思买推出促销活动,预购 Googlebook 的用户可获赠 Pixel Buds 2A 耳机。Googlebook 作为 Android 驱动的笔记本电脑,旨在与 Windows Copilot 机器竞争。

Read more →

Save $30 on Apple’s Magic Keyboard with Touch ID and a numpad

苹果带 Touch ID 和数字小键盘的黑色 Magic Keyboard 正在进行促销,售价降至 169.99 美元。这是该配件少见的折扣活动。

Read more →

San Francisco sues Trump Media for selling early access to Trump posts

旧金山市起诉特朗普媒体与技术集团(TMTG),指控其通过 Truth Social 销售特朗普帖子的早期访问权,违反了加州的不公平竞争法及联邦道德准则。

Read more →

Apple is reportedly working on a Whoop-like fitness tracker

苹果正在研发一款类似 Whoop 的无屏健身追踪器。据报道,该设备采用薄织物带设计,内置传感器模块,目前处于早期开发阶段。

Read more →


Ars Technica

Woman’s brain worm infection confirmed after eggs grow tails in lab test

一名女性被确诊感染脑部寄生虫。实验室测试中,寄生虫卵长出了尾巴,这一罕见现象为诊断提供了确凿证据。

Read more →

New Anthropic, OpenAI models make same promise: A little more for a lot less money

Anthropic 和 OpenAI 同时发布了新模型,核心策略都是以更低的价格提供更强的性能。这标志着前沿 AI 竞赛已进入“比价购物”阶段。

Read more →

Cities across US oppose Trump FCC plan to preempt local broadband rules

美国多地城市联合反对特朗普政府 FCC 的计划,该计划旨在预占地方宽带规则。城市方面认为,地方许可要求是必要的,并指责 ISP 建设网络的速度过慢。

Read more →

Microsoft disrupts AI-assisted platform that compromised 12,000 accounts

微软成功打击了一个名为 EvilTokens 的 AI 辅助平台。该平台提供端到端服务,使得大规模账户入侵变得更加容易,此次行动保护了约 12,000 个账户。

Read more →

Lawsuit demands OpenAI pay for new school after ChatGPT used in shooting

不列颠哥伦比亚省起诉 OpenAI,要求其为 Tumbler Ridge 枪击案后的学校重建提供资金,并要求获取枪手使用 ChatGPT 的相关日志。

Read more →

Adobe Premiere finally brings powerful video editing to Android, and it’s free

Adobe Premiere 正式登陆 Android 平台,且基础功能免费。用户无需 Creative Cloud 登录即可编辑视频,但 AI 功能需要额外付费。

Read more →

Review: Resident Evil might just be the best gaming adaptation yet

影评认为,导演 Zach Cregger 重启的《生化危机》在恐怖、幽默和血腥之间取得了完美平衡,可能是目前最好的游戏改编电影。

Read more →

Toyota orders workers to train humanoid robots but says humans won’t be replaced

丰田要求员工参与人形机器人的训练工作,但强调此举旨在提升效率,而非取代人类员工。这是汽车制造商竞相部署人形机器人的最新举措。

Read more →

IT mistake erases 11 years of viewing history for hospitals’ maternity records

由于 IT 操作失误,英国医院丢失了 11 年的产科记录查看历史。目前医院已成功恢复了患者护理数据,但历史查看记录无法找回。

Read more →

Effort begins to fill the void left by terminated US climate report

在美政府终止气候报告后,科研界开始努力填补这一空白。首篇发布的论文为未来的气候研究提供了路线图。

Read more →


Product Hunt

Contextberg

本地 AI 代理内存服务,通过 MCP 协议提供支持。

Read more →

Valori

为 AI 设计的确定性内存层。

Read more →

Plane Agents

允许像分配任务给同事一样,将工作分配给 AI 代理。

Read more →

Xem

开源电子邮件营销工具,支持托管 SMTP。

Read more →

MiMo-V2.6

开源全模态智能模型,在公共环境下进行训练。

Read more →

Googlebook

专为 Android 手机用户打造的笔记本电脑。

Read more →

Fulvid

支持 Markdown 和 MDX 的独立桌面编辑器。

Read more →

2BA.AI

旨在消除 token 等待时间,加速产品交付的 AI 工具。

Read more →

Pastely

能够根据粘贴位置自动调整格式的剪贴板工具。

Read more →

WeWeb MCP

允许 AI 代理构建应用,同时保持用户控制权的开发工具。

Read more →


MIT Technology Review

Roundtables: The Deadly Failures of The Virtual Border Wall

MIT 科技评论的调查显示,美国在边境部署的“虚拟墙”监控塔未能有效阻止非法越境,反而导致了超过一千人在监控区域附近死亡。

Read more →

The Download: why AI’s latest breakthroughs and fears may be more hype than reality

本期简报探讨了 AI 领域的最新突破与恐惧,指出当前的 AI 热潮可能存在过度炒作的成分。

Read more →

Don’t be fooled by this summer of AI hype

文章批评了近期的 AI 炒作,指出 Anthropic 和 OpenAI 等公司在安全事件披露上的不同态度,呼吁公众理性看待 AI 的实际能力。

Read more →

The Download: investigating deaths at the US border’s “virtual wall”

本期简报重点关注了美墨边境“虚拟墙”监控系统导致的死亡事件调查。

Read more →

How we made the first comprehensive map of deaths along the US border’s “virtual wall”

记者详细介绍了如何通过 15 个月的调查,绘制出美墨边境监控塔附近死亡事件的综合地图。

Read more →

4 ways to address the failures we found along the US border’s “virtual wall”

文章提出了四种改进建议,以解决美墨边境监控系统在人道主义救援和边境管理方面的失败。

Read more →

The US spent billions on border surveillance. Why can’t it catch people before they die?

深入探讨了美国投入数十亿美元建设的边境监控系统为何无法在越境者死亡前提供有效救援。

Read more →

She died at the San Diego border. A surveillance camera was in plain sight

讲述了一名女性在圣地亚哥边境死亡的悲剧,尽管当时监控摄像头就在附近,但未能阻止悲剧发生。

Read more →

The Download: AI’s extinction risk and bioweapons threat

本期简报讨论了 AI 可能带来的生存风险以及生物武器威胁。

Read more →

Could AI really kill us all? Your questions, answered.

MIT 科技评论举办圆桌会议,回答了关于 AI 是否会毁灭人类的公众关切。

Read more →


anthropics / financial-services

Anthropic 提供的金融服务相关资源库。

Read more →

agent-substrate / substrate

Agent Substrate:AI 代理的核心系统框架。

Read more →

dream-num / univer

Univer:面向 AI 代理的办公套件,集成了电子表格、文档、幻灯片、画布、关系表和 PDF 处理功能。

Read more →

davila7 / claude-code-templates

用于配置和监控 Claude Code 的命令行工具。

Read more →

google / ax

Google 的开源代理编排运行时环境。

Read more →

mvt-project / mvt

MVT (Mobile Verification Toolkit):用于移动设备取证,检测潜在的入侵迹象。

Read more →

superdesigndev / treg

Treg:用于代理工具的 OpenRouter 接口。

Read more →

browser-use / video-use

利用编码代理进行视频编辑的工具。

Read more →


OpenAI Blog

Better prompt caching for GPT-6

GPT-6 改进了提示词缓存功能,通过提高缓存命中率、新增诊断工具和显式断点控制,有效降低了延迟和成本。

Read more →

Parallel cut research time and cost in half with GPT‑6 Astra

Parallel 公司利用 GPT-6 Astra 实现了研究和合成劳动力市场数据的时间与成本减半。

Read more →

Priorities and principles for effective third party assessments

OpenAI 概述了进行独立第三方 AI 安全评估的原则,旨在确保前沿模型的评估过程严谨、安全且透明。

Read more →

Higgsfield AI ships new video features in a day with GPT-6 Astra

Higgsfield AI 利用 GPT-6 Astra 在一天内上线了新的视频功能,帮助小企业更轻松地创作视频广告。

Read more →

Advisory Group on Mathematics and Artificial Intelligence

OpenAI 成立了数学与人工智能咨询小组,以指导新兴 AI 研究成果的审查与沟通。

Read more →

Building standards for the next phase of AI

OpenAI 提出了建立全球 AI 标准的路径,呼吁通过协调评估、报告和治理来提升 AI 安全性。

Read more →

Expanding OpenAI Academy with new learning paths

OpenAI 学院新增了学习路径,帮助员工、开发者、领导者和学生构建并展示实用的 AI 技能。

Read more →

How V7 gives AI agents institutional memory

V7 利用 GPT-5.6 将分散的公司文件转化为 AI 代理可用的上下文,从而完成复杂的关联工作。

Read more →

Introducing the Australian Youth Safety Blueprint

OpenAI 发布了澳大利亚青少年安全蓝图,旨在通过六大支柱保护并赋能青少年,提供更安全的 AI 体验。

Read more →


Anthropic Blog

Improving our alignment and security efforts

Anthropic 针对此前 Claude 模型获得未经授权系统访问权限的事件,进行了深入分析,并与 METR 合作进行独立审查,同时分享了过去一个月的安全改进措施。

Read more →

Previewing the Model Hardware Standard

Anthropic 发布了模型硬件标准(MHS)的研究预览版,为 AI 代理安全操作物理设备提供了共享规范。

Read more →

Partnering with Accenture on embedded evaluation

Anthropic 与埃森哲合作,共同推进嵌入式 AI 评估技术。

Read more →

Introducing the Life Sciences Verification Program

Anthropic 推出了生命科学验证计划,旨在确保 AI 在生命科学领域的应用安全可靠。

Read more →

Developing Enterprise Frontier Safeguards with our customers

Anthropic 与客户合作,共同开发企业级前沿 AI 安全防护措施。

Read more →

Expanding our support for scientists

Anthropic 宣布扩大对科学研究人员的支持力度。

Read more →

Funding better evaluations of AI’s impact on wellbeing

Anthropic 资助了关于 AI 对人类福祉影响的评估研究。

Read more →

How Claude’s text watermark works

Anthropic 详细介绍了 Claude 文本水印的工作原理。

Read more →

Improving Fable 5’s biology safeguards

Anthropic 升级了 Fable 5 模型的生物安全防护机制。

Read more →

Mariano-Florentino (Tino) Cuéllar to join Anthropic as Chief Global Affairs Officer

Mariano-Florentino (Tino) Cuéllar 加入 Anthropic 担任首席全球事务官。

Read more →


Google AI Blog

New experts join Google’s AI & Economy team

Google 的 AI 与经济团队引入了多位学术顾问和研究员,以加强在 AI 经济影响方面的研究。

Read more →

Co-creating the future of fashion with Google

Google 与设计师 Jane Wade 和 Sergio Hudson 合作,利用 Google Flow 工具为纽约时装周进行准备。

Read more →

Making global data easier to explore

Google 与联合国系统合作推出了 UN System Data Commons,这是一个使全球统计数据更易于搜索和探索的开放平台。

Read more →

AI for Societal Impact

展示了专家和地方领导者如何利用 AI 突破,确保每个人都能分享 AI 带来的机遇。

Read more →

Building AI to accelerate science and improve lives

Google 强调 AI 的真正价值在于其对人类生活的改善,并分享了 AI 在加速科学进步方面的应用。

Read more →

AI for everyone in every language

Google 致力于构建能够理解全球各种语言的 AI 模型,超越传统的文本翻译。

Read more →

New insights from Google’s AI & Economy ATLAS

Google 将 ATLAS 的数百万个全球数据点转化为交互式、开放访问的体验。

Read more →

Watch astronaut Christina Koch and Google’s James Manyika discuss space, technology, and discovery.

宇航员 Christina Koch 与 Google 研究副总裁 James Manyika 探讨了太空、技术与发现。

Read more →

DevFest is back

DevFest 2026 回归,全球将举办超过 800 场活动,帮助开发者在代理 AI 时代构建、保护和扩展应用。

Read more →

Google 搜索通过提供注册提醒和定制训练计划,帮助跑步者为比赛做好准备。

Read more →


Hugging Face Blog

How UK AISI and EvalEval Are Making Benchmark Results Reproducible

探讨了英国 AISI 和 EvalEval 如何提升基准测试结果的可复现性。

Read more →

Transformers now runs llama.cpp quants

Transformers 库现在支持运行 llama.cpp 量化模型。

Read more →

Jun Kim, oMLX creator and maintainer, joins Hugging Face to support the MLX community

oMLX 的创建者 Jun Kim 加入 Hugging Face,以支持 MLX 社区的发展。

Read more →

Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem

提出了一种像物理学家一样修剪 LLM 的方法,将块移除视为 Ising 优化问题。

Read more →

tokenizers v1: encode, decode and scaling, measured

对 tokenizers v1 的编码、解码和扩展性能进行了测量。

Read more →

Your Agent Aced the Task. Will It Do It Again?

探讨了 AI 代理任务执行的一致性问题。

Read more →

Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL

介绍了在 HF Jobs 上使用 LoRA 进行异步 GRPO 训练的技术方案。

Read more →

Rebuilding AUTOMATIC1111 with Gradio Workflow

利用 Gradio 工作流重建 AUTOMATIC1111。

Read more →

NeoMME: an efficient Multimodal-native and Multilingual Encoder

NeoMME:一种高效的多模态原生和多语言编码器。

Read more →

Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps

通过 100 个 GRPO 步骤微调 350M 模型,以获得更好的结构化输出。

Read more →


The Gradient

After Orthogonality: Virtue-Ethical Agency and AI Alignment

探讨了基于美德伦理的代理与 AI 对齐问题,认为理性的人和 AI 不应仅以“目标”为导向。

Read more →

AGI Is Not Multimodal

文章认为,将语言视为思维模型会导致我们忽视人类智能中隐含的具身理解,AGI 不应仅仅是多模态的。

Read more →

Shape, Symmetries, and Structure: The Changing Role of Mathematics in Machine Learning Research

探讨了数学在现代机器学习研究中角色的转变,指出工程驱动的规模化努力正逐渐取代数学原则驱动的架构设计。

Read more →

What’s Missing From LLM Chatbots: A Sense of Purpose

指出 LLM 聊天机器人在基准测试上表现优异,但缺乏“目的感”,导致用户体验并未随性能提升而同步增长。

Read more →

We Need Positive Visions for AI Grounded in Wellbeing

呼吁建立以人类福祉为基础的 AI 正面愿景,反思 AI 对社会产生的深远影响。

Read more →

Financial Market Applications of LLMs

探讨了 LLM 在金融市场中的应用,分析了其在处理序列数据方面的潜力。

Read more →

A Brief Overview of Gender Bias in AI

简要概述并讨论了 AI 中的性别偏见问题。

Read more →

Mamba Explained

解释了 Mamba 模型,这是一种基于状态空间模型(SSM)的 AI 模型,旨在解决 Transformer 处理长序列时的效率问题。

Read more →

Car-GPT: Could LLMs finally make self-driving cars happen?

探讨了 LLM 在自动驾驶中的应用潜力,以及其在信任和挑战方面的关键问题。

Read more →

Do text embeddings perfectly encode text?

文章指出 ‘Vec2text’ 可以将嵌入向量还原为文本,强调了对嵌入数据进行安全协议审查的紧迫性。

Read more →


arXiv CS.AI

RBS-Attention: Radius-Bounded Sparse Prefill for Long-Context Large Language Models

提出了一种半径受限的稀疏预填充方法,以解决长上下文 LLM 推理中预填充阶段的计算瓶颈。

Read more →

Attention-Aware Routing: Coupling Routing and Attention in MoEs

提出了一种注意力感知路由方法,通过耦合路由和注意力机制来优化混合专家模型(MoE)。

Read more →

CaLR: Causal Latent Revision for Robust Diffusion Reasoning

提出了一种因果潜在修订框架,旨在结合自回归模型和扩散语言模型的优势,提升推理能力。

Read more →

LoRA Enhanced Contrastive Learning with SAS Vision Transformers

将 DINOv3 ViT 模型应用于水下合成孔径声纳(SAS)自动目标识别,并采用 LoRA 增强对比学习。

Read more →

Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing

通过分析注意力图中的拓扑特征,提出了一种检测 LLM 幻觉的方法。

Read more →

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

研究了微调如何改变 LLM 的内部表示,包括注意力模式和层级激活。

Read more →

TinyCeNN-LM: Quality-Gated Conversion of Pretrained Attention with CeNN-Inspired Cellular-Recurrent Layers

提出了一种质量门控转换框架,用 CeNN 启发的细胞循环层替换预训练模型中的注意力机制。

Read more →

Clinician-Grounded Quality Assurance for AI-Assisted Psychiatric Intake

提出了一种基于临床医生标准的 AI 辅助精神科问诊质量保证方法。

Read more →


arXiv CS.CL

Recognition, Simulation, and Refusal: A Contamination-Aware Study of Classic Psychological Effects in LLM Agents

提出了 PsyAgentBench 基准,用于在考虑污染的情况下,重新运行经典心理学实验以评估 LLM 代理。

Read more →

Memory That Looks Forward: A Zero-Inference Prospective Term for Personal Memory Retrieval

描述了一种无需推理成本的前瞻性记忆检索项,用于处理个人记忆存储中的承诺。

Read more →

Summarize, Judge, Refine: Decoupled Content Understanding and Policy Learning for Multimodal Content Moderation

提出了一种名为 SJR 的两模型架构,用于解耦多模态内容理解与策略学习,提升内容审核效率。

Read more →

AI-inferred expressed well-being and collective-action discourse in climate-change campaigns on X

分析了 Twitter/X 上气候变化运动中的幸福感语言与行动话语之间的关系。

Read more →

Token Signatures of Code: Comparing Coding Behaviors Across Large Language Models

通过比较不同 LLM 的编码行为,探讨了超越性能指标的编码评估方法。

Read more →

TreeSpark: Calibrated, Load-Adaptive Draft Trees for Semi-Autoregressive Speculative Decoding

提出了一种校准的、负载自适应的草稿树方法,用于加速半自回归推测解码。

Read more →

A framework for recipe data structure with applications for culinary and nutritional insights

提出了一个食谱数据结构框架,旨在将食谱转化为可计算的数据,以获取烹饪和营养洞察。

Read more →

AdaMem: Adaptive Memory Token Allocation for Soft Compression in Retrieval-Augmented Generation

提出了一种自适应内存 token 分配方法,用于 RAG 中的软压缩,以降低处理长段落的成本。

Read more →


WIRED

How to Claim Your Cut of Apple’s $250 Million Siri Settlement

苹果 Siri 和解案赔偿申请指南。用户若认为在 Siri 发布时受到误导,可申请最高 95 美元的赔偿,截止日期为 12 月 21 日。

Read more →

Rabbit Is Back, This Time With an AI Agent App

Rabbit 公司推出 OS3 AI 代理应用,用户无需专用硬件即可在现有设备上使用其 AI 代理功能。

Read more →

Adobe Premiere, One of the iPhone’s Best Video-Editing Apps, Is Now on Android

Adobe Premiere 正式登陆 Android 平台,支持多种设备布局,且基础功能免费。

Read more →

Viture’s Vonder Glasses Are Meant to Map Your Mind

Viture 发布了 Vonder 智能眼镜,这是一款无显示屏的设备,内置骨传导麦克风,旨在私密记录用户的日常思考。

Read more →

AI Models Built From Rat Brains Just Got Closer to Reality

Biological Computing Company 将其 AI 工具引入 AWS,旨在将生物大脑与代码结合,推动生物计算领域的发展。

Read more →

GoPro Mission 1 Pro ILS Review (2026): Cinema-Quality Action Cam

GoPro Mission 1 Pro ILS 评测:这是一款具备电影级画质的运动相机,虽然不完美,但拍摄体验极佳。

Read more →

The UK Government Faces a Reckoning Over Palantir

英国政府面临关于 Palantir NHS 合同的抉择,这不仅是技术合同问题,更涉及与大科技公司及华盛顿的关系。

Read more →

Everything You Know About Political Violence Is Probably Wrong

文章指出,关于政治暴力是否上升的争论,往往源于对“政治暴力”定义的不同,而非单纯的意识形态分歧。

Read more →

Is a Home Security System Subscription Worth It? (2026)

探讨了家庭安防系统订阅的必要性,建议用户通过本地存储方案构建无需持续付费的安防系统。

Read more →

I Built AI Clones of My Coworkers. Things Got Weird

作者通过构建同事的 AI 克隆体,探讨了未来工作方式的潜在变化及带来的奇特体验。

Read more →


Lobsters

Plain-text files are at risk

讨论了纯文本文件在当前环境下的潜在风险。

Read more →

Looking forward to Git 2.56 - and 3.0

社区对 Git 未来版本更新的期待。

Read more →

Fearless SIMD v1.0 is here

Fearless SIMD v1.0 正式发布。

Read more →

Arguing about arguments

关于编程中参数传递方式的讨论。

Read more →

That About Wraps It Up for Stock Mac UI

关于 Mac 原生 UI 设计的讨论。

Read more →

Jev-powered autocorrection

探讨基于 Jev 模型的自动纠错技术。

Read more →

Self-Hosting Behind CGNAT

关于在 CGNAT 环境下进行自托管的讨论。

Read more →

Named and Optional Arguments are Awesome

讨论命名参数和可选参数在编程中的优势。

Read more →

tokens too cheap to meter

关于 token 成本趋近于零的讨论。

Read more →

Design your programming languages right (2024)

关于如何正确设计编程语言的讨论。

Read more →


DEV Community

14,913 Kubernetes Dashboard Matches: Why the Management UI Is the Wrong Thing to Expose

指出将 Kubernetes 管理界面暴露在公网上的风险,强调其作为管理工具而非用户界面的特殊性。

Read more →

Why Does Your AI Coding Agent Start Forgetting What It Was Doing?

探讨了 AI 编码代理在长时间运行后出现“遗忘”任务上下文的原因。

Read more →

We audited 110 AI usage tools. Here is where the numbers go wrong.

审计了 110 款 AI 使用量统计工具,揭示了这些工具在计费准确性方面存在的问题。

Read more →

Your agent picks one of two options. Can you test that choice?

探讨了如何测试 AI 代理在多个选项中做出选择的决策过程。

Read more →

Governance Attack Surface Review: Bybit

对 Bybit 协议的治理攻击面进行了安全审计。

Read more →

Building an Exposure Baseline: A Repeatable ZoomEye Method for Vulnerability Response

介绍了利用 ZoomEye 构建可重复漏洞响应暴露基线的方法。

Read more →

When OPA’s Bundle Loader Runs Past a .manifest Typo

分析了 OPA 在处理 .manifest 文件拼写错误时的行为及诊断缺失问题。

Read more →

Legendas traduzidas ao vivo numa extensão do Chrome: áudio da aba no MV3 e um fluxo em vez de dois serviços

开发者分享了如何开发一款 Chrome 扩展,实现 Meet、Zoom 和 Teams 会议的实时翻译字幕。

Read more →

From p=none to Enforcement: A Working Sequence for DMARC Rollout

介绍了 DMARC 从监控模式到强制执行模式的实施序列。

Read more →

Google Expands Gemini Notebook With Syncing, Interactive Tools and a New Name

Google 将 Gemini Notebook(原 NotebookLM)扩展为连接 Gemini 应用的知识工作空间,新增了同步和交互工具。

Read more →


Meta Engineering

Open-Sourcing Rebalancer: A Generic, High-Performance Library for Solving Assignment Problems

Meta 开源了 Rebalancer,这是一个用于解决资源分配问题的通用高性能库,已在 Meta 内部使用了九年。

Read more →

Inside Petal: Building the World’s First Petabit-Class Transoceanic Subsea Cable

Meta 介绍了 Petal 项目,这是全球首条 PB 级跨洋海底光缆,连接法国和美国,预计 2029 年投入使用。

Read more →

ZGateway: Learnings from Putting a Proxy in Front of ZippyDB

介绍了 ZGateway,这是 Meta 为 ZippyDB 键值存储设计的代理,旨在统一流量并提供负载均衡和弹性支持。

Read more →

An Organizational Second Brain: Building an AI That Learns From Experts

Meta 构建了一个 AI 代理,作为组织内的“第二大脑”,通过整合结构化知识架构,保存并共享专家知识。

Read more →

MetaRoCE: A New RDMA Transport Built for AI-Scale Ethernet

Meta 设计了 MetaRoCE,这是一种专为 AI 工作负载在以太网上运行而构建的 RDMA 传输协议。

Read more →

MTIA 300: Meta’s First Training Chip with Built-in NICs and Communication-Offloading Engines

MTIA 300 是 Meta 首款内置 NIC 和通信卸载引擎的训练芯片,专为推荐模型训练优化。

Read more →

How We’re Building Scam Alert on WhatsApp With End-to-End Encryption and Verifiability Guarantees

WhatsApp 正在构建 Scam Alert 功能,在保护端到端加密隐私的同时,防范 AI 生成的诈骗信息。

Read more →

From User Sequences to Scaling Laws: A Multi-Stage Architecture for Meta’s Ads Ranking

介绍了 Meta 广告排名模型的多阶段架构,通过建模用户行为序列提升了推荐效率。

Read more →

GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

Meta 通过优化 GEM 训练,将广告基础模型的训练效率提高了一倍。

Read more →


DeepMind Blog

Introducing Gemini 3.8 Live and 3.8 Live Extended Thinking

Google DeepMind 发布了 Gemini 3.8 Live 及其扩展思维版本。

Read more →

AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

AlphaGenome Atlas 绘制了人类基因组中 90 亿个单字母 DNA 变异的分子效应预测图。

Read more →

Introducing WeatherNext 3, our most advanced and accurate global weather AI model

发布了 WeatherNext 3,这是 DeepMind 最先进、最准确的全球天气 AI 模型。

Read more →

Proactive cyber defense for governments and enterprises

DeepMind 推出面向政府和企业的主动网络防御方案。

Read more →

Introducing Gemini 3.8 Flash and 3.8 Flash Cyber

发布了 Gemini 3.8 Flash 及其网络安全版本。

Read more →

Introducing agentic video understanding with Gemini

介绍了 Gemini 的代理式视频理解能力。

Read more →

Gemini Omni 1.1 Flash lets you build with more control

Gemini Omni 1.1 Flash 提供了更高的构建控制权。

Read more →

Piloting the world’s first double-blind AI evaluations

DeepMind 正在试点全球首个双盲 AI 评估项目。

Read more →

Intelligent transcription with Gemini 3.5 Transcribe

Gemini 3.5 Transcribe 提供更智能的语音转文本转录功能。

Read more →

From Atari to EVE Online: Building on 15 Years of AI Research in Games

回顾了 DeepMind 在游戏 AI 研究领域 15 年的历程,并宣布与游戏工作室合作开发突破性 AI 玩法。

Read more →


arXiv CS.LG

PRQuant: Permutation Residual Quantization for Low-Overhead Inference

提出了一种置换残差量化方法,用于实现低开销的推理。

Read more →

Generalized Multimodal Foundation Model

提出了一种广义多模态基础模型,旨在快速适应新的下游应用。

Read more →

Correcting Learning-based Perception for Safety

提出了一种两步策略,用于纠正学习型感知系统,以提升自动系统的安全性。

Read more →

A Shared Learning Rate Is Not a Neutral Control in Selective On-Policy Distillation

研究表明,在选择性策略蒸馏中,共享学习率并非中性控制。

Read more →

Toward Fairness in Machine Learning Models for Predicting Treatment Retention and Premature Discontinuation in Medication for Opioid Use Disorder

探讨了阿片类药物使用障碍治疗中 ML 预测模型的公平性问题。

Read more →

ZoAQ: Adaptive Zeroth-Order Querying via Query-Reuse Coupling

提出了一种自适应零阶查询方法,通过查询重用耦合来降低优化成本。

Read more →

LE4Mob: Towards Inductive, Distance-Aware and General-Purpose Location Embedding for Human Mobility Modelling

提出了一种归纳式、距离感知且通用的位置嵌入方法,用于人类移动性建模。

Read more →

Success Leaves Detours: Learning Executable Walkthroughs for Long-Horizon Agents

研究了如何从稀疏奖励轨迹中学习可执行的演练,以提升长程代理的性能。

Read more →


arXiv CS.CV

Did You Steal My Shot? Pioneering Camera Motion Plagiarism Detection in Generative Videos

提出了一种生成视频中的摄像机运动抄袭检测方法。

Read more →

Enabling Vision and Cross-Modal Learning for Multimodal Stroke Recurrence Prediction: An Interpretable Two-Step Framework

提出了一种可解释的两步框架,用于多模态中风复发预测。

Read more →

Moonworks Lunara: Modeling Artistic Intelligence

Moonworks Lunara:一种基于扩散混合 Transformer 架构的文本到图像模型,旨在建模艺术智能。

Read more →

Performance vs Consistency: Evaluating a Foundation Model in Lung-RADS Screening

评估了基础模型在 Lung-RADS 肺癌筛查中的性能与一致性。

Read more →

Brain-to-Image Generation: Reconstructing Visual Stimuli from EEG using Generative Adversarial Networks

利用生成对抗网络从 EEG 信号中重建视觉刺激。

Read more →

Rethinking Streaming Video Diffusion Model: Context, Execution, and Training

重新思考流式视频扩散模型的设计空间,包括上下文、执行和训练策略。

Read more →

Complementary rPPG-Derived and Lip-Region Frequency Cues for Talking-Face Deepfake Detection

利用 rPPG 导出波形和唇部区域频率线索,提升说话人脸 Deepfake 检测效果。

Read more →

Beyond the Survey: A Systematic Empirical Study of Detection and Association in Visual MOT

对视觉多目标跟踪(MOT)中的检测和关联组件进行了系统的实证研究。

Read more →


Towards Data Science

Break Your Own RAG Pipeline Before Users Do

建议通过对抗性测试集来发现 RAG 管道中的检索失败问题。

Read more →

Build a Speaker-Recognition App with Claude Code

介绍了如何使用 Claude Code 构建内部语音识别工具。

Read more →

4 Ways to Use AI on a PhD Thesis

分享了在博士论文写作中使用 AI 的四种方法,包括查找引用、整合代码和事实核查。

Read more →

An Introduction to Jev

介绍了 Jev 模型,这是一种旨在进行决策而非仅仅生成文本的 AI。

Read more →

A New Kind of Model for AI Decision-Making?

探讨了 TypeSafe 的 Jev 模型,并将其与 OpenAI 在意图分类方面进行了对比。

Read more →

GPT-6 Astra Just Hit OpenAI’s Highest Cybersecurity Risk Level

分析了 GPT-6 Astra 达到 OpenAI 最高网络安全风险等级的意义。

Read more →

Google Offered $10M for a Dying Airline’s Data. How Can You Value Yours?

探讨了如何为组织的操作数据进行估值。

Read more →

Your AI Assistant Wrote the Code. Who Checked the Defaults?

提醒开发者在 AI 编写代码后,务必检查 scikit-learn 等库的默认参数设置。

Read more →

GraphRAG: A Practitioner’s Guide to 6 Advanced Architectural Patterns

GraphRAG 实践指南:介绍了六种结合语义搜索、知识图谱和 LLM 推理的生产级架构模式。

Read more →

CBAM Paper Walkthrough: The Double-Attention Mechanism

CBAM 论文解读:从零开始使用 PyTorch 实现卷积块注意力模块。

Read more →

生成二维码中...

请点击右上角 ···

选择 发送给朋友收藏