Kog is going deeper to squeeze more inference out of GPUs
Kog is going deeper to squeeze more inference out of GPUs
Kog 正在深入挖掘,以从 GPU 中榨取更多推理性能
The race for faster AI inference is on, and markets gave Cerebras and its purpose-built chips a warm welcome in its IPO debut in May. But French startup Kog is betting that there’s a lot more power to be squeezed out of conventional GPUs. The startup hit the front page of Hacker News in May with a tech preview aimed at proving that “extremely fast single-request decoding is possible on the standard datacenter GPUs enterprises already own” — such as the AMD MI300X and Nvidia H200 GPUs it used for its demo. AI 推理速度的竞赛已经打响,市场在 5 月份 Cerebras 及其专用芯片首次公开募股(IPO)时给予了热烈欢迎。但法国初创公司 Kog 押注于传统 GPU 还有巨大的性能潜力可供挖掘。今年 5 月,这家初创公司登上了 Hacker News 的头条,其技术预览旨在证明“在企业现有的标准数据中心 GPU(如其演示中使用的 AMD MI300X 和 Nvidia H200)上,实现极速的单请求解码是可能的”。
Some were disappointed to hear this didn’t extend to GPUs in our laptops, but others saw the potential. With inference speed and costs now being a critical bottleneck, Kog’s promise to unlock new capabilities on existing hardware with software optimization attracted more than onlookers. “We had 200 tangible business leads,” CEO Gaël Delalleau told TechCrunch. 有些人对这一技术无法扩展到笔记本电脑 GPU 感到失望,但另一些人则看到了其潜力。随着推理速度和成本成为关键瓶颈,Kog 通过软件优化在现有硬件上解锁新功能的承诺吸引了众多关注。首席执行官 Gaël Delalleau 告诉 TechCrunch:“我们获得了 200 个切实的商业线索。”
Based on early feedback, the solo founder expects software engineering to be the first use case. Veteran Claude Code users are well aware that they sometimes have to wait hours to get results. Anthropic itself understands that speed is worth money, and charges a price multiple for Claude’s Fast Mode. Kog is hoping to target customers put off by those delays, usually because they rely on AI workflows for professional tasks. 根据早期反馈,这位独立创始人预计软件工程将是首个应用场景。资深的 Claude Code 用户非常清楚,他们有时需要等待数小时才能得到结果。Anthropic 本身也明白速度的价值,并为 Claude 的“快速模式”(Fast Mode)收取溢价。Kog 希望瞄准那些因延迟而望而却步的客户,这些客户通常依赖 AI 工作流来完成专业任务。
But the startup also has design partners that let users generate games and apps with a prompt, and for whom a faster outcome thanks to the Kog Inference Engine (KIE) would mean more revenue, Delalleau said. The company realizes this market is not quite mature yet. While observing demand, Kog learned that its prospective customers aren’t prepared to fine-tune small models. “And that’s why since the launch, we’ve been fully focused on accelerating the development of larger models to meet the demand we’ve seen.” Delalleau 表示,该公司还有一些设计合作伙伴,允许用户通过提示词生成游戏和应用程序,对于他们来说,得益于 Kog 推理引擎(KIE)的更快输出意味着更多的收入。该公司意识到这个市场尚未完全成熟。在观察需求的过程中,Kog 了解到其潜在客户并不准备微调小模型。“这就是为什么自发布以来,我们一直专注于加速大模型的开发,以满足我们所观察到的需求。”
This leaves Kog with a huge leap to make to deliver on its promise of “30x faster LLM inference.” Its demo showed an impressive 3,000 per-request tokens per second (TPS) — but with a purpose-built small model with only some 2 billion parameters, the now open sourced Laneformer 2B. Contradicting skeptics, Delalleau is confident the same approach can work just as well with LLMs, whose size can be a challenge for inference chips. 这使得 Kog 要实现其“大语言模型(LLM)推理速度提升 30 倍”的承诺,还有巨大的跨越要完成。其演示展示了令人印象深刻的每秒 3,000 个令牌(TPS)的单请求处理速度,但这是基于一个仅有约 20 亿参数的专用小模型——即现已开源的 Laneformer 2B。针对怀疑论者,Delalleau 坚信同样的方法在 LLM 上也能奏效,尽管 LLM 的规模对推理芯片来说可能是一个挑战。
“GPUs have a bright future,” he said. For Kog’s CEO, the idea that they aren’t well suited for decoding has become a misconception; newer GPUs have more and more memory bandwidth that only begs to be unlocked. Kog isn’t alone in thinking that software optimization can help GPUs do more than it says on the box. ZML, also from France, released hardware-agnostic software that bypasses Nvidia’s CUDA to support fast inference across competing chips. But Delalleau said Kog is more akin to Stanford University lab Hazy Research, with an even deeper-level focus on GPU acceleration. “GPU 的未来一片光明,”他说。对于 Kog 的首席执行官来说,认为 GPU 不适合解码是一种误解;较新的 GPU 拥有越来越多的内存带宽,等待着被解锁。Kog 并非唯一一家认为软件优化能让 GPU 发挥超越其标称性能的公司。同样来自法国的 ZML 发布了与硬件无关的软件,绕过了 Nvidia 的 CUDA,以支持跨不同芯片的快速推理。但 Delalleau 表示,Kog 更类似于斯坦福大学的 Hazy Research 实验室,对 GPU 加速有着更深层次的关注。
Delalleau himself is not a researcher, and his first startup, TechCrunch50 2009 alum Stribe, has nothing to do with his new one — other than his former co-founder turned VC Kamel Zeroual, whose firm Varsity VC co-led Kog’s seed round. But the startup’s deep-level focus stems from his unique background. Having studied solid-state physics at France’s École Polytechnique, he went on to work in offensive cybersecurity — also known as white hat hacking. Delalleau 本人并非研究人员,他的第一家初创公司(2009 年 TechCrunch50 校友 Stribe)与他现在的新公司并无关联——除了他曾经的联合创始人、现任风险投资人 Kamel Zeroual,其所在的 Varsity VC 领投了 Kog 的种子轮融资。但这家初创公司对深层技术的关注源于他独特的背景。他在法国巴黎综合理工学院学习过固体物理学,随后从事进攻性网络安全工作,也就是所谓的“白帽黑客”。
According to Delalleau, this shaped the mindset he is now encouraging his team to adopt. On the science side, “there’s this mindset of understanding the laws of physics, and the laws of the GPU in order to make the most of them.” As for hacking, the four-time finalist at DEFCON’s CTF tournament said it taught him “to reverse-engineer things at a very low level — down to assembly language and binary code — to understand how it works, and to try to use it to achieve a goal for which it wasn’t necessarily designed.” 据 Delalleau 所述,这塑造了他现在鼓励团队采用的思维方式。在科学方面,“有一种理解物理定律和 GPU 定律的思维,以便充分利用它们。”至于黑客技术,这位曾四次入围 DEFCON CTF 比赛决赛的选手表示,这教会了他“在极低层面进行逆向工程——深入到汇编语言和二进制代码——以理解其工作原理,并尝试利用它来实现它最初未必是为了该目标而设计的功能。”
The downside of this approach is that it is very hands-on and time-consuming. “For every new GPU, we’ll dedicate several weeks or even months, to really dig into the details and conduct GPU engineering research on that hardware.” With a team of 11 people, this puts a limit to the number of chips that Kog can work with, at least for the foreseeable future. 这种方法的缺点是需要大量亲力亲为且非常耗时。“对于每一款新 GPU,我们都会投入数周甚至数月的时间,深入挖掘细节并在该硬件上进行 GPU 工程研究。”由于团队只有 11 人,这限制了 Kog 在可预见的未来能够支持的芯片数量。
In the longer run, Kog hopes to feed its methodology into agent-based pipelines that will let it support more chips and models. As Europe seeks to build its own capability on those two fronts, this could add sovereignty tailwinds for the startup, which is already supported by Scaleway and backed by France’s Bpifrance and French Tech 2030’s program. 从长远来看,Kog 希望将其方法论融入基于智能体的流水线中,从而支持更多的芯片和模型。随着欧洲寻求在这两个领域建立自己的能力,这可能会为这家初创公司带来主权方面的助力,该公司目前已获得 Scaleway 的支持,并由法国国家投资银行(Bpifrance)和“法国科技 2030”(French Tech 2030)计划提供支持。
For now, though, Kog needs to prove to the world that its approach works on LLMs. This will also be key to securing more funding. “Once we’ve implemented our first major model at 10x speed, which I think will be in September, we’ll be able to start demonstrating customer traction and from there, raise our Series A,” Delalleau said. 不过目前,Kog 需要向世界证明其方法在 LLM 上是有效的。这也将是获得更多资金的关键。“一旦我们以 10 倍的速度实现了第一个主要模型(我认为会在 9 月份),我们就能够开始展示客户吸引力,并以此为基础进行 A 轮融资,”Delalleau 说道。