Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.
Do LLMs Catch Bad Startup Math? My First Answer Was Wrong.
大模型能识破创业公司的数学陷阱吗?我最初的结论错了。
Kaggle Benchmarking Challenge Submission. This is a submission for the Kaggle Benchmarking Challenge. What I Benchmarked: I measured Sycophantic Failure Resistance and Numerical Consistency Verification — specifically, whether models perform independent arithmetic checks on unit-economics and growth-rate claims embedded in realistic pitch language, or default to validating a confidently-stated conclusion without verifying the underlying math. Kaggle 基准测试挑战赛投稿。这是我为 Kaggle 基准测试挑战赛提交的内容。我测试的内容是:模型对“阿谀奉承式错误”(Sycophantic Failure)的抵御能力以及数值一致性验证能力。具体来说,就是当模型面对嵌入在真实商业计划书语境中的单位经济效益和增长率声明时,它们是会进行独立的算术核查,还是会默认验证那些自信满满的结论而不去核实背后的数学逻辑。
This interested me because sycophancy — a model agreeing with what’s presented instead of checking it — is a documented, high-stakes failure mode: a business error validated by an AI carries real downstream risk. Before trusting my own results, I built calibration controls: one pitch with correct math models should NOT flag, one with an obvious contradiction they MUST catch. Both passed clean across every model, confirming the harness measures signal, not noise. 我之所以对此感兴趣,是因为“阿谀奉承”(即模型不加核实地认同用户提供的内容)是一种已被记录在案且风险极高的失效模式:如果 AI 验证了一个商业错误,会带来真实的后续风险。在信任自己的测试结果之前,我设置了校准对照组:一个数学逻辑正确的商业计划书模型不应报错,而另一个存在明显矛盾的计划书模型则必须被识别出来。所有模型都顺利通过了这两项测试,这证实了我的测试框架测量的是有效信号,而非噪声。
Models Tested: I benchmarked six models across three labs using a paired frontier-vs-small-tier design — one flagship and one cost-optimized model per lab — to isolate whether model scale predicts arithmetic verification reliability, independent of which lab produced it. 测试模型:我通过“前沿模型 vs. 小型模型”的配对设计,对来自三家实验室的六款模型进行了基准测试——每家实验室选取一款旗舰模型和一款成本优化模型——旨在探究模型规模是否能预测算术验证的可靠性,而不受实验室来源的影响。
OpenAI: gpt-6-astra (frontier) vs. gpt-5.4-nano (small) Google: gemini-3.1-pro-preview (frontier) vs. gemini-3.6-flash (small) Anthropic: claude-opus-5 (frontier) vs. claude-haiku-4-5 (small) OpenAI:gpt-6-astra(前沿)vs. gpt-5.4-nano(小型) Google:gemini-3.1-pro-preview(前沿)vs. gemini-3.6-flash(小型) Anthropic:claude-opus-5(前沿)vs. claude-haiku-4-5(小型)
This design separates two questions that usually get conflated: whether a lab’s flagship model is reliable, and whether that reliability degrades at the small/cheap tier — and if it does, whether the degradation is universal or lab-specific. 这种设计将两个通常被混为一谈的问题分离开来:一是实验室的旗舰模型是否可靠,二是这种可靠性在小型/低成本模型层级是否会下降——如果下降,这种退化是普遍现象还是特定实验室的问题。
Findings: Frontier Consistency: All three frontier models scored 3/3 across both scenarios on every trial, with zero variance. Verification reliability at this tier appears saturated for tasks of this difficulty. 研究发现:前沿模型的一致性:所有三款前沿模型在每次试验的两个场景中均获得 3/3 的满分,且零方差。对于此类难度的任务,该层级的验证可靠性似乎已达到饱和。
The Single-Sample Fallacy: My first pass produced a clean result — gpt-5.4-nano and claude-haiku-4-5 both scored 0/3 on the CAC/LTV scenario, appearing to confirm that small-tier models fail at arithmetic verification. This did not replicate. Re-running gpt-5.4-nano twice more produced 3/3 and 3/3 — meaning the initial “failure” was sampling variance, not a model property. A single trial is statistically insufficient to characterize failure modes in stochastic systems; I would have published a false negative without the rerun. 单样本谬误:我的第一轮测试得出了一个清晰的结果——gpt-5.4-nano 和 claude-haiku-4-5 在 CAC/LTV(获客成本/用户终身价值)场景中均得分为 0/3,这似乎证实了小型模型在算术验证上确实会失败。但这一结果无法复现。在对 gpt-5.4-nano 进行两次重复测试后,结果均为 3/3——这意味着最初的“失败”只是采样方差,而非模型本身的属性。在随机系统中,单次试验在统计学上不足以表征失效模式;如果不是因为进行了复测,我就会发布一个错误的否定结论。
The Robust Finding — Detection/Correction Decoupling: One result held across all three independent trials: claude-haiku-4-5 on the growth-compounding scenario correctly identified the inconsistency in 3/3 runs, but miscalculated the corrected value in 3/3 runs — producing three different wrong answers ($92,000, $92,000, $71,304) against the true value of ~$89,161. This indicates error detection and error correction are not the same capability and don’t necessarily co-occur reliably, even within a single model on a single task type. 稳健的发现——检测与纠错的解耦:有一个结果在所有三次独立试验中都保持一致:claude-haiku-4-5 在增长复利场景中,在 3/3 的运行中都正确识别出了不一致之处,但在 3/3 的运行中都算错了修正后的数值——针对约 $89,161 的真实值,它给出了三个不同的错误答案($92,000、$92,000、$71,304)。这表明错误检测和错误纠正并非同一种能力,且即使在同一个模型的同一类任务中,它们也不一定会可靠地同时出现。
Implication: Benchmark claims based on single-sample runs against non-deterministic systems should be treated as unverified until replicated. The methodologically interesting failures aren’t the ones that appear once — they’re the ones that survive an attempt to falsify them. 启示:基于非确定性系统的单样本运行所做出的基准测试声明,在被复现之前应被视为未经证实。在方法论上,真正值得关注的失败不是那些只出现一次的现象,而是那些在试图证伪后依然存在的现象。
My Benchmark 🔗 Kaggle Task: CAC/LTV Check 🔗 Full notebook — includes all four tasks (both test scenarios plus the two calibration controls), and the full repeat-run transcripts referenced above. 我的基准测试 🔗 Kaggle 任务:CAC/LTV 检查 🔗 完整笔记本 — 包含所有四项任务(两个测试场景加上两个校准对照组),以及上述提到的完整重复运行记录。