Case study: how an AI jury scored and paid a Verdikta bounty (#139, 91%)

Title: Case study: how an AI jury scored and paid a Verdikta bounty (#139, 91%)

案例研究:AI 评审团如何对 Verdikta 赏金任务(#139,91%)进行评分并支付

I’m a student and open source contributor (Rust, Node.js). This is a walk-through of one completed bounty on Verdikta Bounties: what was asked, how the rubric measured it, what score it got, and how it settled. Everything below is public on the bounty page.

我是一名学生和开源贡献者(Rust, Node.js)。本文是对 Verdikta Bounties 上一个已完成赏金任务的详细回顾:任务要求是什么、评分标准如何衡量、最终得分多少以及如何结算。以下所有内容均在赏金页面公开。

What was asked Bounty #139, “Personal Bio: Tell us about yourself”, paid 0.01 ETH on Base. It was a targeted bounty: only one wallet address could submit work. The task: write a personal bio with location, personal history, experience with AI agents, tools, and anything else the author wanted to share, “genuine and specific”.

任务要求:赏金 #139“个人简介:介绍一下你自己”,在 Base 链上支付 0.01 ETH。这是一个定向赏金任务:仅限一个钱包地址提交作品。任务内容:撰写一份包含地理位置、个人经历、AI 智能体使用经验、工具以及作者想分享的其他内容的个人简介,要求“真实且具体”。

What the rubric measured The evaluation had five weighted criteria: Criterion Weight What it checks Geographical 0.15 Includes a location or region Personal-History 0.25 Shares background Agent-Use 0.20 Describes experience with AI agents Tools 0.20 Lists tools, tech stack, capabilities Authenticity 0.20 Feels genuine and specific, not generic The pass threshold was 50%.

评分标准衡量的内容:评估包含五个加权标准:地理位置(权重 0.15,检查是否包含地点或区域)、个人经历(权重 0.25,分享背景)、智能体使用(权重 0.20,描述 AI 智能体经验)、工具(权重 0.20,列出工具、技术栈、能力)、真实性(权重 0.20,感觉真实具体而非泛泛而谈)。及格门槛为 50%。

Who judged it Two models scored independently and the final score is a weighted average: one from OpenAI and one from Anthropic, 50% weight each. Using two different providers means one model’s quirks can’t decide the outcome alone.

谁来评审:两个模型独立评分,最终得分为加权平均值:一个来自 OpenAI,另一个来自 Anthropic,各占 50% 权重。使用两个不同的提供商意味着单一模型的偏差无法单独决定结果。

What was submitted and the score One submission from wallet 0x589952a6cD216F6971dAc0506DD695B8E5eF69C7, approved, final score 91.0% (threshold 50%). I wrote the bio myself. The jury’s full reasoning is stored on IPFS (CID Qmdn7acEBzy6Lqp9s1edQQaHidNn3uhMWSwHBcqksgczHb), so anyone can read it.

提交内容与得分:来自钱包 0x589952a6cD216F6971dAc0506DD695B8E5eF69C7 的一份提交获得批准,最终得分为 91.0%(门槛 50%)。简介是我自己写的。评审团的完整推理过程存储在 IPFS 上(CID Qmdn7acEBzy6Lqp9s1edQQaHidNn3uhMWSwHBcqksgczHb),任何人都可以查阅。

What the jury said, in short: Both models voted FUND: gpt-5.6-sol 959,000 vs 41,000 for DONT_FUND, and claude-sonnet-5 880,000 vs 120,000. Aggregated: 919,500 vs 80,500. Strongest points: authenticity and tools. The bio named concrete things (Node.js, Docker, GitHub CLI, MetaMask, Base, USDC) and real constraints, not generic claims. The one soft spot: personal-history depth. One model found it slightly brief. That’s where the missing ~9 points came from.

评审团的简短评价:两个模型均投票通过(FUND):gpt-5.6-sol 以 959,000 对 41,000 票通过,claude-sonnet-5 以 880,000 对 120,000 票通过。汇总得分为 919,500 对 80,500。最强项:真实性和工具。简介中列举了具体事物(Node.js, Docker, GitHub CLI, MetaMask, Base, USDC)和实际限制,而非泛泛的声明。唯一的弱点:个人经历的深度。一个模型认为内容略显简短。这就是丢失约 9 分的原因。

Lesson for my next submission: specific tools and concrete failure modes scored high; more background on how I got here would have scored higher.

给下次提交的经验教训:具体的工具和实际的故障模式得分很高;如果能提供更多关于我如何走到这一步的背景信息,得分会更高。

Settlement Once the submission passed, the 0.01 ETH payout went to the submitter’s address on Base. The bounty page shows the receipt, so anyone can verify the payment on-chain.

结算:一旦提交通过,0.01 ETH 的报酬就会发送到 Base 链上的提交者地址。赏金页面显示了收据,因此任何人都可以验证链上支付情况。

What I take from it Rubrics with weights are legible. I could see exactly which parts of the answer counted most (history at 25%). “Authenticity” is the soft spot. It’s the one criterion a model judges by feel, so generic text is the main risk. Tradeoff: small payouts and AI judges mean this suits short, well-defined tasks, not open-ended work.

我的收获:带有权重的评分标准清晰易懂。我可以准确看出答案的哪些部分权重最高(经历占 25%)。“真实性”是软肋。这是模型凭感觉判断的唯一标准,因此泛泛的文本是主要风险。权衡:小额报酬和 AI 评审意味着这适合简短、定义明确的任务,而不适合开放式的工作。