快·讯

事件详情 · c9d60483-5862-4c67-9870-223d32e66739

What AI Verification Actually Requires

treenews 其他 人工智能
原文 · SOURCE RECORDS
treenews 10-01 20:29:16 Twitter 此为最新版 · 另有 1 个早前版本
原文 · 8114 字符(点击折叠)

Mira (@miranetwork): What AI Verification Actually Requires Accuracy numbers get thrown around a lot, but they don't tell you much about what's actually in front of you. They measure how a model performed on average, across a test set, in the past. They say nothing about the one output your agent is about to act on right now. That's the gap most teams building with AI haven't priced in yet. Where the confusion tends to come from A model can score 95% on a benchmark and still get this particular question wrong. Maybe it cites the wrong source. Maybe it misreads a date, or lets a bad assumption ride through five steps of reasoning that nobody checked. None of that shows up in the benchmark. It shows up once the model starts acting on the world instead of just describing it. The question that actually matters is much narrower: is this specific output, right now, safe to act on. Trusting the model in general doesn't answer that. Accuracy is an aggregate number. Verification isn't. A model that's right 95% of the time can't tell you whether today's answer falls in the 95 or the 5. Its own confidence doesn't help much either. Models get very confident about wrong answers, and genuinely unsure about the ones they actually nailed. So verification has to happen after generation, run by something other than the model that produced the output. A longer prompt won't do it. Neither will asking the model to check its own work. "Verification" usually covers three different checks. Output verification: is the information correct? Permission verification: was this action authorized? Outcome verification: did it actually happen? Take an agent paying an invoice. Getting the amount right isn't enough on its own: the user actually has to have approved the payment, and once it's sent, someone has to confirm the money got where it was supposed to go. Checking the invoice tells you nothing about whether the payment was authorized. Confirming the payment tells you nothing about whether the invoice was ever legitimate. Different layers, checked by different systems. Lump them all into one "AI safety" bucket and you lose the ability to tell what actually broke. This piece focuses on the first layer: whether the information itself is correct. Start with the claim, not the answer. A generated answer is almost never fully true or fully false. Take this line: "Company X reported $4.2B in revenue, grew 18% YoY, and acquired Company Y in March." Three claims, stapled together. One can be right while the other two are wrong, and you'd never know it from reading the sentence alone. So verification starts by pulling the claims apart and testing each one on its own. It also has to keep track of how the claims relate to each other. If a conclusion depends on two claims underneath it, checking the conclusion without checking those two doesn't really tell you anything. Claim extraction is its own point of failure. Miss a qualifier or mangle a claim, and everything checked after that ends up checking the wrong thing. Different claims need different kinds of checks. 1. Deterministic validation. If code can calculate something exactly, there's no reason to make a model guess at it. Recomputing the math or checking a date against a schema is cheap and reliable, and it's worth doing wherever it applies. 2. Source grounding. Compare the claim against a document or database that actually backs it up. This works well when a solid source exists, and gets harder fast when sources disagree or the data's out of date. A citation by itself doesn't verify anything. The source actually has to say what the claim claims it says. 3. Independent model evaluation. Multiple models check the same claim without seeing each other's answers, which helps when a question needs judgment rather than a simple lookup. They're not fully independent (shared training data means shared blind spots sometimes), but when they disagree, that tells you something a single model checking its own work never could. 4. Human escalation. Some cases genuinely need a person to look at them. Trying to automate every single decision doesn't make a system more capable. It usually just means the system hasn't been given permission to say "I don't know," which is a perfectly fine answer sometimes. The pipeline, roughly six steps. You break the output into claims first, then attach whatever evidence each claim needs to actually be judged. From there, claims go out to evaluators who can't see each other's answers, and you decide upfront what counts as an accepted result. A plain majority might be fine for something low-stakes; anything expensive to get wrong needs more agreement than that. Everything gets logged: who evaluated what, and what evidence they were working from. "The model said so" isn't a record of anything. And the system has to be willing to actually reject something. One that approves everything isn't verifying anything. It's just adding a step. Stricter consensus costs you something too. Unanimous agreement cuts down on wrong answers getting through, but it also rejects correct ones whenever a single evaluator misreads a claim or doesn't have enough context. Loosen the threshold and you cover more ground, at the cost of letting more mistakes through. Every evaluator you add pushes cost and latency up, and bringing in a human raises confidence but caps how much you can scale. There's no universal right answer here. A recommendation engine and an autonomous agent moving real money shouldn't run the same verification policy. What decides the threshold is what a wrong answer would actually cost you. This is roughly how Mira approaches it. Mira treats a model's output as a candidate for verification rather than a finished answer. The output gets broken into independently checkable claims. Those claims go out to multiple models that evaluate them separately. The responses get standardized so they're actually comparable, then checked against a consensus threshold. Clear the bar, and the system issues a cryptographic certificate recording exactly what was checked and how. That means verification work gets spread across independent nodes instead of running through one central pipeline. Claims get sharded, so no single node ever sees the whole output. Evaluator responses stay independent until everything's aggregated at the end, and node operators have real financial skin in the game if they cut corners on honest verification. The certificate doesn't prove a claim is true. What it proves is that a defined process ran and actually hit its threshold. That's a smaller claim than proving truth, but it's a more honest one to make. The numbers, from 78 test cases. Generator alone: 73.1% precision. Two-model consensus: 93.9%. Three-model consensus: 95.6%. Two-of-three consensus: 86.9%. Unanimous three-model agreement accepted 45 cases: 43 correct, two wrong, and 19 rejected that were actually fine. That last number matters. Strict consensus filtered out bad answers. It also turned away some good ones along the way. That's simply the cost of running conservative, and it's worth knowing what that cost is before setting a threshold, because finding out afterward gets expensive. Small sample, structured format. Testing this at scale, on messier and more open-ended output, is the natural next step. The larger point. Every bit of capability handed to AI comes with a matching bit of liability. A model that drafts contracts, moves money, or coordinates other agents raises the cost of whatever mistake slips through unnoticed, simply by doing more with less oversight along the way. AI is going to be wrong sometimes. That was never really in question. What actually matters is whether anything sits between a wrong output and the action it triggers. A longer prompt doesn't fix that. Neither does bolting on another confidence score. What actually helps is a system that can pull a claim apart and check it properly, then say no when it needs to and leave a record of why. That's the part worth actually building, even if it's the least glamorous line item on the roadmap.

源站原文 ↗
译文 · CHINESE

准确性数字经常被提及,但它们并不能告诉你很多关于你面前实际内容的信息。它们衡量的是模型在过去、跨测试集上的平均表现。它们对你代理即将基于其行动的那一个输出毫无说明。这就是大多数使用AI构建的团队尚未计入的差距。 混淆通常来自哪里:一个模型在基准测试中可能得分95%,但仍然会把这道特定题目做错。也许它引用了错误的来源。也许它误读了日期,或者让一个糟糕的假设在五步推理中一路通行而无人检查。这些都不会出现在基准测试中。它会在模型开始对世界采取行动而不仅仅是描述世界时显现出来。真正重要的问题要窄得多:这个特定输出,现在,是否可以安全地采取行动。笼统地信任模型并不能回答这个问题。 准确性是一个聚合数字。验证不是。一个95%时间正确的模型无法告诉你今天的答案属于95还是5。它自己的置信度也没有太大帮助。模型对错误答案非常自信,对真正答对的答案却真正不确定。因此,验证必须在生成之后进行,由产生输出的模型以外的其他东西运行。更长的提示词做不到。让模型检查自己的工作也做不到。 “验证”通常涵盖三种不同的检查。输出验证:信息是否正确?权限验证:此操作是否被授权?结果验证:它是否真的发生了?以一个支付发票的代理为例。金额正确本身不够:用户实际上必须已批准付款,而且一旦发送,必须有人确认钱到了该去的地方。检查发票告诉你关于付款是否被授权一无所知。确认付款告诉你关于发票是否曾经合法一无所知。不同层次,由不同系统检查。把它们全部归入一个“AI安全”桶里,你就失去了分辨实际出了什么问题的能力。本文聚焦第一层:信息本身是否正确。 从声明开始,而不是从答案开始。生成的答案几乎从不是完全正确或完全错误。以这句话为例:“X公司报告了42亿美元营收,同比增长18%,并在3月收购了Y公司。”三个声明,捆绑在一起。一个可以正确而另外两个错误,而你从阅读这句话本身永远不会知道。因此,验证从拆开声明并单独测试每一个开始。它还必须跟踪声明之间如何相互关联。如果一个结论依赖于其下的两个声明,检查结论而不检查那两个声明实际上不会告诉你任何东西。 声明提取本身就是一个失败点。错过一个限定词或弄乱一个声明,之后检查的一切最终都会检查错误的东西。 不同的声明需要不同种类的检查。1. 确定性验证。如果代码可以精确计算某件事,就没有理由让模型去猜测。重新计算数学或对照模式检查日期是廉价且可靠的,值得在适用之处做。2. 来源依据。将声明与真正支持它的文档或数据库进行比较。当存在可靠来源时这很有效,当来源分歧或数据过时时会迅速变难。引用本身不验证任何东西。来源实际上必须说声明声称它说的内容。3. 独立模型评估。多个模型检查同一声明而不看到彼此的答案,这有助于需要判断而非简单查找的问题。它们并非完全独立(共享训练数据意味着有时共享盲点),但当它们分歧时,这告诉你一些单个模型检查自己工作永远无法告诉你的东西。4. 人工升级。有些情况确实需要人来看。试图自动化每一个决定并不会使系统更有能力。它通常只是意味着系统没有被允许说“我不知道”,这有时是一个完全可以的答案。 流程,大致六步。你首先将输出分解为声明,然后为每个声明附加其实际需要被判断的任何证据。从那里,声明发送给无法看到彼此答案的评估者,你预先决定什么算作可接受的结果。简单多数可能对低风险的事情没问题;任何出错代价高昂的事情需要比这更多的共识。所有内容都被记录:谁评估了什么,以及他们依据什么证据。“模型说的”不是任何记录。而且系统必须愿意实际拒绝某些东西。一个批准所有东西的系统没有在验证任何东西。它只是在增加一个步骤。 更严格的共识也会让你付出代价。一致同意减少了错误答案通过,但它也会在单个评估者误读声明或没有足够上下文时拒绝正确答案。放宽阈值你覆盖更多地面,代价是让更多错误通过。你添加的每个评估者都会推高成本和延迟,引入人工会提高置信度但限制你能扩展多少。这里没有普遍正确的答案。推荐引擎和移动真金白银的自主代理不应运行相同的验证策略。决定阈值的是错误答案实际会花费你什么。 这大致是Mira的方法。Mira将模型的输出视为验证的候选者,而不是完成的答案。输出被分解为可独立检查的声明。这些声明发送给多个模型,它们分别评估。响应被标准化以便实际可比,然后对照共识阈值检查。通过标准,系统颁发一个加密证书,记录确切检查了什么以及如何检查。这意味着验证工作分布在独立节点上,而不是通过一个中央流程运行。声明被分片,因此没有单个节点能看到整个输出。评估者响应保持独立,直到所有内容在最后聚合,节点运营商在诚实验证上偷工减料时有真正的财务利害关系。证书不证明声明是真的。它证明的是定义好的流程运行了并实际达到了其阈值。这是一个比证明真理更小的声明,但这是一个更诚实的声明。 数字,来自78个测试案例。仅生成器:73.1%精确率。双模型共识:93.9%。三模型共识:95.6%。三选二共识:86.9%。一致的三模型协议接受了45个案例:43个正确,两个错误,以及19个被拒绝但实际没问题的。最后一个数字很重要。严格共识过滤掉了坏答案。它也在这个过程中拒绝了一些好答案。这仅仅是运行保守的成本,在设定阈值之前知道这个成本是值得的,因为事后发现会变得昂贵。小样本,结构化格式。在更大规模、更杂乱和更开放式的输出上测试这是自然的下一步。 更大的观点。赋予AI的每一点能力都伴随着相应的一点责任。一个起草合同、转移资金或协调其他代理的模型,仅仅通过用更少的监督做更多事,就提高了任何未被注意的错误溜过的成本。AI有时会出错。这从来不是真正的问题。真正重要的是,在错误输出和它触发的动作之间是否有任何东西。更长的提示词不能解决这个问题。附加另一个置信度分数也不能。真正有帮助的是一个系统,它能拆开声明并正确检查,然后在需要时说“不”,并留下为什么的记录。这是真正值得构建的部分,即使它是路线图上最不光彩的一项。

摘要 · AI SUMMARY

AI验证需针对具体输出而非平均准确率。验证分三层:输出、权限和结果,需拆解声明并分别检查。Mira采用多模型独立评估与共识阈值,78个测试案例中,单模型精度73.1%,双模型共识93.9%,三模型共识95.6%,但严格共识拒绝了19个正确结果。验证成本与错误代价需权衡。

以上为 AI 摘要;完整原文请经由源站链接查阅。

← 返回资讯流