原文 · 8114 字符(点击折叠)
Mira (@miranetwork): What AI Verification Actually Requires Accuracy numbers get thrown around a lot, but they don't tell you much about what's actually in front of you. They measure how a model performed on average, across a test set, in the past. They say nothing about the one output your agent is about to act on right now. That's the gap most teams building with AI haven't priced in yet. Where the confusion tends to come from A model can score 95% on a benchmark and still get this particular question wrong. Maybe it cites the wrong source. Maybe it misreads a date, or lets a bad assumption ride through five steps of reasoning that nobody checked. None of that shows up in the benchmark. It shows up once the model starts acting on the world instead of just describing it. The question that actually matters is much narrower: is this specific output, right now, safe to act on. Trusting the model in general doesn't answer that. Accuracy is an aggregate number. Verification isn't. A model that's right 95% of the time can't tell you whether today's answer falls in the 95 or the 5. Its own confidence doesn't help much either. Models get very confident about wrong answers, and genuinely unsure about the ones they actually nailed. So verification has to happen after generation, run by something other than the model that produced the output. A longer prompt won't do it. Neither will asking the model to check its own work. "Verification" usually covers three different checks. Output verification: is the information correct? Permission verification: was this action authorized? Outcome verification: did it actually happen? Take an agent paying an invoice. Getting the amount right isn't enough on its own: the user actually has to have approved the payment, and once it's sent, someone has to confirm the money got where it was supposed to go. Checking the invoice tells you nothing about whether the payment was authorized. Confirming the payment tells you nothing about whether the invoice was ever legitimate. Different layers, checked by different systems. Lump them all into one "AI safety" bucket and you lose the ability to tell what actually broke. This piece focuses on the first layer: whether the information itself is correct. Start with the claim, not the answer. A generated answer is almost never fully true or fully false. Take this line: "Company X reported $4.2B in revenue, grew 18% YoY, and acquired Company Y in March." Three claims, stapled together. One can be right while the other two are wrong, and you'd never know it from reading the sentence alone. So verification starts by pulling the claims apart and testing each one on its own. It also has to keep track of how the claims relate to each other. If a conclusion depends on two claims underneath it, checking the conclusion without checking those two doesn't really tell you anything. Claim extraction is its own point of failure. Miss a qualifier or mangle a claim, and everything checked after that ends up checking the wrong thing. Different claims need different kinds of checks. 1. Deterministic validation. If code can calculate something exactly, there's no reason to make a model guess at it. Recomputing the math or checking a date against a schema is cheap and reliable, and it's worth doing wherever it applies. 2. Source grounding. Compare the claim against a document or database that actually backs it up. This works well when a solid source exists, and gets harder fast when sources disagree or the data's out of date. A citation by itself doesn't verify anything. The source actually has to say what the claim claims it says. 3. Independent model evaluation. Multiple models check the same claim without seeing each other's answers, which helps when a question needs judgment rather than a simple lookup. They're not fully independent (shared training data means shared blind spots sometimes), but when they disagree, that tells you something a single model checking its own work never could. 4. Human escalation. Some cases genuinely need a person to look at them. Trying to automate every single decision doesn't make a system more capable. It usually just means the system hasn't been given permission to say "I don't know," which is a perfectly fine answer sometimes. The pipeline, roughly six steps. You break the output into claims first, then attach whatever evidence each claim needs to actually be judged. From there, claims go out to evaluators who can't see each other's answers, and you decide upfront what counts as an accepted result. A plain majority might be fine for something low-stakes; anything expensive to get wrong needs more agreement than that. Everything gets logged: who evaluated what, and what evidence they were working from. "The model said so" isn't a record of anything. And the system has to be willing to actually reject something. One that approves everything isn't verifying anything. It's just adding a step. Stricter consensus costs you something too. Unanimous agreement cuts down on wrong answers getting through, but it also rejects correct ones whenever a single evaluator misreads a claim or doesn't have enough context. Loosen the threshold and you cover more ground, at the cost of letting more mistakes through. Every evaluator you add pushes cost and latency up, and bringing in a human raises confidence but caps how much you can scale. There's no universal right answer here. A recommendation engine and an autonomous agent moving real money shouldn't run the same verification policy. What decides the threshold is what a wrong answer would actually cost you. This is roughly how Mira approaches it. Mira treats a model's output as a candidate for verification rather than a finished answer. The output gets broken into independently checkable claims. Those claims go out to multiple models that evaluate them separately. The responses get standardized so they're actually comparable, then checked against a consensus threshold. Clear the bar, and the system issues a cryptographic certificate recording exactly what was checked and how. That means verification work gets spread across independent nodes instead of running through one central pipeline. Claims get sharded, so no single node ever sees the whole output. Evaluator responses stay independent until everything's aggregated at the end, and node operators have real financial skin in the game if they cut corners on honest verification. The certificate doesn't prove a claim is true. What it proves is that a defined process ran and actually hit its threshold. That's a smaller claim than proving truth, but it's a more honest one to make. The numbers, from 78 test cases. Generator alone: 73.1% precision. Two-model consensus: 93.9%. Three-model consensus: 95.6%. Two-of-three consensus: 86.9%. Unanimous three-model agreement accepted 45 cases: 43 correct, two wrong, and 19 rejected that were actually fine. That last number matters. Strict consensus filtered out bad answers. It also turned away some good ones along the way. That's simply the cost of running conservative, and it's worth knowing what that cost is before setting a threshold, because finding out afterward gets expensive. Small sample, structured format. Testing this at scale, on messier and more open-ended output, is the natural next step. The larger point. Every bit of capability handed to AI comes with a matching bit of liability. A model that drafts contracts, moves money, or coordinates other agents raises the cost of whatever mistake slips through unnoticed, simply by doing more with less oversight along the way. AI is going to be wrong sometimes. That was never really in question. What actually matters is whether anything sits between a wrong output and the action it triggers. A longer prompt doesn't fix that. Neither does bolting on another confidence score. What actually helps is a system that can pull a claim apart and check it properly, then say no when it needs to and leave a record of why. That's the part worth actually building, even if it's the least glamorous line item on the roadmap.