← All stories
Show the working

An estimating AI has published its accuracy numbers — the benchmark matters more than the number

Provision released Scope Agent to general availability with a report comparing its scope extraction against a general-purpose model on a live hospital project. The claim worth examining is not 97% — it is zero fabricated line items.

AI & Autonomy5 August 2026 · 6 min read · SiteLive News desk
Cover art: SiteLive News

What shipped

Toronto-based Provision announced on 5 August 2026 the general availability of Scope Agent, a system that generates trade scope-of-work packages from drawings and specifications, and published alongside it a benchmark report comparing Scope Agent with Anthropic’s Claude on plumbing scope extraction from a live hospital project. On its own internal validation dataset, measured against human-prepared scope sheets from real North American projects, Provision reports 97.01% match accuracy on scope item extraction and 96.02% subcontractor classification accuracy (top two). On the hospital drawings, it reports Scope Agent extracting 145 plumbing scope items at 91.7% verified accuracy against Claude’s 96 items at 72.9%. In ENR’s reporting the same week, California contractor ProWest describes using the agent first on jobs it had already won, to get subcontracts issued faster, rather than on bids.

The number that actually matters

Buried in the percentages is the finding practitioners should read twice. Of 26 flagged rows in the general-purpose model’s output, Provision says 18 involved invented drawing marks, manufacturer model numbers or schedule items that were absent from the drawings; Scope Agent, it reports, produced zero fabricated items. That is the difference between a tool that is inaccurate and a tool that is unsafe. A missed scope item is a commercial risk you can bound by review; an invented fixture tag that reads as plausible and traces to nothing is a document you might issue to a subcontractor. Which is why the traceability claim — every line cited to a document and section — is the load-bearing feature, not the headline accuracy. If a line cannot be clicked back to a sheet, its accuracy is unverifiable in the only place that counts, which is the bid you are signing.

The honest limits

This is a vendor-run benchmark of a vendor’s product, on a dataset the vendor assembled, published on the day the product went on sale. No independent party replicated it. The comparison is against a general-purpose chat model rather than a competing construction tool, which is the easier fight — general models were never built to read a 2,400-page project set. It covers one trade on one project. And the vendor’s own figures move depending on where you look: the release headlines 97.01% on its internal dataset while Provision’s own published buyer’s guide lists 95% verified accuracy for Scope Agent, and the live-project plumbing test came in at 91.7%. Those are different measurements, not contradictions, but they show how much a single accuracy percentage depends on the definition behind it. The comparison to a human estimator at 91.3% is also the vendor’s framing — an estimator "working to a fixed time budget", a condition Provision defined.

What it means for estimators

The useful shift is that a vendor put a falsifiable claim on the table, with a methodology attached. Ask every AI estimating vendor for the same three things and the field sorts itself quickly: the accuracy definition (item-level match against what baseline?), the fabrication rate, and whether every output line carries a citation to a drawing or spec section. Then run your own measurement, because it is cheap. Take a project you have already closed out, run the tool over the original set, and count two things: scope items the tool missed that later became variations, and items the tool invented. That is your accuracy figure, on your drawings, in your trades — and it is the only one worth quoting internally.

The SiteLive take

Treat AI scope output as a first pass that must be traceable, not as a scope sheet. The number to demand from any vendor is the fabrication rate, and the number to trust is the one you measure on a job you have already built. SiteLive takes the same position on takeoff and pricing — every quantity linked back to the drawing it came from, so an estimator can check the line rather than believe it.

Sources

Share

SiteLive News is edited for people who build. We publish only stories that clear a hard bar — a genuine technical advance, real project data, or a change to how construction, mining, manufacturing and haulage actually work. Every factual claim is grounded in the named sources linked from the piece; analysis is our own and labelled as such. Produced with AI-assisted research under human editorial direction. No sponsored content, no wire rewrites, no filler.

The SiteLive Briefing

The stories that clear our bar — construction, mining, manufacturing, logistics, energy, safety and industrial AI — delivered to your inbox when they publish. No filler, unsubscribe anytime.

You're on the list — the next briefing will land in your inbox.

More from SiteLive News