What shipped
Toronto-based Provision announced on 5 August 2026 the general availability of Scope Agent, a system that generates trade scope-of-work packages from drawings and specifications, and published alongside it a benchmark report comparing Scope Agent with Anthropic’s Claude on plumbing scope extraction from a live hospital project. On its own internal validation dataset, measured against human-prepared scope sheets from real North American projects, Provision reports 97.01% match accuracy on scope item extraction and 96.02% subcontractor classification accuracy (top two). On the hospital drawings, it reports Scope Agent extracting 145 plumbing scope items at 91.7% verified accuracy against Claude’s 96 items at 72.9%. In ENR’s reporting the same week, California contractor ProWest describes using the agent first on jobs it had already won, to get subcontracts issued faster, rather than on bids.
The number that actually matters
Buried in the percentages is the finding practitioners should read twice. Of 26 flagged rows in the general-purpose model’s output, Provision says 18 involved invented drawing marks, manufacturer model numbers or schedule items that were absent from the drawings; Scope Agent, it reports, produced zero fabricated items. That is the difference between a tool that is inaccurate and a tool that is unsafe. A missed scope item is a commercial risk you can bound by review; an invented fixture tag that reads as plausible and traces to nothing is a document you might issue to a subcontractor. Which is why the traceability claim — every line cited to a document and section — is the load-bearing feature, not the headline accuracy. If a line cannot be clicked back to a sheet, its accuracy is unverifiable in the only place that counts, which is the bid you are signing.
The honest limits
This is a vendor-run benchmark of a vendor’s product, on a dataset the vendor assembled, published on the day the product went on sale. No independent party replicated it. The comparison is against a general-purpose chat model rather than a competing construction tool, which is the easier fight — general models were never built to read a 2,400-page project set. It covers one trade on one project. And the vendor’s own figures move depending on where you look: the release headlines 97.01% on its internal dataset while Provision’s own published buyer’s guide lists 95% verified accuracy for Scope Agent, and the live-project plumbing test came in at 91.7%. Those are different measurements, not contradictions, but they show how much a single accuracy percentage depends on the definition behind it. The comparison to a human estimator at 91.3% is also the vendor’s framing — an estimator "working to a fixed time budget", a condition Provision defined.
What it means for estimators
The useful shift is that a vendor put a falsifiable claim on the table, with a methodology attached. Ask every AI estimating vendor for the same three things and the field sorts itself quickly: the accuracy definition (item-level match against what baseline?), the fabrication rate, and whether every output line carries a citation to a drawing or spec section. Then run your own measurement, because it is cheap. Take a project you have already closed out, run the tool over the original set, and count two things: scope items the tool missed that later became variations, and items the tool invented. That is your accuracy figure, on your drawings, in your trades — and it is the only one worth quoting internally.
SITELIVE