What changed in schedule risk
Quantitative schedule risk analysis has for thirty years meant a planner assigning three-point estimates to selected activities and running Monte Carlo over them — a process whose output is governed less by evidence than by the shape of distribution someone chose in a workshop. The new approach replaces the assumption with a base rate. nPlan trains deep-learning models (large language models over activity text plus graph neural networks over the logic network) on what its performance report describes as 539,569 project schedules and 358,871,102 activities as of February 2023, a corpus the company now puts at more than 750,000 programme files. The model learns, activity by activity, what was planned versus what actually happened, then produces a full duration distribution for every activity in a new programme rather than a triangular guess on a handful.
The measured result, in the vendor's own numbers
nPlan publishes its methodology: 70% of projects train the model, 30% evaluate it, repeated across random splits. On a test set of 14,184 projects, the graph model beats PERT on every reported metric — continuous ranked probability score 64.1 against PERT's 703.5 at activity level, and at project level a delay multiplier (forecast duration over actual) of 1.02/1.14/1.41 at P10/P50/P90 against PERT's 1.53/2.04/2.77, where 1.0 is perfect. The more useful number for a delivery team is detection: the model identified 47.2% of activities that actually overran by more than 50%, and 47.9% of projects that overran by more than 30%, from the baseline schedule alone. PERT identified none of the severe activity overruns. Because these were realised delays, they were by definition missed by whatever risk process the projects were running at the time.
The field test that made infrastructure pay attention
Network Rail tested the approach on two of its largest programmes — the Great Western Main Line and the Salisbury to Exeter signalling project, together representing over £3bn of capital expenditure — and reported that cost savings of up to £30m could have been achieved on Great Western alone, primarily by surfacing risks that were invisible to the team because of the sheer size and complexity of the programme data. Network Rail then moved to roll the software out across 40 projects before scaling further. Note the shape of the claim: not that the AI would have built the railway faster, but that it would have flagged specific activities early enough for cheap mitigation instead of expensive recovery.
This is reference-class forecasting, industrialised
None of the underlying logic is new to government. HM Treasury's Green Book supplementary guidance has told UK appraisers for two decades that there is a systematic tendency to optimism and that explicit, empirically based uplifts to cost and duration should be applied from data on past projects. Reference class forecasting — comparing your project to the outcome distribution of similar completed projects — is established practice in the UK, the Netherlands, Denmark, Switzerland and Australia. What the models change is granularity and effort: instead of one uplift applied to a project class by a policy table, you get a learned distribution on every one of tens of thousands of activities, generated in minutes. It is the outside view, applied inside the network logic.
What the model cannot see
The forecast is a function of the programme you feed it. If the logic is thin — missing links, hard constraints substituted for real dependencies, artificial lags, summary-level activities hiding the work — the model produces a well-calibrated forecast of a fiction. It reads structure and text, not intent: it cannot know that the client is about to change the brief, that the crane is single-source, or that the subcontractor pricing the works has no crew. Nor can it separate correlation from cause; it learns that activities of a certain type in a certain position historically overrun, which is a base rate, not a diagnosis. And the accuracy figures above, though cross-validated and unusually transparent for this sector, are self-reported on the vendor's own corpus and have not been independently audited. Ask any vendor for performance on projects like yours, on a holdout you control.
The failure mode is organisational, not statistical
The recurring way these deployments die has nothing to do with the mathematics. A credible P80 completion date is commercially inconvenient: it appears in a data room, an assurance review, a delay claim. Teams learn to run the analysis late, treat the output as a document rather than a decision input, or quietly re-baseline until the forecast agrees with the promise. The projects that get value do the opposite — they run the forecast on the baseline before commitment, mitigate the top-ranked activities while mitigation is still cheap, and re-run it each period against actual progress, treating the divergence as management information rather than an accusation. That is a governance and contracting choice, and no model can make it for you.
The Australian angle
Infrastructure Australia's 2025 Market Capacity Report puts the five-year Major Public Infrastructure Pipeline at $242 billion, its highest level since tracking began, driven by energy transmission and housing — against persistent worker shortages and stagnating construction productivity, with the agency explicitly recommending that governments incentivise productivity-enhancing innovation. Forecasting does not add a single crew to that market. What it does is tell an owner which of its programmes are structurally likely to slip and by how much, early enough to sequence the pipeline rather than let every project compete for the same scarce trades in the same quarter. In a capacity-constrained market, better allocation is the available productivity gain.
What it means for delivery teams
Three practical consequences. First, the value of your historical data just went up: the model that forecasts your next job learns from the as-built actuals of your last one, so recording true activity start and finish dates — not the dates that made the report look tidy — is now a commercial asset. Second, forecasting rewards schedule hygiene, because a well-linked, properly resourced programme yields a forecast worth acting on. Third, the output changes what a risk workshop is for: not arguing about whether a date is achievable, but choosing which of the twenty activities the model ranked highest you are going to attack this month.
The bar to hold vendors to
Any AI that forecasts your programme should be able to show its working: which historical patterns drove the forecast, which activities carry the risk, how the forecast moved when the last period's actuals landed, and how it performed on your completed projects. A number without provenance is not a forecast, it is an opinion with a decimal point.
SITELIVE 