CHEMICALS

Fouling predictions that carry their own confidence.

Fouling models for a preheat train either follow first principles with parameters that drift, or learn from the plant record and fail to extrapolate, and both return a single number with no measure of how far to trust it. A physics informed twin calibrated by Bayesian inference gives every forecast a credible interval, and Bayesian optimisation over it finds a cleaning plan in a handful of evaluations.

All case studies
Process unit stack and pipe racks against the sky, in black and white
Stack and pipework
KEY RESULTS
−43%
prediction error versus a pure data driven model
−58%
training data to reach the target accuracy
−34%
narrower credible intervals after calibration
−68%
fewer twin evaluations to the optimum plan
+38%
longer warning before the inefficient regime
180 days

of routine plant history is enough to calibrate the physical parameters, with no dedicated trials on the asset

0.31 K

mean coil inlet temperature error over a full run against the plant validated reference twin

95%

of measured operating points fall inside the calibrated credible interval, coverage checked against the plant history

THE PROBLEM

A point estimate with no interval is hard to act on.

Fouling models for a preheat train fall into two camps and each fails the operator in a different way. Mechanistic models are faithful to first principles but depend on deposition and removal parameters that drift as feedstock and duty change, and they return a single number with no measure of how much to trust it. Pure data driven models learn from the plant record but need long, clean, stable histories, do not extrapolate beyond the conditions they have seen, and again hand back a point prediction with no error bar. The decisions that hang on these models, when to take a shell offline to clean and how hard to fire the furnace, are high stakes and slow to reverse, while the labelled fouling data is sparse and punctuated by cleans. A bare point estimate with no interval around it is hard to act on with confidence.

01Coil inlet temperature over a 12 month run, measured points against the physics informed twin and a pure data driven model, both trained to month seven. The twin stays on the measurements across the whole run; the data driven model drifts once it leaves the data it was trained on.
Coil inlet temperature over a 12 month run, measured points against the physics informed twin and a pure data driven model, both trained to month seven. The twin stays on the measurements across the whole run; the data driven model drifts once it leaves the data it was trained on.
Months in runPhysics informedData driven
0245.6245.6
1244.4244.4
2243.3243.3
3241.5241.5
4240.1240.1
5238.4238.4
6237237
7235.8235.8
8234.5234.6
9231.8233.5
10228.9232.7
11226.3232
12225231.6
WHAT THE MODEL DOES

A grey box twin, calibrated with its uncertainty.

EntroMetrix builds a physics informed grey box twin of the preheat train. A mechanistic deposition and removal law forms the backbone, and a small machine learning residual learns only the part the physics does not capture, so the model stays data efficient and extrapolates where a black box would not. The physical parameters are then calibrated to the plant record by Bayesian inference, which yields a posterior over each parameter and therefore a credible interval on every forecast. A Gaussian process surrogate is fitted over the twin so that Bayesian optimisation can search cleaning timing and operating conditions in a handful of evaluations rather than an exhaustive sweep. The work is retrospective and offline first, validated on the plant cleaning history before it informs any live decision.

HOW IT WORKS
Grey box structure: a mechanistic core wrapped by a learned residual, so it is interpretable and extrapolates rather than memorising the record.
Mechanistic backbone: a first principles deposition and removal law drives the fouling and the coil inlet temperature decline between cleans.
Learned residual: a small machine learning term carries only the unmodelled part, keeping the fit data efficient.
Bayesian calibration: the physical parameters are inferred as posteriors from the plant record, not fitted to a single value.
Credible intervals: every prediction arrives with a calibrated interval, so a forecast is a range with a confidence.
Gaussian process surrogate: a surrogate over the twin lets Bayesian optimisation find good plans in few evaluations.
02The best cleaning plan found against the number of twin evaluations, as a share of the optimum. Bayesian optimisation over the Gaussian process surrogate reaches the optimum within about thirteen evaluations, where a space filling search is still well short after forty.
The best cleaning plan found against the number of twin evaluations, as a share of the optimum. Bayesian optimisation over the Gaussian process surrogate reaches the optimum within about thirteen evaluations, where a space filling search is still well short after forty.
Twin evaluationsBayesian optimisationSpace filling search
10.10.6
20.50.8
30.60.8
40.70.8
50.8
60.80.9
80.90.9
100.90.9
131.0
161.0
201.00.9
3010.9
4010.9
Optimum1
BUILT ON THE DATA YOU ALREADY HOLD

Built from 180 days of routine history, validated to 0.31 K.

The twin is built from data the plant already holds. Routine daily averages of stream temperatures, flow splits and duties over a single run of the hot end shells are enough to calibrate the mechanistic fouling parameters, because the physics does most of the work and only the residual is learned. Around 180 days of ordinary operating history gives stable long range predictions, so no dedicated trials, feed changes or interventions on the asset are required to stand the model up.

DEMONSTRATED ON A PREHEAT TRAIN

Accuracy and data efficiency are demonstrated, not asserted. Against a plant validated reference the calibrated twin tracks coil inlet temperature to a mean error near 0.31 K over a full run and matches the annual heat recovery cost of the high fidelity model to within about 1.4 per cent, across a network of tens of shells. On the same sparse and clean punctuated record the physics informed hybrid roughly halves the prediction error of a pure data driven model and holds its accuracy when the training history is thinned well below a fifth, exactly the regime a preheat train proof of concept faces.

03Prediction error on the same sparse fouling record, indexed to the worst, with the spread across runs. The physics informed hybrid is lowest: the mechanistic backbone regularises the fit and the residual carries the rest.
Prediction error on the same sparse fouling record, indexed to the worst, with the spread across runs. The physics informed hybrid is lowest: the mechanistic backbone regularises the fit and the residual carries the rest.
Prediction error (index)
Pure data driven100 ± 9
Pure physics61 ± 6
Physics informed27 ± 4
VALIDATION

Validated, calibrated and safe to inspect.

Every output is validated, calibrated and safe to inspect. Coverage of the credible intervals is checked against measured operation so the stated confidence is honest, the calibrated parameters stay physically interpretable, and the whole study runs offline and isolated on plant history. Bayesian optimisation over the Gaussian process surrogate then reaches a good cleaning and operating plan in a fraction of the evaluations an exhaustive search would need, turning a well calibrated twin into a decision the operator can weigh against its own uncertainty.

HOW YOU MAINTAIN CONTROL
Blind validation: calibrated on part of the plant history and tested on runs it did not see, tracking coil inlet temperature to a fraction of a degree.
Calibrated intervals: interval coverage is checked against measured operation, so the stated confidence matches how often the truth lands inside it.
Interpretable parameters: the calibrated quantities are physical deposition, removal and ageing constants an engineer can read and sanity check.
Extrapolates: the mechanistic backbone keeps predictions sensible in operating regions the data never visited.
Offline and isolated: retrospective by design, running on historical data, never in the control loop during the proof of concept.
Continuous recalibration: as each new clean and run is logged the posteriors are updated, so the twin stays current as feedstock and duty move.
04Prediction error as the training history is thinned, with the spread at each point. The physics informed twin holds its accuracy well below a fifth of the record, where a pure data driven model climbs steeply.
Prediction error as the training history is thinned, with the spread at each point. The physics informed twin holds its accuracy well below a fifth of the record, where a pure data driven model climbs steeply.
Training history used (%)Physics informedData driven
105.820.4
155.416.1
204.713.6
304.39.5
453.87.3
653.55.2
853.24.9
1003.74.2
OUTCOMES

Predictions that carry their own confidence.

Accurate where data is thin

The physics informed twin holds its accuracy well below a fifth of the record, where a pure data driven model climbs steeply.

A broad prior, a tight posterior

Bayesian inference turns a broad prior into a tight posterior on each fouling parameter, and the credible interval makes the remaining uncertainty explicit.

A forecast with a band

The twin says where the shell is heading and how sure it is, which a cleaning or firing decision is weighed against.

05A fouling parameter before and after calibration. Bayesian inference turns a broad prior into a tight posterior, and the shaded 95% credible interval makes the remaining uncertainty explicit.
A fouling parameter before and after calibration. Bayesian inference turns a broad prior into a tight posterior, and the shaded 95% credible interval makes the remaining uncertainty explicit.
Deposition constant (relative)PosteriorPrior
0.400.1
1.200.7
1.30.0
1.30.1
1.40.8
1.42.30.8
1.43.9
1.54.5
1.63.90.9
1.62.4
1.61.1
1.70.50.8
1.80.2
1.90.1
20
2.500.1
06A forward forecast of coil inlet temperature with its 90% credible band, measured points behind it. The twin says where the shell is heading and how sure it is, which a cleaning or firing decision is weighed against.
A forward forecast of coil inlet temperature with its 90% credible band, measured points behind it. The twin says where the shell is heading and how sure it is, which a cleaning or firing decision is weighed against.
Months from todayForecast
-12240
-10239.5
-8238.3
-6237
-4236
-3234.6
-2233.9
-1233.2
0233.1
2231.8
4230.7
6229.2
8227.6
10226.2
12224

See what the model finds in your plant.

A short call with our engineering team, your process data stays on site.