Fouling models for a preheat train either follow first principles with parameters that drift, or learn from the plant record and fail to extrapolate, and both return a single number with no measure of how far to trust it. A physics informed twin calibrated by Bayesian inference gives every forecast a credible interval, and Bayesian optimisation over it finds a cleaning plan in a handful of evaluations.
All case studies
of routine plant history is enough to calibrate the physical parameters, with no dedicated trials on the asset
mean coil inlet temperature error over a full run against the plant validated reference twin
of measured operating points fall inside the calibrated credible interval, coverage checked against the plant history
Fouling models for a preheat train fall into two camps and each fails the operator in a different way. Mechanistic models are faithful to first principles but depend on deposition and removal parameters that drift as feedstock and duty change, and they return a single number with no measure of how much to trust it. Pure data driven models learn from the plant record but need long, clean, stable histories, do not extrapolate beyond the conditions they have seen, and again hand back a point prediction with no error bar. The decisions that hang on these models, when to take a shell offline to clean and how hard to fire the furnace, are high stakes and slow to reverse, while the labelled fouling data is sparse and punctuated by cleans. A bare point estimate with no interval around it is hard to act on with confidence.
| Months in run | Physics informed | Data driven |
|---|---|---|
| 0 | 245.6 | 245.6 |
| 1 | 244.4 | 244.4 |
| 2 | 243.3 | 243.3 |
| 3 | 241.5 | 241.5 |
| 4 | 240.1 | 240.1 |
| 5 | 238.4 | 238.4 |
| 6 | 237 | 237 |
| 7 | 235.8 | 235.8 |
| 8 | 234.5 | 234.6 |
| 9 | 231.8 | 233.5 |
| 10 | 228.9 | 232.7 |
| 11 | 226.3 | 232 |
| 12 | 225 | 231.6 |
EntroMetrix builds a physics informed grey box twin of the preheat train. A mechanistic deposition and removal law forms the backbone, and a small machine learning residual learns only the part the physics does not capture, so the model stays data efficient and extrapolates where a black box would not. The physical parameters are then calibrated to the plant record by Bayesian inference, which yields a posterior over each parameter and therefore a credible interval on every forecast. A Gaussian process surrogate is fitted over the twin so that Bayesian optimisation can search cleaning timing and operating conditions in a handful of evaluations rather than an exhaustive sweep. The work is retrospective and offline first, validated on the plant cleaning history before it informs any live decision.
| Twin evaluations | Bayesian optimisation | Space filling search |
|---|---|---|
| 1 | 0.1 | 0.6 |
| 2 | 0.5 | 0.8 |
| 3 | 0.6 | 0.8 |
| 4 | 0.7 | 0.8 |
| 5 | 0.8 | |
| 6 | 0.8 | 0.9 |
| 8 | 0.9 | 0.9 |
| 10 | 0.9 | 0.9 |
| 13 | 1.0 | |
| 16 | 1.0 | |
| 20 | 1.0 | 0.9 |
| 30 | 1 | 0.9 |
| 40 | 1 | 0.9 |
| Optimum | 1 |
The twin is built from data the plant already holds. Routine daily averages of stream temperatures, flow splits and duties over a single run of the hot end shells are enough to calibrate the mechanistic fouling parameters, because the physics does most of the work and only the residual is learned. Around 180 days of ordinary operating history gives stable long range predictions, so no dedicated trials, feed changes or interventions on the asset are required to stand the model up.
Accuracy and data efficiency are demonstrated, not asserted. Against a plant validated reference the calibrated twin tracks coil inlet temperature to a mean error near 0.31 K over a full run and matches the annual heat recovery cost of the high fidelity model to within about 1.4 per cent, across a network of tens of shells. On the same sparse and clean punctuated record the physics informed hybrid roughly halves the prediction error of a pure data driven model and holds its accuracy when the training history is thinned well below a fifth, exactly the regime a preheat train proof of concept faces.
| Prediction error (index) | |
|---|---|
| Pure data driven | 100 ± 9 |
| Pure physics | 61 ± 6 |
| Physics informed | 27 ± 4 |
Every output is validated, calibrated and safe to inspect. Coverage of the credible intervals is checked against measured operation so the stated confidence is honest, the calibrated parameters stay physically interpretable, and the whole study runs offline and isolated on plant history. Bayesian optimisation over the Gaussian process surrogate then reaches a good cleaning and operating plan in a fraction of the evaluations an exhaustive search would need, turning a well calibrated twin into a decision the operator can weigh against its own uncertainty.
| Training history used (%) | Physics informed | Data driven |
|---|---|---|
| 10 | 5.8 | 20.4 |
| 15 | 5.4 | 16.1 |
| 20 | 4.7 | 13.6 |
| 30 | 4.3 | 9.5 |
| 45 | 3.8 | 7.3 |
| 65 | 3.5 | 5.2 |
| 85 | 3.2 | 4.9 |
| 100 | 3.7 | 4.2 |
The physics informed twin holds its accuracy well below a fifth of the record, where a pure data driven model climbs steeply.
Bayesian inference turns a broad prior into a tight posterior on each fouling parameter, and the credible interval makes the remaining uncertainty explicit.
The twin says where the shell is heading and how sure it is, which a cleaning or firing decision is weighed against.
| Deposition constant (relative) | Posterior | Prior |
|---|---|---|
| 0.4 | 0 | 0.1 |
| 1.2 | 0 | 0.7 |
| 1.3 | 0.0 | |
| 1.3 | 0.1 | |
| 1.4 | 0.8 | |
| 1.4 | 2.3 | 0.8 |
| 1.4 | 3.9 | |
| 1.5 | 4.5 | |
| 1.6 | 3.9 | 0.9 |
| 1.6 | 2.4 | |
| 1.6 | 1.1 | |
| 1.7 | 0.5 | 0.8 |
| 1.8 | 0.2 | |
| 1.9 | 0.1 | |
| 2 | 0 | |
| 2.5 | 0 | 0.1 |
| Months from today | Forecast |
|---|---|
| -12 | 240 |
| -10 | 239.5 |
| -8 | 238.3 |
| -6 | 237 |
| -4 | 236 |
| -3 | 234.6 |
| -2 | 233.9 |
| -1 | 233.2 |
| 0 | 233.1 |
| 2 | 231.8 |
| 4 | 230.7 |
| 6 | 229.2 |
| 8 | 227.6 |
| 10 | 226.2 |
| 12 | 224 |
A short call with our engineering team, your process data stays on site.