CHEMICALS

Near optimal crude and production schedules in seconds, not hours.

Crude scheduling is one of the largest combinatorial problems a refinery solves, and it has to be worked again every time a tanker slips, a tank frees or a unit rate changes. Exact solvers reach true optima but scale badly; manual planning is fast but conservative. A reinforcement learning policy trained on the operator's own history returns a feasible schedule in seconds, within a few percent of the mixed integer optimum, and a Pareto set the planner chooses from.

All case studies
Offshore production platform at sea, in black and white
Offshore platform
KEY RESULTS
+3.4%
gross margin against the manual baseline
−8.6%
demurrage and tardiness cost
−7.3%
CDU idle and changeover events
−4.2%
slop and off spec production
−99.2%
time to reschedule after a disturbance
7 to 10 days

planning horizon per schedule, across two crude units and up to eleven process units

under 5%

margin gap of the learned policy against the mixed integer optimum on held out schedules

under 5 s

to return a full reschedule after a tanker slips or a unit rate changes

THE PROBLEM

Exact solvers scale badly, and manual plans leave margin on the table.

Crude scheduling is one of the largest combinatorial problems a refinery solves. Planners must sequence tanker and pipeline receipts, allocate storage and charging tanks, fix blend recipes against sulphur, acidity, density and residue limits, and set the charge to each distillation unit across a horizon of one to two weeks. The schedule has to be worked again every time a tanker slips, a tank frees, a blend target moves or a unit rate changes. Exact mixed integer formulations reach true optima but scale badly, growing from under a minute at five days to tens of hours at ten days, and any disturbance forces a full recomputation. Manual planning is fast but conservative, and leaves margin on the table.

01Margin captured, as a share of the mixed integer optimum, over 2.2 million training steps, drawn as the smoothed curve. The learned policy climbs toward the optimum during training and settles within a few percent of it.
Margin captured, as a share of the mixed integer optimum, over 2.2 million training steps, drawn as the smoothed curve. The learned policy climbs toward the optimum during training and settles within a few percent of it.
Training steps (millions)RL policy
05
0.115
0.125
0.138
0.250
0.354
0.363
0.370
0.475
0.579
0.583
0.685
0.686
0.788
0.892
0.991
194
1.295
1.495.5
1.695.5
1.895.5
295.5
2.295.5
Mixed integer optimum100
WHAT THE SCHEDULER DOES

A scheduling environment on your history, and a policy that stays feasible.

EntroMetrix builds a scheduling environment on the operator's own historical data and trains a reinforcement learning policy to sequence receipts, allocate tanks, set blends and charge the units. The problem is cast as a Markov decision process: the state carries tank inventories, crude qualities and vessel arrivals, the uncertain ones held as Bayesian distributions rather than single values, the action sets receipt, transfer and charge volumes, and the reward is gross margin from a physics informed blend and yield model, net of demurrage, changeover and slop. Every action is passed through a math programming projection, so it lands on the feasible region set by tank capacities, connectivity, blend properties and unit rates before it is applied. Once trained the policy returns a schedule in seconds when a tanker slips or a unit trips, staying within a few percent of the mixed integer optimum, and surfaces a Pareto set the planner chooses from.

HOW IT WORKS
Sequences receipts: orders tanker and pipeline arrivals and their discharge into storage.
Allocates tanks: assigns storage and charging tanks and their inventories across the horizon.
Sets blends: fixes blend recipes to meet sulphur, acidity, density and residue limits.
Charges the units: sets the feed rate and crude to each distillation unit each period.
Physics informed margin: a physics informed blend and yield model turns each crude into product cuts and values, so the reward is real margin net of demurrage, changeover and slop.
Re optimises fast: returns a fresh schedule in seconds when arrivals or unit rates change.
02Margin captured against time to a schedule, on a log scale. The mixed integer optimum is the ceiling; the policy captures almost all of it in seconds, one to two orders of magnitude faster than the solver, while manual planning gives up more margin.
Margin captured against time to a schedule, on a log scale. The mixed integer optimum is the ceiling; the policy captures almost all of it in seconds, one to two orders of magnitude faster than the solver, while manual planning gives up more margin.
Time to a schedule (s), log scaleMargin captured (% of optimum)
Manual planning500088
Mixed integer solver22099.9
RL policy1.595.3
BUILT ON THE DATA YOU ALREADY HOLD

Built from receipt logs and tank movements, within a few percent of the optimum.

The scheduler is built from data the plant already holds. Historical receipt logs, tank movements, blend records, unit rates and the crude assay library are enough to reconstruct the operator's own decisions and to train the policy against them. The logistics and quality limits and the operating window are set with the schedulers before any run.

DEMONSTRATED ON A CRUDE SCHEDULING PROBLEM

On a comparable crude scheduling problem the learned policy reached within a few percent of the mixed integer optimum while returning a schedule in seconds, one to two orders of magnitude faster than a full re solve, and held that margin when a receipt was delayed because it was trained against the Bayesian distribution of arrival times and could re optimise inside the horizon rather than run a stale plan. Feasibility was guaranteed throughout by projecting every action onto the region the plant can actually operate, so proposed schedules stayed inside the accepted limits rather than breaching them as a policy trained on penalties alone would.

03Gross margin, demurrage, unit idle and slop over the same horizon and disturbances, the learned policy against the manual baseline, indexed to the baseline at 100. Margin rises while the three cost and loss objectives fall.
Gross margin, demurrage, unit idle and slop over the same horizon and disturbances, the learned policy against the manual baseline, indexed to the baseline at 100. Margin rises while the three cost and loss objectives fall.
ManualRL policy
Gross margin100103.4
Demurrage10091.4
CDU idle / changeover10092.7
Slop / off spec10095.8
VALIDATION

Planner in the loop, with feasibility held by projection.

Validation is retrospective and blind. Schedules are compared against the operator's own historical schedules on held out periods before any live use, the policy is held inside logistics and process limits by a math programming projection and a constrained objective, and the scheduler runs offline for decision support, with the planner reviewing and approving the Pareto set before a schedule is issued.

HOW YOU MAINTAIN CONTROL
Projection: a math programming layer snaps each action onto the feasible region before it is applied.
Masking: simple logistics limits such as connectivity and tank counts are enforced by construction.
Safe limits: complex process and quality limits are held by a constrained policy that caps cumulative violation.
Blind validation: schedules are compared against historical operator schedules on held out periods.
Planner in the loop: the planner reviews the Pareto set and approves the schedule before it is issued.
Robust to uncertainty: vessel arrivals, crude assays and demand are carried as Bayesian distributions, so the policy holds margin across that uncertainty.
04Realised margin over 48 hours when a tanker arrives late at hour 24, drawn as an hourly path. The learned policy reschedules within seconds and recovers, while a schedule frozen before the disturbance loses margin it cannot win back without a full re solve.
Realised margin over 48 hours when a tanker arrives late at hour 24, drawn as an hourly path. The learned policy reschedules within seconds and recovers, while a schedule frozen before the disturbance loses margin it cannot win back without a full re solve.
Hours into the horizonFrozen scheduleRL policy
0100100
499.599.5
8101101
1298.5100.5
16100.5100.5
20100.8100.8
22102.5102.5
2410096
259293.5
2686.5
308796.5
3486.3
3887
4288
4686.5
488698.7
OUTCOMES

Near optimal schedules in seconds, not hours.

Recovers from a late tanker

When a tanker arrives late the learned policy reschedules within seconds and recovers, while a schedule frozen before the disturbance loses margin it cannot win back without a full re solve.

A steady charge rate

The learned policy holds each distillation unit within a narrow band, while a solver chasing the cheapest paper optimum swings it and switches units to zero throughput.

Safely below the threshold

A policy trained by penalty alone still breaches the accepted threshold; projecting every action onto the feasible set holds the schedule safely below it.

05Charge rate to a distillation unit over the horizon, drawn as an hourly path. The learned policy holds the unit within a narrow band, while a solver chasing the cheapest paper optimum swings it and switches the unit to zero throughput.
Charge rate to a distillation unit over the horizon, drawn as an hourly path. The learned policy holds the unit within a narrow band, while a solver chasing the cheapest paper optimum swings it and switches the unit to zero throughput.
Hours into the horizonMixed integer solverRL policy
0105102
290108
492106
680110
890104
958
1075116
1295102
14100100
15125
160102
170
184098
19100
22100101
24110103
269598
289099
3095102
32124103
3320
34096
3510
3758
3970
4180
43118
4575
47100
4895104
06Cumulative breach of the plant's limits over the horizon against the accepted threshold. A policy trained by penalty alone still breaches it; projecting every action onto the feasible set holds the schedule safely below.
Cumulative breach of the plant's limits over the horizon against the accepted threshold. A policy trained by penalty alone still breaches it; projecting every action onto the feasible set holds the schedule safely below.
Cumulative limit violation
Penalty trained RL128
Projected safe RL16

See what the model finds in your plant.

A short call with our engineering team, your process data stays on site.