Reinforcement Learning in Control: 6 Proven Steps to Success

Share:
Industrial AI
Reinforcement Learning in Control: 6 Proven Steps to Success

Instead of being tuned by hand, some controllers now learn by trial and error inside a plant simulator until they find the best moves. The first industrial trials show that this idea can run real units safely for weeks.

Agent and Reward FKDPP Digital Twin PID and MPC

Classic loops follow fixed equations, while learning agents discover a control policy from experience and rewards. This guide explains the core ideas, a real chemical plant trial and how to calculate discounted return.

Hello everyone, today we are going to learn how reinforcement learning works in process control, how an agent learns from rewards and how it compares with PID and MPC.
reinforcement learning

What Is Reinforcement Learning in Process Control?

Reinforcement learning is a branch of machine learning in which an agent learns a control policy by taking actions on an environment and receiving rewards, aiming to maximise the total reward over time. In process control the environment is a plant or its simulator, and the actions are valve or setpoint moves, extending the ideas in AI and machine learning.

No one writes the control law directly. A reinforcement learning agent discovers it by trying moves and learning which ones lead to good long term results.

JSR chemical plant where an AI agent ran autonomous control
Image credit: Yokogawa

Yokogawa reports that its AI agent controlled a JSR chemical plant autonomously for 35 consecutive days, or 840 hours, from January 17 to February 21, 2022. Only on spec product was made during the trial.

The agent used the FKDPP algorithm, which Yokogawa developed with the Nara Institute of Science and Technology in 2018. It is one of the first confirmed long runs of this kind in a real chemical plant.

Core Terms in Plain Words

TermMeaningProcess Example
AgentThe learner that picks actionsAI controller
EnvironmentWhat the agent acts onDistillation column or simulator
StateWhat the agent observesTemperatures, flows, levels
ActionA move the agent makesChange a valve or setpoint
RewardScore for the resultOn spec product, low energy
PolicyRule from state to actionLearned control strategy

Rewards must be designed with care, because a reinforcement learning agent will exploit any loophole. A reward that ignores energy use may give tight quality with wasteful steam flow.

The discount factor gamma decides how much future rewards count. Values near 1 make the agent plan far ahead, which suits slow process units.

6 Steps to Apply Reinforcement Learning

1
Define the Goal
Pick quality, energy or throughput targets.
2
Build a Simulator
Use a first principles model or digital twin.
3
Design the Reward
Score good outcomes and penalise constraint breaks.
4
Train Offline
Let the agent learn millions of moves safely.
5
Test in Shadow Mode
Compare suggestions with operator actions.
6
Deploy with Guards
Add limits, overrides and operator control.

Training on the real plant would be slow and risky, so most work happens in a digital twin. The simulator must capture delays, noise and disturbances well enough.

Deployment usually sits above existing loops in the distributed control system. The agent writes setpoints, while regulatory loops still move the valves.

Compared with PID and MPC

PID

Fixed feedback law, tuned by hand.

Best for: single loops, fast response
Proven
MPC

Model based optimisation each cycle.

Best for: multivariable units with constraints
Mature
RL Agent

Policy learned from rewards.

Best for: nonlinear units, hard to model
Emerging

PID remains the base layer, as covered in PID controller types and the PID tuning guide. Structures like cascade control stay in place underneath a learning agent.

MPC needs an explicit model, while reinforcement learning needs a good simulator and reward. Advanced schemes are listed in DCS control strategies.

Discounted Return Formula

G = r × (1 minus γⁿ) ÷ (1 minus γ)

r = constant reward per step, γ = discount factor, n = number of steps

Example:
r = 1, γ = 0.9, n = 10
0.9¹⁰ = 0.3487
G = 1 × 0.6513 ÷ 0.1 = 6.513
Infinite horizon limit = r ÷ (1 minus γ) = 10

The formula shows why later rewards matter less when gamma is small. When gamma equals 1, the return is simply r times n.

Discounted Return Calculator

Constant Reward Discounted Return
Result
Discounted return 6.513, infinite horizon limit 10.000

Try gamma of 0.99 to see how a far sighted agent values the long run. Negative rewards, used as penalties, are shown with the word minus.

Advantages
  • Handles nonlinear and hard to model units.
  • Can optimise several goals together.
  • Learns from simulation without plant risk.
  • Adapts when retrained on new data.
Limitations
  • Needs an accurate simulator.
  • Reward design is tricky.
  • Hard to prove stability and safety.
  • Operators may distrust black box moves.

Safety constraints must sit outside the agent, in hard limits and the safety system. Where the model runs is discussed in edge AI vs cloud AI.

For a broader view of AI in control rooms, read AI in PLC, SCADA and DCS. Start with advisory mode before closing the loop.

Deep RL Primer PDF

PDF
Deep Reinforcement Learning for Process Control: A Primer for Beginners
Spielberg and colleagues, University of British Columbia

Literature Review Video

Reinforcement Learning FAQ

What is reinforcement learning?
It is a machine learning method where an agent learns actions by trial and error to maximise reward. The learned rule from state to action is called the policy.
Can it replace PID controllers?
Not in the base layer, where PID is simple and proven. Learning agents usually sit above PID loops and adjust their setpoints.
What did Yokogawa and JSR achieve?
Yokogawa reports an AI agent ran a JSR plant autonomously for 35 consecutive days in 2022. Only on spec product was made during that period.
What is FKDPP?
It is the learning algorithm Yokogawa developed with the Nara Institute of Science and Technology in 2018. It was designed to work with limited plant data.
Why train in a simulator?
Trial and error on a live plant would be slow, costly and unsafe. A digital twin lets the agent practise millions of moves without risk.
What is the discount factor?
Gamma sets how much future rewards count compared with immediate ones. Values close to 1 suit slow processes with long delays.
How is safety ensured?
Hard limits, overrides and the independent safety system stay outside the agent. Operators keep the ability to switch back to normal control.

Related Articles

External References

What We Learn Today

  • Reinforcement learning trains an agent to maximise reward through trial and error.
  • Yokogawa and JSR ran a plant autonomously for 35 days.
  • Train in a simulator and keep safety limits outside the agent.
I hope you like above blog. There is no cost associated in sharing the article in your social media. Thanks for reading!! Happy Learning!!

Leave a Reply

Your email address will not be published. Required fields are marked *