Table of Contents
ToggleInstead of being tuned by hand, some controllers now learn by trial and error inside a plant simulator until they find the best moves. The first industrial trials show that this idea can run real units safely for weeks.
Classic loops follow fixed equations, while learning agents discover a control policy from experience and rewards. This guide explains the core ideas, a real chemical plant trial and how to calculate discounted return.

What Is Reinforcement Learning in Process Control?
Reinforcement learning is a branch of machine learning in which an agent learns a control policy by taking actions on an environment and receiving rewards, aiming to maximise the total reward over time. In process control the environment is a plant or its simulator, and the actions are valve or setpoint moves, extending the ideas in AI and machine learning.
No one writes the control law directly. A reinforcement learning agent discovers it by trying moves and learning which ones lead to good long term results.

Yokogawa reports that its AI agent controlled a JSR chemical plant autonomously for 35 consecutive days, or 840 hours, from January 17 to February 21, 2022. Only on spec product was made during the trial.
The agent used the FKDPP algorithm, which Yokogawa developed with the Nara Institute of Science and Technology in 2018. It is one of the first confirmed long runs of this kind in a real chemical plant.
Core Terms in Plain Words
| Term | Meaning | Process Example |
|---|---|---|
| Agent | The learner that picks actions | AI controller |
| Environment | What the agent acts on | Distillation column or simulator |
| State | What the agent observes | Temperatures, flows, levels |
| Action | A move the agent makes | Change a valve or setpoint |
| Reward | Score for the result | On spec product, low energy |
| Policy | Rule from state to action | Learned control strategy |
Rewards must be designed with care, because a reinforcement learning agent will exploit any loophole. A reward that ignores energy use may give tight quality with wasteful steam flow.
The discount factor gamma decides how much future rewards count. Values near 1 make the agent plan far ahead, which suits slow process units.
6 Steps to Apply Reinforcement Learning
Training on the real plant would be slow and risky, so most work happens in a digital twin. The simulator must capture delays, noise and disturbances well enough.
Deployment usually sits above existing loops in the distributed control system. The agent writes setpoints, while regulatory loops still move the valves.
Compared with PID and MPC
Fixed feedback law, tuned by hand.
Model based optimisation each cycle.
Policy learned from rewards.
PID remains the base layer, as covered in PID controller types and the PID tuning guide. Structures like cascade control stay in place underneath a learning agent.
MPC needs an explicit model, while reinforcement learning needs a good simulator and reward. Advanced schemes are listed in DCS control strategies.
Discounted Return Formula
r = constant reward per step, γ = discount factor, n = number of steps
Example:
r = 1, γ = 0.9, n = 10
0.9¹⁰ = 0.3487
G = 1 × 0.6513 ÷ 0.1 = 6.513
Infinite horizon limit = r ÷ (1 minus γ) = 10
The formula shows why later rewards matter less when gamma is small. When gamma equals 1, the return is simply r times n.
Discounted Return Calculator
Try gamma of 0.99 to see how a far sighted agent values the long run. Negative rewards, used as penalties, are shown with the word minus.
- Handles nonlinear and hard to model units.
- Can optimise several goals together.
- Learns from simulation without plant risk.
- Adapts when retrained on new data.
- Needs an accurate simulator.
- Reward design is tricky.
- Hard to prove stability and safety.
- Operators may distrust black box moves.
Safety constraints must sit outside the agent, in hard limits and the safety system. Where the model runs is discussed in edge AI vs cloud AI.
For a broader view of AI in control rooms, read AI in PLC, SCADA and DCS. Start with advisory mode before closing the loop.
Deep RL Primer PDF
Literature Review Video
Reinforcement Learning FAQ
Related Articles
- Artificial Intelligence and Machine Learning
- AI in PLC, SCADA and DCS
- PID Controller Tuning Guide
- Digital Twin in Industrial Automation
- DCS Control Strategies
External References
- Deep RL for Process Control Primer, Spielberg et al
- Autonomous Control AI Trial at JSR, Yokogawa
- Reinforcement Learning, Wikipedia
What We Learn Today
- Reinforcement learning trains an agent to maximise reward through trial and error.
- Yokogawa and JSR ran a plant autonomously for 35 days.
- Train in a simulator and keep safety limits outside the agent.
