Table of Contents
ToggleA plant historian can collect millions of values every hour, yet most of them only repeat what a straight line between two points already says. Good compression keeps the shape of every trend while storing a small fraction of the raw samples.
Raw scan data grows into terabytes quickly, while most samples add no new information. Learn how deadband and swinging door methods decide what to store, how to set them from instrument accuracy and how much disk they save.

What Is Historian Data Compression?
Historian data compression is the set of rules a process historian uses to decide which incoming values to archive and which to discard, while still allowing every trend to be rebuilt within a stated error. It is a core function of any process historian, from PI and Canary to the historians built into SCADA and DCS packages.
Unlike zip files, this is lossy compression, because discarded samples are gone for good. The historian keeps only the points needed to redraw the signal by straight line interpolation, as described in DCS historian data storage.

Weatherford explains in its CygNet documentation that the method compares each new value with an aperture window built from error bands around the last held value. When a new slope falls outside that window, the held value is archived and the window starts again.
Why Plants Need Historian Data Compression
A 5000 tag plant scanned every second produces more than 150 billion values a year. Storing all of them slows queries, bloats backups and, as shown in SCADA database growing too fast, eventually fills the server.
Historian data compression also improves speed for analytics. AVEVA engineers reported at a PI user conference that even 11 percent compression gave a 20 percent shorter backfill time for Analysis Service calculations.
The same AVEVA presentation describes an offshore oil and gas pilot that cut archived events by about 94 percent. A downstream refinery saw compression between 11 and 95 percent, with half its tested tags beating a 3 to 1 target.
Deadband Versus Swinging Door Compression
Stores a value only when it differs from the last stored value by more than a fixed band.
Keeps a value only when a straight line from the last archived point can no longer pass within the band of all points since.
A first deadband filter at the interface or source that drops repeated values before they travel.
A plain deadband works well when a signal sits still, but on a steady ramp it stores a point every time the band is crossed. Swinging door historian data compression stores only the two ends of that ramp, because the line between them already describes every point in between.
The deadband idea also appears in telemetry, where field devices send only changed values, as in report by exception in DNP3. That saves bandwidth, while the historian applies its own compression later.
The swinging door trending method was described by Edgar Bristol in an ISA conference paper in 1990. It is still the default analog compression in many historians today because it needs only a few calculations per value.
How the Swinging Door Algorithm Works
Imagine two doors hinged at plus and minus CompDev around the last archived value. Every new sample swings the doors closer together, and as long as some straight line can still pass inside both doors, no value is stored.
When a sample forces the doors past parallel, the algorithm archives the last held sample, not the new one. That point becomes the new hinge, so the archived trend is always a chain of straight segments that stays within CompDev of every original value.
Plot one week of raw data and the compressed archive on the same trend before you sign off any setting. If peaks are clipped or noise is stored, adjust CompDev rather than the scan rate.
Exception and Compression in the PI System
In the PI System, values pass two filters. The interface or connector applies the exception test first, then the Data Archive applies swinging door compression, a two stage design explained in PI and AVEVA historian architecture.
5 Smart Historian Data Compression Settings
The AVEVA presentation shows defaults of 0.1 percent for ExcDevPercent and 0.2 percent for CompDevPercent, with ExcMax at 600 s and CompMax at 28800 s. Its main rule is simple: set ExcDev equal to half of CompDev, and keep CompDev at or below the calibration tolerance of the instrument.
Percent settings apply to the tag span, so a wrong span makes them meaningless. Check ranges against the transmitter configuration, using the method in sensor scaling in instrumentation.
ExcDev = CompDev ÷ 2
Example, J type thermocouple loop:
Tolerance about plus or minus 4 °F, so choose CompDev = 2 °F
ExcDev = 2 ÷ 2 = 1 °F
Any stored error stays below the sensor tolerance
Industrial Insight notes that a J type thermocouple measuring up to 1300 °F has a typical tolerance of around plus or minus 4 °F. Storing changes of 0.1 °F from such a sensor records noise, not process behaviour.
Historian Data Compression Accuracy Trade Offs
Too much compression hides real events, while too little wastes storage and slows queries. Industrial Insight gives an example where only 2 of 586 samples of a main steam flow were being stored, which made the trend nearly useless.
The same firm reports that well tuned tags usually keep 90 to 95 percent fidelity while retaining 50 to 75 percent of the original samples. Fast loops need care, since filtering in the field, covered in transmitter damping and response time, already smooths the data before it arrives.
- Far less disk space and smaller backups.
- Faster trends, reports and analytics queries.
- Lower network load from interfaces.
- Noise below instrument accuracy is not stored.
- Discarded samples can never be recovered.
- Over compression clips peaks and short spikes.
- Wrong span makes percent settings meaningless.
- Statistics such as standard deviation change with compression.
Historian Storage Calculator
Use this calculator to estimate yearly storage with and without historian data compression. Bytes per value is your own assumption for the product you use, so treat the result as a planning figure.
Raw storage = Values × Bytes per value
Stored = Raw × Percent kept ÷ 100
Example:
5000 tags, 1 s scan, 10 bytes per value, 10 % kept
Raw = 5000 × 31,536,000 × 10 = 1576.8 GB per year
Stored = 1576.8 × 0.10 = 157.7 GB
Saved = 1419.1 GB per year
Second Worked Example: Pressure Transmitter Tag
A pressure transmitter is ranged 0 to 40 bar, and the complete loop, including the analog card, is good to about 0.25 percent of span. That tolerance is 0.1 bar, so a sensible CompDev is 0.1 bar and ExcDev is 0.05 bar.
The default 0.2 percent CompDev would give 0.08 bar, which is close, so the default is acceptable here. A 12 bit input card has a resolution near 0.025 percent, as shown in analog signal resolution in PLC, so settings far below that only store quantisation steps.
Tuning Historian Data Compression by Tag Type
| Tag Type | Suggested CompDev | Notes |
|---|---|---|
| Temperature | About half of sensor tolerance | Slow signals compress very well |
| Pressure and level | 0.1 to 0.25 % of span | Match loop accuracy |
| Flow | 0.25 to 0.5 % of span | Noisy, avoid storing turbulence |
| Analyser | Near analyser repeatability | Few updates, keep CompMax |
| Digital and status | Zero | Store every change of state |
| Totaliser | Zero or small | Steps matter for reports |
Digital tags, alarm states and sequence data should never be compressed by deviation, because every change matters. Very fast event data belongs in a dedicated system, as explained in sequence of events recording.
Weatherford notes that CygNet always archives a value when its status bits change, even if the number itself sits inside the band. Quality changes such as bad input therefore never vanish from the record.
Create tag templates by instrument type with preset CompDev and ExcDev values. New tags then start with sensible historian data compression instead of system defaults copied blindly.
Historian Data Compression Review Checklist
- Tag span in the historian matches the transmitter range.
- CompDev is at or below instrument tolerance.
- ExcDev is about half of CompDev.
- ExcMax and CompMax are set so flat tags still report.
- Digital and status tags store every change.
- Raw and compressed trends were compared for key tags.
- Clocks on servers and interfaces are synchronised.
Time accuracy matters as much as value accuracy, because a shifted timestamp breaks interpolation. Keep interfaces aligned using SCADA time synchronization with NTP and PTP, and buffer data during outages with store and forward.
Compressed archives also feed models and analytics, including soft sensors, so test a model on compressed data before deployment. Tag naming and templates are planned best in SCADA tag database design.
AVEVA Exception and Compression Presentation
Getting Compression Right Video
Historian Data Compression FAQ
It is the rule set a historian uses to decide which incoming values to archive and which to discard. The goal is to rebuild every trend within a known error band.
Most historians use a deadband for the first filter and swinging door logic for the archive. The discarded samples cannot be recovered, so settings must be chosen with care.
A deadband stores a value whenever it moves more than a fixed amount from the last stored value. On a steady ramp it therefore stores many evenly spaced points.
Swinging door checks whether one straight line still fits all points within the band. On the same ramp it stores only the two ends, which saves far more space.
ExcDev is the exception deadband applied at the interface before data reaches the server. CompDev is the swinging door band applied when the server archives the data for long term storage.
AVEVA recommends setting ExcDev to about half of CompDev. That way the exception stage never removes information that historian data compression at the archive would have kept.
Start from the calibration tolerance of the complete loop in engineering units. Set CompDev at or below that value, never far below the resolution of the input card.
For a 0 to 40 bar transmitter accurate to 0.25 percent, use about 0.1 bar. Then compare raw and compressed trends for a week before you sign off.
They set the longest time allowed between reported or archived values. Common PI defaults are 600 seconds for ExcMax and 28800 seconds, or eight hours, for CompMax.
Without them, a perfectly flat signal might not be stored for many days at a time. Users would then be unsure whether the tag was steady or the interface had stopped.
Yes, over compression clips peaks, hides short upsets and leaves trends with too few points. Industrial Insight found one main steam flow tag storing only 2 of 586 raw samples.
Signs include flat lines during known upsets and missing alarm peaks. Reduce CompDev for those tags and confirm that the configured tag span is correct.
Savings from historian data compression depend on signal behaviour and settings, so there is no single figure. AVEVA reported about 94 percent fewer archived events in an offshore oil and gas pilot.
A refinery in the same study saw compression between 11 and 95 percent by tag. Even small savings speeded up analysis backfills by about 20 percent.
Related Articles
- What Is a Process Historian
- DCS Historian Explained: Data Storage
- SCADA Historian Integration with PI and AVEVA
- Store and Forward in SCADA Historians
- SCADA Database Growing Too Fast
External References
- Exception, Compression and their Impacts on PI System Performance, AVEVA
- Why Compression Is So Important for the PI System, Industrial Insight
- Operational Historian, Wikipedia
What We Learn Today
- Historian data compression keeps only the points needed to redraw a trend within a set error, using deadband filters and swinging door archiving.
- In PI, ExcDev acts at the interface and CompDev at the archive, and AVEVA advises ExcDev at half of CompDev within instrument tolerance.
- Over compression hides real events while under compression stores noise, so compare raw and compressed trends and tune each tag type separately.
