The History of Broadcast Loudness
The Loudness Problem
During the first decade of 2000, a common source of audio related complaints from viewers to the FCC was loudness jumps between programming and commercials. “...loud commercials have been a leading source of complaints to the Commission since the FCC Consumer Call Center began reporting the top consumer complaints in 2002. One common complaint is that a commercial is abruptly louder than the adjacent programming.” [1].
These loudness jumps were also believed to cause viewers to switch channels, damaging broadcasters’ revenue.
Loudness Measurement
Measuring the perceived loudness of audio was recognised as a difficult problem to solve. “The Commission has not regulated the “loudness” of commercials, primarily because of the difficulty of crafting effective rules “due to the subjective nature” of loudness. The Commission has incorporated by reference into its rules various industry standards on digital television, but these standards do not describe a consistent method for industry to measure and control audio loudness.” [2].
The ITU worked on defining an algorithm for measuring subjective loudness, which resulted in ITU-R BS.1770. Further work was then carried out on the practical application of this algorithm by the EBU and the ATSC, resulting in EBU R128 and ATSC A/85.
Loudness Control
These loudness measurements were then adopted into broadcaster QC checks and eventually became legal requirements in many territories.
Dialog Intelligibility
The Intelligibility Problem
Fast forward to today, and now the largest source of audio related complaints to broadcasters and streamers is dialog intelligibility.
“Dialogue intelligibility remains the top complaint in broadcast audio — and it’s getting worse. A recent Xperi survey reports that 84% of consumers have difficulty understanding dialogue in TV shows and films.” [3]. “The issue of poor speech intelligibility has ranked number one on the ‘negative hit-list’ of viewer complaints for quite some time.” [4].
Also, some tests have shown that poor dialog intelligibility negatively affects the “stickiness” of a show (the likelihood that a viewer will watch to the end), and therefore, ultimately, the revenue of the broadcaster/streamer.
Intelligibility Measurement
There is not currently one universally preferred standard for measuring dialog intelligibility. In recent years an algorithm developed by Fraunhofer IDMT is gaining acceptance, called the “Listening Effort Meter” (LE Meter). [5]
The Fraunhofer LE meter uses a Deep Neural Network to identify individual phonemes in an audio clip, with the confidence of the identification leading to a measure of Listening Effort (the more confident, the less effort, and the more intelligible) [6]. This method has proven to be largely language agnostic, and has been validated against human subject listening tests.
Several software and hardware manufacturers have, or are working on, including the LE meter in commercial products (including Steinberg, NUGEN Audio, Telos and RTW).
Intelligibility Control
But the question remains whether some form of dialog intelligibility measure should become a QC check for broadcasters / streamers, or should it become a recommended measure to aid program makers?
In either case there are questions about what the most useful presentation of the raw LE meter readings should be. The LE meter produces a reading many times per second (every 30ms), and these values can vary a lot over a short time frame, making them hard to read. So, as with loudness measurements, some sort of windowing, smoothing and / or ballistics is likely to be required.
Also, for a QC check, ideally there would be a single “yes / no” flag for each program, answering the question “is the dialog intelligibility ok in this program” (or a single number to represent “how good is the dialog intelligibility in this program”).
It is not obvious how this single number for a program should be derived from the raw data. Various statistical measures could be used, but there are some situations when they may not give a full enough picture. One could use the mean or median average of the readings, or the upper or lower quartile. These statistical measures could be taken on the raw LE meter data, or the smoothed data (which may give different results).
However, it is not hard to imagine situations where these might not give the information we really need.
For example, consider a program with excellent dialog intelligibility throughout, except for one 5 minute scene with plot-critical dialog in a noisy environment, which was very hard to understand. In this case, an “average” intelligibility score may well read as being fine, but it might be considered that this should have been a QC fail as viewers may get turned off by the one hard-to-understand scene.
Finding solutions for this sort of problem remains a matter for future research.
Created for audio engineers: Dialog Check
DialogCheck implements technology developed by Fraunhofer IDMT's Oldenburg Branch for Hearing, Speech and Audio Technology HSA, which uses automatic speech recognition and psychoacoustic modeling. The LE-Meter has been validated through systematic listening tests, and its metrics show a high correlation with subjective assessments of listening effort.
Our Dialog Intellegibility Meter:
- Supports up to 9.1.6 channels
- Realtime bar meter
- Timecode-locked history graph with session offset
- Realtime momentary clarity readout
- Median clarity readout
- User-definable Upper and
- Lower Percentile readouts
Want your mix to be heard?
Add Dialog Check to your QC list now:
A paper from the Proceedings of the 2026 NAB BEIT Conference
Written by Dr Paul Tapper, CEO at NUGEN Audio — 20+ years experience in audio-software
| References | |
|---|---|
| [1] | Federal Communications Commission, “FCC 11-84: Implementation of the Commercial Advertisement Loudness Mitigation (CALM) Act”, pp. 2-3 |
| [2] | Federal Communications Commission, “FCC 11-84: Implementation of the Commercial Advertisement Loudness Mitigation (CALM) Act”, p. 3 |
| [3] | Curry, W., Shirley, B., Laverty, T., “Re-Thinking Dialogue Intelligibility: Why High SNR Still Fails Viewers”, https://events.theiet.org/events/re-thinking-dialogue-intelligibility-why-high-snr-still-fails-viewers/ |
| [4] | Baumgartner, H., van Everdingen, R., Schreiner, B., Kahsnitz, M., Krämer, U., “EBU Technical Review: Speech Intelligibility in TV”, p. 4 |
| [5] | Colmer, C., Rennies-Hochmuth, J., “Making dialogs accessible for everyone”, Press Release / Oldenburg / May 27, 2025 |
| [6] | Thornton, M., “Fraunhofer Intelligibility Meter Used In Nuendo 11”, Production Expert, Dec 14, 2020, https://www.production-expert.com/production-expert-1/fraunhofer-intelligibility-meter-used-in-nuendo-11 |
More like this
Share this