Showing posts with label pfpf. Show all posts
Showing posts with label pfpf. Show all posts

30 September 2008

Lars's Paradox, or, Everything You Know Is Wrong

"Listen, there's nothing up with the audio quality. It's 2008, and that's how we make records... Of course, I've heard that there are a few people complaining. But I've been listening to it the last couple of days in my car, and it sounds fuckin' smokin'."
Look, people. The dude isn't fucking deaf. Rick Rubin is also, contrary to popular opinion, not deaf (he owns a rather nice hifi, in fact). Metallica as a band is not deaf. Vlado Meller is not deaf. Millions of music listeners are not deaf. And now quite a few people are coming out the woodwork and saying that Death Magnetic sounds just fine, thank you very much. They too are not deaf.

To suggest otherwise, or to suggest that something is inherently wrong with the way they are listening, is merely fallacious smearing, and honestly, unintelligent. Continuing to insist that music products like Death Magnetic are not of a sufficiently high quality without further proof - especially in the face of #1 sales - is only going to continue the abject apathy that the rest of the music world seems to treat this whole issue with.

Certainly, Rick Rubin knows exactly what he's doing when he produces records like this, and he is quite certain in his belief that it is towards delivering a superior product, as his interview with Michael Fremer made abundantly clear:
Ultimately, if you listen on a car sound system or in the mainstream place where most people listen to music—cars, boomboxes sound systems you get at (chain stores), and if you “A/B” the less compressed version to the more compressed version, you pick the compressed version... Even in a good car stereo. We do shoot-outs all the time. I master with as many as five different mastering engineers mastering the same album and then we “A/B” them and it’s interesting, Vlado wins nine out of ten times, and he claims it’s not him. He’s got technology in that room that’s a 2 million dollar mastering suite that other people don’t have. All I’ll tell you is that my whole job in life is to A/B things, that’s all I do, and for some reason, I don’t know that what he’s doing is necessarily the best, but I haven’t heard anything to beat it and we try.
That the album distorts needlessly is established beyond a reasonable doubt, thanks to mastering engineer Ted Jensen's comments, and comparisons with the vinyl. I haven't bought the album, but I have listened to the free clips from Metallica's web site, and the YouTube GH3 rips, enough to know that I'd prefer the GH3 versions.

But let's have some perspective here. The truth of the matter is that this is a serious counterexample to the entire narrative of the "loudness war": that, despite diverse objective and subjective evidence that modern hypercompressed mastering styles degrade sound quality and music appreciation, the vast majority of music listeners, at all experience levels, at least continue to buy such purportedly terrible masterings, and may even prefer them to less compressed styles. I am going to call this Lars's Paradox, since Lars Ulrich, belligerent bastard that he is, has managed to wade neck-deep into the middle of this like he always tends to do. But whether due to a similar level of belligerence, or devil's-advocacy, or whatnot, I'm actually going to take his side here for a minute.

I believe any fight against hypercompressed mastering in the "loudness war" will founder until this paradox is resolved. More concretely, and extending to other issues, I am claiming the following:
  • Claims of the hypercompressed style resulting in reduced musical enjoyment are completely unproven except on personal, anecdotal, and therefore meaningless, grounds. Real studies need to be done, in real listening environments, to show that the application of hypercompression is a detriment to popular music and the popular music industry.
  • Objective evidence is inaccurate in arguments regarding mastering. Objective evidence cannot prove statements about enjoyment. Such analyses must be more explicit in their relationship between the music, the dynamic range, and the dissonant distortions if they are to be ultimately taken seriously. Waveform plots, ReplayGain, RMS, and pfpf are all highly deficient in one way or another here.
  • (Lars's Paradox) Evidence suggests that the hypercompressed style is preferred by at a large amount, and probably most, of the popular music listening population. Both audio professions and untrained listeners are making this preference. For the uncompressed styles to be taken more seriously, it must be shown concretely that this preference is based on faulty measurements, or is otherwise false in meaningful and important ways.
As long as these points stand, the argument against hypercompression will remain fundamentally flawed, and popular music will continue to be released in the hypercompressed style. Regardless of how many petitions get signed. Marginal releases like on vinyl and high-res formats obviously don't follow this logic as much, nor does classical and experimental music, etc. By and large, those are not popular genres or (yet) popular formats, and this discussion revolves largely around popular music. But there's simply no hope for popular CD/iTunes releases to follow any different mastering style as long as these issues exist with this whole argument.

(I have my own ideas, revolving mostly around psychoacoustics, for resolving the paradox, but they are as yet unfinished.)

Update, October 1:
Debate on the new JusticeForAudio.org forums.

Update, October 2.

27 June 2008

Some notes on pfpf

It's been quite a while since the first update, and I didn't mean this to be a one shot deal, so I might as will give a status update on pfpf. I haven't had the opportunity to work on a new version, but plenty of comments have been made so far:

On Missing Files. First, my hosting provider, storing my screenshots and files, disappeared without any contact information. I just got a new one set up, so all the links work again.

On Download Sizes. Lots of people don't like the 90MB LabVIEW Runtime download, or that registration is required on NI's website to download. I can't get rid of the runtime download entirely, but there is a smaller (28MB) runtime that may work for you - download here.

On Magnitudes. Several people commented that the dynamic range figures seemed too low. A well-mastered pop track may only show up as having only 3-4db of dynamic range on short/medium time scales, and virtually no range on long time scales. Extremely dynamic symphonic works may only have 16db of dynamic range, where by most "normal" evaluations, there should be more like 40-60db. While the choice of scaling has little effect on comparison of pfpf-derived numbers, it has a great effect on their overall interpretation in relation to other decibel figures

Much of this stems from the choice of percentiles used in the variance measurement (from the 50th to the 98th). If this range were to be doubled, the numbers would probably fall more into line with what people normally expect. This could be accomplished by either doubling the 50-97.7 figure.

On Dynamic Range Manipulation. Michael Jamsmith and I independently came up with the idea of running pfpf's histogram plot in reverse to make a Photoshop-like levels adjustment for loudness on a music track. In other words, a reversible 2-pass dynamic range compressor. Certainly something to work on in my copious free time :)

On Resolutions. One persom didn't like that a greater than 1024x768 resolution is required. I'll see what I can do to make the resolution requirements nicer, but I can't promise much. I might just require that people use a 1680x1050 or higher display.

Bob Katz's comments.

An Alternative Proposal - the Sparklemeter. Chromatix on HA has recently proposed measuring dynamic range by comparing the ratio between a PPM and a VU meter (suitably modified), preparing a histogram of the result values, and computing a figure of merit based on the mean and variance of that histogram. The resulting functionality is similar to the medium- and short-timescale figures that pfpf outputs, but the use of exponential-decay meters is new, as is using direct percentile values from the histogram rather than ranges between percentiles. Using VU/PPM meters gives the result great intuitive meaning for audio engineers, although I fear the required modifications may compromise that. Watch that thread.

13 January 2008

pfpf: An Experimental Estimator of Dynamic Range in Music

The dynamic range of a selection of music is dependent on both estimating the time-varying loudness of the music and the timescale used for loudness evaluation. I propose a numerical method of estimating dynamic range that satisfies those dependencies using a modified ITU-R 1770 loudness filter and three moving windows to estimate loudness across three different timescales. The goal is to more accurately measure and compare dynamic range between different music genres and different masterings and processing techniques for the same music.
Summary of algorithm:
  1. Apply ITU-R 1770 filters to convert amplitude to instantaneous loudness.
  2. Estimate loudness across three different timescales by computing 10ms ("short term"), 200ms ("medium term") and 3000ms ("long term") windowed RMS power.
  3. Decouple timescales by scaling 10ms loudness by 200ms loudness, and 200ms loudness by 3000ms loudness.
  4. Threshold loudness at each timescale to remove silence (optional)
  5. Compute histogram for each loudness estimate
  6. Dynamic range = range between 50th and 97.7th percentile, for each timescale

Using the pfpf application

The algorithm is prototyped in a LabVIEW application built for Windows, downloadable here. Unzip it into a new directory. You also need to download the LabVIEW 8.2.1 runtime.

Basic instructions: Run the program and open the folder icon on top to select a WAV file. Press the "Run analysis" button. The file is scanned for instantaneous loudness (indicated by the progress bar) and then a histogram operation is performed to calculate the dynamic range. The output is displayed at the bottom. Additionally, other tabs display plots of instantaneous loudness and histograms.

Interpreting the results:
  1. Long-term dynamic range - loudness changes across multiple seconds, or across multiple measures of a piece of music. Wide swings in orchestration and sustained loud/quiet passages increase this number. Dynamic range compression, in any form, decreases this number. Typical values range from 16db for extremely dynamic orchestral and experimental music to 1-2db for pop/rock singles.
  2. Medium-term dynamic range - loudness changes across hundreds of milliseconds, or single notes. Aggressive dynamic range compression can reduce this.
  3. Short-tern dynamic range - loudness changes across single milliseconds. The use of extremely percussive instruments can increase this. Extremely aggressive dynamic range compression, especially limiting, can decrease it.
  4. ITU-R 1770 loudness - estimate of loudness as per the ITU-R 1770 recommendation.

Example Results: Long/medium/short dynamic range for various tracks:
Musical piece
Long
Medium
Short
Autechre - "Sublimit"
4.0
4.0
10.6
Autechre - "Dial"
1.0
2.8
12.6
Shellac - "Genuine Lullabelle" (long term thresh=-50db)
14.3
7.4
6.9
Merzbow - "I Lead You Towards Glorious Times"
0.64
0.34
0.65
John Mayer - "Waiting On The World To Change"
2.7
3.2
8.9
Battles - "Tonto"
3.2
2.7
5.2
Soundgarden - "Black Hole Sun"
2.5
2.4
4.4
Autechre & The Hafler Trio - "æo³"
14.9
4.0
6.7
Harnoncourt, Beethoven Sym. 9, Chamber Orchestra of Europe
(Harnoncourt)
13.5
4.0
4.4

Screenshots

Configuration and output tab:




Loudness plot tab:



Histogram tab:


Advanced configuration

These options affect the computation of the dynamic range; when they are modified, the results should always include the new configuration. The "Output" string was created for this purpose.

  • Thresholds: If the instantaneous loudness drops under the threshold associated for that time scale, that timescale loudness (and the loudness for any shorter timescale) is clamped to NaN, and ignored in future dynamic range calculations. This is to prevent silence (assumed to be below the listening noise floor and is therefore inaudible) between music from affecting the results. Silence skews the histogram results so as to artificially compress dynamic range across all timescales. Its loudness also varies considerably between different formats (notably vinyl vs CD) and masking it aids in making an accurate comparison of formats.
  • Time scales: Controls the rms window size (in seconds) for each time scale.
  • Percentiles: By default, dynamic range is calculated as the loudness range between the 50% and 97.7% percentiles from histograms at each time scale of loudness. These percentile levels may be adjusted.

Application License

The pfpf application is free for non-commercial use. Do not redistribute it. Source code is available upon request (requires LabVIEW 8.2 or above and the Digital Filter Design toolkit).

Contact Info

Message me (Axon) on HydrogenAudio, or comment below.

Known Issues

  • It is important to take the results with a grain of salt. Transient loudness estimation is a topic of ongoing research, and no truly accurate method has yet to be agreed on. pfpf currently uses a moving-window modification to Leq(RLB), but in the future, a more elaborate loudness estimator, like HEIMDAL, might be used.
  • DC removal is applied at each short block (defaults to 0.01 seconds of signal) that is read, which are composed into the larger medium/long (0.2/3s) blocks. The end result is that the signal receives a 100hz highpass before analysis, removing all bass information. This is anticipated to not be a big deal because of the relatively small contribution that LFE provides to loudness models.
  • Histogram computation is not factored into the progress bar, so there is a noticeable pause between the completion of the progress bar and the display of results.
  • Beware of falling code. Parameters may not be well tested for failure cases or obviously incorrect inputs.

Document Revision History

13 January 2008: Initial revision.