Header

Objective Reviews & Commentary - An Engineer's Perspective
Showing posts with label RMAA. Show all posts
Showing posts with label RMAA. Show all posts

October 12, 2011

Turtle Beach Micro II

tb micro ii benchINTRO: This is the first in a series of inexpensive portable USB DAC reviews I’ll be publishing in the next week or two. The idea is to test the DACs with a high impedance load (such as the O2 Headphone Amp) and a typical headphone load. The DACs are all small, easily portable, and USB powered. The $25 Turtle Beach Audio Advantage Micro II is just such a DAC. According to Turtle Beach, it’s supposed to provide “higher quality sound” compared to internal PC audio. I also have updated how I test and present DAC results.

TURTLE BEACH MICRO II: The Micro II is a small “dongle” with an attached 2 inch “pigtail” USB cable. It has a single 3.5mm jack which serves as line out, headphone out, and an optical digital out (using a supplied 3.5mm-to-Toslink adapter). There’s no volume control, other controls, or inputs of any kind--just a blue LED.

tb micro ii modesWINDOWS INSTALLATION: The Micro II installed smoothly in both XP and Windows 7 without needing any drivers. Windows reported it as a “USB Sound Device”.  The only sample rates and bit depths available are 16/44 and 16/48 as shown to the right in Windows 7.

SUBJECTIVE SOUND QUALITY: There was moderate hiss with my Ultimate Ears IEMs but the Micro II was fairly quiet with less sensitive headphones. The sound quality, however, was seriously odd. Playing familiar well recorded audiophile tracks the Micro II made them sound shrill, glaring, and harsh regardless of what headphones I used. I was really curious to measure the Micro II and find out why it sounded so obviously bad.

POOR DRIVER DESIGN (updated): I first checked the Micro II’s frequency response and it was reasonably flat out to 15 Khz so the poor sound was still a mystery. Then I checked the 1 Khz THD and while it wasn’t great it also wasn’t bad enough to explain the poor sound quality. When I dropped the level to see if the distortion would drop, I found the problem. The Micro II was displaying a horrible linearity problem. Dropping the input from 0 dBFS to –20 dBFS should drop the output by 20 dB as well. But it only drops 8 dB! That’s a massive 12 dB error. The net effect is the Micro II was heavily compressing music—making softer sounds much louder than they should be.

BAD SOUND EXPLAINED: It turns out, as described in the comments to this article, the C-Media CM102 integrated USB sound chip apparently used in the Micro II has a “feature” called Dynamic Range Control (DRC) that defaults to on. Confusingly, there’s an advanced option in the Windows 7 sound options for the Micro II simply labeled “Loudness”. And, worse, it’s enabled by default. You have to uncheck the box to stop the Micro II from heavily compressing anything you play through it.

DRC vs LOUDNESS: The Turtle Beach choice of calling the C-Media’s DRC option “Loudness” is very misleading. In audio, Loudness Compensation involves changing the frequency response at low listening levels to compensate for human hearing. It’s generally based on the Fletcher-Munson Equal Loudness Curves. In this case, however, it has nothing to do with changing the frequency response—only the overall dynamic range. I don’t know what Turtle Beach’s driver calls this feature as it’s an unsigned driver so I didn’t install it on my test bench PC that runs the dScope software. The whole idea is to not need proprietary drivers.

WHAT WERE THEY THINKING (updated)? The “loudness wars” are already out of control without any further help from Turtle Beach. A lot of pop music has a peak to average volume difference of only around 8 – 12 dB as the labels keep compressing music ever further in an effort to have it stand out as being louder. The last thing pop music needs is another approximately 12 da lot more compression but that’s exactly what you get, by default, with the Micro II. For anyone unaware of the option, or who knowingly leaves it on, the Micro II is likely to sound significantly worse than the internal audio of just about any computer its plugged into.

MEASUREMENT SUMMARY: The overall results, even with DRC disabled, are not terribly impressive. The high frequency distortion, in particular, is poor. Here are the results compared to the more expensive FiiO E7 (some tests were run slightly differently but I’ve tried to adjust for that in the E7 numbers):

Measurement TB Micro II Fiio E7
Frequency Response 20hz-15Khz 33 ohms +/- 1.8 dB Fair +/- 0.1 dB Excellent
THD 0 dBFS USB 10K 0.20% Fair 0.05% Good
THD 1 Khz 10K Ohms -3 dBFS 0.022% Good 0.03% Good
THD 1 Khz 33 Ohms -3 dBFS 0.12% Fair 0.03% Good
IMD CCIF USB 0.28% Poor 0.03% Good
IMD SMPTE  0.03% Fair 0.008% Excellent
Noise A-Weighted -93.8 dBu Fair -96.7 dBu Good
Max Output 33 Ohms Vrms/mW 1.26v 52mW Good 57 mW Good
Max Output 10K Ohms Vrms 1.34v Good 1.4v Good
Output Impedance 100hz 0.95 Ohms 0.13 Ohms Excellent
Jitter USB 16/44 Jtest Fair Very Good

BOTTOM LINE: In my opinion the “Loudness” feature enabled by default is an epic fail. Someone either got sloppy or they have very odd priorities for a “high quality” USB DAC. Putting that aside, the rest of the performance of the Micro II still isn’t very impressive. The next several reviews of low priced USB DACs will help put the Micro II’s performance in perspective.

 


TECH SECTION


FREQUENCY RESPONSE: The frequency response with a 10K load (such as a headphone amp) at 16/44 is acceptable but not great. There’s a fraction of a dB of variation below about 30 hz and it’s down –1 dB at about 15 Khz. The steep roll off above 12 Khz is typical of a cheap DAC running at 44 Khz and is due to cost savings in the digital and analog filters. The slight peak around 8 Khz is also disturbing as it indicates either poor DAC filtering and/or potential instability in the headphone amp. Into 33 ohms you can see a slight drop due to the output impedance and a low frequency roll off of about –1.7 dB at 20 hz. That’s borderline audible. Into 16 ohms it would be even worse and more likely to be audible. This indicates less than ideal capacitor coupling in the output:

TB Micro II Frequency Response 100K (blue) & 33 Ohms - 3 dBFS ~400 mV 16-44

 

THD+N vs OUTPUT: This test starts at 10 mV and the rise at below 250 mV is more likely due to more quantization error than noise. The unweighted noise should be below 0.004% but the distortion is more than ten times higher. The lower blue plot is into 10K and the distortion is around 0.025% which is under the worst case guideline of 0.05%. Into 33 ohms, however, it’s over 0.1%  above about 700 mV which could be audible under some circumstances. In both cases, the maximum output level is around 1.3 Vrms. This works out to 52 mW into 32 ohms, 104 mW into 16 ohms and only 6 mW into 300 ohms:

TB Micro II 1 Khz THD N vs Output 10K (blue) & 33 Ohms 16-44

 

THD+N 100 hz 0 dBFS & OUTPUT IMPEDANCE: I now run this test at 100 hz as that’s where output impedance is usually most critical due to the resonance frequency of many headphones. A low 100 hz output impedance keeps the frequency response accurate and provides electrical damping of the driver which can improve the quality of the bass performance. The test is run at 0 dBFS input to reveal any digital overload problems such as the NuForce uDAC-2 exhibits. The Micro II does not reach clipping even at full volume into 15 ohms. The distortion here is mostly in the DAC itself and remains similar into 100K even at lower volume settings. The resulting output voltages at 100K and 15 ohms are used to calculate the output impedance. The Micro II’s distortion is relatively poor at 0 dBFS. It hit 0.14% into 100K and 0.21% into 15 ohms. The output impedance was 0.95 ohms which is acceptably low and it’s slightly lower at 1 khz due to less impact by the output capacitors:

TB Micro II Max Output & Impedance 100 hz THD N 100K & 15 Ohms (blue) BW=22 Khz

 

THD+N vs FREQUENCY: Here’s the 1 Khz THD+Noise plotted from 20 hz to 20 Khz into 10K (yellow) and 33 ohms (blue) at 775 mV (0 dBu). The input is –3 dBFS to prevent any digital overload of the DAC. The increase in low frequency distortion into 33 Ohms is another sign of a capacitor coupled output. The rise from 0.05% to 0.15% is likely the output capacitor’s non-linearity. The drop above about 6 Khz is related to the bandwidth limit of 22 Khz as the harmonics move past the audible band. The sharp rise again above 10 Khz is due to very poor high frequency performance in the DAC (and/or filter) despite the fact the harmonics are cut off:

TB Micro II THD N vs Freq -3 dBFS 0 dBu 10K & 33 Ohms (blue) 16-44

 

SMPTE IMD: The result here is marginal but acceptable at this price. Ideally all distortion products should be below –80 dB but that’s not the case. The spread (or “mountain”) at the base of the 7 Khz signal is another bad sign. This test is run below the DAC’s digital limit and also well below the maximum output (at 0 dBu) into 33 ohms. It’s somewhat better into 10K but not a lot:

TB Micro II SMPTE IMD -3 dBFS ~0 dBu 33 Ohms 16-44

 

CCIF IMD BENCHMARK DAC1 PRE: This is a more challenging test, and again, the goal is to have everything except the 19  and 20 Khz signals below –80 dB. To show how it can look, here’s the result with the Benchmark DAC1 Pre:

Benchmark DAC1 IMD CCIF 33 Ohms ~400mV

CCIF IMD MICRO II 44 Khz 33 Ohms: By comparison, here’s the same test as above from the Micro II. There’s an entire “forest” of distortion products above –80 dB with the 1 khz difference signal at –52 dB which is very likely audible. The digital/analog filters in the Micro II are in real trouble here as might be the DAC itself. The two spikes around 15 Khz exceed –40 dB and may also be audible. This is admittedly a tough test for a cheap DAC running at 16/44 but this is still a much worse than average result made worse by the headphone amp struggling with a 33 ohm load:

TB Micro II CCIF IMD -7 dBFS ~0 dBu 33 Ohms 16-44

 

CCIF IMD MICRO II 44 Khz 10K: Removing the load results in the 1 Khz difference signal improving significantly from –52 dB to about –72 dB dropping the reading by a factor of ten. But note there are still a lot of spikes above –80 dB and even above –60 dB within the audio band. Worst of all, the spikes at 15 Khz are still crossing –40 dB:

TB Micro II CCIF IMD -7 dBFS ~0 dBu 10K 16-44

 

CCIF IMD MICRO II 48 Khz 10K: DACs will typically do better on this test running at 48 Khz but it depends on their filtering and design. In this case, things get significantly better but are still not great with the 10/11 Khz signals still around –40 dB and several other spikes still above –80 dB. If your operating system lets you run a DAC at 48 Khz, 99% of digital music will be re-sampled from 44 Khz up to 48 Khz by the operating system with mixed results. In this case, it’s hard to say which would yield the better result. In XP it’s not an option as the DAC is forced into 44 or 48 Khz depending on the source sampling rate:TB Micro II CCIF IMD -7 dBFS ~0 dBu 10K Ohms 16-48

 

NOISE & LINEARITY: I’ve changed this test slightly to use units of dBu rather than my previous dBr referenced to 400mV. 0 dBu is 775 mV. The Micro II’s –93.8 dBu A-Weighted noise referenced to 400 mV would be 88 dBr (it’s always a difference of 5.7 dB). That’s decent noise performance for a USB powered DAC but falls well short short of what’s required for reasonable silence with the most sensitive IEMs. The goal is –103 dBu. I’m also now showing the absolute (unweighted) noise in microvolts. The linearity was fairly good with an error of only 0.8 dB at –90 dBFS:

TB Micro II 1 Khz -90 dBFS Noise & Linearity Ref 0 dBu 10K

 

-20 dBFS LINEARITY WITH “DRC”: With the “Loudness” option enabled (which is on by default) here’s what happens with a –20 dBFS 1 Khz signal. It’s played back at –8 dB instead of –20 dB. The Micro II is raising the level by a whopping 12 dB as part of its “Dynamic Range Control” feature. This is why, by default, it sounds relatively awful:

TB Micro II 1 Khz -20 dBFS Vol=100% THD N BW=22Khz ref 1.34 Vrms

 

JITTER: Here’s the spectrum from the dScope’s J-Test for jitter. The side bands are average at –110 dB but the “spread” of the signal is relatively poor indicating significant low frequency jitter. The frequency accuracy (clock accuracy) is very good as shown by the frequency reading on the left:

TB Micro II Jitter 11025 hz J-Test 10K ref -3 dBu 16-44

 

RMAA RESULTS: Out of curiosity I tested the Micro II with RMAA with the DRC/Loudness option enabled (10K load). While the frequency response was very similar, the THD spectrum showed some significant differences although the 2nd and 3rd harmonics were similar. The IMD was even worse than the dScope measured and could have been a clipping/level problem. The noise measured a relatively poor –77.5 dB for which I have no explanation. It presumably limited the dynamic range to a similar value. Interestingly, there’s no solid indication of the DRC compression. RMAA missed the huge linearity problem. The results are shown below along with the sound hardware by itself (2nd column) in loopback. For more, see my RMAA article:

TB Micro II RMAA Summary (10K Load)

TB Micro II RMAA Freq Response (no load)

TB Micro II RMAA THD N (no load)

TECH COMMENTS: The good news is the output impedance is below 2 ohms, the noise/linearity is decent, and the frequency response and midrange distortion are semi-acceptable into a 10K load like a headphone amp. Into 32 ohm headphones, however, the distortion rises to unacceptable levels and the high frequency distortion into any load is even worse.

February 18, 2011

Testing Methods

INTRO (revised 4/15/11): Being an engineer by education and profession I might take testing methodology more serious than most. I think it’s important to understand how to do it right.  Even websites that have a relatively large audience (and budget) are publishing invalid, wrong, biased, or meaningless test data. They apparently either don't know any better, don't care to do it right, or they're trying to mislead people on purpose.

RMAA IS NOT AS GREAT AS PEOPLE THINK: One of the biggest "tools" you see results from is RMAA. That's likely because it's free and works with standard PC sound interfaces. But, as I document in my RMAA Article, it has lots of problems and limitations. It’s often impossible to make a valid comparison between RMAA results made by one person to those made by another. And even RMAA results made by the same person may not be trustworthy. One common reason for this is RMAA itself has no concept of absolute levels. So test A might be done at a higher volume setting stressing the device, while test B was done at a much more favorable level. There are just too many uncontrolled variables that can make a bigger difference in the RMAA numbers than the device being tested.

14+ THINGS RMAA DOESN’T TEST: There are many important, and common, audio measurements RMAA does not test for. These include maximum output, output impedance, output power vs THD, square wave tests, slew rate, CCIF IMD, and many more. See the RMAA Article for more.

A REAL WORLD EXAMPLE: A friend built an amp and was raving how it was so detailed he could finally hear differences in RCA interconnect cables. I was skeptical so we measured it and found out it was unstable and ringing (partly oscillating) to varying degrees depending on the input cable capacitance. So his cables did indeed sound different because they caused the amp to produce different amounts of ultrasonic garbage! We traced the problem to a poor ground scheme for the input stage and jacks.

The funny (sad?) part of the story is he seemed disappointed once the grounding was fixed and the amp stopped oscillating. He thought he was onto something great. He heard the severe instability as being "different" and hence "better". It's interesting (sad?) how misleading "designing by ear" can be with all the psychological biases that are typically present.

It's also worth noting he tested the above amp, in its original highly unstable form, with RMAA and it passed with flying colors. The ultrasonic oscillations were well above the cutoff of any soundcard's anti-aliasing filter and hence were invisible to RMAA. And apparently the audible side effects were also not picked up by RMAA. That's just one of many real-world examples why I take RMAA measurements with several big grains of salt.

BAD TESTING IS WORSE THAN NO TESTING: The great thing about the web is you can post something and share it with a very large number of people. If it's something that's obviously subjective such as a favorite restaurant, those reading your review know their tastes might be different that yours. But if something appears factual, and the author seems to know what they're talking about, that's very different. But what if it's really wrong or misleading? Because so few people are in a position to verify the results, and the few who can likely won't take the time, the bad data is very likely to go unchallenged. So it just hangs out there on the web to potentially mislead anyone who happens to find it. This happens more than you might think!

VIRAL TEST RESULTS: There are many examples where someone posts their test results for some piece of new gear and the numbers look great, they rave about it, and soon others run out and buy the same gear and post their own glowing subjective reviews, and before you know it the product has an almost cult-like following of fans on the web. But if you trace some of these back to their source, there's sometimes only a single set of sketchy results that helped start it all. And it can be a long time before anyone tries to verify the results—if ever. All the rest of the glowing reviews are often purely subjective--i.e. people's personal opinions. And those are biased from all the other positive comments and other influences.

ANOTHER REAL EXAMPLE (added 5/14): I tested the AMB Mini3 headphone amp partly because I was impressed by the fairly complete set of impressive measurements listed on the AMB website—many from RMAA. It’s safe to say many others were impressed as well--over the last 3 or 4 years many have decided to build (or buy) a Mini3. But, it turns out, in proper testing the Mini3 didn’t come close to meeting many of it’s measurements. The actual performance was much less impressive and even problematic in some areas.

PROPER TESTING ISN'T SIMPLE: Many think you can just hook up a few cables, run the RMAA calibration routine, click the "Go" button, and get good results. But it's far from that easy. And, ultimately, even used in the best way possible, RMAA still has some serious limitations and doesn't test some important things. So, at best, RMAA is only a partial window into the performance of a device. And making more accurate measurements, and all the ones RMAA can’t make, requires expensive instrumentation and lots of knowledge. Entire books have been written on the subject.

TEST CONDITIONS MAKE A HUGE DIFFERENCE: If I publish a Ford Mustang can go from 0-60 MPH in 5.5 seconds that sounds fairly impressive. But is that on perfectly flat ground or was it down hill? Was there a head or tail wind? How much extra weight was in the car? How accurately was it timed? Was it some guy in the passenger seat with a wristwatch or with professional timing equipment? Was the road wet with lots of wheelspin or dry with good traction? All of these things will significantly change a car's 0-60 time--sometimes dramatically.

MISLEADING INFORMATION: If I run a 0-60 test of a Mustang on dry pavement, and someone else does a 0-60 test on a fairly similar Chevy Camaro but does it in the rain, which one do you think will have the better result? Obviously the Mustang will get a lot more traction at the start and be the clear winner. But if you just saw the two numbers published on the web, with little explanation of how they were obtained, you might easily think the Mustang is a much faster car than the Camaro. This is exactly analogous to what happens when people test audio gear and post their results. While testing a car's acceleration in the rain is hopefully a bad idea to most people, many are not even aware they’re making similar mistakes during audio testing.

THE BEST TEST RESULTS CAN BE READILY COMPARED: What good are test results if you can't make valid comparisons to other results? That's why car magazines try to test cars under as controlled of conditions as possible. They correct for weight, wind, and even things like temperature which affects an engine's horsepower. So when they test a Mustang in January in Detroit and a Camaro in August in California, you can safely compare the results to each other. But most of the audio results being published on the web cannot be compared in similar ways because they're often measured under different or unknown conditions.

THE BEST TEST RESULTS ARE VERIFIED BY OTHERS: In the scientific and medical communities results that can't be verified are completely dismissed as invalid. But in audio many tend to take them as fact. When you conduct testing in a controlled way, it's much easier for others to verify your results. But if they don't know under what conditions the tests were made--what signal levels, loads, settings, with what equipment, etc.--they're almost impossible to verify. And without being able to verify the results, there's no way to know if they're reasonably accurate.

THE FINEST TESTING ADVERTISING WON'T BUY: There are certainly organizations out there with the equipment and knowledge to run proper tests. But, sadly, most of the consumer oriented magazines and sites tend to leave out, or gloss over, anything very negative for fear of losing their advertisers (which very often make the very gear they're testing). A classic example is even some relatively expensive (i.e. $1000) A/V receivers come nowhere close to their advertised power ratings under real world conditions. But tests on these receivers are often done in such a way to avoid revealing just how bad they really are). Why? Most likely because those same manufactures advertise with the same organization publishing the review.


TECH SECTION (for test geeks revised 4/15):

AUDIO ANALYZERS: The best solution is a dedicated instrument designed for audio testing. The two companies with the most market share are Audio Precision and Prism Sound. Audio Precision ( ap.com ) is widely considered the reference standard. Their entry-level product, the APx525, starts around $6,000 for the most basic analog-only 2 channel model with reduced specifications. And the pricing approaches $20,000 with the more popular options for analog and digital measurement. Their better models go up in price from there. Prism Sound ( prismsound.com ) offers the dScope Series III as competition to Audio Precision's analyzers and they're more reasonably priced with similar specifications. When fully configured for analog and digital measurement, the dScope is around $10,000. There are 2 analog-only versions at lower prices similar to the entry-level APx525.

OTHER ANALYZERS: There are a few other choices out there but they're typically either even more expensive and/or have relatively limited capabilities. Some examples are the Rhode & Schwarz UPV/UPL and Agilent U8903A. In my experience, these products are rarely used for consumer gear audio measurements. It makes the most sense to fully leverage the power of today's PCs for the heavy lifting (FFT, etc.) like the dScope and APx5xx products do. Stanford Research took an interesting approach with their relatively new SR1 which I’ve played with at trade shows. They essentially built a PC into a very large bench instrument along with an enhanced “sound card” and made a self contained analyzer for under $9,000. But, to me, the dScope and APx5xx approaches make more sense.  I don’t know of anyone using an SR1 and I’m not sure where it offers an advantage unless perhaps you don’t have a PC.

DISCONTINUED ANALYZERS (added 4/15): There are older discontinued products on the used market but they’re mostly relatively limited in what they can measure and/or have other serious issues. Beware of the older Audio Precision products (i.e. System One, System Two, etc.) as some require a proprietary PC card ISA (think IBM XT circa 1984) interface and are difficult, expensive, or impossible to use with modern PC’s. And you can also find things like the HP 8903A and 8903B distortion analyzers, but like the early AP instruments, they’re big, heavy, clunky and make better boat anchors than audio analyzers. The AP ATS-2 was very limited in performance. And older self contained devices like the Audio Precision ATS-1 usually can’t talk to a PC which means you can’t even do a screen capture of a result.

PRISM dSCOPE: I use the Prism Sound dScope Series III for most of my measurements. It can run pre-defined or user-defined set ups and scripts in a single mouse click. This makes running identical tests on different gear and a whole series of tests much easier and more consistent. It eliminates most sources of human error. The dScope overcomes nearly all the limitations of RMAA as well as providing much higher accuracy, absolute level measurement and much more. It has isolated balanced inputs and outputs and doesn’t suffer from ground loops and problems. You can set the signal generator outputs (analog or digital) to 0.06 dB accuracy and measure/analyze input signals to 0.06 dB accuracy from literally less than 0.000001 volts up to 159 volts RMS with no external dividers or hardware required. The time base is accurate to a few parts per million and can be used for measuring the quality of a digital signal in the digital domain (i.e. true actual jitter, deviation, eye patterns, etc.). Here are the full specs for anyone interested.

AUDIO PRECISION & dSCOPE COMPARISON (added 4/15): Most of the difference between these competing products is in the analog performance and architecture. The best-in-the-world Audio Precision SYS-2722 manages a few dB less residual noise and distortion than the dScope but at a price that will buy you a brand new German luxury car. The dScope essentially matches or exceeds the performance of the less expensive APx5xx line and their older System One and System Two analyzers. Here are some of the more notable specifics:

  • Analog Noise Floor – The dScope is rated at –115 dBu worst case residual noise and typically measures closer to  –116 dBu (about 1.2 microvolts). This is within 1 dB of the newest Audio Precision APx5xx series analyzers, or the older AP analyzers like the System One and System Two units. To put this level of noise in perspective, the self generated Johnson Noise of a single 4.7K resistor is approximately –115 dBu. Put another way, a single 10K resistor can produce more noise than the entire analyzer section of the dScope! The current Audio Precision flagship 2700 Series has residual noise across the audio band of –117.8 dBu This is about 1.0 uV or the noise you get from a single 2.7K resistor. When you consider both products have balanced inputs (which are inherently noisier than the unbalanced variety), an extreme input range from microvolts to around 200 volts, and analog circuitry in close proximity to lots of noisy digital hardware, this is very impressive performance from both companies.
  • Signal Generator Distortion – Over the audio band, the dScope’s analog generator is rated at 0.0007% worst case distortion. The APx5xx is about the same. I have no way to measure just the generator, but even the dScope’s combined THD+N of the generator and analyzer at 1 Khz is typically below 0.0006% at 1 Khz. The costly SYS-27xx is rated at 0.0003% which is significantly better but we’re talking about levels of distortion that are nearly always masked by other factors and any sane person would consider long since inaudible and well past the point of diminishing returns.
  • Measurement Bandwidth – The dScope, APx5xx, and older AP analyzers have a maximum sampling rate of 192 Khz limiting the measurement bandwidth to 96 Khz (or > 90 Khz as AP likes to put it). The signal generator of the dScope is limited to 91 Khz. The much more costly AP SYS-2722 generates and measures out to 200 Khz  The good news is most anything that needs to be done beyond 90 Khz can be done using other equipment such as my 14 bit 100 Mhz digital scope, 14 bit Tektronix 25 Mhz waveform generator, and Agilent DMM. Put another way, there’s little justification in audio for having –120 dB noise measurements or 0.001% THD measurements past 90 Khz.
  • Real Time Analysis – The dScope and flagship AP analyzers have the ability to measure a number of parameters in real time and that’s a big deal for R&D work. Audio Precision’s more reasonably priced products (like the APx515, APx525, APx585, and ATS-2) behave differently—more like fancy sound cards--and they lack this critical capability. A typical example: Most class AB output stages using bipolar transistors have a fairly narrow range of bias current for the lowest distortion. The dScope, using its Continuous Time (CT) detector, can internally measure the THD and display the result in real time. So you simply adjust the bias pot for the lowest distortion. An APx525, by comparison, has to capture an entire sample buffer, send it over the USB link, run a FFT, and finally display the result. So it’s harder to find the “null” point and optimal setting. The dScope can also let you monitor, in real time, residual signals. You can even listen to what the distortion (or signal) sounds like in real time if you want. And the CT detector can be used for noise, IMD, levels, channel balance, crosstalk and more. So let’s say you want to orient the power transformer for the least amount of hum in a product. You can just move it around and watch the 60hz/120hz readings in real time on the dScope. You can also apply a huge variety of filters to the live result. You have to buy the flagship AP series to get this capability but it’s standard in the dScope.

AVERAGING IS A DOUBLE EDGE SWORD (added 4/15): Some companies, especially those selling soundcard based audio software, publish specs that (often in the fine print) include averaging. The residual performance limits of most any analyzer are substantially improved by using averaging as it helps remove inherent noise in the analyzer.  For example, the –115 dBu absolute noise floor of the dScope improves to –123 dBu with the right averaging. The noise is uncorrelated (relatively random) between the averaged passes and tends to cancel itself out. But, beware, averaging isn’t always applicable. If you’re trying to measure the random noise floor of the device under test you don’t want to unfairly average its noise out in the process. Averaging, when done wrong, can also effect the accuracy, in a bad way, of various FFT calculated measurements—especially twin-tone IMD measurements. So averaging is often no substitute for having a low noise floor to begin with—something to remember when trying to use say a PC sound card for high-end measurements. Yeah averaging can improve the apparent loopback performance of your sound hardware, but it’s also unfairly improving whatever you’re trying to test.

HIGH-END BENCH DMM (revised 4/15): A surprising number of people are trying to make audio measurements with typical portable DMM’s. And the readings are often grossly wrong without even realizing it. True RMS measurements are not trivial. In effect, the meter has to accurately measure the “area under the curve” and time average it—see True RMS Measurements for more information. This proves to be rather difficult across a wide range of frequencies if you want to maintain reasonable accuracy at high frequencies and not have the reading “hunt” up and down at low frequencies. The fact is, most DMM’s priced under a few hundred dollars that claim “True RMS” are really only accurate around 60hz—i.e. power line frequencies. Some will measure sine waves accurately across the audio band, but many will not even do that. I have a $150 “True RMS” Extech meter--a relatively well regarded brand--that’s off by nearly 6 dB at 20 Khz compared to 60 hz on a sine wave and is a joke above 1 Khz on non-sinusoidal waveforms. And really complex rapidly changing waveforms like white/pink noise or real music drive such meters crazy. To do it right, you need expensive true RMS circuitry and the ability to optimize the sample rate and averaging for the waveform being measured. Good high end bench DMM’s, like the Agilent 344xx series, let you set these parameters. They also read directly in dB. I use a 6 1/2 digit Agilent true RMS bench DMM that's extremely accurate and flat from 10 hz - 100 Khz for exact levels and other measurements. It has resolution down to 0.1 microvolts so it can even be used to measure noise.

THE RIGHT OSCILLOSCOPE (updated 4/15): Most sound cards and digital audio analyzers, best case, only have a usable bandwidth to around 90 Khz. But it’s useful to know what’s going on beyond 100 Khz—like with my friend’s amp mentioned at the start of this article. The right scope can be essential. But most digital scopes use 8 bit D/A converters which are borderline useless for a lot of audio measurements as they typically only have about 45 dB of dynamic range. And many of the cheaper ones, or PC-based USB digital scopes, have miserably slow waveform update rates which is a serious problem for audio use. They really suffer when analyzing non-repetitive waveforms. And most scopes have grounded single-ended inputs--either power mains ground, or USB PC ground. This often creates noisy problematic ground loops or other problems when making audio measurements. Here are a few examples of when typical digital scopes fail:

  • Music Clipping Behavior - Lot’s of gear exhibits “ugly” behavior at clipping—but sometimes only into a real reactive load like speakers or headphones. You typically don’t want to drive many amps into clipping with sine waves into real loads as you would fry your tweeters, headphones, etc. Using real music, the slow refresh rate on typical digital scopes makes for on-screen waveforms that are jerky, blurry, and don’t come even close to following the music in real time. It’s more like watching a video with a far too slow frame rate and it’s hard to see what’s really going on. Worse, these scopes have a hard time even detecting brief infrequent clipping because their update rates are so slow (typically < 25 updates/second). Only a fraction of the music waveform is being sampled into the scope’s buffer and analyzed—the rest is missed completely. So you can easily miss random clipping. And even if you get lucky and capture a clipping event when you try to zoom in to look for signs of instability, oscillation, “sticking”, “shoot through”, etc, you might not see much because of the 8 bit dynamic range. If it’s a power amp the scope is probably set for 20 volts/division to handle the 80+ V p-p waveform at clipping. A typical 8 bit scope only gives you about 0.6 volts of resolution per bit at that setting. So any clipping behavior less than a few volts will not be terribly visible as it would only represent a couple of A/D steps in the waveform. Good luck with that.
  • FFT Behavior - Typical 8 bit digital scope have a best case 48 dB of dynamic range. That’s further compromised by the resolution of their input gain ranges—i.e. the A/D often isn’t being operated over a full scale range. And, in some cheap scopes, the analog circuitry can further compromise the performance. It costs real money to design and manufacture wide input range, low distortion, flat response, low noise analog front ends with bandwidths out to 60+ Mhz. So cheap scopes usually just aim for “close to 8 bit performance” and call it good enough. If you’re hoping to use one of these scopes for audio FFT work, you’ll find most things of interest completely lost in a very high noise floor.
  • Amplifier Destruction – If you’re trying to evaluate the output of say a bridged amplifier, or an active ground design, with most scopes you have a big problem. Just like with a soundcard, or other PC audio interface, the grounded inputs mean connecting the scope will either damage the amplifier, or at best, make it shut down or perform very poorly. So people do crazy things like un-grounding their scope, or using a battery powered laptop, but even under these circumstances, you have to be really careful not to accidentally ground anything, or even electrocute yourself. Plus there’s often a significant level of parasitic (i.e. stray) capacitive grounding still present which creates a common mode high frequency signal and can cause problems in a variety of ways. And even on non-bridged gear, grounding the outputs through a different ground path often creates other problems and sometimes even damaging oscillation. Trying to float the scope or otherwise work around this can still create problems due to parasitic capacitance creating unwanted feedback loops.

I frequently use four different scopes regularly as each offers various advantages. There isn’t one scope that’s ideal for all audio work so I use the following:

  • 100 Mhz 14 bit digital scope. This improves the theoretical dynamic range from a typical scope’s 48 dB up to a much more useful 84 dB. It’s ideal for evaluating the spectrum beyond 90 Khz well into the megahertz region with a reasonably low noise floor. The downside is the update rate is relatively slow. This scope can also safely “float” for isolation but the inputs are not isolated from each other. This avoids most (but not all) of the grounded input problems.
  • 200 Mhz fast update deep buffer digital scope. This scope updates the entire sample buffer faster than most audio is sampled (i.e. > 44 Khz). This allows it to easily detect infrequent events (like clipping with real music). It’s also excellent for evaluating digital signals like S/PDIF I2S. But it’s only an 8 bit A/D so it has limited dynamic range. The deep buffer allows capturing infrequent problems and zooming in while retaining full resolution.
  • 60 Mhz isolated digital scope. This scope has inputs that are fully isolated from each other and ground with low parasitic capacitance. It can handle several hundred volts of common mode signal without a problem. This scope is ideal of making measurements that normally would require expensive differential probes such as the voltage drop across emitter resistors (i.e. AC current waveform) in a power amp to look for things like shoot through. This scope also works where the ground schemes of other scopes create noise, ground loops, or even oscillation. Fully isolated inputs are very rare in scopes but they can be essential for certain measurements.
  • 100 Mhz analog scope. I have a nice analog scope with digital readout capability for when you really want to see something in true real-time without the limitations of any digital sampling. My 200 Mhz fast digital scope offers a similar “real-time” visual waveform. But it’s 8 bit resolution limits zooming in on say clipping behavior as outlined in the example above. So there are a few things an analog scope can still do better. But, the downside is not being able to capture and “freeze” non-repetitive waveforms and measurements are less accurate.

SIGNAL GENERATION: Most signal generators have much more distortion than even the $30 Sansa Clip+. A typical $200 bench function generator often has around 0.5% THD+N which isn’t very useful for measuring the THD of most of anything these days. While it’s relatively easy to generate reasonably low distortion sinewaves with a PC soundcard or the dScope, you’re limited to a maximum of about 90 Khz. The PC-based solution also usually creates ground loop/isolation problems. If you need something higher in frequency things get more complicated.  I use a 14 bit Tektronix AFG3000 series arbitrary waveform generator (AWG). It can produce sine waves up to 1 Mhz with noise and distortion below –70 dB (0.03% THD+N) and up to 25 Mhz with THD+N < 60 dB. It’s also flat +/- 0.15 dB from 0.001 hz to 5 Mhz and has an 18 nS rise time for impulse, square wave, and slew rate testing. The 14 bit Tek AWG, paired with the 14 bit 100 Mhz digital scope, allows audio testing well beyond the bandwidth restrictions of even the flagship Audio Precision SYS-2722. This is useful for exploring slew rates, open loop characteristics, loop stability, Class-D artifacts, switching power supply artifacts, etc.

RMAA IMPROVED: For RightMark Audio Analyzer testing I use a Benchmark ADC1 for A/D and a Benchmark DAC1 Pre for D/A.  Both work with built-in Windows drivers and hence play nice with RMAA. And both have some of the best specifications available at any price from a USB audio interface. The ADC1 features precise gain adjustments and metering, and the DAC1 has analog output level adjustment with the 24 bit D/A always operating at full resolution. The Benchmark “twins” overcome many of RMAA’s limitations. I use the Agilent DMM (or dScope) to set exact and repeatable levels and also use the correct loads. See the RMAA article for more.

PERFORMANCE EXAMPLES (updated 4/15): Here’s the dScope’s averaged absolute noise floor in dBu:

Prism dScope Unweighted 8 Pass Averaged Residual Analog Absolute Noise in dBu

Here’s the dScope’s residual analog performance (loopback) with a “zoomed” vertical axis operating at the 400 mV reference level I use for many tests. The third harmonic is the worst at –110dB below 400 mV:

Prism dScope Analog Loopback Residual Audible Band Detail

And, while it’s a bit silly, here’s the dScope loopback performance in the digital domain (click to get rid of the blur):

THE ENTIRE CHAIN'S PERFORMANCE: Here's the full deal from the photo at the start of this article. This is the analog output (signal generator) of the dScope, feeding the Benchmark ADC1 to digitize the signal. The ADC1 feeds the Benchmark DAC1 Pre via S/PDIF at 24/96, and the analog line output of the DAC1 is connected to the analyzer input of the dScope. The number on the left is pure THD (just the harmonics, not the noise floor) and the one on the right is the THD plus the noise in the entire signal chain. Note the bandwidth is out to 96 Khz on the analysis side. This is at 1 Khz and 0 dBFS digital while the analog levels are 2 V RMS (click for the full size image): 

BOTTOM LINE: Equipment of this level might be overkill for testing something like a $40 portable player. But it's nearly essential if you're designing or testing high-end audio gear. For example, just a small error in designing the PCB for a DAC or amplifier can seriously degrade the performance due to noise problems, grounding issues, etc. You can't just slap a D/A or audio chip on a PC board and expect to get anywhere near the manufacture's specs without some very careful design work. But the only way to verify you have it close to right, is to have the correct instruments to test with or getting assistance from someone who does. So, for design work, and testing higher-end gear, this level of equipment is invaluable.

February 2, 2011

RightMark Audio Analyzer

ROSE COLORED GLASSES (updated 3/3/11): RightMark Audio Analyzer ( aka RMAA from rightmark.org ), in some ways, is impressive software for the price (free!). A lot of people, websites, and organizations, are using RMAA to publish audio measurements. It's a useful tool when used properly and the user is aware of its many limitations. Considering it’s available for free the developer should be congratulated. But RMAA has many limitations and potential problems. And a lot of people use it incorrectly and, often unknowingly, publish misleading RMAA results.

THE SHORT VERSION FOR NON-RMAA USERS:  You can rarely compare one person's RMAA results to another person's RMAA results. This is especially true for evaluating anything other than a PC audio interface (aka “soundcard”). It's very easy to have a better piece of gear actually show worse results on RMAA because of different test conditions, settings, etc. So don't put too much trust in any RMAA results unless you made them yourself under carefully controlled conditions. Unless you know the absolute levels used, load applied, PC sound hardware used, and settings of the PC and device being tested, you simply can’t trust the results. Changing any of these things can cause a greater change in the results than testing a different audio device.

SETUP IS EVERYTHING: With RMAA there are many things to get wrong. Examples include hidden mixer controls, sound processing settings in the PC sound driver, sample rate conversion by the operating system or sound hardware, level settings—both at the PC and for the device being tested, loading, ground loops, cabling, etc. And if you’re doing loopback testing, you have no way of knowing what performance is related to the A/D, the D/A, or both. Even having a cell phone nearby can corrupt RMAA measurements with no error indication at all—just erroneous measurements. RMAA testing has lots of “blind” aspects that users are often unaware of. It has very poor internal consistency and error checking. Instead of giving a warning something might be wrong, it more often just produces false results.

15 IMPORTANT THINGS RMAA DOES NOT MEASURE (revised April 2011): There are many things that can make an audible difference in sound quality RMAA does not test for. Many are difficult or impossible to measurel with a soundcard-based test setup even using other software, specialized hardware, etc. They include:

  • MAXIMUM CLEAN OUTPUT LEVEL - RMAA has no concept of absolute levels. It can't measure voltages, power outputs, etc. So you have no way of knowing, for example, how loud that portable player or headphone amp can play without obvious distortion. You also have no way of knowing what level you’re testing it at. You could be measuring a device right on the edge of clipping. Or the level might be so low you’re measuring more noise than distortion. RMAA can’t tell you if you’re testing at 2 volts of audio signal or 0.02 volts. 
  • OUTPUT IMPEDANCE - Anything designed for driving headphones including PC's, portable players, headphone amps, USB headphone DACs, audio interfaces, pro gear, etc. has an output impedance. It varies widely and can have a huge impact on the sound quality--especially with certain types of headphones (see Headphones Amp Impedance). RMAA does not measure output impedance at all. This is a huge limitation as it’s one of the biggest factors in the sound quality of headphone sources.
  • DAC LINEARITY – Some DACs exhibit considerable non-linearity at low levels. RMAA doesn’t measure low level linearity, and because it can’t measure absolute levels, everything in RMAA is only relative to an unknown reference. 
  • SQUARE WAVE PERFORMANCE – A 1 Khz square wave reveals a lot of information about analog and digital audio components such as stability, bandwidth, rise time, compensation, and for digital devices, the type of digital filtering used. One or more of these can have an audible impact on device performance. But RMAA doesn't have a square wave test. And even if you use oscilloscope software for a soundcard, it will only measure up to the cut off frequency of the PC’s sound interface. The bandwidth of PC sound hardware is usually too low to see ultrasonic/RF, ringing, oscillations, switching power supply noise, Class-D amp noise, or instabilities. The very limited bandwidth is like looking through rose colored classes. The result can easily appear much better than the real performance.
  • JITTER - Jitter has been proven to be audible in some circumstances. RMAA has no jitter test.
  • VARIOUS TWIN TONE TESTS – RMAA has a single IMD test that’s similar to the SMPTE test but it’s not clear if it follows the SMPTE standard. And there’s no capability for other important twin tone tests such as the popular CCIF 19 Khz/20 Khz which is very revealing of high frequency performance. From what I’ve seen, RMAA doesn’t even calculate SMPTE distortion properly. The results are often artificially optimistic.
  • THD20 – It’s very useful to measure THD at 20 Khz for a variety of reasons. It’s a good indicator of stability, feedback loop design, and more. To perform this measurement properly, you should be able to measure the first 3 harmonics at 40 Khz, 60 Khz and 80 Khz. Even if you get RMAA to work with at a 192 Khz sampling rate, it still can’t perform a proper THD20 measurement. And it has no option to adjust the measurement bandwidth of its THD sweeps.
  • LOW FREQUENCY DISTORTION – It’s also useful to measure THD+N at very low frequencies such as 5 to 10 hz to expose thermal modulation effects, power supply issues, and more. RMAA doesn’t allow arbitrary low frequency testing.
  • POWER vs THD: RMAA cannot perform the classic measurement of output power versus THD. This is a standard benchmark amplifier test that’s widely performed as it’s very revealing of the amplifier’s behavior at different levels. RMAA has no way to perform this critical measurement nor am I aware of any other software that offers this sort of measurement with a sound card.
  • SELECTIVE SPECTRUM TESTS: It’s very useful to perform different spectrum testing with various input signals. RMAA doesn’t give you any control over its signal generator or its analysis. You cannot control the FFT points, averaging, weighting, filtering, etc. This is very limiting for audio testing. You’re stuck with the ‘canned’ results RMAA provides and many of those don’t follow industry standards (or it’s not clear how they’re calculated).
  • RESIDUAL ANALYSIS – Analyzing the residual distortion products can be very revealing—for example it can determine if an amplifier suffers excessive crossover distortion. RMAA has no way to do this.
  • REAL TIME RESULTS - Many problems may only show up briefly or intermittently. And many adjustments (like finding the clipping point) are best performed in real time. For example, maximum output is commonly defined as 1% THD. So with a real time audio analyzer you just raise the level until it reads, in real time, 1% distortion. With RMAA it takes a long time to run a test sequence which makes if very difficult to see what affect adjustments have on the results. It might take 10 or more passes to set the level correctly for even close to 1% THD. And if you have a bad cable, or some other intermittent problem such as outside noise or interference during the test, you will likely never know because you just get a dump of results at the end. I’ve had a cell phone randomly mess up RMAA measurements in ways that just made the product being tested look bad with no other indication the results were invalid.
  • SLEW RATE – RMAA cannot measure slew rate which, especially for audio power amplifiers, can be an important measurement. PC sound hardware is bandwidth limited making these sorts of measurements difficult or impossible with any PC software.
  • DELAY & LATENCY – RMAA cannot measure delay and latency for digital hardware.  This can be critical for live audio recording, studio work, gaming, and for lip sync with video.
  • HARDWARE LIMITATIONS – Nearly all PC sound hardware has some severe limitations that restrict RMAA and any other software. For example bridged and certain other amplifiers cannot have any of their output terminals grounded or connected together. But typical PC sound interfaces have common grounded inputs. To make matters worse, these grounds often go back to the PC’s  AC power grounds which creates even more problems. Some of these issues can seriously harm the device being tested and/or your PC’s sound interface or even the PC itself. USB and Firewire sound hardware is also grounded back to the PC. Most of these devices can only work with up to a few volts of input. Higher levels will cause overload and even damage the hardware. As mentioned earlier, the bandwidth, best case, is usually limited to about 90 Khz and more often only about 20 Khz which precludes many measurements and often hides other problems. There are also inherent issues with bit resolution, gain matching, channel balance matching, absolute levels, and more.

 

ALL THE NASTY DETAILS (non-geeks probably want to skip this section):

If you're not familiar with RMAA, here’s a typical screenshot. Note the 3rd column in the results labeled "RMAA Problems" and the highly questionable numbers. More on that later (click for larger):
 

 
PC HARDWARE ISSUES: RMAA results are only as good as your soundcard and how RMAA is used. Lots of people are running RMAA on whatever sound hardware their computer came with. And that's often a serious limitation. Built-in sound hardware is prone to all sorts of problems--especially the A/D section which is often only used for digitizing voice these days so manufactures don't put much effort into making the built-in A/D signal path perform well.

Nearly all built-in sound hardware amounts to little more than a cheap CODEC chip stuffed onto a crowded and electrically very noisy motherboard. Modern CPUs create spikes of current that can exceed 100 amps. That generates a lot of electrical noise and it's hard to keep it all out of the nearby audio circuitry where it can confuse RMAA measurements.

Even with external sound hardware, other applications or your operating system can play random sounds while you're running the test (like when you get new email). Even CPU loading and hard drive activity can affect the noise floor of built-in audio hardware--especially the line inputs and A/D. And while RMAA should log an error if extra sounds are played during the analysis, it just corrupts the results in a non-obvious way (more on that later).

Some PC's (especially laptops) don't even have a line level input. They only have a microphone input. Trying to use a mic input for line level audio is a disaster and will yield very misleading results. Yet some still do it and publish the results with no indication of how they obtained the results and hence no clue the results might be invalid.

Lots of PC's and sound cards perform various audio processing in the digital domain. It's sometimes impossible to turn all this processing off hence many PC sound devices are not "bit accurate". Some perform internal sample rate conversion, for example, regardless of how they're configured (Creative's Sound Blaster and X-Fi devices, for example, infamously perform conversions). Others attempt to optimize or enhance the audio in various ways. This internal processing affects the accuracy of RMAA in unpredictable ways.

Finally, even lots of outboard PC sound devices have problems--especially with their A/D hardwaer. Some have microphone preamps always in the circuit which seriously compromises their performance for line levels. Others have very inexpensive A/D chips and circuits in them. Some are not bit accurate, have noisy USB-derived power supplies, and/or problematic level controls, etc.

GAIN MATCHING: With many PC sound interfaces, the mixer settings operate in the digital domain. And when you lower the level settings you may lose digital bit resolution. The effective number of bits might be 16 bits with the controls maxed, but if you have to reduce the gain, you may get less than 16 bit resolution. At some point this can become the limiting factor in the test setup (insufficient ENOB). This is especially critical for measuring very low levels like noise, distortion, or crosstalk (channel separation). 

And changing levels is made more difficult because PC sound hardware levels controls are rarely calibrated in known increments (like dB). And even the few I have seen that are calibrated in dB (in software) don’t “track” well with the actual hardware settings. For example, some steps are close to 1 dB but near the end of the range can be as high as 3 dB. The manufacture never intended for their hardware to be used as a lab instrument so they probably didn’t much care about such accuracy as it’s not a big deal for 99% of applications.

So, ultimately, you have to verify levels with an external meter and not just any meter (more on that later). Any changes in the controls or settings usually requires establishing the levels with the external meter all over again.  This can get very tedious when dealing with audio signals that often vary over a wide range during testing. And if you don’t re-adjust the level controls to operate the A/D near full scale, you can lose lots of resolution and dynamic range making the device being tested seem to perform much worse than it really does.

Level setting also depends on the topology of your PC sound hardware. Some of the external USB or Firewire devices have their own physical pots for levels controls. I’ve never seen these calibrated in any meaningful way. They typically cover an extremely wide range of gain (i.e. 65 dB for the E-mu products). These controls, because of their wide range, are extremely “touchy” and it’s just about impossible to make precise small adjustments or get the two channels precisely balanced. And the settings can even change on their own which can really throw off RMAA if you don’t realize it happened during a test.

Better devices use programmable gain amplifiers (PGAs) ahead of the A/D converters and control these PGAs from software. This can be a better solution, but only if you’re given calibrated control in sufficiently fine steps, and the PGAs themselves are accurate. I’ve yet to find any interfaces that meet these criteria. Many of the less expensive ones are not terribly uniform in their steps over their full range. PC sound hardware is not designed with absolute precise values in mind. But that’s exactly what’s needed for a good measurement set up.

Some PC sound hardware uses a fixed gain amplifier and some sort of “digital pot”. These are often less accurate than the PGAs above. And often have fewer, and more coarse, non-linear steps. And the high fixed gain amps can have more noise issues. Digital pots often use a linear rather than logarithmic scale. This can make the step size rather random.

Some PC sound hardware lacks any sort of attenuator or PGA on the inputs. They simply design the circuit so the maximum expected input corresponds to roughly 0 dBFS in the A/D and then do everything in software from there. While this is fine for a person’s voice for Skype, it’s far from ideal for measurement purposes. As the signal level is reduced, the effective resolution of the A/D goes down with it. If your source happens to have a maximum level close to the maximum of the sound hardware input you’re in fairly good shape (say 2 Vrms the Redbook standard for digital audio). But if you’re trying to measure a portable player with a maximum output of 0.4 volts you just lost 14 dB of dynamic range from your measurements.

The gain increments are often too large, variable and/or unknown. For example, if you want to remove a 0.7 dB channel balance error, but only have 2 dB steps to work with, you’re stuck. And the steps often don’t correspond to any convenient number of dB because they’re derived from whatever the hardware supports. Most electronic volume control chips, for example, only have 64, 128 or 256 steps. So what you get is the full range divided by 64 uneven increments—whatever that happens to work out to.

Proper audio analyzers have calibrated analog stepped attenuation built into the input and output circuitry so regardless of the signal level they maintain the resolution of the D/A and/or A/D—to always be within a few dB of full scale. You can also set levels in their software to a resolution of at least 0.01 dB. I’m not aware of any PC audio devices that can do anything even close.

NO ABSOLUTE LEVELS: Even if you run RMAA on a decent PC audio interface with good specs, you still have no idea of the absolute levels. Are you testing too low? Too high? Near the device’s clipping point? This also prevents knowing the true output capability of the device being tested (i.e. dBu, dBv, dBm, volts, watts, milliwatts, etc.). It also means the channels may not be properly balanced which affects RMAA's measurements—the RMAA calibration routine compensates for channel imbalance in your device when you really want to be measuring that imbalance. The actual levels can make a BIG difference in the RMAA numbers you get. Testing a device at 1 volt versus 0.1 volts will yield a 20 dB difference in noise and distortion levels. That’s huge.

AN EXAMPLE: Joe posts his RMAA results for Player A and tested at a relatively low level. This will yield noise numbers significantly worse than the device is capable of. Bill posts his results for Player B but, unknown to Bill, the level was set too high and it was on the edge of clipping. So Bill’s noise numbers look really good but his distortion numbers look bad because the player was on the verge of overload. Hopefully you can see how it’s impossible to draw many valid conclusions from the published results because the tests used very different levels. But that doesn’t show up anywhere in the RMAA report or results. Player B might really be the better player, but Bill had no idea it was starting to clip so the high distortion results make it look like the inferior product.

REFERENCES ARE CRITICAL: The proper way to do audio measurements is to use standard reference values. For example  0 dBu is 0.775 volts RMS and a standard for consumer audio. Without knowing the actual levels of the signals being measured, you have no way to use proper references or even know what the levels are so you can test the next device, or even retest the current device, at the same level. And even if you use an external meter (most of which are designed for 60hz AC power and not even close to accurate across the audio band) you're still likely stuck with the un-calibrated mixer/level controls of your PC audio device. So the instant you change any of the controls, your meter measurements are rendered useless. Without absolute levels, and a known reference, it’s usually impossible to duplicate the results.

NO STANDARDS: RMAA isn’t very clear about what standards it conforms to, if any. Did the guy who wrote the software build in A Weighting to the noise measurement? If you use a higher sampling rate, are the THD values only calculated over the audio spectrum or do they also include ultrasonic noise and distortion? How is the IMD value calculated? There are lots of unanswered questions about how RMAA works internally and it makes comparing RMAA data to real measurements that conform to accepted standards difficult or impossible.

TEST LOADS ARE CRITICAL: Loading is often overlooked. For example, portable MP3 players often measure far better unloaded driving just the line input of your PC audio device than when properly loaded with real headphones or a suitable test load. Nearly all the RMAA measurements I see published never mention anything about what the device was loaded with so I'm guessing it's usually nothing. Unloaded tests are often meaningless. Some measurements, like distortion and crosstalk, are hugely affected by the load. And if you're trying to measure a line level output (rather than a headphone output) most PC sound devices don't have a well defined input impedance. Here again, this makes repeatability, comparing, and verifying results difficult—especially at higher levels of performance.

USING HEADPHONES AS THE LOAD: This can be good and bad. The good news is it’s a realistic non-linear load--especially if they're the headphones the user intends to pair with the device. The bad news is someone else needs the exact same headphones (which they likely don't have) to conduct a fair test that's valid for comparison. And for many devices the headphones will alter the measured frequency response in ways more related to the headphones than the device (see the example below). And, unknown to many, headphones also act as microphones. So, for example, they pick up background noise in your room and this can raise the noise floor measured by RMAA making for misleading noise and distortion measurements. A proper resistive test load is better if you want reproducible results that can be more fairly compared with other results. It also avoids the "microphone effect".

Here are the RMAA results using a 15 ohm resistive load, and 2 different kinds of headphones--the UE SuperFi 5 Pro's have a really wide impedance swing from about 10 ohms to 85 ohms and a pair of Sony MDR-EX76's that only vary by a few ohms at higher frequencies. The levels were within 2dB of clipping at the worst case frequency using a Benchmark ADC1:


The best numbers are with the 15 ohm resistive load. When using the Sony's the THD more than triples and the noise floor rises up a few dB. And what's with those SuperFi 5 numbers? Does the otherwise low distortion source suddenly produce 4% THD? No, but RMAA makes you believe it does! What likely happened is the signal was clipped because of the frequency response swings. But you have no way of directly knowing that, and from the average person’s point of view, the levels were properly calibrated using RMAA’s calibration routine. Here's a "zoomed in" graph showing what effect the Sony headphones have on the frequency response:


So if Joe tested with a resistive 15 ohm load you'd see the white graph in his results. If Bill used his Sony headphones you'd see the green graph published and probably think less of the product being tested. And if Bill used his SuperFi 's here's what you'd see (in green)—a whopping 15 dB of response deviation:


Here's a Sansa e260 MP3 player with no load and a 15 ohm load:


Notice it has more than 8 times as much THD, more than 4 times as much IMD, the crosstalk is far worse, and the frequency response much worse when loaded. Here's the frequency response difference:


There's likely a coupling capacitor in the output amplifier of the Sansa. And, unloaded it has little effect, but with a typical load it's down -3dB at 40 hz. There's also a weird rise at high frequencies likely caused by marginal feedback design in the amplifier. Here's the swept IMD vs frequency:


Unloaded the distortion is down below -72dB worst case, but loaded it's more like -57dB which is likely audible. With no load, as many use RMAA, the Sansa rivals some of the better players out there. Loaded, however, it's closer to the bottom of the pack. It makes a huge difference! These are typical examples. It makes a very big real world difference.

GROUNDING: Even an external PC audio interface is grounded to your electrically noisy PC via the USB/Firewire connection. And the inputs and outputs share a common ground. These both can create potential ground loops, noise sources, or worse, with whatever device you're trying to test. If the device has bridged outputs it may harm the device to ground one side of the output. And many even conventional (non bridged) amplifiers can oscillate or even self destruct if you connect their input and output grounds together externally. You can use transformers to isolate the device, but they introduce lots non-linear distortion on their own so you end up measuring mostly the transformer. Proper audio analyzers have differential and/or floating inputs and outputs that are not only isolated from ground but from each other.

HIDDEN GROUND PROBLEMS: Ground loops, especially where PCs are involved, can create noise in non-obvious ways. It’s very likely a lot of RMAA results are badly corrupted by ground issues but those making the measurements are likely totally unaware of the problem. They just see results that are much worse than what the device is capable of when it’s used normally. You have to remember RMAA was mainly designed to test sound cards in loop back where grounding typically isn’t an issue.

OVERLOAD: If you're trying to measure gear intended to drive speakers, it's likely to overload and possibly damage your PC hardware if you directly connect the output to the line input. Such inputs are usually only designed to handle a few volts maximum and some even less.  But speakers usually require 10+ volts. So you have to make an external divider network which is another potential source of problems and inaccuracy. It also makes it more difficult to compare RMAA measurements made by others unless they’re using the same divider network.  And there’s still the grounding problem mentioned above which can literally damage the gear you're trying to test as well as your PC hardware.

SAMPLING RATE: RMAA’s measurement range is limited by the sampling rate of your audio hardware and what the drivers support. Often it doesn't work (or work right) above a 44 or 48 Khz sampling rate. This restricts the bandwidth to around 20 Khz. Many devices may have instabilities that show up at ultrasonic or even RF frequencies. RMAA is usually completely “blind” to these problems. See Testing Methods for an interesting real world example. And Class-D amplifiers and/or switching power supplies are increasingly being used used in everything from MP3 players to A/V receivers. These switch at high frequencies and often create large amounts of noise that’s entirely missed by typical PC audio inputs. Even if you don't think it's audible, the amount of out-of-band noise can still be a useful indication of the quality of the device being tested. And stability issues often are audible.

BIT DEPTH AND NOISE FLOOR: RMAA is typically limited to 16 bits resolution by either the PC audio hardware and/or driver limitations. But you often can't use the full 16 bit range because of gain limitations (see GAIN MATCHING above). So, in reality, you might end up with less (or similar) useable dynamic range than the device you’re trying to test.  Ideally the test set up should have far better dynamic range than whatever you’re testing. Even if you're lucky enough to get RMAA to work correctly with 24 bit drivers, you won't get anywhere near 24 bit performance from typical 24/96 or 24/192 PC sound hardware--you're lucky to get 17 or 18 bits of effective resolution (ENOB) due to the noise floor, power supply noise, grounding issues, typical A/D and D/A converters, etc.  Often this level of resolution is inferior to what you’re trying to test. ideally your measurement hardware should be at least 10 db better than the gear you’re trying to measure. That often requires expensive hardware which often doesn’t work with RMAA correctly anyway.

FALSE SECURITY: People often do a "loop back" test to first test their PC sound hardware. And, more often than not, RMAA reports fairly impressive results. So they think they're good to go. But there are often problems being masked, or caused, by one or more of the above issues. For example, level, loading, and grounding issues won't show up in a loop back test. And don’t forget the previous list of 15+ things RMAA doesn't test for at all.

RMAA DEFECTS: On top of all the above issues, the RMAA software itself is buggy and prone to problems. The last release was several years ago and development has apparently been abandoned. Some of the FFT/math used to calculate the results is apparently wrong and/or has serious limitations that are neither properly documented nor obvious. In my experience the results often disagree with a professional audi analyzer. Plus there are lots of quirks and some obvious bugs. For example, try right clicking on any of the icons for the combined results graphs--instant total crash and you lose all your results.

Perhaps most important, RMAA often just outputs bad data instead of indicating an error condition with the measurement. To use an extreme example, you can stop playback of the test file half way through the tests, and instead of issuing an error, RMAA goes right on calculating and randomly spews out half truths and half garbage. It's obvious the developer didn't bother to put in many checks for validity of the input and output data. This can take much more subtle forms such as bad cabling, random noise (RFI from a cell phone, sound events from your PC, etc.),  not being detected during the test.

Here's an example of RMAA inventing impossible results out of nothing. It was run with no test file at all and, instead of timing out, it ran without a single error or warning and here's the result:


Wow, 245% IMD, that's pretty bad! I didn't know you could have more than 100%. But the THD is “only” 1.7%. The frequency response doesn't look so good though. I joke, but you get the idea. This doesn't inspire confidence in the software. And then there are messages like this one:


So clipping occurred but 0.000% of the samples were actually clipped? Hmmm. The funny thing is, in this case, there was zero clipping by the ADC (my Benchmark ADC1 has clip lights that stay on if it even clips a single sample until you reset the LEDs). And what's even more interesting is lowering the level by several dB and running the test again (as the error suggests doing) yields the exact same error! It appears to just be a case of RMAA getting confused trying to test this particular device (a device that tests just fine on my Prism dScope). Again, this does nothing to boost confidence in RMAA results.

RMAA DESIGNED BACKWARDS? I’ve seen enough weird results from RMAA I have to wonder if the developer designed it backwards to match the specs of a reference soundcard? Whatever test signals and analysis being used might have been “tweaked” until he got numbers that roughly matched the specs of some respected soundcard. If true, this would explain a lot. Ideally RMAA would just do all the math correctly and the results would be verified against a professional audio analyzer to make sure the math was accurate. But it often doesn’t agree and I’m at a loss to otherwise explain why unless there are either bugs in the calculations or it was designed to produce the expected numbers when testing a certain soundcard.


RMAA TIPS: So how is a person supposed to make RMAA measurements more accurately? First, use the best PC audio hardware you can. A good device would be something like the RME Fireface UC but, in my experience, proprietary low latency drivers designed for multi-track recording don’t work well (or at all) with RMAA. Example popular devices that do work are the E-Mu USB audio interfaces:

  • Creative/E-Mu 0202 - This device has decent A/D and D/A performance but it has un-calibrated level controls that cover a wide range of gain (65 dB!) and are very "touchy". This makes setting the level the same for both channels just about impossible. So you can’t know the actual channel balance of the device you’re trying to test and certain other tests can be inaccurate. The 0202 also has a mic preamp in the signal path on just one channel. Not surprisingly, that channel has higher noise than the line-in only channel. So any device you test with the 0202 is prone to rather lopsided results, weird crosstalk numbers, etc. It's also at the mercy of the sometimes noisy USB power supply. And it has serious issues on some PCs with its proprietary buggy drivers.
  • Creative/E-Mu 0404 - This device is better than the 0202 in that it has its own dedicated power, and both channels have the same signal path (unfortunately both have mic preamps). It also has slightly better metering but still suffers from the "touchy" gain control problem and the inability to exactly match the channels without some external help. See the section below on Levels. Another downside of the 0404 is it requires the same special drivers as the 0202. The drivers haven't been updated in years and especially can have issues with Windows 7. As far as I know there only buggy beta 64 bit drivers.

SETTING LEVELS: This is critical for many reasons. First, it's useful to know if the device you’re testing has a channel imbalance. But if you use the meters in RMAA to say adjust your E-Mu 0404, you will remove any channel imbalance anywhere in the signal chain. You can either feed known matched levels (i.e. verified with a meter) into both inputs and then run the RMAA calibration, or use a "Y" cable to temporarily split one channel into two outputs to set the levels. That way you know any channel imbalance that shows up during testing is the device, not your set up.

MEASURING LEVELS: As discussed earlier, it's important to know what levels you're really testing at. So it's best to use an external true-RMS meter that’s accurate across audio frequencies. Most cheap DMM's, even many that claim true-RMS, are only designed to be accurate around 50 – 400 hz. They often "roll off" dramatically or behave unpredictably at higher, or lower, audio frequencies. True RMS calculation is non-trivial and it’s expensive to have it be accurate at higher frequencies. So most reasonably priced meters don’t bother as they’re not intended for audio use. So if you're testing at say 1 Khz, they might read only half the real value. Looks at the specs for the meter, or if you already have one, test it on the output of your sound interface with test tones across the audio spectrum to see how flat it is (or isn't). It's a hassle to do the math at different frequencies to apply a correction factor but it’s cheaper than buying a more expensive meter.

CHOOSING A TEST LEVEL: If you're testing a device with a volume control, how far do you turn up the volume? I would suggest picking a value that's  well below the maximum output to help assure it won't be close to clipping. But if you test at too low of a value, the noise and distortion numbers will be much worse because noise will dominate the measurement. If you have a meter (see above) and the maximum output level is specified for the device a good starting point is about 70% of that value (roughly –3 dB). But this may also depend on the load being used. if the maximum output is not specified, see the next paragraph.

REFERENCE LEVELS: Here are some guidelines for reference levels:

  • Headphone Outputs - A good guideline for headphone outputs is at least 1mW into 32 ohms as that's used in the manufacture's specs of many devices and a value nearly any device can manage without clipping. P = (V*V)/R which means V = SquareRoot(P*R). So the square root of (.001v * 32 ohms) is 0.179 volts or about 180 mV RMS. This also happens to be about the typical power level most listen to their music at with typical low impedance headphones. So use your meter to set the player/amp to 180 mV RMS while playing a 0dBFS 1 Khz reference file. Then set your levels on your PC sound interface/mixer controls to be just under clipping (0 dB) while monitoring the 180 mV signal. But, beware, such low reference levels may challenge your PC hardware when making noise and crosstalk measurements. Many headphone outputs can deliver 300 – 500 mV into 32 ohms at low distortion so you might want to run a test at several different levels to find where the distortion starts to dramatically increase and the run the test just below that level.
  • Line Level Outputs – The Redbook standard for home digital audio gear is 2 Vrms. That’s what you’ll get from most home CD players playing a 0 dBFS signal. But portable and USB powered line outputs are often lower—typically in the range of 0.5 volts to 1.6 volts. And some gear can be above 2 V rms.
  • Speaker Outputs – As mentioned earlier, you need to use a voltage divider to avoid overloading and likely damaging the line level input of your RMAA sound hardware. This is not something I want to cover as if someone gets it wrong it could cause expensive damage. But for those in the know, the standard level for checking noise and distortion is typically 1 watt into 8 ohms which is 2.8 Vrms.


RUN A LOOPBACK TEST AT THE SAME LEVEL: Regardless of what sound hardware you're using, test it with a "loop back" test by connecting the inputs to outputs as instructed on the RMAA website to know what the "baseline" performance is at that level. Try to do this test at the same absolute level you're going to use for testing your piece of gear (i.e. the 180 mV mentioned above, etc). The signal to noise ratio, distortion, crosstalk, etc. of your PC interface will change at different test levels.

USE A PROPER LOAD: See the earlier section on loading. Ideally use a resistive test load. For devices intended to drive headphones, use resistors in the range of 15 – 600 ohms. A good single value is 32 ohms. Or test at several different impedances if you want. Even small 1/4 watt resistors are fine. They should be carbon film, and not wire wound. Use caution using headphones as a load both because they may be damaged by test tones at high levels, and they can also pick up room noise and create erroneous results due to the their impedance variations with frequency. For line level outputs the (often unknown) load of your RMAA sound hardware is probably sufficient. For speaker outputs the normal load is 8 ohms and it needs to be rated to handle whatever power level you’re going to test at.


VERIFY YOUR CABLES:  You don’t need $100 audiophile cables, but $1 ones from eBay are not a good idea either. Because of the way RMAA works, if you have a cable problem you may never know and just get marginal numbers. So test all the cables in the signal path by listening to a test tone while moving them around. If you hear static, or worse, replace the cable.  And, when testing crosstalk, cables become a huge problem if there’s any extra shared ground wiring.

RUN THE TESTS MULTIPLE TIMES: A good way to help rule out intermittent problems is to run the tests multiple times. If you get consistent results, they’re more likely to be valid. If the results vary by more than 5%, something is likely wrong.

USE A REAL OSCILLOSCOPE:  If possible, it’s worth checking the output, especially driving a real world load, with an oscilloscope that has a bandwidth out to at least 5 Mhz. This may show problems not seen in RMAA using a sound card limited to 20 Khz – 96 Khz. You can also measure the slew rate on a real scope and better evaluate square wave and impulse performance. Beware most scopes have the same grounding issues as PC sound hardware.

HOW I DO IT: For RMAA testing I use a Benchmark ADC1 for the A/D, a Benchmark DAC1 Pre for D/A, a 6 1/2 digit Agilent bench DMM that's extremely accurate and flat from 10 hz - 100 Khz for exact levels and other measurements, and one of several oscilloscopes with 60+ Mhz bandwidth. But, far more often, I use a dScope audio analyzer as it can make far more measurements, conforms to international testing standards, and produces far better and more accurate results. See my blog post Testing Methods for more information.