Header

Objective Reviews & Commentary - An Engineer's Perspective
Showing posts with label Benchmark DAC1 ADC1. Show all posts
Showing posts with label Benchmark DAC1 ADC1. Show all posts

April 8, 2011

Beyerdynamic DT 770 Pro 80 Review

beyer dt770 benchmark

NOT MY USUAL FARE (revised 4/23): Generally I’ve been sticking to reviewing electronics as it’s easier to provide objective measurements and directly compare performance. I happened to give the Beyerdynamic DT 770 Pro 80 headphones a listen at a dealer and liked them enough to bring them home. They turned out to be an interesting example of how different sources can significantly change headphone sound.

REVIEW UPDATE 4/23: After living with the DT 770 Pros for another few weeks, and having a new “reference headphone” to compare them against, I’ve revised several of my comments in this review. I’ve also incorporated feedback from others. As happens with subjective reviews, it’s easy to let some aspects of a product (such as being really comfortable) bias or mask your subjective opinions of other aspects of the product. Motivation to write a long review comes easier when you’re enthusiastic having just listened to something new—hence all the reviews published soon after a product’s purchase. But it’s more realistic to live with gear long enough for the “newness” to wear off. And, in hindsight, that’s what I should have done here.

WHO’s BEYER? Beyerdynamic is the other popular German headphone company. They’re not nearly as well known among consumers as Sennheiser but they still have a loyal following—especially among musicians and professionals. AKG used to be almost German (Austrian) but is now a Harman International USA brand.

CIRCUM WHAT? The DT 770 is a true circumaural headphone. That’s a fancy way of saying they completely cover the ears. Anything resembling a full size headphone is often lumped into this category, but many still rest at least partly on the ears rather than over them. A good example is the popular Sony MDR-V6 (MDR-7506). They have shallow ear pads and rest mostly against your ears rather than your head. The Beyers pretty much leave my ears alone.

THE BASICS: The DT 770 Pro 80s are surprisingly lightweight (about 9 ounces) full sized headphones intended mainly for studio use. But they’ve also found a loyal following among some home listeners. The rated impedance is 80 ohms and they have an “SPL” rating of 96 dB SPL. Those two numbers mean they’re a bit marginal for use with most portable players (see Headphone Impedance). There is a better DT770 version for portable use. More on that later.

CLOSED MINDED: Full size circumaural cans come in two distinct flavors—open and closed. Most of the best sounding headphones are open to varying degrees. This allows the driver to be less affected by whatever is behind it. At least in theory, open back headphones have fewer reflections and resonances to color the sound. The downside is they let the music leak out which might disturb others. They also let more outside sound in. For some, however, closed cans have practical advantages.

ISOLATIONIST TENDENCIES: Most closed full size headphones don’t isolate as much as some might expect. They’re significantly better than open backed headphones but the best sounding and most comfortable closed headphones tend to have less isolation for some good reasons. For one, they don’t “clamp” the head tightly. While this greatly improves comfort it also lets more noise in. And they usually design the ear pads to intentionally “leak” as this helps vent the chamber between the transducer and the ear for more accurate sound. The ones that act like hearing protection muffs sound significantly worse. The special “M” high isolation version of the DT 770 is a good example—they don’t get good reviews for sound quality and are not as comfortable. If you want both great sound and high isolation, in ear monitors (IEMs) such as those from Shure and Etymotic are the best option. But they have their own downsides.

WHY THE DT 770? I need to listen to headphones for extended periods while not bothering others. And I’m often looking to filter out annoying background noise. With closed back headphones you don’t have to crank up the volume as high to drown out other sounds. That helps save your hearing and also makes for less listening fatigue over the long haul. Serious comfort is a must for listening hours at a time. I have several IEMs, and they’re great for flights, the gym, etc. But my ear canals get sore with long and frequent IEM use. They’re also a hassle to take in and out to talk to others, answer the phone, etc. So that leaves full size circumaural closed cans—like the DT770s. And if you want good sound and genuine comfort, the list is fairly short.

SOME COMPETITION (updated 4/10): The DT 770, with a $200 street price, has some significant competition. The ones I’ve heard are:

  • Sennheiser HD 280 Pro – You see these recommended everywhere. They’re reasonably priced with a street price around $100. They’re nicely made, fold up, and provide more isolation than many closed headphones. But, like many closed cans, they don’t really sound very good. The also have high clamping pressure, and the ear pads/ear cups are not really big enough to fully enclose most adult size ears without putting pressure on the ear itself. So many, me included, find them rather uncomfortable for long term regular wear. I own the HD280 and I’ll comment more on the sound quality later. They’re 64 ohms and work OK with some portable gear. Sennheiser doesn’t provide a typical sensitivity rating which makes directly comparing their efficiency to other manufactures difficult. 
  • Denon AH-D2000 – I own these and have spent a lot of time with them. For a closed headphone they’re relatively accurate and neutral sounding. They’re more comfortable than the HD280 but they’re relatively heavy and have oddly shaped ear pads that, despite being very soft and well made, put some pressure on the ears that gets annoying after a while. They also have a relatively stiff cloth covered cord that doesn’t stay out of the way very well. They’re 25 ohms with a very efficient 106 dB sensitivity rating so they work with portable gear but readily reveal flaws in a portable player or lossy MP3 track. They’re also nearly twice the street price of the DT 770s at around $350.
  • Audio-Technica ATH-M50 – Like the Beyer DT 770 Pros, these $150 cans are designed for studio use. They have a fairly loyal following and I think sound much better than the HD280s above. The bad news is they suffer from what I call “Shallow Ear Cup Syndrome” . There’s almost no “depth” to the inside of the ear cup once the earpads compress against your head and this leaves the ear itself squished up against a thinly padded area that protects the driver. They also clamp with enough pressure to bother many people after a while. If you don’t need to wear them for long periods, or have tough ears, they’re worth checking out. I had a pair of these for a short time and discuss their sound later. They also have a relatively low impedance at 38 ohms and high sensitivity of 99 dB making them a decent match for portable listening without using an amp.
  • Sony MDR-V6 (and the very similar MDR-7506) – These are the cheapest of everything here—about $80 street price. And they sound better than the HD280s But they suffer from Shallow Ear Cup Syndrome worse than the M50s above and sound artificially bright. Some suggest different earpads help make them more comfortable. I own a pair and never use them anymore because I can’t stand to wear them for more than an hour or so. They’re rated at 63 ohms and 106 dB sensitivity and work reasonably well with portable gear.

DT 770 CONFUSION: Beyerdynamic’s product managers have not made things easy. While nearly all the above headphones come in just one version, the DT 770 can be had in no less than 7 current versions:

  • DT 770 Pro 80 - Straight cord 80 ohm
  • DT 770 Pro 250 - Coiled cord 250 ohm
  • DT 770 M - High Isolation 80 ohm
  • DT 770 Edition 32 - Premium home version 32 ohm (portable use)
  • DT 770 Edition 250 - Premium home version 250 ohm
  • DT 770 Edition 600 - Premium home version 600 ohm
  • DT 770 Manufaktur 32, 250, 600 - Custom ordered from factory ($300)

And to make it more confusing, the “Edition” versions are also randomly called “Premium” or simply “DT 770”. And, from what I gather, still different DT 770 versions were previously offered. Choice is good, but Beyer might have gone too far.

WHAT’S THE DIFFERENCE (revised 4/9)? BIG DISCLAIMER HERE… I’ve only heard the DT 770 Pro 80 personally. But here’s what I’ve been able to figure out from the Beyer website, other reviews, etc:

The Pro versions use a cheaper looking plastic ear cup, have slightly more clamping pressure, and are made to be more rugged. The street price on the home/premium/edition series are at least $50 more. The “Manufactur” version is a custom factory order program from Beyerdynamic. You can pick options, colors, etc. and in a month or two have exactly the headphones you want. Well, perhaps not exactly, as they don’t offer the 80 ohm version for the Manufactur program. There are four different impedances available and, while I’ve only heard the 80 ohm Pro version, here’s my best guess at some of the differences:

  • 32 Ohm – Designed for use with portable gear like an iPod. They supposedly sound similar to, but not as refined as the 80 ohm version reviewed here.
  • 80 Ohm – The apparent favorite among people who like more bass in their headphones. They’re only available in the pro versions.
  • 250 Ohm – In the pro version, these have a coiled cord and Beyer says they’re intended for mixing music. Reviews and published measurements indicate they have less bass than the 80 ohm and 32 ohm versions. They’re best used with a source intended to drive high impedance headphones.
  • 600 Ohm – These are only available in the home version and, based on published frequency response measurements, have much less bass than the other versions. They’re intended for use with a dedicated headphone amp or other source that can handle their unusually high impedance.

WHICH ONE IS BEST? I chose the Pro 80 for several reasons:

  • They’re what the store stocked, and hence what I listened to
  • I prefer a straight cable (the Pro 250 uses a coiled cable)
  • They’re still (barely) usable with portable gear (the 250 and 600 are not)
  • Beyer doesn’t make the home “Edition/Premium” version in 80 ohm
  • They’re $50+ cheaper and more rugged than the home “Edition/Premium” versions

But your needs might be different. If you like to crank it up and want to use an iPod or other portable, the 32 ohm Edition is the best choice. If you don’t like as much bass emphasis and have a good source that can handle high impedances, the Edition 600 would be a better choice.

THE CABLE: The 80 ohm versions have a fairly heavy duty and very long 3 meter cable (about 10 feet) wired to the left ear cup. That generous length is good if you like to wander around while wearing headphones, or bad if you have to find something to do with the excess. The 250 ohm version has a coiled cord that supposedly extends to the same length. I personally don’t like coiled cords as they tend to weigh more, pull on your head and are more microphonic (rub against things and create unwanted noises in the headphones). Because the straight cable is very slick and smooth the microphonics are fairly low. There’s a 3.5mm mini plug with a screw on 1/4” adapter. The strain reliefs at both the plug and headphones are very robust and well designed. I’m guessing the cable will last a long time which is a good thing as it doesn’t unplug from the ear cup. This is the sort of cable you can run over with your office chair multiple times and not hurt it.

CONSTRUCTION: The black steel head band seems very tough as do the hinges and adjusting clips. The ear pads and headband pad are easily replaceable. Beyer has black velour and black vinyl pads also available and you can find third party ear pads in real leather and with gel instead of foam.

ERGONOMICS & COMFORT: These are seriously comfortable headphones for three reasons. First they don’t suffer from the dreaded “Shallow Ear Cup Syndrome” like many of their peers. There’s actually room for a human size ear inside (see the pic further down). They’re also fairly light at only about 9 ounces. By comparison the Denon AH-D2000 is about 13 ounces. The ear cushions are covered with a soft velour that adds to the comfort as, unlike vinyl, there’s no sweat build up. While the Pro versions supposedly exert about one third more “clamping pressure” they still sit relatively lightly on your head—especially compared to the Sennheiser HD280. And if you want lighter still, you can bend the steel headband as desired. Finally the bottom of the headband is well padded with soft padding. So unlike say several models (AKG comes to mind), there are no pressure points on the top of your head. The only obvious downside is they don’t fold up for travel, and if you’re a DJ, you can’t swivel the earcups to use one at a time.

beyer line up

SOUND QUALITY (revised 4/23): After 48 hours of continuous break in (like breaking in a pair of new shoes it can, in theory, make a difference due to the initial stiffness of the moving diaphragms), I listened to the DT 770s mostly using a Benchmark DAC1 Pre as the source with a variety of material sourced from my PC and a Slim Devices Transporter. I also used an iPod Touch 3G and experimented with the output impedance on the DAC1. See the Tech Section for more. I compared the Beyer’s to three other popular closed back headphones--the Sennheiser HD280, Denon AH-D2000 and Audio-Technical ATH-M50. I also compared them to my Sennheiser HD590 open back headphones. I only had the ATH-M50s a short time so I had less experience with those.

TWO CAMPS: There are two fairly distinct groups of headphone users. A lot of people, especially those using portable players, are used to an “enhanced” sound with emphasized bass and highs. Many manufactures create a “V” shaped frequency response that’s lowest in the middle and rises strongly at both ends of the spectrum. When you select the “Rock” or “Pop” EQ preset on many portable players, you get just such a “V” shape. If that’s what you’re used to, accurate (i.e. relatively “flat”) headphones can sound boring—at least at first. So most portable players, iPods included, come with headphones with at least a somewhat “V shaped” response. But another group of listeners prefers more accuracy. More accurate headphones tend to be much less fatiguing over time and more natural. Plus, you can always add EQ to punch up the bass and/or highs if you want. Headphone manufactures are put in the difficult position of either offering different models to appeal to the different camps, or trying to compromise with a sound that tries to please both groups.

BEYER’s COMPROMISE? The DT 770 is allegedly for professional studio use and one might expect them to be relatively accurate. But I think Beyerdynamic recognizes many musicians, like much of the public, have come to expect enhanced bass and highs. Musicians buy iPods too. So it’s my guess they tweaked the 770 to appeal more to the tastes of pop/rock musicians while still trying to keep them accurate enough to allow use by recording engineers, etc. Sometimes compromise is a good solution, and sometimes it messes up a good thing.

BASS PERFORMANCE (low impedance source): Let’s say you take your boss to lunch in your polite four door sedan. Driving gently, there’s only muted burble from the exhaust hinting the car isn’t your mother’s Camry. Little does your boss know it has 370 horsepower and could pin him in his seat if you wanted to indulge your inner child with your right foot. That’s kind of how the bass is with the Beyers. Listening to a polite string quartet, there’s only a slight hint of the 770’s reputation for being “bass heavy” but play a hip hop track and hang onto your teeth. It’s Jekyll and Hyde. The Beyers have, by far, the most deep bass of all five cans I compared. In overall bass level with pop music, the rank is roughly DT770 > HD280 > ATH-M50 > D2000 > HD590.  But that’s a bit misleading as the DT770, for example, has less upper bass than the say the HD280. The HD280 makes that string quartet sound thick and muddy. The 770 has a very “punchy” bass that only really calls attention to itself when there’s significant deep bass in the music. Male vocals, for example, remain relatively accurate. The overall quality of the bass is better than I expected but it turns out that’s partly related to what they’re connected to.

BASS PART TWO (high impedance source): Based on what I’d read, I expected the bass of the DT 770 to be less controlled than it is. I wrote an article about how source impedance interacts with headphones to change their bass performance and wondered how the 770’s performance would change using a higher impedance source? The answer is quite a bit! Using one standard of 120 ohms, the formerly tight punchy bass of the 770 became more bloated and warm with less deep bass extension. The bass peak now reached up high enough to make male vocals sound thick. In short, the really deep rattle your teeth performance was diminished and replaced by a much less pleasant “boomy” bass peak. I wonder how many who dislike the 770’s bass heard them from a higher impedance source?

MIDRANGE (revised 4/23): The midrange of the DT770, like most closed cans, suffers a little. It’s a bit recessed and sounds a kind of “nasal” on some recordings. But these complaints are relative. The midrange is far more pleasant than the HD280 for example. The M50 sounds warmer (some might say too warm) through the midrange and the $350 Denons still have the best midrange of any closed cans I’ve heard. But the 770’s more laid back mids can sometimes be a benefit on poorly recorded, mastered, and/or ripped material. They’re less “in your face” than the more forward sounding Denons and hence are more forgiving. I rather like the 770’s midrange on a lot of male vocals while on female vocals they sound a bit “dark”. The midrange is somewhat similar to the open HD590 which is also a bit recessed (and some would say “dark”). If you’re a stickler for accurate midrange, I’d suggest the D2000s (or perhaps one of the other Denon models, Shure SRH 840, etc.).

HIGH FREQUENCIES: The 770’s highs are not as bright, revealing, or accurate as the more expensive Denon's. They’re closer to the M50s but the highs have a slight “enhanced edge” to them--not in a harsh way. It’s more an “artificial” sort of enhancement. To use a visual analogy, if you’ve ever seen an over sharpened video or picture, that’s similar to what I’m talking about. Initially it’s impressive, but over time, you realize it’s partly fake detail. The trick worked well enough to get me to take them home from the store. The 770’s highs, despite the “edge”, are still more forgiving of poor source material than the Denon's. And the 770’s highs sound better than the HD280’s. So pick your poison.

SOUNDSTAGE: To my surprise, the Beyers are far more open than the Denons. They also easily beat the M50 and HD280. The Beyer soundstage is deeper and wider and there’s more space between instruments--more like the open backed HD590. More of the ambience comes through. It’s another reason they initially impressed me.

EFFICIENCY (sensitivity): The DT 770s were the least efficient of the five headphones by several dB followed by the HD590s. The other three (Denon, M50 and HD280) were within a few dB of each other at the same volume setting.

PORTABLE COMPATIBILITY: With my iPod Touch 3G, the Beyers sound good on typical heavily compressed pop material but won’t get uber-loud. On really dynamic material, such as Flim & The BB’s Tricycle, the iPod (with its 7 ohm output impedance) clearly lacked bass punch compared to the Benchmark as the bass notes were clipped by the player at higher volumes. The Sansa Clip+ and Fuze did marginally better than the iPod but still fell short. Any portable device conforming to the new EU hearing protection regulations would be seriously underpowered. You can take these headphones with you (if you have room) and they’ll work OK from some portables. But to hear what they can really do, you need a low impedance source with enough beans. See the Tech Section for more details. The 32 ohm version of the 770 would be a much better match for portable use.

MAX SPL: The 770 will play cleanly well into hearing damage levels. If you feed them large amounts of deep bass at high levels, you’ll hear some audible distortion as the drivers run out of excursion (approach their limits of travel). You also get significant intermodulation distortion. This is true to varying degrees for nearly all dynamic headphones but the Beyers are perhaps a bit worse than average in their price class.

SOURCE MATERIAL: The DT 770 is more revealing than the HD280 but less so than the D2000. They’re similar to the M50. If you take them for a test drive I’d suggest using a wide variety of material and not just well recorded audiophile demo stuff. The highs, in particular, don’t get along with some source material.

OVERALL SOUND (updated 4/23): The nearly twice as expensive Denons are much more accurate than the 770s and still the best sounding closed cans I’ve heard on good source material. But the Denons ruthlessly reveal flaws and have a more “in your face” sound. The Beyers sit you a few rows further back in the venue, are more spacious, and try to offer some of the bass and treble “enhancement” many have grown used to. Some are likely to enjoy the punchy deep bass and enhanced highs. The 770's extra bass, at least driven from the Benchmark DAC1, doesn’t get in the way as often as I expected. But those who seek accuracy may find it excessive. The rest of the tonal balance isn’t perfect but, among closed comfortable cans, they’re better than average. The 770s, for example, blow away the well regarded HD280s in virtually every way. So if you like the HD280s, you’ll probably love the 770s.

beyer dt770 earcupTHE GOOD:

  • Very comfortable with room for an actual human ear inside
  • Spacious, engaging sound that’s “enhanced” in ways some might like
  • Serious deep bass without a lot of “muddy boom”
  • Laid back midrange is more forgiving on poor source material
  • Closed back to not disturb others and increase isolation
  • Solid rugged design
  • 80 ohm impedance allows some portable use at lower volumes

THE BAD:

  • Too much deep bass for those who prefer more accuracy
  • Boomy bass with high impedance sources
  • Midrange too laid back for some tastes
  • Somewhat uneven highs create a fake sense of detail
  • Plastic ear cups look kind of low-rent
  • Not the best match for some portable gear for louder levels

BOTTOM LINE: There are not many closed can I can wear for hours on end. The Denons flunk that test where the DT 770s pass it with comfy flying colors. And they’re also more forgiving of poor source material than the Denons. They sound similar to the Audio-Technica M50s but are much more comfortable. They’re well made and better than the HD280s in every way (except portability). With a low impedance source, the solid, deep, punchy bass can be fun with the right music. And some should enjoy the recessed midrange and “enhanced” highs. Beyer did a reasonably good job of trying to keep them from sounding “too boring” for those who are used to enhanced bass and highs. And they did it in a way that still leaves most of the frequency spectrum relatively unscathed. Some might find the DT 770 a perfect compromise but those seeking serious accuracy should look elsewhere.


TECH SECTION:

OBJECTIVE MEASUREMENTS: I’m big on hard numbers, but I don’t have a HATS (Head And Torso Simulator) or even a fake head with an embedded instrumentation grade microphone. So it would be difficult to make meaningful frequency response, distortion, etc. measurements as I would probably have a hard time even reproducing my own measurements let alone comparing those made by others. Headphones (along with speakers) vary widely in sound and the desired frequency response tends to be more subjective.

HEADROOM TO THE RESCUE (added 4/14): Being unable to make my own measurements, the best ones I know of are at HeadRoom. Here are 4 of the headphones discussed in this review:

HeadRoom DT770 vs HD280 vs ATH-M50 vs AH-D2000 Frequency Response

The HD280 has an upper bass dip where the others are flat or have a rise. And the M50 and D2000 have similar bass performance. If you scroll down within the comparison window, you can also view distortion, impedance, isolation and the square wave response. The HD280 has a severe dip in the extreme highs. But, otherwise, the four headphones are more similar than different. The Denons are the flattest overall. Looking at the 50 hz square wave response, the DT770s are the winner:

HeadRoom DT770 vs HD280 vs ATH-M50 vs AH-D2000 Square Wave Response

The “flatter tops” (in blue) of the DT 770 square wave response indicates flatter low frequency response than the M50 or Denon. And, while to many ears, the 770s have exaggerated bass, to the microphone, the bass is more accurate than the other three.

It’s worth noting not everyone’s head and ears are shaped the same as the artificial one used by HeadRoom. And that will change the results. And everyone has their own preferences and hears things differently. But the artificial head, microphone, and test gear at least attempt to provide a fair comparison between headphones. Assuming the measurements were made correctly it’s safe to assume, for example, the DT770 and HD280 have more bass than the other two based on the measurements. And indeed that’s what I heard and most others are likely to hear. So there is some valid correlation. But the graphs alone can’t always tell you which one will sound the best.

HEADPHONE IMPEDANCE: The DT 770 Pro 80 is rated at 80 ohms and measured somewhat above that. At the lowest point they’re about 83 ohms with a bass resonance peak (in free air, not on a head) at about 50 hz and nearly 110 ohms. They also exhibit the typical rise at the high end due to voice coil inductance. The phase shift (in white) is very linear and gradual making them an easy load:

Beyerdynamic DT 770 Pro 80 Impedance and Phase

IMPEDANCE INTERACTION: As discussed in the subjective review, the Beyer’s bass was well controlled when driven from the near zero ohm output of the Benchmark DAC1 Pre. The IEC standard from 1996 (before portable digital players became seriously popular) specified a 120 ohm output impedance. I put 120 ohms in series with the DT 770s on the Benchmark and the sound changed significantly. 120 ohms is enough to seriously degrade the electrical damping of the driver which raises the Q. That, in turn, causes the bass to peak more, at a higher frequency and roll off sooner (less deep bass). The bass also ends up less damped overall and hence less controlled and less “tight”. That’s exactly what I heard. The bass got more muddy, boomy and didn’t reach as low. This is discussed further in the Impedance Article.

The 120 ohm source also creates a frequency response error in the signal driving the headphones. Here’s the signal from the Benchmark at zero ohms in blue, and from a 120 ohm source in yellow. There’s almost 1.5 dB of variation. Note this is the electrical input level, not the acoustic output of the headphones but it will still cause a similar response variation in the acoustic output compared to a zero ohm source (in addition the the headphone’s regular response variations):

Beyerdynamic DT 770 Pro 80 Zero Ohm Source (blue) 120 Ohm IEC Source (yellow)

SOURCE IMPEDANCE SUGGESTIONS:  If you’re not a fan of bloated bass, the lower the output impedance the better with the Beyer’s. This increases the electrical damping and tames their bass boost. When I tried them with my iPod Touch 3G, which is around 7 ohms, the bass lost some punch (also possibly due to the peaks being clipped). And it gets worse as you go up in impedance from there.

SOURCE DRIVE LEVEL: On highly compressed pop music I cranked up the Beyer’s to about as loud as I would normally go and measured the signal on a scope while listening. The peaks were hitting +/- 400 mV which is only about 280 mV RMS. Even most portable devices can manage that. But, playing material with a wider dynamic range at loud levels, I needed up to 3.5 volts peak-to-peak which is about 1.2 volts RMS. The $20 FiiO E5 can manage about 1.3 volts RMS and it got plenty loud even on highly dynamic material but the Beyer’s likely deserve better than the E5.

To translate this into power output, at a more-than-needed 1.5 volts RMS, that’s 70 mW into 32 ohms, 28 mW into 80 ohms, and 9 mW into 250 ohms. Sources rated somewhere in that range should have enough cajones to do the job without clipping or running out of gain.

POWER HANDLING: Beyer rates the 770 Pro at 100 mW. For the 80 ohm version, that’s 2.8 volts RMS and would be uber-loud. The driver, however, runs out of excursion well before that at low frequencies.

SOURCE QUALITY: As mentioned in the review, while not the most revealing cans on the planet, the Beyer’s are well above average. So poor quality sources, poor recordings, lower bit rate MP3’s, poorly designed amps/DACs, etc. are revealed for what they are. The Beyer’s sometimes take a bit of the edge off, but they don’t “gloss over” defects the way a lot cheaper headphones do.

TECH SUMMARY: There were no real technical surprises except I wonder how much of the 770 Pro 80’s reputation for unpleasant bass is from people using higher impedance sources?

March 17, 2011

DAC Listening Challenge Results

NuForce uDAC-2 vs Behringer UCA202 SoundA BIG THANK YOU! I appreciate all the responses to both listening tests. I know listening for subtle differences requires some effort and I want to thank everyone for taking the time and sharing their results.

NOTE TO HEAD-FI MEMBERS (revised 3/28):  Unfortunately, before I posted this article, the administrators at Head-Fi.org started censoring links to this blog. So if you had some trouble finding the results here, I’m sorry. The links were apparently not a problem when I was posting at Head-Fi on several other products and topics. The censorship only started after a paying Head-Fi sponsors (NuForce) publically complained about my review of their product. Head-Fi has gone so far as even deleting posts from other Head-Fi members that reference this blog.

BACKGROUND: NuForce responded to their uDAC-2 measuring poorly by saying it was designed to sound good even if several measurements are notably bad. So I came up with the best way I knew how to judge the NuForce purely on sound quality. I also thought it would be interesting to compare the $29 Behringer UCA202 with a high-end product like the Benchmark DAC1 Pre. The DACs were recorded playing real music under as realistic and similar conditions as possible. Anyone could download the recordings and compare them without knowing which was which—sort of like a “brown bag” wine tasting. Would people like the $3 wine better than the $30 wine? I thought it would be fun to find out!

TEST METHODS: Not having conducted a public web-based listening test before, this was something of a learning experience for me. For the first round, to make the comparison as fair as possible, I used the line outputs of all three DACs . The second round of tests used the headphone outputs driving real headphones and included a modified version of the Behringer UCA202. The modifications were made using a few dollars worth of parts and improve the headphone output of the Behringer. The files were given names of US Presidents in the first test and common trees in the second. The original listening tests, and all the details, can be found here:

A NOTE ABOUT “MASKING” (revised 3/17): Some have questioned the validity of the tests because, as NuForce put it, the output of the DACs have been “re-digitized”.  And some also argue the equipment used to play back the files may not be high enough quality. The concern is these could mask the differences between the DACs. In reality this should not be much of an issue using playback hardware with reasonable fidelity. Here’s why:

  • These trials are only about the differences between the files, not the absolute accuracy. It’s like shopping online for a shirt. You might look at 3 different blue shirts at the same online retailer. Even if your computer display can’t convey every color and detail perfectly, you can still easily tell most of the differences between the shirts if they’re good pictures, all taken the same way, and your have at least a reasonably decent display. For example if your computer distorts the exact shade of blue, you should still be able to tell which of the 3 shirts is the darkest blue. The same is true in comparing the sound files. Even if your headphones might exaggerate the bass, you can still tell differences between the bass in the various files. This “difference effect” has been well documented in research.
  • As to “re-digitized”, the Benchmark ADC1 used to convert the output of the DACs to a CD quality file is already better than most studio gear used to record the music we all listen to. Studio A/D converters typically cost about $100 – $400 per channel of conversion. The ADC1 is a reference-grade A/D with a $900 per channel price tag. It generally has better specifications than nearly all studio gear. If the less expensive studio gear is good to enough to capture the subtle difference between say a high-end Steinway and a high-end Yamaha grand piano, the ADC1 should also capture audible differences between DACs.
  • Most recordings have already been through many more steps and kinds of digital processing yet are still very revealing of subtle differences. So adding one more relatively “pure” step isn’t going to make much difference. The recordings used in these tests are unusually pure and transparent.

PARTICIPANTS: This was a very informal survey. A total of 20 listeners picked their favorite (and sometimes least favorite) tracks. Many ranked them top to bottom. Some only participated in one of the tests and/or one of the songs. Some used ABX and some listened conventionally. More participated in the first (line out) trial than the second (headphone out) trial.

SCORING: Not everyone participated in every trial, some devices were offered in more trials than others, and different listeners provided different sorts of “votes”. Someone who’s better with statistical analysis than I am is reviewing the raw data. But, for now, to summarize the results here’s my attempt at a rough analysis:

  • Top choices scored 2 points
  • Second (“Runner Up”) choices scored 1 point
  • Least favorite (worst) scored –1 point (several indicated only their least favorite and not a favorite)
  • The results are summarized by category (line output vs headphone output)

RESULTS (corrected 2 minor errors 3/16): Here are the total points scored for each downloaded file ranked from most favored (highest score) to least favored using the above scoring method:

Taft - Benchmark Line Out Brick House 9
Harrison - NuForce Line Out Brick House 8
Juniper - Benchmark CX300 Brick House 7
Wilson - Behringer Line Out Brick House 6
Jefferson - Reference CD Just Dance 4
Lincoln - Benchmark Line Out Just Dance 3
Oak - Behringer CX300 Brick House 2
Spruce - Mod UCA202 CX300 Brick House 2
Acorn - Benchmark CX300 Tis of Thee 2
Maple - Mod UCA202 CX300 Tis of Thee 3
Monroe - Reference CD Brick House (!) 1
Hawthorn - NuForce CX300 Tis of Thee 1
Fig - Behringer CX300 Tis of Thee 1
Jackson - NuForce Line Out Just Dance 0
Cypress - NuForce CX300 Brick House 0
Adams - Behringer Line Out Just Dance -2
Pine - NuForce UE SF5 Tis of Thee -4
   
Total Benchmark Line Out 12
Total Benchmark Headphone Out 9
Total NuForce Line Out 8
Total Behringer Line Out 4
Total Modified Behringer Headphone Out 5
Total Behringer Headphone Out 3
Total NuForce Headphone Out CX300 1
Total NuForce Headphone Out Ultimate Ears SF5 -4

COMMENTS ON RESULTS (edited 3/16): Even with the small sample size a few things were fairly clear:

  • The NuForce headphone output with the Ultimate Ears SuperFi 5 Pro headphones was an obvious “fail”. Several expressed a clear dislike of “Pine” and nobody favored it. I think this is due to the large frequency response variations caused by the relatively high output impedance of the NuForce uDAC-2 with these headphones. Several commented the high frequencies were rolled off or “dull”. This consistent with what the measurements would predict.
  • The reference CD tracks didn’t score as well as expected. For “Monroe” it’s likely because the reference track stood out as being different. Early in the forums it was labeled “by far the worst”. And that public comment likely tainted others into also hearing it as the “worst”. Six people rated it as the worst of that group of tracks. I personally think this is a good example of subjective bias--the original piece of music was deemed the “worst”! The Just Dance reference faired much better with a positive score of 4 likely because it wasn’t labeled publically as being bad. And much fewer even bothered to vote on Just Dance so the score of 4 was close to “perfect”.
  • In terms of total votes, the Benchmark stood out as a clear favorite. It was included mainly as a “reference” and not necessarily as fair competition to the other two, much cheaper, DACs. But, some might have expected all the DACs to sound roughly similar via the line outputs based on the measurements. The could be many reasons why this didn’t happen. The most likely is the same “peer bias” that caused a reference track to be rated poorly (see above). But it’s also possible the cheaper DACs do have audible problems.
  • The NuForce did much better via the line outs than the headphone outputs. Given the output impedance problem and higher distortion of the headphone output, this is what the measurements would predict.
  • Any channel balance error was corrected in all tests. So this removed the audible imbalance of the NuForce helping it score better.
  • NuForce’s claim of “better sound” doesn’t seem to be true when using the headphones. It faired poorly even with the more common and “impedance friendly” Sennheiser CX300’s.

Trying to draw other conclusions is a bit more difficult, but some things worth noting:

  • Nearly half the people couldn’t hear any differences at all.
  • The Just Dance track had a lot of inherent distortion that made hearing differences difficult for most. So those results are likely less valid (and there are fewer of them).
  • A number of votes were somewhat randomly distributed between all the middle scoring test files. This may suggest more guessing rather than clear preferences.
  • The line outputs of  the $29 Behringer didn’t do as well as its measurements would suggest. The NuForce and Benchmark were preferred. This could be “peer bias”.
  • The differences among the individual files were more obvious via the headphone outputs rather than the line outputs. This is to be expected due to the impedance interaction.
  • For those who argue tests like this mask differences, it’s interesting to note the test may not have masked any subtle advantages of the high-end Benchmark. This is despite the fact that few listened to the files on similarly high-end gear.

BOTTOM LINE: I think the most obvious thing is roughly half the participants couldn’t tell any difference and another group of the “middle scoring” results is almost random with no clear preferences. Comparing the headphone outputs, however, the differences seemed more obvious. And they were very obvious when using balanced armature headphones (the SuperFi’s) on the NuForce. It’s clear the frequency response variations created by the NuForce’s relatively high output impedance creates audible problems.

It’s apparent subjective opinions of differences are easily swayed by public comments. The reference track from the CD of Brick House was strongly disliked by six listeners. They heard something different and were easily swayed into thinking different was worse based on previous public comments. It just took one person to say something negative and several others followed. This is exactly the sort of subjective bias that affects the majority of online subjective reviews in forums. As further proof, others using ABX (which is blind) voted the same track their favorite.

There’s also significant evidence the NuForce uDAC-2 does not sound better—at least using 2 different types of headphones—as NuForce claims it does. But it seemed to do a respectable job via the line outputs—at least with the channel balance error removed.

Overall I think this has been an interesting experiment. I learned a lot about how to run (and not run!) a listening test. And there were some fairly clear results—some expected and some not. If there’s sufficient interest, I may try to build on what I’ve learned here and conduct future listening tests?

COMMENTS WELCOME: Please feel free to add comments to the end of this article on the results—especially if you participated. I’d also like to know how many are interested in future listening tests? They’re a fair amount of work to put together and only really valid if you get a reasonable number of votes. So it’s something I only want to do if there’s enough interest. Please feel free to make suggestions, etc?


TECH SECTION:

IMPROVING THE RESULTS: I realize this wasn’t the best run study. It was more an informal experiment than anything. I think a larger scale listening test, that had more uniformity, would be needed to verify some of the closer results. If I do this again, I’ll research better methods, and I also welcome input from others with experience in this area?

DIFFERENT LOADS: It’s been suggested it would be also useful to include high impedance headphones and, except for the complexity that adds to the mix, I agree that would be interesting. I chose low impedance headphones as they’re, by far, the most popular—especially for use with a portable entry level DAC.

PASSWORD: An encrypted 7-Zip file was included with the file descriptions to prevent me from cheating or changing anything after people’s votes. The password is:

CY&YUMN5cZ9x2X8BhNj2t

PINE EXPLAINED: The graph below, in blue, shows the frequency response of the NuForce uDAC-2 using the Ultimate Ears headphones used in the trial:

NuForce uDAC-2 -3 dBFS Swept Frequency 15 Ohms (yellow) UE SuperFi 5's (blue) (Ref ~400 mV)

The 4+ dB of response variation seen above is caused by the relatively high output impedance of the uDAC-2 interacting with the SuperFi headphones (typical of balanced armature designs from many manufactures). For more on this see Headphone & Amp Impedance.

OTHER DETAILS: The other details of how the test was run can be found in the original articles:

NuForce uDAC-2 Listening Test

DAC Listening Sequel

February 2, 2011

RightMark Audio Analyzer

ROSE COLORED GLASSES (updated 3/3/11): RightMark Audio Analyzer ( aka RMAA from rightmark.org ), in some ways, is impressive software for the price (free!). A lot of people, websites, and organizations, are using RMAA to publish audio measurements. It's a useful tool when used properly and the user is aware of its many limitations. Considering it’s available for free the developer should be congratulated. But RMAA has many limitations and potential problems. And a lot of people use it incorrectly and, often unknowingly, publish misleading RMAA results.

THE SHORT VERSION FOR NON-RMAA USERS:  You can rarely compare one person's RMAA results to another person's RMAA results. This is especially true for evaluating anything other than a PC audio interface (aka “soundcard”). It's very easy to have a better piece of gear actually show worse results on RMAA because of different test conditions, settings, etc. So don't put too much trust in any RMAA results unless you made them yourself under carefully controlled conditions. Unless you know the absolute levels used, load applied, PC sound hardware used, and settings of the PC and device being tested, you simply can’t trust the results. Changing any of these things can cause a greater change in the results than testing a different audio device.

SETUP IS EVERYTHING: With RMAA there are many things to get wrong. Examples include hidden mixer controls, sound processing settings in the PC sound driver, sample rate conversion by the operating system or sound hardware, level settings—both at the PC and for the device being tested, loading, ground loops, cabling, etc. And if you’re doing loopback testing, you have no way of knowing what performance is related to the A/D, the D/A, or both. Even having a cell phone nearby can corrupt RMAA measurements with no error indication at all—just erroneous measurements. RMAA testing has lots of “blind” aspects that users are often unaware of. It has very poor internal consistency and error checking. Instead of giving a warning something might be wrong, it more often just produces false results.

15 IMPORTANT THINGS RMAA DOES NOT MEASURE (revised April 2011): There are many things that can make an audible difference in sound quality RMAA does not test for. Many are difficult or impossible to measurel with a soundcard-based test setup even using other software, specialized hardware, etc. They include:

  • MAXIMUM CLEAN OUTPUT LEVEL - RMAA has no concept of absolute levels. It can't measure voltages, power outputs, etc. So you have no way of knowing, for example, how loud that portable player or headphone amp can play without obvious distortion. You also have no way of knowing what level you’re testing it at. You could be measuring a device right on the edge of clipping. Or the level might be so low you’re measuring more noise than distortion. RMAA can’t tell you if you’re testing at 2 volts of audio signal or 0.02 volts. 
  • OUTPUT IMPEDANCE - Anything designed for driving headphones including PC's, portable players, headphone amps, USB headphone DACs, audio interfaces, pro gear, etc. has an output impedance. It varies widely and can have a huge impact on the sound quality--especially with certain types of headphones (see Headphones Amp Impedance). RMAA does not measure output impedance at all. This is a huge limitation as it’s one of the biggest factors in the sound quality of headphone sources.
  • DAC LINEARITY – Some DACs exhibit considerable non-linearity at low levels. RMAA doesn’t measure low level linearity, and because it can’t measure absolute levels, everything in RMAA is only relative to an unknown reference. 
  • SQUARE WAVE PERFORMANCE – A 1 Khz square wave reveals a lot of information about analog and digital audio components such as stability, bandwidth, rise time, compensation, and for digital devices, the type of digital filtering used. One or more of these can have an audible impact on device performance. But RMAA doesn't have a square wave test. And even if you use oscilloscope software for a soundcard, it will only measure up to the cut off frequency of the PC’s sound interface. The bandwidth of PC sound hardware is usually too low to see ultrasonic/RF, ringing, oscillations, switching power supply noise, Class-D amp noise, or instabilities. The very limited bandwidth is like looking through rose colored classes. The result can easily appear much better than the real performance.
  • JITTER - Jitter has been proven to be audible in some circumstances. RMAA has no jitter test.
  • VARIOUS TWIN TONE TESTS – RMAA has a single IMD test that’s similar to the SMPTE test but it’s not clear if it follows the SMPTE standard. And there’s no capability for other important twin tone tests such as the popular CCIF 19 Khz/20 Khz which is very revealing of high frequency performance. From what I’ve seen, RMAA doesn’t even calculate SMPTE distortion properly. The results are often artificially optimistic.
  • THD20 – It’s very useful to measure THD at 20 Khz for a variety of reasons. It’s a good indicator of stability, feedback loop design, and more. To perform this measurement properly, you should be able to measure the first 3 harmonics at 40 Khz, 60 Khz and 80 Khz. Even if you get RMAA to work with at a 192 Khz sampling rate, it still can’t perform a proper THD20 measurement. And it has no option to adjust the measurement bandwidth of its THD sweeps.
  • LOW FREQUENCY DISTORTION – It’s also useful to measure THD+N at very low frequencies such as 5 to 10 hz to expose thermal modulation effects, power supply issues, and more. RMAA doesn’t allow arbitrary low frequency testing.
  • POWER vs THD: RMAA cannot perform the classic measurement of output power versus THD. This is a standard benchmark amplifier test that’s widely performed as it’s very revealing of the amplifier’s behavior at different levels. RMAA has no way to perform this critical measurement nor am I aware of any other software that offers this sort of measurement with a sound card.
  • SELECTIVE SPECTRUM TESTS: It’s very useful to perform different spectrum testing with various input signals. RMAA doesn’t give you any control over its signal generator or its analysis. You cannot control the FFT points, averaging, weighting, filtering, etc. This is very limiting for audio testing. You’re stuck with the ‘canned’ results RMAA provides and many of those don’t follow industry standards (or it’s not clear how they’re calculated).
  • RESIDUAL ANALYSIS – Analyzing the residual distortion products can be very revealing—for example it can determine if an amplifier suffers excessive crossover distortion. RMAA has no way to do this.
  • REAL TIME RESULTS - Many problems may only show up briefly or intermittently. And many adjustments (like finding the clipping point) are best performed in real time. For example, maximum output is commonly defined as 1% THD. So with a real time audio analyzer you just raise the level until it reads, in real time, 1% distortion. With RMAA it takes a long time to run a test sequence which makes if very difficult to see what affect adjustments have on the results. It might take 10 or more passes to set the level correctly for even close to 1% THD. And if you have a bad cable, or some other intermittent problem such as outside noise or interference during the test, you will likely never know because you just get a dump of results at the end. I’ve had a cell phone randomly mess up RMAA measurements in ways that just made the product being tested look bad with no other indication the results were invalid.
  • SLEW RATE – RMAA cannot measure slew rate which, especially for audio power amplifiers, can be an important measurement. PC sound hardware is bandwidth limited making these sorts of measurements difficult or impossible with any PC software.
  • DELAY & LATENCY – RMAA cannot measure delay and latency for digital hardware.  This can be critical for live audio recording, studio work, gaming, and for lip sync with video.
  • HARDWARE LIMITATIONS – Nearly all PC sound hardware has some severe limitations that restrict RMAA and any other software. For example bridged and certain other amplifiers cannot have any of their output terminals grounded or connected together. But typical PC sound interfaces have common grounded inputs. To make matters worse, these grounds often go back to the PC’s  AC power grounds which creates even more problems. Some of these issues can seriously harm the device being tested and/or your PC’s sound interface or even the PC itself. USB and Firewire sound hardware is also grounded back to the PC. Most of these devices can only work with up to a few volts of input. Higher levels will cause overload and even damage the hardware. As mentioned earlier, the bandwidth, best case, is usually limited to about 90 Khz and more often only about 20 Khz which precludes many measurements and often hides other problems. There are also inherent issues with bit resolution, gain matching, channel balance matching, absolute levels, and more.

 

ALL THE NASTY DETAILS (non-geeks probably want to skip this section):

If you're not familiar with RMAA, here’s a typical screenshot. Note the 3rd column in the results labeled "RMAA Problems" and the highly questionable numbers. More on that later (click for larger):
 

 
PC HARDWARE ISSUES: RMAA results are only as good as your soundcard and how RMAA is used. Lots of people are running RMAA on whatever sound hardware their computer came with. And that's often a serious limitation. Built-in sound hardware is prone to all sorts of problems--especially the A/D section which is often only used for digitizing voice these days so manufactures don't put much effort into making the built-in A/D signal path perform well.

Nearly all built-in sound hardware amounts to little more than a cheap CODEC chip stuffed onto a crowded and electrically very noisy motherboard. Modern CPUs create spikes of current that can exceed 100 amps. That generates a lot of electrical noise and it's hard to keep it all out of the nearby audio circuitry where it can confuse RMAA measurements.

Even with external sound hardware, other applications or your operating system can play random sounds while you're running the test (like when you get new email). Even CPU loading and hard drive activity can affect the noise floor of built-in audio hardware--especially the line inputs and A/D. And while RMAA should log an error if extra sounds are played during the analysis, it just corrupts the results in a non-obvious way (more on that later).

Some PC's (especially laptops) don't even have a line level input. They only have a microphone input. Trying to use a mic input for line level audio is a disaster and will yield very misleading results. Yet some still do it and publish the results with no indication of how they obtained the results and hence no clue the results might be invalid.

Lots of PC's and sound cards perform various audio processing in the digital domain. It's sometimes impossible to turn all this processing off hence many PC sound devices are not "bit accurate". Some perform internal sample rate conversion, for example, regardless of how they're configured (Creative's Sound Blaster and X-Fi devices, for example, infamously perform conversions). Others attempt to optimize or enhance the audio in various ways. This internal processing affects the accuracy of RMAA in unpredictable ways.

Finally, even lots of outboard PC sound devices have problems--especially with their A/D hardwaer. Some have microphone preamps always in the circuit which seriously compromises their performance for line levels. Others have very inexpensive A/D chips and circuits in them. Some are not bit accurate, have noisy USB-derived power supplies, and/or problematic level controls, etc.

GAIN MATCHING: With many PC sound interfaces, the mixer settings operate in the digital domain. And when you lower the level settings you may lose digital bit resolution. The effective number of bits might be 16 bits with the controls maxed, but if you have to reduce the gain, you may get less than 16 bit resolution. At some point this can become the limiting factor in the test setup (insufficient ENOB). This is especially critical for measuring very low levels like noise, distortion, or crosstalk (channel separation). 

And changing levels is made more difficult because PC sound hardware levels controls are rarely calibrated in known increments (like dB). And even the few I have seen that are calibrated in dB (in software) don’t “track” well with the actual hardware settings. For example, some steps are close to 1 dB but near the end of the range can be as high as 3 dB. The manufacture never intended for their hardware to be used as a lab instrument so they probably didn’t much care about such accuracy as it’s not a big deal for 99% of applications.

So, ultimately, you have to verify levels with an external meter and not just any meter (more on that later). Any changes in the controls or settings usually requires establishing the levels with the external meter all over again.  This can get very tedious when dealing with audio signals that often vary over a wide range during testing. And if you don’t re-adjust the level controls to operate the A/D near full scale, you can lose lots of resolution and dynamic range making the device being tested seem to perform much worse than it really does.

Level setting also depends on the topology of your PC sound hardware. Some of the external USB or Firewire devices have their own physical pots for levels controls. I’ve never seen these calibrated in any meaningful way. They typically cover an extremely wide range of gain (i.e. 65 dB for the E-mu products). These controls, because of their wide range, are extremely “touchy” and it’s just about impossible to make precise small adjustments or get the two channels precisely balanced. And the settings can even change on their own which can really throw off RMAA if you don’t realize it happened during a test.

Better devices use programmable gain amplifiers (PGAs) ahead of the A/D converters and control these PGAs from software. This can be a better solution, but only if you’re given calibrated control in sufficiently fine steps, and the PGAs themselves are accurate. I’ve yet to find any interfaces that meet these criteria. Many of the less expensive ones are not terribly uniform in their steps over their full range. PC sound hardware is not designed with absolute precise values in mind. But that’s exactly what’s needed for a good measurement set up.

Some PC sound hardware uses a fixed gain amplifier and some sort of “digital pot”. These are often less accurate than the PGAs above. And often have fewer, and more coarse, non-linear steps. And the high fixed gain amps can have more noise issues. Digital pots often use a linear rather than logarithmic scale. This can make the step size rather random.

Some PC sound hardware lacks any sort of attenuator or PGA on the inputs. They simply design the circuit so the maximum expected input corresponds to roughly 0 dBFS in the A/D and then do everything in software from there. While this is fine for a person’s voice for Skype, it’s far from ideal for measurement purposes. As the signal level is reduced, the effective resolution of the A/D goes down with it. If your source happens to have a maximum level close to the maximum of the sound hardware input you’re in fairly good shape (say 2 Vrms the Redbook standard for digital audio). But if you’re trying to measure a portable player with a maximum output of 0.4 volts you just lost 14 dB of dynamic range from your measurements.

The gain increments are often too large, variable and/or unknown. For example, if you want to remove a 0.7 dB channel balance error, but only have 2 dB steps to work with, you’re stuck. And the steps often don’t correspond to any convenient number of dB because they’re derived from whatever the hardware supports. Most electronic volume control chips, for example, only have 64, 128 or 256 steps. So what you get is the full range divided by 64 uneven increments—whatever that happens to work out to.

Proper audio analyzers have calibrated analog stepped attenuation built into the input and output circuitry so regardless of the signal level they maintain the resolution of the D/A and/or A/D—to always be within a few dB of full scale. You can also set levels in their software to a resolution of at least 0.01 dB. I’m not aware of any PC audio devices that can do anything even close.

NO ABSOLUTE LEVELS: Even if you run RMAA on a decent PC audio interface with good specs, you still have no idea of the absolute levels. Are you testing too low? Too high? Near the device’s clipping point? This also prevents knowing the true output capability of the device being tested (i.e. dBu, dBv, dBm, volts, watts, milliwatts, etc.). It also means the channels may not be properly balanced which affects RMAA's measurements—the RMAA calibration routine compensates for channel imbalance in your device when you really want to be measuring that imbalance. The actual levels can make a BIG difference in the RMAA numbers you get. Testing a device at 1 volt versus 0.1 volts will yield a 20 dB difference in noise and distortion levels. That’s huge.

AN EXAMPLE: Joe posts his RMAA results for Player A and tested at a relatively low level. This will yield noise numbers significantly worse than the device is capable of. Bill posts his results for Player B but, unknown to Bill, the level was set too high and it was on the edge of clipping. So Bill’s noise numbers look really good but his distortion numbers look bad because the player was on the verge of overload. Hopefully you can see how it’s impossible to draw many valid conclusions from the published results because the tests used very different levels. But that doesn’t show up anywhere in the RMAA report or results. Player B might really be the better player, but Bill had no idea it was starting to clip so the high distortion results make it look like the inferior product.

REFERENCES ARE CRITICAL: The proper way to do audio measurements is to use standard reference values. For example  0 dBu is 0.775 volts RMS and a standard for consumer audio. Without knowing the actual levels of the signals being measured, you have no way to use proper references or even know what the levels are so you can test the next device, or even retest the current device, at the same level. And even if you use an external meter (most of which are designed for 60hz AC power and not even close to accurate across the audio band) you're still likely stuck with the un-calibrated mixer/level controls of your PC audio device. So the instant you change any of the controls, your meter measurements are rendered useless. Without absolute levels, and a known reference, it’s usually impossible to duplicate the results.

NO STANDARDS: RMAA isn’t very clear about what standards it conforms to, if any. Did the guy who wrote the software build in A Weighting to the noise measurement? If you use a higher sampling rate, are the THD values only calculated over the audio spectrum or do they also include ultrasonic noise and distortion? How is the IMD value calculated? There are lots of unanswered questions about how RMAA works internally and it makes comparing RMAA data to real measurements that conform to accepted standards difficult or impossible.

TEST LOADS ARE CRITICAL: Loading is often overlooked. For example, portable MP3 players often measure far better unloaded driving just the line input of your PC audio device than when properly loaded with real headphones or a suitable test load. Nearly all the RMAA measurements I see published never mention anything about what the device was loaded with so I'm guessing it's usually nothing. Unloaded tests are often meaningless. Some measurements, like distortion and crosstalk, are hugely affected by the load. And if you're trying to measure a line level output (rather than a headphone output) most PC sound devices don't have a well defined input impedance. Here again, this makes repeatability, comparing, and verifying results difficult—especially at higher levels of performance.

USING HEADPHONES AS THE LOAD: This can be good and bad. The good news is it’s a realistic non-linear load--especially if they're the headphones the user intends to pair with the device. The bad news is someone else needs the exact same headphones (which they likely don't have) to conduct a fair test that's valid for comparison. And for many devices the headphones will alter the measured frequency response in ways more related to the headphones than the device (see the example below). And, unknown to many, headphones also act as microphones. So, for example, they pick up background noise in your room and this can raise the noise floor measured by RMAA making for misleading noise and distortion measurements. A proper resistive test load is better if you want reproducible results that can be more fairly compared with other results. It also avoids the "microphone effect".

Here are the RMAA results using a 15 ohm resistive load, and 2 different kinds of headphones--the UE SuperFi 5 Pro's have a really wide impedance swing from about 10 ohms to 85 ohms and a pair of Sony MDR-EX76's that only vary by a few ohms at higher frequencies. The levels were within 2dB of clipping at the worst case frequency using a Benchmark ADC1:


The best numbers are with the 15 ohm resistive load. When using the Sony's the THD more than triples and the noise floor rises up a few dB. And what's with those SuperFi 5 numbers? Does the otherwise low distortion source suddenly produce 4% THD? No, but RMAA makes you believe it does! What likely happened is the signal was clipped because of the frequency response swings. But you have no way of directly knowing that, and from the average person’s point of view, the levels were properly calibrated using RMAA’s calibration routine. Here's a "zoomed in" graph showing what effect the Sony headphones have on the frequency response:


So if Joe tested with a resistive 15 ohm load you'd see the white graph in his results. If Bill used his Sony headphones you'd see the green graph published and probably think less of the product being tested. And if Bill used his SuperFi 's here's what you'd see (in green)—a whopping 15 dB of response deviation:


Here's a Sansa e260 MP3 player with no load and a 15 ohm load:


Notice it has more than 8 times as much THD, more than 4 times as much IMD, the crosstalk is far worse, and the frequency response much worse when loaded. Here's the frequency response difference:


There's likely a coupling capacitor in the output amplifier of the Sansa. And, unloaded it has little effect, but with a typical load it's down -3dB at 40 hz. There's also a weird rise at high frequencies likely caused by marginal feedback design in the amplifier. Here's the swept IMD vs frequency:


Unloaded the distortion is down below -72dB worst case, but loaded it's more like -57dB which is likely audible. With no load, as many use RMAA, the Sansa rivals some of the better players out there. Loaded, however, it's closer to the bottom of the pack. It makes a huge difference! These are typical examples. It makes a very big real world difference.

GROUNDING: Even an external PC audio interface is grounded to your electrically noisy PC via the USB/Firewire connection. And the inputs and outputs share a common ground. These both can create potential ground loops, noise sources, or worse, with whatever device you're trying to test. If the device has bridged outputs it may harm the device to ground one side of the output. And many even conventional (non bridged) amplifiers can oscillate or even self destruct if you connect their input and output grounds together externally. You can use transformers to isolate the device, but they introduce lots non-linear distortion on their own so you end up measuring mostly the transformer. Proper audio analyzers have differential and/or floating inputs and outputs that are not only isolated from ground but from each other.

HIDDEN GROUND PROBLEMS: Ground loops, especially where PCs are involved, can create noise in non-obvious ways. It’s very likely a lot of RMAA results are badly corrupted by ground issues but those making the measurements are likely totally unaware of the problem. They just see results that are much worse than what the device is capable of when it’s used normally. You have to remember RMAA was mainly designed to test sound cards in loop back where grounding typically isn’t an issue.

OVERLOAD: If you're trying to measure gear intended to drive speakers, it's likely to overload and possibly damage your PC hardware if you directly connect the output to the line input. Such inputs are usually only designed to handle a few volts maximum and some even less.  But speakers usually require 10+ volts. So you have to make an external divider network which is another potential source of problems and inaccuracy. It also makes it more difficult to compare RMAA measurements made by others unless they’re using the same divider network.  And there’s still the grounding problem mentioned above which can literally damage the gear you're trying to test as well as your PC hardware.

SAMPLING RATE: RMAA’s measurement range is limited by the sampling rate of your audio hardware and what the drivers support. Often it doesn't work (or work right) above a 44 or 48 Khz sampling rate. This restricts the bandwidth to around 20 Khz. Many devices may have instabilities that show up at ultrasonic or even RF frequencies. RMAA is usually completely “blind” to these problems. See Testing Methods for an interesting real world example. And Class-D amplifiers and/or switching power supplies are increasingly being used used in everything from MP3 players to A/V receivers. These switch at high frequencies and often create large amounts of noise that’s entirely missed by typical PC audio inputs. Even if you don't think it's audible, the amount of out-of-band noise can still be a useful indication of the quality of the device being tested. And stability issues often are audible.

BIT DEPTH AND NOISE FLOOR: RMAA is typically limited to 16 bits resolution by either the PC audio hardware and/or driver limitations. But you often can't use the full 16 bit range because of gain limitations (see GAIN MATCHING above). So, in reality, you might end up with less (or similar) useable dynamic range than the device you’re trying to test.  Ideally the test set up should have far better dynamic range than whatever you’re testing. Even if you're lucky enough to get RMAA to work correctly with 24 bit drivers, you won't get anywhere near 24 bit performance from typical 24/96 or 24/192 PC sound hardware--you're lucky to get 17 or 18 bits of effective resolution (ENOB) due to the noise floor, power supply noise, grounding issues, typical A/D and D/A converters, etc.  Often this level of resolution is inferior to what you’re trying to test. ideally your measurement hardware should be at least 10 db better than the gear you’re trying to measure. That often requires expensive hardware which often doesn’t work with RMAA correctly anyway.

FALSE SECURITY: People often do a "loop back" test to first test their PC sound hardware. And, more often than not, RMAA reports fairly impressive results. So they think they're good to go. But there are often problems being masked, or caused, by one or more of the above issues. For example, level, loading, and grounding issues won't show up in a loop back test. And don’t forget the previous list of 15+ things RMAA doesn't test for at all.

RMAA DEFECTS: On top of all the above issues, the RMAA software itself is buggy and prone to problems. The last release was several years ago and development has apparently been abandoned. Some of the FFT/math used to calculate the results is apparently wrong and/or has serious limitations that are neither properly documented nor obvious. In my experience the results often disagree with a professional audi analyzer. Plus there are lots of quirks and some obvious bugs. For example, try right clicking on any of the icons for the combined results graphs--instant total crash and you lose all your results.

Perhaps most important, RMAA often just outputs bad data instead of indicating an error condition with the measurement. To use an extreme example, you can stop playback of the test file half way through the tests, and instead of issuing an error, RMAA goes right on calculating and randomly spews out half truths and half garbage. It's obvious the developer didn't bother to put in many checks for validity of the input and output data. This can take much more subtle forms such as bad cabling, random noise (RFI from a cell phone, sound events from your PC, etc.),  not being detected during the test.

Here's an example of RMAA inventing impossible results out of nothing. It was run with no test file at all and, instead of timing out, it ran without a single error or warning and here's the result:


Wow, 245% IMD, that's pretty bad! I didn't know you could have more than 100%. But the THD is “only” 1.7%. The frequency response doesn't look so good though. I joke, but you get the idea. This doesn't inspire confidence in the software. And then there are messages like this one:


So clipping occurred but 0.000% of the samples were actually clipped? Hmmm. The funny thing is, in this case, there was zero clipping by the ADC (my Benchmark ADC1 has clip lights that stay on if it even clips a single sample until you reset the LEDs). And what's even more interesting is lowering the level by several dB and running the test again (as the error suggests doing) yields the exact same error! It appears to just be a case of RMAA getting confused trying to test this particular device (a device that tests just fine on my Prism dScope). Again, this does nothing to boost confidence in RMAA results.

RMAA DESIGNED BACKWARDS? I’ve seen enough weird results from RMAA I have to wonder if the developer designed it backwards to match the specs of a reference soundcard? Whatever test signals and analysis being used might have been “tweaked” until he got numbers that roughly matched the specs of some respected soundcard. If true, this would explain a lot. Ideally RMAA would just do all the math correctly and the results would be verified against a professional audio analyzer to make sure the math was accurate. But it often doesn’t agree and I’m at a loss to otherwise explain why unless there are either bugs in the calculations or it was designed to produce the expected numbers when testing a certain soundcard.


RMAA TIPS: So how is a person supposed to make RMAA measurements more accurately? First, use the best PC audio hardware you can. A good device would be something like the RME Fireface UC but, in my experience, proprietary low latency drivers designed for multi-track recording don’t work well (or at all) with RMAA. Example popular devices that do work are the E-Mu USB audio interfaces:

  • Creative/E-Mu 0202 - This device has decent A/D and D/A performance but it has un-calibrated level controls that cover a wide range of gain (65 dB!) and are very "touchy". This makes setting the level the same for both channels just about impossible. So you can’t know the actual channel balance of the device you’re trying to test and certain other tests can be inaccurate. The 0202 also has a mic preamp in the signal path on just one channel. Not surprisingly, that channel has higher noise than the line-in only channel. So any device you test with the 0202 is prone to rather lopsided results, weird crosstalk numbers, etc. It's also at the mercy of the sometimes noisy USB power supply. And it has serious issues on some PCs with its proprietary buggy drivers.
  • Creative/E-Mu 0404 - This device is better than the 0202 in that it has its own dedicated power, and both channels have the same signal path (unfortunately both have mic preamps). It also has slightly better metering but still suffers from the "touchy" gain control problem and the inability to exactly match the channels without some external help. See the section below on Levels. Another downside of the 0404 is it requires the same special drivers as the 0202. The drivers haven't been updated in years and especially can have issues with Windows 7. As far as I know there only buggy beta 64 bit drivers.

SETTING LEVELS: This is critical for many reasons. First, it's useful to know if the device you’re testing has a channel imbalance. But if you use the meters in RMAA to say adjust your E-Mu 0404, you will remove any channel imbalance anywhere in the signal chain. You can either feed known matched levels (i.e. verified with a meter) into both inputs and then run the RMAA calibration, or use a "Y" cable to temporarily split one channel into two outputs to set the levels. That way you know any channel imbalance that shows up during testing is the device, not your set up.

MEASURING LEVELS: As discussed earlier, it's important to know what levels you're really testing at. So it's best to use an external true-RMS meter that’s accurate across audio frequencies. Most cheap DMM's, even many that claim true-RMS, are only designed to be accurate around 50 – 400 hz. They often "roll off" dramatically or behave unpredictably at higher, or lower, audio frequencies. True RMS calculation is non-trivial and it’s expensive to have it be accurate at higher frequencies. So most reasonably priced meters don’t bother as they’re not intended for audio use. So if you're testing at say 1 Khz, they might read only half the real value. Looks at the specs for the meter, or if you already have one, test it on the output of your sound interface with test tones across the audio spectrum to see how flat it is (or isn't). It's a hassle to do the math at different frequencies to apply a correction factor but it’s cheaper than buying a more expensive meter.

CHOOSING A TEST LEVEL: If you're testing a device with a volume control, how far do you turn up the volume? I would suggest picking a value that's  well below the maximum output to help assure it won't be close to clipping. But if you test at too low of a value, the noise and distortion numbers will be much worse because noise will dominate the measurement. If you have a meter (see above) and the maximum output level is specified for the device a good starting point is about 70% of that value (roughly –3 dB). But this may also depend on the load being used. if the maximum output is not specified, see the next paragraph.

REFERENCE LEVELS: Here are some guidelines for reference levels:

  • Headphone Outputs - A good guideline for headphone outputs is at least 1mW into 32 ohms as that's used in the manufacture's specs of many devices and a value nearly any device can manage without clipping. P = (V*V)/R which means V = SquareRoot(P*R). So the square root of (.001v * 32 ohms) is 0.179 volts or about 180 mV RMS. This also happens to be about the typical power level most listen to their music at with typical low impedance headphones. So use your meter to set the player/amp to 180 mV RMS while playing a 0dBFS 1 Khz reference file. Then set your levels on your PC sound interface/mixer controls to be just under clipping (0 dB) while monitoring the 180 mV signal. But, beware, such low reference levels may challenge your PC hardware when making noise and crosstalk measurements. Many headphone outputs can deliver 300 – 500 mV into 32 ohms at low distortion so you might want to run a test at several different levels to find where the distortion starts to dramatically increase and the run the test just below that level.
  • Line Level Outputs – The Redbook standard for home digital audio gear is 2 Vrms. That’s what you’ll get from most home CD players playing a 0 dBFS signal. But portable and USB powered line outputs are often lower—typically in the range of 0.5 volts to 1.6 volts. And some gear can be above 2 V rms.
  • Speaker Outputs – As mentioned earlier, you need to use a voltage divider to avoid overloading and likely damaging the line level input of your RMAA sound hardware. This is not something I want to cover as if someone gets it wrong it could cause expensive damage. But for those in the know, the standard level for checking noise and distortion is typically 1 watt into 8 ohms which is 2.8 Vrms.


RUN A LOOPBACK TEST AT THE SAME LEVEL: Regardless of what sound hardware you're using, test it with a "loop back" test by connecting the inputs to outputs as instructed on the RMAA website to know what the "baseline" performance is at that level. Try to do this test at the same absolute level you're going to use for testing your piece of gear (i.e. the 180 mV mentioned above, etc). The signal to noise ratio, distortion, crosstalk, etc. of your PC interface will change at different test levels.

USE A PROPER LOAD: See the earlier section on loading. Ideally use a resistive test load. For devices intended to drive headphones, use resistors in the range of 15 – 600 ohms. A good single value is 32 ohms. Or test at several different impedances if you want. Even small 1/4 watt resistors are fine. They should be carbon film, and not wire wound. Use caution using headphones as a load both because they may be damaged by test tones at high levels, and they can also pick up room noise and create erroneous results due to the their impedance variations with frequency. For line level outputs the (often unknown) load of your RMAA sound hardware is probably sufficient. For speaker outputs the normal load is 8 ohms and it needs to be rated to handle whatever power level you’re going to test at.


VERIFY YOUR CABLES:  You don’t need $100 audiophile cables, but $1 ones from eBay are not a good idea either. Because of the way RMAA works, if you have a cable problem you may never know and just get marginal numbers. So test all the cables in the signal path by listening to a test tone while moving them around. If you hear static, or worse, replace the cable.  And, when testing crosstalk, cables become a huge problem if there’s any extra shared ground wiring.

RUN THE TESTS MULTIPLE TIMES: A good way to help rule out intermittent problems is to run the tests multiple times. If you get consistent results, they’re more likely to be valid. If the results vary by more than 5%, something is likely wrong.

USE A REAL OSCILLOSCOPE:  If possible, it’s worth checking the output, especially driving a real world load, with an oscilloscope that has a bandwidth out to at least 5 Mhz. This may show problems not seen in RMAA using a sound card limited to 20 Khz – 96 Khz. You can also measure the slew rate on a real scope and better evaluate square wave and impulse performance. Beware most scopes have the same grounding issues as PC sound hardware.

HOW I DO IT: For RMAA testing I use a Benchmark ADC1 for the A/D, a Benchmark DAC1 Pre for D/A, a 6 1/2 digit Agilent bench DMM that's extremely accurate and flat from 10 hz - 100 Khz for exact levels and other measurements, and one of several oscilloscopes with 60+ Mhz bandwidth. But, far more often, I use a dScope audio analyzer as it can make far more measurements, conforms to international testing standards, and produces far better and more accurate results. See my blog post Testing Methods for more information.