A top audio engineer explains NPR’s signature sound (2015)
current.org
current.org
https://www.youtube.com/watch?v=e07bI5rz6FY
Edit: this is one of my favourite tiny desk concerts, it sounds so good on headphones https://www.youtube.com/watch?v=47XlUL6sRow
https://www.youtube.com/watch?v=eB4oFu4BtQ8 (Roots. The brass mix gives me goosebumps lol)
https://www.youtube.com/watch?v=jFycqnOpifQ (Nickel Creek. If there really is only one mic someone has sold their soul. The sound stage is perfect.)
“Note: The secret of the 418-S is that it's two mics in one; a cardioid "mid" condenser capsule facing front, and a bidirectional capsule focused on either "side." It's a mid/side stereo mic that needs to be decoded in post, which allows you to control the stereo width of the image, even after the recording is made.”
https://www.npr.org/sections/allsongs/2019/04/02/705579879/t...
The one mic thing is actually ort of a bluegrass tradition, stemming back from radio and how it was recorded. A lot of bluegrass players learn how to balance themselves around a single mic, moving towards or away with solo and leaning in over the instrument to sing.
I used to produce a (very) little bit of live music and the difference in ‘babysitting’ required between an professional performer and someone who may have an amazing or even superior talent but little experience was remarkable.
https://www.npr.org/2007/11/04/15859351/stephin-merritt-two-...
https://www.youtube.com/watch?v=A7My5IpEzVM&ab_channel=NPRMu...
Great thing we have sources like NPR Tiny Desk and KEXP so we can discover and enjoy awesome music.
Just throwing in some semi-random sets I enjoy
KEXP
https://www.youtube.com/watch?v=jhicDUgXyNg Delvon Lamarr Organ Trio
https://www.youtube.com/watch?v=Ub-naDJbKFY Billy Strings & Don Julin
https://www.youtube.com/watch?v=-IIuHL_8Olg amiina
https://www.youtube.com/watch?v=Ev5UXJpUTpY Mokoomba
https://www.youtube.com/watch?v=PpnQlnra_to Kikagaku Moyo
https://www.youtube.com/watch?v=vo1zyMk1n5I José González
NPR
https://www.youtube.com/watch?v=QGAzPtwUJJU Andrew Bird
https://www.youtube.com/watch?v=cwYAeUpH1NM Trombone Shorty
https://www.youtube.com/watch?v=hsNKSbTNd5I DakhaBrakha
I went on a SRV bender a couple of years ago, comparing the studio cut of Lenny to pretty much any bootleg live recording i could find. There’s no comparison to be had, the live ones are just better.
There are a million other examples of course but possibly a surprising one is Miley Cyrus. Her energy during shows just hits highs that seem unavailable in the studio. Her vocals have come into a new era these past few years, especially her covers.
Anyway, thanks for links!
https://www.youtube.com/watch?v=Hxg1dL_x0gw
I think this is the first time i've understood the siren song. I have no idea what they are saying somehow i still want to do it.
Was curious. ~$3600 for the mic set.
As far as mics go, if you don't want to pay thousands for a Neumann, the Austrian Audio OC18 is a fantastic mic with a similarly flat response and has a 3-way switch for different levels of high-pass filtering before the signal even leaves the mic. It's fast becoming my favorite mic to use in the studio.
To clarify a bit, I think that, by "artificial", you mean the boost does not correspond to how the human voice actually sounds, which is true.
But in another sense, it's not artificial. It's a natural side effect of the physics of how microphones work. In building microphones to be directional (favor sounds from, say, in front), they've also made it where the amount of bass picked up is heightened when the mic is very close.
So NPR is artificially (with a high-pass filter) removing a natural side-effect (of directional mics) to avoid getting artificial-sounding boomy bass.
Also, this is one of those accidental invention things where what was originally a side effect has turned into a valued, essential feature. Like guitar amp distortion is part of the electric guitar sound. Or like how resonator guitars (Dobro, National) were invented to be louder but now people like the tone.
This is called the Proximity Effect, isn't it?
He says later on in the article that they try and get people in the studio to not talk directly into the mic but across it. So in some ways they are trying to correct for the issues caused by strong directionality before they get to artificial things like signal filters.
It is so clear that getting a clean analog signal up front is worth a lot.
The article just said it's not left up to the studio.
I don't know if you’ve worked with modern audio software but its truly a tangled stack of complexity, incompatibility, license management, etc etc. It can sound and work great once set up, but its touchy stuff when it comes time to make changes, update the OS, etc. as we all know software is notoriously buggy which is far from ideal in a live scenario.
Software pipelines only began to get into radio some 20/25 years ago. NPR started in 71
Also software can't do magic (and you aren't processing each microphone digitally), you want to be your signal to be as best as it can as close to the source as you can make it to be.
Plugins - software for audio programs - are available but audio engineers are famously persnickety.
Further down the audio path it eventually makes sense to digitize, but if you didn't have the HPF in the mic your noise floor will be worse.
They gave all the equipment to a pack of monkeys for the night. Anything still working in the morning was certified as reliable.
Thanks to near constant use of auto-tune I think most people realize pop vocals are thin.
Edit: clarification to remove accidental contradiction. I initially ended with "... I think most people realize that," which would have essentially translated to, "Most people realize that pop vocals are thinner than most people realize."
These plugins really exist to save time for large studios, not make bad musicians better. Time is money for studios, so they don't want to waste it on multiple retakes when someone can be close enough to make small fixes with melodyne. For session work, market effects still pressure people to, well, not make mistakes like that. A great singer is still going to be in higher demand than a decent one, because then the studios don't have to spend much time at all fixing their vocals.
Also, -noticeable- autotune can be desired. It's a musical choice. In that sense it's no different than using a vocoder, etc. I personally do not like it but that's the beauty of music; there's something for everyone.
Ever since then, it's not bothered me nearly so much when the vocals are tuned. The track hits harder. Yes: it's true the voice loses some of it's natural beauty, but in turn, you get music and voice that follow perfectly.
Now, I really hate hearing it in kids' songs. Sung by kids, for kids, and sounding so flat and blah. Much like laments against modern "beauty" productions, I think excessive autotune presents kids with an unrealistic expectation for their own voices.
But then the whole production sees similar things all over the place and it gets cleaned right up technically. Time, levels, the works right?
And the energy is diminished, could be lost.
Like fashion, this will all cycle in and out. Young people hear the humanity in music made prior to these and other tools and it appeals.
Little things, like a change in tempo, small vocal errors, inconsistency, all add up in a track.
I bet some time from now, could be as little as a decade, maybe two, we will look back at all this and chuckle.
Like you say, there is nothing technically wrong with any of this tech. And it could all be used very differently from how it is today too.
Recently, I have been going back through great live shows. Fantastic! And I still get that tingle from the realization someone delivered it live, to a crowd. And yeah, not so perfect, but oh so very human too.
Good application of it is not, no. When we hear obvious autotune vocals, it's a deliberate aesthetic choice.
I believe what you're talking about is how modern production is about producing "perfect" song recordings, and mapping everything to a click track/beat grid. Now that is totally noticeable compared to music made a few decades ago. I do agree that it makes music sound sterile. This is separate to autotune/melodyne being used.
"I bet some time from now, could be as little as a decade, maybe two, we will look back at all this and chuckle."
Maybe the main industry studios will, but music in general isn't determined by what those folks are doing. There are more indie publishers than ever, and so on.
And yes! The indies are all over the place. Love it.
This week’s video shows a side by side comparison between 3 popular consumer and boutique microphones.
Neumann U87, Røde NT1-A, and Fifine K670.
Are they worth the price difference?? Let’s find out!
The mic cost is almost irrelevant though. A good mic will last decades unless abused. Let's say you want a variety of sounds. You buy a bunch of instrument mics (probably $100 each) and a few matched pairs of all the most popular vocal mics (most of those will run 1-2k per pair). You'll probably not spend over 20k in total. Over 20 years, that's only $1,000 per year or less than $100 per month. In that same 20 years, you will have upgraded your digital equipment several times at an expense far greater than $1,000 per year (upgrading your $3,500 macbook every 4 years is the same amount of money).
If you make your money with those mics, that cost is hardly worth mentioning. It's like people complaining that ergonomic keyboards cost $300. The keyboard will easily last a decade or more (only $1-2 per month to save a lot of future pain). In that same time, you'll probably spend 10k+ on other equipment. Same thing with monitors where $1000 will far outlast that same amount of money put into the computer itself.
I'd say the $250 mic was the best value, but that German mic was niiiice. If you can afford it, then it's probably the one you want.
For simple spoken-word stuff like conferences or streams or whatever, something like a Samson Q2U or an AT2005USB/ATR2100 are less sensitive to unwanted noise and easier for an untrained user to get a good sound out of, while moving into the XLR space gets you access to better dynamic microphones and also some pretty reasonably priced condensers that do quite well (though there's some up-front investment in the audio hardware, of course).
It's actively bad for most people for one reason: capacitive mics pick up everything.
If your room isn't soundproofed, it will be very hard to keep noise out of your recording. Dynamic mics are much less sensitive in this regard.
I would instead recommend a Samson 2Qu or Audio Technica ATR2100-USB on the low end ($70-100) or the Shure MV7 ($250) on the high end for plug-n-play mics.
If you want to move into a cheap audio interface (eg, Focusrite Scarlett Solo + cloudlifter), I'd recommend either the Shure SM7b or the ElectroVoice re20 on the higher end and the Shure sm57 on the cheaper end (good enough for the president to use for the last 40-ish years).
This is a myth that's popular with podcasters. If you get as close to a condenser mic as you must with a less-sensitive dynamic mic* and crank down the gain accordingly, you'll find that condenser mics don't magically capture more ambient noise than than dynamic mics.
* Using a fist as a measure, your mouth should be between 1-2 fists away from the mic.
With a good preprocessor (I use a Symetrix I got from an old radio station), I can crank my dynamic mic (EV RE320) to levels that will pick up anything happening in my entire house, with my office door closed.
It's just that the levels from condenser mics tend to be hotter, so by default you hear more stuff in the room unless you get in close and turn it down.
There's no way to replicate the 'radio' sound if you're 2 or 3 ft away from the mic.
For what it's worth I'm an audio engineering expert, produced albums, broadcast stuff, and used to review professional studio audio equipment for a living for a national magazine.
People like it because it's simple and it looks cool.
If you want something that has the same basic usability, ie plugs directly into USB and is really easy to use with computer audio, buy the Apogee Mic Plus.
I recently experimented with pretty much everything in this category and was very happy with this model, bought a dozen of them for use in a virtual conference series, where I wanted something I could send to non-technical people who'd never be able to navigate a pro audio interface. I've been very happy with it so far.
But then, professional equipment never had economies of scale.
I will also say, the NPR person interviewed seems to have a negative view of the RE20 and SM7b compared to the U87. Despite the low cost, SM7Bs are actually popular studio vocal mics. One was famously used to record Michael Jackson's vocals on Thriller (actually it was an SM7 but I believe the SM7B reissue is almost identical). When recording voice, there normally isn't a "silver bullet" mic that is the best for all voices.
Famously used by Mariah Carey.
In fact, it's coming back with a new SKU (again): https://www.frontendaudio.com/blog/sony-announces-the-c800gp...
And then there are the used vintage mics which can go for $15k+.
At the cheap but well-regarded end of things is the Stellar X2 from TZ TechZone.
It's designed to hold the mic and avoid transmitting vibrations from the mic stand (caused by moving or jostling the stand) to the mic.
Topic drift: I hammer on my students that contracts are much more readable if done in short, single-subject paragraphs without long wall-of-words passages.
Don't lawyers have effective ways to include and reference things, create standard definitions and procedures without pasting the same stuff everywhere?
In some fields, yes — but as a class, lawyers: (A) notoriously prefer reinventing the wheel, and (B) sometimes could be suspected of hoping that the MEGO Factor — Mine Eyes Glaze Over — will cause the other side's contract-draft reviewer to overlook something that the drafter buried in a long, wall-of-words provision. I see that happen pretty regularly.
(In the 1990s I initiated and headed up a project for the American Bar Association Section of IP Law to try to standardize the wording of various building-block clauses for software license agreements. [0] The chief IP counsel of a Fortune X company [X being a very-low number], whom I knew pretty well from the Section, said he was opposed to having any kind of standardized language because, he said (paraphrasing), "I want to be free to be an asshole.")
[0] https://www.oncontracts.com/docs/Rutgers-MSLP-Precursor-to-G...
That could actually turn out to be worse. Take a look at a lot of federal bills. They're written like:
'In 8 USC 552(b)(ii) strike the word "foo" and insert "bar baz"'
You then have to go cross reference everything for every line. It's a nightmare. If the bill was written in a computer readable diff format instead, that could be better.
I no longer work in law ;)
From the oleaginous Francis Urquhart in the wonderful original (British) version of House of Cards: "You might think that. I couldn't possibly comment." [0]
This is one of the last, best uses of ISDN. Guaranteed latency, ultra low jitter, and plenty of high-quality hardware purpose-built for getting the best possible studio audio over 2 bearer channels worth of capacity.
Is there really nothing coming close to that?
Technically you could accomplish the same thing by applying a parametric eq to the master buss, but then you're no longer software agnostic.
It's like photography; sure one can post-process photos in photoshop. But getting everything right before taking the picture, at a hardware level, simplifies things for everyone involved.
Is there a better way?
https://www.inquirer.com/philly/entertainment/WHYY-NPR-Terry...
I should say, the shows mentioned in this article are actually produced by NPR but most of what you hear on a given public radio station isn't. And also, NPR doesn't control the broadcast.
I wrote in to a cable TV show a decade ago to call out the nose hair whistling and mouth sounds. They never replied to me, but they rolled off the highs for the remainder of the shows.
It was like someone I don't know, whispering sweet, unsolicited nothings in my ear. Felt uncomfortably intimate in a way I hated. I was always like, "Lady, I don't know you like that, so cut it out."
https://languagelog.ldc.upenn.edu/nll/?p=17489
"And they want to talk about the crazy ways that young women are speaking. And the first thing they do is attribute it to young women, even though young men are doing it too. So it's a policing of young people, but I think most particularly young women."
Men (and women) have spoken with vocal fry for the past several millennia, but I don't recall reading of anyone being annoyed by it until recently when everybody decided that millennial women speaking like that on the radio was anathema.
This all started in the last 1-2 years. It's not extremely infrequent, I hear these during prime driving times and probably around once/week. I know for sure I have heard it on at least one non-NPR FM station. I wonder anyone else has noticed the same in other markets?
So maybe the problem is really just a defect of my car's radio when toggling.
https://maximumfun.org/podcasts/greatest-generation/
So far, the hosts have done a complete re-watch of TNG and DS9. Just started Voyager recently.
Then when they start talking about the episode it's fun and nostalgic, and they make astute observations that I haven't heard elsewhere.
For anyone even vaguely familiar with audio engineering and recording, these tactics are not profound. Not a bad thing because in the end, less is more.
Worth mentioning that a good mic is arguably the 20% input that contributes to 80%+ of the output/audio quality, as supported by the article.
#6 is really the only non-obvious point. Apparently this is a major subject of debate.
1.) If you can afford it, use the Neumann U87 mic (~$3.5k)
2.) High pass filter (~250hz) on the vocal chain
3.) To avoid plosives, don't speak head-on into the mic. Speak off the side, on a diagonal. Use a pop filter.
4.) Design your studio to minimize reverberation. Make sure the recording space is isolated and there "aren't a lot of solid walls." Absorb sound with baffles, sound panels, etc. Counterintuitively, a larger room with more diffusion is better than the opposite.
5.) Minimize ambient sound. Your mic will pick up everything from fans to CPUs to electronic interference off computer screens. This noise will muddy up the recording.
6.) Minimize processing or compression of the signal before streaming, or in the case of radio, sending to the satellite.
Edit: for clarity
BBC Radio 3 uses no dynamic range compression, so might be most comparable to NPR (although it's likely that each local station applies a ton of compression before the signal hits the air).
Most (other) radio stations apply copious amounts of multiband dynamic range compression on their output - with the nickname of "sausage-making", since the process turns waveforms that look like music into waveforms that look like sausages. In the FM days, louder sounding stations were associated with better signals, so got bigger market share...
Using a decent microphone (AKG D5 in my case) and a little bit of tweaking (just a low cut and some compression is a good start) instantly puts your sound quality in voice chats so far above everyone else using cheap headsets or their laptops' built-in mics.
Anecdotally I've found that sounding more authoritative makes people listen a lot more to what you say, instead of zoning out.
Of course if you had it on and were further away from the mic, you'd thin out lower voices. Just goes to show micing people (or instruments) isn't entirely straightforward.
[0]: (Page 4) https://media.sweetwater.com/store/media/u87ai_u87.pdf
I'm suspicious when percentages that don't have to add up to 100% add up to 100%.
But the average comment section of a news article or blog is so much worse. It has the insanity of 4chan, but with better grammar.
The NPR trick is no solid walls. Eg the Tiny desk room is no normal soundproof studio box, but a normal office room with lots of bookcases, non-empty tables, not many solid surfaces like walls or screens. Having less mics and natural lighting also helps a lot.
He didn't talk about the light. Glass windows are terrible. A normal studio box has a large window, which causes more reflections, more than normal windows at the side. In their case I think they just put some plants there.
Hearing the mouse click as they start playing the various segments and underwriting recordings gives me a slight tic.
I -guess- CNNs can look at e.g. reduced frequency range recordings (like phone calls), and attempt to reconstruct them. However this seems like an arduous mountain to climb, as people's voices are unique. So are their environments and signal chains. I really doubt that something that generalized would work very well at reconstructing a specific person's voice and recording.
This also gets into the problem that it would be constructing a new reality, not recreating it.
I have thought about the constructing a new reality thing - I wouldn't be surprised if models ended up being trained to misspeak words which could get confusing...
[0] https://en.wikipedia.org/wiki/Deep_learning_super_sampling
I mean, after all it's a low-cut filter, isn't it?
I'd guess NPR's view is along the lines that, well, the filter's already there, and we like the way it sounds, so we keep using it.
Maybe you want to preserve the bass of interstitial music or program audio jingles or environmental effects or something. Doing the processing after the mixing means it affects the whole mix. Doing it at the input means you can tailor each element.
Worse, because bass has an outsize effect on the total energy in an audio signal, if there's any sort of dynamic range compression while the bass is still included, the presence of the bass triggers that compression to happen. Later on when the bass is removed, the remaining audio has inexplicable fluctuations in its volume, which can sound super uncomfortable.
This "program level bouncing around in response to a signal which is not part of the program audio" effect can also come from side-chain compression, and arguably filtering after compressing may be a form thereof. Once in a while it's done to great artistic effect in music, but in talk settings it's almost always horrible and disorienting.