96 kHz.org
Advanced Audio Recording

problems with microphones

A common problem in audio data streams are the strong room reflections which create resonances at certain frequencies and can last several hundred milliseconds. These overlap the syllables of subsequent words, obscure intelligibility, and due to signal peaks result in very uneven volume levels. The size of the rooms determines the reflection time and the reflection pattern, which usually occur precisely in the area that causes significant interference that way that they mask the subtile nuances of pronunciation. This is further made worse if you don’t speak directly into the microphone: Microphones generally pick up high frequencies better when sound comes directly at them. This is true even for so called "omnidirectional" microphones. There is a significantly increased on-axis response for high frequencies. If you speak from the side, the high frequencies are reduced or lost. This is even more pronounced if you speak slightly off-center. As a result, speech immediately becomes much harder to understand.

angle sensitivity of cardioid pattern microphones - polar pattern issue

As a result, a great part of the quality is lost already at the beginning of the signal chain. So microphone placement is the key in the first place.

One should always position the microphone so that it covers the entire subject (the mouth when speaking, and the mouth, head, and chest area when singing), while avoiding direct sound transmission to the diaphragm. This applies to both the height and the lateral position: Try to speak or sing over the microphone.

height angle sensitivity of microphones - polar pattern issue when placing to close


issues with audio compression

A great part is also lost during audio compression, both in terms of dynamic range and sound quality. Nodaways streaming formats heavily reduce information which was essential for recognition: In order to reduce data size, these methods take advantage of the characteristics of human hearing and selectively remove acoustic information that cannot be perceived or better: Which is supposed to be irrelevant. As for an example a loud sound masks, for fractions of a second, sounds that occurred shortly before (pre-masking) or shortly after (post-masking). This creates a focus on loud signals, while quiet, masked signals are suppressed. Consequently, strong reflections can lead to a potential loss of important information. Many transient details are lost this way. That makes the audio quality even worse.

Studies show that listening to poor-quality audio requires more concentration and leads to greater fatigue. This is especially the case for long transmissions and presentations. The strategy must therefore be to accurately capture the high frequencies and keep the frequency range and volume as consistent as possible. When recording music and speech, we generally use a setup that allows us to work linearly without extensive dynamic compression, sound correction, or even echo cancellation. To achieve this, the microphones are placed well within the reverberation radius to ensure a sufficient signal-to-noise ratio.
 
 

Headset usage

The recommendation for talking in small rooms is to use a headset micorophone to avoid room reflections and ensure a consistent distance from the microphone and a consistent recording angle. This keeps the sound and volume constant, so the audio processor doesn't have to correct the volume. Furthermore, the compression doesn't focus on the high levels and doesn't mask the finer details. To do this, place the microphone next to your mouth but not directly in front of it. Place it aside directing in 60° to 80° degree to the mouth outside the direction of speech, so that it is not struck by explosive sounds such as B, P, and T.  At this short distance, high frequencies are still reproduced very well without being overemphasized, and vibrations of the diaphragms caused by sudden loud sounds are avoided.

ideal microphone position with headsets

Normally no further processing is necessary then!

 

Audio Processing Issues

Post sound processing should generally be used with caution:

Unfortunately, attempts to suppress or eliminate echoes often lead to very negative results and make speech even less intelligible. Any processing with generalised algorithms can reduce speech intelligibility. Even the best algorithms cannot always perfectly distinguish between important speech information, room reflections, background noise, and speech details.  This also applies to noise reduction, which can cause the beginning of words to be cut off. It's best to avoid gating and denoising if possible.

 To adapt to the data transmission channel (CD, Musiktaxi, the Internet, MP3 streaming), it is therefore usually sufficient to reduce the dynamic range if required.

A useful chain for automatic processing consists of a compressor set to a very fast response of 3 dB within a 40ms ... 60ms window, as well as a compressor set to a very slow response within a 300ms to 500 ms window of also 3 dB, which acts as an AGC (automatic gain control). At the very end, a limiter with a 3 dB headroom can be used to lift up the volumel. This allows a total of 9 dB of equalization of the over all volume level. For typical speech signals, I use less than 2 dB in each of the sections mentioned leading to an over all compensation of 6dB meaning ratio 2.

 

Conclusion and Summary

The best audio quality starts at the source: A well-positioned microphone prevents room reflections and fluctuations in audio level. The cleaner the original signal is, the less compression, noise reduction, or other processing is needed later.


Read more in the earlier article about the advantage of 192kHz.

Read an article about decho canceling.

 

© 2007 J.S.