Glossary
54 terms, defined the way a knowledgeable friend would explain them. Each one says what the thing is, then the detail people actually get wrong. Where a guide goes further, the entry links to it.
A
- aliasing #
- Aliasing is what happens when a process generates frequencies above the Nyquist limit, half the sample rate, and they fold back into the audible range as inharmonic tones. At 44.1 kHz, a 30 kHz harmonic from a saturator reflects off the 22.05 kHz ceiling and lands at 14.1 kHz, unrelated to the note you played. Distortion and clipping plugins are the usual offenders, which is why many offer oversampling.
- attack #
- Attack is the time a compressor takes to ramp its gain reduction to full depth once the signal crosses the threshold. Reduction starts the instant the threshold is crossed; the attack setting only controls how fast it deepens. Drums are where it shows: the snap of a snare lives in roughly its first 10 to 30 ms, so a 1 ms attack flattens it while a 30 ms attack lets it pass before the clamp arrives.
- automation #
- Automation is a recorded curve that moves a parameter over time, so a fader ride or filter sweep plays back identically on every pass. Order of operations is the part to watch: volume automation on a channel is post-insert in most DAWs, so it will not change how hard you drive a compressor sitting on that channel. To automate the level feeding your inserts, automate a gain plugin placed before them.
B
- bit depth #
- Bit depth is the number of bits stored per sample, which sets the gap between the loudest possible level and the noise floor at roughly 6 dB per bit. 16-bit gives about 96 dB of range, 24-bit about 144 dB on paper. The extra bits buy headroom rather than finer detail; you can track around -18 dBFS and the noise floor stays irrelevant. When exporting to 16-bit, dither so quantization error becomes benign noise instead of distortion.
- buffer size #
- Buffer size is how many samples your interface processes per block, and it sets the latency you feel: 256 samples at 44.1 kHz is about 5.8 ms each way, before converter and plugin delay add more. Round-trip figures are what matter for monitoring, usually double the buffer plus a couple of ms. Track at 64 or 128, then raise it to 512 or 1024 for mixing, when latency stops mattering and CPU headroom does.
- bus #
- A bus is a channel that sums the output of several tracks so you can process or level them together, a drum bus and a vocal bus being the usual cases. Buses and sends get conflated: routing a track's output to a bus moves the whole signal, while a send takes an adjustable copy and leaves the original in place. Whether that copy follows the fader is the pre and post fader question.
C
- comb filtering #
- Comb filtering is the row of evenly spaced notches you get when a signal sums with a delayed copy of itself. The spacing follows the delay: a 1 ms offset puts the first null at 500 Hz and repeats every 1000 Hz above it. It is why two mics on one source, or a duplicated track nudged a few samples, can hollow out a tone. Doubling the delay halves the spacing, chewing lower into the spectrum.
- crest factor #
- Crest factor is the gap between a signal's peak level and its RMS level, expressed in dB. A sine wave measures 3 dB; raw multitracked drums can run 18 dB or more, while a heavily limited master might sit near 6. It is the number compression and limiting actually change, which makes it useful for judging how squashed a mix already is before you reach for another limiter.
- cutoff frequency #
- Cutoff frequency is the point where a filter's attenuation reaches 3 dB, the conventional marker for the edge of the passband. A low-pass set to 8 kHz is already 3 dB down at 8 kHz itself, then keeps falling at its slope, typically 12 or 24 dB per octave. That means the filter is audibly working below the number on the dial, and with resonance up it boosts right at the cutoff before the drop.
D
- dBFS #
- dBFS is decibels relative to full scale, the highest value a digital system can represent, so 0 dBFS is the ceiling and every real level reads negative. Producers trip on the analog crossover: many hardware emulation plugins calibrate their internal 0 VU to -18 dBFS (some use -20), so a signal peaking near 0 dBFS drives the model far harder than the original unit was ever run. Keep the average around -18 and the emulation behaves.
- DC offset #
- DC offset is a constant shift that moves the whole waveform off the zero line, effectively a 0 Hz component riding under the audio. It eats headroom silently, since peaks on one side hit the ceiling early, and it puts a click at every edit point because cuts no longer land on zero. Some synth waveforms and cheap interfaces introduce it; a high-pass around 10 to 20 Hz removes it without touching anything audible.
- de-esser #
- A de-esser is a compressor that reacts only to sibilance, the s and t energy sitting between roughly 5 and 8 kHz on most voices. Cut too deep and the esses become a lisp; solo the detection band to find where the ess actually lives, then pull only a few dB. Split-band designs duck just that band, while wideband designs drop the whole vocal level on every ess, dulling whatever else is happening at that moment.
- dither #
- Dither is low-level noise added on purpose before a reduction in bit depth. Without it, truncating to 16 bits rounds each sample to the nearest step, and the rounding error correlates with the signal, which you hear as distortion on fades and reverb tails. Dither trades that for a steady hiss near -96 dBFS, below anything a playback system will expose. Apply it once, at the final bounce, and never before further processing.
- dry/wet #
- Dry/wet is the balance between an effect's unprocessed input (dry) and its processed output (wet). On an insert, the knob does what it says. On a send, set it to 100 percent wet, because the dry signal already exists on the source channel; anything less doubles the dry copy through the return bus and shifts the level every time you touch the send. Parallel compression is the same idea driven from the wet side.
E
- envelope (ADSR) #
- An envelope is a control signal that reshapes a parameter over time, retriggered by each note; ADSR names its four stages: attack, decay, sustain, release. Sustain is the odd one out: a level, not a time. The other three are durations, but sustain is the height the envelope holds while the key stays down. That is why lengthening decay does nothing audible when sustain sits at full, and why a pluck wants sustain at zero with decay doing the work.
- envelope follower #
- An envelope follower tracks the loudness contour of an audio signal and turns it into a control signal, which is how auto-wahs and a lot of ducking effects work. Release is the control to watch. One cycle of an 80 Hz bass lasts 12.5 ms, so a release faster than that makes the follower ride individual waveform cycles instead of the note, and whatever it modulates buzzes at the pitch of the bass. Slow the release until the ripple stops.
F
- FFT #
- The fast Fourier transform is the algorithm that turns a slice of audio into a spectrum, and it sits under every analyzer and most spectral tools. Its resolution is a trade. A 1024-sample window at 44.1 kHz gives frequency bins about 43 Hz wide, which cannot tell two low bass notes apart, since whole octaves down there span less than that. A longer window sharpens frequency and blurs timing, which is why spectrograms of drums look smeared at big FFT sizes. Details live in the DSP section.
- filter slope #
- Filter slope is how quickly a filter attenuates past its cutoff, measured in dB per octave; each pole contributes 6 dB per octave, so a 4-pole filter falls at 24. Slope describes what happens well beyond cutoff, not at it. At the cutoff frequency itself a standard design is only 3 dB down, whatever the slope, so a 24 dB per octave high-pass at 100 Hz still lets plenty of 90 Hz through. The full guide covers where each slope earns its keep.
- formant #
- A formant is a resonant peak in a sound's spectrum whose position is set by the body that made the sound, not by the note being played. Vowels are formant patterns; an adult "ah" puts its first two peaks near 700 and 1100 Hz whatever the pitch. This is why plain pitch shifting sounds wrong on vocals: it drags the formants along with the pitch, shrinking or enlarging the apparent singer. Formant-preserving shifters move the note and leave the peaks where they were.
G
- gain staging #
- Gain staging is setting the level at each point in the chain so every stage receives what it was designed for. Inside a modern DAW the mixer is 32-bit float and will not clip between plugins, so the practice is really about the plugins themselves. Analog-modeled processors are usually calibrated so that around -18 dBFS behaves like the hardware's nominal level; feed them 10 dB hotter and you get more drive than the preset designer intended. Level matters when comparing too, so match outputs before judging.
H
- Haas effect #
- The Haas effect is the ear fusing two arrivals of the same sound into one when the gap is under roughly 35 ms, placing the source at whichever copy arrived first. Pan a dry copy left and a 15 ms delayed copy right and you get width with no level change. The bill arrives in mono: summed, that pair comb filters, with notches every 67 Hz or so for a 15 ms delay. Check mono before committing, or read what nudging does to phase.
- harmonics #
- Harmonics are frequency components at whole-number multiples of a fundamental; a 100 Hz note carries them at 200 and 300 Hz and at every multiple beyond. Their musical positions explain saturation character: the 2nd harmonic lands an octave above the note, the 3rd an octave plus a fifth, so a device favoring even multiples reinforces the note itself while odd multiples add intervals the chord may not want. Which devices favor which is the subject of even and odd harmonics.
- headroom #
- Headroom is the gap between a signal's peaks and the point where the system clips. Digital's trap is that the meter is not the whole story: sample peaks at 0 dBFS can reconstruct above full scale between samples, and lossy encoders push overs further, which is why mastering for streaming usually leaves about 1 dB of true-peak headroom (a ceiling of -1 dBTP). Headroom while mixing is cheaper still, since a 32-bit float bus can take absurd levels; the converter and the codec cannot.
- high-pass filter #
- A high-pass filter passes frequencies above its cutoff and attenuates those below it. The number on the dial misleads: cutoff marks the -3 dB point, not where the low end disappears. A 12 dB per octave high-pass set at 100 Hz leaves 50 Hz only about 12 dB down, still very present on a club system. Steepness is a choice with side effects of its own, covered in filter slopes explained, so check an octave below cutoff before deciding the rumble is gone.
K
- knee #
- Knee is how gradually a compressor moves from no gain reduction to its full ratio around the threshold. A hard knee switches at the threshold exactly; a soft knee spreads the transition across a stated range, so a 6 dB knee starts compressing gently 3 dB below the threshold and only reaches the full ratio 3 dB above it. That is why soft-knee settings show gain reduction earlier than the threshold reading suggests.
L
- latency #
- Latency is the delay between audio entering a system and leaving it. Buffer size sets most of it: 256 samples at 44.1 kHz is 5.8 ms each way, before the converters add their bit. Plugin delay compensation only fixes playback. A lookahead limiter or linear-phase EQ on the master adds its delay to everything you monitor through the DAW while tracking, so bypass those, or record at a small buffer and mix at a big one.
- LFO #
- Low-frequency oscillator: an oscillator cycling below the audible range, usually under 20 Hz, used to modulate other parameters rather than be heard directly. Retrigger is the control worth finding. A free-running LFO sits at a random point in its cycle when each note starts, so a filter sweep lands differently on every hit. Key-synced mode restarts the cycle per note, which is what you want for repeatable wobbles and bass movement.
- limiter #
- A compressor with a ratio of roughly 10:1 or higher and an attack fast enough that nothing passes the ceiling; modern brickwall designs are effectively infinite ratio. True peak and sample peak are not the same number. Lossy encoding can reconstruct inter-sample peaks above your sample-peak ceiling, which is why streaming services recommend a true-peak ceiling of -1 dBTP rather than pushing right up to 0 dBFS.
- lookahead #
- A delay on a dynamics processor's audio path so the detector reads the signal early and starts reducing gain before a peak arrives, usually 1 to 5 ms ahead. Lookahead is real latency: a limiter with 2 ms of lookahead reports at least 2 ms to the DAW. That is harmless on a master bus and a problem while tracking, which is why a limiter's zero-latency mode usually just switches lookahead off.
- LUFS #
- Loudness Units relative to Full Scale, the K-weighted loudness measure from ITU-R BS.1770 that streaming services use for normalization. Integrated LUFS averages the whole track, while short-term reads a 3 second window. Mastering hot anyway buys nothing: Spotify turns a -8 LUFS master down to -14, so the crushed transients stay and the level advantage disappears. Broadcast is a different world, targeting -23 LUFS under EBU R128.
M
- makeup gain #
- The output gain on a compressor that restores the level lost to gain reduction. It is also the main way compressors flatter themselves: louder reads as better even at differences under 1 dB, so a plugin adding 4 dB of makeup wins any bypass comparison regardless of what the compression did. Match bypassed and processed levels before judging, and treat auto-makeup as a rough guess, because most implementations overshoot.
- mid/side #
- A stereo encoding that stores sum and difference instead of left and right: mid is (L+R)/2, side is (L-R)/2. The trap is treating the side channel as the wide stuff and boosting it freely, because everything in side cancels when the mix folds to mono, which keeps only mid. Before reaching for a widener, check what each channel actually contains, since side often holds reverb tails and bleed rather than the parts you meant to widen.
- mono compatibility #
- How much of a mix survives being summed to one channel, which still happens on phone speakers and plenty of club systems. Delay-based widening is the usual casualty: a 10 ms Haas offset spreads a mono source across the stereo field, then comb-filters into a thin version of itself when summed. Anything that exists only in the side channel vanishes outright in mono, so check by listening, not by trusting a correlation meter alone.
N
- Nyquist frequency #
- Half the sample rate, and the highest frequency a digital system can represent: 22.05 kHz in a 44.1 kHz session. Anything a process generates above it folds back below it as aliasing, landing at frequencies with no musical relationship to the source. Saturation is the usual culprit, since it creates harmonics far above the fundamental: distort a 5 kHz tone and its seventh harmonic at 35 kHz reflects back down to 9.1 kHz.
O
- oversampling #
- Running a process at a multiple of the session sample rate, typically 2x to 16x, so the harmonics distortion creates land below the raised Nyquist point instead of folding back as aliasing. Two costs come with it: CPU scales roughly with the multiple, and the downsampling filter adds latency, especially in linear-phase mode. It pays off on clippers and saturators; a delay or a gain utility gets nothing from it.
P
- pan law #
- The attenuation a mixer applies at center pan so a signal keeps the same perceived loudness anywhere in the stereo field. The common center cut is 3 dB, with DAW options running from 2.5 to 6 dB. The mistake is changing it mid-project: every centered track shifts level relative to panned ones, which silently rebalances a mix you thought was done. Pick the law before mixing starts and then leave it alone.
- phase #
- The position of a waveform within its cycle, measured in degrees, where 360 degrees is one full cycle. The standard confusion: the polarity button flips the waveform upside down, which is not a phase shift, although it matches a 180 degree shift on a lone sine wave. Real-world phase problems are timing offsets. Two mics on one snare a few centimetres apart produce comb filtering, and no polarity flip fully repairs that; alignment does.
- pink noise #
- Noise with equal energy per octave, which means its spectrum falls 3 dB per octave as frequency rises. White noise carries equal energy per hertz and sounds far brighter by comparison. Full mixes distribute energy in a roughly pink way, which is why pink noise works as a balancing reference: bring each track up against a fixed pink noise bed until it just pokes through, and you get a usable rough balance in minutes.
- polyphony #
- The number of voices an instrument can sound at once. Voices get used up faster than notes: a four-note chord on a patch with 8-voice unison eats 32 voices, and every release tail holds its voice until it fully fades. When the count runs out the synth steals the oldest voice, which is what that clipped-off pad tail actually is. Raise the voice limit or shorten the release before blaming the preset.
- pre-delay #
- The gap a reverb leaves between the dry signal and the onset of its reflections. Set 20 to 40 ms on a vocal and the consonants land before the tail arrives, which keeps the voice in front of the reverb instead of buried inside it. Left at zero on small rooms, the reverb piles onto the transient itself and reads as smear rather than space.
Q
- Q (resonance) #
- Q is the ratio of a filter band's center frequency to its bandwidth, so a higher Q means a narrower band. A Q of 1.41 spans about one octave; a Q of 10 centered at 200 Hz grabs a band only 20 Hz wide. Many EQs use proportional Q, which widens the band at low gain settings, so the same number behaves differently across plugins. On a resonant filter, higher Q boosts the region around the cutoff instead.
R
- ratio #
- Ratio is how strongly a compressor reduces signal above the threshold, written as input change against output change. At 4:1, a peak that goes 8 dB over the threshold comes out 2 dB over. From roughly 10:1 upward you are limiting. The panel number is a maximum, though: with a soft knee the effective ratio ramps up gradually around the threshold, so a 4:1 driven only lightly into the signal often behaves closer to 2:1.
- release #
- Release is how long a processor takes to let go: on a synth envelope, the fade after note-off; on a compressor, the time gain takes to recover once the signal drops back below the threshold. The compressor version is the one that bites. Set it shorter than one cycle of your lowest frequency and the gain reduction starts tracking the waveform itself, which is distortion. A 40 Hz bass repeats every 25 ms, so releases under that will chew on it.
- round-robin #
- Round-robin is a sampler mode that cycles through several alternate recordings of the same note, so repeated hits do not replay one identical file. A single snare sample fired on every eighth note is the machine-gun effect; even two round robins break it, and drum libraries commonly record 4 to 10 per velocity layer. Watch for samplers that reset or randomize the cycle position, because a bounce can then start on a different variation than the playback you approved.
S
- sample rate #
- Sample rate is how many amplitude snapshots per second a digital system stores, and it caps the highest frequency you can represent at half that figure, the Nyquist limit. 44.1 kHz captures up to 22.05 kHz. Raising the rate does not add detail you can hear; the practical benefit is headroom for saturation and other nonlinear plugins, whose new harmonics would otherwise land past Nyquist and fold back down as aliasing. Oversampling inside the plugin gets you that without a 96 kHz project.
- saturation #
- Saturation is mild nonlinear distortion: the waveform gets squashed and new harmonics appear at whole-number multiples of the input frequencies. Symmetrical clipping adds odd harmonics, asymmetrical circuits add even ones as well, and the second harmonic sits exactly one octave above the note. The reliable trap is level. Saturation adds energy and top end, so the processed signal is louder, and louder reads as better. Match output level to input before you judge it.
- sidechain #
- A sidechain is the detector input of a dynamics processor, split from the audio the processor actually touches. Route the kick into a compressor on the bass and the bass ducks on every kick hit. One step gets skipped: high-pass the sidechain around 60 to 100 Hz on bus compression, so low end stops driving the gain reduction and pumping the whole mix. For rhythmic ducking, a volume shaper is often the better tool than a compressor.
T
- threshold #
- Threshold is the level at which a dynamics processor starts working: above it a compressor turns gain down, below it a gate closes. There are two catches. A soft knee begins compressing below the stated threshold, so a 6 dB knee starts 3 dB early. And lowering the threshold plus adding makeup gain makes everything louder, which skews any bypass comparison toward the processed version. Match levels before deciding the compressor sounds better.
- transient #
- A transient is the short burst of energy at the onset of a sound, the first few milliseconds of a snare crack or a plucked string. Transients carry the peaks: a raw drum bus can peak 15 dB or more above its average level, and that gap is what compressors and clippers spend. Attack times under about 5 ms clamp the transient itself, while 10 to 30 ms lets it pass and squeezes what follows.
- true peak #
- True peak is the estimated level of the reconstructed analog waveform, measured in dBTP rather than dBFS. Sample values can all sit at or below 0 dBFS while the curve between them swings higher, so a master that never touches the clip light can still overload a converter or a lossy encoder. Meters estimate it by oversampling, 4x in the BS.1770 spec, and EBU R128 puts the ceiling at -1 dBTP.
U
- unison #
- Unison stacks a synth voice into several detuned copies per note, usually spread across the stereo field; the classic supersaw is 7 sawtooths. It costs you twice. Polyphony multiplies, so 7-voice unison under a 4-note chord is 28 voices. And the stereo spread is fragile in mono, where detuned copies can cancel instead of reinforcing. Solo the mono fold-down before committing a unison bass, and consider keeping the lowest octave a single voice.
V
- velocity layer #
- A velocity layer is a set of samples mapped to one slice of the MIDI velocity range, so harder key strikes trigger recordings of harder playing rather than the same file turned up. MIDI offers 127 usable velocity values; a 4-layer piano switches samples roughly every 32 of them, and the seam is audible when adjacent layers were recorded at mismatched levels. Crossfading between layers hides the jump at the cost of a slightly softened attack.
W
- wavetable #
- A wavetable is a stored row of single-cycle waveforms that an oscillator reads through, with a position control for scanning between frames; Serum's tables hold up to 256 frames of 2048 samples each. That scanning is what separates it from a plain oscillator. Aliasing is the limit: bright frames played high on the keyboard push harmonics past Nyquist unless the synth band-limits playback per octave, which is why cheap wavetable oscillators fizz in the top range.
- white noise #
- White noise is a random signal with equal power per hertz, a flat spectrum. Each octave spans twice the bandwidth of the one below it, so every octave carries 3 dB more energy than the last, which is why white noise reads as hiss piled at the top. Pink noise, falling 3 dB per octave, is the one that sounds even to human ears. Use white for hats and risers; reach for pink when checking mix balance.