Masking
The quiet tone is still in the signal. Your ear will not report it. So the bits spent on it buy nothing.
The signal, and what a masker hides
1 Hear: a tone, on its own, plainly audible
Twelve components, fixed, so the only thing that moves is the masker. Levels are relative to the loudest of them; nothing on this page is an absolute sound pressure, because a browser cannot know what your speakers do with it.
2 Mask: a louder one beside it, and the first is still there and gone
A masker at 900 Hz, 6 dB, sitting in the critical band that is 147 Hz wide there.
3 Discard: the encoder spends no bits on what the ear cannot reach
components 12kept 10dropped 2bits 176 of 208
Of 12 components, 2 now sit under the threshold this masker raises, so the coded version leaves them out and spends 176 bits where the full signal spends 208. Whether that is audible is the question below, and it is yours rather than this page's.
4 Compare: the same passage at both rates, switched without warning
The third button is the one that settles it. Everything else shows you a picture of a claim; that one plays the components the coder threw away, on their own, so you can hear whether there was anything there.
And the verdict this page will not make for you. A is everything, B is the coded version, and X is one of the two picked by a coin you cannot see. Say which X was, several times over, and the answer stops being an impression: the score line works out how often guessing would have done as well.
Checked when this page loaded: across 20 Hz to 16,000, Zwicker and Terhardt's Bark formula and Traunmüller's — different functions for the same scale — disagree by at most 0.74 Bark, at 15,312 Hz, on a scale spanning 24.1. One Bark is 104 Hz wide at 200 Hz and 1,452 Hz wide at 8 kHz, measured off the same formula this page uses rather than read from a table.
What an ear does with two tones at once
A loud tone does not merely sound louder than a quiet one beside it. It raises the level at which the quiet one becomes detectable at all, and it does so across a band rather than at a single frequency, because the ear resolves frequency in bands rather than points. Inside that band the quiet tone is not faint. It is absent.
That is the whole opportunity. An encoder that knows where the masker is knows that the components underneath it will not be reported by the listener, and it can spend nothing on them. What comes out is not an approximation of what you heard; it is what you heard, with the parts you did not hear left out.
Why the bands, and how wide they are
Nearby in frequency has to mean something measurable, and the measure is the Bark scale: about twenty-four bands across the range of hearing, each roughly the width over which the ear integrates. They are not equal in hertz. One Bark is around a hundred hertz wide down at 200 Hz and around one and a half kilohertz wide up at 8 kHz, which is why a masker at the top of the range hides so much more of the spectrum than one at the bottom. Both numbers on this page are measured off the same formula the model uses, not read from a table beside it.
The page holds two different published formulas for that scale, Zwicker and Terhardt's from 1980 and Traunmüller's from 1990, and checks them against each other when it loads. They are different functions, so agreement is evidence rather than a tautology, and the line at the bottom says how far apart they got rather than claiming they match.
Why this page refuses to tell you what you can hear
Masking depends on how loud the sound actually is at your ear, and the page has no idea. It does not know your volume setting, whether you are on headphones or a laptop speaker, what the room is doing, or what your hearing is like at four kilohertz. A sentence telling you that you cannot hear a given component would be a claim about a person it has not measured.
So it reports only what it did: which components fell under the threshold its model raises, and how many bits that saved. The rest is a listening test, and the two buttons above record your answer rather than predict it. If you did hear a difference, that is a real result about this model being too aggressive for you, and it is the reason a shipping encoder uses a more conservative one than this.
What is real here, and what is not
The spreading model is two straight lines
The threshold falls away from the masker at 27 dB per Bark downward in frequency and 12 dB per Bark upward, which captures the one asymmetry that matters — a masker hides what is above it more readily than what is below. A real coder's upward slope moves with the masker's level and with frequency, and it distinguishes a tonal masker from a noise-like one, because they mask differently. None of that is here. This is enough to show the shape of the argument and nowhere near enough to reproduce an encoder.
There is no absolute threshold of hearing on this page
There is a well-known approximation for it, and it is deliberately absent. It is expressed in dB SPL, and a browser has no idea what sound pressure its output produces, so every absolute number it gave would be uncalibrated by an unknown amount. Beyond that, the text it is usually attributed to is paywalled and nobody here has read it, and citing a document you could not open is the failure this studio spends most of its effort avoiding. What is on the page instead is relative, which is what a listening test can actually check.
Twelve sine components are not music
A real encoder works on a filter bank of hundreds of subbands, frame by frame, on signals that change. This is twelve steady sine tones, chosen so that a masker can be moved across them and the effect seen one component at a time. Nothing here streams, nothing adapts, and the bit counts are twelve times sixteen rather than anything a file would contain.
The playback is scaled so it cannot clip
Thirteen sine tones added together can exceed what the output can carry, and a loud masker took the sum to 1.72 times full scale. Everything over is clamped, and a clamped sine is a sine plus harmonics this model knows nothing about. That would have mattered most where it is least welcome: the coded version has fewer components, so it clipped less than the full one, and above +11 dB there are settings where A distorts and B does not. A listener could have passed the blind test on that difference and the page would have called it evidence about masking. So one gain, worked out from the full set at whatever the masker is doing, scales everything played at that setting. A, B and the discarded part are scaled by the same number, which is the only property the test needs. Raising the masker past that point makes the mix no louder, and that is the reason.
The bit count is a count, not a compression ratio
It is components kept multiplied by a nominal sixteen bits each, plus the masker, which makes the arithmetic something you can check by hand. It is not comparable to an MP3 bitrate: real coding gains come mostly from quantising what is kept more coarsely under the same threshold, not from dropping components outright, and none of that quantisation is modelled here.
What the ABX score does and does not prove
The percentage beside your score is the exact upper tail of a binomial: if every answer were a fair coin, that is how often you would score at least this well. It is not a claim that you can or cannot hear the difference, and a run of five means very little either way. Five out of five is about 3%, nine out of ten is about 1%, and anything at or below one in twenty is the usual line for calling a result something other than luck.
It also refuses to run where it would mean nothing. If the masker is hiding no components, the coded version is the full one and A and B are the same sound; the test says so instead of scoring guesses on two identical signals. Moving the masker mid-run clears the score for the same reason, because trials of different signals do not belong in one total.
A high score is evidence against this page's model rather than against masking. It would mean the two slopes above are discarding components you can still hear, which is exactly why a shipping encoder uses a more conservative threshold than this one.
Temporal masking is missing entirely
A loud sound also masks quieter ones just before and just after it in time, which is a large part of what real encoders exploit and is why a pre-echo is the artefact they fight hardest. Everything on this page is simultaneous. The tones start and stop together.
Sources
- H. Traunmüller, “Analytical expressions for the tonotopic sensory scale”, Journal of the Acoustical Society of America 88(1), 1990, pp. 97–100, and E. Zwicker and E. Terhardt, “Analytical expressions for critical-band rate and critical bandwidth as a function of frequency”, JASA 68(5), 1980, pp. 1523–1525. These are the two papers this page names, cited where they were published instead of through an article about them. Titles, authors, volumes and pages were checked against Crossref. Neither was read. Both are behind the Acoustical Society, and the page takes the attribution on trust while saying so.
- Bark scale, for both formulas the page cross-checks and their attributions to Zwicker and Terhardt (1980) and Traunmüller (1990)
- Julius O. Smith, CCRMA, for the critical-band edges the scale describes