๐ต Absolute Memory for Musical Pitch: The Levitin Effect
๐ Study Overview
Absolute memory for musical pitch: Evidence from the production of learned melodies
Daniel J. Levitin (University of Oregon, Eugene)
Perception & Psychophysics, 1994; 56(4):414-423
๐ฏ Research Question
Do ordinary people (without absolute pitch) possess stable, long-term memory representations for the actual pitches of familiar songs? Or do we only remember melodies (relative pitch information)?
Levitin reframed the fundamental question about AP: instead of asking "Why do so few people have absolute pitch?", perhaps the right question is "Why doesn't everybody?" — since cells that respond to specific frequency bands exist at every level of the auditory system. The information is there; the mystery is why so few can access it consciously.
๐ฌ Methodology
Participants
- Total N = 46 Stanford University students (undergrad + grad)
- Age range: 16-35 years (mean 19.5, mode 18)
- Mixed musical backgrounds (trained and untrained)
- 2 claimed to possess AP (not independently tested)
- 43 completed both trials (3 discontinued after Trial 1)
- Norming study: 250 additional students rated familiarity with 50 songs to select the best-known stimuli
Procedure
- Song selection: 58 CDs of popular songs chosen via norming study (600+ songs available)
- Examples: "Hotel California" (Eagles), "Like A Prayer" (Madonna), "Every Breath You Take" (Police)
- Task: each participant picked 2 familiar songs and sang, hummed or whistled each one from memory, starting wherever they liked (two trials, two different songs)
- Recording: Digital audio tape (DAT) for accurate pitch preservation
- Analysis — only the FIRST TONE was scored: the first three tones each subject sang were coded, but a repeated-measures ANOVA found no effect of tone position [F(90,2) = .58, p = .56], so “the analyses are based on subjects’ first-tone productions” (p. 416). Every accuracy figure on this page therefore refers to that single tone, compared by FFT analysis (Spectro) against the equivalent tone on the CD — not to a whole song sung in tune
- Measurement resolution: subject productions were measured to within 3 cents and quantised to the nearest semitone; the CD target pitches were coded with hardware tuners (Seiko ST-1000, Conn Strobotuner) accurate only to within a semitone
- Octave normalization: octave errors were not penalized — what is scored is pitch class, not absolute pitch. Deviations larger than half an octave were folded to the nearer direction — Levitin’s own example is that singing D3 against a C4 target was coded as +2 semitones, not as a hit. Levitin notes this follows standard practice in AP research (Miyazaki, 1988; Takeuchi & Hulse, 1993; Ward & Burns, 1982)
Why Popular Songs?
Contemporary popular songs are ideal stimuli because:
- Typically encountered in only one version by the original artist
- Heard hundreds of times in the same key
- Unlike folk songs ("Happy Birthday"), which are performed in many different keys
- Provides an objective standard for "correct" pitch
๐ Key Findings
1. Accuracy on Individual Trials (first tone sung)
| Metric | Trial 1 (N=46) | Trial 2 (N=43) |
|---|---|---|
| Exact pitch (0 semitones error) | 26% (12/46) | 23% (10/43) |
| Within 1 semitone | 57% (26/46) | 51% (22/43) |
| Within 2 semitones | 67% (31/46) | 60% (26/43) |
Chance levels: hitting the exact pitch class by chance is 8.3% (1 of 12 classes). The paper does not state chance for the widened criteria; with 12 classes, ±1 semitone admits 3 of them (25%) and ±2 semitones admits 5 (42%) — so the 57% and 67% rows sit far closer to chance than the exact-pitch row does.
2. The "Levitin Effect" - Consistency Across Trials
For the 43 subjects who completed both trials:
- 12% (5/43) sang correct pitch on BOTH trials (chance = 0.7%)
- 40% (17/43) sang correct pitch on at least one trial (chance = 17%)
- 44% (19/43) came within 2 semitones on both trials
- 81% (35/43) came within 2 semitones on at least one trial
The paper reports chance only for the exact-pitch criterion (17% for at least one of two trials, 0.7% for both). The two ±2-semitone rows use a much looser test — see the chance note above — so they are not comparable to the exact-pitch figures.
If people had no pitch memory at all, the errors would be spread uniformly around the circle of pitch classes. They were not: they approximated a circular-normal (von Mises) distribution rather than a uniform one (Rayleigh test: Trial 1 r=.48, p<.001; Trial 2 r=.30, p<.02). The distribution peaks at the correct pitch, but its mean sits below the target — Trial 1 mean = −0.98 semitones (SD 2.36), Trial 2 mean = −0.40 (SD 3.05). That downward displacement is the flat bias described below.
3. Consistency Between Trials
Subjects who hit the pitch on Trial 1 were more likely to hit it again on Trial 2:
- P(correct Trial 2 | correct Trial 1) = 42% vs. 23% overall — a marginal, one-tailed result (z=1.66, p<.05)
- 31 subjects (72%) performed the same way on both trials — but 26 of those 31 were consistent misses; only 5 were consistent hits, so the concordance is dominated by consistency in error
- Yule’s Q = .58 (p=.01). Levitin’s own reading is more guarded than the figure looks: the concordance measures are “reasonably high, but still, many people did not perform consistently” (p. 420)
4. The "Flat Bias" (Lounge Singer Effect)
When subjects made errors, they tended to sing flat — most of the errors fall below the correct pitch, which is why the mean of the error distribution is about a semitone under the target on Trial 1. Levitin notes this mirrors the "lounge singer effect" widely observed by vocal instructors, where amateur singers tend to undershoot pitches. It may also reflect range limitations — many popular singers (Madonna, Sting, Prince) have unusually high voices.
5. No Reliable Relation to Musical Training
No reliable relation was found between performance and:
- Musical training or background
- Gender, handedness, age
- Amount of time spent listening to music
- Amount of time singing (including in the shower or car)
This is a null result in a sample of 43 with no reported power analysis — absence of evidence, not a demonstration of independence. Levitin himself writes only that the ability “seems independent of a subject’s musical background” (p. 421).
๐ก Main Conclusions
"The finding that 1 out of 4 subjects reproduced pitches without error on any given trial, and that 40% performed without error on at least one trial, provides evidence that some degree of absolute memory representation exists in the general population." — Levitin, 1994, p. 418
The Two-Component Theory of Absolute Pitch
To make sense of the evidence, Levitin proposed — his own verb is “posit” — that absolute pitch consists of two distinct component abilities. This is a theory advanced to explain the data, not a pair of measured prevalences: only the first component was tested here.
| Component | Definition | Status in this study |
|---|---|---|
| 1. Pitch Memory | Ability to maintain stable, long-term representations of specific pitches and access them when required | Measured here. 40% hit the exact pitch on at least one of two trials (chance 17%); 12% on both (chance 0.7%); 23–26% on any single trial (chance 8.3%) |
| 2. Pitch Labeling | Ability to attach meaningful labels to pitches (Cโฏ, A440, Do) | Not measured. Its absence was inferred from self-report — all but 2 of the 46 subjects said they did not have AP, and Levitin writes that they “presumably” did not have the ability to label pitches (p. 421). The paper never estimates how common labeling is; the familiar “~1 in 10,000” is the prevalence of AP as a whole, quoted from Profita & Bidder (1988) and Takeuchi & Hulse (1993) |
Key insight: on this model, “true” AP possessors would have both abilities, while pitch memory alone might be widespread among people who never acquired labeling — “possibly because they lack musical training or exposure during a critical period” (p. 421). The hedges are Levitin’s own: he tested neither the labeling component nor the critical period.
Key Implications:
- Pitch memory is not rare: 40% sang the correct pitch on at least one of two trials (chance 17%) and 12% on both (chance 0.7%). Note that the analysis scored only the first tone of each production
- Dual representation (proposed): the results are consistent with memory holding both absolute pitch and relative intervallic information — but Levitin puts it conditionally (“if people do maintain both kinds of information in memory, this would suggest that a dual representation exists in memory for melody”, p. 415) and closes the paper with an explicit warning: “One should be cautious, however, about jumping to conclusions” (p. 421)
- Long-term stability: the representations survived “a long period of time with much intervening distraction” (p. 421) — the study never measured how long, since it did not record when each subject last heard the song
- Challenges the all-or-nothing view: Levitin argues AP is not one mysterious faculty but two separable abilities, one common and one rare. His model is a dissociation between components — though he does raise the wider possibility in the paper: “Perhaps everybody does have AP to some extent”
๐ Why This Study Matters
This study is the most-cited source for the claim that some pitch memory is widespread — though its own abstract frames the result as “a convergence with previous studies” rather than as a first in the general population, not confined to rare AP possessors. The finding — that 40% of ordinary people sang the first note of a familiar song at the correct pitch on at least one of two attempts, against a chance level of 17% — became known as the “Levitin effect”. On any single attempt the hit rate was 23–26% (chance 8.3%), and 12% hit on both attempts (chance 0.7%).
Paradigm Shift:
- Old view: AP is extremely rare (~1 in 10,000) and mysterious
- New view (proposed by Levitin, not measured here): pitch memory is common; pitch labeling is rare
- Question reframed: from “Why do so few people have AP?” to “Why doesn’t everybody?” (p. 414)
Impact on the Field:
- Multi-lab replication: Frieler et al. (2013) ran a replication across European labs and the effect appeared. Any single percentage quoted alongside Levitin’s needs its criterion attached — his headline 40% is “correct on at least one of two trials”, while his per-trial rate was 23–26%
- Inspired new research: Spawned studies on "latent absolute pitch," "quasi-absolute pitch," and pitch shift detection
- Theoretical implications: Supports dual-representation theories of melody (absolute + relative)
โ ๏ธ Limitations & Future Directions
Study Limitations
- Small sample: N=46 from single university (Stanford)
- Self-selected songs: Participants chose familiar songs (possible bias)
- Production vs. identification: Singing accuracy may underestimate pitch memory (vocal production problems)
- Muscle memory cannot be fully separated: Levitin is explicit that “there is always some degree of muscle memory involved in the vocal generation of pitch”, and that the initial pitch of a vocal tone is “by necessity, determined by muscle memory” (p. 421) — and that initial tone is the one scored here. His argument is narrower: muscle memory alone is imprecise. Ward & Burns (1978) denied auditory feedback to trained singers and errors reached 3 semitones; Murry (1990) found average errors of 2.5 semitones, with the worst cases up to 7.5, in the first five waveforms before feedback could act. Zatorre & Beckett (1989) argued that even true AP possessors rely on muscle memory to some extent. Levitin’s conclusion is hedged accordingly: the experiment “seems to have tested, as well as possible, subjects’ memory for particular auditory stimuli” (p. 421)
- Timbre cues: Unclear if pitch is accessed directly or derived from timbral memory of the original recording
Explanations for Errors:
Subjects who came within 1-2 semitones may have had good pitch memory but failed due to:
- Pitch memory with only semitone resolution (still functional AP per Miyazaki 1988)
- Production problems (can't match internal representation vocally)
- Self-monitoring deficits (can't compare own voice to internal representation)
- Exposure to the song in a different key (cassette speed varies by around a semitone; CD players are pitch-stable) — tested and not supported: subjects were asked where they had heard each song before, and a correlational analysis showed no relation between the source of learning and accuracy (p. 420)
- Tonal interference between trials: Singing the first song may establish a tonal center that biases the second production — Tsuzaki (1992) showed that even AP possessors' internal standards are subject to interference from preceding scales
Future Research Needed
- Larger samples across diverse populations
- Recognition tasks (eliminate vocal production confounds)
- Longitudinal studies of pitch memory stability
- Investigate which song features (timbre, tempo, lyrics) aid pitch memory
- Test whether pitch labeling training can convert pitch memory to full AP
๐ Related Research
- Replication: Frieler et al. (2013) — a multi-lab replication of the Levitin effect in Europe; the figures depend on which criterion is compared
- Precursor: Ward (1990) - Informal taped diary showed similar pitch memory effect
- Stability: Halpern (1989) - Showed pitch imagery is stable across occasions (within 2 semitones)
- Pitch shift detection: Schellenberg & Trehub (2003) reported above-chance detection of 1-semitone key shifts in familiar songs. Not audited for this site — and a group effect above chance is not the same as “most listeners”
- Latent AP: Deutsch, Kuyper & Fisher (1987) and Deutsch (1991, 1992) — the tritone paradox shows pitch-class-dependent judgments in listeners who cannot label pitches. These are the papers Levitin cites; Deutsch et al. (2004), covered separately on this site, is the tone-language study, not the tritone work
๐ Access Full Study
๐ Full Citation
Levitin, D. J. (1994). Absolute memory for musical pitch: Evidence from the production of learned melodies. Perception & Psychophysics, 56(4), 414โ423. https://doi.org/10.3758/BF03206733