METHODOLOGY 2009

🎯 A Distribution of Absolute Pitch Ability as Revealed by Computerized Testing

Patrick Bermudez and Robert J. Zatorre

Music Perception (2009) Vol. 27, Issue 2, pp. 89–101

📅 Accepted: June 24, 2009 👥 N=51 musicians 🔬 Computerized behavioral test 🏫 Montreal Neurological Institute, McGill University

🎯 Key Finding

Note-naming performance ran from perfect to random, with a large group in between. Using a computerized test that records accuracy, mean deviation in semitones and reaction time, Bermudez & Zatorre describe across 51 musicians “a range of behavior from perfect to random performance … including an important number of intermediate-level performers.” They are careful about how far that goes: the result is “a distribution of behavior across participants that one would hesitate to describe as bimodal,” and they grant that one could construe two small clusters — 11 participants with nearly perfect performance and 13 essentially at chance — which still leaves 27 of the 51 in between.

The strongest part of the argument is what happens under the harshest scoring. The authors applied to their own data the strictest criterion available — exact chroma only, the one that should most separate the groups — and report that it “still does not create easily resolvable subsets of participants.” The paper is aimed at Athos et al. (2007), who had read a bimodal distribution as evidence that AP could originate in one or a few genes.

📊 Study Design

Participants

  • N=51 musicians (39 females, 12 males)
  • 27 self-reported as AP possessors, 24 as non-possessors (NAP)
  • Average age: 23.1 years (SE = 0.52)
  • Mean age of training onset: 6.1 years (SE = 0.32)
  • Mean total training: 16.4 years (SE = 0.63)
  • Recruited from music faculties of two Montreal universities
  • All gave informed consent; approved by Montreal Neurological Institute ethics
  • 2 participants reclassified based on performance (1 self-reported AP scored 2 SD below AP mean; 1 NAP scored 2 SD above NAP mean)

Stimuli

  • 108 trials (36 notes × 3 intensity levels)
  • Range: C3 to B5 (3 octaves)
  • Based on A = 440 Hz equal temperament
  • Each note presented at 3 intensities: −1, −4, and −7 dB (to prevent loudness cues)
  • Synthetic multiharmonic tones: fundamental + ~9 harmonics (12 dB amplitude decrease between harmonics)
  • Duration: 1 second (50 ms linear onset and offset ramps)
  • 16-bit sampling depth
  • Presented at ~75 dB SPL via headphones

🎮 The Computerized Test Interface

Chroma Response (Step 1)

  • Circular wheel with 12 positions (all pitch classes equidistant from center)
  • Cursor resets to center after each trial (no positional bias)
  • All 12 responses equally accessible (unlike piano keyboard)
  • No timeout: self-paced (allows measurement of natural response speed)
  • Critical innovation: avoids keyboard familiarity confounds

Octave Response (Step 2)

  • After selecting chroma, indicate which octave (C to B range)
  • Color-coded bands in greyscale
  • Emphasized that exact grand staff position not required
  • “Simply click anywhere in the color band representing the octave”
  • Allows analysis of chroma accuracy and octave accuracy separately

Design Innovations (5 Key Advances)

  1. Multiharmonic synthetic stimuli: Equally unfamiliar to all participants (unlike piano/violin tones)
  2. Both chroma and octave judgments collected: Separates pitch class from pitch height
  3. Precise reaction times: 10 ms resolution (identifies strategy differences)
  4. Circular response interface: All 12 responses equidistant (no keyboard bias)
  5. Self-paced: Captures natural response speed (no artificial time pressure)

📈 Results

AP vs NAP Performance

AP Group (n=27*)
77% correct
MAD = 0.38 semitones · RT = 3,346 ms
NAP Group (n=24*)
15% correct
MAD = 2.48 semitones · RT = 7,586 ms

How to read these numbers. “Percent correct” credits only the exact chroma — there is no ±1 semitone tolerance — so chance is 8.3% (1 in 12). MAD is the mean distance in semitones between the answer and the correct note, ignoring the octave: given a possible range of 0 to 6 semitones of deviation, 0 is perfect and 3 is what a completely random responder scores. The NAP group’s 2.48 is therefore above chance, not close to the floor: as the paper puts it, they “are not responding completely arbitrarily, with a very broad mode centered on correct identifications.”

All differences highly significant: accuracy F(1,49) = 217.78, p < .001; MAD F(1,49) = 221.24, p < .001; RT F(1,49) = 30.59, p < .001. (*after reclassification of 2 outliers)

The Spread of Ability

  • Best performers: mean deviation below 0.1 semitone — the paper’s own figure for “the very best”
  • Intermediate performers: between 1.0 and 2.3 semitones of deviation — and frequently at chance on percent correct. The paper’s worked example is a pair of participants who both scored 8.3% (chance) yet had mean deviations of 1.44 and 2.81 semitones and reaction times of 4,905 and 9,124 ms. On the traditional score they are indistinguishable; on MAD and reaction time they are not remotely alike. The paper’s point is narrower: percent correct does not separate intermediate from weak performers as clearly as MAD does — the intermediates are still there, “an important number of intermediate-level performers”, even under the strictest scoring
  • Random performers: MAD around 3 semitones, which is the chance-level value (flat response distribution)
  • Two mini-clusters, and a crowded middle: the authors hesitate to call the distribution bimodal, but concede that “one could construe from our data two small clusters, one of 11 participants with nearly perfect performance and another of 13 participants with essentially random performance (though these are more apparent on the percent correct axis than the MAD axis; Figure 2).” That leaves 27 of the 51 musicians between the two extremes — the contingent a binary AP / non-AP label erases
  • Eight of the best participants were slow: eight of the participants scoring a mean deviation below one semitone had mean reaction times “ranging from 4 to 6 s and therefore many of their responses would have earned a score of 0” under the timed scoring schemes common in the literature. Good ears failed by the stopwatch. And speed alone does not identify AP: all participants with a mean deviation under one semitone answered within 6 s, “as did about half of the random performers”

Pitch Class Dependence (White-Key Advantage)

  • Diatonic notes (C major) identified more accurately and quickly than non-diatonic
  • Marginally significant interaction: F(1,49) = 3.72, p = .06
  • Driven by AP group: white keys significantly more accurate (Tukey HSD)
  • For RT: significant interaction F(1,49) = 23.12, p < .001 (AP faster on white keys)
  • The white-key advantage itself replicates Miyazaki 1988, 1989, 1990; Schellenberg 2001; Takeuchi & Hulse 1991
  • The advantage shrinks as proficiency rises: the ratio of correct white to black notes decreases as AP proficiency increases. The best participants get everything right about equally; leaning on the white keys is a marker of weaker performance
  • Pitch class A identified best overall: highest accuracy and fastest RT in the AP group. NAP participants showed the A advantage too, plausibly using it as a relative reference
  • But this diverges from Miyazaki (1990), who found C and G to be the most quickly and accurately identified. The authors attribute the difference to sample composition: Miyazaki’s musicians were all keyboard players, theirs spanned strings, brass and percussion, of which they write: “Perhaps the only commonality across them is the frequent use of A as a tuning reference” (see Miyazaki 1988)

Reaction Time as Key Dimension

  • Strong correlation: MAD vs log RT: r = .63, p < .0001
  • Better performers respond faster (not trading speed for accuracy)
  • Among NAP participants there is a non-significant trend toward longer response times at lower MAD — compatible with slower, alternative strategies such as computing from one or two memorised reference notes. The paper flags it as non-significant both times it raises it, and is explicit that separating that account from the rival one is not yet possible: “Greater numbers of participants and trials in future work will be required to tease apart those participants who are using a relative pitch strategy from one or a few dependable notes versus those who simply have broad, low-resolution chroma representations centered on the correct pitch”
  • Combined index (MAD × logRT) still shows no sharply separated groups
  • RT captures what % correct misses: two participants both at 8.3% correct (chance) had very different MADs (1.44 vs 2.81) and RTs (4,905 vs 9,124 ms)

Split-Half Reliability

  • MAD: r(49) = .99, p < .001
  • Log RT: r(49) = .98, p < .001
  • Exceptionally high reliability — the test is internally consistent
  • On that basis the authors suggest that “a shortened version would be suitably accurate for screening purposes.” They do not say how short; the reliability was computed by splitting the 108 trials into odd and even halves of 54

Octave: No Difference Between Groups

The test collects the octave judgement separately from the note name, and the result is worth stating plainly, because the page’s design section promises the analysis: AP and NAP participants did not differ in octave identification — accuracy M = 78% ± SE 3.16 vs M = 70% ± SE 3.35, F(1,49) = 2.83, p < .099 — nor in log reaction time for it (F(1,49) = 0.14, p = .71).

The AP advantage is specific to chroma (the note name), not to pitch height. The paper also heads off a misreading of its own histogram: the peaks at −12 and +12 semitones among AP participants are “a misleading impression resulting from the fact that AP participants are able to make octave errors as a direct consequence of being able to correctly identify chroma.” Only someone who got the note right is in a position to get the octave wrong.

Age of Training Onset

  • AP group started training significantly earlier: M = 5.46 years vs NAP M = 6.95 years; t(43) = 2.52, p = .02
  • MAD significantly correlated with training onset: r(43) = .46, p = .01
  • % correct correlated with onset: r(43) = .44, p = .002
  • Log RT correlated with onset: r(42) = .40, p = .007
  • Consistent with early-learning theory (Takeuchi & Hulse 1993)

🎯 The Target: Athos et al. (2007)

About a page of this paper’s introduction (pp. 90–91) is a point-by-point reply to a PNAS study that had reported “some degree of” bimodal distribution of AP ability and concluded, in the authors’ summary of it, “that their data resolve the question of whether AP participants lie at the extreme of a continuum of ability or form a distinct population (clearly the latter, in their opinion), and that AP could have its origins in one or a few genes.” Bermudez & Zatorre raise three objections.

  • Self-selection. The data came from an internet test and survey taken by well over 2,000 self-selected participants, of whom 44% were classified as having a particularly strong form of AP — a proportion far above any population estimate. People who believe they have AP are the ones who take an AP test online
  • The four-second cut-off. That test scored any response slower than four seconds as zero, which “risks exaggerating the separation between the best and worst performers by discarding intermediate performance.” This is not hypothetical here: eight of the best performers in the Montreal sample took 4 to 6 seconds per answer and would have been zeroed out
  • Bimodal behaviour is not evidence of a single gene. The most transferable argument in the paper: “a bimodal distribution of measured behavior does not necessarily map to simple underlying biology and could just as easily suggest a polygenic or multifactorial inheritance model in which an interaction of several genes and environment conspire to produce a variety of phenotypic manifestations.” And: “there is no necessity that it should be attributed to a single gene”

Their verdict is deliberately modest: “Given these and other related issues, it seems imprudent to declare the matter resolved.” Note what they are not claiming — not that AP is settled to be continuous, only that the case for two separate populations has not been made. The same paragraph names Carroll (1975) and Miyazaki (1988) alongside Athos as sources of the “two clearly resolvable populations” claim they are contesting.

💡 Why Scoring Method Matters

How you score an AP test changes what the distribution looks like. The paper walks through the options:

  • Strict % correct: credits only the exact chroma. It discards a large share of performance variance, and is therefore likely to exaggerate — in some cases perhaps even to create — an impression of bimodality
  • Semitone credit (3/4 point for ±1 semitone): the opposite failure. It blurs the distinction between people who are exactly right and people who are consistently close
  • Mean Absolute Deviation (MAD): the mean distance in semitones between the answer and the correct note, ignoring the octave. Scale 0 to 6, with 3 = chance. It should not be confused with a measure of consistency: someone systematically two semitones sharp is perfectly consistent yet scores MAD 2. Recovering that kind of systematic error is what an information-transfer approach does — the paper mentions the alternative but does not apply it
  • MAD + reaction time: the pair the authors find most informative, because time exposes the slow, effortful strategies that accuracy alone hides

The decisive test. Having made the argument that harsh scoring can manufacture bimodality, the authors then applied harsh scoring to their own data — and it did not manufacture anything: “even though scoring strategies that discard important proportions of performance variance are likely to exaggerate and, in some cases, perhaps even create the impression of bimodality in a distribution of responses, applying these harsh scoring approaches to our data still does not create easily resolvable subsets of participants. We can therefore say with some confidence that there is true heterogeneity of AP competence in our sample.”

That is a stronger result than “the bimodality was a scoring artifact.” The scoring argument stands as a warning for reading other studies; it is not the explanation of this one. Here the spread survived the criterion that most favours a clean split.

💬 Critical Analysis

Strengths

  • Novel computerized interface eliminates keyboard familiarity bias
  • Captures both accuracy AND speed (multidimensional assessment)
  • Exceptional split-half reliability (r = .99)
  • Separates chroma and octave judgments
  • Controls for loudness cues (3 intensity levels)
  • Uses synthetic stimuli (no instrument-familiarity confound)
  • Rigorous statistical analysis with reclassification of outliers
  • Directly addresses long-standing controversy about bimodality vs continuum

Limitations

Stated by the authors:

  • Because the test has no time limit, they “cannot know for certain how participants’ responses might have differed had they been attempting to respond inside a typical 2 to 4 s time limit, though we can reasonably expect a higher error rate”
  • The reaction-time trends among intermediate performers are non-significant, and telling a relative-pitch strategy apart from broad, low-resolution chroma representations would need more participants and more trials
  • The test is a note-naming task, so it only sees AP in people already trained to name notes: “The obvious disadvantage is that it cannot identify any form of absolute pitch that manifests itself outside the musical domain or independently of musical instruction.” This is a structural blind spot of nearly the whole AP literature, and the paper says so directly
  • No one has yet separated how much of the white-key advantage is pitch perception and how much is familiarity with the note names: “no work to date has disentangled the relative contributions of pitch and note name/notation processing requirements and familiarity”
  • The response method trades realism for control: “Although somewhat lacking in ethological validity, the chroma response method chosen ensures that the mouse pointer resets to the center…”

Editorial observations (ours, not the paper’s):

  • N = 51 — a moderate sample; a larger one would characterise the intermediate zone better
  • Volunteers recruited from two Montreal music faculties, all with at least 10 years of training. The recruitment ad deliberately sought musicians with and without AP, which the authors credit for the number of intermediates captured — but it is not a population sample
  • No non-musician comparison group, and no participants from outside Montreal
  • Split-half reliability only; no test–retest data
  • Single session, no longitudinal component: this is a snapshot of ability, not a study of how it was acquired

Impact & Legacy

The best measuring instrument in this collection. Of the 21 studies gathered here, this is the one whose test design holds up best — and it anticipates one recommendation the 2025 systematic review would go on to make: a response method that does not depend on familiarity with a piano keyboard or with notation. The match is only partial, though — that same review recommends piano timbre, distractor stimuli between trials, and a 4,000 ms response window, and this test uses synthetic tones, no distractors and no timeout. Ignore the rest of this sentence, which the 2025 systematic review would go on to make: the circular chroma interface (no keyboard bias), no timeout, no ±1 semitone tolerance in the primary score, and no a priori threshold dividing possessors from non-possessors before the data are in. It reports mean deviation and reaction time as continuous measures and lets the reader see where the participants actually fall.

The 108-trial paradigm became a methodological lineage: Bairnsfather et al. (2025) identify a family of derived tasks, noting that this trial count is “shared across multiple paradigms, including those derived from Bermudez and Zatorre (2009)” and Oechslin, and list mean absolute deviation among the alternative scoring methods in use. Sixteen years on, the standardisation this paper asked for has still not arrived: that review reports that “further progress in this area has been hindered by a lack of consensus regarding the AP phenotype.”

The key observation — that measured AP ability forms a continuum with substantial intermediate performance — has been reported repeatedly since. It is not settled, and this paper says so: “this remains a controversial point.” The 2025 systematic review in this collection still calls AP “a rather discrete behavioural trait” and proposes taxometric analysis precisely because no one has tested whether the underlying trait is continuous or categorical (Bairnsfather et al. 2025).

📚 Related Studies

🔗 Access & Resources

📊 Citation

  • DOI: 10.1525/mp.2009.27.2.89
  • Journal: Music Perception, Vol. 27, Issue 2, pp. 89–101
  • ISSN: 0730-7829 (print), 1533-8312 (electronic)
  • Affiliation: Montreal Neurological Institute & BRAMS Laboratory, McGill University