SYSTEMATIC REVIEW 2025

πŸ“Š Conceptual Coherence but Methodological Mayhem: A Systematic Review of Absolute Pitch Phenotyping

Jane E. Bairnsfather, Miriam A. Mosing, Margaret S. Osborne, Sarah J. Wilson

Behavior Research Methods (2025) 57:61

Sample: N=160 studies (23,221 participants) | Type: Systematic Review | DOI: 10.3758/s13428-024-02577-z

🎯 Key Findings

πŸ“š Dataset

160 studies reviewed (1992–2024)
23,221 participants summed — the authors warn this "does not account for potential participant overlap among studies," so the number of distinct people is smaller
~6,500 classified as AP (an estimate; the paper reports 6,493 in the text and 6,520 in Table 3)
10,222 non-AP participants

βœ… Conceptual Agreement

99% agreement on the AP definition
In the paper's words: "AP refers to the ability to identify notes without a reference tone, with some also including the ability to produce notes without reference"
150 of 151 studies agreed

⚠️ Threshold Chaos

Range: 20-100% accuracy
Mean: 77% (SD=20)
59% of studies specified threshold
Enormous variability!

πŸ“Š Actual Performance

AP group: 85.9% mean (raw scores)
Non-AP: 17.0% mean (chance=8.3%)
With semitone credit: 89.1% vs 24.5%

πŸ“– Study Overview

Systematic review published 21 January 2025 — a broad survey of how absolute pitch studies actually measure AP: 160 studies, 23,221 participants, 157 pitch-naming tasks.

It is not a meta-analysis. The authors did not compute formal heterogeneity measures for effect size, and give no reason for it: "although we did not employ formal heterogeneity measures for effect size… the data in this review nevertheless point to a heterogeneous understanding of AP." Its verdict is a diagnosis of an immature field, not a discovery.

The Problem

Despite 30+ years of intensive research and near-universal agreement on what AP is conceptually, the field has no gold-standard task to measure it. Result: methodological mayhem - 160 studies using wildly different methods, thresholds, and scoring systems.

The Impact

This heterogeneity cripples the field:

  • Findings aren't comparable across studies
  • Replication is nearly impossible
  • Genetic research stalls (can't define the phenotype)
  • Same participant could be "AP" in one study, "non-AP" in another

πŸ”¬ Methodology

Search Strategy

Databases: Scopus, PsycInfo, ProQuest Music Periodicals, Music Index, JStor
Search terms: "absolute pitch" OR "perfect pitch"
Original search: 31 October 2019, restricted to work published since 1992 — a 30-year span at the time
Repeated: 31 January 2022 and 23 May 2024, which extends the corpus to 1992–2024

Inclusion Criteria

  • Empirical, peer-reviewed original research
  • AP as primary outcome (in title/abstract)
  • Neurotypical adults with normal hearing
  • Pitch-naming task AND/OR self-report
  • English language only

Data Extracted

Definition How AP was conceptually defined
Task parameters Pitch range, timbre, # trials, stimulus duration, response window
Scoring method Raw accuracy vs. semitone error credit
Thresholds Accuracy cutoffs for AP classification
Performance Mean scores for AP and non-AP groups with 95% CIs

πŸ“ˆ Detailed Results

1. Conceptual Definition (99% Agreement) βœ…

Nearly universal consensus. The paper's formulation: "AP refers to the ability to identify notes without a reference tone, with some also including the ability to produce notes without reference" — the production half is often dropped when this finding is quoted.

Only 1 of 151 studies deviated, defining AP by long-term pitch memory instead (Wayman et al., 1992).

2. Task Heterogeneity πŸŒͺ️

Parameter Reported in Variability
Number of trials 96.8% (152/157) Range: 1-960 trials (median=60)
Stimulus timbre 96% (151/157) 83% used sine or piano tones
Pitch range 89% (139/157) 1 octave to 8+ octaves
Response method 76% (119/157) Written, keyboard, button press, verbal
Stimulus duration 76% (119/157) 100-3000ms (mode=1000ms)
Response window 64% (100/157) 1000ms to self-paced

3. Scoring Methods

Raw scores only: 73% of tasks (110/151)
Semitone error credit: 27% (41/151)

Credit for semitone errors ranged from 0.25 to full point, or varied by participant age. This alone makes studies incomparable.

4. Accuracy Thresholds - The Core Problem ⚠️

Studies specifying threshold: 59% (95/160)
Range: 20% to 100%
Mean (raw scores): 77% (SD=20, median=85%)
Mean (with semitone credit): 71% (SD=16, median=68%)

Note that the two means above are not comparable with each other: the paper reports the two threshold means separately in the text, and plots mean performance in separate figures, "given these metrics are not directly comparable."

The absurdity, using two real thresholds on the same scale: the most lenient raw-score cutoff found was 20% (Maeshima et al., 2018) and the strictest was 100% (Matsuda et al., 2013). A participant scoring 75% is "AP" in the first study and "non-AP" in the second — and "quasi-AP" in a study that defines intermediate categories at all, which only 18 of 95 (19%) did.

Four criticisms the review makes that are easy to miss

  • Tiny samples: the median AP group in an individual study had 16 participants
  • The control group usually has no threshold at all: only 32 of the 95 studies with an AP threshold (34%) also specified one for non-AP performance, so "the performance of the comparison group cannot always be gauged from the published work"
  • The non-AP group is contaminated: "above-chance participants are sometimes included in the 'non-AP group', which reduces the discriminatory power of studies." The evidence is in the numbers above — the non-AP mean is 17.0%, twice chance (8.3%)
  • The method presupposes its own answer: the use of a priori thresholds, in the authors' words, "assumes that AP can meaningfully be divided into discrete categories, perpetuating dichotomous AP/non-AP classification in a somewhat circular manner between measurement and conceptualisation"

5. Task Parameter Effects on Phenotype

Parameter Effect on AP Performance Significance
Pitch range No correlation (r=-0.14, p=.320) Not significant
Timbre Piano > Sine tones t(43.97)=5.06, p<.001 ⭐
Number of trials Apparent negative correlation (r=−0.64) that the authors attribute to an artefact: it is "largely driven by the high degree of variability across studies using 108 trials," a trial count shared by paradigms whose AP thresholds run from 40% to 90% p<.001, but the paper states it does not imply "that actual participant performance decreases as trial numbers increase"
Stimulus duration No correlation (r=0.00, p=.974) Not significant
Response window No correlation (r=-0.26, p=.100) Not significant
Distracter sounds Lower accuracy with distracters t(45.92)=2.40, p=.021 ⭐

6. Publication Trees - Limited Replication

Tasks used: 157 pitch-naming tasks in total — not 157 distinct designs: 95 of them were replications or adaptations of earlier tasks
Based on prior work: 61% (95/157)
Novel/uncited: 39% (62/157)

Critical finding: Limited replication across research groups. Tasks replicated within groups (same lab uses same task), but no cross-lab standardization.

Six influential "source tasks":

  1. Lockhead & Byrd (1981)
  2. Miyazaki (1990)
  3. Baharloo et al. (1998)
  4. Deutsch et al. (2006)
  5. Bermudez & Zatorre (2009)
  6. Oechslin et al. (2010)

But even "replications" introduce modifications (timbre changes, trial number adjustments, etc.).

πŸ’‘ Recommendations for Gold-Standard Task

Based on analysis of 160 studies, the authors propose a standardized pitch-naming task:

Parameter Recommendation Rationale
Timbre Piano tones Contextually relevant, ecologically valid, better performance than sine tones
Pitch range 3 octaves — the authors recommend the span, not specific notes Balances content validity with practical trial length. The only note anchor the paper gives is the central octave C4–B4, which nearly every task already includes
Trials β‰₯5 per chroma (60 minimum) Captures performance variability, allows reliability assessment
Stimulus duration 1000ms Most common, maximizes comparability
Response window 4000ms (excluding stimulus) Sufficient time without rushing, commonly used
Response method Button/key press or screen label Accessible to all, enables RT capture, no music reading required
Distracter stimuli Yes (type unspecified; brown and white noise are what the literature used) Prevents relative pitch strategies across trials
Scoring Report BOTH raw and semitone-credit Enables cross-study comparison
Threshold At/near chance (8.3%) for non-AP Captures full spectrum including intermediate phenotypes (QAP)

Beyond the Gold-Standard Task

Authors advocate for data-driven phenotype characterization:

  • Use taxometric analysis to test discrete vs. continuous models
  • Employ multiple AP tasks (not just one) to capture phenotypic diversity
  • Investigate contextual factors (timbre specificity, range limits)
  • Move away from arbitrary a priori thresholds
  • Develop taxonomy of AP phenotypes empirically

🌍 Implications

For Research

  • Genetic studies: Can't find genes without well-defined phenotype
  • Replication crisis: Heterogeneous methods β†’ non-comparable findings
  • Meta-analyses: the authors did not compute formal heterogeneity measures for effect size, because raw scores and semitone-credit scores "are not directly comparable." Their wording is that heterogeneous methods "make it difficult" to compare studies, not that comparison is impossible
  • Field maturity: High heterogeneity = immature field (Linden & HΓΆnekopp, 2021)

What This Review Says About Learning AP

The authors take a position in two sentences, and neither was on this page before.

  • On training: AP "in most studies cannot be reliably trained" — citing Bittrich et al. (2015), Brady (1970), Cuddy (1968, 1970), Gregersen et al. (1999), Leite et al. (2016), Profita & Bidder (1988) and Sakakibara (2014). They register Van Hedger et al. (2019) as the exception, with an "although see"
  • On a critical period: "there is strong evidence supporting a critical or sensitive period for AP acquisition, including early practice on the piano" — which leads them to suggest AP is "a contextually learned behavioural skill rather than a purely psychophysical phenomenon"

Worth noting for readers of this site: one paper in this collection appears in that first list, on the side of “cannot be reliably trained”: Cuddy (1970). The Gregersen held here is the 1998 editorial, not the Gregersen et al. (1999) the authors cite.

For Clinical/Educational Applications

  • No validated diagnostic tool exists
  • Self-report: an open question, not a verdict. "The validity of self-report as a measure of AP ability is a useful question for further research, though first requires consensus regarding the phenotype that self-reported AP possessors claim to have"
  • Training studies were excluded from this review — "studies that attempted to either teach AP to novices or to pharmacologically alter pitch perception were excluded" (6 removed as AP training; a further 7 as pharmacological). Nothing here speaks to how training studies measure their outcomes

For Understanding AP Itself

The review shows that pitch-naming performance is continuous — "pitch-naming ability is a dimensional trait, with scores lying along a spectrum from chance to ceiling, regardless of participant classification into AP and non-AP groups."

That does not settle whether the AP phenotype is discrete. The authors keep the two apart and say the second question is untested: "data-driven techniques such as taxometric analysis… should be used to assess the extent to which these phenotypes are discrete." In their own introduction they describe AP as "a rather discrete behavioural trait," and note it "has not been established" whether AP and quasi-AP sit on one continuum. The question is open.

  • Performance spans from chance (8.3%) to ceiling (100%)
  • Intermediate phenotypes (QAP, partial AP) exist but poorly characterized
  • Contextual factors matter (timbre, distracters); pitch range showed no significant effect
  • Multiple phenotypes likely: "universal" vs. "limited" AP (Bachem, 1937)

⚠️ Limitations

  • Scope: only studies where AP was the primary focus. Self-report-alone studies were included; what falls outside are studies where AP was not the primary focus — which the authors note are more likely to rely on self-report
  • Language: English-language studies only
  • Population: Neurotypical adults (excludes autism, synesthesia, children)
  • Tasks: Focused on pitch-naming (excludes novel AP measures like pitch production, go/no-go tasks)
  • Publication bias: Grey literature, theses, conference proceedings excluded
  • Not preregistered: the authors state the review was not preregistered

🧠 Theoretical Framework

The Paradox

Conceptual coherence: Everyone agrees what AP is
Methodological mayhem: No one measures it the same way

Why This Matters

"To move AP research to a more mature field of study, we must explore the sources of this heterogeneity and address them from both a methodological and theoretical perspective."

Path Forward

  1. Immediate: Adopt gold-standard task for comparability
  2. Short-term: Use data-driven methods to characterize phenotypic variability
  3. Long-term: Develop empirically validated taxonomy of AP phenotypes

Connection to Musicality Genomics Consortium

This review is timely for the MGC's mission (https://www.mcg.uva.nl/musicgens/) to develop "scalable and robust phenotypes" and harmonize "existing measures of musicality phenotypes."

πŸ”— Connection to Other Research

Earlier Reviews

  • Takeuchi & Hulse (1993): the classic review, cited here as the origin of the pitch-naming task as the standard measure of AP, and again on the decline of accuracy at the extremes of the range. This paper does not say how many studies it covered
  • Ward (1999) and Zatorre (2003) appear in the bibliography, but the paper does not present itself as a continuation of either

Complements Genetic Studies

  • Baharloo 1998: one of the six source tasks whose pitch-naming paradigm this review traces (Fig. 2A). The family-aggregation study proper is Baharloo et al. (2000)
  • Theusch et al. 2009: genome-wide linkage. This review cites it once, in a list of heritability findings; it does not discuss its thresholds
  • Gregersen 2013: AP+synesthesia overlap (phenotype overlap complications)

Validates Heterogeneity Concerns

The paper cites Van Hedger et al. (2020) as "a recent discussion" of whether AP and quasi-AP sit on a pitch-naming continuum, and states the matter "has not been established." What this review documents is that performance is continuous; whether the phenotype is, it proposes as work still to be done.

πŸ“š Citation

Bairnsfather, J. E., Mosing, M. A., Osborne, M. S., & Wilson, S. J. (2025). Conceptual coherence but methodological mayhem: A systematic review of absolute pitch phenotyping. Behavior Research Methods, 57:61. https://doi.org/10.3758/s13428-024-02577-z

πŸš€ Future Directions

  • Urgent: Field-wide adoption of gold-standard task
  • Essential: Taxometric analysis of existing datasets
  • Needed: Multi-task battery to capture phenotypic diversity
  • Critical: International consortium to coordinate phenotyping efforts
  • Ambitious: Large-scale GWAS with well-defined phenotypes