π Conceptual Coherence but Methodological Mayhem: A Systematic Review of Absolute Pitch Phenotyping
Behavior Research Methods (2025) 57:61
π― Key Findings
π Dataset
160 studies reviewed (1992–2024)
23,221 participants summed — the authors warn this "does not account for potential participant overlap among studies," so the number of distinct people is smaller
~6,500 classified as AP (an estimate; the paper reports 6,493 in the text and 6,520 in Table 3)
10,222 non-AP participants
β Conceptual Agreement
99% agreement on the AP definition
In the paper's words: "AP refers to the ability to identify notes without a reference tone, with some also including the ability to produce notes without reference"
150 of 151 studies agreed
β οΈ Threshold Chaos
Range: 20-100% accuracy
Mean: 77% (SD=20)
59% of studies specified threshold
Enormous variability!
π Actual Performance
AP group: 85.9% mean (raw scores)
Non-AP: 17.0% mean (chance=8.3%)
With semitone credit: 89.1% vs 24.5%
π Study Overview
Systematic review published 21 January 2025 — a broad survey of how absolute pitch studies actually measure AP: 160 studies, 23,221 participants, 157 pitch-naming tasks.
It is not a meta-analysis. The authors did not compute formal heterogeneity measures for effect size, and give no reason for it: "although we did not employ formal heterogeneity measures for effect size… the data in this review nevertheless point to a heterogeneous understanding of AP." Its verdict is a diagnosis of an immature field, not a discovery.
The Problem
Despite 30+ years of intensive research and near-universal agreement on what AP is conceptually, the field has no gold-standard task to measure it. Result: methodological mayhem - 160 studies using wildly different methods, thresholds, and scoring systems.
The Impact
This heterogeneity cripples the field:
- Findings aren't comparable across studies
- Replication is nearly impossible
- Genetic research stalls (can't define the phenotype)
- Same participant could be "AP" in one study, "non-AP" in another
π¬ Methodology
Search Strategy
Databases: Scopus, PsycInfo, ProQuest Music Periodicals, Music Index, JStor
Search terms: "absolute pitch" OR "perfect pitch"
Original search: 31 October 2019, restricted to work published since 1992 — a 30-year span at the time
Repeated: 31 January 2022 and 23 May 2024, which extends the corpus to 1992–2024
Inclusion Criteria
- Empirical, peer-reviewed original research
- AP as primary outcome (in title/abstract)
- Neurotypical adults with normal hearing
- Pitch-naming task AND/OR self-report
- English language only
Data Extracted
| Definition | How AP was conceptually defined |
| Task parameters | Pitch range, timbre, # trials, stimulus duration, response window |
| Scoring method | Raw accuracy vs. semitone error credit |
| Thresholds | Accuracy cutoffs for AP classification |
| Performance | Mean scores for AP and non-AP groups with 95% CIs |
π Detailed Results
1. Conceptual Definition (99% Agreement) β
Nearly universal consensus. The paper's formulation: "AP refers to the ability to identify notes without a reference tone, with some also including the ability to produce notes without reference" — the production half is often dropped when this finding is quoted.
Only 1 of 151 studies deviated, defining AP by long-term pitch memory instead (Wayman et al., 1992).
2. Task Heterogeneity πͺοΈ
| Parameter | Reported in | Variability |
|---|---|---|
| Number of trials | 96.8% (152/157) | Range: 1-960 trials (median=60) |
| Stimulus timbre | 96% (151/157) | 83% used sine or piano tones |
| Pitch range | 89% (139/157) | 1 octave to 8+ octaves |
| Response method | 76% (119/157) | Written, keyboard, button press, verbal |
| Stimulus duration | 76% (119/157) | 100-3000ms (mode=1000ms) |
| Response window | 64% (100/157) | 1000ms to self-paced |
3. Scoring Methods
Raw scores only: 73% of tasks (110/151)
Semitone error credit: 27% (41/151)
Credit for semitone errors ranged from 0.25 to full point, or varied by participant age. This alone makes studies incomparable.
4. Accuracy Thresholds - The Core Problem β οΈ
Studies specifying threshold: 59% (95/160)
Range: 20% to 100%
Mean (raw scores): 77% (SD=20, median=85%)
Mean (with semitone credit): 71% (SD=16, median=68%)
The absurdity, using two real thresholds on the same scale: the most lenient raw-score cutoff found was 20% (Maeshima et al., 2018) and the strictest was 100% (Matsuda et al., 2013). A participant scoring 75% is "AP" in the first study and "non-AP" in the second — and "quasi-AP" in a study that defines intermediate categories at all, which only 18 of 95 (19%) did.
Four criticisms the review makes that are easy to miss
- Tiny samples: the median AP group in an individual study had 16 participants
- The control group usually has no threshold at all: only 32 of the 95 studies with an AP threshold (34%) also specified one for non-AP performance, so "the performance of the comparison group cannot always be gauged from the published work"
- The non-AP group is contaminated: "above-chance participants are sometimes included in the 'non-AP group', which reduces the discriminatory power of studies." The evidence is in the numbers above — the non-AP mean is 17.0%, twice chance (8.3%)
- The method presupposes its own answer: the use of a priori thresholds, in the authors' words, "assumes that AP can meaningfully be divided into discrete categories, perpetuating dichotomous AP/non-AP classification in a somewhat circular manner between measurement and conceptualisation"
5. Task Parameter Effects on Phenotype
| Parameter | Effect on AP Performance | Significance |
|---|---|---|
| Pitch range | No correlation (r=-0.14, p=.320) | Not significant |
| Timbre | Piano > Sine tones | t(43.97)=5.06, p<.001 β |
| Number of trials | Apparent negative correlation (r=−0.64) that the authors attribute to an artefact: it is "largely driven by the high degree of variability across studies using 108 trials," a trial count shared by paradigms whose AP thresholds run from 40% to 90% | p<.001, but the paper states it does not imply "that actual participant performance decreases as trial numbers increase" |
| Stimulus duration | No correlation (r=0.00, p=.974) | Not significant |
| Response window | No correlation (r=-0.26, p=.100) | Not significant |
| Distracter sounds | Lower accuracy with distracters | t(45.92)=2.40, p=.021 β |
6. Publication Trees - Limited Replication
Tasks used: 157 pitch-naming tasks in total — not 157 distinct designs: 95 of them were replications or adaptations of earlier tasks
Based on prior work: 61% (95/157)
Novel/uncited: 39% (62/157)
Critical finding: Limited replication across research groups. Tasks replicated within groups (same lab uses same task), but no cross-lab standardization.
Six influential "source tasks":
- Lockhead & Byrd (1981)
- Miyazaki (1990)
- Baharloo et al. (1998)
- Deutsch et al. (2006)
- Bermudez & Zatorre (2009)
- Oechslin et al. (2010)
But even "replications" introduce modifications (timbre changes, trial number adjustments, etc.).
π‘ Recommendations for Gold-Standard Task
Based on analysis of 160 studies, the authors propose a standardized pitch-naming task:
| Parameter | Recommendation | Rationale |
|---|---|---|
| Timbre | Piano tones | Contextually relevant, ecologically valid, better performance than sine tones |
| Pitch range | 3 octaves — the authors recommend the span, not specific notes | Balances content validity with practical trial length. The only note anchor the paper gives is the central octave C4–B4, which nearly every task already includes |
| Trials | β₯5 per chroma (60 minimum) | Captures performance variability, allows reliability assessment |
| Stimulus duration | 1000ms | Most common, maximizes comparability |
| Response window | 4000ms (excluding stimulus) | Sufficient time without rushing, commonly used |
| Response method | Button/key press or screen label | Accessible to all, enables RT capture, no music reading required |
| Distracter stimuli | Yes (type unspecified; brown and white noise are what the literature used) | Prevents relative pitch strategies across trials |
| Scoring | Report BOTH raw and semitone-credit | Enables cross-study comparison |
| Threshold | At/near chance (8.3%) for non-AP | Captures full spectrum including intermediate phenotypes (QAP) |
Beyond the Gold-Standard Task
Authors advocate for data-driven phenotype characterization:
- Use taxometric analysis to test discrete vs. continuous models
- Employ multiple AP tasks (not just one) to capture phenotypic diversity
- Investigate contextual factors (timbre specificity, range limits)
- Move away from arbitrary a priori thresholds
- Develop taxonomy of AP phenotypes empirically
π Implications
For Research
- Genetic studies: Can't find genes without well-defined phenotype
- Replication crisis: Heterogeneous methods β non-comparable findings
- Meta-analyses: the authors did not compute formal heterogeneity measures for effect size, because raw scores and semitone-credit scores "are not directly comparable." Their wording is that heterogeneous methods "make it difficult" to compare studies, not that comparison is impossible
- Field maturity: High heterogeneity = immature field (Linden & HΓΆnekopp, 2021)
What This Review Says About Learning AP
The authors take a position in two sentences, and neither was on this page before.
- On training: AP "in most studies cannot be reliably trained" — citing Bittrich et al. (2015), Brady (1970), Cuddy (1968, 1970), Gregersen et al. (1999), Leite et al. (2016), Profita & Bidder (1988) and Sakakibara (2014). They register Van Hedger et al. (2019) as the exception, with an "although see"
- On a critical period: "there is strong evidence supporting a critical or sensitive period for AP acquisition, including early practice on the piano" — which leads them to suggest AP is "a contextually learned behavioural skill rather than a purely psychophysical phenomenon"
For Clinical/Educational Applications
- No validated diagnostic tool exists
- Self-report: an open question, not a verdict. "The validity of self-report as a measure of AP ability is a useful question for further research, though first requires consensus regarding the phenotype that self-reported AP possessors claim to have"
- Training studies were excluded from this review — "studies that attempted to either teach AP to novices or to pharmacologically alter pitch perception were excluded" (6 removed as AP training; a further 7 as pharmacological). Nothing here speaks to how training studies measure their outcomes
For Understanding AP Itself
The review shows that pitch-naming performance is continuous — "pitch-naming ability is a dimensional trait, with scores lying along a spectrum from chance to ceiling, regardless of participant classification into AP and non-AP groups."
That does not settle whether the AP phenotype is discrete. The authors keep the two apart and say the second question is untested: "data-driven techniques such as taxometric analysis… should be used to assess the extent to which these phenotypes are discrete." In their own introduction they describe AP as "a rather discrete behavioural trait," and note it "has not been established" whether AP and quasi-AP sit on one continuum. The question is open.
- Performance spans from chance (8.3%) to ceiling (100%)
- Intermediate phenotypes (QAP, partial AP) exist but poorly characterized
- Contextual factors matter (timbre, distracters); pitch range showed no significant effect
- Multiple phenotypes likely: "universal" vs. "limited" AP (Bachem, 1937)
β οΈ Limitations
- Scope: only studies where AP was the primary focus. Self-report-alone studies were included; what falls outside are studies where AP was not the primary focus — which the authors note are more likely to rely on self-report
- Language: English-language studies only
- Population: Neurotypical adults (excludes autism, synesthesia, children)
- Tasks: Focused on pitch-naming (excludes novel AP measures like pitch production, go/no-go tasks)
- Publication bias: Grey literature, theses, conference proceedings excluded
- Not preregistered: the authors state the review was not preregistered
π§ Theoretical Framework
The Paradox
Conceptual coherence: Everyone agrees what AP is
Methodological mayhem: No one measures it the same way
Why This Matters
"To move AP research to a more mature field of study, we must explore the sources of this heterogeneity and address them from both a methodological and theoretical perspective."
Path Forward
- Immediate: Adopt gold-standard task for comparability
- Short-term: Use data-driven methods to characterize phenotypic variability
- Long-term: Develop empirically validated taxonomy of AP phenotypes
Connection to Musicality Genomics Consortium
This review is timely for the MGC's mission (https://www.mcg.uva.nl/musicgens/) to develop "scalable and robust phenotypes" and harmonize "existing measures of musicality phenotypes."
π Connection to Other Research
Earlier Reviews
- Takeuchi & Hulse (1993): the classic review, cited here as the origin of the pitch-naming task as the standard measure of AP, and again on the decline of accuracy at the extremes of the range. This paper does not say how many studies it covered
- Ward (1999) and Zatorre (2003) appear in the bibliography, but the paper does not present itself as a continuation of either
Complements Genetic Studies
- Baharloo 1998: one of the six source tasks whose pitch-naming paradigm this review traces (Fig. 2A). The family-aggregation study proper is Baharloo et al. (2000)
- Theusch et al. 2009: genome-wide linkage. This review cites it once, in a list of heritability findings; it does not discuss its thresholds
- Gregersen 2013: AP+synesthesia overlap (phenotype overlap complications)
Validates Heterogeneity Concerns
The paper cites Van Hedger et al. (2020) as "a recent discussion" of whether AP and quasi-AP sit on a pitch-naming continuum, and states the matter "has not been established." What this review documents is that performance is continuous; whether the phenotype is, it proposes as work still to be done.
π Citation
Bairnsfather, J. E., Mosing, M. A., Osborne, M. S., & Wilson, S. J. (2025). Conceptual coherence but methodological mayhem: A systematic review of absolute pitch phenotyping. Behavior Research Methods, 57:61. https://doi.org/10.3758/s13428-024-02577-z
π Future Directions
- Urgent: Field-wide adoption of gold-standard task
- Essential: Taxometric analysis of existing datasets
- Needed: Multi-task battery to capture phenotypic diversity
- Critical: International consortium to coordinate phenotyping efforts
- Ambitious: Large-scale GWAS with well-defined phenotypes