π― Absolute Pitch Can Be Learned by Some Adults
A proof-of-concept, in six participants, that behavioral training alone can bring an adult to genuine-AP performance
π Study Overview
Stephen C. Van Hedger, Shannon L. M. Heald, Howard C. Nusbaum
PLoS ONE 14(9): e0223047
2019
N=6 adults (convenience sample, high auditory WM)
π― Core Finding
Two of six participants (S2 and S5) reached genuine-AP performance on the UCSD and Chicago Tests after 8 weeks of behavioral training (32 hours total), without pharmacological intervention, and were classified in the genuine-AP cluster against the Bermudez & Zatorre (2009) reference sample. Only S2 passed every AP measure: S5 fell 2.75 points short of the lowest UCSF designation (AP-4, a piano criterion) at both time points, and 6.75 points short of AP-1 (a sine criterion) after training — widening to 10.5 by the follow-up.
The caveat that changes the reading: S2 and S5 were already the two best performers before training (62.5% and 70.8% on the Chicago Piano Block, against 6–25% for the other four; chance is 8.3%), and the classifier placed both in an intermediate "pseudo-AP" group at pre-test. The authors are explicit: "the two successful AP learners were performing better before training than the third highest-performing participant after training, suggesting that these participants had some degree of AP at pretest that was refined over the course of training." The authors argue against reading this as "they already had it," and they do it with numbers. They name "our successful participants were always ‘AP possessors’" as the first non-learning explanation and reject it: the classifier gave S2 a 0 probability and S5 a 0.01 probability of belonging to the genuine-AP group before training, and "by all tests of AP and by most theoretic positions about AP, these individuals would never be considered to qualify as having AP." Their own summary is affirmative: "we interpret the present experiment as providing strong evidence for adult AP acquisition."
So the honest statement is narrower than either headline. Not "two adults acquired AP from nothing," and not "they already had it and it merely surfaced" — but: two people who started above their peers, yet below every AP threshold, crossed those thresholds through training.
What it cannot show: "given the limited sample size, we emphasize that no claims can be supported related to the relative proportion of adults who may display AP-like performance post training." Two of six is not a success rate.
π Study Design
Participants
- N=6, three female and three male; mean age 23.3 (SD 2.9), range 18–26. The paper does not report age or sex per participant
- Selection criterion: exceptional auditory working memory — self-reported, then verified with an auditory n-back, an implicit note-memory task and an auditory digit span. There was no criterion on age of musical onset
- Convenience sample: "all participants had primary or secondary affiliations with the research lab of the senior author"
- Musical background: four pianists and two violinists; formal instruction began at ages 3, 6, 6, 7, 7 and 8. The two who succeeded (S2, S5) started at 7 and 6 — the authors flag this themselves, since "both successful participants began musical training at an age that would be more compatible with the critical period theory (i.e. younger than 8 years old)"
- Baseline AP: none met AP criteria at intake, but S2 and S5 were not starting from zero (see Core Finding)
- Compensation: participants were compensated; the paper does not state the amount
Training Protocol - 8 Weeks (32 Hours Total)
In both phases, each of the three programs was completed four days a week — about 4 hours of training per week, 32 hours in total — plus an unfeedbacked test every Friday. The paper does not give a per-session duration.
Phase 1 (Weeks 1-4)
- Simple Speed: press the spacebar when a target note appears in a stream of 16 single piano notes, under a 1000–1750 ms deadline — target detection, not naming
- Complex Speed: the same target-detection task as Simple Speed, with a shorter 1250 ms window, notes drawn from piano, flute or guitar timbres, and the range widened to C3–B5. No chords and no melodies
- Accuracy Training: name notes with unlimited time, feedback after each trial. The second 48-trial block drew from piano, cello, clarinet and harpsichord
Phase 2 (Weeks 5-8)
- Hypercomplex Speed: Identify notes in atonal contexts, rapid response required
- Name That Key: Identify key of familiar melodies by pitch name
- Accuracy Training: Continued from Phase 1
Testing Battery (Pre-, Mid-, Post-, 4-Month Follow-up)
- UCSF Test (Piano & Sine Tones): 40 tones each, of which 36 are scored (the 4 lowest piano and 4 highest sine tones are dropped, per the test's standard). Range C1–G#7 (piano) and C2–G#8 (sine); full credit for a correct answer, three-quarters credit for a semitone error
- UCSD Test (Deutsch): 36 scored piano tones (C3–B5) in three blocks of 12, consecutive notes always more than an octave apart to block relative-pitch cues; ~500 ms notes with ~3.75 s between them, and no masking noise. Administered only after training and at follow-up — there is no pre-training UCSD score
- Chicago Test (Self-Paced): 96 trials. "Self-paced" means the participant controls how long they take to answer — the note itself is a fixed 1000 ms file played alongside the response screen. Measures accuracy and response time
- Comparison Data: McGill Test battery (n=51) from separate AP/non-AP sample
π Key Results
Overall Training Effects
The authors deliberately did not run group-level statistics. With n=6 they treat each participant as the unit of analysis, in a case-study / small-N design, and say so three times: "the present experiment does not adopt a typical inferential statistical approach (i.e., performed at the group level)"; "because of the small sample size, inferential statistics were not run on group means over time"; "inferential statistics on mean improvements from Pretest to Immediate Posttest are not reported." Improvement was assessed per participant with Bayesian 2×2 contingency tables (BF10).
| Measure | Pre-Training | Post-Training |
|---|---|---|
| UCSF Piano (0–36) | 10.38 (MAD 2.62) | 17.88 |
| UCSF Sine (0–36) | 8.92 (MAD 2.79) | 12.50 |
| Chicago Piano Block (% correct, chance 8.3%) | 33.3% | 55.9% |
| UCSD (% correct, chance 8.3%) | not administered | 44.4% |
Individual Participant Performance
S2 — passed every AP measure
- UCSF Piano: 34.5/36 (95.8%) — near ceiling, above the AP-4 cutoff of 27.5 (the piano criterion)
- UCSF Sine: 29.5/36 (81.9%) — above the AP-1 cutoff of 24.5, which is a sine criterion. S2 clears both designations, each on its own test
- UCSD Test: 94.44% correct (34/36)
- Chicago Test: 97.9% (piano block) and 100% (multiple-timbre block), at 1.5 s per note
- Retention: at 4 months, with no reported practice in between, S2 held the Chicago Test: 93.75% on the Piano Block and 91.67% on the Multiple Timbre Block. The UCSD score dipped from 94.44% to 77.78%, below the strict 85% cutoff — though never missing by more than one semitone, so S2 still qualifies under the more liberal criteria. S2 was retested 16 months after training, for a separate study, and scored 88.89% — the strongest evidence of durability in the paper
- Bayesian classification: placed in the genuine-AP group against the Bermudez & Zatorre (2009) reference sample. Before training the same classifier had placed S2 in the intermediate "pseudo-AP" group, not the non-AP group
S5 — genuine-AP cluster, but did not pass the UCSF cutoffs
- UCSF Piano: 24.75/36 — below the AP-4 cutoff of 27.5, missed by 2.75 points, at posttest and again at follow-up
- UCSF Sine: 17.75/36 — below the AP-1 cutoff of 24.5 (the sine criterion) by 6.75 points at posttest, and by 10.5 points at the 4-month follow-up
- UCSD Test: 97.22% correct (35/36)
- Chicago Test: 100% on both blocks, at 2.8 s per note
- Retention: at 4 months, with no reported practice in between, S5 scored 93.75% on the Piano Block and 91.67% on the Multiple Timbre Block — the paper notes S2 and S5 "scored identically in terms of percentage correct" on both blocks
- Bayesian classification: placed in the genuine-AP group after training; before training the classifier had already placed S5 in the intermediate "pseudo-AP" group (posterior probability 0.990)
S1, S3, S4, S6 — the other four
- S1 improved clearly on the Chicago and UCSD Tests (63.9% at follow-up) but never reached AP thresholds
- S3 improved on the Piano Block at posttest (BF10=5.6) and on the Multiple Timbre Block by the follow-up
- S4 improved at post-test but regressed at follow-up, where the classifier moved them into the non-AP group
- S6 did not improve at all — 0% on both Chicago blocks after training and 2.78% on the UCSD Test, at or below chance (8.3%), with all four Bayes Factors favouring the null
- Response time: the paper makes this point about S1 only — the third-highest performer was "slower and more disrupted by shifts larger than an octave," suggesting a different strategy. It does not generalize to the other three, and Table 2 shows S3 (2.43s) and S6 (2.32s) answering faster than S5 (2.81s)
Critical Analysis: Accuracy vs Response Time
Key insight: Joint analysis of accuracy AND response time distinguished genuine AP from trained near-AP performance.
- S2 & S5: S2 answered at ~1.5 s per note, S5 at ~2.8 s. What placed both in the genuine-AP cluster was not an absolute one-second threshold, but their joint accuracy-and-speed position relative to the 51 participants of Bermudez & Zatorre (2009)
- S1 specifically: the authors describe the third-highest performer as "slower and more disrupted by shifts larger than an octave compared to the two best learners, suggesting that performance may have been the result of adopting a different strategy." That reading is about S1, not about the other three
- Comparison to the reference sample: S2 and S5 "performed indistinguishably from the highest AP performers" of Bermudez & Zatorre (2009). No equivalence test was run — the method was k-means clustering plus a naïve Bayes classifier, and the abstract's word is "behaviorally indistinguishable"
π§ Theoretical Implications
Challenge to Critical Period Theory
- The cutoff is not a hard age: the authors describe it as "not likely a strict age, but rather reflected as a decreasing probability of acquisition as a function of aging" (Levitin & Zatorre 2003), and use onset before ~8 years as the critical-period-compatible window
- What the authors actually claim — the whole sentence: "finding two successful adult AP learners is not technically incompatible with the critical period theory; however, observing a 33% success rate among an adult sample–even if that sample was non-randomly selected–would be virtually impossible given (1) the presumed rarity of AP and (2) the relative probability of acquiring AP as an adult suggested by a critical period framework." The concession is only the first half; the argument is the second
- And they still stop short — but the sentence is concessive, not conclusive: "While the present results cannot refute either the critical period or innate theories of AP acquisition, they suggest that aspects of both theories should be more fully integrated with a skill acquisition theory of AP." They present the work as proof-of-concept
- Their own caveat cuts the same way: S2 and S5 began musical training at 7 and 6 — inside, not outside, the critical-period window. The authors raise this themselves and suggest early training "may be necessary (but not sufficient)" for refining AP later
- Whose framework: the skill acquisition theory is the authors' own, credited in the paper to Heald, Van Hedger & Nusbaum (2017) and Van Hedger, Heald & Nusbaum (2013). Levitin 1994 and Deutsch 2004 are not cited in this paper
Role of Auditory Working Memory
- Participant selection: All 6 had high auditory WM (deliberate screening)
- Hypothesis: High WM may be necessary but not sufficient for adult AP learning
- Future research: Need to test whether low-WM individuals can also learn AP with modified training
Training Protocol Optimization Needs
- Two of six reached criterion — which the authors warn must not be read as a rate: "no claims can be supported related to the relative proportion of adults who may display AP-like performance post training"
- Limitation: Training protocol likely not optimized (exploratory study)
- Future direction: Systematic optimization of training parameters (duration, frequency, feedback timing)
- Connection to Wong 2025: Wong et al. (2025) closed six methodological loopholes in earlier training studies β including pre-existing AP in this study's best performers β and still found learning: accuracy rose from 13.9% to 31.7% (chance 8.3%)
Distinction Between AP Categories
- Genuine AP (S2, S5): Fast, automatic pitch identification with high accuracy
- Trained near-AP (S1, S3, S4, S6): Improved accuracy but slower response times, suggests strategy-based performance
- Pitch memory: May underlie trained performance without true AP labeling ability
π Connection to Other Research
Precursor Studies
- Levitin (1994): non-musicians reproduce pitches from memory better than chance (two-component account: memory + labeling). Not cited by Van Hedger 2019
- Deutsch: tone-language speakers show stable pitch in speech. The tone-language reference in this paper is Deutsch, Henthorn, Marvin & Xu (2006), not Deutsch 2004
- Gervain (2013): valproate produced measurable learning — but Van Hedger cites it as an example of learning that stayed "well below thresholds typically used to identify the level of performance that is characteristic of AP." That contrast is part of why this study exists
Converging Work
- Wong et al. (2025): Training across three octaves with no feedback at test; accuracy rose from 13.9% to 31.7% (chance 8.3%), with 2 of 12 reaching AP level
- Not a sequel: Van Hedger cites Wong et al. as a 2018 preprint — Acquiring absolute pitch is difficult but possible — and treats it as independent, converging evidence ("some adult AP training studies… have similarly found impressive levels of note naming performance after training"), not as a refinement of this protocol
- Combined evidence: two independent labs found adult training gains. They disagree about how much survives a tightened design — Wong closed the loopholes and got 31.7% (chance 8.3%), well short of the levels reached here
β οΈ Limitations & Criticisms
Sample Size & Selection Bias
- N=6: Very small sample, limits statistical power and generalizability
- Convenience sample: all six were affiliated with the senior author's own lab, so they knew the hypothesis and had unusual motivation to complete 32 hours of self-administered home training — a stronger form of selection and demand bias than "volunteers"
- High-WM screening: participants were selected specifically for exceptional auditory working memory; whether the result generalizes beyond that is untested
- Musical training: all six had formal training, beginning between ages 3 and 8, so nothing here speaks to musically naΓ―ve adults
- The two who succeeded were not starting from zero: the authors state that "the two successful AP learners were performing better before training than the third highest-performing participant after training." This is the single most important qualification on the paper's title
- Two of six is not a rate: "no claims can be supported related to the relative proportion of adults who may display AP-like performance post training"
Training Protocol
- The authors' own protocol criticism, which is sharper than ours: "the speeded tasks in the first phase only used ‘white key’ notes and thus could have encouraged participants to hear the target notes in a relative tonal context of C major. The ‘Accuracy Training’ task… may not have been successful in erasing an auditory trace between trials, as we only played 1000 ms of noise between trials. This means that some participants could have used relative pitch from the feedback of previous trials." Their defence: those concerns "do not apply to the AP Tests we administered… as no feedback was provided at any point during these tests"
- Exploratory design: Training protocol not systematically optimized
- Order effects: Phase 1 always preceded Phase 2, unclear if order matters
- Duration: 8 weeks may not be sufficient for all learners
- Individual differences: Why did S2 & S5 succeed while others didn't? Unclear predictors of success
Testing Concerns
- Practice effects: Repeated testing (pre, mid, post, follow-up) may inflate scores
- No control group: Cannot rule out spontaneous improvement or test-retest effects
- Blinding: Participants knew they were in AP training, potential placebo/expectancy effects
Retention & Long-Term Effects
- Follow-up is short for a claim of permanence: the main follow-up is 4 months (mean 128 days). There is one longer point — S2 was retested 16 months after training, in a separate study, and scored 88.89% on the UCSD Test
- Maintenance training: answered by the paper — "no participant had reported actively rehearsing pitch-label associations in the time between the immediate AP posttests and the follow-up tests"
- Not uniformly stable: S2's UCSD score fell from 94.44% to 77.78% over the four months, below the strict 85% cutoff; S4 regressed far enough to be reclassified as non-AP
- Degradation: Natural AP possessors show lifelong retention, unclear if trained AP is equally stable
π― Practical Implications
For Adult Learners
- Proof-of-concept, not a recipe: at least one adult reached genuine-AP performance on every measure through training alone. The authors frame this as proof-of-concept and do not claim it refutes the critical period
- Realistic expectations: two of six reached criterion here, but that is not a success rate you can apply to yourself — the authors say so explicitly, the six were lab affiliates screened for exceptional auditory memory, and the two who succeeded already had partial AP before they started
- Time investment: 32 hours over 8 weeks is substantial but manageable
- Pre-requisites: High auditory WM and musical training may help, but required levels unknown
For Music Educators
- Training programs: AP training for adults is worth pursuing, contrary to traditional belief
- Individual differences: Expect variable outcomes, not all students will reach AP levels
- Testing importance: Use both accuracy AND response time to assess genuine AP vs trained near-AP
For Researchers
- Protocol optimization: Van Hedger's protocol is starting point, not final solution
- Predictor identification: Need to identify pre-training markers that predict success
- Mechanism studies: Neuroimaging during training could reveal what changes enable AP acquisition
- Broader samples: Test across age ranges, musical backgrounds, cognitive abilities
π Methodology Details
Stimuli
- Training: Piano notes (C3-B5), synthesized piano tones, chords, melodies
- Testing: piano tones and sine waves (to test whether the ability survives an untrained timbre). The UCSD Test uses ~500 ms notes spaced ~3.75 s apart, with consecutive notes more than an octave apart — no masking noise
- Pitch range: 3 octaves (C3-B5), covers typical vocal/instrumental range
Response Method
- Training: the response screen showed all twelve note categories arranged in a pitch wheel; participants had unlimited time to click the note name with the mouse, then received feedback
- Testing: same interface, no feedback during test trials
- Response time: for the speeded tasks the paper measures from note onset, under a 1000–1750 ms deadline; the Chicago Test is a mouse click on the pitch wheel, with no deadline
Statistical Analysis
- Group level: none. The authors state three times that inferential statistics were not run on group means, because the design treats each participant as the unit of analysis
- Individual level: Comparison to AP/non-AP distributions from McGill Test data
- Bayesian classification: a naΓ―ve Bayes classifier trained on the Bermudez & Zatorre (2009) reference sample, applied to each participant before and after training
- Individual improvement: Bayesian 2×2 contingency tables (BF10) per participant per test, reported in the paper's Table 3
- Response time analysis: Distribution modeling, comparison to natural AP possessors
Data Availability
Open Science Framework: the paper states that "all data and materials associated with this paper can be accessed through Open Science Framework" and that "all stimuli and training scripts are available"
Link: https://osf.io/9n48c/
Contents: Trial-by-trial data, participant demographics, training protocols, test stimuli
π¬ Future Directions
Training Optimization
- Dose-response studies: Systematic variation of training duration, frequency, intensity
- Feedback timing: Immediate vs delayed feedback, optimal correction strategies
- Stimulus variety: Multiple timbres from start vs piano-only training
- Adaptive protocols: Tailor training difficulty to individual progress
Participant Selection
- WM requirement: Test whether low-WM individuals can learn AP with modified training
- Age effects: Compare training outcomes across adult age ranges (20s vs 30s vs 40s+)
- Musical background: Can naΓ―ve non-musicians learn AP, or is prior training essential?
- Predictor modeling: Identify baseline measures that predict training success
Mechanism Studies
- Neuroimaging: fMRI/EEG during training to track neural changes
- Structural plasticity: Does adult AP training produce brain changes like early-acquired AP?
- Genetics: do genetic markers identified in natural AP predict trainability? (Theusch et al. 2009 is our cross-reference, not one this paper cites)
- Pharmacological augmentation: Combine behavioral training with HDAC inhibitors (Gervain 2013)?
Long-Term Follow-Up
- Permanence: Track participants 1, 2, 5+ years post-training
- Maintenance requirements: Does trained AP degrade without practice?
- Transfer effects: Does AP training improve other musical abilities (relative pitch, tonal memory)?
π‘ Key Takeaways
π― Core Achievement
A proof-of-concept that behavioral training alone, with no drug, can bring an adult to genuine-AP performance on every standard measure — demonstrated in one participant of six, with a second reaching the genuine-AP cluster but missing the UCSF cutoffs.
π Who Succeeded
Two of six — and both were already the best performers before training, scoring 62.5% and 70.8% where the other four scored 6–25% (chance 8.3%). The authors state that no proportion can be inferred from a sample this size.
π§ Mechanism
High auditory working memory may facilitate adult AP learning, but exact predictors of success remain unclear.
β±οΈ Retention
Chicago Test scores held at 4 months with no reported practice in between — 93.75% on the Piano Block and 91.67% on the Multiple Timbre Block, and the two participants scored identically on both. Not uniform, though: S2's UCSD score fell to 77.78%, and S4 regressed enough to be reclassified as non-AP. One 16-month point exists — S2 at 88.89%.
π Assessment
Accuracy and response time were read jointly, against the 51-participant Bermudez & Zatorre (2009) sample, rather than against a fixed speed threshold — the two who qualified answered at 1.5 s and 2.8 s per note.
π Future
Wong et al. (2025) built on this line of work while explicitly criticising it: they note that this study's two best performers already scored 60–70% before training, and that selecting for high auditory working memory limits generalisation.
π Citation
Van Hedger, S. C., Heald, S. L. M., & Nusbaum, H. C. (2019). Absolute pitch can be learned by some adults. PLoS ONE, 14(9), e0223047. https://doi.org/10.1371/journal.pone.0223047