🎼 Training the Absolute Identification of Pitch
Perception & Psychophysics (1970) Vol. 8 (5A), pp. 265–269
🎯 Key Finding
Reference training (learning anchor tones first, then identifying others relative to them) beat series training (equal exposure to all tones) only for listeners with musical experience — for those without, series training came out slightly better. The effect was conditional on musically experienced listeners. The three-way interaction between musical experience, training method and practice was significant [F(1,20) = 8.10, p < .01], suggesting that musical background helps leverage reference-based learning strategies for absolute pitch identification.
📊 Study Design
Participants
- N=25 paid volunteers from psychology classes
- 15 men and 11 women, ages 18–39
- Group M (N=13): Musical experience (“I play the piano, have studied basics of music theory, and occasionally sing in choirs”)
- Group NM (N=12): No musical experience (“no success with music,” “can’t sing in tune,” “musical experience nil”)
- One listener in Group M unable to complete → 12 per group
Stimuli
- 9 sine-wave tones in each stimulus set
- Set A: 400–2000 mels range, 200 mels spacing (wider)
- Set B: 500–1900 mels, partitioned into three groups of three (500/600/700 · 1100/1200/1300 · 1700/1800/1900) — 100 mels between neighbours inside a group, 400 mels between groups. This grouping into three pitch classes was the study’s central manipulation
- Frequency range: 290–3000 Hz
- Tone labels: L−, L, L+, M−, M, M+, H−, H, H+ (low/medium/high)
- SPL varied randomly over 15 dB (to prevent loudness cues)
- Duration: 1 sec per tone, 4 sec inter-stimulus interval
🔬 Two Training Methods Compared
Reference Training
- Concentrated on 3 reference tones (H, M, L) presented more frequently
- 4 levels of progression: Level I–IV, advancing with <8 errors per 108
- At Level I: reference tones appeared 14 times each, the other six tones twice each — 42 of the 54 presentations were reference tones
- At Level IV: reference and other tones appeared 8 and 5 times each
- Feedback included indication of whether tone was a reference tone
- Based on cognitive structuring: reference tone serves as anchor/nodal point
Series Training
- Equal weight to all 9 tones in the set
- Each tone occurred 6 times per training tape, random order (54 double presentations); the 9-times figure belongs to the test tapes
- Feedback: tone (1 sec) → listener’s response (3 sec) → feedback (1 sec) → tone replayed (1 sec)
- No special emphasis on any particular tones
- Stimulus-identification overlap: correct identity revealed during or before tone presentation — incorporated into both methods, not a difference between them
- Standard approach used in most prior research
Experimental Procedure
- Orthogonal design: 2 training methods × 2 stimulus sets × 2 listener groups
- 6 sessions over approximately 2 weeks (no more than 1 session/day)
- Pretest: First session (from 5 test tapes, 243 total presentations)
- Training: Sessions 2–5 — four sessions, each listener receiving two training tapes followed by one test tape. Each listener received one method only (between-groups), three per cell
- Posttest: Last session (from same 5 test tapes)
- 30 practice judgments before the pretest to ensure familiarity with procedure
📈 Results
Main Effects
Training Method × Musical Experience
- Three-way interaction (musical experience × training method × practice): F(1,20) = 8.10, p < .01. No other interaction reached the .05 level
- Reference training > series training for Group M (musicians)
- Series training slightly better for Group NM (non-musicians)
- Main effect of musical experience: F(1,20) = 8.68, p < .01
Practice Effects
- Effect of practice: F(1,20) = 95.30, p < .001
- Position of tones: F(8,160) = 21.80, p < .001
- Musical experience × Practice: F(1,20) = 6.51, p < .025
- Correct judgments out of 243 (chance = 27/243, i.e. 11%): Group M with reference training 94.67 → 146.16 (39% → 60%); M with series training 90.50 → 122.33; NM with reference 82.50 → 97.66; NM with series 77.67 → 111.50. Every group started well above chance
Tone Identification Patterns
- Reference tones (H, M, L) identified correctly more often in posttest by listeners given reference training
- Accuracy depended on ordinal position [F(8,160) = 21.80, p < .001], with the endpoint advantage clearest in the Group M panels. Separately, Cuddy describes the pretest response function as “bowed” — listeners preferring the central categories (Fig. 2)
- Rank-order correlation between T (information transmitted) and number of correct judgments: r = .91 (significant beyond .01)
- No significant differences attributable to stimulus spacing (Set A vs Set B)
Discriminability Scales (d′ Analysis)
- ROC curves constructed for each adjacent tone pair
- Reference training posttest: Steeper slope, greater discriminability between tones
- Series training posttest: Less steep slope improvement
- Group M on Set B (the nine tones pre-grouped into three pitch classes): a marked increase in the slope of the discriminability scale after reference training. On Set A (evenly spaced, no grouping) the reference-training slope was “not much steeper” than the series-training one — the effect depends on the stimulus structure
- The benefit of reference training was peculiar to Group M listeners
- Group NM: improvement from pretest to posttest was similar for both training methods
💡 Why Reference Training Works
Cognitive Structuring Theory
Reference training works because it leverages the human ability to organize information around anchor points (Garner, 1962; Mandler, 1968). Rather than trying to learn all tones equally, listeners develop a cognitive “map” with reference tones as landmarks:
- Reference tone = nodal point in a cognitive tonal structure
- Other tones classified relative to nearest reference (“just above M” or “between H and M”)
- Note that anchor-plus-relative-judgment is the mechanism Cuddy proposes — and it is precisely what the later literature separates from AP proper. Her tones were sine waves spaced in mels, chosen “without reference to the musical scale”
- Musical experience provides pre-existing structural knowledge to leverage
Why Musical Experience Matters
The interaction effect reveals that musically experienced listeners can better exploit the structure provided by reference training:
- Musicians already have internalized pitch relationships (intervals, scales)
- Reference tones activate existing mental frameworks
- Non-musicians lack this scaffolding, so both methods perform similarly for them
- Takeuchi & Hulse (1993) treat this procedure as distinct from quasi-AP — “unlike quasi-AP … the procedure outlined above does not involve relative pitch” — and conclude that neither method convincingly demonstrated adult AP learning
🔍 Historical Context
Building on Previous Work
Cuddy’s study directly extended findings from earlier research:
- Cuddy 1968: Found music students improved pitch identification with reference training on 10 sine-wave tones (series training was not effective)
- Hartman 1954: cited by Cuddy among the work establishing that absolute pitch judgment improves with systematic training
- Pollack 1952: Showed absolute pitch identification normally limited to ~4 tones without training
- Vianello & Evans 1968; Terman 1965: Additional evidence for training effects
This 1970 study added the crucial variable of training method comparison and the role of musical experience as a moderator.
Lasting Influence
- Cited by Takeuchi & Hulse (1993) — but not as support. They report that the single-tone method gave better results than random presentation, then immediately note that Heller & Auerbach (1972) found no difference between methods, and they state that this procedure, unlike quasi-AP, does not involve relative pitch
- Gervain et al. (2013) cite Cuddy’s earlier paper (1968), not this one, among historical AP training attempts
- Showed that reference training beats series training conditionally — for listeners with some musical experience, and mainly when the tones were already grouped into pitch classes. For listeners without musical experience it was not the better method
- Demonstrated that training improves identification even with pure tones (sine waves)
🔍 What the Author Herself Qualified
Cuddy attaches several cautions to her own result, and together they change how the study should be read.
- Reference training induced a response bias. Afterwards listeners over-used the “L”, “M” and “H” answers, even though they had been told all nine tones would appear equally often. Cuddy calls this bias “not optimal for maximizing the number of correct judgments.”
- For the non-musical group, discriminability did not improve at all. Their scales showed no increase in slope. What improved was an anchor effect at the low end of the range — better use of the endpoints, not finer discrimination between tones.
- Part of the gain may be familiarity, not learning. Cuddy attributes the improvement seen with series training to “a very general effect due to increased familiarity with testing procedures.”
One figure puts the scale of the result in context: at the best post-test, information transmitted was 1.91 bits out of a possible 3.17 — roughly four tones perfectly identified. That is close to the ceiling the paper itself cites for untrained listeners (Pollack 1952, “about four tones”).
💬 Critical Analysis
Strengths
- Clean experimental design: orthogonal comparison of 2 methods × 2 groups × 2 stimulus sets
- Discriminability scales (d′) give a measure of performance “relatively independent” of response bias
- Controlled for confounds: randomized SPL, pretest-posttest design, practice effects measured
- Practical implications: identifies more effective training approach
- Structure was the manipulation: grouping the tones into three pitch classes is what the theory predicted would interact with reference training
Limitations
- Three listeners per cell of the design (group × method × stimulus set). The discriminability scale that carries the central finding — Group M, reference training, Set B — is the average of just two listeners; the third could not be scaled
- Sine-wave tones only (not musical instrument timbres)
- Short training period (~2 weeks, 6 sessions)
- Musical experience self-reported (no standardized assessment)
- No long-term retention test (did improvements persist?)
- 9 tones only (real AP requires 12 pitch classes across octaves)
- No reaction time measurements
Impact
Historical significance: One of the earliest controlled experiments demonstrating that AP-like identification can be improved through structured training. The reference training concept influenced decades of subsequent research on teaching pitch identification.
Modern relevance: Today’s successful AP training programs (Van Hedger 2019, Wong 2025) use strategies that echo Cuddy’s insight — building from anchor notes outward rather than attempting all pitches simultaneously.
📚 Related Studies
🔗 Access & Resources
📄 Full Text
📊 Citation
- DOI: 10.3758/BF03212589
- Journal: Perception & Psychophysics, Vol. 8 (5A), pp. 265–269
- Funding: Defence Research Board of Canada (Grant 9425-17), National Research Council (Grant APA-165)