Worked cases
What people do with an interlingua: carry an analysis out to every encoding that will take it, and read the corpora research actually ships in without a parser per corpus. Each runs as a tour in the repository, and the output below is what it printed when this page was built.
From a score with no analysis to every encoding
A Bach chorale out of music21's corpus: four voices, a key signature, and not one harmonic label in the file. The chain is analyse, carry, write — and the only part that is HAMON's is the carrying. music21 does the musicology; FlexOHR (DCMLab) is the object model at the far end. HAMON is what lets them be swapped without the harmony quietly degrading on the way.
4 voices, 165 notes, 0 harmony labels
music21 says: f# minor
----------------------
III bVII6 i bVII6 III bVII6 III III6 bVII bVII7 III V6
HAMON
-----
12 labels over one region: F# minor
what each encoding cannot say
-----------------------------
hamon nothing
dcml nothing
dezrann roman fn. x12, 7th/tensions x6, key/region x1
mei roman fn. x12, key/region x1
harte roman fn. x12, key/region x1
jams roman fn. x12
musicxml key/region x1
lilypond roman fn. x12, key/region x1
abc roman fn. x12, key/region x1
musescore roman fn. x12, key/region x1
romantext nothing
humdrum nothing
ireal roman fn. x12, key/region x1
RomanText
---------
Time Signature: 4/4
m1 f#: III
m2 bVII6
m3 i
m4 bVII6
m5 III
HAMON → FlexOHR
---------------
III → quality=major_triad inversion=root_position
bVII6 → quality=major_triad inversion=first
i → quality=minor_triad inversion=root_position
bVII6 → quality=major_triad inversion=first
The split in that table is clean because the input is a pure Roman analysis. The formats built for functional analysis carry every bit of it; the formats built for chord symbols keep no numerals at all, so for them the analysis is the loss. Run the same chain on a lead sheet and the table turns over, which is the argument the ICCCM'26 page makes at length.
The corpora HAMON reads
The case above leans on the same fact: the formats research actually ships in are readable without a bespoke parser per corpus. HAMON reads 22 formats and writes 13. The other 9 are read-only dialects that nothing writes back — each arrived with its own paper, its own table layout and its own parser.
Distant Listening Corpus dcml I V64 I6 V7/V V
Jazz Harmony Treebank treebank Dm7 G7 Cmaj7 A7 Dm7
Key / Modulation (DDMAL) key_modulation I IV V I IV
DiLeMMa pitch array dilemma I V7/V V (4 notes -> 3 harmonies)
22 formats in, 13 out. The other 9 are read-only dialects nothing writes back.
FlexOHR (DCMLab) A7[of:ii] -> C(SPC(A)) dominant_seventh root_position
P1 M3 P5 m7
back -> A major dom7 (the glyphs and [of:ii] stay behind)
Among them: the Distant Listening Corpus (Hentschel, Moss, Neuwirth, Rohrmeier), the Jazz Harmony Treebank (Harasim, Finkensiep et al.), DiLeMMa (Hentschel et al.) whose note-level pitch arrays collapse back to one label per harmony, and the DDMAL key/modulation set. The full list, with authors, licenses and download recipes, is the dataset manifest.