U-MusT · Score Images ⇄ Symbolic Music ⇄ Performance Audio

Interactive demo of the Image-to-Audio (piano) model from U-MusT: A Unified Framework for Cross-Modal Translation of Score Images, Symbolic Music, and Performance Audio (Jung, Kim et al., IEEE TASLP 2026) and the Contin-U full-score inference method. One encoder–decoder Transformer performs every task below; only the target modality changes.

Score images are cropped into musical systems with the fine-tuned ls-yolo system detector and rescaled with the staff-height detector before RQ-VAE tokenization. Audio is decoded with the retrained 44.1 kHz DAC codec.

Checkpoints: image-to-audio / Contin-U run-20250225_062905-9n1554as · OMR run-20250302_101330-hhpxlltr · MIDI-to-audio run-20250330_182257-cogdba9o · device cuda · outputs are for research use (CC BY-NC-SA 4.0 weights).

Running on ZeroGPU. Each generation window (≤ 20 s of audio) takes roughly half a minute of GPU time, and the daily GPU quota is about 2 min for anonymous visitors, 5 min for free accounts and 40 min for PRO. Log in to Hugging Face for longer pieces, or start with a single system / a short MIDI excerpt.

Upload a piano score image (a whole page or a single system). Each detected system is transcribed to Linearized MusicXML (LMX), engraved again with Verovio next to its input crop, and all systems are joined into one downloadable MusicXML file.

Systems to transcribe
0.05 1
Transcribe to see each input system above its engraved transcription.
The joined transcription is engraved here as pages.
Upload a score image to detect its systems.
Example (public-domain engraving from the Mutopia Project)