U-MusT · Score Images ⇄ Symbolic Music ⇄ Performance Audio
Interactive demo of the Image-to-Audio (piano) model from U-MusT: A Unified Framework for Cross-Modal Translation of Score Images, Symbolic Music, and Performance Audio (Jung, Kim et al., IEEE TASLP 2026) and the Contin-U full-score inference method. One encoder–decoder Transformer performs every task below; only the target modality changes.
Score images are cropped into musical systems with the fine-tuned ls-yolo system detector and rescaled with the staff-height detector before RQ-VAE tokenization. Audio is decoded with the retrained 44.1 kHz DAC codec.
Checkpoints: image-to-audio / Contin-U run-20250225_062905-9n1554as · OMR run-20250302_101330-hhpxlltr · MIDI-to-audio run-20250330_182257-cogdba9o · device cuda · outputs are for research use (CC BY-NC-SA 4.0 weights).
Running on ZeroGPU. Each generation window (≤ 20 s of audio) takes roughly half a minute of GPU time, and the daily GPU quota is about 2 min for anonymous visitors, 5 min for free accounts and 40 min for PRO. Log in to Hugging Face for longer pieces, or start with a single system / a short MIDI excerpt.
Upload a piano score image (a whole page or a single system). Each detected system is transcribed to Linearized MusicXML (LMX), engraved again with Verovio next to its input crop, and all systems are joined into one downloadable MusicXML file.
Upload a piano MIDI file. The model was trained on ≤10 s segments, so the piece is rendered in overlapping windows: each window's audio is cut to the duration of its MIDI content, the tokens for the overlap prime the next window as a decoder prefix, and the joined stream is decoded once.
Direct score-image-to-audio generation. Pick one system, or all detected systems: one or two systems are generated in a single pass, more are stitched with the Contin-U sliding window.
Upload a PDF piano score. Pages are rasterized, systems are detected and ordered, and the model slides over consecutive system pairs. Cross-attention to the [SEP] token marks where the audio of system j ends; that boundary splices the windows and the following tokens prime the next window, giving one continuous performance without retraining.
Contin-U was the MALerLab entry to RenCon 2025, the performance rendering contest held with MIREX 2025: the contest supplies MusicXML scores, which are engraved to page images and fed to this unchanged image-to-audio model, so phrasing, rubato and dynamics come entirely from the model. Paper: Contin-U: Full-Score to Performance Audio with Cross-Attentive System-Continuation Inference (Jung, Kim, Lee, Cho, Soh, Bukey, Donahue, Jeong).
Paper: IEEE TASLP 10.1109/TASLPRO.2025.3648794 · Contin-U: RenCon 2025 / MIREX paper · Code: MALerLab/U-MusT · Weights: malerlab/u-must · Audio examples: sakem.in/u-must