AI & ML interests

None defined yet.

Recent Activity

yonghyunk1mย  authored a paper about 14 hours ago
Multi-Task Multi-Frame Visual Piano Transcription
yonghyunk1mย  submitted a paper about 19 hours ago
Multi-Task Multi-Frame Visual Piano Transcription
yonghyunk1mย  updated a Space 1 day ago
PianoVAM/README
View all activity

Organization Card

PianoVAM ๐ŸŽน

A multimodal piano performance dataset, and the models built on it.

From the MAC Lab at KAIST and the Music Informatics Group at Georgia Tech.

PianoVAM captures piano performance across every modality at once. Each of 107 practice recordings on a Yamaha Disklavier aligns top-view video, lossless audio, ground-truth MIDI, and 21-point hand landmarks to the same take. On top of the dataset we build models that read one modality from another, starting with V2N, which transcribes piano to complete MIDI from silent video alone.

PianoVAM

A multimodal piano performance dataset (ISMIR 2025). ๐Ÿ—‚๏ธ Dataset ยท ๐ŸŽฌ Live demo ยท ๐Ÿ“„ Paper ยท ๐Ÿ’ป Code

V2N (Video to Notes)

Visual piano transcription from silent video (ISMIR 2026). ๐ŸŽฌ Live demo ยท ๐Ÿ“„ Paper ยท ๐Ÿ’ป Code

PiaRec & ASDF

Web toolkits for dataset acquisition and fingering annotation (ISMIR 2025 LBD). ๐Ÿ“„ Paper

models 0

None public yet