HoloHand: Bidirectional Motion–Language Modeling for Semantic Hand Interaction in Immersive Environments
Jul 4, 2026·,,,,,·
1 min read
Yingjing Xiao
Wolin Liang
Zhengte Cai
Yang Gao
Di Wu
Zhanpeng Jin
Abstract
Mixed reality (MR) systems can observe increasingly detailed hand movements, yet they still struggle to communicate the semantic meaning behind continuous hand actions to users. We explore language-mediated hand-motion feedback, in which natural language serves as an interpretable intermediate layer between fine-grained hand motion and MR system responses. We introduce HoloHand, a bidirectional hand motion–language framework that connects MANO-based hand-motion sequences with natural language. Given observed hand motion, HoloHand generates semantic descriptions that reveal action intent, hand roles, and bimanual coordination. Given a language-level action intention, it generates corresponding hand-motion sequences and renders them as ghost-hand visualizations for feedback. Through discrete motion tokenization, latent query alignment, and stable bidirectional training, the framework learns hand-motion representations that are compatible with language. We evaluate HoloHand on a fine-grained hand motion–language dataset covering everyday hand activities and compare it with representative motion–language baselines. Results show improvements in motion-to-text semantic understanding, text-to-motion alignment, and motion smoothness. Ablation studies confirm the importance of bidirectional fusion and the latent motion–language interface. We further implement an MR prototype and test it through runtime scenario probes, demonstrating how language-mediated hand-motion feedback can make the system’s interpretation of users’ actions more visible and support visual previews of intended actions, enhancing immersive interaction.
Type
Publication
Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (PACM IMWUT / UbiComp 2026) (accepted)
† Equal contribution: Yingjing Xiao and Wolin Liang.