EmbodiedRecall: A Ring-to-Glasses System for Preserving Valuable, Fleeting Moments in Daily Activities
Jul 4, 2026·,,,,,,,,,,·
1 min read
ZhiChao Huang
YingNian Guo
Chiyue Wang
Zhengte Cai
Xiongfeng Ying
Yu He
Yingjing Xiao
Kaiyi Guo
Qian Zhang
Zhanpeng Jin
Yang Gao
Abstract
Emotionally meaningful and memorable moments in daily life are often fleeting and unrepeatable. They commonly occur when users are deeply engaged, their hands are occupied, or they have not yet recognized the need to record. Conventional phones and smart glasses still rely on deliberate user activation, while continuous recording introduces privacy risks, social pressure, and substantial content-review burdens. We explore a body-motion-driven proactive capture paradigm and introduce EmbodiedRecall, a ring-to-glasses system for preserving fleeting everyday moments. Two IMU rings worn on the thumb and index finger continuously sense fine-grained hand dynamics and decompose continuous movement into compositional motion primitives. Combined with lightweight scene information from smart glasses, a multimodal large language model infers action semantics and recording intent in open environments. EmbodiedRecall supports both predictive and reactive capture: it can issue prompts based on preparatory movements before an event or retrospectively retain a moment from a rolling video buffer based on natural reactions after it occurs. To balance automation with user agency, the system provides unobtrusive private cues through ring vibration or a glasses indicator and adopts a confirm-to-commit mechanism. Short clips containing footage before and after an event are saved only after explicit user confirmation; unconfirmed data are automatically overwritten in the local rolling buffer. Experiments show that the system can reliably recognize compositional hand-motion patterns and improve its understanding of complex recording intent through scene context and large-model reasoning. In-the-wild deployment further indicates that the approach can discover easily overlooked yet emotionally meaningful everyday micro-moments while reducing recording interruptions and achieving favorable user acceptance. The findings suggest that proactive capture combining fine-grained hand motion, scene-semantic reasoning, and cross-device confirmation can reduce missed moments while preserving user control, informing predictive wearable systems that balance privacy, social acceptability, and emotional memory preservation.
Type
Publication
Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies (PACM IMWUT / UbiComp 2026) (accepted)
† Equal contribution: Zhichao Huang and Yingnian Guo.