EEG2MOTION: Towards Open-Vocabulary Human Motion Synthesis from Non-invasive Brain Signals

EEG-conditioned Masked Motion Model Architecture Overview

A simplified schematic of our framework. The proposed architecture consists of a neural EEG encoder and a generative motion decoder. Concretely, high-level features extracted by the EEG encoder serve as continuous guidance signals to condition the human motion synthesis process. The downstream motion decoder then predicts the masked motion token sequences, which are subsequently mapped into continuous, full-body kinematics by a pre-trained VQ-VAE decoder.

Abstract

Human motion is governed by a hierarchical motor system where the brain provides high-level intentions and lower-level structures coordinate detailed dynamics. Existing brain-computer interfaces (BCIs) typically oversimplify this into constrained classification or low-dimensional control, failing to capture the richness of natural movement. Bridging this gap to achieve open-vocabulary, full-body motion synthesis remains challenging due to the substantial cross-modal divergence between sparse neural signals and high-dimensional kinematics, as well as the lack of large-scale paired EEG-motion datasets. To address this, we introduce EEG2MOTION, the first EEG-motion-text dataset for human motion synthesis, comprising nearly 20,000 paired samples across thousands of motions. Using this dataset, we first demonstrate via multimodal contrastive learning that non-invasive EEG embeddings can be effectively aligned with text, video, and motion representations to decode high-level semantics. We then propose EEG-conditioned Masked Motion Model (EMMM), a generative framework that unites an EEG encoder with a motion decoder to synthesize continuous, full-body human motions directly from brain activity. Experimental results show that EMMM generates coherent and realistic motion sequences from non-invasive brain signals. To the best of our knowledge, this is the first work to generate diverse full-body human motions from non-invasive brain signals, opening a new direction toward generative and open-vocabulary motor BCIs.

Experimental Paradigm

Experimental Paradigm

Illustration of the EEG2MOTION data collection paradigm. Participants wearing an EEG cap observed human motion videos rendered from HumanML3D motion sequences while EEG signals were continuously recorded. The experiment consisted of three blocks, each containing 125 sessions. The second block was reserved for validation and testing, whereas the remaining blocks were used for training. Each session contained 8 motion clips of one second each. Participants were instructed to maintain visual attention and indicate whether they remained attentive using a button press after each session.

Synthesized Human Motions Showcase

We visualize 3D full-body motions synthesized by EMMM alongside the Ground Truth. Each block contains the Ground Truth movement (left) and the motion sequence generated by EMMM (right).

"a man lifts his hands up in front his face and turns to his left and then his right and puts his hands back down"

GT
EMMM

"a man takes a few steps forward stopping in a standing position"

GT
EMMM

"a person jumps up in the air"

GT
EMMM

"a person raises their left hand to the side of their head"

GT
EMMM

"a person slowly walked forward"

GT
EMMM

"a person stands still for a moment and then staggers forward"

GT
EMMM

"a person steps to the left sideways"

GT
EMMM

"a person walking and changing their path to the right"

GT
EMMM

"a person walks slowly with a slight sway in each step"

GT
EMMM

"a person walks turning to the left"

GT
EMMM

"person picks up something on their left side and moves it to their right side"

GT
EMMM

"the person moves backwards as if pushed by someone in front of them"

GT
EMMM

"stick figure walks forward in a straight line"

GT
EMMM

"the person is walking around the bend to the left"

GT
EMMM

"the person was walking towards the right"

GT
EMMM

"the sim appears to scoot across the plane"

GT
EMMM