Meta Introduces Revolutionary Technology for Natural Communication with Avatars
Communication between people is like a dance, where each person constantly adjusts what they say, how they say it, and how they gesture. Modeling two-person, or dyadic, conversational dynamics involves understanding the multimodal relationship between vocal, verbal, and visual social cues—and interpersonal behaviors such as listening, visual synchrony, and turn-taking.
Introducing the Dyadic Motion Models Family
Meta Fundamental AI Research (FAIR), in collaboration with the Codec Avatars and Core AI labs, is introducing the Dyadic Motion Models family, which explores new frontiers in social artificial intelligence. These models can transform human speech or speech generated by language models between two people into diverse, expressive full-body gestures and active listening behaviors.
The models process audio and visual inputs to capture the details of conversational dynamics, with the potential to create more natural and interactive virtual agents that can engage in human social interactions across a variety of immersive environments.

Audio-Visual Dyadic Motion Models in Action
Audio-Visual (AV) Dyadic Motion models can jointly generate facial expressions and body gestures. The models use audio—either from two people or output from large language models (LLMs)—as input to create a behavioral component. Imagine visualizing a previously recorded podcast between two speakers—generating the full range of emotions, gestures, and movements arising from their speech.
AV Dyadic Motion models produce the gestures and expressions of one specific speaker while taking audio from both people into account. This allows the models to visualize speech gestures, listening gestures, and turn-taking cues. The models go one step further by also considering visual input from the other person. This enables them to learn visual synchronization cues, such as mirroring a smile or sharing visual attention.

Seamless Interaction Dataset—Unprecedented Data Scale
The Seamless Interaction Dataset is the largest known video dataset of in-person, two-person conversational interactions and is essential for understanding and modeling how people communicate and behave when they are together. The dataset includes:
- More than 4,000 hours of audiovisual behavioral footage featuring over 4,000 participants
- Approximately 1,300 conversational and activity prompts featuring naturalistic and improvised content
- Rich contextualization with metadata on participant relationships and personalities, along with nearly 5,000 video annotations
All conversations were recorded with participants in the same location to preserve the fundamental characteristics of embodied interaction and avoid the drawbacks of remote, video-based communication. Of these recordings, one-third of the interactions are between two people who know each other, such as family members, friends, or colleagues.
Using Professional Actors for Emotional Diversity
It was important for the dataset to capture a broad range of human emotions and attitudes, such as surprise, disagreement, determination, and regret—in other words, the diversity of face-to-face human behavior. These types of interactions are difficult to capture in naturalistic data, so the team hired professional actors with improvisational experience to portray a variety of roles and emotions. These conversations make up approximately one-third of the dataset.

Advanced Model Capabilities
The models were designed to output intermediate codes for facial and body motion, unlocking a wide range of potential applications. This approach makes it possible to adapt these models for use in various contexts, including 2D video generation and 3D Codec Avatar animation, which can be used in immersive VR and AR experiences.
Additionally, the researchers developed these models by incorporating additional control parameters, providing greater flexibility and control over model behavior. This can be particularly useful when users or designers want to adjust an avatar's expressiveness while speaking or listening.
Safety Measures and Privacy Protection
Meta prioritizes privacy, ethics, and quality when collecting and processing data for its research. During the creation of the dataset, participants consented to the collection of their recorded conversations and were advised to avoid sharing personally identifiable information.
A multi-stage quality assurance process was implemented:
- Human reviews: Human reviewers assessed video content for signs of sensitive material
- LLM analysis of text transcripts: Written transcripts were analyzed to detect sensitive material
- Video analysis using Video Language models: Video Language models were used to detect sensitive material
Watermarking Technology for Authenticity
Meta uses AudioSeal and VideoSeal to watermark content generated by its audiovisual behavioral models. These models make it possible to embed secret messages into individual frames of the original content, which can then be extracted by detectors. This makes it possible to verify the authenticity and origin of content even after processing or manipulation.
The Future of Social Technologies
The research models have the potential to transform future social technologies that help improve our everyday lives, entertain us, and bring us closer together. By prioritizing responsible artificial intelligence practices, Meta hopes to continue building trust in its models and creating technologies that benefit everyone.
The dataset and technical report shared by the team can serve as a blueprint for future research into audiovisual behavioral modeling of interactions. The modeling capabilities built on the dataset will help transform social virtual agents, telepresence technologies in VR and AR environments, and multimodal analysis of video content.


