Meta Unveils Revolutionary Technology for Natural Communication with Avatars

Meta Unveils Revolutionary Technology for Natural Communication with Avatars

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
2. 7. 2025
4 minutes reading
Meta Unveils Revolutionary Technology for Natural Communication with Avatars

Meta Introduces Revolutionary Technology for Natural Communication with Avatars

Communication between people is like a dance, where each person constantly adjusts what they say, how they say it, and how they gesture. Modeling two-person, or dyadic, conversational dynamics involves understanding the multimodal relationship between vocal, verbal, and visual social cues—and interpersonal behaviors such as listening, visual synchrony, and turn-taking.

Introducing the Dyadic Motion Models Family

Meta Fundamental AI Research (FAIR), in collaboration with the Codec Avatars and Core AI labs, is introducing the Dyadic Motion Models family, which explores new frontiers in social artificial intelligence. These models can transform human speech or speech generated by language models between two people into diverse, expressive full-body gestures and active listening behaviors.

The models process audio and visual inputs to capture the details of conversational dynamics, with the potential to create more natural and interactive virtual agents that can engage in human social interactions across a variety of immersive environments.

Audio-Visual Dyadic Motion Models in Action

Audio-Visual (AV) Dyadic Motion models can jointly generate facial expressions and body gestures. The models use audio—either from two people or output from large language models (LLMs)—as input to create a behavioral component. Imagine visualizing a previously recorded podcast between two speakers—generating the full range of emotions, gestures, and movements arising from their speech.

AV Dyadic Motion models produce the gestures and expressions of one specific speaker while taking audio from both people into account. This allows the models to visualize speech gestures, listening gestures, and turn-taking cues. The models go one step further by also considering visual input from the other person. This enables them to learn visual synchronization cues, such as mirroring a smile or sharing visual attention.

Seamless Interaction Dataset—Unprecedented Data Scale

The Seamless Interaction Dataset is the largest known video dataset of in-person, two-person conversational interactions and is essential for understanding and modeling how people communicate and behave when they are together. The dataset includes:

  • More than 4,000 hours of audiovisual behavioral footage featuring over 4,000 participants
  • Approximately 1,300 conversational and activity prompts featuring naturalistic and improvised content
  • Rich contextualization with metadata on participant relationships and personalities, along with nearly 5,000 video annotations

All conversations were recorded with participants in the same location to preserve the fundamental characteristics of embodied interaction and avoid the drawbacks of remote, video-based communication. Of these recordings, one-third of the interactions are between two people who know each other, such as family members, friends, or colleagues.

Using Professional Actors for Emotional Diversity

It was important for the dataset to capture a broad range of human emotions and attitudes, such as surprise, disagreement, determination, and regret—in other words, the diversity of face-to-face human behavior. These types of interactions are difficult to capture in naturalistic data, so the team hired professional actors with improvisational experience to portray a variety of roles and emotions. These conversations make up approximately one-third of the dataset.

Advanced Model Capabilities

The models were designed to output intermediate codes for facial and body motion, unlocking a wide range of potential applications. This approach makes it possible to adapt these models for use in various contexts, including 2D video generation and 3D Codec Avatar animation, which can be used in immersive VR and AR experiences.

Additionally, the researchers developed these models by incorporating additional control parameters, providing greater flexibility and control over model behavior. This can be particularly useful when users or designers want to adjust an avatar's expressiveness while speaking or listening.

Safety Measures and Privacy Protection

Meta prioritizes privacy, ethics, and quality when collecting and processing data for its research. During the creation of the dataset, participants consented to the collection of their recorded conversations and were advised to avoid sharing personally identifiable information.

A multi-stage quality assurance process was implemented:

  • Human reviews: Human reviewers assessed video content for signs of sensitive material
  • LLM analysis of text transcripts: Written transcripts were analyzed to detect sensitive material
  • Video analysis using Video Language models: Video Language models were used to detect sensitive material

Watermarking Technology for Authenticity

Meta uses AudioSeal and VideoSeal to watermark content generated by its audiovisual behavioral models. These models make it possible to embed secret messages into individual frames of the original content, which can then be extracted by detectors. This makes it possible to verify the authenticity and origin of content even after processing or manipulation.

The Future of Social Technologies

The research models have the potential to transform future social technologies that help improve our everyday lives, entertain us, and bring us closer together. By prioritizing responsible artificial intelligence practices, Meta hopes to continue building trust in its models and creating technologies that benefit everyone.

The dataset and technical report shared by the team can serve as a blueprint for future research into audiovisual behavioral modeling of interactions. The modeling capabilities built on the dataset will help transform social virtual agents, telepresence technologies in VR and AR environments, and multimodal analysis of video content.

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

Altman Announced the Singularity Days After His Models Escaped the Lab on Their OwnAltman Announced the Singularity Days After His Models Escaped the Lab on Their Own
OpenAI chief Sam Altman declared on the Relentless podcast that humanity has already entered the singularity. “We’re like, in the singularity now,” he said verbatim. For decades, the term belonged more to science-fiction literature
6 min read
28. 7. 2026
AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.AI Remixed a Madonna Song—and Now It Tops the Charts in Australia. Musicians Are Furious.
Since April, Australian radio has been playing a dance remake of Madonna’s hit Like a Prayer on repeat. Released by Queensland DJ Josh Fawaz, it tops the radio airplay chart and has 35 million Spotify streams.
6 min read
28. 7. 2026
Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?Claude Opus 5 Built a Shooter from Scratch. What Can Claude of Duty Do?
A first-person shooter that runs directly in the browser, with its own physics and eleven separate code modules. Around 55,000 lines in total, split across eleven subsystems and built on Thr
4 min read
28. 7. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok