A Breakthrough in Language Communication: AI Headphones for Multi-Speaker Translation with Voice Cloning and 3D Sound
Researchers from the University of Washington have developed a new headphone system called Spatial Speech Translation, which can translate the speech of multiple speakers simultaneously in real time, clone their voices, and preserve the spatial direction of each speaker’s sound.
How was this achieved?
The system uses widely available active noise-canceling headphones equipped with microphones to capture surrounding conversations. Advanced algorithms can distinguish between different speakers - even in noisy environments - and track them as they move through space. Each speaker’s voice is isolated, translated into the listener’s preferred language, and then played through the headphones with a short delay, typically 2-4 seconds. The key innovation is that the translation preserves not only information about who is speaking, but also where the sound is coming from relative to the listener’s position. This creates an immersive 3D audio experience in which each translated voice sounds as though it is coming from its original direction. "For the first time, we have been able to preserve each person’s voice and the direction from which it is coming," says Professor Shyam Gollakota of the Paul G. Allen School of Computer Science & Engineering at the University of Washington. This technology also clones each person’s unique vocal qualities, so translations retain their individual tone and timbre instead of using generic robotic voices.
The prototype combines commercially available headphones (such as the Sony SH-100XM4) with binaural microphones that mimic human hearing. Audio signals are processed by neural networks running on powerful local hardware, such as Apple M2 chips - no cloud connection is required to ensure privacy. The algorithms operate "like radar", continuously scanning the space across a 360-degree range to detect how many people are speaking and update the situation as participants move or join or leave the conversation. The current version supports Spanish, French, and German, but the goal is to expand support to approximately 100 languages. The source code has been released as open source to enable further development of the technology by experts in the field. The researchers anticipate applications in various areas, from travel and tourism to international business negotiations.
The system’s technical innovations include the ability to process multiple speakers simultaneously - unlike previous systems limited to a single speaker; voice cloning, which preserves each speaker’s unique vocal characteristics; spatial (3D) audio, which retains directional cues so that you hear translations from the places where people are actually standing; real-time processing, which provides translations within 2-4 seconds; and on-device processing, which ensures privacy because all processing takes place locally on devices such as laptops or Apple Vision Pro headsets.
This technology represents a significant advancement in overcoming language barriers during group interactions and makes multilingual communication more natural than ever before. Researchers from the University of Washington are thus opening up new possibilities for intercultural communication and collaboration in a globalized world.



