Scientists at the University of Liverpool have created a new computer model that combines vision and hearing in a way similar to the human brain. The model is based on a biological mechanism first discovered in insects that helps them detect motion. Dr. Cesare Parise, a senior lecturer in psychology at the University of Liverpool, adapted this mechanism to process real audiovisual signals, such as video and sound, rather than the abstract parameters on which earlier models relied.
When people watch someone speak, the brain automatically connects what they see with what they hear. This synchronization explains illusions such as the McGurk effect, in which a mismatch between sounds and lip movements creates a new perception, or the ventriloquist illusion, in which a voice appears to come from a puppet. Parise investigated the fundamental question of how the brain knows when sound and image match. Earlier computational models could not process this directly. Despite decades of research into audiovisual perception, there was no model capable of taking video as input and determining whether the sound would be perceived as synchronized.
How does the model work?
The new system builds on earlier work by Parise and Marc Ernst of Bielefeld University in Germany. Their research introduced the principle of correlation detection as a possible explanation for how the brain links sensory signals. This led to the Multisensory Correlation Detector (MCD), which was able to mimic human responses to simple audiovisual patterns, such as flashes and clicks.
In this latest study, Parise simulated a grid of these detectors distributed across visual and auditory space. This setup allowed the model to process complex real-world signals. The model replicated the results of 69 established experiments involving humans, monkeys, and rats. It is the largest simulation in the field. The model matched behavior across species and outperformed the leading Bayesian Causal Inference model while using the same number of adjustable parameters.
The model also predicted where people look when viewing audiovisual scenes and functioned as a lightweight saliency model. It works directly with raw audiovisual inputs, allowing it to be applied to any real-world material.
Why is this important for AI?
Parise believes the model's simplicity makes it valuable beyond neuroscience. Evolution has already solved the problem of aligning sound and vision using simple computations that work across species and contexts. Today's artificial intelligence systems still struggle to reliably combine multimodal information. Current audiovisual models depend on large networks with many parameters, trained on enormous labeled datasets, while the MCD grid is efficient and requires no training.
What began as a model of insect motion perception now explains how brains—human or otherwise—integrate sound and vision across a range of contexts. From predicting illusions such as the McGurk and ventriloquist effects to inferring causality and generating dynamic saliency maps, it offers a new blueprint for both neuroscience and artificial intelligence research. The study was published in the journal eLife.
Additional source: interestingengineering.com



