People with schizophrenia speak differently from others. But the differences are so subtle that doctors often fail to detect them during an appointment. Scientists are therefore experimenting with using artificial intelligence to analyze speech, and the first results look promising. In tests, it distinguished patients from healthy people in roughly 86 percent of cases, while doctors achieved 68 percent accuracy in a comparative test.
Why schizophrenia is so difficult to detect
Schizophrenia disrupts the perception of reality, thinking, and the ability to manage emotions. About 23 million people worldwide live with it, and it usually emerges between late adolescence and the early thirties. No one yet knows what causes it. Scientists suspect a combination of genetics, environment, brain chemistry, and substance use.
The problem is that it cannot be detected through a blood test or by any device. Hallucinations occur only sometimes, social withdrawal tends to be gradual, and delusions may not always become apparent. A doctor therefore sits across from someone who simply sounds different than before and must estimate the extent of that change. Based on this, the doctor then decides whether it is schizophrenia or another mental illness. The consequences of this uncertainty are evident in the numbers. Assessments by two different doctors of the same patient commonly differ by thirty to fifty percent. Moreover, Americans with psychotic disorders, including more than three million people with schizophrenia, receive a diagnosis on average as late as a year and a half after the first symptoms appear. The longer the illness remains untreated, the less well the patient responds to treatment and the greater the risk of brain tissue loss, worsening symptoms, and suicide.
Changes in the voice
Psychiatrists have long known what to look out for in patients. Speech reveals the disorganized thinking typical of schizophrenia. A person jumps from one idea to another along a strange and difficult-to-follow path. The sound of speech changes as well. The voice tends to be more monotonous, almost robotic, and pauses between words become longer. A healthy person also naturally speaks more loudly or softly depending on what they are talking about. In patients with schizophrenia, this contrast diminishes. But detecting these changes by ear, especially at the beginning of the illness, is almost impossible. This is precisely where artificial intelligence may be useful.
First experiment: measuring how speech sounds
A Dutch team used recordings of people who had already been diagnosed with schizophrenia by psychiatrists. They then measured 88 different characteristics in the audio, including volume, pause length, vowel pronunciation, and intonation. From these measurements, the model learned to distinguish how a person with the illness sounds from someone without it. It was then given recordings of patients it had never heard before and classified them correctly in 86.2 percent of cases. It even managed to distinguish between different subtypes of the illness.
The Dutch researchers also noticed where speech differed the most. The speech of patients with schizophrenia is more fragmented, consisting of clusters of syllables separated by longer pauses, and its volume fluctuates more.
Study co-author Alban Voppel, who now works at McGill University in Montreal, adds an important caveat. In clear-cut cases, such assistance is unnecessary because doctors can identify them on their own. The main benefit lies in detecting people at the very earliest stages of the illness, identifying high-risk individuals, and predicting the return of symptoms. According to him, the software appears capable of detecting very subtle details that psychiatrists struggle with.
Second experiment: tracking where speech is heading
Psychiatrist Sunny Tang of the Feinstein Institutes near New York chose a different approach. She was not interested in how speech sounds, but in what is being said. Her team therefore taught a model to assign each word in a conversation transcript a specific address, much like a house has an address. When the software then tracks the position of words in a sentence, it can tell whether the speaker remains within the same area of thought or chaotically jumps to other topics. It is precisely this wandering that indicates disorganized thinking.
The test produced results similar to those of the Dutch team. The model classified the transcripts correctly in 87 percent of cases. Human evaluators who assessed the patients without its assistance achieved an accuracy rate of 68 percent.
Prevention on a smartphone
Beyond diagnosis itself, another possibility is emerging. Doctors could monitor speech over the long term to determine whether treatment is working and whether symptoms are returning.
Today, such monitoring requires regular visits to a doctor’s office, which costs a great deal of time and money. Tang envisions a person speaking into an app or specialized device for a few minutes and receiving a reliable assessment of their condition. She wants the tool to be ready for clinical testing by 2030.
The models still have a great deal to learn
Several factors temper the enthusiasm. A model is only as good as the data it was trained on, and the situation there is more complicated. According to a review of studies led by psychologist Jeffrey Girard of the University of Kansas, most research uses small groups of people that do not reflect the composition of the population. Girard therefore calls for larger and more diverse samples.
Another difficulty is that people without any mental illness may also speak more slowly. Older age or speaking in a foreign language is enough to cause this. Clinical psychologist Sandra Just of the Arctic University in Tromsø points out that the voice also changes under stress, such as in an emergency room, under the influence of medication, or during an ordinary physical illness. Psychiatrist John Torous of Harvard describes this phenomenon as the gap between research and practice. Voice indicators work very well in studies but often fail in real-world settings.
Ideally, a doctor would need to know how a patient spoke before the difficulties emerged. But collecting such baseline material would mean identifying people at higher risk and recording their conversations or having them speak into an app every day, even though they are still entirely healthy. Cognitive neuroscientist Brita Elvevåg, who has studied speech patterns for nearly thirty years, describes this as an extraordinarily complicated problem and is concerned that many people considering such ideas give no thought at all to privacy protection.
Source: scientificamerican.com



