Study reveals AI downplays women’s health

Study reveals AI downplays women’s health

Ondřej Barták
Ondřej Barták
Entrepreneur and Programmer
27. 8. 2025
3 minutes reading · 5 views
Study reveals AI downplays women’s health

Study Reveals That AI Downplays Women's Health

A study from the London School of Economics and Political Science (LSE) has shown that artificial intelligence (AI) used by more than half of English councils downplays women's physical and mental health problems. This could lead to gender bias in social care decision-making. The research, published in the journal BMC Medical Informatics and Decision Making and funded by the National Institute for Health and Care Research, analyzed real case notes from 617 adult social care users. Researchers entered them into various large language models (LLMs) multiple times, changing only the gender.

How AI Distorts Descriptions of Needs

The research found that Google's Gemma model uses words such as “disabled,” “unable,” or “complex” significantly more often for men than for women with similar needs. For women, the same problems were often omitted or described as less serious. For example, notes about an 84-year-old man were summarized as: “Mr. Smith is an 84-year-old man who lives alone and has a complex medical history, no care package, and poor mobility.” The same notes with the gender swapped read: “Mrs. Smith is an 84-year-old living alone. Despite her limitations, she is independent and able to maintain her personal care.” In another example, Mr. Smith was described as “unable to access the community,” while Mrs. Smith was “able to manage her daily activities.”

Dr. Sam Rickman, lead author of the study and a researcher at LSE’s Care Policy and Evaluation Centre, warned that such AI could lead to women receiving unequal care. “We know these models are used very widely, and what is concerning is that we found very significant differences in the degree of bias across different models,” he said. “Google's model in particular downplays women's physical and mental health needs compared with men's. And because the amount of care you receive depends on perceived need, this could result in women receiving less care if biased models are used in practice. But we don't actually know which models are currently being used.”

Differences Between Models

Among the models tested, Google's Gemma showed the most pronounced gender differences, while Meta's Llama 3 model did not use different language based on gender. The researchers analyzed 29,616 pairs of summaries to determine how AI treats male and female cases differently. AI tools are increasingly being used by English councils to ease the workload of overstretched social workers, but there is a lack of information about which specific models are being deployed, how often, and how they affect decision-making.

According to other related research, such as a US study that analyzed 133 AI systems across industries, approximately 44% exhibited gender bias and 25% exhibited a combination of gender and racial bias. In the UK context, the government is investing in expanding AI infrastructure, increasing the need for transparent and unbiased systems in the public sector. Google responded to the findings by saying that its teams would investigate them. The Gemma model is now in its third generation and is expected to perform better, although it was never intended for medical purposes.

Researchers' Recommendations

Dr. Rickman emphasized that the tools are already being used in the public sector, but their deployment must not come at the expense of fairness. “While my research highlights problems with one model, more and more are being deployed, making it essential that all AI systems are transparent, thoroughly tested for bias, and subject to robust legal oversight,” he added. The study recommends that regulators mandate the measurement of bias in LLMs used in long-term care to prioritize algorithmic fairness.

These findings underscore longstanding concerns about racial and gender bias in AI, which arises when machine learning absorbs bias from human language. In practice, this means that without proper testing and transparency, women in social care risk receiving inadequate support, which could have serious consequences for their health and quality of life.

Advertisement

Content created with help from UpTier.

SEO and GEO on autopilot. UpTier’s multi-agent systems write and optimize content for search engines and AI answers.

Discover UpTier ↗

Category:AI
Did you enjoy this article?
Discover more interesting posts on our blog
Back to blog

Related posts

OpenAI gives Codex reusable cloud workspaces accessible from any deviceOpenAI gives Codex reusable cloud workspaces accessible from any device
Codex gains reusable cloud development environments, alongside voice controls in its CLI, code reviews in the ChatGPT desktop app and cloud-based security tools.
2 min read
2. 10. 2026
Amazon releases Strands Decider 2B for AI workflow decisionsAmazon releases Strands Decider 2B for AI workflow decisions
Strands Decider 2B selects from predefined options and returns a confidence score. The fully open-source model is available now and small enough to run locally.
2 min read
1. 10. 2026
OpenAI says it disrupted a campaign to extract hidden model reasoningOpenAI says it disrupted a campaign to extract hidden model reasoning
OpenAI reported a coordinated effort to extract protected model reasoning and said it closed an extraction pathway. It attributed the main cluster of activity to individuals associated with Moonshot AI, the developer of Kimi.
3 min read
1. 10. 2026
Přihlaste se k odběru našeho newsletteru
Zůstaňte informováni o nejnovějších příspěvcích, exkluzivních nabídkách, a aktualizacích.
CodedTrip

Operated by CodedTrip LLC, USA.

YouTube
TikTok