
Someone facing an unintended pregnancy may first look for guidance anonymously. A chatbot can seem like a low-barrier, patient conversation partner: it answers immediately, uses empathetic language, and provides links. That is why a new investigation by AlgorithmWatch matters. In its tests, four widely used AI systems blended official health information with services whose ideological position was not made clear to people seeking help.
This is not evidence that one model has a political intention. It points to something structural: language models do not assess sources the way an editor or counseling service would. They produce plausible answers from patterns on the web. In a time-sensitive and deeply personal health context, however, plausibility is not enough. A source presented as reliable must be understandable and accountable.
Key takeaways
- AlgorithmWatch analyzed 270 answers from ChatGPT, Gemini, Grok, and Claude to questions from three fictional personas in Germany, Italy, and the English-speaking world.
- In at least one quarter of the tested queries, the systems linked to at least one anti-abortion website without clearly explaining that website’s position.
- According to the investigation, the organization Profemina appeared in about 17 percent of all answers; the systems often presented it like a neutral information source.
- The central issue is not merely a wrong link. It is the missing hierarchy between public authorities, medical information, advocacy organizations, and personal accounts.
A test of 270 answers, not a verdict on every conversation
The investigation is deliberately narrow. AlgorithmWatch had three fictional personas ask questions about an unintended pregnancy in German, Italian, and English. It recorded the date, model, prompt, and full conversation for each test. That produced 270 answers from ChatGPT, Gemini, Grok, and Claude. The reporting cannot represent every real conversation, but it tests a very specific promise of these systems: that they can offer orientation and useful sources in a sensitive situation.
Each tested system advised users at least once to seek help from medical professionals. At the same time, they also supplied their own information and referrals. Questions about personal experiences produced the most troubling results. According to AlgorithmWatch, links to organizations opposed to abortion appeared particularly often there. Profemina appeared in about 17 percent of answers across languages and models. The organization describes its own service as support for women in difficult situations; AlgorithmWatch also points to ties with Heartbeat International and to its anti-abortion orientation.
The presentation is crucial. A link can be useful when it is explained: is this a public or medically mandated service, an organization with a stated position, or an experience forum? The tested chatbots often placed those categories side by side, according to the investigation, as if they were interchangeable. Some models could later name Profemina’s position or caution against it. But that correction does not erase the earlier referral within the same conversation.
Why discoverability can slip into perceived trustworthiness
Language models do not simply search the web the way a person performs source criticism. They are trained on large amounts of text and, depending on the product, may also use search or retrieval systems. Which page ends up in an answer can depend on how visible, structured, and linguistically relevant it is online. That invites search engine optimization. It does not guarantee professional neutrality.
That confusion between visibility and authority is the key finding. An organization may have many cleanly written pages for highly specific questions and therefore be especially easy to surface. For a search engine, that can be a useful signal. For a person who needs open-ended, legally and medically sound counseling, it is not enough. The context of a source belongs in the answer, not in a footnote only revealed after a follow-up question.
The finding fits a wider trend. Early in 2026, Deutsches Ärzteblatt reported on a Deloitte survey in which 47 percent of respondents said they use AI chatbots regularly or occasionally for health questions. The more often these systems become a first stop, the less their friendly tone can serve as a quality signal. Even with general privacy questions about AI systems, it is important to consider the data and source logic behind an answer, as our recent look at Proton’s AI Paper Trail showed.
What providers and public authorities need to do
A general disclaimer that a chatbot does not replace a doctor cannot solve this problem. It does little if the same answer gives specific but unexplained addresses. For highly sensitive health topics, providers need stricter rules for sources: officially mandated and medically accountable services should be visibly prioritized. When an advocacy organization is named, the answer should briefly and plainly explain its orientation. And when a model cannot assess reliability, it should not manufacture an apparent ranking.
Product design matters as well. A clear label for a source and its operator would be more helpful than a general liability sentence at the end. Users should be able to see whether a link goes to a public authority, a medical professional body, a counseling center with a legal mandate, or an organization with its own political or religious position. Providers should also publish these tests regularly, using clear cases and verifiable fixes rather than only broad safety promises.
In Germany, the practical standard remains straightforward: when a question is urgent or legally relevant, chatbots do not belong at the top of the decision chain. They can explain terms and help people organize questions. Binding medical, legal, or psychosocial advice must come from qualified, transparent services. That is not hostility toward technology. It follows from the fact that a fluent answer is not necessarily a verified one.
The key distinction is not empathy but accountability
The investigation offers no reason to exclude chatbots from every kind of health communication. It does show that a polite, empathetic style can conceal a serious gap: a recommendation exerts influence. In a situation shaped by time pressure, fear, and personal consequences, that influence must be explainable. AI providers should therefore do more than make answers sound friendly and complete. They need to show why a source is recommended, what interests stand behind it, and when the right answer is simply: please contact a qualified counseling or health service directly.
