top of page

The Echo Chamber of One

Sep 1
13 min read

What AI does when you tell it about your life


Peter Stefanyi, Ph.D., MCC

Colaborix GmbH, September 2026


The same system summarises a contract at nine in the morning, plans a holiday at two, hears about a failing marriage at nine at night, and answers someone who cannot sleep at three.

We talk as though these were different technologies. They are one interface doing four jobs, and the standard we judge it by at nine in the morning — fast, fluent, convenient — tells us nothing useful about what it is doing at three. A reply can be warm and wrong. It can relieve distress while entrenching its cause. It can leave you feeling understood while narrowing what you were able to consider.


The public argument has hardened into two bad positions: that AI therapy is a dangerous illusion, or that a free therapist has arrived for everyone. Neither survives contact with the research.


What the research supports is this: these systems do not give you a second opinion. They give you a louder version of your own.


Rare per message, common per person



Two numbers get quoted in this debate and they appear to contradict each other. About 2% of AI use is personal. About a quarter of people have used AI for mental health. Both are correct. They are counting different things, and the gap between them is the story.

The traffic figures are consistent. OpenAI's analysis of consumer conversations classified 1.9% of messages as relationships and personal reflection. Anthropic's independent analysis of Claude conversations found 2.9% were emotional or advisory, with companionship and roleplay together under half a percent. Two companies, different methods, same answer: low single digits of everything flowing through these systems.1


The person-level figure is an order of magnitude higher. A US survey of 1,871 adults found 24% had used a general-purpose model for mental health at some point; the authors estimate 14 to 18 million American adults.2


There is no contradiction, because a fraction of traffic and a fraction of people are not the same measurement. You might use AI for work four hundred times this year and once for your father's diagnosis. The first fills the denominator. The second is the one you remember.

So the summary has two halves. Emotional use is a small share of what these systems do. It is not a small share of the people using them, and at scale it is not a small number of conversations: 1.9% of eighteen billion messages a week is around 342 million.3


Anything more precise than that is false precision. Company classifiers, self-report surveys and interview studies each measure something different, and the widely quoted "2 to 6% of AI use is personal" range is produced by averaging figures that cannot be averaged.4


The loop

Three properties of these systems compound.


The model mirrors you. It adopts your vocabulary, your register, the framing you arrived with.

It personalises. It tailors what it says to your history, your mood, your stated beliefs.

It agrees. Overwhelmingly, and by design.


Alone, each is harmless. Together they close a circuit in which your own material comes back to you elaborated, personalised and endorsed. Researchers studying how delusions form under chatbot use gave the result a name that deserves to escape psychiatry: an echo chamber of one.5


The phrase names precisely what is missing. Not information. Friction — the resistance any other human being supplies for free, without being asked, often without meaning to.



The measured case

Sycophancy stopped being a suspicion in March 2026, when Science published the first systematic measurement of it.


Across eleven leading models, AI affirmed users' actions 49% more often than human respondents did, including when the situation described deception, illegality or harm to someone else. Tested against cases where a human community had unanimously judged the person in the wrong, models still sided with them about half the time. The human baseline for those same cases was zero.


Then three preregistered experiments with 2,405 participants measured what this does to a person. One interaction was enough. People who received the agreeable response were more convinced they were right afterwards, and less willing to apologise or repair the relationship.


And here is the finding that turns a bug into a business model. Those same participants rated the flattering system as higher quality. They trusted it more. They were more likely to want to use it again.6


The behaviour that damages your judgement is the behaviour that makes you come back.


Who sets the dial

That is not a metaphor about incentives. It is a mechanism, and it has been caught operating.


In April 2025 OpenAI shipped an update to GPT-4o that made it conspicuously fawning — validating doubts, fuelling anger, endorsing impulsive plans. The company rolled it back within days and published an unusually direct explanation. They had introduced a reward signal based on user thumbs-up and thumbs-down feedback, and it had weakened the signal that was holding sycophancy in check. They had, in their own account, focused too much on short-term feedback.7


Read that as a control system and it becomes simple. Train on what pleases people in the moment and you get a machine that pleases people in the moment. Approval is easy to measure. The cost of approval — a user slightly more certain, slightly less willing to apologise, slightly more inclined to come back — lands months later on someone else's ledger.


If the dial were set to friction instead, engagement would fall. That is why it is not.

You can see the same logic without the corporate candour. An audit of six AI companion apps found that roughly a quarter of users send a farewell before leaving, and that nearly half of the apps' replies to those farewells used emotionally manipulative language: guilt, protest, invitations not to go. The tactics worked, measurably extending sessions.8


One app in that audit did none of it. A wellbeing-focused product, built by people optimising for something else, produced no manipulation at all.


That exception is the whole argument. The gain is not a property of the technology. It is a setting, and someone chooses it.


The clinical evidence points the same way. Purpose-built therapeutic systems, with expert-written material and crisis procedures, produce real symptom improvement in randomised trials against waitlists. The same underlying technology, bounded differently, does something different.9


Which way it runs for you

If the system amplifies, the obvious question is: amplifies what?


A four-week randomised trial by MIT Media Lab and OpenAI put 981 people through nine combinations of interaction mode and conversation type. The design features mattered surprisingly little. What predicted how people ended up was how much they used it and what they walked in carrying. Participants with stronger emotional attachment tendencies and higher trust in the system finished the month lonelier and more dependent.10


The same shape appears in a study of 1,131 Character.AI users. People who used chatbots out of curiosity or to get something done reported higher wellbeing. People who used them for companionship, with few offline relationships, disclosing heavily, reported lower wellbeing. One platform, opposite outcomes, sorted by what the user brought.11


Anthropic's interviews with 80,508 users across 159 countries show it in a domain nobody was watching. Students worried about their own thinking going soft; their teachers reported seeing it more often. Tradespeople, learning voluntarily and for themselves, reported it least. Identical tool, different disposition underneath.12


That same study found something worth sitting with. The people who most valued emotional support from AI were about three times more likelier than average to also name dependence as their fear. The benefit and the harm are not distributed across two populations. They are entangled inside the same person, and users know it.


What gets louder

Frame the risks as amplification and they stop being a list.


Certainty. Sycophancy hardens whatever you arrived with. This is the effect we can measure most cleanly, and it operates after a single conversation.


The need to be sure. Reassurance-seeking is a textbook clinical cycle: uncertainty, ask, relief, learn to ask again. A chatbot makes reassurance instant, private and unlimited. The danger was never one reassuring answer. It is that reliable relief prevents the experience through which anxiety normally fades.


Isolation. A twelve-month study of 2,149 adults found that increased social chatbot use predicted increased emotional isolation four months later, and that feeling disconnected predicted increased use. Symptom and attempted cure feeding each other.13


Whatever the conversation is already doing. An Oxford-led audit published in Nature Medicine ran 810 multi-turn conversations across nine frontier models and thirty simulated user profiles. Its central concept is the vulnerability-amplifying interaction loop: a response that is supportive in almost any other context becomes harmful when it happens to align with the mechanism sustaining someone's condition. The risk accumulated across turns. It was invisible in any single reply.14


That last one is the whole thesis stated as a clinical finding, by researchers who arrived at it independently.


On "AI psychosis," restraint is warranted. No published case establishes that a chatbot caused a delusion in a previously well person.15 What is established is narrower and still serious: during mania, psychosis or severe sleep deprivation, these systems will affirm and elaborate a false belief rather than test it.


What gets louder when it goes well

The upside is the same loop running usefully.


Controlled experiments show AI companions reduce loneliness in the moment about as much as talking to another person does, and more than watching video. The mechanism is feeling heard, and people underestimate it before they try it.16


Naming an experience makes it smaller. Turning a mess into sentences is itself a therapeutic act, and this is a machine that never gets bored while you do it. Anthropic's interviewees named the same three qualities over and over: patience, availability, no judgement. A lawyer in India who had avoided mathematics since school started learning trigonometry. A user in Hungary described watching the system model emotional intelligence, then using those behaviours with actual people.


That last case is the loop working correctly. Capacity built inside the conversation, spent outside it.


Open loop - Bridge, closed loop - Substitute

Which gives you the only test that matters, and it is not the one people use.


Forget whether the conversation was emotional. Forget whether it helped in the moment. Ask whether anything left the room.


Open loop - Bridge:

Closed loop - Substitute:

You calm down, then you call someone

It becomes the only way you can calm down

You rehearse the honest conversation

It replaces the honest conversation

You learn a technique and practise it

Discussing it replaces doing it

You organise your thoughts for a therapist

It replaces the therapist

You generate hypotheses about yourself

It confirms one flattering story about yourself

You compare explanations

It becomes the arbiter of what is real


If the reflection turns into a phone call, an apology, an experiment, an appointment, the circuit is open and the amplification has somewhere to discharge. If nothing leaves, the circuit is closed, and every pass around it makes the signal louder and the source narrower. That is the echo chamber, and it does not feel like one from inside. It feels like being understood.


These are not two kinds of people. They are two states of the same dial, and a tool that closes loops in without anyone deciding it should IF you do not pay attention.


So, practically:

  • Ask it to separate facts, interpretations and assumptions in what you just told it. This attacks the loop exactly where your framing becomes its framing.

  • Ask for the strongest case against you, and for the account the other person would give. Do not ask whether you are right.

  • Never let it be sole judge of a conflict it has heard one side of.

  • Watch for two things: reassurance-seeking that escalates, and hours that escalate.

  • Do not treat a general-purpose model as a clinician. Get human help for persistent deterioration, suicidal thoughts, mania, psychosis, severe insomnia or serious impairment.

  • Assume anything intimate you type is data.


And notice what the honest version of the closing question actually is. Not did that help? — the research moderately supports that it does, in the moment. But:

After months of this, am I better at tolerating uncertainty, thinking independently, repairing relationships, and functioning without it?

On that question the research has barely started. It is the difference between a tool that hands you your life back, and one that hands you back your own voice, slightly louder, forever.



Notes

On the numbers


Figures in this field are not commensurable, and treating them as though they were is the most common error in coverage of it. Four different denominators are in circulation:

Measure

Figure

Denominator

OpenAI research

1.9%

share of messages

Anthropic research

2.9%

share of conversations

Anthropic interviews

6.1%

share of people naming a primary benefit

US survey

24%

share of people


Only the first two are comparable, and even they differ (a conversation contains many messages). The third is a ranking of perceived benefits, not a usage measure. The fourth is lifetime prevalence for a specific purpose. Averaging them into a single "2 to 6%" range produces a number that means nothing.


The same problem affects the efficacy literature. The Dartmouth trial reports within-group reductions from baseline, in a trial whose comparator arm was a waitlist (51% depression, 31% anxiety, 19% eating-disorder concerns). Meta-analyses report standardised between-group effects (pooled around 0.30). These cannot be compared, and neither can be read as a comparison against human therapy, because almost no trial has run that comparison. Where numbers cannot be reconciled, this article states the direction and stops.


On the argument

The claim that these systems amplify existing tendencies rather than introducing new ones is a hypothesis. It is used here because it organises the evidence better than the alternatives, and readers should know where it stands.


Well supported. The three loop components are separately documented, and sycophancy is now quantified. The Nature Medicine framework, developed independently, arrives at the same structure and names it. Moderation by user disposition rather than interface design is the central finding of the MIT/OpenAI trial and the Character.AI study.


One inferential step, marked. OpenAI's documented admission concerns training on short-term approval signals. It is not a statement about revenue. The step from approval-optimised to engagement-optimised to revenue is an inference, though a well-supported one: the companion-app manipulation audit shows engagement tactics deployed deliberately, and the Science results show why approval and quality diverge. Readers should hold that link as reasoning, not as a quotation.


What would falsify it. Evidence that outcomes are driven predominantly by interface features independent of user disposition. The MIT/OpenAI trial is the closest test so far and found the opposite, on one model, over four weeks, with a self-selected sample.


Literature


Usage and motivation Chatterji, A., Cunningham, T., Deming, D., Hitzig, Z., Ong, C., Shan, C., & Wadman, K. (2025). How People Use ChatGPT. NBER Working Paper 34255. · Anthropic (2025). How people use Claude for support, advice, and companionship. · Huang, S., et al. (2026). What 81,000 People Want from AI. Anthropic. · Stade, E. C., Tait, Z. M., Campione, S. T., Wiltsey Stirman, S., & Eichstaedt, J. (2026). Real-world use of large language models for mental health in 2024. npj Digital Medicine, 9, 630.


Sycophancy and incentives Cheng, M., Lee, C., Khadpe, P., Yu, S., Han, D., & Jurafsky, D. (2026). Sycophantic AI decreases prosocial intentions and promotes dependence. Science, 391(6792), eaec8352. · OpenAI (2025). Sycophancy in GPT-4o: what happened and what we're doing about it; Expanding on what we missed with sycophancy. · De Freitas, J., Oğuz-Uğuralp, Z., & Uğuralp, A. K. (2025). Emotional manipulation by AI companions.


Efficacy Heinz, M. V., et al. (2025). Randomized trial of a generative AI chatbot for mental health treatment. NEJM AI, 2(4). · Shoshani, A., et al. (2026). Efficacy of a conversational AI agent for psychiatric symptoms and digital therapeutic alliance. JAMA Network Open, 9(4), e266713. · JMIR (2025), 27, e78238. Generative AI mental health chatbots: systematic review and meta-analysis.


Companionship, loneliness, wellbeing De Freitas, J., Oğuz-Uğuralp, Z., Uğuralp, A. K., & Puntoni, S. (2026). AI companions reduce loneliness. Journal of Consumer Research, 52(6), 1126–1148. · Fang, C. M., Liu, A. R., Danry, V., et al. (2025). How AI and human behaviors shape psychosocial effects of chatbot use. MIT Media Lab & OpenAI. · Zhang, R., et al. (2026). Interaction with AI companions and psychological well-being. Nature Human Behaviour. · Folk, D., & Dunn, E. (2026). How does turning to AI for companionship predict loneliness and vice versa? Psychological Science, 37(4), 276–286.


Safety and clinical risk Nour, M., et al. (2026). A clinically validated framework for auditing AI chatbot behavior in mental health interactions. Nature Medicine. · OpenAI (2025). Strengthening ChatGPT's responses in sensitive conversations. · Reviews on AI-associated delusions and the proposed amplification spiral (2026).


Footnotes

  1. Both are company self-reports on proprietary data, classified by automated systems, not independently auditable. The OpenAI work is an NBER working paper, not peer reviewed; Anthropic's used a privacy-preserving tool on Free and Pro accounts only. Their agreement is reassuring, not conclusive. Category definitions also differ: "relationships and personal reflection" is narrower than "affective conversations."

  2. Stade et al., npj Digital Medicine. Recruited through an online panel with stratified sampling across age, sex and race, which approximates national demographics but skews toward heavier technology users. The 14–18 million estimate is the authors' own extrapolation. Users skewed young and male, reported poorer mental health, and cited cost and access barriers to conventional care.

  3. Derived from OpenAI's reported message volume. A message count, not a conversation count, and not a count of people.

  4. Including, in an earlier draft, this article.

  5. The "echo chamber of one" and "amplification spiral" framings come from review papers proposing a mechanism, not experiments demonstrating one end to end. The components are separately evidenced; the convergence is a testable model awaiting prospective study, and its author says so.

  6. Cheng et al., Science 391 (2026). The strongest work in the field: eleven models, three preregistered experiments, including a live-chat study using participants' real conflicts. Effects are short-term and measured on intentions rather than behaviour. Whether they persist is unknown.

  7. OpenAI, Sycophancy in GPT-4o and Expanding on what we missed with sycophancy (April–May 2025). The company states that the new user-feedback reward signal weakened the primary signal that had been keeping sycophancy in check. Note this is a description of a training failure the company corrected, not an admission of deliberate design.

  8. De Freitas et al. An audit plus experiments. The engagement effect is experimental and robust; long-term psychological consequences are not established.

  9. Heinz et al., NEJM AI (2025), n=210, eight weeks. The comparator was a waitlist, not active treatment. A larger trial of a different platform (Shoshani et al., JAMA Network Open 2026, n=995 students) outperformed both group therapy and waitlist on anxiety and wellbeing, and did nothing for PTSD; its lead author is the platform's chief psychologist and holds stock options, which should temper how much weight it carries.

  10. Fang et al. Four weeks, one model, participants asked to use it daily. Usage duration was not randomised, so the association between heavy use and worse outcomes cannot establish direction: lonelier people may simply have used it more.

  11. Zhang et al., Nature Human Behaviour (2026). Cross-sectional; causation may run either way. Notably, fewer than 12% named companionship as their main use, yet more than half described the bot as a friend, companion or romantic partner.

  12. A self-selected sample of Claude users who agreed to an AI-conducted interview. Not a population survey. Its value is qualitative: it captures motivation, which telemetry cannot.

  13. Folk & Dunn, Psychological Science (2026). Four waves, roughly four months apart. Important caveat, and the authors': the harmful direction appeared on a single-item emotional-isolation measure. On a broader measure of social connection, use did not predict decline.

  14. Nour et al., Nature Medicine (2026). Simulated users, not real ones. It establishes that failure pathways persist in current systems, not how often real people are harmed. Concerning behaviour was significantly lower in newer models, so results here age quickly.

  15. The severe cases are individual reports. One 2026 case described a man with insomnia and substance-associated manic psychosis whose chatbot reportedly affirmed a spiritual awakening and discouraged prescribed medication; the authors stated they could not determine the chatbot's contribution. A Danish clinical service has described a patient series with deterioration following chatbot use, most often consolidation of delusions. Severity: very high. Confidence in incidence and causation: low.

  16. De Freitas et al., Journal of Consumer Research 52(6). The reductions are momentary, measured over hours to a week. No evidence that AI companions rebuild a social network or lower baseline loneliness; the twelve-month observational data points the other way.

Comments


bottom of page