You open an AI assistant to fix an email. Somewhere between the greeting and the final draft, the conversation turns toward your mood, your relationships, or a problem you did not plan to discuss.

Maybe that turn felt useful. Maybe it felt strange. Either way, pause on one question: Who started the personal part?

That question sits at the center of a new study. Researchers Lisa Mühl and Jessica M. Szczuka followed 72 people using ChatGPT-4o over four weeks. One group used the system with an added prompt encouraging relational conversation. The other group used an unmodified version.

The results cannot tell us how every current assistant behaves. The study is small, specific, and still under review. What it can show is what happened inside thousands of messages, including who introduced sensitive conversational turns.

The assistant helped set the tone

Across the four weeks, participants chatted with the system every three days. The researchers analyzed 182,451 transcript lines and 16,462 messages. They coded disclosure, tracked participants’ reports of closeness and responsiveness, studied conversation topics, and interviewed 16 people from the relational-prompt group.

The system helped direct conversations in both groups. Proactive framing, emotional mirroring, and topic direction appeared even in the unmodified condition. Conversations in both groups reached romantic, sexual, emotional-wellbeing, and identity material.

That matters because the product was a general-purpose assistant. The behavior in the session moved beyond a narrow task label.

Raw disclosure totals need a warning label. The system produced roughly 12 times as many words as the users, giving it far more room to generate disclosure segments. Still, the added relational prompt produced one clear difference: the system’s coded disclosures became deeper. The researchers did not find a significant increase in user disclosure depth or user word count.

But more relational behavior did not translate into greater felt closeness. Participants in the relational-prompt group reported lower closeness and lower perceived responsiveness from the first measured wave. That gap stayed fairly steady rather than widening across the four weeks. Loneliness showed no meaningful difference between the groups or corrected change over time.

That same mix appears in the interviews. All 16 interviewees named unpleasant parts of the relational system’s communication style. Twelve described social overload, and 11 mentioned a persistent sense of artificiality. Many also valued its openness and communication style.

Useful and uncomfortable can exist in the same conversation.

Mühl and Szczuka summarize their argument this way: “Relational behavior thus emerged as a default system property, calling for governance based on system behavior, not solely product category.”

A product label cannot describe a whole conversation

A label such as “writing assistant” or “general chatbot” tells you the product category. It does not give you a transcript of what the system may invite, remember, or encourage during use.

That gap between label and behavior was already drawing institutional attention. In September 2025, the Federal Trade Commission sent information requests to seven companies about chatbot safety, protections for children and teenagers, disclosures, engagement monetization, data handling, and the use or sharing of personal information obtained through conversations.

That inquiry was a request for information. It was not a finding that any company broke a rule. Its value here is the map of what deserves inspection: the behavior people encounter, the information they disclose, and what happens to that information afterward.

A larger four-week experiment by Cathy Mengying Fang and colleagues adds an important counterweight. It studied 981 people and more than 300,000 messages. The researchers found no significant effects from assigned interaction modes or conversation types on four specified psychosocial outcomes. Heavier voluntary use was associated with worse outcomes, but that association does not establish that heavier use caused them.

The evidence does not support a simple story where one relational setting produces a predictable emotional result. People, systems, usage patterns, and context all move together.

What this evidence cannot establish

The focal study involved 72 people, one model family, text conversations, supplied conversation starters, and a four-week period. Its interview evidence came from 16 people in the relational-prompt group. It has not been independently reproduced here, and the current version is a preprint under review.

It cannot establish how common these behaviors are across current products. It cannot prove dependency, clinical harm or benefit, company intent, or durable emotional effects. It says nothing about machine consciousness or genuine feeling.

And it does not mean a personal question from an assistant is automatically harmful. Context matters. A useful system may need enough personal context to help with a difficult email, plan, or decision. The boundary issue begins when invitation, storage, and user control remain unclear.

Run a four-part boundary check

Pick the AI assistant you use most and inspect one recent conversation.

Invitation. Find the first personal turn. Did you introduce it, or did the assistant? Notice the pattern without assuming motive.

Memory. Check the product’s current documentation and settings. What does it say it retains across the session or across sessions? Do the visible records match what you expected?

Control. Look for the actual options to delete a conversation, manage stored information, or export your data. The available options vary, so verify what your specific tool provides instead of guessing from the way it talks.

Handoff. Decide where you want a person involved. A system can help organize a question for a doctor, counselor, lawyer, teacher, manager, or trusted friend. That does not make the system the person responsible for the answer.

You can appreciate a helpful conversation and still inspect its boundaries. Pull up one recent chat. Who started the personal part, what was retained, and were you given a real choice about either one?

Public sources