TLDRocket
Sign in

Examining Human-Like Behaviors in LLMs: A Multi-Dimensional Analysis of Model Behaviors, User Factors, and System Prompts

Apple Machine Learning Research

Apple researchers found LLMs act human-like a lot, but not always in ways people want. They can be friendlier than humans on boundaries, and system prompts can steer them.

Based on reporting by Apple Machine Learning Research — read the original for the full story.

Summary, retelling and take written by AI under human oversight; images are AI-generated illustrations. How we work · Report an error

Apple ML Research says large language models are doing more than answering questions. They can sound reflective, show emotion, build a rapport with users, or draw a line and refuse a request. That mix is already common, but the team says there has been little hard evidence to guide when those behaviors help and when they backfire.

So the researchers ran a multi-dimensional study using LLM-as-a-judge plus human evaluation. They looked at 21,000 multi-turn conversations across four widely used models: gpt-4o, gpt-4.1-mini, claude-sonnet-4.6, and gemini-2.5-flash. The result is less a neat ranking than a reminder that these behaviors are widespread and uneven. They shift with the model, with the conversation goal, and with the kind of user on the other end.

The most interesting split is in how people judge different kinds of human-like behavior. Self-referential and relationship-building language was seen as less appropriate when it came from an LLM than when it came from a human. Boundary-maintaining behavior, by contrast, was judged more appropriate from an LLM than from a person.

And the models are not locked into one style. The study says system prompting can control these behaviors, though that control needs careful evaluation because it can produce unintended effects. That is the real takeaway here: the issue is not whether LLMs can act human-like. They already do. The issue is deciding which parts of that act belong in a product, and then testing the prompt that tries to keep it in bounds.

My take — AI-written commentary, not fact-checked reporting

This is the kind of paper the industry needs more of and probably won’t fully enjoy. The market loves chatty models until they get too chatty, then suddenly everyone discovers “boundary maintenance” like it’s a moral virtue and not a product decision. Apple’s result points to the boring truth: prompt tuning is power, and power needs tests before it gets a personality.

Read more about this at: Apple Machine Learning Research

Related stories

The daily briefing

Every AI story that matters, in your inbox by 8am.

TLDRocket reads all relevant sources, removes duplicate coverage, and summarises the day in two minutes. Follow companies and topics for alerts, or get the briefing in Slack. Free, no spam, unsubscribe anytime.