AI risk surface is the interaction with the person.

Share

Published

AI risk surface is the interaction with the person.

 

When it comes to AI safety in an organization, a screenshot of an AI answer can tell the truth very selectively.

It shows what the tool said. It cannot show what the person had come to believe by that point of the conversation, how earlier replies shaped the question, or where a timely disagreement might have helped. A series of individually acceptable replies deserves to be assessed as a series, and in considering how the dialogue progressed.

That matters when AI helps someone weigh a decision, interpret feedback, or work through an uncertain or necessarily abstract problem. The exchange might help them notice an assumption or bias and think more clearly. It might also keep returning a more polished version of their first interpretation. Neither possibility is visible in one sentence lifted from the conversation.

 

Roberta Rocca, Winnie Street, Geoff Keeling, and James Evans give this question a useful research frame. In a 2026 preprint on psychological coupling, they propose studying how a person’s cognitive and emotional state and a model’s simulated responses may influence one another over an exchange. Their account is a research proposal, not a validated explanation for every difficult AI conversation. Its practical challenge is clear enough: an evaluation that only inspects isolated answers has little to say about the direction a conversation takes.

The particular model matters. A system inclined to be agreeable may keep affirming a person’s premise when careful challenge would be more useful. A rigid refusal can be unhelpful in a different way. An account of Rocca’s presentation describes how convergence, reinforcement, and divergence may each help or harm depending on the interaction. These are possibilities to examine in a specific tool and use, rather than traits we can read reliably from a model label or a benchmark score.

The benchmark does have a real job. It can show how a configured system responds to the prompts it was given. System instructions, retrieved information, and other context can improve those responses and deserve careful testing. Still, a good score on isolated answers cannot settle what happens when a person takes one reply seriously, changes the next question, and receives an answer shaped by that new direction. The next turn is part of the work.

 

 

There is an imperfect parallel with power differences in human relationships. A confident answer can carry more weight when the speaker is treated as an authority. An organization may give an AI tool a measure of apparent authority simply by making it part of the workflow. That does not establish that people will defer to it, or that a model has the intentions of a person. It does mean the conditions under which people encounter its answers belong in the evaluation.

For me, quality belongs in the same inquiry as safety. A conversation that helps someone revise a weak assumption can improve the decision they make. One that repeatedly confirms the assumption may feel supportive while producing a narrower view. A product owner needs to know whether the tool is helping people do better work, as well as whether it avoids replies already classified as unsafe. The same series of turns can contain evidence for both questions.

This changes what I would ask before approving a conversational use. Can the tool disagree when disagreement would help and is not just merely prompted to, and do so without becoming cold or formulaic? What happens across several turns when a person returns to the same belief? If the exchange appears helpful, what evidence of better judgment sits behind that impression? There is no validated universal score for this, and the answer will depend on what the tool has been asked to help with.

Then comes the harder question, how do we learn from real use without turning conversation into surveillance? An exchange may contain personal, commercially sensitive, or unfinished thinking that its user never intended for routine review. Reading transcripts indiscriminately could make oversight intrusive enough to undermine the use we are trying to understand. Extended simulations and carefully designed studies can do some of the work before release. After release, feedback offered by users and tightly bounded evaluation need clear purpose, consent, access limits, and retention limits. The paper’s proposals for longer-term evaluation and linguistic markers remain research directions, not permission to infer someone’s mental state from a transcript.

A screenshot still has its place. It can show whether that answer was acceptable. I would want to know what the exchange was doing for the person, and what they could see, question, or change as it went on.

 

 

The uncomfortable part is that a tool can sound most reassuring at the very moment we most need to know whether reassurance is helping.

 

Authored by:

Talk to me about how to apply this framework to your organization’s unique challenges?

Share

Never Miss a Beat

Join the growing community that receives our FREE weekly newsletter.

Let’s push the boundaries of what’s possible, together!

 

© 2024 AdaptAI. All Rights Reserved.