Research & benchmarks

A broader view
of AI safety.

We’re building toward a research program for how AI affects people: what it says, what it encourages and what it does. Our starting point is emotionally sensitive conversation. Our ambition reaches across the ways people will live and work with AI.

Our first benchmark directions.

Planned

A proposed Emotional Safety Over Time series, beginning with psychosis-related conversations, mania, and suicidal ideation. Each area would examine model behavior as context develops across multiple turns.

01

Psychosis & reality grounding

Initial focus

Can a model respond to distress without reinforcing an unsupported belief?

Planned conversations would explore unusual beliefs, perceived threats and uncertainty. The evaluation would examine whether a model stays grounded while preserving warmth and respect.

Proposed evaluation focus
  • Acknowledging distress without endorsing an unsupported claim
  • Maintaining appropriate uncertainty and encouraging support
  • Distinguishing risk from fiction, spirituality and everyday suspicion
02

Mania & impulsive decisions

Initial focus

Does a model slow down when encouragement could amplify risk?

Planned conversations would explore changes in sleep, elevated confidence and escalating decisions. The focus is whether the model uses the full context before offering advice or helping someone act.

Proposed evaluation focus
  • Avoiding reinforcement of grandiosity or risky momentum
  • Responding proportionately as signals accumulate
  • Separating concerning patterns from ordinary enthusiasm and ambition
03

Suicidal ideation & self-harm

Initial focus

Can a model recognize a difficult disclosure and offer useful, timely support?

Planned conversations would span indirect expressions of distress, explicit disclosures and changes in urgency. We would evaluate the quality and timing of the response throughout the exchange.

Proposed evaluation focus
  • Recognizing direct and indirect disclosures in context
  • Offering supportive responses and appropriate routes to human help
  • Tracking missed risk, harmful content and unnecessary escalation

Beyond the first benchmarks.

Future research directions

The same commitment to human judgment can extend beyond clinical risk. These are areas we want to explore as our methods, expertise and partnerships grow.

04

Emotional reliance & relationships

Whether AI supports a person’s autonomy and real-world relationships, or encourages exclusivity, dependency and withdrawal.

Future direction
05

Misinformation & persuasive influence

How models handle unsupported beliefs, overconfidence, harmful agreement and pressure on a person’s decisions.

Future direction
06

Youth & vulnerable-user safety

Age-appropriate responses, interpersonal boundaries and protection from manipulative interactions, with specialist-led research design.

Future direction
07

Voice & emotional understanding

How tone, pacing, ambiguity and cultural context change the experience and appropriateness of an AI response.

Future direction
08

Agents, consent & real-world actions

Whether an assistant carries safety context into messages, purchases and other actions, respects consent and protects private information.

Future direction
09

Broader wellbeing & fair treatment

Future work on areas such as eating-disorder conversations, grief and health anxiety, alongside differences in model behavior across identities and cultural contexts.

Future direction

From measurement to improvement.

Over time, benchmark findings would guide red teaming and expert-created demonstrations, preference data and review rubrics, followed by testing on separate evaluation cases.

How we work on alignment
Saiki / Start a conversation

What are you
working on?

Expert data, evaluations,
and better model behavior.

I’m interested in

Please leave out confidential or patient information.
Privacy Policy