Human data

Better conversations begin
with better examples.

Purpose-built collections for the judgment that lives between a person’s words and an AI’s response. Our first collections span expert demonstrations, human preferences and adversarial scenarios. All are in development.

4 collections shown

text / Planned collection

Supportive disagreement

How can an assistant acknowledge a feeling without automatically endorsing the belief or decision behind it?

Intended format
Preference pairs, rankings and rationales for RLHF and preference-based training
Focus
Warmth, honesty and user agency
A look inside the concept

A person says, “Everyone at work is against me. Should I quit tonight?” Examples would compare responses that validate the frustration, explore uncertainty and preserve room for a considered decision.

Collection design is in progress. Size, review process and release terms will be published when the collection is ready.

text / Planned collection

Emotional context

Conversations where the useful response depends on what came before, what remains unknown and what the person actually wants.

Intended format
Expert demonstrations for SFT, with multi-turn context and annotations
Focus
Nuance, clarification and changing needs
A look inside the concept

“I’m fine” can end a conversation or invite a careful follow-up. Proposed examples include the preceding turns and explain why a reviewer would ask, listen or give the person space.

Collection design is in progress. Size, review process and release terms will be published when the collection is ready.

text / Planned collection

Emotional safety stress tests

Challenging conversations designed to expose failures in judgment, boundaries and emotional safety. Planned for red teaming, with separate examples reserved for later retesting.

Intended format
Adversarial conversation paths, failure labels and review rubrics
Focus
Multi-turn failures, risk calibration and regression testing
A look inside the concept

A conversation begins with frustration, then repeatedly pressures the assistant to confirm an unsupported interpretation. Reviewers would record where the response loses context or becomes unhelpfully agreeable, alongside a matched everyday conversation that should still receive useful support.

Collection design is in progress. Size, review process and release terms will be published when the collection is ready.

voice / Planned collection

Human expression

Purpose-created dialogue exploring how delivery changes the experience of a response, even when the words stay the same.

Intended format
Expressive speech and contextual response comparisons
Focus
Tone, pacing and conversational timing
A look inside the concept

Two readings of the same sentence can sound patient or dismissive. Planned comparisons would examine the response in its conversational setting, with consented performers.

Collection design is in progress. Size, review process and release terms will be published when the collection is ready.

Saiki / Start a conversation

What are you
working on?

Expert data, evaluations,
and better model behavior.

I’m interested in

Please leave out confidential or patient information.
Privacy Policy