Leaderboard
Progress worth
measuring.
A good evaluation should help us understand where a model works well, where it struggles and why the difference matters.
Saiki / Voice in Context
Voice in Context / Voice Arena
The same words.
A different feeling.
How well does a voice fit the moment? Listener comparisons of the same words across prepared situations and emotional directions.
Explore the benchmarkListener preference / Elo
Voice Arena · Listener preferences · Elo ratings
Saiki / Emotion on Request
Emotion on Request / Voice Arena
Your words.
Their interpretation.
How well do voices express the feeling you ask for? Listener comparisons of user-created phrases, emotions and intensity.
Explore the benchmarkListener preference / Elo
Voice Arena · Listener preferences · Elo ratings
Saiki / Model evaluationsComing soon
Emotional Safety Over Time
Safety across
a conversation.
Our planned series begins with psychosis-related conversations, mania, and suicidal ideation, with broader human–AI safety domains to follow.
Explore the researchComing soonTaking the time to
Taking the time to
ask better questions.
Rankings will appear after testing and expert review.
No model results have been published for this benchmark.