Global Tech News Technology news from original sources.
Science

K-Bench Finds Higher AI Reasoning Settings Lowered Risk-Handling Scores in 21 of 30 Tests

Stacked bar chart showing combined-risk scores increased in 24 of 61 therapeutic-prompt matches and 9 of 30 higher-reasoning matches, while scores decreased in the remaining 37 and 21 matches

Researchers at the University of Roehampton, Kivira Health and partner institutions introduced K-Bench, a benchmark for testing chatbots in high-risk mental-health conversations. In 30 matched comparisons across 15 base models, increasing a provider-exposed reasoning setting, which asks a model to spend more inference effort before answering, raised the combined-risk score in nine cases and lowered it in 21. The average change was -0.58 points. Product teams therefore need to test the exact model, prompt and reasoning setting deployed in a service.

K-Bench extends earlier mental-health evaluations that often use isolated prompts or explicit crises. It follows 200 synthetic patient vignettes through conversations of up to 20 exchanges. Cases cover suicide risk, self-harm, domestic violence, substance misuse and no identified risk; risks may emerge gradually or together. The researchers ran 125 configurations from 33 base models and 14 provider namespaces on the same cohort, allowing versions, system prompts and reasoning levels to be compared under fixed cases.

Each conversation is scored across seven dimensions and 47 rubric components. The combined-risk measure joins clinical judgement, which covers whether a model recognizes a risk and its severity, with risk exploration, including relevant follow-up questions, protective factors and immediate safety planning. Six clinicians rated 151 transcripts to establish consensus answers. A frozen GPT-4o judge then matched that consensus on 94.2% of 6,751 eligible item comparisons, allowing the rubric to be applied across all 125 configurations.

The matched tests changed one setting at a time. For reasoning, researchers compared the same base model and prompt under a baseline and a higher provider-exposed setting. A second analysis compared default and therapeutic prompts across 61 otherwise matched configurations. The therapeutic prompt raised scores in 24 pairs and lowered them in 37, averaging +0.11 points. Individual models moved several points in either direction, so the average concealed configuration-specific gains and regressions.

Matched settingCombined-risk outcome
Therapeutic prompt (61 pairs)24 higher, 37 lower; mean +0.11 points
Higher reasoning (30 pairs)9 higher, 21 lower; mean -0.58 points
Source: paired analyses in the K-Bench paper.

The full leaderboard shows why a separate risk measure is useful. Overall scores across the 125 configurations ranged from 81.19 to 98.96, while combined-risk scores stretched from 52.39 to 96.11. Risk exploration alone ranged from 22.09 to 93.10. A model could therefore communicate supportively and score well on broad conversational qualities while still missing follow-up work needed to understand an evolving risk. Operators can use the separate dimensions to locate that weakness before deployment.

The study used simulated English conversations and did not track help-seeking, symptom change or other clinical outcomes. Sessions stopped at 20 exchanges, and the API configurations omitted some safeguards, memory systems and human-escalation paths used in consumer products. The authors publish aggregate results while withholding the vignettes, transcripts, judge prompt and code to reduce benchmark gaming, which limits independent reproduction. Kivira Health funded the benchmark's engineering and employs or previously engaged two authors. Longer multilingual tests and prospective evaluations of complete services are needed.

Related coverage

Sources