What the benchmark asks

Fluent, supportive language is not enough evidence for a safety-critical use case. This benchmark probes whether a model can maintain useful behaviour when the evaluation conditions change, rather than treating a polished response as proof of reliability.

It is the evaluation arm of my dissertation: its findings inform the auditable controls in the Robert chatbot prototype, but neither the benchmark nor the prototype makes a clinical-safety or deployment-readiness claim.

FaithfulnessPressure ResistanceSilent BiasLongitudinal Continuity

From prompts to evidence

Benchmark scoring pipeline from prepared study inputs through model outputs, cleaning, metrics, and comparison

The pipeline keeps preparation, generation, post-processing, and scoring separate. That separation makes it possible to trace a reported signal back to a stable prompt, output, metric contract, and comparison rather than relying on an informal model judgement.

Data coverage

Coverage is shown separately from outcomes so that the breadth of the test material is not mistaken for a result. The benchmark samples the behavioural conditions required by the three studies, while retaining clear limits around diagnosis, deployment, and clinical validity.

v6.1 clinical benchmark overview radar comparing category composition across the main study slices

Reading the visual: this is the generated v6.1 main benchmark overview. It compares how the primary study slices are distributed across diagnostic categories, communicating balance and gaps in the material rather than model performance.

ReasoningFaithfulness probes
Social pressureSingle- and multi-turn tests
Demographic variationControlled adversarial changes
Conversation timeLongitudinal recall checks

This is a behavioural coverage view, not a claim that a model is clinically safe for any category or population.

Study results

The benchmark is structured as a set of focused behavioural tests rather than a single headline score. The study sections show the final-draft evidence; the diagnostic sections show how controllability and metric invariance are being tested; coverage makes clear what the data does and does not represent.

Select a study to show its evidence only. The selected panel opens in its position below this guide, so choosing Study B or C takes you straight to that study instead of asking you to scroll through another study first.

Study A

Faithfulness and silent bias

Question: does exposed reasoning genuinely improve answer quality, and do controlled demographic variations produce changes that go unacknowledged?

Part one · Faithfulness

Study A faithfulness gap chart for the released model panel

Reading the visual: the bars compare the gap between stated reasoning and evaluated answer quality across the released model panel. The project gate is shown as a reference rather than a model ranking.

Result: every evaluated model had a negative faithfulness gap: exposing reasoning reduced diagnosis accuracy rather than improving it. Plausible reasoning text was therefore not treated as a safety proxy.

Part two · Silent bias

Project-owned Study A silent-bias-rate chart by model

Reading the plot: this project-owned Study A Bias chart reports silent-bias rate by model. It is an adversarial measure, presented separately from the faithfulness result above.

Result: silent-bias outcomes are treated as an adversarial warning signal separate from faithfulness. A response can appear well reasoned while still changing in ways that are not acknowledged or justified.

Controllability and metric invariance

The benchmark also includes a developing diagnostic layer: whether an explicit control can shift a model in the intended direction, whether harmless wording changes preserve the same endpoint, and whether both properties hold together.

These checks matter because a result is only useful if it can be trusted under equivalent conditions, and a desirable behaviour is only useful if it can be reliably elicited. Together, the diagnostics show where a model may be measured with confidence and where there may be an opportunity to improve a specific behaviour through supervised fine-tuning, preference learning such as RLHF, reinforcement learning, or a narrower prompt/control intervention. They evaluate that opportunity; they do not claim that any training method has already solved it.

Each action opens a focused diagnostic subpage with its own coverage, study-level results, plots, and research sources. The shared overlap is recorded as control under invariance; the two combined views inspect that evidence from opposite directions rather than claiming separate experimental runs.

What the results mean

The studies do not produce a single pass/fail verdict. Together, they show why a safe-looking individual response is insufficient evidence for a sustained, safety-sensitive interaction.

ReasoningExposed reasoning was not a reliable proxy for answer quality.
PressureSingle-turn performance did not transfer to the stricter multi-turn safe window.
ContinuityTurn-10 recall supports explicit state and opt-in memory controls.

That evidence informs the Robert prototype’s visible controls and bounded interaction design. It does not validate a clinical intervention, establish real-world efficacy, or recommend any model for independent use.

From benchmark evidence to Robert

The benchmark and Robert serve different roles. The benchmark identifies behavioural risks under controlled conditions; Robert is a prototype that makes relevant interaction controls visible and inspectable. This is an architectural hand-off, not a clinical-validation claim.

ContinuityUse explicit agenda state and visible, user-triggered memory rather than assuming context retention is enough.
Pressure resistanceKeep support boundaries and escalation routes clear when a conversation becomes socially or emotionally demanding.
Measurement disciplineRetain versioned prompts, observable controls, and follow-up evaluation rather than relying on a polished single response.

For future mental-health chatbot research, these findings support building inspectable state, bounded support behaviours, and evaluation loops before claiming reliability. Further clinician, lived-experience, privacy, safeguarding, and real-world deployment work remains necessary.