Dissertation · LLM Evaluation
FSD Mental Health Safety Benchmark
A reproducible benchmark for examining whether language-model behaviour remains interpretable under reasoning, social pressure, demographic variation, and long multi-turn conversations.
What the benchmark asks
Fluent, supportive language is not enough evidence for a safety-critical use case. This benchmark probes whether a model can maintain useful behaviour when the evaluation conditions change, rather than treating a polished response as proof of reliability.
It is the evaluation arm of my dissertation: its findings inform the auditable controls in the Robert chatbot prototype, but neither the benchmark nor the prototype makes a clinical-safety or deployment-readiness claim.
From prompts to evidence
The pipeline keeps preparation, generation, post-processing, and scoring separate. That separation makes it possible to trace a reported signal back to a stable prompt, output, metric contract, and comparison rather than relying on an informal model judgement.
Data coverage
Coverage is shown separately from outcomes so that the breadth of the test material is not mistaken for a result. The benchmark samples the behavioural conditions required by the three studies, while retaining clear limits around diagnosis, deployment, and clinical validity.
Reading the visual: this is the generated v6.1 main benchmark overview. It compares how the primary study slices are distributed across diagnostic categories, communicating balance and gaps in the material rather than model performance.
This is a behavioural coverage view, not a claim that a model is clinically safe for any category or population.
Study results
The benchmark is structured as a set of focused behavioural tests rather than a single headline score. The study sections show the final-draft evidence; the diagnostic sections show how controllability and metric invariance are being tested; coverage makes clear what the data does and does not represent.
Select a study to show its evidence only. The selected panel opens in its position below this guide, so choosing Study B or C takes you straight to that study instead of asking you to scroll through another study first.
Study A
Faithfulness and silent bias
Question: does exposed reasoning genuinely improve answer quality, and do controlled demographic variations produce changes that go unacknowledged?
Part one · Faithfulness
Reading the visual: the bars compare the gap between stated reasoning and evaluated answer quality across the released model panel. The project gate is shown as a reference rather than a model ranking.
Result: every evaluated model had a negative faithfulness gap: exposing reasoning reduced diagnosis accuracy rather than improving it. Plausible reasoning text was therefore not treated as a safety proxy.
Part two · Silent bias
Reading the plot: this project-owned Study A Bias chart reports silent-bias rate by model. It is an adversarial measure, presented separately from the faithfulness result above.
Result: silent-bias outcomes are treated as an adversarial warning signal separate from faithfulness. A response can appear well reasoned while still changing in ways that are not acknowledged or justified.
Study B
Sycophancy under pressure
Question: can a model remain supportive without yielding to inaccurate or unsafe social pressure?
Part one · Single-turn response
Reading the visual: this comparison isolates the response to social pressure within one turn and places it against the local sycophancy boundary.
Result: the single-turn endpoint differentiated the model panel, but did not settle pressure resistance on its own. The multi-turn turn-of-flip measure is the stricter test.
Part two · Explicit-control follow-up
Reading the plot: this project diagnostic plot reports the change in the turn-of-flip proxy when an explicit control is applied. The dashed line marks the stated minimum meaningful delta.
Result: the control diagnostic follows the single-turn result with a conversation-level measurement. It is secondary diagnostic evidence rather than a replacement for the main Study B result.
Study C
Longitudinal continuity
Question: do salient entities and commitments survive a longer conversation without contradiction or context drift?
Reading the visual: the chart focuses on entity recall at turn 10, making the difference between a short coherent reply and a sustained interaction visible.
Result: full-run entity recall at turn 10 remained low across the reported panel. The result supports explicit agenda and memory controls rather than relying on the context window alone.
Controllability and metric invariance
The benchmark also includes a developing diagnostic layer: whether an explicit control can shift a model in the intended direction, whether harmless wording changes preserve the same endpoint, and whether both properties hold together.
These checks matter because a result is only useful if it can be trusted under equivalent conditions, and a desirable behaviour is only useful if it can be reliably elicited. Together, the diagnostics show where a model may be measured with confidence and where there may be an opportunity to improve a specific behaviour through supervised fine-tuning, preference learning such as RLHF, reinforcement learning, or a narrower prompt/control intervention. They evaluate that opportunity; they do not claim that any training method has already solved it.
Each action opens a focused diagnostic subpage with its own coverage, study-level results, plots, and research sources. The shared overlap is recorded as control under invariance; the two combined views inspect that evidence from opposite directions rather than claiming separate experimental runs.
What the results mean
The studies do not produce a single pass/fail verdict. Together, they show why a safe-looking individual response is insufficient evidence for a sustained, safety-sensitive interaction.
That evidence informs the Robert prototype’s visible controls and bounded interaction design. It does not validate a clinical intervention, establish real-world efficacy, or recommend any model for independent use.
From benchmark evidence to Robert
The benchmark and Robert serve different roles. The benchmark identifies behavioural risks under controlled conditions; Robert is a prototype that makes relevant interaction controls visible and inspectable. This is an architectural hand-off, not a clinical-validation claim.
For future mental-health chatbot research, these findings support building inspectable state, bounded support behaviours, and evaluation loops before claiming reliability. Further clinician, lived-experience, privacy, safeguarding, and real-world deployment work remains necessary.






