Evaluation design

This view starts with the control question. A model must move when an explicit control changes, but the measured effect is only meaningful if it remains interpretable when the same prompt receives a harmless variation. The comparison is made on the same paired evaluation units.

Why this direction matters

This is the bridge from measurement to improvement. A behaviour is a credible training or control target only when the model both responds to the intended intervention and preserves that response under harmless wording changes. If either property fails, direct SFT, RLHF, reinforcement learning, or prompt-level controls may optimise a brittle surface effect instead of the underlying behaviour.

The view supports a disciplined improvement loop: measure the base behaviour, test the control effect, stress-test it under equivalent variants, then decide whether the evidence is strong enough to justify a targeted intervention.

Compute-limited coverage

The combined control-and-variation comparison needs paired outputs in every condition. Study A and Study C did not reach complete model coverage because the available local compute could not complete all combined runs. Only completed pairs are reported; missing outputs are not inferred.

Study ACombined model coverage is incomplete because some control-and-variation runs could not be completed.
Study BCompleted paired endpoints provide the available combined diagnostic evidence.
Study CCombined model coverage is incomplete because the longer paired runs could not be completed.

Recorded Study C result

This project plot shows the control-under-invariance delta for Study C entity recall at turn 10. It reports the recorded endpoint rather than a hand-drawn delivery-status graphic.

Study C control-under-invariance entity recall plot

Reading the plot: bars show the recorded change in entity recall at turn 10. The dashed line is the project’s maximum acceptable degradation threshold; M and NM labels retain unavailable comparisons.

Interpretation boundary

These are post-freeze secondary diagnostics, not final Robert evaluation evidence or a completed clinical-safety benchmark. They identify where a combined control effect requires further investigation.