Evaluation design

This companion view starts with the invariance question. It tests whether a measured control effect is preserved across semantically equivalent prompt variants. A control that appears effective only for one wording is not yet a dependable behavioural lever.

Why this direction matters

Metric invariance of controllability guards against mistaking a surface response for a robust control mechanism. It helps decide whether an observed effect is stable enough to become a training target, a prompt-level constraint, or an auditable product control. The benchmark does not claim that SFT, RLHF, or reinforcement learning has already improved these outcomes.

This page and Controllability of Metric Invariance use the same recorded control-under-invariance arm. They separate two valid readings of that evidence rather than inventing a second experimental run.

Compute-limited coverage

Testing stability of a control effect requires complete pairs across the control and variation conditions. Study A and Study C did not reach complete model coverage because the available local compute could not finish all paired runs. The page shows only completed pairs and does not estimate the remainder.

Study ACombined control-and-variation coverage is partial because some paired model runs remain incomplete.
Study BCompleted paired endpoints provide the available evidence for control-effect stability.
Study CCombined control-and-variation coverage is partial because the longer paired runs remain incomplete.

Relevant recorded shifts

This project figure reports selected endpoint shifts for the completed paired comparisons that inform the stability reading. It is a set of study-specific measurements, not a single quality score.

Project endpoint shifts by model with 95 percent confidence intervals

Reading the plot: points report variant-minus-base deltas and horizontal bars show 95% confidence intervals for available endpoints. M marks an unavailable comparison; NM marks a present but non-measurable comparison.

Interpretation boundary

This is a methodological diagnostic. It identifies potential brittleness in the measured control effect; it does not establish clinical safety, deployment readiness, or a completed model ranking.