Safety Judge Stress Lab
Which content-invariant wrapper transformations produce the highest disagreement across automated safety judges, and can simple ensembles resist them?
Hypothesis
A small ensemble of judges resists content-invariant wrapper attacks meaningfully better than any single judge, but not completely — some wrappers should transfer across models.
Background
Style Over Substance (arXiv 2609.08236) reports up to a 19.9% unsafe-verdict flip rate for a GPT-4o-mini judge under a token refusal wrapper, vs. a 0.5% noise floor — with a verified open repository.
Method
Reproduce the flip-rate baseline; test single-judge vs. majority-vote vs. alternate-rubric ensembles against the same wrapper set.
Measurement Plan
Flip rate, noise floor, and ensemble agreement across wrapper families.
Expected Artifact
A judge-robustness dashboard comparing single-judge and ensemble resistance.
Limitations
Limited wrapper family and judge/dataset scope in the source paper; agent-trajectory judging (not just single replies) remains future work.
Status
This experiment is Proposed — queued, not started. Results, data, and measurements will appear here once it actually runs.