gavel Experiment · Proposed

Safety Judge Stress Lab

Which content-invariant wrapper transformations produce the highest disagreement across automated safety judges, and can simple ensembles resist them?

Inspectable Agent Systems

Hypothesis

A small ensemble of judges resists content-invariant wrapper attacks meaningfully better than any single judge, but not completely — some wrappers should transfer across models.

Background

Style Over Substance (arXiv 2609.08236) reports up to a 19.9% unsafe-verdict flip rate for a GPT-4o-mini judge under a token refusal wrapper, vs. a 0.5% noise floor — with a verified open repository.

Method

Reproduce the flip-rate baseline; test single-judge vs. majority-vote vs. alternate-rubric ensembles against the same wrapper set.

Measurement Plan

Flip rate, noise floor, and ensemble agreement across wrapper families.

Expected Artifact

A judge-robustness dashboard comparing single-judge and ensemble resistance.

Limitations

Limited wrapper family and judge/dataset scope in the source paper; agent-trajectory judging (not just single replies) remains future work.

Status

This experiment is Proposed — queued, not started. Results, data, and measurements will appear here once it actually runs.

Sources

Part of