Grading-prompt robustness
Some rubrics narrow the identity gap; prompt length alone does not
The gap is Amanda Askell’s published-address grade minus the matched Gmail grade, normalized by each rubric’s score scale. More negative values mean harsher grading under Amanda’s identity.
p<.05not significantWhiskers: 95% CI
Simple mitigations
Grading “Claude” (main result)125 chars
-3.79***
Grading “GPT”122 chars
-2.49***
Grading “an AI model”130 chars
-2.49***
Objective + anti-sycophancy260 chars
-2.11***
Published rubrics
MT-Bench: general response quality623 chars
-1.31**
UltraFeedback: constructive feedback1,201 chars
-1.36***
Prometheus: relevance + usefulness1,404 chars
-0.55
Model-written rubrics
5-step checklist (Fable)2,912 chars
-0.55
Tutorial + bias guards (Fable)6,034 chars
-0.40
5 axes + score caps (GPT-5.5 Pro)7,607 chars
-1.18***
Task-specific criteria (GPT-5.5 Pro)5,702 chars
-0.54
−5%−2.5%0+1%
identity gap as percentage of score scale