Role-Prompting Fails to Align Model Capabilities with Assigned Personas
October 1, 2026
The RoleCapBench framework demonstrates that reasoning models maintain high-level capabilities even when prompted to assume low-capability personas, such as a kindergartner. Evaluations show that while stylistic voice alignment is successful, underlying mathematical and reasoning proficiency remains decoupled from the assigned role.
HOW THIS AFFECTS YOU
●
researcherThis highlights a significant gap in current instruction-tuning regarding persona-based capability constraints.
●
policyThis finding matters for safety evaluations where model persona might be used to mask underlying dangerous capabilities.