From Pareto to Preference: Personalized Test-Time Scaling via Amortized Agentic Policy Discovery | HACKOBAR_