●builderRelying on agent trajectories to auto-generate text-based policies may not provide better performance than fixed prompting currently.
●researcherThe disconnect between policy execution and policy learning highlights a core limitation in using natural language as a gradient for agents.