User reports indicate qualitative performance improvements in Opus 5 medium relative to GPT-4o. While specific benchmarks are not provided, the model demonstrates high-reasoning capabilities and unexpected creative outputs in production environments.