●researcherYou must ensure evaluation datasets are independent of the prompt context to avoid contaminated results.
●policyThis suggests caution when using LLMs for automated mental health screening based on contaminated benchmarks.
●healthBe aware that high accuracy in certain LLM depression models may be an artifact of training on interview-specific language.