An LLM-based approach using prompt engineering and token probability classification achieves 0.90 accuracy in detecting outcome switching in clinical trials. The system outperforms baseline text similarity models by identifying discrepancies between primary and reported study outcomes.
HOW THIS AFFECTS YOU
●
researcherThis demonstrates the utility of token-level classification for high-stakes domain auditing.
●
healthYou can use these tools to audit medical literature for reporting biases and statistical spin.