Meddies-PII Multilingual Synthetic Dataset for Clinical De-identification
September 14, 2026
Meddies-PII provides a synthetic corpus of one million clinical documents across 17 languages and 9 PII labels. A trained BIOES token classifier using this data achieves a mean F1 of 0.827 across fifteen external benchmarks.
HOW THIS AFFECTS YOU
●
builderYou can use this large-scale multilingual dataset to train high-accuracy clinical PII extractors.
●
healthThis enables better automated de-identification of multilingual medical records for privacy compliance.