Fine PT-PT Web: 41B Token European Portuguese Corpus
September 9, 2026
A new pipeline for curating European Portuguese web data yields a 41 billion token corpus from 411 TB of raw data. The method uses an early-stage boilerplate removal block to increase document yield by 19.04% compared to standard heuristic filters.
HOW THIS AFFECTS YOU
●
builderYou can leverage this cleaned corpus for specialized language modeling or fine-tuning in the PT-PT market.
●
researcherYou can use this high-quality, dialect-specific dataset to improve LLM performance on European Portuguese.