Representation Learning for Classical Tamil Verse-Commentary Pairs
September 7, 2026
Researchers evaluated recurrent, Transformer, and mBART-style models on 1,262 Classical Tamil verse-commentary pairs. Findings show decoder-only models underperform against a simple 25-word frequency baseline, and token-F1 scores remain low, between 0.02 and 0.20, across tested architectures.
HOW THIS AFFECTS YOU
●
researcherYou can observe how low-resource, structurally complex linguistic pairs impact representation learning performance compared to lexical baselines.