The problem
The models the platform used for moderation and recommendation worked reasonably in English and considerably worse in Spain Spanish. Errors concentrated in nuance: irony, regional colloquialisms, and expressions a Latin American Spanish speaker reads differently.
In Catalan, Galician and Basque, coverage was simply insufficient, and the platform had no way to measure how insufficient, because it had no evaluation sets in those languages.
What we did
Corpshore AI built the evaluation sets first, not the training sets: without a way to measure, any later improvement would have been an impression. The sets were built with native speakers of each language, with written guidelines and cross-review between annotators.
From there, the team annotated training data focused on the cases where the system failed, which is where new data earns its keep rather than in general volume.
Evaluation is continuous: every system version is measured against the same sets, so the platform can state whether it has improved rather than assume it.
Team shape
A team of twelve annotators and three senior evaluators, with native speakers of Castilian, Catalan, Galician and Basque, split between Spain and the corridor.
Timeline
Six weeks for the initial evaluation sets, a 30-day annotation pilot on failure cases, and continuous operation from month three.
Data handling and compliance
The content handled is content published on the platform. Data leaving the European Economic Area rests on standard contractual clauses, and content with identifiable personal data is pseudonymised before it reaches the annotation team.
Results
- +17 pts
- Moderation accuracy in Spain Spanish
- 0 to 4 languages
- Evaluation coverage in co-official languages
- -31 %
- Moderation false positives
- -24 %
- Cases escalated to human review
Measured on the evaluation set built at the outset.
No evaluation set previously existed in Catalan, Galician or Basque.
Legitimate content the system removed in error.
After tuning the system on real failure cases.
Figures relate to the period stated in each case and depend on each client's starting point.
What changed the project was starting with measurement. We had spent two years improving the model without being able to say whether it was improving in Spanish.
