Skip to main content
Corpshore España
Window light casting a grid across a stone table

Training data in Spain Spanish: why performance drops and how to measure it

Corpshore Spain editorial team · · 6 min read

In summary: AI systems typically perform worse in Spain Spanish than in English, and considerably worse in Catalan, Basque and Galician. The first step is not annotating more training data, it is building an evaluation set that tells you whether a version is better or worse.

A Spanish company deploying an AI system over text or voice soon notices a pattern: it works reasonably in English, somewhat worse in Spanish, and considerably worse if the user writes in Catalan, Basque or Galician.

That pattern has a known cause and a fix that is almost always approached in the wrong order.

Why does performance drop in Spain Spanish?

Volume and variety. Most text available for training is in English, and within Spanish the Latin American proportion is far larger than the peninsular one by speaker numbers.

The result is not that the system fails to understand Spanish, it is that it fails on nuance: peninsular colloquialisms, irony, usages a Latin American speaker reads differently, and the register of politeness a Spanish customer expects, which does not match other Spanish-speaking markets exactly.

In Catalan, Basque and Galician the problem is different and more severe: available volume is substantially smaller, and in Basque the structural distance from Romance languages adds a difficulty of its own.

Why start with evaluation rather than training?

Because without a way to measure, any later improvement is an impression. It is the most common and most expensive mistake: thousands of training examples are annotated, the system is tuned, and afterwards nobody can say with data whether it improved in peninsular Spanish or simply changed.

An evaluation set is a group of cases with an agreed correct answer, built by native speakers, with written guidelines and cross-review between annotators. It lets every version be measured against the same yardstick.

Building it first has a second advantage: it shows where the system fails, which allows training data to be annotated against failure cases rather than general volume, which is considerably more efficient.

What does good annotation require?

Native speakers of the specific variety, not of generic Spanish. For Spain Spanish, speakers from Spain; for Catalan, Galician or Basque, native speakers of those languages.

Written, living annotation guidelines. Most of the value in an annotation project is decided by how edge cases are resolved, and those cases have to go into the guideline rather than be resolved differently each time.

And cross-review, because consistency between annotators is what makes a dataset usable. Two annotators labelling the same case differently produce noise, and noise trains just as well as signal.

What data protection obligations arise?

It depends on the data's origin. If user-published content is being annotated, you have to assess whether it contains identifiable personal data and, if so, pseudonymise it before it reaches the annotation team.

If the team works outside the European Economic Area, standard contractual clauses apply and, depending on sensitivity, a transfer impact assessment. Minimisation is again the most effective measure: many annotation projects do not need identifying fields to do the work.

How much data is needed?

Less than usually assumed, if it is well targeted. A solid evaluation set and a few thousand training examples focused on failure cases typically move the result more than a far larger volume collected without criteria.

The useful question is not how many examples, but which specific cases the system fails today and in which language. That question only has an answer if the evaluation set exists, which is where to start.

This article is general information, not legal advice. We work alongside your legal advisors, not in their place.

Frequently asked questions

Why is training on Latin American Spanish not enough?

Because register, expressions and politeness expectations do not match exactly. The system understands, but fails on nuance, which is where a Spanish customer notices the difference.

What is an evaluation set?

A group of cases with an agreed correct answer, built by native speakers with written guidelines and cross-review, allowing each version of the system to be measured against the same reference.

Do Catalan, Basque and Galician need separate annotation?

Yes, if the system will serve those languages. Available volume is substantially smaller and performance is usually considerably worse than in Castilian.

Can real customer data be used for annotation?

With care. You have to assess whether it contains identifiable personal data, pseudonymise where appropriate, and apply transfer instruments if the team sits outside the European Economic Area.

Sources

The data in this article comes from the public sources linked below. If a figure becomes outdated, correct against the source rather than against us.

Does this affect your operation?

Book a discovery call and we will review it against your specific case, or request a proposal with an estimate in euros.