The problem
The group had been treating Latin American Spanish as one thing with accents, training on data translated into regional variants that was grammatically regional but pragmatically Castilian, which improved benchmark scores while leaving real-world performance flat. Its intent, search and moderation models misclassified across variants.
What we did
- Organised 58 annotators into five variant desks: Caribbean, Mexican, Andean, Rioplatense and Castilian
- Extended the label schema to capture pragmatic features, not just semantics
- Drew on Punta Cana's genuinely multi-variant, pan-Latin-American labour pool
- Built a divergence taxonomy across the Spanish variants
Results
Real-world model performance improved where benchmark-driven work had not, because the annotation captured the pragmatic divergence, directness, politeness register and urgency signalling, that was the actual source of misclassification. Full metrics are in the downloadable portfolio.
Figures relate to the period stated and were provided by the client. They depend on that client's starting point, so they are not a forecast for another operation.
We had been treating Latin American Spanish as one thing with accents. The divergence taxonomy their linguistic leads built changed how we think about the entire Latin American market, not just the data.
What stayed in place
The variant divergence taxonomy is a client-owned asset now shaping the group's Latin American product strategy, and the five-variant annotation model is a Corpshore Dominicana standard.
