
Artificial intelligence
Training data and annotation
Datasets with a documented annotation standard and measured inter-annotator agreement, not merely labelled.
In summary: preparation, labelling and quality control of datasets for training and evaluating models, with a documented annotation guide and measured inter-annotator agreement. Work in Spain Spanish and co-official languages is done by native speakers.
The quality of a labelled dataset is not measured by the number of examples, it is measured by consistency between annotators. Two people labelling the same case differently means the annotation guide is ambiguous, and no amount of volume corrects that.
So the first deliverable is not labels, it is the annotation guide, validated on a small sample until inter-annotator agreement reaches the agreed threshold.
What is included
According to data type and task.
- Developing and validating the annotation guide with you
- Labelling of text, audio, image and document data
- Annotation in Spain Spanish and in Catalan, Basque and Galician by native speakers
- Measuring inter-annotator agreement and resolving disagreements
- Evaluation sets independent of the training set
- Sample-based quality control against an agreed threshold
- Traceability of which annotator labelled each case
What is not included
Service limits and ethical limits.
- Model training, which is a separate project
- Acquiring third-party data on our own account
- Labelling data obtained without a valid legal basis, which we do not accept
- Annotation of extreme content without wellbeing protocols agreed in advance
- A performance guarantee for the resulting model, which depends on much besides the data
Where it is delivered from
Spain Spanish and the co-official languages are annotated from Spain, by native speakers. It is the only way to capture register, variant and nuance, which is precisely what a peninsular Spanish dataset should contribute over a generic one.
Volume labelling in general Spanish is delivered from the Latin America corridor, and work on data that must stay in the European Union from Poland.
Provenance and minimisation
The question we ask before accepting a labelling project is where the data comes from and on what legal basis it is processed. That is not a formality: if a dataset contains personal data obtained without a valid basis, labelling it does not solve the problem, it amplifies it.
Where the dataset contains personal data, minimisation gives the best ratio of effort to risk reduction. Many annotation tasks need no identifying fields, and removing them before annotating simplifies processing, transfer and later deletion.
If labelling takes place outside the European Economic Area, standard contractual clauses apply and, depending on sensitivity, a transfer impact assessment.
For annotation of potentially distressing content, the same wellbeing protocols apply as in content moderation: rotation, exposure limits and available support.
Frequently asked questions
How is labelling consistency guaranteed?
By measuring inter-annotator agreement on overlapping cases and adjusting the annotation guide when agreement falls. A recurring disagreement almost always signals ambiguity in the guide, not annotator error.
Can you annotate in Catalan, Basque or Galician?
Yes, from Spain and with native speakers. It is one of the reasons these datasets have value: approximate annotation in a language one does not command produces data that degrades the model rather than improving it.
Who owns the resulting dataset?
You do, along with the annotation guide and the quality metrics. Worth setting out expressly in the contract.
Do you also prepare evaluation sets?
Yes, and they should be independent of the training set and closed before training starts. Evaluating on data seen during training produces excellent and meaningless figures.
Need labelled data in Spanish?
Tell us the task, the volume and whether the dataset contains personal data.