About the role
This role reviews what a model answers and decides whether it meets the criteria agreed with the client. It is judgement work, not template work: most of the value sits in the edge cases and in arguing clearly why a response fails.
What you will do
- Evaluate model responses against quality, bias and safety rubrics.
- Document the reasoning behind each evaluation in an auditable way.
- Identify failure patterns and communicate them to the client with examples.
- Help refine rubrics when uncovered cases appear.
- Take part in calibrations to keep consistency between evaluators.
What we need
- Demonstrable analytical and written argumentation ability.
- Native Spanish and professional reading-level English.
- Valid right to work in Spain.
- Sound judgement on nuances of tone, accuracy and safety.
- Autonomy to work remotely from written procedures.
Useful, but not essential
- Previous experience in model evaluation, moderation or editorial review.
- Background in linguistics, law, journalism or related fields.
- Demonstrated interest in AI system safety.
Equal opportunities
Corpshore selects its teams on competence and fit for the role. We do not discriminate on grounds of age, sex, origin, sexual orientation, gender identity, religion, disability or any other personal or social circumstance. If you need an adjustment during the selection process, say so in your application and we will agree it with you.
