Skip to main content
Corpshore España
A mudejar lattice screen casting a light pattern onto stone

Artificial intelligence

AI system evaluation and safety

Measuring how the system behaves before your customers discover it.

In summary: systematic evaluation of AI systems, with independent test sets, edge case and bias analysis, and the documentation the European framework requires. Evaluation is independent of whoever builds the system, which is where its value comes from.

An AI system that works well in the demo and badly in production is not a technical failure, it is an evaluation failure. The demo runs on the cases the system handles; production brings the ones it does not.

Evaluating seriously means building the test set from the real cases that worry you, the uncomfortable ones included, and measuring before a customer finds them.

What is included

According to the system and its associated risk.

  • Building evaluation sets independent of training data
  • Measuring output quality against agreed criteria
  • Testing on edge cases and adversarial inputs
  • Bias analysis across the dimensions relevant to your case
  • Specific evaluation in Spain Spanish and co-official languages
  • Documentation of results with a reproducible methodology
  • Periodic re-evaluation to detect degradation

What is not included

What an evaluation cannot offer.

  • Certification of conformity with the AI Act, which we do not issue
  • A guarantee that the system will not fail in production
  • Fixing the problems found, which belongs to whoever develops the system
  • Independent audit where we also provide the managed service on that same system
  • A legal opinion on your system's risk classification

Where it is delivered from

Evaluation in Spain Spanish and co-official languages is delivered from Spain, by native speakers, for the same reason as annotation: nuance is exactly what is being measured.

Technical evaluation and test infrastructure are delivered from Poland or Uzbekistan depending on the coordination model.

European framework and timetable

The transparency obligations of Article 50 of the AI Act have applied since 2 August 2026.

Obligations for high-risk systems apply later: December 2027 for Annex III cases and August 2028 for Annex I. That timetable is useful because it allows evaluation documentation to be prepared in good time rather than improvised.

It is worth stating clearly what we do not do: we do not issue conformity certifications and we do not assess your system's legal risk classification. We produce reproducible technical evidence about how it behaves, which is an input to that analysis, not the analysis.

There is also an incompatibility we would rather declare: if we operate your system under a managed service, we are not the right party to audit it independently. An evaluator who agrees to assess their own work should worry you.

Frequently asked questions

Do you issue AI Act conformity certifications?

No. We produce reproducible technical evidence about system behaviour. Conformity and risk classification belong to your organisation and your advisors.

Can you evaluate a system you operate?

Not independently. If we provide the managed service on that system, independent evaluation should be commissioned from a third party. We say so even though it means declining work.

What does evaluating bias mean?

Measuring whether the system behaves systematically differently across dimensions that should not influence the outcome. Which dimensions are relevant depends on your use case and is agreed before starting.

How often should re-evaluation happen?

As often as the system or its inputs change. A stable system on stable data degrades slowly; one updated frequently needs continuous evaluation.

Do you know how your system behaves on the hard cases?

Tell us what the system does and which failures would worry you most.