Article 10 — Data and Data Governance
Training, validation and testing data sets must be of relevant, representative, free of errors and complete to the best extent possible — with documented data-governance practices throughout.
What the Article requires
Article 10 governs the data used to train, validate and test high-risk AI systems. Data must be relevant, representative, free of errors and complete. Data-governance practices must address design choices, data collection processes, data preparation (labelling, cleaning, augmentation), assumptions about what the data represents, prior assessment of availability/quantity/suitability, examination of biases, and identification of any gaps or shortcomings. The Article also creates a special-category-data exception for bias detection and correction.
In engineering terms
Implies a training-data lineage capability — provenance, transformation history, quality metrics — that most enterprises don't have for the AI systems they procure as products. Sovereign deployment of customer-trained or customer-fine-tuned models is the cleanest path to evidence; for off-the-shelf vendor systems, the vendor-side documentation must be obtained and validated.
Compliance checklist
- ✓Documented data-collection processes
- ✓Training-data lineage and provenance records
- ✓Bias examination performed and documented
- ✓Data-preparation steps (labelling, cleaning, augmentation) recorded
- ✓Quality, representativeness and completeness assessments
Terms used here
All Articles in the reference · The EU AI Act compliance architecture
Need audit-survivable evidence for Article 10?
MindMap runs a 90-day path from standing start to audit-survivable evidence. Talk to the engineering team.