AI Governance

Glossary

Model Evaluation

Systematic measurement of an AI model’s capabilities, limitations, risks and performance on defined tasks.

Last reviewed 2026-08-22

In plain language

Evaluation includes academic benchmarks, safety tests, domain-specific trials and real-world monitoring. Benchmarks can saturate or fail to represent local languages, cultures and high-stakes settings. Governance debates now include public evaluation infrastructure and independent access to models. Evaluation results are often incomparable across vendors because protocols differ.

Why it matters for AI governance

Without evaluation, claims about safety, accuracy or fairness are unfalsifiable. Public-interest evaluation capacity is a development and democracy issue, not only a research issue.

Authoritative sources