Measuring What Matters: Validity in AI Evaluations
Contextual Awareness is All You Need Post #4: Reva Schwartz
AI incidents and failures are known to occur when systems are moved from the controlled lab environments in which they are designed and tested to the messy real world. The OECD defines AI incidents as events where AI systems cause harm to people, infrastructure, rights, or the environment. The gap between where AI is tested and where it is used can create serious risks. Validity helps bridge this gap by making sure that AI evaluations actually measure what they claim to.
What is Validity?
Validity can be thought of as “the best available approximation of the truth.” AI practitioners can leverage types of measurement validity to confirm their test is actually measuring what they think it measures. For example, if you claim to test whether an AI system leaks private information, your test should actually detect privacy leaks – not something else entirely. More broadly, AI practitioners engaged in testing and evaluation activities seek statistical validity — the extent to which conclusions drawn from statistical analyses accurately reflect true relationships among the tested variables and that test results are reliable and can be generalized to the broader population.
Why do Current AI Evaluations Fall Short?
Most AI evaluations lack validity. Vendors can make bold claims about their systems being “safe,” “fair,” or “ethical” without proving that their tests actually measure these qualities. There’s no objective way to confirm that current evaluations assess what they claim to assess.
Validity vs. Validation: A Critical Distinction
Many people confuse validation with validity, but these are different concepts. Validation is like checking answers against an answer key, and shows that your AI system operates within specified requirements. For example, validating an AI hiring application would require confirmation that the system can accurately place job applicants into categories based on known information (such as human judgments). Validity goes deeper than validation and considers what the optimal criteria should be in the first place, and then determines whether the assessment process is able to measure those criteria.
Why is Validity Important?
AI evaluations can assess privacy, fairness, and security, but without validity, claims made about those measures cannot be expected to persist across the varying contexts where systems are used. Even when AI systems perform well on a benchmark test, their results can still lack validity – leading practitioners to draw unsound conclusions about the system’s performance. These outcomes can increase the likelihood of downstream risk and may reduce trust in technology
How is Validity Assessed?
There is no single test for demonstrating validity. Instead practitioners must check whether their evaluation results are valid by considering various “challenges” or “threats”. Threats to validity are ways that your test assumptions may be wrong. Threats can be assessed by evaluating the technology under varying contexts.
Are There Types of Validity?
Four types of validity that are relevant to AI measurement and evaluation are described below, using an example that focuses on assessing AI system alignment to the user context.
We start with construct validity, which can demonstrate whether a test is producing inferences that truly measure the concept the test was designed to measure. For example, surveys and polls are commonly used to measure public preferences (which music streaming service do you like best?) and opinions on certain topics (what do you know about financial regulation and how do you think it will affect you?). But, have you ever looked at the results of an opinion poll and thought “it depends on how the question was worded”? That may be due to poor systematization of the underlying construct. The systematization process includes developing technical and neutral definitions for key concepts and pretesting those definitions with different groups and under different contexts to reduce ambiguity. In this case, the test question is systematized to make sure that it’s measuring the concepts it claims to measure (i.e.preference for streaming services).
While the concepts that underlie AI models can be systematized through the same process used in survey design, systematization is rarely included in AI evaluations. Back to our example of “AI alignment to the user context”. For this we would require a clear and unambiguous definition of “aligned to user context”, along with a method to operationalize that systematized concept. With no systematization, the model would rely on data-driven notions of “alignment to user context” which can easily change based on the use case, the user, their behavior and assumptions and expectations from one prompt or session to the next. This poor systematization of real world phenomena can contribute to concept drift and negative outcomes.
Conclusion validity is whether a test can produce credible or reasonable conclusions about observed relationships. Say you want to assess “alignment to user context” by looking at the propensity of certain keywords in AI chatbot output, but your test is unable to account for sarcasm or other types of sentiment (concepts which would also require systematization to demonstrate construct validity). Not being able to account for sarcasm would mean your test fails to demonstrate conclusion validity since it cannot reasonably assess the observed relationship.
Internal validity is whether a test can support a causal conclusion that is free from external or confounding factors. A common challenge with current evaluations is task contamination, where examples of the assessment task have been exposed to the AI model via its training data. This is a bit like having the answer key as you take the test, and can make the system look like it is performing better than it actually is. Since the test effects are attributable to an external factor (the model’s access to examples of the assessment task), the evaluation does not have internal validity.
System evaluations that focus on one specific impact–such as how well a mental health chatbot aligns to the user’s context–cannot be assumed to generalize to other types of impacts or even the same impact in another setting (such as a legal advice chatbot). External validity is relevant here and refers to when test outcomes from one setting, population or time can generalize, or be applied to other settings, populations, and times. For example, an evaluation would demonstrate external validity if it can assess “alignment to the user’s context” in various settings.
By integrating different types of validity into evaluations AI practitioners can ensure the quality of system measurements before deployment, that AI results reflect a causal—not confounded—relationship, and that results can generalize across contexts.
Additional reading:
Chouldechova, A., Atalla, C., Barocas, S., Cooper, A.F., Corvi, E., Dow, P.A., Garcia-Gathright, J.I., Pangakis, N., Reed, S., Sheng, E., Vann, D., Vogel, M., Washington, H., & Wallach, H.M. (2024). A Shared Standard for Valid Measurement of Generative AI Systems' Capabilities, Risks, and Impacts. ArXiv, abs/2412.01934.
Coston, A., Kawakami, A., Zhu, H., Holstein, K., & Heidari, H. (2022). A Validity Perspective on Evaluating the Justified Use of Data-driven Decision-making Algorithms. 2023 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), 690-704.
Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological bulletin, 52(4), 281–302. https://doi.org/10.1037/h0040957
Larsen, K.R., Lukyanenko, R., Mueller, R.M., Storey, V.C., Parsons, J., VanderMeer, D., & Hovorka, D.S. (2025). Validity in Design Science. ArXiv, abs/2503.09466.
OECD (2024), “Defining AI incidents and related terms”, OECD Artificial Intelligence Papers, No. 16, OECD Publishing, Paris, https://doi.org/10.1787/d1a8d965-en.
Recht, Ben (2022) Machine Learning has a Validity Problem. argmin blog https://archives.argmin.net/2022/03/15/external-validity/
Schwartz, R., Chowdhury, R., Kundu, A., Frase, H., Fadaee, M., David, T., Waters, G., Taïk, A., Briggs, M., Hall, P., Jain, S., Yee, K., Thomas, S., Bhandari, S., Sie, L.W., Lu, Q., Holmes, M., & Skeadas, T. (2025). Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects.
Trochim, W. (2007). The Research Methods Knowledge Base.
