Measuring with Proxies and Scenarios in AI Evaluation
Contextual Awareness is All You Need Post #6: Gabriella Waters
The promise of artificial intelligence hinges on whether the technology actually works as intended when people use it in the real world. This seemingly simple requirement reveals the profound limitations of current AI evaluation approaches, where the gap between laboratory/in silico performance and real-world functionality and effectiveness continues to widen.
As AI systems become embedded deeper into critical sectors—from healthcare diagnostics to financial decision-making—the stakes of inadequate evaluation grow exponentially. Yet, our evaluation ecosystem remains anchored to methods that measure what's convenient rather than what's meaningful. This disconnect between measurement and reality demands a fundamental rethinking of how we approach AI evaluation.
The Measurement Challenge: Beyond the Laboratory
Traditional AI benchmarking excels at answering first-order [1] questions about system capabilities: accuracy rates, processing speeds, and immediate outputs. These metrics serve their purpose in controlled environments where variables are constrained and outcomes are predictable only to falter when confronted with the messy complexities of real-world deployment.
The challenge lies in AI's second-order effects [1]—the longer-term outcomes and consequences that emerge when AI is used by people in the real world. These effects include shifts in user behavior, organizational adaptations, and broader societal ramifications that traditional benchmarks cannot capture. A chatbot might score perfectly on toxicity benchmarks while still causing harm through subtle harmful biases that only surface during extended human interaction.
This limitation becomes especially acute when we consider that AI systems are fundamentally socio-technical systems. Their effectiveness depends not just on computational performance but on how humans interpret, trust, and act upon AI-generated outputs. Understanding these dynamics requires evaluation methods that can account for the complex interplay between technology and human behavior.
Proxies: The Art of Indirect Measurement
In contexts where direct measurement would be impractical or impossible, proxy scenarios offer a pathway to meaningful evaluation. Proxy scenarios serve as indicators that correlate with phenomena we care about but are not so easy to observe. In AI evaluation, proxies can be used to bridge the gap between what we are able to measure and what we need to understand.
Consider the challenge of evaluating AI systems for their potential to enable malicious use. The WMDP (Weapons of Mass Destruction Proxy) benchmark addresses this by measuring "proxy information which correlates with, is neighboring to, or is a component of actual hazardous knowledge." Rather than testing dangerous capabilities directly, the benchmark evaluates related knowledge areas that indicate the potential for harm while maintaining safety boundaries.
What does this look like in healthcare AI evaluation? Researchers use surrogate endpoints to assess long-term patient outcomes without waiting years for definitive results. A proxy scenario might evaluate how an AI diagnostic tool affects physician decision-making patterns to serve as an indicator of eventual patient care improvements.
The power of proxy scenarios lies in their ability to make the immeasurable measurable while preserving essential validity. BUT their effectiveness depends critically on understanding the relationship between the proxy and the target phenomenon. A poorly chosen proxy can lead to optimizing for the wrong outcomes.
Scenario-Based Evaluation: Testing Reality's Complexity
While proxy tasks help us evaluate difficult-to-measure phenomena, scenario-based testing can be used to deploy proxies for evaluating AI under realistic conditions in a repeatable and reproducible manner. Scenarios provide a structured method to capture the complexity of real-world use while maintaining experimental control. Effective scenario design requires balancing authenticity with feasibility.
A scenario-based approach recognizes that AI systems behave differently under varying conditions. A language model might perform flawlessly on isolated queries but fail when faced with the multi-turn, contextually dependent conversations that characterize real human interaction. Scenarios allow evaluators to explore these types of behaviors that may only emerge through extended, realistic use, and the conditions under which they happen.
Scenarios can ease the design of repeatable red teaming and field testing paradigms used in AI evaluation By observing how AI systems perform when people use them over extended periods under regular conditions, field testing can capture a broader spectrum of human-AI interaction. This includes unexpected use cases, adaptation effects, and the gradual degradation of performance that can occur as systems encounter edge cases not represented in training data.
Moving upward through the pyramid shows how evaluation methods progressively sacrifice control for authenticity. Just as Maslow's hierarchy suggests that basic needs must be met before higher-level needs can be addressed, the evaluation pyramid suggests that controlled testing should establish baseline system capabilities before progressing to more complex scenarios.
The Integration Challenge: Building Contextual Awareness
Proxy-based scenarios can be used to structure contextually aware AI evaluations, including combining different but complementary measurement approaches. For example, red teaming evaluations can use adversarial scenarios to build out various tasks that approximate AI system vulnerabilities. Red teamers use creative multi-turn prompting and role-playing to induce malicious use cases and generate evidence about whether risks actually manifest in practice.
The key insight is that proxy-based scenarios enable different but complementary functions. Proxy tasks help evaluators structure the key challenges and phenomena they want to investigate and scenarios provide the context —the realistic conditions–under which these phenomena are likely to emerge.
Consider evaluating an AI system for harmful bias. Proxy-based scenarios can be used to mimic tasks or use cases in which harmful bias can’t be directly studied and the conditions in which it occurs (hiring, housing). The combination reveals not just whether harmful bias materialized, but the conditions under which it happened. When combined with field testing, proxy-based scenarios can shed more light on the type of impact, who was impacted and how often.
The Stakes of Getting It Right
The choice between convenient measurement and meaningful evaluation is not exclusively academic—it shapes the trajectory of AI development and deployment. Systems optimized for benchmark performance may fail catastrophically when deployed in contexts their evaluations never considered. Scenario-based evaluation approaches enable the capture of real-world complexity while maintaining experimental control to allow for the development of more robust, beneficial AI systems.
As AI capabilities continue expanding, the evaluation challenge will only intensify. Systems that can perform increasingly complex tasks will require correspondingly sophisticated evaluation methods. The investment in developing proxy tasks and scenario-based testing today will determine whether we can maintain meaningful oversight of AI systems tomorrow.
Evaluation can be used to support better decision-making about AI development and deployment. By combining the precision and authenticity of proxy-based scenarios and the rich contextual data gathered from field testing, we can build evaluation systems that bridge the gap between laboratory performance and real-world impact.
The future of AI evaluation lies not in choosing between different measurement approaches, but in orchestrating them into coherent systems that can capture both the precision we need and the context we cannot afford to ignore. Getting this balance right may determine whether AI becomes a transformative tool for human benefit or another cautionary tale of technological ambition outpacing our wisdom to guide it.
1. Reality Check: A New Evaluation Ecosystem Is Necessary to Understand AI's Real World Effects https://arxiv.org/pdf/2505.18893.pdf

