Contextual Awareness is All You Need: Intro
An introduction to our ten-part series
Welcome to the Civitaas Insights newsletter. Civitaas offers testing-as-a-service to support decision making about advanced technology and to learn about its effects on our culture, society and economy.
My name is Reva Schwartz and I am the co-founder of Civitaas Insights. Over the next few months our team will roll out a ten-part series, sharing what we think is necessary to mature AI into tools that provide value to people in the real world. The series will include writings from myself, Civitaas co-founder Gabriella Waters, and Dr. Rumman Chowdhury, CEO of Humane Intelligence, where Civitaas is being incubated. We will release our newsletters every two weeks between now and early September. This introduction is a sneak peek into what you can expect throughout the series.
Gabriella, Rumman and I are each research scientists who have built, implemented, used, evaluated, and overseen machine learning technologies deployed across various settings. The one constant in our efforts is measurement, without it, we can never know what does or doesn’t work or substantiate claims about the world around us.
I am a linguist and research scientist who has worked in technology evaluation for more than 20 years, initially in the field of human language technology focusing on the speech domain (i.e. machine translation, speech recognition, and speaker recognition). As a research practitioner and former civil servant I’ve worked at and advised agencies like USSS, NIST, DARPA, IARPA, DHS and in the intelligence community on how to build technology that can support subject matter experts in high-risk operational settings and without “ground truth.” Having worked alongside computer scientists and engineers I came to quickly recognize that “better performing” ML technology is typically defined within computational constraints – the algorithms and models used in machine learning, and the data that feed them. While this computational focus has fueled innovation, it also produces various gaps that can lead to mistakes, risks, incidents and negative impacts. These gaps can also create opportunities and “off-label” uses, since many ML users regularly leverage these tools in new and beneficial ways that nobody on the development team considered.
As emerging technologies have rapidly changed many aspects of society, we’ve witnessed negative effects like “enshittification”, increased surveillance, and algorithmic monoculture (or what we like to call “the great flattening”). These harmful outcomes have helped to drive the debate around AI’s role in our lives and subsequently led to various signifiers of what AI “should be” – like trustworthy and responsible, ethical, and “value-aligned”. The broader debate about AI has also:
increased awareness of the role that human and societal factors play in AI’s design, development, deployment, and use;
driven interest in understanding AIs secondary and tertiary effects (these are any long-term outcomes and consequences that may result from AI use in the real world).
We believe that the AI evaluation community is at an inflection point.
The “wicked problems” posed by AI and its secondary effects reside in the real world, and are not really about the inner workings of the technology. Yet, the current AI evaluation toolbox remains firmly entrenched in the computational frame of the AI stack – model-centric, quantitative, static and isolated from the real world. The computational toolbox is designed to measure the first order outputs of technology, which cannot directly transfer to claims about AI’s secondary effects.
For example, many users of AI chatbots have noticed the sycophantic nature of their output, which researchers suggest may arise from approaches to drive AI system “helpfulness”. Computationally-framed evaluation methods seek to 1) discover whether the system produced this type of output, and 2) solve it by tweaking pieces of the computational puzzle – the data, the model, the algorithm. But the sycophantic nature of model/system output lies outside of the computational frame. AI system behavior also relates to the contextual setting from which training data are derived and –to a much larger degree– decisions and pitfalls within the socio-technical contexts where technology is designed and developed [For more see: ACM FAccT 2023 and Selbst et al].
The subject matter expertise and skills necessary to systematize, translate and contextualize AI’s effects in the real world mostly reside in non-computational disciplines. Part of Civitaas' mission is to bridge these various communities and enable evaluation practices that can account for AI’s secondary effects. A collaborative playing field for interdisciplinary testing and evaluation can foster methods for leveraging context instead of reducing human behavior to the narrow choices used in alignment-based model training protocols. A contextual lens is also necessary to systematize and operationalize the complex behavioral and social phenomena that underlie model development, so AI’s effects can be observed and managed along the lifecycle.
Contextual Awareness is All You Need
Over this series we will dive into the many reasons why the current computational style of evaluation persists and discuss how to better account for the messy and contextually rich reality of how AI intersects with people’s everyday lives. We promise that this is possible, not as hard as it seems, and even a lot of fun.
We’ve titled our 10-part series “Contextual Awareness is All You Need” because we think that context–and contextual factors– are the missing link for improving AI evaluation and advancing technologies that provide practical utility [see note 1]. For these reasons, it is context that serves as the primary focus of our 10-part series. We will start with three newsletters that set the stage for “contextually aware measurement.” Dr. Chowdhury will pen newsletters #2 and #3, describing AI testing and evaluation practices and measurement from a contextual/real-world lens. I will be back for newsletter #4 to describe the relationship between context and the measurement validities. Next Gabriella and Rumman will cover experimental design and quasi-experimental design methods for newsletters #5 and #6 respectively.
There has been growing interest in the use of measurement scenarios across parts of the AI evaluation community. Our previous work in the NIST ARIA program provided exposure to scenario-based testing for AI risk measurement. Gabriella will focus on this topic in newsletter #7 and describe the design of proxy-based scenarios for real-world testing paradigms. In newsletter #8 I will describe how context drives systematization and operationalization of real world concepts so they can be modeled. We close out the series with Gabriella focusing on human subject research protocols in newsletter #9, and in newsletter #10, I will describe why interdisciplinary skill sets are necessary for eliciting and capturing context.
We hope you enjoy our upcoming series. Please subscribe to the newsletter and join us in the comments for the broader discussion!
[Note 1] This title is a bit of a provocation to the well-known paper “Attention is All You Need” that introduced the “transformer”, a deep learning architecture that broadly contributed to advances in large language models.
