Fair Play or False Start? Using AI-Generated Data to Assess Evidence for Validity and Measurement Bias

Image
Students taking an exam in classroom

To understand what students know and can do, we need assessments and surveys that meet professional standards for measurement quality. These standards help ensure that scores are supported by evidence of their validity, reliability, and fairness. Conducting studies to gather this evidence often requires collecting large numbers of student responses, which can be complex, time consuming, and costly.

Using AI-generated responses that mimic real student responses could make these studies simpler and cheaper. But before early validation studies can use AI-generated responses, it is important to examine whether the psychometric patterns in AI responses are similar to those in real student responses.  
 

AIR's Study: AI- vs. Human-Generated Data

AIR will study when AI-generated data can replace human data in early validation studies. Our experts will develop and test strategies to generate item responses using large language learning models. We will compare the psychometric qualities of assessments and surveys using AI- versus human-generated data. 

We will also study whether large language models replicate existing measurement bias across different groups of human respondents. At the end of the study, we will share recommendations for using AI-generated data in future measurement studies. 
 

Next Steps

Our team is developing strategies to create AI-generated data and testing them with multiple existing assessments and surveys for both academic and non-academic constructs.