A language model scoring student peer feedback struggled on the same rubric criteria that human raters found hardest
Applying a four-part feedback-quality rubric to 295 peer comments, an open-source language model agreed well with researchers on whether a comment named a problem or suggested a fix, and poorly on the two criteria where the researchers also agreed least.
Agreement between the two researchers (open circles) and between the model and each researcher (maroon bars spanning the two values), as quadratic weighted kappa. The model fell furthest behind on task contributions and behavior, the two criteria where the researchers also agreed least. On action it slightly exceeded the researchers' agreement with each other.
Drawn from Drinkwater Gregg et al. (2025), Table 5, using the kappa values printed in the paper. Original article CC BY-NC 4.0.
When an open-source language model scored first-year engineering students’ peer feedback against a four-part quality rubric, it did best on the criteria where two human researchers agreed most, and worst where they agreed least. On whether a comment identified a gap in a teammate’s performance, the model’s agreement with each researcher (quadratic weighted kappa) was 0.78 to 0.79, against 0.88 between the researchers. On whether a comment suggested an action to fix it, the model reached 0.85 to 0.86, against 0.83. On the two criteria describing a teammate’s contributions to tasks and their behavior, the model fell to between 0.26 and 0.40. The researchers had agreed least there too (0.58 and 0.63).
The disagreements had a pattern. The model read “vague” more broadly than the researchers did: it gave about 80 percent of comments a middle score for task contributions, where the researchers gave that score about 40 percent of the time. Its written justifications were coherent, but it sometimes scored near-identical comments differently.
Why it matters
Rubrics are a common way for education researchers to turn written responses into data, and scoring hundreds of comments twice by hand is slow. This study shows where a local model can share that load and where it cannot yet. The authors’ practical advice is to use a model first as a pilot tester for the rubric itself, since its disagreements point to definitions that leave too much room for judgment, and to verify the process carefully before relying on a model for analysis.
How the lab did it
Katherine Drinkwater Gregg, a PhD candidate in Virginia Tech’s Department of Engineering Education, led the study with fellow PhD candidate Olivia Ryan, Andrew Katz, Mark Huerta, and Susan Sajadi. The data were 295 peer comments written by 118 consenting students in a project-based first-year course. Two researchers scored every comment on four criteria (task contributions, behavior, gap, and action), each from 0 to 2, after revising the rubric when a first round of scoring produced poor agreement. The team then ran qwen-2.5-32b, an open-source model, locally, after a second model stripped names and gendered pronouns from the comments. The model scored one criterion per call, with a fresh conversation for each comment and its temperature set to zero. Getting there took three iterations of the rubric, more than seven prompt revisions, and pilots of five models.
The comments came from one semester of one course whose instructor teaches feedback explicitly, so they may not represent first-year students elsewhere. More than 70 percent of comments suggested no action at all, and the authors note that agreement is easier when there is nothing to judge, which may explain part of the model’s stronger showing on gap and action. Only one model scored the full dataset. Scoring criteria separately let the model count the same phrase twice, and it read grammatically loose comments literally where the researchers inferred the meaning.
The open question is whether sharper rubric definitions can make a model reliable on the subjective criteria, and whether that holds for longer, messier text such as interview transcripts.
The paper
Drinkwater Gregg, K., Ryan, O., Katz, A., Huerta, M., & Sajadi, S. (2025). Expanding possibilities for generative AI in qualitative analysis: Fostering student feedback literacy through the application of a feedback quality rubric. Journal of Engineering Education. https://doi.org/10.1002/jee.70024