Headline results
- 98.5%
- within one point of the official TEA score, pooled across eight modules
- 79.3%
- exact match with the official TEA score; 72.4% on extended responses alone
- 3 sec
- per response, with rubric-aligned feedback attached
TEA’s own agreement standard for two human raters
Two trained TEA raters agree exactly 63% to 75% of the time
Hand scoring runs 15 to 20 minutes per essay
Summary
The Texas Education Agency’s 2023 STAAR redesign expanded constructed-response items sharply, producing six to seven times more written items per test. To score the operational exam at that volume, the state moved to a hybrid model that pairs an automated scoring engine with human raters. Year-round practice writing was left where it was. Teachers still hand-score every practice essay, benchmark, and interim, at a reported 15 to 20 minutes each.
CoGrader brings semi-automated, TEA-aligned scoring into the classroom. The system evaluates student practice writing against the state’s published rubrics and returns a score with rubric-aligned feedback in roughly three seconds. The teacher reviews it, adjusts where needed, and holds final authority over what reaches the student.
This paper describes the system’s design, the validation evidence behind the headline figures, and the limits of what those figures support.
Does CoGrader score student writing the same way official TEA raters score it?
Agreement with human raters
Trained human raters do not match each other every time. The fair test for an automated scorer is whether it agrees with a trained rater about as often as a second trained rater would. The chart compares exact agreement rates. Read the CoGrader extended-response row against the two human ranges, which are for extended responses only.
CoGrader, all eight modules
600 responses, pooled
79.3%CoGrader, extended responses only
283 essays
72.4%Two trained TEA raters, STAAR ECR
Spring 2023, by grade and trait
63% to 75%Two trained NAEP raters, six-point writing
Holistic scale
57% to 67%
Sources: TEA & Cambium Assessment (2023), Table B7, pp. 37 to 38; National Center for Education Statistics, NAEP technical documentation.
Why within one point is the right bar
The within-one-point threshold is not a benchmark we chose. It follows TEA’s own rule for human ECR scoring. Under STAAR’s scoring protocol, each human-scored extended response is read by two trained scorers, and the two scores are added. On each rubric trait, scores that are one point apart both stand. Scores that are further apart go to a third resolution read.
For short constructed responses, TEA requires an exact match between scorers. So for SCR items, exact agreement is the fair comparison.
On STAAR ECRs in spring 2023, two trained TEA raters gave the exact same score on a rubric trait about 63% to 75% of the time, depending on grade and trait. NAEP treats 60% exact agreement as acceptable for a complex six-point writing item. The relevant question for an automated scorer is therefore not whether it matches the official score every time. Trained human raters do not. It is whether the system agrees with a trained rater about as consistently as a second trained rater would.
How the system is built
The CoGrader system combines a current-generation production large language model with three design elements.
Calibrated to TEA’s own standard
Each module uses the publicly released TEA STAAR scoring guides as in-context calibration material. No custom scoring model was trained for this work.
One module per item type and grade band
STAAR does not use one writing rubric. Each item type is scored on its own scale, and the same item type is scored differently across grade bands. So the system is eight task-specific modules, each with its own calibration logic.
A teacher holds final authority
The teacher reviews each score, adjusts where needed, and decides what the student sees.
Where the judgment comes from
Some scoring systems learn from a private collection of student essays. When that happens, a district has no way to check what the system learned, or whether the standard it absorbed matches the one Texas applies.
CoGrader reads TEA’s publicly released scoring guides directly. The standard CoGrader applies is the state’s published standard, and any teacher or administrator can read the same guides CoGrader reads. The scoring engine identity, configuration files, and calibration logic are documented internally and shared with district auditors under a confidentiality agreement.
What the system scores
STAAR does not use one writing rubric. Each item type is scored on its own scale, and the same item type is scored differently across grade bands. The system is therefore eight task-specific modules, each calibrated against the TEA scoring guide for that exact item, grade band, and language.
| Module | Grade band | Language | Rubric scale |
|---|---|---|---|
| ECR Informational | 3 to 5 | English | 0 to 5 |
| ECR Argumentative | 3 to 5 | Spanish | 0 to 5 |
| ECR Informational | 3 to 5 | Spanish | 0 to 5 |
| ECR Informational | 6 to 8 | English | 0 to 5 |
| ECR Argumentative | 6 to 8 | English | 0 to 5 |
| ECR Informational | High school (English I and II) | English | 0 to 5 |
| Short response, reading | All | English | 0, 1, 2 |
| Short response, writing, revising, and editing | All | English | 0, 1 |
Six modules cover Extended Constructed Response (ECR) and two cover Short Constructed Response (SCR). Argumentative coverage is currently partial, with modules at Grades 6 to 8 English and Grades 3 to 5 Spanish. Additional modules are in development.
Results
The system was benchmarked on 600 essays drawn from publicly released TEA scoring guides and TEA-approved rater training material. These scores constitute the formal gold standard TEA uses to train its own human raters.
Pooled across the eight modules, the system’s score landed within one point of the TEA score in 98.5% of cases, and matched it exactly in 79.3% of cases. On extended response alone, the figures are 96.8% within one point and 72.4% exact. On NAEP’s six-point holistic writing scale, two trained raters agree exactly 57% to 67% of the time. CoGrader’s rate sits above that range.
| Module | Items | Exact match | Within one point |
|---|---|---|---|
| SCR, writing | 166 | 89.8% | 100.0% |
| SCR, reading | 151 | 80.8% | 100.0% |
| ECR, all modules | 283 | 72.4% | 96.8% |
Extended response, by module
| ECR module | Essays | Exact match | Within one point |
|---|---|---|---|
| Grades 3 to 5 informational (English) | 72 | 84.7% | 98.6% |
| Grades 6 to 8 informational | 69 | 75.4% | 97.1% |
| Grades 6 to 8 argumentative | 15 | 73.3%* | 100.0% |
| High school informational | 61 | 70.5% | 95.1% |
| Grades 3 to 5 argumentative (Spanish) | 12 | 66.7%* | 91.7% |
| Grades 3 to 5 informational (Spanish) | 54 | 55.6% | 96.3% |
* Small samples; indicative only.
Limits
Two considerations qualify these results.
The corpus is the cleanest possible writing
All 600 responses are TEA scoring examples and rater training material, the clearest illustrations of each score point. Results should be read as a ceiling relative to operational classroom writing, which is messier.
Calibration and validation share a source
The same public scoring guides supply both the in-context calibration material and the validation corpus. We separated them mechanically. Fidelity on responses drawn from outside that material has not yet been measured.
How districts evaluate it
Adoption follows the phased structure districts already use for scoring changes. Nothing is signed before the trial.
Evidence review
The assessment team reads this paper, module by module, including the limitations above.
Scoped trial
A live demonstration on district-selected essays, then a blind agreement test on essays the district’s own raters have already scored.
Implementation decision
Module fit is reviewed alongside operational requirements: integration, FERPA review, audit logging, arbitration workflow, and drift monitoring.
Districts have good reason to scrutinize any scoring instrument that reaches student writing. The approach here is to earn that scrutiny: publish the design, validate against the state’s own gold standard, disclose limitations before others find them, and keep the teacher in final authority over every score.
References
- Alsalem, M. S. (2024). EFL teachers’ perceptions of the use of an AI grading tool (CoGrader) in English writing assessment at Saudi universities: An activity theory perspective. Cogent Education, 11(1), Article 2430865. https://doi.org/10.1080/2331186X.2024.2430865
- Cambium Assessment. (n.d.-a). Communicating to the public about machine scoring [White paper]. https://files.portal.cambiumast.com/corporate-site/documents/CAI-Cambium-CommunicatingPublicMachineScoring-WhitePaper.pdf
- Cambium Assessment. (n.d.-b). Smarter Balanced interim assessments: Automated scoring FAQ. https://smarterbalanced.alohahsap.org/content/contentresources/en/Smarter_Interim_Automated_Scoring_FAQ.pdf
- Chen, Z., Wang, J., Li, Y., Li, H., Shi, C., Zhang, R., & Qu, H. (2025). CoGrader: Transforming instructors’ assessment of project reports through collaborative LLM integration [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.20655
- National Center for Education Statistics. (n.d.-a). NAEP technical documentation: Constructed-response interrater reliability. https://nces.ed.gov/nationsreportcard/tdw/analysis/initial_itemscore.aspx
- National Center for Education Statistics. (n.d.-b). NAEP technical documentation: Scoring. https://nces.ed.gov/nationsreportcard/tdw/scoring/
- Texas Education Agency. (n.d.). STAAR redesign. https://tea.texas.gov/student-assessment/assessment-initiatives/staar-redesign
- Texas Education Agency. (2022, October 20). STAAR redesign constructed response resources and updated calendar of events [To the Administrator Addressed letter]. https://tea.texas.gov/about-tea/news-and-multimedia/correspondence/taa-letters/staar-redesign-constructed-response-resources-and-updated-calendar-of-events
- Texas Education Agency & Cambium Assessment. (2023). STAAR hybrid scoring study, methods and results: Spring 2023 items. https://tea.texas.gov/data-reports/reports-and-studies/2023-staar-hybrid-scoring-study-2.pdf
- Texas Tribune. (2024, April 9). How Texas will use AI to grade this year’s STAAR tests. https://www.texastribune.org/2024/04/09/staar-artificial-intelligence-computer-grading-texas/
Cite this paper
CoGrader Research Team. (2026). Within one point of TEA raters in 98.5% of cases: The design, validity, and limits of CoGrader’s automated STAAR constructed response scoring [Whitepaper]. CoGrader.
The research reported here was supported by the Institute of Education Sciences, U.S. Department of Education, through Grant R305J250071 to Modern Learner Media, LLC. The opinions expressed are those of the authors and do not represent views of the Institute or the U.S. Department of Education.
Test it on your own raters’ essays
A 30-minute discovery call to scope a blind agreement test with your assessment team.
