Whitepaper, June 2026

Within one point of TEA raters in 98.5% of cases

The design, validity, and limits of CoGrader’s automated STAAR constructed response scoring. Benchmarked on 600 responses that TEA scored and published, across eight scoring modules.

PDF, 165 KB, 8 pages. Written by the CoGrader research team.

Headline results

98.5%
within one point of the official TEA score, pooled across eight modules

TEA’s own agreement standard for two human raters

79.3%
exact match with the official TEA score; 72.4% on extended responses alone

Two trained TEA raters agree exactly 63% to 75% of the time

3 sec
per response, with rubric-aligned feedback attached

Hand scoring runs 15 to 20 minutes per essay

Summary

The Texas Education Agency’s 2023 STAAR redesign expanded constructed-response items sharply, producing six to seven times more written items per test. To score the operational exam at that volume, the state moved to a hybrid model that pairs an automated scoring engine with human raters. Year-round practice writing was left where it was. Teachers still hand-score every practice essay, benchmark, and interim, at a reported 15 to 20 minutes each.

CoGrader brings semi-automated, TEA-aligned scoring into the classroom. The system evaluates student practice writing against the state’s published rubrics and returns a score with rubric-aligned feedback in roughly three seconds. The teacher reviews it, adjusts where needed, and holds final authority over what reaches the student.

This paper describes the system’s design, the validation evidence behind the headline figures, and the limits of what those figures support.

Does CoGrader score student writing the same way official TEA raters score it?

Agreement with human raters

Trained human raters do not match each other every time. The fair test for an automated scorer is whether it agrees with a trained rater about as often as a second trained rater would. The chart compares exact agreement rates. Read the CoGrader extended-response row against the two human ranges, which are for extended responses only.

Exact agreement with the official score
  • CoGrader, all eight modules

    600 responses, pooled

    79.3%
  • CoGrader, extended responses only

    283 essays

    72.4%
  • Two trained TEA raters, STAAR ECR

    Spring 2023, by grade and trait

    63% to 75%
  • Two trained NAEP raters, six-point writing

    Holistic scale

    57% to 67%

Sources: TEA & Cambium Assessment (2023), Table B7, pp. 37 to 38; National Center for Education Statistics, NAEP technical documentation.

Why within one point is the right bar

The within-one-point threshold is not a benchmark we chose. It follows TEA’s own rule for human ECR scoring. Under STAAR’s scoring protocol, each human-scored extended response is read by two trained scorers, and the two scores are added. On each rubric trait, scores that are one point apart both stand. Scores that are further apart go to a third resolution read.

For short constructed responses, TEA requires an exact match between scorers. So for SCR items, exact agreement is the fair comparison.

On STAAR ECRs in spring 2023, two trained TEA raters gave the exact same score on a rubric trait about 63% to 75% of the time, depending on grade and trait. NAEP treats 60% exact agreement as acceptable for a complex six-point writing item. The relevant question for an automated scorer is therefore not whether it matches the official score every time. Trained human raters do not. It is whether the system agrees with a trained rater about as consistently as a second trained rater would.

How the system is built

The CoGrader system combines a current-generation production large language model with three design elements.

  • Calibrated to TEA’s own standard

    Each module uses the publicly released TEA STAAR scoring guides as in-context calibration material. No custom scoring model was trained for this work.

  • One module per item type and grade band

    STAAR does not use one writing rubric. Each item type is scored on its own scale, and the same item type is scored differently across grade bands. So the system is eight task-specific modules, each with its own calibration logic.

  • A teacher holds final authority

    The teacher reviews each score, adjusts where needed, and decides what the student sees.

Where the judgment comes from

Some scoring systems learn from a private collection of student essays. When that happens, a district has no way to check what the system learned, or whether the standard it absorbed matches the one Texas applies.

CoGrader reads TEA’s publicly released scoring guides directly. The standard CoGrader applies is the state’s published standard, and any teacher or administrator can read the same guides CoGrader reads. The scoring engine identity, configuration files, and calibration logic are documented internally and shared with district auditors under a confidentiality agreement.

What the system scores

STAAR does not use one writing rubric. Each item type is scored on its own scale, and the same item type is scored differently across grade bands. The system is therefore eight task-specific modules, each calibrated against the TEA scoring guide for that exact item, grade band, and language.

ModuleGrade bandLanguageRubric scale
ECR Informational3 to 5English0 to 5
ECR Argumentative3 to 5Spanish0 to 5
ECR Informational3 to 5Spanish0 to 5
ECR Informational6 to 8English0 to 5
ECR Argumentative6 to 8English0 to 5
ECR InformationalHigh school (English I and II)English0 to 5
Short response, readingAllEnglish0, 1, 2
Short response, writing, revising, and editingAllEnglish0, 1

Six modules cover Extended Constructed Response (ECR) and two cover Short Constructed Response (SCR). Argumentative coverage is currently partial, with modules at Grades 6 to 8 English and Grades 3 to 5 Spanish. Additional modules are in development.

Results

The system was benchmarked on 600 essays drawn from publicly released TEA scoring guides and TEA-approved rater training material. These scores constitute the formal gold standard TEA uses to train its own human raters.

Pooled across the eight modules, the system’s score landed within one point of the TEA score in 98.5% of cases, and matched it exactly in 79.3% of cases. On extended response alone, the figures are 96.8% within one point and 72.4% exact. On NAEP’s six-point holistic writing scale, two trained raters agree exactly 57% to 67% of the time. CoGrader’s rate sits above that range.

ModuleItemsExact matchWithin one point
SCR, writing16689.8%100.0%
SCR, reading15180.8%100.0%
ECR, all modules28372.4%96.8%

Extended response, by module

ECR moduleEssaysExact matchWithin one point
Grades 3 to 5 informational (English)7284.7%98.6%
Grades 6 to 8 informational6975.4%97.1%
Grades 6 to 8 argumentative1573.3%*100.0%
High school informational6170.5%95.1%
Grades 3 to 5 argumentative (Spanish)1266.7%*91.7%
Grades 3 to 5 informational (Spanish)5455.6%96.3%

* Small samples; indicative only.

Limits

Two considerations qualify these results.

  1. The corpus is the cleanest possible writing

    All 600 responses are TEA scoring examples and rater training material, the clearest illustrations of each score point. Results should be read as a ceiling relative to operational classroom writing, which is messier.

  2. Calibration and validation share a source

    The same public scoring guides supply both the in-context calibration material and the validation corpus. We separated them mechanically. Fidelity on responses drawn from outside that material has not yet been measured.

How districts evaluate it

Adoption follows the phased structure districts already use for scoring changes. Nothing is signed before the trial.

  1. Evidence review

    The assessment team reads this paper, module by module, including the limitations above.

  2. Scoped trial

    A live demonstration on district-selected essays, then a blind agreement test on essays the district’s own raters have already scored.

  3. Implementation decision

    Module fit is reviewed alongside operational requirements: integration, FERPA review, audit logging, arbitration workflow, and drift monitoring.

Districts have good reason to scrutinize any scoring instrument that reaches student writing. The approach here is to earn that scrutiny: publish the design, validate against the state’s own gold standard, disclose limitations before others find them, and keep the teacher in final authority over every score.

References

  1. Alsalem, M. S. (2024). EFL teachers’ perceptions of the use of an AI grading tool (CoGrader) in English writing assessment at Saudi universities: An activity theory perspective. Cogent Education, 11(1), Article 2430865. https://doi.org/10.1080/2331186X.2024.2430865
  2. Cambium Assessment. (n.d.-a). Communicating to the public about machine scoring [White paper]. https://files.portal.cambiumast.com/corporate-site/documents/CAI-Cambium-CommunicatingPublicMachineScoring-WhitePaper.pdf
  3. Cambium Assessment. (n.d.-b). Smarter Balanced interim assessments: Automated scoring FAQ. https://smarterbalanced.alohahsap.org/content/contentresources/en/Smarter_Interim_Automated_Scoring_FAQ.pdf
  4. Chen, Z., Wang, J., Li, Y., Li, H., Shi, C., Zhang, R., & Qu, H. (2025). CoGrader: Transforming instructors’ assessment of project reports through collaborative LLM integration [Preprint]. arXiv. https://doi.org/10.48550/arXiv.2507.20655
  5. National Center for Education Statistics. (n.d.-a). NAEP technical documentation: Constructed-response interrater reliability. https://nces.ed.gov/nationsreportcard/tdw/analysis/initial_itemscore.aspx
  6. National Center for Education Statistics. (n.d.-b). NAEP technical documentation: Scoring. https://nces.ed.gov/nationsreportcard/tdw/scoring/
  7. Texas Education Agency. (n.d.). STAAR redesign. https://tea.texas.gov/student-assessment/assessment-initiatives/staar-redesign
  8. Texas Education Agency. (2022, October 20). STAAR redesign constructed response resources and updated calendar of events [To the Administrator Addressed letter]. https://tea.texas.gov/about-tea/news-and-multimedia/correspondence/taa-letters/staar-redesign-constructed-response-resources-and-updated-calendar-of-events
  9. Texas Education Agency & Cambium Assessment. (2023). STAAR hybrid scoring study, methods and results: Spring 2023 items. https://tea.texas.gov/data-reports/reports-and-studies/2023-staar-hybrid-scoring-study-2.pdf
  10. Texas Tribune. (2024, April 9). How Texas will use AI to grade this year’s STAAR tests. https://www.texastribune.org/2024/04/09/staar-artificial-intelligence-computer-grading-texas/

Cite this paper

CoGrader Research Team. (2026). Within one point of TEA raters in 98.5% of cases: The design, validity, and limits of CoGrader’s automated STAAR constructed response scoring [Whitepaper]. CoGrader.

The research reported here was supported by the Institute of Education Sciences, U.S. Department of Education, through Grant R305J250071 to Modern Learner Media, LLC. The opinions expressed are those of the authors and do not represent views of the Institute or the U.S. Department of Education.

Test it on your own raters’ essays

A 30-minute discovery call to scope a blind agreement test with your assessment team.

Book a discovery call
Trusted by educators and built for compliance

Trusted by 100,000+ teachers and educators

Used at 16,000+ schools
Backed by UC Berkeley
  • SOC 2 Type II certified
  • FERPA Compliant
  • COPPA Compliant
Research funded by a U.S. Department of Education IES award

Frequently Asked Questions

About the STAAR scoring validation study.

Is this whitepaper peer reviewed?

No. It is a technical report by the CoGrader research team.

What does “within one point” mean, and why is it the benchmark?

It is TEA’s own standard for human scoring of extended constructed responses. Two trained scorers read each essay, and scores one point apart on a rubric trait both stand. Scores further apart go to a third read. CoGrader met that standard in 98.5% of the 600 responses.

Were the test essays real student writing?

They were real responses that TEA scored and published in its scoring guides and rater training material. They are the clearest examples of each score point, so the results are a ceiling relative to everyday classroom writing.

Does CoGrader replace the teacher’s score?

No. CoGrader drafts a score and rubric-aligned feedback in about three seconds. The teacher reviews it, adjusts where needed, and decides what the student sees.

How can my district check these numbers?

Run a blind agreement test: send essays your own raters have already scored, and compare CoGrader’s scores with theirs before any agreement is signed.