The problem CoGrader is solving

There’s a kind of math that every ELA teacher can do. It involves their grading load: at four minutes per essay, 150 students, it takes ten hours to grade one assignment. And if you are a teacher, you’ve lived these hours on nights and weekends.

This leads to a cycle of neverending catch-up. Essays get back to students weeks after the unit ended. When students skim for the grade, the comments are understandably ignored. It’s been too long for them to be effective.

Teachers work hard to provide feedback, but the constraint is time, and the cost shows up in national data.

On the most recent NAEP writing assessment, 24% of students in grades 8 and 12 performed at the Proficient level, and 3% at Advanced (NCES, 2012).

CoGrader exists to remove the time constraint while empowering the teacher. CoGrader makes it possible for teachers to implement the research-based practices that were previously impossible to actually do.

Finding 1: Feedback is the lever, when it is specific

John Hattie’s 2009 synthesis of more than 800 meta-analyses ranks feedback among the strongest influences on student achievement, with an average effect size of 0.73. For context, an effect size of over .4 translates to more than a year’s growth in a year’s time. His own summary was that the simplest prescription for improving education is “dollops of feedback.” A 2020 meta-analysis by Wisniewski, Zierer, and Hattie found a more modest, but nonetheless notable, effect size of 0.48. They also found a sharper picture of what maximizes the benefits of feedback. Effective Feedback tells a student what they did well, what is missing, and what to do next. Feedback that accomplishes those three goals is what’s required to move the needle. Black and Wiliam’s 1998 review of formative assessment reached the same conclusion from the classroom side, with some of the largest gains reported for any educational intervention.

Bloom’s 1984 “two sigma” study shows the ceiling. Students tutored one to one, with frequent feedback and correction, outperformed 98% of peers who did not receive the intervention. While conventional tutoring, not software, produced that result, that tutoring was enabled by frequent, individual feedback from the students to the tutor, and from the tutor to the students.

CoGrader applies the research as a teaching assistant. It drafts feedback that’s always tied to a criterion on the teacher’s rubric. The feedback always details what the student did, where the essay falls short of the criterion, and one concrete next step the student can take for immediate improvement. This structure mirrors the criteria for effective feedback, derived from Hattie and others’ work. The teacher stays in charge and can revise the feedback before a student receives it.

Finding 2: Timing decides whether feedback gets used

Feedback is best served fresh. It’s more effective when it’s immediate. Kulik and Kulik’s 1988 review of feedback timing found that in classroom settings, immediate feedback generally helps more than delayed feedback. Teachers see the same thing without a meta-analysis: comments returned three weeks after the essay land in the wastepaper basket rather than in revisions. A student may have had 40 class periods since then. They’ve moved on.

Most classrooms follow this pattern. The teacher assigns writing. Teachers spend two to four weeks grading, often during personal time, and the full cycle lasts three to five weeks — the cycle CoGrader is designed to support.

Here is the loop CoGrader is built for. The teacher assigns the essay. Students submit. Within minutes, CoGrader drafts scores and comments against the teacher’s rubric. Then, the teacher reviews, adjusts, and exports the feedback. Students get feedback within a day or two. The teacher has time to plan lessons aligned to individual student’s and small groups of students’ actual needs. After direct, targeted instruction, the next writing task gives students an opportunity to incorporate the learning while the lesson is still fresh. Loop length: one to two days for feedback and another few days to give your students exactly the instruction they require to move their writing to the next level.

An ancillary benefit to this accelerated feedback cycle is that students write more frequently as well. The loop takes one week as opposed to four. That’s the difference between four essays a year and twelve, without even factoring in shorter writing assignments.

Finding 3: Clear Rubrics constrain bias and drift

Holistic grading faces human limitations. Teachers are human. They get tired, become familiar with a class, and inject unintentional bias into their grading. David Quinn’s 2020 experiment demonstrates this with a harrowing finding: teachers rated an identical writing sample lower when it carried signals of a Black author than a White author when using a vague grade-level scale. That changed when the same teachers used a clearly defined rubric. That gap closed. Rubrics mitigate bias. That is why state assessments, STAAR included, score writing against published rubrics. That’s also why Texas has each STAAR essay scored by two trained scorers and adds their scores. This is also why it’s good practice for a PLC to write and use a common rubric to clearly define success for students and teachers alike.

What CoGrader does with it. CoGrader’s rubric agent helps teachers write high quality rubrics and its grading engine maintains consistency applying them. CoGrader grades only against that rubric, whether it be the teacher’s own, or a state rubric such as STAAR’s. In June 2026 we tested CoGrader’s STAAR grading against 600 responses from the Texas Education Agency’s public scoring guides. On 283 essays, CoGrader landed within one point of the official score 97% of the time and matched it exactly 72% of the time. That’s better than STAAR expects human graders to do. Short responses matched exactly 81% of the time for reading and 90% for writing. Exemplar responses are the clearest examples of each score point, so treat the numbers as a ceiling for live classroom writing. The full method is in CoGrader’s STAAR validation.

Finding 4: Volume compounds

Students get better at writing by writing, often, with a response. Graham and Perin’s 2007 meta-analysis, Writing Next, identified 11 instructional practices that measurably affected adolescent writing, such as explicit strategy instruction, specific product goals, and studying models. All of the effective strategies required frequent opportunities for meaningful writing. The efficacy multiplies when students get more practice. However, Applebee and Langer’s 2011 national study found that most of the writing students do in middle and high school is a paragraph or less. They found that extended writing with feedback is rare. The limiting factor is the grading hours each extended assignment demands.

What CoGrader does with it. When feedback on 150 essays takes an evening of review instead of a week of scoring, students write more. In Texas, that means no more year-late post mortems. Students get more STAAR constructed-response practice at the volume students need to improve before April, scored against the state rubric during the school year, when responsive instruction is possible. Teachers can correct misunderstandings and challenge students more often and more frequently, before they sit for STAAR.

Finding 5: Grading load drives burnout and turnover

Hours spent working outside of work is a major driver of teacher turnover. Ingersoll’s 2001 analysis found that school conditions, not retirement or enrollment shifts, account for much of it. RAND’s 2023 State of the American Teacher survey put the average teacher workweek at 53 hours, about 7 hours more than comparable working adults. Additionally, Gallup found in 2022 that 44% of K-12 workers always or very often feel burned out at work, the highest rate of any industry it measured. Trying to keep up with the amount of writing and feedback students need contributes to those hours.

What CoGrader does with it. As a teaching assistant, CoGrader takes the first pass at feedback: reading against the rubric, drafting the comment, and suggesting a score. Teachers spend less time writing repetitive comments, and more time addressing students’ needs. The teacher stays the teacher of record: reviewing, editing, and approving every grade before students see it. The IES-funded CoGrader 2.0 project sets the target at cutting time spent grading and drafting feedback by up to 80%.

CoGrader’s logic model, in one paragraph

Grading a full stack of student essays takes hours a teacher rarely has. CoGrader works as a teaching assistant for that load. By combining a teacher’s custom rubric with student papers, it drafts scores and feedback notes in seconds. The teaching assistant drafts. The teacher reviews, adjusts, and approves every grade before it goes out. Students get feedback while the assignment is still fresh, which makes them more likely to revise and write again soon. In the short term, we expect teachers to spend fewer evenings and weekends grading, scoring to stay more consistent across a class, and students to write more often. Over time, the model predicts that more practice with timely feedback moves students toward grade-level writing standards and lowers the workload that drives teacher burnout. Principals get clearer data on writing progress along the way. These are the outcomes the model predicts; the next section says which ones the evidence supports today.

Where CoGrader’s evidence stands

We would rather state the evidence exactly than round it up.

Federal research funding. The Institute of Education Sciences, the U.S. Department of Education’s research arm, funds CoGrader’s research through grant R305J250071, "CoGrader 2.0: Accelerating Student Writing Proficiency Through AI-Assisted Personalized Feedback at Scale" ($460,190, September 2025 to September 2026). Phase 1 funds design-based research with teachers, principals, and students, and a proof-of-concept report. It is not yet an effects study.

ESSA. CoGrader meets ESSA’s Tier IV definition of evidence-based, Demonstrates a Rationale: a research-grounded logic model plus ongoing research. The tier is self-assessed. Tiers III to I require studies of measured effects.

STAAR validation. The numbers above, with methods and limitations published.

What CoGrader has not shown yet: a measured effect on student writing outcomes. The full picture, with every source linked, is on CoGrader’s research page.

Six questions to ask about any AI grader’s research

These are the questions we want asked of CoGrader. Ask them of everyone.

QuestionWhy it mattersCoGrader’s Answer
Which ESSA evidence tier, and where is the logic model published?A vendor who cannot name a tier usually does not have one.Tier IV, self-assessed, with the logic model published on the research page.
Is there federal research funding?IES and NSF awards come with research protocols and public records.IES grant R305J250071, listed on ies.ed.gov.
Does it grade against the teacher’s rubric?Grading without a rubric drifts. Canned rubrics cannot match state frameworks.The teacher’s own rubric, or a state rubric such as STAAR.
Does a teacher approve grades before students see them?CoGrader is a teaching assistant, not the teacher of record.Yes, every grade.
Are agreement numbers published, with the test set and its limits?A number without a method is marketing.STAAR validation: 600 TEA-scored responses, per-module results, held-out test set, limits stated. See CoGrader’s STAAR validation.
What is the data posture?Student writing is sensitive data.SOC 2 Type 1. Student work is not used to train AI models. See the Privacy Policy and AI Transparency Note.

See the research on your own essays

CoGrader is free to try. Run your next class set through it, compare the scores and comments with your own, and decide from there. No credit card required. Districts that need documentation for procurement can request the evidence package at research@cograder.com.

Try CoGrader free

References

  • Applebee, A. N., & Langer, J. A. (2011). A snapshot of writing instruction in middle schools and high schools. English Journal, 100(6), 14-27.
  • Black, P., & Wiliam, D. (1998). Assessment and classroom learning. Assessment in Education: Principles, Policy & Practice, 5(1), 7-74.
  • Bloom, B. S. (1984). The 2 sigma problem: The search for methods of group instruction as effective as one-to-one tutoring. Educational Researcher, 13(6), 4-16.
  • Gallup (2022, June). K-12 workers have highest burnout rate in U.S.
  • Graham, S., & Perin, D. (2007). Writing Next: Effective strategies to improve writing of adolescents in middle and high schools. Alliance for Excellent Education.
  • Hattie, J. (2009). Visible Learning: A synthesis of over 800 meta-analyses relating to achievement. Routledge.
  • Hattie, J., & Timperley, H. (2007). The power of feedback. Review of Educational Research, 77(1), 81-112.
  • Ingersoll, R. M. (2001). Teacher turnover and teacher shortages: An organizational analysis. American Educational Research Journal, 38(3), 499-534.
  • Institute of Education Sciences (2025). Award R305J250071: CoGrader 2.0. https://ies.ed.gov/use-work/awards/cograder-2-0-accelerating-student-writing-proficiency-through-ai-assisted-personalized-feedback
  • Kulik, J. A., & Kulik, C.-L. C. (1988). Timing of feedback and verbal learning. Review of Educational Research, 58(1), 79-97.
  • National Center for Education Statistics (2012). The Nation’s Report Card: Writing 2011 (NCES 2012-470).
  • Quinn, D. M. (2020). Experimental evidence on teachers’ racial bias in student evaluation: The role of grading scales. Educational Evaluation and Policy Analysis, 42(3), 375-392.
  • RAND Corporation (2023). Doan, S., et al. Teacher well-being and intentions to leave: Findings from the 2023 State of the American Teacher Survey.
  • Texas Education Agency. STAAR constructed-response scoring guides. https://tea.texas.gov/student-assessment/staar
  • Wisniewski, B., Zierer, K., & Hattie, J. (2020). The power of feedback revisited: A meta-analysis of educational feedback research. Frontiers in Psychology, 10, 3087.

Share this post