The International Association for the Evaluation of Educational Achievement, working with UNESCO, released the first TIMSS results measuring math achievement since the pandemic. In a vast majority of participating countries, fourth-grade boys now outperform girls, reversing a decade of gains toward gender parity. Among eighth-graders, the share of countries where boys outscore girls at the advanced level has grown sharply since 2019. Matthias Eck, a UNESCO program specialist and one of the report's authors, told EdSurge the pattern reflects pandemic-era school closures compounding existing stereotypes about who belongs in math class.
The gap isn't about ability. Eck is direct on that point: boys and girls perform equally well on the underlying skills. What moves the numbers is confidence, teacher expectation, and how early those signals reach a student, and by Eck's account, the gap is already visible by age nine or ten. That timeline means a high school math teacher inherits the confidence gap fully formed, not developing. The response has to happen earlier in a student's schooling than your classroom, but you can still watch for it: a capable ninth-grader who under-volunteers, hedges every answer, or opts out of the advanced track despite strong grades is showing the same pattern this data describes at scale.
Digital Promise announced its first grant recipients under the K-12 AI Infrastructure Program on June 29, a $26 million multi-year effort built with Learning Data Insights, DrivenData, Georgetown University's Massive Data Institute, and Catalyst @ Penn GSE. The program's first cycle targets formative assessment, the everyday work of reading a discussion, a written response, or a problem-solving attempt to gauge what a student actually understands while the learning is still happening. One of the four funded projects, led by Learning Equality's Jamie Alexandre, is building an open benchmark for AI to identify science misconceptions in students' free-form written answers. "Without the right foundation, AI will become another barrier to educational progress," Digital Promise president and CEO Jean-Claude Brizard said in the announcement. All outputs will be openly licensed, and roughly 30 grants are planned over the program's life.
Most AI tools already in your building were built to grade what a student produced, not to notice what a student misunderstands. Formative assessment is the harder problem, and it's the one that actually changes instruction mid-unit rather than after the test. An open, shared benchmark means every vendor building a classroom AI tool will eventually get measured against the same standard for reading student thinking, not just student output. That's the gap between a tool that flags a wrong answer and one that tells you why a student got there.
Barbara Treccani of the University of Trento's Department of Psychology and Cognitive Science published a study in Frontiers in Psychology testing whether retrieval practice, the well-documented advantage of self-testing over rereading, holds up outside the lab. Fifth-graders studied real history texts, then either reread the material or took a practice test with accuracy feedback repeated until they answered every question correctly. The testing group retained more. A second manipulation, lengthening the gap between first reading the material and studying it again, showed no measurable benefit. Treccani's team was explicit about why the second result matters as much as the first: it means spacing strategies that work in controlled lab settings don't automatically transfer to a real classroom, and teachers need evidence built on actual school material before adopting them.
Decades of retrieval-practice research were built on lab tasks like word lists, not the kind of dense, multi-paragraph content you actually assign. This study used real history texts with real fifth-graders, and the advantage held. That's the citation you reach for when a colleague or administrator dismisses "the testing effect" as a lab artifact that won't survive contact with an actual unit on the Industrial Revolution. The honest half of the finding matters too: don't assume every popular study-skills strategy transfers to your classroom just because it works in a psychology lab. Test it, the way Treccani did.
Meta announced July 16 that it will notify parents using Instagram's Parental Supervision tools if their teen discusses suicide or self-harm with Meta AI, the company's chatbot. A dedicated detection system flags conversations with a clear reference to self-harm, and every flagged chat is reviewed by a person before a parent is alerted; Meta says it will err toward alerting when a teen's intent is ambiguous. The feature is live now in the U.S., U.K., Australia, and Canada, with global rollout by year's end, and builds on alerts Meta already sends when a teen repeatedly searches self-harm terms on Instagram. Meta also said it will contact emergency services directly, for any user, teen or adult, when a chatbot conversation suggests someone is at immediate risk.
A student in crisis increasingly types it to a chatbot before saying it to you, and that has been true for a while. What's new is that one major platform now routes some of those conversations back to a parent, not just into a log nobody reads. That doesn't reduce your responsibility as the adult in the room five days a week, but it does mean a parent phone call to your building might now reference a conversation you never knew was happening. It's worth knowing this alert exists before a parent brings it to you first.