FacebookBlueskyLinkedInShare

Can Universal Math Screeners Identify Dyscalculia Risk?

Can Universal Math Screeners Identify Dyscalculia Risk?

By Sarah Quesen, Senior Director of Assessment, WestEd 

In a classroom of 25 students, around one is likely to have dyscalculia, a learning disorder that makes it hard to understand and work with numbers. Students with dyscalculia tend to move through the primary grades without anyone flagging the problem. That was true for me in elementary school, where I was told I was “bad at math” and “not a math person.” I was well into adulthood before I realized it was dyscalculia. 

While more than 40 states now mandate or recommend universal screening in the early grades for dyslexia, there is no comparable movement for dyscalculia. The research base also reflects this imbalance. A search of PubMed, the National Library of Medicine’s database of biomedical research, returned roughly 40 times more studies on dyslexia than on dyscalculia.  

To help close this gap, WestEd developed and validated a dyscalculia risk flag for a universal math screener, drawing on data from more than 1,200 K–3 students in Ohio, Pennsylvania, and Texas. The encouraging news is that schools already screening in mathematics may not need a new assessment to screen for dyscalculia risk. Rather than flagging the lowest scorers at a single point in time, the flag is designed to distinguish students who may have dyscalculia from those who are struggling for other reasons. 

Building a Research-Based Criterion for Dyscalculia Risk 

A universal math screener can flag dyscalculia risk in two ways. A norm-referenced risk flag works like a pediatric growth chart. It sets a percentile threshold based on a representative population and flags students who fall in the lowest percentiles for their grade. A criterion-referenced risk flag works like a vision test. It anchors the threshold to an external standard, and a student is flagged by falling short of that standard rather than by ranking below peers. But when we searched the literature, we found no agreed-upon external standard for identifying dyscalculia risk in young children. So, we constructed a standard we could defend—grounding each part of it in the research—and used it to set and evaluate cut scores for the screener. 

We built the criterion from three subtests of the Woodcock-Johnson IV, an individually administered achievement test given to every student in the study at three points throughout the year. Applied Problems measures mathematical reasoning. Calculation measures untimed computation. Math Facts Fluency measures rapid, timed retrieval of basic arithmetic facts. These map onto areas of impairment that the Diagnostic and Statistical Manual of Mental Disorders (DSM-5) lists for the condition. 

A student met the risk criterion by scoring below the 10th percentile on at least two of the three subtests because dyscalculia typically affects multiple areas of mathematics, including problem-solving, computation, and the ability to quickly recall basic math facts. Additionally, one of the low scores had to be on the Math Facts Fluency subtest because children with developmental dyscalculia have consistently been found to struggle with rapid, automatic recall of basic math facts even when they can recall them given unlimited time. A student who struggles with math overall but can recall facts quickly under timed conditions would not meet the criterion. This requirement helps distinguish dyscalculia risk from other challenges in doing math. 

How We Established Dyscalculia Screening Cut Scores

With the criterion in place, we identified the statistically optimal screener cut score for each grade from kindergarten through grade 3 and each screening window. But we did not adopt those cut scores as given because the statistically best cut score is not always the best cut score in practice. The optimal cut for the beginning of 1st grade, for example, caught 96 percent of students who were at risk, but it did so at the cost of incorrectly flagging 40 percent of students who were not at risk.

To improve accuracy, a panel of psychometricians and content experts reviewed the statistical cut scores and adjusted them under two constraints. First, every dyscalculia cut score had to fall below the threshold for the assessment’s lowest performance level. This ensured that the flag identified a smaller group of students with more serious difficulties among those who were already showing signs of needing additional instructional support. Second, every adjustment had to improve the balance between catching students who are at risk and avoiding false alarms. This combination of statistical evidence and expert judgment is a common approach to setting cut scores for educational assessments.

One screening window did not survive this process. At the beginning of kindergarten, 41.6 percent of students met the risk criterion, several times any plausible prevalence, and the share of students flagged fell by half by the middle of the year. These early scores likely reflected differences in children’s preschool experiences, familiarity with testing, and normal variation in early development rather than consistent mathematical learning difficulties, so we excluded the beginning of kindergarten from operational screening. I would encourage anyone building early screening protocols to be skeptical of school-entry risk flags. California is learning this in reading, in which first-weeks kindergarten screening flagged up to 60 percent of students. The state is now moving to delay screening of kindergartners until midyear.

Why One Low Score Is Not Enough

Dyscalculia is a persistent condition, so if the flag captures risk rather than one bad morning, it should recur. Because the screener is given three times a year, we could check to see if this assumption held. About 75 percent of students were never flagged, 14 percent were flagged once, and 10.7 percent were flagged at least twice within a year, a stable group near the prevalence of the condition. The flag behaved like a signal of something persistent, which supports interpreting it as one. That evidence is also what grounds the practical guidance: Prioritize repeatedly flagged students for diagnostic evaluation and treat a single flag as a cue for monitoring alongside targeted instruction.

Students who the screener did not flag were unlikely to meet the criterion, and missing a child is the more consequential screening error. Roughly 4 in 10 students flagged at any single screening met the full criterion when tested individually, which is expected when screening for a condition this uncommon. A flag is a prompt for expert evaluation, not a diagnosis.

Fairness Across Student Groups

We also examined whether the flag carries the same meaning for different groups of students. The relationship between screener scores and risk was consistent across racial groups. But because the groups’ overall score distributions differ, more Black students without dyscalculia risk were flagged than White students. That is one reason why follow-up evaluation has to consider instructional history and opportunity to learn. A low score on a timed math facts task can reflect what a student has not been taught rather than a learning disability.

What This Means for States and Districts

The gap between how many children have dyscalculia and how many are identified will not close on its own. An external criterion grounded in the research literature, transparent cut score setting, and evidence of persistence, applied to assessments that most districts already give, can mean the difference between early intervention and years of unaddressed struggle. It can also help bring parity with the interventions provided to students at risk for dyslexia.

How We Can Help

WestEd’s assessment team helps states, districts, and assessment developers evaluate screening systems, set defensible cut scores, and build identification policies supported by evidence. If your program already runs universal screening and you want to know what your data can support, contact WestEd. 

The study described in this post examined a dyscalculia risk indicator developed for the Amplify mCLASS Math assessment. WestEd conducted the research under contract with Amplify. View the full research paper.  

More Related to This Post