← Research

Difficulty labels under structural audit

Controlled trends, paired trajectories, and a retrospective comparison of mathematics-item tiers.

Original research figures · June 2026Assessment-artifact analysis

A difficulty label and a measured reasoning feature are not interchangeable. These original figures examine how text length, solution structure, and tier assignments relate in one evaluated mathematics bank.

232 generated mathematics items. Tier assignments are intended levels. Correlations, controlled trends, and retrospective relabelling do not establish item difficulty for a human population.

Which item features still track the intended tier after adjustment?

The controlled associations point in different directions. Text length rises with intended tier while several reasoning measures decline.

Reported association
-0.400.40.8Solution stepsSolution steps · Reported association: -0.202Distinct operationsDistinct operations · Reported association: -0.2Relation depthRelation depth · Reported association: -0.2Intermediate stepIntermediate step · Reported association: -0.2Distractor rationaleDistractor rationale · Reported association: 0.186Stem wordsStem words · Reported association: 0.604Stem charactersStem characters · Reported association: 0.672Controlled Spearman correlation

Hover or focus a point to read the value.

Archived domain/length-controlled audit. n=232 except distractor rationale (n=123). Values reproduce the printed estimates; permutation reference bands are not confidence intervals.

View the plotted values
MeasureReported association
Solution steps-0.202
Distinct operations-0.2
Relation depth-0.2
Intermediate step-0.2
Distractor rationale0.186
Stem words0.604
Stem characters0.672
Study context

Analysis of 232 generated mathematics items, controlling for domain and length. Grey bands show the source’s permutation reference intervals. These are item-feature relationships, not measured learner difficulty.

Text length and the reported reasoning measures point in different directions.

The reported controlled associations are positive for text-length features and negative for several solution-structure measures. The analysis of this 232-item bank therefore challenges an interpretation in which the tier ordering simply tracks increasing numbers of reasoning steps.

This is a diagnostic of the evaluated bank and the selected feature set. Concept rarity, abstraction, and human familiarity are not exhausted by step counts. The result motivates examining the tier labels, rather than declaring that longer items must be easier or harder for students.

What is increasing when the tier label increases?

The same four tier groups show opposing trends. Switch between solution steps and text length; each keeps its own measurement scale.

Solution-path length

Solution steps
00.6251.251.8752.5FoundationCoreAdvancedExpertIntended tier Foundation · Solution steps: 2Intended tier Core · Solution steps: 1.73Intended tier Advanced · Solution steps: 1.38Intended tier Expert · Solution steps: 1.29Mean solution stepsIntended tier

Hover or focus a point to read the value.

Printed tier means from the 232-item audit. The endpoints belong to different item groups; they are not changes within a learner.

View the plotted values
Intended tierSolution steps
Foundation2
Core1.73
Advanced1.38
Expert1.29
Study context

The original paired curves compare reported solution steps and stem length across intended tiers. This is a descriptive audit of the evaluated bank, not an official difficulty calibration.

solution steps
The decline concerns the recorded step measure. It does not account for every possible source of mathematical difficulty.
stem length
The increase concerns the number of words. It is a surface measurement, not a direct observation of cognitive demand.

The text gets longer while the measured solution path gets shorter.

The reported mean solution-step count falls from roughly two to roughly 1.3 while stem length rises across the tier sequence. Reading both together exposes a divergence that a single “difficulty” label would hide.

The result is not evidence that every higher-tier question is cognitively easier. It establishes that, in this bank, the two measured properties do not support the same explanation of the ordering.

What does a retrospective change of labels actually demonstrate?

Each row places two reported correlations on the same axis. The thin connector shows how the association differs under the two groupings; it does not represent change in student performance.

Original groupingRetrospective grouping
-0.4-0.200.20.40.6Solution stepsSolution steps · Original grouping: -0.23Solution steps · Retrospective grouping: 0.52Relation depthRelation depth · Original grouping: -0.32Relation depth · Retrospective grouping: 0.37Distinct operationsDistinct operations · Original grouping: -0.04Distinct operations · Retrospective grouping: 0.56Intermediate stepIntermediate step · Original grouping: -0.32Intermediate step · Retrospective grouping: 0.37Distractor rationaleDistractor rationale · Original grouping: 0.01Distractor rationale · Retrospective grouping: 0.46Stem lengthStem length · Original grouping: 0.57Stem length · Retrospective grouping: -0.25Spearman correlation

Hover or focus a point to read the value.

Reported correlations for 232 evaluated mathematics items. Retrospective grouping is not independent validation of difficulty for learners.

View the plotted values
MeasureOriginal groupingRetrospective grouping
Solution steps-0.230.52
Relation depth-0.320.37
Distinct operations-0.040.56
Intermediate step-0.320.37
Distractor rationale0.010.46
Stem length0.57-0.25
Study context

The original scatter and reclassification comparison show the relation between a measured depth score and assigned tiers. This retrospective result does not establish improved student outcomes.

original grouping
Look for overlap among the coloured tiers rather than assuming that horizontal length separates them.
regrouping comparison
The revised bars show coherence with the selected score. They do not show a change in student performance.

A different grouping makes the chosen measure more orderly.

The original labels overlap across the scatter. The revised grouping produces a clearer ordering in the measured score. That contrast is informative about the alignment between a classification and its criterion.

Because the revised labels use the depth measure, a cleaner depth ordering is not independent validation. A separate response-based evaluation would be needed to establish whether the relabelled groups are better calibrated for learners.

Study record

Source record: extended construct-difficulty study, pp. 13 and 15. June 2026.

Interactive charts use verified aggregate values from the reports and archived numerical outputs. Chart notes distinguish these sources; resampling estimates can differ between runs. Where the underlying numerical series is unavailable, the text states the result without reconstructing unverified data. Study context and interpretation notes accompany each analysis. Detailed production procedures and the full internal documents are not included.