← Research

Annotated complexity and evidential necessity

Testing whether the evidence cited by an item is necessary to sustain its answer.

Evaluation study · June 2026Generated assessment records · model-based evaluation

The measurement

The evaluation compared an item’s annotated evidence depth with the mean number of cited sentences found individually necessary under a sentence-removal assessment. The two measurements describe different properties.

How much of the cited evidence is individually necessary?

The graph compares the recorded depth groups with their reported mean evidence dependence. Follow the purple series against the grey diagonal; focus or hover on a point to read its value.

AnnotatedIndividually necessary · mean
0246456Annotated depth 4 · Annotated: 4Annotated depth 5 · Annotated: 5Annotated depth 6 · Annotated: 6Annotated depth 4 · Individually necessary · mean: 2Annotated depth 5 · Individually necessary · mean: 2.3Annotated depth 6 · Individually necessary · mean: 2.4Evidence sentencesAnnotated depth

Hover or focus a point to read the value.

Reported group means from 16 model-evaluated items. Individual scatter and uncertainty are not reconstructed from the image.

View the plotted values
Annotated depthAnnotatedIndividually necessary · mean
442
552.3
662.4
Study context

Sixteen model-evaluated reading items. The scatter preserves the individual observations and reported mean trend. The source’s ceiling language describes this tested setting, not a universal cognitive limit.

The average dependency rises much less than the annotation.

The reported group means are 2.0, 2.3, and 2.4 for annotations of four, five, and six. In the supplementary original figure, individual points spread widely, including zero-valued cases. The relationship is therefore neither one-to-one nor adequately described by the annotation alone.

This changes the interpretation of the preceding calibration curve: an item can carry a larger annotation without a proportionate increase in this operational measure of necessity. Redundancy and the effects of deletion remain possible explanations; the figure does not establish a universal ceiling on cognition.

Annotated depthMean individually necessary sentences
42.0
52.3
62.4

What changed—and what did not

The annotated count increased from four to six. The reported mean count of individually necessary evidence increased from 2.0 to 2.4. A longer evidence annotation did not correspond to a proportionate increase in this measure of dependency.

DiagnosticReported resultInterpretation
Evaluated items16A small, model-evaluated sample.
Mean necessary / annotated ratio0.46Reported mean of item-level ratios; not computed from the three chart means.
Zero individually necessary cited sentences2 of 16Requires inspection; redundancy and the removal procedure affect this measure.

Read the related experiments

The annotation-response result and the response-format comparison provide different views of the same evaluation problem. They are kept together rather than reduced to a single summary bar.

Does a larger depth annotation establish a deeper task?

Agreement and structural checks are separate tests. Neither count measures how much evidence a response actually needs.

Reported count
0102030Key agreementKey agreement · Reported count: 30Structural checksStructural checks · Reported count: 28Items passing / 30

Hover or focus a point to read the value.

Thirty generated items across five annotation levels. These counts concern evaluator agreement and structural checks, not difficulty for human learners.

View the plotted values
MeasureReported count
Key agreement30
Structural checks28
Study context

Thirty generated reading items across five annotation levels. The source’s agreement curve concerns model-based evaluation; it does not establish human difficulty. Read alongside the evidence-removal results below.

annotation response
Horizontal position is the specified depth; vertical position is its recorded realization. Agreement with the diagonal establishes correspondence between these two recorded quantities.
evaluation at each level
The two lines track different criteria. Agreement with the key must not be substituted for structural validity or human difficulty.

Annotation follows the target; necessity remains untested here.

The left curve follows the diagonal across the five levels. On the right, evaluator agreement stays at 100%, while structural validity is lower at two levels. These measurements show annotation compliance and model-based agreement, not how much evidence the answer actually requires.

The next experiment must change the evidence itself. Without that intervention, an exact annotation curve can look like a difficulty result while measuring a property of the record.

Does changing the response format alter evidence dependence?

The purple line follows the reported constructed-response means. The grey reference shows the approximate multiple-choice comparison. The decline across annotation levels is visible without the rest of the original plate.

Constructed responseMultiple-choice reference
01234456Annotated depth 4 · Constructed response: 3.7Annotated depth 5 · Constructed response: 2.7Annotated depth 6 · Constructed response: 2.3Annotated depth 4 · Multiple-choice reference: 2.2Annotated depth 5 · Multiple-choice reference: 2.2Annotated depth 6 · Multiple-choice reference: 2.2Mean necessary sentencesAnnotated depth

Hover or focus a point to read the value.

Nine constructed-response items. Values reproduce the rounded labels in the original figure; the multiple-choice reference is approximate.

View the plotted values
Annotated depthConstructed responseMultiple-choice reference
43.72.2
52.72.2
62.32.2
Study context

Exploratory constructed-response evaluation, nine items. The original comparison is retained; the “doubles” label is not an overall average or a population-wide result.

The observed gain does not continue as the annotation increases.

The first constructed-response group has a higher mean, approximately 3.7. The later groups fall to approximately 2.7 and 2.3. The nine-item experiment therefore contains both a promising contrast and a failure of monotonic growth.

The source title’s “doubles” wording refers to a selected comparison, not the overall average. The useful next question is whether the difference persists in a larger, matched evaluation that holds the task and adjudication conditions constant.

Interpretation boundary

A sentence-removal result measures sensitivity to a specified intervention. It is not a count of mental operations in a learner. Redundant evidence can make individual removals uninformative; removing a sentence can also disturb coherence. These aggregate results therefore motivate further examination, not a universal limit on reading or a guarantee about human difficulty.

Study record

Internal reading-item evaluation, June 2026. Source: evidence-deletion analysis, extended draft §7.6. Evaluators were language models; no human response or learning-outcome result is reported here. The public figures reproduce aggregate values only.