The geometry of observed item features
Spectral structure, reconstruction, stability, and category relationships—read across multiple views.
The same feature representation can yield different answers to different geometric questions. These original figures bring the spectrum, held-out reconstruction, resampling stability, and category structure into one reading sequence.
The DSAT reading-and-writing analyses use 190 generated items. The TOEFL projection uses a separate six-test corpus of 386 multiple-choice items. These are descriptive analyses of item features, not empirical models of human ability.
How many dimensions are present—and by which criterion?
The observed spectrum falls below its reference after the third component. Read that boundary first, then compare it with the reconstruction and stability studies below.
Hover or focus a point to read the value.
190 generated items. First eight stored components of the text-feature analysis. Three exceed the parallel-analysis reference; this criterion is distinct from variance coverage and stability.
View the plotted values
| Component | Observed | Parallel reference · 95th percentile |
|---|---|---|
| 1 | 3.496 | 1.524 |
| 2 | 1.538 | 1.371 |
| 3 | 1.394 | 1.263 |
| 4 | 1.111 | 1.174 |
| 5 | 0.844 | 1.104 |
| 6 | 0.831 | 1.039 |
| 7 | 0.715 | 0.974 |
| 8 | 0.498 | 0.919 |
Study context
Analysis of 190 generated reading and writing items. Parallel analysis and cumulative variance describe the feature representation; component retention is not the same as bootstrap stability.
- spectrum against references
- Compare the observed eigenvalues with the parallel-analysis and broken-stick curves. The reference matters as much as the shape of the scree plot.
- cumulative variance
- Read the added variance as components accumulate. Reaching a variance threshold does not establish that each included direction is reproducible.
A dominant component does not settle the dimensionality question.
The report identifies a dominant first eigenvalue, retains three components under parallel analysis, and reaches 90% cumulative variance at eight components. Retention above a reference and coverage of total variance are not equivalent criteria.
The next step is to ask whether these directions recur when the sample changes and whether they reconstruct held-out observations. Without those checks, a convenient projection can be mistaken for stable structure.
Does retaining more components improve held-out reconstruction?
The smallest model is not the zero-rank baseline: one retained direction reduces held-out error. More directions do not improve it in this archived run.
Hover or focus a point to read the value.
Archived 190-item analysis. Rank 1 minimizes the recorded error. Points reproduce the stored results; connecting segments do not imply measurements between tested ranks.
View the plotted values
| Retained rank | Held-out cell MSE |
|---|---|
| 0 | 0.917 |
| 1 | 0.748 |
| 2 | 0.91 |
| 3 | 1.328 |
| 4 | 1.491 |
| 5 | 1.715 |
| 6 | 1.986 |
| 7 | 1.795 |
| 8 | 1.624 |
| 9 | 1.167 |
| 10 | 0.763 |
Study context
The source compares reconstruction error across candidate ranks for realized-text features. The minimum is specific to this representation and evaluation procedure.
The first retained dimension performs best under this criterion.
Reported held-out reconstruction error is lowest at rank one and is higher over much of the subsequent range. Including more dimensions therefore does not monotonically improve out-of-sample reconstruction in the reported test.
This qualifies what cumulative explained variance can establish. More retained variance in the fitted sample can coexist with poorer held-out reconstruction. The result remains conditional on the features, scaling, sample, and validation procedure.
Which directions survive a change of sample?
Six components are shown separately. Compare each mean with its lower tail: a strong average can hide a direction that changes substantially across samples.
Hover or focus a point to read the value.
190 generated items; six tested components. The connector spans the mean and lower-tail alignment, not a confidence interval. Values use the archived numerical run and may differ from rounded report summaries.
View the plotted values
| Measure | Bootstrap mean | 5th percentile |
|---|---|---|
| Component 1 | 0.992 | 0.983 |
| Component 2 | 0.778 | 0.181 |
| Component 3 | 0.737 | 0.146 |
| Component 4 | 0.834 | 0.596 |
| Component 5 | 0.57 | 0.067 |
| Component 6 | 0.544 | 0.053 |
Study context
Bootstrap stability for realized-text principal components, n=190 items. The report identifies one stable leading axis; variance explained alone would support a different selection criterion.
The leading axis is markedly more reproducible than the later axes.
The first component has high average and lower-tail alignment. Later components have weaker or less consistent support. The source reports one bootstrap-stable leading axis even though another selection criterion retains more components.
This is why spectral size, reconstruction, and stability are presented separately. They are tests of different properties. Agreement among them would strengthen an interpretation; disagreement tells us which claim needs to be narrowed.
Do the category labels correspond to compact regions of feature space?
Three category systems are compared with their own shuffled reference. Read the position relative to zero as well as the gap to the null.
Hover or focus a point to read the value.
190 items. Observed scores exceed the shuffled reference but remain negative. Distinguishability from a null does not imply compact categories.
View the plotted values
| Measure | Observed | Shuffled mean | Shuffled 95th percentile |
|---|---|---|---|
| Question type | -0.044 | -0.133 | -0.116 |
| Skill | -0.022 | -0.098 | -0.074 |
| Domain | -0.015 | -0.039 | -0.022 |
Study context
The original silhouette comparison tests whether labelled categories form cohesive clusters in feature space. The observed and reference values must be read together.
The designed labels do not form clean, separated clusters in this view.
The observed values are small or negative, while some alternative or shuffled groupings have higher cohesion. The figure therefore resists reading the category names as naturally separated geometric populations.
The analysis concerns the selected item representation. It neither invalidates the educational categories in every setting nor establishes that they are interchangeable. It shows that the representation does not recover them as compact clusters.
Which skill groups are close in the measured representation?
The recovered record contains the nine minimum-spanning-tree connections among ten skill centroids. Inspect the topology and the exact distances independently.
Hover or focus a mark to read its value.
Archived minimum-spanning-tree edges. Positions are arranged for readability; visual edge length is not distance. Exact distances are listed below. Thicker lines indicate shorter distances.
Read the connection distances +
| Connection | Distance |
|---|---|
| BOU — FSS | 1.926 |
| CID — CTC | 1.477 |
| CID — INF | 1.253 |
| CID — TSP | 1.413 |
| COE — CTC | 1.404 |
| CTC — SYN | 2.82 |
| FSS — TRA | 1.73 |
| TRA — TSP | 2.788 |
| TRA — WIC | 1.652 |
Study context
A Ward dendrogram and a minimum-spanning-tree view of realized linguistic features. Distances concern the selected item-feature representation, not distances between learners.
- hierarchical linkage
- Read the joining height as distance under the stated clustering procedure. The drawing order of the leaves is not an achievement ranking.
- minimal connections
- The edges summarize proximity among centroids. They do not establish transitions, prerequisites, or causal links between skills.
Local relationships can be inspected without claiming clean global clusters.
The recovered tree links CID and INF at distance 1.253, while the CTC–SYN edge has distance 2.820. Use the listed distances for proximity; node positions here are arranged for readability, not as a geometric projection.
These views complement the weak category-cohesion result. A representation can show local relationships even when the labelled groups overlap. The tree is a descriptive device, not a curriculum or a causal model.
What separates the observed TOEFL item groups in a low-dimensional view?
Each point is one assessment item. Switch between the spatial view and two planar projections to examine how separation and overlap depend on the view.
386 items. Coordinates recomputed from the six source files using the recovered analysis. The item count and rounded explained variance match the archived run: 43%, 22%, 13%. These axes describe item features, not learner ability.
Study context
A surface-feature projection from six full-test evaluations: 386 multiple-choice items. Points describe assessment artifacts, not learner states or learning trajectories.
The sections occupy different regions, with overlap still present.
The cloud shows structured separation rather than a single undifferentiated mass. That pattern is consistent with differences in task form and recorded item properties across the sampled sections.
The axes are not learner ability or growth. This corpus contains 386 multiple-choice items from six full-test evaluations. The projection is a view of those artifacts and should not be read as a psychological map of students.
How cohesive is the chosen partition?
Vary the number of clusters and examine cohesion before giving the partition a substantive interpretation.
Hover or focus a point to read the value.
Archived 190-item analysis. The maximum is at two clusters, but the silhouette remains modest. Selecting a maximum does not establish a natural taxonomy.
View the plotted values
| Number of clusters | Text-feature clustering |
|---|---|
| 2 | 0.192 |
| 3 | 0.157 |
| 4 | 0.161 |
| 5 | 0.169 |
| 6 | 0.153 |
| 7 | 0.155 |
| 8 | 0.136 |
| 9 | 0.135 |
| 10 | 0.139 |
| 11 | 0.141 |
| 12 | 0.134 |
| 13 | 0.129 |
| 14 | 0.129 |
| 15 | 0.127 |
| 16 | 0.137 |
| 17 | 0.131 |
| 18 | 0.122 |
A maximum is a comparison, not a discovery of natural categories.
The archived silhouette curve peaks at two clusters. Larger partitions score lower, although not monotonically.
The selected count depends on this representation, distance and criterion.
Do different methods recover the same categories?
Two clustering methods can return different relationships with the same labels. The comparison tests that dependence directly.
Hover or focus a point to read the value.
Archived text-feature analysis. The two methods use the selected cluster count. Agreement with categories is limited and varies by method; the axes are not student accuracy.
View the plotted values
| Measure | K-means | Ward |
|---|---|---|
| Question type | 0.075 | 0.059 |
| Skill | 0.108 | 0.119 |
| Domain | 0.112 | 0.201 |
Method choice changes the correspondence with labels.
The archived adjusted Rand indices remain modest for both methods. Neither reproduces the category systems perfectly.
These comparisons describe item-feature organization, not learner proficiency.
Study record
Source records: latent-structure follow-up, pp. 3–6; TOEFL structural analysis, p. 3. June 2026.
Interactive charts use verified aggregate values from the reports and archived numerical outputs. Chart notes distinguish these sources; resampling estimates can differ between runs. Where the underlying numerical series is unavailable, the text states the result without reconstructing unverified data. Study context and interpretation notes accompany each analysis. Detailed production procedures and the full internal documents are not included.