Observation, recovery, and the limits of simulation
A visual record of channel comparisons, recovery experiments, reliability, failed models, and scale tests.
These figures examine what different observation regimes can recover when the latent mechanism is specified in a simulation. They also retain the model that failed, so that performance and model adequacy are not conflated.
All learner-state recovery and scale results on this page are simulation-based. The three-million-learner dashboard is a synthetic stress run, not an audience size or a human validation study. Simulation recovery does not establish the truth of the simulated learner model.
Which observation channels distinguish the simulated states?
Each row separates successful detection from false alarms. The two series should be read jointly: high recall does not compensate for an unacceptably high false-alarm rate.
Hover or focus a point to read the value.
Simulation results for the stuck-state task. Higher recall and lower false-alarm rate are desirable; these are not measured student outcomes.
View the plotted values
| Measure | Recall | False-alarm rate |
|---|---|---|
| Heart-rate only | 0.24 | 0.92 |
| Gaze only | 0.92 | 0 |
| Combined channels | 0.92 | 0 |
Study context
Simulation study. The channel-localization and classification results depend on the planted mechanism. They do not demonstrate deployed gaze or physiological monitoring, or performance on actual students.
- detection and false alarms
- Read both bars. Increasing sensitivity while admitting many false alarms can be an unusable trade-off.
- distinct tasks
- Confidence, candidate identification, and pair localization are separate targets; a strong result on one does not transfer automatically to the others.
The channel is useful only if it resolves the distinction being tested.
In this simulation, the heart-rate-only condition performs poorly on the stuck-state task, while gaze-related conditions behave differently. The right panel also shows that confidence detection and two-candidate localization do not have the same performance.
These are results under a planted mechanism, not measurements of actual students or a statement that the product deploys these sensors. The figure motivates a study-design question: whether a channel contains information at the temporal and semantic grain required by the hypothesis.
How much information would a discriminating study require?
Power depends on the contrast being tested. These two curves preserve separate experimental conditions and sample-count grids.
Pair-contrast detection power
Hover or focus a point to read the value.
Archived simulation estimates for this specified contrast. The sample-size axis belongs to this experimental design, not a general product requirement.
View the plotted values
| Sample count | Monte Carlo estimate |
|---|---|
| 20 | 0.757 |
| 30 | 0.834 |
| 50 | 0.967 |
| 80 | 0.999 |
| 120 | 1 |
Second-contrast detection power
Hover or focus a point to read the value.
Archived simulation estimates for this specified contrast. The sample-size axis belongs to this experimental design, not a general product requirement.
View the plotted values
| Sample count | Monte Carlo estimate |
|---|---|
| 20 | 0.455 |
| 40 | 0.717 |
| 64 | 0.884 |
| 100 | 0.961 |
Study context
Study-design analysis combining an item-based testability check with Monte Carlo power and sample-size panels. Projected sample reductions are model-dependent, not measured recruitment savings.
- testability
- The item-based check asks whether the targeted alternatives can be distinguished at all.
- Centre: power curve
- Power increases with available observations under the simulation assumptions. It is not an empirical success curve from recruited students.
- required sample
- The curves express a model-dependent sample requirement. Compare them only for the same target and decision criterion.
Testability comes before the apparent sample-size advantage.
The analysis first distinguishes items on which the contrast can be examined from those on which it cannot, then estimates power and sample requirements under the stated model. An additional channel changes the modelled information available per participant.
A projected reduction in sample size is not an observed saving in a human study. It depends on the assumed mechanism, effect size, noise, and measurement fidelity; those assumptions must be checked before using the projection for planning.
What becomes visible when the simulated observation is richer than right or wrong?
Three separate questions: can the simulated state be recovered, does repeated observation become consistent, and does recovery depend on the strength of the planted distinction? Select one view at a time.
Recovery under two observation regimes
Hover or focus a point to read the value.
30,000 simulated learners; 3.6 million responses. Whiskers are the 2.5th–97.5th bootstrap percentiles of held-out prediction AUC. The process interval collapses to 1 in this run. These are simulated outcomes.
View the plotted values
| Measure | Simulated recovery |
|---|---|
| Accuracy only | 0.653 (0.644–0.661) |
| Process observation | 1 (1–1) |
Consistency with repeated observation
Hover or focus a point to read the value.
Same 30,000-learner simulation. Mean split-half correlation rises with observation count. The archived output retains means only, so no uncertainty band is reconstructed. Consistency does not establish validity.
View the plotted values
| Attempts | Repeated-observation agreement |
|---|---|
| 20 | 0.345 |
| 40 | 0.732 |
| 60 | 0.887 |
| 90 | 0.948 |
| 120 | 0.965 |
Recovery across simulated depth settings
Hover or focus a point to read the value.
Five tested settings in the same simulated model. Depth is a simulator parameter, not a measured learner attribute. The process series remains at 1 under these assumptions; this is not evidence of perfect real-world detection.
View the plotted values
| Simulated depth parameter | Process observation | Accuracy only |
|---|---|---|
| 1 | 1 | 0.56 |
| 1.6 | 1 | 0.6 |
| 2.2 | 1 | 0.641 |
| 2.8 | 1 | 0.687 |
| 3.4 | 1 | 0.719 |
Study context
Simulation study with a planted latent distinction. The original four-panel figure retains recovery, projected states, repeated-attempt reliability, and exploration distributions. It is not a learner-outcome result.
- recovery
- Compare the evaluated AUC for the two information regimes. The target state is known because it was simulated.
- projected states
- The projection displays the planted groups under the process representation. Separation is not independent confirmation of the simulation’s realism.
- reliability
- Repeated-attempt agreement rises with more observations. Reliability and validity remain different properties.
- exploration
- The distributions compare the number of simulated items needed to reveal the planted distinction under the two policies.
The simulation separates states that overlap in the accuracy-only view.
The archived recovery estimate is 0.653 for accuracy alone and 1.000 for process observation. Mean split-half agreement rises from 0.345 at 20 attempts to 0.965 at 120. Across five tested depth settings, accuracy-only recovery varies while the process result remains at 1. These are separate properties of the same simulated model.
Recovery is conditional on the simulation that placed the distinction in the data. This demonstrates behaviour under those assumptions, not that the same latent state has been identified in a real learner population.
Does a more detailed state model actually fit the simulated process?
Inspect model resolution and lag separately. These diagnostics retain the failure and follow-up, rather than selecting only the most favourable result.
Does finer discretization remove the ramp?
Hover or focus a point to read the value.
Rounded archived output. A ratio near 1 would indicate a plateau. Both remain elevated: this test did not isolate drift as the cause.
View the plotted values
| Number of states | Stationary | Drifting |
|---|---|---|
| 6 | 7.12 | 5.97 |
| 20 | 6.68 | 5.37 |
| 60 | 7.55 | 5.04 |
| 120 | 7.7 | 4.58 |
Implied timescale at longer lags
Hover or focus a point to read the value.
Rounded archived follow-up. The drifting curve approaches a plateau while the stationary curve continues rising. This is a model diagnostic, not proof of a learner state.
View the plotted values
| Lag | Stationary | Drifting |
|---|---|---|
| 1 | 12.4 | 12.3 |
| 5 | 24.9 | 23.6 |
| 15 | 66.7 | 57.7 |
| 30 | 132.2 | 97.9 |
| 50 | 218.1 | 128 |
| 80 | 340.8 | 143 |
| 100 | 413 | 147.2 |
Study context
Simulation post-mortem. The discretization and lag-time diagnostics challenge the proposed state model in the tested process. A negative result is retained as part of the research record.
- discretization check
- If coarse grouping were the sufficient explanation, refinement should repair the timescale behaviour. The plotted comparison does not show that repair.
- lag-time diagnostic
- Compare the trajectory with the plateauing reference. The absence of the expected plateau is central to the model criticism.
More states do not produce the expected convergence.
The problematic trajectory does not collapse into a stable plateau merely by refining discretization. The extended-lag view also fails to show the required behaviour for the tested process. The plot retains a negative result instead of presenting complexity as an automatic improvement.
The conclusion is about model adequacy in this experiment. A failed state formulation motivates a different account of the process; it does not invalidate every sequential model or establish a universal fact about human learning.
What changes the precision of an observation?
More observations and richer measurements can change precision differently. Read the simulated error curve for each assumed correlation separately.
Binary-observation error
Hover or focus a point to read the value.
Archived numerical output; no human learner outcomes are measured. The high small-sample error is retained.
View the plotted values
| Sample count | Binary |
|---|---|
| 5 | 9.768 |
| 10 | 5.422 |
| 20 | 0.568 |
| 40 | 0.419 |
| 80 | 0.267 |
| 160 | 0.184 |
Continuous-observation error
Hover or focus a point to read the value.
Archived numerical output; no human learner outcomes are measured. These are distinct assumed correlations, not measured physiological validity.
View the plotted values
| Sample count | Correlation 0.3 | Correlation 0.5 | Correlation 0.7 |
|---|---|---|---|
| 5 | 0.581 | 0.444 | 0.279 |
| 10 | 0.43 | 0.284 | 0.197 |
| 20 | 0.28 | 0.209 | 0.139 |
| 40 | 0.205 | 0.139 | 0.104 |
| 80 | 0.141 | 0.109 | 0.069 |
| 160 | 0.102 | 0.074 | 0.051 |
Observation count and information quality are separate levers.
The archived simulation reports different error curves for binary observations and three assumed continuous correlations.
The assumptions belong to the simulation; these curves do not establish a universal sample-size rule.
Which conclusions survive when the same simulation is made much larger?
The large run tests computational scale under the same simulated assumptions. Its importance is the distinction between more observations and stronger validation—not the number of panels on a dashboard.
| Stress-run scope | Reported scale |
|---|---|
| Simulated learners | 3,000,000 |
| Simulated responses | 360,000,000 |
| Real students in this experiment | None. |
| What the scale establishes | Behaviour under the specified simulator at larger volume. |
| What it does not establish | That the simulator represents real learners. |
Study context
Simulation study: 3,000,000 simulated learners and 360,000,000 simulated responses. The population counts, AUC values, runtime, and reliability in this original dashboard are not real-user counts, production benchmarks, or demonstrated student outcomes.
- density
- The heat map summarizes the simulated process-space distribution and the indicated planted groups.
- Upper centre: recovery
- AUC compares two observation regimes against a known simulated target.
- sensitivity
- The depth sweep changes the planted distinction and follows how the recovery result changes.
- Middle left: reliability
- Agreement rises as more attempts are observed; this concerns repeatability within the simulated setting.
- Middle centre: exploration
- The distributions compare simulated item counts needed to expose the distinction.
- Middle right: accuracy overlap
- The histogram shows why the simulated groups can overlap when reduced to overall accuracy.
- added separation
- The process-axis view shows separation conditional on accuracy within the same simulator.
- Lower centre: population accounting
- The bars are synthetic population counts, not active students or customers.
- machine record
- Runtime and memory describe the reported stress run and its environment, not a production service guarantee.
Scale increases precision; it does not validate the planted mechanism.
The run contains three million simulated learners and 360 million simulated responses. Recovery remains strong for the process representation under the model, while the accuracy-only result changes as the planted distinction becomes shallower. The same research record also reports convergence and separation under repeated simulated observations.
The large counts are a computational stress test, not product adoption. Narrow uncertainty under one simulator cannot eliminate uncertainty about whether the simulator represents real learners. Every performance value on this dashboard remains conditional on that distinction.
Simulation study: 3,000,000 simulated learners and 360,000,000 simulated responses. The population counts, AUC values, runtime, and reliability in this original dashboard are not real-user counts, production benchmarks, or demonstrated student outcomes.
One experiment. Separate questions.
The recovered code is replayed with 30,000 simulated learners. Each graph uses the resulting coordinates, bins or estimates. The historical 3M stress-run record above remains separate.
Where does the simulated population concentrate?
Colour records density on a logarithmic scale. The bright ridge contains more simulated observations than the darker perimeter.
Hover a cell to read its value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result. Colour: log-scaled count, 1–409.
A dense projected region does not by itself show that the planted states can be distinguished.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
Does the observation distinguish the planted state?
Compare the held-out AUC estimates for accuracy alone and process observation. The reference at 0.5 is chance discrimination.
Hover or focus a point to read the value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result. Whiskers are 2.5th–97.5th bootstrap percentiles.
View the plotted values
| Measure | Replay estimate |
|---|---|
| Accuracy only | 0.659 (0.651–0.665) |
| Process observation | 1 (1–1) |
The result depends on the distinction planted by the simulator. It does not demonstrate detection in real students.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
What changes with the strength of the simulated distinction?
Accuracy-only recovery changes across the five tested settings. The process series is evaluated separately under the same simulation.
Hover or focus a point to read the value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result.
View the plotted values
| Simulated depth parameter | Process observation | Accuracy only |
|---|---|---|
| 1 | 1 | 0.558 |
| 1.6 | 1 | 0.596 |
| 2.2 | 1 | 0.647 |
| 2.8 | 1 | 0.683 |
| 3.4 | 1 | 0.721 |
Depth here is a simulator parameter, not an observed learning trajectory.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
How much observation makes the estimate consistent?
Split-half agreement is evaluated at 20, 60 and 120 attempts. The unequal spacing on the horizontal axis is preserved.
Hover or focus a point to read the value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result.
View the plotted values
| Attempts | Replay agreement |
|---|---|
| 20 | 0.348 |
| 60 | 0.888 |
| 120 | 0.965 |
Repeated measurements becoming consistent is a reliability result. It does not establish validity for real learners.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
How widely does the search cost vary?
The histograms compare the full distribution of observations needed by two simulated search policies.
Hover or focus a mark to read its value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result.
These are computational policies in the recovered experiment, not measured reductions in student study time.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
What does an overall score conceal?
The three density-normalized histograms show how different planted groups can overlap in overall accuracy.
Hover or focus a mark to read its value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result.
Read the overlap, rather than treating each group as an isolated average. Group names describe the simulator, not student diagnoses.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
Does another observation separate that overlap?
The horizontal axis is again overall accuracy. Colour gives the local fraction belonging to the planted hidden-state group.
Hover a cell to read its value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result. Colour: 0–1 group fraction.
Compare this with the overlapping histograms. Separation under a constructed observation model is not independent validation of that model.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
Which population produced these patterns?
Counts show the three randomly sampled groups in this replay. They add up to 30,000.
Hover or focus a point to read the value.
Recomputed from the recovered simulation code: 30,000 synthetic learners, 3.6 million responses, seed 7. This is a new smaller replay, not the archived 3-million-learner result.
View the plotted values
| Measure | Group count |
|---|---|
| Group A | 10497 |
| Group B | 8833 |
| Group C | 10670 |
These are synthetic members of a computational experiment, not product users or participants in a human study.
Related study
The archived stress run used 3,000,000 simulated learners and 360,000,000 responses. The graph shown here is a separate 30,000-learner replay of the recovered code. Its coordinates and estimates are not the archived run’s measurements.
Both experiments examine behaviour under a specified simulator. Increasing the population tests computational scale and sampling precision; it does not establish that the model describes real learners.
Study record
Source record: consolidated research report, pp. 3, 5–7. June 2026.
Interactive charts use verified aggregate values from the reports and archived numerical outputs. Chart notes distinguish these sources; resampling estimates can differ between runs. Where the underlying numerical series is unavailable, the text states the result without reconstructing unverified data. Study context and interpretation notes accompany each analysis. Detailed production procedures and the full internal documents are not included.