Roadmap
First, the boundary that governs everything on this page: every validation result we hold comes from synthetic personas driven by a language model. No human being has been measured, and no phase below changes that until the real-user study actually runs.
Second, how this page relates to what it used to say. Two earlier generations of the programme are recorded in the repository's history: a first-generation scoring chain that never ran in production and was decommissioned on 2026-08-04, and a second generation — an earlier corpus engine over an earlier persona library — whose corpora were deprecated on 2026-09-02 once the engine, the persona generator and the persona instruction had all changed. This page describes the third: the current engine, its one registered run, and what follows from it.
Delivered — the current engine and its first registered run
Status: built, run, banked.
- The persona generator, rebuilt (v5). A 3,280-persona library sampled
from the same copula architecture as before, with a rebuilt persona
instruction: every construct named (the earlier sheet named only the top
two moral foundations and one attachment word), the trait ladder read in a
randomised order per persona and stamped on every row, and every band's
wording placed by blind readers at the level it claims — fourteen of the
fifteen upper emotional-intelligence bands place exactly, where the
earlier sheet read a top-quartile persona as an average one
(
calibration/phaseA/wp22/–wp24/). - The wave engine, made resilient. Persona-driver calls run in bulk; a
vendor quota is a wait, not a death; the run resumes from checkpoints at
no re-spend; a stalled batch job is retried rather than fatal; and, after
the confirm run, each chain link's hand-off state is banked so a resume
can no longer re-derive it wrong (
functions/corpus_wave/). - The confirm run (1–3 September 2026): 40 v5 personas, 50 planned sessions each in chains of 20, one driver, $123.16. Registered before it ran; analysed once. The twelve historically unreadable families express above chance (pooled +0.126 against a bar of 0.082, one in 240 by chance), sitting 0.23 below the five classic families. In the served configuration the engine's own per-trait estimate reaches a median recovery of 0.23 against planted truth, ten of 27 families are individually confirmed, one — liberty/oppression — is confirmed backwards and three more read negative (all four withheld from the product), and per-person consistency at fifty sessions is a median of 0.15 — one family meets the 0.70 decision-grade bar. The validation page carries every number with its artifact.
In progress — fixing what the run exposed
Status: underway, September 2026.
- Re-verifying the twelve weakest families against their scoring keys. Four families read backwards and eight read flat. Flipping a sign because a correlation is negative is a refuted lever in this programme — it looked like a large gain in-sample and made the engine worse held out — so the sanctioned method is semantic: independent readers compare each construct's language with the actual shelf actions that score positive and negative on its axis and rule aligned or inverted, with a fix only where two of three agree and only as the smallest text change. This is the method that found and fixed two resilience inversions in August. Some families will come back "aligned — the model simply will not play it", and those get no fix. Until re-verified, the backwards families are withheld from the product.
- A small paired run afterwards, to see whether the touched families' signs actually moved, before anything is deployed on the strength of a text edit.
- The fifty-session milestone. Completed sessions per persona reached a median of 46 on the confirm run, not 50, because 100 sessions died of replay drift across the run's eight restarts. That defect is fixed and tested; the next chained run should reach the milestone without further change.
Next — breadth, then people
A breadth run. The confirm run was sized to answer one pooled question, and it did. What it cannot do is certify each of 27 traits on its own: the per-family intervals are about ±0.27 wide at 40 personas and did not narrow between 27 and 50 sessions, because their width is set by the number of people, not sessions. Roughly 170 personas would pin each family to ±0.15; about 380 to ±0.10. That run is the next spend, once the re-verification above has landed.
Per-person consistency. The honest gap the confirm run measured is between reading a population (median recovery 0.23; ten confirmed families) and reading an individual steadily (median split-half 0.15). Every served reading carries its own interval and a plain confidence word for exactly this reason. Closing the gap is measured work: more information per session, not more sessions — the programme's ceiling analysis shows that at the current per-session information, playing longer asymptotes below the bar.
Real-user validation. Status: planning; consent-blocked. Item-level rows from real players need a consent scope the current catalog does not carry; sessions are retained long enough that a later scope can still cover them. When it runs: recalibration against real-user behaviour with an explicit consent flow, measurement-invariance testing across age × sex × education × geography, and a real cohort distribution to replace the uniform 18–70 age prior. Earliest expected start: late 2026; realistic: 2027. It is the only step that can turn any synthetic result into a claim about people.
What's NOT on the roadmap
- No clinical assessment claims, ever. The 29 traits are normal-range psychological tendencies, not psychiatric diagnoses. The system is not a screening instrument; we do not plan to build one.
- No cultural conditioning yet. Cross-cultural personality variation is real (Schmitt 2008 documents it across 55 cultures). Adding cultural axes adds 5–10 dimensions to the conditioning structure; deferred until the real-user cohort shows which subgroups matter for the SoulMap user base.
- No per-user model. Calibration is global. The stack produces per-user estimates with uncertainty without per-user parameter fitting.
- No personalisation pipelines. SoulMap is a measurement instrument, not a recommender system.
- No real-time online learning. Any calibration that changes served parameters must pass pre-registered gates first; today the served configuration is the deliberate uncalibrated baseline.
- No dimensional expansion beyond 29 in the foreseeable future.
- No demographic features as inputs to the scoring engine. The demographic-blind scoring policy (personas page). Hard line.
- No sign-flipping by statistics. Refuted, with a reopening condition that only an independent channel — human responses, or a second driver family — can meet.
What would falsify the plan
- Run-level: a run whose attrition exceeds the registered 15% bar is void against its own registration. The confirm run passed at 8.45%.
- Test-level: a registered threshold is a threshold. The confirm run's pooled bar (0.0822) was fixed before the data existed and the analysis was run once; a miss would have been published as a miss, as two earlier registered tests were.
- Instrument-level: a family that reads negative on two independent corpora is a blocking finding until traced — the programme's pole-inversion contract — and four families are in that state now.
- Transfer-level: no synthetic result, however strong, counts as evidence about human beings. The falsifier for the whole programme is the first real-user study.
The next page, citations, is the alphabetical bibliography that backs every footnote on every page in this tier.