The science
We map how a person's mind tilts across 29 trait dimensions, named for the constructs behind ten of the most-studied questionnaires in psychology — without ever asking you to fill out a questionnaire. (Calibration measures these as 27 families: two bipolar value dimensions each fold a pair of scales; the projection is spelled out on the methodology page.) Two things must be said before anything else. First: every validation result we hold comes from synthetic personas driven by a language model; no human being has been measured, and whether the results transfer to people is unestablished. Second: the estimator that serves your trait scores is the same one the calibration programme measures, in its deliberate uncalibrated configuration — no fitted parameters are mounted. On the one registered run of the current engine (September 2026; 40 synthetic personas, ~46 completed sessions each), that served estimator recovers planted traits at a median correlation of 0.23, reads ten of 27 families well enough to confirm each on its own, and describes the same person consistently across two halves of their play on only a handful of them (median 0.15). We state both boundaries plainly rather than bury them.
This page is the hub. The deeper material lives in seven sub-pages. If you have ten minutes, read the data layers page; if you have an hour, read all seven in order. If you want the measurement standards we hold ourselves to — and an honest account of which of them we currently meet — methodology is the page you want.
The big picture, in three cards
| Card | What it covers | Read |
|---|---|---|
| The four data layers | Signals → Rules → Constructs → Traits — the four-layer measurement architecture SoulMap was designed around, read with two governing corrections: the rule layer does not execute in production, and a served score today traces to your recorded choices and their table-stamped construct weights. | 9 min |
| Synthetic personas + scenarios | Why we model how traits cluster realistically — not as independent dice rolls — using a Gaussian copula whose dependence structure is assembled from open empirical datasets, peer-reviewed meta-analytic estimates, statistical bridges, and explicit conservative uncertainty where neither bridge nor data exists. | 11 min |
| The calibration programme | The run record of the current engine: how the confirm run was registered and executed, what it cost, what it lost and why, and what each banked artifact may and may not support. | 10 min |
And three more:
| Card | What it covers | Read |
|---|---|---|
| Methodology | How measurement works end to end — the server-dealt evidence shelf, the choice model the product actually serves, the synthetic-persona programme, and the testing discipline: pre-registration, a shuffled null, bootstrap intervals, split-half consistency. | 12 min |
| Validation | What the current engine has demonstrated, at what strength, on what population — one pre-registered run, every number tied to its banked artifact, the misses stated as plainly as the passes. | 9 min |
| Roadmap | Where the programme stands after the confirm run: what is being re-verified, what the next run is for, and what real-user validation would take. | 6 min |
Plus an alphabetical citations roster that backs every footnote on every page in this tier.
Why this exists
Standardised questionnaires are the gold-standard measurement tool in personality psychology. They also have well-known limitations: they ask people to evaluate themselves on abstract claims ("I am the life of the party"), they suffer from acquiescence bias and social desirability bias, and they ask people to remember and average behaviour over weeks or months. Behavioural evidence — what someone actually chooses to do in a situation — is richer than questionnaire evidence, but it is also less standardised: different situations elicit different signals; different people reveal different facets of themselves.
Our research programme applies the inferential machinery that makes questionnaires defensible — explicit measurement models, calibrated uncertainty, pre-registered tests — to short interactive sessions instead. That machinery is validated so far only against synthetic personas; the product serves the measured estimator in its uncalibrated configuration, and we say so rather than blur the two.
Behavioural evidence is richer than questionnaire evidence — but it is also less standardised. The science here is in bridging that gap.What we are NOT trying to do
The boundaries matter as much as the claims:
- We do not diagnose. We are not a clinical assessment tool. The trait estimates we produce describe normal-range variation in psychological tendencies, always with their uncertainty attached; they are not psychiatric diagnoses, they are not screening instruments, and they should not be used as either.
- We do not replace clinical assessment. If you are wondering whether a person is at clinical risk for any condition, the right tool is a licensed clinician, not us. Our calibration data is from non-clinical synthetic populations; we have no claim on clinical sensitivity or specificity.
- Demographic validity is established against real users, not assumed from synthetic data. Every calibration artifact we hold today is synthetic-population-derived. Real-user recalibration with explicit measurement-invariance testing remains future work, and we will not claim demographic validity before it runs.
- We do not promise individual prediction at clinical-grade reliability. Every reading is served with an explicit uncertainty band and a plain confidence word; today every reading is also stamped provisional, because on our own measurements only one trait family meets the 0.70 bar for a decision-grade individual reading. Four families whose readings ran backwards on the registered run are withheld from the product until their scoring has been re-verified. We treat these as findings to publish, not caveats to bury.
This boundary is documented as the demographic-blind scoring policy: demographic information shapes the synthetic ground truth and the persona's behaviour in scenarios, but it is never an input feature to the scoring path that produces a served trait score. Scoring is global; measurement-invariance testing across demographic subgroups is the discipline that would prove the global model fair, and it has not yet been run — we state that as a gap, not a formality. The personas page carries the full rationale; this remains the strongest guarantee we make to the people who read their own results, and it avoids the high-risk-profiling flags of the EU AI Act.
A short reading order
For a journalist or a public visitor with ten minutes: just the data layers page, plus the "What we are NOT trying to do" box above.
For a technically trained reader who wants to audit the full chain behind their own readings: data layers → personas → methodology → validation. About 40 minutes.
For a journalist or analyst writing about the methodology: data layers → calibration → validation. About 25 minutes. The citations page is the bibliography.
Every artifact these pages cite is in the repository under calibration/phaseA/wp26/, with the rows' hash and the commit it was computed from.
Where the data flows
The reflective sessions themselves stay private. Nothing about your session contents leaves the EU; nothing trains another user's model. Your trait estimates are visible in your account; you can export everything as a single JSON file, you can withdraw from the analytics export at any time, and you can delete your account, which cascades deletion across our stores (an independently verified end-to-end erasure test is on the roadmap before the analytics export resumes). The product-level data controls live in the in-app Settings → Privacy & Data page.
This is the science page. The infrastructure that makes the calibration runs work on Google Cloud — the architecture, the operations, the cost engineering, the receipts — sits at /about/cloud-infrastructure.