Synthetic personas + the scenarios engine

The calibration programme runs on a library of 3,280 synthetic personas (2,880 in persona_library_v5.jsonl plus a 400-persona mid-grid supplement; the v5 generation replaced the deprecated v4 library on 2026-08-29) — sampled from a statistical model of how traits cluster in the human population — playing short interactive sessions through the production session engine; the confirm run of September 2026 played 40 of them through 1,831 sessions (calibration/phaseA/wp26/confirm_result_v1.json). Everything measured in this programme so far is synthetic: no human being has been measured, and the validation page leads with that boundary. This page explains why synthetic personas are scientifically defensible at this stage, how we model the trait dependence structure, and the demographic-blind scoring policy that guarantees every reader's SoulMap results are scored by exactly the same rules.


Why synthetic personas

Two reasons: ethics and scale.

Ethics. The conversational evidence we calibrate against is reflective and emotional. Building a Phase-1 calibration corpus from real users would mean asking thousands of people to share intimate reflections specifically so we can fit a statistical model — a use of their data they did not sign up for, and one we are not yet willing to ask for. Synthetic personas let us validate the inferential machinery before any real user data is involved. A future real-user study will recalibrate against real-user data with an explicit consent flow; every calibration to date is synthetic-only.

Scale. A defensible IRT calibration needs enough personas across enough constructs to fit stable item parameters. For the first-generation graded-response design, published guidance put the need at roughly N = 1,500–2,000 personas. The delivered programme's own measurements later showed that at our evidence volumes the binding constraint is rows per persona rather than persona count (the wave-1.3 axis-precision curves; the reliability analysis in the 2026-08-19 audit). Synthetic personas are cheap to produce in this volume; real-user corpora are not — both because an IRB-approved real-user study at this scale takes a year or more, and because the Phase-1 job is to demonstrate that the machinery works before any large-scale study is asked of real people.

The synthetic-persona corpus is the warm-up, not the destination. The destination is real-user calibration with measurement-invariance testing across demographic subgroups, which has not yet run. What the synthetic programme has shown so far is stated with its limits on the validation page — including that no human being has been measured anywhere in it.


We model how traits cluster realistically — not as independent dice rolls

The single largest decision in the persona forge is how the 29 trait dimensions correlate with each other.

A naïve approach would draw each trait independently — sample Conscientiousness from one normal distribution, Agreeableness from another, Extraversion from a third, all independent. The problem with the naïve approach is that real human personality does not work that way: people who score high on Conscientiousness also tend to score moderately high on Agreeableness; people who score high on Honesty-Humility tend to score low on the Dark Triad; people with anxious attachment styles tend to score lower on Emotional Stability. These patterns are not noise — they are the structure that any honest model of human personality has to capture.

We model this dependence structure with a Gaussian copula. A Gaussian copula is a way of saying: draw the traits from a multivariate normal distribution that has whatever correlation structure you want, then map each dimension's distribution into its trait-specific shape (skewed for Honesty-Humility, slightly skewed-positive for Emotional Stability, roughly normal for Conscientiousness). The dependence (which traits cluster together) and the per-trait shape (the marginal distribution) are decoupled — we get to specify each independently.

The Gaussian copula is the right starting point for the bulk of the trait dependence. It is not the right tool for the joint extremes of the Dark Triad — Machiavellianism, Narcissism, Psychopathy — which co-occur at higher rates than a pure Gaussian copula would predict. For the Dark Triad triplet specifically, we use a block-t copula at degrees-of-freedom ν = 5, which produces the heavier joint upper tails we observe in real Dark Triad data. The result: a dependence model that captures both the bulk-of-the-distribution clustering and the tail-clustering that the Dark Triad exhibits.

Three ways to draw correlated traits. Independent dice rolls miss the structure entirely; the Gaussian copula captures the bulk; the block-t copula additionally captures the joint extremes the Dark Triad exhibits.
The Gaussian copula captures the bulk of trait dependence. The block-t copula on the Dark Triad captures the joint extremes the Gaussian copula systematically under-samples.

The four-tier Σ hierarchy, in one paragraph

The Gaussian copula needs a 29 × 29 correlation matrix Σ — the dependence structure we want to recover. Where does Σ come from? Where empirical data exists, we use it; where peer-reviewed meta-analyses exist, we use those; where neither exists, we build statistical bridges; where even the bridges are missing, we use explicit conservative uncertainty. The four tiers, in order of priority: Tier 0 is empirical correlations from open multi-instrument datasets (88 cells of Σ are filled this way; after largest-sample deduplication three sources contribute surviving cells — Anglim & Marty 2024, Anglim et al. 2020, and the Open-Source Psychometrics SD3 dataset — while Anglim 2017 and SAPA were extracted but fully deduplicated away); Tier A is direct meta-analytic estimates from the published literature (29 cells; Schmitt 2008, Lee & Ashton 2020, Roberts 2006, etc.); Tier B is Wright (1934) single-bridge path-implied estimates with a 0.7 attenuation hedge (120 cells; what we get when we have A↔B and B↔C but not A↔C directly); Tier C is a weakly-informative shrinkage prior, N(0, 0.05) truncated to (−0.30, 0.30) (169 cells; what we get when we have nothing else and the honest answer is "we believe the correlation is small"). The assembled 29 × 29 matrix is then projected to nearest positive-definite via Higham 2002, which guarantees a mathematically valid copula at relative Frobenius distance 0.0277 from the source-of-truth correlations (recorded in sigma_v1.yaml). Every cell of Σ carries its tier and source in a provenance manifest; reviewers can trace any correlation back to the dataset or meta-analysis it came from.


Demographic conditioning: at persona generation, never at scoring

This is the section that matters most for anyone reading their own SoulMap results — and for legal compliance.

At persona generation, we condition on age and sex on the latent layer before transforming the multivariate normal draw into trait values. Why: real human personality varies systematically with age and sex. Women score slightly higher on Agreeableness on average; older adults score higher on Conscientiousness; the Dark Triad shows different patterns by sex. A persona corpus that ignored these patterns would not be a realistic distribution; the synthetic-of-synthetic validation would catch this immediately.

At scoring time, demographic information is never an input feature to the downstream scoring engine. Neither the deployed aggregation path that produces served scores nor the research instrument's choice model sees the speaker's age or sex; in the calibration corpus, the persona sheet handed to the role-playing model likewise never conveys age or gender. Scoring is global; measurement-invariance testing across demographic subgroups is the discipline that would prove the global model fair, and it has not yet been run.

This is the demographic-blind scoring policy. It is the standard psychometric playbook (Cheung & Rensvold 2002 for the invariance methodology; Embretson & Reise 2000 chapter 12 for the recalibration discipline) and it is the right answer for both scientific and commercial reasons:

  • Scientifically: per-subgroup parameter sets would mean different "true" trait scores for different demographics, which is exactly what measurement-invariance testing exists to prevent.
  • Practically: the "same trait inferences regardless of who you are" guarantee is the strongest single commitment we make to the people who read their own results. It is also what keeps the system off the high-risk-profiling list of the EU AI Act.

The future real-user recalibration is what would prove this discipline. When it runs, we will run measurement-invariance testing across age × sex × education and report the results — it has not yet run, and we state that as a gap. The policy is: condition at generation to make the corpus realistic, keep demographics out of scoring, and test invariance before claiming fairness rather than after.

The same logic applies to attachment-cluster modelling: we use the published Mickelson, Shaver & Kessler (1997) base rates as the target distribution for the four-style attachment taxonomy at persona generation, but the scoring engine does not predict attachment style — it estimates the two continuous attachment dimensions (Anxiety, Avoidance) and lets the user (or the application layer) decide whether to bin them into clusters.


The scenario side: from a template library to the session engine

A persona is half the corpus. The other half is a scenario — the conversational situation that exercises the persona's traits.

The current instrument is not a scenario-template library but the production session engine itself: each session is a three-turn interactive narrative whose answer options are dealt from a validated evidence shelf of 366 entries (table sha 7bda5d9bf355 since the wording revision of 3 September 2026; the confirm run was banked under its predecessor cfbd7786eb4e, and every row carries the sha it was dealt from). The seventy-five-template, twelve-turn scenario library described in earlier revisions of this page belongs to the decommissioned first-generation pipeline.

The coverage concern the old stratification addressed is real — different content exercises different constructs, and unmanaged sampling starves some of them — so the current programme measures offered-slot exposure on every banked corpus rather than counting shelf entries. On the confirm run, 312 of the 366 entries were offered at least once and every one of the 27 families was genuinely contrasted on at least 294 of the 3,822 rows; the 54 entries never offered are almost all reflection-turn entries, a known gap in how the reflection turn is dealt.

The evidence shelf is versioned and hash-stamped into every corpus row (evidence_table_sha), for the same reason the old design locked its scenario library: fitted parameters must be about the measurement model, not about content drift. Fits made under different shelf or axis versions are treated as non-comparable.


Citation roster

This page draws on the methodology that earned its place in published psychometrics over the past sixty-plus years:

  • Sklar (1959), Joe (2014) — copula theory generally
  • Demarta & McNeil (2005) — t-copula tail dependence
  • Embrechts, McNeil & Straumann (2002) — why correlation alone is insufficient outside elliptical distributions
  • Higham (2002) — nearest-PD projection
  • Schmitt et al. (2008), Lee & Ashton (2020), Friesdorf, Conway & Gawronski (2015), Roberts, Walton & Viechtbauer (2006) — sex / age conditioning meta-analyses
  • Mickelson, Shaver & Kessler (1997) — adult attachment base rates
  • Anglim et al. (2020, 2024) and the Open-Source Psychometrics SD3 dataset — the Tier-0 sources whose cells survive largest-sample deduplication; Anglim et al. (2017) and Condon & Revelle (2017) were extracted but fully deduplicated away
  • Cheung & Rensvold (2002), Embretson & Reise (2000) — measurement invariance + IRT discipline

The complete bibliography with primary-source URLs and one-sentence "what this paper gave us" annotations sits at /about/science/citations.


What this design does not claim

To set expectations honestly:

  • Synthetic personas are not real users. Phase-1 calibration parameters will need recalibration against real-user data in Phase-3. We commit to that explicitly.
  • The dependence model is the best one we can build from the published evidence. It is not a perfect representation of reality; it is a defensible representation of what the published evidence supports.
  • The block-t copula on the Dark Triad uses ν = 5. This is the published value (Demarta & McNeil 2005); we ship Phase-1 at ν = 5 and hold a sensitivity check at ν = 8 in reserve for Phase-2 if the data suggests it.
  • Age conditioning uses a uniform 18–70 prior at Phase-1. Real-user demographics will replace this in Phase-3.
  • Cultural conditioning is deferred. Cross-cultural personality variation is real (Schmitt 2008 is itself a fifty-five-culture study), but adding cultural axes adds another five-to-ten dimensions to the conditioning structure. We defer this to Phase-3 once the real-user cohort tells us which subgroups actually matter for the SoulMap user base.

The next page, the calibration loop, states what the calibration programme has actually shown on this synthetic population — with artifacts, caveats, and the withdrawn claims named.