EVIDENCE

We prove it, not just claim it

We measure decision recovery on 20 public datasets and calibrate every verdict grade against those results.

The numbers below are measured backtest results on public datasets. Paired-pilot comparison against real customer campaigns has not started collecting yet; until it completes, the overall evidence label stays directional.

01What We Measure

Scale-invariant decision recovery

We score the decisions operators actually make, not absolute numbers.

winner

Did we call the winning arm

directional

Is each pairwise direction correct

control-sign

Does each treatment beat control as measured

ρ̄ (Spearman)

Within-group rank correlation

02Measured Results

Market decision recovery

Measured on 11 contrast groups, 3 samples, same model.

Predictorwinnerdirctrl-signρ̄
No-model baseline0.000.000.00−1.00
Single-sample panel0.6360.7840.8570.461
Persona panelShips in product0.7270.8300.9290.625
Structural prior (encodable contrasts)1.0000.8980.9290.961
  • Multidomain generalization: winner holds at 0.750 on e-commerce and income domains outside marketing.
  • On the pricing/choice lane (ModeCanada) the structural prior beats the LLM panel, so we route there.
03Cross-Industry Breadth

Recovering real market decisions across 20 industries

Testing whether the directional methodology generalizes beyond marketing. A leakage-safe structural prior — an auditable proxy for the panel's directional beliefs — is scored on scale-free decision recovery across 20 public datasets (37 contrast groups, 118 cases).

predictorwinnerdirctrl-signρ̄
No-model baseline0.000.010.00−1.00
Structural directional prior0.9460.8990.9080.876
Industries covered (20 datasets)
Email marketingBank telemarketingE-commerce sessionsCensus / incomeRetail promotionTelecom churnTravel pricingPolitical GOTVNonprofit fundraisingBehavioural pricingHospitalityFintech / retirementHealthcareTax / governmentRetail merchandisingPricing psychologyPersuasion / copyEnergy / utilitySales / commitmentNonprofit asks

Beating the floor shows the directions are sound; it does NOT prove the LLM achieves a specific accuracy. The two misses are counter-intuitive reversals the prior deliberately does not fit (fibre-optic churn, new-visitor conversion).

04Failure-Mode QA

We measure the #1 failure mode of synthetic panels

0 groups
Variance collapse

Predicted distributions stay as dispersed as reality. The signature defect of synthetic responses was not observed.

0.21 vs 0.04
Sign flips (panel vs structural)

Sign-flip rates of our two prediction paths — the basis for routing each contrast type to the stronger path.

73.1%
Sentiment (NSMC)

Accuracy at 100% coverage after threshold calibration.

05Trust Gate

Every verdict carries a grade

Rules calibrated on the backtests classify each verdict into three grades.

decision grade

Granted only when the winner holds across repeated samples on a validated contrast type.

directional only

The direction is reliable; confirm with a real A/B before betting on it.

inconclusive

Near-ties and underpowered tests get no call.

believe what repeatsIf the winner wobbles across samples, decision grade is revoked regardless of margin.
paired-pilotPredictions are paired with real outcomes over time — the only path to the human_calibrated label.

Paired-pilot collection is at the starting line. Until it completes, the overall evidence label stays directional.

06Published Boundaries

We publish the boundaries of the measurement

Because the boundaries are public, a graded verdict is one you can trust.

  • Calling a winner between two near-identical creatives is chance-level. The trust gate caps it automatically.
  • Counter-intuitive micro-effects (e.g. new visitors beating returning) are missed by both the structural prior and the LLM.
  • Absolute conversion and spend figures are unvalidated across population and era gaps. We score order-based decisions only.
  • Behavioral-cohort popularity segments sit at winner 0.222, so the entire segments lane is capped at directional.
We publish our falsified hypotheses too

We tested the cross-framing agreement gate against pre-registered criteria. The result was the opposite: agree-subset winner 0.308 vs disagree 0.667. Intra-model framing agreement measures the consistency of shared bias, so the gate is restricted to downgrade-only.

07Reproduce It

Reproduce it yourself

No API key needed — the commands below reproduce every result.

$npm run backtest:panel -- --config panel_persona_v1 --samples 3# decision recovery
$npm run backtest:distribution# variance / sign-flip QA
$npm run backtest:ssr# sentiment (SSR)
$npm run backtest:market-realism# structural prior
$npm run backtest:cross# cross-industry (20 datasets)
$npm run backtest:segments -- --samples 3# segments lane cap
$npm run backtest:segments -- --samples 3 --framings 2# cross-framing falsification
$npm run pilot:status# paired-pilot calibration
Datasets measured (21)
Hillstrom Email A/BUCI Bank MarketingUCI Online ShoppersCensus IncomeNSMC (Korean sentiment)Starbucks Promotion A/BTelco ChurnModeCanada (pricing/choice)GOTV Social PressureCharity Match OfferZero-Price EffectHotel Social Proof401(k) DefaultVaccination PromptTax Social NormChoice Overload (Jam)Decoy EffectThe Word "Because"Home Energy ReportFoot-in-the-DoorEven-a-Penny Ask

Start with a verified direction

Start Free