We prove it, not just claim it
We measure decision recovery on 20 public datasets and calibrate every verdict grade against those results.
The numbers below are measured backtest results on public datasets. Paired-pilot comparison against real customer campaigns has not started collecting yet; until it completes, the overall evidence label stays directional.
Scale-invariant decision recovery
We score the decisions operators actually make, not absolute numbers.
Did we call the winning arm
Is each pairwise direction correct
Does each treatment beat control as measured
Within-group rank correlation
Market decision recovery
Measured on 11 contrast groups, 3 samples, same model.
| Predictor | winner | dir | ctrl-sign | ρ̄ |
|---|---|---|---|---|
| No-model baseline | 0.00 | 0.00 | 0.00 | −1.00 |
| Single-sample panel | 0.636 | 0.784 | 0.857 | 0.461 |
| Persona panelShips in product | 0.727 | 0.830 | 0.929 | 0.625 |
| Structural prior (encodable contrasts) | 1.000 | 0.898 | 0.929 | 0.961 |
- Multidomain generalization: winner holds at 0.750 on e-commerce and income domains outside marketing.
- On the pricing/choice lane (ModeCanada) the structural prior beats the LLM panel, so we route there.
Recovering real market decisions across 20 industries
Testing whether the directional methodology generalizes beyond marketing. A leakage-safe structural prior — an auditable proxy for the panel's directional beliefs — is scored on scale-free decision recovery across 20 public datasets (37 contrast groups, 118 cases).
| predictor | winner | dir | ctrl-sign | ρ̄ |
|---|---|---|---|---|
| No-model baseline | 0.00 | 0.01 | 0.00 | −1.00 |
| Structural directional prior | 0.946 | 0.899 | 0.908 | 0.876 |
Beating the floor shows the directions are sound; it does NOT prove the LLM achieves a specific accuracy. The two misses are counter-intuitive reversals the prior deliberately does not fit (fibre-optic churn, new-visitor conversion).
We measure the #1 failure mode of synthetic panels
Predicted distributions stay as dispersed as reality. The signature defect of synthetic responses was not observed.
Sign-flip rates of our two prediction paths — the basis for routing each contrast type to the stronger path.
Accuracy at 100% coverage after threshold calibration.
Every verdict carries a grade
Rules calibrated on the backtests classify each verdict into three grades.
Granted only when the winner holds across repeated samples on a validated contrast type.
The direction is reliable; confirm with a real A/B before betting on it.
Near-ties and underpowered tests get no call.
Paired-pilot collection is at the starting line. Until it completes, the overall evidence label stays directional.
We publish the boundaries of the measurement
Because the boundaries are public, a graded verdict is one you can trust.
- Calling a winner between two near-identical creatives is chance-level. The trust gate caps it automatically.
- Counter-intuitive micro-effects (e.g. new visitors beating returning) are missed by both the structural prior and the LLM.
- Absolute conversion and spend figures are unvalidated across population and era gaps. We score order-based decisions only.
- Behavioral-cohort popularity segments sit at winner 0.222, so the entire segments lane is capped at directional.
We tested the cross-framing agreement gate against pre-registered criteria. The result was the opposite: agree-subset winner 0.308 vs disagree 0.667. Intra-model framing agreement measures the consistency of shared bias, so the gate is restricted to downgrade-only.
Reproduce it yourself
No API key needed — the commands below reproduce every result.