What each cell measures
Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup —
published survey ground truth, not a model's guess about the subgroup.
Higher = closer to how that real subgroup actually answered.
Demographic conditioning here is not stereotyping: every conditioned
score is checked against what real subgroup members said, not against
assumptions about them — and where the vendor's output diverges from
the real subgroup, the score drops.
- globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions
- opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)
- subpop — ground truth: Suh et al., ACL 2025 — SubPOP: Subpopulation-Level Opinion Prediction
Low-n cells:
cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are
reported for transparency, not as findings.
CIs:
"no CI — single run" marks point estimates from a single run with no
confidence interval yet; treat the uncertainty as unknown, never zero.
Columns:
Score = distributional parity (p_dist) restricted to the subgroup ·
p_cond = conditioning strength vs the unconditioned baseline ·
N = questions answered under that conditioning · Cov. = coverage
dot derived from N (green = high, ≥100 · amber = medium, 50–99 ·
red = low, <50). Topic tables: SPS = Survey Parity Score for the
topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 ·
p_refuse = 1 − mean(|R_model − R_human|).
Multiple comparisons:
with this many subgroup cells, a few extreme cells are expected by
chance alone — read patterns across a dimension, not single cells.
Question-type breakdown
(6 topics)
| Topic | SPS | p_dist | p_rank | p_refuse | N |
| Health & Science
low n — suggestive only
| 0.887 | 0.775 | 1.000 | 1.000 | 1 |
| Economy & Work
low n — suggestive only
| 0.703 | 0.770 | 0.636 | 0.949 | 3 |
| Trust & Wellbeing
low n — suggestive only
| 0.680 | 0.592 | 0.768 | 0.733 | 2 |
| International Relations & Security | 0.669 | 0.653 | 0.685 | 0.906 | 60 |
| Politics & Governance | 0.661 | 0.620 | 0.703 | 0.887 | 30 |
| General Attitudes
low n — suggestive only
| 0.588 | 0.713 | 0.463 | 0.983 | 4 |
Topics with N < 10 are muted and tagged
"low n — suggestive only": too few questions for a stable
3-decimal score.
Not yet measured
This vendor has no demographic-conditioned runs for age, geography,
education, or any other subgroup dimension on globalopinionqa.
No cells are fabricated — scores appear here only when a
conditioned run actually measured them.
How to submit a demographic-conditioned run →
Question-type breakdown
(9 topics)
| Topic | SPS | p_dist | p_rank | p_refuse | N |
| Health & Science
low n — suggestive only
| 0.806 | 0.758 | 0.854 | 0.936 | 1 |
| Social Values & Religion | 0.762 | 0.742 | 0.782 | 0.964 | 15 |
| Technology & Digital Life
low n — suggestive only
| 0.761 | 0.715 | 0.808 | 0.995 | 4 |
| Economy & Work | 0.755 | 0.720 | 0.790 | 0.911 | 21 |
| Trust & Wellbeing
low n — suggestive only
| 0.743 | 0.669 | 0.816 | 0.873 | 1 |
| International Relations & Security | 0.708 | 0.716 | 0.700 | 0.905 | 20 |
| Media & Information
low n — suggestive only
| 0.701 | 0.744 | 0.658 | 0.969 | 2 |
| General Attitudes | 0.682 | 0.649 | 0.716 | 0.954 | 35 |
| Politics & Governance
low n — suggestive only
| 0.576 | 0.494 | 0.658 | 0.670 | 1 |
Topics with N < 10 are muted and tagged
"low n — suggestive only": too few questions for a stable
3-decimal score.
Not yet measured
This vendor has no demographic-conditioned runs for age, geography,
education, or any other subgroup dimension on opinionsqa.
No cells are fabricated — scores appear here only when a
conditioned run actually measured them.
How to submit a demographic-conditioned run →
Question-type breakdown
(9 topics)
| Topic | SPS | p_dist | p_rank | p_refuse | N |
| Politics & Governance
low n — suggestive only
| 0.772 | 0.716 | 0.828 | 0.981 | 4 |
| Technology & Digital Life | 0.745 | 0.710 | 0.780 | 0.932 | 16 |
| Health & Science
low n — suggestive only
| 0.722 | 0.680 | 0.764 | 0.973 | 6 |
| International Relations & Security | 0.702 | 0.649 | 0.755 | 0.982 | 15 |
| Social Values & Religion | 0.690 | 0.614 | 0.765 | 0.962 | 27 |
| Economy & Work | 0.676 | 0.629 | 0.724 | 0.962 | 13 |
| General Attitudes | 0.608 | 0.592 | 0.623 | 0.924 | 17 |
| Identity & Demographics
low n — suggestive only
| 0.506 | 0.394 | 0.618 | 0.995 | 1 |
| Media & Information
low n — suggestive only
| 0.505 | 0.466 | 0.543 | 0.978 | 1 |
Topics with N < 10 are muted and tagged
"low n — suggestive only": too few questions for a stable
3-decimal score.
Not yet measured
This vendor has no demographic-conditioned runs for age, geography,
education, or any other subgroup dimension on subpop.
No cells are fabricated — scores appear here only when a
conditioned run actually measured them.
How to submit a demographic-conditioned run →
No demographic conditioning data has been published for this vendor yet. The
question-type matrix above shows topic-level parity; subgroup rows fill in
once Althing-style conditioned runs land.