Skip to main content

Althing Ensemble (3-model)

3 runs · 3 datasets · 1 model

slug: althing-ensemble-3-model

0.877
Best SPS · opinionsqa

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions
  • opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)
  • subpop — ground truth: Suh et al., ACL 2025 — SubPOP: Subpopulation-Level Opinion Prediction

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

globalopinionqa product

ensemble--3-model-blend--tdefault--tplcurrent--3ba7cfb3

0.813 ± 0.028
SPS · 95% CI [0.785, 0.842] · n = 100
Question-type breakdown (7 topics)
Topic SPS p_dist p_rank p_refuse N
Health & Science low n — suggestive only 0.909 0.891 0.927 1.000 2
Politics & Governance 0.804 0.868 0.740 0.904 22
General Attitudes 0.796 0.842 0.749 0.964 11
Technology & Digital Life low n — suggestive only 0.743 0.893 0.594 0.968 3
International Relations & Security 0.743 0.801 0.685 0.955 50
Economy & Work low n — suggestive only 0.689 0.786 0.593 0.909 5
Trust & Wellbeing low n — suggestive only 0.520 0.561 0.480 0.982 7

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

opinionsqa product

ensemble--3-model-blend--tdefault--tplcurrent--1bef3e62

0.877 ± 0.006
SPS · 95% CI [0.870, 0.883] · n = 684
Question-type breakdown (10 topics)
Topic SPS p_dist p_rank p_refuse N
Social Values & Religion 0.866 0.860 0.872 0.981 37
Media & Information 0.853 0.856 0.850 0.975 63
Health & Science 0.852 0.844 0.859 0.975 47
General Attitudes 0.845 0.840 0.849 0.940 190
Trust & Wellbeing 0.841 0.854 0.829 0.978 25
Economy & Work 0.832 0.829 0.835 0.955 68
International Relations & Security 0.820 0.818 0.823 0.973 149
Technology & Digital Life 0.812 0.821 0.803 0.940 26
Politics & Governance 0.810 0.800 0.820 0.960 40
Identity & Demographics 0.809 0.822 0.797 0.956 39

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on opinionsqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

subpop product

ensemble--3-model-blend--tdefault--tplcurrent--ef75a647

0.858 ± 0.012
SPS · 95% CI [0.845, 0.870] · n = 200
Question-type breakdown (10 topics)
Topic SPS p_dist p_rank p_refuse N
Trust & Wellbeing low n — suggestive only 0.873 0.914 0.833 0.988 2
General Attitudes 0.842 0.847 0.836 0.904 33
Health & Science low n — suggestive only 0.815 0.877 0.753 0.994 6
International Relations & Security 0.814 0.796 0.832 0.992 38
Politics & Governance 0.798 0.792 0.803 0.984 23
Social Values & Religion 0.797 0.781 0.812 0.987 37
Identity & Demographics low n — suggestive only 0.793 0.813 0.774 0.995 1
Economy & Work 0.767 0.747 0.787 0.992 26
Technology & Digital Life 0.766 0.757 0.774 0.993 33
Media & Information low n — suggestive only 0.722 0.737 0.707 0.978 1

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on subpop. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

No demographic conditioning data has been published for this vendor yet. The question-type matrix above shows topic-level parity; subgroup rows fill in once Althing-style conditioned runs land.

← Back to leaderboard