Skip to main content

Althing (Sonnet 4)

1 run · 1 dataset · 1 model

slug: althing-sonnet-4

0.829
Best SPS · opinionsqa

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

opinionsqa product

synthpanel--claude-sonnet-4--tdefault--tplcurrent--8d1cda33

0.829 ± 0.008
SPS · 95% CI [0.821, 0.836] · n = 684
Question-type breakdown (10 topics)
Topic SPS p_dist p_rank p_refuse N
Technology & Digital Life 0.827 0.793 0.861 0.913 26
Media & Information 0.780 0.765 0.795 0.893 63
Social Values & Religion 0.777 0.744 0.809 0.990 37
Health & Science 0.773 0.745 0.801 0.987 47
General Attitudes 0.772 0.742 0.802 0.963 190
Trust & Wellbeing 0.760 0.723 0.797 0.994 25
Economy & Work 0.746 0.710 0.783 0.989 68
Politics & Governance 0.736 0.693 0.779 0.962 40
International Relations & Security 0.735 0.694 0.777 0.986 149
Identity & Demographics 0.732 0.696 0.768 0.984 39

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on opinionsqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

No demographic conditioning data has been published for this vendor yet. The question-type matrix above shows topic-level parity; subgroup rows fill in once Althing-style conditioned runs land.

← Back to leaderboard