Skip to main content

Althing (Haiku 5.5)

3 runs · 1 dataset · 1 model

slug: althing-haiku-5-5

0.815
Best SPS · globalopinionqa

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

globalopinionqa product

althing--claude-haiku-5-5--tdefault--tplstructured--e75ff7a5

0.815 ± 0.029
SPS · 95% CI [0.785, 0.842] · n = 100
Question-type breakdown (6 topics)
Topic SPS p_dist p_rank p_refuse N
Economy & Work low n — suggestive only 0.885 0.880 0.889 0.949 3
Trust & Wellbeing low n — suggestive only 0.841 0.829 0.854 1.000 2
International Relations & Security 0.754 0.777 0.731 0.987 60
Politics & Governance 0.690 0.702 0.679 0.963 30
General Attitudes low n — suggestive only 0.639 0.789 0.489 1.000 4
Health & Science low n — suggestive only 0.463 0.926 0.000 1.000 1

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

globalopinionqa product

althing--claude-haiku-5-5--tdefault--tplstructured--2e811792

0.742 ± 0.034
SPS · 95% CI [0.706, 0.774] · n = 100
Question-type breakdown (6 topics)
Topic SPS p_dist p_rank p_refuse N
Economy & Work low n — suggestive only 0.905 0.841 0.969 0.949 3
Trust & Wellbeing low n — suggestive only 0.723 0.621 0.825 1.000 2
General Attitudes low n — suggestive only 0.722 0.730 0.713 1.000 4
International Relations & Security 0.643 0.646 0.639 0.987 60
Politics & Governance 0.549 0.540 0.559 0.963 30
Health & Science low n — suggestive only 0.295 0.590 0.000 1.000 1

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

globalopinionqa product

althing--claude-haiku-5-5--tdefault--tplcurrent--fed7cc5c

0.695 ± 0.040
SPS · 95% CI [0.654, 0.735] · n = 100
Question-type breakdown (6 topics)
Topic SPS p_dist p_rank p_refuse N
Economy & Work low n — suggestive only 0.958 0.915 1.000 0.926 3
General Attitudes low n — suggestive only 0.776 0.814 0.739 0.925 4
Trust & Wellbeing low n — suggestive only 0.740 0.648 0.832 0.877 2
International Relations & Security 0.670 0.683 0.657 0.837 60
Politics & Governance 0.628 0.670 0.587 0.514 30
Health & Science low n — suggestive only 0.338 0.676 0.000 1.000 1

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

No demographic conditioning data has been published for this vendor yet. The question-type matrix above shows topic-level parity; subgroup rows fill in once Althing-style conditioned runs land.

← Back to leaderboard