Skip to main content

Althing (Gemini Flash Lite)

3 runs · 3 datasets · 1 model

slug: althing-gemini-flash-lite

0.821
Best SPS · subpop

Disaggregated subgroup scorecard. Each card below is one published run for this vendor; expand the question-type and demographic-subgroup sections to see the matrix beneath the headline SPS. Where coverage permits, 95% CI bands accompany the point estimate.

What each cell measures

Each cell compares this vendor's synthetic responses against real human respondents in that demographic subgroup — published survey ground truth, not a model's guess about the subgroup. Higher = closer to how that real subgroup actually answered.

Demographic conditioning here is not stereotyping: every conditioned score is checked against what real subgroup members said, not against assumptions about them — and where the vendor's output diverges from the real subgroup, the score drops.

  • globalopinionqa — ground truth: Durmus et al. 2023, Anthropic — llm_global_opinions
  • opinionsqa — ground truth: Santurkar et al., ICML 2023 — Whose Opinions Do LLMs Reflect? (derived from Pew American Trends Panel)
  • subpop — ground truth: Suh et al., ACL 2025 — SubPOP: Subpopulation-Level Opinion Prediction

Low-n cells: cells with n < 30 are shown muted and tagged low n — suggestive only instead of color-graded; they are reported for transparency, not as findings.

CIs: "no CI — single run" marks point estimates from a single run with no confidence interval yet; treat the uncertainty as unknown, never zero.

Columns: Score = distributional parity (p_dist) restricted to the subgroup · p_cond = conditioning strength vs the unconditioned baseline · N = questions answered under that conditioning · Cov. = coverage dot derived from N (green = high, ≥100 · amber = medium, 50–99 · red = low, <50). Topic tables: SPS = Survey Parity Score for the topic · p_dist = 1 − mean(JSD) · p_rank = (1 + mean(τ)) / 2 · p_refuse = 1 − mean(|R_model − R_human|).

Multiple comparisons: with this many subgroup cells, a few extreme cells are expected by chance alone — read patterns across a dimension, not single cells.

globalopinionqa product

synthpanel--gemini-2.5-flash-lite--tdefault--tplcurrent--ca4dfebb

0.762 ± 0.032
SPS · 95% CI [0.729, 0.792] · n = 100
Question-type breakdown (7 topics)
Topic SPS p_dist p_rank p_refuse N
Technology & Digital Life low n — suggestive only 0.901 0.849 0.952 0.968 3
General Attitudes 0.682 0.744 0.621 1.000 11
Politics & Governance 0.677 0.725 0.630 0.946 22
International Relations & Security 0.672 0.690 0.655 0.979 50
Health & Science low n — suggestive only 0.654 0.881 0.427 1.000 2
Economy & Work low n — suggestive only 0.539 0.606 0.472 0.967 5
Trust & Wellbeing low n — suggestive only 0.403 0.392 0.414 0.980 7

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on globalopinionqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

opinionsqa product

synthpanel--gemini-2.5-flash-lite--tdefault--tplcurrent--5a08bef0

0.816 ± 0.009
SPS · 95% CI [0.807, 0.825] · n = 684
Question-type breakdown (10 topics)
Topic SPS p_dist p_rank p_refuse N
Health & Science 0.797 0.788 0.806 0.945 47
Media & Information 0.777 0.766 0.789 0.989 63
General Attitudes 0.774 0.773 0.775 0.947 190
Social Values & Religion 0.774 0.764 0.784 0.959 37
International Relations & Security 0.760 0.745 0.776 0.950 149
Economy & Work 0.751 0.742 0.759 0.869 68
Trust & Wellbeing 0.745 0.761 0.730 0.953 25
Technology & Digital Life 0.730 0.697 0.764 0.835 26
Identity & Demographics 0.684 0.701 0.668 0.880 39
Politics & Governance 0.680 0.648 0.713 0.892 40

Not yet measured

This vendor has no demographic-conditioned runs for age, geography, education, or any other subgroup dimension on opinionsqa. No cells are fabricated — scores appear here only when a conditioned run actually measured them.

How to submit a demographic-conditioned run →

subpop product

synthpanel--gemini-2.5-flash-lite--tdefault--tplcurrent--650ec804

0.821 ± 0.013
SPS · 95% CI [0.807, 0.834] · n = 200
Question-type breakdown (9 topics)
Topic SPS p_dist p_rank p_refuse N
Trust & Wellbeing low n — suggestive only 0.925 0.896 0.954 0.988 2
Identity & Demographics low n — suggestive only 0.846 0.838 0.854 0.995 1
Health & Science low n — suggestive only 0.828 0.860 0.796 0.994 5
General Attitudes 0.782 0.754 0.811 0.913 37
Economy & Work 0.751 0.714 0.788 0.991 17
Social Values & Religion 0.750 0.697 0.802 0.987 36
International Relations & Security 0.741 0.690 0.792 0.991 33
Technology & Digital Life 0.709 0.681 0.737 0.993 47
Politics & Governance 0.701 0.662 0.740 0.983 22

Topics with N < 10 are muted and tagged "low n — suggestive only": too few questions for a stable 3-decimal score.

Demographic subgroup scorecard (10 cells · 5 dimensions)
Dimension Subgroup Score 95% CI p_cond N Cov.
Geography (US Census region) Northeast 0.740 no CI — single run 0.055 100
South 0.716 no CI — single run 0.051 100
Education College graduate/some postgrad 0.752 no CI — single run 0.049 100
Less than high school 0.691 no CI — single run 0.066 100
Income $100,000 or more 0.748 no CI — single run 0.060 100
Less than $30,000 0.708 no CI — single run 0.065 100
Political party Democrat 0.712 no CI — single run 0.050 100
Republican 0.703 no CI — single run 0.074 100
Sex Female 0.723 no CI — single run 0.050 100
Male 0.748 no CI — single run 0.054 100
Not yet measured: this vendor has no demographic-conditioned runs for Age. Submit a conditioned run to fill in the missing dimension.
← Back to leaderboard