How AI models perceive presidential suitability
A repeated-measures study of subjective LLM judgments, model disagreement, prompt sensitivity, and ranking instability. It does not measure objective candidate quality and is not polling, an official nomination, or an endorsement.
Study snapshot
Experimental design
Five model configurations receive two semantically equivalent prompt variants five times each. Candidate order is deterministically randomized for every run. Models score six criteria on an ordinal 1–10 scale and separately report factual recall, positive reasons, and reservations.
From candidate discovery to human decision
This case study separates five questions that a single ranking would hide. The v4 candidate set was inherited from an exploratory scorecard, checked for breadth, and frozen before scoring. It was not produced by an exhaustive or preregistered search, so every comparison is conditional on this illustrative set.
The prospective study will not preselect a top ten. Its candidate count must come from a preregistered, score-independent sampling frame, eligibility review, coverage, and omission audits—not scores or weights. Open-ended nomination was empirically rejected: six valid five-lane rounds produced 294 search leads and the sixth round still added 21.5% new names. Discovery is therefore blocked until authoritative registries or objective public-role criteria define the frame. Every frozen eligible frame member receives baseline measurement. Top-k membership is computed only afterward and separately for each weight profile; weight-robust, Pareto-undominated, and rank-uncertain candidates remain visible.
Nomination is stage one, not a candidate decision. Agents may emit only source-verification leads tagged with preregistered entry routes, evidence queries, temporal basis, and lane/model provenance. Public salience and current relevance are optional timestamped routes measured from independent sources; model familiarity is not evidence. A lead becomes a frame-qualified nominee only after retrieval, an eligibility-cleared nominee only after human/legal review, and a candidate only after explicit human freeze.
Live nomination smoke: five blind lanes returned 21 merged generated leads and passed every schema and governance gate. Lackó Adrienn did not appear in the blind output. When probed separately by name, zero of five lanes emitted a lead because none could state a defensible allowed-route claim without external source retrieval. This is a test of source-free nomination behavior, not an eligibility or suitability conclusion.
Agentic discovery rehearsal: five independent lanes produced 59 raw nominations, merged to 41 provenance-preserving shadow records, and represented every declared coverage dimension. Injected tests confirmed duplicate merging, ineligible quarantine, missing-coverage and omission escalation, saturation logic, and freeze-tamper detection. The rehearsal passed, but its names remain unverified search leads: it created no eligibility decision, score, binding candidate list, or human freeze.
Protocol extension—not part of the displayed v4 results: the first 30-call pilot across Sol, Terra, and Claude Opus failed its unchanged 90% valid-call gate because Claude returned four policy refusals (26/30 valid); all other operational gates passed. The immutable v2 pilot changed only the cross-lineage control to Gemini 3.1 Pro and passed: 30/30 valid packets, 91.4% synthetic-control agreement, 100% pairwise order consistency, complete abstention fields, exact reproduction, and no blocking audit flag. This unlocks only the protocol gate for the 106-call experiment. Prospective candidate discovery and human list-freeze, methodological review, and legal approval remain mandatory. Claude remains in the full design as a potentially missing-not-at-random rater.
Judgment distributions
Weights are an external value choice. Adjusting them recomputes totals and rank frequencies from retained run packets; it does not turn model judgments into facts.
Preset weightings · Choose a quick starting point if you do not want to set every slider individually.
One gpt-5.6-sol model configuration, two prompt variants, five repetitions per prompt. This is not a multi-model consensus.
Criterion comparison
Variance and sensitivity
Factual recall and model-risk audit
The factual audit compares recalled claims with a limited frozen reference packet. “Not covered” is not treated as hallucination. This is not a complete truth benchmark; human fact checking remains a production blocker.
Reproducibility and release
Raw responses, structured packets, model and prompt identifiers, randomized order, failed and excluded runs, audits, and deterministic aggregation are retained. Chain-of-thought is neither requested nor published.
Release: presidential-selection-v5-20260802T232312+0200
Hash: sha256:fedc10b1c6f52ef4c95378d8814c79ec98c0fdd4d4cf1c0663ef98306bc2b941
Study packet: sha256:0eb50ff62ea6c218f84fee540f6e87bb1fb46a736d30c109bc19bd044c13ad9f
Aggregation: sha256:5ef06f6074518c04bbdb536c941095453b76682667848d8db80aeea97d7321e6
Reproduction match: True
Production blockers
- Human factual and methodological approval is not recorded.
- Hungarian legal and data-protection review is not recorded.
- The study measures LLM perceptions, not objective candidate suitability or public opinion.
- The frozen factual reference packet is incomplete and cannot certify all recalled facts.
- Unresolved high audit flag: single_model_pseudoensemble
- Unresolved high audit flag: unequal_candidate_exposure
- Unresolved high audit flag: degenerate_rank_probabilities
- Unresolved high audit flag: order_effect_uncontrolled