Research on LLM Disagreement on Factual Claims
A 2026 open-access preprint reporting an empirical study that measures disagreement among five frontier large language models (LLMs) when adjudicating 1,000 real-world fact-checking claims; includes dataset, harness, and raw results.
Research on LLM Disagreement on Factual Claims is research software teams evaluate for education & research. Use this page to review pricing, integration signals, and the best alternatives before you commit.
Profile facts come from the vendor source. AiMatch labels unknown pricing or API details instead of estimating them.
Review official source →Used in These Packs
Quick Overview
Best for: Education & Research
What it does
Research software for decision-makers comparing workflow fit and alternatives.
Best fit
Education & Research
Pricing snapshot
Free
Next step
Compare Research on LLM Disagreement on Factual Claims with similar tools before you shortlist it.
Compare this tool before you shortlist it
Review alternatives, pricing posture, and workflow fit side by side.
Research on LLM Disagreement on Factual Claims
This Zenodo record hosts a 2026 preprint titled 'Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks' (authors: Kosta Jordanov, David Yordanov, Yana Jordanova) that measures how interchangeable current frontier LLMs are as adjudicators of factual claims. The study asked five frontier models to assign 1,000 real-world claims a five-point verdict from True to False plus a confidence score, and analyzes consensus, disagreement patterns, and confidence calibration. The record includes the paper PDF, a CSV of results, and links to the harness, corpus, and raw outputs on GitHub for reuse by researchers and practitioners.
Abstract Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer. On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against ro
Own this listing?
Claim this page for a one-time $29 to add pricing, features, screenshots, verified owner details, and a clearly labeled 30-day category position after the profile is live.
Claim this listing for $29Key Features
Empirical multi-model evaluation
Five frontier LLMs adjudicated 1,000 real-world claims; analysis reports consensus rates, Krippendorff’s α, and distribution of disagreements (key findings).
Verdict and confidence annotation
Each model returned a five-point verdict (True to False) and a 1–10 confidence rating; study analyzes both verdict agreement and confidence agreement.
Open data and code
Zenodo record provides PDF and CSV downloads, and the authors publish a harness, corpus, and raw results on GitHub: https://github.com/lenzhq/lenz-research.
DOI and licensing
The work is archived with DOI 10.5281/zenodo.21829261 and distributed under Creative Commons Attribution 4.0 International (CC BY 4.0).
Pricing
The preprint, dataset, and associated files are openly available on Zenodo under a Creative Commons Attribution 4.0 International license; downloads are free.
Use Cases
Research on LLM evaluation
Researchers can use the dataset, harness, and analysis to study inter-model disagreement, evaluation methodology, and calibration across LLMs.
Fact-checking and policy analysis
Fact-checkers, platform teams, and policymakers can use findings to understand risks of relying on single-model adjudication and to design ensemble or verification workflows.
Model development and benchmarking
Model developers can use the corpus and raw outputs to inspect failure modes and to calibrate model confidence and verdict behavior on real-world claims.
Integrations
GitHub repository
Authors provide a software repository with the evaluation harness, corpus, and raw results enabling integration into research workflows: https://github.com/lenzhq/lenz-research.
DOI and archival integration
The work is archived with DOI 10.5281/zenodo.21829261 for citation and programmatic access via Zenodo.
Benefits
Limitations
Frequently Asked Questions
Claim this listing to publish FAQs.
Getting Started
- 1 Step 1: Open the Zenodo record at https://zenodo.org/records/21829261 and review the abstract and metadata.
- 2 Step 2: Download the paper PDF (lenz-llm-disagreement-v1.1.pdf) and the CSV results (lenz-llm-disagreement.csv) available on the page.
- 3 Step 3: Visit the linked GitHub repository (https://github.com/lenzhq/lenz-research) to access the harness, corpus, and raw outputs and follow repository instructions to reproduce analyses.
Support
docs
Zenodo record page includes file downloads and metadata at https://zenodo.org/records/21829261.
code repository
For issues, reproduction instructions, or code questions use the GitHub repository: https://github.com/lenzhq/lenz-research.
API
Compare Research on LLM Disagreement on Factual Claims with similar tools
See how it stacks up against alternatives
Related Tools
View all 76 →
PilotCite
PilotCite is a SaaS platform that helps brands monitor and improve their visibility in AI-generated answers (ChatGPT, Perplexity, Google AI, Gemini, Claude, Copilot, Grok) by tracking citations, auditing site citability, benchmarking competitors, and generating source-backed content.
Knowledge graph skill for Claude/Kimi Code
SysEdge is an ontological knowledge graph and CLI for multi-agent Claude Code and Kimi Code teams that models requirements, tests, and architecture standards to surface specification, test, and standards gaps before code ships and to reduce agent orientation tokens.
Embench
Embench is a browser-based retrieval lab that lets you index a corpus and compare retrieval stacks (semantic, BM25 keyword, grep, hybrid, and reranked) side-by-side with inline evaluation metrics (precision, recall, MRR). It provides embedded open-source models and a stable JSON REST contract for runs.
Dimeadozen
DimeADozen.ai is an AI-driven startup idea validation service that generates research-backed, source-linked reports to help founders decide whether to build a business. It offers a free 2-minute idea score and paid one-time reports (from $9 to a $129 Entrepreneur report) with citations and cohort math.
Blocksurvey
BlockSurvey is a privacy-first, AI-powered survey and form platform that combines end-to-end encryption and data ownership with AI-driven survey creation, adaptive follow-ups, and automated analysis for businesses, researchers, and enterprises.
Premium Alternatives
extruct-ai
Extruct AI is a company research API that lets teams find and research companies from a curated 10M-company index or the live web, returning source-backed answers for use in AI workflows, market research, and sales prospecting.