Research on LLM Disagreement on Factual Claims

Research on LLM Disagreement on Factual Claims

A 2026 open-access preprint reporting an empirical study that measures disagreement among five frontier large language models (LLMs) when adjudicating 1,000 real-world fact-checking claims; includes dataset, harness, and raw results.

Research on LLM Disagreement on Factual Claims is research software teams evaluate for education & research. Use this page to review pricing, integration signals, and the best alternatives before you commit.

Free
#76 in Research (76 tools)
Just launched
Data reviewed Aug 13, 2026

Profile facts come from the vendor source. AiMatch labels unknown pricing or API details instead of estimating them.

Review official source →

Quick Overview

Best for: Education & Research

What it does

Research software for decision-makers comparing workflow fit and alternatives.

Best fit

Education & Research

Pricing snapshot

Free

Next step

Compare Research on LLM Disagreement on Factual Claims with similar tools before you shortlist it.

Compare this tool before you shortlist it

Review alternatives, pricing posture, and workflow fit side by side.

Research on LLM Disagreement on Factual Claims

This Zenodo record hosts a 2026 preprint titled 'Beyond Benchmarks: Disagreement Among Frontier LLMs on Real-World Fact-Checks' (authors: Kosta Jordanov, David Yordanov, Yana Jordanova) that measures how interchangeable current frontier LLMs are as adjudicators of factual claims. The study asked five frontier models to assign 1,000 real-world claims a five-point verdict from True to False plus a confidence score, and analyzes consensus, disagreement patterns, and confidence calibration. The record includes the paper PDF, a CSV of results, and links to the harness, corpus, and raw outputs on GitHub for reuse by researchers and practitioners.

Abstract Frontier large language models achieve comparable accuracy on public benchmarks, a parity often read as evidence that they are interchangeable as assessors of factual claims. We test that assumption directly, on claims not drawn from any public benchmark. Five frontier models were each asked to adjudicate 1,000 real-world claims submitted by users to a fact-checking platform. They were tasked with assigning every claim a verdict on a five-point scale from True to False, as well as reporting their confidence in each answer. On the 997 claims where all five models returned a usable verdict, they fail to reach consensus on 63%, and on 23% the two furthest-apart verdicts differ by at least two verdict categories. The ordinal Krippendorff’s α of 0.77 reflects structured but far from interchangeable judgement. We find that disagreement concentrates in the intermediate verdicts, where claims resist clean adjudication: the definitive poles are unanimous about half the time, against ro

Own this listing?

Claim this page for a one-time $29 to add pricing, features, screenshots, verified owner details, and a clearly labeled 30-day category position after the profile is live.

Claim this listing for $29

Key Features

Empirical multi-model evaluation

Five frontier LLMs adjudicated 1,000 real-world claims; analysis reports consensus rates, Krippendorff’s α, and distribution of disagreements (key findings).

Verdict and confidence annotation

Each model returned a five-point verdict (True to False) and a 1–10 confidence rating; study analyzes both verdict agreement and confidence agreement.

Open data and code

Zenodo record provides PDF and CSV downloads, and the authors publish a harness, corpus, and raw results on GitHub: https://github.com/lenzhq/lenz-research.

DOI and licensing

The work is archived with DOI 10.5281/zenodo.21829261 and distributed under Creative Commons Attribution 4.0 International (CC BY 4.0).

Pricing

Free Tier Available

The preprint, dataset, and associated files are openly available on Zenodo under a Creative Commons Attribution 4.0 International license; downloads are free.

Use Cases

Research on LLM evaluation

Researchers can use the dataset, harness, and analysis to study inter-model disagreement, evaluation methodology, and calibration across LLMs.

Fact-checking and policy analysis

Fact-checkers, platform teams, and policymakers can use findings to understand risks of relying on single-model adjudication and to design ensemble or verification workflows.

Model development and benchmarking

Model developers can use the corpus and raw outputs to inspect failure modes and to calibrate model confidence and verdict behavior on real-world claims.

Integrations

GitHub repository

Authors provide a software repository with the evaluation harness, corpus, and raw results enabling integration into research workflows: https://github.com/lenzhq/lenz-research.

DOI and archival integration

The work is archived with DOI 10.5281/zenodo.21829261 for citation and programmatic access via Zenodo.

Benefits

Quantifies how often frontier LLMs disagree on real-world factual claims, revealing that models are not interchangeable assessors.
Provides open artifacts (paper, CSV, GitHub repository) so others can reproduce, extend, or audit the evaluation.
Identifies patterns of disagreement (concentration in intermediate verdicts) and shows model confidence is not a reliable proxy for panel agreement.

Limitations

Study scope: analysis is based on the 997 claims where all five models returned usable verdicts (not the full 1,000 for some analyses).
Disagreement concentrated in intermediate verdicts — definitive poles reach unanimity far more often, limiting generalization to borderline claims.
Findings reflect the specific panel of five frontier models and the user-submitted claims dataset; results may differ with other models or claim sets.

Frequently Asked Questions

Claim this listing to publish FAQs.

Getting Started

  1. 1 Step 1: Open the Zenodo record at https://zenodo.org/records/21829261 and review the abstract and metadata.
  2. 2 Step 2: Download the paper PDF (lenz-llm-disagreement-v1.1.pdf) and the CSV results (lenz-llm-disagreement.csv) available on the page.
  3. 3 Step 3: Visit the linked GitHub repository (https://github.com/lenzhq/lenz-research) to access the harness, corpus, and raw outputs and follow repository instructions to reproduce analyses.

Support

docs

Zenodo record page includes file downloads and metadata at https://zenodo.org/records/21829261.

code repository

For issues, reproduction instructions, or code questions use the GitHub repository: https://github.com/lenzhq/lenz-research.

API

Available: No

Compare Research on LLM Disagreement on Factual Claims with similar tools

See how it stacks up against alternatives

Related Tools

View all 76 →
Free
PilotCite

PilotCite

PilotCite is a SaaS platform that helps brands monitor and improve their visibility in AI-generated answers (ChatGPT, Perplexity, Google AI, Gemini, Claude, Copilot, Grok) by tracking citations, auditing site citability, benchmarking competitors, and generating source-backed content.

Research
High-growth
Freemium
Knowledge graph skill for Claude/Kimi Code

Knowledge graph skill for Claude/Kimi Code

SysEdge is an ontological knowledge graph and CLI for multi-agent Claude Code and Kimi Code teams that models requirements, tests, and architecture standards to surface specification, test, and standards gaps before code ships and to reduce agent orientation tokens.

Research
High-growth
Freemium
Korvo

Korvo

Korvo is a local-first private research and decision workspace for macOS that organizes files, generates and verifies evidence-backed analyses using connected models (cloud or local), preserves decision history, and supports a two-model critique workflow.

Research
High-growth
Free
Embench

Embench

Embench is a browser-based retrieval lab that lets you index a corpus and compare retrieval stacks (semantic, BM25 keyword, grep, hybrid, and reranked) side-by-side with inline evaluation metrics (precision, recall, MRR). It provides embedded open-source models and a stable JSON REST contract for runs.

Research
High-growth
Freemium
Onecliq

Onecliq

OneCliq is a qualitative insights platform that uses AI to analyze public conversations in real time, turning social and video data into emotion-first, actionable recommendations for strategists, creatives, and research teams.

Research
Freemium
Dimeadozen

Dimeadozen

DimeADozen.ai is an AI-driven startup idea validation service that generates research-backed, source-linked reports to help founders decide whether to build a business. It offers a free 2-minute idea score and paid one-time reports (from $9 to a $129 Entrepreneur report) with citations and cohort math.

Research
Paid
Bearly

Bearly

Bearly is a private AI workspace that provides encrypted, cross-platform tools for research, coding, content creation, team collaboration, and enterprise controls, with support for multiple large language models and developer tools.

Research
High-growth
Freemium
Blocksurvey

Blocksurvey

BlockSurvey is a privacy-first, AI-powered survey and form platform that combines end-to-end encryption and data ownership with AI-driven survey creation, adaptive follow-ups, and automated analysis for businesses, researchers, and enterprises.

Research

Premium Alternatives

Paid
Bearly

Bearly

Bearly is a private AI workspace that provides encrypted, cross-platform tools for research, coding, content creation, team collaboration, and enterprise controls, with support for multiple large language models and developer tools.

Research
High-growth
Paid
extruct-ai

extruct-ai

Extruct AI is a company research API that lets teams find and research companies from a curated 10M-company index or the live web, returning source-backed answers for use in AI workflows, market research, and sales prospecting.

Research
Enterprise-ready

Explore Related Categories

Explore by Outcome