Redactle LLM Leaderboard

Redactle LLM Leaderboard

A public benchmark dashboard that evaluates how well large language models (LLMs) solve Redactle puzzles by running standardized evaluations and publishing ranked results, costs, and performance metrics.

Redactle LLM Leaderboard is research software teams evaluate for research. Use this page to review pricing, integration signals, and the best alternatives before you commit.

Contact for pricing
#91 in Research (91 tools)
Just launched
Data reviewed Sep 3, 2026

Profile facts come from the vendor source. AiMatch labels unknown pricing or API details instead of estimating them.

Review official source →

Quick Overview

Best for: Research

What it does

Research software for decision-makers comparing workflow fit and alternatives.

Best fit

Research

Pricing snapshot

Contact for pricing

Next step

Compare Redactle LLM Leaderboard with similar tools before you shortlist it.

Compare this tool before you shortlist it

Review alternatives, pricing posture, and workflow fit side by side.

Redactle LLM Leaderboard

The Redactle LLM Leaderboard is a public dashboard on redactle.net that measures how different language models perform at solving Redactle puzzles (redacted Wikipedia articles). The page runs standardized evaluations under multiple configurations (for example: 500 words no-hint, 100 words with 3 hints, and a cheating-enabled setting that can query Wikipedia’s API) and reports per-model metrics such as solve rate, score, cost per run, and runtime. The leaderboard is intended for comparing model capability, cost, and speed under the same puzzle-solving rules; it also links to evaluation code to aid reproducibility and further analysis.

The page deliberately omits article titles and per-article excerpts to avoid spoiling puzzles; instead it publishes aggregated results, method notes, and trade-off visualizations so readers can understand relative model performance without revealing puzzle content.

A compact comparison of how language models solve Redactle under distinct evaluation configurations.

Own this listing?

Claim this page for a one-time $29 to add pricing, features, screenshots, verified owner details, and a clearly labeled 30-day category position after the profile is live.

Claim this listing for $29

Key Features

Standardized multi-configuration evaluations

Runs models under several fixed configurations (e.g., 500 words no-hint, 100 words with up to 3 hints, 100 words with research allowed) and reports results per configuration.

Per-model metrics and ranking

Displays solve counts, attempts, solve rate, score, cost per run and run time for many LLM configurations (examples include Gemini, Grok, GPT-5.6, GLM, Claude Sonnet).

Method notes and scoring rules

Explains how scores are computed (accepted guesses above title word-count par, hint penalties, attempt cutoffs like 30 accepted guesses in the 500-word evaluation).

Trade-off visualizations

Includes cost-versus-score and speed-versus-score visualizations and a quadrant view showing better/worse trade-offs across models.

Reproducible evaluation code

Links to evaluation code on GitHub for inspection and reproducibility.

Published result snapshots

Published leaderboard snapshots with aggregate counts and a recent publication date (e.g., Sep 2, 2026).

Pricing

Current pricing details are not available from the vendor source.

Use Cases

Benchmarking LLMs

Compare model accuracy, cost, and latency on a focused task (solving redacted encyclopedia articles) under consistent rules.

Model selection and cost analysis

Choose models or configurations that balance solve rate, runtime, and monetary cost per run for similar generative tasks.

Research and reproducibility

Use the linked evaluation code and method notes to reproduce runs or adapt the benchmark for further experiments.

Curiosity and public reporting

Provide players and the public with an entertaining, empirical view of current LLM capabilities on a puzzle task.

Integrations

GitHub (evaluation code)

Evaluation code is available on GitHub for inspection and reproduction of runs.

Wikipedia API (optional in cheating configuration)

A 'cheating allowed' evaluation configuration lets models research through Wikipedia’s API as part of the test.

Benefits

Provides an apples-to-apples comparison of many LLMs on the same puzzle-solving task
Reports cost and runtime along with accuracy metrics to inform practical choices
Includes method notes and links to evaluation code to aid reproducibility
Avoids spoiling puzzles by omitting per-article titles and excerpts

Limitations

Article titles, excerpts, and per-article outcomes are deliberately omitted to avoid spoiling puzzles, so the leaderboard only shows aggregate metrics.
Incomplete or incompatible model configurations remain visible but are not ranked; some models may be marked 'Not scored: tool incompatible'.
Scores use specific rules (e.g., hint penalties, guess cutoffs) which constrain direct comparison to other benchmarks or tasks.

Frequently Asked Questions

No verified FAQs are available.

Getting Started

  1. 1 Visit the LLM leaderboard page on redactle.net/llm-leaderboard to view published results and news.
  2. 2 Read the method notes (scoring rules and configuration descriptions) on the page to understand how evaluations were run.
  3. 3 Follow the link to the evaluation code on GitHub to inspect or reproduce the benchmark.
  4. 4 Compare models by configuration, cost, and runtime; consult the trade-off visualizations to select candidates for further testing.

Support

docs

Method notes, scoring rules, and trade-off explanations available on the LLM leaderboard page.

github

Evaluation code linked from the leaderboard page for issues, reproduction, and contributions.

support link / donations

Support, Ko-fi donations and an in-site Support link are mentioned for feedback and ad-free access for supporters.

discord

Contact via Discord for feedback (the page invites contacting the author on Discord).

API

Available: No

Compare Redactle LLM Leaderboard with similar tools

See how it stacks up against alternatives

Related Tools

View all 91 →
Contact for pricing
ThoughtDAG

ThoughtDAG

ThoughtDAG is an open-source, desktop-first application that makes LLM context visible, editable, and reproducible by representing context as an editable directed acyclic graph (wires = context) and letting users preview and control exactly what the model receives.

Research
Top source High-growth
Free
PilotCite

PilotCite

PilotCite is a SaaS platform that helps brands monitor and improve their visibility in AI-generated answers (ChatGPT, Perplexity, Google AI, Gemini, Claude, Copilot, Grok) by tracking citations, auditing site citability, benchmarking competitors, and generating source-backed content.

Research
High-growth
Freemium
Knowledge graph skill for Claude/Kimi Code

Knowledge graph skill for Claude/Kimi Code

SysEdge is an ontological knowledge graph and CLI for multi-agent Claude Code and Kimi Code teams that models requirements, tests, and architecture standards to surface specification, test, and standards gaps before code ships and to reduce agent orientation tokens.

Research
High-growth
Freemium
Korvo

Korvo

Korvo is a local-first private research and decision workspace for macOS that organizes files, generates and verifies evidence-backed analyses using connected models (cloud or local), preserves decision history, and supports a two-model critique workflow.

Research
High-growth
Free
Embench

Embench

Embench is a browser-based retrieval lab that lets you index a corpus and compare retrieval stacks (semantic, BM25 keyword, grep, hybrid, and reranked) side-by-side with inline evaluation metrics (precision, recall, MRR). It provides embedded open-source models and a stable JSON REST contract for runs.

Research
High-growth
Free
Research on LLM Disagreement on Factual Claims

Research on LLM Disagreement on Factual Claims

A 2026 open-access preprint reporting an empirical study that measures disagreement among five frontier large language models (LLMs) when adjudicating 1,000 real-world fact-checking claims; includes dataset, harness, and raw results.

Research
High-growth
Free
Knowledgie

Knowledgie

Knowledgie is an AI-powered research assistant that delivers fact-based answers and summaries backed by millions of research papers, with features for searching by question, chatting with PDFs, building a knowledge base, automated citation generation, and cloud PDF storage.

Research
Freemium
Userintuition

Userintuition

User Intuition is an AI-moderated customer research platform that runs voice, video, and chat interviews at scale, using laddering (5–7 levels) to deliver evidence-backed qualitative findings in 24 hours from a 4M+ verified panel or your own customers.

Research

Premium Alternatives

Paid
Bearly

Bearly

Bearly is a private AI workspace that provides encrypted, cross-platform tools for research, coding, content creation, team collaboration, and enterprise controls, with support for multiple large language models and developer tools.

Research
Paid
monkt

monkt

Monkt is a document processing platform that converts PDFs, Word, PowerPoint, Excel, CSV, images and web pages into AI-ready Markdown or structured JSON, with features for batch processing, custom JSON schemas, image understanding, and REST API integration.

Research
Enterprise-ready High-growth
Paid
extruct-ai

extruct-ai

Extruct AI is a company research API that lets teams find and research companies from a curated 10M-company index or the live web, returning source-backed answers for use in AI workflows, market research, and sales prospecting.

Research
Enterprise-ready

Explore Related Categories

Explore by Outcome