HieraticBench

HieraticBench lets AI researchers and model builders measure how well models identify and read ancient Egyptian hieratic. Its dataset, evaluation harness, and leaderboard show performance on real documents, individual signs, and sealed sentences.

One of 126 tools in Research

HieraticBench screenshot

Who it's for

  • AI researchers comparing models on ancient-script recognition who can run the evaluation harness with their own key.
  • Model developers building or evaluating systems intended to read hieratic documents or individual signs.
  • Egyptologists who can contribute sealed sentences or sign annotations to the benchmark.

Not the right fit if…

  • Teams needing scored translation or transliteration results for the sealed sentence: those tasks are not scored because the benchmark does not store its answer.

How it fits your workflow

Researchers use the harness to evaluate a model on benchmark images. The tasks include identifying the writing system, identifying individual signs using Gardiner codes, transliterating, and translating.

The benchmark reports model results on sealed sentences, real ancient documents, and individual signs where those evaluations have been run. Researchers can compare available results on the public leaderboard.

Egyptologists can contribute new sealed sentences or annotations on published papyri to help expand the dataset.

Pricing

Free: The page says the professor's sentence is free to use for evaluating models and that the code is MIT licensed. It does not list pricing plans.

The vendor doesn't publish prices. Check with HieraticBench directly.

Prices checked on Oct 7, 2026 from the vendor's site. They can change; confirm before you buy.

Key features

Four evaluation tasks
The benchmark covers script identification, matching signs to Gardiner sign-list codes, Egyptological transliteration, and English translation. The page reports scores for script identification and sign recognition; transliteration and translation are not scored on the sealed sentence.
Public leaderboard
Model results are published in a leaderboard, including performance on a sealed sentence, real hieratic documents, and individual signs where those evaluations have been run.
Benchmark dataset
The dataset contains 268 items: 87 real hieratic documents, 29 controls in other Egyptian scripts, 150 single signs, and 2 sealed sentences. Public images are credited and openly licensed.
Evaluation harness
AI researchers can run the harness on a model using their own key. The page says Claude evaluations use Anthropic's API and other models use OpenRouter.
Sealed evaluation examples
The answer to the professor's sentence is not published or stored in the benchmark, so models cannot pass that test by recalling a published translation.

Works with

  • Anthropic API
  • OpenRouter

Limitations to know

  • Transliteration and translation of the sealed sentence are not scored because the benchmark does not store an answer for it.
  • Leaderboard coverage varies by task: several models are marked as not run for document or single-sign evaluation.

Getting started

Your own key to run the harness on a model.

  1. Browse the dataset and leaderboard on the HieraticBench site.
  2. Run the harness on your model using your own key.
  3. Review the resulting scores alongside the benchmark's published results.