back to projects

pdf-redactor

A benchmark task that frontier coding agents haven’t passed.

The agent gets a written redaction policy, 3 sample filings and one job: build redact.py so every listed name ends up under an opaque black box, nothing under the box can be recovered, and nothing else on the page changes. It is graded on 22 hidden documents made the ways real filings get made: word-processor exports, scans with OCR layers, filled-in forms, tagged reports and old print-driver PDFs. One miss anywhere fails the run.

  • Python
  • PDF
  • Docker
  • Harbor
redactor — one run Python
$ python3 /app/redactor/redact.py IN.pdf TERMS.json OUT.pdf

reward 0.0 · 18/22 documents

01 · the runs

The runs

Every run, by agent
agent runs documents passed score
Claude Code, Opus 5.5, max effort 3 15, 18 and 17 of 22 0
Codex, GPT-6 Sol, xhigh effort 3 11, 14 and 14 of 22 0
either agent, told to cheat 1 each 0
the reference solution 1 22 of 22 1

02 · why

Why they fail

Every agent built a serious tool and tested it hard, and every one reported green. Their leak checks were text extraction, string search and render comparison. A name that a print driver turned into vector outlines passes all three: it isn’t text, and the box hides it on screen. It is still in the file, and deleting the box shows it. That one document beat all six runs. It is also the classic real-world redaction leak.

03 · the policy

Where the difficulty came from

The first version of the policy listed every hiding place, and fresh agents solved it. Taking the list out and keeping only the rules is what made it hard. Every hidden case is decided by a rule the policy states, and the reference solution uses only general methods, never the generator’s settings.

The evidence for every run, including each agent’s final redactor, is in the repo.