LongListBench: A Benchmark for Long-List Entity Extraction from Complex Business PDFs
Kay.ai · July 2026 · doi.org/10.57967/hf/9768
Abstract
Business PDFs can contain hundreds or thousands of target records, while most document benchmarks emphasize short forms or modest tables. LongListBench measures complete document-to-list extraction across 14 auditable layout, record-assembly, scale, and OCR stressors. It contains 32 synthetic PDFs, 29,599 target records, JSON ground truth, per-document stressor metadata, and OCR transcripts. Four repository-isolated coding-agent configurations are evaluated on the same transcripts. GPT-5.6-Sol and Claude Fable 5 recover 97.9% and 95.1% of records exactly, but reproduce only 8 and 9 of 32 complete document lists. Exact recall is 93.8% and 84.0% on structural challenges versus 99.5% for both on scale controls; secondary field-pair F1 is 99.4% and 96.8%. Heterogeneous policy records separate the systems most sharply, at 73.3% and 14.6% exact recall. All visible content is synthetic; private documents informed structure only.
Links
Cite
@techreport{fedoruk2026longlistbench,
title = {LongListBench: A Benchmark for Long-List Entity Extraction from Complex Business PDFs},
author = {Fedoruk, Anton and Shchoholiev, Serhii and Mehta, Akhil and Rohra, Vishal},
institution = {Kay.ai},
year = {2026},
month = {7},
doi = {10.57967/hf/9768},
url = {https://shchoholiev.com/research/longlistbench}
}