LongListBench: A Benchmark for Long-List Entity Extraction from Complex Business PDFs

Kay.ai · July 2026 · doi.org/10.57967/hf/9768

LongListBench title page

Abstract

Business PDFs can contain hundreds or thousands of target records, while most document benchmarks emphasize short forms or modest tables. LongListBench measures complete document-to-list extraction across 14 auditable layout, record-assembly, scale, and OCR stressors. It contains 32 synthetic PDFs, 29,599 target records, JSON ground truth, per-document stressor metadata, and OCR transcripts. Four repository-isolated coding-agent configurations are evaluated on the same transcripts. GPT-5.6-Sol and Claude Fable 5 recover 97.9% and 95.1% of records exactly, but reproduce only 8 and 9 of 32 complete document lists. Exact recall is 93.8% and 84.0% on structural challenges versus 99.5% for both on scale controls; secondary field-pair F1 is 99.4% and 96.8%. Heterogeneous policy records separate the systems most sharply, at 73.3% and 14.6% exact recall. All visible content is synthetic; private documents informed structure only.

Links

Cite

@techreport{fedoruk2026longlistbench,
  title       = {LongListBench: A Benchmark for Long-List Entity Extraction from Complex Business PDFs},
  author      = {Fedoruk, Anton and Shchoholiev, Serhii and Mehta, Akhil and Rohra, Vishal},
  institution = {Kay.ai},
  year        = {2026},
  month       = {7},
  doi         = {10.57967/hf/9768},
  url         = {https://shchoholiev.com/research/longlistbench}
}

← All research