Epstein Project

Methodology

About this archive

What is here, where it came from, and how far to trust it

What this is

This site indexes public records from official releases of the Jeffrey Epstein case — Department of Justice productions, court filings, and House Oversight Committee releases — and preserves a link back to the source material for every document.

How the text is produced

Where a document carries a machine-readable text layer, that text is used as released. Image-only pages are read with optical character recognition (OCR). Every document page states which applies and, where text was extracted, a confidence figure.

Where it can be wrong

OCR misreads names, dates, handwriting, and degraded or photocopied scans. Some documents in this archive were scanned, printed, faxed, and scanned again before release, and the text reflects that. Government redactions also appear in the text as noise where a black bar was read as characters.

A confidence figure describes how certain the character recognition was, not whether the reading is correct. Check anything you intend to rely on against the original page, which is linked from every document.

Search coverage is not complete

Search matches extracted text, so a document with no readable text cannot be found by searching its contents — only by its Bates number or filename. Work to extract text from the remaining scans is ongoing, and a document's page states its current status.

Names and organisations

Names mentioned in the text are indexed automatically. This is imperfect: the same person may appear under several spellings, OCR fragments produce entries that are not really names, and similarly named people are not distinguished. Treat the index as a way to find documents, not as a finding of fact about anyone.

Corrections

Errors, missing files, and misattributions can be reported through the feedback link in the footer. Methodology last reviewed July 2026.

Open in the archive