Skip to main content

Available Datasets

All files are served from https://data.jmail.world/v1/.

Emails

The primary dataset. Contains all released emails from the Epstein archive. emails.parquet — Full dataset with body text (content_markdown), sender, recipients, subject, dates, and metadata. emails-slim.parquet — Same emails but without body text columns. Much smaller download, ideal for network analysis, sender/recipient graphs, and timeline visualizations.

Key Columns (slim)

Additional Columns (full)

Documents

Metadata for all documents in the archive (DOJ releases, House Oversight, court records).

Document Full-Text Shards

Full extracted text is too large for a single file. Use the sharded files: The Python client handles shard concatenation automatically via client.documents(include_text=True).

Photos

Photo metadata from government releases with AI-generated descriptions.

People

People identified via AWS Rekognition facial recognition.

Photo Faces

Bounding boxes linking detected faces in photos to identified people.

iMessage Conversations

Metadata for iMessage conversations recovered from the archive.

iMessage Messages

Individual iMessage text messages with sender info and timestamps.

Star Counts

Crowd-sourced star/interest counts from jmail.world users.

Release Batches

Metadata about each release batch.