How this archive was built
Method, provenance and limitsFive PDF packages, 1,847 pages in all, were parsed into 1,889 dated diary entries, 124 emails across 48 threads, 9 documents and 108 named correspondents — 643,768 words of verbatim text spanning March 2, 2001 to February 18, 2025. This page explains what that process does, and what it cannot do.
The source material
All of it comes from five packages released publicly by the U.S. Senate Homeland Security & Governmental Affairs Committee under Chairman Rand Paul. Each has its own page with page and word counts:
- Tony's Diary Package — 677 pages, 270,757 words (907 diary, 8 messages, 0 documents)
- Diary Prequel — 263 pages, 214,930 words (980 diary, 2 messages, 0 documents)
- Awards & Nominations — 162 pages, 30,871 words (0 diary, 52 messages, 0 documents)
- Reading Room — 704 pages, 121,535 words (2 diary, 62 messages, 0 documents)
- CBP Release Package — 41 pages, 5,675 words (0 diary, 0 messages, 9 documents)
As works of the U.S. federal government, the records themselves are generally not subject to copyright. Redactions were applied by the releasing body before publication. Nothing has been unredacted here, and nothing has been added.
What the parsing does
Text is extracted from each PDF page, then split into items by the headings the documents themselves use — a date line for a diary entry, a From/To/Subject/Sent block for an email. Each item keeps the page range it spans, so it can always be traced back. Where an email was pasted inside a diary entry, or a reply quotes earlier messages, those are recorded as nested references rather than duplicated as separate items.
One deliberate transformation applies to display only: the source PDFs hard-wrap every line at the page margin, which reads badly at any other width, so lines that reached the margin are rejoined into flowing paragraphs. This changes whitespace and never words. Every page offers a plain-text version with the original line breaks intact — use that when quoting.
What it cannot do
- Attachments. The release records 22 attachment filenames but not the files, except where a document was pasted inline into a message body.
- Scanned pages. The CBP package is mostly images. Those pages were read by OCR, are labelled as such, and can contain character errors; redaction boxes appear as ▮.
- Threading. The emails are printed messages, not mailbox files, so threads are reconstructed from subject lines and participants rather than from message headers.
- Reprinted pages. Two packages print some emails twice — the Reading Room reproduces its pages 2–51 again at 225–274, and 202–223 at 339–360. Of 155 message records, 31 are second printings of an email already in the release, so 124 distinct messages are shown. Nothing is discarded: each message is displayed once and cites every page range it appears on.
- Names in the diary. Person pages match surnames as plain text, which can catch a different person with the same surname.
- Dates. They come from document headings and are only as accurate as those headings.
Reading and citing
Press ⌘K anywhereTap the search icon to search all 643,768 words; @surname restricts to a correspondent, 2020-03 to a period, and "quoted text" to an exact phrase. When citing, use the package name and page numbers shown on the item — those refer to the original PDF, which is the authority, not this site.
Questions
Is any of this text summarized or generated?
No. Every word of body text is extracted verbatim from the released PDFs. Nothing is summarized, paraphrased, rewritten or generated by a model. The only transformation applied to body text is whitespace: lines the PDF broke at the page margin are rejoined so paragraphs flow at the reader's width. Words are never changed. The plain-text version of any page shows the original line breaks.
How can I verify a passage against the original?
Every entry, message and document displays the package name, the source PDF filename, and the page numbers it came from. Open that PDF at those pages and compare.
Why are some pages OCR rather than text?
One package — the CBP Release Package — is mostly scanned images with no text layer. Of its 41 pages, 23 carry recoverable text and 19 of those were read by optical character recognition. Those pages are labelled OCR wherever they appear, redaction boxes show as ▮, and character errors are possible.
How were emails grouped into threads?
The 124 messages were grouped into 48 threads by subject line and participants, since the release contains printed messages rather than mailbox files with real threading headers. Grouping is therefore a reasonable reconstruction, not a record from the mail system.
How are people identified?
Correspondents are keyed on the email addresses in the release, which yields 108 named people. A person page also lists diary entries containing their surname — that is a plain text match, so a common surname can pull in entries about someone else.
Are the dates reliable?
Dates come from the headings in the source documents. They are as reliable as those headings: one entry is headed February 5, 2000 while its content describes events of February 5, 2020, so it sorts under 2000 here. Nothing has been silently corrected — what you see is what the release says.
Who runs this site, and is it affiliated with anyone?
It is an independent archive of U.S. federal public records. It is not affiliated with, endorsed by, or operated by Dr. Anthony Fauci, the National Institutes of Health, NIAID, the Senate, or any government body.