Skip to main content

M01 — Object model and tolerant reading

State: done — Depends on: M00

Goal​

Open any PDF, imperfect ones included, and resolve its objects on demand, without ever loading the whole document into memory.

Scope​

In: the COS object model; the lexer and parser; decoding filters; every form of cross-reference table; lazy resolution; repair; diagnostics; and the corpus harness that later milestones will build on.

Out, explicitly: writing (M03), the notion of a page (M06), decryption of protected documents (M16 — here /Encrypt is detected and a typed exception raised), content stream interpretation (M15).

Design​

Object model — Objects/​

PdfObject as the base, then PdfNull, PdfBoolean, PdfInteger, PdfReal, PdfName, PdfString, PdfArray, PdfDictionary, PdfStream, PdfReference.

  • PdfName is interned and compared by reference; common names are static fields.
  • PdfString keeps its raw bytes and its original notation; conversion to text handles PDFDocEncoding and UTF-16BE with a byte order mark.
  • PdfStream carries its dictionary and a lazy data source: a range of the file, or memory. The bytes are neither read nor decoded before they are asked for, and the encoded form stays available as it is — which is what will later let a stream be copied between documents without recompression.
  • PdfReference resolves through an explicit object source, never through static state.
  • Typed access goes through extensions (GetDictionary, GetArray, GetInteger…) that resolve references on the way — reading a dictionary entry without resolving is the classic defect in this kind of code.

Low-level reading — IO/​

  • PdfLexer: tokenizes over ReadOnlySpan<byte> without allocating. Handles comments, delimiters, literal string escapes (including \ddd and escaped line endings), odd-length hexadecimal strings, and #xx in names.
  • PdfObjectParser: builds objects, with bounded recursion depth for nested arrays and dictionaries.
  • Filters/: FlateDecode with PNG and TIFF predictors, LZWDecode, ASCII85Decode, ASCIIHexDecode, RunLengthDecode, and pass-through for DCTDecode, JPXDecode, CCITTFaxDecode and JBIG2Decode. Filter chains and DecodeParms in both their single and array forms.
  • XRef/: classic tables, cross-reference streams (/W, /Index), object streams (ObjStm), the /Prev chain, and hybrid-reference files (/XRefStm).
  • PdfFileReader: orchestrates indexing, holds the bounded cache, exposes GetObject(PdfObjectId).

Repair​

The index is rebuilt by scanning, once at most per document (PdfFileReader). It is rebuilt at opening when startxref is missing or gives no offset, when the section it names cannot be read, or when the chain indexes no object. A trailer whose /Root leads to no catalog has the catalog looked for among the objects the chain indexed first, and the index rebuilt only when none of them is one. Later, or while opening, it is rebuilt when an object is loaded that is neither at its offset — or whose offset falls outside the file — nor within 512 bytes of it, and when an object the index lacks is asked for while the index may be incomplete: a section of the chain found nowhere, a guard that stopped the chain or a table, a chain that loops, a row of a classic table that cannot be read, a cross-reference stream short of its rows or giving a count out of range, a /Prev or /XRefStm that is not an offset. Otherwise the index is not rebuilt for an object it lacks: that object is null, as the specification says. The scan looks for N G obj headers across the whole file and keeps the last definition of each object, then looks for a trailer or, failing that, an object of /Type /Catalog. Every repair produces a diagnostic entry.

Diagnostics — Diagnostics/​

PdfDiagnostics: a collection of (severity, code, message, position) entries. Codes are stable and documented, because they are part of the public contract. Severities: information, repair, warning, conformance loss.

Slices​

  1. Object model, typed accessors, unit tests. (done)
  2. Lexer and parser, on valid input and then on malformed input. (done)
  3. Filters, with known test vectors. (done)
  4. Classic table, trailer, lazy resolution; opening a hand-written minimal PDF. (done)
  5. Cross-reference streams, object streams, the /Prev chain, hybrid files. (done)
  6. Repair by scanning, and diagnostics. (done)
  7. Hardening: bounds, hostile corpus, memory measurement. (done)
  8. The corpus harness: generation script, manifest, and the corpus-driven test. (done)
  9. Fuzzing of the lexer and parser. (done)

Tests required​

  • Every object type: canonical textual form, and edge cases (a string with unbalanced parentheses, a name containing #, a real with no integer part, an integer beyond long).
  • Lexer: comments, exotic whitespace, premature end of file, odd-length hexadecimal strings.
  • Filters: known vectors for each; PNG predictors on real data; a chain of two filters.
  • Cross-references: classic table, cross-reference stream, object stream, multiple /Prev, hybrid file, a table that lies.
  • Repair: truncated file, wrong startxref, missing table, duplicated object.
  • Hostility: a /Prev cycle, dictionaries nested ten thousand deep, a lying /Length, an outrageous /W in a cross-reference stream, a stream claiming an impossible size.
  • Memory: opening a file of several hundred thousand objects within a stated budget. Stated on 2026-10-02 (#193): 300,000 objects open under 34,200,000 bytes indexed by a classic table and 47,430,000 by a cross-reference stream whose rows Flate stores, 5 % over the 32,571,368 and 45,174,464 measured after a first opening (DocumentReaderTests.Opening_an_index_of_three_hundred_thousand_objects_stays_within_its_memory_budget).

Acceptance conditions​

DocumentsBehaviorVerified by
Every document in the corpusOpens; catalog found, page count matches the manifest, and after a full read every diagnostic the manifest requires is reported — a well-formed document produces no repair and no warning at allEvery_corpus_document_opens_as_its_manifest_describes
damaged/*Recovered as far as an independent tool recovers them, and the damage is never silentDamaged_documents_are_recovered_as_far_as_an_independent_tool_recovers_them
Every document, hostile ones includedNo untyped exception, no unbounded recursion, no allocation chosen by the file; each within its time budget: 20 s for each operation on a corpus document — opening it, reading every object, walking its pages, validating a damaged one (since #193) — and 10 s for each hostile inputCorpusReadingTests, HostileInputTests
The scanned page and the 1000-page journalOpening reads under a quarter of the file: content is not read until asked forOpening_does_not_read_the_content_of
The 1000-page journalIndexing and walking the whole page tree allocates under 4 MBReading_every_page_of_the_largest_document_stays_within_its_memory_budget
A generated index of 300,000 objects (since #193)Opening allocates under 34,200,000 bytes through a classic table and 47,430,000 through a cross-reference streamOpening_an_index_of_three_hundred_thousand_objects_stays_within_its_memory_budget
The corpus itselfCovers at least three distinct producers and every use-case category, and every file present is described by the manifestThe_corpus_covers_several_producers_and_every_use_case, The_corpus_manifest_describes_every_document_present

Documents must come from at least three distinct producers, since the shape of a cross-reference section is a producer's signature. The corpus held 19 documents from four producers when M01 closed; every document added since falls under the same rows.

Corpus​

What the corpus holds​

  • Then: 19 documents from four producers — Chromium, LibreOffice, ReportLab and qpdf's rewrites — covering every use-case category, and the five copies of the invoice damaged on purpose by build_corpus.py.
  • Now: 168 committed documents and 242 remote ones (ADR 32), all under the rows above — every shape of cross-reference section, 56 linearized and 47 incrementally updated committed files, the iPRES 2017 hand-built set, and two remote documents recorded as unsupported for #47 (expect.unsupported), which only Opening_does_not_read_the_content_of skips.

What it lacks​

Nothing blocks: the milestone is closed. What would still widen what the reader is proven on is what only an inbox holds (T10, and the "still wanted" column of docs/corpus-contributions.md):

NeedWhyPriorityLikely source
A copier's file untouched since the copier wrote itEvery committed scan has been through another tool since; copiers write index shapes of their own3A contribution (W03)
A supplier's or a bank's invoice or statement from a Java stack, committedThe real ones are in the remote corpus only, so the main CI job never reads them3A contribution (W04)

Traps​

  • A /Length can be an indirect reference and be wrong: endstream must still be findable.
  • Cross-reference offsets can be off by a byte or two in files from careless tools: look for the object header in a small window around the offset before concluding.
  • An object stream holds objects with no N G obj header: their positions come from its header table.
  • Generation numbers are almost always zero, but not always; never drop them from a key.
  • A file may contain several %%EOF markers: that is the sign of incremental updates, not of damage.

Documentation​

By Diátaxis section, since ADR 47 (2026-09-30):

  • docs/website/docs/tutorials/first-steps.md — writing a one-page document, opening it, and seeing what the reader repaired in a copy cut short, held to the FirstSteps sample by a test.
  • docs/website/docs/guides/find-what-the-reader-repaired.md — checking whether a document needed repairs, filtering what the reader did by severity or code, and keeping the account.
  • docs/website/docs/guides/handle-reader-limits.md — telling whether a document reached a limit, raising the one a report names, or refusing the document rather than reading part of it.
  • docs/website/docs/guides/tune-reader-memory.md — sizing the object cache, keeping more diagnostics, and opening a document without copying it.
  • docs/website/docs/reference/diagnostics.md — PdfDiagnostic and PdfDiagnostics, the severities, the full table of diagnostic codes, and the exceptions.
  • docs/website/docs/reference/reader-limits.md — every limit, its default, its code and what is kept, and the bounds that cannot be lifted.
  • docs/website/docs/concepts/lazy-reading.md — what opening does and deliberately does not do, with figures.
  • docs/website/docs/concepts/diagnostics.md — why the reader repairs what it can and reports what it did, how that account differs from the validator's verdict, and how it reasons about damaged streams.
  • docs/website/docs/concepts/reader-limits.md — how every valid PDF stays readable while every file is treated as hostile, and why the guards are on by default.
  • docs/website/docs/introduction.md — what the library can do today, and what indexing a thousand-page document costs.

Exit criteria​

  • The synthetic corpus opens in full, damaged files included, with the expected diagnostics.
  • No malformed input produces an untyped exception, infinite recursion or an unbounded allocation.
  • Memory at open is proportional to the number of objects, not their size (229 µs and 393 KB for a 1000-page, 4 MB document).
  • Streams are neither read nor decoded until requested (verified by counting reads).
  • An indexing benchmark exists, with MemoryDiagnoser.
  • The real-document corpus exists, with its manifest and its generation script: 19 documents from four producers, covering every use-case category (T03).
  • The acceptance conditions above pass on that corpus, in CI, with no document skipped.
  • A memory budget is enforced in CI: indexing the 1000-page document and walking its page tree allocates under 4 MB — 2.4 MB when M01 closed, 3.2 MB on 2026-10-01 (#38; the throughput budget belongs to M23). Since 2026-10-02 (#193), opening an index of 300,000 objects is held to a budget too, and each operation on a corpus document to 20 s.
  • Integration tests cross-check the corpus against qpdf in a container: page counts, and its verdict on which documents are damaged.
  • The documentation site publishes the reading model, the diagnostics contract and the measurements.
  • The lexer and parser are fuzzed, seeded with the corpus: mutations that flip bits, corrupt digits, truncate, splice and break keywords, each asserted to end in a result or a typed exception, inside a time and an allocation budget. A nightly campaign runs 20 000 mutations on each document it fuzzes: since 2026-09-25, the smallest seed document of every reader structure and a rotating share of the others, so that every seed document is reached within a few nights (FuzzingSeeds, fuzz.yml).

What the milestone found​

Fuzzing earned its place in its first minute of existence. A mutated invoice drove LoadRegularObject and RelocateAndLoad into each other — the index said an object was in one place, the neighborhood search answered with a place that failed to parse the same way, and the two called each other 3 978 times until the stack ran out. A file killing the process is precisely what invariant 4 exists to prevent, and no hand-written test had thought to try it.

Relocation is now three counted attempts — the recorded offset, the neighborhood, a rebuilt index — with no path that calls back into loading.