M02 — Document validation
State: in progress — Depends on: M01 — Placed and named by ADR 36
Goal
Given any document the reader can open, produce a structured, machine-readable verdict on what is wrong with it — separately from whether it could be read at all.
Reading answers "can I get at this?". Validation answers "is this sound?". They are different questions, and every later guarantee is expressed in terms of the second: repair fixes findings, conformance profiles add rules, and a caller ingesting third-party files wants the verdict, not the diagnostics of our parser.
Scope
In: the rule engine, the report, and the structural profile — the rules that hold for any PDF whatever it claims to conform to.
Out, explicitly: PDF/A and PDF/UA rules (M20 — they are profiles for this engine, and designing the engine so they slot in is part of this milestone); fixing anything (M05); rendering-level checks that would require laying out content (never, in this profile).
Where: in the core, namespace AdCodicem.Pdf.Validation, so that the core's PdfRepair (M05) can take
findings as input and the rules can read the reader's internals. M02 ships no package; the PDF/A and PDF/UA
profiles go to the AdCodicem.Pdf.Conformance satellite in M20 (ADR 36).
Design
Findings
A finding is (RuleId, Severity, Location, Message, Remedy).
- RuleId is stable and public from the day it ships —
file.eof-missing,xref.entry-broken,page-tree.count-mismatch,object.key-missing. Callers filter on them, repair keys remedies off them, and renaming one is a breaking change once a stable release carries it. Two segments,family.name, in lowercase kebab case; the family is never a profile's name; one identifier, one rule, one severity; and no identifier equals a reader diagnostic code (ADR 36).PdfValidationRuleIdsholds them as constants. - Severity (
PdfValidationSeverity):Error(the document is broken: the reader cannot vouch that it reads what was written — it rebuilt the index, lost part of the file, or chose what the file does not designate),Warning(the file breaks the specification and is read all the same as it was evidently meant),Information(worth knowing). The bar is ADR 45's, which replaced ADR 36's "readers will disagree" on 2026-09-27. A scale of its own: the reader'sPdfDiagnosticSeveritydescribes what the reader did, not how wrong the file is. - Location (
PdfValidationLocation) points at the object, at a byte offset, and at the page when there is one —PageIndex, from 0 in the page tree's order, written from 1 (slice 3). A finding nobody can locate is not actionable — the lesson already learned from the filter diagnostics in M01. - Remedy is a hint, not an action: it names what repair would do, so that a report reads as a plan.
Engine
PdfValidator runs a profile over a document the caller opened; stateless, thread-safe
PdfValidatorOptions immutable: the profile (Structural by default) and the report's capacity
ValidationProfile an ordered set of rules with a name and a version; Structural is built in
PdfValidationReport the findings in order, bounded, with exact counts by severity
IValidationRule one check, stateless, given a ValidationContext (internal)
ValidationContext the document, its source, and the analyses several rules share, each worked out once:
the probe of the index, the walk of the page tree, the walk of the objects the
trailer reaches — the objects themselves stay in the reader's cache (internal)
The engine is internal until deliberately made public: callers use the built-in profiles and cannot yet write rules; M20's satellite is the first consumer that needs a public rule API (ADR 36).
The validator inherits the reader's discipline: lazy, bounded, and it never loads what it does not inspect. Rules that need content streams say so, and the caller can exclude them.
The structural profile
| Family | Checks |
|---|---|
| File | Header present and plausible, startxref correct, trailer complete, /Root resolves to a catalog, %%EOF present, incremental updates coherent |
| Cross-references | Every entry points at the object it claims, no gaps that matter, generation numbers coherent, object streams intact |
| Object graph | No reference to a non-existent object, no cycle where the specification forbids one, no orphaned page, required entries present for each /Type, with the types and versions the Arlington model gives them (ADR 44) |
| Page tree | /Count matches reality, every kid is reachable, inherited attributes resolve, /MediaBox present and well formed |
| Streams | Declared length matches reality, filters decodable, /DecodeParms coherent with /Filter, image streams have the entries their filter requires |
| Fonts | Embedded or standard-14, /Widths as long as /FirstChar to /LastChar says and coherent with the descriptor, ToUnicode present where text is meant to be extractable, encoding coherent. Widths checked against the embedded program's advances are PDF/A's rule (ISO 19005-2 6.2.11.5), M20's pdfa-font on M08's parser, not the structural profile's |
| Resources | Every name referenced by a content stream exists in the resource dictionary (needs the M15 interpreter — until then, the reverse: resources that exist are well formed) |
| Annotations | Targets resolve, destinations point at real pages, widget annotations belong to a field |
| Metadata | /Info and XMP agree where both exist, dates well formed, /ID present |
| Security | Encryption declared, handler recognized, permissions coherent |
A stream, object or section the reader cut at one of its limits (a limit.* diagnostic, ADR 34) is not
undecodable or malformed: the limit is the reader's, and the file may be valid. The rules on it report at
most that it was not checked whole, as information, never an error; the validator opens documents with
the caller's PdfReaderOptions, so raising a limit lets it check the rest.
Slices
- Findings, report, rule engine, profile plumbing, and a single trivial rule end to end. Done on
2026-09-26: ADR 36,
PdfValidatorand the structural profile withfile.eof-missing, the manifest'sfindings,CorpusValidationTests,ValidationBenchmarks. - File and cross-reference rules; run over the whole corpus and see what they say. Done on 2026-09-27
(#58): twenty rules,
file.*andxref.*, their severities set by ADR 45, which replaced ADR 36's "readers will disagree" with whether the reader can vouch that it reads the file as written; the reader records the file's own structure as it opens it, and looks for a catalog among the indexed objects before rebuilding anything; every document's findings in the manifest, cross-checked against an independent analysis of the raw bytes and against qpdf (ValidationRefereeTests); eight iPRES files no longer unsupported. - Object graph and page tree rules, and the object-shape rules — each dictionary type's required keys,
their types, the keys a version deprecates — generated from the Arlington PDF Model at a pinned commit, at
warning severity at most, with overrides by name where ISO 32000-1 does not require what the model says
(ADR 44, amended on 2026-09-28).
Done on 2026-09-29 (#59), in two pull requests. The first (#122): #51 fixed, thirteen rules —
xref.object-stream-circular, threeobject.*and ninepage-tree.*—, the walks of the page tree and of the objects the trailer reaches, pages counted as qpdf's walk counts them, the page in a finding's location, and fourteen documents no longer unsupported. The second: four rules generated from the model —object.key-missing,object.value-type-wrongandobject.type-value-wrong, warnings, andobject.key-deprecated, information —, bytools/AdCodicem.Pdf.Arlingtonfrom the model vendored at commitc48b363, into tables committed and regenerated by a unit test; the model's notice inNOTICE, packed with the core; overrides where ISO 32000-1 does not require what the model says, and silence where a hand-written rule reports the fault; the last four documents recorded as unsupported until M02 diagnosed, and the catalog's/Typeheld toqpdf --check. A key newer than the version the file declares was left to #123, which the maintainer moved to M20 on 2026-09-29 (ADR 44, Reviewed on 2026-09-29): a newer key still conforms, and what bounds the version is a conformance claim. - Stream and font rules (#60). The rule "filters decodable" can judge a stream by what the reader reports as it
decodes it. Since #56, Flate's reports tell each case apart:
filter.failedfor data left encoded, cut short, or corrupt at a byte its message gives, andfilter.checksum-mismatchfor data that decoded whole under a zlib checksum that disagrees. ASCII85, ASCIIHex and RunLength still pass over what they cannot decode in silence, so the rule waits on #141, and on #144, since a damaged object stream is reported each time the reader decodes it. Its identifier, and the length rule's, take thestreamfamily and none of the reader's codes (ADR 36): notstream.length-invalidorstream.truncated, the obvious names for what the length rule checks, norstream.self-reference. - Annotation, metadata and security rules.
- Determinism, memory budget, and the report's serialized form.
- The milestone's adversarial review (ADR 46,
#137), run by a session that worked on none of M02 once every other issue filed under it is closed, as
docs/milestone-review.mddescribes. It waits on two issues filed under M02 on 2026-09-29, when reviews were adopted: the threat model's first version, for the reader and the validator (#135), and M01's review after the fact (#136), so that M02's review finds a reader that has been through one.
Tests required
- Each rule: a document that triggers it, a document that does not, and a document where it must stay silent because the situation is legal but unusual.
- The report is deterministic and serializes to stable JSON.
- The validator never throws on a document the reader could open — it reports. The one exception is the
PdfLimitExceededExceptiona caller asks for by opening the document withThrowOnLimit(ADR 36). - Rules that need content are skippable, and skipping them is visible in the report.
Acceptance conditions
| Documents | Behavior | Verified by |
|---|---|---|
| Every well-formed corpus document, from every producer — ours, Word, PDF24 and the vendored third-party files | No error-severity finding. A validator that calls Chromium's, LibreOffice's or Word's output broken is wrong, not strict | CorpusValidationTests.Well_formed_documents_have_no_errors |
Every document, clean or not — expect.findings in the manifest, an absent field meaning none | Exactly the rule identifiers the manifest declares, no more and no fewer: a warning nobody declared on a sound file fails as surely as a missing finding on a damaged one | CorpusValidationTests.Every_document_produces_exactly_its_declared_findings |
vendor/pikepdf/handwritten-cyclic-toc.pdf, vendor/pdf-association/handwritten-dict-is-stream.pdf — recorded as unsupported until M02 | Findings for a trailer without /Size, a page without /MediaBox or /Resources, and a page object given a stream body. The reader was silent on all three while qpdf reports them; slice 3 removed their unsupported marker | CorpusValidationTests.Documents_waiting_for_M02_are_now_diagnosed |
The remote documents recorded as unsupported until M02 (docs/corpus-sources.md): the Axapta credit note and the #00 form; two catalogs without /Type, a page tree with a null kid, and two page trees whose kids point at objects the file lacks (remote/opf-format-corpus/jhove-*) | Findings for each — an /Info without endobj, a NUL in a name, a catalog without /Type, a kid that is null or missing — and one page count for a tree with missing kids — the manifest's, which is qpdf's, poppler's and PDFium's, each missing or null kid a blank page, rather than pikepdf's, which skips them; M06's page API then materializes those pages; slice 3 removed their unsupported marker, and M02 closes only on a green Remote corpus run | CorpusValidationTests.Documents_waiting_for_M02_are_now_diagnosed |
The iPRES 2017 hand-built set (remote/ipres2017/*, ADR 33): the 17 files recorded as unsupported until M02 — seven page trees qpdf and the reader count differently, four faults the reader reads without a word (a root typed /Pagez, a page typed /Font, generation 10000 in the table, a trailer without /Size), six recoveries that differ from qpdf's —, and the two references to an object the file lacks, which T27 made supported on 2026-09-27: each now reads as null rather than rebuilding the index | Findings for each, one page count per damaged tree, and a catalog recovered without discarding a sound index when only the trailer's /Root is broken; slices 2 and 3 removed the 17 files' unsupported marker, and M02 closes only on a green Remote corpus run | CorpusValidationTests.Documents_waiting_for_M02_are_now_diagnosed |
Vendored files every reader opens despite an anomaly: stale linearization hint tables, a /Size one too large, an xref stream with no entry for itself, references to objects missing from the xref, font-level XMP that is not well-formed XML, UTF-16LE text strings, two startxref lines, a /MarkInfo pointing at the page tree, a malformed date | Warning or information findings, never error: every reader opens these files, and a reference to an undefined object is null by the specification | CorpusValidationTests.Field_anomalies_are_warnings_not_errors |
Every vendored document that veraPDF finds PDF/A-invalid (conformanceValid: false) | No structural error: they are sound files that break conformance rules, and conformance is M20's business | CorpusValidationTests.Conformance_failures_are_not_structural_errors |
| Any corpus document | Two runs give identical findings in identical order | CorpusValidationTests.Validation_is_deterministic |
| The 1000-page journal | Validation holds within its memory budget and reads no content it does not inspect | CorpusValidationTests.Validating_the_largest_document_stays_within_its_budget |
An external suite. The structural profile has one test suite written by others: the iPRES 2017
hand-built set, 88 files derived from one page, each with one deviation from ISO 32000-1's structure that
its authors describe, from a missing header to a trailer pointing at the wrong cross-reference offset — and
often more than that one: in 45 of the 59 header and body files the edit left a table that no longer leads
to its objects. Since
ADR 33 the remote corpus fetches it
out of its authors' archive. Every one of its 88 files, the 19 named above included, falls under the row
that holds each document to exactly its declared findings; the 16 that qpdf passes or only warns about
also fall under the well-formed row — warnings for those, never errors. Each slice that names a rule adds to
each entry the finding it expects, or the reason the profile stays silent. The 13 content-operator cases (BT, ET, Tf, Tj, cm, their operands
and parentheses) wait for M15's interpreter for their own finding; M02 answers only for the damage around them.
Slice 1 answered for the seven end-of-file cases. t04-002 (no marker), t04-003 (%%EO) and t04-004
(EOF) earn file.eof-missing, a warning since qpdf accepts all three. Four stay silent, each because a
marker lies in the file's last 1,024 bytes, where readers look for it: t04-005, whose marker is followed
by junk on its line, and t04-001, whose marker shares the startxref offset's line, since readers tolerate
what surrounds the marker — a rule requiring it alone on the last line would be a rule of its own; t04-006,
whose extra marker before the last is legal, as every incremental update writes one; and t04-007, whose
only marker comes before the trailer, silent only because the whole file, 637 bytes, is shorter than the
window — the same file two kilobytes long would earn the finding. What the premature marker breaks there,
the trailer and startxref after it, is the business of slice 2's file rules.
Slice 2 answered for the header, cross-reference and trailer cases. The four headers without %PDF-
(t01-004 to t01-007, the replaced dash of t01-005 included) earn file.header-missing, and the three that
name no version of PDF (t01-001 to t01-003) file.header-version-invalid. Of the cross-reference cases,
t03-001 and t03-002, whose startxref names no section, earn file.startxref-wrong; the five whose rows
or subsection headers are wrong (t03-003 to t03-006, t03-009) xref.section-malformed; the entry 60
bytes early (t03-007) xref.entry-shifted, the reader finding the object near it; and generation 10000
(t03-010) xref.generation-mismatch, a warning since the references and the object's header agree against
the entry. t03-008, a row two digits short, stays silent: every reader here reads it as it was meant, and a
rule for the fixed row width is left to #107. Of the trailer cases, the premature markers of t04-006 and
t04-007 move the trailer after them, so that startxref names no section (file.startxref-wrong);
t04-008 has no trailer keyword (file.trailer-missing); t04-009, t04-010, t04-012 and t04-013 have
broken dictionaries (file.trailer-malformed), and t04-011 to t04-014 a /Root that leads to no catalog
(file.root-invalid), the reader finding the catalog among the indexed objects without rebuilding the index;
t04-015 has no /Size (file.size-wrong), t04-016 a catalog numbered above its /Size
(xref.object-past-size, a warning: qpdf and the reader read it); t04-017 and t04-018 lack startxref or
its offset (file.startxref-missing), and t04-019 gives a wrong one (file.startxref-wrong). The 59 header
and body files earn, besides, file.startxref-wrong wherever their edit moved the table: 38 of the 52 body
files and four of the seven header files.
Slice 3 answered for the page tree and page object cases, and for the catalog cases the index leaves readable. A
root that lists itself (t02-02-002) is page-tree.cycle, an error: the reader counts no page there, qpdf reports
the loop, and the page the tree should have listed is page-tree.page-orphaned. A kid the file lacks (t02-02-004,
t02-03-006, and t02-02-003's third) is page-tree.kid-invalid, a warning, each counting as a page with nothing on
it — the manifest's counts are qpdf's walk, where they had been the root's /Count —; t02-02-003's resources
dictionary listed as a kid is read as a page without /Parent, /MediaBox or /Resources. A root without /Kids
(t02-02-005) is page-tree.kids-missing, and one whose /Count is wrong or missing (t02-02-007, t02-02-008)
page-tree.count-mismatch; a page without /Parent, with the wrong one or an array for it (t02-03-003 to
t02-03-005) page-tree.parent-wrong; without /MediaBox or with three numbers (t02-03-008, t02-03-009)
page-tree.mediabox-invalid; without /Resources (t02-03-012) page-tree.resources-missing. A /Pages or a
/Contents naming an object the file lacks (t02-01-004, t02-02-001, t02-03-011) is object.reference-missing, a
warning, and t02-01-004's page, which no tree lists, page-tree.page-orphaned. Two content-stream cases whose
stream keyword is missing, or that have bytes between endstream and endobj (t02-05-01-013, -017), earn
object.endobj-missing for what follows their value; one without endstream (-014) leaves where its object ends
unknown, and its stream is the stream rules' to judge.
The object-shape rules answered for the rest of the catalog, page tree and page object cases. A root typed /Pagez
(t02-02-009), a page typed /Font (t02-03-002) and a catalog of the wrong /Type (t02-01-006) earn
object.type-value-wrong; a page or a root without /Type (t02-03-001, t02-02-006), a catalog without it
(t02-01-005, t02-01-007) or without /Pages (t02-01-003), and t02-02-003's resources dictionary, read as a page
without /Type, object.key-missing. Of the resource and content-stream cases, a font without /BaseFont
(t02-04-01-001) or without /Type (t02-04-01-003) and a stream without /Length (t02-05-01-016) earn
object.key-missing; a real for a /Length (t02-05-01-015) and a content stream whose missing stream keyword
leaves a dictionary where a stream belongs (t02-05-01-013) object.value-type-wrong; a font typed /Page
(t02-04-01-004) object.type-value-wrong. Four stay silent. t02-03-007, whose /Resources names the page tree's
root: the walk types the root as the tree's root first and judges it once, and a dictionary is what /Resources
wants. t02-04-01-005 and -006, a font without /Subtype or with /Subtype /Type7: the subtype is what tells a
font's kinds apart, and the walk leaves unchecked what it cannot type rather than guess, so a finding for them is the
font rules' to consider (slice 4) — as is -002's /BaseFont, which names a font that is not one of the
standard 14.
Three cases no slice had named. t02-01-001 has no catalog at all, and its edit moved the table, so the reader
rebuilds the index as it opens it: it earns file.startxref-wrong, and nothing yet for the missing catalog, since
file.root-invalid judges only the trailer the chain gave — #132, to be paid before M02 closes. t02-01-002, whose
catalog is renumbered 6 while the trailer's /Root still names 5, earns file.root-invalid — the reader takes object
6, the one catalog, which the file does not designate — and xref.entry-broken for the entry of object 5, which no
longer leads to it. t02-03-010, a page without /Contents, stays silent: ISO 32000-1 makes the entry optional, and
a page without it empty (Table 30); what the edit did break, the table's offsets, is file.startxref-wrong's.
Corpus
What the corpus holds
- Sound files from every producer, for the row that allows them no error: Chromium, LibreOffice, ReportLab
and qpdf's rewrites, Word and PDF24, and the 146 vendored third-party files — Distiller from 2 to 10,
PDFMaker, InDesign, LiveCycle, PDFlib, FOP, iText, Quartz and the rest (
docs/corpus-sources.md). - Damage with its findings declared: the five copies of our invoice damaged by
build_corpus.py, which writes theirfindings; the iPRES 2017 hand-built set (88 remote, ADR 33), each file with the deviation its authors describe; the two committed files, and the 24 remote ones — 17 of the iPRES set and seven others —, recorded as unsupported until M02 and named in the acceptance conditions. IBM's QMF manual (T25) and two iPRES files (T27) were too, until their debts were fixed. - Anomalies every reader accepts, for the warning row: stale linearization hints
(
linearization-hints-inconsistent, 26 files),size-off-by-one,utf16le-info-dictionary,markinfo-aliases-pages-node,malformed-font-xmp,invalid-creation-date-year-zero,two-startxref-lines-before-eof. - Conformance without structure: veraPDF's pass and fail fixtures and the 30 committed PDF/A claims, six of them rejected, for the row that keeps conformance out of the structural profile.
- What the later families read:
info-xmp-metadata-mismatchand malformed dates for the metadata rules, non-embedded fonts andno-tounicodefor the font rules,dangling-referenceandcyclic-destinationsfor the graph and annotation rules, 18 encrypted documents for the security rules. - Scale: the 1000-page journal committed.
What it lacks
| Need | Why | Priority | Likely source |
|---|---|---|---|
| A structurally sound file whose Flate stream lost its tail, and one whose LZW stream uses a code it never defined | T32's fix and the stream rule "filters decodable" meet these only in unit tests: the committed damage cuts a whole file, not a stream inside a sound one | 2 | Generated here: a derived variant of our invoice with one stream cut inside its data and its /Length adjusted, recorded in build_corpus.py |
An object stream whose /DecodeParms names an object stored inside it | Since slice 3 the reader reports it (stream.self-reference) and so does the validator (xref.object-stream-circular), but on the hand-built files of ObjectRuleTests only: no corpus document holds it | 3 | Generated here: a hand-built variant, marked as such |
| A widget that belongs to no field, and a destination naming a page the file lacks, from a real producer | The annotation rules meet these only in hand-written and damaged files | 3 | A contribution (W06) |
Traps
- A validator that is too strict is worse than none: it trains its users to ignore it. When in doubt
between
ErrorandWarning, the answer isWarning, and the bar forErroris that the reader cannot vouch that it reads the file as written (ADR 45) — checked against what the reader did, not argued about readers the project does not run. - Several rules want the same resolved objects. Without a shared cache, validation becomes the one place that reads the document several times over.
- Rule identifiers leak into user code the moment they ship. Name them as if they were public API, because they are.
Documentation
By Diátaxis section, since ADR 47 (2026-09-30):
docs/website/docs/tutorials/first-steps.md— opening, inspecting and validating a document, held to theFirstStepssample by a test.docs/website/docs/guides/validate-a-received-document.md— validating a document before accepting it, and acting on the report.docs/website/docs/reference/validation.md— the validator, its options and profiles, the report, findings, severities and locations.docs/website/docs/reference/validation-rules.md— every rule identifier, its severity and its meaning, published on the site and versioned with each release.docs/website/docs/concepts/validation.md— what validation answers that reading does not, where a warning ends and an error begins, and the Arlington model.docs/website/docs/introduction.md— validation added to what the library can do.
Exit criteria
- The rule engine, the report and the structural profile exist, with rules documented by identifier.
- The acceptance conditions above pass on the corpus, in CI, with no document skipped. Since 2026-09-29 the two
documents recorded as unsupported for #47 (M23) are skipped by the laziness test alone (
unsupportedTests), and every validation test holds them to their findings. - T23 (
docs/status.md) is fixed: no syntax error invented at a window's edge — done on 2026-09-26 for objects, streams and classic tables; the rebuild's trailer scan keeps a fixed window (#49, M23). - T25 is fixed, before the cross-reference slice: no cross-reference section dropped in silence — done on
2026-09-26: a section named a few bytes off is found nearby, one that is nowhere is
xref.section-missing, and IBM's QMF manual reads its 429 pages without a rebuild. - T27 is fixed, after T25 and before the object-graph slice: a reference to an object the file lacks is null and a finding, not a rebuild — done on 2026-09-27: null, silently, unless the index may have lost entries; the finding is slice 3's.
- T32 is fixed, before the stream slice at the latest: a Flate stream that lost its tail, or an LZW
stream that stops at a code it never defined, is reported — done on 2026-09-26, the lost tail told from a
lost checksum, and held to qpdf's reading of every corpus document by
FlateRefereeTests. - #55 is fixed, before the stream slice: a stream whose data runs past the parser's window has its declared
length checked, as a shorter one's is — done on 2026-09-29: the file is asked for the bytes after the declared
length, and searched for the first
endstreamwhen they are not one, no further than the nearer of the next object the index as written places (the one rebuilt as the document opened when the file wrote none) and the first object header the file's bytes hold, or the end of the file, once per stream. Nothing a read changes bounds it: in the 401 corpus documents measured, each stream — each copy of an object at its offset — takes the same length whatever was read before it; crafted files still can make it depend on the order (#138, M23). Of the sixteen streams of six documents whose length is wrong past the window, fifteen of five documents now read to theirendstream; the sixteenth, PDFium's object 695, whose tail a block of zeros erased, keeps its declared length, and says so. Each stream is reported once however often it is parsed, what the reader found of it is recorded for this slice's length rule, andStreamLengthRefereeTestsholds the lengths taken to qpdf's. #120 went with it: a/Lengththat gives no length is reported as the file wrote it, no longer as "-1 bytes". - #56 is fixed, before the stream slice: a zlib stream whose checksum is wrong keeps all of its data, and
the disagreement is reported — done on 2026-09-30. Data that faults is read again, and only what the first
reading lost is added to what it kept: a zlib stream's body as raw deflate, which reads to its end when only the
checksum disagreed, and, when that faults too, once more with its input handed over a byte at a time from the
chunk the fault was met in. The 272 damaged Flate streams of the corpus keep exactly what libz keeps. The 86
whose checksum disagrees with whole data keep all of it, and report
filter.checksum-mismatch, a warning; 72 of them had been left encoded, and 14 reported corrupt with only a prefix kept. Of the 137 that turn corrupt, 98 keep what decoded before the fault and report the byte libz meets it at; the 39 in which nothing decodes stay encoded. That is 1,441,772 bytes no longer lost. A header that asks for a preset dictionary is not read again, and qpdf fails on it too. Sound streams are read once, as they were.FlateRefereeTestsholds both reports to qpdf's filtered data: equal where the checksum disagrees, a prefix of ours where the data turns corrupt; and since qpdf keeps nothing of the three committed documents' corrupt streams,CorpusReadingTestsholds those three to libz's own figures, the byte of the fault and the bytes kept. #134 went with it: the referee reads every diagnostic, where the first thousand hid the later streams of IBM's QMF manual. - #51 is fixed with the object-graph slice: an object stream whose
/DecodeParmsnames an object inside it is reported, and the null it produced is not cached — done in slice 3:xref.object-stream-circular, over/Length,/Filter,/DecodeParms,/Nand/First, directly or through another object stream; the reader reportsstream.self-referenceand reads the object again once the stream is read. - The report serializes to stable, documented JSON.
- The object-shape rules are generated from the Arlington model at a pinned commit, its Apache-2.0
notice carried in
NOTICE, and a test fails when the generated tables and the pinned model disagree — done in slice 3:tools/AdCodicem.Pdf.ArlingtongeneratesArlingtonModel.g.csfrom the model vendored atc48b363and its lock;ArlingtonGeneratorTestsregenerates it and compares it byte for byte, checks every vendored file against the lock, and fails on an override the model no longer needs;NOTICEis packed with the core. - A benchmark measures validation of the 1000-page document, with
MemoryDiagnoser:ValidationBenchmarks, since slice 1, measured again as each slice adds rules. -
docs/website/docs/reference/validation-rules.mdlists every rule identifier, its severity and its meaning. - Integration tests compare our verdicts with an independent tool's on every corpus document — the file and
cross-reference families since slice 2:
ValidationRefereeTestsholds every document qpdf rebuilds the index of to a finding of theirs, and every error of theirs to a document qpdf finds fault with; since slice 3, every error of any family, every document whose page tree qpdf repairs to a finding, the manifest's page counts to qpdf's walk of the tree (QpdfRefereeTests) — damaged documents too since 2026-09-29, but a tree qpdf cannot walk and a catalog the reader chose (file.root-invalid) —, and the catalog's/Type, reported where and only where qpdf reports it missing or invalid. - The documentation site publishes the finding model and the full rule table.
- The threat model's first version covers the reader and the validator:
docs/threat-model.md(#135), written on 2026-10-01. The twenty-five debts it filed under M02, #154 to #175 and #181 to #183, are among the issues the milestone's closing waits on. - M01's review is recorded in
docs/reviews/M01.md(#136), against the threat model's reader section (#135) — merged with #198 on 2026-10-01. The nine issues it filed under M02 are among step 4's batches. - The milestone's review is recorded in
docs/reviews/M02.md(#137), and every issue it filed under the milestone is closed.