Skip to main content

M07 — Case-file completion

State: to do — Depends on: M06

Goal​

Everything else a case file is made of before a line of layout exists: split a volume into exhibits and an exhibit into parts, turn scans and photographs into pages without touching their pixels, edit and exchange bookmarks, make links between pieces work once the pieces are one volume, and open a received portfolio as one paginated, bookmarked document.

M06 made one volume out of many without losing anything. M07 is the rest of the daily work on a case file: the court wants one file per exhibit under 32 MB, the client sent a TIFF from a copier and a phone photograph, the pleadings link to piece-12.pdf#page=3, and the opposing party sent a portfolio. Each is object-graph and container work — no layout, no image codec, no font.

Scope​

In:

  • a page-selection grammar — ranges, last page, counting from the end, odd and even, exclusions, page labels, bookmarks — and collation of several inputs, including fronts and reversed backs of a single-sided duplex scan;
  • unreferenced resources pruned on extraction and on every split part;
  • split strategies: by page ranges, by page count, by page-label range, by top-level (or level n) bookmark with each part keeping its outline subtree, by maximum size, by separator page;
  • images to pages without re-encoding: JPEG, JPEG 2000, TIFF (one page per frame; CCITT G3 and G4, LZW, Deflate and PackBits strips passed through), PNG (its compressed data passed through with a PNG predictor), at their true resolution and in their EXIF orientation;
  • outline editing, and outline import and export as JSON;
  • links between pieces: after assembly, GoToR, Launch to a PDF, and relative or file: URIs with PDF open parameters (#page=, #nameddest=) rewritten to internal GoTo; GoToE into attached PDFs; after a split, links that cross parts turned into GoToR; open parameters parsed and written;
  • portfolio unpacking: a received /Collection turned into one paginated volume with one bookmark per member, grouped by folders;
  • the tool's verbs for all of the above: split, collate, images, outline, unpack;
  • the corpus's formats beyond PDF, done once here for every milestone that needs them (the maintainer's decision of 2026-09-27): the manifest admits image files for M07, XML invoices for M14, FDF and XFDF for M16, messages for M18 and DOCX for M31, each with its expectations, under one provenance rule and one remote fetcher.

Out, explicitly:

  • any image codec — blank-page detection on a scanned separator sheet, deskew, joining TIFF strips into one image, transcoding an arithmetic-coded or lossless JPEG — M22 (ADR 42); M07 never changes a sample value and never re-encodes lossily, and the few lossless rewrites it does (below) use the core's own Flate and predictors and are reported;
  • JBIG2 image files, GIF, WebP, HEIC and other image formats as inputs — not in the roadmap for M07; M12.5 takes GIF and WebP for HTML, and a caller can convert first;
  • separators recognized by their text or their barcode — the separator predicate is the caller's until M15 (text) gives it something to call; recognizing a barcode or a patch code is planned by no milestone — M10 only paints codes, and M22 names it out of scope; the roadmap keeps it among its open questions — so a caller brings its own decoder;
  • tagging image pages (a Figure with alternative text) — M13; until then an image part in a tagged volume is reported as untagged, as M06 does for any untagged part;
  • generating an outline from a tagged document's headings or from font sizes — M15, which reads the structure and the text;
  • court-portal presets (file naming, bookmark naming, size caps per portal) — M18, which drives these strategies from data;
  • stamping the parts, Bates numbers — M09; authoring a portfolio, and PDF/R output — not planned;
  • decrypting encrypted inputs or members — M16.

Design​

Page selection and collation​

PdfPageSelection a parsed selection; Resolve(document) -> page indices, in order, bounded
PdfCollation Collate(inputs, group size) and Interleave(fronts, backs, reverse backs)

The grammar, in the spirit of qpdf's and pdfcpu's:

selection = item { "," item }
item = [ "!" ] target [ ":" parity ]
target = page [ "-" page ] 1-3 z r2-r1 5-z
| "label:" label [ ".." label ] label:iv..xii label:A-1..A-9
| "outline:" title { "/" title } outline:Annexes/Pièce 3
page = integer | "z" | "r" integer z is the last page, rN the Nth from the end
parity = "odd" | "even"
  • A range written high to low runs backwards (z-1 reverses a document). ! excludes what it selects from what the items before it selected. An outline: target selects from the item's page to the page before the next item at its level, the rule splitting by bookmark uses.
  • The parser is total: any string yields a selection or a typed parse error with its position, never an exception of another kind; length, number of items and integer values are bounded.
  • M06's verbs used plain ranges (1-3,7,z); this grammar is their superset, so nothing written for M06 changes meaning.
  • Collation takes groups of n pages from each input in turn (qpdf's --collate=n); interleave takes fronts and backs, the backs reversed by default, the order a single-sided duplex scan comes out of a feeder. Unequal counts append the remainder and report it.

Pruning unreferenced resources​

PdfResourceUsage (internal) the resource names a page's content can reach, per category
  • Conservative by construction: every name token in the page's content streams, whatever operator it feeds, keeps the resource of that name in every category. Over-keeping costs bytes; under-keeping breaks the page. Unknown operators, BX/EX sections and inline image dictionaries (/CS, /F names) are therefore covered without special cases.
  • The scan follows what content reaches: form XObjects (their own /Resources, or the page's when they have none), Type 3 glyph procedures, patterns, soft masks in graphics states, and annotation appearance streams, each once (a visited set) and to a bounded depth.
  • It tokenizes with the M01 lexer, does not interpret, allocates nothing per token, and reads each content stream once, decoded under the reader's limits. A real token past a double's range is an infinity: the scan reports it as syntax.number-out-of-range and reads it as null, as the object parser does. A content stream the reader cannot decode keeps all resources and is reported: pruning must never be the step that loses a font.
  • Pruning rewrites the part's resource dictionaries, not the source's; M06's deduplication still shares what remains.

Split strategies​

PdfSplitter Split(document, strategy, options) -> parts, written one at a time to a sink
PdfSplitStrategy Ranges(selections), PageCount(n), LabelRanges, Outline(level), MaxSize(bytes),
Separator(predicate, keep or drop the separator)
PdfSplitOptions immutable: part naming pattern, cross-part links (Remote or Drop), document-level
attachments (FirstPart, EveryPart, None), front matter (OwnPart or FirstPart), pruning
(on by default), the M06 policies that apply per part
PdfSplitPart index, page range, title, file name, bytes written, its report
IPdfSplitSink opens the output stream for each part: files in a folder, a zip, or the caller's own
  • Every part is an M06 extraction: page labels kept by default, each form field carried with the part holding its widgets, structure pruned to the part, named destinations kept where their target is, page-level associated files traveling with their page — so a signed original attached by M06's AttachOriginal follows its exhibit. Document-level attachments go where the options say, and a conformance claim stays only while it stays true (M06).
  • By bookmark (level 1 by default, level n on request): the items at that level that resolve to a page, ordered by page and then by outline order, mark where parts start. Pages before the first mark form a front-matter part. Two items on one page would give an empty part: they are joined into one part carrying both subtrees, and reported. An item without a destination marks nothing; its subtree goes with the part holding its first resolvable descendant. Each part's outline is the subtree of the item that opened it; an item of that subtree whose target lies in another part is treated as a cross-part link.
  • Cross-part links: an internal link whose target lands in another part becomes a GoToR to that part's file name and a named destination created in the target part (Remote, the default), or is removed (Drop); both are reported. Named destinations survive later edits of the target part, which page numbers do not.
  • By size: pages are packed in order; each page's cost is estimated from the encoded length of the objects it brings that the part does not already hold — a resource shared by pages in two parts counts in both. The part is written through a counting sink; if it exceeds the cap, it is cut at the last page that fitted and that page starts the next part — at most once per part, so the procedure ends. A page that alone exceeds the cap becomes a part of its own, reported. Deterministic: the same document and cap give the same parts.
  • By label range: a new part wherever the label style or prefix changes, the ranges of M06's PdfPageLabels.
  • By separator: a predicate on PdfPage. The built-in one is blank by content — no painting operator in the page's content or in the forms it draws, and no annotation with an appearance; a scanned blank sheet is not blank by content, and recognizing it is M22's. Separators are dropped by default.
  • Naming: a pattern over index, title and first label, made safe for every file system — reserved characters and Windows device names removed, Unicode normalized to NFC, length bounded, collisions suffixed — and deterministic. Portal-specific naming is M18's.

Images to pages​

PdfImagePages Add(image source, options) -> pages appended to a document or an assembly part
PdfImagePageOptions immutable: page size (TrueSize, or Fit to a box with margins), orientation (Exif or
AsStored), resolution when the file states none, lossless rewrites allowed or refused,
metadata segments kept (default) or stripped

Every format is read by a container parser in the core — markers, boxes, chunks, IFDs — bounded under ADR 34 and fuzzed. The compressed data is copied forward to the writer with an indirect /Length: an image of any size costs the headers, not the pixels. Only the lossless rewrites of the last column inflate samples, with the core's Flate and PNG predictors, a row at a time — never with an image codec.

InputBecomesRefused, or rewritten losslessly and reported
JPEG, baseline or progressive, 8 bits, 1, 3 or 4 componentsDCTDecode, the file's bytes untouched; DeviceGray, DeviceRGB or DeviceCMYK; an Adobe APP14 CMYK or YCCK file gets /Decode [1 0 1 0 1 0 1 0]; an ICC profile from its APP2 chunks becomes ICCBasedLossless (SOF3), hierarchical and arithmetic-coded JPEG, 12 bits: refused (M22)
JPEG 2000, JP2 file or raw codestreamJPXDecode, the bytes untouched; color taken from the file; alpha through /SMaskInDataA JPX file using Part 2 features readers do not implement: passed through and reported
TIFF, each frame a page (reduced-resolution frames skipped)CCITT G4 (K -1), G3 one- and two-dimensional (K 0, K 1, EndOfLine, EncodedByteAlign from T4Options), Modified Huffman (K 0, byte-aligned) as CCITTFaxDecode; LZW (early change) as LZWDecode; Deflate as FlateDecode; horizontal differencing as Predictor 2; PackBits as RunLengthDecode; palettes as Indexed; an embedded ICC profile as ICCBasedSeveral strips or tiles: one image per strip, stacked, each passed through; FillOrder 2: bytes bit-reversed; PackBits holding the no-op byte 128, which RunLengthDecode reads as the end: the no-ops removed; uncompressed samples: Flate-compressed; JPEG in TIFF with its tables: the tables spliced into one JPEG stream. Old-style LZW and JPEG, planar configuration, floating-point predictor, BigTIFF: refused
PNG, not interlaced, without alphaIts IDAT data concatenated as FlateDecode with Predictor 15; palette as Indexed; one transparent gray, RGB or palette value as a color-key /Mask; iCCP as ICCBased; 16 bits as BitsPerComponent 16Interlaced, alpha channel, or a palette with partial transparency: samples unfiltered and re-deflated, alpha split into an /SMask — lossless; gAMA and cHRM ignored and reported; an animated PNG gives its default image
  • Resolution comes from JFIF density (units 1 or 2; 0 gives an aspect ratio, not a resolution), EXIF XResolution, TIFF XResolution with ResolutionUnit, PNG pHYs, or JP2 resc then resd; when none is stated, or it is zero or absurd, the caller's fallback applies and the choice is reported. The page is the image's pixels divided by its resolution, in points.
  • Orientation: the EXIF or TIFF orientation tag is applied by the page's transformation matrix, never by /Rotate — four of its eight values are mirror images, which /Rotate cannot express — so the image stream stays the file's bytes. The page box is swapped for the transposed orientations.
  • Page size limits: a page past 14,400 units on a side uses UserUnit (PDF 1.6, raised under ADR 40); PNG 16 bits and JPX raise the output to 1.5. Each raise is reported.
  • Metadata: a JPEG's EXIF segment — camera serial, GPS position — stays in the stream by default, since byte identity is the point for evidence; stripping it is an option that changes the bytes and is reported. Removing metadata as a policy is M19's.
  • CMYK photographs, TIFF photometric interpretation MinIsBlack under CCITT (which needs /BlackIs1 true, or the page prints as a negative), and identical images added twice (shared by M06's deduplication) are handled by the table above and tested.

Outlines​

PdfOutline the editable outline of a document or of an assembly: add, insert, move, remove,
restyle; applied when saved, like page edits (M06)
PdfOutlineItem title, target (a page and a fit, a named destination, a remote destination, a URI,
or another action carried as it is), open state, color, bold, italic, children
PdfOutlineJson Export(outline, Utf8JsonWriter) and Import(Utf8JsonReader) -> PdfOutline
  • The JSON form is versioned and published as a JSON Schema on the documentation site: { "version": 1, "items": [ { "title", "page", "fit", "left", "top", "zoom", "open", "color", "bold", "italic", "children" } ] }, with destination for a named destination, file and page or destination for a remote one, uri for a link. page counts from 1, and an export also writes the page's label, ignored on import. An action JSON cannot express — JavaScript, Launch — is exported with its kind only and imported as an item without an action, reported.
  • Written and read with Utf8JsonWriter and Utf8JsonReader, no reflection, so Native AOT and trimming hold; import bounds depth, item count and title length, since the JSON may come from anywhere.
  • Titles are text strings: decoded as #36 fixed them in M03, written in PDFDocEncoding when it can hold them and in UTF-16BE (UTF-8 in 2.0 output) otherwise.
  • Export walks the outline as M06 does — bounded, with a visited set; a cycle is cut and reported.
PdfPieceIdentity the names a part was known by before assembly — its file name and aliases
PdfLinkRewriter after assembly: remote links to a part rewritten as internal; after a split: internal
links across parts rewritten as remote
PdfOpenParameters parse and format PDF open parameters (RFC 8118): page, nameddest, and the view
parameters that map to a destination (zoom, view, viewrect)
  • Which links: GoToR with a file specification; Launch whose file is a PDF (how Word and PDFMaker link to another file — the Cabinet of Horrors' externalLink.pdf does exactly that); URI actions whose URI is relative or file: and whose path names a PDF, with or without open parameters.
  • Matching a file to a part: the file specification is decoded (/UF over /F, then the legacy /DOS, /Mac, /Unix), percent-decoded for URIs, its separators and . and .. segments normalized, and its last segment compared with each part's identities, ignoring case by default, or its whole relative path in strict mode. Two parts answering to one name is ambiguity: the link is left as it is and reported.
  • Which destination: a GoToR explicit destination names a page by number, not by reference — mapped through the target part's page map; a named destination goes through M06's rename map for that part; #page= and #nameddest= alike; no destination means the part's first page.
  • Unresolved remote links are left alone — the file may travel beside the volume — and reported; the caller may ask for them to be removed.
  • GoToE targets are resolved through their chain (/T with /R /C and /N into the embedded files, or /A through a file-attachment annotation; /R /P to the parent), bounded. When the embedded PDF became a part — after unpacking a portfolio, or after the caller merged an attachment — the link becomes internal. Creating a GoToE to a page or a named destination of an attached PDF is part of the API.
  • After a split, the reverse (above): internal links across parts become GoToR with named destinations, so reassembling the parts brings every link back to internal.

Portfolio unpacking​

PdfPortfolio read: the collection's schema (fields, types, order, visibility), sort, initial
document, view, folders; its members with their collection items
PdfPortfolio.Unpack -> an M06 assembly plan: one part per PDF member, one bookmark per member
  • Order: the collection's /Sort fields and directions, ties broken by name-tree order. Folders (PDF 1.7 extension level 3, and PDF 2.0) group members, depth first, folders ordered by name; the folder a member belongs to is written as a <id> prefix on its name-tree key, which is removed from any title. A navigator (Flash, deprecated by PDF 2.0) is reported and ignored.
  • Bookmark titles: the schema field the caller names; otherwise the member's /Desc, otherwise its file name. A member's own outline is nested under its bookmark.
  • Members: PDF members become parts. Image members become image pages, when the caller asks. Other members stay attached to the volume, their relationship Unspecified, and are listed. The portfolio's own cover pages, which typically ask the reader to open the file in Acrobat, are dropped by default and reported.
  • Signed and certified members: unpacking treats members as evidence — each is carried unsigned with its original attached as AFRelationship /Source on its first page (M06's AttachOriginal), and reported; a caller can choose Refuse instead. Nothing is ever voided silently. This default, a certified member included, is decided (the maintainer, 2026-09-27), beside M06's refusal of a certified part.
  • Nested portfolios unpack recursively, to a bounded depth. Encrypted members are refused until M16 and reported; the rest of the portfolio unpacks.
  • Memory: a member whose embedded stream is not compressed is read in place, through a window on the portfolio's own file. A compressed member cannot be read lazily, so it is decoded to a scratch stream — a temporary file by default, one the caller can replace — and never into memory whole, under the reader's decoding limit.
  • Active content inside members — 3D, rich media, JavaScript — is carried as M06 carries it, and reported (ADR 37).

Corpus formats beyond PDF​

The manifest describes PDF files only today: its file patterns end in .pdf, and every acceptance test reads every entry as a PDF. Image files are the first inputs of another format that a milestone must be accepted on, and XML invoices (M14), FDF and XFDF (M16), messages (M18) and DOCX (M31) follow. One manifest describes them all, so each is described once and held to one rule:

  • format, beside file: pdf when absent; jpeg, png, tiff, jp2 (M07), xml (M14), fdf, xfdf (M16), eml, msg (M18), docx (M31). The committed and the remote file patterns are keyed on it, so that a .pdf path under another format, or the reverse, fails the schema.
  • A per-format expectation block in expect, beside the PDF fields, which apply to pdf alone: an image's pixel size, resolution and EXIF orientation, from Pillow; an XML invoice's syntax, profile and KoSIT's verdict; an exchange file's fields and values; a message's headers, body strings and attachment names with their SHA-256; a DOCX's headings, tables and the Word PDF of the same version it is compared with. Each comes from an independent tool, never from the library. M07 defines all five, and a later milestone may add fields to its own block with its first files, never change another's.
  • One provenance rule: a non-PDF file lives where a PDF of its origin would — documents/<use-case>/, vendor/<source>/, remote/<source>/ —, under the same license, personal-data and 2 MB rules, and fetch_remote.py fetches and verifies it as it does a PDF.
  • The PDF acceptance tests filter on format: CorpusReadingTests and every test that opens a document read exactly the entries they read before; each format has its own acceptance tests, in the milestone that reads it.
  • The model in TestSupport/Corpus.cs, CorpusManifestSchemaTests and docs/corpus.md change with the schema. What is written from the standards rather than received — M10's payload set — is test data in TestSupport, not a corpus document.

Report codes and the tool​

  • Each operation reports under its own family — split.*, image.*, outline.*, link.*, portfolio.* — beside M06's assembly.*, in M06's PdfOperationReport, stable and documented.
  • The tool gains split, collate, images (from image files to a PDF), outline export and outline import, and unpack, each with --json and M06's exit codes. Every verb accepts the selection grammar wherever M06's accepted ranges.

Slices​

Each slice ends on a green commit with its acceptance rows passing and its report codes documented.

  1. Page selection and collation. Delivers PdfPageSelection with labels and bookmarks, PdfCollation. Proven by unit tests on every construct, FsCheck (the parser is total; a selection resolves to indices within the document; reversing twice with z-1 is the identity), and by collating fronts and reversed backs that qpdf derives in a container from the corpus's CCITT scans, compared page by page with the original. Leaves splitting.
  2. Pruning. Delivers PdfResourceUsage and pruning on M06's extraction. Proven by unit tests (names only in a nested form, in a Type 3 glyph, in a pattern, in an annotation appearance, in an inline image, in a BX/EX section; a cycle of forms; an undecodable content stream keeping everything) and by pikepdf's remove_unreferenced_resources, in a container, finding nothing left to remove from our extractions. Leaves the parts to the split slices.
  3. Split by ranges, page count and label range. Delivers PdfSplitter, the part writer (labels, fields, structure, named destinations, attachments per part), naming, IPdfSplitSink. Proven by unit tests of each per-part decision and by qpdf, pdftotext and pikepdf on the parts of the corpus's labeled and form documents. Leaves bookmarks, size and separators.
  4. Split by bookmark, and cross-part links. Delivers the outline strategy with its degenerate cases and the conversion of cross-part links to GoToR with named destinations. Proven by unit tests (front matter, two items on one page, items out of page order, items without destinations, a cyclic outline) and by pdf.js and pikepdf reading each part's outline and each rewritten link. Leaves reassembly, which slice 7 proves.
  5. Split by size and by separator. Delivers the size strategy with its estimate and its single re-cut, and the separator strategy with blank-by-content. Proven by unit tests (a cap below one page, a resource shared across the cut, a separator with an empty annotation) and by the parts' sizes on the 1000-page journal and, remotely, on US Code Title 42.
  6. Outlines and JSON. Delivers PdfOutline editing, the JSON form and its published schema. Proven by unit tests (every target kind, styles, deep and wide trees, hostile JSON), FsCheck (export then import is the identity on any outline JSON can express), and pdf.js reading the re-imported outline of every committed document that has one.
  7. Links between pieces. Delivers PdfPieceIdentity, PdfLinkRewriter in both directions, PdfOpenParameters and GoToE. Proven by unit tests on file-specification matching (case, separators, .., percent-encoding, the legacy keys, ambiguity) and on every destination form, by the Cabinet's link-and-target pair, and by splitting then reassembling the VA guide and comparing every link target with the original's. Leaves the corpus's missing cross-linked set (below) as a gap.
  8. Corpus formats beyond PDF. Delivers the manifest's format field, the file patterns keyed on it, the five expectation blocks, the filter in every PDF acceptance test, the model in TestSupport/Corpus.cs, fetch_remote.py for every format, and docs/corpus.md; then the image files of the table below. Proven by CorpusManifestSchemaTests accepting a sample entry of each format and refusing a .pdf path under png and a .png path under pdf; the PDF acceptance tests reading exactly the entries they read before; fetch_remote.py verifying a remote image's hash and size as it verifies a PDF's. Leaves the images themselves.
  9. Images: JPEG and JPEG 2000. Delivers the container parsers, page geometry, EXIF orientation by matrix, ICC and Adobe CMYK handling, streaming. Proven by unit tests (every marker, truncated and oversized segments, all eight orientations, JFIF units 0, 1 and 2, a multi-chunk ICC profile) and by poppler's pdfimages -list and -all, in a container, on a page per JPEG and JPEG 2000 stream of the corpus.
  10. Images: TIFF. Delivers the IFD walker and the TIFF column of the table: frames, strips, CCITT parameters, predictors, PackBits, fill order, photometric interpretation. Proven by unit tests (an IFD chain with a cycle, a strip offset past the end, a million declared frames) and by pages made from every CCITT stream of the corpus wrapped as a TIFF strip, whose image poppler decodes pixel for pixel as it decodes the source's.
  11. Images: PNG. Delivers the chunk walker, IDAT pass-through, color-key masks, and the lossless paths for interlacing and alpha. Proven by unit tests (a bad CRC, an IDAT split across many chunks, tRNS in every color type, 16 bits) and, once the image set below exists, by rendering against Pillow's decoding in a container.
  12. Portfolio unpacking. Delivers PdfPortfolio, Unpack, scratch streams for compressed members, signed and certified members' policy, recursion. Proven by unit tests on schemas, sort orders, folder prefixes, a portfolio that contains itself, and by the Acrobat 9 portfolio in the remote corpus.
  13. The tool's verbs. Delivers split, collate, images, outline, unpack. Proven by in-process handler tests and by CorpusToolTests running the AOT binary against the API's own results.

Tests required​

Unit — tests/AdCodicem.Pdf.Tests:

  • Selection: every construct, reversed ranges, exclusions, labels and bookmarks that do not exist, integers past int; FsCheck: the parser never throws anything but its parse error, and resolution stays within the document.
  • Pruning: every place a resource name can hide; conservative on anything unknown; never on a content stream that failed to decode.
  • Splitting: each strategy's boundaries, degenerate outlines, the size estimate and its single re-cut, naming (reserved characters, device names, NFC, collisions), each per-part decision of M06's policies.
  • Images: each container parser on valid, truncated, oversized and cyclic input; the mapping of every row of the table; resolution fallback; the eight orientations as matrices; page-size limits and UserUnit.
  • Outlines: editing operations; JSON both ways; FsCheck round trip; hostile JSON (depth, size, duplicate keys, invalid UTF-8).
  • Links: file-specification matching, destination mapping, fragments, GoToE chains, ambiguity, unresolved links; the split-then-reassemble round trip on synthetic documents.
  • Portfolios: schemas, sorting, folders, nested portfolios, compressed and uncompressed members, members that are not PDFs, a member that is the portfolio itself.
  • Hostile, throughout: nothing read from an image or a PDF sizes an allocation unchecked; every walk is bounded; every failure is a report entry or a typed exception, inside time and allocation budgets.

Integration — tests/AdCodicem.Pdf.IntegrationTests, every referee in a container (ADR 27):

  • qpdf — --check and --show-npages on every part and every image volume; the fronts and backs that collation starts from, derived by qpdf itself.
  • pikepdf — outlines, link actions and their targets, named destinations, field lists per part, and remove_unreferenced_resources finding nothing to remove.
  • pdf.js — each part's outline, and the outline re-imported from JSON.
  • poppler — pdftotext per page of each part; pdfimages -list for image dimensions, encodings and resolution; pdfimages -all for byte identity and decoded pixels; pdftoppm for oriented pages.
  • Pillow, in a container — the reference decoding and EXIF transposition of the image files, once the corpus holds them.

Acceptance conditions​

DocumentsBehaviorVerified by
vendor/us-federal/finereader8-frb-sr0115-examiner-guidance.pdf (four CCITT pages) and vendor/us-federal/acrobat3-import-irs-1040-1988-scan.pdf (two), split by qpdf into fronts and reversed backsInterleaving restores the original order, each page's image stream byte-identical to the original'sCorpusSplitTests.Collating_fronts_and_reversed_backs_restores_the_scan
vendor/lu-legilux/fop22-legilux-memorial-seal-renewed-timestamps.pdf (40 pages sharing one resource dictionary), finereader8-frb-sr0115-examiner-guidance.pdf, the LibreOffice invoice and reportsExtracting any page carries only what its content uses; pikepdf finds nothing left to removeCorpusSplitTests.Extracted_pages_carry_no_unused_resource
vendor/opf-format-corpus/word9-distiller405-usgs-nwql-volatile-organics-methods.pdf (86 pages, 14 items), vendor/pdf-association/indesign13-pdfua1-german-book-chapter.pdf (21, 7), vendor/opf-format-corpus/pdfmaker707-word-va-esig-developer-guide.pdf (49, 5), pdfmaker7-powerpoint-va-cancer-database-course.pdf (53, 52), pdfmaker707-word-law-library-iraq-legal-history.pdf (33, 2), vendor/eu-publications/3heights-eu-consolidated-regulation-2026.pdf (2, 2)Split by top-level bookmark: each part's page count as qpdf counts it, its outline the source subtree as pdf.js reads it, each page's text the source page's in pdftotext, no unused resourceCorpusSplitTests.Splitting_by_top_level_bookmark_yields_one_part_per_exhibit
word9-distiller405-usgs-nwql-volatile-organics-methods.pdf (two unbookmarked front pages, two items on page 59), indesign13-pdfua1-german-book-chapter.pdf (two items on page 2), vendor/fr-licence-ouverte/pdfmaker-acrobat-cerfa-12156-form.pdf (an item out of page order), vendor/pdf-association/handwritten-utf16le-strings.pdf (three items on its only page)Front matter in a part of its own, same-page items joined, parts in page order, each case reportedCorpusSplitTests.Degenerate_outlines_split_predictably
pdfmaker707-word-va-esig-developer-guide.pdf (80 internal links)Split by bookmark, every link crossing parts is a GoToR naming the right part and a named destination that part holds; reassembled with the parts' identities, every link resolves to the page it did in the originalCorpusLinkTests.Links_across_parts_become_remote_and_come_back_on_reassembly
pdfmaker-acrobat-cerfa-12156-form.pdf (406 fields, a calculation order)Each part carries exactly the fields whose widgets it holds, as pikepdf lists themCorpusSplitTests.Split_parts_carry_only_their_own_fields
word9-distiller405-usgs-nwql-volatile-organics-methods.pdf (unlabeled pages, then lower roman, then decimal)Split by label range: three parts, each keeping its labels in pdf.jsCorpusSplitTests.Splitting_by_label_range_follows_the_labels
documents/stress/reportlab-journal-1000-pages.pdf at 128 KB and 256 KBEvery part under the cap, every page once and in order, the same parts on every runCorpusSplitTests.Splitting_by_size_respects_the_cap
Remote: remote/govinfo/us-code-2023-title42.pdf (9,302 pages) at 32 MB, remote/opf-format-corpus/pdfmaker81-word-va-kernel-systems-guide.pdf (464 pages) by bookmarkSame, with memory following the largest page, not the documentCorpusSplitTests.Splitting_the_largest_documents_holds_its_budget
A scan batch with blank separator pages — not in the corpus (below)Split at each separator, separators dropped, every other page onceCorpusSplitTests.Splitting_at_separators_drops_them
The 56 JPEG streams of the committed corpus taken out as files, four of them once their Flate layer is removed — baseline and progressive, JFIF, EXIF and Adobe APP14, from InDesign, PDFMaker, LibreOffice, FOP, PDFlib and copiersOne page each; pdfimages -list gives the file's width, height and resolution (or the reported fallback), pdfimages -all returns the file byte for byteCorpusImagePageTests.Every_jpeg_becomes_a_page_without_re_encoding
The JP2 stream of vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf; remote, remote/pdfjs/acrobat8-jpx-precincts-issue5475.pdf (several precincts, and no color space in its PDF) and the 125 JPEG 2000 plates of remote/usgs/omnipage-usgs-professional-paper-1-1902.pdfOne page per image, bytes untouched; the 125-page volume built with memory following the largest header, not the imagesCorpusImagePageTests.Every_jpeg_2000_becomes_a_page_without_re_encoding
Every CCITT stream of the committed corpus — 70, G4 with and without EndOfBlock, and the G3 one-dimensional stream with EndOfLine of vendor/pikepdf/scanner-ccitt-endofline.pdf — each wrapped by the test as a single-strip TIFF carrying the stream's own parametersOne page each, its image stream byte-identical to the strip, and decoded by pdfimages pixel for pixel as the source's image isCorpusImagePageTests.Every_ccitt_strip_becomes_a_page_without_re_encoding
Photographs in all eight EXIF orientations, CMYK and ICC-tagged JPEGs, PNGs of every color type, bit depth, interlacing and transparency, multi-frame TIFFs in every compression the table names — not in the corpus (below)Each as the table says: bytes untouched where it says so, lossless rewrites reported, orientation and size as Pillow's transposed decoding shows themCorpusImagePageTests.Every_image_file_becomes_a_page_as_the_table_says
Every committed document with an outline, the UTF-16LE titles of handwritten-utf16le-strings.pdf and the cyclic outline of vendor/pikepdf/handwritten-cyclic-toc.pdf among themExported to JSON and imported onto a copy stripped of its outline, the outline pdf.js reads is the original's, titles and targets alike; the cycle cut and reportedCorpusOutlineTests.Outlines_round_trip_through_json
vendor/opf-format-corpus/pdfmaker9-word-distiller-external-link.pdf, whose link launches text_only_pdfa1b.pdf, assembled with vendor/opf-format-corpus/pdfmaker9-word-distiller-pdfa1b-test-document.pdf declared under that nameThe link becomes an internal GoTo to the second part's first page, in pikepdf and pdf.jsCorpusLinkTests.A_link_to_another_piece_becomes_internal
A set of pieces linking to one another by GoToR with page and named destinations, and by URIs with #page= and #nameddest= — not in the corpus (below)Assembled, every link internal and on the right page; left alone, and reported, when its target is not in the setCorpusLinkTests.Remote_links_between_pieces_become_internal
Remote: remote/opf-format-corpus/acrobat9-portfolio-signed-3d.pdf (five PDF members in folders, one of them a certified XFA form, 3D content)One volume, five bookmarks named from the collection and grouped by folder, the cover dropped, the certified member carried unsigned with its original attached, the 3D content carried, each reported; memory bounded by the scratch policyCorpusPortfolioTests.The_portfolio_unpacks_into_one_bookmarked_volume
A committed portfolio — not in the corpus (below)Same, in the main CI jobCorpusPortfolioTests.A_committed_portfolio_unpacks
Every split, image volume and unpacking aboveTwo runs give identical bytesCorpusSplitTests.Splitting_is_deterministic, CorpusImagePageTests.Image_pages_are_deterministic
The same operations through the toolThe AOT binary produces what the API producesCorpusToolTests.Split_images_outline_and_unpack_match_the_api

The remote rows close only on a green Remote corpus run, recorded in status.md with its date.

Corpus​

What the corpus holds​

  • Outlines to split on: eight committed documents with two or more top-level items, among them every degenerate case the design names — front matter before the first item, two items on one page, an item out of page order, three items on a single page, a cycle; remote, long bookmarked documents (pdfmaker81-word-va-kernel-systems-guide.pdf, 464 pages; jhove-hul-138-pdftex-dissertation.pdf, 120 pages with labels).
  • Links to rewrite: 80, 26 and 15 internal GoTo links in three committed documents; one real Launch to another PDF whose target is itself in the corpus under another name.
  • Labels to split on: the USGS methods report's unlabeled, roman and decimal ranges.
  • Resources to prune: FOP's forty pages sharing one resource dictionary, and LibreOffice's shared root resources; the 1000-page journal for size caps; US Code Title 42 remote.
  • Images inside PDFs: 56 JPEG streams, four of them under a Flate layer — 44 baseline and 8 progressive among those read, JFIF and Adobe markers, one grayscale and none CMYK, three with EXIF resolution tags and only one orientation tag, which says upright; 70 CCITT streams (69 G4, four of them with EndOfBlock, one G3 one-dimensional with EndOfLine); one JP2 committed and 125 remote.
  • Scans to collate: four- and two-page CCITT scans from which qpdf derives fronts and backs at test time.
  • A portfolio: Acrobat 9's, remote, with folders, a schema, a certified member and 3D.

What it lacks​

NeedWhyPriorityLikely source
Image files as corpus documents, and a manifest that can describe them — the schema admits .pdf only todayThe roadmap's acceptance, "every image file in the corpus becomes a page", has no image file to run on; the streams taken out of PDFs cover JPEG, JPEG 2000 and CCITT only1Slice 8's format field, decided for M07, M14, M16, M18 and M31 at once; then the files below
JPEGs in all eight EXIF orientations, with and without JFIF densityOrientation by matrix is the most visible failure an image page can have, and the corpus's only orientation tag says upright1A public source: recurser/exif-orientation-examples (MIT)
PNGs of every color type, bit depth, interlacing and transparencyPNG pass-through and its lossless paths have no input at all1A public source: Willem van Schaik's PngSuite, free for any use
Multi-frame TIFFs: CCITT G3 1-D and 2-D, G4, Modified Huffman, LZW, Deflate with a predictor, PackBits, several strips, FillOrder 2, MinIsBlack, JPEG in TIFFTIFF is what copiers and scanning services deliver; only CCITT strips taken out of PDFs exist1Generated here with libtiff's tiffcp from committed scan pages, recorded in build_corpus.py; a copier's own TIFF as a contribution (W03)
A set of pieces cross-linked by GoToR (page and named destinations) and by URIs with #page= and #nameddest=The rewrite's main case is untested but for one Launch1Generated here: LibreOffice with relative cross-document links exported as GoToR, and Chromium from HTML with href="piece-2.pdf#page=3"
CMYK and YCCK JPEGs with an Adobe APP14 marker, and a JPEG with a multi-chunk ICC profileThe inverted /Decode and ICCBased paths have no real input; the corpus's JPEGs are RGB but one grayscale, and none carries a profile2Generated here with Pillow and libjpeg-turbo, or ImageMagick; a print shop's CMYK photograph as a contribution
JPEG 2000 files: a raw codestream, a JP2 with an alpha channel, a JPX-branded fileOne committed JP2 cannot cover the color and alpha cases2Generated here with OpenJPEG's opj_compress
A committed portfolio, with folders, a schema, PDF and non-PDF membersThe only portfolio is remote, so the unpacking row runs nightly only; the committed-portfolio row above needs it1Generated here with pikepdf, a recorded construction; or a public source under an attribution-only license
A scan batch with blank separator sheetsThe separator strategy has no real batch, and the separator row above needs one1Derived here with qpdf, empty pages inserted between committed scans; a real copier batch with scanned blank sheets as a contribution, for M22
A real single-sided duplex scan, fronts and reversed backs as two filesCollation is proven on halves qpdf derives, not on what a feeder produces3A contribution (W03)
A document with GoToE links into an attached PDFGoToE resolution has no real input3A contribution, or a public source; failing that, hand-written and marked so

Traps​

  • Outlines are not in page order, and two items can share a page. A splitter that assumes the order of the outline is the order of the pages loses pages or makes empty files.
  • A split part is a whole document: labels, fields, structure, named destinations, attachments and claims are decided per part, or the part is a page dump that no longer works.
  • The sum of the parts is larger than the whole: shared resources are paid once per part. A size cap must count them, and a size is known only once the part is written.
  • A GoToR destination names a page by its number, a local destination by reference. And a relative file specification is relative to the document's location, which the library does not know: identities come from the caller.
  • /Rotate cannot mirror. Four EXIF orientations are mirror images; only the matrix expresses all eight.
  • JFIF density unit 0 is an aspect ratio, not a resolution.
  • Adobe CMYK JPEGs are stored inverted; without /Decode [1 0 …] they print as negatives. So does a CCITT TIFF in MinIsBlack without /BlackIs1 true.
  • PackBits and RunLengthDecode differ on one byte: 128 is a no-op in the first and the end of data in the second.
  • TIFF LZW exists in an old, bit-reversed form that LZWDecode cannot read; it is refused, not guessed.
  • Stacked strips can show hairline seams in viewers that anti-alias image edges; joining them needs a decoder, which is M22's.
  • Byte identity keeps EXIF, GPS position included. It is the right default for evidence and the wrong one for publication; the option exists and the documentation says so.
  • A portfolio's members are compressed streams: a lazy reader cannot seek inside them, so they spill to scratch, not to memory.
  • Folder membership hides in name-tree keys, as a <id> prefix a title must not show.
  • Part file names meet Windows device names, reserved characters, combining accents and length limits.

Documentation​

  • docs/website/docs/reference/page-selection.md — the grammar, with examples, shared by the API and the tool.
  • docs/website/docs/guides/splitting.md — the strategies, per-part decisions, cross-part links, naming.
  • docs/website/docs/guides/images-to-pdf.md — supported formats, the table of what is passed through and what is rewritten, resolution, orientation, metadata.
  • docs/website/docs/guides/outlines.md — editing, and the JSON form; its JSON Schema under docs/website/static/schemas/.
  • docs/website/docs/concepts/links-between-pieces.md — identities, rewriting in both directions, open parameters, GoToE.
  • docs/website/docs/guides/portfolios.md — unpacking, titles, folders, signed members, scratch storage.
  • docs/website/docs/reference/assembly-report-codes.md — the split.*, image.*, outline.*, link.* and portfolio.* codes added.
  • docs/website/docs/reference/tool/ — the new verbs.
  • docs/corpus.md — the format field, the five expectation blocks, where a non-PDF file lives, and the image files in the corpus.
  • docs/architecture.md — the container parsers and the resource-usage scan in the core.

Exit criteria​

  • The selection grammar, collation, pruning, the six split strategies, image pages for the four formats, outline editing with JSON, link rewriting in both directions with GoToE and open parameters, and portfolio unpacking exist as designed.
  • No image codec is used, no sample value changes and no lossy step exists; every lossless rewrite is reported with the image it changed.
  • The manifest describes non-PDF formats — images, XML invoices, FDF and XFDF, messages, DOCX — through format and one expectation block each, the PDF acceptance tests filter on it, and fetch_remote.py fetches every format; the image files are in, and the priority-1 gaps above are filled; each remaining gap is recorded in docs/corpus-contributions.md.
  • The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on a green Remote corpus run recorded in status.md.
  • Unit tests cover each behavior, its degenerate cases and its hostile ones; the container parsers and the selection parser are fuzzed, seeded with the corpus, and a nightly campaign runs them.
  • Integration tests confirm parts, image pages, outlines and links through qpdf, pikepdf, pdf.js, poppler and Pillow, each in a container.
  • Benchmarks measure splitting the 1000-page journal and building an image volume, with MemoryDiagnoser; status.md records their memory against the largest page and the largest image header.
  • The tool's new verbs ship in the dotnet tool and the AOT binaries, documented.
  • The documentation site publishes the pages listed above.
  • Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).