Skip to main content

M23 — Optimization, performance, hardening

State: to do — Depends on: M12, M15, M22 — Guards classified per ADR 34, which the piecewise decode reopens as that record foresaw; native memory only where a measurement asks for it, per ADR 35; every lossy step opt-in, named and reported, JBIG2 lossless only, per ADR 42; a core with no dependency, AOT- and trimming-compatible, per ADR 9; referees in containers per ADR 27; the heavy documents remote per ADR 32

Goal​

Deliver the frugality the library promises, with numbers: budgets for time and allocation that fail the build, a decode that holds a window of a stream rather than the stream, caches that hold a budget rather than whatever they were given, files made smaller without changing what they say — and smaller still, by named lossy steps, when the caller asks — and a core proved to run under Native AOT, trimming and browser WebAssembly, and in a constrained container beside the Chromium it replaces.

Every milestone before this one measures what it adds and records the figure in docs/status.md. This one turns the figures into budgets that CI enforces, pays the debts earlier milestones deferred here — #37, #38, #46, #47, #48, #49 and #50 —, adds the operations whose only purpose is size — deduplication, consolidation and subsetting of fonts, recompression, pruning, linearization, and the lossy image steps ADR 42 left to it —, and runs the campaigns that give "hostile input" its meaning: fuzzing over the reader, the writer's round trip and the optimizer, and a measured deployment in a container with 512 MB, a read-only file system and no fontconfig.

The failures this milestone exists to prevent are specific. A change that allocates in the parser's hot loop and merges because nobody ran the benchmark. A court portal that refuses a scanned exhibit over its size cap, and a library that can only split it. An optimizer that merges two subsets sharing a tag but not their glyphs — Illustrator and Konik write exactly that — and changes a word on every page. A "lossless" pass that re-deflates a stream and changes nothing but the conformance claim, or drops the font a form field needs to be filled. A decompression bomb that the guard stops at 256 MB, and a sound 3 GB plate that no setting can read. A rebuild that reads 64 KB at every occurrence of trailer in a file its author filled with the word. A server that holds every name any hostile file ever contained, because the name table is process-wide and never forgets.

Scope​

In:

  • performance budgets in CI: allocation budgets per operation, asserted as tests on every commit; throughput budgets measured on every pull request against its base in the same job, with a tolerance fixed from measured noise; a committed budget file whose every raise is recorded with its reason (#38 closed);
  • the reader's debts: the object cache made least-recently-used and weighted, names interned without an intermediate string and the process-wide name table bounded (#37); cross-reference sections read through windows grown on demand that keep what they read (#47); the rebuild's trailer scan through the same windows, skipping stream data (#49); the cache of decoded object streams held to a byte budget (#50); a piecewise decode, so that MaxDecodedStreamLength bounds what a stream may decode to without the implementation's 2 GB ceiling, and memory held follows a window (#48); memory budgets stated and held on W11's three heavy documents (#46);
  • the optimizer in the core, PdfOptimizer: global deduplication, consolidation of duplicate font subsets, subsetting of fonts embedded in full, lossless recompression, packing into object streams, pruning of unused resources and unreferenced objects, linearization; discards by name — private application data, thumbnails, unembedding of standard fonts —; lossy image steps by name — downsampling, JPEG re-encoding, conversion to gray and to one bit per pixel — on M22's codecs, with a target-size mode; one report listing what each step saved and every change a lossy step or a discard made (ADR 42);
  • linearized output from M03's writer, as an option of any full rewrite and a step of the optimizer, verified by qpdf's --check-linearization;
  • the consumers completed: M06's merge consolidates font subsets on request; M12 applies the caller's named downsampling to images it embeds; M18's portal.file-size shapes a piece by optimization before M07's split; M03's PdfSaveOptions gains Linearize;
  • Native AOT, trimming and browser WebAssembly, proved: an AOT binary processes every corpus document and gives the JIT build's bytes; a trimmed host of the core and of each satellite that declares IsAotCompatible raises no warning; a browser-wasm host processes the committed corpus;
  • the container profile: a chiseled image, a read-only file system, 512 MB, no fontconfig and no ICU, measured on the reference workloads beside headless Chromium under the same limits, the figures published;
  • hardening: the nightly campaign extended to the writer's round trip, the optimizer and, coverage-guided, the reader's own entry points; cancellation honored within a latency budget by every long operation, hostile input included; SECURITY.md brought to what is measured;
  • the command-line tool's optimize verb, and --linearize on every verb that writes.

Out, explicitly:

  • mixed raster content compression — an open question of the roadmap, reopened by color scans too large for a portal cap after this milestone's recompression; pixel deskew and despeckle — an open question;
  • anything that changes what a document says — its text, its structure, its conformance claims, the validity of its signatures — is never an optimization: a step that would do so is refused, or runs only when the caller names it and is reported (the discards and lossy steps below); never silent;
  • color conversion — CMYK to RGB, or any ICC transform to shrink or unify color — M29's color management; M23's conversion to gray uses M22's evaluation, reports where it approximates, and is a lossy step;
  • pattern-matching or lossy JBIG2 — never (ADR 42); JPEG 2000 encoding — not planned (M22);
  • rasterizing a page to shrink it, flattening transparency — not planned; rendering — M25, whose rasters are not needed here (MuPDF is the visual referee);
  • shrinking a document by incremental update — impossible, since an update only appends; the optimizer always writes a full rewrite, and refuses a signed document unless the caller insists (M04's guard);
  • native codecs through P/Invoke (libjpeg-turbo, zlib-ng directly) — excluded by invariant 1 in the core and by ADR 42 in the managed Imaging satellite; a measurement showing a managed codec too slow for its budget would reopen ADR 42 under ADR 35's conditions;
  • WebAssembly for the satellites that carry native code — .Html and .Rendering (SkiaSharp, HarfBuzzSharp); M25 leaves SkiaSharp's WebAssembly assets untested and so does M23;
  • budgets for M24 to M31 — each of those milestones adds its rows to the budget file as it lands, under the rules fixed here.

Design​

Where it lives​

PartWhereWhy
The budget file, the budget benchmarks, the budget checkerbenchmarks/budgets.json, benchmarks/AdCodicem.Pdf.Benchmarks, a Budgets job in ci.ymlBudgets are data with a history; the checker is a small console step, not a test framework
Allocation budgets as teststests/AdCodicem.Pdf.Tests (CorpusPerformanceTests), and each satellite's test projectDeterministic on one runtime, so they run on every commit like any test
The object cache, name interning, windows, the rebuild's scan, the object-stream cache, the piecewise decodeCore, IO/ and Objects/The reader's own internals; the public surface grows only by PdfStream.DecodeTo, OpenDecoded, IPdfDecodeSink, two options, and MaxDecodedStreamLength as a long
PdfOptimizer, its options, steps, plan and reportCore, Optimization/, namespace AdCodicem.Pdf.OptimizationLossless steps need the writer, the fonts and the copier, all core; lossy steps run on the codecs the caller passes, as M22's consumers do
The JPEG encoder (forward DCT, quantization, chroma subsampling)AdCodicem.Pdf.ImagingBeside M22's decoder and baseline re-encoder, which it shares; the core never encodes JPEG
Resampling, gray and bilevel conversionCore, Images/They work on M22's rows and need no codec
LinearizationCore, M03's writer (IO/Writing/)Only the writer knows positions (CLAUDE.md, Known traps)
The AOT host, the trimming host, the WebAssembly hosttests/AdCodicem.Pdf.AotHost, tests/AdCodicem.Pdf.TrimHost, tests/AdCodicem.Pdf.WasmHostTest projects, never packaged; the AOT host is the tool's own binary where the tool covers the operation
The container profiledeploy/container/ (the Dockerfile and its measurement script), a Container profile workflowA deployment the documentation publishes and CI measures
The fuzzing targetstests/AdCodicem.Pdf.Fuzzing (M22's harness) and FuzzingTestsOne harness, one campaign, one record
optimize, --linearizeAdCodicem.Pdf.ToolThe tool ships the Imaging satellite (M22)

Budgets​

benchmarks/budgets.json one entry per (operation, document, metric): the value, the tolerance, the unit, the
measurement it was set from (commit, runner class, date), and the reason for each raise
BudgetChecker reads BenchmarkDotNet's JSON results and the file; exits non-zero naming each budget
exceeded, by how much, and the measurement the budget came from
  • Allocation is a budget of bytes, asserted as a test. On one runtime, allocation is deterministic up to a few hundred bytes of first-call initialization: each test warms the operation once, then measures it with GC.GetAllocatedBytesForCurrentThread, as CorpusReadingTests already does for indexing (4 MB budget, 3.2 MB measured) and DocumentReaderTests for an index of 300,000 objects (32.6 and 45.2 MB budgets, 31.1 and 43.1 MB measured). An operation that runs on several threads — M12's parallel batch, M06's merge with parallel hashing — runs in a child process and is measured with GC.GetTotalAllocatedBytes(precise: true), since the per-thread counter misses what other threads allocate. Tolerance: 5 % over the recorded figure, no more.
  • Throughput is a ratio, measured where the noise cancels. Shared runners vary by more than any regression worth catching. The Budgets job therefore builds the pull request's head and its base, and runs the budget benchmarks of both alternately in the same job, several rounds each, on BenchmarkDotNet's fixed short job; it fails when the ratio of medians exceeds 1 + the tolerance. The tolerance is set by slice 1 from the spread of that ratio over twenty runs of one commit against itself, and recorded; the expectation is under 10 %.
  • Drift across many small regressions is caught by the nightly Benchmarks run on main, which compares absolute figures with the committed baseline under a wider tolerance and fails the workflow beyond it.
  • What is budgeted, each row from the benchmark its milestone wrote: indexing, reading every page and validating the thousand-page journal (ReaderBenchmarks, ValidationBenchmarks); a full rewrite and an incremental update (WriterBenchmarks); a merge of the journal with itself (AssemblyBenchmarks); text extraction (ExtractionBenchmarks); layout of the reference documents and a thousand invoices (LayoutBenchmarks, BatchBenchmarks, whose CI floor M12 set at half the first measurement and which the ratio replaces); decoding per codec (ImagingBenchmarks); and this milestone's OptimizerBenchmarks and StreamDecodingBenchmarks.
  • Raising a budget is an edit of budgets.json with its reason, and a line in status.md; a test fails on an entry whose raise has no reason. A budget set too tight teaches contributors to raise it: each is set from a measurement with its tolerance, never from a wish.
  • benchmarks/README.md changes accordingly: the full suite stays on demand; the budget subset runs on every pull request.

The object cache and names (#37)​

  • Least recently used. The cache keeps its slots in arrays — a dictionary from object identifier to slot, and previous and next indices per slot —, so that a hit moves an entry to the front and an eviction takes the back, O(1) and 0 B per access. ObjectCacheCapacity stays the count it is; a weight is added: each entry costs its length in the file, which the reader knows, and the cache evicts when either the count or PdfReaderOptions.ObjectCacheBudget (bytes, 64 MB by default) is exceeded, so that 8,192 dictionaries of 16 MB each can no longer be held at once. An entry larger than the budget is returned and not cached.
  • Names without an intermediate string. The parser decodes a name's #xx escapes into a stack or pooled buffer, and looks it up through the dictionary's alternate lookup on ReadOnlySpan<char> (StringComparer.Ordinal implements it since .NET 9); a string is allocated only for a name seen for the first time.
  • The process-wide table is bounded. PdfName.Get interns into a static ConcurrentDictionary that grows with every distinct name any file ever contained and never shrinks: in a server reading hostile uploads, a file of a million distinct names is a million strings held for the life of the process, against invariant 2 and invariant 8's spirit. From this slice: the names ISO 32000 defines — the well-known names, and every key the Arlington model lists (ADR 44) — are interned process-wide in a frozen table built once; every other name is interned in a table per document, released with it. Equality is by value already, so no caller sees a difference but memory; reference equality stays the fast path for the names that matter. Measured in slice 2: the process's heap after opening a million-name file and disposing it returns to its level before.

Windows grown on demand (#47, #49)​

  • Cross-reference sections are read through a window that starts small — 4 KB — and grows geometrically to the section's end or MaxXRefSectionLength, appending what it reads rather than reading again from the start: the VHA coding handbook's 273 KB table then costs 273 KB, not 683 KB, and the signed Web Capture file's three sections cost what they weigh, not three 64 KB windows. The laziness test's bound — opening reads less than a quarter of the file — then holds for both documents recorded unsupported for #47, and their markers go.
  • The rebuild's trailer scan uses the same window per occurrence, starting at 1 KB, grown only while the dictionary is open, up to MaxTrailerLength, so that the cost per occurrence is what the trailer weighs; an occurrence of trailer inside a stream's data — which the scan already delimits between stream and endstream — is skipped. The 80,000-occurrence file (625 KB) then reads a bounded amount per keyword, and a sound trailer longer than 64 KB is no longer given a syntax error the file does not have. What the parser meets past a window it did not grow is dropped with the attempt, as T21 and T23 settled.

Decoded object streams under a budget (#50)​

  • The cache of decoded object streams in PdfFileReader gains PdfReaderOptions.ObjectStreamCacheBudget, in decoded bytes (64 MB by default, checked against the corpus's largest object streams in slice 3), evicting the least recently used stream; an evicted stream is decoded again when one of its objects is asked for, and yields the same objects.
  • The rebuild no longer holds what it indexed. It decodes each object stream to index its objects, records each object's stream and index, and releases the decoded bytes under the same budget. The damaged file of four object streams that each decode to 256 MB then holds at most the budget once Open returns, instead of 1 GB.
  • The budget is not a bound in ADR 34's sense: reaching it costs time — a second decode —, never data, so it has no limit.* code and no document is refused by it. A single stream larger than the budget is decoded, used and not kept.

The piecewise decode (#48)​

IPdfDecodeSink Write(ReadOnlySpan<byte> window); Complete(PdfDecodeResult): windows in order, each valid
only during the call
PdfStream.DecodeTo DecodeTo(IPdfDecodeSink, CancellationToken) — push; memory held is one window per stage
PdfStream.OpenDecoded OpenDecoded() -> Stream — pull, read-only, forward-only, over the same pipeline
PdfStream.Decode unchanged: the whole stream in memory, for callers that want it and streams that fit
  • Every filter becomes a stage that consumes an input window and produces output windows into pooled buffers — Flate over the BCL's inflater, LZW with its 4,096-entry table, RunLength, ASCIIHex, ASCII85, the PNG and TIFF predictors with their one row of history, decryption block by block (M16 wrote it so), and M22's image filters, whose row sinks are already this shape. The window is 64 KB per stage; a predictor holds one row as well, whose width /Columns declares and which the decoded-length guard already bounds.
  • The Flate stage keeps what #56 made the whole decode keep. After a fault it reads its input again from the start — the stream's data, or the stages before it run again — and skips what it already handed out, so no window repeats. A zlib checksum that disagrees with whole data is known only once every window is out: Complete carries filter.checksum-mismatch, and a consumer that re-encodes — M03's writer, the optimizer's recompression — never gives such data a fresh checksum in silence. One form of the data can give way to another after it kept bytes: raw deflate read from leading white space, which turns corrupt, to zlib after the white space, which reads further. Data that starts with white space is therefore read once to choose its form before a window goes out, rather than taking back windows already handed out.
  • MaxDecodedStreamLength keeps bounding what one stream may decode to, becomes a long, and PdfReaderLimits.Unbounded sets it to long.MaxValue. Streaming does not make a decompression bomb harmless: a 1 MB stream of nested Flate that decodes to a terabyte costs an internal consumer — the optimizer, the content interpreter, a validator's rule — the time to produce a terabyte, which invariant 4 forbids. What changes is what is held: one window, whatever the length. ADR 34 foresaw that a streaming decode would reopen it, and read that MaxDecodedStreamLength would then bound memory held rather than length; this milestone keeps it on length, for the time it bounds, and records the reason in an ADR at the next free number amending ADR 34.
  • Decode() past Array.MaxLength keeps what an array holds and reports stream.too-large-to-hold, whose message names DecodeTo: that is the implementation's ceiling, not a guard, so no property lifts it and it has no limit.* code. Under Unbounded, DecodeTo reads such a stream whole.
  • Every internal consumer moves to the piecewise path: M03's writer when it re-encodes, the optimizer, M15's content interpreter (content streams are read a window at a time, a token never spanning more than the lexer's bound), M02's stream rules, M06's hashing for deduplication, M14's attachment extraction. Offsets and lengths are long wherever a stream's decoded size is counted.
  • Native memory only if measured. The pooled windows are small and managed; ADR 35 applies only if a benchmark shows a stage cannot meet its budget in managed code, and none is expected to.
  • Cancellation is checked between windows, so a stream decoding for minutes stops within one window's work.

The optimizer​

PdfOptimizer Plan(document, PdfOptimizerOptions) -> PdfOptimizationPlan: what each step would do and
save, without writing; Optimize(document, output, options, IProgress<PdfProgress>?,
CancellationToken) -> PdfOptimizationReport
PdfOptimizerOptions immutable: Steps (lossless, all by default), Discards (none by default), LossySteps (none
by default), TargetSize, Codecs (PdfImageCodecs; the core's by default), Linearize,
SaveOptions (M03's, for version and cross-reference form), SignaturePolicy and
ConformancePolicy (M04's and M09's, Refuse by default)
PdfLosslessStep [Flags] Deduplicate, ConsolidateFontSubsets, SubsetFonts, Recompress, PackObjects,
PruneUnusedResources, DropUnreferenced
PdfDiscard [Flags] PrivateApplicationData, Thumbnails, UnembedStandardFonts
PdfLossyStep a closed set of records: Downsample(threshold ppi, target ppi, method),
ReencodeJpeg(quality, subsampling), ConvertToGray, ConvertToBilevel(threshold)
PdfTargetSize bytes, the lossy steps it may use in order, each with its floor (a lowest resolution, a
lowest quality)
PdfOptimizationReport per step: objects affected, bytes before and after; per image a lossy step or a discard
touched: the object, the pages that draw it, encoding, size, dimensions and effective
resolution before and after, and the step; whether the target was met, and if not, the
smallest size reached; every diagnostic

Three classes of change, recorded in an ADR at the next free number that extends ADR 42's rule beyond images:

  • Lossless — the document renders, extracts, validates and claims exactly what it did. On by default.
  • Discards — data no viewer shows but an application or a person may want: Illustrator's and PDFMaker's /PieceInfo, page /Thumb images and XMP thumbnails, standard fonts embedded where a viewer would substitute them. Never by default; each named; each object discarded in the report.
  • Lossy — the pixels change. Never by default; each named; each image changed in the report (ADR 42).

Lossless steps​

  • Deduplicate. M06's keys, applied to one document rather than to a merge's inputs: a stream's key is SHA-256 over its dictionary without /Length and its encoded bytes; a dictionary's or array's over its canonical form with each reference replaced by its target's key; a cycle falls back to identity. Two additions: a decoded key for streams whose encodings differ — the same logo once in Flate and once uncompressed, the same sRGB profile under four encodings, which M06 already compares decoded — keeping the smaller encoding; and a pass over the whole document, so that a producer's own duplicates (SAP NetWeaver's logo embedded twice, Photoshop's profile twice) are found. Never shared, whatever their bytes: page objects and page content streams (M09 edits content in place), anything carrying /StructParent or /StructParents, annotations, form fields, optional content groups and membership dictionaries, structure elements, outline items, destinations, article beads, signature dictionaries, and streams whose crypt filters differ (/Identity metadata beside encrypted streams). Each of those has an identity its bytes do not carry: two layers with equal dictionaries are two layers.
  • ConsolidateFontSubsets. Two embedded programs are the same face when their names without tag, their font types and every glyph present in both — outlines, advances, and for TrueType the instructions — are equal, compared through M08's parsers; a tag is never trusted, since Illustrator and Konik reuse one tag over different subsets. The union is written once, by M08's subsetter, under a new tag, and each font dictionary keeps its own codes, widths and ToUnicode: a CIDFontType2 through its own CIDToGIDMap into the union; a CIDFontType0 or a Type1C by glyph name or CID, which the union keeps; a simple TrueType only when its codes reach the same glyphs through the union's cmap, and kept apart otherwise. No content stream is rewritten. Under PDF/A-1, CIDSet and CharSet list what the union holds. Two programs sharing a tag whose common glyphs differ are reported (optimize.subset-tag-reused) and kept apart.
  • SubsetFonts. A program embedded in full is cut to the glyphs the document draws with it — found by M15's interpreter over every page's content, form XObjects, tiling patterns, Type 3 glyph procedures, annotation appearances and soft-mask groups, whatever the render mode — under a new tag. Kept whole: a font in the AcroForm's /DR, or named by any field's or free-text annotation's /DA, since a later fill may need any glyph (optimize.font-kept-whole); a font whose fsType forbids subsetting (M08 reads it).
  • Recompress. Every stream re-encoded when the result is smaller, deterministically: uncompressed, LZW, RunLength and ASCII-armored streams to Flate; Flate re-deflated at the BCL's SmallestSize; images given the PNG predictor that a fixed heuristic chooses per image; one-bit images in the smaller of G4 and JBIG2 generic (M22's Smallest policy); JPEG images losslessly re-entropy-coded with optimized Huffman tables by M22's baseline re-encoder, coefficients untouched. A stream whose re-encoding is not smaller keeps its bytes. Never: LZW under any output (Flate only, as M03 writes), JPEG 2000 under PDF/A-1, or a filter the output's version lacks. Embedded files are recompressed like any stream; their /Params /Size and /CheckSum describe the decoded file and stay true.
  • PackObjects. Object and cross-reference streams through M03's ObjectStreams.Generate, except under a PDF/A-1 claim (M03 already refuses them there).
  • PruneUnusedResources. Names in a resource dictionary that no content of its owner uses — found by the same interpretation — removed, per page and per form XObject; M07 did it per extracted page, this does it per document. A resource dictionary shared by several owners keeps the union of their uses.
  • DropUnreferenced. M03's full rewrite drops what nothing reachable refers to already; the optimizer reports it as its own step, with the bytes it saved.

Discards​

  • PrivateApplicationData: the /PieceInfo dictionaries of pages, form XObjects and the catalog, with the /LastModified entries that date them. Illustrator will no longer open the file as editable, and the report says so.
  • Thumbnails: page /Thumb streams and xmp:Thumbnails in XMP (M14's model).
  • UnembedStandardFonts: an embedded program of one of the fourteen standard faces — recognized by M08's parser, not by its name alone — removed, the font dictionary kept with its widths. Refused under a PDF/A, PDF/UA or PDF/X claim, which require every font embedded; run on such a document only under ConformancePolicy.RemoveClaim, the claim then removed and reported at ConformanceLoss (invariant 7).

Lossy image steps and the target size​

  • Where pixels come from. M22's PdfImageReader over the codecs in the options: the core decodes CCITT and the core's filters; PdfImagingCodecs.Default adds JBIG2, JPEG and JPEG 2000. An image the given codecs cannot decode is left as it is and reported (optimize.image-not-decodable).
  • Downsample an image whose highest effective resolution across every placement (M15's inventory) exceeds a threshold, to a target — 150 ppi above 225 by default when the step is named —, by area averaging in integer arithmetic for gray and color and by majority for one-bit images; JPEG 2000 at a coarser resolution level when that lands at or above the target (M22's reduced-resolution decoding), which costs a quarter of the work. Masks (/SMask, a stencil /Mask) are resampled onto the same grid as their image, /Matte kept; a color-key /Mask is kept as ranges. An image shared by a thumbnail and a full page is judged by its largest placement.
  • ReencodeJpeg at a quality, with the IJG scaling of Annex K's tables and 4:2:0 or 4:4:4 subsampling, through the JPEG encoder added to the Imaging satellite; it applies to images that are not bilevel. A JPEG is re-encoded from its decoded samples once: the target-size mode tries each quality from the original, never from its last attempt, so no image suffers two generations.
  • ConvertToGray by M22's evaluation (reported as approximate where M22 approximates); ConvertToBilevel by Otsu's threshold over the image's histogram, clamped to a band (M22's blank-page method), then encoded by the Smallest policy — the largest win on a scanned letter, and the step most likely to lose a light signature, which is why it is named and reported.
  • Never lossy, whatever is named: image masks (already one bit), images inside a signed revision, an image an /Alternates entry names, JBIG2 written with symbols.
  • The target size. The lossless steps first; then each named lossy step in the caller's order, each tried at successive settings down to its floor; after each, the size is re-estimated. Estimation does not write the document: each image is encoded at the candidate setting and only its length kept, the rest of the file weighed once; the final write encodes again at the chosen settings — the same bytes, since encoders are deterministic — so memory stays one image's working set rather than the output. If the floors are reached above the target, the report gives the smallest size reached, and the caller decides (optimize.target-not-met). Deterministic: the same document, target and steps give the same bytes.

Linearization​

  • PdfSaveOptions.Linearize on any full rewrite, and PdfOptimizerOptions.Linearize: the output follows ISO 32000-2 Annex F — the linearization dictionary, the first-page cross-reference section, the first page's objects, the primary hint stream with its page offset and shared object hint tables, then the rest in page order.
  • Two bounded passes in the spirit of ADR 39: the first serializes every object into a counting sink and keeps only its length and the pages that use it — a few bytes per object —; the second writes. The hint stream's length feeds the offsets it describes, so its encoding is iterated to a fixed point, at most three times, then padded. The output is forward-only: a non-seekable stream works, as for any rewrite.
  • A linearized output stays linearized until updated: an incremental update leaves it stale, as every update of a linearized file does and viewers accept (M03's trap); the report says so when the caller updates one.

Signed, encrypted and conforming inputs​

  • Signed: optimization is a full rewrite, which invalidates every signature; M04's guard refuses it with PdfSignatureInvalidationException unless SignaturePolicy is AllowInvalidatingSignatures, which reports each signature it breaks (write.signature-invalidated). Nothing is optimized by update.
  • Encrypted: opened with its password (M16); optimization needs M16's Modify permission under PdfPermissionPolicy.Respect; the output is encrypted as the input was, unless the caller's save options say otherwise.
  • Conforming: a PDF/A, PDF/UA or Factur-X claim is kept only when it stays true, checked by M20's IPdfConformanceChecker when one is registered — every lossless step keeps it by construction and the tests prove it —; a step that would break a claim is refused under ConformancePolicy.Refuse, or run with the claim removed and reported under RemoveClaim.

Consumers completed​

  • M06: PdfAssemblyOptions.ConsolidateFontSubsets, off by default, runs the consolidation over the assembled volume before it is written.
  • M12: PdfRenderOptions.ImageSteps takes the named Downsample and ReencodeJpeg steps, applied as an image is embedded; the image.resolution-excessive report M12 writes then says what was done rather than what could be.
  • M18: portal.file-size shapes a piece first by lossless optimization, then by the lossy steps the caller's case-file options name — never by the preset's choice —, then by M07's split when the preset allows it.
  • M03: PdfSaveOptions.Linearize, above.

Native AOT, trimming and WebAssembly​

  • AOT. The tool's AOT binary (M06) — and, for operations the tool lacks, tests/AdCodicem.Pdf.AotHost published with PublishAot — opens, validates, rewrites, merges with itself and optimizes every committed and remote corpus document; its output is byte-identical to the JIT build's. On linux-x64 on every pull request, linux-arm64 and win-x64 nightly.
  • Trimming. tests/AdCodicem.Pdf.TrimHost references the core and each satellite that declares IsAotCompatible, calls their public entry points, and is published trimmed with TrimmerSingleWarn off and warnings as errors: a warning attributed to any of our assemblies fails the build. The satellites that ship native code (.Html, .Rendering) are checked for AOT the same way; where AngleSharp or a binding warns and the warning cannot be suppressed where it fires with a justification (ADR 29), M12's rule stands — the verb ships in the dotnet tool only — and the measurement is recorded.
  • Browser WebAssembly. tests/AdCodicem.Pdf.WasmHost, a browser-wasm application whose exported functions take a document's bytes and return the outputs' hashes, loaded in headless Chromium driven by Playwright in a container. It opens, validates, rewrites and optimizes losslessly every committed corpus document from memory — there is no file system — and gives the JIT build's decoded content. Bytes are compared where no stream was re-deflated; where the optimizer recompressed, decoded streams are compared, since the browser runtime's deflater may not be the one the x64 runtime ships (to verify, and recorded). What the run proves or refutes about the platform: SHA-256 (managed on browser-wasm, per Microsoft's cross-platform cryptography table); MD5 and AES absent, so that the encrypted documents open under R2 to R6 through the core's managed MD5, RC4 and AES (M06's and M16's, not constant-time) and an AES-GCM file of M16's (R7) is refused with PrimitiveUnavailable; Brotli for WOFF2, whose proof M08 leaves to this run — the host decodes M08's WOFF2 test fonts to the sfnt the JIT build produces —; and the absence of threads.

The container profile​

  • The image: .NET 10's chiseled runtime-deps image, mcr.microsoft.com/dotnet/runtime-deps:10.0-noble-chiseled — not its -extra variant, which adds ICU —, pinned by digest, the AOT tool binary copied in; no shell, no package manager, no fontconfig, no ICU (InvariantGlobalization, which M15 wrote its Unicode tables for); a non-root user; --read-only with no writable mount, since the library writes no temporary file; --memory 512m; no network. The HTML engine's native assets are SkiaSharp.NativeAssets.Linux.NoDependencies, SkiaSharp's build without fontconfig, and HarfBuzzSharp.NativeAssets.Linux (that it links nothing the chiseled image lacks: to verify with ldd in slice 10); fonts come from ADR 11's registry and embedded set, never the system.
  • The workloads: the reference documents rendered from sources/invoice-fr.html, report-fr.html and contract-fr.html; a thousand invoices from one compiled template; opening, validating, extracting and optimizing the thousand-page journal; and, nightly, W11's three heavy documents.
  • The measurement: wall time, CPU time, and peak memory as the container's cgroup records it, per workload, cold and warm; the image's size. Beside it, headless Chromium rendering the same HTML through PuppeteerSharp — the slot the comparison benchmarks reserved — in a container under the same limits, given the same OFL fonts; where Chromium fails under 512 MB, that is the result.
  • The published figures state the runner, the date, the versions and the digest; docs/website/docs/guides/deployment.md gives the profile as a recipe.

Hardening​

  • Fuzzing. The nightly Fuzzing workflow gains targets: the writer's round trip (a mutated document that opens is rewritten, reopened and rewritten again, and the two rewrites are identical), the optimizer (a mutated document optimizes or fails typed, and its output opens with no repair), the piecewise decode against the whole decode (the same bytes, window boundary by window boundary). M22's coverage-guided engine, chosen by its ADR, is applied to the reader's entry points — the lexer, the object parser, the cross-reference readers, the object-stream reader, the rebuild — as it is to the decoders. The seeds are every damaged document, committed and remote. M25's rendering target and M27's parsers join the same campaign when they land. A finding becomes a regression test in HostileInputTests before it is fixed.
  • Cancellation latency. Every long operation — open with rebuild, decode, validate, rewrite, merge, extract, render, optimize — observes its token between units of bounded work (a window, an object, a page, a record), so that a canceled operation returns within 100 ms on the budget runner, hostile input included: a decode bomb under Unbounded, a rebuild of the 80,000-keyword file, a type 4 function of a million operators.
  • SECURITY.md states the threat model as measured: the guards and their defaults, the fuzzing campaign and its record, the container profile, and what a caller must set to read untrusted input in a server.

Bounds, classified (invariant 12, ADR 34)​

BoundClassWhy
What one stream may decode toThe existing guard MaxDecodedStreamLength, now a longA valid file can exceed any value; the cost of a decode is proportional to it
What Decode() returns wholeImplementation ceiling, Array.MaxLength, reported as stream.too-large-to-holdNot a choice: DecodeTo reads past it
A decoding stage's windowInternal constant, 64 KBIndependent of the file; a predictor's row is bounded by the decoded-length guard
Cross-reference windows, trailer windowsThe existing guards MaxXRefSectionLength, MaxTrailerLengthGrowth stops at them; the start size is a constant
Object cache count and weight; object-stream cacheOptions, not guards: ObjectCacheCapacity, ObjectCacheBudget, ObjectStreamCacheBudgetReaching them costs a re-read, never data
The per-document name tableBounded by the names the file holdsEach entry is a name the parser read under MaxObjectLength
The optimizer's deduplication index32 bytes and a number per shared objectBounded by the index the reader holds
The linearization's fixed-point iterationsInternal constant, 3Our own serialization converges or is padded
The target size's settings per stepThe caller's floors, stepped by a constantThe caller's input, not the file's

Diagnostics​

In PdfDiagnosticCodes, disjoint from rule identifiers (ADR 36):

CodeSeverityMeaning
stream.too-large-to-holdWarningDecode() asked for a stream past Array.MaxLength; the message names DecodeTo
optimize.subset-tag-reusedWarningTwo programs under one subset tag whose common glyphs differ; kept apart
optimize.font-kept-wholeInformationA font a form field or a free-text annotation may need; not subset
optimize.image-not-decodableInformationAn image the given codecs cannot decode; left as it is
optimize.lossy-appliedInformationOne per image a lossy step changed, with the step and the figures
optimize.discardedInformationOne per object a discard removed
optimize.target-not-metWarningThe floors were reached above the target; the smallest size reached
optimize.conformance-claim-removedConformanceLossA step run under RemoveClaim broke a claim; the claim is gone
write.linearization-staleInformationAn update to a linearized file leaves its hints stale

The command-line tool​

adpdf optimize FILE -o OUT [--lossless-only] [--steps dedupe,fonts,subset,recompress,pack,prune] [--discard private,thumbnails,standard-fonts] [--downsample 150@225] [--jpeg 75] [--gray] [--bilevel] [--target-size 10MB] [--linearize] [--allow-invalidating-signatures] [--report report.json]; --linearize on every verb that writes. Exit codes are M06's; optimize.target-not-met exits 1 with the output written. The AOT binary produces what the API produces.

Slices​

Each slice ends on a green commit, with its codes documented, its budget rows in benchmarks/budgets.json and its measurements in docs/status.md.

  1. Budgets in CI (#38). Delivers budgets.json, BudgetChecker, the Budgets job with its A/B runs, the allocation budgets of every existing benchmark as CorpusPerformanceTests, the nightly drift check, the tolerance measured and recorded, benchmarks/README.md rewritten, the rule that a raise carries a reason, and an ADR at the next free number recording the change of CI policy — benchmarks, run on demand only until now, gain a budget subset on every pull request —, with the A/B ratio chosen over absolute figures and the reasons. Proved by unit tests of the checker (a result within, at and beyond tolerance; a missing entry; a raise without a reason); a throwaway branch that allocates one array per token in the lexer and one that sleeps a microsecond per object, each failing the job, their runs recorded in status.md. Leaves the reader's debts.
  2. The object cache and names (#37). Delivers LRU with weights, ObjectCacheBudget, span lookup, the frozen process-wide table from the well-known names and the Arlington keys, per-document tables. Proved by unit tests (eviction order under hits; an entry larger than the budget; a name written with #xx escapes interned as the same name written plain; two documents' names equal by value); a property — for any sequence of gets, the cache returns what an uncached reader returns —; ReaderBenchmarks at 0 B per repeated name; the million-name file measured: the heap returns to its level after disposal. Leaves windows.
  3. Windows and the object-stream budget (#47, #49, #50). Delivers growing windows that append, the rebuild's scan through them skipping stream data, ObjectStreamCacheBudget, a rebuild that releases what it indexed. Proved by WindowEdgeTests extended (a section growing across every boundary gives the entries a whole read gives); a trailer inside a content stream not taken; a sound 100 KB trailer rebuilt without a syntax error; the four-stream file holding the budget; CorpusReadingTests.Opening_does_not_read_the_content_of passing on the two remote #47 documents with their markers removed, on a green Remote corpus run. Leaves the piecewise decode.
  4. The piecewise decode (#48). Delivers IPdfDecodeSink, DecodeTo, OpenDecoded, every filter as a stage, MaxDecodedStreamLength as a long, stream.too-large-to-hold, the internal consumers moved, the ADR amending ADR 34, StreamDecodingBenchmarks. Proved by an FsCheck property — for every filter chain, predictor and window size, the piecewise output equals Decode()'s, including windows of one byte —; a synthetic stream of nested Flate decoding to 2.5 GB read whole under Unbounded and its SHA-256 checked, holding one window; the USGS map's image through DecodeTo under the recorded hold (remote); the reader-limits page and ReaderLimitsTests updated, and the manifest schema's readerLimits.maxDecodedStreamLength, capped at 2³¹ − 1 today, raised to what a long holds, CorpusManifestSchemaTests with it. Leaves the optimizer.
  5. The optimizer's frame and lossless structure. Delivers PdfOptimizer, options, plan, report, the three classes and their ADR, Deduplicate, PackObjects, DropUnreferenced, PruneUnusedResources, the signed, encrypted and conforming policies; the manifest's optimization expect fields — distinct decoded streams, duplicate groups, fonts embedded in full, unreferenced objects — and their schema, written by build_corpus.py from pikepdf, pdffonts and qpdf. Proved by unit tests per step on a document that needs it and one that does not (no change, same bytes); two equal optional content groups kept apart; streams differing only by crypt filter kept apart; integration: qpdf --check, pdftotext, MuPDF rasters and veraPDF unchanged on every committed document. Leaves fonts.
  6. Fonts. Delivers ConsolidateFontSubsets and SubsetFonts, M06's option, optimize.subset-tag-reused and optimize.font-kept-whole. Proved by unit tests (two subsets of one face unioned, each dictionary's codes intact; two faces under one tag kept apart; a /DR font kept whole; a PDF/A-1 CIDSet rewritten); integration: pdffonts lists fewer programs, all sub yes; pdftotext's text and MuPDF's rasters identical; veraPDF's verdicts unchanged. Leaves recompression.
  7. Recompression and linearization. Delivers Recompress over every filter, the JPEG entropy re-coding, the G4 and JBIG2 choice, PdfSaveOptions.Linearize, the two passes, write.linearization-stale. Proved by a property — any stream recompressed decodes to its original bytes —; JPEG coefficients read back by libjpeg in the Python container identical; integration: qpdf --check-linearization on every linearized output, including the 26 committed files whose own hints are inconsistent, broken or stale; pikepdf's decoded streams identical. Leaves lossy steps.
  8. Discards, lossy steps and the target size. Delivers the three discards, resampling, gray and bilevel conversion, the JPEG encoder in the Imaging satellite, ReencodeJpeg, the target-size mode, M12's and M18's consumers. Proved by unit tests (no lossy step without a name; a JPEG never re-encoded twice; a mask on its image's grid; a shared image judged by its largest placement; a target met; a target not met, reported); integration: pdfimages -list agrees with the report on every changed image; our JPEGs decode in Pillow; MuPDF's renderings before and after within the structural-similarity threshold slice 8 fixes and records; pdftotext's text identical. Leaves the platforms.
  9. Native AOT, trimming and WebAssembly. Delivers the AOT, trimming and wasm hosts, the nightly matrix, the platform findings recorded. Proved by the AOT and wasm rows below; the trimming host's publish with no warning. Leaves the container.
  10. The container profile. Delivers deploy/container/, the Container profile workflow, the Chromium slot of the comparison benchmarks, the published figures and the deployment guide. Proved by the container row below. Leaves hardening.
  11. Hardening. Delivers the new fuzzing targets, the coverage-guided reader targets, cancellation checks where the latency test finds them missing, SECURITY.md. Proved by CorpusCancellationTests over every long operation and the hostile cases above; fourteen consecutive nights of the campaign with no open finding, recorded. Leaves the whole.
  12. The heavy documents, the verb, the whole (#46). Delivers the memory budgets of W11's documents, measured on the AOT binary as a child process, optimize and --linearize, OptimizerBenchmarks, the documentation. Proved by the remote rows below on a green Remote corpus run; CorpusToolTests.

Tests required​

Unit — tests/AdCodicem.Pdf.Tests, the Imaging satellite's under Imaging/ as M22 placed them:

  • Budgets: the checker's arithmetic and messages; every entry of budgets.json has a unit, a tolerance, a source measurement and, when raised, a reason.
  • Caches: LRU order; weights; the budget with one entry larger than itself; eviction during a rebuild; an evicted object stream re-decoded to the same objects; a property over random access sequences.
  • Names: span lookup allocation-free for a known name; a new name allocated once per document; the frozen table immutable; names from two documents equal by value and hash.
  • Windows: sections growing across each boundary; a section exactly at MaxXRefSectionLength; a trailer split at each byte; trailer inside stream data, inside a string, inside a comment.
  • Piecewise decode: every filter and predictor, alone and chained; windows of 1, 7, 4,096 and 65,536 bytes; ASCII85's ~> split across windows; LZW's early change at a window's edge; a TIFF predictor with 2-bit samples across a window; a Flate stream that lost its tail (T32's report preserved); one that turns corrupt and one whose checksum disagrees (#56's kept bytes and reports preserved); decryption across windows; cancellation between windows.
  • Optimizer: each step on a document that needs it and on one that does not; every "never shared" kind; each discard refused under each claim; the report's figures against the output; determinism — two runs, two cultures, the same bytes.
  • Fonts: TrueType, CFF, Type1C and CID-keyed unions; a simple TrueType whose codes disagree kept apart; fsType forbidding subsetting; glyphs reached only from an annotation's appearance or a Type 3 procedure kept.
  • Lossy: area averaging and majority against hand-computed grids; Otsu's threshold on known histograms; the JPEG encoder's output decoded by our decoder within the quantization error; target-size search order and floors.
  • Linearization: the hint tables of a synthetic document checked field by field; a one-page document; a document whose first page shares every object; the fixed-point padding.
  • Hostile: a decode bomb of nested Flate under default limits (the guard at 256 MB) and under Unbounded (canceled within the latency budget); a million distinct names; 80,000 trailer keywords; four object streams of 256 MB; a font program claiming 65,535 glyphs over 1 KB; an image declaring 65,535 × 65,535 pixels offered to Downsample; a document whose every page shares one resource dictionary of 10,000 entries — each ends in output, a report or a typed exception, within its time and allocation budget.
  • Fuzzing: the new targets join FuzzingTests per commit with fixed seeds and the nightly campaign with the night's seed; every finding a regression test first.

Integration — tests/AdCodicem.Pdf.IntegrationTests, every referee in a container (ADR 27):

  • qpdf — --check on every output; --check-linearization on every linearized output; --show-object --filtered-stream-data for decoded bytes of heavy streams;
  • pikepdf — decoded streams hashed before and after; duplicate groups counted independently; fonts' FontFile* objects counted;
  • poppler — pdftotext before and after, identical for lossless steps and for lossy image steps; pdffonts for embedding and subsetting; pdfimages -list for encodings, sizes and resolutions against the report;
  • MuPDF — mutool draw rasters before and after: identical for lossless steps, within the recorded similarity threshold for lossy ones;
  • veraPDF — the verdict on every claimed level unchanged by lossless optimization and by the lossy steps;
  • Pillow and libjpeg — our JPEGs decoded; coefficients unchanged by entropy re-coding;
  • Playwright with Chromium — the WebAssembly host; headless Chromium through PuppeteerSharp — the container comparison;
  • M14's Factur-X referee — the invoices' XML still validates after their attachments are recompressed.

Acceptance conditions​

"The stress documents" are documents/stress/reportlab-journal-1000-pages.pdf and the benchmarks' synthetic thousand-page document; "the heavy documents" are W11's three remote references, remote/govinfo/us-code-2023-title42.pdf (9,302 pages), remote/usgs/us-topo-washington-west-2023.pdf (one page, a 328,608,000-byte image, opened under its readerLimits) and remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf (125 pages of JPEG 2000, 147 MB). Remote rows close only on a green Remote corpus run, recorded in status.md with its date.

DocumentsBehaviorVerified by
The stress documentsEvery budget of budgets.json holds — allocation on every commit, throughput on every pull request against its base —, and a change beyond a tolerance fails CICorpusPerformanceTests.Budgets_hold_on_the_stress_documents (new), the Budgets job
The heavy documents (#46)Opening, walking every page, validating, extracting text, rewriting and optimizing losslessly each stay within the peak-memory budget recorded in status.md, flat across pages, measured on the AOT binary as a child processCorpusPerformanceTests.Heavy_documents_stay_within_their_memory_budgets (new)
remote/usgs/us-topo-washington-west-2023.pdf (#48)Its image decoded through DecodeTo with the process's managed heap under 64 MB, its SHA-256 equal to qpdf's filtered stream data; Decode() under its raised limit still gives the same bytesCorpusStreamDecodingTests.A_large_image_decodes_a_window_at_a_time (new)
A stream decoding past 2 GB, generated at test time — not in the corpus (below)Read whole under PdfReaderLimits.Unbounded, holding one window, its length and hash exact; cut and reported as limit.decoded-stream under the defaultsCorpusStreamDecodingTests.A_stream_past_two_gigabytes_is_read_whole_under_unbounded (new)
remote/pdfcpu/acrobat-web-capture8-x509-rsa-sha1-signed.pdf and remote/opf-format-corpus/pdfmaker707-word-vha-coding-handbook.pdf, the two documents recorded unsupported for #47Opening reads less than a quarter of each file; their unsupported markers removed, which unblocks M21's and M27's rows on the firstCorpusReadingTests.Opening_does_not_read_the_content_of
Every damaged document, committed and remote (#49)Recovered as before, with no syntax error invented at a trailer longer than a window, reading a bounded amount per trailer keywordCorpusReadingTests.Damaged_documents_are_recovered_as_far_as_an_independent_tool_recovers_them
The 53 committed documents the manifest marks object-streams, and remote/pdfjs/cairo-firefox-objstm-index-overflow-bug1978317.pdf (65,542 objects in one stream) (#50)Every object reads the same under an object-stream budget of 64 KB, forcing eviction, as under the defaultCorpusReadingTests.Every_object_reads_the_same_under_a_small_object_stream_budget (new)
Every committed document the reader opens, but the signed ones, and the remote ones nightlyLossless optimization never grows the file; qpdf --check passes; pdftotext's text, pdfimages -list's images and MuPDF's rasters are unchanged; veraPDF's verdict on every claimed level is unchangedCorpusOptimizationTests.Lossless_optimization_changes_nothing_a_referee_sees (new)
remote/pdfminer/sap-netweaver-invoice-issue1062.pdf (one logo embedded twice), remote/ocrmypdf/photoshop-cc2015-pdfx3-cmyk.pdf (one ICC profile twice), the M06 merges of the corpusEvery duplicate pikepdf's decoded-stream hashes find is written once, and nothing else merged; the bytes saved equal the report'sCorpusOptimizationTests.Duplicate_resources_are_written_once (new)
vendor/us-federal/illustrator-irs-pub1-english.pdf and remote/zugferd-corpus/konik-pdfbox-zugferd1-basic-from-word.pdf (one tag over different subsets), remote/pdf-association/abledocs-pdfua1-tagged-textbook-scan.pdf (a placeholder tag shared), and a merge of committed documents subsetting one face — not in the corpus (below)Programs consolidated only where their common glyphs agree, optimize.subset-tag-reused where they do not; pdffonts lists fewer programs; text and rasters identicalCorpusOptimizationTests.Font_subsets_are_consolidated_only_when_their_glyphs_agree (new)
vendor/eu-publications/pdflib-oj-exchange-rates-greek.pdf, antenna-house-oj-exchange-rates-2019.pdf, vendor/uk-ogl/pdfmaker21-ozev-sample-invoice.pdf, vendor/opf-format-corpus/pdfmaker9-word-distiller-fonts-embedded-in-full.pdf; remote, remote/zugferd-corpus/symtrax-itextsharp-zugferd21-minimum-pdfa3a.pdf (PDF/A-3a)Every program embedded in full subset to the glyphs drawn, tagged, pdffonts reading sub yes; text and rasters identical; the PDF/A claims veraPDF upheld still upheldCorpusOptimizationTests.Fonts_embedded_in_full_are_subset_to_the_glyphs_used (new)
The committed forms whose /DR embeds a font — vendor/fr-licence-ouverte/pdfmaker-acrobat-cerfa-12156-form.pdf (Arial in full, 572 KB, 406 fields) and vendor/us-federal/omniform-usda-rd1924-5-hidden-widgets.pdf (Verdana)Every font a field names kept whole; M16's fill of every field after optimization gives the appearance it gave beforeCorpusOptimizationTests.Fonts_a_field_needs_are_kept_whole (new)
The LZW documents vendor/us-federal/acrobat3-import-irs-1040-1988-scan.pdf, distiller3-irs-ss4-1995-form.pdf, pdfwriter4-usda-dry-whey-standard-2000.pdf; the uncompressed attachments of vendor/zugferd/gnuaccounting-mustang10-zugferd-rc-invoice.pdf and pypdf2-facturx-python-false-pdfa3b.pdf; remote, the uncompressed images of remote/opf-format-corpus/illustrator-distiller601-mac-nida-scholastic-heads-up.pdfEvery stream Flate-encoded where smaller, pikepdf's decoded bytes identical; the Factur-X XML still accepted by M14's referee; /Params still trueCorpusOptimizationTests.Recompressed_streams_decode_to_the_same_bytes (new)
Every committed document, linearized; documents/archival/qpdf-linearized-report.pdf as qpdf's own referenceqpdf --check-linearization reports no error; the 26 committed files whose own hints are inconsistent, broken or stale come out soundCorpusLinearizationTests.Linearized_output_passes_qpdf_s_check (new)
The scans: vendor/us-federal/xerox-workcentre-treasury-imf-report-scan.pdf (JBIG2), acrobat3-import-irs-1040-1988-scan.pdf (CCITT at 400 ppi), documents/scan/reportlab-scanned-receipt.pdf (DCT), vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf (JPX), acrobat11-image-conversion-pdfa1b-image.pdf; remote, the Konica scan remote/ecan/konica-bizhub-c554e-letter-scan.pdf and the heavy documentsNo lossy step runs unless named; with each named, every changed image is in the report as pdfimages -list sees it after; text identical; MuPDF's rasters within the recorded threshold; veraPDF's verdicts unchangedCorpusOptimizationTests.Lossy_steps_run_only_when_named_and_are_reported (new)
remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf, remote/usgs/us-topo-washington-west-2023.pdf; a color scan over a portal's cap — not in the corpus (below)Given a target and the steps allowed, the output fits and says how, or the report gives the smallest size reached; the same bytes on two runs; memory flat per imageCorpusOptimizationTests.The_target_size_is_met_or_the_report_says_why (new)
vendor/us-federal/illustrator-irs-pub1-english.pdf (/PieceInfo, an XMP thumbnail), pdfwriter4-usda-dry-whey-standard-2000.pdf (page thumbnails), vendor/fr-licence-ouverte/fop-dictao-dila-signed-joafe-notice.pdf (standard fonts not embedded)Each discard removes only what it names, each object in the report; UnembedStandardFonts refused on a PDF/A claimCorpusOptimizationTests.Discards_remove_only_what_they_name (new)
The signed committed documents M04 listsOptimization refused with PdfSignatureInvalidationException; under AllowInvalidatingSignatures, each broken signature reportedCorpusOptimizationTests.Signed_documents_are_not_rewritten_unless_the_caller_insists (new)
Every committed and remote documentThe AOT binary opens, validates, rewrites, merges with itself and optimizes it, with the JIT build's bytes, on linux-x64 per pull request and linux-arm64 and win-x64 nightlyCorpusAotTests.The_aot_binary_processes_the_corpus_as_the_jit_build_does (new)
Every committed documentThe browser-wasm host, in headless Chromium, opens, validates, rewrites and optimizes it losslessly, with the JIT build's bytes or, where recompressed, its decoded streamsCorpusWasmTests.The_browser_host_processes_the_corpus_as_the_jit_build_does (new)
The encrypted committed documents, with their recorded passwords; M16's AES-GCM document; M08's WOFF2 test fontsEach document under R2 to R6 opens in the browser-wasm host through the core's managed primitives, its decrypted content the JIT build's; the AES-GCM one refused with PrimitiveUnavailable, never a crash; each WOFF2 font decoded to the JIT build's sfnt, or, should Brotli be missing, reported font-program.unsupported as M08 saysCorpusWasmTests.Encrypted_documents_and_woff2_fonts_behave_in_the_browser (new)
The reference sources, a thousand invoices, the thousand-page journal; the heavy documents nightlyEvery workload completes in the container profile under 512 MB; time and peak memory recorded beside headless Chromium's under the same limits and publishedCorpusContainerProfileTests.The_reference_workloads_run_within_512_mb (new)
Every damaged document as seeds, committed and remoteThe nightly campaign — mutation over the reader, the writer's round trip and the optimizer; coverage-guided over the reader's entry points — finds no untyped exception, hang or unbounded allocation over fourteen consecutive nights before the milestone closesFuzzingTests.Opening_a_mutated_document_either_works_or_fails_with_a_typed_exception, FuzzingTests.Rewriting_a_mutated_document_round_trips_or_fails_typed (new), FuzzingTests.Optimizing_a_mutated_document_either_works_or_fails_typed (new), the Fuzzing workflow's record
The heavy documents and the hostile files aboveEvery long operation canceled mid-way returns within 100 msCorpusCancellationTests.Every_long_operation_stops_within_its_latency_budget (new)
The same operations through the tooloptimize and --linearize produce the API's bytesCorpusToolTests.Optimize_matches_the_api (new)

Corpus​

What the corpus holds​

  • Stress: ReportLab's thousand-page journal, committed; W11's three heavy references, remote — 9,302 pages, one 63 MB page whose image decodes to 328 MB (read under readerLimits), 147 MB of JPEG 2000 —; beside them a 35,000-pixel square CCITT image decoding to 153 MB from 10.5 KB, 65,542 objects in one object stream, 48 incremental updates, an 82-page tagged scan of 10.6 MB (many-pages, single-huge-page, heavy-scan, huge-decoded-image, many-objects, forty-eight-incremental-updates).
  • The #47 documents, remote: the signed Web Capture file and the VHA coding handbook.
  • Duplicates and waste: a logo embedded twice (duplicated-image), an ICC profile twice (duplicate-icc-profile), unreferenced objects (unreferenced-image-objects, orphaned-objects, unreferenced-objects-in-update, unreferenced-indirect-name-objects), unused fonts (fonts-unused), /PieceInfo (pieceinfo-illustrator, pieceinfo-markedpdf), thumbnails (page-thumbnails, large-xmp-with-thumbnail).
  • Fonts: four committed documents embedding fonts in full, and remote, Symtrax's PDF/A-3a invoice and Yousign's signed seal page (full-font-embedding, type0-full-font-embedded); subset tags reused over different programs (reused-subset-tag, placeholder-subset-tag-shared); a Cerfa form whose /DR embeds Arial in full, an OmniForm form whose /DR embeds Verdana, and two signed documents whose /DR embeds Myriad Pro.
  • Compression: LZW in four committed documents, ASCII85 chains, uncompressed attachments, uncompressed images (remote), 53 committed documents with object streams, and many classic tables.
  • Linearization: documents/archival/qpdf-linearized-report.pdf from qpdf, and 58 committed linearized files, 26 of them with hints inconsistent, broken or stale, several linearized then updated.
  • Scans: every codec M22 decodes, committed and remote.
  • Damage: 13 committed and 106 remote damaged documents, the fuzzing seeds.

What it lacks​

NeedWhyPriorityLikely source
A stream that decodes past 2 GBThe Unbounded acceptance needs one; no real file reaches the ceiling, and committing one is pointless1Generated at test time by the test support — nested Flate over zeros, about 2 MB encoded —, never committed, and not a corpus document: docs/corpus.md names it as the exception to committing
A case file merged from several producers' documents, each subsetting the same faceFont consolidation needs real subsets of one face by different subsetters; the corpus's reused tags are the negative case1Generated here: the LibreOffice, Chromium and ReportLab renderings of invoice-fr.html in one OFL face, merged by qpdf, recorded in build_corpus.py
Independent expectations for optimization: per committed document, pikepdf's count of distinct decoded streams and duplicate groups, pdffonts' embedded-in-full fonts, qpdf's unreferenced objects"Every duplicate, and nothing else" must be asserted against a count the file gave an independent tool, not our own optimizer1Generated here: build_corpus.py writes them into new expect fields, whose schema this milestone adds
A color scan too large for a court portal's cap, several pages at 300 ppi or moreThe target-size mode exists for it; the heavy scans are archival plates, not a lawyer's exhibit2A contribution (W03); remote if it cannot be redistributed
Chromium's time and memory on the reference workloads under the container's limitsThe comparison the profile is published against2Generated here, in CI, by the comparison benchmarks' Chromium slot; recorded, not committed
A document of uncompressed images and content, committedRecompression's largest gain is shown on remote files only3Generated here: ReportLab with page compression off
A document linearized by Acrobat with object streams and shared objects across pagesThe hint tables' shared-object section is proven on qpdf's output and old Distiller files3A public source among government publications (W02); remote if needed

Traps​

  • An optimization that changes what a document says is a corruption, not a trade-off: extracted text, structure, conformance and signatures are the acceptance, not size alone.
  • A subset tag lies. Illustrator and Konik write one tag over different programs; consolidation compares glyphs, never names.
  • Identity is not bytes. Two equal optional content groups are two layers; two equal annotations, fields or structure elements are two things. Deduplication shares values, never identities.
  • A form field needs its whole font. Subsetting a font in /DR breaks the next fill, silently, in another tool.
  • CIDSet and CharSet describe the program, and PDF/A-1 checks them after every subsetting and every union.
  • Flate is deterministic per runtime, not across runtimes (M03's trap): a servicing release, or the browser runtime's own deflater, changes every recompressed stream's bytes. Compare decoded bytes across platforms.
  • Re-encoding a JPEG loses a generation each time. Every attempt starts from the original samples.
  • A mask belongs to its image's grid. Downsample one without the other and every edge shifts.
  • An image's resolution is a placement's, not the image's: one XObject drawn as a thumbnail and as a page has two; the largest decides.
  • Inline images cannot be deduplicated or resampled without rewriting content; they are left and counted.
  • The hint tables' offsets: whether an offset after the primary hint stream counts that stream's own length is where Annex F and Acrobat's practice have been reported to differ (to verify against qpdf's implementation); qpdf's --check-linearization is the referee, and the corpus's 26 inconsistent files show how often producers err.
  • A signed document cannot be optimized in place: an update only appends, and a rewrite invalidates.
  • A budget that is too tight trains contributors to raise it; one that is too loose catches nothing. Set each from a measurement with a stated tolerance, and record every raise.
  • Shared runners are noisy: allocation budgets are reliable; time budgets hold only as a ratio measured in one job.
  • GC.GetAllocatedBytesForCurrentThread misses other threads: parallel operations are measured in a child process.
  • A window that grows must keep what it read, or a large section is read again from its start at each step — the VHA handbook's 683 KB for 273 KB.
  • trailer appears inside streams, and a scan that does not know where streams are pays for each occurrence.
  • A stream past 2 GB overflows every int that counts it: positions, lengths, progress.
  • Streaming is not safety: a bomb read piecewise still costs the time to produce it; the length guard stays.
  • A process-wide intern table is a leak in a server that reads hostile input.
  • Chiseled images have no ICU and no fontconfig: globalization must be invariant, and the native graphics assets must not link fontconfig.
  • Chromium in a container needs its sandbox arrangements and shared memory; a comparison that gives it fewer resources than ours, or other fonts, is not a comparison.
  • WebAssembly has no file system, no threads by default, no MD5 and no AES — the core's managed ones stand in, and AES-GCM has no stand-in —, and its address space is 32-bit.

Documentation​

  • docs/website/docs/guides/optimization.md (new) — the lossless steps, the discards, the lossy steps, the target size, the report, signed and conforming inputs, the optimize verb.
  • docs/website/docs/concepts/performance.md (new) — the budgets, how they are measured and enforced, what they promise and what they do not.
  • docs/website/docs/guides/deployment.md (new) — Native AOT, trimming, WebAssembly, the container profile as a recipe with its measured figures beside Chromium's, reader limits for untrusted input.
  • docs/website/docs/reference/reader-limits.md — MaxDecodedStreamLength as a long; docs/website/docs/concepts/reader-limits.md and lazy-reading.md — the piecewise decode, the caches and their budgets, the per-document name table.
  • docs/website/docs/reference/diagnostics.md — the codes above.
  • docs/website/docs/reference/tool/ — optimize, --linearize.
  • docs/website/docs/introduction.md and docs/features/features.json — optimization, aot, hostile-input and cancellation brought to their state.
  • docs/architecture.md — Optimization/, the piecewise decode, the caches; SECURITY.md — the threat model as measured; benchmarks/README.md — budgets in CI.
  • docs/corpus.md, the manifest schema and tests/corpus/README.md — the optimization expect fields.
  • The ADRs: budgets in CI (slice 1), the amendment of ADR 34 (the piecewise decode, and why MaxDecodedStreamLength still bounds length), and the three classes of change (extending ADR 42's rule beyond images).
  • docs/status.md — #37, #38, #46, #47, #48, #49 and #50 closed; the budgets, tolerances and measurements; the platform findings; the campaign's record; the container figures.

Exit criteria​

  • Allocation budgets run on every commit and throughput budgets on every pull request, and each fails CI beyond its tolerance; budgets.json records every figure's source and every raise's reason; the ADR on budgets in CI is accepted.
  • #37, #38, #46, #47, #48, #49 and #50 are closed, each by the behavior its issue asked for, the #47 markers removed from the manifest.
  • The piecewise decode serves every internal consumer; a stream past 2 GB is read whole under Unbounded; the ADR amending ADR 34 is accepted.
  • Every lossless step, discard and lossy step, the target size and linearization are implemented, reported, and refused where a signature or a claim forbids them; the ADR on classes of change is accepted.
  • The AOT binary and the browser-wasm host process the corpus as the JIT build does; the trimmed host raises no warning; the platform findings are recorded.
  • The container profile is published with its figures beside Chromium's.
  • The priority-1 gaps above are filled; each remaining gap is recorded in docs/corpus-contributions.md.
  • The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on a green Remote corpus run recorded in status.md.
  • Unit tests cover each behavior, its degenerate cases and its hostile ones; the FsCheck properties hold; the campaign has run fourteen consecutive nights with no open finding.
  • Integration tests run qpdf, pikepdf, poppler, MuPDF, veraPDF, Pillow and libjpeg, Playwright with Chromium, headless Chromium and M14's Factur-X referee, each in a container.
  • OptimizerBenchmarks and StreamDecodingBenchmarks run with MemoryDiagnoser, and every budget row is in budgets.json; docs/status.md records the measurements.
  • optimize and --linearize ship in the dotnet tool and the AOT binary, documented.
  • The documentation site publishes the pages listed above, and features.json matches what exists.
  • Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).