M23 — Optimization, performance, hardening
State: to do — Depends on: M12, M15, M22 — Guards classified per ADR 34, which the piecewise decode reopens as that record foresaw; native memory only where a measurement asks for it, per ADR 35; every lossy step opt-in, named and reported, JBIG2 lossless only, per ADR 42; a core with no dependency, AOT- and trimming-compatible, per ADR 9; referees in containers per ADR 27; the heavy documents remote per ADR 32
Goal
Deliver the frugality the library promises, with numbers: budgets for time and allocation that fail the build, a decode that holds a window of a stream rather than the stream, caches that hold a budget rather than whatever they were given, files made smaller without changing what they say — and smaller still, by named lossy steps, when the caller asks — and a core proved to run under Native AOT, trimming and browser WebAssembly, and in a constrained container beside the Chromium it replaces.
Every milestone before this one measures what it adds and records the figure in docs/status.md. This one turns the
figures into budgets that CI enforces, pays the debts earlier milestones deferred here — #37, #38, #46, #47, #48,
#49 and #50 —, adds the operations whose only purpose is size — deduplication, consolidation and subsetting of
fonts, recompression, pruning, linearization, and the lossy image steps ADR 42 left to it —, and runs the campaigns
that give "hostile input" its meaning: fuzzing over the reader, the writer's round trip and the optimizer, and a
measured deployment in a container with 512 MB, a read-only file system and no fontconfig.
The failures this milestone exists to prevent are specific. A change that allocates in the parser's hot loop and
merges because nobody ran the benchmark. A court portal that refuses a scanned exhibit over its size cap, and a library
that can only split it. An optimizer that merges two subsets sharing a tag but not their glyphs — Illustrator and Konik
write exactly that — and changes a word on every page. A "lossless" pass that re-deflates a stream and changes
nothing but the conformance claim, or drops the font a form field needs to be filled. A decompression bomb that the
guard stops at 256 MB, and a sound 3 GB plate that no setting can read. A rebuild that reads 64 KB at every
occurrence of trailer in a file its author filled with the word. A server that holds every name any hostile file
ever contained, because the name table is process-wide and never forgets.
Scope
In:
- performance budgets in CI: allocation budgets per operation, asserted as tests on every commit; throughput budgets measured on every pull request against its base in the same job, with a tolerance fixed from measured noise; a committed budget file whose every raise is recorded with its reason (#38 closed);
- the reader's debts: the object cache made least-recently-used and weighted, names interned without an
intermediate string and the process-wide name table bounded (#37); cross-reference sections read through windows
grown on demand that keep what they read (#47); the rebuild's trailer scan through the same windows, skipping stream
data (#49); the cache of decoded object streams held to a byte budget (#50); a piecewise decode, so that
MaxDecodedStreamLengthbounds what a stream may decode to without the implementation's 2 GB ceiling, and memory held follows a window (#48); memory budgets stated and held on W11's three heavy documents (#46); - the optimizer in the core,
PdfOptimizer: global deduplication, consolidation of duplicate font subsets, subsetting of fonts embedded in full, lossless recompression, packing into object streams, pruning of unused resources and unreferenced objects, linearization; discards by name — private application data, thumbnails, unembedding of standard fonts —; lossy image steps by name — downsampling, JPEG re-encoding, conversion to gray and to one bit per pixel — on M22's codecs, with a target-size mode; one report listing what each step saved and every change a lossy step or a discard made (ADR 42); - linearized output from M03's writer, as an option of any full rewrite and a step of the optimizer, verified by
qpdf's
--check-linearization; - the consumers completed: M06's merge consolidates font subsets on request; M12 applies the caller's named
downsampling to images it embeds; M18's
portal.file-sizeshapes a piece by optimization before M07's split; M03'sPdfSaveOptionsgainsLinearize; - Native AOT, trimming and browser WebAssembly, proved: an AOT binary processes every corpus document and gives
the JIT build's bytes; a trimmed host of the core and of each satellite that declares
IsAotCompatibleraises no warning; a browser-wasm host processes the committed corpus; - the container profile: a chiseled image, a read-only file system, 512 MB, no fontconfig and no ICU, measured on the reference workloads beside headless Chromium under the same limits, the figures published;
- hardening: the nightly campaign extended to the writer's round trip, the optimizer and, coverage-guided, the
reader's own entry points; cancellation honored within a latency budget by every long operation, hostile input
included;
SECURITY.mdbrought to what is measured; - the command-line tool's
optimizeverb, and--linearizeon every verb that writes.
Out, explicitly:
- mixed raster content compression — an open question of the roadmap, reopened by color scans too large for a portal cap after this milestone's recompression; pixel deskew and despeckle — an open question;
- anything that changes what a document says — its text, its structure, its conformance claims, the validity of its signatures — is never an optimization: a step that would do so is refused, or runs only when the caller names it and is reported (the discards and lossy steps below); never silent;
- color conversion — CMYK to RGB, or any ICC transform to shrink or unify color — M29's color management; M23's conversion to gray uses M22's evaluation, reports where it approximates, and is a lossy step;
- pattern-matching or lossy JBIG2 — never (ADR 42); JPEG 2000 encoding — not planned (M22);
- rasterizing a page to shrink it, flattening transparency — not planned; rendering — M25, whose rasters are not needed here (MuPDF is the visual referee);
- shrinking a document by incremental update — impossible, since an update only appends; the optimizer always writes a full rewrite, and refuses a signed document unless the caller insists (M04's guard);
- native codecs through P/Invoke (libjpeg-turbo, zlib-ng directly) — excluded by invariant 1 in the core and by ADR 42 in the managed Imaging satellite; a measurement showing a managed codec too slow for its budget would reopen ADR 42 under ADR 35's conditions;
- WebAssembly for the satellites that carry native code —
.Htmland.Rendering(SkiaSharp, HarfBuzzSharp); M25 leaves SkiaSharp's WebAssembly assets untested and so does M23; - budgets for M24 to M31 — each of those milestones adds its rows to the budget file as it lands, under the rules fixed here.
Design
Where it lives
| Part | Where | Why |
|---|---|---|
| The budget file, the budget benchmarks, the budget checker | benchmarks/budgets.json, benchmarks/AdCodicem.Pdf.Benchmarks, a Budgets job in ci.yml | Budgets are data with a history; the checker is a small console step, not a test framework |
| Allocation budgets as tests | tests/AdCodicem.Pdf.Tests (CorpusPerformanceTests), and each satellite's test project | Deterministic on one runtime, so they run on every commit like any test |
| The object cache, name interning, windows, the rebuild's scan, the object-stream cache, the piecewise decode | Core, IO/ and Objects/ | The reader's own internals; the public surface grows only by PdfStream.DecodeTo, OpenDecoded, IPdfDecodeSink, two options, and MaxDecodedStreamLength as a long |
PdfOptimizer, its options, steps, plan and report | Core, Optimization/, namespace AdCodicem.Pdf.Optimization | Lossless steps need the writer, the fonts and the copier, all core; lossy steps run on the codecs the caller passes, as M22's consumers do |
| The JPEG encoder (forward DCT, quantization, chroma subsampling) | AdCodicem.Pdf.Imaging | Beside M22's decoder and baseline re-encoder, which it shares; the core never encodes JPEG |
| Resampling, gray and bilevel conversion | Core, Images/ | They work on M22's rows and need no codec |
| Linearization | Core, M03's writer (IO/Writing/) | Only the writer knows positions (CLAUDE.md, Known traps) |
| The AOT host, the trimming host, the WebAssembly host | tests/AdCodicem.Pdf.AotHost, tests/AdCodicem.Pdf.TrimHost, tests/AdCodicem.Pdf.WasmHost | Test projects, never packaged; the AOT host is the tool's own binary where the tool covers the operation |
| The container profile | deploy/container/ (the Dockerfile and its measurement script), a Container profile workflow | A deployment the documentation publishes and CI measures |
| The fuzzing targets | tests/AdCodicem.Pdf.Fuzzing (M22's harness) and FuzzingTests | One harness, one campaign, one record |
optimize, --linearize | AdCodicem.Pdf.Tool | The tool ships the Imaging satellite (M22) |
Budgets
benchmarks/budgets.json one entry per (operation, document, metric): the value, the tolerance, the unit, the
measurement it was set from (commit, runner class, date), and the reason for each raise
BudgetChecker reads BenchmarkDotNet's JSON results and the file; exits non-zero naming each budget
exceeded, by how much, and the measurement the budget came from
- Allocation is a budget of bytes, asserted as a test. On one runtime, allocation is deterministic up to a few
hundred bytes of first-call initialization: each test warms the operation once, then measures it with
GC.GetAllocatedBytesForCurrentThread, asCorpusReadingTestsalready does for indexing (4 MB budget, 3.2 MB measured) andDocumentReaderTestsfor an index of 300,000 objects (32.6 and 45.2 MB budgets, 31.1 and 43.1 MB measured). An operation that runs on several threads — M12's parallel batch, M06's merge with parallel hashing — runs in a child process and is measured withGC.GetTotalAllocatedBytes(precise: true), since the per-thread counter misses what other threads allocate. Tolerance: 5 % over the recorded figure, no more. - Throughput is a ratio, measured where the noise cancels. Shared runners vary by more than any regression worth
catching. The
Budgetsjob therefore builds the pull request's head and its base, and runs the budget benchmarks of both alternately in the same job, several rounds each, on BenchmarkDotNet's fixed short job; it fails when the ratio of medians exceeds 1 + the tolerance. The tolerance is set by slice 1 from the spread of that ratio over twenty runs of one commit against itself, and recorded; the expectation is under 10 %. - Drift across many small regressions is caught by the nightly
Benchmarksrun onmain, which compares absolute figures with the committed baseline under a wider tolerance and fails the workflow beyond it. - What is budgeted, each row from the benchmark its milestone wrote: indexing, reading every page and validating
the thousand-page journal (
ReaderBenchmarks,ValidationBenchmarks); a full rewrite and an incremental update (WriterBenchmarks); a merge of the journal with itself (AssemblyBenchmarks); text extraction (ExtractionBenchmarks); layout of the reference documents and a thousand invoices (LayoutBenchmarks,BatchBenchmarks, whose CI floor M12 set at half the first measurement and which the ratio replaces); decoding per codec (ImagingBenchmarks); and this milestone'sOptimizerBenchmarksandStreamDecodingBenchmarks. - Raising a budget is an edit of
budgets.jsonwith its reason, and a line instatus.md; a test fails on an entry whose raise has no reason. A budget set too tight teaches contributors to raise it: each is set from a measurement with its tolerance, never from a wish. benchmarks/README.mdchanges accordingly: the full suite stays on demand; the budget subset runs on every pull request.
The object cache and names (#37)
- Least recently used. The cache keeps its slots in arrays — a dictionary from object identifier to slot, and
previous and next indices per slot —, so that a hit moves an entry to the front and an eviction takes the back,
O(1) and 0 B per access.
ObjectCacheCapacitystays the count it is; a weight is added: each entry costs its length in the file, which the reader knows, and the cache evicts when either the count orPdfReaderOptions.ObjectCacheBudget(bytes, 64 MB by default) is exceeded, so that 8,192 dictionaries of 16 MB each can no longer be held at once. An entry larger than the budget is returned and not cached. - Names without an intermediate string. The parser decodes a name's
#xxescapes into a stack or pooled buffer, and looks it up through the dictionary's alternate lookup onReadOnlySpan<char>(StringComparer.Ordinalimplements it since .NET 9); a string is allocated only for a name seen for the first time. - The process-wide table is bounded.
PdfName.Getinterns into a staticConcurrentDictionarythat grows with every distinct name any file ever contained and never shrinks: in a server reading hostile uploads, a file of a million distinct names is a million strings held for the life of the process, against invariant 2 and invariant 8's spirit. From this slice: the names ISO 32000 defines — the well-known names, and every key the Arlington model lists (ADR 44) — are interned process-wide in a frozen table built once; every other name is interned in a table per document, released with it. Equality is by value already, so no caller sees a difference but memory; reference equality stays the fast path for the names that matter. Measured in slice 2: the process's heap after opening a million-name file and disposing it returns to its level before.
Windows grown on demand (#47, #49)
- Cross-reference sections are read through a window that starts small — 4 KB — and grows geometrically to the
section's end or
MaxXRefSectionLength, appending what it reads rather than reading again from the start: the VHA coding handbook's 273 KB table then costs 273 KB, not 683 KB, and the signed Web Capture file's three sections cost what they weigh, not three 64 KB windows. The laziness test's bound — opening reads less than a quarter of the file — then holds for both documents recorded unsupported for #47, and their markers go. - The rebuild's trailer scan uses the same window per occurrence, starting at 1 KB, grown only while the
dictionary is open, up to
MaxTrailerLength, so that the cost per occurrence is what the trailer weighs; an occurrence oftrailerinside a stream's data — which the scan already delimits betweenstreamandendstream— is skipped. The 80,000-occurrence file (625 KB) then reads a bounded amount per keyword, and a sound trailer longer than 64 KB is no longer given a syntax error the file does not have. What the parser meets past a window it did not grow is dropped with the attempt, as T21 and T23 settled.
Decoded object streams under a budget (#50)
- The cache of decoded object streams in
PdfFileReadergainsPdfReaderOptions.ObjectStreamCacheBudget, in decoded bytes (64 MB by default, checked against the corpus's largest object streams in slice 3), evicting the least recently used stream; an evicted stream is decoded again when one of its objects is asked for, and yields the same objects. - The rebuild no longer holds what it indexed. It decodes each object stream to index its objects, records each
object's stream and index, and releases the decoded bytes under the same budget. The damaged file of four object
streams that each decode to 256 MB then holds at most the budget once
Openreturns, instead of 1 GB. - The budget is not a bound in ADR 34's sense: reaching it costs time — a second decode —, never data, so it has
no
limit.*code and no document is refused by it. A single stream larger than the budget is decoded, used and not kept.
The piecewise decode (#48)
IPdfDecodeSink Write(ReadOnlySpan<byte> window); Complete(PdfDecodeResult): windows in order, each valid
only during the call
PdfStream.DecodeTo DecodeTo(IPdfDecodeSink, CancellationToken) — push; memory held is one window per stage
PdfStream.OpenDecoded OpenDecoded() -> Stream — pull, read-only, forward-only, over the same pipeline
PdfStream.Decode unchanged: the whole stream in memory, for callers that want it and streams that fit
- Every filter becomes a stage that consumes an input window and produces output windows into pooled buffers —
Flate over the BCL's inflater, LZW with its 4,096-entry table, RunLength, ASCIIHex, ASCII85, the PNG and TIFF
predictors with their one row of history, decryption block by block (M16 wrote it so), and M22's image filters,
whose row sinks are already this shape. The window is 64 KB per stage; a predictor holds one row as well, whose
width
/Columnsdeclares and which the decoded-length guard already bounds. - The Flate stage keeps what #56 made the whole decode keep. After a fault it reads its input again from the
start — the stream's data, or the stages before it run again — and skips what it already handed out, so no window
repeats. A zlib checksum that disagrees with whole data is known only once every window is out:
Completecarriesfilter.checksum-mismatch, and a consumer that re-encodes — M03's writer, the optimizer's recompression — never gives such data a fresh checksum in silence. One form of the data can give way to another after it kept bytes: raw deflate read from leading white space, which turns corrupt, to zlib after the white space, which reads further. Data that starts with white space is therefore read once to choose its form before a window goes out, rather than taking back windows already handed out. MaxDecodedStreamLengthkeeps bounding what one stream may decode to, becomes along, andPdfReaderLimits.Unboundedsets it tolong.MaxValue. Streaming does not make a decompression bomb harmless: a 1 MB stream of nested Flate that decodes to a terabyte costs an internal consumer — the optimizer, the content interpreter, a validator's rule — the time to produce a terabyte, which invariant 4 forbids. What changes is what is held: one window, whatever the length. ADR 34 foresaw that a streaming decode would reopen it, and read thatMaxDecodedStreamLengthwould then bound memory held rather than length; this milestone keeps it on length, for the time it bounds, and records the reason in an ADR at the next free number amending ADR 34.Decode()pastArray.MaxLengthkeeps what an array holds and reportsstream.too-large-to-hold, whose message namesDecodeTo: that is the implementation's ceiling, not a guard, so no property lifts it and it has nolimit.*code. UnderUnbounded,DecodeToreads such a stream whole.- Every internal consumer moves to the piecewise path: M03's writer when it re-encodes, the optimizer, M15's content
interpreter (content streams are read a window at a time, a token never spanning more than the lexer's bound), M02's
stream rules, M06's hashing for deduplication, M14's attachment extraction. Offsets and lengths are
longwherever a stream's decoded size is counted. - Native memory only if measured. The pooled windows are small and managed; ADR 35 applies only if a benchmark shows a stage cannot meet its budget in managed code, and none is expected to.
- Cancellation is checked between windows, so a stream decoding for minutes stops within one window's work.
The optimizer
PdfOptimizer Plan(document, PdfOptimizerOptions) -> PdfOptimizationPlan: what each step would do and
save, without writing; Optimize(document, output, options, IProgress<PdfProgress>?,
CancellationToken) -> PdfOptimizationReport
PdfOptimizerOptions immutable: Steps (lossless, all by default), Discards (none by default), LossySteps (none
by default), TargetSize, Codecs (PdfImageCodecs; the core's by default), Linearize,
SaveOptions (M03's, for version and cross-reference form), SignaturePolicy and
ConformancePolicy (M04's and M09's, Refuse by default)
PdfLosslessStep [Flags] Deduplicate, ConsolidateFontSubsets, SubsetFonts, Recompress, PackObjects,
PruneUnusedResources, DropUnreferenced
PdfDiscard [Flags] PrivateApplicationData, Thumbnails, UnembedStandardFonts
PdfLossyStep a closed set of records: Downsample(threshold ppi, target ppi, method),
ReencodeJpeg(quality, subsampling), ConvertToGray, ConvertToBilevel(threshold)
PdfTargetSize bytes, the lossy steps it may use in order, each with its floor (a lowest resolution, a
lowest quality)
PdfOptimizationReport per step: objects affected, bytes before and after; per image a lossy step or a discard
touched: the object, the pages that draw it, encoding, size, dimensions and effective
resolution before and after, and the step; whether the target was met, and if not, the
smallest size reached; every diagnostic
Three classes of change, recorded in an ADR at the next free number that extends ADR 42's rule beyond images:
- Lossless — the document renders, extracts, validates and claims exactly what it did. On by default.
- Discards — data no viewer shows but an application or a person may want: Illustrator's and PDFMaker's
/PieceInfo, page/Thumbimages and XMP thumbnails, standard fonts embedded where a viewer would substitute them. Never by default; each named; each object discarded in the report. - Lossy — the pixels change. Never by default; each named; each image changed in the report (ADR 42).
Lossless steps
- Deduplicate. M06's keys, applied to one document rather than to a merge's inputs: a stream's key is SHA-256 over
its dictionary without
/Lengthand its encoded bytes; a dictionary's or array's over its canonical form with each reference replaced by its target's key; a cycle falls back to identity. Two additions: a decoded key for streams whose encodings differ — the same logo once in Flate and once uncompressed, the same sRGB profile under four encodings, which M06 already compares decoded — keeping the smaller encoding; and a pass over the whole document, so that a producer's own duplicates (SAP NetWeaver's logo embedded twice, Photoshop's profile twice) are found. Never shared, whatever their bytes: page objects and page content streams (M09 edits content in place), anything carrying/StructParentor/StructParents, annotations, form fields, optional content groups and membership dictionaries, structure elements, outline items, destinations, article beads, signature dictionaries, and streams whose crypt filters differ (/Identitymetadata beside encrypted streams). Each of those has an identity its bytes do not carry: two layers with equal dictionaries are two layers. - ConsolidateFontSubsets. Two embedded programs are the same face when their names without tag, their font
types and every glyph present in both — outlines, advances, and for TrueType the instructions — are equal, compared
through M08's parsers; a tag is never trusted, since Illustrator and Konik reuse one tag over different subsets. The
union is written once, by M08's subsetter, under a new tag, and each font dictionary keeps its own codes, widths and
ToUnicode: aCIDFontType2through its ownCIDToGIDMapinto the union; aCIDFontType0or aType1Cby glyph name or CID, which the union keeps; a simple TrueType only when its codes reach the same glyphs through the union'scmap, and kept apart otherwise. No content stream is rewritten. Under PDF/A-1,CIDSetandCharSetlist what the union holds. Two programs sharing a tag whose common glyphs differ are reported (optimize.subset-tag-reused) and kept apart. - SubsetFonts. A program embedded in full is cut to the glyphs the document draws with it — found by M15's
interpreter over every page's content, form XObjects, tiling patterns, Type 3 glyph procedures, annotation
appearances and soft-mask groups, whatever the render mode — under a new tag. Kept whole: a font in the
AcroForm's
/DR, or named by any field's or free-text annotation's/DA, since a later fill may need any glyph (optimize.font-kept-whole); a font whosefsTypeforbids subsetting (M08 reads it). - Recompress. Every stream re-encoded when the result is smaller, deterministically: uncompressed, LZW,
RunLength and ASCII-armored streams to Flate; Flate re-deflated at the BCL's
SmallestSize; images given the PNG predictor that a fixed heuristic chooses per image; one-bit images in the smaller of G4 and JBIG2 generic (M22'sSmallestpolicy); JPEG images losslessly re-entropy-coded with optimized Huffman tables by M22's baseline re-encoder, coefficients untouched. A stream whose re-encoding is not smaller keeps its bytes. Never: LZW under any output (Flate only, as M03 writes), JPEG 2000 under PDF/A-1, or a filter the output's version lacks. Embedded files are recompressed like any stream; their/Params /Sizeand/CheckSumdescribe the decoded file and stay true. - PackObjects. Object and cross-reference streams through M03's
ObjectStreams.Generate, except under a PDF/A-1 claim (M03 already refuses them there). - PruneUnusedResources. Names in a resource dictionary that no content of its owner uses — found by the same interpretation — removed, per page and per form XObject; M07 did it per extracted page, this does it per document. A resource dictionary shared by several owners keeps the union of their uses.
- DropUnreferenced. M03's full rewrite drops what nothing reachable refers to already; the optimizer reports it as its own step, with the bytes it saved.
Discards
PrivateApplicationData: the/PieceInfodictionaries of pages, form XObjects and the catalog, with the/LastModifiedentries that date them. Illustrator will no longer open the file as editable, and the report says so.Thumbnails: page/Thumbstreams andxmp:Thumbnailsin XMP (M14's model).UnembedStandardFonts: an embedded program of one of the fourteen standard faces — recognized by M08's parser, not by its name alone — removed, the font dictionary kept with its widths. Refused under a PDF/A, PDF/UA or PDF/X claim, which require every font embedded; run on such a document only underConformancePolicy.RemoveClaim, the claim then removed and reported atConformanceLoss(invariant 7).
Lossy image steps and the target size
- Where pixels come from. M22's
PdfImageReaderover the codecs in the options: the core decodes CCITT and the core's filters;PdfImagingCodecs.Defaultadds JBIG2, JPEG and JPEG 2000. An image the given codecs cannot decode is left as it is and reported (optimize.image-not-decodable). - Downsample an image whose highest effective resolution across every placement (M15's inventory) exceeds a
threshold, to a target — 150 ppi above 225 by default when the step is named —, by area averaging in integer
arithmetic for gray and color and by majority for one-bit images; JPEG 2000 at a coarser resolution level when that
lands at or above the target (M22's reduced-resolution decoding), which costs a quarter of the work. Masks
(
/SMask, a stencil/Mask) are resampled onto the same grid as their image,/Mattekept; a color-key/Maskis kept as ranges. An image shared by a thumbnail and a full page is judged by its largest placement. - ReencodeJpeg at a quality, with the IJG scaling of Annex K's tables and 4:2:0 or 4:4:4 subsampling, through the JPEG encoder added to the Imaging satellite; it applies to images that are not bilevel. A JPEG is re-encoded from its decoded samples once: the target-size mode tries each quality from the original, never from its last attempt, so no image suffers two generations.
- ConvertToGray by M22's evaluation (reported as approximate where M22 approximates); ConvertToBilevel by
Otsu's threshold over the image's histogram, clamped to a band (M22's blank-page method), then encoded by the
Smallestpolicy — the largest win on a scanned letter, and the step most likely to lose a light signature, which is why it is named and reported. - Never lossy, whatever is named: image masks (already one bit), images inside a signed revision, an image an
/Alternatesentry names, JBIG2 written with symbols. - The target size. The lossless steps first; then each named lossy step in the caller's order, each tried at
successive settings down to its floor; after each, the size is re-estimated. Estimation does not write the document:
each image is encoded at the candidate setting and only its length kept, the rest of the file weighed once; the
final write encodes again at the chosen settings — the same bytes, since encoders are deterministic — so memory
stays one image's working set rather than the output. If the floors are reached above the target, the report gives
the smallest size reached, and the caller decides (
optimize.target-not-met). Deterministic: the same document, target and steps give the same bytes.
Linearization
PdfSaveOptions.Linearizeon any full rewrite, andPdfOptimizerOptions.Linearize: the output follows ISO 32000-2 Annex F — the linearization dictionary, the first-page cross-reference section, the first page's objects, the primary hint stream with its page offset and shared object hint tables, then the rest in page order.- Two bounded passes in the spirit of ADR 39: the first serializes every object into a counting sink and keeps only its length and the pages that use it — a few bytes per object —; the second writes. The hint stream's length feeds the offsets it describes, so its encoding is iterated to a fixed point, at most three times, then padded. The output is forward-only: a non-seekable stream works, as for any rewrite.
- A linearized output stays linearized until updated: an incremental update leaves it stale, as every update of a linearized file does and viewers accept (M03's trap); the report says so when the caller updates one.
Signed, encrypted and conforming inputs
- Signed: optimization is a full rewrite, which invalidates every signature; M04's guard refuses it with
PdfSignatureInvalidationExceptionunlessSignaturePolicyisAllowInvalidatingSignatures, which reports each signature it breaks (write.signature-invalidated). Nothing is optimized by update. - Encrypted: opened with its password (M16); optimization needs M16's
Modifypermission underPdfPermissionPolicy.Respect; the output is encrypted as the input was, unless the caller's save options say otherwise. - Conforming: a PDF/A, PDF/UA or Factur-X claim is kept only when it stays true, checked by M20's
IPdfConformanceCheckerwhen one is registered — every lossless step keeps it by construction and the tests prove it —; a step that would break a claim is refused underConformancePolicy.Refuse, or run with the claim removed and reported underRemoveClaim.
Consumers completed
- M06:
PdfAssemblyOptions.ConsolidateFontSubsets, off by default, runs the consolidation over the assembled volume before it is written. - M12:
PdfRenderOptions.ImageStepstakes the namedDownsampleandReencodeJpegsteps, applied as an image is embedded; theimage.resolution-excessivereport M12 writes then says what was done rather than what could be. - M18:
portal.file-sizeshapes a piece first by lossless optimization, then by the lossy steps the caller's case-file options name — never by the preset's choice —, then by M07's split when the preset allows it. - M03:
PdfSaveOptions.Linearize, above.
Native AOT, trimming and WebAssembly
- AOT. The tool's AOT binary (M06) — and, for operations the tool lacks,
tests/AdCodicem.Pdf.AotHostpublished withPublishAot— opens, validates, rewrites, merges with itself and optimizes every committed and remote corpus document; its output is byte-identical to the JIT build's. On linux-x64 on every pull request, linux-arm64 and win-x64 nightly. - Trimming.
tests/AdCodicem.Pdf.TrimHostreferences the core and each satellite that declaresIsAotCompatible, calls their public entry points, and is published trimmed withTrimmerSingleWarnoff and warnings as errors: a warning attributed to any of our assemblies fails the build. The satellites that ship native code (.Html,.Rendering) are checked for AOT the same way; where AngleSharp or a binding warns and the warning cannot be suppressed where it fires with a justification (ADR 29), M12's rule stands — the verb ships in the dotnet tool only — and the measurement is recorded. - Browser WebAssembly.
tests/AdCodicem.Pdf.WasmHost, abrowser-wasmapplication whose exported functions take a document's bytes and return the outputs' hashes, loaded in headless Chromium driven by Playwright in a container. It opens, validates, rewrites and optimizes losslessly every committed corpus document from memory — there is no file system — and gives the JIT build's decoded content. Bytes are compared where no stream was re-deflated; where the optimizer recompressed, decoded streams are compared, since the browser runtime's deflater may not be the one the x64 runtime ships (to verify, and recorded). What the run proves or refutes about the platform: SHA-256 (managed on browser-wasm, per Microsoft's cross-platform cryptography table); MD5 and AES absent, so that the encrypted documents open under R2 to R6 through the core's managed MD5, RC4 and AES (M06's and M16's, not constant-time) and an AES-GCM file of M16's (R7) is refused withPrimitiveUnavailable; Brotli for WOFF2, whose proof M08 leaves to this run — the host decodes M08's WOFF2 test fonts to the sfnt the JIT build produces —; and the absence of threads.
The container profile
- The image: .NET 10's chiseled
runtime-depsimage,mcr.microsoft.com/dotnet/runtime-deps:10.0-noble-chiseled— not its-extravariant, which adds ICU —, pinned by digest, the AOT tool binary copied in; no shell, no package manager, no fontconfig, no ICU (InvariantGlobalization, which M15 wrote its Unicode tables for); a non-root user;--read-onlywith no writable mount, since the library writes no temporary file;--memory 512m; no network. The HTML engine's native assets areSkiaSharp.NativeAssets.Linux.NoDependencies, SkiaSharp's build without fontconfig, andHarfBuzzSharp.NativeAssets.Linux(that it links nothing the chiseled image lacks: to verify withlddin slice 10); fonts come from ADR 11's registry and embedded set, never the system. - The workloads: the reference documents rendered from
sources/invoice-fr.html,report-fr.htmlandcontract-fr.html; a thousand invoices from one compiled template; opening, validating, extracting and optimizing the thousand-page journal; and, nightly, W11's three heavy documents. - The measurement: wall time, CPU time, and peak memory as the container's cgroup records it, per workload, cold and warm; the image's size. Beside it, headless Chromium rendering the same HTML through PuppeteerSharp — the slot the comparison benchmarks reserved — in a container under the same limits, given the same OFL fonts; where Chromium fails under 512 MB, that is the result.
- The published figures state the runner, the date, the versions and the digest;
docs/website/docs/guides/deployment.mdgives the profile as a recipe.
Hardening
- Fuzzing. The nightly
Fuzzingworkflow gains targets: the writer's round trip (a mutated document that opens is rewritten, reopened and rewritten again, and the two rewrites are identical), the optimizer (a mutated document optimizes or fails typed, and its output opens with no repair), the piecewise decode against the whole decode (the same bytes, window boundary by window boundary). M22's coverage-guided engine, chosen by its ADR, is applied to the reader's entry points — the lexer, the object parser, the cross-reference readers, the object-stream reader, the rebuild — as it is to the decoders. The seeds are everydamageddocument, committed and remote. M25's rendering target and M27's parsers join the same campaign when they land. A finding becomes a regression test inHostileInputTestsbefore it is fixed. - Cancellation latency. Every long operation — open with rebuild, decode, validate, rewrite, merge, extract,
render, optimize — observes its token between units of bounded work (a window, an object, a page, a record), so that
a canceled operation returns within 100 ms on the budget runner, hostile input included: a decode bomb under
Unbounded, a rebuild of the 80,000-keyword file, a type 4 function of a million operators. - SECURITY.md states the threat model as measured: the guards and their defaults, the fuzzing campaign and its record, the container profile, and what a caller must set to read untrusted input in a server.
Bounds, classified (invariant 12, ADR 34)
| Bound | Class | Why |
|---|---|---|
| What one stream may decode to | The existing guard MaxDecodedStreamLength, now a long | A valid file can exceed any value; the cost of a decode is proportional to it |
What Decode() returns whole | Implementation ceiling, Array.MaxLength, reported as stream.too-large-to-hold | Not a choice: DecodeTo reads past it |
| A decoding stage's window | Internal constant, 64 KB | Independent of the file; a predictor's row is bounded by the decoded-length guard |
| Cross-reference windows, trailer windows | The existing guards MaxXRefSectionLength, MaxTrailerLength | Growth stops at them; the start size is a constant |
| Object cache count and weight; object-stream cache | Options, not guards: ObjectCacheCapacity, ObjectCacheBudget, ObjectStreamCacheBudget | Reaching them costs a re-read, never data |
| The per-document name table | Bounded by the names the file holds | Each entry is a name the parser read under MaxObjectLength |
| The optimizer's deduplication index | 32 bytes and a number per shared object | Bounded by the index the reader holds |
| The linearization's fixed-point iterations | Internal constant, 3 | Our own serialization converges or is padded |
| The target size's settings per step | The caller's floors, stepped by a constant | The caller's input, not the file's |
Diagnostics
In PdfDiagnosticCodes, disjoint from rule identifiers (ADR 36):
| Code | Severity | Meaning |
|---|---|---|
stream.too-large-to-hold | Warning | Decode() asked for a stream past Array.MaxLength; the message names DecodeTo |
optimize.subset-tag-reused | Warning | Two programs under one subset tag whose common glyphs differ; kept apart |
optimize.font-kept-whole | Information | A font a form field or a free-text annotation may need; not subset |
optimize.image-not-decodable | Information | An image the given codecs cannot decode; left as it is |
optimize.lossy-applied | Information | One per image a lossy step changed, with the step and the figures |
optimize.discarded | Information | One per object a discard removed |
optimize.target-not-met | Warning | The floors were reached above the target; the smallest size reached |
optimize.conformance-claim-removed | ConformanceLoss | A step run under RemoveClaim broke a claim; the claim is gone |
write.linearization-stale | Information | An update to a linearized file leaves its hints stale |
The command-line tool
adpdf optimize FILE -o OUT [--lossless-only] [--steps dedupe,fonts,subset,recompress,pack,prune] [--discard private,thumbnails,standard-fonts] [--downsample 150@225] [--jpeg 75] [--gray] [--bilevel] [--target-size 10MB] [--linearize] [--allow-invalidating-signatures] [--report report.json]; --linearize on every
verb that writes. Exit codes are M06's; optimize.target-not-met exits 1 with the output written. The AOT binary
produces what the API produces.
Slices
Each slice ends on a green commit, with its codes documented, its budget rows in benchmarks/budgets.json and its
measurements in docs/status.md.
- Budgets in CI (#38). Delivers
budgets.json,BudgetChecker, theBudgetsjob with its A/B runs, the allocation budgets of every existing benchmark asCorpusPerformanceTests, the nightly drift check, the tolerance measured and recorded,benchmarks/README.mdrewritten, the rule that a raise carries a reason, and an ADR at the next free number recording the change of CI policy — benchmarks, run on demand only until now, gain a budget subset on every pull request —, with the A/B ratio chosen over absolute figures and the reasons. Proved by unit tests of the checker (a result within, at and beyond tolerance; a missing entry; a raise without a reason); a throwaway branch that allocates one array per token in the lexer and one that sleeps a microsecond per object, each failing the job, their runs recorded instatus.md. Leaves the reader's debts. - The object cache and names (#37). Delivers LRU with weights,
ObjectCacheBudget, span lookup, the frozen process-wide table from the well-known names and the Arlington keys, per-document tables. Proved by unit tests (eviction order under hits; an entry larger than the budget; a name written with#xxescapes interned as the same name written plain; two documents' names equal by value); a property — for any sequence of gets, the cache returns what an uncached reader returns —;ReaderBenchmarksat 0 B per repeated name; the million-name file measured: the heap returns to its level after disposal. Leaves windows. - Windows and the object-stream budget (#47, #49, #50). Delivers growing windows that append, the rebuild's scan
through them skipping stream data,
ObjectStreamCacheBudget, a rebuild that releases what it indexed. Proved byWindowEdgeTestsextended (a section growing across every boundary gives the entries a whole read gives); a trailer inside a content stream not taken; a sound 100 KB trailer rebuilt without a syntax error; the four-stream file holding the budget;CorpusReadingTests.Opening_does_not_read_the_content_ofpassing on the two remote #47 documents with their markers removed, on a greenRemote corpusrun. Leaves the piecewise decode. - The piecewise decode (#48). Delivers
IPdfDecodeSink,DecodeTo,OpenDecoded, every filter as a stage,MaxDecodedStreamLengthas along,stream.too-large-to-hold, the internal consumers moved, the ADR amending ADR 34,StreamDecodingBenchmarks. Proved by an FsCheck property — for every filter chain, predictor and window size, the piecewise output equalsDecode()'s, including windows of one byte —; a synthetic stream of nested Flate decoding to 2.5 GB read whole underUnboundedand its SHA-256 checked, holding one window; the USGS map's image throughDecodeTounder the recorded hold (remote); the reader-limits page andReaderLimitsTestsupdated, and the manifest schema'sreaderLimits.maxDecodedStreamLength, capped at 2³¹ − 1 today, raised to what alongholds,CorpusManifestSchemaTestswith it. Leaves the optimizer. - The optimizer's frame and lossless structure. Delivers
PdfOptimizer, options, plan, report, the three classes and their ADR,Deduplicate,PackObjects,DropUnreferenced,PruneUnusedResources, the signed, encrypted and conforming policies; the manifest's optimizationexpectfields — distinct decoded streams, duplicate groups, fonts embedded in full, unreferenced objects — and their schema, written bybuild_corpus.pyfrom pikepdf,pdffontsand qpdf. Proved by unit tests per step on a document that needs it and one that does not (no change, same bytes); two equal optional content groups kept apart; streams differing only by crypt filter kept apart; integration: qpdf--check, pdftotext, MuPDF rasters and veraPDF unchanged on every committed document. Leaves fonts. - Fonts. Delivers
ConsolidateFontSubsetsandSubsetFonts, M06's option,optimize.subset-tag-reusedandoptimize.font-kept-whole. Proved by unit tests (two subsets of one face unioned, each dictionary's codes intact; two faces under one tag kept apart; a/DRfont kept whole; a PDF/A-1CIDSetrewritten); integration:pdffontslists fewer programs, allsub yes; pdftotext's text and MuPDF's rasters identical; veraPDF's verdicts unchanged. Leaves recompression. - Recompression and linearization. Delivers
Recompressover every filter, the JPEG entropy re-coding, the G4 and JBIG2 choice,PdfSaveOptions.Linearize, the two passes,write.linearization-stale. Proved by a property — any stream recompressed decodes to its original bytes —; JPEG coefficients read back by libjpeg in the Python container identical; integration: qpdf--check-linearizationon every linearized output, including the 26 committed files whose own hints are inconsistent, broken or stale; pikepdf's decoded streams identical. Leaves lossy steps. - Discards, lossy steps and the target size. Delivers the three discards, resampling, gray and bilevel
conversion, the JPEG encoder in the Imaging satellite,
ReencodeJpeg, the target-size mode, M12's and M18's consumers. Proved by unit tests (no lossy step without a name; a JPEG never re-encoded twice; a mask on its image's grid; a shared image judged by its largest placement; a target met; a target not met, reported); integration:pdfimages -listagrees with the report on every changed image; our JPEGs decode in Pillow; MuPDF's renderings before and after within the structural-similarity threshold slice 8 fixes and records; pdftotext's text identical. Leaves the platforms. - Native AOT, trimming and WebAssembly. Delivers the AOT, trimming and wasm hosts, the nightly matrix, the platform findings recorded. Proved by the AOT and wasm rows below; the trimming host's publish with no warning. Leaves the container.
- The container profile. Delivers
deploy/container/, theContainer profileworkflow, the Chromium slot of the comparison benchmarks, the published figures and the deployment guide. Proved by the container row below. Leaves hardening. - Hardening. Delivers the new fuzzing targets, the coverage-guided reader targets, cancellation checks where the
latency test finds them missing,
SECURITY.md. Proved byCorpusCancellationTestsover every long operation and the hostile cases above; fourteen consecutive nights of the campaign with no open finding, recorded. Leaves the whole. - The heavy documents, the verb, the whole (#46). Delivers the memory budgets of W11's documents, measured on
the AOT binary as a child process,
optimizeand--linearize,OptimizerBenchmarks, the documentation. Proved by the remote rows below on a greenRemote corpusrun;CorpusToolTests.
Tests required
Unit — tests/AdCodicem.Pdf.Tests, the Imaging satellite's under Imaging/ as M22 placed them:
- Budgets: the checker's arithmetic and messages; every entry of
budgets.jsonhas a unit, a tolerance, a source measurement and, when raised, a reason. - Caches: LRU order; weights; the budget with one entry larger than itself; eviction during a rebuild; an evicted object stream re-decoded to the same objects; a property over random access sequences.
- Names: span lookup allocation-free for a known name; a new name allocated once per document; the frozen table immutable; names from two documents equal by value and hash.
- Windows: sections growing across each boundary; a section exactly at
MaxXRefSectionLength; a trailer split at each byte;trailerinside stream data, inside a string, inside a comment. - Piecewise decode: every filter and predictor, alone and chained; windows of 1, 7, 4,096 and 65,536 bytes; ASCII85's
~>split across windows; LZW's early change at a window's edge; a TIFF predictor with 2-bit samples across a window; a Flate stream that lost its tail (T32's report preserved); one that turns corrupt and one whose checksum disagrees (#56's kept bytes and reports preserved); decryption across windows; cancellation between windows. - Optimizer: each step on a document that needs it and on one that does not; every "never shared" kind; each discard refused under each claim; the report's figures against the output; determinism — two runs, two cultures, the same bytes.
- Fonts: TrueType, CFF,
Type1Cand CID-keyed unions; a simple TrueType whose codes disagree kept apart;fsTypeforbidding subsetting; glyphs reached only from an annotation's appearance or a Type 3 procedure kept. - Lossy: area averaging and majority against hand-computed grids; Otsu's threshold on known histograms; the JPEG encoder's output decoded by our decoder within the quantization error; target-size search order and floors.
- Linearization: the hint tables of a synthetic document checked field by field; a one-page document; a document whose first page shares every object; the fixed-point padding.
- Hostile: a decode bomb of nested Flate under default limits (the guard at 256 MB) and under
Unbounded(canceled within the latency budget); a million distinct names; 80,000trailerkeywords; four object streams of 256 MB; a font program claiming 65,535 glyphs over 1 KB; an image declaring 65,535 × 65,535 pixels offered toDownsample; a document whose every page shares one resource dictionary of 10,000 entries — each ends in output, a report or a typed exception, within its time and allocation budget. - Fuzzing: the new targets join
FuzzingTestsper commit with fixed seeds and the nightly campaign with the night's seed; every finding a regression test first.
Integration — tests/AdCodicem.Pdf.IntegrationTests, every referee in a container (ADR 27):
- qpdf —
--checkon every output;--check-linearizationon every linearized output;--show-object--filtered-stream-datafor decoded bytes of heavy streams; - pikepdf — decoded streams hashed before and after; duplicate groups counted independently; fonts'
FontFile*objects counted; - poppler —
pdftotextbefore and after, identical for lossless steps and for lossy image steps;pdffontsfor embedding and subsetting;pdfimages -listfor encodings, sizes and resolutions against the report; - MuPDF —
mutool drawrasters before and after: identical for lossless steps, within the recorded similarity threshold for lossy ones; - veraPDF — the verdict on every claimed level unchanged by lossless optimization and by the lossy steps;
- Pillow and libjpeg — our JPEGs decoded; coefficients unchanged by entropy re-coding;
- Playwright with Chromium — the WebAssembly host; headless Chromium through PuppeteerSharp — the container comparison;
- M14's Factur-X referee — the invoices' XML still validates after their attachments are recompressed.
Acceptance conditions
"The stress documents" are documents/stress/reportlab-journal-1000-pages.pdf and the benchmarks' synthetic
thousand-page document; "the heavy documents" are W11's three remote references, remote/govinfo/us-code-2023-title42.pdf
(9,302 pages), remote/usgs/us-topo-washington-west-2023.pdf (one page, a 328,608,000-byte image, opened under its
readerLimits) and remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf (125 pages of JPEG 2000, 147 MB). Remote
rows close only on a green Remote corpus run, recorded in status.md with its date.
| Documents | Behavior | Verified by |
|---|---|---|
| The stress documents | Every budget of budgets.json holds — allocation on every commit, throughput on every pull request against its base —, and a change beyond a tolerance fails CI | CorpusPerformanceTests.Budgets_hold_on_the_stress_documents (new), the Budgets job |
| The heavy documents (#46) | Opening, walking every page, validating, extracting text, rewriting and optimizing losslessly each stay within the peak-memory budget recorded in status.md, flat across pages, measured on the AOT binary as a child process | CorpusPerformanceTests.Heavy_documents_stay_within_their_memory_budgets (new) |
remote/usgs/us-topo-washington-west-2023.pdf (#48) | Its image decoded through DecodeTo with the process's managed heap under 64 MB, its SHA-256 equal to qpdf's filtered stream data; Decode() under its raised limit still gives the same bytes | CorpusStreamDecodingTests.A_large_image_decodes_a_window_at_a_time (new) |
| A stream decoding past 2 GB, generated at test time — not in the corpus (below) | Read whole under PdfReaderLimits.Unbounded, holding one window, its length and hash exact; cut and reported as limit.decoded-stream under the defaults | CorpusStreamDecodingTests.A_stream_past_two_gigabytes_is_read_whole_under_unbounded (new) |
remote/pdfcpu/acrobat-web-capture8-x509-rsa-sha1-signed.pdf and remote/opf-format-corpus/pdfmaker707-word-vha-coding-handbook.pdf, the two documents recorded unsupported for #47 | Opening reads less than a quarter of each file; their unsupported markers removed, which unblocks M21's and M27's rows on the first | CorpusReadingTests.Opening_does_not_read_the_content_of |
Every damaged document, committed and remote (#49) | Recovered as before, with no syntax error invented at a trailer longer than a window, reading a bounded amount per trailer keyword | CorpusReadingTests.Damaged_documents_are_recovered_as_far_as_an_independent_tool_recovers_them |
The 53 committed documents the manifest marks object-streams, and remote/pdfjs/cairo-firefox-objstm-index-overflow-bug1978317.pdf (65,542 objects in one stream) (#50) | Every object reads the same under an object-stream budget of 64 KB, forcing eviction, as under the default | CorpusReadingTests.Every_object_reads_the_same_under_a_small_object_stream_budget (new) |
| Every committed document the reader opens, but the signed ones, and the remote ones nightly | Lossless optimization never grows the file; qpdf --check passes; pdftotext's text, pdfimages -list's images and MuPDF's rasters are unchanged; veraPDF's verdict on every claimed level is unchanged | CorpusOptimizationTests.Lossless_optimization_changes_nothing_a_referee_sees (new) |
remote/pdfminer/sap-netweaver-invoice-issue1062.pdf (one logo embedded twice), remote/ocrmypdf/photoshop-cc2015-pdfx3-cmyk.pdf (one ICC profile twice), the M06 merges of the corpus | Every duplicate pikepdf's decoded-stream hashes find is written once, and nothing else merged; the bytes saved equal the report's | CorpusOptimizationTests.Duplicate_resources_are_written_once (new) |
vendor/us-federal/illustrator-irs-pub1-english.pdf and remote/zugferd-corpus/konik-pdfbox-zugferd1-basic-from-word.pdf (one tag over different subsets), remote/pdf-association/abledocs-pdfua1-tagged-textbook-scan.pdf (a placeholder tag shared), and a merge of committed documents subsetting one face — not in the corpus (below) | Programs consolidated only where their common glyphs agree, optimize.subset-tag-reused where they do not; pdffonts lists fewer programs; text and rasters identical | CorpusOptimizationTests.Font_subsets_are_consolidated_only_when_their_glyphs_agree (new) |
vendor/eu-publications/pdflib-oj-exchange-rates-greek.pdf, antenna-house-oj-exchange-rates-2019.pdf, vendor/uk-ogl/pdfmaker21-ozev-sample-invoice.pdf, vendor/opf-format-corpus/pdfmaker9-word-distiller-fonts-embedded-in-full.pdf; remote, remote/zugferd-corpus/symtrax-itextsharp-zugferd21-minimum-pdfa3a.pdf (PDF/A-3a) | Every program embedded in full subset to the glyphs drawn, tagged, pdffonts reading sub yes; text and rasters identical; the PDF/A claims veraPDF upheld still upheld | CorpusOptimizationTests.Fonts_embedded_in_full_are_subset_to_the_glyphs_used (new) |
The committed forms whose /DR embeds a font — vendor/fr-licence-ouverte/pdfmaker-acrobat-cerfa-12156-form.pdf (Arial in full, 572 KB, 406 fields) and vendor/us-federal/omniform-usda-rd1924-5-hidden-widgets.pdf (Verdana) | Every font a field names kept whole; M16's fill of every field after optimization gives the appearance it gave before | CorpusOptimizationTests.Fonts_a_field_needs_are_kept_whole (new) |
The LZW documents vendor/us-federal/acrobat3-import-irs-1040-1988-scan.pdf, distiller3-irs-ss4-1995-form.pdf, pdfwriter4-usda-dry-whey-standard-2000.pdf; the uncompressed attachments of vendor/zugferd/gnuaccounting-mustang10-zugferd-rc-invoice.pdf and pypdf2-facturx-python-false-pdfa3b.pdf; remote, the uncompressed images of remote/opf-format-corpus/illustrator-distiller601-mac-nida-scholastic-heads-up.pdf | Every stream Flate-encoded where smaller, pikepdf's decoded bytes identical; the Factur-X XML still accepted by M14's referee; /Params still true | CorpusOptimizationTests.Recompressed_streams_decode_to_the_same_bytes (new) |
Every committed document, linearized; documents/archival/qpdf-linearized-report.pdf as qpdf's own reference | qpdf --check-linearization reports no error; the 26 committed files whose own hints are inconsistent, broken or stale come out sound | CorpusLinearizationTests.Linearized_output_passes_qpdf_s_check (new) |
The scans: vendor/us-federal/xerox-workcentre-treasury-imf-report-scan.pdf (JBIG2), acrobat3-import-irs-1040-1988-scan.pdf (CCITT at 400 ppi), documents/scan/reportlab-scanned-receipt.pdf (DCT), vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf (JPX), acrobat11-image-conversion-pdfa1b-image.pdf; remote, the Konica scan remote/ecan/konica-bizhub-c554e-letter-scan.pdf and the heavy documents | No lossy step runs unless named; with each named, every changed image is in the report as pdfimages -list sees it after; text identical; MuPDF's rasters within the recorded threshold; veraPDF's verdicts unchanged | CorpusOptimizationTests.Lossy_steps_run_only_when_named_and_are_reported (new) |
remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf, remote/usgs/us-topo-washington-west-2023.pdf; a color scan over a portal's cap — not in the corpus (below) | Given a target and the steps allowed, the output fits and says how, or the report gives the smallest size reached; the same bytes on two runs; memory flat per image | CorpusOptimizationTests.The_target_size_is_met_or_the_report_says_why (new) |
vendor/us-federal/illustrator-irs-pub1-english.pdf (/PieceInfo, an XMP thumbnail), pdfwriter4-usda-dry-whey-standard-2000.pdf (page thumbnails), vendor/fr-licence-ouverte/fop-dictao-dila-signed-joafe-notice.pdf (standard fonts not embedded) | Each discard removes only what it names, each object in the report; UnembedStandardFonts refused on a PDF/A claim | CorpusOptimizationTests.Discards_remove_only_what_they_name (new) |
| The signed committed documents M04 lists | Optimization refused with PdfSignatureInvalidationException; under AllowInvalidatingSignatures, each broken signature reported | CorpusOptimizationTests.Signed_documents_are_not_rewritten_unless_the_caller_insists (new) |
| Every committed and remote document | The AOT binary opens, validates, rewrites, merges with itself and optimizes it, with the JIT build's bytes, on linux-x64 per pull request and linux-arm64 and win-x64 nightly | CorpusAotTests.The_aot_binary_processes_the_corpus_as_the_jit_build_does (new) |
| Every committed document | The browser-wasm host, in headless Chromium, opens, validates, rewrites and optimizes it losslessly, with the JIT build's bytes or, where recompressed, its decoded streams | CorpusWasmTests.The_browser_host_processes_the_corpus_as_the_jit_build_does (new) |
| The encrypted committed documents, with their recorded passwords; M16's AES-GCM document; M08's WOFF2 test fonts | Each document under R2 to R6 opens in the browser-wasm host through the core's managed primitives, its decrypted content the JIT build's; the AES-GCM one refused with PrimitiveUnavailable, never a crash; each WOFF2 font decoded to the JIT build's sfnt, or, should Brotli be missing, reported font-program.unsupported as M08 says | CorpusWasmTests.Encrypted_documents_and_woff2_fonts_behave_in_the_browser (new) |
| The reference sources, a thousand invoices, the thousand-page journal; the heavy documents nightly | Every workload completes in the container profile under 512 MB; time and peak memory recorded beside headless Chromium's under the same limits and published | CorpusContainerProfileTests.The_reference_workloads_run_within_512_mb (new) |
Every damaged document as seeds, committed and remote | The nightly campaign — mutation over the reader, the writer's round trip and the optimizer; coverage-guided over the reader's entry points — finds no untyped exception, hang or unbounded allocation over fourteen consecutive nights before the milestone closes | FuzzingTests.Opening_a_mutated_document_either_works_or_fails_with_a_typed_exception, FuzzingTests.Rewriting_a_mutated_document_round_trips_or_fails_typed (new), FuzzingTests.Optimizing_a_mutated_document_either_works_or_fails_typed (new), the Fuzzing workflow's record |
| The heavy documents and the hostile files above | Every long operation canceled mid-way returns within 100 ms | CorpusCancellationTests.Every_long_operation_stops_within_its_latency_budget (new) |
| The same operations through the tool | optimize and --linearize produce the API's bytes | CorpusToolTests.Optimize_matches_the_api (new) |
Corpus
What the corpus holds
- Stress: ReportLab's thousand-page journal, committed; W11's three heavy references, remote — 9,302 pages, one
63 MB page whose image decodes to 328 MB (read under
readerLimits), 147 MB of JPEG 2000 —; beside them a 35,000-pixel square CCITT image decoding to 153 MB from 10.5 KB, 65,542 objects in one object stream, 48 incremental updates, an 82-page tagged scan of 10.6 MB (many-pages,single-huge-page,heavy-scan,huge-decoded-image,many-objects,forty-eight-incremental-updates). - The #47 documents, remote: the signed Web Capture file and the VHA coding handbook.
- Duplicates and waste: a logo embedded twice (
duplicated-image), an ICC profile twice (duplicate-icc-profile), unreferenced objects (unreferenced-image-objects,orphaned-objects,unreferenced-objects-in-update,unreferenced-indirect-name-objects), unused fonts (fonts-unused),/PieceInfo(pieceinfo-illustrator,pieceinfo-markedpdf), thumbnails (page-thumbnails,large-xmp-with-thumbnail). - Fonts: four committed documents embedding fonts in full, and remote, Symtrax's PDF/A-3a invoice and Yousign's
signed seal page (
full-font-embedding,type0-full-font-embedded); subset tags reused over different programs (reused-subset-tag,placeholder-subset-tag-shared); a Cerfa form whose/DRembeds Arial in full, an OmniForm form whose/DRembeds Verdana, and two signed documents whose/DRembeds Myriad Pro. - Compression: LZW in four committed documents, ASCII85 chains, uncompressed attachments, uncompressed images (remote), 53 committed documents with object streams, and many classic tables.
- Linearization:
documents/archival/qpdf-linearized-report.pdffrom qpdf, and 58 committed linearized files, 26 of them with hints inconsistent, broken or stale, several linearized then updated. - Scans: every codec M22 decodes, committed and remote.
- Damage: 13 committed and 106 remote damaged documents, the fuzzing seeds.
What it lacks
| Need | Why | Priority | Likely source |
|---|---|---|---|
| A stream that decodes past 2 GB | The Unbounded acceptance needs one; no real file reaches the ceiling, and committing one is pointless | 1 | Generated at test time by the test support — nested Flate over zeros, about 2 MB encoded —, never committed, and not a corpus document: docs/corpus.md names it as the exception to committing |
| A case file merged from several producers' documents, each subsetting the same face | Font consolidation needs real subsets of one face by different subsetters; the corpus's reused tags are the negative case | 1 | Generated here: the LibreOffice, Chromium and ReportLab renderings of invoice-fr.html in one OFL face, merged by qpdf, recorded in build_corpus.py |
Independent expectations for optimization: per committed document, pikepdf's count of distinct decoded streams and duplicate groups, pdffonts' embedded-in-full fonts, qpdf's unreferenced objects | "Every duplicate, and nothing else" must be asserted against a count the file gave an independent tool, not our own optimizer | 1 | Generated here: build_corpus.py writes them into new expect fields, whose schema this milestone adds |
| A color scan too large for a court portal's cap, several pages at 300 ppi or more | The target-size mode exists for it; the heavy scans are archival plates, not a lawyer's exhibit | 2 | A contribution (W03); remote if it cannot be redistributed |
| Chromium's time and memory on the reference workloads under the container's limits | The comparison the profile is published against | 2 | Generated here, in CI, by the comparison benchmarks' Chromium slot; recorded, not committed |
| A document of uncompressed images and content, committed | Recompression's largest gain is shown on remote files only | 3 | Generated here: ReportLab with page compression off |
| A document linearized by Acrobat with object streams and shared objects across pages | The hint tables' shared-object section is proven on qpdf's output and old Distiller files | 3 | A public source among government publications (W02); remote if needed |
Traps
- An optimization that changes what a document says is a corruption, not a trade-off: extracted text, structure, conformance and signatures are the acceptance, not size alone.
- A subset tag lies. Illustrator and Konik write one tag over different programs; consolidation compares glyphs, never names.
- Identity is not bytes. Two equal optional content groups are two layers; two equal annotations, fields or structure elements are two things. Deduplication shares values, never identities.
- A form field needs its whole font. Subsetting a font in
/DRbreaks the next fill, silently, in another tool. CIDSetandCharSetdescribe the program, and PDF/A-1 checks them after every subsetting and every union.- Flate is deterministic per runtime, not across runtimes (M03's trap): a servicing release, or the browser runtime's own deflater, changes every recompressed stream's bytes. Compare decoded bytes across platforms.
- Re-encoding a JPEG loses a generation each time. Every attempt starts from the original samples.
- A mask belongs to its image's grid. Downsample one without the other and every edge shifts.
- An image's resolution is a placement's, not the image's: one XObject drawn as a thumbnail and as a page has two; the largest decides.
- Inline images cannot be deduplicated or resampled without rewriting content; they are left and counted.
- The hint tables' offsets: whether an offset after the primary hint stream counts that stream's own length is
where Annex F and Acrobat's practice have been reported to differ (to verify against qpdf's implementation); qpdf's
--check-linearizationis the referee, and the corpus's 26 inconsistent files show how often producers err. - A signed document cannot be optimized in place: an update only appends, and a rewrite invalidates.
- A budget that is too tight trains contributors to raise it; one that is too loose catches nothing. Set each from a measurement with a stated tolerance, and record every raise.
- Shared runners are noisy: allocation budgets are reliable; time budgets hold only as a ratio measured in one job.
GC.GetAllocatedBytesForCurrentThreadmisses other threads: parallel operations are measured in a child process.- A window that grows must keep what it read, or a large section is read again from its start at each step — the VHA handbook's 683 KB for 273 KB.
trailerappears inside streams, and a scan that does not know where streams are pays for each occurrence.- A stream past 2 GB overflows every
intthat counts it: positions, lengths, progress. - Streaming is not safety: a bomb read piecewise still costs the time to produce it; the length guard stays.
- A process-wide intern table is a leak in a server that reads hostile input.
- Chiseled images have no ICU and no fontconfig: globalization must be invariant, and the native graphics assets must not link fontconfig.
- Chromium in a container needs its sandbox arrangements and shared memory; a comparison that gives it fewer resources than ours, or other fonts, is not a comparison.
- WebAssembly has no file system, no threads by default, no MD5 and no AES — the core's managed ones stand in, and AES-GCM has no stand-in —, and its address space is 32-bit.
Documentation
docs/website/docs/guides/optimization.md(new) — the lossless steps, the discards, the lossy steps, the target size, the report, signed and conforming inputs, theoptimizeverb.docs/website/docs/concepts/performance.md(new) — the budgets, how they are measured and enforced, what they promise and what they do not.docs/website/docs/guides/deployment.md(new) — Native AOT, trimming, WebAssembly, the container profile as a recipe with its measured figures beside Chromium's, reader limits for untrusted input.docs/website/docs/reference/reader-limits.md—MaxDecodedStreamLengthas along;docs/website/docs/concepts/reader-limits.mdandlazy-reading.md— the piecewise decode, the caches and their budgets, the per-document name table.docs/website/docs/reference/diagnostics.md— the codes above.docs/website/docs/reference/tool/—optimize,--linearize.docs/website/docs/introduction.mdanddocs/features/features.json—optimization,aot,hostile-inputandcancellationbrought to their state.docs/architecture.md—Optimization/, the piecewise decode, the caches;SECURITY.md— the threat model as measured;benchmarks/README.md— budgets in CI.docs/corpus.md, the manifest schema andtests/corpus/README.md— the optimizationexpectfields.- The ADRs: budgets in CI (slice 1), the amendment of ADR 34 (the piecewise decode, and why
MaxDecodedStreamLengthstill bounds length), and the three classes of change (extending ADR 42's rule beyond images). docs/status.md— #37, #38, #46, #47, #48, #49 and #50 closed; the budgets, tolerances and measurements; the platform findings; the campaign's record; the container figures.
Exit criteria
- Allocation budgets run on every commit and throughput budgets on every pull request, and each fails CI beyond its
tolerance;
budgets.jsonrecords every figure's source and every raise's reason; the ADR on budgets in CI is accepted. - #37, #38, #46, #47, #48, #49 and #50 are closed, each by the behavior its issue asked for, the #47 markers removed from the manifest.
- The piecewise decode serves every internal consumer; a stream past 2 GB is read whole under
Unbounded; the ADR amending ADR 34 is accepted. - Every lossless step, discard and lossy step, the target size and linearization are implemented, reported, and refused where a signature or a claim forbids them; the ADR on classes of change is accepted.
- The AOT binary and the browser-wasm host process the corpus as the JIT build does; the trimmed host raises no warning; the platform findings are recorded.
- The container profile is published with its figures beside Chromium's.
- The priority-1 gaps above are filled; each remaining gap is recorded in
docs/corpus-contributions.md. - The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on a
green
Remote corpusrun recorded instatus.md. - Unit tests cover each behavior, its degenerate cases and its hostile ones; the FsCheck properties hold; the campaign has run fourteen consecutive nights with no open finding.
- Integration tests run qpdf, pikepdf, poppler, MuPDF, veraPDF, Pillow and libjpeg, Playwright with Chromium, headless Chromium and M14's Factur-X referee, each in a container.
-
OptimizerBenchmarksandStreamDecodingBenchmarksrun withMemoryDiagnoser, and every budget row is inbudgets.json;docs/status.mdrecords the measurements. -
optimizeand--linearizeship in the dotnet tool and the AOT binary, documented. - The documentation site publishes the pages listed above, and
features.jsonmatches what exists. - Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).