Skip to main content

M22 — Imaging and OCR

State: to do — Depends on: M08, M15 — Codecs, lossy steps and the OCR boundary per ADR 42; every decoder bound classified per ADR 34; unsafe code only where a measurement asks for it, per ADR 35; a dependency-free satellite per ADR 9; active content never produced, per ADR 37

Goal​

Decode the pixels of every image a PDF can carry — CCITT in the core; JBIG2, JPEG and JPEG 2000 in the AdCodicem.Pdf.Imaging satellite; their color spaces, masks and functions evaluated —, re-encode bilevel scans losslessly, and give a scan the text it does not have: an invisible, positioned text layer written from the words an OCR engine recognized, so that it can be searched, extracted, redacted and archived as PDF/A-2u. Tell a blank scanned sheet from a page that carries only a signature, and a sideways page from an upright one.

A case file is mostly scans. The exhibits a lawyer receives come off copiers and fax servers as JBIG2, CCITT, JPEG and JPEG 2000; M15 can say that such a page has no text, but nothing after it can act on the page until its pixels can be read: M19 cannot remove a name from a scanned letter, M21 cannot convert a JPEG 2000 image that breaks PDF/A-2's clause, nor bring a scan to level u, M23 cannot downsample, M25 cannot rasterize, M07 cannot recognize a scanned separator sheet, and M18's portal presets cannot make a piece searchable. ADR 42 put the decoders in one fuzzed, bounded codec set so that all of them share it.

The failures this milestone exists to prevent are specific. A decoder that can be driven outside its buffers — the FORCEDENTRY exploit of 2021 went through a JBIG2 decoder's symbol count. A decoder that is almost right — a JPEG off by one where a reference implementation is exact, a CCITT row shifted by one run, a JPEG 2000 plate whose tiles meet with a seam. A re-encoding that changes what a scan says — the pattern-matching JBIG2 of the copiers of 2013, which replaced digits. A text layer whose words land on the wrong line, split in the middle, or without the spaces extractors need. A blank-page detector that drops the page bearing only a signature. A color conversion that gives a different byte on Windows and on Linux.

Scope​

In:

  • CCITT decoding in the core: T.4 one- and two-dimensional and T.6, every parameter of ISO 32000-2 §7.4.6, damaged rows recovered and reported;
  • the image seam in the core: a PdfImage read model over image XObjects and inline images; PdfImageCodecs, which M19 defined, completed with decoders and encoders by filter; IPdfImageDecoder, IPdfImageEncoder and a row sink; PdfImageReader, which runs the filters, the codec, /Decode, color-key masks, stencil and soft masks and /Matte, and delivers rows;
  • color spaces and functions evaluated in the core: CalGray, CalRGB, Lab, ICCBased, Indexed, Separation and DeviceN, and functions of types 0, 2, 3 and 4, converting samples to gray, RGB or CMYK for export and analysis, deterministically on every platform;
  • the AdCodicem.Pdf.Imaging satellite (ADR 42): the JBIG2 decoder (every region type of ISO/IEC 14492 in PDF's embedded organization, with /JBIG2Globals); the JPEG decoder (baseline, extended and progressive Huffman, arithmetic-coded sequential and progressive; gray, YCbCr, RGB, CMYK and YCCK; restart intervals); the JPEG 2000 decoder (Part 1 codestreams, raw or in JP2 and JPX files, with every progression order, tiles, precincts and code-block style, and /SMaskInData); the lossless CCITT G4 and JBIG2 generic-region encoders; and the JPEG block wipe that M19 left to this milestone;
  • the consumers completed: M15's image export gains decoded output with color evaluated; M19 redacts the pixels of CCITT, JBIG2, JPEG and JPEG 2000 images; M21 re-encodes a JPEG 2000 image outside part 2's clause; M07 joins a multi-strip TIFF into one image and transcodes an arithmetic-coded JPEG file, on request; M02's stream rule on decodable filters judges image filters;
  • the text layer: a PdfRecognizedPage model; readers for hOCR 1.2, ALTO v2 to v4 and Tesseract's TSV; a glyphless font; PdfTextLayer, which writes invisible (render mode 3), positioned words, optionally tagged; IOcrEngine and PdfOcr, which pick the pages without text, decode their images, call the caller's engine and write the layer; the IPdfPageRasterizer seam M25 fills;
  • blank-page detection on pixels, beside M07's detection by content, as a predicate M07's split takes;
  • page orientation from the text layer or from the engine, applied as /Rotate;
  • the command-line tool's text-layer and blank verbs, and decoded output for inventory --images;
  • a fuzzing harness for every decoder and for the recognized-text readers, run per commit and nightly from the first slice.

Out, explicitly:

  • running an OCR engine — the caller's, behind IOcrEngine (ADR 42); a first-party engine satellite stays an open question of the roadmap. A sample shows an adapter; no package ships one;
  • rasterizing a page that is not one image — vector content, several images, text over an image — so that an engine can read it: M25 fills IPdfPageRasterizer; until then such a page is reported and skipped;
  • every lossy step — downsampling, JPEG re-encoding, conversion to gray or to one bit per pixel — M23, where each is opt-in and reported (ADR 42); M22's only encoders are lossless;
  • pattern-matching or lossy JBIG2 — never (ADR 42); JPEG 2000 encoding — not planned; HTJ2K (ISO/IEC 15444-15) — not planned, refused with a diagnostic;
  • pixel deskew, despeckle and binarization, and orientation found from pixels alone — an open question of the roadmap; mixed raster content compression — an open question;
  • color management — ICC transforms, conversion to an output intent — M29's satellite; M22's conversions serve export and analysis and say where they approximate;
  • reading barcodes from pixels — no milestone plans it: barcode and patch-code recognition is an open question of the roadmap; M07's separator predicate can call a caller's decoder over M22's rows;
  • tagging beyond one paragraph per recognized block — headings, lists, tables inferred from a scan — the roadmap's open question on heuristic tagging;
  • lossless JPEG (SOF3), hierarchical processes and 12-bit precision — refused with a diagnostic, reopened by a document that needs them; JBIG2's color extension and JPEG 2000 features outside §7.4.9 — refused;
  • rendering — M25, which composes these decoders with M15's interpreter.

Design​

Where it lives​

PartWhereWhy
CCITT G3 and G4 decoderCore, IO/Filters/, internalADR 42: small, needed by M07's pages and by fax-era scans
PdfImage, PdfImageCodecs, IPdfImageDecoder, IPdfImageEncoder, IPdfImageRowSink, PdfImageReader, PdfBlankPageDetectorCore, Images/, namespace AdCodicem.Pdf.ImagesThe seam M19 named; M07, M15, M19, M21, M23 and M25 consume it; blank detection needs only rows
PdfColorSpace evaluation, PdfFunctionCore, Graphics/, namespace AdCodicem.Pdf.GraphicsEvaluation needs no codec: M15 reads the spaces, M19 writes an overlay's color into samples, M25 evaluates shadings with the same functions
JBIG2, JPEG and JPEG 2000 decoders; G4 and JBIG2 generic encoders; the JPEG block wipeAdCodicem.Pdf.ImagingADR 42: the large decoders, and their attack surface, in a package a caller adds knowingly
PdfRecognizedPage, the hOCR, ALTO and TSV readers, the glyphless font, PdfTextLayer, IOcrEngine, PdfOcr, PdfPageOrientationCore, Ocr/, namespace AdCodicem.Pdf.OcrADR 42: the layer is written in the core, engines live outside it
IPdfPageRasterizer, PdfRasterCore, Content/, beside M15's device seamOne seam for every raster of a page — OCR here, M24's visual comparison, M25's implementation —, in the core so that neither Compare nor Rendering depends on the other
The fuzzing harnesstests/AdCodicem.Pdf.FuzzingA test project, never packaged
text-layer, blank, inventory --images --decodeAdCodicem.Pdf.ToolThe tool ships the Imaging satellite

The satellite depends on the core alone: no Skia, no native library, IsAotCompatible from its first commit. It is a new package, so #42 (an API baseline per package, which M10 had to close) applies to it. Its public surface is small: PdfImagingCodecs.Default — a PdfImageCodecs value holding its decoders and encoders — and the options records; the codecs themselves are internal. The MQ arithmetic decoder, which JBIG2 and JPEG 2000 share, is written once there.

The image seam​

PdfImage read model over an image XObject or an inline image: width, height, color space, bits per
component, filters and their parameters, /Decode, /ImageMask, /Mask (a stencil stream or a
color-key array), /SMask with /Matte, /SMaskInData, /Interpolate, /Intent. Placements are M15's
PdfImageCodecs immutable (M19): decoders and encoders by filter name. PdfImageCodecs.Core holds the core's
filters and CCITT; PdfImagingCodecs.Default adds the satellite's. Passed in an operation's
options, never registered statically (invariant 8)
IPdfImageDecoder Decode(PdfImageDecodeContext, ReadOnlySpan<byte> encoded, IPdfImageRowSink sink). Public, so that
a caller may bring a codec of its own, native included, under the documented contract below
IPdfImageEncoder Encode(PdfImageLayout, rows, IBufferWriter<byte>) -> filter name and parameters
IPdfImageRowSink BeginImage(PdfImageLayout); WriteRow(y, samples); EndImage(PdfImageDecodeResult): rows in order,
top down, each written once
PdfImageReader Open(image, codecs, options) -> rows with /Decode applied, bit depth expanded as asked (1, 8 or
16 bits), optionally converted to gray, RGB or CMYK; masks delivered as images of their own, at
their own size, or resampled onto the image's grid when the caller composites
PdfImageRows a pooled sink for a caller that wants the whole image, bounded like any decoded stream
PdfImageDecodeResult rows delivered; rows missing, filled as the referees fill them and reported; approximations
  • Rows, not bitmaps. Every decoder writes rows in order into a sink, and holds only what its format obliges it to: CCITT two reference rows; a baseline JPEG one row of MCUs; a JPEG 2000 image one row of tiles; a JBIG2 page one stripe, or the page when it is not striped; a progressive JPEG the coefficients of the whole image, since its last scan may refine its first block. That is each codec's working set, and it is what the guard below bounds. The shape is the one M23's piecewise decoding (#48) generalizes to every filter.
  • The encoded data reaches the codec through M01's pipeline — a FlateDecode or an ASCII85Decode before a DCTDecode —, bounded as every filter is (ADR 34). The image filter must be last (§7.4); one found elsewhere is reported and the image yields no rows.
  • PdfStream.Decode() keeps stopping at an image filter, as it has since M01: its callers want data, not pixels, and M15's pdfimages -all path wants the encoded bytes. Pixels are PdfImageReader's.
  • Masks are images. A stencil mask (/ImageMask, or a /Mask stream) is one bit deep with its own /Decode; a color-key /Mask is a set of ranges compared with the samples before /Decode; an /SMask is a gray image of any size, un-premultiplied by /Matte when present; /SMaskInData makes the JPEG 2000 decoder deliver its opacity channel as the mask.
  • The contract of IPdfImageDecoder is invariant 4's: the encoded bytes are hostile, nothing is allocated from a value read without a checked bound, every loop ends by what it produces, and failure is a PdfImageDecodeResult, not an exception. A caller's own codec is held to it by documentation; ours by the fuzzing campaign.

CCITT in the core​

  • T.4 Modified Huffman (/K 0), T.4 Modified READ (/K > 0, one-dimensional rows every K) and T.6 (/K < 0); /EndOfLine, /EncodedByteAlign, /EndOfBlock, /BlackIs1, /Columns (1728 by default), /Rows, /DamagedRowsBeforeError.
  • Table-driven over a bit reader on a span: the changing elements of two rows in pooled arrays of Columns + 2 entries, the row expanded into the sink's buffer, 0 B allocated per row.
  • A damaged row — a code no table knows, a run past the row's end, a row that ends early — is replaced by the previous row (white for the first) and counted; decoding resynchronizes at the next EOL when the stream has them, and stops where it has none. After /DamagedRowsBeforeError consecutive damaged rows the image ends there. Missing rows are white. One image.rows-damaged or image.data-truncated diagnostic carries the counts; never an exception.
  • /Rows absent or disagreeing with the image's /Height: the height wins, as every viewer does, and the disagreement is reported. /EncodedByteAlign pads differently for /K < 0 and /K ≥ 0, and producers get it wrong both ways: the declared reading is tried first, the other only when the first fails in its first rows, and that is reported.

JBIG2​

  • The embedded organization of ISO/IEC 14492 as §7.4.7 uses it: no file header, the segments of /JBIG2Globals read before the page's own, one page per image. Segment types: symbol dictionary; text region, immediate and intermediate; pattern dictionary; halftone region; generic region, arithmetic with typical prediction (TPGDON) and MMR; generic refinement region; page information; end of stripe; end of page; tables; extension segments ignored when T.88 marks them unnecessary, refused when it marks them necessary.
  • Arithmetic and Huffman decoding of symbol dictionaries and text regions with refinement and aggregation; the standard tables B.1 to B.15 and custom tables; the combination operators; the page's default pixel and default operator; intermediate regions composed onto the page.
  • Striped pages: a page information segment of height 0xFFFFFFFF grows by stripes, and each end-of-stripe segment releases the rows above it to the sink, so a striped page holds one stripe.
  • Counts are 64-bit, and checked before they size anything. The number of symbols a text region may address is the sum of the exports of every dictionary it refers to plus its own; FORCEDENTRY (CVE-2021-30860) overflowed exactly that sum in 32 bits, and a later check trusted the overflowed value. Every count — symbols, instances, referred-to segments, table lines, pattern counts — is computed in 64 bits and checked against the working-set guard before an array exists.
  • The arithmetic decoder reads 0xFF past the end of its data, as T.88 specifies. Every loop therefore ends by what it produces — a region's rows, a dictionary's declared symbols — and never by what it consumes.
  • The Xerox copier scans in the corpus draw their text through symbol dictionaries and text regions: pattern matching, the mechanism whose substitutions made copiers of 2013 change digits. The decoder reproduces what the file says; the encoder below never writes a symbol.

JPEG​

  • Processes: baseline (SOF0), extended sequential Huffman (SOF1) and progressive (SOF2) at 8 bits; arithmetic-coded sequential and progressive (SOF9, SOF10), which libjpeg-turbo writes and M07 asked to transcode. Lossless (SOF3), hierarchical (SOF5 to SOF7, SOF13 to SOF15) and 12-bit precision are refused with image.process-unsupported.
  • The reference is libjpeg-turbo with its defaults, which Pillow uses and most readers link: the accurate integer inverse DCT, triangular ("fancy") upsampling of subsampled chroma, and the integer YCbCr tables. The algorithms are the standard's and IJG's documentation; the target being a reference implementation's output, "correct" means identical to it, not close. Other decoders differ from it — MuPDF bundles its own libjpeg — and against them the difference is measured and recorded, never asserted as ours to close.
  • Color: one component is gray; three are YCbCr unless an Adobe APP14 marker says transform 0, the component identifiers read R, G, B, or /ColorTransform is 0; four are CMYK, or YCCK when the Adobe marker says transform 2. Where the marker and /ColorTransform disagree, the marker wins and the dictionary entry is ignored, as §7.4.8 says; without either, three components are YCbCr and the others not. Adobe's inverted CMYK is delivered as stored; the image's /Decode decides (M07 writes [1 0 1 0 1 0 1 0] for such files).
  • Damage: restart intervals resynchronize a damaged scan at its next marker; data that ends early leaves the rest of the image as libjpeg-turbo leaves it; both are reported. A stream that is not a JPEG at all — the PNG a producer labeled DCTDecode, in the remote corpus — yields no rows and image.data-invalid.
  • EXIF orientation means nothing inside a PDF: the placement matrix decides, and M07 puts a JPEG file's orientation there. The decoder never rotates.
  • The block wipe (M19's open decision, below): the entropy-coded data decoded to quantized coefficients; the 8 × 8 blocks of every component that intersect a mark given one DC value — the overlay's color, quantized — and zero AC coefficients; the whole re-encoded as a baseline JPEG with the file's quantization tables and optimized Huffman tables. Every block outside the marks keeps its coefficients exactly — no generation loss, as jpegtran's -wipe offers. The marked region grows to MCU boundaries, up to 16 pixels, and the report says by how much. Progressive and arithmetic-coded inputs come out baseline Huffman. The baseline re-encoder written here is the one M23's lossless entropy re-coding reuses.

JPEG 2000​

  • Codestream: SIZ, CAP, COD, COC, QCD, QCC, RGN, POC, PPM, PPT, TLM, PLM, PLT, CRG, COM, SOT, SOP, EPH; the five progression orders (LRCP, RLCP, RPCL, PCRL, CPRL) and progression changes; tile-parts in any order; precincts; every code-block style (selective arithmetic bypass, context reset, termination on each pass, vertically causal context, predictable termination, segmentation symbols); scalar derived and expounded quantization; region of interest by max-shift; the reversible 5/3 and irreversible 9/7 wavelets and component transforms; subsampled components upsampled to the image's grid.
  • Files: JP2 and JPX boxes — jp2h, ihdr, bpcc, colr (enumerated sRGB, grayscale and sYCC; a restricted ICC profile), pclr and cmap (palettes), cdef (opacity, premultiplied or not), res — and raw codestreams.
  • PDF's rules (§7.4.9): a /ColorSpace in the image dictionary overrides the file's color; /SMaskInData 1 or 2 turns the opacity channel into the mask; /Decode is ignored except for an image mask; the bit depth is the codestream's, so 16-bit samples reach the sink as 16-bit samples.
  • Tier-1 is the hot loop: the MQ decoder and the three coding passes over pooled code-block buffers reused across blocks, 0 B allocated per code block once warm.
  • The 9/7 wavelet is floating point, and still deterministic: single precision in a fixed order of operations, no fused multiply-add unless written explicitly, the same bytes on x64 and ARM64 — tested on both. The reversible path is integer and exact. OpenJPEG, the referee, is not bit-exact with other irreversible decoders either, so the tolerance against it is measured on the 9/7 path, and zero on the 5/3 path.
  • Reduced resolution: decoding may stop at a coarser resolution level — a quarter, a sixteenth of the pixels —, which M23's downsampling and M25's thumbnails want; the sink is told the smaller layout, and the result says so.
  • HTJ2K codestreams (a CAP marker announcing Part 15) are refused with image.process-unsupported.

Color spaces and functions​

M15 reads each color space's family and parameters, and until now exported Lab, Separation and DeviceN samples as their components. M22 evaluates them.

  • PdfColorSpace gains ToGray, ToRgb and ToCmyk over rows of components, through a lookup table built once per image and color space where the input depth allows it (every 8-bit case), so that the hot loop is a table read.
  • CIE-based spaces: CalGray, CalRGB and Lab through CIE XYZ to sRGB, with Bradford adaptation from the space's white point to D65. ICCBased: a profile recognized as sRGB IEC 61966-2.1 — by its header and description, the readers M20 wrote — converts as the identity; any other through its /Alternate, or the device space of its /N, reported as image.color-approximated. Indexed through its lookup string. Separation and DeviceN through their tint transform into the alternate space, /All and /None honored, a DeviceN's /Process and /Colorants read. DeviceCMYK to RGB by the device formulas of §10, reported as approximate. A color-managed conversion is M29's.
  • Functions: type 0 (sampled — /Encode, /Decode, 1 to 32 bits per sample; /Order 3, cubic spline interpolation, evaluated as linear, as pdf.js and poppler do, and reported with the conversion's image.color-approximated, since §7.10.2 ignores it only when a /Size is under 4), type 2 (exponential), type 3 (stitching, the last interval closed at its upper bound) and type 4 (the PostScript calculator). A type 4 program is compiled once into a compact form and run on a fixed stack of 100 operands — §7.10.5 requires room for at least 100 entries of every implementation, requires no more, and makes an overflow an error —, with if and ifelse as jumps; having no loops, it costs its length. roll, index and copy are bounds-checked; a failure yields the range's lower bounds and one function.invalid per function.
  • Transcendental functions are ours. The power of a gamma curve, Lab's cube root, and type 4's exp, ln, log, sin, cos and atan are computed by routines written from IEEE basic operations, which are correctly rounded everywhere; Math.Pow and its kin defer to the platform's C library, whose last bit differs between Linux, Windows, macOS and WebAssembly, and a last bit becomes a byte at a rounding boundary. Invariant 6 holds only if the same file converts to the same bytes on every platform, and the test runs on x64 and ARM64.
  • A type 3 function is evaluated iteratively; one that reaches itself is cut and reported. An Indexed lookup shorter than (hival + 1) × components gives black for the missing entries, and an index above hival clamps, as viewers do; both are image.indexed-out-of-range.

Lossless encoders​

  • CCITT G4 (T.6) for any 1-bit image: /K -1, /BlackIs1 chosen to match the source's polarity, /EncodedByteAlign false.
  • JBIG2 generic region: one immediate lossless generic region per image, arithmetic-coded with template 0, its nominal adaptive pixels and typical prediction; no globals. Never a symbol dictionary or a text region: that is where pattern matching lives (ADR 42).
  • A policy chooses: Smallest encodes both and keeps the smaller, ties going to G4, the older and more widely read. M19 calls it for redacted scans, M23 for recompression, M07 to join a multi-strip bilevel TIFF.
  • Flate with the PNG predictors, from M03's deflater, stays the lossless encoder for everything that is not bilevel: M21 re-encodes a JPEG 2000 image outside part 2's clause with it, and M19 a redacted JPEG 2000 image.

Pixels under a redaction​

M19 left to this milestone how each codec's pixels are edited. The choice is recorded in an ADR at the next free number:

FilterEdited throughWhat stays exact
CCITT, JBIG2Decoded, the samples under each mark set, re-encoded in the smaller of G4 and JBIG2 genericEvery sample outside the marks; a JBIG2 image drawn from symbols comes out as one generic region
DCTThe block wipeEvery coefficient outside the marked MCUs; decoded pixels outside them and their one-pixel border, since chroma upsampling reads neighboring blocks
JPXDecoded, re-encoded as FlateEvery decoded sample outside the marks; the file grows, and the report says by how much
Flate, LZW, RunLength, uncompressedM19's own pathUnchanged

The text layer​

PdfRecognizedPage an engine's words: the image's size in pixels and its resolution; the frame the boxes are in —
ImagePixels of a named image, or PageVisual at a resolution (M15's hOCR and ALTO exports); the
page's orientation and its confidence; blocks > lines > words, each with a box, a baseline and
angle where known, a confidence and a language
PdfRecognizedText ReadHocr, ReadAlto, ReadTesseractTsv (Stream, PdfRecognizedTextOptions) -> pages
PdfTextLayer Add(page, recognized, PdfTextLayerOptions) -> what was written, and its report entries
PdfTextLayerOptions immutable: tagging (None, Paragraphs); language when the engine gives none; minimum word
confidence; a page with text (Skip, ReplaceOwnLayer, ReplaceInvisibleText); orientation
(FromEngine, FromTextLayer, None)
IOcrEngine Info (name, version, model) and RecognizeAsync(PdfOcrImage, PdfOcrRequest, CancellationToken)
PdfOcrImage the page image as the engine asks for it: Gray8, Bilevel or Rgb24 rows, or encoded PNG or TIFF
PdfOcr RunAsync(document, engine, PdfOcrOptions, IProgress<PdfProgress>?, CancellationToken)
-> PdfOcrReport: per page, recognized or skipped and why, words, mean confidence, the engine's
name, version and model
IPdfPageRasterizer Rasterize(page, resolution, cancellationToken) -> PdfRaster: the seam M25 fills, through which a
page that is not one image is rendered for the engine; M24's visual comparison consumes the
same seam
PdfRaster width, height, stride, format (Gray8, Rgb24, Rgba32), a pooled buffer, the page frame it covers
  • Frames. Tesseract's hOCR, ALTO and TSV give pixels of the image the engine saw, origin top left, y down. A box maps to default user space through the page image's placement matrix — rotated, skewed and cropped placements included, the Xerox 5755's deskewed matrix among them —, or through the page's visual frame at a resolution for M15's exports: crop box, /Rotate and /UserUnit applied once, as M15 converts them. hOCR's ocr_page box, scan_res and ppageno and ALTO's MeasurementUnit (pixel, mm10, inch1200) are read; an engine that resized the image is scaled back from the image's size in the file.
  • The font: one Type0 font per document, its CIDFontType2 descendant written by M08's TrueType writer — an empty glyph per code, one advance for all (500 units per 1,000 em), /DW equal to it, so that the dictionary and the program agree as PDF/A checks —; codes assigned to distinct character sequences in the order they first appear, and ToUnicode mapping each to its text, supplementary planes and combining sequences included. The .notdef glyph is never drawn.
  • The words: one text object per line in render mode 3; each word placed by a text matrix at its box's start on the line's baseline — hOCR's baseline, else the box's bottom less a descender estimate — along the line's angle (hOCR's textangle, ALTO's ROTATION), at a size from the line's x_size or height, and scaled horizontally (Tz) so that its advance equals its box's width. A space code is placed in the gap before the next word of the line, so every extractor sees the space, not only those that infer one from a gap.
  • Where it goes: a new content stream appended to the page's /Contents, preceded, when M09's balance analysis says the page leaves its state altered, by a q stream at the front — so that the layer starts in the page's initial state —, and bracketed by a marked-content sequence under a private tag the library recognizes. That is how ReplaceOwnLayer removes exactly its own layer, idempotently as M09's stamps are, and how M19 names it.
  • Tagging. Paragraphs, in an untagged document: a structure tree is created — Document, then one P per recognized block with its Lang — with /MarkInfo /Marked true, and the page image made an /Artifact. In a tagged document, the page's layer is appended as a Part when no structure element owns the page's content; otherwise it is written inside an /Artifact sequence and ocr.structure-not-extended says so. Headings, lists and tables are not inferred.
  • A page that has text — M15's Text or OcrLayer — is skipped by default. ReplaceOwnLayer removes a layer this library wrote; ReplaceInvisibleText removes any invisible text through M19's content editing pipeline, a copier's layer included, and reports it.
  • PdfOcr sends the engine the image the page consists of — one image covering the crop box as M15's ImageOnly threshold measures it, under a text layer being replaced or none —, decoded by the codecs given, in the form and at the resolution the engine asks for; mixed raster content (a JPEG background under JBIG2 or CCITT masks) is composited first. Any other page goes to IPdfPageRasterizer when one is given, and is otherwise skipped as ocr.page-skipped. Pages are processed one at a time, with M03's progress (stages Decoding, Recognizing, WritingTextLayer) and cancellation between them; whether the engine runs pages in parallel is the caller's affair, and the document's writes stay in page order.
  • Determinism. The layer is a function of the recognized words. The report records the engine's name, version and model as IOcrEngine.Info gives them, since invariant 6 holds only for an engine the caller pins (ADR 42).
  • The readers are hostile-input parsers: XmlReader with DTD processing prohibited and no resolver, a depth and a size bound — MaxInputLength, 64 MB by default, with a typed exception, since recognized text is the caller's input rather than the file's —; hOCR's title micro-syntax parsed by hand and bounded; a TSV row with the wrong number of fields skipped and counted.

Blank pages​

  • PdfBlankPageDetector.Examine(page, codecs, options) -> PdfBlankVerdict: Blank, NotBlank or Undetermined, with the ink ratio, the method and the evidence. Detection removes nothing.
  • By content first, M07's rule extended with M15's facts: nothing painted inside the crop box, or only invisible text, white fills and artifacts. A page with visible text is never blank.
  • By pixels when the page is one scanned image — M15's ImageOnly, or OcrLayer whose layer is empty: the image read row by row; a margin excluded, 5 % of each side by default, where scanner edges and punched holes are; a pixel counted as ink when its luminance after /Decode falls below a threshold — Otsu's over the page's histogram, clamped to a band, for gray and color; the sample itself for one bit —; connected ink components smaller than a speck, 0.3 mm across at the image's resolution by default, discarded, by a union-find over the runs of two rows in bounded memory; and the page blank when the ink left is under a ratio, 0.1 % by default. Integer arithmetic throughout, so the verdict is the same on every machine.
  • The verdict is a predicate for M07's split by separator (PdfPagePredicates.BlankByPixels); removing blank pages is M06's page removal, reported page by page.

Orientation​

  • PdfPageOrientation.Detect(page) -> (0, 90, 180 or 270, confidence, evidence) from M15's lines in the displayed frame: glyphs counted by direction, vertical writing modes left out, a rotation proposed only when enough glyphs agree and one direction dominates — the count and the ratio fixed by slice 11 against the corpus and recorded —, and Undetermined otherwise.
  • Apply sets /Rotate on the page itself, never on an ancestor since the attribute is inherited, to the value that makes the dominant direction read left to right, and reports the old and new values (orientation.applied). Content, annotations and the text layer are untouched: /Rotate turns them together.
  • A change of /Rotate in a signed document is a page change under M04's classification: refused on a certified document unless the caller insists, reported otherwise.
  • With PdfOcr, FromEngine takes the orientation the engine reports — Tesseract's orientation and script detection gives one —, writes the words along the rotated baselines, and sets /Rotate in the same revision.

Bounds, classified (invariant 12, ADR 34)​

BoundClassWhy
An image's decoded samplesThe existing guard MaxDecodedStreamLengthAn image's samples are its stream's decoded data
A decoder's working set beyond the rows it hands out — progressive JPEG coefficients, a row of JPEG 2000 tiles, a JBIG2 page or stripe with its symbol bitmaps and intermediate regionsGuard: PdfReaderLimits.MaxImageWorkingSet, 1 GiB by default, code limit.image-working-setA valid 256 MB image in one JPEG 2000 tile holds four times its samples as 32-bit coefficients; the default admits every image the decoded-length default admits
Scans in a progressive JPEGGuard: PdfReaderLimits.MaxJpegScans, code limit.jpeg-scans, its default set above every scan script the corpus and the common encoders writeEach scan is a pass over the image and the standard sets no limit, so a valid file can exceed any number
JPEG 2000 tiles, layers, decomposition levels, code-block size, coding passes, componentsInternal constants: the standard's own maximaThey are field widths and ranges ISO/IEC 15444-1 defines; no valid codestream exceeds them
JBIG2 symbol counts, referred-to segments, table lines, pattern countsComputed in 64 bits and checked against the working setCounts the file declares; what they size is the working set
The PostScript calculator's stackInternal constant, 100§7.10.5: no processor is required to hold more, and overflowing is an error, so no conforming writer can count on more
Stitching depth; color-space nesting (Indexed over Separation over ICCBased)Evaluated iteratively, cycles cutA chain is no longer than the objects in the file, each read once
hOCR, ALTO and TSV inputPdfRecognizedTextOptions.MaxInputLength, a typed exceptionThe caller's input, not the file's: not a reader guard

Each guard is a PdfReaderLimits property with its limit.* code, its rows in ReaderLimitsTests, its line on the reader-limits page, and its key in the manifest schema's readerLimits, which CorpusManifestSchemaTests holds to PdfReaderLimits.

Diagnostics and findings​

In PdfDiagnosticCodes, disjoint from rule identifiers (ADR 36) and from M12's image.format-unsupported, image.animation-dropped and image.resolution-excessive:

CodeSeverityMeaning
image.rows-damagedWarningCCITT rows that did not decode, replaced by the previous row; the count
image.data-truncatedWarningThe data ended before the image did; rows or blocks missing, filled as the referees fill them
image.data-invalidWarningData the decoder could not read at all; the image yields no rows
image.process-unsupportedWarningA JPEG process, a JBIG2 segment or a JPEG 2000 capability the decoders do not implement
image.color-approximatedInformationA conversion through an alternate space or the device formulas
image.mask-mismatchWarningA mask whose size or depth disagrees with its image; resampled onto its grid
image.indexed-out-of-rangeWarningA lookup shorter than hival needs, or indices above it
function.invalidWarningA function that cannot be evaluated as written: a stack overflow, a bound out of range, a cycle
ocr.page-skippedInformationA page not recognized, and why: it has text; it is not one image and no rasterizer was given
ocr.word-outside-pageWarningA recognized word whose box falls outside the crop box; dropped
ocr.input-invalidWarningRecognized text the reader could not use: a malformed bbox, a TSV row with missing fields
ocr.structure-not-extendedWarningA tagged page whose structure could not take the layer; the layer written as an artifact
orientation.applied, orientation.undeterminedInformationA page turned, with its evidence; a page whose text could not decide
limit.image-working-set, limit.jpeg-scansWarningThe guards above; the message names the property that lifts each

Validation: M02's stream rule on decodable filters, which could not judge image filters, judges them now — CCITT always, the others when the validator's options carry PdfImagingCodecs.Default. Its identifier stays M02's; a finding on an image names the codec that decided. Each document whose image data does not decode gains it in expect.findings.

Consumers completed​

  • M15: PdfImageExtractor gains a Decoded output — PNG for gray, RGB and palette samples of 1 to 16 bits, TIFF for CMYK and for separations kept as components —, with pdfimages -png and -tiff as referees; the Lab, Separation and DeviceN samples it exported as components are converted, or kept as components on request.
  • M19: given PdfImagingCodecs.Default, the redactor edits CCITT, JBIG2, DCT and JPX pixels as the table above says; the documents it recorded as unsupported for pixel redaction until M22 lose that marker.
  • M21: the JPEG 2000 remedy row — an image outside part 2's clause re-encoded as Flate, reported as a change of encoding; a scan reaches level u once a text layer is written first.
  • M07: PdfImagePageOptions.JoinStrips decodes the strips of a TIFF frame and re-encodes them as one image — G4 or JBIG2 for bilevel, Flate otherwise —, losslessly and reported; an arithmetic-coded JPEG file is transcoded to Flate on request, and refused as before otherwise.
  • M02: the stream rule above.

The command-line tool​

text-layer FILE (--hocr F… | --alto F… | --tsv F…) [--frame image|page --dpi N] [--pages SEL] [--tag] [--replace own|invisible] [--orient] -o OUT; blank FILE [--list | --remove | --split] [--ink RATIO] [--margin PCT]; inventory FILE --images DIR --decode [png|tiff]. Each takes M06's page selection and exit codes and writes its report as JSON on request; the AOT binary produces what the API produces. The tool never runs an engine: it writes a layer from the files an engine wrote.

Slices​

Each slice ends on a green commit, with the codes, guards and rules it introduces documented, its benchmark recorded in docs/status.md, and its decoder in the nightly fuzzing campaign from the commit that adds it.

  1. The seam and CCITT. Delivers PdfImage, PdfImageCodecs completed, IPdfImageDecoder, IPdfImageRowSink, PdfImageRows, PdfImageReader with /Decode, masks and bit expansion over the core's filters, the CCITT decoder, MaxImageWorkingSet classified, the image.* codes it needs. Delivers the fuzzing harness and an ADR on its engine: coverage-guided — SharpFuzz with libFuzzer is the candidate, its fitness for .NET 10 established here — beside FuzzingTests' seeded mutations, which extend to encoded image streams (FuzzingSeeds gains the corpus's image streams). Proved by unit tests per code-table row (terminating and make-up codes, EOL, fill bits, the two-dimensional modes), per parameter, and per fault (damaged rows with and without EOL, a run past Columns, /Rows against /Height); an FsCheck property — for any byte sequence and parameters the decoder writes exactly Height rows of Columns bits and reads nothing outside its input —; random bitmaps encoded by libtiff in the container, committed as fixtures, decoded back to themselves; integration: every CCITT image in the corpus equals MuPDF's and poppler's decoding, compared as hashes of rows; CcittBenchmarks at 0 B per row. Leaves the satellite.
  2. The Imaging satellite and JBIG2 generic regions. Delivers the project and package with its API baseline (#42), PdfImagingCodecs.Default, the MQ decoder shared with slice 6, generic regions (templates 0 to 3, adaptive pixels, TPGDON, MMR through the CCITT decoder), page information, end of stripe, end of page, immediate and intermediate regions, /JBIG2Globals. Proved by unit tests from T.88's worked examples, a striped page of unknown height, a generic region whose data ends early; integration: the generic regions of vendor/us-federal/xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdf and docusign-pdfkit-gsa-sf30-contract-modification.pdf equal jbig2dec's (MuPDF's container) and poppler's decoding. Leaves symbols.
  3. JBIG2 symbols, text, refinement and halftones. Delivers symbol dictionaries (arithmetic and Huffman, with refinement and aggregation), text regions, generic refinement, pattern dictionaries, halftone regions, the standard and custom Huffman tables, and the 64-bit counting discipline. Proved by unit tests — a text region whose three dictionaries' exports sum past 2³², refused before any allocation (FORCEDENTRY's shape); a refinement with typical prediction; a halftone on a skewed grid —; integration: the Xerox IMF scan (symbol dictionaries in /JBIG2Globals, arithmetic text regions) and the Xerox 5755 MRC scan equal jbig2dec's and poppler's; the segment types no committed document uses, from the conformance streams not in the corpus (below). Leaves JPEG.
  4. JPEG, sequential. Delivers the marker parser, Huffman and arithmetic sequential decoding, the integer inverse DCT, upsampling, color conversion, the Adobe and JFIF rules, restart intervals, the damage rules. Proved by unit tests per marker and per fault (a table redefined between scans, a missing restart marker, a scan naming a component the frame lacks); a property over fixtures cjpeg encoded from generated samples at every sampling factor and committed — each decodes byte-identical to djpeg's output, committed beside it —; integration: every committed baseline JPEG equals Pillow's decoding; MuPDF's and poppler's differences measured and recorded; JpegBenchmarks, and a JPEG row in the comparison benchmarks against SkiaSharp's decoder, recorded. Leaves progressive.
  5. JPEG, progressive, and the block wipe. Delivers spectral selection and successive approximation, arithmetic progressive, the coefficient buffer under MaxImageWorkingSet, MaxJpegScans, the coefficient-domain wipe and the baseline re-encoder with optimized Huffman tables, and the ADR on pixels under a redaction. Proved by unit tests (every scan script jpegtran -progressive and mozjpeg write, an image cut after its first scan, 10⁵ scans reaching the guard); integration: the eight progressive images of the committed invoices and forms equal Pillow's decoding; after a wipe, the quantized coefficients read back through libjpeg in the Python container are identical outside the marked MCUs, and the decoded samples identical outside them and their one-pixel border. Leaves JPEG 2000.
  6. JPEG 2000, the core path. Delivers the codestream parser, tier-2 for every progression order, tier-1 on the shared MQ decoder, dequantization, both wavelets and component transforms, tiles and tile-parts, subsampled components, the JP2 boxes and color, raw codestreams. Proved by unit tests per marker and per fault; the committed vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf — 2717 × 3701, twelve tiles of 1024², RPCL, six layers, the 9/7 wavelet, SOP and EPH markers, segmentation symbols — within the tolerance of OpenJPEG's opj_decompress and MuPDF; the conformance codestreams not in the corpus (below); the determinism test on x64 and ARM64 runners. Leaves the rest of Part 1.
  7. JPEG 2000, complete. Delivers precincts and POC, PPM and PPT, every code-block style, region of interest, palettes and cdef with /SMaskInData, 16-bit samples, sYCC, reduced-resolution decoding, HTJ2K refused; JpxBenchmarks. Proved by unit tests (a POC that revisits a resolution, packed headers split across tile-parts, a 256-entry palette with 16-bit outputs); integration, remote: the rows below against OpenJPEG. Leaves color.
  8. Color spaces, functions and decoded export. Delivers the conversions, the four function types, the type 4 compiler, masks composited and un-premultiplied by /Matte, M15's Decoded export, inventory --images --decode. Proved by unit tests with known values (the white point of Lab D50 to sRGB white; CalRGB with the sRGB matrix and gamma against the formulas; each type 4 operator; the stitching boundary rule; a lookup cut short); an FsCheck property — our transcendental routines are within one unit in the last place of Math's on random inputs, and give the same bits on x64 and ARM64 —; integration: pdfimages -png and -tiff equal ours for device and indexed spaces on every committed image, MuPDF's RGB within the tolerance for CIE-based and tint-transformed spaces; the decoded-sample hash added to M15's expect.images rows where the referees agree, with its schema. Leaves the encoders.
  9. Lossless encoders and the consumers. Delivers the G4 and JBIG2 generic encoders, the Smallest policy, M19's pixel path, M21's JPEG 2000 remedy, M07's strip joining and arithmetic-JPEG transcoding, and M02's stream rule over image filters. Proved by FsCheck — any bitmap encoded and decoded is itself —; our encoders' output decoded by libtiff and jbig2dec in the container is the input; integration: the M19, M21 and M07 rows below, veraPDF on M21's output. Leaves the text layer.
  10. The text layer. Delivers PdfRecognizedPage, the three readers, the glyphless font, PdfTextLayer with its frames, tagging and replacement, IOcrEngine, PdfOcrImage, IPdfPageRasterizer and PdfRaster, PdfOcr, the ocr.* codes, the text-layer verb; adds the manifest's ocr field and its schema. Proved by unit tests (a word at every rotation, an ALTO file in each unit, a TSV with a missing column, an hOCR with an entity expansion and a 10 MB title, supplementary-plane text, a page already tagged, a page whose content leaves the state altered); an FsCheck property — any set of words placed in a random frame (rotated, cropped, /UserUnit) comes back from M15's extraction as the same words at the same boxes within the tolerance —; IOcrEngine substituted with NSubstitute to assert the orchestration — one call per page needing one, cancellation between pages, progress reported, pages written in order —; integration: the text-layer rows below through pdftotext and PyMuPDF, OCRmyPDF's renderer over the same hOCR as a second implementation, veraPDF at level 2u after M21's conversion, and its PDF/UA-1 profile on the tagged variant; TextLayerBenchmarks. Leaves blank pages and orientation.
  11. Blank pages and orientation. Delivers PdfBlankPageDetector, PdfPagePredicates.BlankByPixels, PdfPageOrientation, the blank verb and text-layer --orient; adds expect.blankPages and expect.orientation to the manifest and its schema. Proved by unit tests (a white page with dark scanner edges, a page carrying one signature, a speckled page, a page whose /Decode inverts, a color page with show-through, a page of vertical Japanese, a landscape table on a portrait page); an FsCheck property — a mark larger than the speck anywhere inside the margins turns a blank verdict into not blank, and one smaller than the speck does not —; integration: ink ratios against an independent NumPy computation over poppler's decoded images; Tesseract's orientation detection on MuPDF's rendering of each turned page reading 0°. Leaves the whole.
  12. Budgets, campaign and the whole. Delivers ImagingBenchmarks across the codecs with MemoryDiagnoser (megabytes per second, bytes per row), the working-set checkpoints on the heavy scans, the decoding rows of the comparison benchmarks recorded, the fuzzing campaign's record, the documentation. Proved by the budget rows below, the remote rows on a green Remote corpus run, and the campaign's record in status.md.

Tests required​

Unit — tests/AdCodicem.Pdf.Tests, the Imaging satellite's tests under Imaging/ as M10 placed its satellite's:

  • CCITT: every code-table entry; /K negative, zero and positive; /EndOfLine, /EncodedByteAlign both ways, /EndOfBlock false, /BlackIs1, /DamagedRowsBeforeError; rows cut, runs past the row, an EOL missing.
  • JBIG2: every segment type; arithmetic and Huffman paths; refinement and aggregation; striped and unstriped pages; every combination operator; globals referred to by a page; a segment referring to one that does not exist.
  • JPEG: every process decoded and every one refused; every sampling factor; restart intervals; each Adobe and JFIF case of the color rule; scans that end early; the wipe on MCU boundaries and inside one.
  • JPEG 2000: every marker, progression order, code-block style and quantization style; tile-parts out of order; POC; PPM and PPT; a palette; cdef; /SMaskInData 0, 1 and 2; reduced resolution; HTJ2K refused.
  • Color and functions: each family's conversion against known values; each function type; every type 4 operator; stitching boundaries; lookups cut short; cycles.
  • Text layer: each reader against its specification's examples; each frame; each tagging case; replacement of our own layer twice over giving the same bytes; the glyphless font read back by M08's parser with widths consistent.
  • Blank pages and orientation: each case of slice 11.
  • Hostile: for every decoder, an image declaring 65,535 × 65,535 pixels over ten bytes of data; a JBIG2 symbol count summing past 2³²; a JPEG with 10⁵ scans; a JPEG 2000 codestream claiming 65,535 tiles and 32 decomposition levels; a type 4 program of a million operators; an Indexed space whose lookup is empty; an hOCR entity expansion; a TSV of a hundred million rows — each ends in rows, a report or a typed exception within its time and allocation budget. The guards are reached, raised and thrown as ReaderLimitsTests does for the reader's.
  • Fuzzing: every decoder and reader joins FuzzingTests per commit and the nightly campaign, seeded with the corpus's encoded image streams and the conformance streams; the coverage-guided engine of slice 1 runs each decoder nightly on a time budget; a finding becomes a regression test in HostileInputTests before it is fixed.
  • Determinism: decoded samples, conversions and text layers byte-identical across two runs, two cultures and the x64 and ARM64 runners.

Integration — tests/AdCodicem.Pdf.IntegrationTests, every referee in a container (ADR 27):

  • MuPDF — mutool extract and PyMuPDF's pixmaps, for decoded samples of every filter, its bundled jbig2dec for JBIG2 and OpenJPEG for JPEG 2000; mutool draw for renderings;
  • poppler — pdfimages -list, -png, -tiff, -all; pdftotext with -bbox-layout for the text layer;
  • Pillow over libjpeg-turbo, and libjpeg-turbo's cjpeg, djpeg and jpegtran, for JPEG; a coefficient reader over libjpeg for the wipe;
  • OpenJPEG — opj_decompress, and opj_compress to generate fixtures;
  • libtiff — G3 and G4 fixtures and decoding of our G4 output; jbig2dec standalone for our JBIG2 output;
  • PyMuPDF — words and search_for quads over the text layer;
  • OCRmyPDF — its hOCR renderer, a second implementation of the layer, over the same sidecars;
  • Tesseract — its orientation and script detection on renderings, the orientation referee;
  • veraPDF — PDF/A-2u on scans given a layer and converted by M21, PDF/UA-1 on the tagged variant, PDF/A on M21's JPEG 2000 remedy;
  • qpdf — --check on every document this milestone writes.

Where the referees disagree — MuPDF's bundled libjpeg and libjpeg-turbo on subsampled chroma, two CCITT decoders on a damaged row — a row passes when we agree with the reading the manifest records and the reason it records for the other. Disagreeing with every referee fails. The tolerances — JPEG against MuPDF, the 9/7 wavelet against OpenJPEG, CIE conversions against MuPDF, word boxes against pdftotext and PyMuPDF — are fixed by the slices that introduce them and recorded in status.md.

Acceptance conditions​

"Every committed image" means the 423 image XObjects of the 53 committed documents that open with images, and their inline images: 169 Flate, 96 uncompressed, 71 CCITT, 57 DCT (four behind Flate or ASCII85), 28 JBIG2, one JPX, one LZW. The remote rows close only on a green Remote corpus run, recorded in status.md with its date.

DocumentsBehaviorVerified by
The 71 committed CCITT images in nine documents — vendor/us-federal/acrobat3-import-irs-1040-1988-scan.pdf (G4 at 400 ppi), vendor/pikepdf/scanner-ccitt-endofline.pdf (G3 with /EndOfLine and a /Decode array), finereader8-frb-sr0115-examiner-guidance.pdf, vendor/uk-ogl/indesign-acrobat-hmcts-n208-form.pdf, pagemaker-distiller5-hmrc-iht205-form.pdf, and the image masks of vendor/opf-format-corpus/word9-distiller405-usgs-nwql-volatile-organics-methods.pdf, distiller952-pscript5-kb-pdf-risk-inventory.pdf, pdfmaker707-word-va-esig-developer-guide.pdf and pdfmaker8-word-va-lms-process-reference.pdf; remote, remote/ocrmypdf/tiff2pdf-35000px-ccitt-image.pdf (35,000² in 10.5 KB, 153 MB decoded), remote/pdfjs/itext5-ccitt-g4-mask-issue4379.pdf, the Kodak and Epson scans of remote/ocrmypdf/, the Ricoh scan remote/pdfjs/ricoh-3heights-scan-issue5747.pdf, Konica's CCITT masks, and the JHOVE tiff2pdf, Apex, Pixel Translations and Paper Capture scansSamples identical to MuPDF's and poppler's decoding, compared as hashes of rows; the 35,000² image decoded holding two rows, within its time budgetCorpusImageDecodingTests.Ccitt_images_decode_to_the_referees_pixels (new)
The 28 committed JBIG2 images — vendor/us-federal/xerox-workcentre-treasury-imf-report-scan.pdf (symbol dictionaries in /JBIG2Globals, text regions, striped pages), xerox-workcentre-5755-ocr-hud-fonsi-mrc.pdf (symbols, image masks over a JPEG background), xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdf and docusign-pdfkit-gsa-sf30-contract-modification.pdf (generic regions); remote, remote/pdf-association/abledocs-pdfua1-tagged-textbook-scan.pdf, remote/opf-format-corpus/jhove-hul-136-acrobat7-paper-capture-scanned-report.pdf (314 pages), jhove-hul-117-acrobat101-student-design-report.pdf, jhove-hul-35-atypon-pdfplus-journal-article.pdfSamples identical to jbig2dec's (MuPDF) and poppler's; the segment types each file uses recorded, and each type the corpus lacks covered by the conformance streams not in the corpus (below)CorpusImageDecodingTests.Jbig2_images_decode_to_the_referees_pixels (new)
The 57 committed JPEG images — 45 baseline in three components, eight progressive (the Mustang, GnuAccounting and factur-x Python invoices, DILA's signed notice, the LibreOffice Cerfa, OpenOffice's PNG page), flate-over-dct and ASCII85 chains, Acrobat 11's image whose /Height was altered —; remote, CMYK and YCCK in remote/ocrmypdf/photoshop-cc2015-pdfx3-cmyk.pdf (/ColorTransform over CMYK) and remote/pdf-association/abledocs-pdfua1-textbook-chapter.pdf, the damaged DCT inside Flate of remote/pdfjs/canon-scan-junk-after-eof-scan-bad.pdfSamples byte-identical to Pillow's decoding over libjpeg-turbo; MuPDF's and poppler's differences within the tolerance recorded; the altered image and the damaged scan decoded as far as their data goes and reportedCorpusImageDecodingTests.Jpeg_images_decode_as_libjpeg_turbo_does (new)
JPEG 2000: vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf; remote, remote/pdfjs/acrobat8-jpx-precincts-issue5475.pdf (precincts, no /ColorSpace), the soft-masked JPEG 2000 of the AbleDocs chapter, remote/pdf-association/indesign-cs6-pdfua1-brochure.pdf, the portfolio's embedded JPEG 2000 in remote/opf-format-corpus/acrobat9-portfolio-signed-3d.pdf, the 125 plates of remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf, and remote/pdfbox/distiller6-zeroed-object-stream-pdfbox3947.pdfSamples within the 9/7 tolerance of OpenJPEG's (through MuPDF and opj_decompress), and identical on the reversible path; /SMaskInData opacity equal to MuPDF's mask; each plate decoded holding one row of tiles, within the working-set guardCorpusImageDecodingTests.Jpeg_2000_images_decode_as_openjpeg_does (new)
Color: the Indexed images over CalRGB (word9-distiller405-usgs-nwql-volatile-organics-methods.pdf), DeviceCMYK (vendor/us-federal/designer-distiller23-uscis-i9-javascript-form.pdf) and ICCBased (pdfmaker7-powerpoint-va-cancer-database-course.pdf), the 48 ICCBased images of 16 documents, vendor/pdf-association/handwritten-inline-image-abbreviations.pdf (a named CalRGB, /Decode arrays), handwritten-indexed-color-out-of-range.pdf; remote, DeviceN under a type 4 tint transform in the Acrobat 9 portfolio; Lab and Separation images and /Matte soft masks — not in the corpus (below)PNG and TIFF exports equal pdfimages -png and -tiff for device and indexed spaces; RGB within the tolerance of MuPDF's for CIE-based and tint-transformed spaces; the out-of-range lookup reported, as the referees paint itCorpusImageExportTests.Decoded_images_export_as_the_referees_do (new)
Masks: the 105 soft masks of ten committed documents (Word's text drawn as images, the iBooks Author pages, OpenOffice's PNG, the OZEV invoice, the DH factsheet), the CCITT and JBIG2 image masks above, the stencil /Mask stream of the remote iText G4 file; a color-key mask and a /Matte soft mask — not in the corpus (below)Alpha and stencil samples identical to MuPDF's masks, each at its own sizeCorpusImageExportTests.Masks_decode_with_their_images (new)
Every committed image, and every remote oneIdentical samples in two runs, under two cultures, and on the x64 and ARM64 runnersCorpusImageDecodingTests.Decoding_is_deterministic_everywhere (new)
M19's scans: xerox-workcentre-treasury-imf-report-scan.pdf (JBIG2), acrobat3-import-irs-1040-1988-scan.pdf (CCITT G4), documents/scan/reportlab-scanned-receipt.pdf (DCT), imagemagick-false-pdfa1b-jpx.pdf (JPX), and the OCR'd finereader8-frb-sr0115-examiner-guidance.pdf and xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdfRedacted with PdfImagingCodecs.Default: every sample under a mark uniform, every other exact as the redaction table says — coefficients outside the marked MCUs for DCT, by the libjpeg reader —; the OCR glyphs under the marks gone; the words under the marks absent from Tesseract's reading of MuPDF's rendering; the unsupported marker M19 recorded removedCorpusImageRedactionTests.Scanned_pixels_under_a_mark_are_removed_and_no_others (new)
vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf, and every JPEG 2000 document M21 recorded as not convertible because an image breaks part 2's clauseConverted by M21 to PDF/A-2b through M22's decoder: an image outside the clause re-encoded as Flate, its samples identical to its decoding, one inside it left as it was; veraPDF accepts; the report lists each change of encoding; M21's unsupported marker removedCorpusPdfAConversionTests.Jpeg_2000_images_outside_the_clause_are_re_encoded (new)
Every document whose image data does not decode — the damaged scans above, remote remote/pdf-differences/unidentified-wine-merchant-sheet-png-under-dctdecode.pdf (a PNG under DCTDecode), and every hand-written hostile fileNo untyped exception, no hang, each within its time and allocation budget; rows delivered as far as the data goes; M02's stream rule reports each in expect.findingsCorpusImageDecodingTests.Damaged_image_data_decodes_as_far_as_it_goes (new)
The committed image-only scans — the receipt, the Xerox IMF scan, the IRS 1040 import, the G3 scanner page, ImageMagick's JPX page — given hOCR, ALTO and TSV from a pinned Tesseract — not in the corpus (below)Every recognized word found by pdftotext and PyMuPDF with its text, at its box within the tolerance after the frame conversion; PyMuPDF's search_for finds each line's text; OCRmyPDF's rendering of the same hOCR agrees on the words and their orderCorpusTextLayerTests.Recognized_words_are_found_where_the_engine_saw_them (new)
The OCR'd committed scans — vendor/us-federal/hp-mfp-acrobat-ocr-nih-report.pdf (Acrobat), the Xerox 5335 and 5755 copier layers (render mode 3, Tz up to 2000 %), finereader8-frb-sr0115-examiner-guidance.pdf (text under the image), the scanned page of vendor/fr-licence-ouverte/pdfmaker-acrobat-cerfa-12156-form.pdf; remote, the Konica, Ricoh, Canon and OmniPage scansM15's hOCR and ALTO exports, re-applied with ReplaceInvisibleText, give back the same words, as pdftotext reads them, at the same boxes within the tolerance; applied twice with ReplaceOwnLayer, the same bytesCorpusTextLayerTests.Exported_layers_reapply_to_the_same_words (new)
The image-only scans above with their layer, converted by M21veraPDF accepts PDF/A-2u; with Paragraphs, veraPDF's PDF/UA-1 profile reports no failure and M20's verdict is ConformsPendingReview; pdfinfo -struct-text gives each block's text in orderCorpusTextLayerTests.Scans_with_a_text_layer_reach_pdf_a_2u (new)
Blank pages: remote/ecan/konica-bizhub-c554e-letter-scan.pdf, whose scanned blank page the manifest is to record; the committed image-only scans, none blank; a copier batch with blank backs and separator sheets, and near-blank pages — a page number alone, a signature alone, punched holes, show-through — not in the corpus (below)Verdicts equal expect.blankPages; no page with visible text or a signature called blank; ink ratios equal NumPy's over poppler's decoded images within 0.01 percentage points; M07's split by BlankByPixels gives the parts the manifest recordsCorpusBlankPageTests.Blank_scanned_pages_are_found_and_nothing_else (new)
Orientation: docusign-pdfkit-gsa-sf30-contract-modification.pdf (/Rotate 270, upright as displayed), the Xerox 5755's deskewed placement, remote remote/ocrmypdf/epson-scan-indirect-rotate.pdf (an indirect /Rotate) and remote/maine-legislature/ricoh-docusign-itextsharp-state-contract-amendment.pdf; scans turned by 90, 180 and 270° without /Rotate, with a layer — not in the corpus (below)The upright ones unchanged; the turned ones given the /Rotate after which Tesseract's orientation detection on MuPDF's rendering reads 0°; expect.orientation met; nothing but /Rotate changed, object by objectCorpusOrientationTests.Pages_are_turned_upright_from_their_text_layer (new)
Every decoder and reader, seeded with every committed image stream, the conformance streams, and every damaged/* documentA nightly campaign, mutation and coverage-guided, finds no untyped exception, hang or unbounded allocation over fourteen consecutive nights before the milestone closes; its executions, coverage and findings, each fixed with a regression test, recorded in status.mdFuzzingTests.Decoding_mutated_image_data_either_works_or_reports (new), the campaign's record
remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf (147 MB of JPEG 2000), the Paper Capture scan (314 pages of CCITT and JBIG2), the 35,000² CCITT imageDecoding every image of every page holds its working set within the budgets recorded in status.md, flat across pages; 0 B allocated per row once warm, the rows exceptedCorpusImageDecodingTests.Heavy_scans_decode_within_their_budget (new), ImagingBenchmarks (new)
The same operations through the toolThe AOT binary's text-layer, blank and inventory --images --decode produce what the API producesCorpusToolTests.Text_layer_blank_and_decoded_images_match_the_api (new)

Corpus​

What the corpus holds​

  • CCITT: 71 committed images in nine documents, G4 and one G3 with /EndOfLine (ccitt, ccitt-g4, ccitt-g3, ccitt-endofline, ccitt-image-mask); remote, 25 documents between them, the 35,000² tiff2pdf image (huge-decoded-image, high-compression-ratio), and the JHOVE scans from tiff2pdf, Apex, Pixel Translations and Acrobat Paper Capture.
  • JBIG2: 28 committed images in four documents — two with /JBIG2Globals, symbol dictionaries and arithmetic text regions, two with generic regions only (jbig2, jbig2-globals, jbig2-without-globals, jbig2-image-masks, jbig2-bitonal-full-page-image); remote, the AbleDocs scan, the 314-page Paper Capture report, two JHOVE papers.
  • JPEG: 57 committed images — baseline, progressive, behind Flate and ASCII85 (dct-image, jpeg-image, flate-over-dct); remote CMYK and YCCK (dct-cmyk-images, dct-colortransform-on-cmyk, cmyk-jpeg), a corrupt DCT (dct-corrupt, dct-in-flate), a PNG under DCTDecode (png-labeled-dctdecode).
  • JPEG 2000: one committed (a JP2 in RPCL with tiles, layers, SOP and EPH); remote, eight documents (jpx, jp2-container, jpx-no-colorspace, jpx-multiple-precincts, jpx-soft-mask, large-jpx-plate, embedded-jpx).
  • Color and masks: Indexed over DeviceRGB, DeviceCMYK, CalRGB and ICCBased, ICCBased with /N 3, 105 soft masks, CCITT and JBIG2 image masks, /Decode arrays, an Indexed lookup out of range (indexed-color-out-of-range); remote, a stencil /Mask stream (explicit-mask), DeviceN with a type 4 function (devicen-nchannel, type4-function), a pattern color space.
  • OCR layers: from Acrobat, two Xerox copiers, FineReader 8, a Cerfa page; remote, Konica, Ricoh, Canon, OmniPage, AbleDocs (ocr-layer, copier-native-ocr, invisible-text-render-mode-3, text-under-image, ocr-text-painted-under-page-image).
  • Scans without text: five committed image-only documents (image-only, no-text); remote, the Kodak, Epson and tiff2pdf scans.
  • Pages placed or turned: rotate-270, indirect-rotate, rotated-deskewed-image-matrix, rotated-page.
  • Blank pages: one remote (blank-page, the Konica scan).
  • Damage: image-dimensions-wrong, dct-corrupt, object-stream-destroyed (a JPEG 2000 document), the hand-written hostile files.

What it lacks​

NeedWhyPriorityLikely source
hOCR, ALTO and TSV of the committed image-only scans from a pinned Tesseract (French and English models), with a manifest field naming them and the engine's versionThe text-layer and PDF/A-2u rows need recognized words, and documents are committed, not generated at test time: Tesseract's output changes with its version1Generated here: Tesseract from the distribution's packages in a container, its version and models recorded in build_corpus.py, the output reviewed and committed under tests/corpus/sources/ocr/
JBIG2 streams with the segment types no committed document uses: refinement, pattern dictionaries and halftone regions, Huffman-coded symbol dictionaries and text regions, custom tables, MMR generic regions, intermediate regionsThe FORCEDENTRY class of fault lives in these paths, and a decoder proven on two segment types is not proven; the fuzzing campaign needs them as seeds1A public source (W20): the ITU-T T.88 test bitstream and the conformance streams in jbig2dec's and pdf.js's test suites, their terms read first, wrapped one per page by a recorded pikepdf transformation; pdf.js's JBIG2 bug-report PDFs to the remote corpus
JPEG 2000 beyond the one committed file: reversible 5/3, raw codestreams, subsampled components, palettes, cdef opacity with /SMaskInData 1 and 2, 16-bit, sYCC, four components, every code-block style and progression order, POC, PPM and PPT, region of interestOne committed JPX cannot prove a decoder, and the remote ones are tested only nightly1Generated here with OpenJPEG's opj_compress over our own images, wrapped byte for byte by img2pdf (pypi); a public source: the ISO/IEC 15444-4 conformance codestreams in OpenJPEG's data repository, terms read first
Scans turned by 90, 180 and 270° without /Rotate, with and without a layer, and one with a landscape table on portrait pagesThe orientation rows have nothing to turn: every committed rotated page is already upright as displayed1Derived here: a recorded pikepdf transformation turning committed scans' content, Tesseract's sidecars made on the turned rendering; a contribution (W03): copier output fed the wrong way
A copier batch with blank backs and scanned separator sheets, and near-blank pages — a page number alone, a signature alone, punched holes, show-through, speckleOne blank page cannot show the detector's two errors, and M07's separator split waits for a real batch1A contribution (W03), as M07 already asked; derived here meanwhile: blank sheets scanned from our own paper, and near-blank pages cut from committed scans, each recorded
expect fields: decoded-sample hashes on M15's expect.images rows where the referees agree, blankPages, orientation, and the ocr sidecar field beside expectExpectations must come from the file and independent tools; the unit suite asserts pixels without a container only if the hashes are in the manifest1Generated here: build_corpus.py records the hashes from MuPDF and poppler and the reason where they disagree; the pages from a person's review, recorded with it
JPEG in CMYK and YCCK committed, restart intervals, 4:2:2, 4:1:1 and 4:4:0 sampling, arithmetic coding, a progressive image cut shortThe committed JPEGs are all YCbCr at common sampling; CMYK and damage are remote only; arithmetic coding is absent2Generated here with libjpeg-turbo's cjpeg and jpegtran (-restart, -sample, -arithmetic), wrapped by img2pdf
CCITT with /BlackIs1 true, /EncodedByteAlign true, two-dimensional G3 (/K > 0), /EndOfBlock false, and damaged rows under /DamagedRowsBeforeErrorThe committed CCITT is G4 but for one G3 page; every other parameter has unit tests only2Generated here with libtiff's tiffcp and fax2tiff, wrapped by img2pdf or a recorded pikepdf construction; a public source: fax-server output in public records
Images in Lab, Separation and DeviceN under type 0, 2 and 3 tint transforms, CalGray, 2- and 4-bit samples, a color-key /Mask array, /Matte soft masksThe color and mask rows name spaces and masks no committed image uses2Generated here: a recorded pikepdf construction over committed images' samples; a public source: the Ghent Workgroup's output suites, terms read first
An ALTO file from a commercial engine with its page imageThe ALTO reader is proven on Tesseract's dialect and M15's own export only2A public source (W20): the Library of Congress's Chronicling America, public-domain page images with ABBYY's ALTO — to the remote corpus when over 2 MB
Multi-strip TIFFs, bilevel and gray, and a TIFF with FillOrder 2, to joinM07's strip joining is proven on synthetic files only3Generated here with libtiff's tiffcp -r

Traps​

  • FORCEDENTRY was a count. A JBIG2 text region's symbol count summed in 32 bits and trusted after it overflowed. Every count a segment declares is 64-bit and checked against the working set before anything is sized by it.
  • The arithmetic decoders never run dry: T.88 and 15444-1 feed 0xFF past the end of the data. A loop that waits for the input to end does not end.
  • Pattern-matching JBIG2 changes what a scan says, and the corpus's Xerox files are made of it. Decode it as it is; never write it.
  • A JPEG decoder that is "close" is wrong when the reference is exact: the integer inverse DCT and the upsampling rounding decide every byte. Different references disagree among themselves; say which one is the reference, and measure the others.
  • Chroma upsampling reads the neighboring blocks: pixels bordering a wiped MCU change even though their coefficients do not. The redaction guarantee is on coefficients, and one pixel of border on samples.
  • /ColorTransform and the Adobe marker can disagree, Adobe's CMYK is inverted, and a JPEG's EXIF orientation is ignored inside a PDF. The Photoshop PDF/X-3 file carries the first case.
  • CCITT's polarity is stated twice: /BlackIs1 and the image's /Decode. The scanner page sets /EndOfLine and a /Decode array; inverting twice is not inverting.
  • /Rows and /Height disagree, and /EncodedByteAlign pads differently for /K < 0 and /K ≥ 0 — producers get both wrong.
  • JPEG 2000 color comes from three places — /ColorSpace, the JP2 colr box, the component count — and the dictionary wins when present; opacity may be premultiplied; subsampled components must be upsampled before any color transform.
  • Floating point is deterministic only if written to be: no implicit fused multiply-add, a fixed order of operations, no Math.Pow whose last bit the platform's C library chooses.
  • A progressive JPEG holds the whole image as coefficients until its last scan; the working set is twice the samples, not a row.
  • An image mask paints with the current color, a soft mask is an image of its own size, and /Matte premultiplies — extraction and compositing must not confuse the three.
  • Extractors infer spaces differently. A text layer without explicit spaces reads as one word in one tool and as words in another; an OCR layer with Tz at 2000 % (the Xerox copiers') still gives each word the box the engine saw.
  • hOCR is y-down and in the image's pixels; PDF is y-up and in points, through a placement that may rotate, skew or crop the image, and a page that may be rotated, cropped and scaled by /UserUnit.
  • An XML file from an engine is hostile input too: an entity expansion in an hOCR file is a denial of service like any other.
  • Invisible text is what M19's sanitizer removes as hidden.invisible-text and hidden.ocr-layer: our layer is marked so that it can be named, kept or removed on purpose.
  • A blank sheet is not white: scanners leave dark edges, punched holes, show-through from the back and JPEG noise; and a page bearing only a signature or a page number is not blank.
  • Orientation from text fails without text, and a page may hold text in two directions — a landscape table on a portrait page. Undecided is an answer.
  • An engine's output depends on its version, models and threads. The layer is deterministic for given words; the words are the engine's, and the report says which engine gave them.

Documentation​

  • docs/website/docs/concepts/images.md (new): the image seam, rows and working sets, masks, color evaluation and where it approximates, the codec set and the Imaging satellite, the contract of a caller's own codec.
  • docs/website/docs/guides/scans.md (new): decoding and exporting scans, blank pages, orientation, and joining TIFF strips.
  • docs/website/docs/guides/ocr-text-layer.md (new): recognized text from hOCR, ALTO and TSV, frames, tagging, replacing a layer, IOcrEngine with the Tesseract adapter sample, reaching PDF/A-2u through M21, and what the library does not do (run an engine, deskew).
  • docs/website/docs/reference/reader-limits.md and diagnostics.md: MaxImageWorkingSet, MaxJpegScans, and the image.*, function.*, ocr.* and orientation.* codes.
  • The redaction guide M19 wrote: pixels of every codec, per the redaction table.
  • docs/website/docs/reference/tool/: text-layer, blank, inventory --images --decode.
  • docs/website/docs/introduction.md and docs/features/features.json: the imaging entry brought to its state.
  • docs/architecture.md: the core's Images/, Graphics/ and Ocr/, the Imaging satellite's contents and dependencies, the fuzzing harness. SECURITY.md: decoders as attack surface, and the fuzzing that covers them.
  • docs/corpus.md, the manifest schema and tests/corpus/README.md: the ocr field, the decoded-sample hashes, blankPages, orientation, the new readerLimits keys; docs/corpus-sources.md: the sidecars' engine and version, the conformance streams' terms.
  • The ADRs of slices 1 and 5: the fuzzing engine, and pixels under a redaction.
  • docs/status.md: the tolerances, the working-set measurements, the campaign's record.

Exit criteria​

  • CCITT decodes in the core; JBIG2, JPEG and JPEG 2000 in the Imaging satellite, with every process, segment type and capability the design lists, and each one it refuses refused with a diagnostic.
  • Color spaces and functions convert samples identically on x64 and ARM64, and approximations are reported.
  • The G4 and JBIG2 generic encoders are lossless, verified by independent decoders, and no code path writes a JBIG2 symbol.
  • M15's decoded export, M19's pixel redaction, M21's JPEG 2000 remedy, M07's strip joining and M02's stream rule over image filters work through the codec set, and the unsupported markers that waited for M22 are gone.
  • The text layer is written from hOCR, ALTO and TSV, tagged on request, replaced idempotently, and a scan with a layer reaches PDF/A-2u through M21, confirmed by veraPDF.
  • Blank pages and orientation are detected as designed, their thresholds fixed and recorded.
  • MaxImageWorkingSet and MaxJpegScans are guards with their codes, tests, documentation and schema keys; every other bound is classified where it is declared.
  • The ADRs on the fuzzing engine and on pixels under a redaction are accepted.
  • The priority-1 gaps above are filled; each remaining gap is recorded in docs/corpus-contributions.md.
  • The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on a green Remote corpus run recorded in status.md.
  • Unit tests cover each behavior, its degenerate cases and its hostile ones; the FsCheck properties hold; every decoder and reader has run fourteen consecutive nights in the fuzzing campaign with no open finding.
  • Integration tests run MuPDF, poppler, Pillow and libjpeg-turbo, OpenJPEG, libtiff, jbig2dec, PyMuPDF, OCRmyPDF, Tesseract, veraPDF and qpdf, each in a container.
  • CcittBenchmarks, JpegBenchmarks, JpxBenchmarks, TextLayerBenchmarks and ImagingBenchmarks run with MemoryDiagnoser; status.md records the working sets and the comparison rows.
  • The tool's verbs ship in the dotnet tool and the AOT binaries, documented.
  • The documentation site publishes the pages listed above.
  • Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).