M22 — Imaging and OCR
State: to do — Depends on: M08, M15 — Codecs, lossy steps and the OCR boundary per ADR 42; every decoder bound classified per ADR 34; unsafe code only where a measurement asks for it, per ADR 35; a dependency-free satellite per ADR 9; active content never produced, per ADR 37
Goal
Decode the pixels of every image a PDF can carry — CCITT in the core; JBIG2, JPEG and JPEG 2000 in the
AdCodicem.Pdf.Imaging satellite; their color spaces, masks and functions evaluated —, re-encode bilevel scans
losslessly, and give a scan the text it does not have: an invisible, positioned text layer written from the words
an OCR engine recognized, so that it can be searched, extracted, redacted and archived as PDF/A-2u. Tell a blank
scanned sheet from a page that carries only a signature, and a sideways page from an upright one.
A case file is mostly scans. The exhibits a lawyer receives come off copiers and fax servers as JBIG2, CCITT, JPEG and JPEG 2000; M15 can say that such a page has no text, but nothing after it can act on the page until its pixels can be read: M19 cannot remove a name from a scanned letter, M21 cannot convert a JPEG 2000 image that breaks PDF/A-2's clause, nor bring a scan to level u, M23 cannot downsample, M25 cannot rasterize, M07 cannot recognize a scanned separator sheet, and M18's portal presets cannot make a piece searchable. ADR 42 put the decoders in one fuzzed, bounded codec set so that all of them share it.
The failures this milestone exists to prevent are specific. A decoder that can be driven outside its buffers — the FORCEDENTRY exploit of 2021 went through a JBIG2 decoder's symbol count. A decoder that is almost right — a JPEG off by one where a reference implementation is exact, a CCITT row shifted by one run, a JPEG 2000 plate whose tiles meet with a seam. A re-encoding that changes what a scan says — the pattern-matching JBIG2 of the copiers of 2013, which replaced digits. A text layer whose words land on the wrong line, split in the middle, or without the spaces extractors need. A blank-page detector that drops the page bearing only a signature. A color conversion that gives a different byte on Windows and on Linux.
Scope
In:
- CCITT decoding in the core: T.4 one- and two-dimensional and T.6, every parameter of ISO 32000-2 §7.4.6, damaged rows recovered and reported;
- the image seam in the core: a
PdfImageread model over image XObjects and inline images;PdfImageCodecs, which M19 defined, completed with decoders and encoders by filter;IPdfImageDecoder,IPdfImageEncoderand a row sink;PdfImageReader, which runs the filters, the codec,/Decode, color-key masks, stencil and soft masks and/Matte, and delivers rows; - color spaces and functions evaluated in the core: CalGray, CalRGB, Lab, ICCBased, Indexed, Separation and DeviceN, and functions of types 0, 2, 3 and 4, converting samples to gray, RGB or CMYK for export and analysis, deterministically on every platform;
- the
AdCodicem.Pdf.Imagingsatellite (ADR 42): the JBIG2 decoder (every region type of ISO/IEC 14492 in PDF's embedded organization, with/JBIG2Globals); the JPEG decoder (baseline, extended and progressive Huffman, arithmetic-coded sequential and progressive; gray, YCbCr, RGB, CMYK and YCCK; restart intervals); the JPEG 2000 decoder (Part 1 codestreams, raw or in JP2 and JPX files, with every progression order, tiles, precincts and code-block style, and/SMaskInData); the lossless CCITT G4 and JBIG2 generic-region encoders; and the JPEG block wipe that M19 left to this milestone; - the consumers completed: M15's image export gains decoded output with color evaluated; M19 redacts the pixels of CCITT, JBIG2, JPEG and JPEG 2000 images; M21 re-encodes a JPEG 2000 image outside part 2's clause; M07 joins a multi-strip TIFF into one image and transcodes an arithmetic-coded JPEG file, on request; M02's stream rule on decodable filters judges image filters;
- the text layer: a
PdfRecognizedPagemodel; readers for hOCR 1.2, ALTO v2 to v4 and Tesseract's TSV; a glyphless font;PdfTextLayer, which writes invisible (render mode 3), positioned words, optionally tagged;IOcrEngineandPdfOcr, which pick the pages without text, decode their images, call the caller's engine and write the layer; theIPdfPageRasterizerseam M25 fills; - blank-page detection on pixels, beside M07's detection by content, as a predicate M07's split takes;
- page orientation from the text layer or from the engine, applied as
/Rotate; - the command-line tool's
text-layerandblankverbs, and decoded output forinventory --images; - a fuzzing harness for every decoder and for the recognized-text readers, run per commit and nightly from the first slice.
Out, explicitly:
- running an OCR engine — the caller's, behind
IOcrEngine(ADR 42); a first-party engine satellite stays an open question of the roadmap. A sample shows an adapter; no package ships one; - rasterizing a page that is not one image — vector content, several images, text over an image — so that an
engine can read it: M25 fills
IPdfPageRasterizer; until then such a page is reported and skipped; - every lossy step — downsampling, JPEG re-encoding, conversion to gray or to one bit per pixel — M23, where each is opt-in and reported (ADR 42); M22's only encoders are lossless;
- pattern-matching or lossy JBIG2 — never (ADR 42); JPEG 2000 encoding — not planned; HTJ2K (ISO/IEC 15444-15) — not planned, refused with a diagnostic;
- pixel deskew, despeckle and binarization, and orientation found from pixels alone — an open question of the roadmap; mixed raster content compression — an open question;
- color management — ICC transforms, conversion to an output intent — M29's satellite; M22's conversions serve export and analysis and say where they approximate;
- reading barcodes from pixels — no milestone plans it: barcode and patch-code recognition is an open question of the roadmap; M07's separator predicate can call a caller's decoder over M22's rows;
- tagging beyond one paragraph per recognized block — headings, lists, tables inferred from a scan — the roadmap's open question on heuristic tagging;
- lossless JPEG (SOF3), hierarchical processes and 12-bit precision — refused with a diagnostic, reopened by a document that needs them; JBIG2's color extension and JPEG 2000 features outside §7.4.9 — refused;
- rendering — M25, which composes these decoders with M15's interpreter.
Design
Where it lives
| Part | Where | Why |
|---|---|---|
| CCITT G3 and G4 decoder | Core, IO/Filters/, internal | ADR 42: small, needed by M07's pages and by fax-era scans |
PdfImage, PdfImageCodecs, IPdfImageDecoder, IPdfImageEncoder, IPdfImageRowSink, PdfImageReader, PdfBlankPageDetector | Core, Images/, namespace AdCodicem.Pdf.Images | The seam M19 named; M07, M15, M19, M21, M23 and M25 consume it; blank detection needs only rows |
PdfColorSpace evaluation, PdfFunction | Core, Graphics/, namespace AdCodicem.Pdf.Graphics | Evaluation needs no codec: M15 reads the spaces, M19 writes an overlay's color into samples, M25 evaluates shadings with the same functions |
| JBIG2, JPEG and JPEG 2000 decoders; G4 and JBIG2 generic encoders; the JPEG block wipe | AdCodicem.Pdf.Imaging | ADR 42: the large decoders, and their attack surface, in a package a caller adds knowingly |
PdfRecognizedPage, the hOCR, ALTO and TSV readers, the glyphless font, PdfTextLayer, IOcrEngine, PdfOcr, PdfPageOrientation | Core, Ocr/, namespace AdCodicem.Pdf.Ocr | ADR 42: the layer is written in the core, engines live outside it |
IPdfPageRasterizer, PdfRaster | Core, Content/, beside M15's device seam | One seam for every raster of a page — OCR here, M24's visual comparison, M25's implementation —, in the core so that neither Compare nor Rendering depends on the other |
| The fuzzing harness | tests/AdCodicem.Pdf.Fuzzing | A test project, never packaged |
text-layer, blank, inventory --images --decode | AdCodicem.Pdf.Tool | The tool ships the Imaging satellite |
The satellite depends on the core alone: no Skia, no native library, IsAotCompatible from its first commit. It is
a new package, so #42 (an API baseline per package, which M10 had to close) applies to it. Its public surface is
small: PdfImagingCodecs.Default — a PdfImageCodecs value holding its decoders and encoders — and the options
records; the codecs themselves are internal. The MQ arithmetic decoder, which JBIG2 and JPEG 2000 share, is written
once there.
The image seam
PdfImage read model over an image XObject or an inline image: width, height, color space, bits per
component, filters and their parameters, /Decode, /ImageMask, /Mask (a stencil stream or a
color-key array), /SMask with /Matte, /SMaskInData, /Interpolate, /Intent. Placements are M15's
PdfImageCodecs immutable (M19): decoders and encoders by filter name. PdfImageCodecs.Core holds the core's
filters and CCITT; PdfImagingCodecs.Default adds the satellite's. Passed in an operation's
options, never registered statically (invariant 8)
IPdfImageDecoder Decode(PdfImageDecodeContext, ReadOnlySpan<byte> encoded, IPdfImageRowSink sink). Public, so that
a caller may bring a codec of its own, native included, under the documented contract below
IPdfImageEncoder Encode(PdfImageLayout, rows, IBufferWriter<byte>) -> filter name and parameters
IPdfImageRowSink BeginImage(PdfImageLayout); WriteRow(y, samples); EndImage(PdfImageDecodeResult): rows in order,
top down, each written once
PdfImageReader Open(image, codecs, options) -> rows with /Decode applied, bit depth expanded as asked (1, 8 or
16 bits), optionally converted to gray, RGB or CMYK; masks delivered as images of their own, at
their own size, or resampled onto the image's grid when the caller composites
PdfImageRows a pooled sink for a caller that wants the whole image, bounded like any decoded stream
PdfImageDecodeResult rows delivered; rows missing, filled as the referees fill them and reported; approximations
- Rows, not bitmaps. Every decoder writes rows in order into a sink, and holds only what its format obliges it to: CCITT two reference rows; a baseline JPEG one row of MCUs; a JPEG 2000 image one row of tiles; a JBIG2 page one stripe, or the page when it is not striped; a progressive JPEG the coefficients of the whole image, since its last scan may refine its first block. That is each codec's working set, and it is what the guard below bounds. The shape is the one M23's piecewise decoding (#48) generalizes to every filter.
- The encoded data reaches the codec through M01's pipeline — a
FlateDecodeor anASCII85Decodebefore aDCTDecode—, bounded as every filter is (ADR 34). The image filter must be last (§7.4); one found elsewhere is reported and the image yields no rows. PdfStream.Decode()keeps stopping at an image filter, as it has since M01: its callers want data, not pixels, and M15'spdfimages -allpath wants the encoded bytes. Pixels arePdfImageReader's.- Masks are images. A stencil mask (
/ImageMask, or a/Maskstream) is one bit deep with its own/Decode; a color-key/Maskis a set of ranges compared with the samples before/Decode; an/SMaskis a gray image of any size, un-premultiplied by/Mattewhen present;/SMaskInDatamakes the JPEG 2000 decoder deliver its opacity channel as the mask. - The contract of
IPdfImageDecoderis invariant 4's: the encoded bytes are hostile, nothing is allocated from a value read without a checked bound, every loop ends by what it produces, and failure is aPdfImageDecodeResult, not an exception. A caller's own codec is held to it by documentation; ours by the fuzzing campaign.
CCITT in the core
- T.4 Modified Huffman (
/K 0), T.4 Modified READ (/K> 0, one-dimensional rows everyK) and T.6 (/K< 0);/EndOfLine,/EncodedByteAlign,/EndOfBlock,/BlackIs1,/Columns(1728 by default),/Rows,/DamagedRowsBeforeError. - Table-driven over a bit reader on a span: the changing elements of two rows in pooled arrays of
Columns + 2entries, the row expanded into the sink's buffer, 0 B allocated per row. - A damaged row — a code no table knows, a run past the row's end, a row that ends early — is replaced by the
previous row (white for the first) and counted; decoding resynchronizes at the next EOL when the stream has them,
and stops where it has none. After
/DamagedRowsBeforeErrorconsecutive damaged rows the image ends there. Missing rows are white. Oneimage.rows-damagedorimage.data-truncateddiagnostic carries the counts; never an exception. /Rowsabsent or disagreeing with the image's/Height: the height wins, as every viewer does, and the disagreement is reported./EncodedByteAlignpads differently for/K< 0 and/K≥ 0, and producers get it wrong both ways: the declared reading is tried first, the other only when the first fails in its first rows, and that is reported.
JBIG2
- The embedded organization of ISO/IEC 14492 as §7.4.7 uses it: no file header, the segments of
/JBIG2Globalsread before the page's own, one page per image. Segment types: symbol dictionary; text region, immediate and intermediate; pattern dictionary; halftone region; generic region, arithmetic with typical prediction (TPGDON) and MMR; generic refinement region; page information; end of stripe; end of page; tables; extension segments ignored when T.88 marks them unnecessary, refused when it marks them necessary. - Arithmetic and Huffman decoding of symbol dictionaries and text regions with refinement and aggregation; the standard tables B.1 to B.15 and custom tables; the combination operators; the page's default pixel and default operator; intermediate regions composed onto the page.
- Striped pages: a page information segment of height 0xFFFFFFFF grows by stripes, and each end-of-stripe segment releases the rows above it to the sink, so a striped page holds one stripe.
- Counts are 64-bit, and checked before they size anything. The number of symbols a text region may address is the sum of the exports of every dictionary it refers to plus its own; FORCEDENTRY (CVE-2021-30860) overflowed exactly that sum in 32 bits, and a later check trusted the overflowed value. Every count — symbols, instances, referred-to segments, table lines, pattern counts — is computed in 64 bits and checked against the working-set guard before an array exists.
- The arithmetic decoder reads 0xFF past the end of its data, as T.88 specifies. Every loop therefore ends by what it produces — a region's rows, a dictionary's declared symbols — and never by what it consumes.
- The Xerox copier scans in the corpus draw their text through symbol dictionaries and text regions: pattern matching, the mechanism whose substitutions made copiers of 2013 change digits. The decoder reproduces what the file says; the encoder below never writes a symbol.
JPEG
- Processes: baseline (SOF0), extended sequential Huffman (SOF1) and progressive (SOF2) at 8 bits;
arithmetic-coded sequential and progressive (SOF9, SOF10), which libjpeg-turbo writes and M07 asked to transcode.
Lossless (SOF3), hierarchical (SOF5 to SOF7, SOF13 to SOF15) and 12-bit precision are refused with
image.process-unsupported. - The reference is libjpeg-turbo with its defaults, which Pillow uses and most readers link: the accurate integer inverse DCT, triangular ("fancy") upsampling of subsampled chroma, and the integer YCbCr tables. The algorithms are the standard's and IJG's documentation; the target being a reference implementation's output, "correct" means identical to it, not close. Other decoders differ from it — MuPDF bundles its own libjpeg — and against them the difference is measured and recorded, never asserted as ours to close.
- Color: one component is gray; three are YCbCr unless an Adobe APP14 marker says transform 0, the component
identifiers read
R,G,B, or/ColorTransformis 0; four are CMYK, or YCCK when the Adobe marker says transform 2. Where the marker and/ColorTransformdisagree, the marker wins and the dictionary entry is ignored, as §7.4.8 says; without either, three components are YCbCr and the others not. Adobe's inverted CMYK is delivered as stored; the image's/Decodedecides (M07 writes[1 0 1 0 1 0 1 0]for such files). - Damage: restart intervals resynchronize a damaged scan at its next marker; data that ends early leaves the
rest of the image as libjpeg-turbo leaves it; both are reported. A stream that is not a JPEG at all — the PNG a
producer labeled
DCTDecode, in the remote corpus — yields no rows andimage.data-invalid. - EXIF orientation means nothing inside a PDF: the placement matrix decides, and M07 puts a JPEG file's orientation there. The decoder never rotates.
- The block wipe (M19's open decision, below): the entropy-coded data decoded to quantized coefficients; the
8 × 8 blocks of every component that intersect a mark given one DC value — the overlay's color, quantized — and
zero AC coefficients; the whole re-encoded as a baseline JPEG with the file's quantization tables and optimized
Huffman tables. Every block outside the marks keeps its coefficients exactly — no generation loss, as jpegtran's
-wipeoffers. The marked region grows to MCU boundaries, up to 16 pixels, and the report says by how much. Progressive and arithmetic-coded inputs come out baseline Huffman. The baseline re-encoder written here is the one M23's lossless entropy re-coding reuses.
JPEG 2000
- Codestream: SIZ, CAP, COD, COC, QCD, QCC, RGN, POC, PPM, PPT, TLM, PLM, PLT, CRG, COM, SOT, SOP, EPH; the five progression orders (LRCP, RLCP, RPCL, PCRL, CPRL) and progression changes; tile-parts in any order; precincts; every code-block style (selective arithmetic bypass, context reset, termination on each pass, vertically causal context, predictable termination, segmentation symbols); scalar derived and expounded quantization; region of interest by max-shift; the reversible 5/3 and irreversible 9/7 wavelets and component transforms; subsampled components upsampled to the image's grid.
- Files: JP2 and JPX boxes —
jp2h,ihdr,bpcc,colr(enumerated sRGB, grayscale and sYCC; a restricted ICC profile),pclrandcmap(palettes),cdef(opacity, premultiplied or not),res— and raw codestreams. - PDF's rules (§7.4.9): a
/ColorSpacein the image dictionary overrides the file's color;/SMaskInData1 or 2 turns the opacity channel into the mask;/Decodeis ignored except for an image mask; the bit depth is the codestream's, so 16-bit samples reach the sink as 16-bit samples. - Tier-1 is the hot loop: the MQ decoder and the three coding passes over pooled code-block buffers reused across blocks, 0 B allocated per code block once warm.
- The 9/7 wavelet is floating point, and still deterministic: single precision in a fixed order of operations, no fused multiply-add unless written explicitly, the same bytes on x64 and ARM64 — tested on both. The reversible path is integer and exact. OpenJPEG, the referee, is not bit-exact with other irreversible decoders either, so the tolerance against it is measured on the 9/7 path, and zero on the 5/3 path.
- Reduced resolution: decoding may stop at a coarser resolution level — a quarter, a sixteenth of the pixels —, which M23's downsampling and M25's thumbnails want; the sink is told the smaller layout, and the result says so.
- HTJ2K codestreams (a
CAPmarker announcing Part 15) are refused withimage.process-unsupported.
Color spaces and functions
M15 reads each color space's family and parameters, and until now exported Lab, Separation and DeviceN samples as their components. M22 evaluates them.
PdfColorSpacegainsToGray,ToRgbandToCmykover rows of components, through a lookup table built once per image and color space where the input depth allows it (every 8-bit case), so that the hot loop is a table read.- CIE-based spaces: CalGray, CalRGB and Lab through CIE XYZ to sRGB, with Bradford adaptation from the space's
white point to D65. ICCBased: a profile recognized as sRGB IEC 61966-2.1 — by its header and description, the
readers M20 wrote — converts as the identity; any other through its
/Alternate, or the device space of its/N, reported asimage.color-approximated. Indexed through its lookup string. Separation and DeviceN through their tint transform into the alternate space,/Alland/Nonehonored, a DeviceN's/Processand/Colorantsread. DeviceCMYK to RGB by the device formulas of §10, reported as approximate. A color-managed conversion is M29's. - Functions: type 0 (sampled —
/Encode,/Decode, 1 to 32 bits per sample;/Order3, cubic spline interpolation, evaluated as linear, as pdf.js and poppler do, and reported with the conversion'simage.color-approximated, since §7.10.2 ignores it only when a/Sizeis under 4), type 2 (exponential), type 3 (stitching, the last interval closed at its upper bound) and type 4 (the PostScript calculator). A type 4 program is compiled once into a compact form and run on a fixed stack of 100 operands — §7.10.5 requires room for at least 100 entries of every implementation, requires no more, and makes an overflow an error —, withifandifelseas jumps; having no loops, it costs its length.roll,indexandcopyare bounds-checked; a failure yields the range's lower bounds and onefunction.invalidper function. - Transcendental functions are ours. The power of a gamma curve, Lab's cube root, and type 4's
exp,ln,log,sin,cosandatanare computed by routines written from IEEE basic operations, which are correctly rounded everywhere;Math.Powand its kin defer to the platform's C library, whose last bit differs between Linux, Windows, macOS and WebAssembly, and a last bit becomes a byte at a rounding boundary. Invariant 6 holds only if the same file converts to the same bytes on every platform, and the test runs on x64 and ARM64. - A type 3 function is evaluated iteratively; one that reaches itself is cut and reported. An Indexed lookup shorter
than
(hival + 1) × componentsgives black for the missing entries, and an index abovehivalclamps, as viewers do; both areimage.indexed-out-of-range.
Lossless encoders
- CCITT G4 (T.6) for any 1-bit image:
/K -1,/BlackIs1chosen to match the source's polarity,/EncodedByteAlignfalse. - JBIG2 generic region: one immediate lossless generic region per image, arithmetic-coded with template 0, its nominal adaptive pixels and typical prediction; no globals. Never a symbol dictionary or a text region: that is where pattern matching lives (ADR 42).
- A policy chooses:
Smallestencodes both and keeps the smaller, ties going to G4, the older and more widely read. M19 calls it for redacted scans, M23 for recompression, M07 to join a multi-strip bilevel TIFF. - Flate with the PNG predictors, from M03's deflater, stays the lossless encoder for everything that is not bilevel: M21 re-encodes a JPEG 2000 image outside part 2's clause with it, and M19 a redacted JPEG 2000 image.
Pixels under a redaction
M19 left to this milestone how each codec's pixels are edited. The choice is recorded in an ADR at the next free number:
| Filter | Edited through | What stays exact |
|---|---|---|
| CCITT, JBIG2 | Decoded, the samples under each mark set, re-encoded in the smaller of G4 and JBIG2 generic | Every sample outside the marks; a JBIG2 image drawn from symbols comes out as one generic region |
| DCT | The block wipe | Every coefficient outside the marked MCUs; decoded pixels outside them and their one-pixel border, since chroma upsampling reads neighboring blocks |
| JPX | Decoded, re-encoded as Flate | Every decoded sample outside the marks; the file grows, and the report says by how much |
| Flate, LZW, RunLength, uncompressed | M19's own path | Unchanged |
The text layer
PdfRecognizedPage an engine's words: the image's size in pixels and its resolution; the frame the boxes are in —
ImagePixels of a named image, or PageVisual at a resolution (M15's hOCR and ALTO exports); the
page's orientation and its confidence; blocks > lines > words, each with a box, a baseline and
angle where known, a confidence and a language
PdfRecognizedText ReadHocr, ReadAlto, ReadTesseractTsv (Stream, PdfRecognizedTextOptions) -> pages
PdfTextLayer Add(page, recognized, PdfTextLayerOptions) -> what was written, and its report entries
PdfTextLayerOptions immutable: tagging (None, Paragraphs); language when the engine gives none; minimum word
confidence; a page with text (Skip, ReplaceOwnLayer, ReplaceInvisibleText); orientation
(FromEngine, FromTextLayer, None)
IOcrEngine Info (name, version, model) and RecognizeAsync(PdfOcrImage, PdfOcrRequest, CancellationToken)
PdfOcrImage the page image as the engine asks for it: Gray8, Bilevel or Rgb24 rows, or encoded PNG or TIFF
PdfOcr RunAsync(document, engine, PdfOcrOptions, IProgress<PdfProgress>?, CancellationToken)
-> PdfOcrReport: per page, recognized or skipped and why, words, mean confidence, the engine's
name, version and model
IPdfPageRasterizer Rasterize(page, resolution, cancellationToken) -> PdfRaster: the seam M25 fills, through which a
page that is not one image is rendered for the engine; M24's visual comparison consumes the
same seam
PdfRaster width, height, stride, format (Gray8, Rgb24, Rgba32), a pooled buffer, the page frame it covers
- Frames. Tesseract's hOCR, ALTO and TSV give pixels of the image the engine saw, origin top left, y down. A
box maps to default user space through the page image's placement matrix — rotated, skewed and cropped
placements included, the Xerox 5755's deskewed matrix among them —, or through the page's visual frame at a
resolution for M15's exports: crop box,
/Rotateand/UserUnitapplied once, as M15 converts them. hOCR'socr_pagebox,scan_resandppagenoand ALTO'sMeasurementUnit(pixel,mm10,inch1200) are read; an engine that resized the image is scaled back from the image's size in the file. - The font: one
Type0font per document, itsCIDFontType2descendant written by M08's TrueType writer — an empty glyph per code, one advance for all (500 units per 1,000 em),/DWequal to it, so that the dictionary and the program agree as PDF/A checks —; codes assigned to distinct character sequences in the order they first appear, andToUnicodemapping each to its text, supplementary planes and combining sequences included. The.notdefglyph is never drawn. - The words: one text object per line in render mode 3; each word placed by a text matrix at its box's start on
the line's baseline — hOCR's
baseline, else the box's bottom less a descender estimate — along the line's angle (hOCR'stextangle, ALTO'sROTATION), at a size from the line'sx_sizeor height, and scaled horizontally (Tz) so that its advance equals its box's width. A space code is placed in the gap before the next word of the line, so every extractor sees the space, not only those that infer one from a gap. - Where it goes: a new content stream appended to the page's
/Contents, preceded, when M09's balance analysis says the page leaves its state altered, by aqstream at the front — so that the layer starts in the page's initial state —, and bracketed by a marked-content sequence under a private tag the library recognizes. That is howReplaceOwnLayerremoves exactly its own layer, idempotently as M09's stamps are, and how M19 names it. - Tagging.
Paragraphs, in an untagged document: a structure tree is created —Document, then onePper recognized block with itsLang— with/MarkInfo /Marked true, and the page image made an/Artifact. In a tagged document, the page's layer is appended as aPartwhen no structure element owns the page's content; otherwise it is written inside an/Artifactsequence andocr.structure-not-extendedsays so. Headings, lists and tables are not inferred. - A page that has text — M15's
TextorOcrLayer— is skipped by default.ReplaceOwnLayerremoves a layer this library wrote;ReplaceInvisibleTextremoves any invisible text through M19's content editing pipeline, a copier's layer included, and reports it. PdfOcrsends the engine the image the page consists of — one image covering the crop box as M15'sImageOnlythreshold measures it, under a text layer being replaced or none —, decoded by the codecs given, in the form and at the resolution the engine asks for; mixed raster content (a JPEG background under JBIG2 or CCITT masks) is composited first. Any other page goes toIPdfPageRasterizerwhen one is given, and is otherwise skipped asocr.page-skipped. Pages are processed one at a time, with M03's progress (stagesDecoding,Recognizing,WritingTextLayer) and cancellation between them; whether the engine runs pages in parallel is the caller's affair, and the document's writes stay in page order.- Determinism. The layer is a function of the recognized words. The report records the engine's name, version
and model as
IOcrEngine.Infogives them, since invariant 6 holds only for an engine the caller pins (ADR 42). - The readers are hostile-input parsers:
XmlReaderwith DTD processing prohibited and no resolver, a depth and a size bound —MaxInputLength, 64 MB by default, with a typed exception, since recognized text is the caller's input rather than the file's —; hOCR'stitlemicro-syntax parsed by hand and bounded; a TSV row with the wrong number of fields skipped and counted.
Blank pages
PdfBlankPageDetector.Examine(page, codecs, options) -> PdfBlankVerdict:Blank,NotBlankorUndetermined, with the ink ratio, the method and the evidence. Detection removes nothing.- By content first, M07's rule extended with M15's facts: nothing painted inside the crop box, or only invisible text, white fills and artifacts. A page with visible text is never blank.
- By pixels when the page is one scanned image — M15's
ImageOnly, orOcrLayerwhose layer is empty: the image read row by row; a margin excluded, 5 % of each side by default, where scanner edges and punched holes are; a pixel counted as ink when its luminance after/Decodefalls below a threshold — Otsu's over the page's histogram, clamped to a band, for gray and color; the sample itself for one bit —; connected ink components smaller than a speck, 0.3 mm across at the image's resolution by default, discarded, by a union-find over the runs of two rows in bounded memory; and the page blank when the ink left is under a ratio, 0.1 % by default. Integer arithmetic throughout, so the verdict is the same on every machine. - The verdict is a predicate for M07's split by separator (
PdfPagePredicates.BlankByPixels); removing blank pages is M06's page removal, reported page by page.
Orientation
PdfPageOrientation.Detect(page) -> (0, 90, 180 or 270, confidence, evidence)from M15's lines in the displayed frame: glyphs counted by direction, vertical writing modes left out, a rotation proposed only when enough glyphs agree and one direction dominates — the count and the ratio fixed by slice 11 against the corpus and recorded —, andUndeterminedotherwise.Applysets/Rotateon the page itself, never on an ancestor since the attribute is inherited, to the value that makes the dominant direction read left to right, and reports the old and new values (orientation.applied). Content, annotations and the text layer are untouched:/Rotateturns them together.- A change of
/Rotatein a signed document is a page change under M04's classification: refused on a certified document unless the caller insists, reported otherwise. - With
PdfOcr,FromEnginetakes the orientation the engine reports — Tesseract's orientation and script detection gives one —, writes the words along the rotated baselines, and sets/Rotatein the same revision.
Bounds, classified (invariant 12, ADR 34)
| Bound | Class | Why |
|---|---|---|
| An image's decoded samples | The existing guard MaxDecodedStreamLength | An image's samples are its stream's decoded data |
| A decoder's working set beyond the rows it hands out — progressive JPEG coefficients, a row of JPEG 2000 tiles, a JBIG2 page or stripe with its symbol bitmaps and intermediate regions | Guard: PdfReaderLimits.MaxImageWorkingSet, 1 GiB by default, code limit.image-working-set | A valid 256 MB image in one JPEG 2000 tile holds four times its samples as 32-bit coefficients; the default admits every image the decoded-length default admits |
| Scans in a progressive JPEG | Guard: PdfReaderLimits.MaxJpegScans, code limit.jpeg-scans, its default set above every scan script the corpus and the common encoders write | Each scan is a pass over the image and the standard sets no limit, so a valid file can exceed any number |
| JPEG 2000 tiles, layers, decomposition levels, code-block size, coding passes, components | Internal constants: the standard's own maxima | They are field widths and ranges ISO/IEC 15444-1 defines; no valid codestream exceeds them |
| JBIG2 symbol counts, referred-to segments, table lines, pattern counts | Computed in 64 bits and checked against the working set | Counts the file declares; what they size is the working set |
| The PostScript calculator's stack | Internal constant, 100 | §7.10.5: no processor is required to hold more, and overflowing is an error, so no conforming writer can count on more |
| Stitching depth; color-space nesting (Indexed over Separation over ICCBased) | Evaluated iteratively, cycles cut | A chain is no longer than the objects in the file, each read once |
| hOCR, ALTO and TSV input | PdfRecognizedTextOptions.MaxInputLength, a typed exception | The caller's input, not the file's: not a reader guard |
Each guard is a PdfReaderLimits property with its limit.* code, its rows in ReaderLimitsTests, its line on the
reader-limits page, and its key in the manifest schema's readerLimits, which CorpusManifestSchemaTests holds to
PdfReaderLimits.
Diagnostics and findings
In PdfDiagnosticCodes, disjoint from rule identifiers (ADR 36) and from M12's image.format-unsupported,
image.animation-dropped and image.resolution-excessive:
| Code | Severity | Meaning |
|---|---|---|
image.rows-damaged | Warning | CCITT rows that did not decode, replaced by the previous row; the count |
image.data-truncated | Warning | The data ended before the image did; rows or blocks missing, filled as the referees fill them |
image.data-invalid | Warning | Data the decoder could not read at all; the image yields no rows |
image.process-unsupported | Warning | A JPEG process, a JBIG2 segment or a JPEG 2000 capability the decoders do not implement |
image.color-approximated | Information | A conversion through an alternate space or the device formulas |
image.mask-mismatch | Warning | A mask whose size or depth disagrees with its image; resampled onto its grid |
image.indexed-out-of-range | Warning | A lookup shorter than hival needs, or indices above it |
function.invalid | Warning | A function that cannot be evaluated as written: a stack overflow, a bound out of range, a cycle |
ocr.page-skipped | Information | A page not recognized, and why: it has text; it is not one image and no rasterizer was given |
ocr.word-outside-page | Warning | A recognized word whose box falls outside the crop box; dropped |
ocr.input-invalid | Warning | Recognized text the reader could not use: a malformed bbox, a TSV row with missing fields |
ocr.structure-not-extended | Warning | A tagged page whose structure could not take the layer; the layer written as an artifact |
orientation.applied, orientation.undetermined | Information | A page turned, with its evidence; a page whose text could not decide |
limit.image-working-set, limit.jpeg-scans | Warning | The guards above; the message names the property that lifts each |
Validation: M02's stream rule on decodable filters, which could not judge image filters, judges them now — CCITT
always, the others when the validator's options carry PdfImagingCodecs.Default. Its identifier stays M02's; a
finding on an image names the codec that decided. Each document whose image data does not decode gains it in
expect.findings.
Consumers completed
- M15:
PdfImageExtractorgains aDecodedoutput — PNG for gray, RGB and palette samples of 1 to 16 bits, TIFF for CMYK and for separations kept as components —, withpdfimages -pngand-tiffas referees; the Lab, Separation and DeviceN samples it exported as components are converted, or kept as components on request. - M19: given
PdfImagingCodecs.Default, the redactor edits CCITT, JBIG2, DCT and JPX pixels as the table above says; the documents it recorded as unsupported for pixel redaction until M22 lose that marker. - M21: the JPEG 2000 remedy row — an image outside part 2's clause re-encoded as Flate, reported as a change of encoding; a scan reaches level u once a text layer is written first.
- M07:
PdfImagePageOptions.JoinStripsdecodes the strips of a TIFF frame and re-encodes them as one image — G4 or JBIG2 for bilevel, Flate otherwise —, losslessly and reported; an arithmetic-coded JPEG file is transcoded to Flate on request, and refused as before otherwise. - M02: the stream rule above.
The command-line tool
text-layer FILE (--hocr F… | --alto F… | --tsv F…) [--frame image|page --dpi N] [--pages SEL] [--tag] [--replace own|invisible] [--orient] -o OUT; blank FILE [--list | --remove | --split] [--ink RATIO] [--margin PCT]; inventory FILE --images DIR --decode [png|tiff]. Each takes M06's page selection and exit codes
and writes its report as JSON on request; the AOT binary produces what the API produces. The tool never runs an
engine: it writes a layer from the files an engine wrote.
Slices
Each slice ends on a green commit, with the codes, guards and rules it introduces documented, its benchmark recorded
in docs/status.md, and its decoder in the nightly fuzzing campaign from the commit that adds it.
- The seam and CCITT. Delivers
PdfImage,PdfImageCodecscompleted,IPdfImageDecoder,IPdfImageRowSink,PdfImageRows,PdfImageReaderwith/Decode, masks and bit expansion over the core's filters, the CCITT decoder,MaxImageWorkingSetclassified, theimage.*codes it needs. Delivers the fuzzing harness and an ADR on its engine: coverage-guided — SharpFuzz with libFuzzer is the candidate, its fitness for .NET 10 established here — besideFuzzingTests' seeded mutations, which extend to encoded image streams (FuzzingSeedsgains the corpus's image streams). Proved by unit tests per code-table row (terminating and make-up codes, EOL, fill bits, the two-dimensional modes), per parameter, and per fault (damaged rows with and without EOL, a run pastColumns,/Rowsagainst/Height); an FsCheck property — for any byte sequence and parameters the decoder writes exactlyHeightrows ofColumnsbits and reads nothing outside its input —; random bitmaps encoded by libtiff in the container, committed as fixtures, decoded back to themselves; integration: every CCITT image in the corpus equals MuPDF's and poppler's decoding, compared as hashes of rows;CcittBenchmarksat 0 B per row. Leaves the satellite. - The Imaging satellite and JBIG2 generic regions. Delivers the project and package with its API baseline
(#42),
PdfImagingCodecs.Default, the MQ decoder shared with slice 6, generic regions (templates 0 to 3, adaptive pixels, TPGDON, MMR through the CCITT decoder), page information, end of stripe, end of page, immediate and intermediate regions,/JBIG2Globals. Proved by unit tests from T.88's worked examples, a striped page of unknown height, a generic region whose data ends early; integration: the generic regions ofvendor/us-federal/xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdfanddocusign-pdfkit-gsa-sf30-contract-modification.pdfequal jbig2dec's (MuPDF's container) and poppler's decoding. Leaves symbols. - JBIG2 symbols, text, refinement and halftones. Delivers symbol dictionaries (arithmetic and Huffman, with
refinement and aggregation), text regions, generic refinement, pattern dictionaries, halftone regions, the
standard and custom Huffman tables, and the 64-bit counting discipline. Proved by unit tests — a text region
whose three dictionaries' exports sum past 2³², refused before any allocation (FORCEDENTRY's shape); a refinement
with typical prediction; a halftone on a skewed grid —; integration: the Xerox IMF scan (symbol dictionaries in
/JBIG2Globals, arithmetic text regions) and the Xerox 5755 MRC scan equal jbig2dec's and poppler's; the segment types no committed document uses, from the conformance streams not in the corpus (below). Leaves JPEG. - JPEG, sequential. Delivers the marker parser, Huffman and arithmetic sequential decoding, the integer
inverse DCT, upsampling, color conversion, the Adobe and JFIF rules, restart intervals, the damage rules.
Proved by unit tests per marker and per fault (a table redefined between scans, a missing restart marker, a
scan naming a component the frame lacks); a property over fixtures
cjpegencoded from generated samples at every sampling factor and committed — each decodes byte-identical todjpeg's output, committed beside it —; integration: every committed baseline JPEG equals Pillow's decoding; MuPDF's and poppler's differences measured and recorded;JpegBenchmarks, and a JPEG row in the comparison benchmarks against SkiaSharp's decoder, recorded. Leaves progressive. - JPEG, progressive, and the block wipe. Delivers spectral selection and successive approximation, arithmetic
progressive, the coefficient buffer under
MaxImageWorkingSet,MaxJpegScans, the coefficient-domain wipe and the baseline re-encoder with optimized Huffman tables, and the ADR on pixels under a redaction. Proved by unit tests (every scan scriptjpegtran -progressiveand mozjpeg write, an image cut after its first scan, 10⁵ scans reaching the guard); integration: the eight progressive images of the committed invoices and forms equal Pillow's decoding; after a wipe, the quantized coefficients read back through libjpeg in the Python container are identical outside the marked MCUs, and the decoded samples identical outside them and their one-pixel border. Leaves JPEG 2000. - JPEG 2000, the core path. Delivers the codestream parser, tier-2 for every progression order, tier-1 on the
shared MQ decoder, dequantization, both wavelets and component transforms, tiles and tile-parts, subsampled
components, the JP2 boxes and color, raw codestreams. Proved by unit tests per marker and per fault; the
committed
vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf— 2717 × 3701, twelve tiles of 1024², RPCL, six layers, the 9/7 wavelet, SOP and EPH markers, segmentation symbols — within the tolerance of OpenJPEG'sopj_decompressand MuPDF; the conformance codestreams not in the corpus (below); the determinism test on x64 and ARM64 runners. Leaves the rest of Part 1. - JPEG 2000, complete. Delivers precincts and POC, PPM and PPT, every code-block style, region of interest,
palettes and
cdefwith/SMaskInData, 16-bit samples, sYCC, reduced-resolution decoding, HTJ2K refused;JpxBenchmarks. Proved by unit tests (a POC that revisits a resolution, packed headers split across tile-parts, a 256-entry palette with 16-bit outputs); integration, remote: the rows below against OpenJPEG. Leaves color. - Color spaces, functions and decoded export. Delivers the conversions, the four function types, the type 4
compiler, masks composited and un-premultiplied by
/Matte, M15'sDecodedexport,inventory --images --decode. Proved by unit tests with known values (the white point of Lab D50 to sRGB white; CalRGB with the sRGB matrix and gamma against the formulas; each type 4 operator; the stitching boundary rule; a lookup cut short); an FsCheck property — our transcendental routines are within one unit in the last place ofMath's on random inputs, and give the same bits on x64 and ARM64 —; integration:pdfimages -pngand-tiffequal ours for device and indexed spaces on every committed image, MuPDF's RGB within the tolerance for CIE-based and tint-transformed spaces; the decoded-sample hash added to M15'sexpect.imagesrows where the referees agree, with its schema. Leaves the encoders. - Lossless encoders and the consumers. Delivers the G4 and JBIG2 generic encoders, the
Smallestpolicy, M19's pixel path, M21's JPEG 2000 remedy, M07's strip joining and arithmetic-JPEG transcoding, and M02's stream rule over image filters. Proved by FsCheck — any bitmap encoded and decoded is itself —; our encoders' output decoded by libtiff and jbig2dec in the container is the input; integration: the M19, M21 and M07 rows below, veraPDF on M21's output. Leaves the text layer. - The text layer. Delivers
PdfRecognizedPage, the three readers, the glyphless font,PdfTextLayerwith its frames, tagging and replacement,IOcrEngine,PdfOcrImage,IPdfPageRasterizerandPdfRaster,PdfOcr, theocr.*codes, thetext-layerverb; adds the manifest'socrfield and its schema. Proved by unit tests (a word at every rotation, an ALTO file in each unit, a TSV with a missing column, an hOCR with an entity expansion and a 10 MBtitle, supplementary-plane text, a page already tagged, a page whose content leaves the state altered); an FsCheck property — any set of words placed in a random frame (rotated, cropped,/UserUnit) comes back from M15's extraction as the same words at the same boxes within the tolerance —;IOcrEnginesubstituted with NSubstitute to assert the orchestration — one call per page needing one, cancellation between pages, progress reported, pages written in order —; integration: the text-layer rows below through pdftotext and PyMuPDF, OCRmyPDF's renderer over the same hOCR as a second implementation, veraPDF at level 2u after M21's conversion, and its PDF/UA-1 profile on the tagged variant;TextLayerBenchmarks. Leaves blank pages and orientation. - Blank pages and orientation. Delivers
PdfBlankPageDetector,PdfPagePredicates.BlankByPixels,PdfPageOrientation, theblankverb andtext-layer --orient; addsexpect.blankPagesandexpect.orientationto the manifest and its schema. Proved by unit tests (a white page with dark scanner edges, a page carrying one signature, a speckled page, a page whose/Decodeinverts, a color page with show-through, a page of vertical Japanese, a landscape table on a portrait page); an FsCheck property — a mark larger than the speck anywhere inside the margins turns a blank verdict into not blank, and one smaller than the speck does not —; integration: ink ratios against an independent NumPy computation over poppler's decoded images; Tesseract's orientation detection on MuPDF's rendering of each turned page reading 0°. Leaves the whole. - Budgets, campaign and the whole. Delivers
ImagingBenchmarksacross the codecs withMemoryDiagnoser(megabytes per second, bytes per row), the working-set checkpoints on the heavy scans, the decoding rows of the comparison benchmarks recorded, the fuzzing campaign's record, the documentation. Proved by the budget rows below, the remote rows on a greenRemote corpusrun, and the campaign's record instatus.md.
Tests required
Unit — tests/AdCodicem.Pdf.Tests, the Imaging satellite's tests under Imaging/ as M10 placed its satellite's:
- CCITT: every code-table entry;
/Knegative, zero and positive;/EndOfLine,/EncodedByteAlignboth ways,/EndOfBlockfalse,/BlackIs1,/DamagedRowsBeforeError; rows cut, runs past the row, an EOL missing. - JBIG2: every segment type; arithmetic and Huffman paths; refinement and aggregation; striped and unstriped pages; every combination operator; globals referred to by a page; a segment referring to one that does not exist.
- JPEG: every process decoded and every one refused; every sampling factor; restart intervals; each Adobe and JFIF case of the color rule; scans that end early; the wipe on MCU boundaries and inside one.
- JPEG 2000: every marker, progression order, code-block style and quantization style; tile-parts out of order;
POC; PPM and PPT; a palette;
cdef;/SMaskInData0, 1 and 2; reduced resolution; HTJ2K refused. - Color and functions: each family's conversion against known values; each function type; every type 4 operator; stitching boundaries; lookups cut short; cycles.
- Text layer: each reader against its specification's examples; each frame; each tagging case; replacement of our own layer twice over giving the same bytes; the glyphless font read back by M08's parser with widths consistent.
- Blank pages and orientation: each case of slice 11.
- Hostile: for every decoder, an image declaring 65,535 × 65,535 pixels over ten bytes of data; a JBIG2 symbol
count summing past 2³²; a JPEG with 10⁵ scans; a JPEG 2000 codestream claiming 65,535 tiles and 32 decomposition
levels; a type 4 program of a million operators; an Indexed space whose lookup is empty; an hOCR entity expansion;
a TSV of a hundred million rows — each ends in rows, a report or a typed exception within its time and
allocation budget. The guards are reached, raised and thrown as
ReaderLimitsTestsdoes for the reader's. - Fuzzing: every decoder and reader joins
FuzzingTestsper commit and the nightly campaign, seeded with the corpus's encoded image streams and the conformance streams; the coverage-guided engine of slice 1 runs each decoder nightly on a time budget; a finding becomes a regression test inHostileInputTestsbefore it is fixed. - Determinism: decoded samples, conversions and text layers byte-identical across two runs, two cultures and the x64 and ARM64 runners.
Integration — tests/AdCodicem.Pdf.IntegrationTests, every referee in a container
(ADR 27):
- MuPDF —
mutool extractand PyMuPDF's pixmaps, for decoded samples of every filter, its bundled jbig2dec for JBIG2 and OpenJPEG for JPEG 2000;mutool drawfor renderings; - poppler —
pdfimages -list,-png,-tiff,-all;pdftotextwith-bbox-layoutfor the text layer; - Pillow over libjpeg-turbo, and libjpeg-turbo's
cjpeg,djpegandjpegtran, for JPEG; a coefficient reader over libjpeg for the wipe; - OpenJPEG —
opj_decompress, andopj_compressto generate fixtures; - libtiff — G3 and G4 fixtures and decoding of our G4 output; jbig2dec standalone for our JBIG2 output;
- PyMuPDF — words and
search_forquads over the text layer; - OCRmyPDF — its hOCR renderer, a second implementation of the layer, over the same sidecars;
- Tesseract — its orientation and script detection on renderings, the orientation referee;
- veraPDF — PDF/A-2u on scans given a layer and converted by M21, PDF/UA-1 on the tagged variant, PDF/A on M21's JPEG 2000 remedy;
- qpdf —
--checkon every document this milestone writes.
Where the referees disagree — MuPDF's bundled libjpeg and libjpeg-turbo on subsampled chroma, two CCITT
decoders on a damaged row — a row passes when we agree with the reading the manifest records and the reason it
records for the other. Disagreeing with every referee fails. The tolerances — JPEG against MuPDF, the 9/7 wavelet
against OpenJPEG, CIE conversions against MuPDF, word boxes against pdftotext and PyMuPDF — are fixed by the slices
that introduce them and recorded in status.md.
Acceptance conditions
"Every committed image" means the 423 image XObjects of the 53 committed documents that open with images, and
their inline images: 169 Flate, 96 uncompressed, 71 CCITT, 57 DCT (four behind Flate or ASCII85), 28 JBIG2, one
JPX, one LZW. The remote rows close only on a green Remote corpus run, recorded in status.md with its date.
| Documents | Behavior | Verified by |
|---|---|---|
The 71 committed CCITT images in nine documents — vendor/us-federal/acrobat3-import-irs-1040-1988-scan.pdf (G4 at 400 ppi), vendor/pikepdf/scanner-ccitt-endofline.pdf (G3 with /EndOfLine and a /Decode array), finereader8-frb-sr0115-examiner-guidance.pdf, vendor/uk-ogl/indesign-acrobat-hmcts-n208-form.pdf, pagemaker-distiller5-hmrc-iht205-form.pdf, and the image masks of vendor/opf-format-corpus/word9-distiller405-usgs-nwql-volatile-organics-methods.pdf, distiller952-pscript5-kb-pdf-risk-inventory.pdf, pdfmaker707-word-va-esig-developer-guide.pdf and pdfmaker8-word-va-lms-process-reference.pdf; remote, remote/ocrmypdf/tiff2pdf-35000px-ccitt-image.pdf (35,000² in 10.5 KB, 153 MB decoded), remote/pdfjs/itext5-ccitt-g4-mask-issue4379.pdf, the Kodak and Epson scans of remote/ocrmypdf/, the Ricoh scan remote/pdfjs/ricoh-3heights-scan-issue5747.pdf, Konica's CCITT masks, and the JHOVE tiff2pdf, Apex, Pixel Translations and Paper Capture scans | Samples identical to MuPDF's and poppler's decoding, compared as hashes of rows; the 35,000² image decoded holding two rows, within its time budget | CorpusImageDecodingTests.Ccitt_images_decode_to_the_referees_pixels (new) |
The 28 committed JBIG2 images — vendor/us-federal/xerox-workcentre-treasury-imf-report-scan.pdf (symbol dictionaries in /JBIG2Globals, text regions, striped pages), xerox-workcentre-5755-ocr-hud-fonsi-mrc.pdf (symbols, image masks over a JPEG background), xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdf and docusign-pdfkit-gsa-sf30-contract-modification.pdf (generic regions); remote, remote/pdf-association/abledocs-pdfua1-tagged-textbook-scan.pdf, remote/opf-format-corpus/jhove-hul-136-acrobat7-paper-capture-scanned-report.pdf (314 pages), jhove-hul-117-acrobat101-student-design-report.pdf, jhove-hul-35-atypon-pdfplus-journal-article.pdf | Samples identical to jbig2dec's (MuPDF) and poppler's; the segment types each file uses recorded, and each type the corpus lacks covered by the conformance streams not in the corpus (below) | CorpusImageDecodingTests.Jbig2_images_decode_to_the_referees_pixels (new) |
The 57 committed JPEG images — 45 baseline in three components, eight progressive (the Mustang, GnuAccounting and factur-x Python invoices, DILA's signed notice, the LibreOffice Cerfa, OpenOffice's PNG page), flate-over-dct and ASCII85 chains, Acrobat 11's image whose /Height was altered —; remote, CMYK and YCCK in remote/ocrmypdf/photoshop-cc2015-pdfx3-cmyk.pdf (/ColorTransform over CMYK) and remote/pdf-association/abledocs-pdfua1-textbook-chapter.pdf, the damaged DCT inside Flate of remote/pdfjs/canon-scan-junk-after-eof-scan-bad.pdf | Samples byte-identical to Pillow's decoding over libjpeg-turbo; MuPDF's and poppler's differences within the tolerance recorded; the altered image and the damaged scan decoded as far as their data goes and reported | CorpusImageDecodingTests.Jpeg_images_decode_as_libjpeg_turbo_does (new) |
JPEG 2000: vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf; remote, remote/pdfjs/acrobat8-jpx-precincts-issue5475.pdf (precincts, no /ColorSpace), the soft-masked JPEG 2000 of the AbleDocs chapter, remote/pdf-association/indesign-cs6-pdfua1-brochure.pdf, the portfolio's embedded JPEG 2000 in remote/opf-format-corpus/acrobat9-portfolio-signed-3d.pdf, the 125 plates of remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf, and remote/pdfbox/distiller6-zeroed-object-stream-pdfbox3947.pdf | Samples within the 9/7 tolerance of OpenJPEG's (through MuPDF and opj_decompress), and identical on the reversible path; /SMaskInData opacity equal to MuPDF's mask; each plate decoded holding one row of tiles, within the working-set guard | CorpusImageDecodingTests.Jpeg_2000_images_decode_as_openjpeg_does (new) |
Color: the Indexed images over CalRGB (word9-distiller405-usgs-nwql-volatile-organics-methods.pdf), DeviceCMYK (vendor/us-federal/designer-distiller23-uscis-i9-javascript-form.pdf) and ICCBased (pdfmaker7-powerpoint-va-cancer-database-course.pdf), the 48 ICCBased images of 16 documents, vendor/pdf-association/handwritten-inline-image-abbreviations.pdf (a named CalRGB, /Decode arrays), handwritten-indexed-color-out-of-range.pdf; remote, DeviceN under a type 4 tint transform in the Acrobat 9 portfolio; Lab and Separation images and /Matte soft masks — not in the corpus (below) | PNG and TIFF exports equal pdfimages -png and -tiff for device and indexed spaces; RGB within the tolerance of MuPDF's for CIE-based and tint-transformed spaces; the out-of-range lookup reported, as the referees paint it | CorpusImageExportTests.Decoded_images_export_as_the_referees_do (new) |
Masks: the 105 soft masks of ten committed documents (Word's text drawn as images, the iBooks Author pages, OpenOffice's PNG, the OZEV invoice, the DH factsheet), the CCITT and JBIG2 image masks above, the stencil /Mask stream of the remote iText G4 file; a color-key mask and a /Matte soft mask — not in the corpus (below) | Alpha and stencil samples identical to MuPDF's masks, each at its own size | CorpusImageExportTests.Masks_decode_with_their_images (new) |
| Every committed image, and every remote one | Identical samples in two runs, under two cultures, and on the x64 and ARM64 runners | CorpusImageDecodingTests.Decoding_is_deterministic_everywhere (new) |
M19's scans: xerox-workcentre-treasury-imf-report-scan.pdf (JBIG2), acrobat3-import-irs-1040-1988-scan.pdf (CCITT G4), documents/scan/reportlab-scanned-receipt.pdf (DCT), imagemagick-false-pdfa1b-jpx.pdf (JPX), and the OCR'd finereader8-frb-sr0115-examiner-guidance.pdf and xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdf | Redacted with PdfImagingCodecs.Default: every sample under a mark uniform, every other exact as the redaction table says — coefficients outside the marked MCUs for DCT, by the libjpeg reader —; the OCR glyphs under the marks gone; the words under the marks absent from Tesseract's reading of MuPDF's rendering; the unsupported marker M19 recorded removed | CorpusImageRedactionTests.Scanned_pixels_under_a_mark_are_removed_and_no_others (new) |
vendor/opf-format-corpus/imagemagick-false-pdfa1b-jpx.pdf, and every JPEG 2000 document M21 recorded as not convertible because an image breaks part 2's clause | Converted by M21 to PDF/A-2b through M22's decoder: an image outside the clause re-encoded as Flate, its samples identical to its decoding, one inside it left as it was; veraPDF accepts; the report lists each change of encoding; M21's unsupported marker removed | CorpusPdfAConversionTests.Jpeg_2000_images_outside_the_clause_are_re_encoded (new) |
Every document whose image data does not decode — the damaged scans above, remote remote/pdf-differences/unidentified-wine-merchant-sheet-png-under-dctdecode.pdf (a PNG under DCTDecode), and every hand-written hostile file | No untyped exception, no hang, each within its time and allocation budget; rows delivered as far as the data goes; M02's stream rule reports each in expect.findings | CorpusImageDecodingTests.Damaged_image_data_decodes_as_far_as_it_goes (new) |
| The committed image-only scans — the receipt, the Xerox IMF scan, the IRS 1040 import, the G3 scanner page, ImageMagick's JPX page — given hOCR, ALTO and TSV from a pinned Tesseract — not in the corpus (below) | Every recognized word found by pdftotext and PyMuPDF with its text, at its box within the tolerance after the frame conversion; PyMuPDF's search_for finds each line's text; OCRmyPDF's rendering of the same hOCR agrees on the words and their order | CorpusTextLayerTests.Recognized_words_are_found_where_the_engine_saw_them (new) |
The OCR'd committed scans — vendor/us-federal/hp-mfp-acrobat-ocr-nih-report.pdf (Acrobat), the Xerox 5335 and 5755 copier layers (render mode 3, Tz up to 2000 %), finereader8-frb-sr0115-examiner-guidance.pdf (text under the image), the scanned page of vendor/fr-licence-ouverte/pdfmaker-acrobat-cerfa-12156-form.pdf; remote, the Konica, Ricoh, Canon and OmniPage scans | M15's hOCR and ALTO exports, re-applied with ReplaceInvisibleText, give back the same words, as pdftotext reads them, at the same boxes within the tolerance; applied twice with ReplaceOwnLayer, the same bytes | CorpusTextLayerTests.Exported_layers_reapply_to_the_same_words (new) |
| The image-only scans above with their layer, converted by M21 | veraPDF accepts PDF/A-2u; with Paragraphs, veraPDF's PDF/UA-1 profile reports no failure and M20's verdict is ConformsPendingReview; pdfinfo -struct-text gives each block's text in order | CorpusTextLayerTests.Scans_with_a_text_layer_reach_pdf_a_2u (new) |
Blank pages: remote/ecan/konica-bizhub-c554e-letter-scan.pdf, whose scanned blank page the manifest is to record; the committed image-only scans, none blank; a copier batch with blank backs and separator sheets, and near-blank pages — a page number alone, a signature alone, punched holes, show-through — not in the corpus (below) | Verdicts equal expect.blankPages; no page with visible text or a signature called blank; ink ratios equal NumPy's over poppler's decoded images within 0.01 percentage points; M07's split by BlankByPixels gives the parts the manifest records | CorpusBlankPageTests.Blank_scanned_pages_are_found_and_nothing_else (new) |
Orientation: docusign-pdfkit-gsa-sf30-contract-modification.pdf (/Rotate 270, upright as displayed), the Xerox 5755's deskewed placement, remote remote/ocrmypdf/epson-scan-indirect-rotate.pdf (an indirect /Rotate) and remote/maine-legislature/ricoh-docusign-itextsharp-state-contract-amendment.pdf; scans turned by 90, 180 and 270° without /Rotate, with a layer — not in the corpus (below) | The upright ones unchanged; the turned ones given the /Rotate after which Tesseract's orientation detection on MuPDF's rendering reads 0°; expect.orientation met; nothing but /Rotate changed, object by object | CorpusOrientationTests.Pages_are_turned_upright_from_their_text_layer (new) |
Every decoder and reader, seeded with every committed image stream, the conformance streams, and every damaged/* document | A nightly campaign, mutation and coverage-guided, finds no untyped exception, hang or unbounded allocation over fourteen consecutive nights before the milestone closes; its executions, coverage and findings, each fixed with a regression test, recorded in status.md | FuzzingTests.Decoding_mutated_image_data_either_works_or_reports (new), the campaign's record |
remote/usgs/omnipage-usgs-professional-paper-1-1902.pdf (147 MB of JPEG 2000), the Paper Capture scan (314 pages of CCITT and JBIG2), the 35,000² CCITT image | Decoding every image of every page holds its working set within the budgets recorded in status.md, flat across pages; 0 B allocated per row once warm, the rows excepted | CorpusImageDecodingTests.Heavy_scans_decode_within_their_budget (new), ImagingBenchmarks (new) |
| The same operations through the tool | The AOT binary's text-layer, blank and inventory --images --decode produce what the API produces | CorpusToolTests.Text_layer_blank_and_decoded_images_match_the_api (new) |
Corpus
What the corpus holds
- CCITT: 71 committed images in nine documents, G4 and one G3 with
/EndOfLine(ccitt,ccitt-g4,ccitt-g3,ccitt-endofline,ccitt-image-mask); remote, 25 documents between them, the 35,000² tiff2pdf image (huge-decoded-image,high-compression-ratio), and the JHOVE scans from tiff2pdf, Apex, Pixel Translations and Acrobat Paper Capture. - JBIG2: 28 committed images in four documents — two with
/JBIG2Globals, symbol dictionaries and arithmetic text regions, two with generic regions only (jbig2,jbig2-globals,jbig2-without-globals,jbig2-image-masks,jbig2-bitonal-full-page-image); remote, the AbleDocs scan, the 314-page Paper Capture report, two JHOVE papers. - JPEG: 57 committed images — baseline, progressive, behind Flate and ASCII85 (
dct-image,jpeg-image,flate-over-dct); remote CMYK and YCCK (dct-cmyk-images,dct-colortransform-on-cmyk,cmyk-jpeg), a corrupt DCT (dct-corrupt,dct-in-flate), a PNG underDCTDecode(png-labeled-dctdecode). - JPEG 2000: one committed (a JP2 in RPCL with tiles, layers, SOP and EPH); remote, eight documents
(
jpx,jp2-container,jpx-no-colorspace,jpx-multiple-precincts,jpx-soft-mask,large-jpx-plate,embedded-jpx). - Color and masks: Indexed over DeviceRGB, DeviceCMYK, CalRGB and ICCBased, ICCBased with
/N 3, 105 soft masks, CCITT and JBIG2 image masks,/Decodearrays, an Indexed lookup out of range (indexed-color-out-of-range); remote, a stencil/Maskstream (explicit-mask), DeviceN with a type 4 function (devicen-nchannel,type4-function), a pattern color space. - OCR layers: from Acrobat, two Xerox copiers, FineReader 8, a Cerfa page; remote, Konica, Ricoh, Canon,
OmniPage, AbleDocs (
ocr-layer,copier-native-ocr,invisible-text-render-mode-3,text-under-image,ocr-text-painted-under-page-image). - Scans without text: five committed image-only documents (
image-only,no-text); remote, the Kodak, Epson and tiff2pdf scans. - Pages placed or turned:
rotate-270,indirect-rotate,rotated-deskewed-image-matrix,rotated-page. - Blank pages: one remote (
blank-page, the Konica scan). - Damage:
image-dimensions-wrong,dct-corrupt,object-stream-destroyed(a JPEG 2000 document), the hand-written hostile files.
What it lacks
| Need | Why | Priority | Likely source |
|---|---|---|---|
| hOCR, ALTO and TSV of the committed image-only scans from a pinned Tesseract (French and English models), with a manifest field naming them and the engine's version | The text-layer and PDF/A-2u rows need recognized words, and documents are committed, not generated at test time: Tesseract's output changes with its version | 1 | Generated here: Tesseract from the distribution's packages in a container, its version and models recorded in build_corpus.py, the output reviewed and committed under tests/corpus/sources/ocr/ |
| JBIG2 streams with the segment types no committed document uses: refinement, pattern dictionaries and halftone regions, Huffman-coded symbol dictionaries and text regions, custom tables, MMR generic regions, intermediate regions | The FORCEDENTRY class of fault lives in these paths, and a decoder proven on two segment types is not proven; the fuzzing campaign needs them as seeds | 1 | A public source (W20): the ITU-T T.88 test bitstream and the conformance streams in jbig2dec's and pdf.js's test suites, their terms read first, wrapped one per page by a recorded pikepdf transformation; pdf.js's JBIG2 bug-report PDFs to the remote corpus |
JPEG 2000 beyond the one committed file: reversible 5/3, raw codestreams, subsampled components, palettes, cdef opacity with /SMaskInData 1 and 2, 16-bit, sYCC, four components, every code-block style and progression order, POC, PPM and PPT, region of interest | One committed JPX cannot prove a decoder, and the remote ones are tested only nightly | 1 | Generated here with OpenJPEG's opj_compress over our own images, wrapped byte for byte by img2pdf (pypi); a public source: the ISO/IEC 15444-4 conformance codestreams in OpenJPEG's data repository, terms read first |
Scans turned by 90, 180 and 270° without /Rotate, with and without a layer, and one with a landscape table on portrait pages | The orientation rows have nothing to turn: every committed rotated page is already upright as displayed | 1 | Derived here: a recorded pikepdf transformation turning committed scans' content, Tesseract's sidecars made on the turned rendering; a contribution (W03): copier output fed the wrong way |
| A copier batch with blank backs and scanned separator sheets, and near-blank pages — a page number alone, a signature alone, punched holes, show-through, speckle | One blank page cannot show the detector's two errors, and M07's separator split waits for a real batch | 1 | A contribution (W03), as M07 already asked; derived here meanwhile: blank sheets scanned from our own paper, and near-blank pages cut from committed scans, each recorded |
expect fields: decoded-sample hashes on M15's expect.images rows where the referees agree, blankPages, orientation, and the ocr sidecar field beside expect | Expectations must come from the file and independent tools; the unit suite asserts pixels without a container only if the hashes are in the manifest | 1 | Generated here: build_corpus.py records the hashes from MuPDF and poppler and the reason where they disagree; the pages from a person's review, recorded with it |
| JPEG in CMYK and YCCK committed, restart intervals, 4:2:2, 4:1:1 and 4:4:0 sampling, arithmetic coding, a progressive image cut short | The committed JPEGs are all YCbCr at common sampling; CMYK and damage are remote only; arithmetic coding is absent | 2 | Generated here with libjpeg-turbo's cjpeg and jpegtran (-restart, -sample, -arithmetic), wrapped by img2pdf |
CCITT with /BlackIs1 true, /EncodedByteAlign true, two-dimensional G3 (/K > 0), /EndOfBlock false, and damaged rows under /DamagedRowsBeforeError | The committed CCITT is G4 but for one G3 page; every other parameter has unit tests only | 2 | Generated here with libtiff's tiffcp and fax2tiff, wrapped by img2pdf or a recorded pikepdf construction; a public source: fax-server output in public records |
Images in Lab, Separation and DeviceN under type 0, 2 and 3 tint transforms, CalGray, 2- and 4-bit samples, a color-key /Mask array, /Matte soft masks | The color and mask rows name spaces and masks no committed image uses | 2 | Generated here: a recorded pikepdf construction over committed images' samples; a public source: the Ghent Workgroup's output suites, terms read first |
| An ALTO file from a commercial engine with its page image | The ALTO reader is proven on Tesseract's dialect and M15's own export only | 2 | A public source (W20): the Library of Congress's Chronicling America, public-domain page images with ABBYY's ALTO — to the remote corpus when over 2 MB |
Multi-strip TIFFs, bilevel and gray, and a TIFF with FillOrder 2, to join | M07's strip joining is proven on synthetic files only | 3 | Generated here with libtiff's tiffcp -r |
Traps
- FORCEDENTRY was a count. A JBIG2 text region's symbol count summed in 32 bits and trusted after it overflowed. Every count a segment declares is 64-bit and checked against the working set before anything is sized by it.
- The arithmetic decoders never run dry: T.88 and 15444-1 feed 0xFF past the end of the data. A loop that waits for the input to end does not end.
- Pattern-matching JBIG2 changes what a scan says, and the corpus's Xerox files are made of it. Decode it as it is; never write it.
- A JPEG decoder that is "close" is wrong when the reference is exact: the integer inverse DCT and the upsampling rounding decide every byte. Different references disagree among themselves; say which one is the reference, and measure the others.
- Chroma upsampling reads the neighboring blocks: pixels bordering a wiped MCU change even though their coefficients do not. The redaction guarantee is on coefficients, and one pixel of border on samples.
/ColorTransformand the Adobe marker can disagree, Adobe's CMYK is inverted, and a JPEG's EXIF orientation is ignored inside a PDF. The Photoshop PDF/X-3 file carries the first case.- CCITT's polarity is stated twice:
/BlackIs1and the image's/Decode. The scanner page sets/EndOfLineand a/Decodearray; inverting twice is not inverting. /Rowsand/Heightdisagree, and/EncodedByteAlignpads differently for/K< 0 and/K≥ 0 — producers get both wrong.- JPEG 2000 color comes from three places —
/ColorSpace, the JP2colrbox, the component count — and the dictionary wins when present; opacity may be premultiplied; subsampled components must be upsampled before any color transform. - Floating point is deterministic only if written to be: no implicit fused multiply-add, a fixed order of
operations, no
Math.Powwhose last bit the platform's C library chooses. - A progressive JPEG holds the whole image as coefficients until its last scan; the working set is twice the samples, not a row.
- An image mask paints with the current color, a soft mask is an image of its own size, and
/Mattepremultiplies — extraction and compositing must not confuse the three. - Extractors infer spaces differently. A text layer without explicit spaces reads as one word in one tool and
as words in another; an OCR layer with
Tzat 2000 % (the Xerox copiers') still gives each word the box the engine saw. - hOCR is y-down and in the image's pixels; PDF is y-up and in points, through a placement that may rotate,
skew or crop the image, and a page that may be rotated, cropped and scaled by
/UserUnit. - An XML file from an engine is hostile input too: an entity expansion in an hOCR file is a denial of service like any other.
- Invisible text is what M19's sanitizer removes as
hidden.invisible-textandhidden.ocr-layer: our layer is marked so that it can be named, kept or removed on purpose. - A blank sheet is not white: scanners leave dark edges, punched holes, show-through from the back and JPEG noise; and a page bearing only a signature or a page number is not blank.
- Orientation from text fails without text, and a page may hold text in two directions — a landscape table on a portrait page. Undecided is an answer.
- An engine's output depends on its version, models and threads. The layer is deterministic for given words; the words are the engine's, and the report says which engine gave them.
Documentation
docs/website/docs/concepts/images.md(new): the image seam, rows and working sets, masks, color evaluation and where it approximates, the codec set and the Imaging satellite, the contract of a caller's own codec.docs/website/docs/guides/scans.md(new): decoding and exporting scans, blank pages, orientation, and joining TIFF strips.docs/website/docs/guides/ocr-text-layer.md(new): recognized text from hOCR, ALTO and TSV, frames, tagging, replacing a layer,IOcrEnginewith the Tesseract adapter sample, reaching PDF/A-2u through M21, and what the library does not do (run an engine, deskew).docs/website/docs/reference/reader-limits.mdanddiagnostics.md:MaxImageWorkingSet,MaxJpegScans, and theimage.*,function.*,ocr.*andorientation.*codes.- The redaction guide M19 wrote: pixels of every codec, per the redaction table.
docs/website/docs/reference/tool/:text-layer,blank,inventory --images --decode.docs/website/docs/introduction.mdanddocs/features/features.json: theimagingentry brought to its state.docs/architecture.md: the core'sImages/,Graphics/andOcr/, the Imaging satellite's contents and dependencies, the fuzzing harness.SECURITY.md: decoders as attack surface, and the fuzzing that covers them.docs/corpus.md, the manifest schema andtests/corpus/README.md: theocrfield, the decoded-sample hashes,blankPages,orientation, the newreaderLimitskeys;docs/corpus-sources.md: the sidecars' engine and version, the conformance streams' terms.- The ADRs of slices 1 and 5: the fuzzing engine, and pixels under a redaction.
docs/status.md: the tolerances, the working-set measurements, the campaign's record.
Exit criteria
- CCITT decodes in the core; JBIG2, JPEG and JPEG 2000 in the Imaging satellite, with every process, segment type and capability the design lists, and each one it refuses refused with a diagnostic.
- Color spaces and functions convert samples identically on x64 and ARM64, and approximations are reported.
- The G4 and JBIG2 generic encoders are lossless, verified by independent decoders, and no code path writes a JBIG2 symbol.
- M15's decoded export, M19's pixel redaction, M21's JPEG 2000 remedy, M07's strip joining and M02's stream rule
over image filters work through the codec set, and the
unsupportedmarkers that waited for M22 are gone. - The text layer is written from hOCR, ALTO and TSV, tagged on request, replaced idempotently, and a scan with a layer reaches PDF/A-2u through M21, confirmed by veraPDF.
- Blank pages and orientation are detected as designed, their thresholds fixed and recorded.
-
MaxImageWorkingSetandMaxJpegScansare guards with their codes, tests, documentation and schema keys; every other bound is classified where it is declared. - The ADRs on the fuzzing engine and on pixels under a redaction are accepted.
- The priority-1 gaps above are filled; each remaining gap is recorded in
docs/corpus-contributions.md. - The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on a
green
Remote corpusrun recorded instatus.md. - Unit tests cover each behavior, its degenerate cases and its hostile ones; the FsCheck properties hold; every decoder and reader has run fourteen consecutive nights in the fuzzing campaign with no open finding.
- Integration tests run MuPDF, poppler, Pillow and libjpeg-turbo, OpenJPEG, libtiff, jbig2dec, PyMuPDF, OCRmyPDF, Tesseract, veraPDF and qpdf, each in a container.
-
CcittBenchmarks,JpegBenchmarks,JpxBenchmarks,TextLayerBenchmarksandImagingBenchmarksrun withMemoryDiagnoser;status.mdrecords the working sets and the comparison rows. - The tool's verbs ship in the dotnet tool and the AOT binaries, documented.
- The documentation site publishes the pages listed above.
- Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).