M15 — Extraction and analysis
State: to do — Depends on: M06, M08, M11, M14 — Extraction per ADR 15 (the tagged structure first, heuristics with a stated confidence); image data left encoded where a decoder is needed, by ADR 42
Goal
Read what a PDF contains — its text in reading order with positions, fonts and sizes, its tables with the confidence the detection earned, its images, metadata, outline, attachments and structure —, find text in it with the quads a highlight or a redaction needs, export it to Markdown, JSON, ALTO and hOCR anchored to pages, page labels and Bates numbers, and describe everything a document carries in one inventory; one page in memory at a time, and never noise where there is no text.
A case file is read as much as it is written. An exhibit received as a scan must say that it has no text, so
that someone runs OCR on it; a contract fed to a retrieval pipeline must come back in reading order, cited by
page label and Bates number; a party's name must be found on the page and at the place a redaction (M19) will
remove it. Everything after this milestone that looks inside a page stands on what is built here: M18 links
"pièce n° 12" by finding it, M19 redacts what search finds, M22 writes text layers in the formats exported here,
M24 compares the words extracted here, and M25 rasterizes through the interpreter written here
(ADR 16). The failures this milestone exists
to prevent are the ordinary ones: a two-column annex read across the columns; a ligature extracted as nothing
and an fi as fi; an OCR layer skipped because it is invisible, or a hidden layer's text returned as if it
were shown; a scan whose copier stamped a font with no ToUnicode returning a line of gibberish; a Type 3 font's
text lost; Arabic returned backwards; a word split by a hyphen at a line break that search cannot find; a table
invented and presented as fact.
Scope
In:
- the content stream interpreter in the core — graphics and text state, every operator of ISO 32000-2 §8 and
§9 interpreted, form XObjects, Type 3 glyph procedures, inline and XObject images, shadings, marked content,
annotation appearances placed by their
/Matrixand/BBox— over M08'sPdfContentReader, with a device seam that M19's editing pipeline and M25's rasterizer plug into; - fonts read for extraction: simple fonts' encodings (base encodings,
/Differences, built-in encodings of embedded programs, the symbolic TrueType rules of §9.6.5.4), glyph names to Unicode,ToUnicodeCMaps and their faults, Type 1 programs (FontFile), bare and CID-keyed CFF through M08's parser, Type 3 fonts, CID fonts with/W,/W2and embedded CMaps, the standard 14 through M08's metrics, and the predefined CJK CMaps with the Adobe collections' CID-to-Unicode maps, which M08 deferred here; - positioned glyphs with Unicode text, font, size, color, render mode, box and origin, the marked content they belong to, and visibility flags — hidden by optional content (M11's evaluator), invisible by render mode, outside the crop box, clipped away, covered by an image drawn later;
- words, lines, blocks and columns, right-to-left runs in logical order, vertical writing,
ActualTexthonored, duplicated "fake bold" glyphs merged, hyphens at line ends kept and understood; - reading order: the structure tree first where the document is tagged, then a deterministic layout analysis, each page saying which it used and with what confidence (ADR 15);
- the public read model of the structure tree — the internal reader M13 left to this milestone —, with role maps, attributes, alternative descriptions, marked-content and object references;
- table detection — tagged, ruled and aligned — with a confidence and its evidence, and tables continued across pages linked;
- page analysis: what kind of text a page has (none, image only, an OCR layer, vector outlines, unreadable, text), its content bounding box (which M09's crop asked for), and running headers and footers recognized; M07's split by separator page given a predicate on a page's text, which M07 left to this milestone;
- images listed as
pdfimages -listlists them, with the effective resolution of each placement, and extracted aspdfimages -allextracts them — data passed through where a decoder would be needed (M22), samples written as PNG or TIFF where none is; - metadata, outline and attachments, composed from M14's
PdfXmpMetadata, M07'sPdfOutlineand M06'sPdfAttachmentsrather than read again; - annotations' text: the text under a markup annotation's quads, a comment summary quoting it (M11's deferral), and the text of visible annotation appearances, which extraction includes as poppler does;
- an outline generated from a tagged document's headings or from type sizes, with a confidence, returned as M07's model for M07's editing API to apply (M07's deferral);
- search — literal, and regular expressions without backtracking by default —, folding case, diacritics, ligatures and compatibility forms, across hyphenated line breaks, returning page, page label, quads and structure element;
- exports: plain text, Markdown, JSON with a published schema, ALTO v4 and hOCR 1.2, chunked by page or by section with anchors — page index, page label, Bates number, the section's path — for retrieval pipelines;
- a feature inventory of a document: fonts, images and their resolution, color spaces and output intents, annotations, forms, signatures, attachments, layers, conformance claims, active content and space by category;
- the
content.*validation rules the M02 structural profile waited for — operators out of place, operands of the wrong kind, a resource name the resources lack — answering the iPRES 2017 content cases M02 left open; - the visible check of a received Factur-X or ZUGFeRD hybrid: its invoice number, dates and totals on the
page against its embedded XML — a slice in the
AdCodicem.Pdf.FacturXsatellite that runs M14's visible-consistency matcher over this milestone's search, which M14 deferred here (the maintainer's decision of 2026-09-27; hence M14 among the dependencies); - the command-line tool's
text,markdown,searchandinventoryverbs.
Out, explicitly:
- OCR and the invisible text layer written onto a scan — M22, behind
IOcrEngine(ADR 42). M15 says a page has no text and how much of it an image covers; M22's writer reads the hOCR and ALTO exported here; - decoding JBIG2, JPEG 2000, CCITT and JPEG to pixels, and evaluating color spaces and functions for image
export — M22 (ADR 42); until then those images are extracted as their encoded data, which is what
pdfimages -alldoes too; - rendering pixels — M25, on this interpreter; comparison and zone templates — M24; redaction and rewriting content — M19, on this interpreter and M11's marked-content filter;
- decrypting — M16. The reader raises
PdfEncryptedExceptionuntil then, so encrypted corpus documents are excepted from this milestone's acceptance, or read through a decrypted twin where a row needs one; - form field values, XFA datasets and usage rights beyond their presence — M16; the inventory counts them;
- executing JavaScript or rendering XFA — never (ADR 37); the inventory reports active content, M19 removes it;
- reconstructing Word, Excel or HTML from a PDF — never (ADR 37): the Markdown and JSON exports are the machine route;
- embeddings, vector stores and calls to a language model — never in the library; chunks are the hand-off;
- heuristic tagging of untagged documents and editing a received structure tree — open questions of the roadmap; mathematical formulae rebuilt as MathML — not planned (M30 writes MathML, nobody reads it back);
Design
Where it lives
In the core, with no dependency (invariant 1): System.Text.Json with source generation for the JSON exports
(AOT-safe), XmlWriter for ALTO and hOCR, RegexOptions.NonBacktracking, ZLibStream for the PNGs.
| Part | Where | Why |
|---|---|---|
| The interpreter, the graphics state, the device seam | Content/ | M08's reader and builder live there; M19 and M25 extend it |
Encodings, glyph names, ToUnicode and predefined CMaps, Type 1 and Type 3 reading | Fonts/ | Beside M08's parser, whose CFF and TrueType readers it reuses |
| Glyphs, words, lines, blocks, reading order, tables, search | Text/, namespace AdCodicem.Pdf.Text | The public model callers use |
| The structure tree read model | Structure/ | Beside M13's writer, whose PdfStructureType and PdfStructureAttributes it shares |
| Images, exports, chunks, the inventory | Extraction/, namespace AdCodicem.Pdf.Extraction | Composition of the above with M06, M07, M11 and M14's models |
| Unicode property and folding tables | Text/Unicode/, generated at build time | See below |
| The visible check of a received hybrid | AdCodicem.Pdf.FacturX | M14's matcher over this milestone's search; the satellite reads the core, never the reverse |
Unicode data is generated, not borrowed from the platform. string.Normalize and culture-aware comparison
call ICU on Linux — its version differs between distributions — and behave otherwise under
InvariantGlobalization, which AOT and container deployments commonly set, and which M23's container profile
may. Search folding and the logical ordering of
right-to-left text would then give different answers on different machines, against invariant 6. The tables
this milestone needs — case folding (full mappings, so that ß matches ss), canonical and compatibility
decompositions, the combining classes, Bidi_Class — are generated at build time as static span data from the
Unicode Character Database of the version CharUnicodeInfo implements, by the generator M08 wrote for Script,
with the same test that fails when the two versions part. One generator, each table compiled into the assembly
that reads it (ADR 43): these join
Script in the core, Bidi_Class among them, since this milestone reorders extracted right-to-left text.
The interpreter and the device seam
PdfContentInterpreter (internal) runs a page, a form XObject, a Type 3 glyph procedure or an annotation
appearance through M08's PdfContentReader; iterative, one frame per nested content
IPdfContentDevice (internal until M25, its first outside consumer, needs it public) receives what the
content paints: glyph runs, paths (their segments and paint operator), clips, images
(XObject or inline, with the CTM), shadings, marked-content begin and end, transparency
groups and soft masks, annotation appearances
PdfInterpreterState (internal) a value type on a pooled stack: CTM, clip bounding box, colors and color
spaces, line state, ExtGState entries, text state, the optional-content visibility of
the current section — beside M08's PdfGraphicsState, which the builder writes with
- A page's content streams are one sequence, and an inline image is skipped by its declared length — M08's reader does both; the interpreter never tokenizes a stream on its own.
- Resources resolve by inheritance through M06's page model — the trap
CLAUDE.mdnames —, a form XObject's own resources first and the page's when it has none, as PDF 1.1 allowed. - Optional content is decided by M11's
PdfLayerVisibilityfor an event and a configuration — View under/Dby default — on marked sections/OC … BDC, on XObjects'/OCand on annotations'/OC. Hidden content is interpreted, since its state operators still change what follows (M11's trap), and its paints reach the device flagged as hidden; the extraction options decide whether they are kept. - Type 3 glyphs run in their own frame, the glyph procedure's space mapped by
/FontMatrixand the text rendering matrix;d0andd1set the width and, ford1, the box; a glyph procedure that shows text in another font is followed, one that reaches its own font again is a cycle, cut and reported. - Devices are small: M15 writes a text device (glyphs), an image device (placements), a ruling device (the axis-aligned segments and thin rectangles tables are drawn with), a bounds device (the union of what is painted, for M09's crop), a resource-usage device (for the inventory), and a tracing device used only by the tests, whose log compares with MuPDF's trace (below). The interface carries full path geometry and blend state although M15's devices need only boxes, because M25's does.
- The hot loop allocates nothing (invariant 3): operands are the reader's spans, states are copied by value into a pooled stack, glyphs go into a pooled per-page buffer, and no string exists until a caller asks for one.
Bounds, classified (invariant 12, ADR 34)
| Bound | Kind | Why |
|---|---|---|
| Operations interpreted for one page, its forms, Type 3 glyphs and appearances included | Guard, PdfReaderLimits.MaxContentOperations, limit.content-operations | A valid page can draw a form that draws another form twice, thirty levels deep: a few kilobytes that execute a billion operations. A GIS map legitimately executes millions. The default is set from the heaviest corpus page — the remote USGS map — with a margin, and recorded |
| Glyphs collected for one page | Guard, PdfReaderLimits.MaxPageGlyphs, limit.page-glyphs | Layout analysis holds a page's glyphs; one TJ in a 256 MB stream can show two hundred million of them, validly. The default is set well above the densest corpus page. Past it, the rest of the page's text is returned in content order, unanalyzed, and the page says so |
Graphics state depth — q, and the implicit save of a form's Do | Guard, PdfReaderLimits.MaxGraphicsStateDepth, limit.graphics-state-depth | ISO 32000-2 sets no limit (PDF 1.7's Annex C cited Acrobat's 28 as an implementation limit, and 2.0 dropped it); each level holds a state. Since a form's Do saves a state, the same guard bounds form nesting |
| A form XObject, pattern or Type 3 glyph that reaches itself | Internal: the cycle is cut and reported | Content that paints itself forever is invalid; the visited set is the current chain |
| Type 1 charstrings: 24 operands, 10 nested subroutine calls | Internal constants | The Adobe Type 1 Font Format's own limits; a program past them is invalid |
| Type 1 charstring operations | M08's guard, MaxCharstringOperations | Same reason as for CFF |
usecmap chains and structure-tree walks | No bound needed | Walked iteratively with a visited set; the index bounds the objects reachable |
A CMap's code space, 1 to 4 bytes; a bfrange spanning 2³² codes | Format limit; ranges stored, never expanded | ISO 32000-2 §9.7.6.2; a range costs one entry whatever it spans |
| Samples of an image written as PNG | Checked arithmetic on width × components × bits, under MaxDecodedStreamLength; rows streamed | The dimensions come from the file |
| A regular expression's match time | A search option, PdfSearchOptions.MatchTimeout, reported as search.timeout | Not a reader bound: it limits the caller's pattern on hostile text, and only when the pattern needs backtracking |
The three new guards join PdfReaderLimits, the manifest's readerLimits keys and its schema, and
ReaderLimitsTests, as M08's did.
Fonts read for extraction
A code becomes text by the first of these that answers, per ISO 32000-2 §9.10.2, and the glyph records which answered — the confidence of a page's text is the share of its glyphs mapped by a source that cannot guess:
| Font | Code to Unicode, in order | Widths | Glyph box |
|---|---|---|---|
Simple Type1, MMType1, TrueType | An enclosing ActualText; ToUnicode; the encoding — base encoding, /Differences, else the program's built-in encoding — to a glyph name, then the Adobe Glyph List, uniXXXX, uXXXX[XX], ligature names (f_i), suffixes stripped (a.sc); for a symbolic TrueType, the §9.6.5.4 rules through its (3,0) or (1,0) cmap and post names | /Widths, else the program's (hsbw or hmtx), else M08's AFM metrics for the standard 14 | /Ascent and /Descent, else the /FontBBox, times the size |
Type3 | ActualText; ToUnicode; /Differences names as above | /Widths × /FontMatrix | d1's box, else /FontBBox, through /FontMatrix |
Type0 over CIDFontType0 or 2 | ActualText; ToUnicode; for a predefined ordering (Adobe-Japan1, -GB1, -CNS1, -Korea1, -KR), code to CID through the encoding CMap and CID to Unicode through the collection's map; for Identity over an embedded TrueType, the program's cmap reversed | /W and /DW; vertical /W2 and /DW2 | Descriptor, as above |
- Type 1 programs are read here: the cleartext part for
/Encoding(Standard or an explicit array),/FontMatrixand/FontBBox; the eexec part decrypted (key 55665, then 4330 for charstrings,lenIVread) for glyph names andhsbwwidths./Length1is often wrong: the boundary is found by searching foreexec, and binary or hexadecimal encryption is told apart by the first four bytes, as the format specifies. PFB segment headers inside aFontFile— some producers leave them — are recognized and skipped. ToUnicodeis parsed as a CMap —begincodespacerange,bfchar,bfrangewith a destination string or an array, destinations in UTF-16BE with surrogate pairs and several code points for a ligature,usecmap— and its faults are read through, each reported once per font: a destination byte-swapped (UTF-16LE), detected when the glyph names disagree with most of the map and agree with it swapped; U+0020 mapped to U+0009; a visible glyph mapped to U+00AD, the soft hyphen — the Axapta credit note's minus signs —, returned as U+002D in text, since a soft hyphen that is drawn is a hyphen; destinations that are control characters; a code mapped to nothing.- Unmappable codes — no source answers, or names such as
/g37and/cid123— become U+FFFD in the glyph model, counted per font (extraction.unmappable-codes), and lower the page's text confidence. Below a threshold the page's text isUnreadable: exports carry no text for it and say why, rather than returning what a code table guessed. - Predefined CMaps and the CID-to-Unicode maps come from Adobe's
cmap-resources(BSD-3-Clause, notice inNOTICE), compiled at build time into a compact binary form and Brotli-compressed. They weigh several megabytes as published; where they ship — resources in the core, loaded on first use, or a data satellite the core finds when present — is decided by an ADR in slice 3, on the measured size. If they live outside the core, a CID font whose ordering needs them and finds none earnsextraction.cmap-unavailable, naming the CMap and the package, and its text isUnreadable, never guessed.Identity-HandIdentity-V, which almost every modern producer uses, need no data.
The glyph model
PdfTextExtractor document.Text: ExtractPage(index, options) -> PdfTextPage; Pages(options), lazily
PdfTextExtractionOptions immutable: reading order (Auto, Structure, Layout, Content); optional-content event
and configuration; what is kept — hidden, invisible render modes, outside the crop
box, annotation appearances, furniture; limits
PdfTextPage disposable, its buffers pooled: glyphs, words, lines, blocks; the reading-order
source and confidence; text kind; content box; GetText(); GetTextUnder(quads);
tables on demand; the page's diagnostics
PdfGlyph readonly struct: text (a range of the page's character buffer), code, font, size,
origin, advance, quad, fill and stroke color, render mode, visibility flags, mapping
source, marked content (MCID, artifact, ActualText span, Lang), source location
PdfTextWord, PdfTextLine, PdfTextBlock readonly structs: ranges of glyphs, a quad, a direction
PdfPageTextKind Blank, ImageOnly, VectorOnly, OcrLayer, Unreadable, Text
- Coordinates in the model are the page's default user space, scaled by
/UserUnitand with an inverted/MediaBoxnormalized; exports convert to a visual frame — the crop box, after/Rotate, origin top left — in points (JSON), 1/1200 inch (ALTO) or pixels at a stated resolution (hOCR). - The source location — the operator's index in the page's sequence and the glyph's index in its string —
costs two integers and is what M19 needs to remove a glyph from a
TJ. - Visibility flags are facts, not decisions:
HiddenByOptionalContent;InvisibleRenderMode(modes 3 and 7 — OCR layers are mode 3, and they are the text a scan has);OutsideCropBox;Clipped(the glyph lies wholly outside the clip's bounding box, which over-approximates the clip, so the flag is never wrong);CoveredByImage(an opaque image without mask placed later covers the glyph — FineReader's layer under its page image);FromAnnotation. The default keeps what every referee extracts — render modes 3 and 7 and covered text included, since that is how OCR layers are written — and drops what optional content hides and what lies outside the crop box; each is an option. ActualTexton marked content or a structure element replaces the glyphs it spans, kept as one span, so that the text and the boxes stay tied; control characters in it (the InDesign chapter's) are dropped and reported.
Words, lines and right-to-left text
- Words break at a space glyph — code 32 or a
ToUnicodespace —, at a gap between one glyph's advance and the next glyph's origin wider than a fraction of the font's space width (0.25 em when it has none), at a change of direction or baseline. The thresholds are fixed in slice 4 against the referees and recorded. Word spaces drawn only asTJdisplacements (the Atypon article, Ghostscript's kerning) therefore separate words; a negative displacement never does. - Lines gather glyphs of one writing direction whose baselines agree within a fraction of the size, a superscript or subscript joining the line it rises from. Direction is taken from the text rendering matrix, so a y-down text matrix under a flipping CTM (wkhtmltopdf) reads upright, and rotated text reads along its baseline.
- "Fake bold" — the same glyph drawn two or three times a fraction of a point apart (the GBK report) — is
merged: identical text, font and size within 0.1 em on one line keep one glyph, counted in
extraction.duplicate-glyphs-merged. - Right to left: within a line, glyphs are taken in position order, and maximal runs of strong
right-to-left characters (
Bidi_ClassR and AL), with the neutrals between them, are reversed into logical order, digit runs inside them keeping their own; the line's direction is its first strong character's. It is the inverse of display ordering, not UAX #9 itself — which ADR 43 places inAdCodicem.Pdf.Html—, and needs only theBidi_Classtable generated in the core; it depends only on positions, so a producer that wrote its glyphs in logical order and one that wrote them in visual order read alike. Arabic presentation forms are kept in the text and folded by search. - Vertical writing —
Identity-V, a-VCMap,/W2— makes vertical lines, read top to bottom, columns right to left. - Hyphens at a line end stay in the text; the line records that it ends in one, whether U+002D, U+2010 or
U+00AD, and whether the next line starts in lower case — which is what search and the plain-text option
JoinHyphenatedWordsuse. A soft hyphen that is not drawn is never text.
The structure tree read model
PdfStructureTree document.Structure: IsTagged (/MarkInfo /Marked and a root), the root's kids, role map,
class map, PDF 2.0 namespaces (read), ID tree, parent tree lookup by page and MCID
PdfStructureNode type as written, standard type through the role map (PdfStructureType, M13's), attributes
(PdfStructureAttributes, M13's), Alt, ActualText, E, Lang, T, ID, page, kids —
nodes, marked-content references (page, MCID, /Stm), object references — enumerated
lazily and iteratively
M13's writer calls its handle PdfStructureElement; the read side takes other names so that the two never
meet in one signature. A role map that cycles, a kid reached twice, an MCID claimed by two elements, a page
whose /StructParents has no entry, /StructParents without a tree — each is reported and read around.
Reading order
- Tagged (the default when
IsTaggedand the tree resolves): content in the depth-first order of the structure's content references;ActualTextandAltas the element gives them; artifacts out of the text, pagination artifacts kept as furniture; annotations placed where theirOBJRsits. Content that is neither referenced nor an artifact — the untagged pieces of a tagged document — is placed by the layout analysis and lowers the page's confidence (extraction.untagged-content). A structure that does not cover at least the page's visible glyphs, by a threshold set in slice 5, falls back to the layout, reported. - Layout: lines into blocks by spacing, overlap and size; blocks ordered by a recursive XY-cut (Nagy and Seth, 1984) — cut across the widest full gap, bands top to bottom, columns left to right, or right to left when the page's text is mostly right to left —, iterative, with ties broken by content order, so the result is deterministic. The confidence falls with gaps barely wider than the threshold, blocks that straddle a gutter, and cuts that disagree with content order.
- Content order, on request, for callers who want the producer's order.
Furniture — running headers and footers — is recognized across pages: a line in the top or bottom band of the page whose text, digits folded, recurs at the same place on most pages of a window of eight; pagination artifacts, M09's marks among them, are furniture with certainty. Exports leave furniture out by default and keep it on request.
Page text kind: Blank; ImageOnly (no glyph, images over at least half the crop box, measured on a
32 × 32 grid); VectorOnly (no glyph, paths only — the poster whose text is outlines); OcrLayer (every glyph
invisible or covered, over an image of the page); Unreadable (the text confidence below its threshold);
Text. The manifest's hasExtractableText: false is Blank, ImageOnly or VectorOnly, and those pages
return no text at all.
Tables
PdfTableDetector Detect(page, options) -> PdfTable[]; tables continued across pages linked
PdfTable page, box, rows, columns, cells (row, column, spans, header flag, glyph range, text),
Source (Tagged, Ruled, Aligned), Confidence in [0, 1], Evidence, continuation of
- Tagged:
Table,THead,TBody,TFoot,TR,TH,TD, withRowSpan,ColSpan,ScopeandHeaders. Confidence 1 when every row's spans fill the same number of columns; less, with the rows that do not, otherwise. - Ruled: horizontal and vertical segments — strokes, and filled rectangles thinner than a threshold — from the ruling device, snapped and merged into a grid; cells are the grid's, merged where no rule separates them.
- Aligned: columns from the left, right and center edges of words that line up across three lines or more, numbers right-aligned; a header row by weight, fill or a rule beneath.
- Continuations: a table at the top of a page whose header row repeats the previous page's last table's, text and columns, is its continuation; Markdown and JSON write them as one table with each row's page.
- Confidence is calibrated, not asserted (ADR 15): slice 7 measures it on the corpus tables whose truth is
known and records the calibration in
status.md; a caller never gets a grid without the number.
Search
PdfTextSearch document.Search(query, options) -> IEnumerable<PdfSearchHit>, page by page, lazily
PdfSearchQuery Literal(text) | Regex(pattern, RegexOptions)
PdfSearchOptions folding (case, diacritics, compatibility — ligatures, presentation forms, full-width, the
spaces of U+00A0, U+202F and U+2000–U+200A —, punctuation), across line breaks, whole words,
pages, what is searched (hidden, annotations, furniture), MatchTimeout, MaxHits
PdfSearchHit page, page label, quads (one per line fragment, in the order M11 writes), text as found,
structure node, glyph range, context before and after
- Each page's text is folded once into a pooled buffer with a map from folded offset to glyph index, so a hit maps back to glyphs and quads; no allocation per glyph, one per hit.
- Across line breaks: a line break folds to a space; a hyphen at a line end followed by a lower-case letter is
optional to a literal query —
réglementationfindsréglemen-⏎tation, ande-mailstill findse-⏎mail—, and dropped from the text a regular expression runs over. The difference is documented. - Regular expressions run with
RegexOptions.NonBacktrackingby default, whose time is linear in the text; a pattern that needs backreferences or lookaround runs with backtracking underMatchTimeout(two seconds a page by default), and a page that times out is skipped and reported (search.timeout), never retried. - Quads come from glyph boxes: the font's ascent and descent, or a Type 3 glyph's box, through the text rendering matrix — rotated and vertical text get rotated and vertical quads.
Images
PdfImageExtractor List(document) -> PdfImagePlacement rows; Write(placement, sink)
PdfImagePlacement page, object (or inline), width, height, color space, components, bits, filters,
interpolation, image mask, soft mask, effective resolution from the CTM, encoded size
- The listing is
pdfimages -list's, row for row, so that it can be held to it. - Extraction chooses
pdfimages -all's file types: DCT as.jpgand JPX as.jp2, byte for byte; JBIG2 as.jb2ewith its globals as.jb2g; CCITT as.ccittwith its parameters as.params, or a single-strip TIFF wrapping the unchanged data, on request; everything else decoded — Flate, LZW, RunLength, uncompressed — to PNG for gray, RGB and indexed samples at 1 to 16 bits, and TIFF for CMYK. Masks and soft masks are their own images. Lab, Separation and DeviceN samples are written as their components and reported, until M22 evaluates color spaces. PNG rows are written as they are decoded; the CRC-32 is ours (the BCL's is not in-box), and the deflater is M03's, so the bytes are as deterministic as the writer's.
Exports and chunks
PdfExport ToText, ToMarkdown, ToJson, ToAlto, ToHocr — into an IBufferWriter<byte> or a Stream,
page by page
PdfChunker Chunks(document, PdfChunkingOptions) -> IEnumerable<PdfChunk>
PdfChunk deterministic id, text or Markdown, anchors: pages, page labels, Bates numbers, section
path, boxes per page; PdfAnchorProvider lets a caller add its own (M18's piece numbers)
- Plain text: reading order, pages separated by a form feed as
pdftotextdoes; aLayoutoption keeps columns apart with spaces, best effort. - Markdown: CommonMark with GFM tables. Headings from
H,H1–H6and role-mapped types, else from type size and weight clusters where their confidence clears a threshold; lists fromL,LI,LblandLBody, else from bullets and enumerators with a hanging indent; tables as GFM when rectangular, as an HTML block when cells span; figures aswhen images are requested; links from link annotations. Text from the file is escaped so that it cannot make structure — a line that reads# heading, a|in a cell, a<script>in a paragraph stay text. Furniture is left out; page boundaries are comments carrying the anchors. - JSON: the whole model at a chosen detail — blocks, lines, words or glyphs —, with document metadata, the outline in M07's JSON form, attachments, page labels, tables, images and annotations; its schema is versioned and published on the documentation site beside M07's outline schema; numbers formatted invariant, with fixed decimals.
- ALTO v4:
MeasurementUnitinch1200by default (orpixelat a stated resolution);Page,PrintSpace,TextBlock,TextLine,StringwithCONTENT, position,WCandSTYLEREFS,SP,HYP; tables asComposedBlockwithTYPE="table". - hOCR 1.2: XHTML with
ocr_page(bbox,ppageno,scan_res),ocr_carea,ocr_par,ocr_line(bbox,baseline,x_size),ocrx_word(bbox,x_wconf,x_font,x_fsize). ALTO, hOCR and TSV are the formats M22's text-layer writer reads, chosen so that the two milestones meet. - Chunking: by page or by section (tagged headings, else the outline, else the heuristic headings), up to a
size in characters with an overlap in words — no tokenizer, which would be a dependency and a model. The
Bates anchor comes from M09's
PdfBatesRangeMapwhen the caller has one, else from pagination artifacts of subtype/Bates(PDF 2.0) and from M09's own marks, else from a caller's pattern over the footer band. The page label is M06's.
The inventory
PdfInventory.Build(document, options) -> PdfInventoryReport, stable JSON, lazy, bounded:
| Section | Content | Held to |
|---|---|---|
| Document | Header and catalog versions, /Extensions, producer and creator (Info and XMP), linearization, revisions and signatures with their coverage (M04), tagging, conformance claims (pdfaid, pdfuaid, pdfx, fx), output intents | qpdf --json, XMP, veraPDF's feature report |
| Pages | Count, distinct boxes, rotations, UserUnit, text kind per page | pdfinfo -box |
| Fonts | As pdffonts: name, type, encoding, embedded, subset, ToUnicode, object; plus glyphs used and codes unmappable | pdffonts, M08's fonts verb |
| Images | Placements with effective resolution, as pdfimages -list | pdfimages -list |
| Color | Color spaces in use, ICC profiles (header and description), separations and DeviceN colorants, patterns, shadings | veraPDF's feature report |
| Interactive | Annotations by subtype (M11), links and their targets, AcroForm fields by type, XFA present, usage rights present, signature fields signed and unsigned | M11's listing, pikepdf |
| Attachments and layers | M06's attachments with their relationships; M11's groups and default states; a portfolio's /Collection | pdfdetach -list, M11 |
| Active content (ADR 37) | Document JavaScript, /OpenAction, additional actions at every level, JavaScript, Launch, SubmitForm, ImportData, GoToR and GoToE actions with their targets, URIs by host, rich media, 3D, movie, sound and screen annotations — each located, none run | pdfinfo -js, a pikepdf walk |
| Space | Bytes by category — content, fonts, images, form XObjects, structure, metadata, annotations and fields, attachments, signatures, cross-reference overhead, unreferenced objects, earlier revisions' superseded objects — from the index's offsets and a reachability walk, decoding nothing but object streams; the categories sum to the file's length | A pikepdf script of the test support, written independently |
The inventory states facts; judging them is M02's rules, M19's action.* and hidden.* families, and M20's
profiles.
Validation rules
M02's structural profile gains the rules its "Resources" row and the iPRES content cases waited for, in the
content and resource families (disjoint from diagnostic codes, which ValidationRuleIdTests checks — M08's
content.* diagnostics and M12's resource.* ones among them). Each iPRES content-operator case gets its finding
or the recorded reason the profile stays silent, as M02 did for the end-of-file cases.
| Rule | Checks |
|---|---|
content.operator-outside-object | A text-showing or text-state operator outside BT…ET, a path operator inside it, BT nested |
content.object-unclosed | A text object or marked section still open at the end of the page's content |
content.operand-invalid | An operator given operands of the wrong number or type — Tf without a name, a size that is a lone dot, a string written as an array for Tj |
content.operands-unused | Operands left pending at an operator that takes none, or at the end of the content |
content.font-not-set | Text shown before any Tf in the text state |
content.self-reference | A form XObject, pattern or Type 3 glyph that reaches itself |
resource.name-undefined | A name the content uses — font, XObject, ExtGState, color space, pattern, shading, properties — that the effective resources lack |
Severities follow M02's bar (ADR 45) — Error only where the reader cannot vouch that it reads the page as
written, established with the referees' outputs — and are fixed per rule by slice 12. Every well-formed corpus document stays free of
error findings.
Diagnostics
In PdfDiagnosticCodes, disjoint from validation rule identifiers (ADR 36):
| Code | Severity | Meaning |
|---|---|---|
content.cycle-cut | Warning | A form, pattern or Type 3 glyph reached itself; the inner occurrence was skipped |
content.operator-misplaced, content.operand-mismatch | Warning | The interpreter read around an operator out of place or its operands |
extraction.unmappable-codes | Warning | Codes no source maps, per font, with counts |
extraction.tounicode-suspect | Warning | A ToUnicode byte-swapped, mapping spaces to tabs, drawn glyphs to soft hyphens, or codes to controls; read through |
extraction.cmap-unavailable | Warning | A predefined CMap the ordering needs is not present |
extraction.structure-incomplete | Information, or Warning when the order falls back | The structure tree does not cover the page's content, or cannot be followed |
extraction.untagged-content | Information | Content of a tagged document neither referenced nor an artifact |
extraction.duplicate-glyphs-merged | Information | Overprinted glyphs merged |
extraction.text-unreadable | Warning | A page's text confidence is below the threshold; no text is returned for it |
extraction.image-dimensions-mismatch | Warning | An image's data disagrees with its declared size; what exists is written |
font-program.* | as M08's | Faults in Type 1 programs, M08's family |
search.timeout | Warning | A backtracking pattern timed out on a page, which was skipped |
limit.content-operations, limit.page-glyphs, limit.graphics-state-depth | Warning | The guards above; the message names the PdfReaderLimits property |
Outline generation and comment summaries
PdfOutlineGenerator.FromStructure(document)makes one bookmark per heading element;FromLayout(document)one per heading the size-and-weight clustering finds, each with its confidence, the caller choosing a floor. Both return M07'sPdfOutline, targets at the heading's page and top, for M07's API to apply.PdfCommentSummary.Build(document): every markup annotation with its page, subtype, author, dates, contents, replies and the text under its quads (GetTextUnder), in JSON and Markdown.
The command-line tool
text FILE [--pages SEL] [--format text|json|alto|hocr] [--layout] [--order auto|structure|layout|content],
markdown FILE [--chunk page|section] [--max-chars N], search FILE QUERY [--regex] [--fold …] [--json],
inventory FILE [--json] [--no-content] [--images DIR] — the last writing every listed image as pdfimages -all
would, since M07 already gave images its meaning (image files to pages). Each takes M06's page selection and exit
codes, and the AOT binary produces what the API produces.
Slices
Each slice ends on a green commit, with the codes, rules and guards it introduces documented and its benchmark,
if it has one, recorded in docs/status.md.
- The interpreter and the device seam. Delivers
PdfContentInterpreter,PdfInterpreterState,IPdfContentDevice, forms with inherited resources, inline images, marked content, optional content through M11's evaluator, annotation appearances by their matrices, the bounds device (a page's content box, which M09's crop now calls), the three guards with their classification, thecontent.*diagnostics. Proved by unit tests per operator and state (nested forms, a form drawing itself, a form drawing another twice thirty levels deep, 10⁶ nestedq, a pattern in a pattern, hidden sections that still move the text position); an FsCheck property — for any sequence M08's builder writes, the CTM and text matrix at each paint equal those of a naive recursive interpreter in the test support —; integration: the tracing device's image placements and text-run matrices equalmutool draw -F trace's on every committed page MuPDF reads, and the content box equals the union MuPDF's bounding-box output (mutool draw -F bbox) reports within a point;InterpreterBenchmarksat 0 B per operator. Leaves text. - Simple fonts, encodings and Type 1. Delivers base encodings and
/Differences(M08's tables), the glyph name rules, the symbolic TrueType rules, the Type 1 reader,ToUnicodewith its fault handling, the mapping source per glyph. Proved by unit tests per rule of §9.6.5 and §9.10.2, every Adobe Glyph List form, abfrangecrossing a byte boundary, arrays, surrogates,usecmapcycles, eachToUnicodefault, a/Length1that lies, PFB headers insideFontFile; FsCheck — any byte sequence is a CMap or a diagnostic, and a range lookup equals the naive expansion on small ranges —; integration: on every committed page, the multiset of characters each font yields equals poppler's and PyMuPDF's where they agree — a comparison blind to order, which isolates mapping from reading order —; the Type 1 programs' glyph names and widths equal fontTools't1Liband FreeType's. Leaves Type 3 and CID fonts. - Type 3, CID fonts and the predefined CMaps. Delivers Type 3 glyph procedures with
d0,d1and their matrices, embedded CMaps,Identity-Hand-V,/Wand/W2,CIDToGIDMap, the ADR on the CMap data and its packaging as decided, the CID-to-Unicode maps. Proved by unit tests (a Type 3 font with a rotated/FontMatrix, a glyph procedure that shows text,handwritten-type3-recursion.pdf, a-VCMap, a code space of mixed lengths); the character multisets of the Type 3, CID-keyed CFF and CJK documents below against MuPDF, which ships the CMaps, andpdftotextwithpoppler-data; the package-size measurement recorded. Leaves positions. - Glyphs, words and lines. Delivers
PdfGlyphwith boxes and visibility flags,ActualTextspans, word and line building, duplicate merging, right-to-left ordering with the generatedBidi_Classtable, vertical lines, hyphen facts, the frame conversions, the tolerances. Addsexpect.hiddenTextto the manifest and its schema. Proved by unit tests (each visibility case, a superscript,Tzat 2000 %, a y-down text matrix, a word split across twoTJ, logical order from visual and from logical input); FsCheck — for any text M08's simple path lays out with random gaps either side of the threshold, the words recovered are the words written —; integration: word boxes againstpdftotext -bbox-layoutand PyMuPDF's words, at least 98 % of words matched with an intersection over union of 0.9, the rest listed and explained; the right-to-left rows below. Leaves order. - The structure tree and tagged reading order. Delivers
PdfStructureTreeandPdfStructureNode, the MCID map through pages and XObjects,OBJR,ActualTextandAltfrom elements, artifacts, the fallback rule. Proved by unit tests (a role map cycle, a kid reached twice, a duplicated MCID,/StructParentswithout a tree, an MCR with/Stm); integration: each element's text equalspdfinfo -struct-text's and a pikepdf walk's on the 66 tagged committed documents that open; the tagged rows below. Leaves untagged order. - Layout, reading order and page kinds. Delivers blocks, the XY-cut with its confidence, furniture, page text
kinds and
Unreadable. Addsexpect.readingOrder(fragments that must appear in this order) andexpect.textlessPages— the 1-based numbers of the pages without extractable text, for a document whose pages differ, such as the remote Konica scan whose blank page sits among OCR-layered ones;hasExtractableText: falsestill means every page, and a scanned page with an OCR layer, as in Cerfa 12156, has text — to the manifest and its schema. Proved by FsCheck — for multi-column pages M08's builder generates with gutters above the threshold the order is recovered, and the confidence does not rise as the gutter narrows —; integration: everytextContainsandhasExtractableTextrow below; the untagged twin of the two-column report in its source's order. Leaves tables. - Tables. Delivers
PdfTableDetectorin its three modes, continuations, the calibration. Addsexpect.tables(grid and header texts) for the documents whose truth is known. Proved by unit tests (spans, a table without rules, rules without a table — a form's boxes —, a header repeated on the next page); FsCheck — ruled grids with random spans generated by M08's builder come back exactly —; integration: pdfplumber'sextract_tableson the ruled tables, the structure on the tagged ones, the source on ours. Leaves search. - Search. Delivers
PdfTextSearchwith the generated folding tables, the offset map, literal and regular queries, line-break handling, quads. Proved by FsCheck — folding is idempotent; any string placed by M08's builder across a line break, hyphenated or not, in any case and with or without its diacritics, is found with quads covering exactly its glyphs —; a catastrophic pattern ((a+)+$over a page ofa) returning within its budget; integration: everytextContainsstring found where PyMuPDF'ssearch_forfinds it; one of M11's highlights authored on each hit, covering it in pdf.js's rendering (M11's harness). Leaves images. - Received hybrids: visible values against the XML. Delivers, in
AdCodicem.Pdf.FacturX, M14's visible-consistency matcher run over this milestone's search on a received hybrid: the invoice number (BT-1), its date (BT-2), the total with VAT (BT-112) and the amount due (BT-115) looked for on its pages in the forms its language writes them, each not found reported (facturx.visible-value-not-found, warning, naming the business term), each found with its page and quads;facturx validateruns it on a PDF. Proved by unit tests over hand-made pages (a total split across two text runs, a date written in words, a value repeated in a footer); integration: the rows below, each value located where PyMuPDF'ssearch_forfinds it. Leaves images. - Images, metadata, outline, attachments and annotations' text. Delivers
PdfImageExtractor, the PNG and TIFF writers,GetTextUnder,PdfCommentSummary,PdfOutlineGenerator, and the composition of M06, M07 and M14's models into the extraction result. Addsexpect.images(rows ofpdfimages -list), written bybuild_corpus.py. Proved by unit tests (a mask, a soft mask, a 1-bit indexed image, a/Decodearray, declared dimensions larger than the data); integration:pdfimages -listand-all,pdfdetach, LibreOffice's own outline, PyMuPDF's words under the Reader X highlight. Leaves exports. - Exports and chunks. Delivers
PdfExport, the JSON schema,PdfChunker, anchors andPdfAnchorProvider. Proved by unit tests (escaping every Markdown construct; an FsCheck property that any string written as text parses back through a CommonMark reference as one text node; JSON round trip); integration:cmark-gfm's AST, Python'sjsonschema,xmllintagainst the Library of Congress's ALTO 4 schema, hOCR tools'hocr-check; anchors on a Bates-numbered set made by M09. Leaves the inventory. - Inventory and content rules. Delivers
PdfInventory, the space accounting, thecontent.*andresource.name-undefinedrules with their severities. Proved by unit tests per section and per rule (each of the content faults in a hand-built page); integration: each inventory section against the tool named in its row above; the iPRES content cases' findings. Leaves the tool and the budgets. - The tool, budgets and the whole. Delivers the four verbs;
ExtractionBenchmarks,SearchBenchmarksandExportBenchmarkswithMemoryDiagnoser; memory checkpoints at 10, 100 and 1,000 pages; a text row in the comparison benchmarks against PdfPig's word extraction and iText's location strategy, recorded, not asserted; the documentation. Proved byCorpusToolTests, the budget rows, determinism across every export, and a greenRemote corpusrun recorded instatus.md.
Tests required
Unit — tests/AdCodicem.Pdf.Tests:
- The interpreter: every operator's effect on the state; forms, patterns and Type 3 glyphs nested and cyclic;
hidden sections; annotation appearances under every
/Matrix; each guard reached, raised and thrown, asReaderLimitsTestsdoes for the reader's. - Fonts: every rule of the mapping table; every
ToUnicodefault; Type 1 programs with each encryption, a lying/Length1, a charstring past its limits; Type 3 with every matrix; CID fonts with each CMap form, vertical metrics, aCIDToGIDMapstream. - Glyphs and lines: every visibility flag; rotation,
/UserUnit, an inverted media box,Tz,Ts,Tc,Twon one- and two-byte codes; right to left with digits and neutrals; vertical text; duplicated glyphs; hyphens. - Structure: every fault named in slice 5; PDF 2.0 namespaces read; a tree of a million nodes walked in bounded memory.
- Tables, search, exports and the inventory: as their slices name, degenerate inputs included — an empty page, a page of one glyph, a table of one cell, a search for the empty string, a document with no text at all.
- Hostile: a page drawing a form 2³⁰ times, a
TJof 10⁸ glyphs, aToUnicodeof a million ranges, a structure tree whose kids loop, a Markdown injection in every construct, a pattern built to backtrack — each ends in a result, a report or a typed exception within time and allocation budgets. The interpreter, the CMap parser, the Type 1 reader and the glyph-name mapper join the nightly fuzzing campaign, seeded with the corpus's content streams and font programs. - Determinism: every export of every committed document, twice, byte for byte; extraction under a shuffled font cache and a different culture gives the same bytes.
Integration — tests/AdCodicem.Pdf.IntegrationTests, every referee in a container
(ADR 27):
- poppler —
pdftotext(default,-layout,-bbox-layout,-cropbox),pdfinfo -struct-text,-js,-box,pdffonts,pdfimages -listand-all,pdfdetach, withpoppler-datafor the CMaps; - MuPDF —
mutool draw -F stextand-F trace, which ships the CMaps; PyMuPDF — words, characters,search_forwith quads, image listing, text in a rectangle; - pdfplumber — characters and
extract_tableswith its lines and text strategies; - pikepdf — structure walks, content parsing, the space accounting script, active content walks;
- FreeType and fontTools — Type 1 and CFF programs, in M08's Python image;
- veraPDF — its feature report, for the inventory's fonts, color spaces, embedded files and claims;
- qpdf —
--jsonand--show-xref, for the inventory's document section and space accounting; - cmark-gfm, xmllint with the ALTO 4 schema, hocr-tools and Python's jsonschema — the exports.
Where the referees disagree a row passes when we agree with the reading the manifest records and the reason
it records for the others — the corpus already knows that poppler reverses two letters of the IRS Arabic
edition's "الضرائب" where xpdf does not, and that xpdf 4.00 extracts nothing from the compacted-syntax file where
poppler and PDFium do. Disagreeing with every referee fails. The tolerances — positions within a point, word
boxes at an intersection over union of 0.9, quads within 1.5 points on each corner — are fixed by slices 4 and 8
against the referees and recorded in status.md, with how PyMuPDF's and poppler's frames are converted into ours
on the rotated corpus pages.
Acceptance conditions
"Every committed document" means every document under documents/ and vendor/ the reader opens: the eighteen
encrypted ones wait for M16, except where a row names a decrypted twin.
| Documents | Behavior | Verified by |
|---|---|---|
Every committed document with textContains — 109 that open, from every producer in the corpus —, and the 179 remote ones | Every string found in the extracted text, whitespace folded; expect.hiddenText absent from it; no untyped exception | CorpusExtractionTests.Extracted_text_contains_what_the_manifest_declares |
Every document with hasExtractableText: false — 14 committed: documents/scan/reportlab-scanned-receipt.pdf, the Xerox JBIG2 scan vendor/us-federal/xerox-workcentre-treasury-imf-report-scan.pdf, the 1999 CCITT import acrobat3-import-irs-1040-1988-scan.pdf, vendor/pikepdf/scanner-ccitt-endofline.pdf, ImageMagick's JPX page, Acrobat 11's Image Conversion page, Word's text drawn as images vendor/fr-licence-ouverte/word365-dila-text-drawn-as-images.pdf, make-pdf-javascript-openaction.pdf, the hand-written hostile files —; 45 remote, among them the Print to PDF poster whose text is outlines (remote/govuk/print-to-pdf-border-force-poster-updated.pdf) | Every page Blank, ImageOnly or VectorOnly; no text returned, not a character; the page's image coverage stated | CorpusExtractionTests.Documents_without_text_report_none_rather_than_noise |
OCR layers: vendor/us-federal/hp-mfp-acrobat-ocr-nih-report.pdf (Acrobat), xerox-workcentre-5335-ocr-hud-fonsi-linearized.pdf and xerox-workcentre-5755-ocr-hud-fonsi-mrc.pdf (the copiers' own, render mode 3, horizontal scaling up to 2000 %), finereader8-frb-sr0115-examiner-guidance.pdf (text under the page image), vendor/fr-licence-ouverte/pdfmaker-acrobat-cerfa-12156-form.pdf (one scanned page with live OCR text); remote, the Konica, Ricoh, Canon and OmniPage scans | OcrLayer pages, glyphs flagged invisible or covered and extracted; textContains met; word boxes as poppler's within the tolerance despite the scaling | CorpusExtractionTests.Ocr_layers_are_extracted_and_marked |
documents/report/chromium-report-fr.pdf and libreoffice-report-fr.pdf (tagged), and the untagged twin of the Chromium report — not in the corpus (below); remote, the two-column Astrophysical Journal article remote/opf-format-corpus/jhove-hul-129-latex-distiller705-journal-article.pdf | The two-column annex's four paragraphs in the source's order, from the tags on the tagged files and from the layout on the twin, the twin's confidence above the threshold slice 6 sets; the article's expect.readingOrder fragments in order | CorpusReadingOrderTests.Multi_column_pages_read_in_column_order |
| The 66 tagged committed documents that open — Word, LibreOffice, Chromium, InDesign, PDFlib, Antenna House, PDFMaker — and remote, the 82-page AbleDocs PDF/UA-1 scan whose OCR text lies under its page images | Reading order is the structure's; each element's text equals pdfinfo -struct-text's; artifacts are out of the text and kept as furniture; ActualText replaces what it spans | CorpusStructureTests.Tagged_documents_read_in_structure_order |
Layers: vendor/uk-ogl/pdfmaker21-ozev-sample-invoice.pdf (Acrobat's Watermark, on), vendor/opf-format-corpus/pdfmaker707-word-va-esig-developer-guide.pdf (PDFMaker's HeaderFooter, on), vendor/pdf-association/handwritten-utf16le-strings.pdf (two groups off, drawing boxes); a document with text in a group off by default and in a group shown only when printing — not in the corpus (below); remote, remote/pdf20examples/handwritten-pdf20-utf8-strings.pdf and the USGS map's 31 layers | Text under a group hidden for View is absent by default and present, flagged, on request; for the Print event, the print-only text appears; the watermark "Sample" and the PDFMaker headers are extracted — the headers as furniture; the hidden boxes add nothing to the content box; each verdict agrees with pdf.js's display and print renderings through M11's harness | CorpusExtractionTests.Hidden_layers_are_honored |
Fonts of every kind: Type 3 in vendor/opf-format-corpus/distiller3-dea-cfr-damaged.pdf and vendor/verapdf/pdfa2b-content-pass.pdf; Type 1 in pdfmaker707-word-law-library-iraq-legal-history.pdf (five subsets) and the DEA file (AvantGarde-Demi); bare CFF from Distiller 4 to 9.5 and InDesign; CID-keyed CFF in indesign-irs-pub1-chinese-traditional.pdf and the Cerfa's OCR fonts; no ToUnicode in illustrator-irs-pub1-english.pdf (with unmapped ligatures), quartz-word-mac2011-lorem-ipsum.pdf (MacRoman TrueType), pdfmaker9-word-distiller-pdfa1b-test-document.pdf, pdfwriter4-usda-dry-whey-standard-2000.pdf, openpdf-jasperreports-financial-statement.pdf; a non-embedded Type 1 with /Differences and no base (distiller3-mac-msha-crusher-dust-card-1997.pdf); remote, the AFP bank statement's and the z/OS card statement's bitmap Type 3 fonts, Google Docs' Type 3, the 1994 Distiller 1.0.2 report, pdfTeX's Type 1 subsets, Ghostscript's Type1C | textContains met; each font's character multiset equals the referees' where they agree; ligatures expanded (fi is fi); the seventh Type 1 program M08 counted, inside the RC4-encrypted vendor/us-federal/distiller2-irs-9465-1996-rc4-40.pdf, read once M16 opens it | CorpusFontReadingTests.Every_font_kind_extracts_as_the_referees_read_it |
Scripts: Arabic vendor/us-federal/indesign-irs-pub1-arabic.pdf, Russian and Traditional Chinese editions, distiller6-cdc-west-nile-chinese-traditional.pdf, the Greek vendor/eu-publications/pdflib-oj-exchange-rates-greek.pdf; the decrypted twin of the vertical Japanese vendor/jp-nta/indesign-distiller18-nta-gift-tax-vertical.pdf — not in the corpus (below); remote, WeasyPrint's Arabic, the Census and USDA Hebrew | Logical order; the Arabic edition's "الضرائب" as written, where poppler reverses two letters; vertical Japanese read top to bottom, columns right to left, through the Adobe-Japan1 map; textContains met | CorpusScriptTests.Right_to_left_vertical_and_cjk_text_reads_in_logical_order |
Producer faults: PDF24's ToUnicode that maps a glyph never drawn (documents/invoice/pdf24-invoice-fr.pdf), the InDesign chapter's control character in ActualText, handwritten-compacted-syntax.pdf, handwritten-content-stream-indirect-refs.pdf, handwritten-inline-image-abbreviations.pdf, the damaged pages that end inside BT; remote, Photoshop's byte-swapped ToUnicode, wkhtmltopdf's spaces mapped to tabs and y-down matrix, Axapta's soft-hyphen minus signs, the GBK report's duplicated runs, Atypon's word spaces by position, Ghostscript's spaces as kerning, Symtrax's and 4s4u's letter-spaced titles, typeset.sh's text before Tf | Each read through, with its diagnostic where the fault is detectable; textContains met; Axapta's amounts carry their minus signs; duplicated runs appear once | CorpusExtractionTests.Producer_faults_are_read_through_and_reported |
The invoice through five producers — documents/invoice/chromium-invoice-fr.pdf, libreoffice-invoice-fr.pdf, word-invoice-fr.pdf (tagged), word-print-driver-invoice-fr.pdf, pdf24-invoice-fr.pdf (untagged) —, the five damaged copies of the Chromium one, and ReportLab's reportlab-invoice.pdf (ruled, untagged) | The five line items with their five columns — Désignation, Quantité, Prix unitaire, TVA, Montant HT — as sources/invoice-fr.html has them, and ReportLab's four items in four columns as build_corpus.py writes them; confidence 1 from the tags, and from the rules or the layout the confidence the calibration states; the damaged copies extract as their original, the truncated tail excepted, whose loss M05 reports | CorpusTableTests.The_invoice_line_items_are_found_with_their_columns |
Other tables: vendor/jasper-modular/openpdf-jasperreports-financial-statement.pdf (a vector grid), vendor/pdf-association/pdflib-pps-kraxi-pdfa2a-pdfua1-invoice.pdf (tagged, THead and TBody), LibreOffice's report's "Tableau des mesures" (libreoffice-report-fr.pdf: its header row repeated on page 3, repeated-table-headers — Chromium's rendering fits the table on one page), Excel's printed tables print-to-pdf-excel-dod-fcf-rates-2021.pdf and print-to-pdf-excel-agec2025.pdf, the EU exchange-rate notices; remote, the central bank's seven-column bulletin, the AFP statement's wrapped header, the Census abstract's tables | Ruled tables cell for cell as pdfplumber's lines strategy finds them; tagged ones as their structure; the report's table one table across its pages; expect.tables met; every table's confidence within the calibration | CorpusTableTests.Tables_are_detected_with_the_confidence_they_state |
Every textContains string of every document, committed and remote; the hyphenated breaks of a document not in the corpus (below); the invoice's "etude noailles" and "TOTAL TTC" | Found on the page and at the quads PyMuPDF's search_for reports, within the tolerance; across hyphenated line breaks; folded as asked | CorpusSearchTests.Search_finds_what_pymupdf_finds_where_it_finds_it |
documents/invoice/qpdf-invoice-with-facturx-xml.pdf and every document with expect.attachments — nine committed, among them the Mustang, weclapp, Docentric and factur-x Python invoices and BFO's embedded PDF/A; twelve remote | Names as declared; bytes equal to what pdfdetach extracts, and for the derived invoice to the XML build_corpus.py embedded; read streamed | CorpusExtractionTests.Attachments_are_extracted_byte_for_byte |
| The 416 image XObjects of the 48 committed documents that open with images — Flate, uncompressed, CCITT, DCT, JBIG2, JPX — and their inline images | Rows equal pdfimages -list's — size, color, components, bits, encoding, resolution within 1 ppi; passed-through files equal pdfimages -all's byte for byte; PNGs decode (Pillow, in the container) to the samples its files hold; expect.images met | CorpusImageTests.Images_are_listed_and_extracted_as_pdfimages_does |
The tagged committed documents; libreoffice-report-fr.pdf | Markdown whose cmark-gfm tree has the structure's headings, lists and tables in order; JSON valid under the published schema; ALTO valid under the ALTO 4 schema and hOCR passing hocr-check, both carrying the model's words and boxes; an outline generated from the report's headings equal to the one LibreOffice wrote | CorpusExportTests.Exports_are_valid_and_carry_the_model, CorpusOutlineTests.Generated_outlines_match_the_headings |
| A Bates-numbered set made by M09 from corpus documents; M06's documents with page labels | Every chunk's anchors equal M09's range map and M06's labels; the chunks' text, concatenated without overlap, equals the document's text | CorpusExportTests.Chunks_carry_page_label_and_bates_anchors |
vendor/opf-format-corpus/reader10-openoffice32-annotated-object-streams.pdf; the widget appearances of vendor/fr-licence-ouverte/fop-dictao-dila-signed-joafe-notice.pdf; remote, the Foxit signatures whose text is only in their appearance | The highlight's marked text equals PyMuPDF's words under its quads; the comment summary quotes it; appearance text found as poppler finds it, flagged FromAnnotation | CorpusAnnotationTextTests.Text_under_and_inside_annotations_is_found |
| Every committed document that opens | The inventory's fonts equal pdffonts, images pdfimages -list, attachments pdfdetach -list, annotations M11's listing, layers M11's model, JavaScript and actions pdfinfo -js and a pikepdf walk (the Didier Stevens OpenAction, the USCIS I-9's scripts, the XFA forms), claims the XMP and veraPDF's feature report; the space categories sum to the file's length and each equals the independent script's | CorpusInventoryTests.The_inventory_agrees_with_independent_tools |
Remote, the 18 iPRES 2017 content-stream cases (remote/ipres2017/t02-05-01-*); every well-formed corpus document | Each operator case its content.* or resource.* finding, or the recorded reason for silence (a missing cm is legal); expect.findings extended and met; no error finding on a well-formed document | CorpusValidationTests.Content_rules_answer_the_ipres_content_cases, CorpusValidationTests.Every_document_produces_exactly_its_declared_findings |
vendor/pdf-association/handwritten-type3-recursion.pdf, handwritten-indexed-color-out-of-range.pdf, every damaged/* document | No untyped exception, no hang: each within its time and allocation budget, each anomaly reported | CorpusExtractionTests.Hostile_content_never_breaks_the_interpreter |
documents/stress/reportlab-journal-1000-pages.pdf; remote, the 9,302-page remote/govinfo/us-code-2023-title42.pdf and the 63 MB page of remote/usgs/us-topo-washington-west-2023.pdf | Retained memory flat across 10, 100 and 1,000 pages and independent of document length, within budgets recorded in status.md; 0 B allocated per glyph once warm, output excepted; the map's page read under its readerLimits and the new guards' defaults | CorpusExtractionTests.Extracting_the_largest_documents_holds_its_budget, ExtractionBenchmarks |
The corpus's hybrids — the Mustang, Docentric and GnuAccounting invoices, the factur-x Python library's, M14's own Factur-X invoices, and vendor/zugferd/itext-pdfbox-weclapp-facturx-en16931-invoice.pdf, whose visible total differs from its XML by a cent (visible-total-differs-from-xml); remote, the FNFE-MPE, intarsys, Symtrax, Konik and DWC ones | Each business term the page shows found, on its page and at the quads PyMuPDF finds; a term the page does not show reported, as the recorded reading of each page expects; the weclapp invoice's total reported facturx.visible-value-not-found | CorpusFacturXTests.Visible_values_match_the_embedded_xml (new) |
| Every export of every committed document | Two runs, and runs under two cultures, give identical bytes | CorpusExportTests.Extraction_is_deterministic |
| The same operations through the tool | The AOT binary produces what the API produces | CorpusToolTests.Text_markdown_search_and_inventory_match_the_api |
The remote rows close only on a green Remote corpus run, recorded in status.md with its date. The rows that
name a decrypted twin close on the twin; the original joins them when M16 opens it.
Corpus
What the corpus holds
- Text expectations:
textContainson 109 committed documents that open and 179 remote ones;hasExtractableText: falseon 14 committed and 45 remote;attachmentson 9 committed and 12 remote. - Scans and OCR layers: image-only scans in JBIG2 (Xerox), CCITT G3 and G4, JPEG and JPX; OCR layers from
Acrobat, two Xerox copiers, ABBYY FineReader 8 (text under the image), a Cerfa page; remote, Konica, Ricoh,
Canon, OmniPage, Acrobat Paper Capture, Pixel Translations and AbleDocs (
ocr-layer,invisible-text,invisible-text-render-mode-3,text-under-image,ocr-text-painted-under-page-image,image-only). - Text that is not text: Word's text drawn as images (
text-as-images), a poster in vector outlines (text-as-paths, remote), render mode 7 in Photoshop's file (text-render-mode-7-clip, remote). - Tagged documents: 66 committed that open, from Word, LibreOffice, Chromium, InDesign, PDFlib, Antenna House,
PDFMaker and FOP (
tagged,tagged-structure,actual-text,figure-alt-text,table-structure,list-structure,toc-structure,role-map,duplicate-mcid,struct-parents-without-struct-tree). - Fonts: embedded Type 1 (six programs in two committed documents that open), Type 3 (three committed
documents, eleven remote), bare and CID-keyed CFF, TrueType with and without
ToUnicode, non-embedded standard 14 and non-standard fonts, symbolic fonts (type3-font,type3-bitmap-font,type3-named-encoding,embedded-type1,embedded-type1c,cidfonttype0c,no-tounicode,differences-encoding,truetype-subset-macroman-no-tounicode). - Scripts: Arabic, Russian, Greek, Traditional Chinese twice, vertical Japanese (encrypted); remote, Hebrew twice, shaped Arabic, GBK font names, PDF 2.0 UTF-8 strings with bidirectional controls.
- Producer faults:
tounicode-byte-swapped,tounicode-space-mapped-to-tab,soft-hyphen-minus,tounicode-duplicate-codes,duplicated-text-runs,word-spaces-as-kerning,word-spaces-by-positioning-only,letter-spaced-text,ligature-fi,ligature-glyphs,unmapped-ligatures,missing-glyph,control-char-in-extracted-text,text-before-tf,y-down-text-matrix,one-tj-per-glyph,text-shifted-by-three-in-pdftotext,compacted-syntax-no-whitespace-between-tokens. - Layout:
two-column-text(Chromium's tagged report and LibreOffice's two renderings),repeated-table-headers(LibreOffice's renderings),vector-table-grid,wrapped-table-header, rotated pages (rotate-270,/Rotate 90),UserUnit, an inverted media box (remote), article threads (remote). - Layers: nine committed documents with
/OCProperties(one of them encrypted) and eight remote, none with text hidden by default. - Content cases: the 18 iPRES 2017 content-stream files (remote) and the hand-written hostile files.
- Scale: the 1000-page journal; remote, 9,302 pages and a 63 MB page.
What it lacks
| Need | Why | Priority | Likely source |
|---|---|---|---|
| An untagged twin of the two-column report, and untagged multi-column documents whose order is known — two and three columns, a full-width heading between column sets | Every committed multi-column page is tagged, so the roadmap's two-column acceptance would test the structure and never the layout analysis; the one untagged two-column article is remote | 1 | Generated here: Chromium from sources/report-fr.html with tagging turned off (the DevTools protocol's printToPDF with generateTaggedPDF: false), and a recorded pikepdf transformation that removes the structure tree from the tagged file; a public source: US Federal Register issues (three columns, public domain) |
| Text in an optional-content group off by default, and in a group shown only when printing | The committed hidden groups draw boxes, not text; "hidden layers honored" has nothing to hide | 1 | Generated here: a recorded pikepdf construction over a committed page, recorded in build_corpus.py; a public source: GIS and CAD exports, the remote US Topo map's layers to be examined first |
| Hyphenated line breaks — hard hyphens at line ends and soft hyphens —, in French and English | Search across hyphenated breaks is an acceptance, and no committed document hyphenates | 1 | Generated here: LibreOffice with automatic hyphenation (hyphen-fr, hyphen-en-us) from an ODT source, and Chromium from HTML with ­ |
The manifest fields this milestone reads: expect.readingOrder, expect.tables, expect.images, expect.hiddenText, expect.textlessPages | Expectations must come from the file and an independent tool, never from our reader | 1 | Generated here: build_corpus.py records images from pdfimages -list; the others from the generated sources and from a person's reading of third-party pages, each recorded with its reason |
The decrypted twin of indesign-distiller18-nta-gift-tax-vertical.pdf, with textContains | The only vertical Japanese without ToUnicode — the case the predefined CMaps exist for — is AES-128-encrypted under an empty user password, and no expectation records its text | 1 | Derived here: qpdf --decrypt, a recorded transformation, the text recorded from MuPDF and confirmed by a reader |
textContains for the 22 committed documents that have none, documents/stress/reportlab-journal-1000-pages.pdf among them | "Text matches the manifest" says nothing about them | 2 | Generated here: build_corpus.py records strings from the generators' sources; for third-party files, from poppler and PyMuPDF where they agree |
| Korean, and CJK text in non-embedded fonts with predefined CMaps that we may commit | The Adobe-Korea1 map has no committed case; the non-embedded CJK fonts are Foxit's, remote | 2 | A contribution (W09); a public source: Korean government forms under the KOGL Type 1 license |
Tables with merged cells, ruled and unruled, untagged; and a tagged table with RowSpan and ColSpan | Spans are where table detection fails, and no committed table has them | 2 | Generated here: LibreOffice Calc with merged cells, ReportLab's SPAN, Chromium from HTML with rowspan and colspan |
| A real invoice or statement received from a supplier, untagged, with a table, that we may commit | The line-item detection is proven on our own invoice and on published samples; real ones are remote (W04) | 2 | A contribution (W04) |
| Text made invisible by other means — white on white, outside the crop box, clipped away | The visibility flags exist for M19's sanitization and for prompt injection hidden in documents fed to language models; the corpus has only render mode 3 and 7 | 2 | Generated here: a recorded ReportLab construction; a public source: published prompt-injection samples, examined before any is vendored |
| Bates numbers applied by another tool — Acrobat's, an e-discovery platform's | Anchors are proven on M09's marks only | 3 | A contribution |
| A math-heavy article with inline formulae, from LaTeX and from Word | Extraction of formulae is not promised, but its failure mode — order and spacing — should be seen | 3 | A public source: arXiv papers under CC BY; the remote pdfTeX dissertations first |
Traps
- A glyph is not a character, and
ToUnicodelies. Byte-swapped, spaces as tabs, minus signs as soft hyphens, controls, gaps: every one of them is in the corpus. Read through, report, never trust blindly. - Content order is not reading order. Producers draw headers last, tables by column, and OCR layers line by line in any order. The structure tree, when present, is the author's; otherwise the layout decides.
- A space is often not a glyph. Word boundaries by position,
TJdisplacements,Twon code 32 only; and a letter-spaced title drawn with space glyphs between its letters is extracted with them, as every referee extracts it. - Invisible is not hidden. Render mode 3 is how OCR layers are written, and FineReader paints its text under the page image: that text is the scan's only text. Optional content switched off is hidden. The flags keep the difference; the defaults follow it.
- Hidden content still changes the graphics state (M11's trap), so the interpreter runs it and drops only its paints.
ActualTextreplaces, and can contain anything, control characters included.- Type 3 glyph space is not text space:
/FontMatrixscales it, often to 0.001, sometimes flipped, and the widths are in glyph space too. /Widthswins over the program's widths for positioning — it is what viewers advance by —, and a missing/Widthsfor a standard 14 font is legal before PDF 2.0./Length1in a Type 1FontFileis often wrong, and some producers leave PFB segment headers inside.- Right-to-left text arrives in either order. Reordering by position makes both read alike; trusting content order reverses one of them. The referees disagree, and poppler is wrong on the IRS Arabic edition.
- Vertical CMaps and
/W2change the advance direction; a vertical page read as horizontal is one character per line. - Rotation, crop box origin,
/UserUnitand inverted media boxes all move coordinates; exports convert once, from one frame, and PyMuPDF's and poppler's frames are converted the same way before any comparison. - The predefined CMaps are big. Shipping them in the core costs every caller; shipping them apart must never degrade silently.
- Unicode normalization depends on the platform unless the tables are ours.
- A caller's regular expression is hostile input's companion: without
NonBacktrackingor a timeout, a pattern and a page can take the process. - Markdown built from a file's text is a structure the file chooses unless every character is escaped.
- Referees are not ground truth. poppler and xpdf disagree on Arabic, xpdf 4.00 extracts nothing from compacted syntax, PyMuPDF and poppler place boxes differently; the manifest records whose reading is right and why.
- A table seen is not a table. A form's boxes, a ruled letterhead and aligned addresses all look like grids; the confidence and the evidence are the answer, not a stricter detector.
- A rule identifier must not equal a diagnostic code: M08 already uses
content.operands-dropped, M12resource.not-found.
Documentation
docs/website/docs/concepts/text-extraction.md(new): glyphs, words, lines and blocks; where text comes from and how sure the library is; visibility; reading order, tags first; page text kinds.docs/website/docs/guides/extracting-text.md(new): pages, options, plain text,GetTextUnder, fonts that do not map, scans that have no text.docs/website/docs/guides/search.md(new): queries, folding, line breaks, regular expressions and their timeout, quads for M11's highlights and M19's redaction.docs/website/docs/guides/exports.md(new): Markdown, JSON, ALTO, hOCR, chunking and anchors for retrieval pipelines — and what the library does not do (embeddings, language models).docs/website/docs/guides/tables.md(new): the three detection modes, confidence and evidence, continuations.docs/website/docs/guides/factur-x.md: reading a received hybrid gains the visible check.docs/website/docs/guides/inventory.md(new): each section, active content, space by category.docs/website/docs/reference/extraction-json-schema.mdand the schema itself, beside M07's.docs/website/docs/reference/diagnostics.mdandreader-limits.md: theextraction.*,content.*andsearch.*codes, and the three guards.docs/website/docs/reference/validation-rules.md: thecontent.*andresource.name-undefinedrules; M02's "Resources" row brought in line.docs/website/docs/reference/tool/:text,markdown,search,inventory.docs/website/docs/introduction.mdanddocs/features/features.json: theextractionandsearch-exportentries brought to their state.docs/architecture.md:Text/,Extraction/, the interpreter and its device seam for M19 and M25, the generated Unicode tables, and — if slice 3's ADR puts the predefined CMaps in a data satellite — its row in §2's package table.docs/corpus.mdand the manifest schema: the newexpectfields;docs/corpus-sources.md: the referees' disagreements recorded.- The ADR on the predefined CMaps' packaging, written in slice 3.
docs/status.md: the budgets, the tolerances and the table calibration; #39 closed if M03 and M07, which bring pdftotext, veraPDF andpdftoppminto the suite, have not closed it; the referees this milestone adds named indocs/corpus.md.
Exit criteria
- The interpreter runs every page of every document the reader opens, with the device seam M19 and M25 need, and the three guards classified and documented.
- Type 1, Type 3, CFF, TrueType and CID fonts map codes to text by the §9.10.2 order, and every
ToUnicodefault in the corpus is read through and reported. - The predefined CMaps are packaged as the slice 3 ADR decides, their size measured and recorded.
- Glyphs, words, lines, blocks, reading order with its confidence, page text kinds, tables with their calibrated confidence, the structure read model, search, images, exports, chunks and the inventory exist as designed, with their public API documented.
- The
content.*andresource.name-undefinedrules are in the structural profile, documented, and every manifest entry declares what they report on it. -
AdCodicem.Pdf.FacturXchecks a received hybrid's visible invoice number, dates and totals against its XML, with M14's matcher over this milestone's search. - #36's text-string decoding is closed — by M11 at the latest; if not, before slice 4.
- The priority-1 gaps above are filled; each remaining gap is recorded in
docs/corpus-contributions.md. - The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on
a green
Remote corpusrun recorded instatus.md. - Unit tests cover each behavior, its degenerate cases and its hostile ones; the FsCheck properties hold; the interpreter, the CMap parser and the Type 1 reader run in the nightly fuzzing campaign.
- Integration tests run poppler, MuPDF, PyMuPDF, pdfplumber, pikepdf, FreeType, fontTools, veraPDF, qpdf, cmark-gfm, xmllint, hocr-tools and jsonschema, each in a container.
-
InterpreterBenchmarks,ExtractionBenchmarks,SearchBenchmarksandExportBenchmarksrun withMemoryDiagnoser;status.mdrecords the budgets and the text row of the comparison benchmarks. - The tool's verbs ship in the dotnet tool and the AOT binaries, documented.
- The documentation site publishes the pages listed above.
- Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).