Skip to main content

M08 — Fonts, text and content streams

State: to do — Depends on: M03 — Fonts per ADR 11; shaping and text analysis left to .Html by ADR 43

Goal​

Write text that is correct, embedded, extractable and accessible: any font the caller registers — or the embedded OFL set — parsed, subsetted and embedded so that every glyph on the page reads back as the text that was meant, with a diagnostic naming every character no font could draw.

Everything that puts marks on a page after this milestone goes through it: M09's stamps and Bates numbers, M10's barcodes and their captions, M11's annotation appearances, M12's whole engine, M16's field appearances. The failures it exists to prevent are the ordinary ones: a stamp in Helvetica that silently voids a PDF/A claim, a ligature that extracts as nothing, a Cyrillic surname printed as empty boxes with no word to anyone, a subset tag that two different subsets share, and a content stream appended to Chromium's output that draws upside down because the page never restored its own cm.

Scope​

In:

  • a font program parser in the core: the sfnt container and collections (ttcf), head, hhea, maxp, OS/2, post, name, hmtx, cmap (formats 0, 2, 4, 6, 12, 13 and 14 — every one a font in use carries), glyf and loca, and CFF — name-keyed and CID-keyed, from an OpenType file or bare from a PDF's FontFile3 — with a Type 2 charstring interpreter for widths, subroutine use and accented composites;
  • WOFF and WOFF2 decoding, including WOFF2's glyf, loca and hmtx transforms, bounded by ADR 34 guards;
  • subsetting of TrueType and CFF outlines, deterministic, with a subset tag derived from the glyph set;
  • embedding as Type0 over CIDFontType2 or CIDFontType0, Identity-H, with /W, /DW, a ToUnicode CMap that expresses ligatures and supplementary-plane characters, and /CIDSet when the output claims PDF/A-1;
  • the standard 14 fonts: their metrics, their built-in encodings, WinAnsi text, and the Adobe Glyph List;
  • a font registry — explicit, immutable, thread-safe — with family resolution by the CSS Fonts 4 matching algorithm, per-cluster fallback by script coverage, synthetic bold and oblique, and, when nothing covers a character, a visible .notdef (or, under a PDF/A-2+ or PDF/UA target, a drawn box) and a diagnostic naming the code points;
  • text on a page: a simple path that maps characters to glyphs one to one for the library's own text (stamps, headers, captions), and a glyph-run path that takes glyphs already shaped — the contract M12's HarfBuzz output will use — with ActualText where ToUnicode cannot say what a cluster means;
  • content streams: a builder for every operator of ISO 32000-2 §8 and §9 that a generator needs, checked against the graphics-object state machine, with resources named without collision, form XObjects and graphics-state dictionaries; and a tokenizing reader — the one M07's pruning started from the M01 lexer, promoted — with inline images and the balance analysis (q/Q, BT/ET, marked content) that M09 needs before it appends anything to a page;
  • a forward-only generation API, PdfDocumentBuilder: pages written as they are finished, fonts finalized last — the public face of M03's internal writer, which M03 left to "M06 or M08" to decide;
  • the embedded OFL font set (debt #34): Liberation Sans, Serif and Mono as WOFF2 in a data-only satellite, AdCodicem.Pdf.Fonts, and Liberation Sans regular in the core — the maintainer's decision of 2026-09-27, confirmed in an ADR with the measured sizes;
  • the command-line tool's fonts verb, which lists a document's fonts as pdffonts does and says what our parser made of each embedded program.

Out, explicitly:

  • shaping — ligature substitution, kerning, mark positioning, complex scripts, OpenType features — and bidirectional analysis, line breaking and hyphenation: AdCodicem.Pdf.Html (M12.2), by ADR 43 and architecture §3.3. The core's simple path maps one character to one glyph and says so when a script needs more; everything else arrives pre-shaped through the glyph-run path;
  • @font-face, unicode-range, local() and fetching web fonts: M12.6, through the resolver of ADR 38; M08 decodes the bytes it is given;
  • vertical writing (Identity-V, vmtx, vhea): M30;
  • reading Type 1 (FontFile), Type 3 and simple fonts' encodings for extraction, and the predefined CJK CMaps: M15. M08's CFF parser reads bare CFF programs' metrics, which M15 builds on;
  • variable fonts (fvar, gvar, CFF2) and color fonts: open questions of the roadmap; a CFF2 program is refused with a diagnostic, a variable glyf font is used at its default instance and reported;
  • consolidating subsets across merged documents and unembedding: M23; embedding missing fonts into received documents: M21;
  • the logical structure: M13. M08 writes marked content with ActualText, Lang, Alt and E, and nothing that needs a structure tree;
  • image XObjects: M07 (passed through) and M12.5;
  • encrypting what is written: M16.

Design​

Where it lives​

Fonts/ and Content/ in the core (architecture §3), no dependency: the BCL's BrotliDecoder for WOFF2, ZLibStream for WOFF and Flate, SHA256 for subset tags, StringInfo for grapheme clusters. The one table the BCL lacks — the Unicode Script and Script_Extensions properties — is generated at build time as static span data from the Unicode Character Database of the version CharUnicodeInfo implements, with a test that fails when the two versions part. That generator is the only one (ADR 43): M12 and M15 add their tables to it, each compiled into the assembly that reads it, the core holding Script and, from M15, Bidi_Class. The OFL set lives in AdCodicem.Pdf.Fonts, its sans regular in the core (below).

Font programs​

PdfFontProgram (internal) one face of an sfnt, a collection member, or a bare CFF: its table directory,
read lazily from a PdfFontSource; tables parsed on first use and never held whole
unless subsetting asks for glyph data
PdfFontSource bytes, a file path read on demand, or a PDF stream (FontFile2, FontFile3) decoded under
the document's reader limits
PdfFontFace public, immutable: family and style names (every localization), weight, width, slope,
unitsPerEm, ascender, descender, line gap, cap and x heights, underline and strikeout,
italic angle, bounding box, embedding permissions (OS/2 fsType), coverage
PdfFontCoverage the code points the cmap maps, as sorted ranges: Covers(codePoint) is a binary search,
no allocation
  • Metrics follow the OpenType specification: hhea ascender and descender unless OS/2 sets USE_TYPO_METRICS; cap and x heights from OS/2 version 2 and later, measured on H and x otherwise; advances from hmtx, the last advanceWidth repeated past numberOfHMetrics.
  • cmap: the (3,10) subtable, then (0,4) and (0,6), then (3,1), then (3,0) for a symbol font, then (1,0); format 14 for variation sequences, whose default glyphs are the base's. A subtable that runs past its table is read up to where the table ends and reported — the corpus holds four such programs (below).
  • glyf is read glyph by glyph through loca; composite glyphs are walked iteratively, with a visited set, so a component cycle is reported, not followed. A loca whose offsets go backwards or past glyf makes the glyphs it misplaces empty, reported.
  • CFF: INDEX, DICT, charset (the three predefined sets included), Encoding (with its supplements), FDArray, FDSelect (formats 0 and 3), private DICTs, local and global subroutines. The Type 2 charstring interpreter runs a glyph to find its width (defaultWidthX, nominalWidthX), the subroutines it calls, and the two components of an endchar that carries seac-like arguments; it keeps the specification's own limits — 48 operands, subroutine calls nested 10 deep — and a guard on the operations a glyph may execute. hintmask and cntrmask read their bytes from the stem count. random, which no deterministic subsetter can honor, makes the glyph's subroutines count as all used, reported.
  • Parsing never throws on a program's content. A missing required table, a table past the end of the file, a checksum that disagrees (information: fonts in the wild get it wrong), a maxp that contradicts loca — each is a font-program.* diagnostic, and what can be read is read. PdfFontFormatException is kept for a source that is not a font at all.

Bounds, classified (invariant 12, ADR 34)​

BoundKindWhy
A font program's length, after WOFF or WOFF2 decoding — 64 MBGuard, PdfReaderLimits.MaxFontLength, limit.font-lengthA valid face can exceed it: Arial Unicode MS is 23 MB, a pan-CJK OpenType collection well over 64 MB. One record governs every read of hostile bytes, a font read from a PDF or one a caller registers (ADR 34, amended 2026-09-27), so the registry takes a PdfReaderLimits too; there is no separate PdfFontLimits
Operations a CFF glyph may execute, subroutines included — 65,536Guard, PdfReaderLimits.MaxCharstringOperations, limit.charstring-operationsThe Type 2 format bounds nesting, not work: ten levels of subroutines that each call the next a hundred times are valid and take 10²⁰ steps
48 operands, 10 nested subroutine callsInternal constantThe Type 2 charstring specification's own limits (Adobe TN 5177, appendix B); a program past them is invalid
Composite glyph depthNone neededThe walk is iterative with a visited set, so its work is bounded by numGlyphs, at most 65,535 by the format
A WOFF2 table that decompresses past its declared lengthInternal constant: stop at the declared lengthThe declared totalSfntSize and origLength are checked against MaxFontLength before anything is allocated; data past them makes the file invalid
Glyphs per subset — 65,535Format limit, not a boundA two-byte CID; a subset that would pass it opens a second subset of the same face

WOFF and WOFF2​

  • WOFF 1.0: the table directory, each table inflated with ZLibStream into a buffer of its declared origLength, which is checked first; metadata and private blocks ignored.
  • WOFF 2.0: the header, the variable-length table directory, one Brotli stream decoded with the BCL's BrotliDecoder into a pooled buffer of the declared size; the glyf and loca transform (triplet-encoded points, the bounding-box bitmap, composite and instruction streams) and the hmtx transform reconstructed; collections read. Reconstruction checks every count against the stream it indexes, so a hostile point count ends in a diagnostic, not a read past a buffer.
  • The output is an ordinary sfnt in memory, handed to the same parser. WOFF2 is decoded, never encoded here.

Subsetting​

FontSubsetter (internal) per font resource: the (glyph, text) pairs used, then Finish() -> the program
  • Numbering: .notdef stays glyph 0; each pair (source glyph, the text it stands for) gets the next glyph number at its first use, and the CID written in the content stream is that number, so content can be written before the subset exists and /CIDToGIDMap is /Identity. The same glyph standing for two texts — µ for U+00B5 and U+03BC, a space glyph standing for U+202F, é precomposed and decomposed — gets two numbers whose outlines are the same bytes, so ToUnicode can tell them apart.
  • TrueType: composite components are added at Finish, after every glyph the content used, and their indices rewritten in the composite data; loca in short or long form as the offsets require, hmtx with numberOfHMetrics recomputed, maxp, head, hhea updated; cvt , fpgm and prep kept, since instructions reference them; a minimal (3,1) cmap, name, OS/2 and post version 3 kept although ISO 32000 does not need them for a CIDFontType2 — the OpenType Sanitizer, which Chromium runs and our referee runs, refuses a font without them; GSUB, GPOS, GDEF, kern, DSIG, hdmx, LTSH, VDMX dropped. Tables in tag order, padded, checksummed, checkSumAdjustment recomputed.
  • CFF: always written CID-keyed — a name-keyed font converted with registry–ordering–supplement Adobe–Identity–0, one Font DICT holding the original private DICT and FontMatrix, FDSelect format 3 — so it embeds as FontFile3 /CIDFontType0C, which PDF/A-1 accepts, rather than /OpenType, which needs PDF 1.6. The charset maps new glyph to CID one to one. Subroutines the kept glyphs never call are replaced by a bare return, which keeps every index and bias as it was: no charstring is rewritten.
  • Embedding permissions: fsType restricted-license (bit 1) refuses the face for embedding, reported, and resolution moves to the next face; no-subsetting (bit 8) embeds the whole program; bitmap-only (bit 9) is refused. IgnoreEmbeddingRestrictions exists for a caller who holds the license, and is recorded in the document's diagnostics when used.
  • The subset tag is six capital letters from the SHA-256 of the face's PostScript name and the sorted list of source glyphs: two different subsets never share a tag unless the hash collides, and two runs give the same tag — never a random one. The corpus shows what shared tags cost (Illustrator and Konik reuse one tag for different subsets).
  • Determinism: same face, same pairs in the same order, same bytes. Nothing in the output depends on the clock, the platform or a hash table's iteration order.

Embedding​

PdfFontResource one embedded face in one document: object numbers reserved at creation, the subsetter,
the (glyph, text) -> CID map; Finish() writes the program and the dictionaries
  • Type0 with /Encoding /Identity-H and /DescendantFonts [CIDFontType2 or CIDFontType0], /CIDSystemInfo (Adobe, Identity, 0), /DW set to the most frequent width and /W in its compact forms (c [w…] for runs, cfirst clast w for equal widths). Widths are in thousandths of an em, rounded to an integer, and layout uses the rounded values, so what the library measures is what a viewer advances; veraPDF's width-consistency rule is the referee of whether the rounding stays inside its tolerance.
  • FontDescriptor: /FontName with its tag, /Flags (symbolic, as for every CIDFont), /FontBBox, /ItalicAngle, /Ascent, /Descent, /CapHeight, /StemV estimated from the weight class, /FontFile2 or /FontFile3. /CIDSet, deprecated in PDF 2.0, is written — complete — only when the output claims PDF/A-1, which requires it.
  • ToUnicode: codes sorted, beginbfchar and beginbfrange blocks of at most 100 entries (the CMap format's limit), destinations in UTF-16BE with surrogate pairs, several code points per destination for a ligature. Every code maps to something: PDF/A-2u, 3u and PDF/UA require it.
  • Finalization: content streams reference the font by name as they are written; the program, /W and ToUnicode are written at the end, into the numbers reserved at creation — the forward-only writer's rule (architecture §3.2), which lets a thousand pages share one subset without a second pass.
  • The version table of ADR 40 gains no row: everything here is PDF 1.3 or earlier.

The standard 14 fonts​

  • Widths, bounding boxes, ascent, descent and cap heights for the 14 fonts, generated at build time from Adobe's Core 14 AFM files, whose license allows redistribution with its notice (carried in NOTICE), with their KPX pairs kept for M15 and M21. Built-in encodings — Standard, WinAnsi, MacRoman, PDFDoc, Symbol, ZapfDingbats — from ISO 32000-2 Annex D, and the Adobe Glyph List (BSD-3-Clause) for glyph names.
  • Writing a standard 14 font: a simple Type1 font with /WinAnsiEncoding for the Latin twelve, the built-in encoding for Symbol and ZapfDingbats, and /FirstChar, /LastChar, /Widths always — PDF 2.0 requires them even for these fonts. A character WinAnsi lacks falls back, per cluster, to a registry face.
  • They are never the default. A standard 14 font is used only when the caller names it; it is never embedded, so a PDF/A or PDF/UA target, or PDF 2.0 output (which deprecates it), gets a ConformanceLoss or a warning naming the font — never silence (invariant 7).

The registry​

PdfFontRegistry immutable, thread-safe; Resolve(query, text) -> runs of (face, text range)
PdfFontRegistryBuilder Add(source, overrides), AddCollection(source), AddDirectory(path) (opt-in, sorted),
AddOflSet() (from slice 10), SetGenericFamily(generic, families),
SetScriptFallbacks(script, families), Limits (PdfReaderLimits)
PdfFontQuery families (CSS order, generic families included), weight, stretch, style,
synthesis allowed (weight, style)
PdfFontRun face, synthetic bold, synthetic oblique, the text range it covers
  • Nothing is scanned by default (ADR 11): a container has no fonts, and output must not depend on the machine. AddDirectory exists for a caller who wants the system's fonts, and adds them in sorted path order.
  • Names: typographic family and subfamily (name IDs 16 and 17) before the legacy ones (1 and 2), every platform and language registered as an alias — a Japanese face answers to its native name — compared with ASCII case folding, as CSS compares family names. The first face registered for a (family, weight, stretch, style) wins; a later duplicate is reported.
  • Matching is CSS Fonts 4 §5.2, step by step: stretch, then style (italic, then oblique, then normal), then weight (400 tries 500 before lighter weights, 500 tries 400, lighter and heavier searched in the specified order). Collections register each member.
  • Fallback, per grapheme cluster (StringInfo, UAX #29): the query's families in order, then the registry's fallback families for the cluster's script, then every face in registration order. A cluster falls back whole — a base with its combining marks, an emoji with its modifiers — so a mark is never drawn from another face than its base. Characters of script Common or Inherited (spaces, digits, punctuation) stay in the current run's face when it covers them, so the hyphen and initials' full stops of a Cyrillic name stay in its run rather than switching faces between letters. Default-ignorable code points (joiners, variation selectors, the soft hyphen) are not looked up; their text goes with the preceding cluster, so extraction keeps it.
  • Whitespace a face lacks — U+202F, which French typography puts before ;, :, !, ? and », the other fixed-width spaces of U+2000 to U+200A — is drawn with the face's own space glyph, advanced to the width the character calls for and mapped to the character itself through the (glyph, text) pair: no fallback face for a space, and no character lost on extraction.
  • Nothing covers the cluster: the primary face's .notdef, mapped in ToUnicode to the missing text, so the page shows a box and extraction still returns what was meant; under a PDF/A-2, 3 or 4 or a PDF/UA target, which forbid referencing .notdef, a box drawn as a path of the same advance inside a /Span whose /ActualText is the missing text. Either way, one text.code-point-not-covered warning per resource, listing the code points with their counts, at most 64 of them and the number left out.
  • Synthesis: bold asked of a family with no bold face is drawn fill-and-stroke (2 Tr, a line width of 1/30 of the size), oblique with a text matrix skewed by 14°, the angle CSS uses; each reported once (text.style-synthesized, information), and refused when the query forbids it.
  • The registry holds names, styles, coverage ranges and metrics — a few kilobytes a face — and reads glyph data from the source only when a document embeds it, through a small cache of open programs with a fixed capacity.

Text on a page​

PdfTextStyle face query, size, color, rendering mode, character and word spacing, horizontal scale
PdfSimpleText (internal) the core's own path: clusters -> glyphs through the cmap, one to one;
Measure(text) -> advance; greedy wrapping at spaces for a stamp's box
PdfGlyphRun a shaped run: glyph ids, advances and offsets, a cluster index per glyph, the source
text — what HarfBuzz returns, which M12 passes through unchanged
  • The simple path does no kerning, no ligature substitution, no reordering. A cluster of a script that needs shaping or bidirectional reordering — Arabic, Hebrew, Syriac, Thaana, N'Ko, the Indic and South-East Asian scripts, Mongolian, Tibetan — is drawn in logical order with the face's default glyphs and reported (text.shaping-required, warning, naming the script). A caller who needs those scripts before M12 passes a glyph run.
  • ShowGlyphs writes a TJ whose numbers carry the difference between the run's advances and the widths the font dictionary declares, and a Ts rise around glyphs with a vertical offset. The first glyph of a cluster carries the cluster's text in ToUnicode; a cluster that one glyph per character cannot express — several glyphs for one character, a reordered cluster — is written inside /Span << /ActualText … >> BDC … EMC, and its other glyphs map to U+FFFD rather than to nothing.
  • Spaces are drawn, as the space glyph, never as a gap in a TJ: extraction reads a gap as a space only when it guesses right (the corpus's Atypon article places words by position alone).
  • Accessibility belongs to M13, but its raw material is written here: Lang on a span whose language differs from the document's, ActualText, Alt and E (expansion of an abbreviation) on marked content.

Content streams​

PdfContentBuilder writes one content stream into an IBufferWriter<byte>: every operator of the graphics
state, paths, clipping, color, text state and positioning, text showing, XObjects,
shadings, marked content, inline images of known length; checked against the graphics-
object state machine of ISO 32000-2 figure 9
PdfResourceScope the names a stream uses, per category: allocated as /F1, /X1, /GS1… skipping every name
the page or form already has and every name token its existing content uses
PdfFormXObjectBuilder a form XObject: /BBox, /Matrix, /Resources, an optional transparency /Group, its content
PdfGraphicsState an ExtGState value (opacity, blend mode, line parameters), deduplicated by value
PdfContentReader reads a content stream, or a page's streams as one sequence, as (operator, operands):
pooled operands, no interpretation
PdfContentBalance what a sequence leaves open: q depth and its lowest point, an open text object, open
marked-content sequences, top-level cm
  • The state machine refuses, with InvalidOperationException, a text operator outside BT…ET, q, Q, cm or Do inside a text object, a painting operator with no path, BT inside BT, and a stream finished with q, BT or marked content open. These are the caller's mistakes, not a file's.
  • Numbers: integers as they are; computed reals formatted invariant, at most four decimals, no exponent, -0 written 0, through IUtf8SpanFormattable into the buffer — no string, no allocation. Four decimals of a point is a hundred-thousandth of a millimeter. Text strings for a Type0 font as hexadecimal two-byte codes; for a simple font, literal with (, ) and \ escaped.
  • The reader is M01's lexer driven over a decoded stream, with the one construct the object lexer does not know: an inline image. BI … ID is parsed as a dictionary, abbreviations expanded; the data length comes from /L or /Length when present, from the image's dimensions when it has no filter, and otherwise from the first EI preceded by white space and followed by white space or the end — the ambiguity ISO 32000-2 added /L to remove — reported when it had to be guessed. A real token past a double's range is an infinity, which the reader reports as syntax.number-out-of-range and reads as null, as the object parser does. Operands pending before an operator are bounded by an internal constant of 256, since no operator takes more than 33 (scn with a pattern and 32 components); what exceeds it is dropped and reported. BX/EX sections and unknown operators pass through.
  • A page's streams are one sequence: ISO 32000 lets an operator's operands straddle the boundary between two streams of one page (the remote PDFpen file splits a graphics state across them), so the reader joins them with a delimiter, never parses them one by one.
  • The balance analysis exists for M09: before appending to a page, it says how many Q the content pops below its starting depth, how many q it leaves open, whether it ends inside a text object or inside marked content, and whether it changes the CTM outside any q. The corpus has every case: Chromium and ReportLab open with a top-level cm that is never undone, PDFKit's page ends with a q open, Distiller 4's congressional bill and PDFMaker 7's legal history end pages inside BT because their content is damaged.

Generating a document​

PdfDocumentBuilder Create(output, PdfGenerationOptions) — AddPage(size) -> PdfPageBuilder — Finish()
PdfPageBuilder Content (a PdfContentBuilder), Resources; Finish() writes the page and forgets it
PdfGenerationOptions immutable: output version (ADR 40), conformance target, document identifier and dates
supplied by the caller, metadata, compression
  • The builder is the public face of M03's forward-only writer, which stays internal: a page's content stream is compressed on the fly behind an indirect /Length and flushed when the page finishes; fonts, the page tree and the catalog are written at Finish into numbers reserved ahead. Memory follows the heaviest page plus the fonts' (glyph, text) maps — never the page count.
  • Editing an existing document goes through M03's change set: PdfPage.AppendContent takes the same PdfContentBuilder and font resources, finalized when the document is saved. M09's stamper, which touches every page, applies its marks instead as a plan the save asks for page by page, so that its memory does not grow with the page count; both share the builder, the resource scope and the font resources.
  • M03's cancellation and progress convention applies to Finish, and progress to pages written.

The OFL set (#34)​

  • What: faces under the SIL Open Font License that give the library a first attempt that works in a container with no fonts — the generic families serif, sans-serif and monospace, and metric-compatible substitutes for Helvetica and Arial, Times and Times New Roman, Courier and Courier New, which M21 embeds in place of fonts a received document left out. Liberation Sans, Serif and Mono (four styles each) are the set: their advances are Arial's, Times New Roman's and Courier New's, which are the standard 14's — the acceptance checks exactly that against the widths Word, Distiller and PDFMaker wrote. Coverage beyond Latin, Greek and Cyrillic, and CJK above all (a single CJK weight is over 15 MB), is registered by the caller or waits for a package of its own.
  • Where, decided by the maintainer on 2026-09-27: a data-only satellite, AdCodicem.Pdf.Fonts, holding the set as WOFF2 files (roughly half the size of the TTFs, decoded once per process by slice 9's decoder) and referenced by AdCodicem.Pdf.Html; the core itself carries one face — Liberation Sans regular — so that the core alone can stamp a PDF/A document without falling back to a non-embedded font. Slice 10's ADR records it with the measured sizes of the package and of the core's face, and the license checks below.
  • License: each face's OFL text and copyright in NOTICE; the Reserved Font Names respected — the set ships the faces unmodified, and embedding a subset in a document is use, not a Modified Version, which the slice confirms against the OFL FAQ and records.

Diagnostics​

In PdfDiagnosticCodes, disjoint from validation rule identifiers (ADR 36):

CodeSeverityMeaning
font-program.table-missingWarningA table the use needs is absent; what can be done without it is done
font-program.table-truncatedWarningA table, or a cmap subtable, runs past its end; read up to the end
font-program.checksum-mismatchInformationA table's checksum disagrees; common and harmless
font-program.glyph-malformedWarningA glyph's outline or charstring cannot be read; drawn empty
font-program.composite-cycleWarningA composite glyph reaches itself; the cycle is cut
font-program.unsupportedWarningCFF2, a bitmap-only sfnt, a variable font used at its default instance
font-program.embedding-restrictedWarningfsType forbids embedding; the face is skipped
limit.font-length, limit.charstring-operationsWarningThe guards above; the message names the PdfReaderLimits property
text.code-point-not-coveredWarningNo face covers these code points; .notdef or a box drawn
text.fallback-usedInformationThese clusters were drawn from a fallback face
text.style-synthesizedInformationBold or oblique drawn synthetically
text.shaping-requiredWarningThe simple path drew a script that needs shaping or reordering
text.standard-font-not-embeddedConformanceLoss, or WarningA standard 14 font under a PDF/A or PDF/UA target, or in PDF 2.0 output
content.inline-image-end-guessedWarningAn inline image's end was found by searching for EI
content.operands-droppedWarningOperands past the internal bound before an operator

The command-line tool​

fonts FILE [--json]: every font a page, form XObject, annotation appearance or Type 3 glyph reaches — name, type, encoding, embedded, subset, ToUnicode, object — as pdffonts prints them, plus what our parser made of each embedded program (format, glyph count, diagnostics). It is the tool's first verb built on M08, and its output is what the acceptance compares with poppler's.

Slices​

Each slice ends on a green commit, with the diagnostics it introduces documented and its benchmark, if it has one, recorded in docs/status.md.

  1. The sfnt container and its metrics. Delivers PdfFontProgram over the table directory and collections, head, hhea, maxp, OS/2, post, name, hmtx, every cmap format, PdfFontFace, PdfFontCoverage, MaxFontLength, the font-program.* codes. Proved by unit tests on the test fonts and on hand-made hostile tables (a directory entry past the file, overlapping tables, a format 12 group count larger than its table, a numberOfHMetrics of zero); an FsCheck property — any byte sequence parses to a face or to PdfFontFormatException, never another exception —; the corpus's embedded TrueType programs read through FontFile2 — 259 in the documents the reader opens today, 28 more in encrypted ones from M16; integration: FreeType (freetype-py, FT_LOAD_NO_SCALE) and fontTools in a Python container agree with every advance and every cmap mapping; mutation fuzzing seeded with the test fonts. Leaves outlines.
  2. Outlines: glyf, loca and CFF. Delivers composite walking, the CFF structures, the Type 2 interpreter with its limits and MaxCharstringOperations, bare CFF from FontFile3. Proved by the corpus's 58 bare CFF, eight CID-keyed CFF and one OpenType program outside the encrypted documents, each glyph's advance equal to FreeType's; the programs fontTools refuses (below) read with a diagnostic, never an exception; a component cycle and a subroutine bomb in unit tests; remote: Oracle Outside In's broken loca, ArcMap's malformed FDSelect, typeset.sh's undecodable program. Leaves writing.
  3. Content streams. Delivers PdfContentBuilder with the state machine and number formatting, PdfResourceScope, PdfFormXObjectBuilder, PdfGraphicsState, PdfContentReader with inline images, PdfContentBalance; M07's resource scan moved onto the reader. Proved by an FsCheck property — every operator sequence the builder accepts reads back through PdfContentReader as the same operations —; a table of sequences the state machine refuses; ContentBenchmarks at 0 B allocated per operator; integration: pikepdf's parse_content_stream reads every generated stream as we wrote it, and our reader and pikepdf's agree on the operations of every committed document's pages that pikepdf can parse; the balance analysis agrees with a pikepdf walk on the cases above. Leaves text.
  4. The standard 14 fonts and simple text. Delivers the generated AFM metrics, encodings and glyph list, simple Type1 font dictionaries, PdfSimpleText over them, PdfDocumentBuilder and PdfPageBuilder, and the first generated documents. Proved by unit tests per font and encoding; the producers' /Widths for non-embedded Times and Helvetica in the corpus equal ours; integration: pdftotext and PyMuPDF extract the generated text, qpdf --check passes. Leaves embedding.
  5. TrueType subsetting. Delivers FontSubsetter for glyf fonts — (glyph, text) numbering, duplicates, composites, tables, checksums, the tag — and FontBenchmarks.Subset_truetype. Proved by each subset glyph drawing, through fontTools' recording pen, exactly what its source glyph draws; determinism over repeated runs and shuffled registration; integration: ots-sanitize accepts every subset, FreeType's advances equal the source's. Leaves CFF.
  6. CFF subsetting. Delivers CID-keyed output, name-keyed conversion with its FontMatrix, subroutine blanking, seac components. Proved by the same outline equality through fontTools' T2 decompiler with subroutines expanded; a face at 2,048 units per em drawn at the right size; ots-sanitize and FreeType in containers. Leaves the PDF objects.
  7. Embedding, ToUnicode and the glyph-run path. Delivers PdfFontResource, Type0 with its descendant, /W and /DW, the descriptor, /CIDSet for PDF/A-1, ToUnicode, finalization into reserved numbers, ShowGlyphs with TJ adjustments, rises and ActualText. The shaped runs the tests pass are HarfBuzz's — the shaper M12 will pass them from — computed once with uharfbuzz, HarfBuzz's Python binding, pinned in the Python referee container, and committed as test data beside the test fonts with each font's checksum, so that no test project takes a native dependency (ADR 43 keeps HarfBuzzSharp in .Html). Proved by generated documents with ligatures, marks and supplementary-plane characters extracting exactly; integration: pdffonts reports every font embedded, subset and with Unicode; veraPDF's PDF/A-2u profile reports no failure under clause 6.2.11 (fonts) — the rest of the profile is M14's —; pdftotext and PyMuPDF extract the exact text. Leaves choosing fonts.
  8. The registry and fallback. Delivers PdfFontRegistry, its builder, names and aliases, CSS matching, per-cluster fallback by script, whitespace substitution, .notdef and the drawn box, synthesis, the text.* diagnostics; the generated Script tables. Proved by unit tests over the matching rules of CSS Fonts 4 one by one, fallback of a combining sequence, a Common character inside a Cyrillic run, a missing character under each target; FsCheck — for any string, the runs cover it exactly once, in order —; the acceptance document of the roadmap, extracted exactly by pdftotext and PyMuPDF. Leaves web formats.
  9. WOFF and WOFF2. Delivers both decoders with the transforms and the collection directory, under MaxFontLength. Proved by each test font compressed by fontTools and by Google's woff2_compress decoding to what Google's woff2_decompress produces, in a container — byte for byte for every table but the reconstructed glyf and loca, whose flag packing decoders may choose, and glyph by glyph through fontTools' recording pen for those —; the WOFF 1.0 tables byte-identical to the original's; hostile WOFF2 files — a point count past its stream, a declared size of 4 GB, a Brotli stream that runs past its table — ending in diagnostics; a fuzzing campaign over the decoder in the nightly job. Leaves the set.
  10. The OFL set (#34). Delivers the ADR, the AdCodicem.Pdf.Fonts package with its API baseline, the core's Liberation Sans regular, AddOflSet, the generic families and the standard 14 substitutes, NOTICE. Proved by the set's advances for Arial, Times New Roman and Courier New's metric twins equal the /Widths Word, Distiller and PDFMaker wrote for those fonts in the corpus; the package size measured and recorded; a stamp-sized document generated with no font registered embeds its font and passes the slice 7 font clauses. Leaves the tool and the budgets.
  11. The fonts verb, budgets and documentation. Delivers the verb, FontBenchmarks and ContentBenchmarks with MemoryDiagnoser, the memory checkpoints of a 1,000-page generated document, the remote corpus run, the site. Proved by CorpusToolTests.Fonts_matches_pdffonts, the budget tests, and a green Remote corpus run recorded in docs/status.md.

Tests required​

Unit — tests/AdCodicem.Pdf.Tests:

  • The parser: every table on well-formed and hostile input; every cmap format, including a format 4 whose last segment is not 0xFFFF and a format 12 with overlapping groups; metrics under USE_TYPO_METRICS and without; a collection member by index; the checksum rule; the guards reached, raised and thrown, as ReaderLimitsTests does for the reader's.
  • CFF: each predefined charset, both FDSelect formats, a CID-keyed font, seac-like endchar, hintmask counting, the operand and nesting limits, the operation guard, random.
  • Subsetting: numbering by first use; one glyph for two texts; composites added last with rewritten indices; loca switching from short to long; a subset that passes 65,535 glyphs opening a second resource; the tag's determinism; fsType in each state.
  • Embedding: /W compaction; rounding shared by layout and /W; ToUnicode with surrogates, multi-code-point destinations and blocks at their limit of 100; /CIDSet present exactly under a PDF/A-1 target.
  • The registry: each step of CSS matching; names in several languages; duplicates; per-cluster fallback; whitespace substitution; .notdef and the drawn box under every target; synthesis allowed and refused; FsCheck over arbitrary strings — runs partition the text, and every code point is either drawn by a face that covers it or listed in the diagnostic.
  • Text: the simple path's measurement against the sum of advances; wrapping at spaces; ShowGlyphs with advances that differ from /W, with vertical offsets, with a many-to-one and a one-to-many cluster.
  • Content: the state machine's refusals; every operator's syntax; number formatting at its edges (−0, 1e-5, 1e9, rounding at the fourth decimal); resource names that collide with existing names and with name tokens in content; inline images with and without /L, including the contradictory-keys file; the balance analysis on each shape the corpus has.
  • WOFF and WOFF2: every transform, collections, and the hostile shapes above; a WOFF2 decoded through a seam that reports Brotli unavailable ends in font-program.unsupported, not an exception.
  • Hostile inputs everywhere: FsCheck over random bytes for every parser entry point — a face, a PdfFontFormatException, or a diagnostic, and never another exception, a hang, or an allocation the input chose.

Integration — tests/AdCodicem.Pdf.IntegrationTests, each referee in a container (ADR 27):

  • FreeType (freetype-py) and fontTools, pinned in a Python image: advances and cmap of every corpus font program; outline equality of every subset glyph with its source.
  • OpenType Sanitizer (ots-sanitize, from the pinned opentype-sanitizer package): every subset we write.
  • Google's woff2 tools: the reference decoding of every WOFF2 test file.
  • poppler — pdffonts (embedded, subset, Unicode, type) and pdftotext; PyMuPDF as the second extractor, on a different engine; pikepdf for content-stream parsing.
  • veraPDF: the font clauses of PDF/A-2u (6.2.11) and of PDF/UA-1 (7.21) on every generated document.
  • qpdf — --check on every generated document.

Acceptance conditions​

DocumentsBehaviorVerified by
Every embedded TrueType (FontFile2, 259 programs), bare CFF (FontFile3 /Type1C, 58), CID-keyed CFF (/CIDFontType0C, 8) and OpenType program (/OpenType: the GPO's Myriad Pro in vendor/us-federal/itext-govinfo-us-code-certified.pdf) in the 101 committed documents that embed one and are not encrypted, read through their PDF; the 36 programs of the 17 encrypted ones join when M16 decrypts themEvery glyph's advance equals FreeType's, unscaled, to the font unit; every cmap mapping equals fontTools' where fontTools reads the table; no exceptionCorpusFontTests.Every_embedded_font_program_reports_the_advances_freetype_reads, FreeTypeRefereeTests
The programs fontTools refuses: the TrueType subsets whose cmap subtable is six bytes short in vendor/opf-format-corpus/pdfmaker5-distiller5-va-select-agents.pdf (two), distiller505-va-portland-adjuvant-guidelines.pdf and distiller505-va-portland-animal-care-guidelines.pdf; the CFF programs of distiller3-dea-cfr-damaged.pdf, distiller4-congress-hr1904-enrolled-bill.pdf and distiller952-pscript5-kb-pdf-risk-inventory.pdfRead with a font-program.* diagnostic naming what is wrong, every advance FreeType gives equal to ours; where the fault is fontTools' rather than the program's, the test records it with its reasonCorpusFontTests.Damaged_font_programs_are_read_as_far_as_they_go
Remote: remote/pdfjs/outside-in-pdfa1a-broken-loca-issue17671.pdf (a broken loca under a PDF/A-1a claim veraPDF upholds), remote/pdfjs/arcmap-cff-fdselect-bug1146106.pdf (a malformed FDSelect), remote/pdfjs/typeset-sh-corrupt-flate-issue11651.pdf (a font program that does not decode)Diagnostics, never an exception; the glyphs the program still describes read as FreeType reads themCorpusFontTests.Damaged_font_programs_are_read_as_far_as_they_go
The non-embedded standard 14 fonts whose producers wrote /Widths: Times and Helvetica in vendor/opf-format-corpus/groff-distiller405-mac-usgs-gps-noise-spectra.pdf, Times in vendor/us-federal/finereader8-frb-sr0115-examiner-guidance.pdf, vendor/us-federal/hp-mfp-acrobat-ocr-nih-report.pdf and vendor/opf-format-corpus/distiller6-pscript5-va-bariatric-bibliography.pdf; vendor/uk-ogl/pagemaker-distiller5-hmrc-iht205-form.pdf once M16 opens it (RC4-128)Our AFM-derived width equals the producer's for every code the document uses, within one unitCorpusFontTests.Standard_14_widths_are_the_producers_widths
Non-embedded Arial, Times New Roman and Courier New with /Widths from Word, Distiller and PDFMaker: vendor/uk-ogl/word2019-ccs-contract-schedule.pdf, distiller6-pscript5-va-bariatric-bibliography.pdf, pdfmaker7-powerpoint-va-cancer-database-course.pdf, pdfmaker707-word-va-esig-developer-guide.pdf, pdfmaker707-word-law-library-iraq-legal-history.pdf, distiller952-pscript5-kb-pdf-risk-inventory.pdf; pagemaker-distiller5-hmrc-iht205-form.pdf once M16 opens itThe OFL set's metric-compatible faces advance as those widths say, within one unit, for every code the document uses (Word writes 0 for the codes it does not): the substitutes M21 will embed change no line break. A disagreement is recorded with its reason, never absorbed into the toleranceCorpusFontTests.The_ofl_set_substitutes_without_moving_a_glyph
A document we generate — the roadmap's sample: accented French (Pièce n° 2 — « Été », Œuvre, fi and fl from shaped runs), a CJK sample, a Cyrillic name absent from the primary face, a supplementary-plane character, a U+202F before : — with the test fonts, in 1.7 and 2.0 outputpdftotext and PyMuPDF extract exactly the input text; the Cyrillic and CJK clusters come from their fallback faces and text.fallback-used lists themCorpusTextTests.Generated_text_extracts_exactly_as_written, PopplerTextRefereeTests, PyMuPdfTextRefereeTests
The same documentpdffonts shows every font embedded, subset and with a ToUnicode; each subset holds exactly the glyphs used, their components and .notdef (fontTools); ots-sanitize accepts each; veraPDF finds no failure under PDF/A-2u 6.2.11 or PDF/UA-1 7.21CorpusTextTests.Every_generated_font_is_embedded_and_subset, PdffontsRefereeTests, OtsSanitizeRefereeTests, VeraPdfFontRefereeTests
The same document with a character no registered face covers, under no target and under a PDF/A-2u targetA .notdef box, then a drawn box with ActualText; text.code-point-not-covered names the code point; veraPDF finds no .notdef reference under the targetCorpusTextTests.An_uncovered_character_is_visible_and_reported
Generated twice, with the fonts registered in two orders that do not change resolutionIdentical bytesCorpusTextTests.Generation_is_deterministic
Every committed document's pages that pikepdf's content parser readsOur reader yields the same operators and operands; the balance analysis finds exactly what a pikepdf walk finds — the top-level cm of Chromium and ReportLab, PDFKit's open q (vendor/node-signpdf/pdfkit-node-signpdf-unsigned-placeholder.pdf), the damaged pages that end inside BT in distiller4-congress-hr1904-enrolled-bill.pdf and pdfmaker707-word-law-library-iraq-legal-history.pdfCorpusContentTests.Content_reads_as_pikepdf_reads_it, PikepdfContentRefereeTests
vendor/pdf-association/handwritten-inline-image-abbreviations.pdf (inline images whose keys contradict each other, which pikepdf's parser refuses), handwritten-content-stream-indirect-refs.pdf, and remote remote/ocrmypdf/pdfpen-imprints-split-content.pdf (a graphics state split across two streams)Read without an exception; each anomaly reported; the split operands joinedCorpusContentTests.Awkward_content_is_read_whole
Every committed document, through the toolfonts --json lists the fonts pdffonts lists, with the same type, embedded, subset and Unicode columnsCorpusToolTests.Fonts_matches_pdffonts
A generated document of 10, 100 and 1,000 pages of the report's textRetained memory flat across the checkpoints, within a budget set from the first measurement and recorded in docs/status.md; 0 B allocated per character mapped once the faces are warmCorpusTextTests.Generating_a_thousand_pages_holds_memory_flat, FontBenchmarks, ContentBenchmarks

The remote rows close only on a green Remote corpus run, recorded in status.md with its date.

Corpus​

What the corpus holds​

  • Font programs, measured with pikepdf on 2026-09-26 and counted by font descriptor: 118 committed documents embed 368 programs — 287 TrueType, 61 bare CFF, 12 CID-keyed CFF, one OpenType and seven Type 1 (M15's); 17 of those documents are encrypted and their 36 programs wait for M16 — from Word, PDFMaker, Distiller 2 to 9.5, InDesign, Illustrator, LibreOffice, Chromium, PDFlib, FOP, Ghostscript, Quartz and PDFBox. Features embedded-truetype, truetype-subset, type0-subset, cidfonttype2*, cidfonttype0*, embedded-type1c, type1c-subset, full-font-embedding, fonts-embedded-in-full, truetype-subsets-no-os2-table, symbolic-truetype, cidtogidmap-stream, cidset.
  • Damaged or odd programs: the four Distiller 5 subsets with short cmap subtables and the three CFF programs above; remote, truetype-loca-malformed, cff-fdselect-malformed, font-program-undecodable.
  • Scripts in real producers' output: Cyrillic (indesign-irs-pub1-russian.pdf), Greek (pdflib-oj-exchange-rates-greek.pdf), traditional Chinese in CID-keyed CFF and in an Arial Unicode MS TrueType subset (indesign-irs-pub1-chinese-traditional.pdf, distiller6-cdc-west-nile-chinese-traditional.pdf), vertical Japanese (indesign-distiller18-nta-gift-tax-vertical.pdf, AES-128, from M16), Arabic (indesign-irs-pub1-arabic.pdf); remote, Hebrew and WeasyPrint's Arabic.
  • Standard 14 and core Windows fonts left unembedded, some sixty documents, with /Widths in the ones named above: the reference for the AFM metrics and for the OFL set's metric compatibility.
  • What producers get wrong with text, for the traps: reused-subset-tag (Illustrator; remote, Konik), unmapped-ligatures (Illustrator), ligature-glyphs (remote, Ghostscript), word-spaces-by-positioning-only (remote, Atypon), tounicode-duplicate-codes and soft-hyphen-minus (remote, Axapta), tounicode-space-mapped-to-tab (remote, wkhtmltopdf), missing-glyph (PDF24).
  • Content streams: top-level cm in Chromium's and ReportLab's output, an open q from PDFKit, pages ending inside BT, unbalanced-q-Q, inline-image*, content-array, multiple-content-streams, graphics-state-split-across-streams (remote), bx-ex-compatibility-section, unknown-content-operator.

What it lacks​

NeedWhyPriorityLikely source
The acceptance sample — accented French with fi and fl ligatures, CJK, a Cyrillic name, a supplementary-plane character, U+202F — as Chromium and LibreOffice render itThe roadmap's first acceptance is judged on our output alone; a real producer's rendering of the same text says what "extracts exactly" means for them, and whether their ligatures and fallbacks do better than ours2Generated here: one HTML and one ODT source in tests/corpus/sources, through Chromium and LibreOffice, with OFL fonts
A committed document embedding a CFF-flavored OpenType face from a current producerOnly the GPO's Myriad Pro is /OpenType; the 12 CID-keyed CFF programs are from InDesign, Distiller and PDFMaker; our CFF subsetter's output has no modern peer to be compared with2Generated here: Chromium and LibreOffice with Source Serif 4 (OFL)
Glyphs of the supplementary planes with a ToUnicode that writes surrogate pairsOnly strings, not glyphs, carry them in the corpus (remote handwritten-pdf20-utf8-strings.pdf), and the Power BI emoji font is remote2Generated here: Chromium with a mathematical-alphanumerics or CJK Extension B face under the OFL
A document whose text came from a TrueType collection member (.ttc)Collections are a registry feature; no corpus document says which producers handle them and how their subsets look3A contribution: Word on Windows with MS Gothic, or Quartz with a system collection
A document with a restricted-license (fsType bit 1) face that a producer embedded anyway, or refused toThe embedding-permission rule has no real case to be compared with3A public source (GovDocs1), else a contribution

Test fonts are not corpus documents: they live in tests/fonts/ under their OFL licenses, each under 2 MB — a CJK face subsetted by fontTools to the sample's ideographs and a margin, the subsetting recorded — with a WOFF and a WOFF2 of each made by fontTools and Google's woff2_compress, and hostile fonts made by recorded mutations.

Traps​

  • A page's own cm is not undone for you. Chromium opens every page with a scale of 0.24 and a vertical flip (0.23999999 0 0 -0.23999999 0 841.91998 cm) outside any q, and never restores it. Content appended after it draws upside down and at a quarter of its size. The balance analysis exists for this.
  • .notdef is forbidden by PDF/A-2 and later and by PDF/UA, even in invisible text. The visible box the roadmap asks for must be a path under those targets.
  • A glyph is not a character. One glyph for two characters (a ligature), two glyphs for one (a decomposition), one glyph for two different characters (µ): ToUnicode maps a code to one string, so the code must be chosen per (glyph, text), not per glyph.
  • CID-keyed CFF has two FontMatrixes. A name-keyed face at 2,048 units per em converted to CID-keyed must carry its matrix into the Font DICT and leave the top DICT's default, or every glyph is drawn two thousand times too large or too small.
  • Subroutine numbers are biased, by 107, 1,131 or 32,768 depending on the count. Renumbering subroutines changes the bias; blanking the unused ones changes nothing.
  • hmtx is shorter than the glyph count when trailing glyphs share the last advance.
  • A subset tag is not decoration. Two different subsets under one tag let a merge (M06) or an optimizer (M23) treat them as one font; Illustrator and Konik's output in the corpus do exactly that.
  • Widths are rounded twice or not at all. If layout advances by the exact width and /W says the rounded one, a line of 200 glyphs drifts by up to 100 thousandths of an em.
  • A space is a glyph. Word spacing done only by position extracts as run-together words; Tw applies only to the single-byte code 32, never to a two-byte Type0 code, so justification of CID text goes through TJ.
  • Many faces lack U+202F — the narrow no-break space of French typography. Dropping it changes the extracted text; falling back to another face for a space changes the line.
  • The OpenType Sanitizer refuses what PDF allows: a CIDFontType2 program needs no cmap, name, OS/2 or post, but a subset without them fails the referee Chromium also runs.
  • EI inside image data ends an inline image too early; ISO 32000-2 added /L for that reason, and older producers never write it.
  • Operands may straddle two content streams of one page. Tokenizing stream by stream breaks such pages.
  • A registry that reads the machine's fonts is not deterministic, and a container has none: nothing is scanned unless the caller asks.
  • Brotli in browser WebAssembly: M23 runs the core there. The .NET 10 API reference marks BrotliDecoder with no UnsupportedOSPlatform("browser"), unlike ZstandardStream (checked on 2026-09-27), but only M23's browser run proves it; should it be missing on a runtime, WOFF2 fails with a font-program.unsupported diagnostic, never a PlatformNotSupportedException.
  • A face's fsType is a license statement, not a technical limit; the library respects it by default and records it when told not to.

Documentation​

  • docs/website/docs/concepts/fonts.md (new): the registry, matching, per-cluster fallback, synthesis, the standard 14 and why they are never the default, embedding and subsetting, the OFL set.
  • docs/website/docs/concepts/text-and-content-streams.md (new): the simple path and the glyph-run path, what ToUnicode and ActualText guarantee, the content builder and its state machine, the reader and the balance analysis.
  • docs/website/docs/guides/generating-a-document.md (new): PdfDocumentBuilder end to end, with memory figures.
  • docs/website/docs/guides/registering-fonts.md (new): files, collections, WOFF2, generic families, script fallbacks, embedding permissions.
  • docs/website/docs/reference/diagnostics.md and reader-limits.md: the font-program.*, text.* and content.* codes, and the two new guards.
  • docs/website/docs/reference/tool/index.md: the fonts verb.
  • docs/architecture.md §3.3 as built, with the (glyph, text) numbering; AdCodicem.Pdf.Fonts in §2 and the package tables (README, introduction, docs/releasing.md) moved from "Planned, M08" to shipped; NOTICE for the AFM files, the Adobe Glyph List and each OFL face; an ADR for the OFL set.
  • docs/features/features.json: the fonts and typography entries brought to their state.
  • docs/status.md: #34 closed, the measurements.

Exit criteria​

  • The parser reads sfnt, collections, every cmap format, glyf and CFF, and WOFF and WOFF2, with every bound classified and the two guards in PdfReaderLimits.
  • TrueType and CFF subsetting are deterministic, keep outlines identical, and pass ots-sanitize.
  • Type0 embedding writes /W, ToUnicode and, for PDF/A-1, /CIDSet; every generated glyph extracts to the text it was drawn for.
  • The registry resolves by CSS Fonts 4 matching and falls back per cluster; no uncovered character is silent, and none reaches .notdef under a PDF/A-2+ or PDF/UA target.
  • The standard 14 fonts are available and never chosen by default; their use under a target is reported.
  • The content builder, reader and balance analysis exist; PdfDocumentBuilder generates with memory flat across pages.
  • #34 is closed: Liberation Sans, Serif and Mono ship as WOFF2 in AdCodicem.Pdf.Fonts, Liberation Sans regular in the core, licensed and measured as the ADR records.
  • The acceptance conditions above pass on the corpus, in CI, with no document skipped, and the remote rows on a green Remote corpus run recorded in status.md.
  • Unit tests cover every behavior above, its degenerate and its hostile cases; FsCheck properties hold for parsing, content round trips and fallback coverage; the font parsers join the fuzzing campaign.
  • Integration tests run FreeType, fontTools, ots-sanitize, woff2, poppler, PyMuPDF, pikepdf, veraPDF and qpdf in containers.
  • FontBenchmarks and ContentBenchmarks measure parsing, mapping, subsetting and writing with MemoryDiagnoser; the memory budget of the 1,000-page generation is enforced in CI and recorded.
  • The documentation site publishes the pages above; features.json and status.md are brought in line.
  • Every page of Documentation is written in its Diátaxis section, one mode per page (ADR 47).