Diagnostics
Real PDFs are frequently not conforming. A reader that refuses them fails where every reader on the market succeeds; a reader that silently patches them leaves you guessing. This library does the third thing: it repairs what it can and returns an account of what it did.
The report is a returned value, not a log: it can be inspected, serialized, asserted on in your own tests, and used to decide whether to accept a file into your system. Find out what the reader repaired shows how; Diagnostics, in the reference, lists every severity and code.
What the reader did, not what is wrong
Diagnostics say what the reader did to read a file. To know what is wrong with the file itself — whether or
not the reader could work around it —, validate it: findings are a separate verdict, with
their own severities and rule identifiers that never reuse a diagnostic code. A Repair says what the reader did,
not how wrong the file is.
Where a stream ends
The reader parses an object through a window of the file — 8 KB to start with — and reads it again
through a larger one when it runs past the edge. What the smaller window saw there, such as a string
without its end or a stream without its endstream, is dropped with that attempt, so the report says what
the object holds, not where a window happened to end. When MaxObjectLength allows no larger window, the
object is kept as far as it was read and the guard is reported: what the parse met before the edge — a stray
token, a key that is not a name — is reported beside it, and nothing of the token the edge cut, nor of what that token
leaves open, is.
A stream's declared length is checked against the endstream that must follow it. When the data runs past
the window, the reader asks the file for the few bytes after the declared length rather than reading the
data, and they tell it whether endobj follows too. When no endstream is there, the first one after the
start of the data ends it, and where it is looked for depends on where the declared length ends:
- Inside the window, or nowhere — a
/Lengththat gives no length —, the firstendstreamafter the data is taken wherever it lies, past the header of the next object too: the reader looks in the window, and in a larger one while none is found, as far asMaxObjectLengthallows. - Past the window, the reader searches the file from the start of the data up to the next object —
the one the file's index places as the file wrote it (as the reader rebuilt it when the document opened, if
the file wrote none), or the first object header,
N G obj, the file's bytes hold, whichever comes first —, or up to the end of the file. Stopping at the next object keeps a later object'sendstreamfrom ending this stream, which is what a stretch of zeros that erased the end of one and the objects after it would otherwise do, or an object whose entry the index lost; when none lies before it, the declared length is kept, and the report says so. Text in a damaged stream's data that reads as an object header stops the search as well, and the declared length is kept: an object number and a generation the parser takes, however many zeros lead them, white space between the three tokens, however much, and white space, a delimiter, or the start or the end of the search around them. A comment between the tokens makes no header here, as it makes none for the scan that rebuilds a damaged index; nor does a regular character glued before the number, though that scan takes the digits after it for one.
The search is not one of the reader's limits: a valid file's endstream follows its length, so only a
damaged file is searched, and once for each stream. Where it stops depends on the file's bytes and on its
index as the file wrote it, which the reader rebuilding or correcting its own index later does not change:
in the 401 documents of the project's test corpus measured, each stream takes the same length whatever was
read before it — though a damaged index can still serve another copy of an object, at another offset,
according to what was read first. A file crafted for it can still make what a stream holds depend on which
object was asked for first. What the searches of one document read together is bounded: none starts once
they have read four times the file's length. They read twice the file at most — 0.35 of it at most in the
test corpus — unless their stretches overlap: one object's header inside another's dictionary, or a header
the search does not take — a regular character glued before its number — that the index as the file wrote
it does not place. Past the bound, a stream keeps its declared length without a search, and the report says
so.
Damaged stream data
Damaged stream data decodes as far as it goes, and what decoded is kept, with a report. A Flate stream whose tail was lost, as a file cut short or a producer that stopped writing leaves it, is a warning: what decoded before the end is all there is. A Flate stream that lost only the zlib checksum after its last block decoded whole, unchecked, and is a repair, as a stream with no zlib header at all is. A Flate stream that turns corrupt is a warning too: decoding stops at the fault, and what decoded before the byte the fault lies in is kept; what that one byte decoded ahead of the fault can be lost with it, though no stream in the project's test corpus lost any. The report says at which byte of the encoded data the fault was found and how many bytes were kept, and does not vouch for them: the damage may lie before the point where decoding found it, since damaged data can go on decoding, wrongly, for a while. A stream in which nothing decodes before that byte is left encoded, and reported as one that could not be decoded. Finding what decoded before a fault reads the stream's data again, the last few kilobytes before the fault a byte at a time; a sound stream is read once.
A Flate stream that decodes to its end, but whose zlib checksum disagrees with what it decoded to, is a
warning of its own, filter.checksum-mismatch: all of its data is kept, as other readers keep it, and some of
it may be wrong — the checksum does not say where. In the project's test corpus, none of the 86 such streams
could be shown intact, and 39 were shown damaged, 28 of them by a change of line endings that, undone, makes
the checksum agree.
An LZW stream that uses a code it has not defined stops there, since what follows cannot be read reliably. An LZW stream without its end-of-data code is taken as complete, as other readers take it.
Decoding is lazy like the rest of reading, so a stream's damage is reported when the stream is decoded, not when the
document opens — and a stream you built in memory and decode without a PdfDiagnostics of your own has nowhere to
report, and reports nothing. A document's stream never decodes in silence.
The reader's limits are not faults
Growing windows, and decoding, are bounded. The limit.* codes report the reader's own limits, not
faults of the file: a valid document can reach them — a large-format scan decodes past the 256 MB a stream
may decode to by default. What fits within the limit is kept, and the limit replaces what the parser met where it
stopped. Reader limits explains why they exist and why they are on by default.
Exceptions, by contrast
Exceptions are reserved for what makes the operation impossible: the input is not a PDF, or it is encrypted and
cannot be opened. Anything the reader can work around is a diagnostic, never an exception — unless you ask for one:
with PdfReaderOptions.ThrowOnLimit, reaching a reader limit throws from whichever operation reached it, for an
application that would rather refuse a document than read part of it. The
reference lists them.