Find out what the reader repaired
This guide shows how to find out what the reader had to work around to read a document you received, and how to act on it. It assumes you can open a document; if not, start with the tutorial.
Check whether anything happened
using AdCodicem.Pdf.Diagnostics;
using AdCodicem.Pdf.Documents;
using var document = PdfDocument.Open("received-from-supplier.pdf");
if (document.WasRepaired)
{
// The index was rebuilt by scanning the file: the reader found the objects itself.
}
if (document.Diagnostics.HasRepairs || document.Diagnostics.HasWarnings)
{
// The file was not conforming, or something may not read as you expect.
}
WasRepaired is true only when the whole index had to be rebuilt. Diagnostics holds everything else the reader
worked around. Both say what the reader did so far: a read after Open can still rebuild the index, and add to
the diagnostics.
List what the reader did
Each entry writes itself out with its severity, its code, where in the file it applies and what happened:
foreach (var entry in document.Diagnostics)
{
Console.WriteLine(entry); // Repair xref.offset-adjusted at 1874: …
}
To format it your own way, use its members:
foreach (var entry in document.Diagnostics)
{
Console.WriteLine($"{entry.Severity} {entry.Code} at {entry.Position}: {entry.Message}");
}
Filter on what matters to you
Filter on the code, never on the message: codes are stable from one release to the next, messages are not.
PdfDiagnosticCodes holds every code as a constant.
if (document.Diagnostics.Contains(PdfDiagnosticCodes.XRefRebuilt))
{
// Every object was found by scanning; no offset in the file was trusted.
}
foreach (var entry in document.Diagnostics)
{
if (entry.Severity == PdfDiagnosticSeverity.Warning)
{
Console.WriteLine(entry);
}
}
Catch what decoding reports
Opening reads the index and the catalog and decodes no page's content, image or font, so a damaged stream is
reported when it is decoded, into the same Diagnostics. To keep the reports of one stream apart — to know which
image of a page lost its tail —, pass a list of your own:
using AdCodicem.Pdf.Objects;
if (document.GetObject(new PdfObjectId(12)) is PdfStream stream)
{
var reports = new PdfDiagnostics();
var data = stream.Decode(reports);
if (reports.Contains(PdfDiagnosticCodes.FilterFailed))
{
// What came back is what decoded before the damage.
}
}
Decide whether to accept the file
Diagnostics say what the reader did, not whether the file is sound. To decide whether to accept it, validate it too: Validate a document before accepting it. A policy that refuses a file whose index was lost, and accepts the others with their account, might read:
if (document.WasRepaired)
{
Reject(path, "The file's index was lost; ask the sender for a sound copy.");
}
else
{
Store(path, document.Diagnostics); // keep the account beside the file
}
Apply it once you have read what you need from the document, or once you have validated it — not straight after
Open. The index can be rebuilt after Open returned, when a read asks for an object the index lacks while the
index may have lost entries, or for one that is neither where the index says nor near it
(Lazy reading lists when): WasRepaired turns true then, and a policy applied
before that read accepts a file it would have refused after it.
Keep the account
The diagnostics are a returned value, not a log: store them beside the file, serialize them, or assert on them in
your tests. A document keeps 1,000 entries at most by default; SuppressedCount says how many it dropped past that,
and PdfReaderOptions.DiagnosticCapacity raises the bound
(Tune the reader's memory).
See also
- Diagnostics, in the reference: every severity and code.
- Diagnostics, explained: why the reader repairs and reports rather than refuses.