Skip to main content

Find out what the reader repaired

This guide shows how to find out what the reader had to work around to read a document you received, and how to act on it. It assumes you can open a document; if not, start with the tutorial.

Check whether anything happened​

using AdCodicem.Pdf.Diagnostics;
using AdCodicem.Pdf.Documents;

using var document = PdfDocument.Open("received-from-supplier.pdf");

if (document.WasRepaired)
{
// The index was rebuilt by scanning the file: the reader found the objects itself.
}

if (document.Diagnostics.HasRepairs || document.Diagnostics.HasWarnings)
{
// The file was not conforming, or something may not read as you expect.
}

WasRepaired is true only when the whole index had to be rebuilt. Diagnostics holds everything else the reader worked around. Both say what the reader did so far: a read after Open can still rebuild the index, and add to the diagnostics.

List what the reader did​

Each entry writes itself out with its severity, its code, where in the file it applies and what happened:

foreach (var entry in document.Diagnostics)
{
Console.WriteLine(entry); // Repair xref.offset-adjusted at 1874: …
}

To format it your own way, use its members:

foreach (var entry in document.Diagnostics)
{
Console.WriteLine($"{entry.Severity} {entry.Code} at {entry.Position}: {entry.Message}");
}

Filter on what matters to you​

Filter on the code, never on the message: codes are stable from one release to the next, messages are not. PdfDiagnosticCodes holds every code as a constant.

if (document.Diagnostics.Contains(PdfDiagnosticCodes.XRefRebuilt))
{
// Every object was found by scanning; no offset in the file was trusted.
}

foreach (var entry in document.Diagnostics)
{
if (entry.Severity == PdfDiagnosticSeverity.Warning)
{
Console.WriteLine(entry);
}
}

Catch what decoding reports​

Opening reads the index and the catalog and decodes no page's content, image or font, so a damaged stream is reported when it is decoded, into the same Diagnostics. To keep the reports of one stream apart — to know which image of a page lost its tail —, pass a list of your own:

using AdCodicem.Pdf.Objects;

if (document.GetObject(new PdfObjectId(12)) is PdfStream stream)
{
var reports = new PdfDiagnostics();
var data = stream.Decode(reports);

if (reports.Contains(PdfDiagnosticCodes.FilterFailed))
{
// What came back is what decoded before the damage.
}
}

Decide whether to accept the file​

Diagnostics say what the reader did, not whether the file is sound. To decide whether to accept it, validate it too: Validate a document before accepting it. A policy that refuses a file whose index was lost, and accepts the others with their account, might read:

if (document.WasRepaired)
{
Reject(path, "The file's index was lost; ask the sender for a sound copy.");
}
else
{
Store(path, document.Diagnostics); // keep the account beside the file
}

Apply it once you have read what you need from the document, or once you have validated it — not straight after Open. The index can be rebuilt after Open returned, when a read asks for an object the index lacks while the index may have lost entries, or for one that is neither where the index says nor near it (Lazy reading lists when): WasRepaired turns true then, and a policy applied before that read accepts a file it would have refused after it.

Keep the account​

The diagnostics are a returned value, not a log: store them beside the file, serialize them, or assert on them in your tests. A document keeps 1,000 entries at most by default; SuppressedCount says how many it dropped past that, and PdfReaderOptions.DiagnosticCapacity raises the bound (Tune the reader's memory).

See also​

  • Diagnostics, in the reference: every severity and code.
  • Diagnostics, explained: why the reader repairs and reports rather than refuses.