Skip to main content

Class PdfString

Namespace: AdCodicem.Pdf.Objects
Assembly: AdCodicem.Pdf.dll

Represents a PDF string: a sequence of bytes whose interpretation depends on where it appears.

public sealed class PdfString : PdfObject

Inheritance​

object ← PdfObject ← PdfString

Inherited Members​

PdfObject.Resolve(), object.Equals(object?), object.Equals(object?, object?), object.GetHashCode(), object.GetType(), object.ReferenceEquals(object?, object?), object.ToString()

Extension Methods​

PdfObjectExtensions.AsArray(PdfObject?), PdfObjectExtensions.AsBoolean(PdfObject?), PdfObjectExtensions.AsDictionary(PdfObject?), PdfObjectExtensions.AsInteger(PdfObject?), PdfObjectExtensions.AsName(PdfObject?), PdfObjectExtensions.AsNumber(PdfObject?), PdfObjectExtensions.AsStream(PdfObject?), PdfObjectExtensions.AsString(PdfObject?), PdfObjectExtensions.AsText(PdfObject?), PdfObjectExtensions.Resolved(PdfObject?)

Remarks​

The original notation is kept so that a document can be rewritten without gratuitous differences. Text strings are either PDFDocEncoded or UTF-16BE with a byte order mark; binary strings, such as the entries of a document identifier, must never be run through a text conversion.

Constructors​

PdfString(ReadOnlyMemory<byte>, bool)​

Initializes a string from its raw bytes.

public PdfString(ReadOnlyMemory<byte> bytes, bool hexadecimal = false)

Parameters​

bytes ReadOnlyMemory<byte>

The bytes, already unescaped.

hexadecimal bool

Whether the string was written in hexadecimal notation.

Properties​

Bytes​

Gets the raw bytes of the string.

public ReadOnlyMemory<byte> Bytes { get; }

Property Value​

ReadOnlyMemory<byte>

IsHexadecimal​

Gets a value indicating whether the string was written in hexadecimal notation.

public bool IsHexadecimal { get; }

Property Value​

bool

Length​

Gets the number of bytes in the string.

public int Length { get; }

Property Value​

int

Methods​

FromText(string)​

Creates a text string: text in ASCII, a byte per character, when it holds only printable ASCII, from space to tilde; otherwise in UTF-16BE with a byte order mark, in hexadecimal notation.

public static PdfString FromText(string text)

Parameters​

text string

Returns​

PdfString

Remarks​

The choice is made for the whole text: one character outside printable ASCII — an accent, a tab, a line break — writes all of it in two bytes per UTF-16 code unit, after the two of the mark, so é alone takes four bytes. PDFDocEncoding, which holds it in one, is not used.

ToString()​

Returns a string that represents the current object.

public override string ToString()

Returns​

string

A string that represents the current object.

ToText()​

Interprets the string as text.

public string ToText()

Returns​

string

Remarks​

A UTF-16 byte order mark selects UTF-16; otherwise the bytes are read as Latin-1, which matches PDFDocEncoding over the printable range. Full PDFDocEncoding is tracked as issue #36.