The PDF document library: a low level reader that resolves cross reference tables, object streams and stream filters to extract page text through ToUnicode CMaps, together with a streaming generator that lays out text, tables and images.