01

A table on a page is not always a data table

A PDF can display a grid of numbers without storing rows and columns as spreadsheet cells. Extraction software may infer boundaries from lines and text positions. Merged headings, wrapped labels and borderless layouts can make those boundaries ambiguous.

02

Decide what must be preserved

Before conversion, identify the header row, units, footnotes and any totals. A value such as 00123 may be an identifier whose leading zeroes matter. Commas and periods can represent different decimal or thousands conventions. Dates can also be ambiguous; record the source convention before cleaning the data.

03

Check more than the first row

Compare the start, middle and end of the extracted table with the PDF. Look for shifted columns and repeated page headers inserted as data. Recalculate totals where appropriate, but remember that a matching grand total alone does not prove every row is correct. Keep the source as a reference for later checks.

04

Availability and alternatives

PDF to Excel is coming soon in Folio. PDF to Text is available for selectable text, but it does not promise spreadsheet columns or typed numeric cells. For a small table, carefully transcribing and checking values may be more reliable than assuming a text export is ready for calculations.

Explore the tools available today

This feature is in development. Merge, split, compress, extract text or turn an image into a PDF now.

Explore free tools →