← all articles

Parsing PDFs for an AI pipeline without losing the table

The PDF problem nobody warns you about

Every RAG pipeline tutorial starts the same way: load a folder of PDFs, chunk them, embed them, done. Then you point it at a real document, a vendor contract with a pricing table on page 4, and the answer comes back wrong. Not hallucinated wrong. Wrong because the table never made it into the index in a usable form.

This isn’t a model problem. It’s an extraction problem, and it happens before your LLM ever sees a token. If you’re building anything that ingests PDFs for retrieval, structured extraction, or agentic workflows, you need to understand what a PDF actually is, because it is not a text format. It’s a page description format, and that distinction is the root of every table-parsing headache you’ll hit.

Why PDFs don’t contain “text” the way you think

A PDF stores instructions for a rendering engine: draw this glyph at coordinate (312, 480), draw that glyph at (340, 480), draw a line here. There’s no inherent concept of “this is a paragraph” or “this is a table cell.” Words that appear next to each other on the page might be stored in completely different order in the file, sometimes reversed, sometimes column-by-column instead of row-by-row, depending on how the PDF was generated.

When you run a naive text extractor (think pypdf or pdftotext in default mode) it reads glyphs in file order and mushes them into a string. For prose, this usually works fine because most PDF generators write text in reading order. For tables, it frequently doesn’t. A three-column table can come out as three vertical text blocks concatenated, or as a single scrambled line where row 1 column 1, row 2 column 1, row 3 column 1 all appear before row 1 column 2 shows up. Feed that into a chunker and then an LLM, and the model is trying to reconstruct a table from a shuffled deck. Sometimes it gets close. Often it invents numbers to fill the gaps, because that’s what language models do when the input pattern looks like a table but the values don’t line up.

The three extraction strategies people actually use

Coordinate-based reconstruction. Tools like pdfplumber and PyMuPDF (fitz) give you access to the bounding box of every character or word on the page. You can cluster words by their x/y position to infer rows and columns, then rebuild the table structure yourself. This works well on PDFs with clean, ruled tables (visible grid lines) because you can snap text to the nearest cell boundary. It falls apart on tables that use whitespace alignment instead of lines, or that span multiple pages with inconsistent column widths. You end up writing a fair amount of heuristic code, and every new PDF template can break your heuristics.

Layout-model based extraction. Tools built on document layout models, LayoutLM-style architectures, or table-detection models like Microsoft’s Table Transformer, treat the PDF page as an image and run object detection to find table regions, then a second pass to segment rows and columns. This is what’s under the hood in most “smart” PDF parsing services. It handles messier layouts better than pure coordinate math because it’s trained on scanned and rendered documents, not just clean vector text. The cost is that it’s slower (you’re running inference per page) and it can still misread merged cells, multi-line cell content, or tables with no visible borders.

Vision-model extraction. The newer approach, and the one gaining ground fast, is to render each PDF page as an image and send it to a multimodal LLM with a prompt asking for the table content in markdown or JSON. This sidesteps the coordinate-reconstruction problem entirely because the model is reading the page the way a human would, visually. It’s noticeably better at tables with merged headers, rotated text, or unusual formatting that breaks rule-based extractors. The tradeoffs are real though: you’re paying per-page vision API costs, latency goes up because each page is a separate model call, and you’re trusting the model not to silently drop a row or misread a decimal point in a way that looks plausible. I’ve seen a vision model turn “1,024” into “1,624” on a low-resolution scan and produce a table that reads perfectly clean, no obvious sign anything was off.

None of these three is strictly better. Which one you reach for depends on your document source, and most production pipelines end up using more than one.

What actually breaks in production

Three failure modes show up again and again once you’re past the demo stage.

Scanned PDFs with no text layer force you into OCR before any of the above even applies. Tesseract is free and fine for clean scans, but table structure recognition on top of OCR output is where quality drops off a cliff, skewed scans, faint grid lines, and multi-column layouts all degrade the row/column inference.

Multi-page tables are a structural problem, not a parsing accuracy problem. A table that continues across a page break usually repeats the header row on the new page, or doesn’t, inconsistently, and your extraction code needs explicit logic to detect and merge continuation tables. Most off-the-shelf tools treat each page independently and just give you two separate tables.

Nested or merged cells (a header that spans three columns, a cell with a sub-list inside it) don’t map cleanly to a simple 2D grid at all. Whatever format you extract into, CSV, markdown table, JSON, you have to decide how to flatten that structure, and the decision changes what your downstream LLM can correctly reason about. A markdown table with a merged header awkwardly repeated across three columns is more useful to an LLM at query time than a “technically accurate” nested JSON structure the model then has to unpack in its own reasoning.

What to actually do about it

Don’t run one extraction pass on the whole document and hope. Split your pipeline into a layout detection step and a content extraction step, even if you’re doing it cheaply. Detect where the tables are first (bounding boxes, page numbers), then run a targeted extraction just on those regions instead of relying on your general text extractor to also handle tabular data correctly. This alone fixes a large share of the “table turned into word soup” problem, because you stop asking one tool to do two different jobs.

For the table regions themselves, match the tool to the document type. Digitally generated PDFs with ruled tables (most invoices, most exported reports) do fine with coordinate-based tools like pdfplumber, they’re cheap, fast, and deterministic, which matters when you’re processing thousands of documents and can’t afford per-page API costs. Scanned documents or PDFs with irregular, borderless tables are where the added cost of a vision-model pass earns its keep, because the alternative is a coordinate reconstruction that’s confidently wrong.

Whatever you extract, store the table as markdown, not as a flattened paragraph of comma-separated values shoved into a text chunk. LLMs parse markdown tables meaningfully better than they parse “Column A: value, value, value” style flattening, because the row and column structure is explicit in the syntax rather than implied by punctuation.

And validate. Row count and column count consistency checks across a sample of extracted tables will catch a shocking number of silent failures, merged rows, dropped columns, header duplication, before they ever reach your retrieval index. This is the boring, unglamorous part of building an AI pipeline that people skip, and it’s the part that determines whether your system gives correct answers about the third row of a pricing table six months from now or just sounds confident while being wrong.

PDF parsing isn’t solved by picking the right library. It’s solved by treating table extraction as its own step in the pipeline, choosing the extraction strategy per document type instead of one tool for everything, and checking your output before you trust it downstream.

For more breakdowns of what actually works when you’re building with AI tools day to day, head back to the homepage.

for builders
Running agents or scrapers at scale?

AI pipelines that crawl, research, or automate the web hit rate limits and geo-blocks fast. Singapore Mobile Proxy runs real 4G/5G mobile IPs that carriers still trust.

see plans →
read on
More from the Gazette

Tool reviews, model and pricing news, and build guides for people shipping real things with AI.

browse all articles →