← Back to all projects

Structure-Aware RAG (Empirical)

An exploratory comparison of plain-text and layout-aware parsing for questions over financial tables.

Sources reviewed: 2026-09-12

01 / CONTEXT

The problem

Flattening a PDF table can separate a number from its row label or year. Retrieval may then return text that looks relevant but cannot support the answer.

02 / OWNERSHIP

My contribution

  • Led the comparison questions and interpretation of the results.
  • Used AI-assisted experiments and reviewed the parsing examples and failure analysis.
Working approach: product-led, AI-assisted implementation, followed by review and iteration.
03 / ENGINEERING JUDGMENT

Decisions & tradeoffs

01

Compare document representations

The choice
Contrast plain-text extraction with Markdown-based table representation while documenting the downstream setup.
The tradeoff
The published evidence is preliminary and confined to one source document.
How it is checked
The technical report includes parsing examples, a result table, and question-level failure categories.

What the work shows

The useful result is a concrete account of how layout loss, retrieval misses, arithmetic errors, and semantic ambiguity lead to different failure modes.

FROM CLAIM TO SOURCE

Follow the evidence

What this does not establish

  • The public question bank contains 30 questions; the historical summary reports 8 scored items with partial credit. The original per-question CSV is not published, and the automated scorer does not implement the report’s manual ±1% rule. No overall accuracy improvement is claimed here.
  • The study uses a commercial parser and local generation. It does not establish that all processing stays local or that the work has been peer reviewed.

Tools & methods

PythonRAGDocument parsingError analysis