The problem
Flattening a PDF table can separate a number from its row label or year. Retrieval may then return text that looks relevant but cannot support the answer.
My contribution
- Led the comparison questions and interpretation of the results.
- Used AI-assisted experiments and reviewed the parsing examples and failure analysis.
Working approach: product-led, AI-assisted implementation, followed by review and iteration.
Decisions & tradeoffs
01
Compare document representations
- The choice
- Contrast plain-text extraction with Markdown-based table representation while documenting the downstream setup.
- The tradeoff
- The published evidence is preliminary and confined to one source document.
- How it is checked
- The technical report includes parsing examples, a result table, and question-level failure categories.
What the work shows
The useful result is a concrete account of how layout loss, retrieval misses, arithmetic errors, and semantic ambiguity lead to different failure modes.
Follow the evidence
What this does not establish
- The public question bank contains 30 questions; the historical summary reports 8 scored items with partial credit. The original per-question CSV is not published, and the automated scorer does not implement the report’s manual ±1% rule. No overall accuracy improvement is claimed here.
- The study uses a commercial parser and local generation. It does not establish that all processing stays local or that the work has been peer reviewed.