LLM (semi) structured training and inference
Key
- EVAPORATE: [2304.09433] Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 1 ^fe342d
-
We identify two fundamentally different strategies […]: prompt the LLM to directly extract values from documents or prompt the LLM to synthesize code that performs the extraction.
- schema found automatically
-
Our key insight is to generate many candidate functions and ensemble their extractions using weak supervision. EVAPORATE-CODE+ not only outperforms the state-of-the art systems, but does so using a sublinear pass over the documents with the LLM.
-
To reduce variance, we synthesize many candidate functions, then estimate their quality and aggregate their extractions using weak supervision. Described separately below:
-
- Tables2Traces: Distilling Tabular Data to Improve LLM Reasoning in Healthcare | OpenReview2
- Create training instances from tabular medical data
- Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy 3
- OCR errors taxonomy incl. interleaved columns, recover good text w/ finetuned LLama 1B
Structured I/O
Benchmarks
- [2604.25359v1] The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models 4 (8/10)
- Takeaway: often schema compliance but silently wrong values
- Neat overview of what’s wrong with other benchmarks
- Three modalities (image+audio)
- Thorough evaluation metrics, includes related work for e.g. constraint adherence or e.g. comparing JSONs
- [2505.20139] StructEval: Benchmarking LLMs’ Capabilities to Generate Structural Outputs 5 (10/10)
- generation (text-to-format) vs conversion tasks (format1-to-format2)
- A LOT of formats, text-based and visual-based ones
- text: JSON,CSV,TOML etc
- visual: SVG, latex, mergmaid, tikz, typst, vue, react etc.
- Evaluation
- for visual ones: VQA / ask a vision model whether the generated picture fullfills criterium.
- for text: dot-path
- Cool task generation prompt (“pick a super-creative and random domain”)
- Takeaways
- generation harder than conversion, vision harder than text
- [2411.19504] TQA-Bench: Evaluating LLMs for Multi-Table Question Answering6 (10/10)
- Tests multi-table reasoning with tables represented as Markdown, CSV, JSON, HTML
- Tasks like aggregation (count, sum avg), lookup, top-selection, correlation
- Results
- finds Markdown to be the best and JSON to be the worst for multi-table QA
- CSV/JSON’s relative ordering changes based on task type etc.
- Table-specific models underperform; distillation has a negative impact for long contexts
- Instruct+Code models the best, Chat-models consistently bad
- Multi-table reasoning is harder than single-table to the extent that merging 2 tables into 1 can improve scores
- finds Markdown to be the best and JSON to be the worst for multi-table QA
- Useful bits
- token budget defined by markdown representation and then converted to whatever (that will take more tokens)
- lists sources of relational databases, has cool algorithm to sample them
- [2501.10868v3] JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models7
- 10k JSON schemas from various sources
- Comparing constrained decoding frameworks, very very thorough
- Including efficiency (time to compile / first token / …)
- Evaluate coverage of features, compliance etc. — but not content
Surveys
- Large Language Models for Structured and Semi-Structured Data, Recommender Systems and Knowledge Base Engineering: A Survey of Recent Techniques and Architectures8 Section 5
- Lists many open topics
- Tables
Papers
- Using tables for training data
- Tables2Traces: Distilling Tabular Data to Improve LLM Reasoning in Healthcare | OpenReview2 GOOD
- Generate reasoning traces based on tables to leverage them as much better training data in a medical context
- Table-LLM-Specialist: Language Model Specialists for Tables using Iterative Fine-tuning - ACL Anthology11 tables can create tasks that are generative and classificative in nature, use that to create a Generator-Validator to create training data
- Tables2Traces: Distilling Tabular Data to Improve LLM Reasoning in Healthcare | OpenReview2 GOOD
- Models
- TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios12 LLama 8B finetuned for querying+editing tables
- SpreadsheetLLM: Encoding Spreadsheets for Large Language Models13 is GOOD
- Spreadsheets != tables
- Uses an inverted index (
bob: C1,C4,A0), homogeneous columns (e.g. all dates), automatically find the useful table bit in the sheet. Genuinely a lot of cool ideas. (Also SpreadsheetCompressor to save tokens)
- [2508.15910v1] Evaluating Structured Decoding for Text-to-Table Generation: Evidence from Three Datasets 14
- text-to-table
- discussion on limited suitability of usual NLP metrics on table level
- does neat postprocessing to find and check the markdown tables returned by the LLM
- Table Meets LLM: Can Large Language Models Understand Structured Table Data? A Benchmark and Empirical Study | Proceedings of the 17th ACM International Conference on Web Search and Data Mining 2024, simple table tasks, seems one of the first / important ones.
Etc.
- Blog posts by the same company on table formats/representations
- Which Table Format Do LLMs Understand Best? (Results for 11 Formats), includes list of formats and sample data
- only gpt4.1-nano
- ‘Markdown-KV’ won, XML was second, CSV was the second worst
- Which Nested Data Format Do LLMs Understand Best? JSON vs. YAML vs. XML vs. Markdown
- JSON, CSV, Markdown, YAML
- 3 models tested, results varied
- Which Table Format Do LLMs Understand Best? (Results for 11 Formats), includes list of formats and sample data
- OpenIE is IE that extracts all information from a document, without being provided any schema or vocabulary (compared to by EVAPORATE1)
IE / Document extraction
- [2607.29677] ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction 15 , which is basically “given a document and a user-defined schema, the agent faithfully follows the schema to produce the correct output with source evidence as grounding metadata” and this is evaluated. Really cool, including classifying documents by how exactly are they problematic
Tools
- Table extraction
Corrupted tables/columns/…
META
- Closely related: Conversation Disentanglement that does this for e.g. slack messages
Papers
- Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy 3 TODO GOOD
- Simulates PDF-to-text issues and trains a 1B LLama to correct them
- OCR error categorization: column-level, word-level, character level
- Includes column reading order as one scenario as well as character-level and word level
- Eval:
- document retrieval on their corrected text vs bad one
- TODO VQA?
Retrieval / RAG
- [2603.23047] Parametric Knowledge and Retrieval Behavior in RAG Fine-Tuning for Electronic Design Automation16
- Fine-tune a 7B Qwen2.5-7B-Instruct+LORA RAG
- different adapters
- triplet-based (response, context, query)
- long quotes
- Metrics for retrieval incl. using triplet-based (response, context, query)
- neat intro and related work
- Fine-tune a 7B Qwen2.5-7B-Instruct+LORA RAG
- When Good OCR Is Not Enough: Benchmarking OCR Robustness for Retrieval-Augmented Generation - ACL Anthology17
- Good OCR bad retrieval scenarios
-
We introduce an OCR benchmark for industrial RAG systems covering 11 challenging document types, including extreme layouts, high-resolution pages, complex or watermarked backgrounds, historical documents with non-standard reading orders, visually decorated text, and documents containing tables and mathematical formulas.
Venues
- This footnote from Section 2.1 of 18:
Consider, for example, the International Conference on Document Analysis and Recognition (ICDAR), the International Journal on Document Analysis and Recognition (IJDAR), the International Workshop on Document Analyis Systems (DAS), the International Workshop on Graphics Recognition (GREC), some of which host serial competitions, such as the Robust Reading Competition linking the document analysis and computer vision communities [Ka13; Ya17], cf. https://rrc.cvc.uab.es/.
- Robust Reading Competition does really neat competitions
“Robust Reading” refers to the research area dealing with the interpretation of written communication in unconstrained settings. // e.g. scans maps etc. + VQA
- Start | The ACM Symposium on Document Engineering
-
Document engineering is the computer science discipline that investigates systems for documents in any form and in all media. As with the relationship between software engineering and software, document engineering is concerned with principles, tools and processes that improve our ability to create, manage, and maintain documents.
- Proceedings, e.g. 2025: Proceedings of the 2025 ACM Symposium on Document Engineering | ACM Conferences
- Sessions on IR, classification, analysis
-
Less relevant but interesting
Tabular foundation models
TL;DR: pretrained on large tabular datasets, predict via in-context learning (dynamically adapting to a task without optimizing any parameters). Do classification regression etc.
- Tabular Foundation Models the WIP book on the basics.
- Models
Long tail of weird stuff
- Exotic formats
- NestedText — Structured Data for Humans YAML without typing and therefore simpler, validation expected to be done by reader
- TOON tries to use fewer tokens than JSON/TOML/… (but CSV wins for small datasets)
Angles
- Extraction
- cluster similar BBK docs (by issuer) and do sth like EVAPORATE1 to quickly extract data from them
- Data
- RAG over tables represented differently? (compare: 16)
- Tables (everything below as 0-shot or as finetuning)
- Tables:
- long form vs wide form transformations
- natural text
- either have column names in each example, or only once in table header — does this matter?
- augmented tables a la tables2traces2
- Multi-table reasoning with a focus on distance of the needed rows
- 6 says merging two tables into one improves scores — could that be purely a context/distance thing?
- Corner cases
- very long column names or cells
- very sparse table
- not just masked, but wrong column/table names
- bad extractions
- can LLM detect/un-destroy wrongly extracted tables?
- Tables:
- Finetuning
- Finetune on tasks like table-and-question-about-it/expected-answer
- bad extraction -> training on this bad data -> ??
- e.g. extraction forgot about a column, how much does this make everything worse?
- compare to humans: can they recover the original table content?
- bad extraction -> training on this bad data -> ??
- repetitions
- logs etc. have a lot of boilerplate and LLMs don’t like repetitions: does this make llm training worse?
- Finetune on tasks like table-and-question-about-it/expected-answer
- Corrupted columns
- Can a LLM reason over interlevaed columns?
- Does training on interleaved data matter?
Concretely
1. EVAPORATE-ish on BBK documents
- Cluster incoming .PDFs, apply known cheap per-cluster strategies -> cheap extraction of structured data
- Direct inspiration: EVAPORATE1
- Cluster similar BBK docs (eval: should cluster by Emittent)
- Create strategies for each cluster:
- Let LLM write code to extract from cluster
- Use light model to detect pages with tables, then use $extractor on them
- Do full LLM extraction
- Each new document either assigned to cluster, or becomes its own cluster if it’s different
- Continuous evaluation by feeding to most expensive model to see if cluster strategy still works
2. Test different table representations
- Both 0-shot and LoRA-finetuning of <9B models
- Play with:
- Long form vs wide form transformations
- Different representations (JSON, XML, CSV, Markdown, YAML, etc.)
- 0-shot JSON vs YAML is not novel but finetuning may be
- Augment table based on context (add units, expand short column names, make missing values explicit) 2
- Table to natural text -> table to training examples! 2
- Variation of this: RAG over different table formats
- EXTREMELY inspired by Oestreich et al ("Parametric Knowledge and Retrieval Behavior in RAG Fine-Tuning for Electronic Design Automation"16)
3. Bad/destroyed/… tables
- RQ1: Can LLMs parse/recover data from imperfect tables?
- Scenario: table extraction ignored a column now all values are shifted-by-1
- How much do clear good column names help?
- How much do similar data types help?
- SpreadsheetBench13 uses them as signal
- Connection to Infai’s obfuscation & RAG research
- Eval: can humans do this? + compare to LLM
- Scenario: table extraction ignored a column now all values are shifted-by-1
- RQ2: How much does training on wrong tables mess everything up?
4. LLMs & broken/interleaved columns
- Two PDF columns interleaved intead of correctly parsed during text extraction
- 0-shot:
- Can LLMs disentangle these columns? REVISE3 did this
- Reason over these columns?
- IE over these columns?
- Training:
- How much do such columns make training worse?
- By % of such bad data
- Can we enrich LLMs with such imperfect data?
-
“How much linguistic and factual information remains learnable when document serialization destroys discourse order but preserves the underlying tokens?” ChatGPT says paper title
-
- How much do such columns make training worse?
- Scenarios
- Two different languages (scripts?) in the columns
- Same lang, different domain/topic
- Same lang, same topic
- Same lang, same document
-
<_(@aroralanguagemodels2025) “Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes” (2025) / Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, Christopher Ré: http://arxiv.org/abs/2304.09433 /
10.48550/arXiv.2304.09433_> ↩︎ ↩︎ ↩︎ ↩︎ -
<_(@werling2026tablestraces) “Tables2Traces: Distilling tabular data to improve LLM reasoning in healthcare” (2026) / Mikkel Werling, Nabeel Seedat, Jiashuo Liu, Lars Grønlykke, Carsten Utoft Niemann, Mihaela vander Schaar, Rudi Agius: https://openreview.net/forum?id=cqNAjXUBOV / `` _> ↩︎ ↩︎ ↩︎ ↩︎ ↩︎
-
<_(@shimreviseframework2026) “Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy” (2026) / Gyuho Shim, Seongtae Hong, Heuiseok Lim: http://arxiv.org/abs/2604.08115 /
10.48550/arXiv.2604.08115_> ↩︎ ↩︎ ↩︎ -
<_(@singhstructuredoutput2026) “The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models” (2026) / Abhinav Kumar Singh, Harsha Vardhan Khurdula, Yoeven D. Khemlani, Vineet Agarwal: http://arxiv.org/abs/2604.25359 /
10.48550/arXiv.2604.25359_> ↩︎ -
<_(@yangstructevalbenchmarking2026) “StructEval: Benchmarking LLMs’ Capabilities to Generate Structural Outputs” (2026) / Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Zhiheng Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, Wenhu Chen: http://arxiv.org/abs/2505.20139 /
10.48550/arXiv.2505.20139_> ↩︎ -
<_(@qiutqabenchevaluating2026) “TQA-Bench: Evaluating LLMs for Multi-Table Question Answering” (2026) / Zipeng Qiu, Chenyue Li, You Peng, Guangxin He, Binhang Yuan, Chen Wang: http://arxiv.org/abs/2411.19504 /
10.48550/arXiv.2411.19504_> ↩︎ ↩︎ -
<_(@gengjsonschemabenchrigorous2025) “JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models” (2025) / Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, Harsha Nori: http://arxiv.org/abs/2501.10868 /
10.48550/arXiv.2501.10868_> ↩︎ -
<_(@smajiclargelanguage2025) “Large Language Models for Structured and Semi-Structured Data, Recommender Systems and Knowledge Base Engineering: A Survey of Recent Techniques and Architectures” (2025) / Alma Smajić, Ratomir Karlović, Mieta Bobanović Dasko, Ivan Lorencin: https://www.mdpi.com/2079-9292/14/15/3153 /
10.3390/electronics14153153_> ↩︎ -
<_(@lulargelanguage2024) “Large Language Model for Table Processing: A Survey” (2024) / Weizheng Lu, Jing Zhang, Ju Fan, Zihao Fu, Yueguo Chen, Xiaoyong Du: https://arxiv.org/abs/2402.05121 /
10.48550/ARXIV.2402.05121_> ↩︎ -
<_(@wutabulardata2025) “Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges” (2025) / Xiaofeng Wu, Alan Ritter, Wei Xu: https://arxiv.org/abs/2508.00217 /
10.48550/ARXIV.2508.00217_> ↩︎ -
<_(@xing-etal-2025-table) “Table-LLM-specialist: Language model specialists for tables using iterative fine-tuning” (2025) / Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Dongmei Zhang, Surajit Chaudhuri: https://aclanthology.org/2025.emnlp-main.1795/ /
10.18653/v1/2025.emnlp-main.1795_> ↩︎ -
<_(@zhangtablellmenabling2024) “TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios” (2024) / Xiaokang Zhang, Sijia Luo, Bohan Zhang, Zeyao Ma, Jing Zhang, Yang Li, Guanlin Li, Zijun Yao, Kangli Xu, Jinchang Zhou, Daniel Zhang-Li, Jifan Yu, Shu Zhao, Juanzi Li, Jie Tang: https://arxiv.org/abs/2403.19318 /
10.48550/ARXIV.2403.19318_> ↩︎ -
<_(@dongspreadsheetllmencoding2024) “SpreadsheetLLM: Encoding Spreadsheets for Large Language Models” (2024) / Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Shiyu Xia, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, Dongmei Zhang: https://arxiv.org/abs/2407.09025 /
10.48550/ARXIV.2407.09025_> ↩︎ ↩︎ -
<_(@oestreichevaluatingstructured2025) “Evaluating Structured Decoding for Text-to-Table Generation: Evidence from Three Datasets” (2025) / Julian Oestreich, Lydia Müller: https://arxiv.org/abs/2508.15910 /
10.48550/ARXIV.2508.15910_> ↩︎ -
<_(@zhangextractbenchbenchmark2026) “ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction” (2026) / Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo: http://arxiv.org/abs/2607.29677 /
10.48550/arXiv.2607.29677_> ↩︎ -
<_(@oestreichparametricknowledge2026) “Parametric Knowledge and Retrieval Behavior in RAG Fine-Tuning for Electronic Design Automation” (2026) / Julian Oestreich, Maximilian Bley, Frank Binder, Lydia Müller, Maksym Sydorenko, André Alcalde: http://arxiv.org/abs/2603.23047 /
10.48550/arXiv.2603.23047_> ↩︎ ↩︎ ↩︎ -
<_(@sun-etal-2026-good) “When good OCR is not enough: Benchmarking OCR robustness for retrieval-augmented generation” (2026) / Lin Sun, Wangdexian, Jingang Huang, Linglin Zhang, Change Jia, Zhengwei Cheng, Xiangzheng Zhang: https://aclanthology.org/2026.acl-industry.60/ /
10.18653/v1/2026.acl-industry.60_> ↩︎ -
<_(@perakisclassifyingfigures2023) “Classifying figures and illustrations in electronics datasheets: A comparative evaluation of recent computer vision models on a custom collection of 4000 technical documents” (2023) / Lymperis Perakis, Julian Balling, Frank Binder, Gerhard Heyer, Franz Kreupl: https://dl.gi.de/handle/20.500.12116/43114 /
10.18420/INF2023_186_> ↩︎