serhii.net

In the middle of the desert you can say anything you want

23 Jul 2026

LLM (semi) structured training and inference

Document AI / IE / Unstructured

Getting things out of documents

Benchmarks

Unstructured and semi-structured

Key

  • EVAPORATE: [2304.09433] Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes 2
    • We propose and evaluate EVAPORATE […]. We identify two fundamentally different strategies for implementing this system: prompt the LLM to directly extract values from documents or prompt the LLM to synthesize code that performs the extraction.

    • schema found automatically
    • Our key insight is to generate many candidate functions and ensemble their extractions using weak supervision. EVAPORATE-CODE+ not only outperforms the state-of-the art systems, but does so using a sublinear pass over the documents with the LLM.

    • To reduce variance, we synthesize many candidate functions, then estimate their quality and aggregate their extractions using weak supervision.

Benchmarks

Surveys

Processing

Table formats/representations

Etc.

Etc.

Venues for relevant literature

Less relevant but interesting topics

Tabular foundation models

TL;DR: pretrained on large tabular datasets, predict via in-context learning (dynamically adapting to a task without optimizing any parameters). Do classification regression etc.

Long tail of weird stuff

TODO


  1. <_(@zhangextractbenchbenchmark2026) “ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction” (2026) / Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo: http://arxiv.org/abs/2607.29677 / 10.48550/arXiv.2607.29677 _> ↩︎

  2. <_(@aroralanguagemodels2025) “Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes” (2025) / Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, Christopher Ré: http://arxiv.org/abs/2304.09433 / 10.48550/arXiv.2304.09433 _> ↩︎

  3. <_(@singhstructuredoutput2026) “The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models” (2026) / Abhinav Kumar Singh, Harsha Vardhan Khurdula, Yoeven D. Khemlani, Vineet Agarwal: http://arxiv.org/abs/2604.25359 / 10.48550/arXiv.2604.25359 _> ↩︎

  4. <_(@wutabulardata2025) “Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges” (2025) / Xiaofeng Wu, Alan Ritter, Wei Xu: https://arxiv.org/abs/2508.00217 / 10.48550/ARXIV.2508.00217 _> ↩︎

  5. <_(@smajiclargelanguage2025) “Large Language Models for Structured and Semi-Structured Data, Recommender Systems and Knowledge Base Engineering: A Survey of Recent Techniques and Architectures” (2025) / Alma Smajić, Ratomir Karlović, Mieta Bobanović Dasko, Ivan Lorencin: https://www.mdpi.com/2079-9292/14/15/3153 / 10.3390/electronics14153153 _> ↩︎

  6. <_(@lulargelanguage2024) “Large Language Model for Table Processing: A Survey” (2024) / Weizheng Lu, Jing Zhang, Ju Fan, Zihao Fu, Yueguo Chen, Xiaoyong Du: https://arxiv.org/abs/2402.05121 / 10.48550/ARXIV.2402.05121 _> ↩︎

  7. <_(@zhangtablellmenabling2024) “TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios” (2024) / Xiaokang Zhang, Sijia Luo, Bohan Zhang, Zeyao Ma, Jing Zhang, Yang Li, Guanlin Li, Zijun Yao, Kangli Xu, Jinchang Zhou, Daniel Zhang-Li, Jifan Yu, Shu Zhao, Juanzi Li, Jie Tang: https://arxiv.org/abs/2403.19318 / 10.48550/ARXIV.2403.19318 _> ↩︎

  8. <_(@dongspreadsheetllmencoding2024) “SpreadsheetLLM: Encoding Spreadsheets for Large Language Models” (2024) / Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Shiyu Xia, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, Dongmei Zhang: https://arxiv.org/abs/2407.09025 / 10.48550/ARXIV.2407.09025 _> ↩︎

  9. <_(@oestreichevaluatingstructured2025) “Evaluating Structured Decoding for Text-to-Table Generation: Evidence from Three Datasets” (2025) / Julian Oestreich, Lydia Müller: https://arxiv.org/abs/2508.15910 / 10.48550/ARXIV.2508.15910 _> ↩︎

  10. <_(@perakisclassifyingfigures2023) “Classifying figures and illustrations in electronics datasheets: A comparative evaluation of recent computer vision models on a custom collection of 4000 technical documents” (2023) / Lymperis Perakis, Julian Balling, Frank Binder, Gerhard Heyer, Franz Kreupl: https://dl.gi.de/handle/20.500.12116/43114 / 10.18420/INF2023_186 _> ↩︎

  11. <_(@oestreichparametricknowledge2026) “Parametric Knowledge and Retrieval Behavior in RAG Fine-Tuning for Electronic Design Automation” (2026) / Julian Oestreich, Maximilian Bley, Frank Binder, Lydia Müller, Maksym Sydorenko, André Alcalde: http://arxiv.org/abs/2603.23047 / 10.48550/arXiv.2603.23047 _> ↩︎

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus