serhii.net

In the middle of the desert you can say anything you want

UNLISTED

23 Jul 2026

LLM (semi) structured training and inference

Key

Structured I/O

Benchmarks

  • [2604.25359v1] The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models 4 (8/10)
    • Takeaway: often schema compliance but silently wrong values
    • Neat overview of what’s wrong with other benchmarks
    • Three modalities (image+audio)
    • Thorough evaluation metrics, includes related work for e.g. constraint adherence or e.g. comparing JSONs
  • [2505.20139] StructEval: Benchmarking LLMs’ Capabilities to Generate Structural Outputs 5 (10/10)
    • generation (text-to-format) vs conversion tasks (format1-to-format2)
    • A LOT of formats, text-based and visual-based ones
      • text: JSON,CSV,TOML etc
      • visual: SVG, latex, mergmaid, tikz, typst, vue, react etc.
    • Evaluation
      • for visual ones: VQA / ask a vision model whether the generated picture fullfills criterium.
      • for text: dot-path
    • Cool task generation prompt (“pick a super-creative and random domain”)
    • Takeaways
      • generation harder than conversion, vision harder than text
  • [2411.19504] TQA-Bench: Evaluating LLMs for Multi-Table Question Answering6 (10/10)
    • Tests multi-table reasoning with tables represented as Markdown, CSV, JSON, HTML
    • Tasks like aggregation (count, sum avg), lookup, top-selection, correlation
    • Results
      • finds Markdown to be the best and JSON to be the worst for multi-table QA
        • CSV/JSON’s relative ordering changes based on task type etc.
      • Table-specific models underperform; distillation has a negative impact for long contexts
      • Instruct+Code models the best, Chat-models consistently bad
      • Multi-table reasoning is harder than single-table to the extent that merging 2 tables into 1 can improve scores
    • Useful bits
      • token budget defined by markdown representation and then converted to whatever (that will take more tokens)
      • lists sources of relational databases, has cool algorithm to sample them
  • [2501.10868v3] JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models7
    • 10k JSON schemas from various sources
    • Comparing constrained decoding frameworks, very very thorough
      • Including efficiency (time to compile / first token / …)
      • Evaluate coverage of features, compliance etc. — but not content

Surveys

Papers

Etc.

IE / Document extraction

Tools

Corrupted tables/columns/…

META

  • Closely related: Conversation Disentanglement that does this for e.g. slack messages

Papers

Retrieval / RAG

Venues

Less relevant but interesting

Tabular foundation models

TL;DR: pretrained on large tabular datasets, predict via in-context learning (dynamically adapting to a task without optimizing any parameters). Do classification regression etc.

Long tail of weird stuff


Angles

  • Extraction
    • cluster similar BBK docs (by issuer) and do sth like EVAPORATE1 to quickly extract data from them
  • Data
    • RAG over tables represented differently? (compare: 16)
  • Tables (everything below as 0-shot or as finetuning)
    • Tables:
      • long form vs wide form transformations
      • natural text
      • either have column names in each example, or only once in table header — does this matter?
      • augmented tables a la tables2traces2
    • Multi-table reasoning with a focus on distance of the needed rows
      • 6 says merging two tables into one improves scores — could that be purely a context/distance thing?
    • Corner cases
      • very long column names or cells
      • very sparse table
      • not just masked, but wrong column/table names
    • bad extractions
      • can LLM detect/un-destroy wrongly extracted tables?
  • Finetuning
    • Finetune on tasks like table-and-question-about-it/expected-answer
      • bad extraction -> training on this bad data -> ??
        • e.g. extraction forgot about a column, how much does this make everything worse?
        • compare to humans: can they recover the original table content?
    • repetitions
      • logs etc. have a lot of boilerplate and LLMs don’t like repetitions: does this make llm training worse?
  • Corrupted columns
    • Can a LLM reason over interlevaed columns?
    • Does training on interleaved data matter?

Concretely

1. EVAPORATE-ish on BBK documents

  • Cluster incoming .PDFs, apply known cheap per-cluster strategies -> cheap extraction of structured data
    • Direct inspiration: EVAPORATE1
  • Cluster similar BBK docs (eval: should cluster by Emittent)
  • Create strategies for each cluster:
    • Let LLM write code to extract from cluster
    • Use light model to detect pages with tables, then use $extractor on them
    • Do full LLM extraction
  • Each new document either assigned to cluster, or becomes its own cluster if it’s different
  • Continuous evaluation by feeding to most expensive model to see if cluster strategy still works

2. Test different table representations

  • Both 0-shot and LoRA-finetuning of <9B models
  • Play with:
    • Long form vs wide form transformations
    • Different representations (JSON, XML, CSV, Markdown, YAML, etc.)
      • 0-shot JSON vs YAML is not novel but finetuning may be
    • Augment table based on context (add units, expand short column names, make missing values explicit) 2
    • Table to natural text -> table to training examples! 2

3. Bad/destroyed/… tables

  • RQ1: Can LLMs parse/recover data from imperfect tables?
    • Scenario: table extraction ignored a column now all values are shifted-by-1
      • How much do clear good column names help?
      • How much do similar data types help?
        • SpreadsheetBench13 uses them as signal
    • Connection to Infai’s obfuscation & RAG research
    • Eval: can humans do this? + compare to LLM
  • RQ2: How much does training on wrong tables mess everything up?

4. LLMs & broken/interleaved columns

  • Two PDF columns interleaved intead of correctly parsed during text extraction
  • 0-shot:
    • Can LLMs disentangle these columns? REVISE3 did this
    • Reason over these columns?
    • IE over these columns?
  • Training:
    • How much do such columns make training worse?
      • By % of such bad data
    • Can we enrich LLMs with such imperfect data?
      • “How much linguistic and factual information remains learnable when document serialization destroys discourse order but preserves the underlying tokens?” ChatGPT says paper title

  • Scenarios
    • Two different languages (scripts?) in the columns
    • Same lang, different domain/topic
    • Same lang, same topic
    • Same lang, same document

  1. <_(@aroralanguagemodels2025) “Language Models Enable Simple Systems for Generating Structured Views of Heterogeneous Data Lakes” (2025) / Simran Arora, Brandon Yang, Sabri Eyuboglu, Avanika Narayan, Andrew Hojel, Immanuel Trummer, Christopher Ré: http://arxiv.org/abs/2304.09433 / 10.48550/arXiv.2304.09433 _> ↩︎ ↩︎ ↩︎ ↩︎

  2. <_(@werling2026tablestraces) “Tables2Traces: Distilling tabular data to improve LLM reasoning in healthcare” (2026) / Mikkel Werling, Nabeel Seedat, Jiashuo Liu, Lars Grønlykke, Carsten Utoft Niemann, Mihaela vander Schaar, Rudi Agius: https://openreview.net/forum?id=cqNAjXUBOV / `` _> ↩︎ ↩︎ ↩︎ ↩︎ ↩︎

  3. <_(@shimreviseframework2026) “Revise: A Framework for Revising OCRed text in Practical Information Systems with Data Contamination Strategy” (2026) / Gyuho Shim, Seongtae Hong, Heuiseok Lim: http://arxiv.org/abs/2604.08115 / 10.48550/arXiv.2604.08115 _> ↩︎ ↩︎ ↩︎

  4. <_(@singhstructuredoutput2026) “The Structured Output Benchmark: A Multi-Source Benchmark for Evaluating Structured Output Quality in Large Language Models” (2026) / Abhinav Kumar Singh, Harsha Vardhan Khurdula, Yoeven D. Khemlani, Vineet Agarwal: http://arxiv.org/abs/2604.25359 / 10.48550/arXiv.2604.25359 _> ↩︎

  5. <_(@yangstructevalbenchmarking2026) “StructEval: Benchmarking LLMs’ Capabilities to Generate Structural Outputs” (2026) / Jialin Yang, Dongfu Jiang, Lipeng He, Sherman Siu, Yuxuan Zhang, Disen Liao, Zhuofeng Li, Huaye Zeng, Yiming Jia, Haozhe Wang, Benjamin Schneider, Chi Ruan, Wentao Ma, Zhiheng Lyu, Yifei Wang, Yi Lu, Quy Duc Do, Ziyan Jiang, Ping Nie, Wenhu Chen: http://arxiv.org/abs/2505.20139 / 10.48550/arXiv.2505.20139 _> ↩︎

  6. <_(@qiutqabenchevaluating2026) “TQA-Bench: Evaluating LLMs for Multi-Table Question Answering” (2026) / Zipeng Qiu, Chenyue Li, You Peng, Guangxin He, Binhang Yuan, Chen Wang: http://arxiv.org/abs/2411.19504 / 10.48550/arXiv.2411.19504 _> ↩︎ ↩︎

  7. <_(@gengjsonschemabenchrigorous2025) “JSONSchemaBench: A Rigorous Benchmark of Structured Outputs for Language Models” (2025) / Saibo Geng, Hudson Cooper, Michał Moskal, Samuel Jenkins, Julian Berman, Nathan Ranchin, Robert West, Eric Horvitz, Harsha Nori: http://arxiv.org/abs/2501.10868 / 10.48550/arXiv.2501.10868 _> ↩︎

  8. <_(@smajiclargelanguage2025) “Large Language Models for Structured and Semi-Structured Data, Recommender Systems and Knowledge Base Engineering: A Survey of Recent Techniques and Architectures” (2025) / Alma Smajić, Ratomir Karlović, Mieta Bobanović Dasko, Ivan Lorencin: https://www.mdpi.com/2079-9292/14/15/3153 / 10.3390/electronics14153153 _> ↩︎

  9. <_(@lulargelanguage2024) “Large Language Model for Table Processing: A Survey” (2024) / Weizheng Lu, Jing Zhang, Ju Fan, Zihao Fu, Yueguo Chen, Xiaoyong Du: https://arxiv.org/abs/2402.05121 / 10.48550/ARXIV.2402.05121 _> ↩︎

  10. <_(@wutabulardata2025) “Tabular Data Understanding with LLMs: A Survey of Recent Advances and Challenges” (2025) / Xiaofeng Wu, Alan Ritter, Wei Xu: https://arxiv.org/abs/2508.00217 / 10.48550/ARXIV.2508.00217 _> ↩︎

  11. <_(@xing-etal-2025-table) “Table-LLM-specialist: Language model specialists for tables using iterative fine-tuning” (2025) / Junjie Xing, Yeye He, Mengyu Zhou, Haoyu Dong, Shi Han, Dongmei Zhang, Surajit Chaudhuri: https://aclanthology.org/2025.emnlp-main.1795/ / 10.18653/v1/2025.emnlp-main.1795 _> ↩︎

  12. <_(@zhangtablellmenabling2024) “TableLLM: Enabling Tabular Data Manipulation by LLMs in Real Office Usage Scenarios” (2024) / Xiaokang Zhang, Sijia Luo, Bohan Zhang, Zeyao Ma, Jing Zhang, Yang Li, Guanlin Li, Zijun Yao, Kangli Xu, Jinchang Zhou, Daniel Zhang-Li, Jifan Yu, Shu Zhao, Juanzi Li, Jie Tang: https://arxiv.org/abs/2403.19318 / 10.48550/ARXIV.2403.19318 _> ↩︎

  13. <_(@dongspreadsheetllmencoding2024) “SpreadsheetLLM: Encoding Spreadsheets for Large Language Models” (2024) / Haoyu Dong, Jianbo Zhao, Yuzhang Tian, Junyu Xiong, Shiyu Xia, Mengyu Zhou, Yun Lin, José Cambronero, Yeye He, Shi Han, Dongmei Zhang: https://arxiv.org/abs/2407.09025 / 10.48550/ARXIV.2407.09025 _> ↩︎ ↩︎

  14. <_(@oestreichevaluatingstructured2025) “Evaluating Structured Decoding for Text-to-Table Generation: Evidence from Three Datasets” (2025) / Julian Oestreich, Lydia Müller: https://arxiv.org/abs/2508.15910 / 10.48550/ARXIV.2508.15910 _> ↩︎

  15. <_(@zhangextractbenchbenchmark2026) “ExtractBench: A Benchmark for Schema-Guided Enterprise Document Extraction” (2026) / Boyang Zhang, Adrian Lyjak, Eli Stewart, Zhaoqi Li, Simon Suo: http://arxiv.org/abs/2607.29677 / 10.48550/arXiv.2607.29677 _> ↩︎

  16. <_(@oestreichparametricknowledge2026) “Parametric Knowledge and Retrieval Behavior in RAG Fine-Tuning for Electronic Design Automation” (2026) / Julian Oestreich, Maximilian Bley, Frank Binder, Lydia Müller, Maksym Sydorenko, André Alcalde: http://arxiv.org/abs/2603.23047 / 10.48550/arXiv.2603.23047 _> ↩︎ ↩︎ ↩︎

  17. <_(@sun-etal-2026-good) “When good OCR is not enough: Benchmarking OCR robustness for retrieval-augmented generation” (2026) / Lin Sun, Wangdexian, Jingang Huang, Linglin Zhang, Change Jia, Zhengwei Cheng, Xiangzheng Zhang: https://aclanthology.org/2026.acl-industry.60/ / 10.18653/v1/2026.acl-industry.60 _> ↩︎

  18. <_(@perakisclassifyingfigures2023) “Classifying figures and illustrations in electronics datasheets: A comparative evaluation of recent computer vision models on a custom collection of 4000 technical documents” (2023) / Lymperis Perakis, Julian Balling, Frank Binder, Gerhard Heyer, Franz Kreupl: https://dl.gi.de/handle/20.500.12116/43114 / 10.18420/INF2023_186 _> ↩︎

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus