serhii.net

In the middle of the desert you can say anything you want

UNLISTED

29 Sep 2023

Text corpora deduplication

Problem happened in the context of a large-ish text corpus.

We know that [2107.06499] Deduplicating Training Data Makes Language Models Better, problem is how do we do that.

If the documents are separate it’s relatively easy:

Also: malteos/awesome-document-similarity: A curated list of resources on document similarity measures (papers, tutorials, code, …) - Large expert-curated database for benchmarking document similarity detection in biomedical literature search | Database | Oxford Academic

Problem: I’m dealing with basically single large documents that may contain 0…n copies of smaller ones inside.

Attempt 1: model it as longest repeating substring

pyalgs · PyPI has an implementation of LCS

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus