Text corpora deduplication
Problem happened in the context of a large-ish text corpus.
We know that [2107.06499] Deduplicating Training Data Makes Language Models Better, problem is how do we do that.
If the documents are separate it’s relatively easy:
- probabilistic estimation of metrics: ekzhu/datasketch: MinHash, LSH, LSH Forest, Weighted MinHash, HyperLogLog, HyperLogLog++, LSH Ensemble and HNSW
- looks like a python-based implementation of a lot of algos: text-dedup · PyPI
- documentation for the older version has easier examples: text-dedup · PyPI
- text-dedup/benchmarks/wiki40.ipynb at main · ChenghaoMou/text-dedup
- relevant: w-shingling - Wikipedia
Also: malteos/awesome-document-similarity: A curated list of resources on document similarity measures (papers, tutorials, code, …) - Large expert-curated database for benchmarking document similarity detection in biomedical literature search | Database | Oxford Academic
Problem: I’m dealing with basically single large documents that may contain 0…n copies of smaller ones inside.
Attempt 1: model it as longest repeating substring
pyalgs · PyPI has an implementation of LCS
Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus