serhii.net

In the middle of the desert you can say anything you want

UNLISTED

25 Apr 2023

BBK handbook for doing ALL THE THINGS

DID THIS IN THE BRANCH

Plan

Content

  1. BBK do all things
  2. How to add criterium

Approach

  • copy a lot from READMEs

General

Installation

  • requirements.txt contains the needed packages; as usual, virtualenv etc.
  • You need (by default) spacy’s german model (you’ll get a readable error if you don’t have it): python -m spacy download de_core_news_lg
    • An Internet connection for huggingface to download model checkpoints is a good idea, but you can download them beforehand too.

Running stuff

  • Most things get run as modules, that is python3 -m anhaltai_bbk.your.module.name; for this, PYTHONPATH should point to the folder containing ./anhaltai_bbk . Either explicitly (PYTHONPATH=/home/sh/hsa/BBK/bundesbank-emissionspruefung/src python3 -m anhaltai_bbk.data.bio.multilabel_conv) or just by cd-ing to this folder before running the command.
  • Most things have a valid -help, and: -P for “drop to shell pdb on exception”.

Data conversion

Basic process

  1. Convert Konfuzio export dirs to RawDatasets
  2. Convert the RawDatasets first to binary BIO-format, then to multilabel BIO format
  3. Use that for training (under the hood: preprocessing)

Getting data from Konfuzio

… is out of scope for this guide, but previously the konf_converter was used for this:

python3 -m anhaltai_bbk.konf_to_hf.downloader -i 12270 -o . 

We updated the konfuzio_sdk multiple times since the last time it was used, so YMMV. For now, we assume to have the data downloaded as Konfuzio dataset.

Relevant:

Konfuzio-to-RawDataset conversion

RawDataset format

RawDataset / RawDocument / … is the main format we keep the data in.

{'documents': [
	{'document_id': 26026, # or name or whatever
	'document_filename': 'whatever.pdf', 
	'text': 'My name is Bob', # for Konfuzio's parsed text for the document
	'annotations': [
		{
			'tag': 'PERSON',
			'spans': 11, 14  # [begin, end); Inside a list to handle gaps later on
			'annotation_text': 'Bob',
			'bbox': {'x0': 123, 'y0': 234, ..}, # bounding box of envelope
			'bbox_spans': [{'x0': ..,}, ...], # bboxes of each span in 'spans'
			'confidence': 1.0, # 'None' posible, but 1.0 is set by default
		}, 
		{...},
		]
	}, 
	{
		// document 2 
	}
}

Our README about the format: src/anhaltai_bbk/konf_to_hf · main · Künstliche Intelligenz / Künstliche Intelligenz - Projekte / Bundesbank - Emissionsprüfung · GitLab

Conversion

TL;DR

TL;DR -i for input konfuzio project (or dir containing multiple K. projects!), -o for output directory, -S for creating test splits by annotator name. E.g. from within ./data/bbk, assuming ./0-konfuzio contains our three konfuzio projects:

PYTHONPATH=/home/sh/hsa/BBK/bundesbank-emissionspruefung/src python3 -m anhaltai_bbk.konf_to_hf.converter -i 0-konfuzio/  -o 1-raw -S 

This automatically remaps the annotation types defined in anhaltai_bbk/konf_to_hf/converter/annotation_mapper.py, splits by annotator name into test_XX splits, and puts Basisprospekte in a separate split.

(venv) 16:49:43 ~/hsa/BBK/bundesbank-emissionspruefung/data/bbk/ 130
> tree 1-raw
1-raw
├── __metadata.json
├── separated_Basisprospekte.json
├── test_JB.json
├── test_PR.json
├── test_RB.json
├── test_SG.json
└── train.json
Details
  • -i can also be a directory containing multiple Konfuzio projekte

Otherwise, the most important flags are:

--output-dir [OUTPUT_DIR], -o [OUTPUT_DIR]
					Target directory for the converted .json files.
					(/tmp/Konfuzio/converted)
--input INPUT, -i INPUT
					Directory with the downloaded Konfuzio data (the annotated data)
					(/tmp/Konfuzio/exported), or directory containing multiple such
					directories that will be dealt with as a single dataset.
--split-by-filename, -S
					Split by filenames/annotator names. Look for documents like
					'filename_AB.pdf' and 'filename_BC.pdf ',assume they are part of the
					test sets, and return N test sets, one for each suffix: 'AB:
					filename_AB,anotherfile_AB.pdf; 'BC: filename_BC.pdf, ...'
--split-ratio SPLIT_RATIO, -s SPLIT_RATIO
					Ratio of the TEST dataset (0.0<=x<1.0). A value of 0 disables
					splitting. Example: 0.25 would create a train/test split of 0.75/0.25
					respectively.
--no_map_annos        If set, will not remap anno types as required by BBK
  • -S: You almost always want to use -S, that splits by the dataset into train and test splits based on annotator name.
    • If you don’t, you can pass a float to -s / split-ratio to split the dataset to train/test
  • --no_map_annos if you don’t want the annotation remapping from anhaltai_bbk/konf_to_hf/converter/annotation_mapper.py (there’s a list here1)

RawDataset -> BIO Dataset conversion

The anhaltai_bbk.data README contains a more detailed description of everything in this section: src/anhaltai_bbk/data · main · Künstliche Intelligenz / Künstliche Intelligenz - Projekte / Bundesbank - Emissionsprüfung · GitLab

BIO-Dataset format(s)

This part is about converting from the offset-based RawDataset into one based on per-token BIO-like-tags2.

We have three main formats, of which only the last two are important:

  • non-binary BIO dataset (never used, probably not supported)
  • binary BIO dataset
  • multilabel BIO dataset 3 (the main one we use)

The different formats are different because of the problem of overlaps: for example, “Bob Smith” can be a PERSON, “Bob” can be a “first name” - but then the token “Bob” is both a “first name” and part of a PERSON, so we have overlapping types.

  • Non-binary BIO-format puts all annotation types in the same string, if there are any overlaps either crashes or overwrites (=losing some of the types).
  • Binary format would create two datasets: PERSON and FIRST_NAME, each having a different value for “Bob”. Then two different models would be trained.
  • Multilabel BIO-dataset, the newest and best one, is a single dataset with two different fields, one per annotation type. A single model is trained.

We used Binary datasets before, now use multilabel ones.

Conversion

First RawDataset to binary BIO, then binary BIO to multilabel BIO.

TL;DR
python3 -m anhaltai_bbk.data.bio -i ./1-raw -o ./3.5-binary -b -D 

-b for binary mode, -D for creating a train/dev split from train.

python3 -m anhaltai_bbk.data.bio.multilabel_conv -i ./3.5-binary -o ./3.7-multilabel

At the end of the prev. command, you should have the dataset you can directly use for training in your -output dir.

Details

First, BIOConverter.

  • BIOConverter in non-binary mode is not well documented, basically always use -b.
  • The -input can accept either a RawDataset .json, or a directory containing multiple such files. In the latter case, the filenames will be taken a split names. __metadata.json will be ignored.
  • -D for train/dev splitter: a WeightedStratifier, thoroughly documented in the anhaltai_bbk.data README, 4 is used. It splits the train split into a train and dev in such a way that the dev split contains a similar distribution of the various annotation types as train. Splitting is done per-document. (If train has 20 PERSON and 40 CURRENCY annotations, it would be split into a train of 15 PERSON and 30 CURRENCY, dev would have 5/10.) The intent is to make the some rare types be present both in train and in dev. Strongly recommended because training relies on a dev dataset for metrics and early stopping.
  • -T {yes,no} - whether the custom tokenizer is enabled (by default it is). It has special custom rules to better tokenize financial language.
  • -g GAP_MODE ways to fill the gaps. fill is the default good option for binary mode. There’s also skip that skips any annotations with gaps.
  • -t TAG_MODE: one of BIO or BILUO

You will see messages about skipped annotations etc., most of these are normal and happen when the annotation is misaligned to the tokens (an annotation would cut a token in the middle, etc.). Big part of this is taken care of by the custom tokenizer, but not all of them.

The Multilabel converter is much easier. It requires python 3.9 and converts binary BIO to multilabel BIO.

Training


  1. Annotations remapped in konf_converter:

    {
            "final_is_sonder_is_early_redemption_eligible": [
                "redemption_at_maturity_eligible",
                "sonderkuendigung_eligible",
                "early_redemption_eligible",
            ],
            "final_redemption_is_sonderkuendigung_eligible": [
                "redemption_at_maturity_eligible",
                "sonderkuendigung_eligible",
            ],
            "final_redemption_is_early_redemption_eligible": [
                "redemption_at_maturity_eligible",  # Same as in the prev. one!
                "early_redemption_eligible",
            ],
            "early_redemption_is_sonder_eligible": [
                "early_redemption_eligible",
                "sonderkuendigung_eligible"
            ]
        }
    	```
    
     ↩︎
  2. Linguistic Features · spaCy Usage Documentation ↩︎

  3. formerly “Gent-mode” ↩︎

  4. src/anhaltai_bbk/data · main · Künstliche Intelligenz / Künstliche Intelligenz - Projekte / Bundesbank - Emissionsprüfung · GitLab ↩︎

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus