BBK handbook for doing ALL THE THINGS
DID THIS IN THE BRANCH
Plan
Content
- BBK do all things
- How to add criterium
Approach
- copy a lot from READMEs
General
Installation
requirements.txtcontains the needed packages; as usual, virtualenv etc.- You need (by default) spacy’s german model (you’ll get a readable error if you don’t have it):
python -m spacy download de_core_news_lg- An Internet connection for huggingface to download model checkpoints is a good idea, but you can download them beforehand too.
Running stuff
- Most things get run as modules, that is
python3 -m anhaltai_bbk.your.module.name; for this,PYTHONPATHshould point to the folder containing./anhaltai_bbk. Either explicitly (PYTHONPATH=/home/sh/hsa/BBK/bundesbank-emissionspruefung/src python3 -m anhaltai_bbk.data.bio.multilabel_conv) or just bycd-ing to this folder before running the command. - Most things have a valid
-help, and:-Pfor “drop to shell pdb on exception”.
Data conversion
Basic process
- Convert Konfuzio export dirs to RawDatasets
- Convert the RawDatasets first to binary BIO-format, then to multilabel BIO format
- Use that for training (under the hood: preprocessing)
Getting data from Konfuzio
… is out of scope for this guide, but previously the konf_converter was used for this:
python3 -m anhaltai_bbk.konf_to_hf.downloader -i 12270 -o .
We updated the konfuzio_sdk multiple times since the last time it was used, so YMMV. For now, we assume to have the data downloaded as Konfuzio dataset.
Relevant:
- Konfuzio Data Layer: Explanations — Konfuzio documentation
Konfuzio-to-RawDataset conversion
RawDataset format
RawDataset / RawDocument / … is the main format we keep the data in.
{'documents': [
{'document_id': 26026, # or name or whatever
'document_filename': 'whatever.pdf',
'text': 'My name is Bob', # for Konfuzio's parsed text for the document
'annotations': [
{
'tag': 'PERSON',
'spans': 11, 14 # [begin, end); Inside a list to handle gaps later on
'annotation_text': 'Bob',
'bbox': {'x0': 123, 'y0': 234, ..}, # bounding box of envelope
'bbox_spans': [{'x0': ..,}, ...], # bboxes of each span in 'spans'
'confidence': 1.0, # 'None' posible, but 1.0 is set by default
},
{...},
]
},
{
// document 2
}
}
Our README about the format: src/anhaltai_bbk/konf_to_hf · main · Künstliche Intelligenz / Künstliche Intelligenz - Projekte / Bundesbank - Emissionsprüfung · GitLab
Conversion
TL;DR
TL;DR -i for input konfuzio project (or dir containing multiple K. projects!), -o for output directory, -S for creating test splits by annotator name. E.g. from within ./data/bbk, assuming ./0-konfuzio contains our three konfuzio projects:
PYTHONPATH=/home/sh/hsa/BBK/bundesbank-emissionspruefung/src python3 -m anhaltai_bbk.konf_to_hf.converter -i 0-konfuzio/ -o 1-raw -S
This automatically remaps the annotation types defined in anhaltai_bbk/konf_to_hf/converter/annotation_mapper.py, splits by annotator name into test_XX splits, and puts Basisprospekte in a separate split.
(venv) 16:49:43 ~/hsa/BBK/bundesbank-emissionspruefung/data/bbk/ 130
> tree 1-raw
1-raw
├── __metadata.json
├── separated_Basisprospekte.json
├── test_JB.json
├── test_PR.json
├── test_RB.json
├── test_SG.json
└── train.json
Details
-ican also be a directory containing multiple Konfuzio projekte
Otherwise, the most important flags are:
--output-dir [OUTPUT_DIR], -o [OUTPUT_DIR]
Target directory for the converted .json files.
(/tmp/Konfuzio/converted)
--input INPUT, -i INPUT
Directory with the downloaded Konfuzio data (the annotated data)
(/tmp/Konfuzio/exported), or directory containing multiple such
directories that will be dealt with as a single dataset.
--split-by-filename, -S
Split by filenames/annotator names. Look for documents like
'filename_AB.pdf' and 'filename_BC.pdf ',assume they are part of the
test sets, and return N test sets, one for each suffix: 'AB:
filename_AB,anotherfile_AB.pdf; 'BC: filename_BC.pdf, ...'
--split-ratio SPLIT_RATIO, -s SPLIT_RATIO
Ratio of the TEST dataset (0.0<=x<1.0). A value of 0 disables
splitting. Example: 0.25 would create a train/test split of 0.75/0.25
respectively.
--no_map_annos If set, will not remap anno types as required by BBK
-S: You almost always want to use -S, that splits by the dataset into train and test splits based on annotator name.- If you don’t, you can pass a float to
-s/ split-ratio to split the dataset to train/test
- If you don’t, you can pass a float to
--no_map_annosif you don’t want the annotation remapping fromanhaltai_bbk/konf_to_hf/converter/annotation_mapper.py(there’s a list here1)
RawDataset -> BIO Dataset conversion
The anhaltai_bbk.data README contains a more detailed description of everything in this section: src/anhaltai_bbk/data · main · Künstliche Intelligenz / Künstliche Intelligenz - Projekte / Bundesbank - Emissionsprüfung · GitLab
BIO-Dataset format(s)
This part is about converting from the offset-based RawDataset into one based on per-token BIO-like-tags2.
We have three main formats, of which only the last two are important:
- non-binary BIO dataset (never used, probably not supported)
- binary BIO dataset
- multilabel BIO dataset 3 (the main one we use)
The different formats are different because of the problem of overlaps: for example, “Bob Smith” can be a PERSON, “Bob” can be a “first name” - but then the token “Bob” is both a “first name” and part of a PERSON, so we have overlapping types.
- Non-binary BIO-format puts all annotation types in the same string, if there are any overlaps either crashes or overwrites (=losing some of the types).
- Binary format would create two datasets: PERSON and FIRST_NAME, each having a different value for “Bob”. Then two different models would be trained.
- Multilabel BIO-dataset, the newest and best one, is a single dataset with two different fields, one per annotation type. A single model is trained.
We used Binary datasets before, now use multilabel ones.
Conversion
First RawDataset to binary BIO, then binary BIO to multilabel BIO.
TL;DR
python3 -m anhaltai_bbk.data.bio -i ./1-raw -o ./3.5-binary -b -D
-b for binary mode, -D for creating a train/dev split from train.
python3 -m anhaltai_bbk.data.bio.multilabel_conv -i ./3.5-binary -o ./3.7-multilabel
At the end of the prev. command, you should have the dataset you can directly use for training in your -output dir.
Details
First, BIOConverter.
- BIOConverter in non-binary mode is not well documented, basically always use
-b. - The
-input can accept either a RawDataset .json, or a directory containing multiple such files. In the latter case, the filenames will be taken a split names.__metadata.jsonwill be ignored. -Dfor train/dev splitter: aWeightedStratifier, thoroughly documented in theanhaltai_bbk.dataREADME, 4 is used. It splits thetrainsplit into atrainanddevin such a way that thedevsplit contains a similar distribution of the various annotation types astrain. Splitting is done per-document. (If train has 20 PERSON and 40 CURRENCY annotations, it would be split into atrainof 15 PERSON and 30 CURRENCY,devwould have 5/10.) The intent is to make the some rare types be present both in train and in dev. Strongly recommended because training relies on adevdataset for metrics and early stopping.-T {yes,no}- whether the custom tokenizer is enabled (by default it is). It has special custom rules to better tokenize financial language.-g GAP_MODEways to fill the gaps.fillis the default good option for binary mode. There’s alsoskipthat skips any annotations with gaps.-t TAG_MODE: one ofBIOorBILUO
You will see messages about skipped annotations etc., most of these are normal and happen when the annotation is misaligned to the tokens (an annotation would cut a token in the middle, etc.). Big part of this is taken care of by the custom tokenizer, but not all of them.
The Multilabel converter is much easier. It requires python 3.9 and converts binary BIO to multilabel BIO.
Training
-
Annotations remapped in konf_converter:
↩︎{ "final_is_sonder_is_early_redemption_eligible": [ "redemption_at_maturity_eligible", "sonderkuendigung_eligible", "early_redemption_eligible", ], "final_redemption_is_sonderkuendigung_eligible": [ "redemption_at_maturity_eligible", "sonderkuendigung_eligible", ], "final_redemption_is_early_redemption_eligible": [ "redemption_at_maturity_eligible", # Same as in the prev. one! "early_redemption_eligible", ], "early_redemption_is_sonder_eligible": [ "early_redemption_eligible", "sonderkuendigung_eligible" ] } ``` -
formerly “Gent-mode” ↩︎
-
src/anhaltai_bbk/data · main · Künstliche Intelligenz / Künstliche Intelligenz - Projekte / Bundesbank - Emissionsprüfung · GitLab ↩︎