TWM 3 - vector models
Vector stuff - theory
Bag of words
Mathematically - it’s a vector that contains 0 (absent) or 1 (present).
(D) Termgewicht = “Wert eines Vektors in einer Dimension = Stärke der Assoziation zwischen Term und Dokument”
Termgewichte
- Vector (=column) of size $d_{|dict|}$
- Relevanz eines Termes
BoolEsche Termgewichte
- (L) Boolesch!
- Einfachste ${0,1}^{|dict|}$
- Vorhandensein oder Fehlen eines Terms
Termhäufigkeitsgewichtung
- the more often we see it, the more important it is
- e.g. just counts
Statistical ways
- We care about the middle part of the Zipf Verteilung - not the “the"s and not the long tail
- Ways to get it:
- Filter during preprocessing
- Weight correspondingly
IDF and friends
IDF - Inverse Document Frequency
-
High -> rare term, low -> frequent
-
(D) Dokumenthäufigkeit $df$ -> number of documents, where we see the word.
Inverse Dockumentenhäufigkeit eines terms: $$idf_t = log(\frac{N}{df_t})$$ (where N is the total number of documents)
So - $log(\frac{10.000}{1})=4$ means document found once in 10k docs. - $log(10.000/10.000)=0$ - found in each document.
Wikipedia: IDF is a measure of how much information the word provides (’the’ provides almost none)
TF-IDF
- (D) TF-IDF: Combination of Term Frequency and IDF
$$w_{t,d} = tf_{t,d}*ids_t$$
High weight for terms that:
- Found often in THIS document
- Rare in the other documents
TF
- Term frequency - how often in document, a lot of schemas possible. Natural (count), binary (0/1), log., etc.
- $freq_{t,d}$
Termgewichtungsschema
- Termgewichtung, generalized (Sl.19)
- Three bits
- Termhäufigkeit $freq_{t,d}$
- Dokumentenhäufigkeit $collect_t$
- Normierung $norm$
$$w_{t,d} = \frac{freq_{t,d}*collect_t}{norm}$$
Varianten
- Termhäufigkeit
- natural,binary,log,etc.
- Dokumentenhäufigkeit $collect_t$ (Sl.21)
- 1 (none)
- $log\frac{N}{df_t}$ invers (=IDF)
- probabilistic etc.
- Normierung
- 1 / keine
- $\sum_i^1w_{i,d}$
TODO: what’s the logic in i,d in norm options?
Vector stuff - practice
Saving space
- Don’t save the 0s to save space, just save word-1,word2-2,word5-3 etc.
Libraries
- Gensim for text mining
- sklearn for everything
Similarity
- Sim. of docs == sim of their vectors
- can measure this vector similarity in a lot of ways
- Similarity <-> distance
- $sim(d_1,d_2) = \frac{1}{1+dist(d_1,d_2)}$
- Similarity is between $[0;1]$
- Distance is between $[0;\infty]$
Euc. distance
-
$\sqrt{\sum^n_{i=1}(d_{1,i}-d_{2,i})^2}$
-
- Problem with the last one: CD is close, BC is far, but we don’t see it
(Sl.35)
-
Problems:
- TODO
- length of vectors influences the distance
- document pasted after itself is now totally different
- presence and absence of terms are gleichbehandelt
- TODO ?
-
Solution: use the angle instead of Euc.
-
Normalization to length 1 helps!
Problems with vectors
Lab3 solution
3-1
Dokument 1: im garten geben es viel blume und baum Dokument 2: im park sein auch baum aber auch strauch Dokument 3: anna mögen baum und strauch und blume Dokument 4: sie sein gern im garten und park und sie gehen gern spazieren und walken
Berechnen Sie den euklidischen Abstand und die Kosinusähnlichkeit zwischen Dokument 1 und 3 für die Häufigkeitsvektoren und die normalisierten TF-IDF-Vektoren.
Wörterbuch
Geben Sie das (alphabetisch sortierte) Wörterbuch an! Dokumentvorkommen
[('aber', 1),
('anna', 1),
('auch', 1), // ERROR! was 2 (Dokumentvorkommen, nicht W.)
('baum', 3),
('blume', 2),
('es', 1),
('garten', 2),
('geben', 1),
('gehen', 1),
('gern', 1), // Error: war 2
('im', 3),
('mögen', 1),
('park', 2),
('sein', 2),
('sie', 1), // ERROR! War 2
('spazieren', 1),
('strauch', 2),
('und', 3), // ERROR! Dockumentvorkommen! (was wf: 6)
('viel', 1),
('walken', 1)]
Häufigkeitvektoren, TF-IDSF and normalized TF-IDFs
Geben Sie die Häufigkeitsvektoren, die TF-IDF-Vektoren sowie die normalisierten TF-IDF-Vektoren für die Dokumente 1 und 3 an, und zwar sowohl in vollständiger als auch in spärlicher Form!
Doc1
Dokument 1: im garten geben es viel blume und baum
TF
Natürlich, binär:
im: 1
garten: 1
geben: 1
es: 1
viel: 1
blume: 1
und: 1
baum: 1
IDF
IDF sentence 1 im, 0.28768207245178085 garten, 0.6931471805599453 geben, 1.3862943611198906 es, 1.3862943611198906 viel, 1.3862943611198906 blume, 0.6931471805599453 und, 0.28768207245178085 baum, 0.28768207245178085
TF-IDF
im: log(10.29) -> 0.29 garten: log(10.69) -> 0.69 geben: log(11.39) -> 1.39 es: log(11.39) -> 1.39 viel: log(11.39) -> 1.39 blume: log(10.69) -> 0.69 und: log(10.29) -> 0.29 baum: log(10.29) -> 0.29
Doc3
Dokument 3: anna mögen baum und strauch und blume
TF: anna: 1 mögen: 1 baum: 1 und: 2 strauch: 1 blume: 1
TF-IDF: anna: log(11.39) -> 1.39 mögen: log(11.39) -> 1.39 baum: log(10.29) -> 0.29 und: log(20.29) -> 0.58 strauch: log(10.69) -> 0.69 und: log(20.29) -> 0.58 blume: log(1*0.69) -> 0.69
j,
VECTORS - ANSWER
Sentence 1
== NICHT NORMIERT == TF-IDF - sparlich [(0, 0.28768207245178085), (1, 0.6931471805599453), (2, 1.3862943611198906), (3, 1.3862943611198906), (4, 1.3862943611198906), (5, 0.6931471805599453), (6, 0.28768207245178085), (7, 0.28768207245178085)]
TF-IDF - vollständig [(0, 0.28768207245178085), (1, 0.6931471805599453), (2, 1.3862943611198906), (3, 1.3862943611198906), (4, 1.3862943611198906), (5, 0.6931471805599453), (6, 0.28768207245178085), (7, 0.28768207245178085), (9, 0), (10, 0), (11, 0), (13, 0), (15, 0), (16, 0), (17, 0), (23, 0), (25, 0), (32, 0), (34, 0), (36, 0)]
== NORMIERT == TF-IDF - sparlich [(0, 0.10893121906625666), (1, 0.2624611493070678), (2, 0.5249222986141356), (3, 0.5249222986141356), (4, 0.5249222986141356), (5, 0.2624611493070678), (6, 0.10893121906625666), (7, 0.10893121906625666)]
TF-IDF - vollständig [(0, 0.10893121906625666), (1, 0.2624611493070678), (2, 0.5249222986141356), (3, 0.5249222986141356), (4, 0.5249222986141356), (5, 0.2624611493070678), (6, 0.10893121906625666), (7, 0.10893121906625666), (9, 0.0), (10, 0.0), (11, 0.0), (13, 0.0), (15, 0.0), (16, 0.0), (17, 0.0), (23, 0.0), (25, 0.0), (32, 0.0), (34, 0.0), (36, 0.0)]
Sentence 3
== NICHT NORMIERT == TF-IDF - sparlich [(16, 1.3862943611198906), (17, 1.3862943611198906), (7, 0.28768207245178085), (6, 0.5753641449035617), (15, 0.6931471805599453), (5, 0.6931471805599453)]
TF-IDF - vollständig [(16, 1.3862943611198906), (17, 1.3862943611198906), (7, 0.28768207245178085), (6, 0.5753641449035617), (15, 0.6931471805599453), (5, 0.6931471805599453), (0, 0), (1, 0), (2, 0), (3, 0), (4, 0), (9, 0), (10, 0), (11, 0), (13, 0), (23, 0), (25, 0), (32, 0), (34, 0), (36, 0)]
== NORMIERT == TF-IDF - sparlich [(16, 0.6068613490456398), (17, 0.6068613490456398), (7, 0.12593510835844393), (6, 0.25187021671688786), (15, 0.3034306745228199), (5, 0.3034306745228199)]
TF-IDF - vollständig [(16, 0.6068613490456398), (17, 0.6068613490456398), (7, 0.12593510835844393), (6, 0.25187021671688786), (15, 0.3034306745228199), (5, 0.3034306745228199), (0, 0.0), (1, 0.0), (2, 0.0), (3, 0.0), (4, 0.0), (9, 0.0), (10, 0.0), (11, 0.0), (13, 0.0), (23, 0.0), (25, 0.0), (32, 0.0), (34, 0.0), (36, 0.0)]
Euclidean distance
Nicht normiert: 2.0488859691682535
(0.28768207245178085-1.3862943611198906)^2 = 1.206948960812582 (0.6931471805599453-1.3862943611198906)^2 = 0.4804530139182014 (1.3862943611198906-0.28768207245178085)^2 = 1.206948960812582 (1.3862943611198906-0.5753641449035617)^2 = 0.6576078155726619 (1.3862943611198906-0.6931471805599453)^2 = 0.4804530139182014 (0.28768207245178085-0)^2 = 0.08276097481015168 (0.28768207245178085-0)^2 = 0.08276097481015168
Normiert: 0.8214397067657133
(0.10893121906625666-0.6068613490456398)^2 = 0.24793441434128538 (0.2624611493070678-0.6068613490456398)^2 = 0.11861149757996829 (0.5249222986141356-0.12593510835844393)^2 = 0.15919077798813155 (0.5249222986141356-0.25187021671688786)^2 = 0.0745574394284213 (0.5249222986141356-0.3034306745228199)^2 = 0.04905853954260871 (0.2624611493070678-0.3034306745228199)^2 = 0.0016785019964041463 (0.10893121906625666-0.0)^2 = 0.0118660104872608 (0.10893121906625666-0.0)^2 = 0.0118660104872608