serhii.net

In the middle of the desert you can say anything you want

UNLISTED

22 Apr 2023

TWM 3 - vector models

Vector stuff - theory

Bag of words

Mathematically - it’s a vector that contains 0 (absent) or 1 (present).

(D) Termgewicht = “Wert eines Vektors in einer Dimension = Stärke der Assoziation zwischen Term und Dokument”

Termgewichte

  • Vector (=column) of size $d_{|dict|}$
  • Relevanz eines Termes

BoolEsche Termgewichte

  • (L) Boolesch!
  • Einfachste ${0,1}^{|dict|}$
  • Vorhandensein oder Fehlen eines Terms

Termhäufigkeitsgewichtung

  • the more often we see it, the more important it is
  • e.g. just counts

Statistical ways

2023-04-22-190638_688x481_scrot.png

  • We care about the middle part of the Zipf Verteilung - not the “the"s and not the long tail
  • Ways to get it:
    • Filter during preprocessing
    • Weight correspondingly

IDF and friends

IDF - Inverse Document Frequency

  • High -> rare term, low -> frequent

  • (D) Dokumenthäufigkeit $df$ -> number of documents, where we see the word.

Inverse Dockumentenhäufigkeit eines terms: $$idf_t = log(\frac{N}{df_t})$$ (where N is the total number of documents)

So - $log(\frac{10.000}{1})=4$ means document found once in 10k docs. - $log(10.000/10.000)=0$ - found in each document.

Wikipedia: IDF is a measure of how much information the word provides (’the’ provides almost none)

TF-IDF

(tf–idf - Wikipedia)

  • (D) TF-IDF: Combination of Term Frequency and IDF

$$w_{t,d} = tf_{t,d}*ids_t$$

High weight for terms that:

  • Found often in THIS document
  • Rare in the other documents

TF

  • Term frequency - how often in document, a lot of schemas possible. Natural (count), binary (0/1), log., etc.
  • $freq_{t,d}$ 2023-04-22-191914_1011x488_scrot.png

Termgewichtungsschema

  • Termgewichtung, generalized (Sl.19)
  • Three bits
    • Termhäufigkeit $freq_{t,d}$
    • Dokumentenhäufigkeit $collect_t$
    • Normierung $norm$

$$w_{t,d} = \frac{freq_{t,d}*collect_t}{norm}$$

Varianten

  • Termhäufigkeit
    • natural,binary,log,etc.
  • Dokumentenhäufigkeit $collect_t$ (Sl.21)
    • 1 (none)
    • $log\frac{N}{df_t}$ invers (=IDF)
    • probabilistic etc.
  • Normierung
    • 1 / keine
    • $\sum_i^1w_{i,d}$

2023-04-22-192530_1173x567_scrot.png

TODO: what’s the logic in i,d in norm options?

Vector stuff - practice

Saving space

  • Don’t save the 0s to save space, just save word-1,word2-2,word5-3 etc.

Libraries

  • Gensim for text mining
  • sklearn for everything

Similarity

  • Sim. of docs == sim of their vectors
  • can measure this vector similarity in a lot of ways
  • Similarity <-> distance
    • $sim(d_1,d_2) = \frac{1}{1+dist(d_1,d_2)}$
    • Similarity is between $[0;1]$
    • Distance is between $[0;\infty]$

Euc. distance

  • $\sqrt{\sum^n_{i=1}(d_{1,i}-d_{2,i})^2}$

  • 2023-04-22-193509_871x663_scrot.png

    • Problem with the last one: CD is close, BC is far, but we don’t see it

(Sl.35)

  • Problems:

    • TODO
    • length of vectors influences the distance
      • document pasted after itself is now totally different
    • presence and absence of terms are gleichbehandelt
      • TODO ?
  • Solution: use the angle instead of Euc.

  • Normalization to length 1 helps! 2023-04-22-193751_1235x749_scrot.png

Problems with vectors

2023-04-22-193931_1168x575_scrot.png

Lab3 solution

3-1

Dokument 1: im garten geben es viel blume und baum Dokument 2: im park sein auch baum aber auch strauch Dokument 3: anna mögen baum und strauch und blume Dokument 4: sie sein gern im garten und park und sie gehen gern spazieren und walken

Berechnen Sie den euklidischen Abstand und die Kosinusähnlichkeit zwischen Dokument 1 und 3 für die Häufigkeitsvektoren und die normalisierten TF-IDF-Vektoren.

Wörterbuch

Geben Sie das (alphabetisch sortierte) Wörterbuch an! Dokumentvorkommen

[('aber', 1),
 ('anna', 1),
 ('auch', 1), // ERROR! was 2 (Dokumentvorkommen, nicht W.)
 ('baum', 3),
 ('blume', 2),
 ('es', 1),
 ('garten', 2),
 ('geben', 1),
 ('gehen', 1),
 ('gern', 1), // Error: war 2
 ('im', 3),
 ('mögen', 1),
 ('park', 2),
 ('sein', 2),
 ('sie', 1), // ERROR! War 2
 ('spazieren', 1),
 ('strauch', 2),
 ('und', 3), // ERROR! Dockumentvorkommen! (was wf: 6)
 ('viel', 1),
 ('walken', 1)]

Häufigkeitvektoren, TF-IDSF and normalized TF-IDFs

Geben Sie die Häufigkeitsvektoren, die TF-IDF-Vektoren sowie die normalisierten TF-IDF-Vektoren für die Dokumente 1 und 3 an, und zwar sowohl in vollständiger als auch in spärlicher Form!

Doc1

Dokument 1: im garten geben es viel blume und baum

TF

Natürlich, binär:

im: 1
garten: 1
geben: 1
es: 1
viel: 1
blume: 1
und: 1
baum: 1
IDF

IDF sentence 1 im, 0.28768207245178085 garten, 0.6931471805599453 geben, 1.3862943611198906 es, 1.3862943611198906 viel, 1.3862943611198906 blume, 0.6931471805599453 und, 0.28768207245178085 baum, 0.28768207245178085

TF-IDF

im: log(10.29) -> 0.29 garten: log(10.69) -> 0.69 geben: log(11.39) -> 1.39 es: log(11.39) -> 1.39 viel: log(11.39) -> 1.39 blume: log(10.69) -> 0.69 und: log(10.29) -> 0.29 baum: log(10.29) -> 0.29

Doc3

Dokument 3: anna mögen baum und strauch und blume

TF: anna: 1 mögen: 1 baum: 1 und: 2 strauch: 1 blume: 1

TF-IDF: anna: log(11.39) -> 1.39 mögen: log(11.39) -> 1.39 baum: log(10.29) -> 0.29 und: log(20.29) -> 0.58 strauch: log(10.69) -> 0.69 und: log(20.29) -> 0.58 blume: log(1*0.69) -> 0.69

j,

VECTORS - ANSWER

Sentence 1

== NICHT NORMIERT == TF-IDF - sparlich [(0, 0.28768207245178085), (1, 0.6931471805599453), (2, 1.3862943611198906), (3, 1.3862943611198906), (4, 1.3862943611198906), (5, 0.6931471805599453), (6, 0.28768207245178085), (7, 0.28768207245178085)]

TF-IDF - vollständig [(0, 0.28768207245178085), (1, 0.6931471805599453), (2, 1.3862943611198906), (3, 1.3862943611198906), (4, 1.3862943611198906), (5, 0.6931471805599453), (6, 0.28768207245178085), (7, 0.28768207245178085), (9, 0), (10, 0), (11, 0), (13, 0), (15, 0), (16, 0), (17, 0), (23, 0), (25, 0), (32, 0), (34, 0), (36, 0)]

== NORMIERT == TF-IDF - sparlich [(0, 0.10893121906625666), (1, 0.2624611493070678), (2, 0.5249222986141356), (3, 0.5249222986141356), (4, 0.5249222986141356), (5, 0.2624611493070678), (6, 0.10893121906625666), (7, 0.10893121906625666)]

TF-IDF - vollständig [(0, 0.10893121906625666), (1, 0.2624611493070678), (2, 0.5249222986141356), (3, 0.5249222986141356), (4, 0.5249222986141356), (5, 0.2624611493070678), (6, 0.10893121906625666), (7, 0.10893121906625666), (9, 0.0), (10, 0.0), (11, 0.0), (13, 0.0), (15, 0.0), (16, 0.0), (17, 0.0), (23, 0.0), (25, 0.0), (32, 0.0), (34, 0.0), (36, 0.0)]

Sentence 3

== NICHT NORMIERT == TF-IDF - sparlich [(16, 1.3862943611198906), (17, 1.3862943611198906), (7, 0.28768207245178085), (6, 0.5753641449035617), (15, 0.6931471805599453), (5, 0.6931471805599453)]

TF-IDF - vollständig [(16, 1.3862943611198906), (17, 1.3862943611198906), (7, 0.28768207245178085), (6, 0.5753641449035617), (15, 0.6931471805599453), (5, 0.6931471805599453), (0, 0), (1, 0), (2, 0), (3, 0), (4, 0), (9, 0), (10, 0), (11, 0), (13, 0), (23, 0), (25, 0), (32, 0), (34, 0), (36, 0)]

== NORMIERT == TF-IDF - sparlich [(16, 0.6068613490456398), (17, 0.6068613490456398), (7, 0.12593510835844393), (6, 0.25187021671688786), (15, 0.3034306745228199), (5, 0.3034306745228199)]

TF-IDF - vollständig [(16, 0.6068613490456398), (17, 0.6068613490456398), (7, 0.12593510835844393), (6, 0.25187021671688786), (15, 0.3034306745228199), (5, 0.3034306745228199), (0, 0.0), (1, 0.0), (2, 0.0), (3, 0.0), (4, 0.0), (9, 0.0), (10, 0.0), (11, 0.0), (13, 0.0), (23, 0.0), (25, 0.0), (32, 0.0), (34, 0.0), (36, 0.0)] ​

Euclidean distance

Nicht normiert: 2.0488859691682535

(0.28768207245178085-1.3862943611198906)^2 = 1.206948960812582 (0.6931471805599453-1.3862943611198906)^2 = 0.4804530139182014 (1.3862943611198906-0.28768207245178085)^2 = 1.206948960812582 (1.3862943611198906-0.5753641449035617)^2 = 0.6576078155726619 (1.3862943611198906-0.6931471805599453)^2 = 0.4804530139182014 (0.28768207245178085-0)^2 = 0.08276097481015168 (0.28768207245178085-0)^2 = 0.08276097481015168

Normiert: 0.8214397067657133

(0.10893121906625666-0.6068613490456398)^2 = 0.24793441434128538 (0.2624611493070678-0.6068613490456398)^2 = 0.11861149757996829 (0.5249222986141356-0.12593510835844393)^2 = 0.15919077798813155 (0.5249222986141356-0.25187021671688786)^2 = 0.0745574394284213 (0.5249222986141356-0.3034306745228199)^2 = 0.04905853954260871 (0.2624611493070678-0.3034306745228199)^2 = 0.0016785019964041463 (0.10893121906625666-0.0)^2 = 0.0118660104872608 (0.10893121906625666-0.0)^2 = 0.0118660104872608

Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus