CBDM V.5 - MapReduce
Next / part 2 with examples: 230522-1436 CBDM 6 MapReduce-2 and hadoop
MapReduce
Parallele Programmierung
- Großer Datenmengen -> berechnung auf mehrere Knoten
- -> Divide-and-Conquer: Aufteilung in kleine Sub-Tasks, unabhnnängig ausgeführt, und Kombination der Ergebnisse
Funkt. programmierung
- map (fn, list) -> list
- e.g. python
lambda
- e.g. python
- fold (fn, startWert, list) -> Wert:
- Applies a function element-by-element, aggregates the result, and returns the final value
MapReduce
Example: parse txt rows to get max temp. pro each year
- Data flow:
- *input: get a list
- map: get the important bit from each element
- e.g. “pro k1/v1 pair wird eine k2/v2 pair erzeugt”
(1940,3),(1950,5), ...
- e.g. “pro k1/v1 pair wird eine k2/v2 pair erzeugt”
- shuffle&sort:
- e.g. sort k2/v2 pairs, group based on k2
(1949, [10,15]), (1950, [03,30]), ...
- e.g. sort k2/v2 pairs, group based on k2
- etc.
Nel mezzo del deserto posso dire tutto quello che voglio.
comments powered by Disqus