Technische Informationsbibliothek (TIB)18 Module
Text mining with r
18 Einträge
Aktualisiert
Beschreibung und Lernergebnis
Mentoren
GKGábor Kismihók
CECarolin Eisentraut

Beschreibung

This learning path provides a comprehensive introduction to text mining concepts and their practical application using R. It covers essential techniques such as text preprocessing (lower case conversion, punctuation and stopword removal, tokenization), stemming, lemmatization, feature extraction (Bag of Words, TF-IDF), and advanced models like Word2Vec, Doc2Vec, Sentiment Analysis, Latent Semantic Analysis, and Latent Dirichlet Allocation, all with a focus on implementation in R.

Lernziele

  • Understand fundamental text mining concepts.

  • Apply text preprocessing techniques in R (lower case conversion, punctuation and stopword removal, tokenization).

  • Implement stemming and lemmatization in R.

  • Utilize feature extraction methods like Bag of Words and TF-IDF in R.

  • Work with advanced text mining models such as Word2Vec and Doc2Vec in R.

  • Perform sentiment analysis using R.

  • Apply Latent Semantic Analysis (LSA) and Latent Dirichlet Allocation (LDA) in R for topic modeling.

Enthaltene Inhalte
Entdecke die Module, die in diesem Lernpfad enthalten sind.
1

Overview of text mining

Linkinhalt
2

Lower case conversion, remove punctuation and stopwords, text tokenization in r

Linkinhalt
3

Stemming and lemmatization

Linkinhalt
4

Stemming and lemmatization in r

Linkinhalt
5

Bag of word

Linkinhalt
6

Bag of word in r

Linkinhalt
7

Tf-idf

Linkinhalt
8

Tf-idf in r

Linkinhalt
9

Part of speech tagging

Linkinhalt
10

Part of speech tagging in r

Linkinhalt
11

Word2vec

Linkinhalt
12

Word2vec in r

Linkinhalt
13

Doc2vec

Linkinhalt
14

Sentiment analysis in r

Linkinhalt
15

Latent semantic analysis

Linkinhalt
16

Latent semantic analysis in r

Linkinhalt
17

Latent dirichlet allocation

Linkinhalt
18

Latent dirichlet allocation in r

Linkinhalt