PyDA Course
🔤
Advanced ⏱ 50 min 📖 2 Lessons

Module 2: Tokenization & Word Frequency

Split raw text into tokens, count word frequencies, and build the vocabulary for our language model.

tokenizationword-frequencydictvocabularynlp

Lessons

Loading a Text Corpus Bigram Probability Tables