GEN 15341

Statistical Models of Text and Language

Stanford University · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

This course examines the statistical foundations of text and language, emphasizing explicit probabilistic models rather than black-box NLP techniques. Language follows well-defined statistical laws that govern word frequency, predictability, and variation. Understanding these properties enables quantitative text analysis, measurement of information content, and development of interpretable models used in linguistics, information retrieval, and computational text processing. As large-scale textual data continues to grow, statistical methods are crucial for detecting patterns, analyzing linguistic trends, and constructing efficient, interpretable models. Key topics include: Word Frequency Distributions (Zipf's and Heaps' laws); Entropy & Information Theory (redundancy and uncertainty in language); Probabilistic Language Models (n-grams, smoothing, perplexity); Markov Models & Hidden Markov Chains (stochastic text sequences); Text Similarity & Distance Metrics (measuring divergence in text); Corpus Statistics & Sampling (estimating linguistic trends); Random Processes in Text Generation (stochastic models of language). By the end of the course, students will develop a strong foundation in statistical text analysis, equipping them with essential tools for computational linguistics, AI, search technologies, and digital humanities in an increasingly data-driven world.

Sections

Current meeting, instructor, credit, and enrollment details

Updated 5 hours ago

001

Availability not recently verified
Class #stanford-15341Fall 2026UGRD3 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?