Stanford Root

Schedule

Stanford Root

Schedule

STATS 292

Statistical Models of Text and Language

UNITS:3
GRADING:Letter or Credit/No Credit
LEVEL:Graduate
GER:—

This course examines the statistical foundations of text and language, emphasizing explicit probabilistic models rather than black-box NLP techniques. Language follows well-defined statistical laws that govern word frequency, predictability, and variation. Understanding these properties enables quantitative text analysis, measurement of information content, and development of interpretable models used in linguistics, information retrieval, and computational text processing. As large-scale textual data continues to grow, statistical methods are crucial for detecting patterns, analyzing linguistic trends, and constructing efficient, interpretable models. Key topics include: Word Frequency Distributions (Zipf's and Heaps' laws); Entropy & Information Theory (redundancy and uncertainty in language); Probabilistic Language Models (n-grams, smoothing, perplexity); Markov Models & Hidden Markov Chains (stochastic text sequences); Text Similarity & Distance Metrics (measuring divergence in text); Corpus Statistics & Sampling (estimating linguistic trends); Random Processes in Text Generation (stochastic models of language). By the end of the course, students will develop a strong foundation in statistical text analysis, equipping them with essential tools for computational linguistics, AI, search technologies, and digital humanities in an increasingly data-driven world.

Syllabus for selected term:
View Spring 2027 Syllabus

Sections

1 Term
Lecture 1Open
ID: 12683
0 / 45 enrolled
DAYS:Tuesday, Thursday
TIME:1:30 PM – 2:50 PM
LOCATION:TBD
INSTRUCTOR:
Donoho, David
3units

STATS 292: Statistical Models of Text and Language

3 units · Letter or Credit/No Credit

This course examines the statistical foundations of text and language, emphasizing explicit probabilistic models rather than black-box NLP techniques. Language follows well-defined statistical laws that govern word frequency, predictability, and variation. Understanding these properties enables quantitative text analysis, measurement of information content, and development of interpretable models used in linguistics, information retrieval, and computational text processing. As large-scale textual data continues to grow, statistical methods are crucial for detecting patterns, analyzing linguistic trends, and constructing efficient, interpretable models. Key topics include: Word Frequency Distributions (Zipf's and Heaps' laws); Entropy & Information Theory (redundancy and uncertainty in language); Probabilistic Language Models (n-grams, smoothing, perplexity); Markov Models & Hidden Markov Chains (stochastic text sequences); Text Similarity & Distance Metrics (measuring divergence in text); Corpus Statistics & Sampling (estimating linguistic trends); Random Processes in Text Generation (stochastic models of language). By the end of the course, students will develop a strong foundation in statistical text analysis, equipping them with essential tools for computational linguistics, AI, search technologies, and digital humanities in an increasingly data-driven world.

Offered in Spring 2027 at Stanford University.

Spring 2027 sections

  • Lecture — Tuesday Thursday 1:30 PM – 2:50 PM — Donoho, David (Graduate)

More STATS courses

  • STATS 260C: Workshop in Biomedical Data Science (BMDS 280C)
  • STATS 261: Intermediate Biostatistics: Analysis of Discrete Data (BMDS 241, EPI 261)
  • STATS 262: Intermediate Biostatistics: Regression, Prediction, Survival Analysis (EPI 262)
  • STATS 263: Design of Experiments (STATS 363)
  • STATS 264: Foundations of Statistical and Scientific Inference (BMDS 243, EPI 264)
  • STATS 270: Bayesian Statistics (STATS 370)
  • STATS 298: Industrial Research for Statisticians
  • STATS 299: Independent Study
  • STATS 300A: Theory of Statistics I
  • STATS 300B: Theory of Statistics II
  • STATS 300C: Theory of Statistics III
  • STATS 301: Statistics Teaching Practicum

All STATS courses · All departments