ORIE 5741

Learning with Big Messy Data

Cornell University · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

Modern data sets, whether collected by scientists, engineers, medical researchers, government, financial firms, social networks, or software companies, are often big, messy, and extremely useful. This course addresses scalable robust methods for learning from big messy data. We'll cover techniques for learning with data that is messy --- consisting of real numbers, integers, booleans, categoricals, ordinals, graphs, text, sets, and more, with missing entries and with outliers --- and that is big --- which means we can only use algorithms whose complexity scales linearly in the size of the data. We will cover techniques for cleaning data, supervised and unsupervised learning, finding similar items, model validation, and feature engineering.

Sections

Current meeting, instructor, credit, and enrollment details

Updated 7 hours ago

001

Availability not recently verified
Class #cornell_2-ORIE5741Fall 2026UGRD4 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?