CSE 589

Vision and Language

Pennsylvania State University-Main Campus · UGRD · Fall 2026

1 section
Add to a schedule

Catalog description

Multi-modal data processing with vision and language inputs is ubiquitous in various applications and faces the challenges from visual and language representation, efficient learning and cross-modal reasoning, etc. This course covers the latest techniques in deep learning for various vision and language tasks. Neural network architectures such as convolutional neural network, recurrent neural network and transformer will be leveraged to build deep learning models for cross-modal retrieval, captioning, visual question answering and referring expression in both image and video domain. Efficient learning techniques such as weakly-supervised learning, unsupervised learning and few-shot/zero-shot learning will be utilized to maximize the label usage in these tasks. It will also introduce visual structure representation in the form of scene graphs and incorporate external knowledge graphs for cross-modal reasoning among text and image/video. This course prepares students to conduct graduate level research using deep learning, computer vision and natural language processing techniques.

Sections

Current meeting, instructor, credit, and enrollment details

Updated 7 hours ago

001

Availability not recently verified
Class #pennsylvania_main_campus-CSE589Fall 2026UGRD3 credits
Days & times
No scheduled meeting time
Meeting dates
Location
Instructor
Staff
Class numbers and section codes come from the registrar.
Spot missing or incorrect course data?