CSE 589
Vision and Language
Pennsylvania State University-Mont Alto Campus · UGRD · Fall 2026
Catalog description
Multi-modal data processing with vision and language inputs is ubiquitous in various applications and faces the challenges from visual and language representation, efficient learning and cross-modal reasoning, etc. This course covers the latest techniques in deep learning for various vision and language tasks. Neural network architectures such as convolutional neural network, recurrent neural network and transformer will be leveraged to build deep learning models for cross-modal retrieval, captioning, visual question answering and referring expression in both image and video domain. Efficient learning techniques such as weakly-supervised learning, unsupervised learning and few-shot/zero-shot learning will be utilized to maximize the label usage in these tasks. It will also introduce visual structure representation in the form of scene graphs and incorporate external knowledge graphs for cross-modal reasoning among text and image/video. This course prepares students to conduct graduate level research using deep learning, computer vision and natural language processing techniques.
Sections
Current meeting, instructor, credit, and enrollment details
001
Availability not recently verified- Days & times
- No scheduled meeting time
- Meeting dates
- —
- Location
- —
- Instructor
- Staff