Lecture 9: Image Data Processing (Part I)
(Last updated: Sep 27, 2026)
This lecture introduces the theory for image data processing, including the convolution operation, image filtering, and convolutional neural network.
Check the GenAI usage policy if you are using the course materials with GenAI for self-study and fact-checking.
Preparation
Read the required course readings.
Lecture
Below are the slides:
- Slides for Lecture 9-1: Image Data Processing (Image Filtering)
- Slides for Lecture 9-2: Image Data Processing (Convolutional Neural Network)
Required Course Readings
- Section 5.4 (Convolutional neural networks) in book Computer Vision: Algorithms and Applications (Szeliski, 2022). You can ignore the following subsections: U-Nets and Feature Pyramid Networks, Mobile networks, 5.4.4 Model zoos, 5.4.6 Adversarial examples, and 5.4.7 Self-supervised learning.
Optional Course Readings
- Section 5.3 (Deep neural networks) in book Computer Vision: Algorithms and Applications (Szeliski, 2022).
Exercises
See the instruction in the syllabus about how to use the exercises.
- Describe the reason why we need the residual block? How is it implemented? What problem does the block solve? What are the intuitions behind using it?
- Describe the reason of using leaky ReLU activation function. What problem does it solve? How is it implemented?
- If we give you a simple image in a numpy array, and a kernel also in a numpy array, how to compute the result of convolution using a particular stride and zero-padding (either by hand or writing python code)? If we just ask you to calculate the shape of the output, how to do that?
- Given an image (with arbitrary width, height, and number of channels) and a convolutional neural net block (with arbitrary kernel size, stride, and zero-padding), how to compute the number of training parameters? How about a fully-connected layer?
- Describe how the max pooling layer works. Are there training parameters in the layer? If yes, what are these training parameters?
- Describe how the normalization layer works (mathematically). Are there training parameters in the layer? If yes, what are these training parameters? Also, explain why we need Batch Normalization.
- Explain how Class Activation Mapping works conceptually. How can we construct the class activation map (using information from which layer in the network)?
Additional Resources
The following documentations explains the convolutional layer in great details:
The following website has an interactive visualization for understanding image filtering:
The following website has an interactive visualization for understanding Convolutional Neural Network:
The following website shows a live training demo using the MNIST dataset:
Below are tutorials of Convolutional Neural Network:
- Convolutional Neural Networks tutorial (more technical)
- Convolutional Neural Networks tutorial (more intuitive)
The following webpage has many examples of Computer Vision applications: