Lecture 5: Deep Learning Overview

(Last updated: Sep 27, 2026)

This course introduces deep learning techniques, such as deep neural networks, loss functions, and gradient descent.

Check the GenAI usage policy if you are using the course materials with GenAI for self-study and fact-checking.

Preparation

Read the required course readings.

Lecture

Below are the slides:

Required Course Readings

  • Section 2.4.4 (Neural Networks, including 2.4.4.1 and 2.4.4.2) in the ML4Design lecture notes (Bozzon, 2023)
  • Section 10.7 (Fitting a Neural Network, including 10.7.1, 10.7.2, 10.7.3, and 10.7.4) in book An Introduction to Statistical Learning (James et al., 2013). Do not worry about the chain rule math in 10.7.1 (focus on understanding why we need backpropagation).

Optional Course Readings

  • Section 5.2 (Capacity, Overfitting and Underfitting, including 5.2.1, and 5.2.2) in book Deep Learning (Goodfellow et al., 2016).
  • Section 6.2 (Shrinkage Methods, including 6.2.1, 6.2.2, and 6.2.3) in book An Introduction to Statistical Learning (James et al., 2013)
  • Section 7.4 (Dataset Augmentation) in book Deep Learning (Goodfellow et al., 2016).

Exercises

See the instruction in the syllabus about how to use the exercises.

  • Explain the difference between linear regression and logistic regression. What are they used for? How to think about them from the artificial neuron perspective? What activation functions do they use? What loss functions do they use? Why not just using linear regression to predict the probabilities for the binary classification task?
  • What activation functions to choose if the data points are linearly separable in a high-dimensional feature space? If the data points are not linearly separable, what are the activation functions to use?
  • What is the role of a loss function? Why do we need it in the training process?
  • If we ask you to compute the gradient descent for 10 steps, how to write Python code to compute the gradient descent updates, just like what we did in the class exercise?
  • Why do we need lasso and ridge regularization? How do they work (mathematically, what is the math that is added to the loss function)? What do they mean geometrically? What are their behaviors in terms of shrinking the coefficients toward zero?

Additional Resources

Below is a well-known paper that gives a nice overview of deep learning:

Below is a collection of projects by OpenAI (which creates ChatGPT):

Below is a collection of project by DeepMind (which creates AlphaGo):

Below is a collection of AI experiments by Google:


This site uses Just the Docs, a documentation theme for Jekyll.