Lecture 3: Structured Data Processing (Part I)

(Last updated: Sep 27, 2026)

This lecture explains the theory of Decision Tree and Random Forest models that are used in the structured data processing module.

Check the GenAI usage policy if you are using the course materials with GenAI for self-study and fact-checking.

Preparation

Read the required course readings.

Lecture

Below are the slides:

Required Course Readings

Optional Course Readings

  • Section 5.4 (Estimators, Bias and Variance) in book Deep Learning (Goodfellow et al., 2016).
  • Section 2.2.2 (The Bias-Variance Trade-Off) and 12.2 (Principal Components Analysis, including 12.2.1, 12.2.2 ) in book An Introduction to Statistical Learning (James et al., 2013)

Exercises

See the instruction in the syllabus about how to use the exercises.

  • Explain the procedure for training a decision tree. How does the training data look like? How to pick the feature to split a node? Which metric to use for node splitting? After splitting a node, what to do for splitting other nodes (using what kind of logic in programming)? What is the tree doing when we think about how the model cuts a high-dimensional space (with data points) into which kind of structure? How to prevent the tree from overfitting? What to do when we have features with continuous values?
  • What is a good strategy to do hyper-parameter tuning? How should you split the dataset? Which part of the dataset split should be used for hyper-parameter tuning?
  • Explain the difference between a feedforward neural network and a recurrent neural network. What are the differences between their inputs? Are there any differences in the neural net architecture?
  • Explain the procedure of constructing a random forest. What are the underlying small models in a random forest? How to train each small model? How to aggregate the output from these small models?
  • Compared with a decision tree model, what are the advantages of a random forest in terms of bias and variance tradeoff? Also, describe the underlying statistical theory that makes random forest work well.
  • If we give you the probabilities of seeing each side of a two-sided coin, how to compute entropy by writing Python code? What about a dice that has four sides and six sides?
  • Describe PCA conceptually. Given a coordinate system with many data points, what does PCA do to the coordinate system? Is PCA doing something to the origin of the coordinate system, and in which way? Think about what will happen after PCA if you look at the data on the principal component axes.
  • Explain what under-fitting and over-fitting mean in terms of bias, variance, and model complexity. What does a high-bias and low-variance model mean intuitively? Also, what does a low-bias and high-variance model mean? Think about the situation when we run the experiment infinite number of times (so we have infinite number of datasets and models).

Additional Resources

Below are videos from StatQuest that explains the Decision Tree model and PCA nicely:


This site uses Just the Docs, a documentation theme for Jekyll.