Lecture 2: Data Science Fundamentals
(Last updated: Sep 27, 2026)
This lecture recaps the fundamentals of data science, such as table operations, classification, and regression.
Check the GenAI usage policy if you are using the course materials with GenAI for self-study and fact-checking.
Preparation
Read the required course readings.
Lecture
Below are the slides:
- Slides for Lecture 2-1: Data Science Fundamentals (Preprocessing)
- Slides for Lecture 2-2: Data Science Fundamentals (Modeling)
Below is the link to the online notebook:
Follow the steps on the notebook page to set up the notebook.
Required Course Readings
- The following sections in book An Introduction to Statistical Learning (James et al., 2013)
- 2.2.1 (Measuring the Quality of Fit)
- 3.1.1 (Estimating the Coefficients)
- 3.1.3 (Assessing the Accuracy of the Model)
- 9.1.1 (What Is a Hyperplane?)
- 9.1.2 (Classification Using a Separating Hyperplane)
Optional Course Readings
- Section 5.3 (Hyperparameters and Validation Sets, including 5.3.1) in book Deep Learning (Goodfellow et al., 2016).
- Section 4.5.1 (Rosenblatt’s Perceptron Learning Algorithm) in book The Elements of Statistical Learning (Hastie et al., 2009)
Exercises
See the instruction in the syllabus about how to use the exercises.
- If we give you two numpy arrays: one is the prediction of a model (either regression or classification), and one is the ground truth, how to compute the evaluation metrics (either F-score or R-squared) by writing Python code?
- Explain the intuition of precision, recall, and f-score. What does a high-precision and a low-recall model mean? Conversely, what does a low-precision and a high-recall model mean?
- Explain the intuition of the R-squared metric. What does it mean geometrically if we plot the regression line on a 2D plot, where the x-axis is the feature, and the y-axis is the prediction or ground truth? What does it mean when having a bad R-squared value? Can we always use R-squared to determine if there is a pattern in the data, and why?
- Describe how to construct a linear classifier. How to represent the linear classifier using math equations? Give an example of the metric for determining whether the linear classifier work well or not. What is the function that you need to optimize (in mathematical form)? Do the same exercise for the linear regression model.
- Describe the procedure of computing permutation feature importance. What to do if we have high-correlated features? Why do we need to run the permutation and compute the importance for multiple times for one feature?
- Explain why do we need to map raw data into data points in a high-dimensional space? What do the axes in the high-dimensional space mean?
- What are the typical assumptions for linear regression?
Additional Resources
Below are website for data visualization inspirations:
- Seaborn: Statistical Data Visualization
- Exploratory Data Analysis by the US EPA
- Examples of Data Exploration by the Statistics Netherlands
- Examples of Data Visualization
Below are interesting data science case studies:
The textbook below contains more information about how to select models:
- Section 11.8 Comparing Different Models in book: Introduction to Statistics and Data Analysis
The websites below contains exercises for Python pandas: