Introduction To Artificial
Intelligence and Machine
Learning
Aniketh Chenjeri, Andrew Doyle, Swarnim Ghimire, Mr. Igor
Tomcej
2026-08-20
Contents
Foreword
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
4
How this book is structured
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
5
Part I: Supervised Learning
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
7
1
Linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
8
1.1
What is linear regression?
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
8
1.2
The limits of a two-point estimate
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
9
1.3
Define linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
10
1.4
Fit the line
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
10
1.5
Measure and interpret error
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
10
1.6
Optimize with gradient descent
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
12
1.7
Understand gradients
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
13
1.8
Problem packet for January 16, 2026
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
16
1.9
Problem packet
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
17
Theory questions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
17
Practice problems
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
17
2
Pseudoinverse and multiple linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
2.1
Linear algebra primer
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
Sets
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
Common number sets
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
Relationships between sets
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
Vectors, vector addition, and scalar multiplication
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
19
Matrices, notation, and dimensions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
20
Add and subtract matrices
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
21
Dot product
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
22
Matrix multiplication
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
22
Transpose
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
23
Identity matrix
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
23
Determinant
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
24
Inverse matrix
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
25
2.2
Define multiple linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
26
2.3
Solve with the normal equation
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
26
Build the matrix expression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
27
2.4
Use gradient descent with multiple variables
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
28
2.5
Scale the features
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
28
Examine a constructed failure case
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
29
Compare three training runs
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
30
Connect scaling to the loss geometry
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
31
Apply scaling to other datasets
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
32
2.6
Interpret weights after scaling
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
32
2.7
Linear algebra practice problems
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
32
Sets and number sets
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
32
Vectors and vector operations
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
33
Notation and dimensions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
33
Dot product
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
34
Matrix multiplication
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
34
Transpose
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
34
Determinant
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
34
Identity matrix and inverse
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
35
Mixed practice
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
35
2.8
Problem packet
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
35
Data representation and the design matrix
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
35
Matrix operations reference
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
36
Normal equation with derivation
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
36
MSE and gradient with derivation
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
36
Apply RSS and MSE
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
37
Collinearity and remedies
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
37
Appendix: Programming Reference
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
3
Python and libraries
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
3.1
NumPy
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
Import NumPy
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
Create arrays and vectors
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
Represent matrices with two-dimensional arrays
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
Call methods and functions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
Understand views and shared memory
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
39
Transpose, ndim, and shape
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
39
Apply elementwise operations
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
39
Generate random numbers
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
39
Reproduce results with the Generator API
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
40
Calculate the mean, variance, and standard deviation
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
40
Calculate along an axis
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
40
Plot with Matplotlib
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
40
Keep results reproducible
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
41
Appendix: Math Fundamentals
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4
Calculus
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4.1
Limits
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4.2
Derivatives
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
Limit definition of a derivative
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4.3
Gradients
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
Vector-valued functions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
Gradient definition
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
Partial derivatives and rules
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
5
Linear algebra
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
6
Statistics and probability
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
Reference
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
43
Glossary of Definitions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
44
4
Foreword
Cherry Creek first offered Introduction to Artificial Intelligence and Machine
Learning during the 2024-2025 school year. The course gave students a gentle
introduction to the field. It avoided most of the mathematics and focused on
libraries such as TensorFlow, NumPy, pandas, and scikit-learn.
The 2025-2026 school year began without a stable course plan, so we rebuilt the
syllabus. We kept the original goal of an approachable introduction to machine
learning, but added the rigor that the subject requires. We still teach each
mathematical idea through intuition before notation.
That balance is hard because machine learning draws from programming, algebra,
calculus, statistics, and probability. Leaving out the mathematics would also leave
out the reasons that the algorithms work.
This textbook contains the lessons, lecture notes, and exercises for the course. The
course has no calculus prerequisite, but mathematics appears throughout artificial
intelligence and machine learning. We could omit that mathematics and cover more
topics, or include it and move more slowly. We chose depth.
You do not need calculus to begin this book. You need basic algebra and some
experience with Python. When a lesson uses unfamiliar mathematics, we first
explain what each operation represents. We calculate difficult derivatives for you,
then ask you to interpret and use the result. This practical approach gives you a
base for further study, but it does not replace a full course in calculus, linear
algebra, or statistics.
The following people contributed to this book:
•
Primary Writers:
‣
Aniketh Chenjeri (CCHS ‘26)
‣
Andrew Doyle (CCHS ‘26)
‣
Mr. Igor Tomcej
•
Reviewers:
‣
Hariprasad Gridharan (CCHS ‘25, Cornell ‘29)
‣
Siddharth Menon (CCHS ‘26)
‣
Ani Gadepalli (CCHS ‘26)
5
How this book is structured
Note:
This book is still in development. We release chapters after review, so the
version you are reading is not final.
Early editions may contain errors. If you find an unclear explanation or a
technical mistake, report it to your teacher or teaching assistant.
Your report will help us correct the book for future classes.
We begin with supervised learning because its central task is concrete: use examples
with known answers to predict answers for new examples. These lessons introduce
the mathematics as we need it. Exercises ask you to interpret derivatives that we
provide, apply linear algebra, and measure a model’s errors. The statistics and
probability sections explain why those measurements work and what they cannot
tell you.
We then study unsupervised learning through k-means clustering. The final part
introduces neural networks. Neural networks power systems such as image
classifiers and large language models, including ChatGPT, Gemini, and Claude. Our
goal is to make the basic mechanism behind those systems understandable.
The appendices contain reference material for the libraries, programming concepts,
and mathematics used in the lessons. Use them when a symbol, function, or
operation is unfamiliar.
The book assumes some familiarity with Python, programming, and Algebra 2. This
is primarily a programming course, but you will also explain the theory and
mathematics behind your code.
We drew on the following books:
1.
Introduction to Statistical Learning by Trevor Hastie, Robert Tibshirani, Jerome
Friedman
2.
The Elements of Statistical Learning by Trevor Hastie, Robert Tibshirani, Jerome
Friedman
3.
Various books by Justin Skyzak
4.
Mathematics for Machine Learning by Marc Peter Deisenroth, A. Aldo Faisal,
Cheng Soon Ong
5.
The Matrix Cookbook by Joseph Montanez
6.
The Deep Learning Book by Ian Goodfellow, Yoshua Bengio, and Aaron Courville
6
Table of contents
Foreword
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
4
How this book is structured
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
5
Part I: Supervised Learning
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
7
1
Linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
8
1.1
What is linear regression?
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
8
1.2
The limits of a two-point estimate
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
9
1.3
Define linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
10
1.4
Fit the line
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
10
1.5
Measure and interpret error
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
10
1.6
Optimize with gradient descent
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
12
1.7
Understand gradients
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
13
1.8
Problem packet for January 16, 2026
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
16
1.9
Problem packet
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
17
2
Pseudoinverse and multiple linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
2.1
Linear algebra primer
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
18
2.2
Define multiple linear regression
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
26
2.3
Solve with the normal equation
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
26
2.4
Use gradient descent with multiple variables
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
28
2.5
Scale the features
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
28
2.6
Interpret weights after scaling
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
32
2.7
Linear algebra practice problems
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
32
2.8
Problem packet
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
35
Appendix: Programming Reference
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
3
Python and libraries
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
3.1
NumPy
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
38
Appendix: Math Fundamentals
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4
Calculus
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4.1
Limits
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4.2
Derivatives
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
4.3
Gradients
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
5
Linear algebra
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
6
Statistics and probability
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
42
Reference
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
43
Glossary of Definitions
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
.
44
7
Part I: Supervised Learning
In supervised learning, you receive inputs
𝑋
and their known responses
𝑌
. A
model estimates the relationship between them. We write that relationship as:
𝑌
=
𝑓
(
𝑋
)
+
𝜀
Here,
𝑓
is the unknown relationship that we want to estimate. The term
𝜀
represents random error that
𝑋
does not explain. We assume that this error has an
average value of zero.
A training dataset pairs each input with its known response:
𝒯︀
=
{
(
𝑥
1
,
𝑦
1
)
,
(
𝑥
2
,
𝑦
2
)
,
…
,
(
𝑥
𝑛
,
𝑦
𝑛
)
}
A learning algorithm uses those pairs to estimate
𝑓
. We call the estimate
̂
𝑓
. The
model then predicts a response for a new input with
̂
𝑦
=
̂
𝑓
(
𝑋
)
.
Definition 0.1.
A
feature
is a measurable property used as an input
𝑥
. To predict a house
price, a model might use the house’s size, number of bedrooms, and location as
features.
Definition 0.2.
A
label
is the known output
𝑦
paired with a training example. In the house
example, the sale price is the label.
The supervised learning lessons cover these methods and ideas:
•
Linear regression with the pseudoinverse
gives us a direct way to estimate
parameters. Multiple linear regression extends the method to several features.
•
Gradient descent
and
stochastic gradient descent
improve parameters
through repeated updates.
•
The bias-variance tradeoff
helps us reason about underfitting, overfitting, and
prediction error.
•
Polynomial regression
lets a linear model represent curved relationships by
adding transformed features.
•
Ridge and lasso regression
use
𝐿
2
and
𝐿
1
regularization to limit model
complexity. Lasso can also remove unhelpful features.
•
Logistic regression
predicts the probability of a category, such as yes or no.
•
k-nearest neighbors
, or
k-NN
, predicts from nearby training examples.
These names are a map, not a prerequisite. Each lesson introduces its method from
the beginning.
8
1
Linear regression
1.1
What is linear regression?
Supervised learning starts with known inputs
𝑋
and known outputs
𝑌
. We use
those examples to fit an equation that predicts
𝑌
from
𝑋
.
One of the simplest approaches assumes a linear relationship between an input and
an output. Real relationships are rarely perfectly linear, but a linear model gives you
a useful baseline. A more complex model earns its complexity only if it predicts
better than that baseline.
Solving linear regression by hand requires linear algebra. This course has no linear
algebra prerequisite, so we build the needed ideas as we use them. We will later use
scikit-learn to fit models, but first we need to understand what a fitted line means
and how a computer can find one.
Example 1.1.1.
Suppose you use two thermometers to record the same temperatures in Celsius
and Fahrenheit. Your measurements look like this:
Celsius (x)
Fahrenheit (y)
0
31.8
5
41.9
10
49.2
15
60.1
20
67.4
25
78.9
30
87.5
Plot the measurements:
The exact conversion equation is:
𝑦
=
1
.
8
𝑥
+
3
2
9
Now pretend that you do not know the conversion equation. You can estimate
the slope from two measurements with
𝑚
=
𝑦
2
−
𝑦
1
𝑥
2
−
𝑥
1
. The table also gives an
estimated intercept of
3
1
.
8
at
𝑥
=
0
. Together, these estimates produce the
model:
̂
𝑦
=
1
.
8
5
𝑥
+
3
1
.
8
Note:
The
hat
in
̂
𝑦
marks a predicted value. It distinguishes a model’s
prediction from a measured value
𝑦
. In our model,
̂
𝑦
=
1
.
8
5
𝑥
+
3
1
.
8
̂
𝑦
is the predicted Fahrenheit temperature for a Celsius input
𝑥
.
Statistics and machine learning use the same notation for estimated
coefficients, such as
̂
𝑤
and
̂
𝑏
.
You have made a statistical model from data. The two-point method worked for this
clean example, but real measurements expose its limits.
1.2
The limits of a two-point estimate
Suppose you measure temperatures around campus instead of using the clean table
above.
The AI class collected the following data in 2025:
The points do not fall on one perfect line. A thermometer may be poorly calibrated.
Direct sunlight or a gust of wind may affect one measurement. These sources of
noise make any chosen pair of points unreliable.
Linear regression uses all the observations instead of trusting one pair. It chooses
the line with the smallest total error across the dataset. The line does not pass
through every point because noisy data rarely allows that.
10
1.3
Define linear regression
Definition 1.3.1.
Simple linear regression
models the relationship between one input
𝑥
and
one output
𝑦
with
̂
𝑦
=
𝑤
𝑥
+
𝑏
. The weight
𝑤
controls the slope, and the
intercept
𝑏
controls where the line crosses the vertical axis. For the
temperature conversion, the exact values are
𝑤
=
1
.
8
and
𝑏
=
3
2
.
The following synthetic dataset imitates a larger set of noisy measurements:
1.4
Fit the line
A computer can estimate
𝑤
and
𝑏
in several ways. One method, gradient descent,
starts with initial values and updates them to reduce error. Another method solves
for the best values with linear algebra. Scikit-learn’s
LinearRegression
estimator
uses a linear algebra solver rather than gradient descent.
Both methods need a precise definition of error. We will define that measurement
before studying gradient descent.
1.5
Measure and interpret error
Example 1.5.1.
The table pairs five observed values
𝑦
with predictions
̂
𝑦
:
𝑥
𝑦
̂
𝑦
1
3
2.5
2
5
5.2
3
4
4.1
4
7
6.8
11
5
6
6.3
The difference
𝑦
𝑖
−
̂
𝑦
𝑖
is the residual for observation
𝑖
. Residual sum of squares
combines all residuals into one measurement.
Definition 1.5.1.
Residual sum of squares
, or
RSS
, is:
RSS
=
∑
𝑁
𝑖
=
0
(
𝑦
𝑖
−
̂
𝑦
𝑖
)
2
For each observation, subtract the prediction from the measured value. Square each
residual, then add the squares. Squaring serves three purposes:
1.
Squaring makes all residuals positive, so large underpredictions and
overpredictions both contribute to the total error.
2.
It gives large residuals more influence. A residual of 4 contributes 16, while a
residual of 2 contributes 4.
3.
It produces a smooth loss function that we can optimize with derivatives.
The following diagram shows each residual as the vertical distance between a point
and the line:
12
RSS already gives one total, but its size grows with the number of observations.
Mean squared error divides that total by
𝑁
so datasets of different sizes are easier to
compare.
Definition 1.5.2.
Mean squared error
, or
MSE
, is:
𝐿
(
𝑤
,
𝑏
)
=
1
𝑁
∑
𝑁
𝑖
=
0
(
𝑦
𝑖
−
̂
𝑦
𝑖
)
2
For the model
̂
𝑦
𝑖
=
𝑤
𝑥
𝑖
+
𝑏
, we can expand the equation:
𝐿
(
𝑤
,
𝑏
)
=
1
𝑁
∑
𝑁
𝑖
=
0
(
𝑦
𝑖
−
(
𝑤
𝑥
𝑖
+
𝑏
)
)
2
The factor
1
𝑁
averages the squared residuals.
The following Python function calculates MSE directly:
#
Assume X is a spreadsheet of our data and X[i] is the ith row
def
mse
(
W
,
b
,
X
,
y
)
:
n
=
len
(
X
)
total_error
=
0
for
i
in
range
(
n
)
:
prediction
=
W
*
X
[
i
]
+
b
total_error
+=
(
y
[
i
]
-
prediction
)
*
*
2
return
total_error
/
n
1.6
Optimize with gradient descent
MSE tells us how well a pair of values
𝑤
and
𝑏
fits the data. We still need a method
that finds the values with the lowest MSE.
Gradient descent improves an initial guess through repeated updates. Each update
uses a derivative.
A derivative measures how a function changes when its input changes by a small
amount. On a graph, the derivative gives the slope at one point. A positive
derivative means the function rises as the input increases. A negative derivative
means the function falls.
For linear regression, we ask how the loss changes when
𝑤
changes. The derivative
of the loss with respect to
𝑤
answers that question.
13
If the derivative is positive, increasing
𝑤
increases the loss, so we decrease
𝑤
. If the
derivative is negative, increasing
𝑤
decreases the loss, so we increase
𝑤
. In both
cases, moving opposite the derivative lowers the loss locally.
Gradient descent applies this update rule:
𝑤
←
𝑤
−
𝜂
𝜕
𝐿
𝜕
𝑤
The learning rate
𝜂
controls the size of each update. A very small learning rate takes
many updates. A learning rate that is too large can jump past the minimum and
cause the loss to increase.
Update
𝑏
with the same pattern:
𝑏
←
𝑏
−
𝜂
𝜕
𝐿
𝜕
𝑏
Repeat both updates until the loss stops decreasing by a meaningful amount.
You do not need to derive the formulas in this course. We will provide the partial
derivatives:
𝜕
𝐿
𝜕
𝑤
=
−
2
𝑁
∑
𝑁
𝑖
=
1
(
𝑦
𝑖
−
(
𝑤
𝑥
𝑖
+
𝑏
)
)
𝑥
𝑖
𝜕
𝐿
𝜕
𝑏
=
−
2
𝑁
∑
𝑁
𝑖
=
1
(
𝑦
𝑖
−
(
𝑤
𝑥
𝑖
+
𝑏
)
)
The corresponding NumPy updates are:
dw
=
(
-
2
.
0
/
N
)
*
np
.
sum
(
error
*
x_scaled
,
dtype
=
np
.
float64
)
w
-=
lr
*
dw
db
=
(
-
2
.
0
/
N
)
*
np
.
sum
(
error
,
dtype
=
np
.
float64
)
b
-=
lr
*
db
1.7
Understand gradients
Before we use a gradient with two inputs, consider the one-input function
𝑓
(
𝑥
)
=
𝑥
2
:
14
Suppose
𝑓
(
𝑤
)
=
𝑤
2
represents the loss for the model
̂
𝑦
=
𝑤
𝑥
. If
𝑤
=
2
, the loss is
4. Training seeks the value of
𝑤
at the lowest point of the loss function.
For this parabola, we can find the minimum from its vertex. A more complex loss
function may not have such an obvious minimum:
15
Gradient descent searches for a low point without relying on a visual graph. For a
function with one input, it uses the derivative.
The derivative is the slope of the tangent line at a point. A tangent line matches the
function’s direction at that one point. The purple line below is a tangent line:
Figure 1: Example tangent line in purple for a loss function
At
𝑥
=
0
.
4
, the tangent line has slope
𝑚
=
−
2
.
2
6
5
2
8
1
5
8
0
5
7
. The negative value
tells us that the function falls as
𝑥
increases near this point.
We write the derivative of
𝑓
as
𝑓
′
. For a loss function with one weight
𝑤
, the
update becomes:
𝑤
←
𝑤
−
𝜂
𝑓
′
(
𝑤
)
A real model often has several parameters. Its loss function therefore has several
inputs. A gradient collects the slope for each input so we can update every
parameter.
Definition 1.7.1.
A
gradient
generalizes the derivative to a function with several inputs. For
𝑓
(
𝑥
,
𝑦
)
=
𝑥
2
+
𝑦
2
, the gradient is a vector that contains one partial derivative
for each input:
∇
𝑓
=
(
𝜕
𝑓
𝜕
𝑥
,
𝜕
𝑓
𝜕
𝑦
)
=
(
2
𝑥
,
2
𝑦
)
Each component of
∇
𝑓
is a
partial derivative
. A partial derivative measures how
the function changes with respect to one input while the other inputs stay fixed. For
the linear regression loss, the relevant inputs are
𝑤
and
𝑏
:
𝜕
𝐿
𝜕
𝑤
=
−
2
𝑁
∑
𝑁
𝑖
=
1
(
𝑦
𝑖
−
(
𝑤
𝑥
𝑖
+
𝑏
)
)
𝑥
𝑖
𝜕
𝐿
𝜕
𝑏
=
−
2
𝑁
∑
𝑁
𝑖
=
1
(
𝑦
𝑖
−
(
𝑤
𝑥
𝑖
+
𝑏
)
)
Together, these two partial derivatives form the gradient of
𝐿
(
𝑤
,
𝑏
)
.
The following graph shows
𝐿
(
𝑤
,
𝑏
)
for the model used in the next lab:
16
Figure 2: The loss surface for our model:
̂
𝑦
=
1
.
8
5
𝑥
+
3
1
.
4
7
The red line traces the parameter values that gradient descent tried on its way
toward the minimum. Watch the
gradient descent training video
to see the updates
in real time.
1.8
Problem packet for January 16, 2026
Problem 1.
You are modeling Fahrenheit from Celsius. You try three baselines:
•
Constant model:
̂
𝑦
=
𝑐
•
Linear model:
̂
𝑦
=
𝑊
𝑥
+
𝑏
•
Piecewise linear with one breakpoint at
𝑥
=
1
0
:
̂
𝑦
=
𝑊
1
𝑥
+
𝑏
1
for
𝑥
≤
1
0
, and
̂
𝑦
=
𝑊
2
𝑥
+
𝑏
2
for
𝑥
>
1
0
Without using code, rank the baselines from strongest to weakest. Define a
strong
baseline
in terms of training and evaluation. Explain how it affects comparisons
with later models, and give at least two problems caused by a weak baseline.
Problem 2.
In the text,
̂
𝑦
is predicted output, and
̂
𝑊
,
̂
𝑏
are estimated parameters.
Explain the difference between an observed
𝑦
, a prediction
̂
𝑦
, the true but unknown
relationship
𝑓
, estimated parameters
̂
𝑊
and
̂
𝑏
, and measurement noise. Then
answer: if you re-collect the dataset tomorrow with the
same thermometers
,
which of these are expected to change and why?
Problem 3.
The chapter lists reasons to square residuals, explain why these reasons
are important. Then argue for one scenario where squaring residuals is a bad idea,
and name a better alternative loss. (You can use the internet to find the answer to
this question.)
Problem 4.
You train the same model form
̂
𝑦
=
𝑊
𝑥
+
𝑏
on two datasets A and B.
17
•
Dataset A has
𝑁
=
2
0
points, RSS = 120.
•
Dataset B has
𝑁
=
2
0
0
points, RSS = 950.
1.
Compute
MSE
for both.
2.
Explain why the
RSS
value can mislead you across dataset sizes.
3.
Give one situation where
RSS
is still useful or preferred (be specific).
Problem 5.
Suppose your fitted line gives small MSE, but when you plot residuals
𝑟
𝑖
=
𝑦
𝑖
−
̂
𝑦
𝑖
versus
𝑥
𝑖
, you see a clear U-shape. Explain what this implies about: the
linearity assumption, whether the bias term
̂
𝑏
is wrong, what kind of model change
would address it, and why
MSE
alone did not warn you. Give at least two possible
model changes.
Problem 6.
Given:
𝑥
=
[
−
5
,
0
,
5
,
1
0
]
,
𝑦
=
[
2
0
.
0
,
3
1
.
8
,
4
0
.
0
,
5
5
.
0
]
, Model:
̂
𝑦
=
1
.
6
𝑥
+
3
1
.
8
Compute:
̂
𝑦
for each
𝑥
residuals
𝑟
𝑖
=
𝑦
𝑖
−
̂
𝑦
𝑖
RSS
and
MSE
Identify
which point contributes most to
RSS
and explain why
Problem 7.
Given:
𝑥
=
[
−
5
,
0
,
5
,
1
0
]
𝑦
=
[
2
0
.
0
,
3
1
.
8
,
4
0
.
0
,
5
5
.
0
]
Model:
̂
𝑦
=
1
.
6
𝑥
+
3
1
.
8
Compute:
̂
𝑦
for each
𝑥
residuals
𝑟
𝑖
=
𝑦
𝑖
−
̂
𝑦
RSS
and
MSE
Identify which
point contributes most to
RSS
and explain why
Problem 8.
Dataset:
𝑥
=
[
0
,
5
,
1
0
,
1
5
,
2
0
]
𝑦
=
[
3
2
.
0
,
4
1
.
0
,
5
0
.
5
,
6
0
.
0
,
6
8
.
0
]
Two
candidate models: A:
̂
𝑦
=
1
.
8
𝑥
+
3
2
𝐵
:
accent(y, hat) = 1.9 x + 31 Compute RSS for
both and decide which is better under RSS/MSE. Then answer: which model is
more plausible physically
, and can plausibility disagree with MSE here?
1.9
Problem packet
Theory questions
Problem 1.
The text describes linear models as a “baseline.” Explain the importance
of establishing a baseline model before moving on to more complex machine
learning algorithms.
Problem 2.
In the equation
̂
𝑦
=
1
.
8
5
𝑥
+
3
1
.
8
, explain what the hat notation
̂
(
)
means and why it distinguishes predictions from observations.
Problem 3.
The lesson provides three specific reasons for squaring residuals in the
RSS formula. List them and explain why making the loss function “smooth and
differentiable” is beneficial for optimization.
Problem 4.
What is the mathematical difference between Residual Sum of Squares
(
RSS
) and Mean Squared Error (
MSE
)? Why is
MSE
generally preferred when
working with datasets of varying sizes?
Practice problems
Problem 5.
You are given the coefficients
𝑎
=
1
,
𝑏
=
4
, and
𝑐
=
2
for the function
𝑓
(
𝑥
)
=
𝑥
2
+
4
𝑥
+
2
. Using the derivative
𝑓
′
(
𝑥
)
=
2
𝑥
+
4
, write a Python function
to find the minimum of
𝑓
(
𝑥
)
using gradient descent. Start at
𝑥
=
1
0
, use
𝜂
=
0
.
1
,
and run for 10 iterations.
Problem 6.
Calculate the
RSS
and
MSE
by hand for the following dataset given
the model
̂
𝑦
=
2
𝑥
+
1
:
•
𝑥
=
[
1
,
2
,
3
]
•
𝑦
=
[
3
,
6
,
7
]
Problem 7.
Given
𝑥
=
[
1
,
2
,
3
]
and
𝑦
=
[
2
,
3
,
4
]
, and initial parameters
𝑊
=
0
and
𝑏
=
0
, compute:
•
The predicted values
̂
𝑦
•
The residuals
(
𝑦
𝑖
−
̂
𝑦
𝑖
)
•
The current
MSE
18
Problem 8.
Using the data and initial parameters from Problem 9, perform one full
batch gradient descent update to find
𝑊
new
and
𝑏
new
. Use
𝜂
=
0
.
1
and the formulas:
𝜕
𝐿
𝜕
𝑊
=
−
2
𝑁
∑
𝑁
𝑖
=
1
(
𝑦
𝑖
−
(
𝑊
𝑥
𝑖
+
𝑏
)
)
𝑥
𝑖
𝜕
𝐿
𝜕
𝑏
=
−
2
𝑁
∑
𝑁
𝑖
=
1
(
𝑦
𝑖
−
(
𝑊
𝑥
𝑖
+
𝑏
)
)
Note: Use the sign convention from the provided Python code where the
gradient is subtracted.
Problem 9.
A thermometer model is trained to
̂
𝑦
=
1
.
8
5
𝑥
+
3
1
.
8
. If the actual
temperature is
0
°
𝐶
and the observed Fahrenheit reading is
3
1
.
8
, what is the
residual? If the actual temperature is
3
0
°
𝐶
and the observed reading is
8
7
.
5
, what is
the residual?
Problem 10.
Write a Python function
get_error(y_true, y_pred)
that returns the
Mean Squared Error using only the standard library (no numpy). Assume both
inputs are lists of equal length.
Complete the remaining problems in the
introduction assignments notebook
.
2
Pseudoinverse and multiple linear regression
The previous model used one feature
𝑋
to predict
𝑌
. A house-price model may
need size, bedroom count, and location at the same time.
Linear algebra gives us notation and operations for working with these features
together.
2.1
Linear algebra primer
Note:
This section introduces several symbols and operations. For another
visual explanation, watch 3Blue1Brown’s
Essence of Linear Algebra
series.
Some videos cover material beyond this course.
Sets
In mathematics, a
set
is a collection of distinct objects. A programming set follows
the same two basic rules: each value appears once, and its position does not matter.
Common number sets
Number sets tell us which values a variable may contain:
•
Natural numbers
ℕ
are the positive counting numbers
{
1
,
2
,
3
,
…
}
. We use them
for counts that cannot be zero or negative.
•
Integers
ℤ
are whole numbers, including zero and negative numbers:
{
…
,
−
2
,
−
1
,
0
,
1
,
2
,
…
}
.
•
Rational numbers
ℚ
can be written as a fraction
𝑝
𝑞
, where
𝑝
,
𝑞
∈
ℤ
and
𝑞
≠
0
.
The set includes terminating decimals such as
0
.
7
5
.
•
Real numbers
ℝ
are all the points on the number line. This set includes rational
numbers and irrational numbers such as
𝜋
and
√
2
.
•
Complex numbers
ℂ
contain a real part and an imaginary part. They appear in
fields such as signal processing but are not used in this lesson.
Relationships between sets
These number sets are nested. Every natural number is an integer, every integer is
rational, and every rational number is real. The symbol
⊂
means “is a subset of”:
19
ℕ
⊂
ℤ
⊂
ℚ
⊂
ℝ
Vectors, vector addition, and scalar multiplication
Suppose you want to predict a house price from these features:
•
living area in square feet,
•
number of bedrooms, and
•
age in years.
Write the three feature names in a fixed order:
(
square footage
,
number of bedrooms
,
age
)
.
The values
(
2
5
0
0
,
4
,
1
0
)
describe a 2,500-square-foot house with 4 bedrooms that is
10 years old. A new house with the same size and bedroom count has values
(
2
5
0
0
,
4
,
0
)
.
This ordered list of values is a vector. Because it contains three values, it is a 3-
vector.
We can write the same vector vertically:
(
2
5
0
0
4
1
0
)
A
scalar
is one number. For example,
2
0
0
is a scalar, while
(
2
0
0
,
3
0
0
,
2
5
)
is a
vector.
Algebra often uses symbols such as
𝑥
and
𝑦
for scalars. We use bold symbols, such
as
𝒙
and
𝒚
, for vectors.
Definition 2.1.1.
For a positive integer
𝑛
, an
n-vector
is an ordered list of
𝑛
real numbers. The
symbol
ℝ
𝑛
denotes the set of all real n-vectors.
You can also picture a vector as an arrow. For
𝑛
=
2
, the vector
𝒗
=
(
𝑣
1
𝑣
2
)
points
from the origin
(
0
,
0
)
to the point
(
𝑣
1
,
𝑣
2
)
. A 3-vector works the same way in three-
dimensional space.
In physics, vectors can represent quantities with magnitude and direction, such as
velocity and force. In machine learning, vectors usually store related numerical
values. A dataset may contain vectors with thousands of entries.
Example 2.1.1.
Suppose there are 100 students in the AI class. We can keep track of all their
grades on the first test by using a 100-vector
𝑬
=
(
𝐸
1
𝐸
2
⋮
𝐸
1
0
0
)
Here
𝐸
1
is the first exam grade of the first student,
𝐸
2
the first exam grade of
the second student, and so on.
We will use two basic vector operations: vector addition and scalar multiplication.
20
Definition 2.1.2.
The sum
𝒗
+
𝒘
of two vectors is defined only when
𝒗
and
𝒘
are
𝑛
-vectors. In
that case, we define their sum by the rule
(
𝑣
1
𝑣
2
⋮
𝑣
𝑛
)
+
(
𝑤
1
𝑤
2
⋮
𝑤
𝑛
)
=
(
𝑣
1
+
𝑤
1
𝑣
2
+
𝑤
2
⋮
𝑣
𝑛
+
𝑤
𝑛
)
.
Definition 2.1.3.
To multiply an n-vector by a scalar
𝑐
, multiply every component by
𝑐
:
𝑐
(
𝑣
1
𝑣
2
⋮
𝑣
𝑛
)
=
(
𝑐
𝑣
1
𝑐
𝑣
2
⋮
𝑐
𝑣
𝑛
)
.
Example 2.1.2.
Practice problem: Let
𝒗
=
(
2
−
1
3
)
and
𝒘
=
(
5
4
−
2
)
. Compute
𝒗
+
𝒘
and
−
2
𝒗
.
𝒗
+
𝒘
=
(
2
+
5
−
1
+
4
3
+
(
−
2
)
)
=
(
7
3
1
)
.
−
2
𝒗
=
(
−
4
2
−
6
)
.
Matrices, notation, and dimensions
You may have seen matrices in Algebra 2. We begin with the definition:
Definition 2.1.4.
A matrix is a rectangular array of numbers or other mathematical objects with
elements or entries arranged in rows and columns. A matrix with
𝑝
rows and
𝑑
columns is called a
𝑝
×
𝑑
matrix.
The symbol
∈
means “is an element of.” In linear algebra, this notation can describe
both a matrix’s dimensions and the type of its entries.
The statement
𝑋
∈
ℝ
𝑝
×
𝑑
gives two facts about
𝑋
:
1.
Every entry is a real number.
2.
The matrix has
𝑝
rows and
𝑑
columns.
A general
𝑝
×
𝑑
matrix looks like this:
𝑋
=
(
𝑥
1
1
𝑥
2
1
⋮
𝑥
𝑝
1
𝑥
1
2
𝑥
2
2
⋮
𝑥
𝑝
2
…
…
⋱
…
𝑥
1
𝑑
𝑥
2
𝑑
⋮
𝑥
𝑝
𝑑
)
Subscripts identify an entry’s row and column. In
𝑥
1
1
, the first
1
identifies the first
row and the second
1
identifies the first column. The entry
𝑥
𝑝
𝑑
lies in row
𝑝
and
column
𝑑
.
21
If a dataset contains 100 houses and 5 features per house, then
𝑋
∈
ℝ
1
0
0
×
5
. The
matrix contains 500 real-valued entries arranged in 100 rows and 5 columns.
A feature is one measurement that describes an observation. For a house,
examples include living area and bedroom count. Each feature occupies one
column of
𝑋
.
A label is the target value for one observation. House price is a possible label.
The label vector has one entry for each row of
𝑋
.
Example 2.1.3.
Practice problem: Suppose
𝐴
∈
ℝ
2
×
3
and
𝐵
∈
ℝ
3
×
4
. What is the shape of
𝐴
𝐵
?
Since the inner dimensions match (
3
), the product is defined and
𝐴
𝐵
∈
ℝ
2
×
4
.
Add and subtract matrices
Definition 2.1.5.
The sum of 2 matrices
𝐴
=
(
𝑎
1
1
𝑎
2
1
⋮
𝑎
𝑝
1
𝑎
1
2
𝑎
2
2
⋮
𝑎
𝑝
2
⋮
⋮
⋱
⋮
𝑎
1
𝑑
𝑎
2
𝑑
⋮
𝑎
𝑝
𝑑
)
and
𝐵
=
(
𝑏
1
1
𝑏
2
1
⋮
𝑏
𝑝
1
𝑏
1
2
𝑏
2
2
⋮
𝑏
𝑝
2
⋮
⋮
⋱
⋮
𝑏
1
𝑑
𝑏
2
𝑑
⋮
𝑏
𝑝
𝑑
)
is
defined only when
𝐴
and
𝐵
are of the same size. In that case, we define their
sum by the rule
𝐴
+
𝐵
=
(
𝑎
1
1
+
𝑏
1
1
𝑎
2
1
+
𝑏
2
1
⋮
𝑎
𝑝
1
+
𝑏
𝑝
1
𝑎
1
2
+
𝑏
1
2
𝑎
2
2
+
𝑏
2
2
⋮
𝑎
𝑝
2
+
𝑏
𝑝
2
⋮
⋮
⋱
⋮
𝑎
1
𝑑
+
𝑏
1
𝑑
𝑎
2
𝑑
+
𝑏
2
𝑑
⋮
𝑎
𝑝
𝑑
+
𝑏
𝑝
𝑑
)
.
Example 2.1.4.
Let
𝐴
=
(
1
3
2
4
)
and
𝐵
=
(
5
7
6
8
)
. Then
𝐴
+
𝐵
=
(
1
+
5
3
+
7
2
+
6
4
+
8
)
=
(
6
1
0
8
1
2
)
.
Definition 2.1.6.
The difference of 2 matrices
𝐴
=
(
𝑎
1
1
𝑎
2
1
⋮
𝑎
𝑝
1
𝑎
1
2
𝑎
2
2
⋮
𝑎
𝑝
2
⋮
⋮
⋱
⋮
𝑎
1
𝑑
𝑎
2
𝑑
⋮
𝑎
𝑝
𝑑
)
and
𝐵
=
(
𝑏
1
1
𝑏
2
1
⋮
𝑏
𝑝
1
𝑏
1
2
𝑏
2
2
⋮
𝑏
𝑝
2
⋮
⋮
⋱
⋮
𝑏
1
𝑑
𝑏
2
𝑑
⋮
𝑏
𝑝
𝑑
)
is also defined only when
𝐴
and
𝐵
are of the same size (in other words
𝐴
∈
ℝ
𝑝
×
𝑑
and
𝐵
∈
ℝ
𝑝
×
𝑑
). In that case, we define their difference by using sums but
multiplying the second matrix by the scalar
−
1
:
First, multiply the second matrix by
−
1
:
−
𝐵
=
(
−
𝑏
1
1
−
𝑏
2
1
⋮
−
𝑏
𝑝
1
−
𝑏
1
2
−
𝑏
2
2
⋮
−
𝑏
𝑝
2
⋮
⋮
⋱
⋮
−
𝑏
1
𝑑
−
𝑏
2
𝑑
⋮
−
𝑏
𝑝
𝑑
)
Then, add the matrices:
22
𝐴
−
𝐵
=
𝐴
+
(
−
𝐵
)
=
(
𝑎
1
1
−
𝑏
1
1
𝑎
2
1
−
𝑏
2
1
⋮
𝑎
𝑝
1
−
𝑏
𝑝
1
𝑎
1
2
−
𝑏
1
2
𝑎
2
2
−
𝑏
2
2
⋮
𝑎
𝑝
2
−
𝑏
𝑝
2
⋮
⋮
⋱
⋮
𝑎
1
𝑑
−
𝑏
1
𝑑
𝑎
2
𝑑
−
𝑏
2
𝑑
⋮
𝑎
𝑝
𝑑
−
𝑏
𝑝
𝑑
)
Dot product
Definition 2.1.7.
The dot product of two vectors
𝒂
=
(
𝑎
1
𝑎
2
…
𝑎
𝑛
)
and
𝒃
=
(
𝑏
1
𝑏
2
…
𝑏
𝑛
)
is defined as:
𝒂
⋅
𝒃
=
∑
𝑛
𝑖
=
1
𝑎
𝑖
𝑏
𝑖
=
𝑎
1
𝑏
1
+
𝑎
2
𝑏
2
+
…
+
𝑎
𝑛
𝑏
𝑛
Example 2.1.5.
Let
𝒂
=
(
1
2
3
)
and
𝒃
=
(
4
5
6
)
. Pair the corresponding components, multiply
each pair, and add the products:
𝒂
⋅
𝒃
=
(
1
×
4
)
+
(
2
×
5
)
+
(
3
×
6
)
=
4
+
1
0
+
1
8
=
3
2
Example 2.1.6.
Practice problem: Let
𝒖
=
(
2
−
1
4
)
and
𝒗
=
(
3
0
−
2
)
. Compute
𝒖
⋅
𝒗
.
𝒖
⋅
𝒗
=
(
2
)
(
3
)
+
(
−
1
)
(
0
)
+
(
4
)
(
−
2
)
=
6
+
0
−
8
=
−
2
.
Matrix multiplication uses this same pair, multiply, and add pattern.
Matrix multiplication
Definition 2.1.8.
Suppose that we have
𝐴
∈
ℝ
𝑟
×
𝑑
and
𝐵
∈
ℝ
𝑑
×
𝑠
. Then the product of
𝐴
and
𝐵
is denoted
𝐴
𝐵
. The
(
𝑖
,
𝑗
)
th element of
(
𝐴
𝐵
)
is computed by multiplying each
element of the
𝑖
th row of
𝐴
by the corresponding element of the
𝑗
th column of
𝐵
. That is,
(
𝐴
𝐵
)
𝑖
𝑗
=
∑
𝑑
𝑘
=
1
𝑎
𝑖
𝑘
𝑏
𝑘
𝑗
.
Example 2.1.7.
Consider these two matrices:
𝑨
=
(
1
3
2
4
)
and
𝑩
=
(
5
7
6
8
)
.
Then
23
𝑨
𝑩
=
(
1
3
2
4
)
(
5
7
6
8
)
=
(
1
×
5
+
2
×
7
3
×
5
+
4
×
7
1
×
6
+
2
×
8
3
×
6
+
4
×
8
)
=
(
1
9
4
3
2
2
5
0
)
.
Example 2.1.8.
Practice problem: Let
𝐴
=
(
2
3
−
1
4
0
1
)
and
𝐵
=
(
1
−
2
5
2
0
−
1
)
. Compute
𝐴
𝐵
.
First check dimensions:
𝐴
∈
ℝ
2
×
3
and
𝐵
∈
ℝ
3
×
2
, so
𝐴
𝐵
∈
ℝ
2
×
2
.
𝐴
𝐵
=
(
2
×
1
+
(
−
1
)
×
(
−
2
)
+
0
×
5
3
×
1
+
4
×
(
−
2
)
+
1
×
5
2
×
2
+
(
−
1
)
×
0
+
0
×
(
−
1
)
3
×
2
+
4
×
0
+
1
×
(
−
1
)
)
=
(
4
0
4
5
)
.
The product has
𝑟
rows and
𝑠
columns. We can compute
𝐴
𝐵
only when the number
of columns in
𝐴
equals the number of rows in
𝐵
.
Transpose
Definition 2.1.9.
The
transpose
of a matrix swaps its rows and columns. Row 1 becomes
column 1, row 2 becomes column 2, and so on. A superscript
𝑇
denotes this
operation.
If
𝑋
∈
ℝ
𝑝
×
𝑑
, then
𝑋
𝑇
∈
ℝ
𝑑
×
𝑝
. Entry
(
𝑖
,
𝑗
)
of the transpose comes from entry
(
𝑗
,
𝑖
)
of the original matrix:
(
𝑋
𝑇
)
𝑖
𝑗
=
𝑋
𝑗
𝑖
.
Example 2.1.9.
For example, if we have a matrix
𝑋
∈
ℝ
3
×
2
we can take its transpose
𝑋
𝑇
∈
ℝ
2
×
3
by swapping the rows and columns:
𝑋
=
(
𝑥
1
1
𝑥
2
1
𝑥
3
1
𝑥
1
2
𝑥
2
2
𝑥
3
2
)
𝑋
𝑇
=
(
𝑥
1
1
𝑥
1
2
𝑥
2
1
𝑥
2
2
𝑥
3
1
𝑥
3
2
)
Example 2.1.10.
Practice problem: If
𝐶
=
(
0
2
3
5
−
1
4
)
, compute
𝐶
𝑇
.
𝐶
𝑇
=
(
0
3
−
1
2
5
4
)
.
Identity matrix
Multiplying a scalar by 1 leaves the scalar unchanged. The identity matrix has the
same role in matrix multiplication.
Definition 2.1.10.
24
The identity matrix
𝐼
𝑛
(or just
𝐼
when the size is clear) is a square
𝑛
×
𝑛
matrix with 1s on the diagonal and 0s everywhere else:
𝐼
3
=
(
1
0
0
0
1
0
0
0
1
)
For any matrix
𝐴
of compatible size:
𝐴
𝐼
=
𝐼
𝐴
=
𝐴
Example 2.1.11.
Verify the property with a
2
×
2
matrix:
(
2
4
3
5
)
(
1
0
0
1
)
=
(
2
×
1
+
3
×
0
4
×
1
+
5
×
0
2
×
0
+
3
×
1
4
×
0
+
5
×
1
)
=
(
2
4
3
5
)
The product equals the original matrix.
Determinant
The determinant is a scalar calculated from a square matrix. A determinant of zero
tells us that the matrix has no inverse.
Definition 2.1.11.
For a
2
×
2
matrix
𝐴
=
(
𝑎
𝑐
𝑏
𝑑
)
, the
determinant
is defined as:
det
(
𝐴
)
=
𝑎
𝑑
−
𝑏
𝑐
The determinant is often written as
|
𝐴
|
or
det
(
𝐴
)
.
Example 2.1.12.
For
𝐴
=
(
3
1
2
4
)
:
det
(
𝐴
)
=
(
3
)
(
4
)
−
(
2
)
(
1
)
=
1
2
−
2
=
1
0
The absolute value of the determinant tells us how matrix multiplication scales area.
If
det
(
𝐴
)
=
2
, multiplication by
𝐴
doubles area. If
det
(
𝐴
)
=
0
, a two-dimensional
region collapses to a line or a point. That loss of information prevents an inverse.
Definition 2.1.12.
A matrix
𝐴
is
invertible
(has an inverse) if and only if
det
(
𝐴
)
≠
0
.
Example 2.1.13.
Check whether
𝐵
=
(
1
2
2
4
)
has an inverse:
det
(
𝐵
)
=
(
1
)
(
4
)
−
(
2
)
(
2
)
=
4
−
4
=
0
Because
det
(
𝐵
)
=
0
, the matrix has no inverse. The second row is twice the
first, so the rows do not contain independent information.
25
Inverse matrix
For a nonzero scalar, multiplication by its reciprocal produces 1. An inverse matrix
follows the same pattern and produces the identity matrix.
Definition 2.1.13.
For a square matrix
𝐴
, its
inverse
𝐴
−
1
is the matrix such that:
𝐴
𝐴
−
1
=
𝐴
−
1
𝐴
=
𝐼
Not every matrix has an inverse. A matrix that has an inverse is called
invertible
or
non-singular
.
Compute the inverse of a
2
×
2
matrix
A
2
×
2
matrix has a direct formula:
Definition 2.1.14.
If
𝐴
=
(
𝑎
𝑐
𝑏
𝑑
)
and
det
(
𝐴
)
≠
0
, then:
𝐴
−
1
=
1
det
(
𝐴
)
(
𝑑
−
𝑐
−
𝑏
𝑎
)
Swap the diagonal entries, negate the off-diagonal entries, and divide every
entry by the determinant.
Example 2.1.14.
Compute the inverse of
𝐴
=
(
3
1
2
4
)
.
Step 1:
Compute the determinant.
det
(
𝐴
)
=
(
3
)
(
4
)
−
(
2
)
(
1
)
=
1
2
−
2
=
1
0
Since
det
(
𝐴
)
=
1
0
≠
0
, the inverse exists.
Step 2:
Apply the formula.
𝐴
−
1
=
1
1
0
(
4
−
1
−
2
3
)
=
(
4
1
0
−
1
1
0
−
2
1
0
3
1
0
)
=
(
0
.
4
−
0
.
1
−
0
.
2
0
.
3
)
Step 3:
Verify by computing
𝐴
𝐴
−
1
.
𝐴
𝐴
−
1
=
(
3
1
2
4
)
(
0
.
4
−
0
.
1
−
0
.
2
0
.
3
)
=
(
3
(
0
.
4
)
+
2
(
−
0
.
1
)
1
(
0
.
4
)
+
4
(
−
0
.
1
)
3
(
−
0
.
2
)
+
2
(
0
.
3
)
1
(
−
0
.
2
)
+
4
(
0
.
3
)
)
=
(
1
.
2
−
0
.
2
0
.
4
−
0
.
4
−
0
.
6
+
0
.
6
−
0
.
2
+
1
.
2
)
=
(
1
0
0
1
)
=
𝐼
✓
Example 2.1.15.
Compute the inverse of
𝐴
=
(
1
3
2
4
)
.
Step 1:
det
(
𝐴
)
=
(
1
)
(
4
)
−
(
2
)
(
3
)
=
4
−
6
=
−
2
Step 2:
Apply the formula:
26
𝐴
−
1
=
1
−
2
(
4
−
3
−
2
1
)
=
(
−
2
3
2
1
−
1
2
)
Work with larger matrices
For larger matrices, Gaussian elimination and cofactor expansion require many
steps.
Cofactor expansion produces 6 terms for a
3
×
3
determinant and 24 terms for a
4
×
4
determinant. The number grows as
𝑛
!
, so a
5
×
5
determinant has 120 terms
before simplification.
In practice, numerical libraries calculate matrix inverses with more efficient
algorithms:
import
numpy
as
np
A
=
np
.
array
(
[
[
3
,
2
]
,
[
1
,
4
]
]
)
A_inv
=
np
.
linalg
.
inv
(
A
)
print
(
A_inv
)
#
[[ 0.4 -0.2]
#
[-0.1 0.3]]
In this course, you will calculate inverses by hand only for
2
×
2
matrices. For
larger matrices, use NumPy and interpret the result.
2.2
Define multiple linear regression
The previous model predicted a house price from one feature. A model that uses
size, bedroom count, and location needs a separate weight for each feature.
Multiple linear regression estimates all these weights together.
Definition 2.2.1.
Give each of the
𝑝
features its own weight. The prediction is:
̂
𝑦
=
𝛽
0
+
𝑋
1
𝛽
1
+
𝑋
2
𝛽
2
+
…
+
𝑋
𝑝
𝛽
𝑝
In this equation:
•
𝛽
0
is the intercept. It replaces
𝑏
from the previous chapter.
•
𝑋
1
,
𝑋
2
,
…
,
𝑋
𝑝
are the features, such as size and bedroom count.
•
𝛽
1
,
𝛽
2
,
…
,
𝛽
𝑝
are the feature weights.
Library documentation often writes the same calculation as one matrix
multiplication:
̂
𝑦
=
𝑋
𝛽
2.3
Solve with the normal equation
The previous chapter used gradient descent to approach the weights through
repeated updates. Linear algebra can solve for the least-squares weights directly
when the required inverse exists.
Definition 2.3.1.
The
normal equation
finds the weights that minimize RSS:
27
̂
𝛽
=
(
𝑋
𝑇
𝑋
)
−
1
𝑋
𝑇
𝑦
If
𝑋
𝑇
𝑋
is invertible, the equation gives one set of weights with the lowest
possible RSS.
Start with the scalar equation
𝑦
=
𝛽
𝑥
. If
𝑥
is nonzero, dividing both sides by
𝑥
gives
𝛽
=
𝑥
−
1
𝑦
. The normal equation plays a related role for a matrix
𝑋
, but a
rectangular matrix has no ordinary inverse.
Build the matrix expression
For multiple observations and features, write the system as
𝑦
=
𝑋
𝛽
. A dataset
usually has more observations than features, so
𝑋
is rectangular. We cannot invert
it directly.
1.
Form the Gram matrix
𝑋
𝑇
𝑋
.
Multiplying
𝑋
𝑇
by
𝑋
creates a symmetric square matrix. If its columns contain
independent information, this matrix is invertible.
Example 2.3.1.
Calculate the coefficients for this small dataset:
Size
Bedrooms
Price
1,000
2
100,000
2,000
4
200,000
3,000
3
300,000
4,000
5
400,000
Put the features in
𝑋
and the prices in
𝑦
. The first column of ones lets the
matrix multiplication include an intercept.
𝑋
=
(
1
1
1
1
1,000
2,000
3,000
4,000
2
4
3
5
)
,
𝑦
=
(
100,000
200,000
300,000
400,000
)
First, transpose
𝑋
:
𝑋
𝑇
=
(
1
1,000
2
1
2,000
4
1
3,000
3
1
4,000
5
)
Next, multiply
𝑋
𝑇
by
𝑋
:
𝑋
𝑇
𝑋
=
(
4
10,000
1
4
10,000
30,000,000
39,000
1
4
39,000
5
4
)
Also multiply
𝑋
𝑇
by the price vector:
𝑋
𝑇
𝑦
=
(
1,000,000
3,000,000,000
3,900,000
)
28
Finally, multiply
(
𝑋
𝑇
𝑋
)
−
1
by
𝑋
𝑇
𝑦
. The arithmetic is long, so we give the
result:
̂
𝛽
=
(
0
1
0
0
0
)
•
The intercept is
𝛽
0
=
0
dollars.
•
The size weight is
𝛽
1
=
1
0
0
dollars per square foot.
•
The bedroom weight is
𝛽
2
=
0
dollars. In this constructed dataset, bedroom
count adds no information after the model knows the size.
The final equation is
̂
𝑦
=
0
+
1
0
0
𝑥
1
+
0
𝑥
2
, which reduces to
̂
𝑦
=
1
0
0
𝑥
1
.
2.4
Use gradient descent with multiple variables
Matrix inversion becomes expensive for datasets with many features. Gradient
descent offers an iterative alternative.
Definition 2.4.1.
The prediction remains
̂
𝑦
=
𝑋
𝛽
. Define one batch loss over all
𝑛
houses:
𝐽
(
𝛽
)
=
(
1
2
𝑛
)
‖
𝑋
𝛽
−
𝑦
‖
2
The gradient gives the direction in which the loss increases fastest:
∇
𝛽
𝐽
=
(
1
𝑛
)
𝑋
𝑇
(
𝑋
𝛽
−
𝑦
)
Subtract a fraction of that gradient from every weight:
𝛽
≔
𝛽
−
𝛼
∗
∇
𝛽
𝐽
Example 2.4.1.
Apply batch gradient descent in five steps:
1.
Start with an initial
𝛽
, often a vector of zeros.
2.
Compute the predictions
𝑋
𝛽
.
3.
Measure the error
𝑋
𝛽
−
𝑦
.
4.
Use the gradient to update all the weights.
5.
Repeat until the loss stops changing much.
Each run needs many updates, but gradient descent works with large datasets
and does not require matrix inversion.
2.5
Scale the features
Feature scale changes how gradient descent updates each weight. In a housing
dataset, square footage might be about 3,000 while bedroom count might be about 3.
Both features matter, but their numerical scales differ by roughly a factor of 1,000.
Without scaling, one weight can receive much larger updates than another. Training
may then move slowly, oscillate across the minimum, or diverge.
29
Figure 3: A three-dimensional view of the loss. Unscaled features produce an
elongated valley and an oscillating path. Scaled features produce rounder contours
and a more direct path.
Definition 2.5.1.
Feature scaling
transforms feature columns to comparable numerical ranges.
A common method is
standardization
, also called z-score scaling:
𝑥
′
𝑖
,
𝑗
=
𝑥
𝑖
,
𝑗
−
𝜇
𝑗
𝜎
𝑗
with
𝜇
𝑗
=
(
1
𝑛
)
∑
𝑛
𝑖
=
1
𝑥
𝑖
,
𝑗
𝜎
𝑗
=
√
(
1
𝑛
)
∑
𝑛
𝑖
=
1
(
𝑥
𝑖
,
𝑗
−
𝜇
𝑗
)
2
Here,
𝜇
𝑗
is the mean of feature
𝑗
. The standard deviation
𝜎
𝑗
measures the
typical distance from that mean.
Examine a constructed failure case
Imagine a model with only two input features:
•
𝑥
1
= square footage (roughly 800 to 4500)
•
𝑥
2
= bedrooms (roughly 1 to 5)
Many implementations initialize
𝛽
=
0
. At that point, the prediction error is
−
𝑦
,
and each gradient component is proportional to:
𝜕
𝐽
𝜕
𝛽
𝑗
∝
−
(
1
𝑛
)
∑
𝑛
𝑖
=
1
𝑦
𝑖
𝑥
𝑖
,
𝑗
Feature magnitude therefore affects gradient magnitude. Values in the thousands
tend to produce larger updates than values between 1 and 5.
Example 2.5.1.
Compare the feature magnitudes:
•
A typical home might have 2,500 square feet and 3 bedrooms.
30
•
The raw magnitude ratio is
2
5
0
0
3
≈
8
3
3
, so the square-footage gradient can
be hundreds of times larger.
This difference does not show that bedroom count is unimportant. It shows
that the units affect the optimizer.
Figure 4: Raw and standardized feature values. After z-score scaling, both features
occupy comparable numerical ranges.
Read the figure from left to right:
•
Left panel: the model sees one axis with values in the thousands and another near
single digits.
•
Right panel: both axes are centered around 0 with similar spreads, so the
gradients are more balanced.
Figure 5: Step-0 gradient magnitudes (log scale). Raw features create an extreme
update imbalance; scaled features reduce that gap.
The gradient-magnitude plot shows the cause of the unstable path. When feature
scales differ, one weight receives much larger updates.
Compare three training runs
The next figure compares three runs on the same dataset:
31
Figure 6: Three gradient descent runs. Unscaled data trains slowly with a tiny
learning rate and diverges with a larger one. Scaled data trains faster with the larger
learning rate.
Example 2.5.2.
Read the curves as follows:
1.
Raw data with a tiny learning rate:
Training is stable but slow.
2.
Raw data with a larger learning rate:
The loss grows, so training
diverges.
3.
Scaled data with a larger learning rate:
The loss falls steadily.
Scaling often increases the range of learning rates that produce stable training.
Connect scaling to the loss geometry
Each point on the loss surface represents a set of parameter values. Unscaled
features often produce elongated contours. Scaled features make the contours more
circular.
Figure 7: Loss contours and parameter paths. Unscaled features create an elongated
valley and an oscillating path. Scaled features produce rounder contours and a more
direct path.
Compare the two panels:
•
Left: narrow contours cause the optimizer to cross the valley repeatedly.
•
Right: contours are more circular, so the path can head toward the minimum more
directly.
32
Apply scaling to other datasets
Use these rules for models trained with gradient-based methods:
•
If features use different units, scaling is usually necessary.
•
Standardization is a useful default for linear models and neural networks.
•
Fit
𝜇
and
𝜎
only on the training data. Reuse those values for the validation data
and the test data.
If you calculate new scaling values from the test data, information from the test set
affects preprocessing. This data leakage makes the evaluation unreliable.
2.6
Interpret weights after scaling
A model trained on standardized features learns weights in standardized units.
Those coefficients are not measured in dollars per square foot or dollars per
bedroom. Convert them back before comparing them with coefficients from a model
trained on raw features.
Definition 2.6.1.
If we standardize each feature with
𝑥
′
𝑗
=
𝑥
𝑗
−
𝜇
𝑗
𝜎
𝑗
and train a model with weights
𝛽
′
and intercept
𝛽
′
0
, then the equivalent weights in the original feature units
are:
𝛽
𝑗
=
𝛽
′
𝑗
𝜎
𝑗
𝛽
0
=
𝛽
′
0
−
∑
𝑗
(
𝛽
′
𝑗
∗
𝜇
𝑗
𝜎
𝑗
)
Example 2.6.1.
Suppose your trained scaled model has:
•
𝛽
′
0
=
4
2
0
,
0
0
0
•
𝛽
′
sqft
=
1
2
6
,
5
0
0
and
𝜎
sqft
=
1
1
0
0
•
𝛽
′
bed
=
1
9
,
0
0
0
and
𝜎
bed
=
0
.
9
•
𝜇
sqft
=
2
5
0
0
•
𝜇
bed
=
3
.
4
Convert back:
•
𝛽
sqft
=
1
2
6
,
5
0
0
1
1
0
0
≈
1
1
5
dollars per square foot
•
𝛽
bed
=
1
9
,
0
0
0
0
.
9
≈
2
1
,
1
1
1
dollars per additional bedroom
After this conversion, the coefficients again describe changes in the original
feature units.
2.7
Linear algebra practice problems
Sets and number sets
Problem 1.
Classify each of the following values into the most specific number set
(
ℕ
,
ℤ
,
ℚ
,
ℝ
, or
ℂ
):
•
7
•
−
3
•
0
.
7
5
•
√
2
33
•
3
+
2
𝑖
Problem 2.
Classify each value into the most specific number set (
ℕ
,
ℤ
,
ℚ
,
ℝ
, or
ℂ
):
•
0
•
−
7
3
•
√
9
•
5
+
0
𝑖
Problem 3.
True or False: Every natural number is also a rational number. Explain
your reasoning using the subset relationships.
Problem 4.
True or False: Every real number is also a rational number. If false, give
a counterexample.
Problem 5.
A machine learning dataset contains the following columns: “number
of bedrooms” (values like 2, 3, 4) and “house price” (values like $245,000.50). Which
number set would you use to describe each column?
Problem 6.
If
𝐴
⊂
𝐵
and
𝐵
⊂
𝐶
, what can you conclude about the relationship
between
𝐴
and
𝐶
? Apply this to explain why
ℕ
⊂
ℝ
.
Problem 7.
Give an example of a number that is in
ℝ
but not in
ℚ
. Why does this
distinction matter for computer representations of numbers?
Vectors and vector operations
Problem 8.
A data point for a student has the following features: GPA (3.5), SAT
score (1400), and number of extracurriculars (4). Write this as a 3-vector in column
notation.
Problem 9.
Given two vectors
𝒂
=
(
2
5
−
1
)
and
𝒃
=
(
3
−
2
4
)
, compute
𝒂
+
𝒃
.
Problem 10.
Compute
3
𝒗
where
𝒗
=
(
4
−
2
7
)
.
Problem 11.
Compute
−
2
𝒂
+
𝒃
where
𝒂
=
(
3
1
−
4
)
and
𝒃
=
(
−
1
5
2
)
.
Problem 12.
Given
𝒖
=
(
1
2
)
and
𝒘
=
(
4
6
)
, compute
2
𝒖
+
3
𝒘
.
Problem 13.
State whether each expression is defined. If it is, compute it.
1.
𝒑
+
𝒒
where
𝒑
=
(
1
2
3
)
and
𝒒
=
(
4
5
6
)
2.
𝒓
+
𝒔
where
𝒓
=
(
1
2
)
and
𝒔
=
(
3
4
5
)
Problem 14.
Why can’t you add the vectors
𝒑
=
(
1
2
3
)
and
𝒒
=
(
4
5
)
? In a machine
learning context, what would this situation represent?
Notation and dimensions
Problem 15.
If a matrix
𝑀
∈
ℝ
5
0
×
7
, how many rows does it have? How many
columns? How many total entries?
Problem 16.
If
𝐴
∈
ℝ
4
×
2
, how many entries are in
𝐴
?
Problem 17.
Write the general form of a matrix
𝐴
∈
ℝ
2
×
3
using subscript notation
for each element.
Problem 18.
Write the general form of a matrix
𝐵
∈
ℝ
3
×
2
using subscript notation.
Problem 19.
You have a dataset of 1000 images, where each image is represented
by 784 pixel values. What is the shape of the data matrix
𝑋
if each row is one
image? Write it in the form
𝑋
∈
ℝ
𝑝
×
𝑑
.
34
Problem 20.
Given
𝑋
∈
ℝ
3
×
4
, what element is located at row 2, column 3? Write it
using subscript notation.
Problem 21.
A machine learning model takes in data matrix
𝑋
∈
ℝ
𝑛
×
𝑑
and
outputs predictions
̂
𝑦
∈
ℝ
𝑛
. Explain in plain English what
𝑛
and
𝑑
represent.
Dot product
Problem 22.
Compute the dot product of
𝒂
=
(
2
3
1
)
and
𝒃
=
(
4
−
1
5
)
.
Problem 23.
Compute the dot product of
𝒖
=
(
−
2
0
3
)
and
𝒗
=
(
5
4
−
1
)
.
Problem 24.
If
𝒙
=
(
1
0
0
)
and
𝒚
=
(
0
1
0
)
, compute
𝒙
⋅
𝒚
.
Problem 25.
A house has features
𝒙
=
(
1
2
0
0
0
3
)
(intercept, square footage,
bedrooms) and the model weights are
𝜷
=
(
5
0
0
0
0
1
0
0
5
0
0
0
)
. Compute the predicted price
using the dot product.
Problem 26.
Compute
𝒗
⋅
𝒗
where
𝒗
=
(
3
4
)
.
Problem 27.
Given
𝒑
=
(
2
−
1
)
and
𝒒
=
(
4
3
)
, compute
𝒑
⋅
𝒒
.
Problem 28.
Why must two vectors have the same dimension for the dot product
to be defined? Give a practical example where this constraint matters.
Matrix multiplication
Problem 29.
Given
𝐴
=
(
1
3
2
4
)
and
𝐵
=
(
5
7
6
8
)
, compute the element in row 1,
column 2 of
𝐴
𝐵
.
Problem 30.
If
𝐴
∈
ℝ
3
×
4
and
𝐵
∈
ℝ
4
×
2
, what is the shape of the product
𝐴
𝐵
?
Problem 31.
Can you multiply
𝑃
∈
ℝ
2
×
3
by
𝑄
∈
ℝ
2
×
3
? Explain why or why not.
Problem 32.
Compute the full matrix product:
(
2
1
0
3
)
(
1
2
4
5
)
Problem 33.
Compute the full matrix product:
(
1
0
−
1
3
2
1
)
(
2
−
1
4
1
0
−
2
)
Problem 34.
Let
𝐴
∈
ℝ
2
×
3
and
𝐵
∈
ℝ
3
×
1
. What is the shape of
𝐴
𝐵
?
Transpose
Problem 35.
Compute the transpose of
𝐴
=
(
1
4
2
5
3
6
)
.
Problem 36.
If
𝑀
∈
ℝ
1
0
×
3
, what is the shape of
𝑀
𝑇
?
Problem 37.
Given the column vector
𝒗
=
(
2
5
8
)
, write
𝒗
𝑇
.
Problem 38.
Verify that
(
𝐴
𝑇
)
𝑇
=
𝐴
for
𝐴
=
(
1
3
5
2
4
6
)
.
Problem 39.
Let
𝐴
=
(
1
−
3
2
0
)
and
𝐵
=
(
4
2
−
1
5
)
. Compute
(
𝐴
+
𝐵
)
𝑇
.
Determinant
Problem 40.
Compute the determinant of
𝐴
=
(
5
2
3
4
)
.
35
Problem 41.
Compute the determinant of
𝐷
=
(
7
3
−
1
2
)
.
Problem 42.
Compute the determinant of
𝐵
=
(
−
2
3
6
−
9
)
. Does this matrix have an
inverse?
Problem 43.
For the matrix
𝐶
=
(
𝑎
𝑏
2
𝑎
2
𝑏
)
, compute the determinant. What does this
tell you about matrices where one column is a multiple of the other?
Problem 44.
If
det
(
𝐴
)
=
5
, what is
det
(
2
𝐴
)
for a
2
×
2
matrix? (Hint: work out an
example.)
Problem 45.
The determinant has a geometric interpretation: it tells us how a
matrix scales area. If
det
(
𝐴
)
=
3
, what happens to the area of a unit square when
transformed by
𝐴
? What if
det
(
𝐴
)
=
−
2
?
Identity matrix and inverse
Problem 46.
Write the
2
×
2
identity matrix
𝐼
2
and the
3
×
3
identity matrix
𝐼
3
.
Problem 47.
Compute
𝐼
2
𝐴
for
𝐴
=
(
−
1
2
4
0
)
.
Problem 48.
Verify that
𝐴
𝐼
2
=
𝐴
for
𝐴
=
(
3
2
7
5
)
.
Problem 49.
Compute the inverse of
𝐴
=
(
4
2
3
2
)
by hand using the formula
𝐴
−
1
=
1
det
(
𝐴
)
(
𝑑
−
𝑐
−
𝑏
𝑎
)
.
Problem 50.
Compute the inverse of
𝐵
=
(
5
7
2
3
)
by hand. Show all steps.
Problem 51.
Compute the inverse of
𝐶
=
(
1
3
2
5
)
by hand.
Problem 52.
Attempt to compute the inverse of
𝐷
=
(
6
4
3
2
)
. What happens and
why?
Problem 53.
If
𝐴
−
1
=
(
2
−
3
−
1
2
)
, find
𝐴
.
Problem 54.
For which values of
𝑘
is the matrix
𝐴
=
(
1
2
𝑘
4
)
invertible?
Mixed practice
Problem 55.
Let
𝒂
=
(
2
−
1
0
)
and
𝒃
=
(
1
3
−
2
)
. Compute
(
𝒂
+
𝒃
)
⋅
𝒃
.
Problem 56.
Let
𝐴
=
(
1
0
3
2
−
1
4
)
. Compute
𝐴
𝑇
and then compute
𝐴
𝑇
𝐴
.
Problem 57.
Let
𝐵
=
(
2
−
4
1
−
2
)
. Determine whether
𝐵
is invertible and explain
why.
2.8
Problem packet
The packet provides each calculus derivation. Apply or interpret the result instead
of deriving it again. Show the requested calculations, and explain each
interpretation in a complete sentence.
Data representation and the design matrix
Problem 1.
Given the dataset below, write the design matrix
𝑋
including an
intercept column and the label vector
𝑦
. Then interpret the meaning of each column
in one sentence.
Problem 2.
Size (ft squared)
Bedrooms
Price ($)
“1,000”
2
“100,000”
“2,000”
4
“200,000”
36
“3,000”
3
“300,000”
“4,000”
5
“400,000”
State the shape of
𝑋
in the form
𝑋
∈
ℝ
𝑝
×
𝑑
and explain what
𝑝
and
𝑑
represent in
this dataset.
Matrix operations reference
Problem 3.
Using the design matrix from Problem 1, compute
𝑋
𝑇
𝑋
and
𝑋
𝑇
𝑦
.
Then explain in one sentence what each result measures.
Normal equation with derivation
The loss is the residual sum of squares:
𝐿
(
𝛽
)
=
(
𝑦
−
𝑋
𝛽
)
𝑇
(
𝑦
−
𝑋
𝛽
)
Worked derivation:
𝐿
(
𝛽
)
=
𝑦
𝑇
𝑦
−
2
𝛽
𝑇
𝑋
𝑇
𝑦
+
𝛽
𝑇
𝑋
𝑇
𝑋
𝛽
Taking the gradient with respect to
𝛽
:
∇
𝛽
𝐿
(
𝛽
)
=
−
2
𝑋
𝑇
𝑦
+
2
𝑋
𝑇
𝑋
𝛽
Setting to zero:
−
2
𝑋
𝑇
𝑦
+
2
𝑋
𝑇
𝑋
𝛽
=
0
⟶
𝑋
𝑇
𝑋
𝛽
=
𝑋
𝑇
𝑦
The solution is:
̂
𝛽
=
(
𝑋
𝑇
𝑋
)
−
1
𝑋
𝑇
𝑦
Problem 4.
Using the formula above, compute
̂
𝛽
for Problem 1. Interpret each
coefficient in a single clear sentence.
Problem 5.
Explain why
̂
𝛽
is a minimizer of the loss using geometric or algebraic
intuition.
MSE and gradient with derivation
ℒ︀
(
𝛽
)
=
1
𝑁
(
𝑦
−
𝑋
𝛽
)
𝑇
(
𝑦
−
𝑋
𝛽
)
Worked derivation:
∇
𝛽
ℒ︀
(
𝛽
)
=
1
𝑁
(
−
2
𝑋
𝑇
𝑦
+
2
𝑋
𝑇
𝑋
𝛽
)
=
−
2
𝑁
𝑋
𝑇
(
𝑦
−
𝑋
𝛽
)
The gradient descent update with learning rate
𝜂
is:
𝛽
←
𝛽
−
𝜂
∇
𝛽
ℒ︀
(
𝛽
)
=
𝛽
+
2
𝜂
𝑁
𝑋
𝑇
(
𝑦
−
𝑋
𝛽
)
Problem 6.
Apply the formula. Given
𝑁
=
3
,
𝜂
=
0
.
1
:
𝑋
=
(
1
1
1
2
4
1
1
2
0
)
,
𝛽
=
(
5
0
5
2
)
,
𝑦
=
(
6
5
7
5
5
2
)
Compute
𝑦
−
𝑋
𝛽
, then
𝑋
𝑇
(
𝑦
−
𝑋
𝛽
)
, and the update increment
Δ
𝛽
.
Problem 7.
Explain in two sentences what the vector
𝑋
𝑇
(
𝑦
−
𝑋
𝛽
)
represents.
37
Apply RSS and MSE
Problem 8.
Using the results from Problem 6 (where
𝑒
=
(
3
,
1
,
−
3
)
𝑇
), explain if the
model under- or over-predicts for each observation and state what an MSE of 6.33
means.
Collinearity and remedies
Problem 9.
In two sentences, explain how near-collinearity affects the stability of
̂
𝛽
. Propose one practical remedy, and explain it in one sentence.
38
Appendix: Programming
Reference
3
Python and libraries
3.1
NumPy
Import NumPy
Import NumPy with its conventional alias,
np
:
import
numpy
as
np
Create arrays and vectors
In NumPy an array is a generic term for a multidimensional set of numbers. One-
dimensional NumPy arrays act like vectors. The following code creates two one-
dimensional arrays and adds them elementwise. If you attempted the same with
plain Python lists you would not get elementwise addition.
x
=
np
.
array
(
[
3
,
4
,
5
]
)
y
=
np
.
array
(
[
4
,
9
,
7
]
)
print
(
x
+
y
)
#
array([ 7, 13, 12])
Represent matrices with two-dimensional arrays
Matrices in NumPy are typically represented as two-dimensional arrays. The object
returned by np.array has attributes such as ndim for the number of dimensions,
dtype for the data type, and shape for the size of each axis.
x
=
np
.
array
(
[
[
1
,
2
]
,
[
3
,
4
]
]
)
print
(
x
)
#
array([[1, 2], [3, 4]])
print
(
x
.
ndim
)
#
2
print
(
x
.
dtype
)
#
e.g. dtype('int64')
print
(
x
.
shape
)
#
(2, 2)
If any element passed into np.array is a floating point number, NumPy upcasts the
whole array to a floating point dtype.
print
(
np
.
array
(
[
[
1
,
2
]
,
[
3
.
0
,
4
]
]
)
.
dtype
)
#
dtype('float64')
print
(
np
.
array
(
[
[
1
,
2
]
,
[
3
,
4
]
]
,
float
)
.
dtype
)
#
dtype('float64')
Call methods and functions
Methods are functions bound to objects. Calling x.sum() calls the sum method with
x as the implicit first argument. The module-level function np.sum(x) does the same
computation but is not bound to x.
x
=
np
.
array
(
[
1
,
2
,
3
,
4
]
)
print
(
x
.
sum
(
)
)
#
method on the array object
print
(
np
.
sum
(
x
)
)
#
module-level function
The reshape method returns a new view with the same data arranged into a new
shape. You pass a tuple that specifies the new dimensions.
x
=
np
.
array
(
[
1
,
2
,
3
,
4
,
5
,
6
]
)
print
(
"
beginning x:
\n
"
,
x
)
x_reshape
=
x
.
reshape
(
(
2
,
3
)
)
print
(
"
reshaped x:
\n
"
,
x_reshape
)
39
NumPy uses zero-based indexing. The first row and first column entry of x_reshape
is accessed with x_reshape[0, 0]. The entry in the second row and third column is
x_reshape[1, 2]. The third element of the original one-dimensional x is x[2].
print
(
x_reshape
[
0
,
0
]
)
#
1
print
(
x_reshape
[
1
,
2
]
)
#
6
print
(
x
[
2
]
)
#
third element of x
Understand views and shared memory
Reshaping often returns a view rather than a copy. Modifying a view will modify
the original array because they share the same memory. This behavior is important
when you expect independent copies.
print
(
"
x before modification:
\n
"
,
x
)
print
(
"
x_reshape before modification:
\n
"
,
x_reshape
)
x_reshape
[
0
,
0
]
=
5
print
(
"
x_reshape after modification:
\n
"
,
x_reshape
)
print
(
"
x after modification:
\n
"
,
x
)
If you need an independent copy, call x.copy() explicitly.
x
=
np
.
array
(
[
1
,
2
,
3
,
4
,
5
,
6
]
)
x_copy
=
x
.
copy
(
)
x_reshape_copy
=
x_copy
.
reshape
(
(
2
,
3
)
)
x_reshape_copy
[
0
,
0
]
=
99
print
(
"
x remains unchanged:
\n
"
,
x
)
print
(
"
x_reshape_copy changed:
\n
"
,
x_reshape_copy
)
Tuples are immutable sequences in Python and will raise a TypeError if you try to
modify an element. This differs from NumPy arrays and Python lists.
my_tuple
=
(
3
,
4
,
5
)
#
my_tuple[0] = 2 # would raise TypeError: 'tuple' object does not
support item assignment
Transpose, ndim, and shape
You can request several attributes at once. The transpose T flips axes and is useful
for matrix algebra.
print
(
x_reshape
.
shape
,
x_reshape
.
ndim
,
x_reshape
.
T
)
#
For example: ((2, 3), 2, array([[5, 4], [2, 5], [3, 6]]))
Apply elementwise operations
NumPy supports elementwise arithmetic and universal functions such as np.sqrt.
Raising an array to a power is elementwise.
print
(
np
.
sqrt
(
x
)
)
#
elementwise square root
print
(
x
*
*
2
)
#
elementwise square
print
(
x
*
*
0
.
5
)
#
alternative for square root
Generate random numbers
NumPy provides random number generation. The signature for rng.normal is
normal(loc=0.0, scale=1.0, size=None). The arguments loc and scale are keyword
arguments for mean and standard deviation and size controls the shape of the
output.
x
=
np
.
random
.
normal
(
size
=
50
)
print
(
x
)
#
random sample from N(0,1), different each run
To create a dependent array, add a random variable with a different mean to each
element.
y
=
x
+
np
.
random
.
normal
(
loc
=
50
,
scale
=
1
,
size
=
50
)
print
(
np
.
corrcoef
(
x
,
y
)
)
#
correlation matrix between x and y
40
Reproduce results with the Generator API
To produce identical random numbers across runs, use np.random.default_rng with
an integer seed to create a Generator object and then call its methods. The
Generator API is the recommended approach for reproducibility.
rng
=
np
.
random
.
default_rng
(
1303
)
print
(
rng
.
normal
(
scale
=
5
,
size
=
2
)
)
rng2
=
np
.
random
.
default_rng
(
1303
)
print
(
rng2
.
normal
(
scale
=
5
,
size
=
2
)
)
#
Both prints produce the same arrays because the same seed was used.
When you use rng.standard_normal or rng.normal you are using the Generator
instance, which ensures reproducibility if you control the seed.
rng
=
np
.
random
.
default_rng
(
3
)
y
=
rng
.
standard_normal
(
10
)
print
(
np
.
mean
(
y
)
,
y
.
mean
(
)
)
Calculate the mean, variance, and standard deviation
NumPy provides np.mean, np.var, and np.std as module-level functions. Arrays also
have methods mean, var, and std. By default np.var divides by n. If you need the
sample variance that divides by n minus 1, provide ddof=1.
rng
=
np
.
random
.
default_rng
(
3
)
y
=
rng
.
standard_normal
(
10
)
print
(
np
.
var
(
y
)
,
y
.
var
(
)
,
np
.
mean
(
(
y
-
y
.
mean
(
)
)
*
*
2
)
)
print
(
np
.
sqrt
(
np
.
var
(
y
)
)
,
np
.
std
(
y
)
)
#
Use np.var(y, ddof=1) for sample variance dividing by n-1.
Calculate along an axis
NumPy arrays are row-major ordered. The first axis, axis=0, refers to rows and the
second axis, axis=1, refers to columns. Passing axis into reduction methods lets you
compute means, sums, and other statistics along rows or columns.
rng
=
np
.
random
.
default_rng
(
3
)
X
=
rng
.
standard_normal
(
(
10
,
3
)
)
print
(
X
)
#
10 by 3 matrix
print
(
X
.
mean
(
axis
=
0
)
)
#
column means
print
(
X
.
mean
(
0
)
)
#
same as previous
When you compute X.mean(axis=1) you obtain a one-dimensional array of row
means. When you compute X.sum(axis=0) you obtain column sums.
Plot with Matplotlib
Matplotlib is the standard plotting library. A plot consists of a figure and one or
more axes. The subplots function returns a tuple containing the figure and the axes.
The axes object has a plot method and other methods to customize titles, labels, and
markers.
from
matplotlib
.
pyplot
import
subplots
fig
,
ax
=
subplots
(
figsize
=
(
8
,
8
)
)
rng
=
np
.
random
.
default_rng
(
3
)
x
=
rng
.
standard_normal
(
100
)
y
=
rng
.
standard_normal
(
100
)
ax
.
plot
(
x
,
y
)
#
default line plot
ax
.
plot
(
x
,
y
,
'
o
'
)
#
scatter-like circles
#
To save: fig.savefig("scatter.png")
#
To display in an interactive session: import matplotlib.pyplot as
plt; plt.show()
41
Keep results reproducible
Use
np.random.default_rng
with a fixed seed when an exercise requires repeatable
results. NumPy versions can still produce small differences in some random outputs.
When you calculate variance, set
ddof=1
for sample variance and use the default
ddof=0
for population variance.
42
Appendix: Math Fundamentals
4
Calculus
4.1
Limits
4.2
Derivatives
Limit definition of a derivative
4.3
Gradients
Vector-valued functions
Gradient definition
Partial derivatives and rules
5
Linear algebra
6
Statistics and probability
43
Reference
44
Glossary of Definitions
Definition 0.1
p. 7
Definition 0.2
p. 7
Definition 1.3.1
p. 10
Definition 1.5.1
p. 11
Definition 1.5.2
p. 12
Definition 1.7.1
p. 15
Definition 2.1.1
p. 19
Definition 2.1.2
p. 20
Definition 2.1.3
p. 20
Definition 2.1.4
p. 20
Definition 2.1.5
p. 21
Definition 2.1.6
p. 21
Definition 2.1.7
p. 22
Definition 2.1.8
p. 22
Definition 2.1.9
p. 23
Definition 2.1.10
p. 23
Definition 2.1.11
p. 24
Definition 2.1.12
p. 24
Definition 2.1.13
p. 25
Definition 2.1.14
p. 25
Definition 2.2.1
p. 26
Definition 2.3.1
p. 26
Definition 2.4.1
p. 28
Definition 2.5.1
p. 29
Definition 2.6.1
p. 32
45
←
Page 1 of ?
→
−
100%
+