Introduction To Artificial Intelligence and Machine LearningAniketh Chenjeri, Andrew Doyle, Swarnim Ghimire, Mr. Igor Tomcej2026-08-20ContentsForeword .................................................................................................. ⁠4How this book is structured .............................................................................. ⁠5Part I: Supervised Learning ............................................................................... ⁠71Linear regression ....................................................................................... ⁠81.1What is linear regression? ......................................................................... ⁠81.2The limits of a two-point estimate ................................................................ ⁠91.3Define linear regression .......................................................................... ⁠101.4Fit the line ......................................................................................... ⁠101.5Measure and interpret error ...................................................................... ⁠101.6Optimize with gradient descent .................................................................. ⁠121.7Understand gradients ............................................................................. ⁠131.8Problem packet for January 16, 2026 ............................................................. ⁠161.9Problem packet ................................................................................... ⁠17Theory questions ................................................................................ ⁠17Practice problems ............................................................................... ⁠172Pseudoinverse and multiple linear regression ....................................................... ⁠182.1Linear algebra primer ............................................................................ ⁠18Sets ............................................................................................... ⁠18Common number sets ........................................................................... ⁠18Relationships between sets ..................................................................... ⁠18Vectors, vector addition, and scalar multiplication ............................................ ⁠19Matrices, notation, and dimensions ............................................................ ⁠20Add and subtract matrices ...................................................................... ⁠21Dot product ...................................................................................... ⁠22Matrix multiplication ........................................................................... ⁠22Transpose ........................................................................................ ⁠23Identity matrix .................................................................................. ⁠23Determinant ..................................................................................... ⁠24Inverse matrix ................................................................................... ⁠252.2Define multiple linear regression ................................................................ ⁠262.3Solve with the normal equation .................................................................. ⁠26Build the matrix expression ..................................................................... ⁠272.4Use gradient descent with multiple variables ................................................... ⁠282.5Scale the features ................................................................................. ⁠28Examine a constructed failure case ............................................................. ⁠29Compare three training runs ................................................................... ⁠30Connect scaling to the loss geometry .......................................................... ⁠31Apply scaling to other datasets ................................................................. ⁠322.6Interpret weights after scaling ................................................................... ⁠322.7Linear algebra practice problems ................................................................ ⁠32Sets and number sets ............................................................................ ⁠32Vectors and vector operations .................................................................. ⁠33Notation and dimensions ....................................................................... ⁠33Dot product ...................................................................................... ⁠34Matrix multiplication ........................................................................... ⁠34Transpose ........................................................................................ ⁠34Determinant ..................................................................................... ⁠34Identity matrix and inverse ..................................................................... ⁠35Mixed practice ................................................................................... ⁠352.8Problem packet ................................................................................... ⁠35Data representation and the design matrix .................................................... ⁠35Matrix operations reference .................................................................... ⁠36Normal equation with derivation ............................................................... ⁠36MSE and gradient with derivation .............................................................. ⁠36Apply RSS and MSE ............................................................................. ⁠37Collinearity and remedies ....................................................................... ⁠37Appendix: Programming Reference ..................................................................... ⁠383Python and libraries .................................................................................. ⁠383.1NumPy ............................................................................................ ⁠38Import NumPy .................................................................................. ⁠38Create arrays and vectors ....................................................................... ⁠38Represent matrices with two-dimensional arrays ............................................. ⁠38Call methods and functions ..................................................................... ⁠38Understand views and shared memory ......................................................... ⁠39Transpose, ndim, and shape .................................................................... ⁠39Apply elementwise operations ................................................................. ⁠39Generate random numbers ..................................................................... ⁠39Reproduce results with the Generator API ..................................................... ⁠40Calculate the mean, variance, and standard deviation ......................................... ⁠40Calculate along an axis .......................................................................... ⁠40Plot with Matplotlib ............................................................................. ⁠40Keep results reproducible ....................................................................... ⁠41Appendix: Math Fundamentals ......................................................................... ⁠424Calculus ............................................................................................... ⁠424.1Limits ............................................................................................. ⁠424.2Derivatives ........................................................................................ ⁠42Limit definition of a derivative ................................................................. ⁠424.3Gradients .......................................................................................... ⁠42Vector-valued functions ......................................................................... ⁠42Gradient definition .............................................................................. ⁠42Partial derivatives and rules .................................................................... ⁠425Linear algebra ......................................................................................... ⁠426Statistics and probability .............................................................................. ⁠42Reference ................................................................................................. ⁠43Glossary of Definitions .................................................................................. ⁠444ForewordCherry Creek first offered Introduction to Artificial Intelligence and Machine Learning during the 2024-2025 school year. The course gave students a gentle introduction to the field. It avoided most of the mathematics and focused on libraries such as TensorFlow, NumPy, pandas, and scikit-learn.The 2025-2026 school year began without a stable course plan, so we rebuilt the syllabus. We kept the original goal of an approachable introduction to machine learning, but added the rigor that the subject requires. We still teach each mathematical idea through intuition before notation.That balance is hard because machine learning draws from programming, algebra, calculus, statistics, and probability. Leaving out the mathematics would also leave out the reasons that the algorithms work.This textbook contains the lessons, lecture notes, and exercises for the course. The course has no calculus prerequisite, but mathematics appears throughout artificial intelligence and machine learning. We could omit that mathematics and cover more topics, or include it and move more slowly. We chose depth.You do not need calculus to begin this book. You need basic algebra and some experience with Python. When a lesson uses unfamiliar mathematics, we first explain what each operation represents. We calculate difficult derivatives for you, then ask you to interpret and use the result. This practical approach gives you a base for further study, but it does not replace a full course in calculus, linear algebra, or statistics.The following people contributed to this book:•Primary Writers:‣Aniketh Chenjeri (CCHS ‘26)‣Andrew Doyle (CCHS ‘26)‣Mr. Igor Tomcej•Reviewers:‣Hariprasad Gridharan (CCHS ‘25, Cornell ‘29)‣Siddharth Menon (CCHS ‘26)‣Ani Gadepalli (CCHS ‘26)5How this book is structuredNote: This book is still in development. We release chapters after review, so the version you are reading is not final.Early editions may contain errors. If you find an unclear explanation or a technical mistake, report it to your teacher or teaching assistant.Your report will help us correct the book for future classes.We begin with supervised learning because its central task is concrete: use examples with known answers to predict answers for new examples. These lessons introduce the mathematics as we need it. Exercises ask you to interpret derivatives that we provide, apply linear algebra, and measure a model’s errors. The statistics and probability sections explain why those measurements work and what they cannot tell you.We then study unsupervised learning through k-means clustering. The final part introduces neural networks. Neural networks power systems such as image classifiers and large language models, including ChatGPT, Gemini, and Claude. Our goal is to make the basic mechanism behind those systems understandable.The appendices contain reference material for the libraries, programming concepts, and mathematics used in the lessons. Use them when a symbol, function, or operation is unfamiliar.The book assumes some familiarity with Python, programming, and Algebra 2. This is primarily a programming course, but you will also explain the theory and mathematics behind your code.We drew on the following books:1.Introduction to Statistical Learning by Trevor Hastie, Robert Tibshirani, Jerome Friedman2.The Elements of Statistical Learning by Trevor Hastie, Robert Tibshirani, Jerome Friedman3.Various books by Justin Skyzak4.Mathematics for Machine Learning by Marc Peter Deisenroth, A. Aldo Faisal, Cheng Soon Ong5.The Matrix Cookbook by Joseph Montanez6.The Deep Learning Book by Ian Goodfellow, Yoshua Bengio, and Aaron Courville6Table of contentsForeword ................................................................................ ⁠4How this book is structured ............................................................ ⁠5Part I: Supervised Learning ............................................................ ⁠71Linear regression ..................................................................... ⁠81.1What is linear regression? ....................................................... ⁠81.2The limits of a two-point estimate .............................................. ⁠91.3Define linear regression ........................................................ ⁠101.4Fit the line ...................................................................... ⁠101.5Measure and interpret error ................................................... ⁠101.6Optimize with gradient descent ............................................... ⁠121.7Understand gradients .......................................................... ⁠131.8Problem packet for January 16, 2026 .......................................... ⁠161.9Problem packet ................................................................. ⁠172Pseudoinverse and multiple linear regression ..................................... ⁠182.1Linear algebra primer .......................................................... ⁠182.2Define multiple linear regression .............................................. ⁠262.3Solve with the normal equation ............................................... ⁠262.4Use gradient descent with multiple variables ................................. ⁠282.5Scale the features ............................................................... ⁠282.6Interpret weights after scaling ................................................. ⁠322.7Linear algebra practice problems .............................................. ⁠322.8Problem packet ................................................................. ⁠35Appendix: Programming Reference .................................................. ⁠383Python and libraries ................................................................ ⁠383.1NumPy .......................................................................... ⁠38Appendix: Math Fundamentals ....................................................... ⁠424Calculus ............................................................................. ⁠424.1Limits ........................................................................... ⁠424.2Derivatives ...................................................................... ⁠424.3Gradients ....................................................................... ⁠425Linear algebra ....................................................................... ⁠426Statistics and probability ........................................................... ⁠42Reference .............................................................................. ⁠43Glossary of Definitions ................................................................ ⁠447Part I: Supervised LearningIn supervised learning, you receive inputs 𝑋 and their known responses 𝑌. A model estimates the relationship between them. We write that relationship as:𝑌=𝑓(𝑋)+𝜀Here, 𝑓 is the unknown relationship that we want to estimate. The term 𝜀 represents random error that 𝑋 does not explain. We assume that this error has an average value of zero.A training dataset pairs each input with its known response:𝒯︀={(𝑥1,𝑦1),(𝑥2,𝑦2),…,(𝑥𝑛,𝑦𝑛)}A learning algorithm uses those pairs to estimate 𝑓. We call the estimate ̂𝑓. The model then predicts a response for a new input with ̂𝑦=̂𝑓(𝑋).Definition 0.1.A feature is a measurable property used as an input 𝑥. To predict a house price, a model might use the house’s size, number of bedrooms, and location as features.Definition 0.2.A label is the known output 𝑦 paired with a training example. In the house example, the sale price is the label.The supervised learning lessons cover these methods and ideas:•Linear regression with the pseudoinverse gives us a direct way to estimate parameters. Multiple linear regression extends the method to several features.•Gradient descent and stochastic gradient descent improve parameters through repeated updates.•The bias-variance tradeoff helps us reason about underfitting, overfitting, and prediction error.•Polynomial regression lets a linear model represent curved relationships by adding transformed features.•Ridge and lasso regression use 𝐿2 and 𝐿1 regularization to limit model complexity. Lasso can also remove unhelpful features.•Logistic regression predicts the probability of a category, such as yes or no.•k-nearest neighbors, or k-NN, predicts from nearby training examples.These names are a map, not a prerequisite. Each lesson introduces its method from the beginning.81Linear regression1.1What is linear regression?Supervised learning starts with known inputs 𝑋 and known outputs 𝑌. We use those examples to fit an equation that predicts 𝑌 from 𝑋.One of the simplest approaches assumes a linear relationship between an input and an output. Real relationships are rarely perfectly linear, but a linear model gives you a useful baseline. A more complex model earns its complexity only if it predicts better than that baseline.Solving linear regression by hand requires linear algebra. This course has no linear algebra prerequisite, so we build the needed ideas as we use them. We will later use scikit-learn to fit models, but first we need to understand what a fitted line means and how a computer can find one. Example 1.1.1.Suppose you use two thermometers to record the same temperatures in Celsius and Fahrenheit. Your measurements look like this:Celsius (x)Fahrenheit (y)031.8541.91049.21560.12067.42578.93087.5Plot the measurements:The exact conversion equation is:𝑦=1.8𝑥+329Now pretend that you do not know the conversion equation. You can estimate the slope from two measurements with 𝑚=𝑦2−𝑦1𝑥2−𝑥1. The table also gives an estimated intercept of 31.8 at 𝑥=0. Together, these estimates produce the model:̂𝑦=1.85𝑥+31.8Note: The hat in ̂𝑦 marks a predicted value. It distinguishes a model’s prediction from a measured value 𝑦. In our model,̂𝑦=1.85𝑥+31.8̂𝑦 is the predicted Fahrenheit temperature for a Celsius input 𝑥.Statistics and machine learning use the same notation for estimated coefficients, such as ̂𝑤 and ̂𝑏.You have made a statistical model from data. The two-point method worked for this clean example, but real measurements expose its limits. 1.2The limits of a two-point estimateSuppose you measure temperatures around campus instead of using the clean table above.The AI class collected the following data in 2025:The points do not fall on one perfect line. A thermometer may be poorly calibrated. Direct sunlight or a gust of wind may affect one measurement. These sources of noise make any chosen pair of points unreliable. Linear regression uses all the observations instead of trusting one pair. It chooses the line with the smallest total error across the dataset. The line does not pass through every point because noisy data rarely allows that.101.3Define linear regressionDefinition 1.3.1.Simple linear regression models the relationship between one input 𝑥 and one output 𝑦 with ̂𝑦=𝑤𝑥+𝑏. The weight 𝑤 controls the slope, and the intercept 𝑏 controls where the line crosses the vertical axis. For the temperature conversion, the exact values are 𝑤=1.8 and 𝑏=32. The following synthetic dataset imitates a larger set of noisy measurements: 1.4Fit the lineA computer can estimate 𝑤 and 𝑏 in several ways. One method, gradient descent, starts with initial values and updates them to reduce error. Another method solves for the best values with linear algebra. Scikit-learn’s LinearRegression estimator uses a linear algebra solver rather than gradient descent.Both methods need a precise definition of error. We will define that measurement before studying gradient descent.1.5Measure and interpret error Example 1.5.1.The table pairs five observed values 𝑦 with predictions ̂𝑦:𝑥𝑦̂𝑦132.5255.2344.1476.811566.3The difference 𝑦𝑖−̂𝑦𝑖 is the residual for observation 𝑖. Residual sum of squares combines all residuals into one measurement.Definition 1.5.1.Residual sum of squares, or RSS, is:RSS=∑𝑁𝑖=0(𝑦𝑖−̂𝑦𝑖)2For each observation, subtract the prediction from the measured value. Square each residual, then add the squares. Squaring serves three purposes:1.Squaring makes all residuals positive, so large underpredictions and overpredictions both contribute to the total error.2.It gives large residuals more influence. A residual of 4 contributes 16, while a residual of 2 contributes 4.3.It produces a smooth loss function that we can optimize with derivatives. The following diagram shows each residual as the vertical distance between a point and the line:12RSS already gives one total, but its size grows with the number of observations. Mean squared error divides that total by 𝑁 so datasets of different sizes are easier to compare.Definition 1.5.2.Mean squared error, or MSE, is:𝐿(𝑤,𝑏)=1𝑁∑𝑁𝑖=0(𝑦𝑖−̂𝑦𝑖)2For the model ̂𝑦𝑖=𝑤𝑥𝑖+𝑏, we can expand the equation:𝐿(𝑤,𝑏)=1𝑁∑𝑁𝑖=0(𝑦𝑖−(𝑤𝑥𝑖+𝑏))2The factor 1𝑁 averages the squared residuals.The following Python function calculates MSE directly:# Assume X is a spreadsheet of our data and X[i] is the ith rowdef mse(W, b, X, y): n = len(X) total_error = 0 for i in range(n): prediction = W * X[i] + b total_error += (y[i] - prediction) ** 2 return total_error / n1.6Optimize with gradient descentMSE tells us how well a pair of values 𝑤 and 𝑏 fits the data. We still need a method that finds the values with the lowest MSE.Gradient descent improves an initial guess through repeated updates. Each update uses a derivative.A derivative measures how a function changes when its input changes by a small amount. On a graph, the derivative gives the slope at one point. A positive derivative means the function rises as the input increases. A negative derivative means the function falls.For linear regression, we ask how the loss changes when 𝑤 changes. The derivative of the loss with respect to 𝑤 answers that question.13If the derivative is positive, increasing 𝑤 increases the loss, so we decrease 𝑤. If the derivative is negative, increasing 𝑤 decreases the loss, so we increase 𝑤. In both cases, moving opposite the derivative lowers the loss locally.Gradient descent applies this update rule:𝑤←𝑤−𝜂𝜕𝐿𝜕𝑤The learning rate 𝜂 controls the size of each update. A very small learning rate takes many updates. A learning rate that is too large can jump past the minimum and cause the loss to increase.Update 𝑏 with the same pattern:𝑏←𝑏−𝜂𝜕𝐿𝜕𝑏Repeat both updates until the loss stops decreasing by a meaningful amount.You do not need to derive the formulas in this course. We will provide the partial derivatives:𝜕𝐿𝜕𝑤=−2𝑁∑𝑁𝑖=1(𝑦𝑖−(𝑤𝑥𝑖+𝑏))𝑥𝑖𝜕𝐿𝜕𝑏=−2𝑁∑𝑁𝑖=1(𝑦𝑖−(𝑤𝑥𝑖+𝑏))The corresponding NumPy updates are:dw = (-2.0 / N) * np.sum(error * x_scaled, dtype=np.float64)w -= lr * dwdb = (-2.0 / N) * np.sum(error, dtype=np.float64)b -= lr * db1.7Understand gradientsBefore we use a gradient with two inputs, consider the one-input function 𝑓(𝑥)=𝑥2:14Suppose 𝑓(𝑤)=𝑤2 represents the loss for the model ̂𝑦=𝑤𝑥. If 𝑤=2, the loss is 4. Training seeks the value of 𝑤 at the lowest point of the loss function.For this parabola, we can find the minimum from its vertex. A more complex loss function may not have such an obvious minimum:15Gradient descent searches for a low point without relying on a visual graph. For a function with one input, it uses the derivative.The derivative is the slope of the tangent line at a point. A tangent line matches the function’s direction at that one point. The purple line below is a tangent line:Figure 1: Example tangent line in purple for a loss functionAt 𝑥=0.4, the tangent line has slope 𝑚=−2.26528158057. The negative value tells us that the function falls as 𝑥 increases near this point.We write the derivative of 𝑓 as 𝑓′. For a loss function with one weight 𝑤, the update becomes:𝑤←𝑤−𝜂𝑓′(𝑤)A real model often has several parameters. Its loss function therefore has several inputs. A gradient collects the slope for each input so we can update every parameter.Definition 1.7.1.A gradient generalizes the derivative to a function with several inputs. For 𝑓(𝑥,𝑦)=𝑥2+𝑦2, the gradient is a vector that contains one partial derivative for each input:∇𝑓=(𝜕𝑓𝜕𝑥,𝜕𝑓𝜕𝑦)=(2𝑥,2𝑦)Each component of ∇𝑓 is a partial derivative. A partial derivative measures how the function changes with respect to one input while the other inputs stay fixed. For the linear regression loss, the relevant inputs are 𝑤 and 𝑏:𝜕𝐿𝜕𝑤=−2𝑁∑𝑁𝑖=1(𝑦𝑖−(𝑤𝑥𝑖+𝑏))𝑥𝑖𝜕𝐿𝜕𝑏=−2𝑁∑𝑁𝑖=1(𝑦𝑖−(𝑤𝑥𝑖+𝑏))Together, these two partial derivatives form the gradient of 𝐿(𝑤,𝑏).The following graph shows 𝐿(𝑤,𝑏) for the model used in the next lab:16Figure 2: The loss surface for our model: ̂𝑦=1.85𝑥+31.47The red line traces the parameter values that gradient descent tried on its way toward the minimum. Watch the gradient descent training video to see the updates in real time.1.8Problem packet for January 16, 2026Problem 1. You are modeling Fahrenheit from Celsius. You try three baselines:•Constant model: ̂𝑦=𝑐•Linear model: ̂𝑦=𝑊𝑥+𝑏•Piecewise linear with one breakpoint at 𝑥=10: ̂𝑦=𝑊1𝑥+𝑏1 for 𝑥≤10, and ̂𝑦=𝑊2𝑥+𝑏2 for 𝑥>10Without using code, rank the baselines from strongest to weakest. Define a strong baseline in terms of training and evaluation. Explain how it affects comparisons with later models, and give at least two problems caused by a weak baseline.Problem 2. In the text, ̂𝑦 is predicted output, and ̂𝑊, ̂𝑏 are estimated parameters. Explain the difference between an observed 𝑦, a prediction ̂𝑦, the true but unknown relationship 𝑓, estimated parameters ̂𝑊 and ̂𝑏, and measurement noise. Then answer: if you re-collect the dataset tomorrow with the same thermometers, which of these are expected to change and why?Problem 3. The chapter lists reasons to square residuals, explain why these reasons are important. Then argue for one scenario where squaring residuals is a bad idea, and name a better alternative loss. (You can use the internet to find the answer to this question.)Problem 4. You train the same model form ̂𝑦=𝑊𝑥+𝑏 on two datasets A and B.17•Dataset A has 𝑁=20 points, RSS = 120.•Dataset B has 𝑁=200 points, RSS = 950.1.Compute MSE for both.2.Explain why the RSS value can mislead you across dataset sizes.3.Give one situation where RSS is still useful or preferred (be specific).Problem 5. Suppose your fitted line gives small MSE, but when you plot residuals 𝑟𝑖=𝑦𝑖−̂𝑦𝑖 versus 𝑥𝑖, you see a clear U-shape. Explain what this implies about: the linearity assumption, whether the bias term ̂𝑏 is wrong, what kind of model change would address it, and why MSE alone did not warn you. Give at least two possible model changes.Problem 6. Given: 𝑥=[−5,0,5,10], 𝑦=[20.0,31.8,40.0,55.0], Model: ̂𝑦=1.6𝑥+31.8 Compute: ̂𝑦 for each 𝑥 residuals 𝑟𝑖=𝑦𝑖−̂𝑦𝑖 RSS and MSE Identify which point contributes most to RSS and explain whyProblem 7. Given: 𝑥=[−5,0,5,10] 𝑦=[20.0,31.8,40.0,55.0] Model: ̂𝑦=1.6𝑥+31.8 Compute: ̂𝑦 for each 𝑥 residuals 𝑟𝑖=𝑦𝑖−̂𝑦 RSS and MSE Identify which point contributes most to RSS and explain whyProblem 8. Dataset: 𝑥=[0,5,10,15,20] 𝑦=[32.0,41.0,50.5,60.0,68.0] Two candidate models: A: ̂𝑦=1.8𝑥+32𝐵:accent(y, hat) = 1.9 x + 31 Compute RSS for both and decide which is better under RSS/MSE. Then answer: which model is more plausible physically, and can plausibility disagree with MSE here?1.9Problem packetTheory questionsProblem 1. The text describes linear models as a “baseline.” Explain the importance of establishing a baseline model before moving on to more complex machine learning algorithms.Problem 2. In the equation ̂𝑦=1.85𝑥+31.8, explain what the hat notation ̂() means and why it distinguishes predictions from observations.Problem 3. The lesson provides three specific reasons for squaring residuals in the RSS formula. List them and explain why making the loss function “smooth and differentiable” is beneficial for optimization.Problem 4. What is the mathematical difference between Residual Sum of Squares (RSS) and Mean Squared Error (MSE)? Why is MSE generally preferred when working with datasets of varying sizes?Practice problemsProblem 5. You are given the coefficients 𝑎=1, 𝑏=4, and 𝑐=2 for the function 𝑓(𝑥)=𝑥2+4𝑥+2. Using the derivative 𝑓′(𝑥)=2𝑥+4, write a Python function to find the minimum of 𝑓(𝑥) using gradient descent. Start at 𝑥=10, use 𝜂=0.1, and run for 10 iterations.Problem 6. Calculate the RSS and MSE by hand for the following dataset given the model ̂𝑦=2𝑥+1:•𝑥=[1,2,3]•𝑦=[3,6,7]Problem 7. Given 𝑥=[1,2,3] and 𝑦=[2,3,4], and initial parameters 𝑊=0 and 𝑏=0, compute:•The predicted values ̂𝑦•The residuals (𝑦𝑖−̂𝑦𝑖)•The current MSE18Problem 8. Using the data and initial parameters from Problem 9, perform one full batch gradient descent update to find 𝑊new and 𝑏new. Use 𝜂=0.1 and the formulas:𝜕𝐿𝜕𝑊=−2𝑁∑𝑁𝑖=1(𝑦𝑖−(𝑊𝑥𝑖+𝑏))𝑥𝑖𝜕𝐿𝜕𝑏=−2𝑁∑𝑁𝑖=1(𝑦𝑖−(𝑊𝑥𝑖+𝑏))Note: Use the sign convention from the provided Python code where the gradient is subtracted.Problem 9. A thermometer model is trained to ̂𝑦=1.85𝑥+31.8. If the actual temperature is 0°𝐶 and the observed Fahrenheit reading is 31.8, what is the residual? If the actual temperature is 30°𝐶 and the observed reading is 87.5, what is the residual?Problem 10. Write a Python function get_error(y_true, y_pred) that returns the Mean Squared Error using only the standard library (no numpy). Assume both inputs are lists of equal length.Complete the remaining problems in the introduction assignments notebook.2Pseudoinverse and multiple linear regressionThe previous model used one feature 𝑋 to predict 𝑌. A house-price model may need size, bedroom count, and location at the same time.Linear algebra gives us notation and operations for working with these features together.2.1Linear algebra primerNote: This section introduces several symbols and operations. For another visual explanation, watch 3Blue1Brown’s Essence of Linear Algebra series. Some videos cover material beyond this course.SetsIn mathematics, a set is a collection of distinct objects. A programming set follows the same two basic rules: each value appears once, and its position does not matter.Common number setsNumber sets tell us which values a variable may contain:•Natural numbers ℕ are the positive counting numbers {1,2,3,…}. We use them for counts that cannot be zero or negative.•Integers ℤ are whole numbers, including zero and negative numbers: {…,−2,−1,0,1,2,…}.•Rational numbers ℚ can be written as a fraction 𝑝𝑞, where 𝑝,𝑞∈ℤ and 𝑞≠0. The set includes terminating decimals such as 0.75.•Real numbers ℝ are all the points on the number line. This set includes rational numbers and irrational numbers such as 𝜋 and √2.•Complex numbers ℂ contain a real part and an imaginary part. They appear in fields such as signal processing but are not used in this lesson.Relationships between setsThese number sets are nested. Every natural number is an integer, every integer is rational, and every rational number is real. The symbol ⊂ means “is a subset of”:19ℕ⊂ℤ⊂ℚ⊂ℝVectors, vector addition, and scalar multiplicationSuppose you want to predict a house price from these features:•living area in square feet,•number of bedrooms, and•age in years.Write the three feature names in a fixed order:(square footage,number of bedrooms,age).The values (2500,4,10) describe a 2,500-square-foot house with 4 bedrooms that is 10 years old. A new house with the same size and bedroom count has values (2500,4,0).This ordered list of values is a vector. Because it contains three values, it is a 3-vector.We can write the same vector vertically:(2500410)A scalar is one number. For example, 200 is a scalar, while (200,300,25) is a vector.Algebra often uses symbols such as 𝑥 and 𝑦 for scalars. We use bold symbols, such as 𝒙 and 𝒚, for vectors.Definition 2.1.1.For a positive integer 𝑛, an n-vector is an ordered list of 𝑛 real numbers. The symbol ℝ𝑛 denotes the set of all real n-vectors.You can also picture a vector as an arrow. For 𝑛=2, the vector 𝒗=(𝑣1𝑣2) points from the origin (0,0) to the point (𝑣1,𝑣2). A 3-vector works the same way in three-dimensional space.In physics, vectors can represent quantities with magnitude and direction, such as velocity and force. In machine learning, vectors usually store related numerical values. A dataset may contain vectors with thousands of entries.Example 2.1.1.Suppose there are 100 students in the AI class. We can keep track of all their grades on the first test by using a 100-vector𝑬=(𝐸1𝐸2⋮𝐸100)Here 𝐸1 is the first exam grade of the first student, 𝐸2 the first exam grade of the second student, and so on.We will use two basic vector operations: vector addition and scalar multiplication.20Definition 2.1.2.The sum 𝒗+𝒘 of two vectors is defined only when 𝒗 and 𝒘 are 𝑛-vectors. In that case, we define their sum by the rule (𝑣1𝑣2⋮𝑣𝑛)+(𝑤1𝑤2⋮𝑤𝑛)=(𝑣1+𝑤1𝑣2+𝑤2⋮𝑣𝑛+𝑤𝑛).Definition 2.1.3.To multiply an n-vector by a scalar 𝑐, multiply every component by 𝑐: 𝑐(𝑣1𝑣2⋮𝑣𝑛)=(𝑐𝑣1𝑐𝑣2⋮𝑐𝑣𝑛).Example 2.1.2.Practice problem: Let 𝒗=(2−13) and 𝒘=(54−2). Compute 𝒗+𝒘 and −2𝒗.𝒗+𝒘=(2+5−1+43+(−2))=(731).−2𝒗=(−42−6).Matrices, notation, and dimensionsYou may have seen matrices in Algebra 2. We begin with the definition:Definition 2.1.4.A matrix is a rectangular array of numbers or other mathematical objects with elements or entries arranged in rows and columns. A matrix with 𝑝 rows and 𝑑 columns is called a 𝑝×𝑑 matrix.The symbol ∈ means “is an element of.” In linear algebra, this notation can describe both a matrix’s dimensions and the type of its entries.The statement 𝑋∈ℝ𝑝×𝑑 gives two facts about 𝑋:1.Every entry is a real number.2.The matrix has 𝑝 rows and 𝑑 columns.A general 𝑝×𝑑 matrix looks like this:𝑋=(𝑥11𝑥21⋮𝑥𝑝1𝑥12𝑥22⋮𝑥𝑝2……⋱…𝑥1𝑑𝑥2𝑑⋮𝑥𝑝𝑑)Subscripts identify an entry’s row and column. In 𝑥11, the first 1 identifies the first row and the second 1 identifies the first column. The entry 𝑥𝑝𝑑 lies in row 𝑝 and column 𝑑.21If a dataset contains 100 houses and 5 features per house, then 𝑋∈ℝ100×5. The matrix contains 500 real-valued entries arranged in 100 rows and 5 columns.A feature is one measurement that describes an observation. For a house, examples include living area and bedroom count. Each feature occupies one column of 𝑋.A label is the target value for one observation. House price is a possible label. The label vector has one entry for each row of 𝑋.Example 2.1.3.Practice problem: Suppose 𝐴∈ℝ2×3 and 𝐵∈ℝ3×4. What is the shape of 𝐴𝐵?Since the inner dimensions match (3), the product is defined and 𝐴𝐵∈ℝ2×4.Add and subtract matricesDefinition 2.1.5.The sum of 2 matrices 𝐴=(𝑎11𝑎21⋮𝑎𝑝1𝑎12𝑎22⋮𝑎𝑝2⋮⋮⋱⋮𝑎1𝑑𝑎2𝑑⋮𝑎𝑝𝑑) and 𝐵=(𝑏11𝑏21⋮𝑏𝑝1𝑏12𝑏22⋮𝑏𝑝2⋮⋮⋱⋮𝑏1𝑑𝑏2𝑑⋮𝑏𝑝𝑑) is defined only when 𝐴 and 𝐵 are of the same size. In that case, we define their sum by the rule 𝐴+𝐵=(𝑎11+𝑏11𝑎21+𝑏21⋮𝑎𝑝1+𝑏𝑝1𝑎12+𝑏12𝑎22+𝑏22⋮𝑎𝑝2+𝑏𝑝2⋮⋮⋱⋮𝑎1𝑑+𝑏1𝑑𝑎2𝑑+𝑏2𝑑⋮𝑎𝑝𝑑+𝑏𝑝𝑑).Example 2.1.4.Let 𝐴=(1324) and 𝐵=(5768). Then 𝐴+𝐵=(1+53+72+64+8)=(610812).Definition 2.1.6.The difference of 2 matrices 𝐴=(𝑎11𝑎21⋮𝑎𝑝1𝑎12𝑎22⋮𝑎𝑝2⋮⋮⋱⋮𝑎1𝑑𝑎2𝑑⋮𝑎𝑝𝑑) and 𝐵=(𝑏11𝑏21⋮𝑏𝑝1𝑏12𝑏22⋮𝑏𝑝2⋮⋮⋱⋮𝑏1𝑑𝑏2𝑑⋮𝑏𝑝𝑑) is also defined only when 𝐴 and 𝐵 are of the same size (in other words 𝐴∈ℝ𝑝×𝑑 and 𝐵∈ℝ𝑝×𝑑). In that case, we define their difference by using sums but multiplying the second matrix by the scalar −1:First, multiply the second matrix by −1:−𝐵=(−𝑏11−𝑏21⋮−𝑏𝑝1−𝑏12−𝑏22⋮−𝑏𝑝2⋮⋮⋱⋮−𝑏1𝑑−𝑏2𝑑⋮−𝑏𝑝𝑑)Then, add the matrices:22𝐴−𝐵=𝐴+(−𝐵)=(𝑎11−𝑏11𝑎21−𝑏21⋮𝑎𝑝1−𝑏𝑝1𝑎12−𝑏12𝑎22−𝑏22⋮𝑎𝑝2−𝑏𝑝2⋮⋮⋱⋮𝑎1𝑑−𝑏1𝑑𝑎2𝑑−𝑏2𝑑⋮𝑎𝑝𝑑−𝑏𝑝𝑑)Dot productDefinition 2.1.7.The dot product of two vectors 𝒂=(𝑎1𝑎2…𝑎𝑛) and 𝒃=(𝑏1𝑏2…𝑏𝑛) is defined as:𝒂⋅𝒃=∑𝑛𝑖=1𝑎𝑖𝑏𝑖=𝑎1𝑏1+𝑎2𝑏2+…+𝑎𝑛𝑏𝑛Example 2.1.5.Let 𝒂=(123) and 𝒃=(456). Pair the corresponding components, multiply each pair, and add the products:𝒂⋅𝒃=(1×4)+(2×5)+(3×6)=4+10+18=32Example 2.1.6.Practice problem: Let 𝒖=(2−14) and 𝒗=(30−2). Compute 𝒖⋅𝒗.𝒖⋅𝒗=(2)(3)+(−1)(0)+(4)(−2)=6+0−8=−2.Matrix multiplication uses this same pair, multiply, and add pattern.Matrix multiplicationDefinition 2.1.8.Suppose that we have 𝐴∈ℝ𝑟×𝑑 and 𝐵∈ℝ𝑑×𝑠. Then the product of 𝐴 and 𝐵 is denoted 𝐴𝐵. The (𝑖,𝑗)th element of (𝐴𝐵) is computed by multiplying each element of the 𝑖th row of 𝐴 by the corresponding element of the 𝑗th column of 𝐵. That is, (𝐴𝐵)𝑖𝑗=∑𝑑𝑘=1𝑎𝑖𝑘𝑏𝑘𝑗.Example 2.1.7.Consider these two matrices:𝑨=(1324)and𝑩=(5768).Then23𝑨𝑩=(1324)(5768)=(1×5+2×73×5+4×71×6+2×83×6+4×8)=(19432250).Example 2.1.8.Practice problem: Let 𝐴=(23−1401) and 𝐵=(1−2520−1). Compute 𝐴𝐵.First check dimensions: 𝐴∈ℝ2×3 and 𝐵∈ℝ3×2, so 𝐴𝐵∈ℝ2×2.𝐴𝐵=(2×1+(−1)×(−2)+0×53×1+4×(−2)+1×52×2+(−1)×0+0×(−1)3×2+4×0+1×(−1))=(4045).The product has 𝑟 rows and 𝑠 columns. We can compute 𝐴𝐵 only when the number of columns in 𝐴 equals the number of rows in 𝐵.TransposeDefinition 2.1.9.The transpose of a matrix swaps its rows and columns. Row 1 becomes column 1, row 2 becomes column 2, and so on. A superscript 𝑇 denotes this operation.If 𝑋∈ℝ𝑝×𝑑, then 𝑋𝑇∈ℝ𝑑×𝑝. Entry (𝑖,𝑗) of the transpose comes from entry (𝑗,𝑖) of the original matrix: (𝑋𝑇)𝑖𝑗=𝑋𝑗𝑖.Example 2.1.9.For example, if we have a matrix 𝑋∈ℝ3×2 we can take its transpose 𝑋𝑇∈ℝ2×3 by swapping the rows and columns:𝑋=(𝑥11𝑥21𝑥31𝑥12𝑥22𝑥32)𝑋𝑇=(𝑥11𝑥12𝑥21𝑥22𝑥31𝑥32)Example 2.1.10.Practice problem: If 𝐶=(0235−14), compute 𝐶𝑇.𝐶𝑇=(03−1254).Identity matrixMultiplying a scalar by 1 leaves the scalar unchanged. The identity matrix has the same role in matrix multiplication.Definition 2.1.10.24The identity matrix 𝐼𝑛 (or just 𝐼 when the size is clear) is a square 𝑛×𝑛 matrix with 1s on the diagonal and 0s everywhere else:𝐼3=(100010001)For any matrix 𝐴 of compatible size: 𝐴𝐼=𝐼𝐴=𝐴Example 2.1.11.Verify the property with a 2×2 matrix:(2435)(1001)=(2×1+3×04×1+5×02×0+3×14×0+5×1)=(2435)The product equals the original matrix.DeterminantThe determinant is a scalar calculated from a square matrix. A determinant of zero tells us that the matrix has no inverse.Definition 2.1.11.For a 2×2 matrix 𝐴=(𝑎𝑐𝑏𝑑), the determinant is defined as:det(𝐴)=𝑎𝑑−𝑏𝑐The determinant is often written as |𝐴| or det(𝐴).Example 2.1.12.For 𝐴=(3124):det(𝐴)=(3)(4)−(2)(1)=12−2=10The absolute value of the determinant tells us how matrix multiplication scales area. If det(𝐴)=2, multiplication by 𝐴 doubles area. If det(𝐴)=0, a two-dimensional region collapses to a line or a point. That loss of information prevents an inverse.Definition 2.1.12.A matrix 𝐴 is invertible (has an inverse) if and only if det(𝐴)≠0.Example 2.1.13.Check whether 𝐵=(1224) has an inverse:det(𝐵)=(1)(4)−(2)(2)=4−4=0Because det(𝐵)=0, the matrix has no inverse. The second row is twice the first, so the rows do not contain independent information.25Inverse matrixFor a nonzero scalar, multiplication by its reciprocal produces 1. An inverse matrix follows the same pattern and produces the identity matrix.Definition 2.1.13.For a square matrix 𝐴, its inverse 𝐴−1 is the matrix such that:𝐴𝐴−1=𝐴−1𝐴=𝐼Not every matrix has an inverse. A matrix that has an inverse is called invertible or non-singular.Compute the inverse of a 2×2 matrixA 2×2 matrix has a direct formula:Definition 2.1.14.If 𝐴=(𝑎𝑐𝑏𝑑) and det(𝐴)≠0, then:𝐴−1=1det(𝐴)(𝑑−𝑐−𝑏𝑎)Swap the diagonal entries, negate the off-diagonal entries, and divide every entry by the determinant.Example 2.1.14.Compute the inverse of 𝐴=(3124).Step 1: Compute the determinant.det(𝐴)=(3)(4)−(2)(1)=12−2=10Since det(𝐴)=10≠0, the inverse exists.Step 2: Apply the formula.𝐴−1=110(4−1−23)=(410−110−210310)=(0.4−0.1−0.20.3)Step 3: Verify by computing 𝐴𝐴−1.𝐴𝐴−1=(3124)(0.4−0.1−0.20.3)=(3(0.4)+2(−0.1)1(0.4)+4(−0.1)3(−0.2)+2(0.3)1(−0.2)+4(0.3))=(1.2−0.20.4−0.4−0.6+0.6−0.2+1.2)=(1001)=𝐼✓Example 2.1.15.Compute the inverse of 𝐴=(1324).Step 1: det(𝐴)=(1)(4)−(2)(3)=4−6=−2Step 2: Apply the formula:26𝐴−1=1−2(4−3−21)=(−2321−12)Work with larger matricesFor larger matrices, Gaussian elimination and cofactor expansion require many steps.Cofactor expansion produces 6 terms for a 3×3 determinant and 24 terms for a 4×4 determinant. The number grows as 𝑛!, so a 5×5 determinant has 120 terms before simplification.In practice, numerical libraries calculate matrix inverses with more efficient algorithms:import numpy as npA = np.array([[3, 2], [1, 4]])A_inv = np.linalg.inv(A)print(A_inv) # [[ 0.4 -0.2] # [-0.1 0.3]]In this course, you will calculate inverses by hand only for 2×2 matrices. For larger matrices, use NumPy and interpret the result.2.2Define multiple linear regressionThe previous model predicted a house price from one feature. A model that uses size, bedroom count, and location needs a separate weight for each feature.Multiple linear regression estimates all these weights together.Definition 2.2.1.Give each of the 𝑝 features its own weight. The prediction is:̂𝑦=𝛽0+𝑋1𝛽1+𝑋2𝛽2+…+𝑋𝑝𝛽𝑝In this equation:•𝛽0 is the intercept. It replaces 𝑏 from the previous chapter.•𝑋1,𝑋2,…,𝑋𝑝 are the features, such as size and bedroom count.•𝛽1,𝛽2,…,𝛽𝑝 are the feature weights.Library documentation often writes the same calculation as one matrix multiplication:̂𝑦=𝑋𝛽2.3Solve with the normal equationThe previous chapter used gradient descent to approach the weights through repeated updates. Linear algebra can solve for the least-squares weights directly when the required inverse exists.Definition 2.3.1.The normal equation finds the weights that minimize RSS:27̂𝛽=(𝑋𝑇𝑋)−1𝑋𝑇𝑦If 𝑋𝑇𝑋 is invertible, the equation gives one set of weights with the lowest possible RSS.Start with the scalar equation 𝑦=𝛽𝑥. If 𝑥 is nonzero, dividing both sides by 𝑥 gives 𝛽=𝑥−1𝑦. The normal equation plays a related role for a matrix 𝑋, but a rectangular matrix has no ordinary inverse.Build the matrix expressionFor multiple observations and features, write the system as 𝑦=𝑋𝛽. A dataset usually has more observations than features, so 𝑋 is rectangular. We cannot invert it directly.1.Form the Gram matrix 𝑋𝑇𝑋.Multiplying 𝑋𝑇 by 𝑋 creates a symmetric square matrix. If its columns contain independent information, this matrix is invertible.Example 2.3.1.Calculate the coefficients for this small dataset:SizeBedroomsPrice1,0002100,0002,0004200,0003,0003300,0004,0005400,000Put the features in 𝑋 and the prices in 𝑦. The first column of ones lets the matrix multiplication include an intercept.𝑋=(11111,0002,0003,0004,0002435),𝑦=(100,000200,000300,000400,000)First, transpose 𝑋:𝑋𝑇=(11,000212,000413,000314,0005)Next, multiply 𝑋𝑇 by 𝑋:𝑋𝑇𝑋=(410,0001410,00030,000,00039,0001439,00054)Also multiply 𝑋𝑇 by the price vector:𝑋𝑇𝑦=(1,000,0003,000,000,0003,900,000)28Finally, multiply (𝑋𝑇𝑋)−1 by 𝑋𝑇𝑦. The arithmetic is long, so we give the result:̂𝛽=(01000)•The intercept is 𝛽0=0 dollars.•The size weight is 𝛽1=100 dollars per square foot.•The bedroom weight is 𝛽2=0 dollars. In this constructed dataset, bedroom count adds no information after the model knows the size.The final equation is ̂𝑦=0+100𝑥1+0𝑥2, which reduces to ̂𝑦=100𝑥1.2.4Use gradient descent with multiple variablesMatrix inversion becomes expensive for datasets with many features. Gradient descent offers an iterative alternative.Definition 2.4.1.The prediction remains ̂𝑦=𝑋𝛽. Define one batch loss over all 𝑛 houses:𝐽(𝛽)=(12𝑛)‖𝑋𝛽−𝑦‖2The gradient gives the direction in which the loss increases fastest:∇𝛽𝐽=(1𝑛)𝑋𝑇(𝑋𝛽−𝑦)Subtract a fraction of that gradient from every weight:𝛽≔𝛽−𝛼∗∇𝛽𝐽Example 2.4.1.Apply batch gradient descent in five steps:1.Start with an initial 𝛽, often a vector of zeros.2.Compute the predictions 𝑋𝛽.3.Measure the error 𝑋𝛽−𝑦.4.Use the gradient to update all the weights.5.Repeat until the loss stops changing much.Each run needs many updates, but gradient descent works with large datasets and does not require matrix inversion.2.5Scale the featuresFeature scale changes how gradient descent updates each weight. In a housing dataset, square footage might be about 3,000 while bedroom count might be about 3. Both features matter, but their numerical scales differ by roughly a factor of 1,000.Without scaling, one weight can receive much larger updates than another. Training may then move slowly, oscillate across the minimum, or diverge.29Figure 3: A three-dimensional view of the loss. Unscaled features produce an elongated valley and an oscillating path. Scaled features produce rounder contours and a more direct path.Definition 2.5.1.Feature scaling transforms feature columns to comparable numerical ranges. A common method is standardization, also called z-score scaling:𝑥′𝑖,𝑗=𝑥𝑖,𝑗−𝜇𝑗𝜎𝑗with𝜇𝑗=(1𝑛)∑𝑛𝑖=1𝑥𝑖,𝑗𝜎𝑗=√(1𝑛)∑𝑛𝑖=1(𝑥𝑖,𝑗−𝜇𝑗)2Here, 𝜇𝑗 is the mean of feature 𝑗. The standard deviation 𝜎𝑗 measures the typical distance from that mean.Examine a constructed failure caseImagine a model with only two input features:•𝑥1 = square footage (roughly 800 to 4500)•𝑥2 = bedrooms (roughly 1 to 5)Many implementations initialize 𝛽=0. At that point, the prediction error is −𝑦, and each gradient component is proportional to:𝜕𝐽𝜕𝛽𝑗∝−(1𝑛)∑𝑛𝑖=1𝑦𝑖𝑥𝑖,𝑗Feature magnitude therefore affects gradient magnitude. Values in the thousands tend to produce larger updates than values between 1 and 5.Example 2.5.1.Compare the feature magnitudes:•A typical home might have 2,500 square feet and 3 bedrooms.30•The raw magnitude ratio is 25003≈833, so the square-footage gradient can be hundreds of times larger.This difference does not show that bedroom count is unimportant. It shows that the units affect the optimizer.Figure 4: Raw and standardized feature values. After z-score scaling, both features occupy comparable numerical ranges.Read the figure from left to right:•Left panel: the model sees one axis with values in the thousands and another near single digits.•Right panel: both axes are centered around 0 with similar spreads, so the gradients are more balanced.Figure 5: Step-0 gradient magnitudes (log scale). Raw features create an extreme update imbalance; scaled features reduce that gap.The gradient-magnitude plot shows the cause of the unstable path. When feature scales differ, one weight receives much larger updates.Compare three training runsThe next figure compares three runs on the same dataset:31Figure 6: Three gradient descent runs. Unscaled data trains slowly with a tiny learning rate and diverges with a larger one. Scaled data trains faster with the larger learning rate.Example 2.5.2.Read the curves as follows:1.Raw data with a tiny learning rate: Training is stable but slow.2.Raw data with a larger learning rate: The loss grows, so training diverges.3.Scaled data with a larger learning rate: The loss falls steadily.Scaling often increases the range of learning rates that produce stable training.Connect scaling to the loss geometryEach point on the loss surface represents a set of parameter values. Unscaled features often produce elongated contours. Scaled features make the contours more circular.Figure 7: Loss contours and parameter paths. Unscaled features create an elongated valley and an oscillating path. Scaled features produce rounder contours and a more direct path.Compare the two panels:•Left: narrow contours cause the optimizer to cross the valley repeatedly.•Right: contours are more circular, so the path can head toward the minimum more directly.32Apply scaling to other datasetsUse these rules for models trained with gradient-based methods:•If features use different units, scaling is usually necessary.•Standardization is a useful default for linear models and neural networks.•Fit 𝜇 and 𝜎 only on the training data. Reuse those values for the validation data and the test data.If you calculate new scaling values from the test data, information from the test set affects preprocessing. This data leakage makes the evaluation unreliable.2.6Interpret weights after scalingA model trained on standardized features learns weights in standardized units. Those coefficients are not measured in dollars per square foot or dollars per bedroom. Convert them back before comparing them with coefficients from a model trained on raw features.Definition 2.6.1.If we standardize each feature with 𝑥′𝑗=𝑥𝑗−𝜇𝑗𝜎𝑗 and train a model with weights 𝛽′ and intercept 𝛽′0, then the equivalent weights in the original feature units are:𝛽𝑗=𝛽′𝑗𝜎𝑗𝛽0=𝛽′0−∑𝑗(𝛽′𝑗∗𝜇𝑗𝜎𝑗)Example 2.6.1.Suppose your trained scaled model has:•𝛽′0=420,000•𝛽′sqft=126,500 and 𝜎sqft=1100•𝛽′bed=19,000 and 𝜎bed=0.9•𝜇sqft=2500•𝜇bed=3.4Convert back:•𝛽sqft=126,5001100≈115 dollars per square foot•𝛽bed=19,0000.9≈21,111 dollars per additional bedroomAfter this conversion, the coefficients again describe changes in the original feature units.2.7Linear algebra practice problemsSets and number setsProblem 1. Classify each of the following values into the most specific number set (ℕ, ℤ, ℚ, ℝ, or ℂ):•7•−3•0.75•√233•3+2𝑖Problem 2. Classify each value into the most specific number set (ℕ, ℤ, ℚ, ℝ, or ℂ):•0•−73•√9•5+0𝑖Problem 3. True or False: Every natural number is also a rational number. Explain your reasoning using the subset relationships.Problem 4. True or False: Every real number is also a rational number. If false, give a counterexample.Problem 5. A machine learning dataset contains the following columns: “number of bedrooms” (values like 2, 3, 4) and “house price” (values like $245,000.50). Which number set would you use to describe each column?Problem 6. If 𝐴⊂𝐵 and 𝐵⊂𝐶, what can you conclude about the relationship between 𝐴 and 𝐶? Apply this to explain why ℕ⊂ℝ.Problem 7. Give an example of a number that is in ℝ but not in ℚ. Why does this distinction matter for computer representations of numbers?Vectors and vector operationsProblem 8. A data point for a student has the following features: GPA (3.5), SAT score (1400), and number of extracurriculars (4). Write this as a 3-vector in column notation.Problem 9. Given two vectors 𝒂=(25−1) and 𝒃=(3−24), compute 𝒂+𝒃.Problem 10. Compute 3𝒗 where 𝒗=(4−27).Problem 11. Compute −2𝒂+𝒃 where 𝒂=(31−4) and 𝒃=(−152).Problem 12. Given 𝒖=(12) and 𝒘=(46), compute 2𝒖+3𝒘.Problem 13. State whether each expression is defined. If it is, compute it.1.𝒑+𝒒 where 𝒑=(123) and 𝒒=(456)2.𝒓+𝒔 where 𝒓=(12) and 𝒔=(345)Problem 14. Why can’t you add the vectors 𝒑=(123) and 𝒒=(45)? In a machine learning context, what would this situation represent?Notation and dimensionsProblem 15. If a matrix 𝑀∈ℝ50×7, how many rows does it have? How many columns? How many total entries?Problem 16. If 𝐴∈ℝ4×2, how many entries are in 𝐴?Problem 17. Write the general form of a matrix 𝐴∈ℝ2×3 using subscript notation for each element.Problem 18. Write the general form of a matrix 𝐵∈ℝ3×2 using subscript notation.Problem 19. You have a dataset of 1000 images, where each image is represented by 784 pixel values. What is the shape of the data matrix 𝑋 if each row is one image? Write it in the form 𝑋∈ℝ𝑝×𝑑.34Problem 20. Given 𝑋∈ℝ3×4, what element is located at row 2, column 3? Write it using subscript notation.Problem 21. A machine learning model takes in data matrix 𝑋∈ℝ𝑛×𝑑 and outputs predictions ̂𝑦∈ℝ𝑛. Explain in plain English what 𝑛 and 𝑑 represent.Dot productProblem 22. Compute the dot product of 𝒂=(231) and 𝒃=(4−15).Problem 23. Compute the dot product of 𝒖=(−203) and 𝒗=(54−1).Problem 24. If 𝒙=(100) and 𝒚=(010), compute 𝒙⋅𝒚.Problem 25. A house has features 𝒙=(120003) (intercept, square footage, bedrooms) and the model weights are 𝜷=(500001005000). Compute the predicted price using the dot product.Problem 26. Compute 𝒗⋅𝒗 where 𝒗=(34).Problem 27. Given 𝒑=(2−1) and 𝒒=(43), compute 𝒑⋅𝒒.Problem 28. Why must two vectors have the same dimension for the dot product to be defined? Give a practical example where this constraint matters.Matrix multiplicationProblem 29. Given 𝐴=(1324) and 𝐵=(5768), compute the element in row 1, column 2 of 𝐴𝐵.Problem 30. If 𝐴∈ℝ3×4 and 𝐵∈ℝ4×2, what is the shape of the product 𝐴𝐵?Problem 31. Can you multiply 𝑃∈ℝ2×3 by 𝑄∈ℝ2×3? Explain why or why not.Problem 32. Compute the full matrix product:(2103)(1245)Problem 33. Compute the full matrix product:(10−1321)(2−1410−2)Problem 34. Let 𝐴∈ℝ2×3 and 𝐵∈ℝ3×1. What is the shape of 𝐴𝐵?TransposeProblem 35. Compute the transpose of 𝐴=(142536).Problem 36. If 𝑀∈ℝ10×3, what is the shape of 𝑀𝑇?Problem 37. Given the column vector 𝒗=(258), write 𝒗𝑇.Problem 38. Verify that (𝐴𝑇)𝑇=𝐴 for 𝐴=(135246).Problem 39. Let 𝐴=(1−320) and 𝐵=(42−15). Compute (𝐴+𝐵)𝑇.DeterminantProblem 40. Compute the determinant of 𝐴=(5234).35Problem 41. Compute the determinant of 𝐷=(73−12).Problem 42. Compute the determinant of 𝐵=(−236−9). Does this matrix have an inverse?Problem 43. For the matrix 𝐶=(𝑎𝑏2𝑎2𝑏), compute the determinant. What does this tell you about matrices where one column is a multiple of the other?Problem 44. If det(𝐴)=5, what is det(2𝐴) for a 2×2 matrix? (Hint: work out an example.)Problem 45. The determinant has a geometric interpretation: it tells us how a matrix scales area. If det(𝐴)=3, what happens to the area of a unit square when transformed by 𝐴? What if det(𝐴)=−2?Identity matrix and inverseProblem 46. Write the 2×2 identity matrix 𝐼2 and the 3×3 identity matrix 𝐼3.Problem 47. Compute 𝐼2𝐴 for 𝐴=(−1240).Problem 48. Verify that 𝐴𝐼2=𝐴 for 𝐴=(3275).Problem 49. Compute the inverse of 𝐴=(4232) by hand using the formula 𝐴−1=1det(𝐴)(𝑑−𝑐−𝑏𝑎).Problem 50. Compute the inverse of 𝐵=(5723) by hand. Show all steps.Problem 51. Compute the inverse of 𝐶=(1325) by hand.Problem 52. Attempt to compute the inverse of 𝐷=(6432). What happens and why?Problem 53. If 𝐴−1=(2−3−12), find 𝐴.Problem 54. For which values of 𝑘 is the matrix 𝐴=(12𝑘4) invertible?Mixed practiceProblem 55. Let 𝒂=(2−10) and 𝒃=(13−2). Compute (𝒂+𝒃)⋅𝒃.Problem 56. Let 𝐴=(1032−14). Compute 𝐴𝑇 and then compute 𝐴𝑇𝐴.Problem 57. Let 𝐵=(2−41−2). Determine whether 𝐵 is invertible and explain why.2.8Problem packetThe packet provides each calculus derivation. Apply or interpret the result instead of deriving it again. Show the requested calculations, and explain each interpretation in a complete sentence.Data representation and the design matrixProblem 1. Given the dataset below, write the design matrix 𝑋 including an intercept column and the label vector 𝑦. Then interpret the meaning of each column in one sentence.Problem 2.Size (ft squared)BedroomsPrice ($)“1,000”2“100,000”“2,000”4“200,000”36“3,000”3“300,000”“4,000”5“400,000”State the shape of 𝑋 in the form 𝑋∈ℝ𝑝×𝑑 and explain what 𝑝 and 𝑑 represent in this dataset.Matrix operations referenceProblem 3. Using the design matrix from Problem 1, compute 𝑋𝑇𝑋 and 𝑋𝑇𝑦. Then explain in one sentence what each result measures.Normal equation with derivationThe loss is the residual sum of squares:𝐿(𝛽)=(𝑦−𝑋𝛽)𝑇(𝑦−𝑋𝛽)Worked derivation:𝐿(𝛽)=𝑦𝑇𝑦−2𝛽𝑇𝑋𝑇𝑦+𝛽𝑇𝑋𝑇𝑋𝛽Taking the gradient with respect to 𝛽:∇𝛽𝐿(𝛽)=−2𝑋𝑇𝑦+2𝑋𝑇𝑋𝛽Setting to zero:−2𝑋𝑇𝑦+2𝑋𝑇𝑋𝛽=0⟶𝑋𝑇𝑋𝛽=𝑋𝑇𝑦The solution is:̂𝛽=(𝑋𝑇𝑋)−1𝑋𝑇𝑦Problem 4. Using the formula above, compute ̂𝛽 for Problem 1. Interpret each coefficient in a single clear sentence.Problem 5. Explain why ̂𝛽 is a minimizer of the loss using geometric or algebraic intuition.MSE and gradient with derivationℒ︀(𝛽)=1𝑁(𝑦−𝑋𝛽)𝑇(𝑦−𝑋𝛽)Worked derivation:∇𝛽ℒ︀(𝛽)=1𝑁(−2𝑋𝑇𝑦+2𝑋𝑇𝑋𝛽)=−2𝑁𝑋𝑇(𝑦−𝑋𝛽)The gradient descent update with learning rate 𝜂 is:𝛽←𝛽−𝜂∇𝛽ℒ︀(𝛽)=𝛽+2𝜂𝑁𝑋𝑇(𝑦−𝑋𝛽)Problem 6. Apply the formula. Given 𝑁=3, 𝜂=0.1:𝑋=(111241120),𝛽=(5052),𝑦=(657552)Compute 𝑦−𝑋𝛽, then 𝑋𝑇(𝑦−𝑋𝛽), and the update increment Δ𝛽.Problem 7. Explain in two sentences what the vector 𝑋𝑇(𝑦−𝑋𝛽) represents.37Apply RSS and MSEProblem 8. Using the results from Problem 6 (where 𝑒=(3,1,−3)𝑇), explain if the model under- or over-predicts for each observation and state what an MSE of 6.33 means.Collinearity and remediesProblem 9. In two sentences, explain how near-collinearity affects the stability of ̂𝛽. Propose one practical remedy, and explain it in one sentence.38Appendix: Programming Reference3Python and libraries3.1NumPyImport NumPyImport NumPy with its conventional alias, np:import numpy as npCreate arrays and vectorsIn NumPy an array is a generic term for a multidimensional set of numbers. One-dimensional NumPy arrays act like vectors. The following code creates two one-dimensional arrays and adds them elementwise. If you attempted the same with plain Python lists you would not get elementwise addition.x = np.array([3, 4, 5])y = np.array([4, 9, 7])print(x + y) # array([ 7, 13, 12])Represent matrices with two-dimensional arraysMatrices in NumPy are typically represented as two-dimensional arrays. The object returned by np.array has attributes such as ndim for the number of dimensions, dtype for the data type, and shape for the size of each axis.x = np.array([[1, 2], [3, 4]])print(x) # array([[1, 2], [3, 4]])print(x.ndim) # 2print(x.dtype) # e.g. dtype('int64')print(x.shape) # (2, 2)If any element passed into np.array is a floating point number, NumPy upcasts the whole array to a floating point dtype.print(np.array([[1, 2], [3.0, 4]]).dtype) # dtype('float64')print(np.array([[1, 2], [3, 4]], float).dtype) # dtype('float64')Call methods and functionsMethods are functions bound to objects. Calling x.sum() calls the sum method with x as the implicit first argument. The module-level function np.sum(x) does the same computation but is not bound to x.x = np.array([1, 2, 3, 4])print(x.sum()) # method on the array objectprint(np.sum(x)) # module-level functionThe reshape method returns a new view with the same data arranged into a new shape. You pass a tuple that specifies the new dimensions.x = np.array([1, 2, 3, 4, 5, 6])print("beginning x:\n", x)x_reshape = x.reshape((2, 3))print("reshaped x:\n", x_reshape)39NumPy uses zero-based indexing. The first row and first column entry of x_reshape is accessed with x_reshape[0, 0]. The entry in the second row and third column is x_reshape[1, 2]. The third element of the original one-dimensional x is x[2].print(x_reshape[0, 0]) # 1print(x_reshape[1, 2]) # 6print(x[2]) # third element of xUnderstand views and shared memoryReshaping often returns a view rather than a copy. Modifying a view will modify the original array because they share the same memory. This behavior is important when you expect independent copies.print("x before modification:\n", x)print("x_reshape before modification:\n", x_reshape)x_reshape[0, 0] = 5print("x_reshape after modification:\n", x_reshape)print("x after modification:\n", x)If you need an independent copy, call x.copy() explicitly.x = np.array([1, 2, 3, 4, 5, 6])x_copy = x.copy()x_reshape_copy = x_copy.reshape((2, 3))x_reshape_copy[0, 0] = 99print("x remains unchanged:\n", x)print("x_reshape_copy changed:\n", x_reshape_copy)Tuples are immutable sequences in Python and will raise a TypeError if you try to modify an element. This differs from NumPy arrays and Python lists.my_tuple = (3, 4, 5)# my_tuple[0] = 2 # would raise TypeError: 'tuple' object does not support item assignmentTranspose, ndim, and shapeYou can request several attributes at once. The transpose T flips axes and is useful for matrix algebra.print(x_reshape.shape, x_reshape.ndim, x_reshape.T)# For example: ((2, 3), 2, array([[5, 4], [2, 5], [3, 6]]))Apply elementwise operationsNumPy supports elementwise arithmetic and universal functions such as np.sqrt. Raising an array to a power is elementwise.print(np.sqrt(x)) # elementwise square rootprint(x ** 2) # elementwise squareprint(x ** 0.5) # alternative for square rootGenerate random numbersNumPy provides random number generation. The signature for rng.normal is normal(loc=0.0, scale=1.0, size=None). The arguments loc and scale are keyword arguments for mean and standard deviation and size controls the shape of the output.x = np.random.normal(size=50)print(x) # random sample from N(0,1), different each runTo create a dependent array, add a random variable with a different mean to each element.y = x + np.random.normal(loc=50, scale=1, size=50)print(np.corrcoef(x, y)) # correlation matrix between x and y40Reproduce results with the Generator APITo produce identical random numbers across runs, use np.random.default_rng with an integer seed to create a Generator object and then call its methods. The Generator API is the recommended approach for reproducibility.rng = np.random.default_rng(1303)print(rng.normal(scale=5, size=2))rng2 = np.random.default_rng(1303)print(rng2.normal(scale=5, size=2))# Both prints produce the same arrays because the same seed was used.When you use rng.standard_normal or rng.normal you are using the Generator instance, which ensures reproducibility if you control the seed.rng = np.random.default_rng(3)y = rng.standard_normal(10)print(np.mean(y), y.mean())Calculate the mean, variance, and standard deviationNumPy provides np.mean, np.var, and np.std as module-level functions. Arrays also have methods mean, var, and std. By default np.var divides by n. If you need the sample variance that divides by n minus 1, provide ddof=1.rng = np.random.default_rng(3)y = rng.standard_normal(10)print(np.var(y), y.var(), np.mean((y - y.mean())**2))print(np.sqrt(np.var(y)), np.std(y))# Use np.var(y, ddof=1) for sample variance dividing by n-1.Calculate along an axisNumPy arrays are row-major ordered. The first axis, axis=0, refers to rows and the second axis, axis=1, refers to columns. Passing axis into reduction methods lets you compute means, sums, and other statistics along rows or columns.rng = np.random.default_rng(3)X = rng.standard_normal((10, 3))print(X) # 10 by 3 matrixprint(X.mean(axis=0)) # column meansprint(X.mean(0)) # same as previousWhen you compute X.mean(axis=1) you obtain a one-dimensional array of row means. When you compute X.sum(axis=0) you obtain column sums.Plot with MatplotlibMatplotlib is the standard plotting library. A plot consists of a figure and one or more axes. The subplots function returns a tuple containing the figure and the axes. The axes object has a plot method and other methods to customize titles, labels, and markers.from matplotlib.pyplot import subplotsfig, ax = subplots(figsize=(8, 8))rng = np.random.default_rng(3)x = rng.standard_normal(100)y = rng.standard_normal(100)ax.plot(x, y) # default line plotax.plot(x, y, 'o') # scatter-like circles# To save: fig.savefig("scatter.png")# To display in an interactive session: import matplotlib.pyplot as plt; plt.show()41Keep results reproducibleUse np.random.default_rng with a fixed seed when an exercise requires repeatable results. NumPy versions can still produce small differences in some random outputs. When you calculate variance, set ddof=1 for sample variance and use the default ddof=0 for population variance.42Appendix: Math Fundamentals4Calculus4.1Limits4.2DerivativesLimit definition of a derivative4.3GradientsVector-valued functionsGradient definitionPartial derivatives and rules5Linear algebra6Statistics and probability43Reference44Glossary of DefinitionsDefinition 0.1p. 7Definition 0.2p. 7Definition 1.3.1p. 10Definition 1.5.1p. 11Definition 1.5.2p. 12Definition 1.7.1p. 15Definition 2.1.1p. 19Definition 2.1.2p. 20Definition 2.1.3p. 20Definition 2.1.4p. 20Definition 2.1.5p. 21Definition 2.1.6p. 21Definition 2.1.7p. 22Definition 2.1.8p. 22Definition 2.1.9p. 23Definition 2.1.10p. 23Definition 2.1.11p. 24Definition 2.1.12p. 24Definition 2.1.13p. 25Definition 2.1.14p. 25Definition 2.2.1p. 26Definition 2.3.1p. 26Definition 2.4.1p. 28Definition 2.5.1p. 29Definition 2.6.1p. 3245