This handout condenses the entire Unit 2 syllabus into a single, exam-focused reference. It covers:
No code is required for the exam — only concepts, examples, and statistical reasoning.
Statistics is the science of collecting, organizing, summarizing, analyzing, and interpreting data to make informed decisions. In engineering, statistics is used for:
Key Definition: Statistics is not just about numbers — it is a framework for decision-making under uncertainty. Engineers use it to separate signal from noise.
| Type | Purpose | Example |
|---|---|---|
| Descriptive | Summarize and describe the data at hand | Mean rainfall of Sydney = 85 mm |
| Inferential | Draw conclusions about a population from a sample | Is Sydney’s rainfall increasing over time? |
Descriptive statistics form the foundation of Exploratory Data Analysis (EDA). Inferential statistics build on EDA for hypothesis testing and modeling.
Exam Tip: A statistic estimates a parameter. The sample mean (\(\bar{x}\)) estimates the population mean (\(\mu\)). The sample standard deviation (\(s\)) estimates the population standard deviation (\(\sigma\)).
Statistical learning refers to a vast set of tools for modeling and understanding complex datasets. It blends statistics with machine learning and is central to modern data science.
Statistical learning = using data to estimate the unknown function \(f\) that maps inputs \(X\) to output \(Y\).
We observe a quantitative response \(Y\) and \(p\) predictors \(X_1, X_2, \ldots, X_p\). We assume:
\[Y = f(X) + \varepsilon\]
Where:
The goal of statistical learning is to estimate \(f\).
| Aspect | Supervised Learning | Unsupervised Learning |
|---|---|---|
| Inputs | \(X\) and \(Y\) (labeled data) | Only \(X\) (no labels) |
| Goal | Predict \(Y\) from \(X\) | Discover structure in \(X\) |
| Examples | Regression, classification, LDA | Clustering, PCA |
| Engineering Use | Predicting strength of concrete | Grouping similar soil samples |
Supervised = “learning with a teacher” (we know the correct
answers).
Unsupervised = “learning without a teacher” (we
discover patterns ourselves).
There are two main reasons to estimate \(f\):
When inputs \(X\) are available but output \(Y\) is not, we predict \(Y\) using:
\[\hat{Y} = \hat{f}(X)\]
Example: Predict a patient’s risk of adverse drug reaction from blood sample characteristics.
We want to understand how \(Y\) changes as \(X_1, \ldots, X_p\) change. We care about the form of \(f\), not just prediction accuracy.
Example: Which media (TV, radio, newspaper) actually contribute to sales?
Two-step approach:
Advantages:
Disadvantages:
Make no explicit assumptions about the functional form of \(f\). Instead, they seek an estimate that closely follows the data without excessive wiggliness.
Advantages:
Disadvantages:
Overfitting: When a model follows the random noise in the training data too closely, it performs poorly on new data. Flexible models with many parameters are prone to overfitting.
This is the bias-variance trade-off:
R is a free, open-source programming language for statistical computing and data visualization. It is widely adopted in:
| Data Type | Description | Example |
|---|---|---|
| Numerical | Continuous or discrete numbers | Temperature = 23.5°C, Rainfall = 12 mm |
| Categorical | Labels or groups | Soil type = {Clay, Sand, Silt} |
| Time-Series | Observations indexed by time | Daily rainfall from 2011–2014 |
| Spatial | Data with geographic coordinates | Rainfall at different weather stations |
| Ordinal | Ordered categories | Concrete grade = {M20, M25, M30} |
| Binary | Two categories | Yes/No, Pass/Fail |
| Structure | Description | Can Hold Mixed Types? |
|---|---|---|
| Vector | 1D sequence of same type | No |
| Factor | Categorical variable | No (levels) |
| Matrix | 2D, same type | No |
| Data Frame | 2D, columns can differ | Yes |
| List | Ordered collection of anything | Yes |
| Tibble | Modern data frame | Yes |
A data frame is the workhorse of R data analysis — it is like an Excel sheet where each column can be a different type (numeric, character, factor, date).
\[\bar{x} = \frac{1}{n}\sum_{i=1}^{n} x_i\]
The middle value when data is sorted. If \(n\) is even, average of the two middle values.
The most frequently occurring value.
Exam Tip: When data is skewed (e.g., rainfall, income), the median is a better measure of center than the mean. When data is symmetric, mean = median = mode.
\[\text{Range} = \text{Max} - \text{Min}\]
\[s^2 = \frac{1}{n-1}\sum_{i=1}^{n}(x_i - \bar{x})^2\]
\[s = \sqrt{s^2}\]
\[\text{IQR} = Q_3 - Q_1\]
This summary is the basis of the boxplot.
| Shape | Description | Mean vs. Median |
|---|---|---|
| Symmetric | Bell-shaped, equal tails | Mean ≈ Median |
| Right-skewed (positive) | Long right tail | Mean > Median |
| Left-skewed (negative) | Long left tail | Mean < Median |
Rainfall example: Most days have little rain, but a few days have heavy downpours. This creates a right-skewed distribution. The mean is pulled upward by extreme values, so the median is a better summary.
\[\text{CV} = \frac{s}{\bar{x}} \times 100\%\]
| Distribution | Type | Use Case |
|---|---|---|
| Binomial | Discrete | Number of successes in \(n\) trials |
| Poisson | Discrete | Number of events in a time interval |
| Normal | Continuous | Heights, errors, temperatures |
| Exponential | Continuous | Time between events |
| Log-normal | Continuous | Rainfall, income |
| Weibull | Continuous | Reliability, failure times |
Exam Tip: For risk assessment, engineers use return periods. A 100-year flood has a 1% chance of occurring in any given year. This is based on probability distributions fitted to historical data.
Real-world data is messy. Common issues:
Garbage In, Garbage Out (GIGO): No statistical method can fix bad data. Cleaning is often 60–80% of the analysis effort.
| Strategy | When to Use | Drawback |
|---|---|---|
| Remove rows | Few missing values, MCAR | Loses information |
| Mean/Median imputation | Simple, quick | Reduces variance, biases |
| Mode imputation | Categorical data | Same as above |
| Regression imputation | MAR, correlated variables | Complex, can overfit |
| Multiple imputation | Best practice | Computationally intensive |
A missingness heatmap shows where NAs occur (red = missing, gray = present). This helps identify patterns.
Duplicate rows over-represent certain observations and bias results.
distinct() in R to remove exact duplicatesCaution: Not all duplicates are errors. Multiple orders from the same customer on the same day may be legitimate. Understand the context before removing.
An observation is an outlier if:
\[x < Q_1 - 1.5 \times IQR \quad \text{or} \quad x > Q_3 + 1.5 \times IQR\]
\[z = \frac{x - \bar{x}}{s}\]
Exam Tip: In rainfall data, extreme values (storms) are often real and important. Removing them would destroy the analysis of extreme events. Always investigate before removing outliers.
Exploratory Data Analysis (EDA) is the process of examining data before formal modeling to:
EDA is detective work. You look for clues before building a case (model).
Examine one variable at a time:
Examine relationships between two variables:
Examine three or more variables:
The Iris dataset (Fisher, 1936) contains 150 flowers from 3 species:
Four measurements per flower (in cm):
Why it’s a perfect teaching tool:
| Term | Description | Function |
|---|---|---|
| Sepal | Outer, leaf-like protective structures | Protect flower bud before bloom |
| Petal | Inner, often colorful structures | Attract pollinators |
ggplot2 implements the Grammar of Graphics — a systematic approach to building plots layer-by-layer.
Every plot has:
aes) — map variables to
visual properties (x, y, color, size)geom_point(), geom_bar(), etc.)Layered philosophy: Data → Aesthetics → Geoms → Scales →
Facets → Themes
Start simple, then add layers iteratively.
geom_point())Purpose: Visualize relationship between two continuous variables.
What to look for:
Enhancements:
geom_smooth() adds trend linecolor/fill for groupingalpha for transparency (handles overplotting)size for point emphasisExample (mtcars): Weight vs. MPG — heavier cars have lower fuel efficiency (negative relationship).
Best for: Hypothesis generation, trend detection, identifying data gaps.
geom_histogram())Purpose: Show distribution of a single continuous variable by binning values.
What to look for:
Key parameter: binwidth — too wide
hides detail, too narrow creates noise.
Example (mtcars): Distribution of MPG — most cars get 15–25 MPG, with a few outliers.
Best for: Initial data exploration, detecting data entry errors, understanding central tendency.
geom_density())Purpose: Smoothed version of histogram — estimates probability density function.
What to look for: Same as histograms, but easier to compare multiple groups.
Key parameter: adjust — controls
smoothness (higher = smoother).
Example (mtcars): Density of MPG by cylinder count — 4-cylinder cars have higher MPG, 8-cylinder lower.
When to prefer over histograms:
geom_boxplot())Purpose: Display five-number summary (min, Q1, median, Q3, max) and outliers.
What to look for:
Whiskers: Extend to 1.5 × IQR; beyond that are plotted as points.
Example (mtcars): MPG by cylinder count — 4-cylinder cars have higher median MPG.
Best for: Comparing distributions across several categorical groups compactly.
geom_violin())Purpose: Combine boxplot with mirrored density plot — shows full shape of distribution and summary statistics.
What to look for:
Key parameters:
trim = FALSE keeps the tailsdraw_quantiles overlays quartile linesExample (mtcars): Violin plot of MPG by cylinders — shows the shape of each group’s distribution.
Combination tip: Overlay
geom_boxplot(width = 0.1) on top of a violin to keep
summary visible.
Best for: Rich group comparisons when you suspect complex distributions (e.g., bimodal data).
geom_tile())Purpose: Encode a matrix of numeric values using color intensity.
What to look for:
Best for:
Example: Average MPG by gear and cylinder — darker colors indicate higher fuel efficiency.
Note: Use coord_fixed() for square
tiles.
geom_tile() +
scale_fill_gradient2)Purpose: Visualize pairwise correlations between multiple numeric variables.
What to look for:
Best for:
Why diverging colors? They make the sign and
magnitude immediately obvious. scale_fill_gradient2 is
perfect (midpoint = 0).
Example (mtcars): Correlation matrix of all variables — weight and MPG are strongly negatively correlated.
| Plot | Geom | Data Types | Primary Utility |
|---|---|---|---|
| Scatter | geom_point() |
2 continuous | Relationship, trends, clustering, outliers |
| Histogram | geom_histogram() |
1 continuous | Distribution shape, skewness, modality |
| Density | geom_density() |
1 continuous (+ groups) | Smoothed distribution, compare groups |
| Boxplot | geom_boxplot() |
1 continuous + 1 categorical | Summarize and compare distributions |
| Violin | geom_violin() |
1 continuous + 1 categorical | Full distribution + summary |
| Heatmap | geom_tile() |
2 categorical (values) | 2-D aggregated metrics / patterns |
| Correlation | geom_tile() |
Many numeric → matrix | Collinearity, feature relationships |
Correlation measures the strength and direction of a linear relationship between two variables.
\[r = \frac{\sum (x_i - \bar{x})(y_i - \bar{y})}{\sqrt{\sum (x_i - \bar{x})^2 \sum (y_i - \bar{y})^2}}\]
Correlation ≠ Causation: Just because two variables are correlated does not mean one causes the other. There may be a confounding variable, or the relationship may be coincidental.
| \(|r|\) Value | Interpretation |
|---|---|
| 0.0 – 0.2 | Very weak / negligible |
| 0.2 – 0.4 | Weak |
| 0.4 – 0.6 | Moderate |
| 0.6 – 0.8 | Strong |
| 0.8 – 1.0 | Very strong |
Spearman’s rank correlation (\(\rho\)) measures monotonic relationships (not necessarily linear).
A correlation matrix shows pairwise correlations among multiple variables.
Multicollinearity occurs when predictors are highly correlated with each other.
Problems:
Detection:
Solution:
When exploring feature relationships:
Regression analysis estimates relationships between a dependent variable (response) and one or more independent variables (predictors).
Goals:
\[Y = \beta_0 + \beta_1 X + \varepsilon\]
Finds the line that minimizes the sum of squared errors (SSE):
\[\text{SSE} = \sum_{i=1}^{n}(y_i - \hat{y}_i)^2\]
Predicting Petal.Length from Sepal.Length:
\[Y = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \cdots + \beta_p X_p + \varepsilon\]
Each \(\beta_j\) represents the change in \(Y\) for a one-unit change in \(X_j\), holding all other predictors constant.
Penalizes for adding unnecessary predictors:
\[R^2_{adj} = 1 - \frac{(1-R^2)(n-1)}{n-p-1}\]
Predicting Petal.Length from Sepal.Length, Sepal.Width, and Petal.Width:
\[Y = \beta_0 + \beta_1 X + \beta_2 X^2 + \beta_3 X^3 + \cdots + \varepsilon\]
Use when: Relationship between \(X\) and \(Y\) is curved.
Example: Plant growth vs. fertilizer — too little or too much fertilizer reduces growth (quadratic relationship).
Caution: High-degree polynomials can overfit. Use only when justified by theory or data.
\[P(Y=1) = \frac{1}{1 + e^{-(\beta_0 + \beta_1 X_1 + \cdots + \beta_p X_p)}}\]
Use when: Response variable is binary (0/1, yes/no, pass/fail).
Output: Predicted probability between 0 and 1.
Example (Iris): Predicting whether a flower is setosa (1) or not (0) based on petal length.
Interpretation:
LDA is both a dimensionality reduction technique and a classification method. Developed by Fisher (1936) — the same Fisher who gave us the iris dataset!
LDA seeks a projection vector \(w\) that maximizes:
\[J(w) = \frac{w^T S_B w}{w^T S_W w}\]
Where:
Intuition: LDA finds the “best viewing angle” to separate the species. Imagine rotating a 3D scatter plot until the three species are most clearly separated — that’s what LDA does mathematically.
\[\text{Minimize: } \sum_{i=1}^{n}(y_i - \hat{y}_i)^2 + \lambda \sum_{j=1}^{p} \beta_j^2\]
Use when: Predictors are highly correlated (multicollinearity).
Effect: Shrinks coefficients toward zero but never exactly zero.
Parameter \(\lambda\): Controls amount of shrinkage (higher = more shrinkage).
\[\text{Minimize: } \sum_{i=1}^{n}(y_i - \hat{y}_i)^2 + \lambda \sum_{j=1}^{p} |\beta_j|\]
Use when: You need feature selection (some coefficients become exactly zero).
Effect: Performs variable selection by shrinking some coefficients to zero.
Difference from Ridge: Lasso can eliminate variables; Ridge only shrinks them.
Estimates conditional quantiles (e.g., median, 10th percentile, 90th percentile) of the response, rather than the mean.
Use when:
Advantage: Robust to outliers.
Example: 90th percentile regression shows how extreme rainfall changes with temperature, while median regression shows how typical rainfall changes.
| Method | Type | Response | Key Feature | Best Use Case |
|---|---|---|---|---|
| Simple Linear | Regression | Continuous | One predictor | Basic relationships |
| Multiple Linear | Regression | Continuous | Multiple predictors | Complex relationships |
| Polynomial | Regression | Continuous | Curved relationships | Non-linear patterns |
| Logistic | Classification | Binary | Probability estimation | Binary outcomes |
| LDA | Classification | Categorical | Dimension reduction | Multi-class problems |
| Ridge | Regression | Continuous | L2 regularization | Correlated predictors |
| Lasso | Regression | Continuous | L1 regularization | Feature selection |
| Quantile | Regression | Continuous | Robust to outliers | Non-normal errors |
Measures influence of each observation on the model. Values > \(4/n\) are potentially influential.
AvgTemp = (MinTemp + MaxTemp) / 2logRainfall = log(Rainfall + 1) to handle
skewnessmonth, year, year_monthtotal_discount_in_dollars = Sales × Discountselling_price = Sales / Quantityprofit_margin = (Profit / Sales) × 100order_fulfillment_time = Ship Date - Order Date| Measure | Formula | Robust? |
|---|---|---|
| Mean | \(\bar{x} = \frac{1}{n}\sum x_i\) | No |
| Median | Middle value | Yes |
| Variance | \(s^2 = \frac{1}{n-1}\sum(x_i - \bar{x})^2\) | No |
| Standard Deviation | \(s = \sqrt{s^2}\) | No |
| IQR | \(Q_3 - Q_1\) | Yes |
| Range | Max - Min | No |
| CV | \(\frac{s}{\bar{x}} \times 100\%\) | No |
| Method | Formula | Assumption |
|---|---|---|
| IQR | \(x < Q_1 - 1.5 \times IQR\) or \(x > Q_3 + 1.5 \times IQR\) | None |
| Z-score | \(\|z\| > 3\) where \(z = (x - \bar{x})/s\) | Normality |
| Assumption | Check | Fix |
|---|---|---|
| Linearity | Residuals vs. Fitted | Add polynomial terms |
| Independence | Residuals vs. order | Time-series models |
| Normality | Q-Q plot | Transform \(Y\) |
| Equal variance | Scale-Location | Weighted least squares |
| \(|r|\) | Interpretation |
|---|---|
| 0.0–0.2 | Very weak |
| 0.2–0.4 | Weak |
| 0.4–0.6 | Moderate |
| 0.6–0.8 | Strong |
| 0.8–1.0 | Very strong |
| Goal | Plot |
|---|---|
| Relationship between 2 continuous variables | Scatter |
| Distribution of 1 continuous variable | Histogram or Density |
| Compare distributions across groups | Boxplot or Violin |
| 2-D aggregated values | Heatmap |
| Pairwise correlations | Correlation matrix |
| Time trend | Line plot |
| Part-to-whole | Stacked bar or Pie |
| Scenario | Recommended Method |
|---|---|
| Continuous response, 1 predictor | Simple Linear Regression |
| Continuous response, multiple predictors | Multiple Linear Regression |
| Curved relationship | Polynomial Regression |
| Binary response | Logistic Regression |
| Multi-class classification | LDA |
| Multicollinearity | Ridge Regression |
| Feature selection | Lasso Regression |
| Outliers / non-normal errors | Quantile Regression |
1. Understand concepts, don’t memorize formulas. Know why each method is used, not just how to compute it.
2. Practice interpreting output. Given a regression summary, explain what each coefficient means in context.
3. Link plots to purposes. For each plot type, know when to use it and what to look for.
4. Know the assumptions. Every statistical method has assumptions. Know what they are and how to check them.
5. Think about the workflow. EDA → Cleaning → Visualization → Modeling → Diagnostics → Interpretation.
6. Use examples. Bring in the iris, mtcars, rainfall, and superstore examples to illustrate concepts.
7. Distinguish correlation from causation. This is a common exam question.
8. Know the difference between parametric and non-parametric. And when to use each.
9. Be able to explain R-squared and Adjusted R-squared. Why is Adjusted R-squared preferred for model comparison?
10. Understand LDA intuitively. It finds the best viewing angle to separate classes.
| Term | Definition |
|---|---|
| Aesthetics | Mapping of variables to visual properties in ggplot2 |
| AAGR | Average Annual Growth Rate |
| Bias | Error from overly simple models (underfitting) |
| Boxplot | Plot showing five-number summary |
| Categorical | Data consisting of labels or groups |
| Cooks Distance | Measure of influence of an observation |
| Correlation | Strength and direction of linear relationship |
| CV | Coefficient of Variation |
| Density Plot | Smoothed histogram |
| EDA | Exploratory Data Analysis |
| Facet | Sub-plot by categorical variable |
| Geom | Geometric object in ggplot2 |
| GIGO | Garbage In, Garbage Out |
| Heatmap | Color-encoded matrix of values |
| Histogram | Binned distribution of a continuous variable |
| Homoscedasticity | Constant variance of errors |
| IQR | Interquartile Range |
| LDA | Linear Discriminant Analysis |
| Lasso | L1 regularization for feature selection |
| Least Squares | Method minimizing sum of squared errors |
| Logistic Regression | Regression for binary response |
| MCAR | Missing Completely At Random |
| Multicollinearity | High correlation among predictors |
| Non-parametric | Methods making no distributional assumptions |
| Outlier | Observation far from other data points |
| Overfitting | Model follows noise too closely |
| Parametric | Methods assuming a functional form |
| Pearson | Correlation for linear relationships |
| Polynomial Regression | Regression with curved terms |
| Quantile Regression | Regression for conditional quantiles |
| R-squared | Proportion of variance explained |
| Ridge | L2 regularization for correlated predictors |
| Scatter Plot | Plot of two continuous variables |
| Spearman | Rank-based correlation |
| Standard Deviation | Average distance from the mean |
| STL | Seasonal-Trend decomposition using Loess |
| Supervised Learning | Learning with labeled data |
| Unsupervised Learning | Learning without labels |
| Variance | Average squared deviation from mean |
| VIF | Variance Inflation Factor |
| Violin Plot | Boxplot + density plot |
| Z-score | Standardized value \((x-\bar{x})/s\) |
This handout has covered the complete Unit 2 syllabus:
Remember the layered philosophy of data analysis:
Data → Cleaning → EDA → Visualization → Modeling → Diagnostics →
Interpretation → Decision
Start simple, check assumptions,
and iteratively add complexity to tell a compelling data story.