class: center, middle, inverse, title-slide .title[ # .b[Statistics &
Exploratory Data Analysis] ] .subtitle[ ## .b.f3[14hr | 24CV201T | July: 2026-27] ] .author[ ### .mv0.lh-solid[Dr. Ankit Deshmukh
Assistant Professor SoT, PDEU] ] .date[ ### 06 July 2026 ] --- class: left top inverse <style> .title-slide { background-image: url("./images/u2/Background-02.png"); } </style>
<!-- ------------------------- Start your slides ------------------------- --> # .gold[Content of Unit -2] .center.f3.co-warn[Building a Strong Foundation in Data Exploration, Statistical Thinking, and Engineering Decision-Making] .center[<img src="images/u2/lie.jpg" style="width:100%;border-radius:20px;border:2px solid gold;">] .footnote[[How to Lie With Statistics](https://www.shortform.com/pdf/how-to-lie-with-statistics-pdf-darrell-huff)] --- # Content of Unit -2 | Session | Duration | Topics Covered | | ------------- | :------: | -------------------------------------------------------------------------------------------------------------------------------------------------------- | | **1** | 2 Hours | Introduction to Statistics, Role of Statistics in Engineering, Exploratory Data Analysis (EDA), EDA Workflow, Importing and Inspecting Data in R | | **2** | 2 Hours | Engineering Data Types: Numerical, Categorical, Time-Series, Spatial Data, Data Structures in R | | **3** | 2 Hours | Descriptive Statistics: Mean, Median, Mode, Range, Variance, Standard Deviation, Quartiles, IQR, Summary Statistics | | **4** | 2 Hours | Probability Foundations for Risk Assessment: Probability Concepts, Random Variables, Probability Distributions, Engineering Applications | | **5** | 2 Hours | Visual Analytics with **ggplot2**: Scatter Plots, Histograms, Density Plots, Boxplots, Violin Plots, Heatmaps, Correlation Matrices | | **6** | 2 Hours | Data Cleaning Procedures: Missing Values, Duplicate Records, Outlier Detection (IQR & Z-score), Data Validation and Quality Checks | | **7** | 2 Hours | Identifying Feature Relationships, Correlation Analysis, Environmental Dataset Analysis (Rainfall Variability), Complete EDA Workflow and Interpretation | --- # Statistical learning Statistical learning refers to a set of tools for modeling and understanding complex datasets. It is a recently developed area in statistics and blends with parallel developments in computer science and, in particular, machine learning. - With the explosion of "Big Data" problems, statistical learning has become a very hot field in many scientific areas as well as marketing, finance, and other business disciplines. - People with statistical learning skills are in high demand. .center.co-tip[Statistical learning refers to a vast set of tools for understanding data.] These tools can be classified as supervised or unsupervised. 1. **Supervised statistical learning** involves building a statistical model for predicting, or estimating, an output based on one or more inputs. - Problems of this nature occur in fields as diverse as business, medicine, astrophysics, and public policy. 1. **Unsupervised statistical learning** there are inputs but no supervising output; nevertheless we can learn relationships and structure from such data. --- # .blue[**R**] - A tools for Statistical Data Analysis .pr-70[ > R is a programming language for statistical computing and data visualization. It has been widely adopted in the fields of data mining, bioinformatics, data analysis, and data science. > Widely adopted by universities, research organizations, government agencies, and industries for data science and analytics. ] .pl-30[ <img src="https://external-content.duckduckgo.com/iu/?u=https%3A%2F%2Fwww.clipartmax.com%2Fpng%2Fmiddle%2F13-137348_logo-r-programming.png&f=1&nofb=1&ipt=ec1e0f59a9546744ca8a44d9b6ecdfd13ac4e8848dad7c23b6875469b8d0c308" style="width:40%;"> ] ### Key Contributors - Ross Ihaka - Co-creator of R - Robert Gentleman - Co-creator of R ### Keypoints - 1992: Developed by Ross Ihaka and Robert Gentleman at the University of Auckland, New Zealand. - Inspired by the S programming language, developed at Bell Laboratories. - 1995: Released as free and open-source software under the GNU General Public License (GPL). - 1997: The Comprehensive R Archive Network (CRAN) was established to distribute R packages. - 2000: R version 1.0.0 officially released, marking its maturity as a statistical computing platform. - Today, R has 20,000+ packages on CRAN for statistics, machine learning, visualization, finance, healthcare, geospatial analysis, and engineering. --- ## Wage Dataset - A Sample data .center[<img src="images/u2/Screenshot 2026-08-12 114943.png" style="width:70%;">] Wage data, which contains income survey information for males from the central Atlantic region of the United States. - Left: wage as a function of age. On average, wage increases with age until about 60 years of age, at which point it begins to decline. - Center: wage as a function of year. There is a slow but steady increase of approximately $10,000 in the average wage between 2003 and 2009. - Right: Boxplots displaying wage as a function of education. On average, wage increases with the level of education. --- # A sample dataset URL: https://github.com/selva86/datasets/blob/master/Wage.csv ``` r url <- "./Wage.csv" data <- read.csv(url) ``` ### Get the basic information about the data ``` r head(data, n = 3) ``` ``` ## year age sex maritl race education region ## 1 2006 18 1. Male 1. Never Married 1. White 1. < HS Grad 2. Middle Atlantic ## 2 2004 24 1. Male 1. Never Married 1. White 4. College Grad 2. Middle Atlantic ## 3 2003 45 1. Male 2. Married 1. White 3. Some College 2. Middle Atlantic ## jobclass health health_ins logwage wage ## 1 1. Industrial 1. <=Good 2. No 4.318063 75.04315 ## 2 2. Information 2. >=Very Good 2. No 4.255273 70.47602 ## 3 1. Industrial 1. <=Good 1. Yes 4.875061 130.98218 ``` ``` r dim(data) ``` ``` ## [1] 3000 12 ``` --- ## What Is Statistical Learning? .f3[Statistical learning theory is a framework for machine learning drawing from the fields of statistics and functional analysis.] ### A Brief History of Statistical Learning - **Early 1800s - Linear Methods:** Legendre and Gauss developed least squares, forming the foundation of linear regression for quantitative prediction. - **1930s-1970s - Classification & GLMs:** Fisher introduced Linear Discriminant Analysis (1936), followed by logistic regression and Generalized Linear Models (GLMs). - **1980s - Non-linear Learning:** Advances in computing enabled Classification and Regression Trees (CART) and Generalized Additive Models (GAMs), with cross-validation supporting model selection. - **Modern Era - Statistical Learning & ML:** Statistical learning evolved into a broader field covering supervised and unsupervised learning, accelerated by powerful computing and accessible software such as R. > Suppose that we are statistical consultants hired by a client to provide advice on how to improve sales of a particular product. - Which media contribute to sales? - Which media generate the biggest boost in sales? or - How much increase in sales is associated with a given increase in TV advertising? --- ## Y = f(X) + `\(\varepsilon\)` We observe a quantitative response `\(Y\)` and `\(p\)` different predictors, `\(X_1, X_2, \ldots, X_p\)`. We assume that there is some relationship between `\(Y\)` and `\(X = (X_1, X_2, \ldots, X_p)\)`, which can be written in the very general form in eq. below, .b.red[where <i>f</i> is some fixed but unknown function] and `\(\varepsilon\)` is a random error. $$ Y = f(X) + \varepsilon, $$ .center[<img src="images/u2/Screenshot 2026-08-12 162343.png" style="width:60%;">] --- ## Fixed but Unknown Function `\(f\)` may involve more than one input variable .pl-40[ > .f3[We can say, statistical learning refers to a set of approaches for estimating **_f_**.] The figure below plot `income` as a function of `years of education` and `seniority`. Here `\(f\)` is a two-dimensional surface that must be estimated based on the observed data. - The plot displays income as a function of years of education and seniority in the Income data set. - The blue surface represents the true underlying relationship between income and years of education and seniority, which is known since the data are simulated. - The red dots indicate the observed values of these quantities for 30 individuals. ] .pr-60[ .center[<img src="images/u2/Screenshot 2026-08-12 163449.png" style="width:100%;">] ] --- ## The way to find the function `\(f\)` is the key of statistical learning There are two main reasons that we may wish to estimate `\(f\)`: .green.b[prediction] and .blue.b[inference]. ### Prediction In many situations, a set of inputs `\(X\)` are readily available, but the output `\(Y\)` cannot be easily obtained. In this setting, since the error term averages to zero, we can predict `\(Y\)` using: `$$\hat{Y} = f(\hat{X})$$` where `\(\hat{f}\)` represents our estimate for `\(f\)`, and `\(\hat{Y}\)` represents the resulting prediction for `\(Y\)`. > As an example, suppose that `\(X_1, \ldots , X_p\)` are characteristics of a patient's blood sample that can be easily measured in a lab, and Y is a variable encoding the patient's risk for a severe adverse reaction to a particular drug. ### Inference We are often interested in understanding the way that `\(Y\)` is affected as `\(X_1, \ldots, X_p\)` change. In this situation we wish to estimate `\(f\)`, but our goal is not necessarily to make predictions for `\(Y\)`. --- # How do we estimate `\(f\)` - We explore many linear and non-linear approaches for estimating `\(f\)`. - These methods share certain characteristics and we will overview of these shared characteristics. - We will always assume that we have observed a set of `\(n\)` different data points. - These observations are called the training data because we will use these training data observations to train, or teach, our method how to estimate `\(f\)`. .co-tip[Our goal is to apply a statistical learning method to the training data in order to estimate the unknown function _**f**_.] --- # 1. Parametric Methods - Overview Two-step model-based approach: 1. Assume functional form of `\(f\)` 2. Fit/train the model using training data ### Step 1: Assume linear form Simple assumption: `\(f\)` is **linear** in `\(X\)`: `$$f(X) = \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \ldots + \beta_p X_p$$` .b[Why to Use:] Instead of estimating arbitrary `\(p\)`-dimensional `\(f(X)\)`, we only estimate `\(p+1\)` coefficients `\(\beta_0, \beta_1, \ldots, \beta_p\)` ### Step 2: Fit the model and estimation method Estimate parameters `\(\beta_0, \beta_1, \ldots, \beta_p\)` such that: `$$Y \approx \beta_0 + \beta_1 X_1 + \beta_2 X_2 + \ldots + \beta_p X_p$$` - .b[Most common:] Ordinary Least Squares - .b[Alternatives:] Other approaches for parameter estimation --- # Parametric Approach - Benefits .pull-left[ **Definition:** Model-based approach reducing estimation of `\(f\)` to estimating a set of parameters - Easier to estimate `\(\beta_0, \beta_1, \ldots, \beta_p\)` than arbitrary `\(f\)` - Linear model `\(f(X) = \beta_0 + \beta_1 X_1 + \ldots + \beta_p X_p\)` ### Parametric Approach: Trade-offs - Chosen model usually won't match true unknown `\(f\)` - Choose **flexible models** fitting many functional forms - Flexible models -> more parameters -> **overfitting** .footnote[Overfitting = following errors/noise too closely] ] .pull-right[ <img src="images/u2/Screenshot 2026-08-13 103607.png" style="width:100%;"> _A linear model fit by least squares to the Income data. The observations are shown in red, and the yellow plane indicates the least squares fit to the data._ ] --- ## Choosing the Right Function: Balancing Simplicity and Flexibility. .center[<img src="images/u2/fit.png" style="width:80%;">] --- ## Approaches applied to the Income data .pull-left[ - A linear model of the form: `$$\text{income} \approx \beta_0 + \beta_1 \times \text{education} + \beta_2 \times \text{seniority}$$` <img src="images/u2/Screenshot 2026-08-13 103607.png" style="width:80%;"> Assumed a linear relationship between the response and the two predictors, the entire fitting problem reduces to estimating `\(\beta_0, \beta_1, and \beta_2\)`, which we do using least squares linear regression. ] .pull-right[ - Non-parametric Methods <img src="images/u2/Screenshot 2026-08-14 105942.png" style="width:100%;"> A smooth thin-plate spline fit to the Income data is shown in yellow; the observations are displayed in red. ] --- # Overview of Non-Parametric Methods - Non-parametric methods make no explicit assumptions about the functional form of `\(f\)`, instead seeking an estimate that closely follows the data without excessive wiggliness. - A key advantage is their flexibility, allowing them to accurately fit a much wider range of possible shapes for the true `\(f\)` than parametric approaches. - While parametric models risk poor fits if their assumed form deviates from the true `\(f\)`, non-parametric methods inherently avoid this model-misspecification danger. - The major drawback is that, without reducing `\(f\)` to a small set of parameters, these methods require substantially more observations than parametric models to produce an accurate estimate. --- # Iris dataset <img src="images/u2/iris.png" style="width:80%;"> ``` r summary(iris) ``` ``` ## Sepal.Length Sepal.Width Petal.Length Petal.Width ## Min. :4.300 Min. :2.000 Min. :1.000 Min. :0.100 ## 1st Qu.:5.100 1st Qu.:2.800 1st Qu.:1.600 1st Qu.:0.300 ## Median :5.800 Median :3.000 Median :4.350 Median :1.300 ## Mean :5.843 Mean :3.057 Mean :3.758 Mean :1.199 ## 3rd Qu.:6.400 3rd Qu.:3.300 3rd Qu.:5.100 3rd Qu.:1.800 ## Max. :7.900 Max. :4.400 Max. :6.900 Max. :2.500 ## Species ## setosa :50 ## versicolor:50 ## virginica :50 ## ## ## ``` --- # LDA Projection of Iris Species .pull-left[ ``` r library(MASS) # The proper fit: Species ~ all 4 measurements fit_lda <- lda(Species ~ ., data = iris) fit_lda # Predict LDA scores lda_scores <- predict(fit_lda)$x # Base R plot of the two discriminants plot(lda_scores[, 1], lda_scores[, 2], main = "LDA Projection of Iris Species", xlab = "Linear Discriminant 1 (LD1)", ylab = "Linear Discriminant 2 (LD2)", pch = 19, col = as.numeric(iris$Species) ) ``` ] .pull-right[ ``` r # Predict species for the training data predicted <- predict(fit_lda)$class # Confusion Matrix (Predicted vs Actual) confusion_table <- table(Predicted = predicted, Actual = iris$Species) confusion_table # Calculate accuracy accuracy <- sum(diag(confusion_table)) / sum(confusion_table) cat("Overall Accuracy:", round(accuracy * 100, 2), "%") ``` ] --- # LDA Projection of Iris Species ``` ## Call: ## lda(Species ~ ., data = iris) ## ## Prior probabilities of groups: ## setosa versicolor virginica ## 0.3333333 0.3333333 0.3333333 ## ## Group means: ## Sepal.Length Sepal.Width Petal.Length Petal.Width ## setosa 5.006 3.428 1.462 0.246 ## versicolor 5.936 2.770 4.260 1.326 ## virginica 6.588 2.974 5.552 2.026 ## ## Coefficients of linear discriminants: ## LD1 LD2 ## Sepal.Length 0.8293776 -0.02410215 ## Sepal.Width 1.5344731 -2.16452123 ## Petal.Length -2.2012117 0.93192121 ## Petal.Width -2.8104603 -2.83918785 ## ## Proportion of trace: ## LD1 LD2 ## 0.9912 0.0088 ``` --- # Base R plot of the two discriminants <img src="Unit_2_Statistics_and_Exploratory_Data_Analysis_files/figure-html/unnamed-chunk-9-1.png" width="100%" /> --- # Prediction with IRIS class with LDA ``` r # Predict species for the training data predicted <- predict(fit_lda)$class # Confusion Matrix (Predicted vs Actual) confusion_table <- table(Predicted = predicted, Actual = iris$Species) confusion_table ``` ``` ## Actual ## Predicted setosa versicolor virginica ## setosa 50 0 0 ## versicolor 0 48 1 ## virginica 0 2 49 ``` ``` r # Calculate accuracy accuracy <- sum(diag(confusion_table)) / sum(confusion_table) cat("Overall Accuracy:", round(accuracy * 100, 2), "%") ``` ``` ## Overall Accuracy: 98 % ``` --- # Linear Discriminant Analysis (LDA): > Linear Discriminant Analysis (LDA) is a supervised statistical method used for two main purposes: dimensionality reduction (projecting data onto a lower-dimensional space) and classification (predicting which group a new observation belongs to). **Model Training:** Fits a Linear Discriminant Analysis (LDA) model using all four numeric measurements (Sepal.Length, Sepal.Width, Petal.Length, Petal.Width) as predictors to classify the three Species in the iris dataset, then prints the model details (prior probabilities and group means). **Score Extraction:** Computes the linear discriminant scores for each training observation by applying the fitted model, and stores only the projected coordinates ($x) for the first two discriminant axes (LD1 and LD2) into a variable for plotting. **Visualization:** Generates a scatter plot of the observations in the reduced 2-dimensional LDA space (LD1 vs. LD2), coloring each point by its actual species and adding a legend to distinguish the three groups. **Classification & Confusion Matrix:** Predicts the species class for the same training data using the fitted LDA model, then builds a confusion matrix (predicted vs. actual labels) to tabulate how many observations were correctly or incorrectly classified. **Accuracy Evaluation:** Calculates the overall classification accuracy by dividing the sum of the diagonal entries (correct predictions) in the confusion matrix by the total number of observations, and prints the result as a cleanly formatted percentage. ---