ggplot2 implements the Grammar of Graphics, a systematic approach to building plots layer-by-layer. Every plot follows a consistent structure:
aes) – map variables to
visual properties (x, y, colour, size)geom_point(), geom_bar(), …)This document walks through seven essential chart types, explains when to use each, and shows how to implement them.
geom_point() is the workhorse. Additional layers like
geom_smooth() add trend lines.
ggplot(mtcars, aes(x = wt, y = mpg)) +
geom_point(aes(color = factor(cyl)), size = 3, alpha = 0.8) +
geom_smooth(method = "lm", se = TRUE, color = "black", linetype = "dashed") +
scale_color_viridis_d(name = "Cylinders") +
labs(
title = "Vehicle Weight vs. Fuel Efficiency",
subtitle = "Coloured by cylinder count with linear trend line",
x = "Weight (1000 lbs)",
y = "Miles per Gallon"
)Key arguments: alpha for transparency
(handles over-plotting), size for point emphasis, and
color/fill for grouping.
geom_histogram() counts observations per bin. Choose
binwidth carefully – too wide hides detail, too narrow
creates noise.
ggplot(mtcars, aes(x = mpg)) +
geom_histogram(binwidth = 2, fill = "steelblue", color = "white", alpha = 0.9) +
labs(
title = "Distribution of Miles per Gallon",
subtitle = "Binwidth = 2",
x = "MPG",
y = "Count"
)Tip: Overlay a density curve (next section) by
setting aes(y = after_stat(density)) to compare the
histogram with a smooth estimate.
geom_density() uses a kernel smoother. The
adjust parameter controls smoothness (higher =
smoother).
ggplot(mtcars, aes(x = mpg, fill = factor(cyl))) +
geom_density(alpha = 0.5, adjust = 1.5) +
scale_fill_viridis_d(name = "Cylinders") +
labs(
title = "Density of MPG by Cylinder Count",
subtitle = "Transparent overlap highlights differences in distribution shape",
x = "MPG",
y = "Density"
)When to prefer over histograms: When sample size is small, or you need a clean visual for overlapping groups.
geom_boxplot() – whiskers extend to 1.5×IQR; beyond that
are plotted as points.
ggplot(mtcars, aes(x = factor(cyl), y = mpg, fill = factor(cyl))) +
geom_boxplot(outlier.color = "red", outlier.size = 2.5, alpha = 0.8) +
scale_fill_viridis_d(name = "Cylinders") +
labs(
title = "MPG Distribution by Cylinder Count",
subtitle = "Red points indicate potential outliers",
x = "Cylinders",
y = "MPG"
) +
theme(legend.position = "none")Utility in practice: Quickly see if groups have different medians, spreads, and whether outliers exist.
geom_violin() with trim = FALSE keeps the
tails. Add draw_quantiles to overlay quartile lines.
ggplot(mtcars, aes(x = factor(cyl), y = mpg, fill = factor(cyl))) +
geom_violin(trim = FALSE, draw_quantiles = c(0.25, 0.5, 0.75), alpha = 0.7) +
scale_fill_viridis_d(name = "Cylinders") +
labs(
title = "Violin Plot of MPG by Cylinders",
subtitle = "Quartile lines (25%, 50%, 75%) are superimposed",
x = "Cylinders",
y = "MPG"
) +
theme(legend.position = "none")Combination tip: Overlay
geom_boxplot(width = 0.1) on top of a violin to keep the
summary visible while showing the density.
geom_tile() creates rectangular cells. Use
geom_text() to overlay actual numbers if the grid is
small.
# Create aggregated data: average MPG by gear and cylinder
heat_data <- mtcars %>%
group_by(gear, cyl) %>%
summarise(avg_mpg = mean(mpg), .groups = "drop")
ggplot(heat_data, aes(x = factor(gear), y = factor(cyl), fill = avg_mpg)) +
geom_tile(color = "white", linewidth = 1) +
geom_text(aes(label = round(avg_mpg, 1)), color = "black", size = 5) +
scale_fill_viridis_c(option = "plasma", name = "Avg MPG") +
labs(
title = "Average MPG by Gear and Cylinders",
subtitle = "Darker colours indicate higher fuel efficiency",
x = "Number of Gears",
y = "Cylinders"
)Note: Use coord_fixed() if you need
square tiles, especially when both axes have the same scale.
Compute cor() on numeric columns, melt()
into long format, then plot with a diverging colour
scale (midpoint = 0).
corr_matrix <- round(cor(mtcars), 2)
corr_melt <- melt(corr_matrix, varnames = c("Var1", "Var2"))
ggplot(corr_melt, aes(x = Var1, y = Var2, fill = value)) +
geom_tile(color = "white") +
geom_text(aes(label = value), size = 3.5, color = "black") +
scale_fill_gradient2(
low = "#2c3e50", # dark blue for negative
mid = "white", # zero
high = "#e74c3c", # red for positive
midpoint = 0,
limit = c(-1, 1),
name = "Correlation"
) +
labs(
title = "Correlation Matrix of mtcars Variables",
subtitle = "Values range from -1 (negative) to +1 (positive)"
) +
theme(
axis.text.x = element_text(angle = 45, hjust = 1),
axis.title = element_blank()
)Why diverging colours? They make the sign and
magnitude immediately obvious. scale_fill_gradient2 is
perfect for this.
| Plot | Geom | Data Types | Primary Utility |
|---|---|---|---|
| Scatter | geom_point() |
2 continuous | Relationship, trends, clustering, outliers |
| Histogram | geom_histogram() |
1 continuous | Distribution shape, skewness, modality |
| Density | geom_density() |
1 continuous (+ fill groups) | Smoothed distribution, compare multiple groups |
| Boxplot | geom_boxplot() |
1 continuous + 1 categorical | Summarise and compare distributions (medians, IQR) |
| Violin | geom_violin() |
1 continuous + 1 categorical | Full distribution + summary for group comparisons |
| Heatmap | geom_tile() |
2 categorical (values) | Visualise 2-D aggregated metrics / patterns |
| Correlation | geom_tile() |
Many numeric → matrix | Spot collinearity, feature relationships at scale |
Mastering these seven visualisations covers 90% of day-to-day exploratory data analysis needs. Remember the layered philosophy of ggplot2:
Data → Aesthetics → Geoms → Scales → Facets → Themes
Start simple, then iteratively add labels, colors, and statistical layers to tell a compelling data story.