Recently Published
Project 2 — Data Tidying: State Unemployment
Tidying a BLS table whose two-level headers spread five measures across two years, then comparing state unemployment rates and their change from 2023 to 2024.
Project 2 — Data Tidying: Electricity Production by Country
Tidying a Wikipedia electricity table with multi-level headers, then asking which countries are most sustainable and how they group by dominant energy source.
Project 2 — Data Tidying: Life Expectancy by Country
Tidying the World Bank life expectancy table from wide to long with tidyr, then comparing countries over time and across the COVID-19 period.
Chess Elo Calculations
Calculating each player's Elo expected score from the Project 1 chess tournament, and listing the biggest over- and underperformers.
Tidying and Transforming Data: Airline Arrival Delays
Tidying a small airline delay table in R, then comparing two airlines overall and city by city — a worked example of Simpson's paradox.
Project 1: Chess Tournament Results
Parsing a formatted chess tournament text file in R into a tidy CSV with each player's name, state, points, pre-rating and average opponent pre-rating.
SQL Window Functions: Stock Price Averages
Year-to-date and six-day moving averages of Apple, Microsoft and Amazon closing prices, calculated with SQL window functions and loaded into R.
Global Baseline Estimate Recommender
A Global Baseline Estimate movie recommender in R, checked against the course spreadsheet and applied to my Week 2 movie ratings.
Evaluating Classification Model Performance
Null error rate, confusion matrices, and accuracy, precision, recall and F1 at thresholds 0.2, 0.5 and 0.8 for a penguin sex classifier, with use cases for each threshold.
SQL and R: Movie Ratings
Six friends rate six 2025 movies: a normalized SQLite database loaded into R, with a tested strategy for missing ratings and a check on whether standardizing ratings helps.
Assignment 1 — Loading and Transforming English Premier League Match Data
Loading 25 seasons of EPL match data into R, selecting and renaming columns, decoding the full-time result target variable, and checking a COVID-era home-advantage study against the cleaned data.