Make-up 2026
You will tackle 4 projects on publicly available datasets. For each one:
-
Download the dataset from the Kaggle link below.
-
Run a quick EDA: variable types, distributions, missing values, first visualisations. Conclude in writing on the nature of the problem and the metric you will use.
-
Choose the approach you find most appropriate — it is up to you to decide after the EDA.
-
Implement your solution, evaluate it honestly (cross-validation, held-out test set…), and comment on your methodological choices.
Submission format: a single Jupyter notebook (.ipynb) containing all 4 projects in clearly identified sections. The notebook must run end-to-end without errors.
1. Used-car classifieds
You receive a dump from a major US used-car classifieds website (~400 000 listings, ~25 columns). For each listing you have the manufacturer, model, year, mileage, declared condition, fuel type, transmission, paint colour, the US state of the listing, and the asking price.
Listings are hand-typed: many missing values, outlier prices (0 …), very high-cardinality categorical columns, and possible duplicates.
It is your job to decide what to predict and how to encode the categorical columns cleanly before modelling.
🔗 Kaggle dataset — Used Cars Dataset (Craigslist)
2. Spaceship Titanic — missing passengers
The starship Spaceship Titanic was carrying around 8 700 passengers to three habitable exoplanets when it hit a spatial anomaly.
You are given travel records: cabin, home planet, age, on-board spending across the ship's amenities, and a target column
Transported. Several columns contain missing values, and passenger full names hint at family groupings you may want to exploit.
🔗 Kaggle dataset — Spaceship Titanic
3. Kuzushiji-MNIST — ancient cursive characters
KMNIST is a set of 70 000 grayscale 28×28 images of cursive Japanese characters (kuzushiji) extracted from historical books of the Edo period.
Like MNIST, there are 10 classes (ten different hiragana), but the visual difficulty is higher: the same class can have very different strokes, and some characters look similar.
Adopt a workflow similar to what you used on plain MNIST, adapting it as needed.
🔗 Kaggle dataset — Kuzushiji-MNIST
4. Alien vs Predator — who is who?
A database of movie stills mixes photographs from the two sci-fi franchises Alien and Predator. The set contains only a few hundred colour images, organised in two folders (
alien and predator), with a train/validation split already provided.
The challenge: train a strong visual classifier from a very small amount of data.
🔗 Kaggle dataset — Alien vs Predator Images
General advice
- The EDA drives everything: take the time to understand the data before coding a model.
- Document your assumptions, your preprocessing choices and your metric.
- A well-evaluated simple baseline beats a hastily built complex model.
- You may use Kaggle Notebooks to benefit from the environment and the GPU, then download the final
and upload it here..ipynb