Make-upMake-up 2026
You will tackle **4 projects** on publicly available datasets. For each one:
1. Download the dataset from the Kaggle link below.
2. Run a **quick EDA**: variable types, distributions, missing values, first visualisations. Conclude in writing on the nature of the problem and the metric you will use.
3. Choose the approach you find most appropriate — it is up to you to decide after the EDA.
4. Implement your solution, evaluate it honestly (cross-validation, held-out test set…), and comment on your methodological choices.
**Submission format:** a **single Jupyter notebook (.ipynb)** containing all 4 projects in clearly identified sections. The notebook must run end-to-end without errors.
---
### 1. Used-car classifieds
You receive a dump from a major US used-car classifieds website (~400 000 listings, ~25 columns). For each listing you have the manufacturer, model, year, mileage, declared condition, fuel type, transmission, paint colour, the US state of the listing, and the asking price.
Listings are hand-typed: many missing values, outlier prices (0 $, 9 999 999 $…), very high-cardinality categorical columns, and possible duplicates.
It is your job to decide what to predict and how to encode the categorical columns cleanly before modelling.
🔗 [Kaggle dataset — Used Cars Dataset (Craigslist)](https://www.kaggle.com/datasets/austinreese/craigslist-carstrucks-data)
---
### 2. Spaceship Titanic — missing passengers
The starship *Spaceship Titanic* was carrying around 8 700 passengers to three habitable exoplanets when it hit a spatial anomaly.
You are given travel records: cabin, home planet, age, on-board spending across the ship's amenities, and a target column `Transported`. Several columns contain missing values, and passenger full names hint at family groupings you may want to exploit.
🔗 [Kaggle dataset — Spaceship Titanic](https://www.kaggle.com/competitions/spaceship-titanic/data)
---
### 3. Kuzushiji-MNIST — ancient cursive characters
KMNIST is a set of 70 000 grayscale 28×28 images of cursive Japanese characters (*kuzushiji*) extracted from historical books of the Edo period.
Like MNIST, there are 10 classes (ten different hiragana), but the visual difficulty is higher: the same class can have very different strokes, and some characters look similar.
Adopt a workflow similar to what you used on plain MNIST, adapting it as needed.
🔗 [Kaggle dataset — Kuzushiji-MNIST](https://www.kaggle.com/datasets/anokas/kuzushiji)
---
### 4. Alien vs Predator — who is who?
A database of movie stills mixes photographs from the two sci-fi franchises *Alien* and *Predator*. The set contains only a few hundred colour images, organised in two folders (`alien` and `predator`), with a train/validation split already provided.
The challenge: train a strong visual classifier from a very small amount of data.
🔗 [Kaggle dataset — Alien vs Predator Images](https://www.kaggle.com/datasets/pmigdal/alien-vs-predator-images)
---
**General advice**
- The EDA drives everything: take the time to understand the data before coding a model.
- Document your assumptions, your preprocessing choices and your metric.
- A well-evaluated simple baseline beats a hastily built complex model.
- You may use Kaggle Notebooks to benefit from the environment and the GPU, then download the final `.ipynb` and upload it here.
max 50 MB · types : .ipynb