Make-up

Make-up 2026

You will tackle 4 projects on publicly available datasets. For each one:

  1. Download the dataset from the Kaggle link below.

  2. Run a quick EDA: variable types, distributions, missing values, first visualisations. Conclude in writing on the nature of the problem and the metric you will use.

  3. Choose the approach you find most appropriate — it is up to you to decide after the EDA.

  4. Implement your solution, evaluate it honestly (cross-validation, held-out test set…), and comment on your methodological choices.

 

Submission format: a single Jupyter notebook (.ipynb) containing all 4 projects in clearly identified sections. The notebook must run end-to-end without errors.

 


 

1. Used-car classifieds

You receive a dump from a major US used-car classifieds website (~400 000 listings, ~25 columns). For each listing you have the manufacturer, model, year, mileage, declared condition, fuel type, transmission, paint colour, the US state of the listing, and the asking price.

Listings are hand-typed: many missing values, outlier prices (0 ,9999999, 9 999 999 …), very high-cardinality categorical columns, and possible duplicates.

It is your job to decide what to predict and how to encode the categorical columns cleanly before modelling.

🔗 Kaggle dataset — Used Cars Dataset (Craigslist)

 


 

2. Spaceship Titanic — missing passengers

The starship Spaceship Titanic was carrying around 8 700 passengers to three habitable exoplanets when it hit a spatial anomaly.

You are given travel records: cabin, home planet, age, on-board spending across the ship's amenities, and a target column

Transported
. Several columns contain missing values, and passenger full names hint at family groupings you may want to exploit.

🔗 Kaggle dataset — Spaceship Titanic

 


 

3. Kuzushiji-MNIST — ancient cursive characters

KMNIST is a set of 70 000 grayscale 28×28 images of cursive Japanese characters (kuzushiji) extracted from historical books of the Edo period.

Like MNIST, there are 10 classes (ten different hiragana), but the visual difficulty is higher: the same class can have very different strokes, and some characters look similar.

Adopt a workflow similar to what you used on plain MNIST, adapting it as needed.

🔗 Kaggle dataset — Kuzushiji-MNIST

 


 

4. Alien vs Predator — who is who?

A database of movie stills mixes photographs from the two sci-fi franchises Alien and Predator. The set contains only a few hundred colour images, organised in two folders (

alien
and
predator
), with a train/validation split already provided.

The challenge: train a strong visual classifier from a very small amount of data.

🔗 Kaggle dataset — Alien vs Predator Images

 


 

General advice

  • The EDA drives everything: take the time to understand the data before coding a model.
  • Document your assumptions, your preprocessing choices and your metric.
  • A well-evaluated simple baseline beats a hastily built complex model.
  • You may use Kaggle Notebooks to benefit from the environment and the GPU, then download the final
    .ipynb
    and upload it here.