Published

Machine Learning · Class Kaggle Competition

Emotion Recognition

A reproducible four-class emotion classification pipeline designed around subject-level generalization, not just a strong validation score.

PythonPyTorchscikit-learnPandasNumPy
Year
2026
Type
Machine learning competition
Role
ML pipeline design & implementation
Architecture view

Subject-aware inference pipeline

The same preprocessing contract is preserved from cross-validation through final inference.

01 / Context

What I built
and why.

Overview

This project predicts one of four emotion regions on the valence–arousal plane from a 1,793-dimensional feature vector. The complete workflow—from cleaning and validation to ensembling and submission generation—lives in one reproducible training pipeline.

The challenge

A random split can look convincing while still failing on people the model has never seen. The central engineering problem was therefore not only classification performance, but building an evaluation setup that made subject-level generalization visible.

The approach

I paired label-balanced StratifiedKFold evaluation with GroupKFold splits based on person_id. Inputs are cleaned, clipped and standardized inside each training fold. Two complementary MLP architectures are trained across multiple seeds, with class-aware sampling, regularization and probability calibration used to make inference more stable.

02 / Engineering

Key decisions

The choices that shaped the system—not only the technologies that appear in it.

01

Validate the real failure mode

GroupKFold keeps samples from the same subject together, providing a stricter view of performance on unseen people.

02

Keep preprocessing fold-safe

Scaling is fit only on the training side of each fold to prevent information from leaking into validation.

03

Reduce prediction variance

Multiple seeds, EMA weights and probability renormalization create a more stable final ensemble than a single training run.

Technical toolkit

Built with

PythonPyTorchscikit-learnPandasNumPy
03 / Reflection

What the project taught me

  • A validation strategy is part of the model, not a final reporting step.
  • Reproducible experiments make architecture and calibration choices easier to compare.
  • Subject-aware evaluation can reveal risks hidden by otherwise strong aggregate scores.

This page uses a system diagram to document the current technical structure. Product screenshots and experiment outputs can be added here as each project evolves.