This project predicts one of four emotion regions on the valence–arousal plane from a 1,793-dimensional feature vector. The complete workflow—from cleaning and validation to ensembling and submission generation—lives in one reproducible training pipeline.
Machine Learning · Class Kaggle Competition
Emotion Recognition
A reproducible four-class emotion classification pipeline designed around subject-level generalization, not just a strong validation score.
Subject-aware inference pipeline
The same preprocessing contract is preserved from cross-validation through final inference.
What I built
and why.
A random split can look convincing while still failing on people the model has never seen. The central engineering problem was therefore not only classification performance, but building an evaluation setup that made subject-level generalization visible.
I paired label-balanced StratifiedKFold evaluation with GroupKFold splits based on person_id. Inputs are cleaned, clipped and standardized inside each training fold. Two complementary MLP architectures are trained across multiple seeds, with class-aware sampling, regularization and probability calibration used to make inference more stable.
Key decisions
The choices that shaped the system—not only the technologies that appear in it.
Validate the real failure mode
GroupKFold keeps samples from the same subject together, providing a stricter view of performance on unseen people.
Keep preprocessing fold-safe
Scaling is fit only on the training side of each fold to prevent information from leaking into validation.
Reduce prediction variance
Multiple seeds, EMA weights and probability renormalization create a more stable final ensemble than a single training run.
Built with
What the project taught me
- A validation strategy is part of the model, not a final reporting step.
- Reproducible experiments make architecture and calibration choices easier to compare.
- Subject-aware evaluation can reveal risks hidden by otherwise strong aggregate scores.
This page uses a system diagram to document the current technical structure. Product screenshots and experiment outputs can be added here as each project evolves.