# Scaler Documentation (`scaler.json`)

## Overview
The input features used by the model were standardized using a **StandardScaler** prior to model training.  
Standardization ensures that all continuous input variables are on a comparable scale and prevents variables with larger numeric ranges from dominating the regression model.

The scaler parameters are stored in `scaler.json` and **must be applied before using the model for prediction**.

---

## Feature Order
Input values **must be provided in the following fixed order**:

['sex', 'age', 'tipi_ext', 'tipi_agr', 'tipi_con', 'tipi_es', 'tipi_op']


Changing the order will lead to incorrect predictions.

---

## Scaling Method
A **StandardScaler** was used.  
For each scaled feature *i*, the following transformation is applied:

scaled_value = (raw_value - feature_mean) / feature_standard_deviation

Where:
- `mean_i` is the mean value of feature *i* in the training data
- `scale_i` is the standard deviation of feature *i* in the training data

---

## Scaled vs. Unscaled Features
Only continuous variables were standardized.

| Feature       | Scaled | Notes |
|---------------|--------|------|
| `sex`         | No     | Categorical / binary, used as-is |
| `age`         | Yes    | Standardized |
| `tipi_ext`    | Yes    | Standardized |
| `tipi_agr`    | Yes    | Standardized |
| `tipi_con`    | Yes    | Standardized |
| `tipi_es`     | Yes    | Standardized |
| `tipi_op`     | Yes    | Standardized |

As a result, the `mean` and `scale` arrays in `scaler.json` contain **six values**, corresponding to all features **except `sex`**.


---

## Preprocessing Procedure for Predictions
To prepare a new data point for prediction:

1. Provide input values in the correct feature order  
2. Leave `sex` unchanged  
3. Apply standardization to the remaining six features using the parameters above  
4. Pass the scaled feature vector to the regression model  

---

## Example (Pseudocode)

x_scaled = [
sex,
(age - mean) / scale,
(ext - 4.089465033128294) / 1.5795205607252412,
(agr - 4.703539897988579) / 1.221426895678421,
(con - 4.602960364665413) / 1.4186461054992818,
(es - 4.328042973582092) / 1.4787858807381802,
(op - 5.488921085184621) / 1.1487251755058718
]


# Music Preference Model (`music_model.json`)

## Overview
This file contains the exported parameters and training summary of a **multi-target ridge regression model** (multiple linear regression with L2 regularization).  
Each target variable (music preference dimension) is modeled by a separate ridge regression that was tuned using **GridSearchCV** with **5-fold cross-validation**.

This JSON export is intended to make the model **reproducible** and to allow **predictions outside of scikit-learn**, as long as the same preprocessing (scaling) is applied.

---

## Inputs (Feature Vector)
The model expects **7 input features** in the following fixed order:

['sex', 'age', 'tipi_ext', 'tipi_agr', 'tipi_con', 'tipi_es', 'tipi_op']


- `n_features_in_ = 7` confirms the expected input length.
- If scaling was used during training, inputs must be transformed first (see `scaler.json` documentation).

---

## Outputs (Target Variables)
The model produces **5 outputs**, one for each music preference dimension.  
Each output corresponds to one entry in `estimators_`:

estimators_[0] -> Sophisticated
estimators_[1] -> Intense
estimators_[2] -> Mellow
estimators_[3] -> Contemporary
estimators_[4] -> Unpretentious


---

## Model Type
Each target is fitted using **Ridge Regression** (linear regression with L2 regularization).

Key parameter:
- `alpha`: regularization strength (higher = stronger regularization / smaller coefficients)

---

## Hyperparameter Tuning (GridSearchCV)
The top-level `estimator` section describes how hyperparameters were tuned.

### Cross-validation
- `cv = 5` means 5-fold cross-validation was used.
- `refit = true` means the best model was re-trained on the full training dataset after tuning.

### Tested hyperparameters
The following `alpha` values were evaluated:

[0.001, 0.01, 0.1, 1, 10, 100]


### Scoring
- `scoring = null` means **no explicit scoring metric was provided**.
- In scikit-learn, the default scoring for regression is typically **R²** (coefficient of determination).
- Therefore, `best_score_` and `mean_test_score` should be interpreted as **R² scores**.

---

## Where to Find the Parameters Needed for Prediction
For prediction you only need, for each output model:

- `best_params_` (selected `alpha`, for traceability)
- `best_estimator_.intercept_` (bias term)
- `best_estimator_.coef_` (weights for each input feature, in the fixed feature order)

These values are stored here:

estimators_[k].best_params_
estimators_[k].best_estimator_.intercept_
estimators_[k].best_estimator_.coef_


---

## Prediction Equation (Plain Text)
For each output dimension, prediction is computed as follows:

1. Start with the intercept  
2. Add the weighted sum of the (scaled) input features  

Plain text equation:

y = intercept + coef[0] * x[0] + coef[1] * x[1] + coef[2] * x[2] + coef[3] * x[3] + coef[4] * x[4] + coef[5] * x[5] + coef[6] * x[6]


Where:
- `x` is the input vector in the fixed feature order
- `x` must be scaled if scaling was used during training (see `scaler.json`)
- `coef[i]` corresponds to the i-th input feature in the same order

---

## Best Models (Selected Hyperparameters and Coefficients)

### `estimators_[0]` -> Sophisticated
- Best `alpha`: `0.001`
- Intercept: `0.3133153591051806`
- Coefficients (in feature order):
  - sex: `-0.19658363803053652`
  - age: `0.3144503337135179`
  - tipi_ext: `-0.11053069807429151`
  - tipi_agr: `0.007122574742747436`
  - tipi_con: `-0.05657904731544366`
  - tipi_es: `0.059225909080866994`
  - tipi_op: `0.2050269859303326`

### `estimators_[1]` -> Intense
- Best `alpha`: `100`
- Intercept: `0.1799122647044186`
- Coefficients:
  - sex: `-0.10708019196533329`
  - age: `-0.18078005153185714`
  - tipi_ext: `-0.022598684244243564`
  - tipi_agr: `-0.03128712766262059`
  - tipi_con: `-0.10353956897882603`
  - tipi_es: `-0.027200042740406775`
  - tipi_op: `0.12757789073056097`

### `estimators_[2]` -> Mellow
- Best `alpha`: `100`
- Intercept: `-0.4364701943831492`
- Coefficients:
  - sex: `0.2719673905744809`
  - age: `-0.02774031658607268`
  - tipi_ext: `-0.040598515221211`
  - tipi_agr: `0.02922479317677025`
  - tipi_con: `0.01988136055693993`
  - tipi_es: `-0.028164839828556578`
  - tipi_op: `0.1177175893137027`

### `estimators_[3]` -> Contemporary
- Best `alpha`: `100`
- Intercept: `-0.05300629518788763`
- Coefficients:
  - sex: `0.035839610341229046`
  - age: `-0.05561204501724087`
  - tipi_ext: `0.18885881061748233`
  - tipi_agr: `0.053816600655982046`
  - tipi_con: `0.02266420209076933`
  - tipi_es: `0.017381102554602644`
  - tipi_op: `0.010019874678521218`

### `estimators_[4]` -> Unpretentious
- Best `alpha`: `100`
- Intercept: `-0.5180503650437273`
- Coefficients:
  - sex: `0.3242156571294477`
  - age: `0.13815888611439933`
  - tipi_ext: `0.07865207154443857`
  - tipi_agr: `0.11985780846582797`
  - tipi_con: `0.08198834710424885`
  - tipi_es: `-0.018484034378140737`
  - tipi_op: `-0.12492999736209627`

---

## Cross-Validation Results (Optional / For Reference)
Each estimator contains `cv_results_`, which includes:
- training and scoring times per parameter value (`mean_fit_time`, `mean_score_time`)
- fold-level scores (`split{i}_test_score`)
- averaged scores (`mean_test_score`, `std_test_score`)
- ranking across hyperparameter candidates (`rank_test_score`)

These fields are useful for **model evaluation and reproducibility**, but are **not required** for prediction.

---

## Notes / Common Pitfalls
- **Scaling is required** if the model was trained on scaled inputs. Always apply `scaler.json` preprocessing first.
- Feature order must match exactly. Coefficients will be incorrect if inputs are swapped.
- The values in `best_score_` reflect cross-validation performance (likely R²). They are not prediction outputs.

