{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# (0) Introduction \n\nIn this notebook we are trying to predict the score \"sii\" which gives an indication of how severe is the addiction of the patient to internet use. Where 0 stands for None, 1 for Mild, 2 for Moderate, and 3 for Severe. \n\nWe have two data sources, dictionary data on tests and surveys. For some individuals, we also have time series with sensor data from a wrist watch.\n\nThe steps to process the data were: \n\n1. Data exploration and cleaning\n2. Feature engineering\n3. Model selection\n4. Prediction and submission\n\nThe models used were: \n- [SGD: Stochastic gradient descent](https://scikit-learn.org/1.5/modules/sgd.html)\n- [XGB: XGBoost](https://xgboost.readthedocs.io/en/stable/)\n- [LGBM: Light GBM](https://lightgbm.readthedocs.io/en/stable/)\n- [RF: Random Forests](https://scikit-learn.org/1.5/modules/generated/sklearn.ensemble.RandomForestClassifier.html)\n- [CB: CatBoost](https://catboost.ai/)  \n\nThe results are summarized below: \n\n| Model   | Feat. Eng. | Avg. Quadratic kappa | \n| ------- | ---------- | -------------------- | \n| SGD     | No         | 0.0249               | \n| XGB     | No         | 0.8938               | \n| LGBM    | No         | 0.8980               | \n| RF      | No         | 0.8933               | \n| CB      | No         | 0.8833               | \n|   Public score       | 0.127                |\n\n| Model   | Feat. Eng. | Avg. Quadratic kappa | \n| ------- | ---------- | -------------------- | \n| SGD     | Yes        | 0.0249               | \n| XGB     | Yes        | 0.8932               | \n| LGBM    | Yes        | 0.8936               | \n| RF      | Yes        | 0.8885               | \n| CB      | Yes        | 0.8803               | \n|   Public score       | **0.129**                |\n\nWhere the Avg. Quadratic Kappa is the result of the cross-validation using the [Quadratic Kappa score](https://datatab.net/tutorial/weighted-cohens-kappa), which is the metric used to score this competition. \n\nThe weights for the ensemble were chosen based on the relative performance of each model when the Feature Engineering setting was enabled. Despite the results without feature engineering showing superior performance on the provided test data, incorporating feature engineering yielded slightly better outcomes on the public score.\n\nFuture work will explore two-step models that initially calculate the probability of scoring zero or non-zero. Subsequently, those identified as non-zero will be further classified. This approach will aggregate non-zero data points, effectively reducing class imbalance in the initial model phase. Another area of focus will be to enhance the resources and options available for parameter tuning of models like XGBoost and LightGBM, following the strategies employed in high-ranking solutions. Additionally, integrating neural network models such as Pytorch.TabNet into our analysis could provide significant benefits. Generally, the use of ensemble methods, which featured prominently in many high-ranking solutions, will continue to be a key strategy. This method of prediction should not only be maintained but potentially expanded to leverage its success in achieving high quadratic kappa scores.","metadata":{}},{"cell_type":"markdown","source":"## (0.a.) Package loading ","metadata":{}},{"cell_type":"code","source":"# Packages\nimport os\nimport pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nfrom pandas.core.common import random_state\nfrom imblearn.over_sampling import SMOTE\nfrom copy import copy\npd.set_option(\"display.max_rows\", None)\n\n# Parallel processing library\nfrom joblib import Parallel, delayed\nfrom concurrent.futures import ThreadPoolExecutor, ProcessPoolExecutor\nfrom tqdm import tqdm\n\n# SKLearn\nfrom sklearn.linear_model import LinearRegression\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import accuracy_score, confusion_matrix \nfrom sklearn.metrics import classification_report, average_precision_score, cohen_kappa_score, make_scorer\nfrom sklearn.model_selection import train_test_split, KFold, cross_val_score\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.impute import SimpleImputer\n\n# Other models \nimport xgboost as xgb\nimport lightgbm as lgb\nfrom catboost import CatBoostClassifier","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-28T00:17:54.397177Z","iopub.execute_input":"2024-12-28T00:17:54.397654Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (1) Data exploration and cleaning \n","metadata":{}},{"cell_type":"markdown","source":"## (1.a) Data dictionary\n\nThe tabular data in train.csv and test.csv comprises measurements from a variety of instruments. The fields within each instrument are described in data_dictionary.csv. These instruments are:\n- Demographics - Information about age and sex of participants. \n- Internet Use - Number of hours of using computer/internet per day.\n- Children's Global Assessment Scale - Numeric scale used by mental health clinicians to rate the general functioning of youths under the age of 18.\n- Physical Measures - Collection of blood pressure, heart rate, height, weight and waist, and hip measurements.\n- FitnessGram Vitals and Treadmill - Measurements of cardiovascular fitness assessed using the NHANES treadmill protocol.\n- FitnessGram Child - Health related physical fitness assessment measuring five different parameters including aerobic capacity, muscular strength, muscular endurance, flexibility, and body composition.\n- Bio-electric Impedance Analysis - Measure of key body composition elements, including BMI, fat, muscle, and water content.\n- Physical Activity Questionnaire - Information about children's participation in vigorous activities over the last 7 days.\n- Sleep Disturbance Scale - Scale to categorize sleep disorders in children.\n- Actigraphy - Objective measure of ecological physical activity through a research-grade biotracker.\n- Parent-Child Internet Addiction Test - 20-item scale that measures characteristics and behaviors associated with compulsive use of the Internet including compulsivity, escapism, and dependency.\n\nThe target ``sii`` for this competition is derived from the field ``PCIAT-PCIAT_Total`` as described in the data dictionary: 0 for None, 1 for Mild, 2 for Moderate, and 3 for Severe. Additionally, each participant has been assigned a unique identifier id.","metadata":{}},{"cell_type":"code","source":"# Data dictionary\ntrainDData = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntestDData = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### (1.a.1) Data profiling \n\nFor a better analysis, the package DataProfiling was used. But since it takes a lot of time to generate, the file has been saved as ``EDA_before.html`` and added as an attachement. \n\nWe will eliminate the columns PCIAT, because there are not in the test set and are summarised in the value ``sii`` which is our response variable as well. ","metadata":{}},{"cell_type":"code","source":"# First, we will get rid of all the columns that are not in the test set. \nnot_in_test_set = []\nfor col in trainDData.columns:\n    if col not in list(testDData.columns) and col !=\"sii\":\n        not_in_test_set.append(col)\n\ntrainDData = trainDData.drop(not_in_test_set, axis=1, inplace=False)\nprint(not_in_test_set)\ntrainDData.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Dropping NaNs in the response variable\nprint(\"Registries before removing the unknown sii\", len(trainDData.index))\ntrainDData_nona = trainDData.dropna(axis=0, subset=['sii'], inplace=False)\nprint(\"Registries after removing the unknown sii\", len(trainDData_nona.index))\nprint(\"Registries removed\", len(trainDData.index)-len(trainDData_nona.index))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"Registries by value of sii\")\nprint(trainDData_nona.groupby(trainDData_nona[\"sii\"])[\"id\"].count())\ntrainDData_nona[\"sii\"].plot.hist()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Removing the 999 occurrences in the \"CGAS-CGAS_Score\"\ncol = \"CGAS-CGAS_Score\"\ntrainDData_nona[col] = trainDData_nona[col].replace(to_replace=999, value=trainDData_nona[col].median(), inplace=False)\ntestDData[col] = testDData[col].replace(to_replace=999, value=trainDData_nona[col].median(), inplace=False)\ntrainDData_nona[col].plot.hist()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Removing the 0 occurrences in the \"Physical-Weight\"\ncol = \"Physical-Weight\"\ntrainDData_nona[col] = trainDData_nona[col].replace(to_replace=0, value=trainDData_nona[col].median(), inplace=False)\ntestDData[col] = testDData[col].replace(to_replace=0, value=trainDData_nona[col].median(), inplace=False)\ntrainDData_nona[col].plot.hist()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Removing the 0 occurrences in the \"Physical-BMI\"\ncol = \"Physical-BMI\"\ntrainDData_nona[col] = trainDData_nona[col].replace(to_replace=0, value=trainDData_nona[col].median(), inplace=False)\ntestDData[col] = testDData[col].replace(to_replace=0, value=trainDData_nona[col].median(), inplace=False)\ntrainDData_nona[col].plot.hist()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### (1.a.b.) Categorical columns\n\nOnly the dictionary data has categorical columns related to the season. For these columns with seasons, we will create less columns if instead of using a onehot encoder, we just convert each season into a number. {\"Spring\": 1, \"Summer\": 2, \"Fall\": 3, \"Winter\" : 4}, if the value is missing, then we add a zero first with an imputer pipeline. ","metadata":{}},{"cell_type":"code","source":"# Check categorical variables \ncat_train = trainDData_nona.select_dtypes(['object'])\ncat_test = testDData.select_dtypes(['object'])\ncat_train.dtypes","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# --------------------------------------------------\n# Pipeline for season categorical values\n# --------------------------------------------------\n\ncategorical_transformer = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='constant', fill_value=0)) # Impute missing values with value zero\n])","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Transform categories for the TRAIN set\n\n# Fit and transform the categorical features (excluding 'id')\ncategorical_cols_train = cat_train.columns.drop('id', errors='ignore') # Drop 'id' if present, otherwise do nothing\nencoded_data_train = categorical_transformer.fit_transform(cat_train[categorical_cols_train])\n\n# Create a new DataFrame from the encoded data\nencoded_cat_train = pd.DataFrame(encoded_data_train, columns = categorical_cols_train)\n\n# Reset index of encoded_cat_train to match cat_train\nencoded_cat_train.index = cat_train.index \n\n# Concatenate with other columns (if needed) and the id column if present.\nif 'id' in cat_train.columns:\n    encoded_cat_train['id'] = cat_train['id']\n    print(\"Registries in encoded cat_train\", len(encoded_cat_train.index))\n    #encoded_cat_train.head()\n\n# Assign the number of the season\nseasonDict = {\"Spring\": 1, \"Summer\": 2, \"Fall\": 3, \"Winter\" : 4}\nfor col in encoded_cat_train.columns:\n    if col != 'id':\n        encoded_cat_train[col] = encoded_cat_train[col].replace(seasonDict, inplace=False)\nencoded_cat_train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Transform categories for the TEST set\n\n# Fit and transform the categorical features (excluding 'id')\ncategorical_cols_test = cat_test.columns.drop('id', errors='ignore') # Drop 'id' if present, otherwise do nothing\nencoded_data_test = categorical_transformer.fit_transform(cat_test[categorical_cols_test])\n\n# Create a new DataFrame from the encoded data\nencoded_cat_test = pd.DataFrame(encoded_data_test, columns = categorical_cols_test)\n\n# Reset index of encoded_cat_test to match cat_test\nencoded_cat_test.index = cat_test.index \n\n# Concatenate with other columns (if needed) and the id column if present.\nif 'id' in cat_test.columns:\n    encoded_cat_test['id'] = cat_test['id']\n    print(\"Registries in encoded cat_test\", len(encoded_cat_test.index))\n    #encoded_cat_test.head()\n\nseasonDict = {\"Spring\": 1, \"Summer\": 2, \"Fall\": 3, \"Winter\" : 4}\nfor col in encoded_cat_test.columns:\n    if col != 'id':\n        encoded_cat_test[col] = encoded_cat_test[col].replace(seasonDict)\nencoded_cat_test.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# We will return the numerical columns\nnum_train = trainDData_nona.select_dtypes(['number'])\nnum_test = testDData.select_dtypes(['number'])\nif 'id' in cat_train.columns:\n    num_train['id'] = cat_train['id']\n    num_test['id'] = cat_test['id']\nprint(num_train.columns)\nprint(\"Registries in the numerical part of the database\",len(num_train.index))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Merge the numerical categories with the *now* numerical seasons \ncat_clean_train = pd.merge(encoded_cat_train, num_train, on='id', how='left')\ncat_clean_test = pd.merge(encoded_cat_test, num_test, on='id', how='left')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (1.b.) Time series data\n\nDuring their participation in the HBN study, some participants were given an accelerometer to wear for up to 30 days continually while at home and going about their regular daily lives.\n\n- series_{train|test}.parquet/id={id} - Series to be used as training data, partitioned by id. Each series is a continuous recording of accelerometer data for a single subject spanning many days.\n-  id - The patient identifier corresponding to the id field in train/test.csv.\n-  step - An integer timestep for each observation within a series.\n-  X, Y, Z - Measure of acceleration, in g, experienced by the wrist-worn watch along each standard axis.\n-  enmo - As calculated and described by the wristpy package, ENMO is the Euclidean Norm Minus One of all accelerometer signals (along each of the x-, y-, and z-axis, measured in g-force) with negative values rounded to zero. Zero values are indicative of periods of no motion. While no standard measure of acceleration exists in this space, this is one of the several commonly computed features.\n-  anglez - As calculated and described by the wristpy package, Angle-Z is a metric derived from individual accelerometer components and refers to the angle of the arm relative to the horizontal plane.\n-  non-wear_flag - A flag (0: watch is being worn, 1: the watch is not worn) to help determine periods when the watch has been removed, based on the GGIR definition, which uses the standard deviation and range of the accelerometer data.\n-  light - Measure of ambient light in lux. See ​​here for details.\n-  battery_voltage - A measure of the battery voltage in mV.\n-  time_of_day - Time of day representing the start of a 5s window that the data has been sampled over, with format %H:%M:%S.%9f.\n-  weekday - The day of the week, coded as an integer with 1 being Monday and 7 being Sunday.\n-  quarter - The quarter of the year, an integer from 1 to 4.\n-  relative_date_PCIAT - The number of days (integer) since the PCIAT test was administered (negative days indicate that the actigraphy data has been collected before the test was administered).","metadata":{}},{"cell_type":"markdown","source":"### (1.b.1.) Data profiling\n\nThere are a total of 966 registries with Time Series Data. However, these files are heavy (> 19 GB) and data profiling for a combined DataFrame is not possible in reasonable time (or resources provided by Kaggle). Instead, the data profiling of a couple of study participants were analysed. The reports can be found in the files ``'id001f3379.html'`` and ``'id00115b9f.html'``. \n\nAfter analysing the reports, arrived to the following conclusions: \n\n- The column ``step`` is not required, as it only serves as an id for the time step and steps are not homogeneous.\n- The data should not be considered when the value of ``non-wear_flag`` is equal to $1$. \n- The data can provide insights through statistical measures (mean, median, standard deviation, min and max) of each feature, for each participant in such a way that each id has only one entry in the data.\n - The feature ``time_of_day`` can give insights in wake-up and sleep time via its minimum and maximum value, respectively.\n - The feature ``weekday`` can be used to aggregate the data into columns and calculate statistics per day of the week. To address potential seasonality. \n\nThese aspects are considered when loading the files in the function ``custom_summary``","metadata":{}},{"cell_type":"code","source":"# Group data by weekday and calculate statistics. \n# Most people only do the test for one season, so we will store this information in one column\ndef custom_summary(x):\n    d = {}\n    vals,counts = np.unique(x['quarter'], return_counts=True)\n    index = np.argmax(counts)\n    d['quarter'] = vals[index]\n    # Check for seasons and add corresponding columns\n    for day in range(1,8):  # Iterate through all weekdays (1-7)\n    \n      d[f'percentage_wear_{day}'] = 1 - np.mean(x[(x['weekday'] == day)]['non-wear_flag'])\n    \n      # Filter data where 'non-wear_flag' is 0, and normalise for the quarter and the weekday\n      filtered_x = x[(x['weekday'] == day) & (x['non-wear_flag'] == 0)]\n      if not filtered_x.empty:\n        for col in ['X', 'Y', 'Z', 'anglez','battery_voltage']:\n          d[f'{col}_mean_{day}'] = np.mean(filtered_x[col])\n          d[f'{col}_std_{day}'] = np.std(filtered_x[col])\n          d[f'{col}_median_{day}'] = np.median(filtered_x[col])\n          d[f'{col}_min_{day}'] = np.min(filtered_x[col])\n          d[f'{col}_max_{day}'] = np.max(filtered_x[col])\n    \n        for col in ['time_of_day']:\n          # This can help us capture some sleeping and waking up patterns\n          d[f'{col}_min_{day}'] = np.min(filtered_x[col])\n          d[f'{col}_max_{day}'] = np.max(filtered_x[col])\n    \n        for col in ['enmo', 'light']:\n          d[f'{col}_mean_{day}'] = np.mean(filtered_x[col])\n          d[f'{col}_std_{day}'] = np.std(filtered_x[col])\n          d[f'{col}_median_{day}'] = np.median(filtered_x[col])\n          d[f'{col}_max_{day}'] = np.max(filtered_x[col])\n      else:\n        #print(f\"Warning: No data for quarter {season} and day {day} with non-wear_flag = 0. Skipping calculations.\")\n        # You might want to add placeholder values or handle this differently\n        # For example:\n        for col in ['X', 'Y', 'Z', 'anglez','battery_voltage']:\n          d[f'{col}_mean_{day}'] = 9999\n          d[f'{col}_std_{day}'] = 9999\n          d[f'{col}_median_{day}'] = 9999\n          d[f'{col}_min_{day}'] = 9999\n          d[f'{col}_max_{day}'] = 9999\n    \n        for col in ['time_of_day']:\n            # This can help us capture some sleeping and waking up patterns\n            d[f'{col}_min_{day}'] = 9999\n            d[f'{col}_max_{day}'] = 9999\n    \n        for col in ['enmo', 'light']:\n            d[f'{col}_mean_{day}'] = 9999\n            d[f'{col}_std_{day}'] = 9999\n            d[f'{col}_median_{day}'] = 9999\n            d[f'{col}_max_{day}'] = 9999\n    \n    df = pd.Series(d)\n    df = df.to_frame()\n    # Comment following line if want to see all the columns in the resulting dataframe\n    df = df.T\n                      \n    return df\n\ndef loadParquetFile(directory, fileName):\n    \"\"\"\n    read parquet file \n    \"\"\"\n    path = os.path.join(directory, fileName, \"part-0.parquet\")\n    df = pd.read_parquet(path)\n    statDF = custom_summary(df)\n    return statDF, fileName.split(\"=\")[1] # get ids\n    \n\ndef loadTimeSeriesData(directory):\n    \"\"\"\n    input : root direction\n    \"\"\"\n    filesIds = os.listdir(directory) # get list of folder name (files ids)\n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: loadParquetFile(directory, fname), filesIds), total=len(filesIds)))\n    statistic, ids = zip(*results) # pack into Statistic and Ids \n    data = pd.concat(statistic, axis=0)\n    data.reset_index(inplace=True, drop=True)\n    data[\"id\"] = ids # add ids into dataframe\n    return data","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\ntrainTSData = loadTimeSeriesData((\"/kaggle/input/child-mind-institute-problematic-internet-use/\" + \"series_train.parquet\"))\ntestTSData = loadTimeSeriesData((\"/kaggle/input/child-mind-institute-problematic-internet-use/\" +\"series_test.parquet\"))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"trainTSData.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"testTSData.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"for col in trainTSData:\n    if col != 'id':\n        trainTSData[col] = trainTSData[col].replace(to_replace=9999, value=trainTSData[col].median(), inplace=False)\n        testTSData[col] = testTSData[col].replace(to_replace=9999, value=trainTSData[col].median(), inplace=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (1.c.) Combine the data","metadata":{}},{"cell_type":"code","source":"combinedTrainDF =  pd.merge(cat_clean_train, trainTSData, how=\"left\", on='id')\ncombinedTestDF = pd.merge(cat_clean_test, testTSData, how=\"left\", on='id')","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"num_train = combinedTrainDF.select_dtypes(['number'])\nnum_test = combinedTestDF.select_dtypes(['number'])\nprint(num_train.columns)\nprint(\"Registries in the numerical part of the database\",len(num_train.index))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------\n# Pipeline for numerical values\n# ---------------------------------\nnumerical_transformer = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='median')),\n])\n\n# Fit and transform the numerical features (excluding 'id')\nnumerical_cols = num_train.columns.drop('id', errors='ignore')\nimputed_data_train = numerical_transformer.fit_transform(num_train[numerical_cols]) # .toarray() removed as it's not needed for SimpleImputer\n\n# Create a new DataFrame from the imputed data\nimputed_num_train = pd.DataFrame(imputed_data_train, columns=numerical_cols) # Use numerical_cols directly\n\n# Reset index of imputed_num_train to match num_train\nimputed_num_train.index = cat_train.index\n\n# Concatenate with other columns (if needed) and the id column if present.\nif 'id' in cat_train.columns:\n    imputed_num_train['id'] = cat_train['id']\n    print(\"Registries imputed num_train\", len(imputed_num_train.index))\n\nimputed_num_train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ---------------------------------\n# Pipeline for numerical values\n# ---------------------------------\nnumerical_transformer = Pipeline(steps=[\n    ('imputer', SimpleImputer(strategy='most_frequent')),\n])\n\n# Fit and transform the numerical features (excluding 'id')\nnumerical_cols = num_test.columns.drop('id', errors='ignore')\nimputed_data_test = numerical_transformer.fit_transform(num_test[numerical_cols]) # .toarray() removed as it's not needed for SimpleImputer\n\n# Create a new DataFrame from the imputed data\nimputed_num_test = pd.DataFrame(imputed_data_test, columns=numerical_cols) # Use numerical_cols directly\n\n# Reset index of imputed_num_train to match num_train\nimputed_num_test.index = cat_test.index\n\n# Concatenate with other columns (if needed) and the id column if present.\nif 'id' in cat_test.columns:\n    imputed_num_test['id'] = cat_test['id']\n    print(\"Registries imputed num_test\", len(imputed_num_test.index))\n\nimputed_num_test.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Concatenate encoded_cat_train and imputed_num_train by the 'id' column\nfinal_train =  copy(imputed_num_train)\nfinal_test = copy(imputed_num_test)\nprint(\"Number of registries after pipeline:\", len(final_train.index))\nfinal_train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (2) Feature engineering\n\nFor the feature engineering we will try to add thresholds of how critical some metrics such as $BMI$ are, depending on sex and age of the participant. The new most relevant features are: \n\n- \"BMI-Classify\": metric that takes values of 0-4 depending on the age and sex of the participant. Where 0 = Severe thiness, 1 = Thinness, 2 = Normal, 3 = Overweight, 4 = Obesity.\n- \"Waist-Height-Ratio-Classify\": metric that takes values of 0-4 depending on the age and sex of the participant. Where 0 = Extremely slim, 1 = Slim, 2 = Healthy, 3 = Overweight, 4 = Very overweight, 5 = Obese.\n- \"FATPercent-Classify\": metric that takes values of 0-4 depending on the sex of the participant. Where 0 = Essential fat, 1 = Athletes, 2 = Fitness, 3 = Acceptable, 4 = Not-acceptable.\n\nOther metrics were created by multiplying or dividing two features.     ","metadata":{}},{"cell_type":"code","source":"# Kudos to Johnson Chong for figuring out the threshold values from the WHO website. https://www.kaggle.com/code/johnsonhk88/cmi-problematic-internet-use-ml#Model-Prediction-for-Submission\ndef getBMIInfo(df):\n    \"\"\"\n    classify Weight caterogy by BMI reference from World Health Organization  \n    \"\"\"\n    # print(df)\n    bmi = df[\"Physical-BMI\"]\n    age = df[\"Basic_Demos-Age\"]\n    gender = df[\"Basic_Demos-Sex\"]\n    bodyweightClass = 0\n    if df[\"Basic_Demos-Sex\"] == 0: # for female\n        if age < 8 :\n            adaptThreshold1 = 12.0\n            adaptThreshold2 = 13.0\n            adaptThreshold3 = 17.7\n            adaptThreshold4 = 0.5*(age-5) + 19.0 \n            if bmi < adaptThreshold1:\n                bodyweightClass = 0 # Severe thinness\n            elif bmi <adaptThreshold2:\n                bodyweightClass = 1 # Thiness\n            elif bmi <adaptThreshold3:\n                bodyweightClass = 2 # Normal \n            elif bmi <adaptThreshold4:\n                 bodyweightClass = 3 # Overweight \n            else :\n                 bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n            \n        elif age < 15:\n            # apply linear regression \n            adaptThreshold1 = 0.357 * (age-8) + 12.0\n            adaptThreshold2 = 0.429 * (age-8) + 13.0\n            adaptThreshold3 = 0.828 * (age-8) + 17.7\n            adaptThreshold4 = 1.1*(age-8) + 20.5 \n            if bmi < adaptThreshold1:\n                bodyweightClass = 0 # Severe thinness\n            elif bmi <adaptThreshold2:\n                bodyweightClass = 1 # Thinness\n            elif bmi <adaptThreshold3:\n                bodyweightClass = 2 # Normal \n            elif bmi <adaptThreshold4:\n                 bodyweightClass = 3 # Overweight \n            else :\n                 bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n        \n        elif age < 19:\n            # apply linear regression \n            adaptThreshold1 = 14.7\n            adaptThreshold2 = 16.5\n            adaptThreshold3 = 0.5 * (age-15) + 23.5\n            adaptThreshold4 = 0.325 *(age-15) + 28.2 \n            if bmi < adaptThreshold1:\n                bodyweightClass = 0 # Severe thinness\n            elif bmi <adaptThreshold2:\n                bodyweightClass = 1 # Thinness\n            elif bmi <adaptThreshold3:\n                bodyweightClass = 2 # Normal \n            elif bmi <adaptThreshold4:\n                 bodyweightClass = 3 # Overweight \n            else :\n                 bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n        else: \n            #use adult\n            if bmi < 18.5: \n                bodyweightClass = 0 # Thinness \n            elif bmi < 24.9:\n                bodyweightClass = 2 # Normal \n            elif bmi < 29.9:\n                 bodyweightClass = 3 # Overweight \n            else:\n                bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n\n\n    else: # for male\n        if age < 8 :\n            adaptThreshold1 = 12.5\n            adaptThreshold2 = 13.2\n            adaptThreshold3 = 0.266 *(age-5) + 16.7\n            adaptThreshold4 = 0.5*(age-5) + 18.2 \n            if bmi < adaptThreshold1:\n                bodyweightClass = 0 # Severe thinness\n            elif bmi <adaptThreshold2:\n                bodyweightClass = 1 # Thinness \n            elif bmi <adaptThreshold3:\n                bodyweightClass = 2 # normal \n            elif bmi <adaptThreshold4:\n                 bodyweightClass = 3 # Overweight \n            else :\n                 bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n            \n        elif age < 15:\n            # apply linear regression \n            adaptThreshold1 = 0.314 * (age-8) + 12.5\n            adaptThreshold2 = 0.4 * (age-8) + 13.2\n            adaptThreshold3 = 0.742 * (age-8) + 17.5\n            adaptThreshold4 = 1.04 *(age-8) + 19.7 \n            if bmi < adaptThreshold1:\n                bodyweightClass = 0 # Severe thinness\n            elif bmi <adaptThreshold2:\n                bodyweightClass = 1 # Thinness \n            elif bmi <adaptThreshold3:\n                bodyweightClass = 2 # Normal \n            elif bmi <adaptThreshold4:\n                 bodyweightClass = 3 # Overweight \n            else :\n                 bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n        \n        elif age < 19:\n            # apply linear regression \n            adaptThreshold1 = 15.7\n            adaptThreshold2 = 0.375 * (age-15) + 16.0\n            adaptThreshold3 = 0.7 * (age-15) + 22.7\n            adaptThreshold4 = 0.675 *(age-15) + 27.0\n            if bmi < adaptThreshold1:\n                bodyweightClass = 0 # Severe thinness\n            elif bmi <adaptThreshold2:\n                bodyweightClass = 1 # Thinness\n            elif bmi <adaptThreshold3:\n                bodyweightClass = 2 # Normal\n            elif bmi <adaptThreshold4:\n                 bodyweightClass = 3 # Overweight \n            else :\n                 bodyweightClass = 4 # Obesity\n\n            return bodyweightClass\n        else: \n            #use adult\n            if bmi < 18.5: \n                bodyweightClass = 0 # Thinness \n            elif bmi < 24.9:\n                bodyweightClass = 2 # Normal \n            elif bmi < 29.9:\n                 bodyweightClass = 3 # Overweight \n            else:\n                bodyweightClass = 4 # Obesity\n            return bodyweightClass","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def getWaistHeightInfo(df):\n    age = df[\"Basic_Demos-Age\"]\n    gender = df[\"Basic_Demos-Sex\"]\n    ratio = df[\"Physical-Waist_Circumference\"] / df[\"Physical-Height\"]\n    ratioClass = 0\n    if age < 15:\n        if ratio < 0.34:\n            ratioClass =0 # Extremely Slim\n        elif ratio < 0.45:\n            ratioClass  = 1 # Slim\n        elif ratio < 0.51:\n            ratioClass = 2 # Healthy\n        elif ratio < 0.63:\n            ratioClass = 3 # Over weight\n        else:\n            ratioClass = 5 # Obese\n        return ratioClass \n\n    else : # for adult \n        if gender == 0: # for female\n            if ratio < 0.34:\n                ratioClass =0 # Extremely slim\n            elif ratio < 0.41:\n                ratioClass  = 1 # Slim\n            elif ratio < 0.48:\n                ratioClass = 2 # Healthy\n            elif ratio < 0.53:\n                ratioClass = 3 # Over weight\n            elif ratio < 0.57:\n                ratioClass = 4 # Very overweight\n            else:\n                ratioClass = 5 # Obese\n            return ratioClass\n        else : # for male\n            if ratio < 0.34:\n                ratioClass =0 # Extremely Slim\n            elif ratio < 0.42:\n                ratioClass  = 1 # Slim\n            elif ratio < 0.52:\n                ratioClass = 2 # Healthy\n            elif ratio < 0.57:\n                ratioClass = 3 # Over weight\n            elif ratio < 0.62:\n                ratioClass = 4 # Very Overweight\n            else:\n                ratioClass = 5 # Obese\n            return ratioClass\n\ndef getBodyFatPercentInfo(df):\n    \"\"\"\n    classify the user health status by body fat percentage\n    \"\"\"\n    gender = df[\"Basic_Demos-Sex\"]\n    fatPercent = df[\"BIA-BIA_Fat\"]\n    fatclass =0\n    if gender == 0: # for female\n        if fatPercent < 13:  \n            fatclass = 0 # Essential Fat\n        elif fatPercent < 20:\n            fatclass = 1 # Athletes\n        elif fatPercent < 24:\n            fatclass = 2 # Fitness\n        elif fatPercent < 31:\n            fatclass = 3 # Acceptable \n        else:\n            fatclass = 4 # Not-acceptable \n        return fatclass\n    else: # for male\n        if fatPercent < 5:  \n            fatclass = 0 # Essential Fat\n        elif fatPercent < 13:\n            fatclass = 1 # Athletes\n        elif fatPercent < 17:\n            fatclass = 2 # Fitness\n        elif fatPercent < 24:\n            fatclass = 3 # Acceptable \n        else:\n            fatclass = 4 # Not-acceptable\n        return fatclass","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def featureEngineering(df):\n    df[\"BMI-AGE\"] = df[\"Physical-BMI\"] * df[\"Basic_Demos-Age\"]\n    df[\"BMI-Classify\"]= df.apply(getBMIInfo, axis=1)\n    df[\"Waist-Height-Ratio\"] = df[\"Physical-Waist_Circumference\"] / df[\"Physical-Height\"]\n    df[\"Waist-Height-Ratio-Classify\"] = df.apply(getWaistHeightInfo, axis=1) \n    df['Internet-Hours-Age'] = df['PreInt_EduHx-computerinternet_hoursday'] * df['Basic_Demos-Age']\n    df['BMI-Internet-Hours'] = df['Physical-BMI'] * df['PreInt_EduHx-computerinternet_hoursday']\n    df[\"FATPercent-Classify\"]= df.apply(getBodyFatPercentInfo, axis= 1)\n    df['ICW_TBW'] = df['BIA-BIA_ICW'] / df['BIA-BIA_TBW']\n    df['FFMI_BFP'] = df['BIA-BIA_FFMI'] / df['BIA-BIA_Fat']   \n    df['FMI_BFP'] = df['BIA-BIA_FMI'] / df['BIA-BIA_Fat']\n    df['LST_TBW'] = df['BIA-BIA_LST'] / df['BIA-BIA_TBW']\n    df['BFP_BMR'] = df['BIA-BIA_Fat'] * df['BIA-BIA_BMR']\n    df['BFP_DEE'] = df['BIA-BIA_Fat'] * df['BIA-BIA_DEE']\n    df['BMR_Weight'] = df['BIA-BIA_BMR'] / df['Physical-Weight']\n    df['BMR_Weight'] =df['BMR_Weight'].fillna(0)\n    df['DEE_Weight'] = df['BIA-BIA_DEE'] / df['Physical-Weight']\n    df['DEE_Weight'] = df['DEE_Weight'].fillna(0)\n    df['SMM_Height'] = df['BIA-BIA_SMM'] / df['Physical-Height']\n    df['Muscle_to_Fat'] = df['BIA-BIA_SMM'] / df['BIA-BIA_FMI']\n    df['Hydration_Status'] = df['BIA-BIA_TBW'] / df['Physical-Weight']\n    df['BFP_BMI'] = df['BIA-BIA_Fat'] / df['BIA-BIA_BMI']\n    \n    return df","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Apply the functions\nfinal_train = featureEngineering(final_train)\nfinal_test = featureEngineering(final_test)\n# Example of engineered feature\nfinal_train['BMI-Classify'].plot.hist()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (3) Model selection \n\nThe models used were: \n- [SGD: Stochastic gradient descent](https://scikit-learn.org/1.5/modules/sgd.html)\n- [XGB: XGBoost](https://xgboost.readthedocs.io/en/stable/)\n- [LGBM: Light GBM](https://lightgbm.readthedocs.io/en/stable/)\n- [RF: Random Forests](https://scikit-learn.org/1.5/modules/generated/sklearn.ensemble.RandomForestClassifier.html)\n- [CB: CatBoost](https://catboost.ai/)  \n\nWe will compare the models using a K-fold cross validation to compare scores. We will repeat the process with a parameter grid to get the best fitting parameters for each model. But before the fitting, we will resample the data to mitigate the effect of unbalanced prediction categories. ","metadata":{}},{"cell_type":"code","source":"# Get the labels for later use and remove the id from the final_train\nif 'id' in final_train.columns:\n    labels = final_train[[\"id\"]]\n    final_train.drop(\"id\", axis=1, inplace=True)\nlabels.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"final_train.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"final_test.head()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Remove the response variable from the x group and create a y variable\nif \"sii\" in final_train.columns:\n    y = final_train[\"sii\"]\n    final_train.drop([\"sii\"], axis = 1, inplace = True)\nx = final_train","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Resampling \nsmote = SMOTE(random_state=42)\nX_train_resampled, y_train_resampled = smote.fit_resample(x, y)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Checking the 'sii' in the resampling\ny_train_resampled.hist().plot()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Splitting dataset into test/train\nX_train, X_test, y_train, y_test = train_test_split(X_train_resampled, y_train_resampled, test_size = 0.20, random_state = 0)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (3.a.) Stochastic Gradient Descent Classifier\n\n","metadata":{}},{"cell_type":"code","source":"# Create the classifier object\nclf = SGDClassifier(loss='log_loss', alpha=0.05, max_iter=10000, random_state=42)\n\n# Train the classifier\nclf.fit(X_train, y_train)\n \n# Make predictions\ny_pred = clf.predict(X_test)\n\n# Perform cross-validation\ncv_scores = cross_val_score(clf, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\nprint(\"Cross-validation scores:\", cv_scores)\nprint(\"Mean Kappa score:\", cv_scores.mean())\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (3.b.) XGBoost","metadata":{}},{"cell_type":"code","source":"%%time\nbest_params = {'learning_rate': 0.1, 'max_depth': 7, 'n_estimators': 200}\n\n# Create the XGBoost classifier\nxgb_model = xgb.XGBClassifier(objective='multi:softmax', num_class=4, random_state=42, **best_params) # num_class set to 4 as 'sii' takes 4 values\n\n# Train the model\nxgb_model.fit(X_train, y_train)\n\n# Make predictions on the validation set\ny_pred = xgb_model.predict(X_test)\n\n# Perform cross-validation\ncv_scores = cross_val_score(xgb_model, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\nprint(\"Cross-validation scores:\", cv_scores)\nprint(\"Mean Kappa score:\", cv_scores.mean())\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"'''\n%%time\n\nbest_kappa_score = 0\nbest_params = {}\n\n\nparam_grid = {\n    'learning_rate': [0.1, 0.2],\n    'max_depth': [7],\n    'n_estimators': [200]\n}\n\nfor lr in param_grid['learning_rate']:\n  for md in param_grid['max_depth']:\n    for ne in param_grid['n_estimators']:\n        print(f\"Testing parameters: learning_rate={lr}, max_depth={md}, n_estimators={ne}\")\n        xgb_model = xgb.XGBClassifier(objective='multi:softmax', num_class=4, random_state=42, learning_rate=lr, max_depth=md, n_estimators=ne)\n        cv_scores = cross_val_score(xgb_model, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\n        kappa_score = cv_scores.mean()\n        print(\"Mean Kappa score:\", cv_scores.mean())\n        if kappa_score > best_kappa_score:\n            best_kappa_score = kappa_score\n            best_params = {'learning_rate': lr, 'max_depth': md, 'n_estimators': ne}\n        print(\"-\"*50)\n\nprint(f\"Best kappa score: {best_kappa_score}\")\nprint(f\"Best parameters: {best_params}\")\n# Train the model with the best parameters\nbest_xgb_model = xgb.XGBClassifier(objective='multi:softmax', num_class=4, random_state=42, **best_params)\nbest_xgb_model.fit(X_train, y_train)\n\n# Make predictions on the validation set\ny_pred = best_xgb_model.predict(X_test)\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()\n'''","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (3.c.) LightGBM","metadata":{}},{"cell_type":"code","source":"%%time\nbest_params = {'learning_rate': 0.1, 'max_depth': 7, 'n_estimators': 200}\n\n# Create the XGBoost classifier\nlgb_model = lgb.LGBMClassifier(objective='multiclass', num_class=4, random_state=42, verbose= -100, verbose_eval = -1, **best_params)\n\n# Train the model\nlgb_model.fit(X_train, y_train)\n\n# Make predictions on the validation set\ny_pred = lgb_model.predict(X_test)\n\n# Perform cross-validation\ncv_scores = cross_val_score(lgb_model, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\nprint(\"Cross-validation scores:\", cv_scores)\nprint(\"Mean Kappa score:\", cv_scores.mean())\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"'''\n%%time\nbest_kappa_score = 0\nbest_params = {}\n\nparam_grid = {\n    'learning_rate': [0.01, 0.1, 0.2],\n    'max_depth': [3, 5, 7],\n    'n_estimators': [50, 100, 200]\n}\n\nfor lr in param_grid['learning_rate']:\n  for md in param_grid['max_depth']:\n    for ne in param_grid['n_estimators']:\n        print(f\"Testing parameters: learning_rate={lr}, max_depth={md}, n_estimators={ne}\")\n        lgb_model = lgb.LGBMClassifier(objective='multiclass', num_class=4, random_state=42, verbose= -100, verbose_eval = -1, learning_rate=lr, max_depth=md, n_estimators=ne)\n        cv_scores = cross_val_score(lgb_model, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\n        kappa_score = cv_scores.mean()\n        print(\"Mean Kappa score:\", cv_scores.mean())\n        if kappa_score > best_kappa_score:\n            best_kappa_score = kappa_score\n            best_params = {'learning_rate': lr, 'max_depth': md, 'n_estimators': ne}\n        print(\"-\"*50)\n\nprint(f\"Best kappa score: {best_kappa_score}\")\nprint(f\"Best parameters: {best_params}\")\n\n# Make predictions on the validation set\nbest_lgb_model = lgb.LGBMClassifier(objective='multiclass', num_class=4, random_state=42, verbose= -100, verbose_eval = -1, **best_params)\nbest_lgb_model.fit(X_train, y_train)\n\n# Make predictions on the validation set\ny_pred = best_lgb_model.predict(X_test)\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()\n'''","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (3.d.) Random forest","metadata":{}},{"cell_type":"code","source":"%%time\nbest_params = {'max_features': 'sqrt', 'max_depth': 20, 'n_estimators': 100}\n\n# Create the classifier object\nrfc = RandomForestClassifier(random_state=42, **best_params)\n\n# Train the classifier\nrfc.fit(X_train, y_train)\n \n# Make predictions\ny_pred = rfc.predict(X_test)\n\n# Perform cross-validation\ncv_scores = cross_val_score(rfc, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\nprint(\"Cross-validation scores:\", cv_scores)\nprint(\"Mean Kappa score:\", cv_scores.mean())\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"'''\n%%time\nbest_kappa_score = 0\nbest_params = {}\n\n\nparam_grid = {\n    'max_features': ['sqrt', 'log2'],\n    'max_depth': [10, 20, 30, None],\n    'n_estimators': [50, 100, 200]\n}\n\nfor mf in param_grid['max_features']:\n  for md in param_grid['max_depth']:\n    for ne in param_grid['n_estimators']:\n        print(f\"Testing parameters: max_features={mf}, max_depth={md}, n_estimators={ne}\")\n        rf_model = RandomForestClassifier(max_features=mf, max_depth=md, n_estimators=ne, random_state=42)\n        cv_scores = cross_val_score(rf_model, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\n        kappa_score = cv_scores.mean()\n        print(\"Mean Kappa score:\", kappa_score)\n        if kappa_score > best_kappa_score:\n            best_kappa_score = kappa_score\n            best_params = {'max_features': mf, 'max_depth': md, 'n_estimators': ne}\n        print(\"-\"*50)\n\n\nbest_params = {'max_features': 'sqrt', 'max_depth': 20, 'n_estimators': 100}\n\nprint(f\"Best kappa score: {best_kappa_score}\")\nprint(f\"Best parameters: {best_params}\")\n\n# Make predictions on the validation set\nbest_rf_model = RandomForestClassifier(random_state=42, **best_params)\nbest_rf_model.fit(X_train, y_train)\n\n# Make predictions on the validation set\ny_pred = best_rf_model.predict(X_test)\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()\n'''","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## (3.e) Catboost","metadata":{}},{"cell_type":"code","source":"%%time\nbest_params = {'learning_rate': 0.1, 'max_depth': 7, 'n_estimators': 200}\n\n# Create the classifier object\ncb = CatBoostClassifier(loss_function='MultiClass', verbose=False, **best_params)\n\n# Train the classifier\ncb.fit(X_train, y_train)\n \n# Make predictions\ny_pred = cb.predict(X_test)\n\n# Perform cross-validation\ncv_scores = cross_val_score(cb, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\nprint(\"Cross-validation scores:\", cv_scores)\nprint(\"Mean Kappa score:\", cv_scores.mean())\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()\n","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\n'''\n%%time\nbest_kappa_score = 0\nbest_params = {}\n\nparam_grid = {\n    'learning_rate': [0.01, 0.1, 0.2],\n    'max_depth': [3, 5, 7],\n    'n_estimators': [50, 100, 200]\n}\n\nfor lr in param_grid['learning_rate']:\n  for md in param_grid['max_depth']:\n    for ne in param_grid['n_estimators']:\n        print(f\"Testing parameters: learning_rate={lr}, max_depth={md}, n_estimators={ne}\")\n        cb_model = CatBoostClassifier(loss_function='MultiClass', verbose=False, random_state=42, learning_rate=lr, max_depth=md, n_estimators=ne)\n        cv_scores = cross_val_score(cb_model, X_train, y_train, cv=5, scoring=make_scorer(cohen_kappa_score, weights='quadratic'))\n        kappa_score = cv_scores.mean()\n        print(\"Mean Kappa score:\", cv_scores.mean())\n        if kappa_score > best_kappa_score:\n            best_kappa_score = kappa_score\n            best_params = {'learning_rate': lr, 'max_depth': md, 'n_estimators': ne}\n        print(\"-\"*50)\n\nprint(f\"Best kappa score: {best_kappa_score}\")\nprint(f\"Best parameters: {best_params}\")\n\n# Make predictions on the validation set\nbest_cb_model = cb_model = CatBoostClassifier(loss_function='MultiClass', verbose=False, random_state=42, **best_params)\nbest_cb_model.fit(X_train, y_train)\n\n# Make predictions on the validation set\ny_pred = best_cb_model.predict(X_test)\n\n# Evaluate the model\naccuracy = accuracy_score(y_test, y_pred)\nprint(f'Accuracy: {accuracy}')\nkappa_score = cohen_kappa_score(y_test, y_pred, weights='quadratic')\nprint(f'Kappa score: {kappa_score}')\nprint(confusion_matrix(y_test, y_pred))\nprint(classification_report(y_test, y_pred, zero_division=0))\nprint(\"-\"*50)\n\n# Plot the confusion matrix using Seaborn\nconf_matrix = confusion_matrix(y_test, y_pred)\nplt.figure(figsize=(6, 6))\nsns.heatmap(conf_matrix, annot=True, fmt=\"d\", cmap=\"Blues\", cbar=False)\nplt.xlabel('Predicted')\nplt.ylabel('Actual')\nplt.title('Confusion Matrix')\nplt.show()\n'''","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# (4) Prediction","metadata":{}},{"cell_type":"markdown","source":"Instead of choosing just one mode. We will create an ensemble of models by assigning a weight to each model proportional to its performance. We leave out SGD since the performance is very poor compared to the other models","metadata":{}},{"cell_type":"code","source":"if \"id\" in final_test.columns:\n    labels = final_test[\"id\"]\n    final_test.drop([\"id\"], axis = 1, inplace = True)","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def roundoff(arr, thresholds=[0.5, 1.5, 2.5]):\n    return np.where(arr < thresholds[0], 0, \n                np.where(arr < thresholds[1], 1, \n                    np.where(arr < thresholds[2], 2, 3)))","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Weights: XGBoost, LGBM, RF, CatBoost\nmodelWeight = [0.3, 0.4, 0.2, 0.1] \nXGBPredict = xgb_model.predict(final_test) # best_xgb_model\nLGBMPredict = lgb_model.predict(final_test) # best_lgb_model\nRFPredict = rfc.predict(final_test) # best_rf_model\nCBPredict = cb.predict(final_test) # best_cb_model\ny_pred = modelWeight[0] * XGBPredict + \\\n         modelWeight[1] * LGBMPredict + \\\n         modelWeight[2] * RFPredict + \\\n         modelWeight[3] * np.squeeze(CBPredict)\ny_pred = roundoff(y_pred)\ny_pred","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submit_df = pd.DataFrame(y_pred, columns = ['sii'])\nsubmit_df['id'] = testDData['id']\nsubmit_df = submit_df[['id', 'sii']]\nsubmit_df","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"submit_df.to_csv('submission.csv', index=False)","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}