{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"},{"sourceId":7453542,"sourceType":"datasetVersion","datasetId":921302}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Based\n- https://www.kaggle.com/code/honganzhu/cmi-piu-competition?scriptVersionId=201912528 Version44 LB0.492\n\n If you find this notebook useful, please upvote this and the based one.","metadata":{}},{"cell_type":"markdown","source":"# Description of Imported Libraries\n\n- **NumPy (`np`)**: Used for efficient numerical operations, including linear algebra and array manipulation.\n- **Pandas (`pd`)**: Provides data structures like DataFrames for handling structured data, essential for data preprocessing.\n- **Polars (`pl`)**: A faster alternative to pandas for DataFrame operations, particularly useful for large datasets.\n- **Matplotlib & Seaborn (`plt`, `sns`)**: Visualization libraries. Matplotlib is used for basic plots, while Seaborn builds on it to create more advanced statistical visualizations.\n- **LightGBM, XGBoost, CatBoost**: Machine learning libraries used for gradient boosting, which is efficient for both regression and classification tasks.\n- **Colorama**: Enhances console output with colored text, making it easier to highlight important results or warnings.\n- **SciPy (`minimize`)**: Provides optimization routines, such as adjusting thresholds to maximize performance metrics like kappa scores.\n- **OS**: Used for file path manipulations and system-related functions.\n- **Scikit-learn (`sklearn`)**: A powerful machine learning library, providing utilities for cross-validation, metrics, and model cloning.\n- **YDF**: A specialized library for machine learning tasks, likely including decision forests.\n- **ThreadPoolExecutor & TQDM**: Tools for parallelizing tasks and displaying progress bars for long-running loops, improving efficiency and usability.\n- **Warnings**: Filters out unwanted warnings to keep the output clean, useful when dealing with noisy outputs from multiple libraries.\n- **IPython display (`clear_output`)**: A utility for clearing the Jupyter notebook output, often used to avoid clutter in long-running scripts.\n","metadata":{}},{"cell_type":"markdown","source":"## Notes About The Problem Setup\nby Limeng\n\nIn particular, the target `sii` is missing for a portion of the participants in the training set. You may wish to apply non-supervised learning techniques to this data. The `sii` value is present for all instances in the test set <br>\n`parquet` files containing the accelerometer (actigraphy) series and csv files (`train.csv` and `test.csv`) containing the remaining tabular data <br>\n-> physical activity metrics (`ENMO` derived) and sleep metrics (`angle-Z` derived) <br>\nparquet files at a glimpse: (cols) step, X, Y, Z, enmo, anglez, non-wear_flag, light, battery_voltage, time_of_day, weekday, quarter, relative_date_PCIAT <br>\n\nhave both categorical (`...Season`) and numerical data. ? Use encoders to convert categorical ones into numerical ones <br>\nSubmission file format: `id` and `sii` <br>","metadata":{}},{"cell_type":"code","source":"!pip -q install /kaggle/input/pytorchtabnet/pytorch_tabnet-4.1.0-py3-none-any.whl","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:27:38.537894Z","iopub.execute_input":"2024-11-15T18:27:38.538301Z","iopub.status.idle":"2024-11-15T18:27:54.219080Z","shell.execute_reply.started":"2024-11-15T18:27:38.538261Z","shell.execute_reply":"2024-11-15T18:27:54.217665Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from pytorch_tabnet.tab_model import TabNetRegressor\nimport torch","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:27:54.221526Z","iopub.execute_input":"2024-11-15T18:27:54.221971Z","iopub.status.idle":"2024-11-15T18:27:58.598492Z","shell.execute_reply.started":"2024-11-15T18:27:54.221921Z","shell.execute_reply":"2024-11-15T18:27:58.597439Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nimport re\nfrom sklearn.base import clone\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.model_selection import StratifiedKFold\nfrom scipy.optimize import minimize\nfrom concurrent.futures import ThreadPoolExecutor\nfrom tqdm import tqdm\nimport polars as pl\nimport polars.selectors as cs\nimport matplotlib.pyplot as plt\nfrom matplotlib.ticker import MaxNLocator, FormatStrFormatter, PercentFormatter\nimport seaborn as sns\n\nfrom sklearn.preprocessing import StandardScaler\nimport matplotlib.pyplot as plt\nfrom keras.models import Model\nfrom keras.layers import Input, Dense\nfrom keras.optimizers import Adam\nimport torch\nimport torch.nn as nn\nimport torch.optim as optim\n\nfrom colorama import Fore, Style\nfrom IPython.display import clear_output\nimport warnings\nfrom lightgbm import LGBMRegressor\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nfrom sklearn.ensemble import VotingRegressor, RandomForestRegressor, GradientBoostingRegressor\nfrom sklearn.impute import SimpleImputer, KNNImputer\nfrom sklearn.pipeline import Pipeline\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:27:58.599991Z","iopub.execute_input":"2024-11-15T18:27:58.600681Z","iopub.status.idle":"2024-11-15T18:28:15.545875Z","shell.execute_reply.started":"2024-11-15T18:27:58.600624Z","shell.execute_reply":"2024-11-15T18:28:15.544442Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"target_labels = ['None', 'Mild', 'Moderate', 'Severe']","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:29:01.571128Z","iopub.execute_input":"2024-11-15T18:29:01.572790Z","iopub.status.idle":"2024-11-15T18:29:01.578923Z","shell.execute_reply.started":"2024-11-15T18:29:01.572738Z","shell.execute_reply":"2024-11-15T18:29:01.577304Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Note that train.csv has more columns than test.csv, so we may need to select certain common columns.","metadata":{}},{"cell_type":"markdown","source":"# Loading Data","metadata":{}},{"cell_type":"code","source":"#season_dtype = pl.Enum(['Spring', 'Summer', 'Fall', 'Winter'])\n#this is my code\ntrain = (\n    pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\n)\n\ntest = (\n    pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\n)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:29:03.252057Z","iopub.execute_input":"2024-11-15T18:29:03.252616Z","iopub.status.idle":"2024-11-15T18:29:03.356120Z","shell.execute_reply.started":"2024-11-15T18:29:03.252530Z","shell.execute_reply":"2024-11-15T18:29:03.354664Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"import pandas as pd\nimport os\n\n# Define paths for both train and test datasets\ntrain_base_path = \"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\"\ntest_base_path = \"/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet\"\n\n# Function to load and filter based on the first N IDs\ndef load_and_filter_data(base_path, num_ids=200):\n    # List all subfolders to get IDs (assuming each folder is named \"id=xxxxxx\")\n    all_subfolders = sorted([f for f in os.listdir(base_path) if f.startswith(\"id=\")])\n\n    # Extract only the IDs and select the first N\n    all_ids = [folder.split('=')[1] for folder in all_subfolders]\n    selected_ids = all_ids[:num_ids]\n\n    # Initialize list to store each DataFrame\n    data_chunks = []\n\n    # Load each Parquet file for the selected IDs\n    for id_ in selected_ids:\n        # Construct the path to the Parquet file\n        file_path = os.path.join(base_path, f\"id={id_}\", \"part-0.parquet\")\n\n        # Check if the file exists and load it\n        if os.path.exists(file_path):\n            chunk = pd.read_parquet(file_path)\n            chunk['id'] = id_  # Add 'id' column\n\n            # Reorder columns to place 'id' at the leftmost position\n            chunk = chunk[['id'] + [col for col in chunk.columns if col != 'id']]\n\n            data_chunks.append(chunk)\n            del chunk\n\n    # Concatenate all loaded chunks into a single DataFrame\n    combined_data = pd.concat(data_chunks, ignore_index=True)\n    return combined_data\n\n# Load and filter data for both train and test datasets\nfiltered_train_data = load_and_filter_data(train_base_path, num_ids=500)\nfiltered_test_data = load_and_filter_data(test_base_path, num_ids=500)\n\n# Now both train and test datasets have the 'id' column at the leftmost position\n\n# Display the first few rows of each filtered dataset\nprint(\"Filtered Train Data:\")\nprint(filtered_train_data.head())\n\nprint(\"\\nFiltered Test Data:\")\nprint(filtered_test_data.head())\n\n# Filter the original train DataFrame (assuming you have an original train dataset 'train')\n# and merge it with the first 200 IDs\nfiltered_train_original = train[train['id'].isin(filtered_train_data['id'].unique())]\n#filtered_test_original = test[test['id'].isin(filtered_test_data['id'].unique())]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:29:05.194961Z","iopub.execute_input":"2024-11-15T18:29:05.195501Z","iopub.status.idle":"2024-11-15T18:30:41.417837Z","shell.execute_reply.started":"2024-11-15T18:29:05.195440Z","shell.execute_reply":"2024-11-15T18:30:41.415638Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#this is my code\n# Convert sii to categorical\nfiltered_train_original['sii'] = filtered_train_original['sii'].astype('category')\n\n#test['sii'] = test['sii'].astype('category')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:41.421286Z","iopub.execute_input":"2024-11-15T18:30:41.421794Z","iopub.status.idle":"2024-11-15T18:30:41.433687Z","shell.execute_reply.started":"2024-11-15T18:30:41.421737Z","shell.execute_reply":"2024-11-15T18:30:41.432147Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"filtered_train_original.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:41.435524Z","iopub.execute_input":"2024-11-15T18:30:41.436009Z","iopub.status.idle":"2024-11-15T18:30:41.458497Z","shell.execute_reply.started":"2024-11-15T18:30:41.435965Z","shell.execute_reply":"2024-11-15T18:30:41.457246Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Separate the target variable\n#this is my code\nfrom sklearn.model_selection import train_test_split\n\n# Define your feature matrix X and target vector y\nX = filtered_train_original.drop(columns=['sii'])  # Replace 'target_column' with the name of your target\ny = filtered_train_original['sii']\n\n# Split the data\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=42)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:50.338967Z","iopub.execute_input":"2024-11-15T18:30:50.339608Z","iopub.status.idle":"2024-11-15T18:30:50.358314Z","shell.execute_reply.started":"2024-11-15T18:30:50.339522Z","shell.execute_reply":"2024-11-15T18:30:50.356831Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#this is my code\n# Impute numerical columns with the mean\n# Selecting columns with data type 'object' or 'category'\nnumerical_columns = X_train.select_dtypes(include=['number']).columns\n\n# Calculate the mean of each column in the training set\nmeans = X_train[numerical_columns].mean()\n\n# Use the means from the training set to fill missing values in both training and testing sets\nX_train[numerical_columns] = X_train[numerical_columns].fillna(means)\nX_test[numerical_columns] = X_test[numerical_columns].fillna(means)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:51.423922Z","iopub.execute_input":"2024-11-15T18:30:51.424354Z","iopub.status.idle":"2024-11-15T18:30:51.502631Z","shell.execute_reply.started":"2024-11-15T18:30:51.424313Z","shell.execute_reply":"2024-11-15T18:30:51.501354Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#this is my code\n# Selecting columns with data type 'object' or 'category'\ncategorical_columns = X_train.select_dtypes(include=['object', 'category']).columns\n# Select categorical columns, excluding 'id'\ncategorical_columns = categorical_columns.drop('id')\n\n# Calculate the mode (most frequent value) for each categorical column in the training set\ncategorical_modes = X_train[categorical_columns].mode().iloc[0]\n\n# Impute missing values in categorical columns for both training and testing sets\nX_train[categorical_columns] = X_train[categorical_columns].fillna(categorical_modes)\nX_test[categorical_columns] = X_test[categorical_columns].fillna(categorical_modes)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:52.221843Z","iopub.execute_input":"2024-11-15T18:30:52.222332Z","iopub.status.idle":"2024-11-15T18:30:52.252726Z","shell.execute_reply.started":"2024-11-15T18:30:52.222284Z","shell.execute_reply":"2024-11-15T18:30:52.251505Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#this is my code\n# Calculate the mode of y_train\nmode_sii = y_train.mode().iloc[0]\n\n# Fill missing values in both y_train and y_test with the mode from y_train\ny_train = y_train.fillna(mode_sii)\ny_test = y_test.fillna(mode_sii)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:52.785744Z","iopub.execute_input":"2024-11-15T18:30:52.786180Z","iopub.status.idle":"2024-11-15T18:30:52.796603Z","shell.execute_reply.started":"2024-11-15T18:30:52.786140Z","shell.execute_reply":"2024-11-15T18:30:52.795058Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#this is my code\nfrom sklearn.preprocessing import MinMaxScaler\nimport pandas as pd\n\n# Identify numerical columns\n\n# Initialize the scaler\nscaler = MinMaxScaler()\n\n# Fit the scaler on the numerical columns in the training set and transform\nX_train[numerical_columns] = scaler.fit_transform(X_train[numerical_columns])\n\n# Transform the numerical columns in the test set using the same scaler\nX_test[numerical_columns] = scaler.transform(X_test[numerical_columns])\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:53.534455Z","iopub.execute_input":"2024-11-15T18:30:53.534934Z","iopub.status.idle":"2024-11-15T18:30:53.576647Z","shell.execute_reply.started":"2024-11-15T18:30:53.534888Z","shell.execute_reply":"2024-11-15T18:30:53.575442Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#this is my code\nfrom sklearn.preprocessing import OneHotEncoder\nimport pandas as pd\n\n# Initialize the OneHotEncoder with handle_unknown='ignore'\nonehot_encoder = OneHotEncoder(drop='first', sparse=False, handle_unknown='ignore')\n\n# Fit and transform on training data, then transform on test data\nX_train_encoded = onehot_encoder.fit_transform(X_train[categorical_columns])\nX_test_encoded = onehot_encoder.transform(X_test[categorical_columns])\n\n# Convert back to DataFrame and add back to the original data\nX_train_encoded = pd.DataFrame(X_train_encoded, columns=onehot_encoder.get_feature_names_out(categorical_columns), index=X_train.index)\nX_test_encoded = pd.DataFrame(X_test_encoded, columns=onehot_encoder.get_feature_names_out(categorical_columns), index=X_test.index)\n\n# Drop original categorical columns and concatenate the encoded columns\nX_train = X_train.drop(columns=categorical_columns).join(X_train_encoded)\nX_test = X_test.drop(columns=categorical_columns).join(X_test_encoded)\n#here the train part is normalized and one hot encoded, please go to the time series part","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:54.184618Z","iopub.execute_input":"2024-11-15T18:30:54.185663Z","iopub.status.idle":"2024-11-15T18:30:54.227424Z","shell.execute_reply.started":"2024-11-15T18:30:54.185616Z","shell.execute_reply":"2024-11-15T18:30:54.226441Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train.dtypes","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:54.865639Z","iopub.execute_input":"2024-11-15T18:30:54.866688Z","iopub.status.idle":"2024-11-15T18:30:54.877410Z","shell.execute_reply.started":"2024-11-15T18:30:54.866639Z","shell.execute_reply":"2024-11-15T18:30:54.876095Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:55.473444Z","iopub.execute_input":"2024-11-15T18:30:55.473900Z","iopub.status.idle":"2024-11-15T18:30:55.481822Z","shell.execute_reply.started":"2024-11-15T18:30:55.473858Z","shell.execute_reply":"2024-11-15T18:30:55.480510Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"print(\"test data columns:\\n\", X_test.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:56.310443Z","iopub.execute_input":"2024-11-15T18:30:56.310863Z","iopub.status.idle":"2024-11-15T18:30:56.317675Z","shell.execute_reply.started":"2024-11-15T18:30:56.310823Z","shell.execute_reply":"2024-11-15T18:30:56.316113Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# common columns\ncommon_cols = []\nfor col in X_test.columns:\n    if col in X_train.columns:\n        common_cols.append(col)\nprint(\"Common columns:\\n\", common_cols)\nprint(\"Total # of common columns:\", len(common_cols))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:57.048735Z","iopub.execute_input":"2024-11-15T18:30:57.049163Z","iopub.status.idle":"2024-11-15T18:30:57.057184Z","shell.execute_reply.started":"2024-11-15T18:30:57.049123Z","shell.execute_reply":"2024-11-15T18:30:57.056031Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"there are too many features, indicating that feature selection may be needed","metadata":{}},{"cell_type":"code","source":"# Get total number of ids utilized in training\nlen(X_train['id'].unique())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:30:58.521215Z","iopub.execute_input":"2024-11-15T18:30:58.521677Z","iopub.status.idle":"2024-11-15T18:30:58.532315Z","shell.execute_reply.started":"2024-11-15T18:30:58.521631Z","shell.execute_reply":"2024-11-15T18:30:58.531147Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## EDA & Feature Engineering\n\nby Limeng","metadata":{}},{"cell_type":"markdown","source":"For training data","metadata":{}},{"cell_type":"code","source":"train = pd.concat([X_train, y_train], axis=1)\nlist(train.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:00.636515Z","iopub.execute_input":"2024-11-15T18:31:00.636985Z","iopub.status.idle":"2024-11-15T18:31:00.651110Z","shell.execute_reply.started":"2024-11-15T18:31:00.636940Z","shell.execute_reply":"2024-11-15T18:31:00.649843Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"vc = X_train['Basic_Demos-Enroll_Season'].value_counts()\nplt.pie(vc['count'], labels=vc['Basic_Demos-Enroll_Season'])\nplt.title('Season of enrollment (X_train)')\nplt.show()","metadata":{}},{"cell_type":"code","source":"vc = X_train['Basic_Demos-Sex'].value_counts()\nplt.pie(vc.values, labels=['boys', 'girls'])\nplt.title('Sex of participant (X_train)')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:02.229120Z","iopub.execute_input":"2024-11-15T18:31:02.229608Z","iopub.status.idle":"2024-11-15T18:31:02.455756Z","shell.execute_reply.started":"2024-11-15T18:31:02.229547Z","shell.execute_reply":"2024-11-15T18:31:02.454114Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train.columns = X_train.columns.str.strip()\n_, axs = plt.subplots(2, 1, sharex=True)\n\nfor sex in range(2):\n    ax = axs.ravel()[sex]\n    \n    # filter rows based on the 'Basic_Demos-Sex' column value\n    filtered_data = X_train.loc[X_train['Basic_Demos-Sex'] == sex]\n    \n    # compute value counts for the 'Basic_Demos-Age' column and sort by age\n    vc = filtered_data['Basic_Demos-Age'].value_counts().sort_index()\n    \n    # plot the age distribution for the current 'sex' value\n    ax.bar(vc.index, vc.values, \n           color=['lightblue', 'coral'][sex], \n           label=['boys', 'girls'][sex])\n    \n    # set x-axis to use integer values only\n    ax.xaxis.set_major_locator(MaxNLocator(integer=True))\n    ax.set_ylabel('count')\n    ax.legend()\n\nplt.suptitle('Age distribution (X_train)')\naxs.ravel()[1].set_xlabel('years')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:06.267420Z","iopub.execute_input":"2024-11-15T18:31:06.267915Z","iopub.status.idle":"2024-11-15T18:31:06.726920Z","shell.execute_reply.started":"2024-11-15T18:31:06.267870Z","shell.execute_reply":"2024-11-15T18:31:06.725510Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train.columns = train.columns.str.strip()\n_, axs = plt.subplots(2, 1, sharex=True, sharey=True)\n\nfor sex in range(2):\n    ax = axs.ravel()[sex]\n    \n    # filter data based on 'Basic_Demos-Sex' and get counts for 'sii'\n    filtered_data = train.loc[train['Basic_Demos-Sex'] == sex]\n    vc = filtered_data['sii'].value_counts(normalize=True).sort_index()  # Normalized for percentage\n    \n    ax.bar(vc.index, vc.values, color=['lightblue', 'coral'][sex], label=['boys', 'girls'][sex])\n    \n    # set x-ticks and labels\n    ax.set_xticks(np.arange(len(target_labels)))\n    ax.set_xticklabels(target_labels)\n    \n    # format y-axis as a percentage\n    ax.yaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))\n    ax.set_ylabel('Percentage')\n    ax.legend()\n\nplt.suptitle('Target Distribution (Train)')\naxs.ravel()[1].set_xlabel('Severity Impairment Index (sii)')\n\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:06.904747Z","iopub.execute_input":"2024-11-15T18:31:06.905189Z","iopub.status.idle":"2024-11-15T18:31:07.283942Z","shell.execute_reply.started":"2024-11-15T18:31:06.905150Z","shell.execute_reply":"2024-11-15T18:31:07.282749Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"Select features with high correlation with the target `sii`\n\nCorrelation-based feature selection: [tutorial](https://medium.com/@sariq16/correlation-based-feature-selection-in-a-data-science-project-3ca08d2af5c6)","metadata":{}},{"cell_type":"code","source":"train['sii'].dtype","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:09.559301Z","iopub.execute_input":"2024-11-15T18:31:09.559791Z","iopub.status.idle":"2024-11-15T18:31:09.568541Z","shell.execute_reply.started":"2024-11-15T18:31:09.559748Z","shell.execute_reply":"2024-11-15T18:31:09.567266Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"numeric_columns = train.select_dtypes(include='number').columns.difference(['sii'])\n\n# one-hot encode the categorical target column 'sii'\nsii_dummies = pd.get_dummies(train['sii'], prefix='sii')\n\n# combine the numeric columns with the one-hot encoded target column\ntrain_encoded = pd.concat([train[numeric_columns], sii_dummies], axis=1)\n\ncorr_matrix = train_encoded.corr()\ncorr_with_target = corr_matrix.loc[numeric_columns, sii_dummies.columns]\nprint(\"Correlation of numeric columns with each sii category:\\n\", corr_with_target)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:10.440199Z","iopub.execute_input":"2024-11-15T18:31:10.440702Z","iopub.status.idle":"2024-11-15T18:31:10.473871Z","shell.execute_reply.started":"2024-11-15T18:31:10.440653Z","shell.execute_reply":"2024-11-15T18:31:10.472693Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"max_corr_magnitude = corr_with_target.abs().max(axis=1)\n# can change the value of k\nk = 50\n# combine max correlation values with feature names and sort by correlation magnitude\ntop_k_features = max_corr_magnitude.nlargest(k)\nselected_features = train[top_k_features.index]\nprint(\"\\nSelected top target-correlated features:\\n\", selected_features.columns)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:11.100598Z","iopub.execute_input":"2024-11-15T18:31:11.101061Z","iopub.status.idle":"2024-11-15T18:31:11.113781Z","shell.execute_reply.started":"2024-11-15T18:31:11.101018Z","shell.execute_reply":"2024-11-15T18:31:11.112360Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"top_k_features","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:31:11.771440Z","iopub.execute_input":"2024-11-15T18:31:11.771946Z","iopub.status.idle":"2024-11-15T18:31:11.782890Z","shell.execute_reply.started":"2024-11-15T18:31:11.771902Z","shell.execute_reply":"2024-11-15T18:31:11.781512Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"`PCIAT-` are time series data\n`Season_` are time series data","metadata":{}},{"cell_type":"code","source":"# create a heatmap based on the top 20 features\ntop_20_features = max_corr_magnitude.nlargest(20)\nselected_corr_matrix = train[top_20_features.index].corr()\n\nplt.figure(figsize=(12, 10))\nsns.heatmap(selected_corr_matrix, annot=True, fmt=\".3f\", cmap=\"coolwarm\", cbar=True, square=True)\nplt.title(\"Top 20 Most Correlated Features Heatmap\")\nplt.xticks(rotation=45, ha='right')\nplt.yticks(rotation=0)\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:40:31.657823Z","iopub.execute_input":"2024-11-15T18:40:31.658265Z","iopub.status.idle":"2024-11-15T18:40:33.359497Z","shell.execute_reply.started":"2024-11-15T18:40:31.658224Z","shell.execute_reply":"2024-11-15T18:40:33.358046Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# plot the top 20 features VS the target\nplt.figure(figsize=(12, 8))\ntop_20_features.sort_values().plot(kind='barh', color='skyblue')\nplt.xlabel(\"Correlation Magnitude with Target\")\nplt.ylabel(\"Features\")\nplt.title(\"Top 20 Features with Highest Correlation to Target\")\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:48:11.857177Z","iopub.execute_input":"2024-11-15T18:48:11.858034Z","iopub.status.idle":"2024-11-15T18:48:12.482244Z","shell.execute_reply.started":"2024-11-15T18:48:11.857982Z","shell.execute_reply":"2024-11-15T18:48:12.480846Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# feel free to add/remove features here\nselected_features = selected_features.columns.tolist()\nselected_features.insert(0, 'id')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:32:50.320474Z","iopub.execute_input":"2024-11-15T18:32:50.321354Z","iopub.status.idle":"2024-11-15T18:32:50.326816Z","shell.execute_reply.started":"2024-11-15T18:32:50.321290Z","shell.execute_reply":"2024-11-15T18:32:50.325416Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"selected_features","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:32:54.511041Z","iopub.execute_input":"2024-11-15T18:32:54.512071Z","iopub.status.idle":"2024-11-15T18:32:54.521965Z","shell.execute_reply.started":"2024-11-15T18:32:54.512006Z","shell.execute_reply":"2024-11-15T18:32:54.520563Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_test = X_test[selected_features]\nX_test.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-15T18:33:14.241002Z","iopub.execute_input":"2024-11-15T18:33:14.241623Z","iopub.status.idle":"2024-11-15T18:33:14.312127Z","shell.execute_reply.started":"2024-11-15T18:33:14.241531Z","shell.execute_reply":"2024-11-15T18:33:14.310960Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"","metadata":{"trusted":true},"outputs":[],"execution_count":null}]}