{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.metrics import classification_report, confusion_matrix, accuracy_score\nfrom sklearn.preprocessing import OneHotEncoder, StandardScaler\nfrom sklearn.decomposition import PCA\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.ensemble import RandomForestRegressor\nfrom sklearn.metrics import mean_squared_error, r2_score\nfrom sklearn.compose import ColumnTransformer\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-10-23T05:37:10.487699Z","iopub.execute_input":"2024-10-23T05:37:10.488190Z","iopub.status.idle":"2024-10-23T05:37:16.107843Z","shell.execute_reply.started":"2024-10-23T05:37:10.488122Z","shell.execute_reply":"2024-10-23T05:37:16.106303Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Understanding the Task\n\nIn this competition, we aim to predict the **Severity Impairment Index (SII)**, a measure derived from the **Parent-Child Internet Addiction Test (PCIAT)**. The SII reflects the degree of problematic internet use in participants aged 5-22, based on various clinical and behavioral data. \n\nThe dataset comes from the **Healthy Brain Network (HBN)** study, which includes comprehensive clinical and research screenings. Our task is to leverage two primary data sources:\n- **Actigraphy Data**: Time-series data captured by a wrist-worn accelerometer over a period of up to 30 days.\n- **Tabular Data**: Demographic, physical, behavioral, and mental health measurements gathered from various clinical instruments, including:\n  - **Demographics** (age, sex)\n  - **Physical Activity Questionnaires**\n  - **Fitness and Physical Measures**\n  - **Sleep Disturbance Scales**\n  - **Internet Usage Behavior** (daily hours of use)\n\nThe target variable, **SII**, ranges from:\n- `0` for None\n- `1` for Mild\n- `2` for Moderate\n- `3` for Severe\n\nThe primary challenge is that the majority of participants have missing values in many fields, and a portion of the training data lacks SII labels. The test data contains complete SII values, which are withheld during the competition.\n\n### Data Files\n- **series_train.parquet / series_test.parquet**: Continuous accelerometer readings (actigraphy) for each participant. Each file contains:\n  - **X, Y, Z**: Raw accelerometer readings.\n  - **ENMO**: Calculated metric of movement intensity.\n  - **AngleZ**: Arm position relative to the horizontal plane.\n  - **non-wear_flag**: Whether the accelerometer was not worn (0 = worn, 1 = not worn).\n  - **Light**: Ambient light level in lux.\n  - **Time of Day**: Time of recording, broken down into seconds.\n  \n- **train.csv / test.csv**: Tabular data including demographics, internet usage, and clinical measures.\n- **data_dictionary.csv**: Descriptions of all fields and clinical instruments used in the study.\n\n### Goal\nUsing the combination of actigraphy data and tabular clinical measures, your goal is to predict each participant's **Severity Impairment Index (SII)**, identifying patterns that link physical activity, fitness, and internet usage behaviors with internet addiction severity.\n\n","metadata":{}},{"cell_type":"markdown","source":"# Importing Libraries","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport matplotlib.gridspec as gridspec\nimport seaborn as sns\nimport warnings\nwarnings.filterwarnings('ignore', category=FutureWarning)\n\nsns.set(style=\"whitegrid\")","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.109775Z","iopub.execute_input":"2024-10-23T05:37:16.110332Z","iopub.status.idle":"2024-10-23T05:37:16.517877Z","shell.execute_reply.started":"2024-10-23T05:37:16.110288Z","shell.execute_reply":"2024-10-23T05:37:16.516706Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Data","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\ndata_dict = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/data_dictionary.csv')\ntrain.head()","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.519183Z","iopub.execute_input":"2024-10-23T05:37:16.519718Z","iopub.status.idle":"2024-10-23T05:37:16.668987Z","shell.execute_reply.started":"2024-10-23T05:37:16.519666Z","shell.execute_reply":"2024-10-23T05:37:16.667799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Preprocessing","metadata":{"execution":{"iopub.status.busy":"2024-10-16T00:06:57.272757Z","iopub.execute_input":"2024-10-16T00:06:57.273994Z","iopub.status.idle":"2024-10-16T00:06:57.279541Z","shell.execute_reply.started":"2024-10-16T00:06:57.273938Z","shell.execute_reply":"2024-10-16T00:06:57.278060Z"}}},{"cell_type":"code","source":"train.describe().transpose()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.671593Z","iopub.execute_input":"2024-10-23T05:37:16.671997Z","iopub.status.idle":"2024-10-23T05:37:16.865795Z","shell.execute_reply.started":"2024-10-23T05:37:16.671954Z","shell.execute_reply":"2024-10-23T05:37:16.864466Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.867467Z","iopub.execute_input":"2024-10-23T05:37:16.867895Z","iopub.status.idle":"2024-10-23T05:37:16.906048Z","shell.execute_reply.started":"2024-10-23T05:37:16.867850Z","shell.execute_reply":"2024-10-23T05:37:16.904732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train[\"sii\"].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.907597Z","iopub.execute_input":"2024-10-23T05:37:16.908085Z","iopub.status.idle":"2024-10-23T05:37:16.926514Z","shell.execute_reply.started":"2024-10-23T05:37:16.908023Z","shell.execute_reply":"2024-10-23T05:37:16.925149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"threshold = 0.5 * len(train)\ncolumns_with_data = train.columns[train.isnull().sum() < threshold]\ntrain = train[columns_with_data]\ntrain = train.fillna(0)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.928359Z","iopub.execute_input":"2024-10-23T05:37:16.928886Z","iopub.status.idle":"2024-10-23T05:37:16.955185Z","shell.execute_reply.started":"2024-10-23T05:37:16.928827Z","shell.execute_reply":"2024-10-23T05:37:16.953927Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"target_column = 'sii'\ntrain_data_cleaned = train.dropna(subset=[target_column])\n\ntrain_data_cleaned.head()\ntrain_data_cleaned.info()","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.956659Z","iopub.execute_input":"2024-10-23T05:37:16.957014Z","iopub.status.idle":"2024-10-23T05:37:16.982023Z","shell.execute_reply.started":"2024-10-23T05:37:16.956975Z","shell.execute_reply":"2024-10-23T05:37:16.980713Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# EDA (Exploratory Data Analysis)","metadata":{}},{"cell_type":"code","source":"categorical_columns = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', \n                       'FGC-Season', 'BIA-Season', 'PCIAT-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n\nplt.figure(figsize=(18, 20))\nsns.set(style=\"whitegrid\")\n\nfor i, col in enumerate(categorical_columns, 1):\n    plt.subplot(4, 2, i)  # Create subplots: 4 rows, 2 columns, plot index i\n    sns.boxplot(x=col, y='sii', data=train_data_cleaned, palette=\"Set2\")\n    \n\n    plt.xlabel(f'{col}', fontsize=12)\n    plt.ylabel('Severity Impairment Index (sii)', fontsize=12)\n    plt.title(f\"Distribution of 'sii' by {col}\", fontsize=14)\n    \n    plt.xticks(rotation=30 if train_data_cleaned[col].nunique() > 5 else 0)\n    \nplt.tight_layout()\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:16.983610Z","iopub.execute_input":"2024-10-23T05:37:16.984087Z","iopub.status.idle":"2024-10-23T05:37:20.115807Z","shell.execute_reply.started":"2024-10-23T05:37:16.984032Z","shell.execute_reply":"2024-10-23T05:37:20.114383Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_cols = train_data_cleaned.select_dtypes(include=['float64', 'int64']).columns\nplots_per_row = 4  \nn_rows = (len(numerical_cols) + plots_per_row - 1) // plots_per_row\n\nplt.figure(figsize=(22, 5 * n_rows))\nsns.set(style=\"whitegrid\")\n\nfor i, col in enumerate(numerical_cols):\n    plt.subplot(n_rows, plots_per_row, i + 1)\n    sns.boxplot(x='sii', y=col, data=train_data_cleaned, palette=\"Set3\", showfliers=False)\n    sns.stripplot(x='sii', y=col, data=train_data_cleaned, color='black', size=3, alpha=0.5, jitter=True)\n    \n    plt.title(f\"'{col}' vs 'sii'\", fontsize=13)\n    plt.xlabel('Severity Impairment Index (sii)', fontsize=12)\n    plt.ylabel(col, fontsize=12)\n    \n    plt.tight_layout(pad=1.0)\n\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:37:20.121294Z","iopub.execute_input":"2024-10-23T05:37:20.121782Z","iopub.status.idle":"2024-10-23T05:38:51.296496Z","shell.execute_reply.started":"2024-10-23T05:37:20.121734Z","shell.execute_reply":"2024-10-23T05:38:51.295176Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(20, 20))\n\nfor i, col in enumerate(numerical_cols, 1):\n    plt.subplot((len(numerical_cols) + plots_per_row - 1) // plots_per_row, plots_per_row, i)  # Create subplots\n    sns.histplot(train_data_cleaned[col], kde=True, bins=30, color='blue')\n    plt.xlabel(col, fontsize=12)\n    plt.ylabel('Frequency', fontsize=12)\n    plt.title(f'Distribution of {col}', fontsize=14)\n\nplt.tight_layout()\nplt.show()\n","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:38:51.298157Z","iopub.execute_input":"2024-10-23T05:38:51.298569Z","iopub.status.idle":"2024-10-23T05:39:17.824957Z","shell.execute_reply.started":"2024-10-23T05:38:51.298528Z","shell.execute_reply":"2024-10-23T05:39:17.823702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# ","metadata":{}},{"cell_type":"markdown","source":"# Data Preprocessing and Feature Engineering 2","metadata":{}},{"cell_type":"markdown","source":"## Split Features","metadata":{}},{"cell_type":"code","source":"X = train_data_cleaned.drop(columns=[target_column])  # Features\ny = train_data_cleaned[target_column]","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:17.826770Z","iopub.execute_input":"2024-10-23T05:39:17.827160Z","iopub.status.idle":"2024-10-23T05:39:17.835040Z","shell.execute_reply.started":"2024-10-23T05:39:17.827118Z","shell.execute_reply":"2024-10-23T05:39:17.833859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Categorical Columns and Encode Them","metadata":{}},{"cell_type":"code","source":"categorical_cols = X.select_dtypes(include=['object', 'category']).columns\nif len(categorical_cols) > 0:\n    X = pd.get_dummies(X, columns=categorical_cols, drop_first=True)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:17.836827Z","iopub.execute_input":"2024-10-23T05:39:17.837257Z","iopub.status.idle":"2024-10-23T05:39:17.888238Z","shell.execute_reply.started":"2024-10-23T05:39:17.837207Z","shell.execute_reply":"2024-10-23T05:39:17.887209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Train Test Split","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(X, y, test_size=0.2, random_state=42, stratify=y)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:17.889503Z","iopub.execute_input":"2024-10-23T05:39:17.889855Z","iopub.status.idle":"2024-10-23T05:39:17.927596Z","shell.execute_reply.started":"2024-10-23T05:39:17.889817Z","shell.execute_reply":"2024-10-23T05:39:17.926115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Scaling Features","metadata":{}},{"cell_type":"code","source":"scaler = StandardScaler()\nX_train_scaled = scaler.fit_transform(X_train)\nX_val_scaled = scaler.transform(X_val)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:17.929574Z","iopub.execute_input":"2024-10-23T05:39:17.930035Z","iopub.status.idle":"2024-10-23T05:39:19.032096Z","shell.execute_reply.started":"2024-10-23T05:39:17.929979Z","shell.execute_reply":"2024-10-23T05:39:19.030580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Random Forest Classifier","metadata":{}},{"cell_type":"code","source":"rf_model = RandomForestClassifier(n_estimators=100, random_state=42)\nrf_model.fit(X_train_scaled, y_train)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:19.033766Z","iopub.execute_input":"2024-10-23T05:39:19.034248Z","iopub.status.idle":"2024-10-23T05:39:20.835052Z","shell.execute_reply.started":"2024-10-23T05:39:19.034189Z","shell.execute_reply":"2024-10-23T05:39:20.833853Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Predictions on Validation Set ","metadata":{}},{"cell_type":"code","source":"y_pred = rf_model.predict(X_val_scaled)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:20.836343Z","iopub.execute_input":"2024-10-23T05:39:20.836737Z","iopub.status.idle":"2024-10-23T05:39:20.873591Z","shell.execute_reply.started":"2024-10-23T05:39:20.836694Z","shell.execute_reply":"2024-10-23T05:39:20.872328Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Evaluate the Model","metadata":{}},{"cell_type":"code","source":"accuracy = accuracy_score(y_val, y_pred)\nprint(f\"Accuracy: {accuracy:.4f}\")","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:20.875189Z","iopub.execute_input":"2024-10-23T05:39:20.875593Z","iopub.status.idle":"2024-10-23T05:39:20.886261Z","shell.execute_reply.started":"2024-10-23T05:39:20.875550Z","shell.execute_reply":"2024-10-23T05:39:20.885040Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(classification_report(y_val, y_pred))","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:20.887606Z","iopub.execute_input":"2024-10-23T05:39:20.888060Z","iopub.status.idle":"2024-10-23T05:39:20.909470Z","shell.execute_reply.started":"2024-10-23T05:39:20.888016Z","shell.execute_reply":"2024-10-23T05:39:20.908519Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Feature Importance","metadata":{}},{"cell_type":"code","source":"importances = rf_model.feature_importances_\nfeature_importance_df = pd.DataFrame({\n    'feature': X_train.columns,\n    'importance': importances\n}).sort_values(by='importance', ascending=False)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:20.910928Z","iopub.execute_input":"2024-10-23T05:39:20.911331Z","iopub.status.idle":"2024-10-23T05:39:20.932022Z","shell.execute_reply.started":"2024-10-23T05:39:20.911288Z","shell.execute_reply":"2024-10-23T05:39:20.930721Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(feature_importance_df.head(10))","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:39:20.933444Z","iopub.execute_input":"2024-10-23T05:39:20.933859Z","iopub.status.idle":"2024-10-23T05:39:20.943469Z","shell.execute_reply.started":"2024-10-23T05:39:20.933815Z","shell.execute_reply":"2024-10-23T05:39:20.942097Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"season_mapping = {'Winter': 0, 'Fall': 0, 'Spring': 1, 'Summer': 1}\ntest['Basic_Demos-Enroll_Season'] = test['Basic_Demos-Enroll_Season'].map(season_mapping)\ntest['CGAS-Season'] = test['CGAS-Season'].map(season_mapping)\ntest['Physical-Season'] = test['Physical-Season'].map(season_mapping)\ntest['SDS-Season'] = test['SDS-Season'].map(season_mapping)\ntest['PreInt_EduHx-Season'] = test['PreInt_EduHx-Season'].map(season_mapping)\ntest['FGC-Season'] = test['FGC-Season'].map(season_mapping)\ntest['BIA-Season'] = test['BIA-Season'].map(season_mapping)\ntest['Basic_Demos-Enroll_Season'] = test['Basic_Demos-Enroll_Season'].map(season_mapping)\ntest['CGAS-Season'] = test['CGAS-Season'].map(season_mapping)\ntest['Physical-Season'] = test['Physical-Season'].map(season_mapping)\ntest = test.reindex(columns=train.columns, fill_value=0)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:44:52.243756Z","iopub.execute_input":"2024-10-23T05:44:52.244217Z","iopub.status.idle":"2024-10-23T05:44:52.266118Z","shell.execute_reply.started":"2024-10-23T05:44:52.244169Z","shell.execute_reply":"2024-10-23T05:44:52.264733Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_percentage = train.isnull().sum() / len(train)\ncolumns_to_drop = missing_percentage[missing_percentage > threshold].index\nnum_cols = train.select_dtypes(include=['float64', 'int64']).columns\ntest[num_cols] = test[num_cols].fillna(test[num_cols].median())\ncat_cols = train.select_dtypes(include=['object']).columns\ntest[cat_cols] = test[cat_cols].fillna(test[cat_cols].mode().iloc[0])\ntest.drop(columns_to_drop, axis=1, inplace=True)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:25.919932Z","iopub.execute_input":"2024-10-23T05:45:25.920528Z","iopub.status.idle":"2024-10-23T05:45:25.984067Z","shell.execute_reply.started":"2024-10-23T05:45:25.920472Z","shell.execute_reply":"2024-10-23T05:45:25.982822Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"threshold = 0.5  # \nmissing_percentage = train.isnull().sum() / len(train)\ncolumns_to_drop = missing_percentage[missing_percentage > threshold].index\ntest = test.reindex(columns=train.columns, fill_value=0)\ntest.drop(columns_to_drop, axis=1, inplace=True)\ntest[num_cols] = test[num_cols].fillna(test[num_cols].median())\ntest[cat_cols] = test[cat_cols].fillna(test[cat_cols].mode().iloc[0])\nprint(test.isnull().sum().sum())","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:28.000099Z","iopub.execute_input":"2024-10-23T05:45:28.001338Z","iopub.status.idle":"2024-10-23T05:45:28.066760Z","shell.execute_reply.started":"2024-10-23T05:45:28.001268Z","shell.execute_reply":"2024-10-23T05:45:28.065145Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_final = test.drop(columns=['id', 'sii'])\n","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:37.024948Z","iopub.execute_input":"2024-10-23T05:45:37.026367Z","iopub.status.idle":"2024-10-23T05:45:37.036064Z","shell.execute_reply.started":"2024-10-23T05:45:37.026264Z","shell.execute_reply":"2024-10-23T05:45:37.034352Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_final = pd.get_dummies(X_test_final, drop_first=True)\n","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:38.487922Z","iopub.execute_input":"2024-10-23T05:45:38.488461Z","iopub.status.idle":"2024-10-23T05:45:38.501846Z","shell.execute_reply.started":"2024-10-23T05:45:38.488402Z","shell.execute_reply":"2024-10-23T05:45:38.500446Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_final = X_test_final.reindex(columns=X_train.columns, fill_value=0)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:39.487490Z","iopub.execute_input":"2024-10-23T05:45:39.488739Z","iopub.status.idle":"2024-10-23T05:45:39.496228Z","shell.execute_reply.started":"2024-10-23T05:45:39.488678Z","shell.execute_reply":"2024-10-23T05:45:39.494856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"y_test_pred = rf_model.predict(X_test_final)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:43.647847Z","iopub.execute_input":"2024-10-23T05:45:43.648420Z","iopub.status.idle":"2024-10-23T05:45:43.740493Z","shell.execute_reply.started":"2024-10-23T05:45:43.648361Z","shell.execute_reply":"2024-10-23T05:45:43.739115Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Test Predictions:\", y_test_pred)","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:45.184035Z","iopub.execute_input":"2024-10-23T05:45:45.184633Z","iopub.status.idle":"2024-10-23T05:45:45.192728Z","shell.execute_reply.started":"2024-10-23T05:45:45.184576Z","shell.execute_reply":"2024-10-23T05:45:45.191182Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({\n    'id': test['id'],\n    'sii': y_test_pred\n})\n","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:51.330782Z","iopub.execute_input":"2024-10-23T05:45:51.331316Z","iopub.status.idle":"2024-10-23T05:45:51.338901Z","shell.execute_reply.started":"2024-10-23T05:45:51.331265Z","shell.execute_reply":"2024-10-23T05:45:51.337320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)\nsubmission","metadata":{"execution":{"iopub.status.busy":"2024-10-23T05:45:53.819201Z","iopub.execute_input":"2024-10-23T05:45:53.819773Z","iopub.status.idle":"2024-10-23T05:45:53.841831Z","shell.execute_reply.started":"2024-10-23T05:45:53.819721Z","shell.execute_reply":"2024-10-23T05:45:53.840523Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}