{"metadata":{"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false},"kernelspec":{"display_name":"Python 3","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.10.14"},"papermill":{"default_parameters":{},"duration":24.623649,"end_time":"2024-09-30T17:28:12.521762","environment_variables":{},"exception":null,"input_path":"__notebook__.ipynb","output_path":"__notebook__.ipynb","parameters":{},"start_time":"2024-09-30T17:27:47.898113","version":"2.6.0"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n### 1. **Importing Necessary Libraries**\n   - Libraries like `numpy`, `pandas` are used for data manipulation.\n   - `re` is for regular expressions, which might be useful for string operations.\n   - `matplotlib.pyplot` and `seaborn` are used for visualizations.\n   - Machine learning libraries from `sklearn` like `train_test_split`, `Pipeline`, and model evaluation metrics like `confusion_matrix` are used for model training and validation.\n   - Libraries like `XGBoost`, `LightGBM`, and `CatBoost` are imported for different classifier models.\n   - `warnings` is imported to suppress warning messages for cleaner output.\n\n### 2. **Setting Up Matplotlib and Seaborn for Visualizations**\n   - Configure `matplotlib` for LaTeX-style text rendering and set font to `serif`.\n   - Use `sns.set_theme()` to set the Seaborn theme for consistent and polished plots.\n\n### 3. **Loading Data**\n   - Use `pd.read_csv()` to load training, testing, and a data dictionary CSV files from a specified folder path.\n   - The `train.csv` file contains training data with an 'sii' target variable.\n   - The `test.csv` file contains test data for predictions.\n\n### 4. **Visualizing Target Variable ('sii') Distribution**\n   - A count plot (`sns.countplot`) is created to visualize the distribution of the target variable ('sii') in the training data. This helps understand the class imbalance if any.\n\n### 5. **Identifying Missing Columns in the Test Set**\n   - Compare the columns of the training and test sets to identify missing columns in the test set.\n   - The missing columns (usually the target variable 'sii' or others) are removed before model training.\n\n### 6. **Visualizing Missing Data**\n   - A heatmap (`sns.heatmap`) is plotted to visualize the missing values in the training data. This helps identify where imputations are needed.\n\n### 7. **Handling Missing Data and Dropping Unnecessary Columns**\n   - The subset of the training data without missing 'sii' labels is selected for training.\n   - Unnecessary columns that don't appear in both train and test datasets are dropped.\n\n### 8. **Visualizing the Distribution of Numerical Features**\n   - Histograms (`X_train[numerical_cols].hist()`) are plotted for numerical features in the training data to check for distributions and potential outliers.\n   - Boxplots (`sns.boxplot`) are plotted to visualize the spread and outliers in the numerical features.\n\n### 9. **Outlier Handling via Clipping Function**\n   - A function `clipped()` is created to clip outliers based on quantiles. This reduces the impact of extreme values on the model performance.\n\n### 10. **Processing Function for Data Preprocessing**\n   - The `processing()` function handles both numerical and categorical features:\n     - **Numerical Features**: Outliers are clipped, and missing values are replaced with the mean.\n     - **Categorical Features**: One-hot encoding is applied for categorical variables, and `KNNImputer` is used to handle missing values in categorical features.\n   - The `ColumnTransformer` is used to apply transformations on categorical and numerical columns in the dataset.\n\n### 11. **Visualizing Feature Correlation**\n   - A heatmap of the correlation matrix (`sns.heatmap(corr)`) is plotted for the numerical features. This helps understand relationships between features and detect multicollinearity.\n\n### 12. **Processing Training Data**\n   - The training data is processed using the `processing()` function which applies imputations, outlier handling, and transformations.\n\n### 13. **Visualizing Processed Features' Distribution**\n   - After preprocessing, histograms are plotted for the processed features to confirm if the distributions are normalized and any issues like missing values or skewness are handled.\n\n### 14. **Quadratic Weighted Kappa (QWK) Score Function**\n   - A custom metric `quadratic_weighted_kappa()` is defined to calculate the quadratic weighted kappa score, which is useful for evaluating models on ordinal classification tasks (like predicting 'sii').\n\n### 15. **Model Hyperparameter Search Using GridSearchCV**\n   - A function `ParamSearchModels()` is defined to perform hyperparameter tuning using GridSearchCV on several models:\n     - Models like `XGBClassifier` are initialized with a grid of hyperparameters.\n     - The model is trained and validated using cross-validation (`ShuffleSplit`), with the QWK score as the evaluation metric.\n     - The best hyperparameters and scores are recorded for each model.\n\n### 16. **Training the Model Using XGBoost**\n   - The XGBoost classifier is trained using the best parameters obtained from the grid search. This model is used for predictions on the test set.\n\n### 17. **Processing Test Data**\n   - The test data is processed using the same `processing()` function to ensure that transformations applied to the training data are also applied to the test data.\n\n### 18. **Making Predictions**\n   - Predictions on the processed test data are made using the trained XGBoost model. These predictions are for the target variable 'sii'.\n\n### 19. **Visualizing Predictions**\n   - A histogram (`sns.histplot`) of the predicted 'sii' values is plotted to visualize the distribution of predictions and ensure they align with expectations.\n\n### 20. **Preparing the Submission File**\n   - A submission file is prepared for the Kaggle competition. The test data predictions are saved in a CSV format with columns 'id' and 'sii'.\n   - The first few rows of the submission file are displayed using `print()` for verification.\n","metadata":{"papermill":{"duration":0.009424,"end_time":"2024-09-30T17:27:50.613263","exception":false,"start_time":"2024-09-30T17:27:50.603839","status":"completed"},"tags":[]}},{"cell_type":"markdown","source":"# Import necessary libraries","metadata":{"papermill":{"duration":0.010071,"end_time":"2024-09-30T17:27:50.632837","exception":false,"start_time":"2024-09-30T17:27:50.622766","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Import necessary libraries\nimport numpy as np  # For numerical operations\nimport pandas as pd  # For handling dataframes\nimport re  # Regular expressions, though unused here\nimport matplotlib.pyplot as plt  # For plotting and visualizations\nimport seaborn as sns  # For advanced visualizations","metadata":{"_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","papermill":{"duration":2.26432,"end_time":"2024-09-30T17:27:52.906162","exception":false,"start_time":"2024-09-30T17:27:50.641842","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:50:56.807410Z","iopub.execute_input":"2024-10-04T16:50:56.807905Z","iopub.status.idle":"2024-10-04T16:51:00.014832Z","shell.execute_reply.started":"2024-10-04T16:50:56.807861Z","shell.execute_reply":"2024-10-04T16:51:00.013373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configure matplotlib settings for better plots","metadata":{"papermill":{"duration":0.00828,"end_time":"2024-09-30T17:27:52.923142","exception":false,"start_time":"2024-09-30T17:27:52.914862","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Configure matplotlib settings for better plots\nplt.rc('text', usetex=True)  # Enables LaTeX rendering for text in plots\nplt.rc('font', family='serif')  # Sets font to 'serif'\nplt.rc('xtick', labelsize=12)  # Sets font size for x-axis labels\nplt.rc('ytick', labelsize=12)  # Sets font size for y-axis labels","metadata":{"papermill":{"duration":0.01724,"end_time":"2024-09-30T17:27:52.948853","exception":false,"start_time":"2024-09-30T17:27:52.931613","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:00.018410Z","iopub.execute_input":"2024-10-04T16:51:00.019177Z","iopub.status.idle":"2024-10-04T16:51:00.027908Z","shell.execute_reply.started":"2024-10-04T16:51:00.019100Z","shell.execute_reply":"2024-10-04T16:51:00.025280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set Seaborn theme for consistent style\nsns.set_theme()","metadata":{"papermill":{"duration":0.016419,"end_time":"2024-09-30T17:27:52.973917","exception":false,"start_time":"2024-09-30T17:27:52.957498","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:00.030537Z","iopub.execute_input":"2024-10-04T16:51:00.031565Z","iopub.status.idle":"2024-10-04T16:51:00.048368Z","shell.execute_reply.started":"2024-10-04T16:51:00.031484Z","shell.execute_reply":"2024-10-04T16:51:00.046862Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ignore warnings for cleaner output","metadata":{"papermill":{"duration":0.008316,"end_time":"2024-09-30T17:27:52.991054","exception":false,"start_time":"2024-09-30T17:27:52.982738","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Ignore warnings for cleaner output\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"papermill":{"duration":0.016042,"end_time":"2024-09-30T17:27:53.015506","exception":false,"start_time":"2024-09-30T17:27:52.999464","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:00.052195Z","iopub.execute_input":"2024-10-04T16:51:00.053996Z","iopub.status.idle":"2024-10-04T16:51:00.066801Z","shell.execute_reply.started":"2024-10-04T16:51:00.053817Z","shell.execute_reply":"2024-10-04T16:51:00.064823Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import necessary modules from sklearn for data preprocessing, model evaluation, and training","metadata":{"papermill":{"duration":0.008388,"end_time":"2024-09-30T17:27:53.032461","exception":false,"start_time":"2024-09-30T17:27:53.024073","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Import necessary modules from sklearn for data preprocessing, model evaluation, and training\nfrom sklearn.impute import SimpleImputer, KNNImputer  # For missing data imputation\nfrom sklearn.compose import ColumnTransformer  # For applying different transformations to specific columns\nfrom sklearn.pipeline import Pipeline  # To create pipelines for consistent data transformations\nfrom sklearn.preprocessing import OneHotEncoder, LabelEncoder, StandardScaler  # For encoding and scaling\nfrom sklearn.model_selection import train_test_split, cross_val_score, ShuffleSplit, GridSearchCV  # For model evaluation and grid search\nfrom xgboost import XGBClassifier  # XGBoost Classifier\nfrom lightgbm import LGBMClassifier  # LightGBM Classifier\nfrom catboost import CatBoostClassifier  # CatBoost Classifier\nfrom sklearn.metrics import confusion_matrix, make_scorer, cohen_kappa_score  # For model evaluation metrics","metadata":{"papermill":{"duration":1.687915,"end_time":"2024-09-30T17:27:54.728818","exception":false,"start_time":"2024-09-30T17:27:53.040903","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:00.068743Z","iopub.execute_input":"2024-10-04T16:51:00.069685Z","iopub.status.idle":"2024-10-04T16:51:02.601402Z","shell.execute_reply.started":"2024-10-04T16:51:00.069621Z","shell.execute_reply":"2024-10-04T16:51:02.599812Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define folder path where the dataset is stored","metadata":{"papermill":{"duration":0.008284,"end_time":"2024-09-30T17:27:54.745689","exception":false,"start_time":"2024-09-30T17:27:54.737405","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Define folder path where the dataset is stored\nfolder_path = '/kaggle/input/child-mind-institute-problematic-internet-use/'","metadata":{"papermill":{"duration":0.015802,"end_time":"2024-09-30T17:27:54.770046","exception":false,"start_time":"2024-09-30T17:27:54.754244","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:02.603313Z","iopub.execute_input":"2024-10-04T16:51:02.604108Z","iopub.status.idle":"2024-10-04T16:51:02.612034Z","shell.execute_reply.started":"2024-10-04T16:51:02.604052Z","shell.execute_reply":"2024-10-04T16:51:02.609405Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load training and test datasets, with 'id' as the index column","metadata":{"papermill":{"duration":0.008384,"end_time":"2024-09-30T17:27:54.787012","exception":false,"start_time":"2024-09-30T17:27:54.778628","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Load training and test datasets, with 'id' as the index column\ntrain_df = pd.read_csv(folder_path + 'train.csv', index_col='id')\ntest_df = pd.read_csv(folder_path + 'test.csv', index_col='id')\ndata_dict = pd.read_csv(folder_path + 'data_dictionary.csv')  # Load data dictionary","metadata":{"papermill":{"duration":0.094844,"end_time":"2024-09-30T17:27:54.890509","exception":false,"start_time":"2024-09-30T17:27:54.795665","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:02.613836Z","iopub.execute_input":"2024-10-04T16:51:02.614328Z","iopub.status.idle":"2024-10-04T16:51:02.742142Z","shell.execute_reply.started":"2024-10-04T16:51:02.614283Z","shell.execute_reply":"2024-10-04T16:51:02.740852Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize class distribution of target variable 'sii'","metadata":{"papermill":{"duration":0.008478,"end_time":"2024-09-30T17:27:54.907544","exception":false,"start_time":"2024-09-30T17:27:54.899066","status":"completed"},"tags":[]}},{"cell_type":"code","source":"import matplotlib as mpl\nmpl.rcParams['text.usetex'] = False","metadata":{"papermill":{"duration":0.015685,"end_time":"2024-09-30T17:27:54.931708","exception":false,"start_time":"2024-09-30T17:27:54.916023","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:02.743970Z","iopub.execute_input":"2024-10-04T16:51:02.744403Z","iopub.status.idle":"2024-10-04T16:51:02.754453Z","shell.execute_reply.started":"2024-10-04T16:51:02.744360Z","shell.execute_reply":"2024-10-04T16:51:02.752696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize class distribution of target variable 'sii'\nplt.figure(figsize=(6, 4))\nsns.countplot(data=train_df, x='sii')\nplt.title('Class Distribution of SII')\nplt.show()","metadata":{"papermill":{"duration":0.28771,"end_time":"2024-09-30T17:27:55.227886","exception":false,"start_time":"2024-09-30T17:27:54.940176","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:02.756841Z","iopub.execute_input":"2024-10-04T16:51:02.758765Z","iopub.status.idle":"2024-10-04T16:51:03.183093Z","shell.execute_reply.started":"2024-10-04T16:51:02.758713Z","shell.execute_reply":"2024-10-04T16:51:03.180936Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training set has additional columns (like 'PCIAT') compared to test set\n# We find the extra columns in the training data that need to be removed before training\ntrain_cols = train_df.columns.tolist()  # List of columns in training data\ntest_cols = test_df.columns.tolist()  # List of columns in test data","metadata":{"papermill":{"duration":0.016513,"end_time":"2024-09-30T17:27:55.253478","exception":false,"start_time":"2024-09-30T17:27:55.236965","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:03.187890Z","iopub.execute_input":"2024-10-04T16:51:03.188319Z","iopub.status.idle":"2024-10-04T16:51:03.195173Z","shell.execute_reply.started":"2024-10-04T16:51:03.188280Z","shell.execute_reply":"2024-10-04T16:51:03.193861Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_cols = set(train_cols) - set(test_cols)\nprint(f\"Columns missing in test set: {missing_cols}\")","metadata":{"papermill":{"duration":0.017548,"end_time":"2024-09-30T17:27:55.279966","exception":false,"start_time":"2024-09-30T17:27:55.262418","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:03.196679Z","iopub.execute_input":"2024-10-04T16:51:03.197107Z","iopub.status.idle":"2024-10-04T16:51:03.208756Z","shell.execute_reply.started":"2024-10-04T16:51:03.197067Z","shell.execute_reply":"2024-10-04T16:51:03.207396Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Find columns in training data that are not in test data\ndrop_cols = list(set(train_cols) - set(test_cols))  \ndrop_cols.remove('sii')  # Do not remove the target variable 'sii'","metadata":{"papermill":{"duration":0.017183,"end_time":"2024-09-30T17:27:55.306200","exception":false,"start_time":"2024-09-30T17:27:55.289017","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:03.210172Z","iopub.execute_input":"2024-10-04T16:51:03.210582Z","iopub.status.idle":"2024-10-04T16:51:03.220321Z","shell.execute_reply.started":"2024-10-04T16:51:03.210515Z","shell.execute_reply":"2024-10-04T16:51:03.218957Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize missing data in the training set","metadata":{"papermill":{"duration":0.008622,"end_time":"2024-09-30T17:27:55.323740","exception":false,"start_time":"2024-09-30T17:27:55.315118","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Visualize missing data in the training set\nplt.figure(figsize=(10, 6))\nsns.heatmap(train_df.isnull(), cbar=False, cmap='viridis')\nplt.title('Missing Data in Training Set')\nplt.show()","metadata":{"papermill":{"duration":1.256181,"end_time":"2024-09-30T17:27:56.588719","exception":false,"start_time":"2024-09-30T17:27:55.332538","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:03.222614Z","iopub.execute_input":"2024-10-04T16:51:03.223242Z","iopub.status.idle":"2024-10-04T16:51:04.855933Z","shell.execute_reply.started":"2024-10-04T16:51:03.223191Z","shell.execute_reply":"2024-10-04T16:51:04.854349Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a subset of the training data where 'sii' labels are available\ntrain_df_subset = train_df[train_df['sii'].notna()]\ntrain_df_subset.drop(columns=drop_cols, inplace=True)  # Drop unnecessary columns","metadata":{"papermill":{"duration":0.02448,"end_time":"2024-09-30T17:27:56.624013","exception":false,"start_time":"2024-09-30T17:27:56.599533","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:04.857905Z","iopub.execute_input":"2024-10-04T16:51:04.858372Z","iopub.status.idle":"2024-10-04T16:51:04.872178Z","shell.execute_reply.started":"2024-10-04T16:51:04.858327Z","shell.execute_reply":"2024-10-04T16:51:04.870656Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Separate features (X) and labels (y) for training","metadata":{"papermill":{"duration":0.010629,"end_time":"2024-09-30T17:27:56.645824","exception":false,"start_time":"2024-09-30T17:27:56.635195","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Separate features (X) and labels (y) for training\nX_train = train_df_subset.iloc[:, :-1]  # Features\ny = train_df_subset['sii'].astype('int')  # Target (label 'sii')\nX_test = test_df  # Test set (no labels)","metadata":{"papermill":{"duration":0.020399,"end_time":"2024-09-30T17:27:56.677018","exception":false,"start_time":"2024-09-30T17:27:56.656619","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:04.873511Z","iopub.execute_input":"2024-10-04T16:51:04.873921Z","iopub.status.idle":"2024-10-04T16:51:04.882894Z","shell.execute_reply.started":"2024-10-04T16:51:04.873883Z","shell.execute_reply":"2024-10-04T16:51:04.881268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Identify numerical and categorical columns using 'data_dictionary.csv'","metadata":{"papermill":{"duration":0.010414,"end_time":"2024-09-30T17:27:56.698328","exception":false,"start_time":"2024-09-30T17:27:56.687914","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Identify numerical and categorical columns using 'data_dictionary.csv'\nnumerical_cols = data_dict[\n    (data_dict['Type'] == 'float') | (data_dict['Type'] == 'int')\n]['Field'].tolist()\nnumerical_cols.remove('PCIAT-PCIAT_Total')  # Remove a specific column not in test data","metadata":{"papermill":{"duration":0.021103,"end_time":"2024-09-30T17:27:56.731553","exception":false,"start_time":"2024-09-30T17:27:56.710450","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:04.885159Z","iopub.execute_input":"2024-10-04T16:51:04.885737Z","iopub.status.idle":"2024-10-04T16:51:04.896158Z","shell.execute_reply.started":"2024-10-04T16:51:04.885684Z","shell.execute_reply":"2024-10-04T16:51:04.894734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Categorical columns also contain 'PCIAT', which needs to be removed","metadata":{"papermill":{"duration":0.010549,"end_time":"2024-09-30T17:27:56.753154","exception":false,"start_time":"2024-09-30T17:27:56.742605","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Categorical columns also contain 'PCIAT', which needs to be removed\ncategorical_cols = data_dict[data_dict['Type'] == 'categorical int']['Field'].tolist()\ncategorical_cols = [col for col in categorical_cols if col not in drop_cols]  # Filter out unnecessary columns","metadata":{"papermill":{"duration":0.01968,"end_time":"2024-09-30T17:27:56.783976","exception":false,"start_time":"2024-09-30T17:27:56.764296","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:04.897782Z","iopub.execute_input":"2024-10-04T16:51:04.898201Z","iopub.status.idle":"2024-10-04T16:51:04.909625Z","shell.execute_reply.started":"2024-10-04T16:51:04.898156Z","shell.execute_reply":"2024-10-04T16:51:04.908179Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot distribution of numerical columns","metadata":{"papermill":{"duration":0.011169,"end_time":"2024-09-30T17:27:56.806041","exception":false,"start_time":"2024-09-30T17:27:56.794872","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Plot distribution of numerical columns\nX_train[numerical_cols].hist(figsize=(12, 10), bins=20, color='steelblue', edgecolor='black')\nplt.suptitle('Histograms of Numerical Features', fontsize=16)\nplt.show()","metadata":{"papermill":{"duration":5.889044,"end_time":"2024-09-30T17:28:02.705915","exception":false,"start_time":"2024-09-30T17:27:56.816871","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:04.911421Z","iopub.execute_input":"2024-10-04T16:51:04.911911Z","iopub.status.idle":"2024-10-04T16:51:13.258888Z","shell.execute_reply.started":"2024-10-04T16:51:04.911868Z","shell.execute_reply":"2024-10-04T16:51:13.257322Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Boxplot of numerical columns","metadata":{"papermill":{"duration":0.013166,"end_time":"2024-09-30T17:28:02.732022","exception":false,"start_time":"2024-09-30T17:28:02.718856","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Boxplot of numerical columns\nplt.figure(figsize=(12, 6))\nsns.boxplot(data=X_train[numerical_cols], orient=\"h\")\nplt.title('Boxplots of Numerical Features')\nplt.show()","metadata":{"papermill":{"duration":0.769986,"end_time":"2024-09-30T17:28:03.514984","exception":false,"start_time":"2024-09-30T17:28:02.744998","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:13.260786Z","iopub.execute_input":"2024-10-04T16:51:13.261288Z","iopub.status.idle":"2024-10-04T16:51:14.377713Z","shell.execute_reply.started":"2024-10-04T16:51:13.261235Z","shell.execute_reply":"2024-10-04T16:51:14.375991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print numerical and categorical columns (for debugging purposes)\nprint(numerical_cols)\nprint(categorical_cols)","metadata":{"papermill":{"duration":0.022959,"end_time":"2024-09-30T17:28:03.552227","exception":false,"start_time":"2024-09-30T17:28:03.529268","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:14.379567Z","iopub.execute_input":"2024-10-04T16:51:14.380031Z","iopub.status.idle":"2024-10-04T16:51:14.387885Z","shell.execute_reply.started":"2024-10-04T16:51:14.379986Z","shell.execute_reply":"2024-10-04T16:51:14.386253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function to cap numerical values to the specified quantiles","metadata":{"papermill":{"duration":0.013641,"end_time":"2024-09-30T17:28:03.579950","exception":false,"start_time":"2024-09-30T17:28:03.566309","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Function to cap numerical values to the specified quantiles\ndef clipped(X):\n    X_clipped = X.copy()\n    X_clipped.clip(\n        lower=X_clipped.quantile(0.05),  # Cap lower values at the 5th percentile\n        upper=X_clipped.quantile(0.95),  # Cap upper values at the 95th percentile\n        axis=1\n    )\n    return X_clipped\n\n# Function to process the data (for both training and test sets)\ndef processing(X):\n    X_numerical = X[numerical_cols]  # Subset for numerical columns\n    X_categorical = X[categorical_cols]  # Subset for categorical columns\n\n    # Fix erroneous zero values in numerical columns by replacing them with mean values\n    X_numerical = clipped(X_numerical)  # Apply clipping to limit outliers\n    X_numerical.replace(0, X_numerical.mean(axis=0), inplace=True)  # Replace 0 with column mean\n\n    # ColumnTransformer for handling different data types\n    ct = ColumnTransformer(\n        transformers=[\n            ('ohe', OneHotEncoder(sparse_output=False), [categorical_cols[0]]),  # One-hot encode the 'sex' column\n            ('imp_cat', KNNImputer(n_neighbors=2, weights='uniform'), categorical_cols[1:])  # Impute missing categorical values\n        ]\n    )\n\n    # Apply transformation to categorical columns\n    ct.set_output(transform='pandas')\n    X_categorical = ct.fit_transform(X_categorical).astype('int')  # Ensure categorical data is integer-encoded\n\n    # Pipeline to process numerical columns (e.g., imputation)\n    pipeline = Pipeline([\n        ('imp_num', KNNImputer(n_neighbors=2, weights='uniform')),  # Impute missing values in numerical columns\n    ])\n    pipeline.set_output(transform='pandas')\n    X_numerical = pipeline.fit_transform(X_numerical)  # Apply the pipeline to numerical columns\n\n    # Combine processed numerical and categorical data\n    X_processed = X_categorical.join(X_numerical)\n    return X_processed\n\n# Custom evaluation metric (Quadratic Weighted Kappa)\ndef quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')","metadata":{"papermill":{"duration":0.026857,"end_time":"2024-09-30T17:28:03.621631","exception":false,"start_time":"2024-09-30T17:28:03.594774","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:14.390620Z","iopub.execute_input":"2024-10-04T16:51:14.391135Z","iopub.status.idle":"2024-10-04T16:51:14.404978Z","shell.execute_reply.started":"2024-10-04T16:51:14.391091Z","shell.execute_reply":"2024-10-04T16:51:14.403425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Hyperparameter tuning with larger search space","metadata":{"papermill":{"duration":0.014482,"end_time":"2024-09-30T17:28:03.650364","exception":false,"start_time":"2024-09-30T17:28:03.635882","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# Hyperparameter tuning with larger search space\ndef ParamSearchModels(X, y):\n    algorithms = {\n        'XGBoost': {\n            'model': XGBClassifier(),\n            'params': {\n                'n_estimators': [50, 100, 200, 300],  # Increased range of estimators\n                'max_depth': [3, 5, 7, 10],  # Trying more depths\n                'learning_rate': [0.01, 0.05, 0.1, 0.2],  # Trying smaller learning rates\n                'gamma': [0.01, 0.05, 0.1],  # Gamma for controlling regularization\n                'subsample': [0.8, 0.9, 1],  # Subsampling for robustness\n                'colsample_bytree': [0.7, 0.8, 1],  # Column subsampling\n                'alpha': [1, 5, 10, 20],  # L1 regularization\n                'lambda': [1, 5, 10]  # L2 regularization\n            }\n        },\n        # You can add other models like LightGBM or CatBoost here\n    }\n\n    scores = []\n    QWK_scorer = make_scorer(quadratic_weighted_kappa, greater_is_better=True)\n    cv = ShuffleSplit(n_splits=5, test_size=0.2)\n\n    for algorithm, config in algorithms.items():\n        GS = GridSearchCV(config['model'], config['params'], cv=cv, return_train_score=False, scoring=QWK_scorer, n_jobs=-1)\n        GS.fit(X, y)\n        scores.append({\n            'model': algorithm,\n            'best_score': GS.best_score_,\n            'best_params': GS.best_params_\n        })\n\n    return pd.DataFrame(scores, columns=['model', 'best_score', 'best_params'])\n\n# Preprocess the training and test data\nX_train_processed = processing(X_train)\nX_test_processed = processing(X_test)\n\n# Update hyperparameters and retrain the model\nparams = {\n    'alpha': 10,  # L1 regularization\n    'lambda': 10,  # L2 regularization\n    'learning_rate': 0.05,  # Lower learning rate for better generalization\n    'gamma': 0.1,  # Regularization parameter\n    'max_depth': 7,  # Moderate tree depth to prevent overfitting\n    'n_estimators': 200,  # More trees for better learning\n    'subsample': 0.9,  # Subsampling to avoid overfitting\n    'colsample_bytree': 0.8  # Column subsampling to prevent overfitting\n}\n","metadata":{"papermill":{"duration":2.083558,"end_time":"2024-09-30T17:28:05.748246","exception":false,"start_time":"2024-09-30T17:28:03.664688","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:14.407077Z","iopub.execute_input":"2024-10-04T16:51:14.407693Z","iopub.status.idle":"2024-10-04T16:51:17.139364Z","shell.execute_reply.started":"2024-10-04T16:51:14.407639Z","shell.execute_reply":"2024-10-04T16:51:17.137966Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# XGBoost classifier with new parameters","metadata":{"papermill":{"duration":0.013662,"end_time":"2024-09-30T17:28:05.776037","exception":false,"start_time":"2024-09-30T17:28:05.762375","status":"completed"},"tags":[]}},{"cell_type":"code","source":"# XGBoost classifier with new parameters\nxgb_clf = XGBClassifier(**params)\nxgb_clf.fit(X_train_processed, y)","metadata":{"papermill":{"duration":4.703056,"end_time":"2024-09-30T17:28:10.492885","exception":false,"start_time":"2024-09-30T17:28:05.789829","status":"completed"},"tags":[],"execution":{"iopub.status.busy":"2024-10-04T16:51:17.140674Z","iopub.execute_input":"2024-10-04T16:51:17.141077Z","iopub.status.idle":"2024-10-04T16:51:19.441775Z","shell.execute_reply.started":"2024-10-04T16:51:17.141037Z","shell.execute_reply":"2024-10-04T16:51:19.440534Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"param_grid_xgb = {\n    'n_estimators': [100, 200, 300],\n    'max_depth': [5, 10, 15],\n    'learning_rate': [0.01, 0.05, 0.1],\n    'subsample': [0.7, 0.8, 1.0],\n    'colsample_bytree': [0.7, 0.8, 1.0],\n    'min_child_weight': [1, 3, 5],\n    'gamma': [0, 0.1, 0.2],\n    'alpha': [0, 1, 5],\n    'lambda': [1, 2, 3]\n}\n\n# Hyperparameter search using RandomizedSearchCV\nfrom sklearn.model_selection import RandomizedSearchCV\nxgb_clf = XGBClassifier()\nrandom_search = RandomizedSearchCV(\n    xgb_clf, param_distributions=param_grid_xgb, n_iter=100, scoring='neg_log_loss', cv=5, n_jobs=-1, verbose=3\n)\nrandom_search.fit(X_train_processed, y)\nbest_params = random_search.best_params_\nprint(f\"Best hyperparameters found: {best_params}\")","metadata":{"execution":{"iopub.status.busy":"2024-10-04T16:51:19.443192Z","iopub.execute_input":"2024-10-04T16:51:19.443667Z","iopub.status.idle":"2024-10-04T17:07:54.042442Z","shell.execute_reply.started":"2024-10-04T16:51:19.443617Z","shell.execute_reply":"2024-10-04T17:07:54.041143Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import StackingClassifier\nfrom sklearn.linear_model import LogisticRegression\n\n# Define base learners\nestimators = [\n    ('xgb', XGBClassifier(**params)),\n    ('lgbm', LGBMClassifier(n_estimators=100)),\n    ('catboost', CatBoostClassifier(verbose=0))\n]\n\n# Create the stacking model with a Logistic Regression meta-learner\nstacking_model = StackingClassifier(\n    estimators=estimators, \n    final_estimator=LogisticRegression(),\n    cv=5\n)\n\n# Fit the stacking model\nstacking_model.fit(X_train_processed, y)\n\n# Make predictions\nstacking_preds = stacking_model.predict(X_test_processed)\n\n# Prepare and save the submission file\nsubmission = pd.DataFrame({'id': X_test.index, 'sii': stacking_preds})\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2024-10-04T17:07:54.047971Z","iopub.execute_input":"2024-10-04T17:07:54.049413Z","iopub.status.idle":"2024-10-04T17:09:42.762429Z","shell.execute_reply.started":"2024-10-04T17:07:54.049349Z","shell.execute_reply":"2024-10-04T17:09:42.761114Z"},"trusted":true},"execution_count":null,"outputs":[]}]}