{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"\n### 1. **Importing Necessary Libraries**\n   - Libraries like `numpy`, `pandas` are used for data manipulation.\n   - `re` is for regular expressions, which might be useful for string operations.\n   - `matplotlib.pyplot` and `seaborn` are used for visualizations.\n   - Machine learning libraries from `sklearn` like `train_test_split`, `Pipeline`, and model evaluation metrics like `confusion_matrix` are used for model training and validation.\n   - Libraries like `XGBoost`, `LightGBM`, and `CatBoost` are imported for different classifier models.\n   - `warnings` is imported to suppress warning messages for cleaner output.\n\n### 2. **Setting Up Matplotlib and Seaborn for Visualizations**\n   - Configure `matplotlib` for LaTeX-style text rendering and set font to `serif`.\n   - Use `sns.set_theme()` to set the Seaborn theme for consistent and polished plots.\n\n### 3. **Loading Data**\n   - Use `pd.read_csv()` to load training, testing, and a data dictionary CSV files from a specified folder path.\n   - The `train.csv` file contains training data with an 'sii' target variable.\n   - The `test.csv` file contains test data for predictions.\n\n### 4. **Visualizing Target Variable ('sii') Distribution**\n   - A count plot (`sns.countplot`) is created to visualize the distribution of the target variable ('sii') in the training data. This helps understand the class imbalance if any.\n\n### 5. **Identifying Missing Columns in the Test Set**\n   - Compare the columns of the training and test sets to identify missing columns in the test set.\n   - The missing columns (usually the target variable 'sii' or others) are removed before model training.\n\n### 6. **Visualizing Missing Data**\n   - A heatmap (`sns.heatmap`) is plotted to visualize the missing values in the training data. This helps identify where imputations are needed.\n\n### 7. **Handling Missing Data and Dropping Unnecessary Columns**\n   - The subset of the training data without missing 'sii' labels is selected for training.\n   - Unnecessary columns that don't appear in both train and test datasets are dropped.\n\n### 8. **Visualizing the Distribution of Numerical Features**\n   - Histograms (`X_train[numerical_cols].hist()`) are plotted for numerical features in the training data to check for distributions and potential outliers.\n   - Boxplots (`sns.boxplot`) are plotted to visualize the spread and outliers in the numerical features.\n\n### 9. **Outlier Handling via Clipping Function**\n   - A function `clipped()` is created to clip outliers based on quantiles. This reduces the impact of extreme values on the model performance.\n\n### 10. **Processing Function for Data Preprocessing**\n   - The `processing()` function handles both numerical and categorical features:\n     - **Numerical Features**: Outliers are clipped, and missing values are replaced with the mean.\n     - **Categorical Features**: One-hot encoding is applied for categorical variables, and `KNNImputer` is used to handle missing values in categorical features.\n   - The `ColumnTransformer` is used to apply transformations on categorical and numerical columns in the dataset.\n\n### 11. **Visualizing Feature Correlation**\n   - A heatmap of the correlation matrix (`sns.heatmap(corr)`) is plotted for the numerical features. This helps understand relationships between features and detect multicollinearity.\n\n### 12. **Processing Training Data**\n   - The training data is processed using the `processing()` function which applies imputations, outlier handling, and transformations.\n\n### 13. **Visualizing Processed Features' Distribution**\n   - After preprocessing, histograms are plotted for the processed features to confirm if the distributions are normalized and any issues like missing values or skewness are handled.\n\n### 14. **Quadratic Weighted Kappa (QWK) Score Function**\n   - A custom metric `quadratic_weighted_kappa()` is defined to calculate the quadratic weighted kappa score, which is useful for evaluating models on ordinal classification tasks (like predicting 'sii').\n\n### 15. **Model Hyperparameter Search Using GridSearchCV**\n   - A function `ParamSearchModels()` is defined to perform hyperparameter tuning using GridSearchCV on several models:\n     - Models like `XGBClassifier` are initialized with a grid of hyperparameters.\n     - The model is trained and validated using cross-validation (`ShuffleSplit`), with the QWK score as the evaluation metric.\n     - The best hyperparameters and scores are recorded for each model.\n\n### 16. **Training the Model Using XGBoost**\n   - The XGBoost classifier is trained using the best parameters obtained from the grid search. This model is used for predictions on the test set.\n\n### 17. **Processing Test Data**\n   - The test data is processed using the same `processing()` function to ensure that transformations applied to the training data are also applied to the test data.\n\n### 18. **Making Predictions**\n   - Predictions on the processed test data are made using the trained XGBoost model. These predictions are for the target variable 'sii'.\n\n### 19. **Visualizing Predictions**\n   - A histogram (`sns.histplot`) of the predicted 'sii' values is plotted to visualize the distribution of predictions and ensure they align with expectations.\n\n### 20. **Preparing the Submission File**\n   - A submission file is prepared for the Kaggle competition. The test data predictions are saved in a CSV format with columns 'id' and 'sii'.\n   - The first few rows of the submission file are displayed using `print()` for verification.\n","metadata":{}},{"cell_type":"markdown","source":"# Import necessary libraries","metadata":{}},{"cell_type":"code","source":"# Import necessary libraries\nimport numpy as np  # For numerical operations\nimport pandas as pd  # For handling dataframes\nimport re  # Regular expressions, though unused here\nimport matplotlib.pyplot as plt  # For plotting and visualizations\nimport seaborn as sns  # For advanced visualizations","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-09-30T17:10:06.646504Z","iopub.execute_input":"2024-09-30T17:10:06.646930Z","iopub.status.idle":"2024-09-30T17:10:07.685938Z","shell.execute_reply.started":"2024-09-30T17:10:06.646886Z","shell.execute_reply":"2024-09-30T17:10:07.684770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Configure matplotlib settings for better plots","metadata":{}},{"cell_type":"code","source":"# Configure matplotlib settings for better plots\nplt.rc('text', usetex=True)  # Enables LaTeX rendering for text in plots\nplt.rc('font', family='serif')  # Sets font to 'serif'\nplt.rc('xtick', labelsize=12)  # Sets font size for x-axis labels\nplt.rc('ytick', labelsize=12)  # Sets font size for y-axis labels","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:07.688277Z","iopub.execute_input":"2024-09-30T17:10:07.688919Z","iopub.status.idle":"2024-09-30T17:10:07.695160Z","shell.execute_reply.started":"2024-09-30T17:10:07.688866Z","shell.execute_reply":"2024-09-30T17:10:07.694075Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Set Seaborn theme for consistent style\nsns.set_theme()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:07.696631Z","iopub.execute_input":"2024-09-30T17:10:07.697045Z","iopub.status.idle":"2024-09-30T17:10:07.706593Z","shell.execute_reply.started":"2024-09-30T17:10:07.696995Z","shell.execute_reply":"2024-09-30T17:10:07.705641Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Ignore warnings for cleaner output","metadata":{}},{"cell_type":"code","source":"# Ignore warnings for cleaner output\nimport warnings\nwarnings.filterwarnings('ignore')","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:07.709253Z","iopub.execute_input":"2024-09-30T17:10:07.709614Z","iopub.status.idle":"2024-09-30T17:10:07.716675Z","shell.execute_reply.started":"2024-09-30T17:10:07.709576Z","shell.execute_reply":"2024-09-30T17:10:07.715598Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Import necessary modules from sklearn for data preprocessing, model evaluation, and training","metadata":{}},{"cell_type":"code","source":"# Import necessary modules from sklearn for data preprocessing, model evaluation, and training\nfrom sklearn.impute import SimpleImputer, KNNImputer  # For missing data imputation\nfrom sklearn.compose import ColumnTransformer  # For applying different transformations to specific columns\nfrom sklearn.pipeline import Pipeline  # To create pipelines for consistent data transformations\nfrom sklearn.preprocessing import OneHotEncoder, LabelEncoder, StandardScaler  # For encoding and scaling\nfrom sklearn.model_selection import train_test_split, cross_val_score, ShuffleSplit, GridSearchCV  # For model evaluation and grid search\nfrom xgboost import XGBClassifier  # XGBoost Classifier\nfrom lightgbm import LGBMClassifier  # LightGBM Classifier\nfrom catboost import CatBoostClassifier  # CatBoost Classifier\nfrom sklearn.metrics import confusion_matrix, make_scorer, cohen_kappa_score  # For model evaluation metrics","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:07.718150Z","iopub.execute_input":"2024-09-30T17:10:07.718529Z","iopub.status.idle":"2024-09-30T17:10:08.888570Z","shell.execute_reply.started":"2024-09-30T17:10:07.718491Z","shell.execute_reply":"2024-09-30T17:10:08.887557Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Define folder path where the dataset is stored","metadata":{}},{"cell_type":"code","source":"# Define folder path where the dataset is stored\nfolder_path = '/kaggle/input/child-mind-institute-problematic-internet-use/'","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:08.889898Z","iopub.execute_input":"2024-09-30T17:10:08.890516Z","iopub.status.idle":"2024-09-30T17:10:08.895367Z","shell.execute_reply.started":"2024-09-30T17:10:08.890474Z","shell.execute_reply":"2024-09-30T17:10:08.894198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Load training and test datasets, with 'id' as the index column","metadata":{}},{"cell_type":"code","source":"# Load training and test datasets, with 'id' as the index column\ntrain_df = pd.read_csv(folder_path + 'train.csv', index_col='id')\ntest_df = pd.read_csv(folder_path + 'test.csv', index_col='id')\ndata_dict = pd.read_csv(folder_path + 'data_dictionary.csv')  # Load data dictionary","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:08.896913Z","iopub.execute_input":"2024-09-30T17:10:08.897401Z","iopub.status.idle":"2024-09-30T17:10:08.999378Z","shell.execute_reply.started":"2024-09-30T17:10:08.897343Z","shell.execute_reply":"2024-09-30T17:10:08.998222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize class distribution of target variable 'sii'","metadata":{}},{"cell_type":"code","source":"import matplotlib as mpl\nmpl.rcParams['text.usetex'] = False","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:09.000767Z","iopub.execute_input":"2024-09-30T17:10:09.001157Z","iopub.status.idle":"2024-09-30T17:10:09.006103Z","shell.execute_reply.started":"2024-09-30T17:10:09.001119Z","shell.execute_reply":"2024-09-30T17:10:09.005009Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Visualize class distribution of target variable 'sii'\nplt.figure(figsize=(6, 4))\nsns.countplot(data=train_df, x='sii')\nplt.title('Class Distribution of SII')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:09.007518Z","iopub.execute_input":"2024-09-30T17:10:09.007850Z","iopub.status.idle":"2024-09-30T17:10:09.336412Z","shell.execute_reply.started":"2024-09-30T17:10:09.007814Z","shell.execute_reply":"2024-09-30T17:10:09.335253Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Training set has additional columns (like 'PCIAT') compared to test set\n# We find the extra columns in the training data that need to be removed before training\ntrain_cols = train_df.columns.tolist()  # List of columns in training data\ntest_cols = test_df.columns.tolist()  # List of columns in test data","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:09.341135Z","iopub.execute_input":"2024-09-30T17:10:09.341501Z","iopub.status.idle":"2024-09-30T17:10:09.346870Z","shell.execute_reply.started":"2024-09-30T17:10:09.341463Z","shell.execute_reply":"2024-09-30T17:10:09.345795Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"missing_cols = set(train_cols) - set(test_cols)\nprint(f\"Columns missing in test set: {missing_cols}\")","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:09.348487Z","iopub.execute_input":"2024-09-30T17:10:09.348892Z","iopub.status.idle":"2024-09-30T17:10:09.359171Z","shell.execute_reply.started":"2024-09-30T17:10:09.348840Z","shell.execute_reply":"2024-09-30T17:10:09.357681Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Find columns in training data that are not in test data\ndrop_cols = list(set(train_cols) - set(test_cols))  \ndrop_cols.remove('sii')  # Do not remove the target variable 'sii'","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:09.360601Z","iopub.execute_input":"2024-09-30T17:10:09.361063Z","iopub.status.idle":"2024-09-30T17:10:09.369439Z","shell.execute_reply.started":"2024-09-30T17:10:09.360997Z","shell.execute_reply":"2024-09-30T17:10:09.368288Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Visualize missing data in the training set","metadata":{}},{"cell_type":"code","source":"# Visualize missing data in the training set\nplt.figure(figsize=(10, 6))\nsns.heatmap(train_df.isnull(), cbar=False, cmap='viridis')\nplt.title('Missing Data in Training Set')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:09.370824Z","iopub.execute_input":"2024-09-30T17:10:09.371218Z","iopub.status.idle":"2024-09-30T17:10:10.818534Z","shell.execute_reply.started":"2024-09-30T17:10:09.371180Z","shell.execute_reply":"2024-09-30T17:10:10.816785Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Create a subset of the training data where 'sii' labels are available\ntrain_df_subset = train_df[train_df['sii'].notna()]\ntrain_df_subset.drop(columns=drop_cols, inplace=True)  # Drop unnecessary columns","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:10.819965Z","iopub.execute_input":"2024-09-30T17:10:10.820343Z","iopub.status.idle":"2024-09-30T17:10:10.830533Z","shell.execute_reply.started":"2024-09-30T17:10:10.820305Z","shell.execute_reply":"2024-09-30T17:10:10.829518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Separate features (X) and labels (y) for training","metadata":{}},{"cell_type":"code","source":"# Separate features (X) and labels (y) for training\nX_train = train_df_subset.iloc[:, :-1]  # Features\ny = train_df_subset['sii'].astype('int')  # Target (label 'sii')\nX_test = test_df  # Test set (no labels)","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:10.832104Z","iopub.execute_input":"2024-09-30T17:10:10.832865Z","iopub.status.idle":"2024-09-30T17:10:10.840862Z","shell.execute_reply.started":"2024-09-30T17:10:10.832814Z","shell.execute_reply":"2024-09-30T17:10:10.839783Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Identify numerical and categorical columns using 'data_dictionary.csv'","metadata":{}},{"cell_type":"code","source":"# Identify numerical and categorical columns using 'data_dictionary.csv'\nnumerical_cols = data_dict[\n    (data_dict['Type'] == 'float') | (data_dict['Type'] == 'int')\n]['Field'].tolist()\nnumerical_cols.remove('PCIAT-PCIAT_Total')  # Remove a specific column not in test data","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:10.842236Z","iopub.execute_input":"2024-09-30T17:10:10.843322Z","iopub.status.idle":"2024-09-30T17:10:10.853104Z","shell.execute_reply.started":"2024-09-30T17:10:10.843267Z","shell.execute_reply":"2024-09-30T17:10:10.852109Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Categorical columns also contain 'PCIAT', which needs to be removed","metadata":{}},{"cell_type":"code","source":"# Categorical columns also contain 'PCIAT', which needs to be removed\ncategorical_cols = data_dict[data_dict['Type'] == 'categorical int']['Field'].tolist()\ncategorical_cols = [col for col in categorical_cols if col not in drop_cols]  # Filter out unnecessary columns","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:10.854489Z","iopub.execute_input":"2024-09-30T17:10:10.854846Z","iopub.status.idle":"2024-09-30T17:10:10.862409Z","shell.execute_reply.started":"2024-09-30T17:10:10.854809Z","shell.execute_reply":"2024-09-30T17:10:10.861431Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plot distribution of numerical columns","metadata":{}},{"cell_type":"code","source":"# Plot distribution of numerical columns\nX_train[numerical_cols].hist(figsize=(12, 10), bins=20, color='steelblue', edgecolor='black')\nplt.suptitle('Histograms of Numerical Features', fontsize=16)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:10.864043Z","iopub.execute_input":"2024-09-30T17:10:10.864824Z","iopub.status.idle":"2024-09-30T17:10:18.854837Z","shell.execute_reply.started":"2024-09-30T17:10:10.864773Z","shell.execute_reply":"2024-09-30T17:10:18.853738Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Boxplot of numerical columns","metadata":{}},{"cell_type":"code","source":"# Boxplot of numerical columns\nplt.figure(figsize=(12, 6))\nsns.boxplot(data=X_train[numerical_cols], orient=\"h\")\nplt.title('Boxplots of Numerical Features')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:18.856180Z","iopub.execute_input":"2024-09-30T17:10:18.856508Z","iopub.status.idle":"2024-09-30T17:10:19.885986Z","shell.execute_reply.started":"2024-09-30T17:10:18.856473Z","shell.execute_reply":"2024-09-30T17:10:19.884971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Print numerical and categorical columns (for debugging purposes)\nprint(numerical_cols)\nprint(categorical_cols)","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:19.887384Z","iopub.execute_input":"2024-09-30T17:10:19.887731Z","iopub.status.idle":"2024-09-30T17:10:19.893680Z","shell.execute_reply.started":"2024-09-30T17:10:19.887696Z","shell.execute_reply":"2024-09-30T17:10:19.892478Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Function to cap numerical values to the specified quantiles","metadata":{}},{"cell_type":"code","source":"# Function to cap numerical values to the specified quantiles\ndef clipped(X):\n    X_clipped = X.copy()\n    X_clipped.clip(\n        lower=X_clipped.quantile(0.05),  # Cap lower values at the 5th percentile\n        upper=X_clipped.quantile(0.95),  # Cap upper values at the 95th percentile\n        axis=1\n    )\n    return X_clipped\n\n# Function to process the data (for both training and test sets)\ndef processing(X):\n    X_numerical = X[numerical_cols]  # Subset for numerical columns\n    X_categorical = X[categorical_cols]  # Subset for categorical columns\n\n    # Fix erroneous zero values in numerical columns by replacing them with mean values\n    X_numerical = clipped(X_numerical)  # Apply clipping to limit outliers\n    X_numerical.replace(0, X_numerical.mean(axis=0), inplace=True)  # Replace 0 with column mean\n\n    # ColumnTransformer for handling different data types\n    ct = ColumnTransformer(\n        transformers=[\n            ('ohe', OneHotEncoder(sparse_output=False), [categorical_cols[0]]),  # One-hot encode the 'sex' column\n            ('imp_cat', KNNImputer(n_neighbors=2, weights='uniform'), categorical_cols[1:])  # Impute missing categorical values\n        ]\n    )\n\n    # Apply transformation to categorical columns\n    ct.set_output(transform='pandas')\n    X_categorical = ct.fit_transform(X_categorical).astype('int')  # Ensure categorical data is integer-encoded\n\n    # Pipeline to process numerical columns (e.g., imputation)\n    pipeline = Pipeline([\n        ('imp_num', KNNImputer(n_neighbors=2, weights='uniform')),  # Impute missing values in numerical columns\n    ])\n    pipeline.set_output(transform='pandas')\n    X_numerical = pipeline.fit_transform(X_numerical)  # Apply the pipeline to numerical columns\n\n    # Combine processed numerical and categorical data\n    X_processed = X_categorical.join(X_numerical)\n    return X_processed\n\n# Custom evaluation metric (Quadratic Weighted Kappa)\ndef quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:19.895427Z","iopub.execute_input":"2024-09-30T17:10:19.895790Z","iopub.status.idle":"2024-09-30T17:10:19.908324Z","shell.execute_reply.started":"2024-09-30T17:10:19.895753Z","shell.execute_reply":"2024-09-30T17:10:19.907244Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Hyperparameter tuning with larger search space","metadata":{}},{"cell_type":"code","source":"# Hyperparameter tuning with larger search space\ndef ParamSearchModels(X, y):\n    algorithms = {\n        'XGBoost': {\n            'model': XGBClassifier(),\n            'params': {\n                'n_estimators': [50, 100, 200, 300],  # Increased range of estimators\n                'max_depth': [3, 5, 7, 10],  # Trying more depths\n                'learning_rate': [0.01, 0.05, 0.1, 0.2],  # Trying smaller learning rates\n                'gamma': [0.01, 0.05, 0.1],  # Gamma for controlling regularization\n                'subsample': [0.8, 0.9, 1],  # Subsampling for robustness\n                'colsample_bytree': [0.7, 0.8, 1],  # Column subsampling\n                'alpha': [1, 5, 10, 20],  # L1 regularization\n                'lambda': [1, 5, 10]  # L2 regularization\n            }\n        },\n        # You can add other models like LightGBM or CatBoost here\n    }\n\n    scores = []\n    QWK_scorer = make_scorer(quadratic_weighted_kappa, greater_is_better=True)\n    cv = ShuffleSplit(n_splits=5, test_size=0.2)\n\n    for algorithm, config in algorithms.items():\n        GS = GridSearchCV(config['model'], config['params'], cv=cv, return_train_score=False, scoring=QWK_scorer, n_jobs=-1)\n        GS.fit(X, y)\n        scores.append({\n            'model': algorithm,\n            'best_score': GS.best_score_,\n            'best_params': GS.best_params_\n        })\n\n    return pd.DataFrame(scores, columns=['model', 'best_score', 'best_params'])\n\n# Preprocess the training and test data\nX_train_processed = processing(X_train)\nX_test_processed = processing(X_test)\n\n# Update hyperparameters and retrain the model\nparams = {\n    'alpha': 10,  # L1 regularization\n    'lambda': 10,  # L2 regularization\n    'learning_rate': 0.05,  # Lower learning rate for better generalization\n    'gamma': 0.1,  # Regularization parameter\n    'max_depth': 7,  # Moderate tree depth to prevent overfitting\n    'n_estimators': 200,  # More trees for better learning\n    'subsample': 0.9,  # Subsampling to avoid overfitting\n    'colsample_bytree': 0.8  # Column subsampling to prevent overfitting\n}\n","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:19.909942Z","iopub.execute_input":"2024-09-30T17:10:19.910330Z","iopub.status.idle":"2024-09-30T17:10:22.237813Z","shell.execute_reply.started":"2024-09-30T17:10:19.910292Z","shell.execute_reply":"2024-09-30T17:10:22.236911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# XGBoost classifier with new parameters","metadata":{}},{"cell_type":"code","source":"# XGBoost classifier with new parameters\nxgb_clf = XGBClassifier(**params)\nxgb_clf.fit(X_train_processed, y)","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:22.239114Z","iopub.execute_input":"2024-09-30T17:10:22.239459Z","iopub.status.idle":"2024-09-30T17:10:24.281444Z","shell.execute_reply.started":"2024-09-30T17:10:22.239421Z","shell.execute_reply":"2024-09-30T17:10:24.280277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Process the test data and make predictions","metadata":{}},{"cell_type":"code","source":"# Process the test data and make predictions\nX_test_processed = processing(X_test)\ntest_preds = xgb_clf.predict(X_test_processed).tolist()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:24.282878Z","iopub.execute_input":"2024-09-30T17:10:24.283481Z","iopub.status.idle":"2024-09-30T17:10:24.355617Z","shell.execute_reply.started":"2024-09-30T17:10:24.283426Z","shell.execute_reply":"2024-09-30T17:10:24.354629Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Plotting the prediction distribution","metadata":{}},{"cell_type":"code","source":"# Plotting the prediction distribution\nplt.figure(figsize=(6, 4))\nsns.histplot(test_preds, bins=10, kde=False, color='orange')\nplt.title('Test Predictions Distribution')\nplt.xlabel('Predicted SII')\nplt.ylabel('Frequency')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:24.356911Z","iopub.execute_input":"2024-09-30T17:10:24.357355Z","iopub.status.idle":"2024-09-30T17:10:24.729731Z","shell.execute_reply.started":"2024-09-30T17:10:24.357315Z","shell.execute_reply":"2024-09-30T17:10:24.728621Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Prepare the submission file","metadata":{}},{"cell_type":"code","source":"# Prepare the submission file\nsubmission = pd.DataFrame({'id': X_test.index, 'sii': test_preds})\nsubmission.to_csv('submission.csv', index=False)\n\n# Display the first few rows of the submission\nprint(submission.head())","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:10:24.731501Z","iopub.execute_input":"2024-09-30T17:10:24.731991Z","iopub.status.idle":"2024-09-30T17:10:24.744797Z","shell.execute_reply.started":"2024-09-30T17:10:24.731918Z","shell.execute_reply":"2024-09-30T17:10:24.743476Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Import plot_importance from xgboost\nfrom xgboost import plot_importance\n\n# Feature importance (optional but helpful for model interpretation)\nplot_importance(xgb_clf)\nplt.title('Feature Importance - XGBoost')\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2024-09-30T17:25:16.775776Z","iopub.execute_input":"2024-09-30T17:25:16.776557Z","iopub.status.idle":"2024-09-30T17:25:17.668114Z","shell.execute_reply.started":"2024-09-30T17:25:16.776511Z","shell.execute_reply":"2024-09-30T17:25:17.667015Z"},"trusted":true},"execution_count":null,"outputs":[]}]}