{"metadata":{"kernelspec":{"name":"python3","display_name":"Python 3","language":"python"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30786,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Child Mind Institute — Problematic Internet Use\n","metadata":{}},{"cell_type":"markdown","source":"1. [Problem Statement](#1) - Statement of the problem and the dataset\n\n2. [Load and Check Data](#2) - Loading the dataset and checking for any missing/duplicate values\n\n3. [Data Preprocessing](#3) - Data Imputation, Standardization, Normalization, Encoding, Feature Selection\n\n4. [Modeling](#4) - Model Selection, Training, and Testing","metadata":{}},{"cell_type":"markdown","source":"## <a id = \"1\">1. Problem Statement</a>\nCan you predict the level of problematic internet usage exhibited by children and adolescents, based on their physical activity? The goal of this competition is to develop a predictive model that analyzes children's physical activity and fitness data to identify early signs of problematic internet use. Identifying these patterns can help trigger interventions to encourage healthier digital habits.\n","metadata":{}},{"cell_type":"code","source":"# Import normal packages\n\nimport os\nimport numpy as np\nimport pandas as pd\npd.set_option('display.max_rows', 150)\nimport seaborn as sns\nimport matplotlib.pyplot as plt\nfrom matplotlib.ticker import PercentFormatter\n# Import Statistics packages\nimport scipy.stats as stats\n# Import Progress bar\nfrom tqdm import tqdm\n# Import concurrent futures (for parallel processing)\nfrom concurrent.futures import ThreadPoolExecutor\n# Import sklearn - preprocessing\nfrom sklearn.preprocessing import OrdinalEncoder\nfrom sklearn.preprocessing import OneHotEncoder\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.preprocessing import Normalizer\nfrom sklearn.preprocessing import MinMaxScaler\n# Import sklearn - model selection\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.model_selection import RandomizedSearchCV\n# Import Model\nimport lightgbm as lgb\nfrom lightgbm import early_stopping, log_evaluation\nimport xgboost as xgb\nfrom xgboost import XGBClassifier\nfrom catboost import CatBoostClassifier\n# Import sklearn - metrics\nfrom sklearn.metrics import cohen_kappa_score\n\n\n# Other Import \nfrom IPython.display import clear_output\nfrom scipy.optimize import minimize\nfrom sklearn.base import clone\nfrom colorama import Fore, Style\nfrom sklearn.model_selection import StratifiedKFold","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:43.520397Z","iopub.execute_input":"2024-10-11T00:19:43.521073Z","iopub.status.idle":"2024-10-11T00:19:43.529740Z","shell.execute_reply.started":"2024-10-11T00:19:43.521035Z","shell.execute_reply":"2024-10-11T00:19:43.528634Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <a id = \"2\">2.Load and Check Data</a>\n> Load the dataset and check for any missing/duplicate values\n\n### 2.0 Introduction to Dataset\n#### 2.0.1 Accelerometer Data (Time Series Data)\n\nThese data are stored in `series_train.parquet` and `series_test.parquet`. Each participant has a separate folder containing accelerometer data recorded from their wrist-worn devices.\n\nEach file contains the following fields:\n\n- **id**: The unique identifier for each participant, corresponding to the `id` field in `train.csv` and `test.csv`.\n- **step**: The time step of the observation, indicating the time sequence of the data.\n- **X, Y, Z**: Accelerometer values on the three axes, measured in units of g (gravitational acceleration).\n- **enmo**: The Euclidean Norm Minus One, representing the intensity of acceleration; negative values are set to zero, and a value of 0 indicates stillness.\n- **anglez**: The angle of the arm relative to the horizontal plane.\n- **non-wear_flag**: Flag indicating whether the wrist device was worn (0 = worn, 1 = not worn).\n- **light**: Ambient light intensity in lux.\n- **battery_voltage**: Battery voltage in millivolts (mV).\n- **time_of_day**: The specific time window of data sampling (5-second windows).\n- **weekday**: The day of the week, where 1 = Monday and 7 = Sunday.\n- **quarter**: The quarter of the year (1-4).\n- **relative_date_PCIAT**: The number of days relative to the PCIAT test, with negative values indicating data collected before the test.\n\n#### Usage:\n\n- **Feature Extraction**: These time series data need to be transformed into features for model training. For example, you can extract the average acceleration, maximum acceleration, periods of activity, periods of stillness, etc.\n- **Activity Pattern Recognition**: These data can be used to identify activity patterns and determine whether there is excessive sedentary time, which may relate to problematic internet use.\n\n#### 2.0.2 Tabular Data (`train.csv` and `test.csv`)\n\nThis dataset contains most of the non-time-series data, detailing the participants' background, physical measurements, internet use behavior, etc. This data is more straightforward for direct input into machine learning models.\n\nKey fields include:\n\n- **id**: The unique identifier for each participant, matching the `id` in the accelerometer data.\n- **sii**: The target variable representing the severity of problematic internet use (only available in `train.csv`), with the following values:\n  - 0: None\n  - 1: Mild\n  - 2: Moderate\n  - 3: Severe\n\n#### Demographics:\n\n- **age**: The participant's age.\n- **sex**: The participant's gender.\n\n#### Internet Use:\n\n- **internet_hours**: The number of hours per day spent using the internet/computer.\n\n#### Children's Global Assessment Scale (CGAS):\n\n- A measure for evaluating functional status in adolescents under 18 years old.\n\n#### Physical Measures:\n\n- Includes blood pressure, heart rate, height, weight, waist circumference, hip circumference, etc.\n\n#### FitnessGram Vitals and Treadmill:\n\n- Cardiovascular health measurements.\n\n#### FitnessGram Child:\n\n- Health-related fitness parameters in children, including cardiovascular endurance, muscle strength, muscular endurance, flexibility, and body composition.\n\n#### Bio-electric Impedance Analysis:\n\n- Body composition data, including BMI, body fat, muscle mass, and water content.\n\n#### Physical Activity Questionnaire:\n\n- Information on recent participation in vigorous physical activities in the past 7 days.\n\n#### Sleep Disturbance Scale:\n\n- Used to classify sleep disorders in children.\n\n#### Parent-Child Internet Addiction Test (PCIAT):\n\nThis questionnaire consists of 20 questions measuring addictive behaviors, including compulsive use, reality avoidance, dependency, etc.\n\n- **PCIAT_Total**: This field summarizes the scores of all questions and is used to derive the `sii` target variable in `train.csv`.\n\n#### Usage:\n\n- **Direct Model Input**: Many fields, such as age, gender, height, weight, and internet usage time, can be directly input into models.\n- **Data Cleaning & Processing**: Be mindful of potential missing values, as many fields may not be complete for every participant. You may consider filling or discarding missing values, or applying imputation techniques such as multiple imputation.\n\n\n### 2.1 Load Data\n\n- Read the data (Original Data and Parquet Data)\n\n- Check the data (info, head, shape)","metadata":{}},{"cell_type":"code","source":"# Read data\nfolder_path = r'/kaggle/input/child-mind-institute-problematic-internet-use'\ndata_dictionary_df = pd.read_csv(os.path.join(folder_path, 'data_dictionary.csv'))\ntrain_df = pd.read_csv(os.path.join(folder_path, 'train.csv'))\ntest_df = pd.read_csv(os.path.join(folder_path, 'test.csv'))\nsample_submission_df = pd.read_csv(os.path.join(folder_path, 'sample_submission.csv'))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:43.532369Z","iopub.execute_input":"2024-10-11T00:19:43.532867Z","iopub.status.idle":"2024-10-11T00:19:43.626832Z","shell.execute_reply.started":"2024-10-11T00:19:43.532827Z","shell.execute_reply":"2024-10-11T00:19:43.625790Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check Data Shape\nprint('Train Shape:', train_df.shape)\nprint('Test Shape:', test_df.shape)\n\n# Display data\ntrain_df.head()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:43.628604Z","iopub.execute_input":"2024-10-11T00:19:43.629124Z","iopub.status.idle":"2024-10-11T00:19:43.670884Z","shell.execute_reply.started":"2024-10-11T00:19:43.629000Z","shell.execute_reply":"2024-10-11T00:19:43.669764Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load Parquet Files\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n\n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n\n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    \n    return df\n\noutput_folder_path = r'/kaggle/working/'\nif not os.path.exists(os.path.join(os.path.join(output_folder_path,'train_time_series.csv'))) and not os.path.exists(os.path.join(os.path.join(output_folder_path,'test_time_series.csv'))):\n    # Load Time Series Data\n    train_time_series_df = load_time_series(os.path.join(os.path.join(folder_path,'series_train.parquet/')))\n    test_time_series_df = load_time_series(os.path.join(os.path.join(folder_path,'series_test.parquet/')))\n    # Save Data)\n    train_time_series_df.to_csv(os.path.join(os.path.join(output_folder_path,'train_time_series.csv')), index=False)\n    test_time_series_df.to_csv(os.path.join(os.path.join(output_folder_path,'test_time_series.csv')), index=False)\n\n# Read Data\ntrain_time_series_df = pd.read_csv(os.path.join(os.path.join(output_folder_path,'train_time_series.csv')))\ntest_time_series_df = pd.read_csv(os.path.join(os.path.join(output_folder_path,'test_time_series.csv')))\n# Merge Data\ntrain_df = train_df.merge(train_time_series_df, on='id', how='left')\ntest_df = test_df.merge(test_time_series_df, on='id', how='left')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:43.673353Z","iopub.execute_input":"2024-10-11T00:19:43.673719Z","iopub.status.idle":"2024-10-11T00:19:43.720492Z","shell.execute_reply.started":"2024-10-11T00:19:43.673680Z","shell.execute_reply":"2024-10-11T00:19:43.719587Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 2.2 Check Data\n\n- Check Duplicate Values\n\n- Check missing values\n\n- Check Outliers","metadata":{}},{"cell_type":"code","source":"# Check Duplicate Rows\nprint(train_df.duplicated().sum())\nprint(test_df.duplicated().sum())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:43.721656Z","iopub.execute_input":"2024-10-11T00:19:43.721966Z","iopub.status.idle":"2024-10-11T00:19:43.792239Z","shell.execute_reply.started":"2024-10-11T00:19:43.721933Z","shell.execute_reply":"2024-10-11T00:19:43.791235Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Check Outliers\ntrain_df.describe()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:43.793468Z","iopub.execute_input":"2024-10-11T00:19:43.793815Z","iopub.status.idle":"2024-10-11T00:19:44.338722Z","shell.execute_reply.started":"2024-10-11T00:19:43.793773Z","shell.execute_reply":"2024-10-11T00:19:44.337647Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def plot_missing_values(df, ax, title):\n    \"\"\"\n    Plots a horizontal bar chart showing the proportion of missing and available values for each feature on a specific subplot (ax).\n\n    Parameters:\n    df: pd.DataFrame or cudf.DataFrame\n        The input dataframe (either pandas or cudf). The function calculates the missing value ratio and plots it.\n\n    ax: matplotlib axis\n        The axis on which the plot is drawn.\n    title: str\n        Title for the plot.\n    \"\"\"\n\n    # Step 1: Calculate the number of missing values for each feature    \n    missing_count = df.isnull().sum()\n    total = len(df)\n\n    # Step 2: Calculate the missing value ratio and available value ratio\n    missing_ratio = missing_count / total  # Ratio of missing values for each feature\n    available_ratio = 1 - missing_ratio  # Ratio of available (non-missing) values\n\n    # Step 3: Create a dataframe to store feature names, missing ratios, and available ratios\n    missing_data = pd.DataFrame({\n        'feature': missing_count.index,  # Feature names\n        'missing_ratio': missing_ratio,  # Ratio of missing values\n        'available_ratio': available_ratio  # Ratio of available values\n    }).sort_values(by='missing_ratio', ascending=False)  # Sort by missing ratio in descending order\n\n    # Step 4: Plot the stacked horizontal bar chart\n    ax.barh(np.arange(len(missing_data)), missing_data['missing_ratio'], color='coral', label='missing')\n    ax.barh(np.arange(len(missing_data)),\n             missing_data['available_ratio'],\n             left=missing_data['missing_ratio'],  # Position this bar to the right of the missing ratio bar\n             color='darkseagreen', label='available')\n\n\n    # Step 5: Set y-axis labels to feature names and format x-axis as percentage\n    ax.set_yticks(np.arange(len(missing_data)))\n    ax.set_yticklabels(missing_data['feature'])\n    ax.xaxis.set_major_formatter(PercentFormatter(xmax=1, decimals=0))  # Format x-axis as percentages\n    ax.set_xlim(0, 1)  # Set x-axis limits from 0 to 1\n    ax.legend()  # Add a legend\n    ax.set_title(title)  # Set plot title\n\n# Create the figure and two subplots side by side\nfig, axes = plt.subplots(1, 2, figsize=(16, 30))\n\n# Plot for train_df and test_df on different axes\nplot_missing_values(train_df, axes[0], 'Train Data Missing Values')\nplot_missing_values(train_df[train_df['sii'] == 1], axes[1], 'Test Data Missing Values')\n\n# Display the plot\nplt.tight_layout()\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:44.340668Z","iopub.execute_input":"2024-10-11T00:19:44.341014Z","iopub.status.idle":"2024-10-11T00:19:49.031048Z","shell.execute_reply.started":"2024-10-11T00:19:44.340977Z","shell.execute_reply":"2024-10-11T00:19:49.029972Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <a id=\"3\">3. Data Preprocessing</a>\n> Encoding, Standardization / Normalization, Feature Selection","metadata":{}},{"cell_type":"markdown","source":"### 3.0 Special Features","metadata":{}},{"cell_type":"code","source":"# Find Different Columns\ndifferent_columns = set(train_df.columns).difference(test_df.columns)\ndifferent_columns.remove('sii')\nprint('Different Columns:', different_columns)\n# Drop Different Columns\ntrain_df = train_df.drop(columns=different_columns)\nprint(f\"Dropped Columns is same as Test Data Columns (except sii): {set(train_df.columns).difference(test_df.columns)}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.032561Z","iopub.execute_input":"2024-10-11T00:19:49.033450Z","iopub.status.idle":"2024-10-11T00:19:49.042737Z","shell.execute_reply.started":"2024-10-11T00:19:49.033409Z","shell.execute_reply":"2024-10-11T00:19:49.041523Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.1. Encoding\n\n- Use One-Hot Encoding for Unordered Data\n\n- Use Label Encoding for Ordered Data\n","metadata":{}},{"cell_type":"code","source":"\"\"\"\n\nOne Hot Encoding\n\n\"\"\"\n\nunordered_columns = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season',\n       'Fitness_Endurance-Season', 'FGC-Season', 'BIA-Season', 'PAQ_A-Season',\n       'PAQ_C-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n\n# For Unordered Columns, fillna with 'Unknown'\ntrain_df[unordered_columns] = train_df[unordered_columns].fillna('Unknown')\ntest_df[unordered_columns] = test_df[unordered_columns].fillna('Unknown')\n\n# Convert unordered Columns to Encoding - One Hot\none_hot_encoder = OneHotEncoder(sparse_output=False)\n# For Train Data\ntrain_encoded = one_hot_encoder.fit_transform(train_df[unordered_columns])\ntrain_encoded_df = pd.DataFrame(train_encoded, columns=one_hot_encoder.get_feature_names_out(unordered_columns))\ntrain_df_rest_pandas = train_df.drop(columns=unordered_columns)\ntrain_encoded_df = pd.concat([train_df_rest_pandas.reset_index(drop=True), train_encoded_df.reset_index(drop=True)], axis=1)\ntrain_encoded_df\n\n# For Test Data\n# Convert unordered Columns to Encoding - One Hot\ntest_encoded = one_hot_encoder.transform(test_df[unordered_columns])\ntest_encoded_df = pd.DataFrame(test_encoded, columns=one_hot_encoder.get_feature_names_out(unordered_columns))\ntest_df_rest_pandas = test_df.drop(columns=unordered_columns)\ntest_encoded_df = pd.concat([test_df_rest_pandas.reset_index(drop=True), test_encoded_df.reset_index(drop=True)], axis=1)\ntest_encoded_df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.046291Z","iopub.execute_input":"2024-10-11T00:19:49.046778Z","iopub.status.idle":"2024-10-11T00:19:49.148267Z","shell.execute_reply.started":"2024-10-11T00:19:49.046730Z","shell.execute_reply":"2024-10-11T00:19:49.147268Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"\"\"\"\nOrdinal Encoding\n\"\"\"\nordered_columns = ['Basic_Demos-Sex','FGC-FGC_CU_Zone','FGC-FGC_GSND_Zone','FGC-FGC_GSD_Zone','FGC-FGC_PU_Zone','FGC-FGC_SRL_Zone','FGC-FGC_SRR_Zone','FGC-FGC_TL_Zone','BIA-BIA_Activity_Level_num','BIA-BIA_Frame_num','PreInt_EduHx-computerinternet_hoursday']\n\n# Convert ordered Columns to Encoding - Label\nordinal_encoder = OrdinalEncoder(handle_unknown='use_encoded_value', unknown_value=np.nan)\ntrain_encoded_df[ordered_columns] = ordinal_encoder.fit_transform(train_encoded_df[ordered_columns])\ntest_encoded_df[ordered_columns] = ordinal_encoder.transform(test_encoded_df[ordered_columns])\n\n# For Ordered Columns, fillna with '-1' (Unknown) Note: '-1' is used as unknown value in Ordinal Encoding, but it does not suitable for Linear Models\ntrain_encoded_df[ordered_columns] = train_encoded_df[ordered_columns].fillna(-1)\ntest_encoded_df[ordered_columns] = test_encoded_df[ordered_columns].fillna(-1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.149458Z","iopub.execute_input":"2024-10-11T00:19:49.149798Z","iopub.status.idle":"2024-10-11T00:19:49.185543Z","shell.execute_reply.started":"2024-10-11T00:19:49.149763Z","shell.execute_reply":"2024-10-11T00:19:49.184414Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"### 3.2. Standardization / Normalization","metadata":{}},{"cell_type":"code","source":"# Unordered Numerical Columns\nunordered_numerical_columns = list(set(train_encoded_df.columns) - set(ordered_columns) - set(unordered_columns) - set([col for col in train_encoded_df.columns if 'Season' in col]) - set(['id', 'sii']))\n# Results\nnormality_results = {}\n# Normality Check\ndef check_normality(column_data):\n    # 1. Shapiro-Wilk Test\n    shapiro_test = stats.shapiro(column_data.dropna())\n    shapiro_p_value = shapiro_test.pvalue\n\n    # 2. D'Agostino's K-squared Test\n    dagostino_test = stats.normaltest(column_data.dropna())\n    dagostino_p_value = dagostino_test.pvalue\n    \n    # 3. Normality Check\n    if shapiro_p_value > 0.05 and dagostino_p_value > 0.05:\n        return True  \n    else:\n       return False \n\n# Iterate over numerical columns\nfor col in unordered_numerical_columns:\n    normality_results[col] = check_normality(train_encoded_df[col])\n\n# Print Results\nnormal_columns = [col for col, is_normal in normality_results.items() if is_normal]\nnon_normal_columns = [col for col, is_normal in normality_results.items() if not is_normal]\nprint(f\"Normal Columns: {normal_columns}\")\nprint(f\"Not Normal Columns: {non_normal_columns}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.186960Z","iopub.execute_input":"2024-10-11T00:19:49.187389Z","iopub.status.idle":"2024-10-11T00:19:49.462431Z","shell.execute_reply.started":"2024-10-11T00:19:49.187341Z","shell.execute_reply":"2024-10-11T00:19:49.461291Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"scaler = MinMaxScaler()\n# MinMax Scaler\ntrain_encoded_df[unordered_numerical_columns] = scaler.fit_transform(train_encoded_df[unordered_numerical_columns])\ntest_encoded_df[unordered_numerical_columns] = scaler.transform(test_encoded_df[unordered_numerical_columns])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.463850Z","iopub.execute_input":"2024-10-11T00:19:49.464731Z","iopub.status.idle":"2024-10-11T00:19:49.514809Z","shell.execute_reply.started":"2024-10-11T00:19:49.464681Z","shell.execute_reply":"2024-10-11T00:19:49.513620Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <a id = \"4\">4. Modeling</a>","metadata":{}},{"cell_type":"code","source":"# Prepare Data\ntrain = train_encoded_df.copy()\ntrain = train.drop(columns=['id'])\ntrain = train[train['sii'].notna()]\ntest = test_encoded_df.copy()\ntest = test.drop(columns=['id'])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.516156Z","iopub.execute_input":"2024-10-11T00:19:49.516487Z","iopub.status.idle":"2024-10-11T00:19:49.535894Z","shell.execute_reply.started":"2024-10-11T00:19:49.516452Z","shell.execute_reply":"2024-10-11T00:19:49.534893Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\ndef threshold_Rounder(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0,\n                 np.where(oof_non_rounded < thresholds[1], 1,\n                      np.where(oof_non_rounded < thresholds[2], 2, 3)))\n\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    rounded_p = threshold_Rounder(oof_non_rounded, thresholds)\n    return -quadratic_weighted_kappa(y_true, rounded_p)\n\ndef TrainML(model_class, test_data):\n    X = train.drop(['sii'], axis=1)\n    y = train['sii']\n    \n    SKF = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=SEED)\n    \n    train_S = []\n    test_S = []\n    \n    oof_non_rounded = np.zeros(len(y), dtype=float) \n    oof_rounded = np.zeros(len(y), dtype=int) \n    test_preds = np.zeros((len(test_data), n_splits))\n\n    for fold, (train_idx, test_idx) in enumerate(tqdm(SKF.split(X, y), desc=\"Training Folds\", total=n_splits)):\n        X_train, X_val = X.iloc[train_idx], X.iloc[test_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[test_idx]\n\n        model = clone(model_class)\n        model.fit(X_train, y_train)\n\n        y_train_pred = model.predict(X_train)\n        y_val_pred = model.predict(X_val)\n\n        oof_non_rounded[test_idx] = y_val_pred\n        y_val_pred_rounded = y_val_pred.round(0).astype(int)\n        oof_rounded[test_idx] = y_val_pred_rounded\n\n        train_kappa = quadratic_weighted_kappa(y_train, y_train_pred.round(0).astype(int))\n        val_kappa = quadratic_weighted_kappa(y_val, y_val_pred_rounded)\n\n        train_S.append(train_kappa)\n        test_S.append(val_kappa)\n        test_preds[:, fold] = model.predict(test_data) \n\n        print(f\"Fold {fold+1} - Train QWK: {train_kappa:.4f}, Validation QWK: {val_kappa:.4f}\")\n        clear_output(wait=True)\n\n    print(f\"Mean Train QWK --> {np.mean(train_S):.4f}\")\n    print(f\"Mean Validation QWK ---> {np.mean(test_S):.4f}\")\n\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead') # Nelder-Mead | # Powell\n    assert KappaOPtimizer.success, \"Optimization did not converge.\"\n\n    oof_tuned = threshold_Rounder(oof_non_rounded, KappaOPtimizer.x)\n    tKappa = quadratic_weighted_kappa(y, oof_tuned)\n    print(f\"----> || Optimized QWK SCORE :: {Fore.CYAN}{Style.BRIGHT} {tKappa:.3f}{Style.RESET_ALL}\")\n    tpm = test_preds.mean(axis=1)\n    tpTuned = threshold_Rounder(tpm, KappaOPtimizer.x)\n    submission = pd.DataFrame({\n        'id': sample_submission_df['id'],\n        'sii': tpTuned\n    })\n    return submission, tKappa\n\n\n\nbest_lgb_params = {'learning_rate': 0.03884249148676395, 'max_depth': 12, 'num_leaves': 413, 'min_data_in_leaf': 14,\n           'feature_fraction': 0.7987976913702801, 'bagging_fraction': 0.7602261703576205, 'bagging_freq': 2, \n           'lambda_l1': 4.735462555910575, 'lambda_l2': 4.735028557007343e-06} # CV : 0.4094 | LB : 0.471\n\noptuna_lgb_params = {'learning_rate': 0.17433484207621536, 'num_leaves': 338, 'max_depth': 13, 'min_data_in_leaf': 44, 'feature_fraction': 0.650608698747034, 'bagging_fraction': 0.6120239572683143, 'bagging_freq': 1, 'lambda_l1': 8.092937699098033e-06, 'lambda_l2': 0.0007858401079131629}\n\nSEED = 42\nn_splits = 5\n\nlight_1 = lgb.LGBMRegressor(**best_lgb_params,random_state=SEED, verbose=-1,n_estimators=200)\nSubmission,model = TrainML(light_1,test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:19:49.537403Z","iopub.execute_input":"2024-10-11T00:19:49.537745Z","iopub.status.idle":"2024-10-11T00:20:05.580703Z","shell.execute_reply.started":"2024-10-11T00:19:49.537711Z","shell.execute_reply":"2024-10-11T00:20:05.579417Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Save submission\nSubmission.to_csv('submission.csv', index=False)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-10-11T00:20:05.582244Z","iopub.execute_input":"2024-10-11T00:20:05.582703Z","iopub.status.idle":"2024-10-11T00:20:05.589015Z","shell.execute_reply.started":"2024-10-11T00:20:05.582653Z","shell.execute_reply":"2024-10-11T00:20:05.588143Z"}},"outputs":[],"execution_count":null}]}