{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"},{"sourceId":208433087,"sourceType":"kernelVersion"}],"dockerImageVersionId":30761,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div style=\"color:navy;background-color:lightgreen;padding:1.2%;border-radius:12px 12px;font-size:1em;text-align:center\">📱Child Mind Institute — Problematic Internet Use</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:10px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>✍️ Description of Notebook -1</p></div>\n\n- In this challenge, the value of **sii (target)** is unknown for **1224 rows** of train.csv file.\n\n- In this notebook, only the missing values ​​of sii for the 1224 rows (mentioned above) are calculated with high accuracy.\n \n- First, \"Features Imputation\" is temporarily performed for the train.csv file, and then the train.csv file is separated into two parts. The first part contains 2736 rows and the value of sii is known in it, and therefore it is the **train part** of calculations. The second part contains 1224 rows in which the value of sii is uncertain and is the calculation **test part**.\n\n- In the next step, regression is performed using LGBM and sii values are calculated with high accuracy.\n\n- In the last step, only the sii column of the train.csv file is completed using the calculated values, and then it is sent as an output with the name **train_sii.csv**.\n\n- You can use the **train_sii.csv** file instead of train.csv in your notebooks, and in this way the information of 1224 rows will be usable.\n\n\n# <div style=\"color:yellow;display:inline-block;border-radius:10px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>✍️ Description of Notebook -2</p></div>\n\n- Notebook -1 >> [[1] CMI 🏁 Target(sii) Imputation Via LGBM](https://www.kaggle.com/code/mehrankazeminia/1-cmi-target-sii-imputation-via-lgbm)\n\n-  In this notebook, using the output of the **first notebook**, the necessary data for regression is set.\n\n-  Then regression is done using different algorithms and finally several better results are **Ensembling** together.\n","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:darkred; align-items: center;\">၊၊||၊ Relating Physical Activity to Problematic Internet Use</span>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport polars as pl\nimport os, time, copy \nimport gc, json, random\nfrom pathlib import Path\n\nimport itertools\nfrom scipy import stats\nfrom scipy.optimize import minimize\nfrom scipy.spatial.distance import cdist\n\nimport seaborn as sns\nfrom matplotlib import colors\nimport matplotlib.pyplot as plt\nfrom colorama import Style, Fore\n%matplotlib inline\n\n# ............................................\nimport warnings\nwarnings.filterwarnings('ignore')\n!ls ../input/*","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-11-23T08:26:34.888940Z","iopub.execute_input":"2024-11-23T08:26:34.889574Z","iopub.status.idle":"2024-11-23T08:26:36.654509Z","shell.execute_reply.started":"2024-11-23T08:26:34.889519Z","shell.execute_reply":"2024-11-23T08:26:36.653252Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"# Notebook -1 >> >> \ndtrain = pd.read_csv('/kaggle/input/1-cmi-target-sii-imputation-via-lgbm/train_sii.csv', index_col='id')\n\n# ...............................................................................................................\n# dtrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv', index_col='id')\ndtest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv', index_col='id')\nsub_sample = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\n\n# ..............................................................................................................\ndtrain_col = dtrain.columns.tolist()\ndtest_col = dtest.columns.tolist()\n\ndtrain.shape, dtest.shape, sub_sample.shape","metadata":{"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-11-23T08:26:36.656691Z","iopub.execute_input":"2024-11-23T08:26:36.658119Z","iopub.status.idle":"2024-11-23T08:26:36.765001Z","shell.execute_reply.started":"2024-11-23T08:26:36.658073Z","shell.execute_reply":"2024-11-23T08:26:36.763966Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Duplicates\n\n- We see \"Duplicates\" in the \"Train\" file. But because there is another file called \"series_train.parquet\", we can't delete duplicate rows at the moment.","metadata":{}},{"cell_type":"code","source":"# print('Duplicates in dtrain:', dtrain.duplicated().sum())\n# print('Duplicates in dtest:', dtest.duplicated().sum())\n\n# dtrain.drop_duplicates(inplace=True)\n# dtrain.shape, dtest.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:36.766142Z","iopub.execute_input":"2024-11-23T08:26:36.766443Z","iopub.status.idle":"2024-11-23T08:26:36.770762Z","shell.execute_reply.started":"2024-11-23T08:26:36.766415Z","shell.execute_reply":"2024-11-23T08:26:36.769690Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Missing Values","metadata":{}},{"cell_type":"code","source":"train = dtrain.copy()\ntest = dtest.copy()\n\ntrain_col = train.columns.tolist()\ntest_col = test.columns.tolist()\n\nmissing_values = train.isnull().mean() * 100\nmissing_values.plot(kind='barh', figsize=(10, 25), color=['lightgreen','violet','skyblue','pink'])\n\nplt.title('Percentage of Missing Values', fontsize=18, color='gray')\nplt.xlabel('Percentage', fontsize=18, color='gray')\nplt.ylabel('Features', fontsize=18, color='gray')\nplt.gca().set_facecolor('lightcyan')\nplt.xticks(rotation=0)\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:36.773029Z","iopub.execute_input":"2024-11-23T08:26:36.773638Z","iopub.status.idle":"2024-11-23T08:26:37.967052Z","shell.execute_reply.started":"2024-11-23T08:26:36.773599Z","shell.execute_reply":"2024-11-23T08:26:37.965899Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Features","metadata":{}},{"cell_type":"code","source":"features = test_col.copy()\nprint('Number of Features :', len(features))\n\n# Numerical Features\nnum_features = [f for f in features if train[f].dtype==float or f=='Basic_Demos-Age']\nprint('The number of numerical features :', len(num_features))\n\n# Categorical Features\ncat_features = [f for f in features if f not in num_features]\nprint('The number of categorical features :', len(cat_features))\n\n# Target Features\ntarget_col = [f for f in train_col if f not in test_col]\nprint('The number of target features :', len(target_col), '\\n')\n\n# Unique Number\n# pd.set_option('display.max_rows', 500)\n# pd.DataFrame(data= {'Unique number in train': train[features].nunique(), \n#                     'Unique number in test': test[features].nunique()}).sort_values(by=['Unique number in train']) ","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:37.968270Z","iopub.execute_input":"2024-11-23T08:26:37.968616Z","iopub.status.idle":"2024-11-23T08:26:37.978306Z","shell.execute_reply.started":"2024-11-23T08:26:37.968576Z","shell.execute_reply":"2024-11-23T08:26:37.977174Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Features Imputation (for train & test)\n\n- For numerical features, the **\"KNNImputer\"** library is used, and for categorical features, a new category called **\"unknown\"** is created for missing values.","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import KNNImputer\nimputer_num = KNNImputer(n_neighbors=2, weights=\"uniform\")\n\n# .........................................................................\ndf_data = pd.concat([train[features], test], axis=0)\n\nimputer_num.fit(df_data[num_features])\ndf_data[num_features] = imputer_num.transform(df_data[num_features])\n\nfor col in cat_features:\n    df_data[col] = df_data[col].fillna('unknown')\n    df_data[col] = df_data[col].astype('category')\n\n# .........................................................................\ndf_data.shape, df_data.isnull().mean() * 100 ","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:37.979589Z","iopub.execute_input":"2024-11-23T08:26:37.980037Z","iopub.status.idle":"2024-11-23T08:26:43.071848Z","shell.execute_reply.started":"2024-11-23T08:26:37.979995Z","shell.execute_reply":"2024-11-23T08:26:43.070729Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Pandas get_dummies & Preprocessing - Scaler (for train & test)","metadata":{}},{"cell_type":"code","source":"df_code = pd.get_dummies(df_data, columns=cat_features)\n\n# ......................................................................\n# StandardScaler\n# from sklearn.preprocessing import StandardScaler\n\n# scaler = StandardScaler()\n# df_code[num_features] = scaler.fit_transform(df_code[num_features])\n\n# ......................................................................\n# MinMaxScaler\nfrom sklearn.preprocessing import MinMaxScaler\n\nscaler = MinMaxScaler()\ndf_code[num_features] = scaler.fit_transform(df_code[num_features])\n\n# ......................................................................\ndf_code.shape","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:43.073352Z","iopub.execute_input":"2024-11-23T08:26:43.073797Z","iopub.status.idle":"2024-11-23T08:26:43.109940Z","shell.execute_reply.started":"2024-11-23T08:26:43.073753Z","shell.execute_reply":"2024-11-23T08:26:43.108830Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Time Series Aggregation\n\n- Time series statistics (e.g., mean, standard deviation) from the **Actigraphy** data are computed and merged into the main dataset to create additional features for model training.","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nfrom IPython.display import clear_output\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:43.111252Z","iopub.execute_input":"2024-11-23T08:26:43.111593Z","iopub.status.idle":"2024-11-23T08:26:43.122052Z","shell.execute_reply.started":"2024-11-23T08:26:43.111563Z","shell.execute_reply":"2024-11-23T08:26:43.121025Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# .....................................................................................................................\ndef process_file(filename, dirname):\n    data = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    data.drop('step', axis=1, inplace=True)\n    return data.describe().values.reshape(-1), filename.split('=')[1]\n\n# .....................................................................................................................\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    stats, indexes = zip(*results)\n    \n    data = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    data['id'] = indexes\n    return data\n\n# .....................................................................................................................\ntrain_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet')\ntest_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet')\n\ntime_series_cols = train_ts.columns.tolist()\ntime_series_cols.remove('id')\n\n# .....................................................................................................................\ntrain_ts.shape, test_ts.shape, len(time_series_cols)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:26:43.123586Z","iopub.execute_input":"2024-11-23T08:26:43.123938Z","iopub.status.idle":"2024-11-23T08:28:08.502289Z","shell.execute_reply.started":"2024-11-23T08:26:43.123897Z","shell.execute_reply":"2024-11-23T08:28:08.501223Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Preprocessing - Scaler (for time series)","metadata":{}},{"cell_type":"code","source":"df_data_ts = pd.concat([train_ts, test_ts], axis=0)\n\n# .........................................................................................\n# StandardScaler\n# from sklearn.preprocessing import StandardScaler\n\n# scaler = StandardScaler()\n# df_data_ts[time_series_cols] = scaler.fit_transform(df_data_ts[time_series_cols])\n\n# .........................................................................................\n# MinMaxScaler\nfrom sklearn.preprocessing import MinMaxScaler\n\nscaler = MinMaxScaler()\ndf_data_ts[time_series_cols] = scaler.fit_transform(df_data_ts[time_series_cols])\n\n# .........................................................................................\ndf_data_ts.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:28:08.504933Z","iopub.execute_input":"2024-11-23T08:28:08.505229Z","iopub.status.idle":"2024-11-23T08:28:08.532737Z","shell.execute_reply.started":"2024-11-23T08:28:08.505200Z","shell.execute_reply":"2024-11-23T08:28:08.531503Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Data Setting","metadata":{}},{"cell_type":"code","source":"df_code = df_code.reset_index()\n\ntrain_df = df_code[:3960].copy()\ntest_df = df_code[3960:].copy()\n\ntrain_df.shape, test_df.shape, df_code.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:28:08.534123Z","iopub.execute_input":"2024-11-23T08:28:08.534580Z","iopub.status.idle":"2024-11-23T08:28:08.547667Z","shell.execute_reply.started":"2024-11-23T08:28:08.534534Z","shell.execute_reply":"2024-11-23T08:28:08.546589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_ts = df_data_ts[:996].copy()\ntest_ts = df_data_ts[996:].copy()\n\ntrain_ts.shape, test_ts.shape, df_data_ts.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:28:08.548886Z","iopub.execute_input":"2024-11-23T08:28:08.549180Z","iopub.status.idle":"2024-11-23T08:28:08.561721Z","shell.execute_reply.started":"2024-11-23T08:28:08.549151Z","shell.execute_reply":"2024-11-23T08:28:08.560790Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = pd.merge(train_df, train_ts, how=\"left\", on='id')\ntest_df = pd.merge(test_df, test_ts, how=\"left\", on='id')\n\n# ..................................................................\nfor col in time_series_cols:\n    \n    train_df[col] = train_df[col].fillna(train_df[col].median())\n    test_df[col] = test_df[col].fillna(test_df[col].median())\n\n# ..................................................................\ntrain_df.shape, test_df.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:28:08.563465Z","iopub.execute_input":"2024-11-23T08:28:08.564255Z","iopub.status.idle":"2024-11-23T08:28:08.678012Z","shell.execute_reply.started":"2024-11-23T08:28:08.564205Z","shell.execute_reply":"2024-11-23T08:28:08.676942Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Features Correlation","metadata":{}},{"cell_type":"code","source":"df_corr = train_df.copy()\ndf_corr['sii'] = train['sii'].values.copy()\n\n# .............................................................\ncorr_sii = df_corr.corr(numeric_only=True)['sii']\ncorr_sii = corr_sii[(corr_sii > 0.02) | (corr_sii < -0.02)]\n\ncorr_list = corr_sii.keys().tolist()\ncorr_list.remove('sii')\n\nlen(corr_list)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:28:08.679888Z","iopub.execute_input":"2024-11-23T08:28:08.680267Z","iopub.status.idle":"2024-11-23T08:28:09.104597Z","shell.execute_reply.started":"2024-11-23T08:28:08.680233Z","shell.execute_reply":"2024-11-23T08:28:09.103576Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y  = train['sii'].copy()\nX  = train_df[corr_list].copy()\nXX = test_df[corr_list].copy()\n\ny.shape, X.shape, XX.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:28:09.105692Z","iopub.execute_input":"2024-11-23T08:28:09.106006Z","iopub.status.idle":"2024-11-23T08:28:09.124457Z","shell.execute_reply.started":"2024-11-23T08:28:09.105976Z","shell.execute_reply":"2024-11-23T08:28:09.123418Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:darkcyan;\">၊၊||၊ LightGBM - Cross Validation</span>\n\n<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nfrom bayes_opt import BayesianOptimization\nfrom sklearn.neighbors import KNeighborsRegressor\n\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import RandomizedSearchCV\n\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import train_test_split\n\nfrom sklearn.model_selection import RepeatedKFold\nfrom sklearn.model_selection import RepeatedStratifiedKFold\n\nfrom sklearn.metrics import log_loss\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.metrics import ConfusionMatrixDisplay","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:28:15.975158Z","iopub.execute_input":"2024-11-23T08:28:15.975652Z","iopub.status.idle":"2024-11-23T08:28:15.983379Z","shell.execute_reply.started":"2024-11-23T08:28:15.975616Z","shell.execute_reply":"2024-11-23T08:28:15.982128Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- R-Squared (R² or the coefficient of determination) is a statistical measure in a regression model that determines the proportion of variance in the dependent variable that can be explained by the independent variable. In other words, r-squared shows how well the data fit the regression model (the goodness of fit).","metadata":{}},{"cell_type":"code","source":"base_model = lgb.LGBMRegressor(random_state=420, verbose=-1)\nprint('\\nr2_score :', cross_val_score(base_model, X ,y ,cv=5, scoring='r2')) ","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:28:22.266734Z","iopub.execute_input":"2024-11-23T08:28:22.267976Z","iopub.status.idle":"2024-11-23T08:28:25.463369Z","shell.execute_reply.started":"2024-11-23T08:28:22.267917Z","shell.execute_reply":"2024-11-23T08:28:25.462238Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Auxiliary Functions","metadata":{}},{"cell_type":"code","source":"# Rounding using thresholds\n# Raw Predictions: pred_raw\n# Rounded Predictions: pred\n# Thresholds: t\n\ndef round_t(pred_raw, t):\n    pred = np.where(pred_raw < t[0], 0, np.where(pred_raw < t[1], 1, np.where(pred_raw < t[2], 2, 3)))\n    return pred\n\ndef qw_kappa(y_true, pred):\n    return -cohen_kappa_score(y_true, pred, weights='quadratic')\n\n# Thanks to: @ambrosm\ndef fun(t, y_true, pred_raw):\n    pred = round_t(pred_raw, t)\n    return -cohen_kappa_score(y_true, pred, weights='quadratic')\n\ndef optimized_thresholds(fun, y_true, pred_raw):\n    res = minimize(fun, x0=[0.5, 1.5, 2.5], args=(y_true, pred_raw), method='Nelder-Mead')\n    assert res.success\n    return res.x.round(2) # optimized_thresholds\n\nt = [0.6, 1.07, 2.52] # Optimized Thresholds (initial)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:28:33.354765Z","iopub.execute_input":"2024-11-23T08:28:33.355188Z","iopub.status.idle":"2024-11-23T08:28:33.363929Z","shell.execute_reply.started":"2024-11-23T08:28:33.355148Z","shell.execute_reply":"2024-11-23T08:28:33.362651Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Hyperparameters","metadata":{}},{"cell_type":"code","source":"# ::::::::::::::::::::::::::::::::::::::::::::::::\nlgbm_params1 = {  \n    \n    'metric'              :'rmse',\n    'objective'           :'regression',\n    'learning_rate'       : 0.04,\n    'max_depth'           : 12,\n    'num_leaves'          : 59,\n    'subsample'           : 0.70,\n    'colsample_bytree'    : 0.50,\n    'min_child_weight'    : 12, \n    'min_child_samples'   : 14,    \n    'reg_alpha'           : 0.23,\n    'reg_lambda'          : 0.36,\n}\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::\nlgbm_params2 = {  \n    \n    'metric'              :'rmse',\n    'objective'           :'regression',\n    'learning_rate'       : 0.05,\n    'max_depth'           : 9,\n    'num_leaves'          : 59,\n    'subsample'           : 0.80,\n    'colsample_bytree'    : 0.50,\n    'min_child_weight'    : 12, \n    'min_child_samples'   : 14,  \n    'reg_alpha'           : 0.23,\n    'reg_lambda'          : 0.36,\n}\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::\nlgbm_params3 = {  \n    \n    'metric'              :'rmse',\n    'objective'           :'regression',\n    'learning_rate'       : 0.046,\n    'max_depth'           : 12,\n    'num_leaves'          : 478,\n    'min_data_in_leaf'    : 13,\n    'feature_fraction'    : 0.893,\n    'bagging_fraction'    : 0.784,\n    'bagging_freq'        : 4,\n    'lambda_l1'           : 10, \n    'lambda_l2'           : 0.01, \n}\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::\n\nmodel1 = lgb.LGBMRegressor(**lgbm_params1, n_estimators=10000, random_state=421, early_stopping_rounds=350, verbose=-1)\n\nmodel2 = lgb.LGBMRegressor(**lgbm_params2, n_estimators=10000, random_state=422, early_stopping_rounds=350, verbose=-1)\n\nmodel3 = lgb.LGBMRegressor(**lgbm_params3, n_estimators=10000, random_state=423, early_stopping_rounds=350, verbose=-1)\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::\n\nmodel_list = [model1, model2, model3]\n","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:28:39.998604Z","iopub.execute_input":"2024-11-23T08:28:39.999061Z","iopub.status.idle":"2024-11-23T08:28:40.009821Z","shell.execute_reply.started":"2024-11-23T08:28:39.999018Z","shell.execute_reply":"2024-11-23T08:28:40.008511Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Detecting less important features","metadata":{}},{"cell_type":"code","source":"from lightgbm import plot_importance   \nmodel_b = lgb.LGBMRegressor(**lgbm_params1, verbose=-1)\nmodel_b.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:28:48.716691Z","iopub.execute_input":"2024-11-23T08:28:48.717161Z","iopub.status.idle":"2024-11-23T08:28:49.695615Z","shell.execute_reply.started":"2024-11-23T08:28:48.717124Z","shell.execute_reply":"2024-11-23T08:28:49.694437Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"plot_importance(model_b, figsize=(12, 30), color=['lightgreen','violet','skyblue','pink'], height=0.6, max_num_features=100,\n                title='LightGBM - Feature importance', xlabel='Value', ylabel='Name Feature');\n\nplt.gca().set_facecolor('lightcyan')","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:28:58.214537Z","iopub.execute_input":"2024-11-23T08:28:58.215038Z","iopub.status.idle":"2024-11-23T08:28:59.854455Z","shell.execute_reply.started":"2024-11-23T08:28:58.214995Z","shell.execute_reply":"2024-11-23T08:28:59.853377Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"name_list = X.columns.tolist()\nweight_list = model_b.feature_importances_ \n\n# ............................................................\nname_weight = []\nfor i in range(len(name_list)):\n    name_weight.append([name_list[i], weight_list[i]])\n    \nname_weight_sort = sorted(name_weight, key=lambda x: x[1])   \n\n# ............................................................\nprint('Total number of features :', len(name_weight_sort))\nweight_list, name_weight_sort[:10] \n","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:29:11.718592Z","iopub.execute_input":"2024-11-23T08:29:11.719033Z","iopub.status.idle":"2024-11-23T08:29:11.733156Z","shell.execute_reply.started":"2024-11-23T08:29:11.718998Z","shell.execute_reply":"2024-11-23T08:29:11.732028Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=123)\nmodel_t = lgb.LGBMRegressor(**lgbm_params1, verbose=-1)\n\nmodel_t.fit(X_train, y_train)\npredict_t = model_t.predict(X_test)\n\nscore_t = cohen_kappa_score(y_test, round_t(predict_t, t), weights='quadratic')\nprint('Score based on all (', len(name_list), ') features :', score_t, '\\n')\n\n# .................................................................................................\nX_t = X.copy()\nXX_t = XX.copy()\n\nfor i in range(len(name_weight_sort)):\n    \n    name_i = name_weight_sort[i][0]\n    weight_i = name_weight_sort[i][1]\n    \n    X_i = X_t.drop(columns=name_i, inplace=False)\n    X_train, X_test, y_train, y_test = train_test_split(X_i, y, test_size=0.20, random_state=123)\n    \n    model_t.fit(X_train, y_train)\n    predict_i = model_t.predict(X_test)\n    \n    score_i = cohen_kappa_score(y_test, round_t(predict_i, t), weights='quadratic')\n    \n    if (score_t < score_i):\n        X_t.drop(columns=name_i, inplace=True)\n        XX_t.drop(columns=name_i, inplace=True)\n        score_t = score_i\n        \n        print('New Score :',round(score_i, 5), '| Weight :', weight_i, '| Name-deleted :', name_i)\n    \n# .................................................................................................\nprint('\\nX_t Shape :', X_t.shape, 'XX_t Shape :', XX_t.shape) ","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:34:41.843953Z","iopub.execute_input":"2024-11-23T08:34:41.845315Z","iopub.status.idle":"2024-11-23T08:36:23.976033Z","shell.execute_reply.started":"2024-11-23T08:34:41.845257Z","shell.execute_reply":"2024-11-23T08:36:23.974881Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# X = X_t.copy()\n# XX = XX_t.copy()","metadata":{"execution":{"iopub.status.busy":"2024-11-23T04:28:10.734054Z","iopub.execute_input":"2024-11-23T04:28:10.734386Z","iopub.status.idle":"2024-11-23T04:28:10.738684Z","shell.execute_reply.started":"2024-11-23T04:28:10.734359Z","shell.execute_reply":"2024-11-23T04:28:10.737603Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Cross-validation","metadata":{}},{"cell_type":"code","source":"score_mean = 0\nth = np.zeros(3)\npred = np.zeros(len(XX))\nrkf = RepeatedKFold(n_splits=3, n_repeats=8, random_state=424)\n\nfor fold, (train_idx, valid_idx) in enumerate(rkf.split(X)):  \n    X_train, y_train = X.iloc[train_idx], y.iloc[train_idx]\n    X_valid, y_valid = X.iloc[valid_idx], y.iloc[valid_idx]  \n\n    print(f'\\n:::::::::::::::::: Fold ~ {fold+1} :::::::::::::::::::')\n    N = random.randrange(4) \n         \n    if (N==0):\n        print('LGBMRegressor - 1')\n        model1.fit(X_train, y_train,             \n                  eval_set=[(X_valid, y_valid)])        \n        oof = model1.predict(X_valid)\n        prd = model1.predict(XX)\n        \n        th_opt = optimized_thresholds(fun, y_valid, oof)\n        print('th_opt =', th_opt)\n \n    if (N==1):\n        print('LGBMRegressor - 2')\n        model2.fit(X_train, y_train,             \n                  eval_set=[(X_valid, y_valid)])      \n        oof = model2.predict(X_valid)\n        prd = model2.predict(XX)\n        \n        th_opt = optimized_thresholds(fun, y_valid, oof)\n        print('th_opt =', th_opt)\n \n    if (N==2 or N==3):\n        print('LGBMRegressor - 3')\n        model3.fit(X_train, y_train,             \n                  eval_set=[(X_valid, y_valid)])        \n        oof = model3.predict(X_valid)\n        prd = model3.predict(XX) \n        \n        th_opt = optimized_thresholds(fun, y_valid, oof)\n        print('th_opt =', th_opt)\n    \n    score = cohen_kappa_score(y_valid, round_t(oof, th_opt), weights='quadratic')\n    print('SCORE:', round(score, 4))\n    \n    th += np.array(th_opt)                          \n    score_mean += score \n    pred += prd\n    \nscore_mean = score_mean / rkf.get_n_splits(X, y)   \nt = np.round(th / rkf.get_n_splits(X, y), 2)\npreds_lgbm_raw = pred / rkf.get_n_splits(X, y)\npreds_lgbm = round_t(preds_lgbm_raw, t)\n\nprint('\\n', '='* 40)\nprint(' .'* 20)\nprint(' SCORE(mean):', score_mean, '\\n')\nprint(' Optimized thresholds:', t)\nprint(' .'* 20)\nprint('='* 40, '\\n')","metadata":{"execution":{"iopub.status.busy":"2024-11-23T08:36:54.420978Z","iopub.execute_input":"2024-11-23T08:36:54.421529Z","iopub.status.idle":"2024-11-23T08:38:01.373487Z","shell.execute_reply.started":"2024-11-23T08:36:54.421487Z","shell.execute_reply":"2024-11-23T08:38:01.372349Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.set()\nplt.figure(figsize=(6, 3))\nplt.hist(preds_lgbm_raw, bins=50)\n\nplt.gca().set_facecolor('lightcyan')\nplt.suptitle('Prediction-raw Histogram', y=0.96, fontsize=16, c='navy')\n\nround(min(preds_lgbm_raw), 3), round(max(preds_lgbm_raw), 3)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:38:07.990060Z","iopub.execute_input":"2024-11-23T08:38:07.991206Z","iopub.status.idle":"2024-11-23T08:38:08.412424Z","shell.execute_reply.started":"2024-11-23T08:38:07.991152Z","shell.execute_reply":"2024-11-23T08:38:08.411313Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"sns.set()\nplt.figure(figsize=(6, 3))\nplt.hist(preds_lgbm, bins=25)\n\nplt.gca().set_facecolor('lightgreen')\nplt.suptitle('Prediction Histogram', y=0.96, fontsize=16, c='navy')\n\nmin(preds_lgbm), max(preds_lgbm)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:38:14.492807Z","iopub.execute_input":"2024-11-23T08:38:14.493279Z","iopub.status.idle":"2024-11-23T08:38:14.806621Z","shell.execute_reply.started":"2024-11-23T08:38:14.493243Z","shell.execute_reply":"2024-11-23T08:38:14.805316Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Submission (LGBM - Cross Validation)","metadata":{}},{"cell_type":"code","source":"sub_lgbm = sub_sample.copy()\nsub_lgbm['sii'] = preds_lgbm\n\n# ............................................\n# sub_lgbm.to_csv('submission.csv', index=False)\nprint(sub_lgbm['sii'].value_counts())\n# Public Score : \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:38:32.109027Z","iopub.execute_input":"2024-11-23T08:38:32.109584Z","iopub.status.idle":"2024-11-23T08:38:33.313966Z","shell.execute_reply.started":"2024-11-23T08:38:32.109532Z","shell.execute_reply":"2024-11-23T08:38:33.312430Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ LightGBM - Single Model</p></div>","metadata":{}},{"cell_type":"code","source":"# ...........................................................................................\nParams = {\n    'learning_rate': 0.046,\n    'max_depth': 12,\n    'num_leaves': 478,\n    'min_data_in_leaf': 13,\n    'feature_fraction': 0.893,\n    'bagging_fraction': 0.784,\n    'bagging_freq': 4,\n    'lambda_l1': 10,  \n    'lambda_l2': 0.01  \n}\n\n# ...........................................................................................\nmodel_lg = lgb.LGBMRegressor(**Params, random_state=420, verbose=-1, n_estimators=300)\nmodel_lg.fit(X, y)             \n\npreds_lg_raw = model_lg.predict(XX)\npreds_lg = round_t(preds_lg_raw, t)\n\n# ...........................................................................................\nsub_lg = sub_sample.copy()\nsub_lg['sii'] = preds_lg\n\n# sub_lg.to_csv('submission.csv', index=False)\nprint(sub_lg['sii'].value_counts())\n# Public Score : \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:38:47.859611Z","iopub.execute_input":"2024-11-23T08:38:47.860168Z","iopub.status.idle":"2024-11-23T08:38:51.451513Z","shell.execute_reply.started":"2024-11-23T08:38:47.860113Z","shell.execute_reply":"2024-11-23T08:38:51.450040Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ XGBoost</p></div>","metadata":{}},{"cell_type":"code","source":"# ...........................................................................................\nXGB_Params = {\n    'learning_rate': 0.05,\n    'max_depth': 6,\n    'n_estimators': 200,\n    'subsample': 0.8,\n    'colsample_bytree': 0.8,\n    'reg_alpha': 1,  \n    'reg_lambda': 5,  \n    'random_state': 420,\n    'tree_method': 'exact'\n}\n\n# ...........................................................................................\nmodel_xgb = XGBRegressor(**XGB_Params)\nmodel_xgb.fit(X, y)             \n\npreds_xgb_raw = model_xgb.predict(XX)\npreds_xgb = round_t(preds_xgb_raw, t)\n\n# ...........................................................................................\nsub_xgb = sub_sample.copy()\nsub_xgb['sii'] = preds_xgb\n\n# sub_xgb.to_csv('submission.csv', index=False)\nprint(sub_xgb['sii'].value_counts())\n# Public Score : \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:39:38.789869Z","iopub.execute_input":"2024-11-23T08:39:38.790406Z","iopub.status.idle":"2024-11-23T08:39:46.991865Z","shell.execute_reply.started":"2024-11-23T08:39:38.790358Z","shell.execute_reply":"2024-11-23T08:39:46.990501Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ CatBoost</p></div>","metadata":{}},{"cell_type":"code","source":"X_c  = df_data[:3960].copy()\nXX_c = df_data[3960:].copy()\n\n# ...........................................................................................\nfeatures = X_c.columns.tolist()\nprint('Number of Features :', len(features))\n\n# Numerical Features\nnum_features = [f for f in features if X_c[f].dtype==float or f=='Basic_Demos-Age']\nprint('The number of numerical features :', len(num_features))\n\n# Categorical Features\ncat_features = [f for f in features if f not in num_features]\nprint('The number of categorical features :', len(cat_features))\n\n# ...........................................................................................","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:40:09.000706Z","iopub.execute_input":"2024-11-23T08:40:09.001312Z","iopub.status.idle":"2024-11-23T08:40:09.028907Z","shell.execute_reply.started":"2024-11-23T08:40:09.001257Z","shell.execute_reply":"2024-11-23T08:40:09.027399Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ...........................................................................................\nCatBoost_Params = {\n    'learning_rate': 0.05,\n    'depth': 6,\n    'iterations': 200,\n    'random_seed': 420,\n    'cat_features': cat_features,\n    'verbose': 0,\n    'l2_leaf_reg': 10  \n}\n\n# ...........................................................................................\nmodel_cat = CatBoostRegressor(**CatBoost_Params)\nmodel_cat.fit(X_c, y) \n\npreds_cat_raw = model_cat.predict(XX_c)\npreds_cat = round_t(preds_cat_raw, t)\n\n# ...........................................................................................\nsub_cat = sub_sample.copy()\nsub_cat['sii'] = preds_cat\n\n# sub_cat.to_csv('submission.csv', index=False)\nprint(sub_cat['sii'].value_counts())\n# Public Score : \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:40:20.666776Z","iopub.execute_input":"2024-11-23T08:40:20.667356Z","iopub.status.idle":"2024-11-23T08:40:24.155924Z","shell.execute_reply.started":"2024-11-23T08:40:20.667297Z","shell.execute_reply":"2024-11-23T08:40:24.154412Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ LinearSVR</p></div>","metadata":{}},{"cell_type":"code","source":"from sklearn.svm import LinearSVR\n\n# .............................................................\nmodel_svr = LinearSVR(max_iter= 1000, epsilon= 0.1)\nmodel_svr.fit(X, y) \n\npreds_svr_raw = model_svr.predict(XX)\npreds_svr = round_t(preds_svr_raw, t)\n\n# ...........................................................................................\nsub_svr = sub_sample.copy()\nsub_svr['sii'] = preds_svr\n\n# sub_svr.to_csv('submission.csv', index=False)\nprint(sub_svr['sii'].value_counts())\n# Public Score : \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:40:32.666926Z","iopub.execute_input":"2024-11-23T08:40:32.667476Z","iopub.status.idle":"2024-11-23T08:40:34.246090Z","shell.execute_reply.started":"2024-11-23T08:40:32.667422Z","shell.execute_reply":"2024-11-23T08:40:34.244333Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ KNeighborsRegressor</p></div>","metadata":{}},{"cell_type":"code","source":"from sklearn.neighbors import KNeighborsRegressor\nX_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.20, random_state=123)\n\n# ...............................................................................................\ndef best_knn(coeff): \n    neigh_t = KNeighborsRegressor(n_neighbors=coeff)\n    neigh_t.fit(X_train, y_train)\n    \n    predict_t = neigh_t.predict(X_test)\n    score = cohen_kappa_score(y_test, round_t(predict_t, t), weights='quadratic')\n    return score\n\n# ...............................................................................................\nresults = {}\nfor i in range(3, 21, 2):       \n    results[i] = best_knn(i)  \n    \nsns.set()\nplt.figure(figsize=(8, 5))\nplt.gca().set_facecolor('lightcyan')\nplt.suptitle('Results based on changes in \"n_neighbors\"', y=0.92, fontsize=14, c='gray')\n\nplt.plot(list(results.keys()), list(results.values()))\nplt.show() ","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:41:20.025879Z","iopub.execute_input":"2024-11-23T08:41:20.026429Z","iopub.status.idle":"2024-11-23T08:41:21.032576Z","shell.execute_reply.started":"2024-11-23T08:41:20.026368Z","shell.execute_reply":"2024-11-23T08:41:21.031407Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# ...................................................\nneigh = KNeighborsRegressor(n_neighbors=7)\nneigh.fit(X, y)\n\npreds_knn_raw = neigh.predict(XX)\npreds_knn = round_t(preds_knn_raw, t)\n\n# ...................................................\nsub_knn = sub_sample.copy()\nsub_knn['sii'] = preds_knn\n\n# sub_knn.to_csv('submission.csv', index=False)\nprint(sub_knn['sii'].value_counts())\n# Public Score : \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:41:56.691875Z","iopub.execute_input":"2024-11-23T08:41:56.692983Z","iopub.status.idle":"2024-11-23T08:41:57.949154Z","shell.execute_reply.started":"2024-11-23T08:41:56.692916Z","shell.execute_reply":"2024-11-23T08:41:57.947616Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ Ensembling</p></div>","metadata":{}},{"cell_type":"code","source":"# ..........................................................................................................\npreds_ens_raw = (preds_lgbm_raw + preds_lgbm_raw + preds_lg_raw + preds_xgb_raw + preds_cat_raw ) / 5\npreds_ens = round_t(preds_ens_raw, t)\n\n# ..........................................................................................................\nsub_ens = sub_sample.copy()\nsub_ens['sii'] = preds_ens\n\nsub_ens.to_csv('submission.csv', index=False)\nprint(sub_ens['sii'].value_counts())\n# Public Score: \n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-23T08:44:32.532643Z","iopub.execute_input":"2024-11-23T08:44:32.533373Z","iopub.status.idle":"2024-11-23T08:44:33.750621Z","shell.execute_reply.started":"2024-11-23T08:44:32.533267Z","shell.execute_reply":"2024-11-23T08:44:33.748978Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}}]}