{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"},{"sourceId":208433087,"sourceType":"kernelVersion"}],"dockerImageVersionId":30761,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <div style=\"color:navy;background-color:lightgreen;padding:1.2%;border-radius:12px 12px;font-size:1em;text-align:center\">📱Child Mind Institute — Problematic Internet Use</div>\n\n# <div style=\"color:yellow;display:inline-block;border-radius:10px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>✍️ Description of Notebook -1</p></div>\n\n- In this challenge, the value of **sii (target)** is unknown for **1224 rows** of train.csv file.\n\n- In this notebook, only the missing values ​​of sii for the 1224 rows (mentioned above) are calculated with high accuracy.\n \n- First, \"Features Imputation\" is temporarily performed for the train.csv file, and then the train.csv file is separated into two parts. The first part contains 2736 rows and the value of sii is known in it, and therefore it is the **train part** of calculations. The second part contains 1224 rows in which the value of sii is uncertain and is the calculation **test part**.\n\n- In the next step, regression is performed using LGBM and sii values are calculated with high accuracy.\n\n- In the last step, only the sii column of the train.csv file is completed using the calculated values, and then it is sent as an output with the name **train_sii.csv**.\n\n- You can use the **train_sii.csv** file instead of train.csv in your notebooks, and in this way the information of 1224 rows will be usable.\n\n\n# <div style=\"color:yellow;display:inline-block;border-radius:10px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>✍️ Description of Notebook -2</p></div>\n\n- Notebook -1 >> [[1] CMI 🏁 Target(sii) Imputation Via LGBM](https://www.kaggle.com/code/mehrankazeminia/1-cmi-target-sii-imputation-via-lgbm)\n\n-  In this notebook, using the output of the **first notebook**, the necessary data for regression is set.\n\n-  Then regression is done using different algorithms and finally several better results are **Ensembling** together.\n\n# <div style=\"color:yellow;display:inline-block;border-radius:10px;background-color:lightgray;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:darkred;overflow:hidden;font-size:70%;letter-spacing:0.5px;margin:0\"><b> </b>✍️ Description of Notebook -3</p></div>\n\n- Notebook -2 >> [[2] CMI 🏁 LGBM XGB CAT SVR KNN](https://www.kaggle.com/code/mehrankazeminia/2-cmi-lgbm-xgb-cat-svr-knn)\n\n- The best result of my second notebook is Public Score = 0.452. In the third notebook, I combine this result with several public notebooks to get a better score.\n\n- To combine the results of different notebooks in this competition, there are a few things to consider. First, we only see the results of twenty rows of the challenge test set. Second, the regression score of all notebooks is less than 0.5.\n\n- These issues mean that combining the results must be done by repeating the **trial and error method**, and for example, using correlation to identify suitable notebooks or ..... cannot be effective.\n","metadata":{}},{"cell_type":"markdown","source":"# <span style=\"color:darkred; align-items: center;\">၊၊||၊ Relating Physical Activity to Problematic Internet Use</span>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport polars as pl\nimport os, time, copy \nimport gc, json, random\nfrom pathlib import Path\n\nimport itertools\nfrom scipy import stats\nfrom scipy.optimize import minimize\nfrom scipy.spatial.distance import cdist\n\nimport seaborn as sns\nfrom matplotlib import colors\nimport matplotlib.pyplot as plt\nfrom colorama import Style, Fore\n%matplotlib inline\n\n# ............................................\nimport warnings\nwarnings.filterwarnings('ignore')\n!ls ../input/*","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-12-09T13:25:58.625318Z","iopub.execute_input":"2024-12-09T13:25:58.626006Z","iopub.status.idle":"2024-12-09T13:26:02.602419Z","shell.execute_reply.started":"2024-12-09T13:25:58.625936Z","shell.execute_reply":"2024-12-09T13:26:02.601167Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"# Notebook -1 >> >> \ndtrain = pd.read_csv('/kaggle/input/1-cmi-target-sii-imputation-via-lgbm/train_sii.csv', index_col='id')\n\n# ...............................................................................................................\n# dtrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv', index_col='id')\ndtest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv', index_col='id')\nsub_sample = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\n\n# ..............................................................................................................\ndtrain_col = dtrain.columns.tolist()\ndtest_col = dtest.columns.tolist()\n\ndtrain.shape, dtest.shape, sub_sample.shape","metadata":{"_kg_hide-output":false,"execution":{"iopub.status.busy":"2024-12-09T13:26:02.604881Z","iopub.execute_input":"2024-12-09T13:26:02.605533Z","iopub.status.idle":"2024-12-09T13:26:02.709895Z","shell.execute_reply.started":"2024-12-09T13:26:02.605495Z","shell.execute_reply":"2024-12-09T13:26:02.708837Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Duplicates & Missing Values\n\n- We see \"Duplicates\" in the \"Train\" file. But because there is another file called \"series_train.parquet\", we can't delete duplicate rows at the moment.","metadata":{}},{"cell_type":"code","source":"# print('Duplicates in dtrain:', dtrain.duplicated().sum())\n# print('Duplicates in dtest:', dtest.duplicated().sum())\n\n# dtrain.drop_duplicates(inplace=True)\n# dtrain.shape, dtest.shape\n\n# .......................................................................\ntrain = dtrain.copy()\ntest = dtest.copy()\n\ntrain_col = train.columns.tolist()\ntest_col = test.columns.tolist()\n\n# .......................................................................\nlen(train_col), train.isnull().mean()[:20] * 100","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:26:02.710934Z","iopub.execute_input":"2024-12-09T13:26:02.711198Z","iopub.status.idle":"2024-12-09T13:26:02.730497Z","shell.execute_reply.started":"2024-12-09T13:26:02.711172Z","shell.execute_reply":"2024-12-09T13:26:02.729447Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Features","metadata":{}},{"cell_type":"code","source":"features = test_col.copy()\nprint('Number of Features :', len(features))\n\n# Numerical Features\nnum_features = [f for f in features if train[f].dtype==float or f=='Basic_Demos-Age']\nprint('The number of numerical features :', len(num_features))\n\n# Categorical Features\ncat_features = [f for f in features if f not in num_features]\nprint('The number of categorical features :', len(cat_features))\n\n# Target Features\ntarget_col = [f for f in train_col if f not in test_col]\nprint('The number of target features :', len(target_col), '\\n')\n\n# Unique Number\npd.set_option('display.max_rows', 500)\npd.DataFrame(data= {'U_number train_cat': train[cat_features].nunique(), \n                    'U_number test_cat': test[cat_features].nunique()}).sort_values(by=['U_number train_cat']) ","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:26:02.733553Z","iopub.execute_input":"2024-12-09T13:26:02.734277Z","iopub.status.idle":"2024-12-09T13:26:02.769552Z","shell.execute_reply.started":"2024-12-09T13:26:02.734242Z","shell.execute_reply":"2024-12-09T13:26:02.768529Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Features Imputation (for train & test)\n\n- For numerical features, the **\"KNNImputer\"** library is used, and for categorical features, a new category called **\"unknown\"** is created for missing values.","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import KNNImputer\nimputer_num = KNNImputer(n_neighbors=2, weights=\"uniform\")\n\n# .........................................................................\ndf_data = pd.concat([train[features], test], axis=0)\n\nimputer_num.fit(df_data[num_features])\ndf_data[num_features] = imputer_num.transform(df_data[num_features])\n\nfor col in cat_features:\n    df_data[col] = df_data[col].fillna('unknown')\n    df_data[col] = df_data[col].astype('category')\n\n# .........................................................................\ndf_data.shape, df_data.isnull().mean()[:20] * 100 ","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:26:02.770969Z","iopub.execute_input":"2024-12-09T13:26:02.771260Z","iopub.status.idle":"2024-12-09T13:26:08.209046Z","shell.execute_reply.started":"2024-12-09T13:26:02.771231Z","shell.execute_reply":"2024-12-09T13:26:08.207752Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Pandas get_dummies & Preprocessing (for train & test)","metadata":{}},{"cell_type":"code","source":"df_code = pd.get_dummies(df_data, columns=cat_features)\n\n# ......................................................................\n# StandardScaler\n# from sklearn.preprocessing import StandardScaler\n\n# scaler = StandardScaler()\n# df_code[num_features] = scaler.fit_transform(df_code[num_features])\n\n# ......................................................................\n# MinMaxScaler\nfrom sklearn.preprocessing import MinMaxScaler\n\nscaler = MinMaxScaler()\ndf_code[num_features] = scaler.fit_transform(df_code[num_features])\n\n# ......................................................................\ndf_code.shape","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:26:08.210621Z","iopub.execute_input":"2024-12-09T13:26:08.211107Z","iopub.status.idle":"2024-12-09T13:26:08.247479Z","shell.execute_reply.started":"2024-12-09T13:26:08.211060Z","shell.execute_reply":"2024-12-09T13:26:08.246505Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Time Series Aggregation\n\n- Time series statistics (e.g., mean, standard deviation) from the **Actigraphy** data are computed and merged into the main dataset to create additional features for model training.","metadata":{}},{"cell_type":"code","source":"from tqdm import tqdm\nfrom IPython.display import clear_output\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:26:08.248974Z","iopub.execute_input":"2024-12-09T13:26:08.249412Z","iopub.status.idle":"2024-12-09T13:26:08.261313Z","shell.execute_reply.started":"2024-12-09T13:26:08.249365Z","shell.execute_reply":"2024-12-09T13:26:08.260117Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# .....................................................................................................................\ndef process_file(filename, dirname):\n    data = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    data.drop('step', axis=1, inplace=True)\n    return data.describe().values.reshape(-1), filename.split('=')[1]\n\n# .....................................................................................................................\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    stats, indexes = zip(*results)\n    \n    data = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    data['id'] = indexes\n    return data\n\n# .....................................................................................................................\ntrain_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet')\ntest_ts = load_time_series('/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet')\n\ntime_series_cols = train_ts.columns.tolist()\ntime_series_cols.remove('id')\n\n# .....................................................................................................................\ntrain_ts.shape, test_ts.shape, len(time_series_cols)","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:26:08.262902Z","iopub.execute_input":"2024-12-09T13:26:08.263225Z","iopub.status.idle":"2024-12-09T13:27:36.129871Z","shell.execute_reply.started":"2024-12-09T13:26:08.263195Z","shell.execute_reply":"2024-12-09T13:27:36.128784Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Preprocessing (for time series)","metadata":{}},{"cell_type":"code","source":"df_data_ts = pd.concat([train_ts, test_ts], axis=0)\n\n# .........................................................................................\n# StandardScaler\n# from sklearn.preprocessing import StandardScaler\n\n# scaler = StandardScaler()\n# df_data_ts[time_series_cols] = scaler.fit_transform(df_data_ts[time_series_cols])\n\n# .........................................................................................\n# MinMaxScaler\nfrom sklearn.preprocessing import MinMaxScaler\n\nscaler = MinMaxScaler()\ndf_data_ts[time_series_cols] = scaler.fit_transform(df_data_ts[time_series_cols])\n\n# .........................................................................................\ndf_data_ts.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.131133Z","iopub.execute_input":"2024-12-09T13:27:36.131547Z","iopub.status.idle":"2024-12-09T13:27:36.157741Z","shell.execute_reply.started":"2024-12-09T13:27:36.131478Z","shell.execute_reply":"2024-12-09T13:27:36.156721Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Data Setting","metadata":{}},{"cell_type":"code","source":"df_code = df_code.reset_index()\n\ntrain_df = df_code[:3960].copy()\ntest_df = df_code[3960:].copy()\n\ntrain_df.shape, test_df.shape, df_code.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.162604Z","iopub.execute_input":"2024-12-09T13:27:36.163109Z","iopub.status.idle":"2024-12-09T13:27:36.177759Z","shell.execute_reply.started":"2024-12-09T13:27:36.163073Z","shell.execute_reply":"2024-12-09T13:27:36.176365Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_ts = df_data_ts[:996].copy()\ntest_ts = df_data_ts[996:].copy()\n\ntrain_ts.shape, test_ts.shape, df_data_ts.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.179004Z","iopub.execute_input":"2024-12-09T13:27:36.179302Z","iopub.status.idle":"2024-12-09T13:27:36.194215Z","shell.execute_reply.started":"2024-12-09T13:27:36.179274Z","shell.execute_reply":"2024-12-09T13:27:36.193091Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_df = pd.merge(train_df, train_ts, how=\"left\", on='id')\ntest_df = pd.merge(test_df, test_ts, how=\"left\", on='id')\n\n# ..................................................................\nfor col in time_series_cols:\n    \n    train_df[col] = train_df[col].fillna(train_df[col].median())\n    test_df[col] = test_df[col].fillna(test_df[col].median())\n\n# ..................................................................\ntrain_df.shape, test_df.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.195632Z","iopub.execute_input":"2024-12-09T13:27:36.196073Z","iopub.status.idle":"2024-12-09T13:27:36.317818Z","shell.execute_reply.started":"2024-12-09T13:27:36.196041Z","shell.execute_reply":"2024-12-09T13:27:36.316742Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Shapiro–Wilk test","metadata":{}},{"cell_type":"code","source":"cols = train_df.columns.tolist()\ncols.remove('id')\n\nprint('Number of all columns :', len(cols))\n\n# ...........................................................\ncols_select = []\nalpha = 0.05\n\nfor col in cols:\n    _, p_value = stats.shapiro(train_df[col])\n    \n    if (p_value <= alpha): \n        cols_select.append(col) \n    else:\n        print('\\t- Column deleted :', col)\n\n# ...........................................................\nprint('\\nNumber of selected columns :', len(cols_select))\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.319193Z","iopub.execute_input":"2024-12-09T13:27:36.319555Z","iopub.status.idle":"2024-12-09T13:27:36.427868Z","shell.execute_reply.started":"2024-12-09T13:27:36.319522Z","shell.execute_reply":"2024-12-09T13:27:36.426694Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 📱Features Correlation","metadata":{}},{"cell_type":"code","source":"df_corr = train_df.copy()\ndf_corr['sii'] = train['sii'].values.copy()\n\n# .............................................................\ncorr_sii = df_corr.corr(numeric_only=True)['sii']\ncorr_sii = corr_sii[(corr_sii > 0.03) | (corr_sii < -0.03)]\n\ncorr_list = corr_sii.keys().tolist()\ncorr_list.remove('sii')\n\nlen(corr_list)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.429022Z","iopub.execute_input":"2024-12-09T13:27:36.429321Z","iopub.status.idle":"2024-12-09T13:27:36.859009Z","shell.execute_reply.started":"2024-12-09T13:27:36.429292Z","shell.execute_reply":"2024-12-09T13:27:36.857759Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"y  = train['sii'].copy()\nX  = train_df[corr_list].copy()\nXX = test_df[corr_list].copy()\n\ny.shape, X.shape, XX.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:36.860274Z","iopub.execute_input":"2024-12-09T13:27:36.860593Z","iopub.status.idle":"2024-12-09T13:27:36.876392Z","shell.execute_reply.started":"2024-12-09T13:27:36.860562Z","shell.execute_reply":"2024-12-09T13:27:36.875416Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"# <span style=\"color:darkcyan;\">၊၊||၊ LightGBM - Cross Validation</span>\n\n<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}},{"cell_type":"code","source":"import lightgbm as lgb\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nfrom bayes_opt import BayesianOptimization\nfrom sklearn.neighbors import KNeighborsRegressor\n\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.model_selection import RandomizedSearchCV\n\nfrom sklearn.model_selection import cross_val_score\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import train_test_split\n\nfrom sklearn.model_selection import RepeatedKFold\nfrom sklearn.model_selection import RepeatedStratifiedKFold\n\nfrom sklearn.metrics import log_loss\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.metrics import ConfusionMatrixDisplay","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:27:36.877907Z","iopub.execute_input":"2024-12-09T13:27:36.878654Z","iopub.status.idle":"2024-12-09T13:27:38.222562Z","shell.execute_reply.started":"2024-12-09T13:27:36.878604Z","shell.execute_reply":"2024-12-09T13:27:38.221741Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"- R-Squared (R² or the coefficient of determination) is a statistical measure in a regression model that determines the proportion of variance in the dependent variable that can be explained by the independent variable. In other words, r-squared shows how well the data fit the regression model (the goodness of fit).","metadata":{}},{"cell_type":"code","source":"base_model = lgb.LGBMRegressor(random_state=420, verbose=-1)\nprint('\\nr2_score :', cross_val_score(base_model, X ,y ,cv=5, scoring='r2')) ","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:27:38.223955Z","iopub.execute_input":"2024-12-09T13:27:38.224436Z","iopub.status.idle":"2024-12-09T13:27:40.913659Z","shell.execute_reply.started":"2024-12-09T13:27:38.224403Z","shell.execute_reply":"2024-12-09T13:27:40.912625Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Auxiliary Functions","metadata":{}},{"cell_type":"code","source":"# Rounding using thresholds\n# Raw Predictions: pred_raw\n# Rounded Predictions: pred\n# Thresholds: t\n\ndef round_t(pred_raw, t):\n    pred = np.where(pred_raw < t[0], 0, np.where(pred_raw < t[1], 1, np.where(pred_raw < t[2], 2, 3)))\n    return pred\n\ndef qw_kappa(y_true, pred):\n    return -cohen_kappa_score(y_true, pred, weights='quadratic')\n\n# Thanks to: @ambrosm\ndef fun(t, y_true, pred_raw):\n    pred = round_t(pred_raw, t)\n    return -cohen_kappa_score(y_true, pred, weights='quadratic')\n\ndef optimized_thresholds(fun, y_true, pred_raw):\n    res = minimize(fun, x0=[0.5, 1.5, 2.5], args=(y_true, pred_raw), method='Nelder-Mead')\n    assert res.success\n    return res.x.round(2) # optimized_thresholds\n\nt = [0.6, 1.07, 2.52] # Optimized Thresholds (initial)","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:27:40.914997Z","iopub.execute_input":"2024-12-09T13:27:40.915295Z","iopub.status.idle":"2024-12-09T13:27:40.925024Z","shell.execute_reply.started":"2024-12-09T13:27:40.915257Z","shell.execute_reply":"2024-12-09T13:27:40.923685Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Hyperparameters","metadata":{}},{"cell_type":"code","source":"# ::::::::::::::::::::::::::::::::::::::::::::::::\nlgbm_params1 = {  \n    \n    'metric'              :'rmse',\n    'objective'           :'regression',\n    'learning_rate'       : 0.04,\n    'max_depth'           : 12,\n    'num_leaves'          : 59,\n    'subsample'           : 0.70,\n    'colsample_bytree'    : 0.50,\n    'min_child_weight'    : 12, \n    'min_child_samples'   : 14,    \n    'reg_alpha'           : 0.23,\n    'reg_lambda'          : 0.36,\n}\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::\nlgbm_params2 = {  \n    \n    'metric'              :'rmse',\n    'objective'           :'regression',\n    'learning_rate'       : 0.05,\n    'max_depth'           : 9,\n    'num_leaves'          : 59,\n    'subsample'           : 0.80,\n    'colsample_bytree'    : 0.50,\n    'min_child_weight'    : 12, \n    'min_child_samples'   : 14,  \n    'reg_alpha'           : 0.23,\n    'reg_lambda'          : 0.36,\n}\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::\nlgbm_params3 = {  \n    \n    'metric'              :'rmse',\n    'objective'           :'regression',\n    'learning_rate'       : 0.046,\n    'max_depth'           : 12,\n    'num_leaves'          : 23,\n    'min_data_in_leaf'    : 13,\n    'feature_fraction'    : 0.893,\n    'bagging_fraction'    : 0.784,\n    'bagging_freq'        : 4,\n    'lambda_l1'           : 10, \n    'lambda_l2'           : 0.01, \n}\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::\n\nmodel1 = lgb.LGBMRegressor(**lgbm_params1, n_estimators=10000, random_state=421, early_stopping_rounds=350, verbose=-1)\n\nmodel2 = lgb.LGBMRegressor(**lgbm_params2, n_estimators=10000, random_state=422, early_stopping_rounds=350, verbose=-1)\n\nmodel3 = lgb.LGBMRegressor(**lgbm_params3, n_estimators=10000, random_state=423, early_stopping_rounds=350, verbose=-1)\n\n# ::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::::\n\nmodel_list = [model1, model2, model3]\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:27:40.927019Z","iopub.execute_input":"2024-12-09T13:27:40.927461Z","iopub.status.idle":"2024-12-09T13:27:40.942012Z","shell.execute_reply.started":"2024-12-09T13:27:40.927413Z","shell.execute_reply":"2024-12-09T13:27:40.940923Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Cross-validation","metadata":{}},{"cell_type":"code","source":"score_mean = 0\nth = np.zeros(3)\npred = np.zeros(len(XX))\nrkf = RepeatedKFold(n_splits=3, n_repeats=8, random_state=424)\n\nfor fold, (train_idx, valid_idx) in enumerate(rkf.split(X)):  \n    X_train, y_train = X.iloc[train_idx], y.iloc[train_idx]\n    X_valid, y_valid = X.iloc[valid_idx], y.iloc[valid_idx]  \n\n    print(f'\\n:::::::::::::::::: Fold ~ {fold+1} :::::::::::::::::::')\n    N = random.randrange(4) \n         \n    if (N==0):\n        print('LGBMRegressor - 1')\n        model1.fit(X_train, y_train,             \n                  eval_set=[(X_valid, y_valid)])        \n        oof = model1.predict(X_valid)\n        prd = model1.predict(XX)\n        \n        th_opt = optimized_thresholds(fun, y_valid, oof)\n        print('th_opt =', th_opt)\n \n    if (N==1):\n        print('LGBMRegressor - 2')\n        model2.fit(X_train, y_train,             \n                  eval_set=[(X_valid, y_valid)])      \n        oof = model2.predict(X_valid)\n        prd = model2.predict(XX)\n        \n        th_opt = optimized_thresholds(fun, y_valid, oof)\n        print('th_opt =', th_opt)\n \n    if (N==2 or N==3):\n        print('LGBMRegressor - 3')\n        model3.fit(X_train, y_train,             \n                  eval_set=[(X_valid, y_valid)])        \n        oof = model3.predict(X_valid)\n        prd = model3.predict(XX) \n        \n        th_opt = optimized_thresholds(fun, y_valid, oof)\n        print('th_opt =', th_opt)\n    \n    score = cohen_kappa_score(y_valid, round_t(oof, th_opt), weights='quadratic')\n    print('SCORE:', round(score, 4))\n    \n    th += np.array(th_opt)                          \n    score_mean += score \n    pred += prd\n    \nscore_mean = score_mean / rkf.get_n_splits(X, y)   \nt = np.round(th / rkf.get_n_splits(X, y), 2)\npreds_lgbm_raw = pred / rkf.get_n_splits(X, y)\npreds_lgbm = round_t(preds_lgbm_raw, t)\n\nprint('\\n', '='* 40)\nprint(' .'* 20)\nprint(' SCORE(mean):', score_mean, '\\n')\nprint(' Optimized thresholds:', t)\nprint(' .'* 20)\nprint('='* 40, '\\n')","metadata":{"execution":{"iopub.status.busy":"2024-12-09T13:27:40.943602Z","iopub.execute_input":"2024-12-09T13:27:40.944065Z","iopub.status.idle":"2024-12-09T13:28:35.598693Z","shell.execute_reply.started":"2024-12-09T13:27:40.944018Z","shell.execute_reply":"2024-12-09T13:28:35.597381Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"## 🍀 Submission - LGBM","metadata":{}},{"cell_type":"code","source":"sub_lgbm = sub_sample.copy()\nsub_lgbm['sii'] = preds_lgbm\n\n# ..............................................\nsub_lgbm.to_csv('submission0.csv', index=False)\n# Public Score : 0.452\nprint(preds_lgbm)\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:35.600187Z","iopub.execute_input":"2024-12-09T13:28:35.600759Z","iopub.status.idle":"2024-12-09T13:28:36.820315Z","shell.execute_reply.started":"2024-12-09T13:28:35.600713Z","shell.execute_reply":"2024-12-09T13:28:36.818805Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ Notebook imported - 1</p></div>\n\n### [Multi-Target Prediction](https://www.kaggle.com/code/tubotubo/starter-notebook-multi-target-prediction?scriptVersionId=197456007)\n\n- Thanks to : @tubotubo","metadata":{}},{"cell_type":"code","source":"import warnings\nfrom functools import partial\nfrom pathlib import Path\n\nimport matplotlib.pyplot as plt\nimport numpy as np\nimport optuna\nimport polars as pl\nimport polars.selectors as cs\nfrom catboost import CatBoostRegressor, MultiTargetCustomMetric\nfrom numpy.typing import ArrayLike, NDArray\nfrom polars.testing import assert_frame_equal\nfrom sklearn.base import BaseEstimator\nfrom sklearn.metrics import cohen_kappa_score\nfrom sklearn.model_selection import StratifiedKFold\n\nwarnings.filterwarnings(\"ignore\", message=\"Failed to optimize method\")\n\nDATA_DIR = Path(\"/kaggle/input/child-mind-institute-problematic-internet-use\")\nTARGET_COLS = [\n    \"PCIAT-PCIAT_01\",\n    \"PCIAT-PCIAT_02\",\n    \"PCIAT-PCIAT_03\",\n    \"PCIAT-PCIAT_04\",\n    \"PCIAT-PCIAT_05\",\n    \"PCIAT-PCIAT_06\",\n    \"PCIAT-PCIAT_07\",\n    \"PCIAT-PCIAT_08\",\n    \"PCIAT-PCIAT_09\",\n    \"PCIAT-PCIAT_10\",\n    \"PCIAT-PCIAT_11\",\n    \"PCIAT-PCIAT_12\",\n    \"PCIAT-PCIAT_13\",\n    \"PCIAT-PCIAT_14\",\n    \"PCIAT-PCIAT_15\",\n    \"PCIAT-PCIAT_16\",\n    \"PCIAT-PCIAT_17\",\n    \"PCIAT-PCIAT_18\",\n    \"PCIAT-PCIAT_19\",\n    \"PCIAT-PCIAT_20\",\n    \"PCIAT-PCIAT_Total\",\n    \"sii\",\n]\n\nFEATURE_COLS = [\n    \"Basic_Demos-Enroll_Season\",\n    \"Basic_Demos-Age\",\n    \"Basic_Demos-Sex\",\n    \"CGAS-Season\",\n    \"CGAS-CGAS_Score\",\n    \"Physical-Season\",\n    \"Physical-BMI\",\n    \"Physical-Height\",\n    \"Physical-Weight\",\n    \"Physical-Waist_Circumference\",\n    \"Physical-Diastolic_BP\",\n    \"Physical-HeartRate\",\n    \"Physical-Systolic_BP\",\n    \"Fitness_Endurance-Season\",\n    \"Fitness_Endurance-Max_Stage\",\n    \"Fitness_Endurance-Time_Mins\",\n    \"Fitness_Endurance-Time_Sec\",\n    \"FGC-Season\",\n    \"FGC-FGC_CU\",\n    \"FGC-FGC_CU_Zone\",\n    \"FGC-FGC_GSND\",\n    \"FGC-FGC_GSND_Zone\",\n    \"FGC-FGC_GSD\",\n    \"FGC-FGC_GSD_Zone\",\n    \"FGC-FGC_PU\",\n    \"FGC-FGC_PU_Zone\",\n    \"FGC-FGC_SRL\",\n    \"FGC-FGC_SRL_Zone\",\n    \"FGC-FGC_SRR\",\n    \"FGC-FGC_SRR_Zone\",\n    \"FGC-FGC_TL\",\n    \"FGC-FGC_TL_Zone\",\n    \"BIA-Season\",\n    \"BIA-BIA_Activity_Level_num\",\n    \"BIA-BIA_BMC\",\n    \"BIA-BIA_BMI\",\n    \"BIA-BIA_BMR\",\n    \"BIA-BIA_DEE\",\n    \"BIA-BIA_ECW\",\n    \"BIA-BIA_FFM\",\n    \"BIA-BIA_FFMI\",\n    \"BIA-BIA_FMI\",\n    \"BIA-BIA_Fat\",\n    \"BIA-BIA_Frame_num\",\n    \"BIA-BIA_ICW\",\n    \"BIA-BIA_LDM\",\n    \"BIA-BIA_LST\",\n    \"BIA-BIA_SMM\",\n    \"BIA-BIA_TBW\",\n    \"PAQ_A-Season\",\n    \"PAQ_A-PAQ_A_Total\",\n    \"PAQ_C-Season\",\n    \"PAQ_C-PAQ_C_Total\",\n    \"SDS-Season\",\n    \"SDS-SDS_Total_Raw\",\n    \"SDS-SDS_Total_T\",\n    \"PreInt_EduHx-Season\",\n    \"PreInt_EduHx-computerinternet_hoursday\",\n]","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:36.823161Z","iopub.execute_input":"2024-12-09T13:28:36.823731Z","iopub.status.idle":"2024-12-09T13:28:37.034003Z","shell.execute_reply.started":"2024-12-09T13:28:36.823677Z","shell.execute_reply":"2024-12-09T13:28:37.032942Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Load data\ntrain = pl.read_csv(DATA_DIR / \"train.csv\")\ntest = pl.read_csv(DATA_DIR / \"test.csv\")\ntrain_test = pl.concat([train, test], how=\"diagonal\")\n\nIS_TEST = test.height <= 100\n\nassert_frame_equal(train, train_test[: train.height].select(train.columns))\nassert_frame_equal(test, train_test[train.height :].select(test.columns))\n\n# ....................................................................................\n# Cast string columns to categorical\ntrain_test = train_test.with_columns(cs.string().cast(pl.Categorical).fill_null(\"NAN\"))\ntrain = train_test[: train.height]\ntest = train_test[train.height :]\n\n# ignore rows with null values in TARGET_COLS\ntrain_without_null = train_test.drop_nulls(subset=TARGET_COLS)\nX = train_without_null.select(FEATURE_COLS)\nX_test = test.select(FEATURE_COLS)\ny = train_without_null.select(TARGET_COLS)\ny_sii = y.get_column(\"sii\").to_numpy()  # ground truth\ncat_features = X.select(cs.categorical()).columns","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:37.035208Z","iopub.execute_input":"2024-12-09T13:28:37.035512Z","iopub.status.idle":"2024-12-09T13:28:37.284741Z","shell.execute_reply.started":"2024-12-09T13:28:37.035482Z","shell.execute_reply":"2024-12-09T13:28:37.283654Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class MultiTargetQWK(MultiTargetCustomMetric):\n    def get_final_error(self, error, weight):\n        return np.sum(error)  # / np.sum(weight)\n\n    def is_max_optimal(self):\n        # if True, the bigger the better\n        return True\n\n    def evaluate(self, approxes, targets, weight):\n        # approxes: 予測値 (shape: [ターゲット数, サンプル数])\n        # targets: 実際の値 (shape: [ターゲット数, サンプル数])\n        # weight: サンプルごとの重み (Noneも可)\n\n        approx = np.clip(approxes[-1], 0, 3).round().astype(int)\n        target = targets[-1]\n\n        qwk = cohen_kappa_score(target, approx, weights=\"quadratic\")\n\n        return qwk, 1\n\n    def get_custom_metric_name(self):\n        return \"MultiTargetQWK\"\n\n\nclass OptimizedRounder:\n    \"\"\"\n    A class for optimizing the rounding of continuous predictions into discrete class labels using Optuna.\n    The optimization process maximizes the Quadratic Weighted Kappa score by learning thresholds that separate\n    continuous predictions into class intervals.\n\n    Args:\n        n_classes (int): The number of discrete class labels.\n        n_trials (int, optional): The number of trials for the Optuna optimization. Defaults to 100.\n\n    Attributes:\n        n_classes (int): The number of discrete class labels.\n        labels (NDArray[np.int_]): An array of class labels from 0 to `n_classes - 1`.\n        n_trials (int): The number of optimization trials.\n        metric (Callable): The Quadratic Weighted Kappa score metric used for optimization.\n        thresholds (List[float]): The optimized thresholds learned after calling `fit()`.\n\n    Methods:\n        fit(y_pred: NDArray[np.float_], y_true: NDArray[np.int_]) -> None:\n            Fits the rounding thresholds based on continuous predictions and ground truth labels.\n\n            Args:\n                y_pred (NDArray[np.float_]): Continuous predictions that need to be rounded.\n                y_true (NDArray[np.int_]): Ground truth class labels.\n\n            Returns:\n                None\n\n        predict(y_pred: NDArray[np.float_]) -> NDArray[np.int_]:\n            Predicts discrete class labels by rounding continuous predictions using the fitted thresholds.\n            `fit()` must be called before `predict()`.\n\n            Args:\n                y_pred (NDArray[np.float_]): Continuous predictions to be rounded.\n\n            Returns:\n                NDArray[np.int_]: Predicted class labels.\n\n        _normalize(y: NDArray[np.float_]) -> NDArray[np.float_]:\n            Normalizes the continuous values to the range [0, `n_classes - 1`].\n\n            Args:\n                y (NDArray[np.float_]): Continuous values to be normalized.\n\n            Returns:\n                NDArray[np.float_]: Normalized values.\n\n    References:\n        - This implementation uses Optuna for threshold optimization.\n        - Quadratic Weighted Kappa is used as the evaluation metric.\n    \"\"\"\n\n    def __init__(self, n_classes: int, n_trials: int = 100):\n        self.n_classes = n_classes\n        self.labels = np.arange(n_classes)\n        self.n_trials = n_trials\n        self.metric = partial(cohen_kappa_score, weights=\"quadratic\")\n\n    def fit(self, y_pred: NDArray[np.float_], y_true: NDArray[np.int_]) -> None:\n        y_pred = self._normalize(y_pred)\n\n        def objective(trial: optuna.Trial) -> float:\n            thresholds = []\n            for i in range(self.n_classes - 1):\n                low = max(thresholds) if i > 0 else min(self.labels)\n                high = max(self.labels)\n                th = trial.suggest_float(f\"threshold_{i}\", low, high)\n                thresholds.append(th)\n            try:\n                y_pred_rounded = np.digitize(y_pred, thresholds)\n            except ValueError:\n                return -100\n            return self.metric(y_true, y_pred_rounded)\n\n        optuna.logging.disable_default_handler()\n        study = optuna.create_study(direction=\"maximize\")\n        study.optimize(\n            objective,\n            n_trials=self.n_trials,\n        )\n        self.thresholds = [study.best_params[f\"threshold_{i}\"] for i in range(self.n_classes - 1)]\n\n    def predict(self, y_pred: NDArray[np.float_]) -> NDArray[np.int_]:\n        assert hasattr(self, \"thresholds\"), \"fit() must be called before predict()\"\n        y_pred = self._normalize(y_pred)\n        return np.digitize(y_pred, self.thresholds)\n\n    def _normalize(self, y: NDArray[np.float_]) -> NDArray[np.float_]:\n        # normalize y_pred to [0, n_classes - 1]\n        return (y - y.min()) / (y.max() - y.min()) * (self.n_classes - 1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:37.286757Z","iopub.execute_input":"2024-12-09T13:28:37.287232Z","iopub.status.idle":"2024-12-09T13:28:37.303604Z","shell.execute_reply.started":"2024-12-09T13:28:37.287184Z","shell.execute_reply":"2024-12-09T13:28:37.302274Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# setting catboost parameters\nparams = dict(\n    loss_function=\"MultiRMSE\",\n    eval_metric=MultiTargetQWK(),\n    iterations=1 if IS_TEST else 100000,\n    learning_rate=0.1,\n    depth=5,\n    early_stopping_rounds=50,\n)\n\n# Cross-validation\nskf = StratifiedKFold(n_splits=5, shuffle=True, random_state=52)\nmodels: list[CatBoostRegressor] = []\ny_pred = np.full((X.height, len(TARGET_COLS)), fill_value=np.nan)\nfor train_idx, val_idx in skf.split(X, y_sii):\n    X_train: pl.DataFrame\n    X_val: pl.DataFrame\n    y_train: pl.DataFrame\n    y_val: pl.DataFrame\n    X_train, X_val = X[train_idx], X[val_idx]\n    y_train, y_val = y[train_idx], y[val_idx]\n\n    # train model\n    model = CatBoostRegressor(**params)\n    model.fit(\n        X_train.to_pandas(),\n        y_train.to_pandas(),\n        eval_set=(X_val.to_pandas(), y_val.to_pandas()),\n        cat_features=cat_features,\n        verbose=False,\n    )\n    models.append(model)\n\n    # predict\n    y_pred[val_idx] = model.predict(X_val.to_pandas())\n\nassert np.isnan(y_pred).sum() == 0\n# Optimize thresholds\noptimizer = OptimizedRounder(n_classes=4, n_trials=300)\ny_pred_total = y_pred[:, TARGET_COLS.index(\"PCIAT-PCIAT_Total\")]\noptimizer.fit(y_pred_total, y_sii)\ny_pred_rounded = optimizer.predict(y_pred_total)\n\n# Calculate QWK\nqwk = cohen_kappa_score(y_sii, y_pred_rounded, weights=\"quadratic\")\nprint(f\"Cross-Validated QWK Score: {qwk}\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:37.305147Z","iopub.execute_input":"2024-12-09T13:28:37.305515Z","iopub.status.idle":"2024-12-09T13:28:47.510687Z","shell.execute_reply.started":"2024-12-09T13:28:37.305484Z","shell.execute_reply":"2024-12-09T13:28:47.509733Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class AvgModel:\n    def __init__(self, models: list[BaseEstimator]):\n        self.models = models\n\n    def predict(self, X: ArrayLike) -> NDArray[np.int_]:\n        preds: list[NDArray[np.int_]] = []\n        for model in self.models:\n            pred = model.predict(X)\n            preds.append(pred)\n\n        return np.mean(preds, axis=0)\n\n# ...............................................................................................\navg_model = AvgModel(models)\ntest_pred = avg_model.predict(X_test.to_pandas())[:, TARGET_COLS.index(\"PCIAT-PCIAT_Total\")]\ntest_pred_rounded = optimizer.predict(test_pred)\ntest.select(\"id\").with_columns(\n    pl.Series(\"sii\", pl.Series(\"sii\", test_pred_rounded)),\n).write_csv(\"submission1.csv\")\n\n# Public Score : 0.440\nprint(test_pred_rounded)\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:47.511890Z","iopub.execute_input":"2024-12-09T13:28:47.512214Z","iopub.status.idle":"2024-12-09T13:28:48.743147Z","shell.execute_reply.started":"2024-12-09T13:28:47.512182Z","shell.execute_reply":"2024-12-09T13:28:48.741571Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ Notebook imported - 2</p></div>\n\n### [CMI - Detecting Problematic Digital Behavior](https://www.kaggle.com/code/lennarthaupts/cmi-detecting-problematic-digital-behavior)\n\n- Thanks to : @lennarthaupts","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport os\nfrom sklearn.base import clone\nfrom sklearn.metrics import cohen_kappa_score, make_scorer, confusion_matrix\nfrom sklearn.model_selection import StratifiedKFold, KFold\nfrom scipy.optimize import minimize\nfrom scipy import stats\nfrom concurrent.futures import ThreadPoolExecutor\nfrom tqdm import tqdm\nfrom sklearn.feature_selection import RFECV\nfrom sklearn.linear_model import ElasticNetCV, LassoCV, Lasso\nimport warnings\nfrom lightgbm import LGBMRegressor\nfrom xgboost import XGBRegressor\nfrom catboost import CatBoostRegressor\nimport optuna\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nwarnings.filterwarnings('ignore')\n\n# ......................................................................................................\nSEED = 42\nn_splits = 10\noptimize_params = False\nn_trials = 25 # n_trials for optuna \nbase_thresholds = [30, 50, 80]\ny_model = \"PCIAT-PCIAT_Total\" # Score, target for the model\ny_comp = \"sii\" # Index, target of the competition\n\n# ......................................................................................................\n# Load datasets\ntrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:48.747283Z","iopub.execute_input":"2024-12-09T13:28:48.747818Z","iopub.status.idle":"2024-12-09T13:28:48.833729Z","shell.execute_reply.started":"2024-12-09T13:28:48.747774Z","shell.execute_reply":"2024-12-09T13:28:48.832589Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def time_features(df):\n    # Convert time_of_day to hours\n    df[\"hours\"] = df[\"time_of_day\"] // (3_600 * 1_000_000_000)\n    # Basic features \n    features = [\n        df[\"non-wear_flag\"].mean(),\n        df[\"enmo\"][df[\"enmo\"] >= 0.05].sum(),\n    ]\n    \n    # Define conditions for night, day, and no mask (full data)\n    night = ((df[\"hours\"] >= 22) | (df[\"hours\"] <= 5))\n    day = ((df[\"hours\"] <= 20) & (df[\"hours\"] >= 7))\n    no_mask = np.ones(len(df), dtype=bool)\n    \n    # List of columns of interest and masks\n    keys = [\"enmo\", \"anglez\", \"light\", \"battery_voltage\"]\n    masks = [no_mask, night, day]\n    \n    # Helper function for feature extraction\n    def extract_stats(data):\n        return [\n            data.mean(), \n            data.std(), \n            data.max(), \n            data.min(), \n            data.diff().mean(), \n            data.diff().std()\n        ]\n    \n    # Iterate over keys and masks to generate the statistics\n    for key in keys:\n        for mask in masks:\n            filtered_data = df.loc[mask, key]\n            features.extend(extract_stats(filtered_data))\n\n    return features\n\n# Code for parallelized computation of time series data from: Sheikh Muhammad Abdullah \n# https://www.kaggle.com/code/abdmental01/cmi-best-single-model\ndef process_file(filename, dirname):\n    # Process file and extract time features\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return time_features(df), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    # Load time series from directory in parallel\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    \n    return df","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:48.842109Z","iopub.execute_input":"2024-12-09T13:28:48.842462Z","iopub.status.idle":"2024-12-09T13:28:48.854830Z","shell.execute_reply.started":"2024-12-09T13:28:48.842431Z","shell.execute_reply":"2024-12-09T13:28:48.853440Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\")\ntest_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet\")\n\ntime_series_cols = train_ts.columns.tolist()\n\ntime_series_cols = train_ts.columns.tolist()\ntime_series_cols.remove(\"id\")\n\ntrain = pd.merge(train, train_ts, how=\"left\", on='id')\ntest = pd.merge(test, test_ts, how=\"left\", on='id')\n\ntrain = train.drop('id', axis=1)\ntest = test.drop('id', axis=1)\n\ntrain = train[train[y_comp].notna()] # Keep rows where target is available\ntrain.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:28:48.856294Z","iopub.execute_input":"2024-12-09T13:28:48.856789Z","iopub.status.idle":"2024-12-09T13:30:50.007254Z","shell.execute_reply.started":"2024-12-09T13:28:48.856739Z","shell.execute_reply":"2024-12-09T13:30:50.005752Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Features to exclude, because they're not in test\nexclude = ['PCIAT-Season', 'PCIAT-PCIAT_01', 'PCIAT-PCIAT_02', 'PCIAT-PCIAT_03',\n           'PCIAT-PCIAT_04', 'PCIAT-PCIAT_05', 'PCIAT-PCIAT_06', 'PCIAT-PCIAT_07',\n           'PCIAT-PCIAT_08', 'PCIAT-PCIAT_09', 'PCIAT-PCIAT_10', 'PCIAT-PCIAT_11',\n           'PCIAT-PCIAT_12', 'PCIAT-PCIAT_13', 'PCIAT-PCIAT_14', 'PCIAT-PCIAT_15',\n           'PCIAT-PCIAT_16', 'PCIAT-PCIAT_17', 'PCIAT-PCIAT_18', 'PCIAT-PCIAT_19',\n           'PCIAT-PCIAT_20', 'PCIAT-PCIAT_Total', 'sii']\n\nfeatures = [f for f in train.columns if f not in exclude]\n\n# Categorical features\ncat_c = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', 'Fitness_Endurance-Season', \n          'FGC-Season', 'BIA-Season', 'PAQ_A-Season', 'PAQ_C-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n\nfor col in cat_c:\n    a_map = {}\n    all_unique = set(train[col].unique()) | set(test[col].unique())\n    for i, value in enumerate(all_unique):\n        a_map[value] = i\n\n    train[col] = train[col].map(a_map)\n    test[col] = test[col].map(a_map)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:30:50.009194Z","iopub.execute_input":"2024-12-09T13:30:50.009574Z","iopub.status.idle":"2024-12-09T13:30:50.039844Z","shell.execute_reply.started":"2024-12-09T13:30:50.009542Z","shell.execute_reply":"2024-12-09T13:30:50.038591Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"class Impute_With_Model:\n    \n    def __init__(self, na_frac=0.5, min_samples=0):\n        self.model_dict = {}\n        self.mean_dict = {}\n        self.features = None\n        self.na_frac = na_frac\n        self.min_samples = min_samples\n        \n    def find_features(self, data, feature, tmp_features):\n        missing_rows = data[feature].isna()\n        na_fraction = data[missing_rows][tmp_features].isna().mean(axis=0)\n        valid_features = np.array(tmp_features)[na_fraction <= self.na_frac]\n        return valid_features\n\n    def fit_models(self, model, data, features):\n        self.features = features\n        n_data = data.shape[0]\n        for feature in features:\n            self.mean_dict[feature] = np.mean(data[feature])\n        for feature in tqdm(features):\n            if data[feature].isna().sum() > 0:\n                model_clone = clone(model)\n                X = data[data[feature].notna()].copy()\n                tmp_features = [f for f in features if f != feature]\n                tmp_features = self.find_features(data, feature, tmp_features)\n                if len(tmp_features) >= 1 and X.shape[0] > self.min_samples:\n                    for f in tmp_features:\n                        X[f] = X[f].fillna(self.mean_dict[f])\n                    model_clone.fit(X[tmp_features], X[feature])\n                    self.model_dict[feature] = (model_clone, tmp_features.copy())\n                else:\n                    self.model_dict[feature] = (\"mean\", np.mean(data[feature]))\n            \n    def impute(self, data):\n        imputed_data = data.copy()\n        for feature, model in self.model_dict.items():\n            missing_rows = imputed_data[feature].isna()\n            if missing_rows.any():\n                if model[0] == \"mean\":\n                    imputed_data[feature].fillna(model[1], inplace=True)\n                else:\n                    tmp_features = [f for f in self.features if f != feature]\n                    X_missing = data.loc[missing_rows, tmp_features].copy()\n                    for f in tmp_features:\n                        X_missing[f] = X_missing[f].fillna(self.mean_dict[f])\n                    imputed_data.loc[missing_rows, feature] = model[0].predict(X_missing[model[1]])\n        return imputed_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:30:50.041437Z","iopub.execute_input":"2024-12-09T13:30:50.041911Z","iopub.status.idle":"2024-12-09T13:30:50.057045Z","shell.execute_reply.started":"2024-12-09T13:30:50.041854Z","shell.execute_reply":"2024-12-09T13:30:50.055720Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Code for finding optimal thresholds copied from: Michael Semenoff\n# https://www.kaggle.com/code/michaelsemenoff/cmi-actigraphy-feature-engineering-selection\ndef round_with_thresholds(raw_preds, thresholds):\n    return np.where(raw_preds < thresholds[0], int(0),\n                    np.where(raw_preds < thresholds[1], int(1),\n                             np.where(raw_preds < thresholds[2], int(2), int(3))))\n\ndef optimize_thresholds(y_true, raw_preds, start_vals=[0.5, 1.5, 2.5]):\n    def fun(thresholds, y_true, raw_preds):\n        rounded_preds = round_with_thresholds(raw_preds, thresholds)\n        return -cohen_kappa_score(y_true, rounded_preds, weights='quadratic')\n\n    res = minimize(fun, x0=start_vals, args=(y_true, raw_preds), method='Nelder-Mead')\n    assert res.success\n    return res.x\n\n# .............................................................................................\ndef calculate_weights(series):\n    # Create bins for the target variable and assign weights based on frequency\n    bins = pd.cut(series, bins=10, labels=False)\n    weights = bins.value_counts().reset_index()\n    weights.columns = ['target_bins', 'count']\n    weights['count'] = 1 / weights['count']\n    weight_map = weights.set_index('target_bins')['count'].to_dict()\n    weights = bins.map(weight_map)\n    return weights / weights.mean() \n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:30:50.058811Z","iopub.execute_input":"2024-12-09T13:30:50.059280Z","iopub.status.idle":"2024-12-09T13:30:50.076551Z","shell.execute_reply.started":"2024-12-09T13:30:50.059233Z","shell.execute_reply":"2024-12-09T13:30:50.075574Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def cross_validate(model_, data, features, score_col, index_col, cv, sample_weights=False, verbose=False):\n    \"\"\"\n    Perform cross-validation with a given model and compute the out-of-fold \n    predictions and Cohen's Kappa score for each fold.\n\n    Returns:\n    float: Mean Kappa score across all folds.\n    array: Out-of-fold score predictions for the entire dataset.\n    \"\"\"\n    kappa_scores = [] \n    oof_score_predictions = np.zeros(len(data))  \n\n    score_to_index_thresholds = base_thresholds  \n\n    for fold_idx, (train_idx, val_idx) in enumerate(cv.split(data, data[index_col])):\n        X_train, X_val = data[features].iloc[train_idx], data[features].iloc[val_idx]\n        y_train_score = data[score_col].iloc[train_idx]  \n        y_val_score = data[score_col].iloc[val_idx]      \n        y_val_index = data[index_col].iloc[val_idx]     \n        \n        # Train model with sample weights if provided\n        if sample_weights:\n            weights = calculate_weights(y_train_score)\n            model_.fit(X_train, y_train_score, sample_weight=weights)\n        else:\n            model_.fit(X_train, y_train_score)\n        \n        y_pred_val_score = model_.predict(X_val)\n        \n        oof_score_predictions[val_idx] = y_pred_val_score \n\n        y_pred_val_index = round_with_thresholds(y_pred_val_score, score_to_index_thresholds)\n\n        kappa_score = cohen_kappa_score(y_val_index, y_pred_val_index, weights='quadratic')\n        kappa_scores.append(kappa_score)\n        \n        if verbose:\n            print(f\"Fold {fold_idx}: Kappa Score = {kappa_score}\")\n    \n    if verbose:\n        print(f\"## Mean CV Kappa Score: {np.mean(kappa_scores)} ##\")\n    \n    return np.mean(kappa_scores), oof_score_predictions","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:30:50.078140Z","iopub.execute_input":"2024-12-09T13:30:50.078525Z","iopub.status.idle":"2024-12-09T13:30:50.095349Z","shell.execute_reply.started":"2024-12-09T13:30:50.078483Z","shell.execute_reply":"2024-12-09T13:30:50.094280Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def objective(trial, model_type, X, features, score_col, index_col, cv, sample_weights=False):\n    # Parameter space to explore if model is xgboost\n    if model_type == 'xgboost':\n        params = {\n            'objective': trial.suggest_categorical('objective', ['reg:squarederror','reg:tweedie', 'reg:pseudohubererror']),\n            'random_state': SEED,\n            'n_estimators': trial.suggest_int('n_estimators', 300, 600),\n            'max_depth': trial.suggest_int('max_depth', 3, 8),\n            'learning_rate': trial.suggest_loguniform('learning_rate', 0.02, 0.1),\n            'subsample': trial.suggest_float('subsample', 0.5, 1.0),\n            'colsample_bytree': trial.suggest_float('colsample_bytree', 0.5, 1.0),\n            'gamma': trial.suggest_float('gamma', 0.0, 5.0),\n            'reg_alpha': trial.suggest_loguniform('reg_alpha', 1e-5, 1e-1),\n            'reg_lambda': trial.suggest_loguniform('reg_lambda', 1e-5, 1e-1)\n        }\n        if params['objective'] == 'reg:tweedie':\n            params['tweedie_variance_power'] = trial.suggest_float('tweedie_variance_power', 1, 2)\n        model = XGBRegressor(**params, use_label_encoder=False)\n    \n    # Parameter space to explore if model is lightgbm\n    elif model_type == 'lightgbm':\n        params = {\n            'objective': trial.suggest_categorical('objective', ['poisson', 'tweedie', 'regression']),\n            'random_state': SEED,\n            'verbosity': -1,\n            'n_estimators': trial.suggest_int('n_estimators', 100, 600),\n            'max_depth': trial.suggest_int('max_depth', 3, 10),\n            'learning_rate': trial.suggest_loguniform('learning_rate', 0.01, 0.3),\n            'subsample': trial.suggest_float('subsample', 0.5, 1.0),\n            'colsample_bytree': trial.suggest_float('colsample_bytree', 0.5, 1.0),\n        }\n        if params['objective'] == 'tweedie':\n            params['tweedie_variance_power'] = trial.suggest_float('tweedie_variance_power', 1, 2)\n        model = LGBMRegressor(**params)\n    \n    # Parameter space to explore if model is catboost\n    elif model_type == 'catboost':\n        params = {\n            'loss_function': trial.suggest_categorical('objective', ['Tweedie:variance_power=1.5', 'Poisson', 'RMSE']),\n            'random_state': SEED,\n            'iterations': trial.suggest_int('iterations', 200, 600),\n            'depth': trial.suggest_int('depth', 4, 10),\n            'learning_rate': trial.suggest_loguniform('learning_rate', 0.01, 0.1),\n            'l2_leaf_reg': trial.suggest_loguniform('l2_leaf_reg', 1e-5, 1e-1),\n            'subsample': trial.suggest_float('subsample', 0.5, 1.0),\n            'bagging_temperature': trial.suggest_float('bagging_temperature', 0.0, 1.0),\n            'random_strength': trial.suggest_float('random_strength', 1e-3, 10.0),\n            'colsample_bylevel': trial.suggest_float('colsample_bylevel', 0.5, 1.0),\n            'min_data_in_leaf': trial.suggest_int('min_data_in_leaf', 1, 100),\n        }\n        model = CatBoostRegressor(**params, verbose=0)\n    \n    else:\n        raise ValueError(f\"Unsupported model_type: {model_type}\")\n\n    score, _ = cross_validate(model, X, features, score_col, index_col, cv, sample_weights=True, verbose=False)\n\n    return score\n\ndef run_optimization(X, features, score_col, index_col, model_type, n_trials=30, cv=None, sample_weights=False):\n    study = optuna.create_study(direction=\"maximize\")\n    study.optimize(lambda trial: objective(trial, model_type, X, features, score_col, index_col, cv, sample_weights), \n                   n_trials=n_trials)\n    \n    print(f\"Best params for {model_type}: {study.best_params}\")\n    print(f\"Best score: {study.best_value}\")\n    return study.best_params","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:30:50.097208Z","iopub.execute_input":"2024-12-09T13:30:50.098246Z","iopub.status.idle":"2024-12-09T13:30:50.116968Z","shell.execute_reply.started":"2024-12-09T13:30:50.098193Z","shell.execute_reply":"2024-12-09T13:30:50.115712Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Replace if subsets for features have been selected\nlgb_features = features\nxgb_features = features\ncat_features = features\n\n# ...........................................................................\n# Parameters for LGBM, XGB and CatBoost\nlgb_params = {\n    'objective': 'tweedie', \n    'n_estimators': 242, \n    'max_depth': 4, \n    'learning_rate': 0.029229916231368648, \n    'subsample': 0.9435713052516868, \n    'colsample_bytree': 0.6372563562692964, \n    'tweedie_variance_power': 1.7598875942201002\n}\n\nxgb_params = {\n    'objective': 'reg:tweedie', \n    'n_estimators': 554, \n    'max_depth': 3, \n    'learning_rate': 0.020148793517835852, \n    'subsample': 0.7245109070247534, \n    'colsample_bytree': 0.7516980111036932, \n    'gamma': 1.4405479996512962, \n    'reg_alpha': 0.00015467164351805926, \n    'reg_lambda': 0.011510449488765364, \n    'tweedie_variance_power': 1.2525085567413385\n}\n\ncat_params = {\n    'objective': 'RMSE', \n    'iterations': 476, \n    'depth': 6, \n    'learning_rate': 0.01508021072978329, \n    'l2_leaf_reg': 0.009219274204258077, \n    'subsample': 0.909899776448952, \n    'bagging_temperature': 0.4068004305795976, \n    'random_strength': 0.13085860045085365, \n    'colsample_bylevel': 0.5000595287404359,\n    'min_data_in_leaf': 27\n}\n\n# ...........................................................................\nmodel = Lasso(alpha=0.3, random_state=SEED)\nimputer = Impute_With_Model(na_frac=0.3)\nimputer.fit_models(model, train, features)\ntrain = imputer.impute(train)\ntest = imputer.impute(test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:30:50.118654Z","iopub.execute_input":"2024-12-09T13:30:50.119170Z","iopub.status.idle":"2024-12-09T13:31:13.597863Z","shell.execute_reply.started":"2024-12-09T13:30:50.119114Z","shell.execute_reply":"2024-12-09T13:31:13.596873Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"kf = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=SEED)\n\nif optimize_params:\n    # LightGBM Optimization\n    lgb_params = run_optimization(train, lgb_features, y_model, y_comp, 'lightgbm', n_trials=n_trials, cv=kf, sample_weights=True)\n\n    # XGBoost Optimization\n    xgb_params = run_optimization(train, xgb_features, y_model, y_comp, 'xgboost', n_trials=n_trials, cv=kf, sample_weights=True)\n\n    # CatBoost Optimization\n    cat_params = run_optimization(train, cat_features, y_model, y_comp, 'catboost', n_trials=n_trials, cv=kf, sample_weights=True)\n\n# ....................................................................................................................................\n# Define models\nlgb_model = LGBMRegressor(**lgb_params, random_state=SEED, verbosity=-1)\nxgb_model = XGBRegressor(**xgb_params, random_state=SEED, verbosity=0)\ncat_model = CatBoostRegressor(**cat_params, random_state=SEED, verbose=0)\n\nweights = calculate_weights(train[y_model])\n\n# Cross-validate LGBM model\nscore_lgb, oof_lgb = cross_validate(lgb_model, train, lgb_features, y_model, y_comp, kf, verbose=True, sample_weights=True)\nlgb_model.fit(train[lgb_features], train[y_model], sample_weight=weights)\ntest_lgb = lgb_model.predict(test[lgb_features])\n\n# Cross-validate XGBoost model\nscore_xgb, oof_xgb = cross_validate(xgb_model, train, xgb_features, y_model, y_comp, kf, verbose=True, sample_weights=True)\nxgb_model.fit(train[xgb_features], train[y_model], sample_weight=weights)\ntest_xgb = xgb_model.predict(test[xgb_features])\n\n# Cross-validate CatBoost model\nscore_cat, oof_cat = cross_validate(cat_model, train, cat_features, y_model, y_comp, kf, verbose=True, sample_weights=True)\ncat_model.fit(train[cat_features], train[y_model], sample_weight=weights)\ntest_cat = cat_model.predict(test[cat_features])\n\nprint(f'Overall Mean Kappa: {np.mean([score_lgb, score_xgb, score_cat])}')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:31:13.599250Z","iopub.execute_input":"2024-12-09T13:31:13.599626Z","iopub.status.idle":"2024-12-09T13:33:39.499553Z","shell.execute_reply.started":"2024-12-09T13:31:13.599591Z","shell.execute_reply":"2024-12-09T13:33:39.498212Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Optimize thresholds for each model's OOF predictions\nlgb_thresholds = optimize_thresholds(train[y_comp], oof_lgb, start_vals=base_thresholds)\nprint(f\"LGBM optimized thresholds: {lgb_thresholds}\")\n\nxgb_thresholds = optimize_thresholds(train[y_comp], oof_xgb, start_vals=base_thresholds)\nprint(f\"XGBoost optimized thresholds: {xgb_thresholds}\")\n\ncat_thresholds = optimize_thresholds(train[y_comp], oof_cat, start_vals=base_thresholds)\nprint(f\"CatBoost optimized thresholds: {cat_thresholds}\")\n\n# Apply the optimized thresholds to OOF predictions\noof_lgb = round_with_thresholds(oof_lgb, lgb_thresholds)\noof_xgb = round_with_thresholds(oof_xgb, xgb_thresholds)\noof_cat = round_with_thresholds(oof_cat, cat_thresholds)\nvoted_oof = stats.mode(np.array([oof_lgb, oof_xgb, oof_cat]), axis=0).mode.flatten().astype(int)\n\n# Calculate Kappa score for voted OOF predictions\nkappa_score = cohen_kappa_score(train[y_comp], voted_oof, weights='quadratic')\nprint(f\"Voted ensemble Kappa score: {kappa_score}\")\n\n# Apply the optimized thresholds to test predictions\ntest_lgb = round_with_thresholds(test_lgb, lgb_thresholds)\ntest_xgb = round_with_thresholds(test_xgb, xgb_thresholds)\ntest_cat = round_with_thresholds(test_cat, cat_thresholds)\nvoted_test = stats.mode(np.array([test_lgb, test_xgb, test_cat]), axis=0).mode.flatten().astype(int)\n\nsub_mission = pd.read_csv(\"/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv\")\nsub_mission[y_comp] = voted_test\nsub_mission.to_csv(\"submission2.csv\", index=False)\n# Public Score : 0.457\nprint(voted_test)\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:33:39.501359Z","iopub.execute_input":"2024-12-09T13:33:39.501749Z","iopub.status.idle":"2024-12-09T13:33:41.942578Z","shell.execute_reply.started":"2024-12-09T13:33:39.501708Z","shell.execute_reply":"2024-12-09T13:33:41.940889Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ Notebook imported - 3</p></div>\n\n### [CMI Problematic Internet Use](https://www.kaggle.com/code/vishnupriyagarige/cmi-problematic-internet-use/notebook)\n\n- Thanks to : @vishnupriyagarige","metadata":{}},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom sklearn.preprocessing import StandardScaler, LabelEncoder\n\nfrom xgboost import XGBRegressor\nimport xgboost as xgb\nfrom lightgbm import LGBMRegressor,LGBMClassifier\nimport lightgbm as lgb\nfrom catboost import CatBoostRegressor\nfrom sklearn.ensemble import VotingRegressor\n\nfrom sklearn.base import clone\nfrom sklearn.model_selection import KFold\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import cohen_kappa_score\nfrom scipy.optimize import minimize\n\n# ................................................................................................................\ntrain_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\nsample_data = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:33:41.945528Z","iopub.execute_input":"2024-12-09T13:33:41.946191Z","iopub.status.idle":"2024-12-09T13:33:42.165867Z","shell.execute_reply.started":"2024-12-09T13:33:41.946127Z","shell.execute_reply":"2024-12-09T13:33:42.164709Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"from tqdm import tqdm\nfrom IPython.display import clear_output\nfrom concurrent.futures import ThreadPoolExecutor\n\ndef process_file(filename, dirname):\n    data = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    data.drop('step', axis=1, inplace=True)\n    return data.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    data = pd.DataFrame(stats, columns=[f\"stat_{i}\" for i in range(len(stats[0]))])\n    data['id'] = indexes\n    \n    return data\n\n# ......................................................................................................................\ntrain_parquet = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\")\ntest_parquet = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet\")\ntime_series_cols = train_parquet.columns.tolist()\ntime_series_cols.remove(\"id\")","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:33:42.167733Z","iopub.execute_input":"2024-12-09T13:33:42.168190Z","iopub.status.idle":"2024-12-09T13:36:04.494896Z","shell.execute_reply.started":"2024-12-09T13:33:42.168142Z","shell.execute_reply":"2024-12-09T13:36:04.493571Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data = pd.merge(train_data, train_parquet, how=\"left\", on='id')\ntest_data = pd.merge(test_data, test_parquet, how=\"left\", on='id')\n\n# ............................................................................\ntrain_data = train_data.drop('id',axis=1)\ntest_data = test_data.drop('id',axis=1)\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.496369Z","iopub.execute_input":"2024-12-09T13:36:04.496767Z","iopub.status.idle":"2024-12-09T13:36:04.517952Z","shell.execute_reply.started":"2024-12-09T13:36:04.496731Z","shell.execute_reply":"2024-12-09T13:36:04.516587Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Example columns in your train and test data\ntrain_columns = train_data.columns.tolist()  # Train columns including target\ntest_columns = test_data.columns.tolist()    # Test columns\n\n# Identify common feature columns between train and test (excluding target)\ncommon_columns = [col for col in train_columns if col in test_columns]\n\n# Include the target column explicitly in the final train set\ncommon_columns.append('sii')\n\n# Now, reduce the training data to only the common feature columns + target\ntrain_data = train_data[common_columns]\n\n# Print the resulting columns in the training data\nprint(\"Train data columns:\", len(train_data.columns))\nprint(\"Test data columns:\", len(test_data.columns))","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.521918Z","iopub.execute_input":"2024-12-09T13:36:04.522313Z","iopub.status.idle":"2024-12-09T13:36:04.532992Z","shell.execute_reply.started":"2024-12-09T13:36:04.522276Z","shell.execute_reply":"2024-12-09T13:36:04.531819Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"train_data = train_data.dropna(subset='sii')\n\nnum_cols = list(train_data.select_dtypes(exclude=['object']).columns.difference(['sii']))\ncat_cols = list(train_data.select_dtypes(include=['object']).columns)\n\nnum_cols_test = list(test_data.select_dtypes(exclude=['object']).columns)\ncat_cols_test = list(test_data.select_dtypes(include=['object']).columns)\n\n#num_cols_test = [col for col in num_cols_test if col not in ['id']]\n\n# ...............................................................................................\nfor col in cat_cols:\n    train_data[col] = train_data[col].fillna('missing')\n    train_data[col] = train_data[col].astype('category')\n    \n    test_data[col] = test_data[col].fillna('missing')\n    test_data[col] = test_data[col].astype('category')\n\n# ...............................................................................................\nfor col in num_cols:\n    train_data[col] = train_data[col].fillna(train_data[col].median())\n    test_data[col] = test_data[col].fillna(test_data[col].median())\n","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.534549Z","iopub.execute_input":"2024-12-09T13:36:04.534944Z","iopub.status.idle":"2024-12-09T13:36:04.716580Z","shell.execute_reply.started":"2024-12-09T13:36:04.534910Z","shell.execute_reply":"2024-12-09T13:36:04.715239Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"#  object datatype columns encoding:\n\nlabelencoder = LabelEncoder()\nfor col_name in cat_cols:\n    train_data[col_name]=labelencoder.fit_transform(train_data[col_name]).astype(int)\n        \nfor col_name in cat_cols_test:\n    test_data[col_name]=labelencoder.transform(test_data[col_name]).astype(int)\n\n# ..........................................................................................\nscaler = StandardScaler()\ntrain_data[num_cols] = scaler.fit_transform(train_data[num_cols])\ntest_data[num_cols] = scaler.transform(test_data[num_cols])\n\n# ..........................................................................................\nfrom sklearn.model_selection import train_test_split\nX = train_data.drop(['sii'], axis=1)\ny = train_data['sii']\ntest = test_data","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.717921Z","iopub.execute_input":"2024-12-09T13:36:04.718244Z","iopub.status.idle":"2024-12-09T13:36:04.794469Z","shell.execute_reply.started":"2024-12-09T13:36:04.718212Z","shell.execute_reply":"2024-12-09T13:36:04.793210Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"params_lgb = {'learning_rate': 0.060210090165748686, 'n_estimators': 676, 'max_depth': 5, 'num_leaves': 193, 'min_child_weight': 1.4676211795900709, 'subsample': 0.9176759029661259, 'colsample_bytree': 0.6228483814299844, 'lambda_l1': 5.972758940714118, 'lambda_l2': 0.08520209502101517}\n#Best QWK score: 0.4036\nparams_xgb = {'learning_rate': 0.07655257571702724, 'n_estimators': 688, 'max_depth': 7, 'min_child_weight': 10, 'subsample': 0.8740939473627481, 'colsample_bytree': 0.9986562622011108, 'gamma': 0.005098593898531702, 'reg_alpha': 9.637641942675724, 'reg_lambda': 0.014395773764050573}\n#Best QWK score: 0.4086\nparams_cat = {'iterations': 587, 'learning_rate': 0.055230940995657174, 'depth': 4, 'l2_leaf_reg': 0.00018791609018454318, 'subsample': 0.6500825893922675, 'colsample_bylevel': 0.9880985604359044, 'random_strength': 0.12043764855944512, 'bagging_temperature': 0.0008351502400011265, 'min_data_in_leaf': 21}\n#Best QWK score: 0.3859","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.796030Z","iopub.execute_input":"2024-12-09T13:36:04.796359Z","iopub.status.idle":"2024-12-09T13:36:04.804279Z","shell.execute_reply.started":"2024-12-09T13:36:04.796329Z","shell.execute_reply":"2024-12-09T13:36:04.802778Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\ndef threshold_Rounder(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0,\n                    np.where(oof_non_rounded < thresholds[1], 1,\n                             np.where(oof_non_rounded < thresholds[2], 2, 3)))\n\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    rounded_p = threshold_Rounder(oof_non_rounded, thresholds)\n    return -quadratic_weighted_kappa(y_true, rounded_p)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.806014Z","iopub.execute_input":"2024-12-09T13:36:04.806519Z","iopub.status.idle":"2024-12-09T13:36:04.818060Z","shell.execute_reply.started":"2024-12-09T13:36:04.806467Z","shell.execute_reply":"2024-12-09T13:36:04.816917Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"n_splits = 5\ndef train_model(model_class, test_data):\n    \n    X = train_data.drop(['sii'], axis=1)\n    y = train_data['sii']\n    test = test_data\n\n    SKF = StratifiedKFold(n_splits=5, shuffle=True, random_state=42)\n    \n    models = []\n    train_pred = []\n    test_pred = []\n    \n    oof_non_rounded = np.zeros(len(y), dtype=float) \n    oof_rounded = np.zeros(len(y), dtype=int) \n    test_preds = np.zeros((len(test_data), n_splits))\n    for fold, (train_idx, test_idx) in enumerate(tqdm(SKF.split(X, y), desc=\"Training Folds\", total=n_splits)):\n        X_train, X_val = X.iloc[train_idx], X.iloc[test_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[test_idx]\n\n        model = clone(model_class)\n        model.fit(X_train, y_train)\n\n        y_train_pred = model.predict(X_train)\n        y_val_pred = model.predict(X_val)\n\n        oof_non_rounded[test_idx] = y_val_pred\n        y_val_pred_rounded = y_val_pred.round(0).astype(int)\n        oof_rounded[test_idx] = y_val_pred_rounded\n\n        train_kappa = quadratic_weighted_kappa(y_train, y_train_pred.round(0).astype(int))\n        val_kappa = quadratic_weighted_kappa(y_val, y_val_pred_rounded)\n\n        train_pred.append(train_kappa)\n        test_pred.append(val_kappa)\n        \n        test_preds[:, fold] = model.predict(test)\n        \n        print(f\"Fold {fold+1} - Train QWK: {train_kappa:.4f}, Validation QWK: {val_kappa:.4f}\")\n        clear_output(wait=True)\n        models.append(model)\n\n    print(f\"Mean Train QWK : {np.mean(train_pred):.4f}\")\n    print(f\"Mean Validation QWK : {np.mean(test_pred):.4f}\")\n\n    KappaOPtimizer = minimize(evaluate_predictions, x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), method='Nelder-Mead') \n    assert KappaOPtimizer.success, \"Optimization did not converge.\"\n    \n    oof_tuned = threshold_Rounder(oof_non_rounded, KappaOPtimizer.x)\n    tKappa = quadratic_weighted_kappa(y, oof_tuned)\n\n    print(f\"Optimized QWK SCORE :: {Fore.CYAN}{Style.BRIGHT} {tKappa:.3f}{Style.RESET_ALL}\")\n\n    pred_mean = test_preds.mean(axis=1)\n    pred = threshold_Rounder(pred_mean, KappaOPtimizer.x)\n    \n    submission = pd.DataFrame({'id': sample_data['id'],'sii': pred})\n\n    return submission","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.819708Z","iopub.execute_input":"2024-12-09T13:36:04.820133Z","iopub.status.idle":"2024-12-09T13:36:04.837272Z","shell.execute_reply.started":"2024-12-09T13:36:04.820080Z","shell.execute_reply":"2024-12-09T13:36:04.835975Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"!pip install colorama\nfrom colorama import Fore, Style, init\n\n# ............................................................................\n# Create model instances\nlgb_model = LGBMRegressor(**params_lgb, random_state=42, verbosity=-1)\nXGB_Model = XGBRegressor(**params_xgb, random_state=42)\nCatBoost_Model = CatBoostRegressor(**params_cat,verbose=0)\n\n# Combine models using Voting Regressor\nvoting_model = VotingRegressor(estimators=[('lightgbm', lgb_model),\n                                           ('xgboost', XGB_Model),\n                                           ('catboost', CatBoost_Model)])","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:04.838900Z","iopub.execute_input":"2024-12-09T13:36:04.839374Z","iopub.status.idle":"2024-12-09T13:36:49.870863Z","shell.execute_reply.started":"2024-12-09T13:36:04.839323Z","shell.execute_reply":"2024-12-09T13:36:49.869356Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Train the ensemble model\nSub_vote = train_model(voting_model, test)\n\n# ...............................................\n# Saving submission:\nSub_vote.to_csv('submission3.csv', index=False)\nvote = Sub_vote['sii'].values.copy()\n# Public Score : 0.464\nprint(vote)\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:36:49.873058Z","iopub.execute_input":"2024-12-09T13:36:49.873612Z","iopub.status.idle":"2024-12-09T13:37:38.660542Z","shell.execute_reply.started":"2024-12-09T13:36:49.873556Z","shell.execute_reply":"2024-12-09T13:37:38.658571Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ Notebook imported - 4</p></div>\n\n### [CMI | Best Single Model](https://www.kaggle.com/code/abdmental01/cmi-best-single-model)\n\n- Thanks to : @abdmental01","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport polars as pl\nimport pandas as pd\nfrom sklearn.base import clone\nfrom copy import deepcopy\nimport optuna\nfrom scipy.optimize import minimize\nimport os\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport re\nfrom colorama import Fore, Style\n\nfrom tqdm import tqdm\nfrom IPython.display import clear_output\nfrom concurrent.futures import ThreadPoolExecutor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None\n\nimport lightgbm as lgb\nfrom catboost import CatBoostRegressor, CatBoostClassifier\nfrom xgboost import XGBRegressor\nfrom sklearn.ensemble import VotingRegressor\nfrom sklearn.model_selection import *\nfrom sklearn.metrics import *\n\nSEED = 42\nn_splits = 5","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:37:38.663184Z","iopub.execute_input":"2024-12-09T13:37:38.663769Z","iopub.status.idle":"2024-12-09T13:37:38.676067Z","shell.execute_reply.started":"2024-12-09T13:37:38.663695Z","shell.execute_reply":"2024-12-09T13:37:38.674801Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop('step', axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(executor.map(lambda fname: process_file(fname, dirname), ids), total=len(ids)))\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    \n    return df\n\ntrain = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/train.csv')\ntest = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/test.csv')\nsample = pd.read_csv('/kaggle/input/child-mind-institute-problematic-internet-use/sample_submission.csv')\n\ntrain_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_train.parquet\")\ntest_ts = load_time_series(\"/kaggle/input/child-mind-institute-problematic-internet-use/series_test.parquet\")\ntime_series_cols = train_ts.columns.tolist()\ntime_series_cols.remove(\"id\")\n\ntrain = pd.merge(train, train_ts, how=\"left\", on='id')\ntest = pd.merge(test, test_ts, how=\"left\", on='id')\n\ntrain = train.drop('id', axis=1)\ntest = test.drop('id', axis=1)\n\nfeaturesCols = ['Basic_Demos-Enroll_Season', 'Basic_Demos-Age', 'Basic_Demos-Sex',\n                'CGAS-Season', 'CGAS-CGAS_Score', 'Physical-Season', 'Physical-BMI',\n                'Physical-Height', 'Physical-Weight', 'Physical-Waist_Circumference',\n                'Physical-Diastolic_BP', 'Physical-HeartRate', 'Physical-Systolic_BP',\n                'Fitness_Endurance-Season', 'Fitness_Endurance-Max_Stage',\n                'Fitness_Endurance-Time_Mins', 'Fitness_Endurance-Time_Sec',\n                'FGC-Season', 'FGC-FGC_CU', 'FGC-FGC_CU_Zone', 'FGC-FGC_GSND',\n                'FGC-FGC_GSND_Zone', 'FGC-FGC_GSD', 'FGC-FGC_GSD_Zone', 'FGC-FGC_PU',\n                'FGC-FGC_PU_Zone', 'FGC-FGC_SRL', 'FGC-FGC_SRL_Zone', 'FGC-FGC_SRR',\n                'FGC-FGC_SRR_Zone', 'FGC-FGC_TL', 'FGC-FGC_TL_Zone', 'BIA-Season',\n                'BIA-BIA_Activity_Level_num', 'BIA-BIA_BMC', 'BIA-BIA_BMI',\n                'BIA-BIA_BMR', 'BIA-BIA_DEE', 'BIA-BIA_ECW', 'BIA-BIA_FFM',\n                'BIA-BIA_FFMI', 'BIA-BIA_FMI', 'BIA-BIA_Fat', 'BIA-BIA_Frame_num',\n                'BIA-BIA_ICW', 'BIA-BIA_LDM', 'BIA-BIA_LST', 'BIA-BIA_SMM',\n                'BIA-BIA_TBW', 'PAQ_A-Season', 'PAQ_A-PAQ_A_Total', 'PAQ_C-Season',\n                'PAQ_C-PAQ_C_Total', 'SDS-Season', 'SDS-SDS_Total_Raw',\n                'SDS-SDS_Total_T', 'PreInt_EduHx-Season',\n                'PreInt_EduHx-computerinternet_hoursday', 'sii']\n\nfeaturesCols += time_series_cols\n\ntrain = train[featuresCols]\ntrain = train.dropna(subset='sii')\n\ncat_c = ['Basic_Demos-Enroll_Season', 'CGAS-Season', 'Physical-Season', 'Fitness_Endurance-Season', \n          'FGC-Season', 'BIA-Season', 'PAQ_A-Season', 'PAQ_C-Season', 'SDS-Season', 'PreInt_EduHx-Season']\n\ndef update(df):\n    for c in cat_c: \n        df[c] = df[c].fillna('Missing')\n        df[c] = df[c].astype('category')\n    return df\n        \ntrain = update(train)\ntest = update(test)\n\ndef create_mapping(column, dataset):\n    unique_values = dataset[column].unique()\n    return {value: idx for idx, value in enumerate(unique_values)}\n\n\"\"\"This Mapping Works Fine For me I also Check Each Values in Train and test Using Logic. There no Data Lekage.\"\"\"\n\nfor col in cat_c:\n    mapping_train = create_mapping(col, train)\n    mapping_test = create_mapping(col, test)\n    \n    train[col] = train[col].replace(mapping_train).astype(int)\n    test[col] = test[col].replace(mapping_test).astype(int)\n\nprint(f'Train Shape : {train.shape} || Test Shape : {test.shape}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:37:38.677945Z","iopub.execute_input":"2024-12-09T13:37:38.678318Z","iopub.status.idle":"2024-12-09T13:40:00.176941Z","shell.execute_reply.started":"2024-12-09T13:37:38.678282Z","shell.execute_reply":"2024-12-09T13:40:00.175701Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"def quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\ndef threshold_Rounder(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0,\n                    np.where(oof_non_rounded < thresholds[1], 1,\n                             np.where(oof_non_rounded < thresholds[2], 2, 3)))\n\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    rounded_p = threshold_Rounder(oof_non_rounded, thresholds)\n    return -quadratic_weighted_kappa(y_true, rounded_p)\n\ndef TrainML(model_class, test_data):\n    \n    X = train.drop(['sii'], axis=1)\n    y = train['sii']\n\n    SKF = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=SEED)\n    \n    train_S = []\n    test_S = []\n    \n    oof_non_rounded = np.zeros(len(y), dtype=float) \n    oof_rounded = np.zeros(len(y), dtype=int) \n    test_preds = np.zeros((len(test_data), n_splits))\n\n    for fold, (train_idx, test_idx) in enumerate(tqdm(SKF.split(X, y), desc=\"Training Folds\", total=n_splits)):\n        X_train, X_val = X.iloc[train_idx], X.iloc[test_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[test_idx]\n\n        model = clone(model_class)\n        model.fit(X_train, y_train)\n\n        y_train_pred = model.predict(X_train)\n        y_val_pred = model.predict(X_val)\n\n        oof_non_rounded[test_idx] = y_val_pred\n        y_val_pred_rounded = y_val_pred.round(0).astype(int)\n        oof_rounded[test_idx] = y_val_pred_rounded\n\n        train_kappa = quadratic_weighted_kappa(y_train, y_train_pred.round(0).astype(int))\n        val_kappa = quadratic_weighted_kappa(y_val, y_val_pred_rounded)\n\n        train_S.append(train_kappa)\n        test_S.append(val_kappa)\n        \n        test_preds[:, fold] = model.predict(test_data)\n        \n        print(f\"Fold {fold+1} - Train QWK: {train_kappa:.4f}, Validation QWK: {val_kappa:.4f}\")\n        clear_output(wait=True)\n\n    print(f\"Mean Train QWK --> {np.mean(train_S):.4f}\")\n    print(f\"Mean Validation QWK ---> {np.mean(test_S):.4f}\")\n\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead') # Nelder-Mead | # Powell\n    assert KappaOPtimizer.success, \"Optimization did not converge.\"\n    \n    oof_tuned = threshold_Rounder(oof_non_rounded, KappaOPtimizer.x)\n    tKappa = quadratic_weighted_kappa(y, oof_tuned)\n\n    print(f\"----> || Optimized QWK SCORE :: {Fore.CYAN}{Style.BRIGHT} {tKappa:.3f}{Style.RESET_ALL}\")\n\n    tpm = test_preds.mean(axis=1)\n    tpTuned = threshold_Rounder(tpm, KappaOPtimizer.x)\n    \n    submission = pd.DataFrame({\n        'id': sample['id'],\n        'sii': tpTuned\n    })\n\n    return submission,model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:40:00.178961Z","iopub.execute_input":"2024-12-09T13:40:00.179474Z","iopub.status.idle":"2024-12-09T13:40:00.199247Z","shell.execute_reply.started":"2024-12-09T13:40:00.179417Z","shell.execute_reply":"2024-12-09T13:40:00.197808Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Params7 = {'learning_rate': 0.03884249148676395, 'max_depth': 12, 'num_leaves': 413, 'min_data_in_leaf': 14,\n           'feature_fraction': 0.7987976913702801, 'bagging_fraction': 0.7602261703576205, 'bagging_freq': 2, \n           'lambda_l1': 4.735462555910575, 'lambda_l2': 4.735028557007343e-06} # CV : 0.4094 | LB : 0.471\n\nLight = lgb.LGBMRegressor(**Params7,random_state=SEED, verbose=-1,n_estimators=200)\nSubmission,model = TrainML(Light,test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:40:00.201103Z","iopub.execute_input":"2024-12-09T13:40:00.201623Z","iopub.status.idle":"2024-12-09T13:40:15.482612Z","shell.execute_reply.started":"2024-12-09T13:40:00.201568Z","shell.execute_reply":"2024-12-09T13:40:15.481481Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"Submission.to_csv('submission4.csv', index=False)\nsingle = Submission['sii'].values.copy()\n# Public Score : 0.471\nprint(single)\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:40:15.484140Z","iopub.execute_input":"2024-12-09T13:40:15.484505Z","iopub.status.idle":"2024-12-09T13:40:16.729878Z","shell.execute_reply.started":"2024-12-09T13:40:15.484469Z","shell.execute_reply":"2024-12-09T13:40:16.728011Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>\n\n## <div style=\"color:yellow;display:inline-block;border-radius:5px;background-color:darkcyan;font-block:Nexa;overflow:hidden\"><p style=\"padding:15px;color:white;overflow:hidden;font-size:85%;letter-spacing:0.5px;margin:0\"><b> </b>၊၊||၊ End to End</p></div>","metadata":{}},{"cell_type":"code","source":"score_list = [0.452, 0.440, 0.457, 0.464, 0.471]\npreds_list = [preds_lgbm, test_pred_rounded, voted_test, vote, single]\n\n# ..........................................................................\nresult_corr = pd.DataFrame({\n    'pred-0': preds_list[0],\n    'pred-1': preds_list[1],\n    'pred-2': preds_list[2],\n    'pred-3': preds_list[3],\n    'pred-4': preds_list[4],\n})\n\ndf_corr = result_corr.corr()\nsns.heatmap(df_corr, center=0, annot=True, cmap='Pastel2', linewidths=1)\nplt.title('Correlation between results')\nplt.show()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:40:16.731981Z","iopub.execute_input":"2024-12-09T13:40:16.732388Z","iopub.status.idle":"2024-12-09T13:40:17.183892Z","shell.execute_reply.started":"2024-12-09T13:40:16.732348Z","shell.execute_reply":"2024-12-09T13:40:17.182646Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"preds_list = [preds_lgbm, test_pred_rounded, voted_test, vote, single]\n\n# ..........................................................................\npr_blend = stats.mode(preds_list)[0] \ncount_pr_blend = stats.mode(preds_list)[1]\n\nsub_blend = sub_sample.copy()\nsub_blend['sii'] = pr_blend.copy()    \nsub_blend.to_csv('submission.csv', index=False)\n# Public Score: 0.483\n\n# ..........................................................................\nprint(pr_blend)\nprint(count_pr_blend)\n!ls","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-12-09T13:40:17.185292Z","iopub.execute_input":"2024-12-09T13:40:17.185690Z","iopub.status.idle":"2024-12-09T13:40:18.406647Z","shell.execute_reply.started":"2024-12-09T13:40:17.185629Z","shell.execute_reply":"2024-12-09T13:40:18.404966Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"border-bottom: 50px solid lightgreen\"></p>\n\n<p style=\"border-bottom: 15px solid darkcyan\"></p>","metadata":{}}]}