{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.14","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30761,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<p style=\"background-color:#33FFAA ; font-family:'Times New Romans'; color:#000000; font-size:200%; text-align:center; border: 3px solid #00EEEE; border-radius:10px; padding: 10px;\">Child Mind Institute | Single LightGBM Regressor</p>","metadata":{}},{"cell_type":"markdown","source":"### Predicting Severity Impairment Index (SII) using the MHN data.\n\nThe aim of this [competition](http://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data) is to predict the Severity Impairment Index (sii), a standard measure for the level of problematic internet use among children and adolescents, based on physical activity data and other features. \n\nThe sii values are derived from `PCIAT-PCIAT_Total`, the sum of scores from the Parent-Child Internet Addiction Test (PCIAT: 20 questions, scored 0-5), which makes sii an ordinal categorical variable with four levels, where the order of categories is meaningful. It is defined as:\n- 0: None (PCIAT-PCIAT_Total from 0 to 30)\n- 1: Mild (PCIAT-PCIAT_Total from 31 to 49)\n- 2: Moderate (PCIAT-PCIAT_Total from 50 to 79)\n- 3: Severe (PCIAT-PCIAT_Total 80 and more) \n\nThe test.csv file contains 20 test samples in the correct format to help for find the solutions. The full test set comprises about 3800 instances.\n\nDataset is divided into two sources:\n * **parquet** files: containing the accelerometer (actigraphy) series,and\n * **csv** files containing the remaining tabular data.\n\nThe majority of measures are missing for most participants. In particular, **the target sii is missing for a portion of the participants in the training set**. You may wish to apply non-supervised learning techniques to this data. The sii value is present for all instances in the test set.\n\nFor more info about the data, read the data page of the challage [here](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/data).","metadata":{}},{"cell_type":"markdown","source":"# Loading Libraries","metadata":{}},{"cell_type":"code","source":"import os\nimport numpy as np\nimport pandas as pd\n\nfrom colorama import Fore, Style\nfrom IPython.display import clear_output\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nfrom tqdm import tqdm\nfrom concurrent.futures import ThreadPoolExecutor\nfrom scipy.optimize import minimize\nimport optuna\n\n\nfrom sklearn.base import clone\nfrom sklearn.ensemble import VotingRegressor\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.metrics import cohen_kappa_score\n\nimport lightgbm as lgb\nfrom catboost import CatBoostRegressor, CatBoostClassifier\nfrom xgboost import XGBRegressor\n\nimport warnings\nwarnings.filterwarnings('ignore')\npd.options.display.max_columns = None\n\nSEED = 42\nn_splits = 5","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2024-11-22T20:52:59.806061Z","iopub.execute_input":"2024-11-22T20:52:59.806507Z","iopub.status.idle":"2024-11-22T20:53:02.766894Z","shell.execute_reply.started":"2024-11-22T20:52:59.806463Z","shell.execute_reply":"2024-11-22T20:53:02.765905Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#33FFAA ; font-family:'Times New Romans'; color:#000000; font-size:170%; text-align:center; border: 3px solid #00EEEE; border-radius:10px; padding: 10px;\">Reading Data Files</p>","metadata":{}},{"cell_type":"markdown","source":"# Reading Data files","metadata":{}},{"cell_type":"code","source":"%%time\nDrop_Cols = ['step']#, 'battery_voltage','time_of_day','weekday', 'quarter', 'X', 'Y', 'Z',  'non-wear_flag']\n\ndef process_file(filename, dirname):\n    df = pd.read_parquet(os.path.join(dirname, filename, 'part-0.parquet'))\n    df.drop(Drop_Cols, axis=1, inplace=True)\n    return df.describe().values.reshape(-1), filename.split('=')[1]\n    # return df.describe().drop('count', axis=0).values.reshape(-1), filename.split('=')[1]\n\ndef load_time_series(dirname) -> pd.DataFrame:\n\n    ids = os.listdir(dirname)\n    \n    with ThreadPoolExecutor() as executor:\n        results = list(tqdm(\n            executor.map(lambda fname: process_file(fname, dirname), ids),\n            total=len(ids))\n        )\n    \n    stats, indexes = zip(*results)\n    \n    df = pd.DataFrame(stats, columns=[f\"Stat_{i}\" for i in range(len(stats[0]))])\n    df['id'] = indexes\n    \n    return df\n    \n# Reading data files\ndata_path = '/kaggle/input/child-mind-institute-problematic-internet-use'\ntrain = pd.read_csv(f'{data_path}/train.csv')\ntest = pd.read_csv(f'{data_path}/test.csv')\nsample = pd.read_csv(f'{data_path}/sample_submission.csv')\n\ntrain_ts = load_time_series(f'{data_path}/series_train.parquet')\ntest_ts = load_time_series(f'{data_path}/series_test.parquet')\n\ntrain_orig = pd.merge(train, train_ts, how=\"left\", on='id')\ntest_orig = pd.merge(test, test_ts, how=\"left\", on='id')\n","metadata":{"execution":{"iopub.status.busy":"2024-11-22T20:53:10.748837Z","iopub.execute_input":"2024-11-22T20:53:10.749767Z","iopub.status.idle":"2024-11-22T20:54:32.267874Z","shell.execute_reply.started":"2024-11-22T20:53:10.749704Z","shell.execute_reply":"2024-11-22T20:54:32.266833Z"},"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# save original data\ntrain = train_orig.copy()\ntest = test_orig.copy()","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T20:54:44.007721Z","iopub.execute_input":"2024-11-22T20:54:44.008667Z","iopub.status.idle":"2024-11-22T20:54:44.023868Z","shell.execute_reply.started":"2024-11-22T20:54:44.008619Z","shell.execute_reply":"2024-11-22T20:54:44.022509Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#33FFAA ; font-family:'Times New Romans'; color:#000000; font-size:170%; text-align:center; border: 3px solid #00EEEE; border-radius:10px; padding: 10px;\">Basic Preprocess</p>","metadata":{}},{"cell_type":"markdown","source":"# Basic preprocessing","metadata":{}},{"cell_type":"code","source":"# Drop PCIAT Cols from the training data (they don't exist in the test data)\n\npciat_Cols = [col for col in train.columns if 'PCIAT' in col]\ntrain = train.drop(pciat_Cols, axis=1)\n\ntrain.shape, test.shape","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T20:54:48.293891Z","iopub.execute_input":"2024-11-22T20:54:48.294292Z","iopub.status.idle":"2024-11-22T20:54:48.306464Z","shell.execute_reply.started":"2024-11-22T20:54:48.294255Z","shell.execute_reply":"2024-11-22T20:54:48.305382Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# Prepare feature values\n\ntrain = train.dropna(subset='sii')\n\ncat_Cols = [col for col in train.columns if 'Season' in col]\n\ndef update(df):\n    for c in cat_Cols:\n        if df[c].dtype.name == 'category':\n            # Add 'Missing' to the categories if it's not already present\n            if 'Missing' not in df[c].cat.categories:\n                df[c] = df[c].cat.add_categories('Missing')\n\n        # Fill missing values with 'Missing'\n        df[c] = df[c].fillna('Missing')\n\n        # Ensure the column is of 'category' dtype\n        df[c] = df[c].astype('category')\n    return df\n\n\ntrain = update(train)\ntest = update(test)\n\n\"\"\"\n    This Mapping Works Fine For me, I also \n    check each values in train and test using \n    logic. There no Data Lekage.\n\"\"\"\n\ndef create_mapping(column, dataset):\n    unique_values = dataset[column].unique()\n    return {value: idx for idx, value in enumerate(unique_values)}\n    \nfor col in cat_Cols:\n    mapping_train = create_mapping(col, train)\n    mapping_test = create_mapping(col, test)\n\n    train[col] = train[col].replace(mapping_train).astype(int)\n    test[col] = test[col].replace(mapping_test).astype(int)\n\nprint(f'Train Shape : {train.shape} || Test Shape : {test.shape}')","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T20:54:57.097913Z","iopub.execute_input":"2024-11-22T20:54:57.098824Z","iopub.status.idle":"2024-11-22T20:54:57.171451Z","shell.execute_reply.started":"2024-11-22T20:54:57.098771Z","shell.execute_reply":"2024-11-22T20:54:57.170292Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#33FFAA ; font-family:'Times New Romans'; color:#000000; font-size:170%; text-align:center; border: 3px solid #00EEEE; border-radius:10px; padding: 10px;\">Modeling | Single LightGBM Regressor</p>","metadata":{}},{"cell_type":"markdown","source":"# Modeling and Training","metadata":{}},{"cell_type":"code","source":"train = train.drop('id', axis=1)\ntest_id = test['id'].copy()\ntest = test.drop('id', axis=1)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T20:55:03.564559Z","iopub.execute_input":"2024-11-22T20:55:03.565315Z","iopub.status.idle":"2024-11-22T20:55:03.573558Z","shell.execute_reply.started":"2024-11-22T20:55:03.565254Z","shell.execute_reply":"2024-11-22T20:55:03.572489Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n# Functions for training the evaluating the selected model \n\ndef quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\ndef threshold_Rounder(oof_non_rounded, thresholds):\n    return np.where(oof_non_rounded < thresholds[0], 0,\n                    np.where(oof_non_rounded < thresholds[1], 1,\n                             np.where(oof_non_rounded < thresholds[2], 2, 3)))\n\ndef evaluate_predictions(thresholds, y_true, oof_non_rounded):\n    rounded_p = threshold_Rounder(oof_non_rounded, thresholds)\n    return -quadratic_weighted_kappa(y_true, rounded_p)\n\ndef TrainML(model_class, test_data):\n    \n    X = train.drop(['sii'], axis=1)\n    y = train['sii']\n\n    SKF = StratifiedKFold(n_splits=n_splits, shuffle=True, random_state=SEED)\n    \n    train_S = []\n    test_S = []\n    \n    oof_non_rounded = np.zeros(len(y), dtype=float) \n    oof_rounded = np.zeros(len(y), dtype=int) \n    test_preds = np.zeros((len(test_data), n_splits))\n\n    for fold, (train_idx, test_idx) in enumerate(tqdm(SKF.split(X, y), desc=\"Training Folds\", total=n_splits)):\n        X_train, X_val = X.iloc[train_idx], X.iloc[test_idx]\n        y_train, y_val = y.iloc[train_idx], y.iloc[test_idx]\n\n        model = clone(model_class)\n        model.fit(X_train, y_train)\n\n        y_train_pred = model.predict(X_train)\n        y_val_pred = model.predict(X_val)\n\n        oof_non_rounded[test_idx] = y_val_pred\n        y_val_pred_rounded = y_val_pred.round(0).astype(int)\n        oof_rounded[test_idx] = y_val_pred_rounded\n\n        train_kappa = quadratic_weighted_kappa(y_train, y_train_pred.round(0).astype(int))\n        val_kappa = quadratic_weighted_kappa(y_val, y_val_pred_rounded)\n\n        train_S.append(train_kappa)\n        test_S.append(val_kappa)\n        \n        test_preds[:, fold] = model.predict(test_data)\n        \n        print(f\"Fold {fold+1} - Train QWK: {train_kappa:.4f}, Validation QWK: {val_kappa:.4f}\")\n        clear_output(wait=True)\n\n    print(f\"Mean Train QWK --> {np.mean(train_S):.4f}\")\n    print(f\"Mean Validation QWK ---> {np.mean(test_S):.4f}\")\n\n    KappaOPtimizer = minimize(evaluate_predictions,\n                              x0=[0.5, 1.5, 2.5], args=(y, oof_non_rounded), \n                              method='Nelder-Mead') # Nelder-Mead | # Powell\n    assert KappaOPtimizer.success, \"Optimization did not converge.\"\n    \n    oof_tuned = threshold_Rounder(oof_non_rounded, KappaOPtimizer.x)\n    tKappa = quadratic_weighted_kappa(y, oof_tuned)\n\n    print(f\"----> || Optimized QWK SCORE :: {Fore.CYAN}{Style.BRIGHT} {tKappa:.3f}{Style.RESET_ALL}\")\n\n    tpm = test_preds.mean(axis=1)\n    tpTuned = threshold_Rounder(tpm, KappaOPtimizer.x)\n    \n    submission = pd.DataFrame({\n        'id': test_id,     #sample['id'],\n        'sii': tpTuned\n    })\n\n    return submission, model","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T20:55:45.984385Z","iopub.execute_input":"2024-11-22T20:55:45.984874Z","iopub.status.idle":"2024-11-22T20:55:45.998896Z","shell.execute_reply.started":"2024-11-22T20:55:45.984834Z","shell.execute_reply":"2024-11-22T20:55:45.997662Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"# test_id == sample['id']","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n\n#Train and predict sii for test data \n\nLGB_Params = {\n    'learning_rate': 0.07, \n    'random_state': SEED, \n    'n_estimators': 200,\n    'max_depth': 8, \n    'num_leaves': 300, \n    'min_data_in_leaf': 17,\n    'feature_fraction': 0.7689, \n    'bagging_fraction': 0.6879, \n    'bagging_freq': 2, \n    'lambda_l1': 4.74, \n    'lambda_l2': 4.743e-06,\n    'verbose': -1,\n    # CV : 0.4094 | LB : 0.471\n}\n\nModel = lgb.LGBMRegressor(**LGB_Params)\n\nSubmission, model = TrainML(Model,test)","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T21:17:31.523161Z","iopub.execute_input":"2024-11-22T21:17:31.524391Z","iopub.status.idle":"2024-11-22T21:17:43.192512Z","shell.execute_reply.started":"2024-11-22T21:17:31.524316Z","shell.execute_reply":"2024-11-22T21:17:43.191376Z"}},"outputs":[],"execution_count":null},{"cell_type":"code","source":"%%time\n\n# feature_importance_df = pd.DataFrame({\n#     'Feature': model.booster_.feature_name(),\n#     'Importance': model.booster_.feature_importance(importance_type='gain')\n# })\n\n# feature_importance_df = feature_importance_df.sort_values(by='Importance', ascending=False)\n\n# plt.figure(figsize=(20, 40))\n# sns.barplot(x='Importance', y='Feature', data=feature_importance_df.head(100)) \n# plt.title(\"Top Feature Importance\")\n# plt.show()","metadata":{"trusted":true},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#33FFAA ; font-family:'Times New Romans'; color:#000000; font-size:170%; text-align:center; border: 3px solid #00EEEE; border-radius:10px; padding: 10px;\">Submission</p>","metadata":{}},{"cell_type":"markdown","source":"# Submit resluts","metadata":{}},{"cell_type":"code","source":"%%time\n\nSubmission.to_csv('submission.csv', index=False)\nprint(Submission['sii'].value_counts())","metadata":{"trusted":true,"execution":{"iopub.status.busy":"2024-11-22T21:17:06.989178Z","iopub.execute_input":"2024-11-22T21:17:06.989628Z","iopub.status.idle":"2024-11-22T21:17:06.999780Z","shell.execute_reply.started":"2024-11-22T21:17:06.989589Z","shell.execute_reply":"2024-11-22T21:17:06.998721Z"}},"outputs":[],"execution_count":null},{"cell_type":"markdown","source":"<p style=\"background-color:#33FFAA ; font-family:'Times New Romans'; color:#000000; font-size:170%; text-align:center; border: 3px solid #00EEEE; border-radius:10px; padding: 10px;\">Your Turn</p>","metadata":{}},{"cell_type":"markdown","source":"# Your Turn","metadata":{}},{"cell_type":"markdown","source":"# What you can do next:\n* experiment with the parameter tuning for the lbgm model\n* try other models, e.g. XGBoostRegressor, CatBoostRegressor (already imported)\n* **Most Important** Feature Engineering: Engineer features to get rid of the bad and create ew features that improve prediction.\n\nKindly, upvote if you find this helpful, and don't forget to upvote the base [notebook](https://www.kaggle.com/code/abdmental01/cmi-best-single-model).","metadata":{}},{"cell_type":"markdown","source":"# Credit:\nBase Notebook: https://www.kaggle.com/code/abdmental01/cmi-best-single-model","metadata":{}}]}