{"metadata":{"kernelspec":{"display_name":"Python 3 (ipykernel)","language":"python","name":"python3"},"language_info":{"codemirror_mode":{"name":"ipython","version":3},"file_extension":".py","mimetype":"text/x-python","name":"python","nbconvert_exporter":"python","pygments_lexer":"ipython3","version":"3.11.8"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":81933,"databundleVersionId":9643020,"sourceType":"competition"}],"dockerImageVersionId":30775,"isInternetEnabled":false,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Training on labeled data - boosted classifiers\n\nHere we train only the subset of ```train.csv``` that contains the ```sii``` labels on a few boosted classifier algorithms (XGBoost, CatBoost and LightGBM). For a simple analysis, we do not incorporate the actigraphy data. In this notebook, we perform the following tasks-\n1. **Data preprocessing:** The training data with ```sii``` is separated out and the numerical and categorical columns are processed separately. Since some of the numerical features have very large outliers, they are *winsorized*. During the processing, the numerical and categorical features are imputed using various strategies (simple, knn and iterative).\n2. **Hyperparameter tuning:** GridSearchCV and ShuffleSplit are used on the training data with various estimators to obtain the optimal hyperparameters.","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport re\nimport matplotlib.pyplot as plt\n%matplotlib inline\nimport seaborn as sns\n\nplt.rc('text', usetex=True)\nplt.rc('font', family='serif')\nplt.rc('xtick',labelsize=12)\nplt.rc('ytick',labelsize=12)\n\nsns.set_theme()\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\nfrom tqdm import tqdm \nfrom sklearn.impute import SimpleImputer, KNNImputer\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.preprocessing import OneHotEncoder, StandardScaler\nfrom sklearn.model_selection import ShuffleSplit, GridSearchCV\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier\nfrom sklearn.metrics import make_scorer, cohen_kappa_score","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"folder_path = '/kaggle/input/child-mind-institute-problematic-internet-use/'\n\ntrain_df = pd.read_csv(folder_path + 'train.csv', index_col='id')\ntest_df = pd.read_csv(folder_path + 'test.csv', index_col='id')\ndata_dict = pd.read_csv(folder_path + 'data_dictionary.csv')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test set has fewer attributes than training. Other than ```sii```, it is also missing the different ```PCIAT``` columns. This makes sense since the ```sii``` scores are possibly based on the ```PCIAT``` scores and the task of the learning model is to use other factors to predict the score.","metadata":{}},{"cell_type":"code","source":"train_cols = train_df.columns.tolist()\ntest_cols = test_df.columns.tolist()\n\nset(train_cols) - set(test_cols) # the difference between two lists can be taken if converted to sets","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Preprocessing","metadata":{}},{"cell_type":"markdown","source":"We first remove the different ```PCIAT``` columns from the training dataset. Then, we consider the subset of the data that has ```sii``` values associated with them. The numerical and categorical columns are then separated out. Here, the categorical columns are those features which are labeled as ```categorical int``` in ```data_dictionary.csv```.","metadata":{}},{"cell_type":"code","source":"drop_cols = list(set(train_cols) - set(test_cols)) # columns for dropping from training set\ndrop_cols.remove('sii')\n\n''' \nSelecting the subset that contains 'sii' labels and create X_train and y (labels)\n'''\ntrain_df_subset = train_df[train_df.loc[:, 'sii'].notna()]\ntrain_df_subset.drop(columns=drop_cols, inplace=True)\n\nX_train = train_df_subset.iloc[:, 0:-1]\ny = train_df_subset.loc[:, 'sii'].astype('int')\nX_test = test_df ","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We separate out the numerical and categorical columns using the ```Type``` field in the ```data_dictionary.csv``` file. However, we note that ```PCIAT-PCIAT_Total``` field must be manually removed. Furthermore, from the categorical columns, we remove all the different ```PCIAT``` fields. Furthermore, the 'season of participation' columns of each category has been removed.","metadata":{}},{"cell_type":"code","source":"numerical_cols = data_dict[\n    (data_dict['Type']=='float') | (data_dict['Type']=='int')\n]['Field'].tolist()\nnumerical_cols.remove('PCIAT-PCIAT_Total')\n\ncategorical_cols = data_dict[data_dict['Type']=='categorical int']['Field'].tolist() # this contains PCIAT; needs to be removed\n\ncategorical_cols = [item for item in categorical_cols if (item not in drop_cols)]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(numerical_cols)\nprint(f'There are {len(numerical_cols)} numerical features.')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(categorical_cols)\nprint(f'There are {len(categorical_cols)} numerical features.')","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Useful functions\n\nWe define some useful functions here. \n1. ```clipped(X, r)```: This function takes in the numerical columns of the training data as input and winsorizes only those columns for which the standard deviation is much greater than the mean, expressed as the ratio $r$.\n2. ```processing(X)```: This function takes the training dataset ```X```, splits it into numerical and categorical features and performs the necessary processing steps. These include winsorization, replacing occurrences of zeros with means and imputation.\n3. ```ParamSearchModels(X, y)```: This function takes in the processed data and performs ```ShuffleSplit()``` and ```GridSearchCV``` for the estimators and parameter sets specified.","metadata":{}},{"cell_type":"code","source":"def clipped(X, r):\n    ''' \n    This function caps the numerical columns to the specified upper and lower quantiles on respective ends. It does so for those featutres\n    for which the standard deviation is much greater or equal the mean. This is passed as a desired ratio of std/mean.\n    '''\n    X_clipped = X.copy()\n\n    ratio = X_clipped.std() / X_clipped.mean()\n    idx_list = ratio[ ratio.where(ratio >= r).notna() ].index.tolist()\n\n    for idx in idx_list:\n        X_clipped[idx].clip(\n            lower=X_clipped.quantile(0.05),\n            upper=X_clipped.quantile(0.95),\n            inplace=True\n        )\n\n    return X_clipped\n\ndef processing(X):\n    ''' \n    The training data is first split into numerical and categorical features.\n    '''\n    X_numerical = X[numerical_cols]\n    X_categorical = X[categorical_cols]\n\n    # Some numerical features have erroneous zero values - e.g. BMI, blood pressure etc\n    # Winsorize X_numerical and replace occurrences of zero with feature mean\n\n    #X_numerical = clipped(X_numerical, 1.5)\n    X_numerical.replace(0, X_numerical.mean(axis=0), inplace=True)\n\n    ''' \n    The 'sex' column is one-hot encoded while the other categorical features are kept as ordinal. The NaN\n    entries in the categorical features are imputed with mode.\n\n    The numerical columns are imputed using mean and then standardized.\n    '''\n    ct = ColumnTransformer(\n        transformers=[\n            ('ohe', OneHotEncoder(sparse_output=False), [categorical_cols[0]]),\n            ('imp_cat', KNNImputer(n_neighbors=2, weights='uniform'), categorical_cols[1:])\n        ]\n    )\n\n    ct.set_output(transform='pandas')\n    X_categorical = ct.fit_transform(X_categorical)\n    X_categorical = X_categorical.astype('int')\n\n    pipeline = Pipeline(\n        [\n            ('imp_num',  KNNImputer(n_neighbors=2, weights='uniform')),\n            ('ss', StandardScaler())\n        ]\n    )\n    pipeline.set_output(transform='pandas')\n    X_numerical = pipeline.fit_transform(X_numerical)\n\n    X_processed = X_categorical.join(X_numerical)\n\n    return X_processed\n\n\ndef quadratic_weighted_kappa(y_true, y_pred):\n    return cohen_kappa_score(y_true, y_pred, weights='quadratic')\n\ndef ParamSearchModels(X, y):\n\n    algorithms = {\n        #'LightGBM': {\n        #    'model'    : LGBMClassifier(),\n        #    'params'   : {\n        #        'n_estimators' : [50, 100, 150],\n        #        'max_depth'    : [5, 10, 15],\n        #        'learning_rate'    : [0.01, 0.1]\n        #    }\n        #},\n        \n        'CatBoost' : {\n            'model' : CatBoostClassifier(verbose=False),\n            'params' : {\n                'iterations' : [100, 200],\n                'learning_rate' : [0.01, 0.1]\n            }\n            \n        },\n        'XGBoost' : {\n            'model'    : XGBClassifier(),\n            'params'   : {\n                'n_estimators' : [50, 100, 150],\n                'max_depth'    : [5, 10, 15],\n                'subsample'    : [0.5, 0.75],\n                'gamma'        : [0.01, 0.05, 0.1],\n                'eta'          : [0.001, 0.01, 0.1],\n                'lambda'       : [1, 5, 10],\n                'alpha'        : [1, 5, 10]\n            }\n        }\n    }\n\n    scores = []\n    QWK_scorer = make_scorer(quadratic_weighted_kappa, greater_is_better=True)\n    cv = ShuffleSplit(n_splits = 5, test_size=0.2)\n    \n    for algorithm, config in tqdm(algorithms.items()):\n        GS = GridSearchCV( \n            config['model'], \n            config['params'], \n            cv=cv, \n            return_train_score=False,\n            scoring = QWK_scorer,\n            n_jobs = -1\n        )\n        GS.fit(X, y)\n        scores.append(\n            {\n                'model'       : algorithm,\n                'best_score'  : GS.best_score_,\n                'best_params' : GS.best_params_ \n            }\n        )\n\n    return pd.DataFrame(scores, columns=['model', 'best_score', 'best_params'])\n\ndef make_predictions(estimator, params, X_train, X_test):\n    clf = estimator(**params)\n    clf.fit(X_train, y)\n\n    y_pred = clf.predict(X_test).tolist()\n    test_id = X_test.index.tolist()\n\n    return pd.DataFrame({'id': test_id, 'sii': y_pred})","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_train_processed = processing(X_train)\n\n#GS_params_data = ParamSearchModels(X_train_processed, y)\n\n#GS_params_data","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#GS_params_data.iloc[1, 2]","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_test_processed = processing(X_test)\nparams = {\n    'alpha': 10,\n    'eta': 0.1,\n    'gamma': 0.05,\n    'lambda': 10,\n    'max_depth': 5,\n    'n_estimators': 150,\n    'subsample': 0.5\n}\n\npreds = make_predictions(\n    XGBClassifier,\n    params,\n    X_train_processed,\n    X_test_processed\n)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"preds.to_csv('submission.csv', index=False)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}