{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Using StratifiedGroupKFold\n\nFirst of all, I want to thanks all notebooks shared in this competition so far. I'm learning a lot and hopefully making progress in my studies. I will share the links I've used as reference to make my notebook below. Please, upvote if you like and feel free to share any ideas/suggestions to increase any performance.\n\n## Sources:\n\n- https://www.kaggle.com/code/desalegngeb/tps08-logisticregression-and-some-fe by [@des.](https://www.kaggle.com/desalegngeb)\n- https://www.kaggle.com/code/maxsarmento/lb-0-58978-standing-on-the-shoulder-of-giants by [@maxsarmento](https://www.kaggle.com/maxsarmento)\n- https://www.kaggle.com/code/themikejones/tps-aug-22-votingclassifier by [@themikejones](https://www.kaggle.com/themikejones)\n- https://www.kaggle.com/code/cdeotte/forward-selection-oof-ensemble-0-942-private by [@cdeote](https://www.kaggle.com/cdeotte)\n- https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense by [@ambrosm](https://www.kaggle.com/ambrosm)\n\nSome notes:\n\n1. I really don't know why no one is using this method to cross-validate results, since to me it's a perfect fit to this problem (groups of product_code) . Maybe I'm missing something?\n2. This is my first time using Chris Deotte's Ensemble method, so feel free to point out any issue or mistake to create the lists and folds (I had to do several [shenanigans]() to make it work).\n    \n    \n","metadata":{}},{"cell_type":"markdown","source":"# 0.0 Imports","metadata":{}},{"cell_type":"code","source":"!pip install feature-engine","metadata":{"execution":{"iopub.status.busy":"2022-08-12T05:18:30.597691Z","iopub.execute_input":"2022-08-12T05:18:30.598783Z","iopub.status.idle":"2022-08-12T05:18:45.170210Z","shell.execute_reply.started":"2022-08-12T05:18:30.598715Z","shell.execute_reply":"2022-08-12T05:18:45.169141Z"},"_kg_hide-output":true,"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import pandas as pd\nimport numpy as np\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport pandas_profiling as pdp\n\nfrom datetime import datetime\n\nfrom IPython.core.display  import HTML\nfrom IPython.display       import Image\n\nfrom itertools import combinations\n\nfrom sklearn.preprocessing   import RobustScaler, MinMaxScaler, OneHotEncoder, LabelEncoder, PolynomialFeatures, StandardScaler, PowerTransformer\nfrom sklearn.experimental    import enable_iterative_imputer\nfrom sklearn.impute          import IterativeImputer, KNNImputer, SimpleImputer\nfrom sklearn.ensemble        import RandomForestClassifier\nfrom sklearn.ensemble        import ExtraTreesClassifier\nfrom sklearn.neighbors       import KNeighborsClassifier\nfrom sklearn.ensemble        import RandomForestClassifier\nfrom sklearn.linear_model    import LogisticRegression\nfrom sklearn.linear_model    import LogisticRegressionCV\nfrom sklearn.metrics         import confusion_matrix\n\nfrom sklearn.preprocessing   import normalize\n\nfrom xgboost                 import XGBClassifier, DMatrix\nfrom lightgbm                import LGBMClassifier\nfrom catboost                import CatBoostClassifier\nfrom sklearn.svm             import SVC\n\n\nfrom sklearn.model_selection import train_test_split, cross_val_score, GridSearchCV, RandomizedSearchCV\nfrom sklearn.model_selection import KFold, GroupKFold, StratifiedKFold, StratifiedGroupKFold\nfrom sklearn.model_selection import cross_val_score\n\n\n\nfrom sklearn.metrics         import accuracy_score\nfrom sklearn.metrics         import f1_score\nfrom sklearn.metrics         import roc_auc_score, roc_curve\n\nfrom scikitplot              import metrics         as mt\n\nfrom sklearn.pipeline        import Pipeline\nfrom sklearn.compose         import ColumnTransformer\n\nfrom colorama                import Fore, Back, Style\n\n\n\n\nimport warnings\nwarnings.filterwarnings('ignore')\n\n\n!git clone https://github.com/analokmaus/kuma_utils.git\n    \nfrom feature_engine.encoding import WoEEncoder\n\nfrom kuma_utils.preprocessing.imputer import LGBMImputer\n\n\n# Neural Network\n\nimport tensorflow as tf\ntf.config.threading.set_intra_op_parallelism_threads(6)\ntf.config.threading.set_inter_op_parallelism_threads(2)\n\nfrom tensorflow import keras\nfrom tensorflow.keras import layers, callbacks","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:08.256659Z","start_time":"2022-08-11T20:57:07.210075Z"},"execution":{"iopub.status.busy":"2022-08-12T05:19:39.989190Z","iopub.execute_input":"2022-08-12T05:19:39.989719Z","iopub.status.idle":"2022-08-12T05:19:40.026917Z","shell.execute_reply.started":"2022-08-12T05:19:39.989673Z","shell.execute_reply":"2022-08-12T05:19:40.025684Z"},"_kg_hide-output":true,"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 1.0 Data Description","metadata":{}},{"cell_type":"code","source":"train = pd.read_csv('../input/tabular-playground-series-aug-2022/train.csv', index_col='id')\ntest = pd.read_csv('../input/tabular-playground-series-aug-2022/test.csv', index_col='id')\n\ndf_sub = pd.DataFrame({'id': test.index,\n                       'target': np.zeros(test.shape[0])})\n\nprint('Train set shape:', train.shape)\nprint('Test set shape:', test.shape)\ntrain.head(2)","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:09.290519Z","start_time":"2022-08-11T20:57:09.222864Z"},"execution":{"iopub.status.busy":"2022-08-12T05:20:15.861605Z","iopub.execute_input":"2022-08-12T05:20:15.862603Z","iopub.status.idle":"2022-08-12T05:20:16.189878Z","shell.execute_reply.started":"2022-08-12T05:20:15.862560Z","shell.execute_reply":"2022-08-12T05:20:16.188610Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_cols = [col for col in train.columns if train[col].dtype == 'object' and train[col].nunique() < 20]\nnumerical_cols = [col for col in train.columns if train[col].dtype in ['int64', 'float64']]","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:09.989426Z","start_time":"2022-08-11T20:57:09.984261Z"},"execution":{"iopub.status.busy":"2022-08-12T05:20:30.968943Z","iopub.execute_input":"2022-08-12T05:20:30.969396Z","iopub.status.idle":"2022-08-12T05:20:30.985948Z","shell.execute_reply.started":"2022-08-12T05:20:30.969358Z","shell.execute_reply":"2022-08-12T05:20:30.984427Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_cols","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:10.335293Z","start_time":"2022-08-11T20:57:10.331909Z"},"execution":{"iopub.status.busy":"2022-08-12T05:20:32.039941Z","iopub.execute_input":"2022-08-12T05:20:32.040914Z","iopub.status.idle":"2022-08-12T05:20:32.050935Z","shell.execute_reply.started":"2022-08-12T05:20:32.040851Z","shell.execute_reply":"2022-08-12T05:20:32.049257Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_cols","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:10.592858Z","start_time":"2022-08-11T20:57:10.589413Z"},"execution":{"iopub.status.busy":"2022-08-12T05:20:32.865429Z","iopub.execute_input":"2022-08-12T05:20:32.865914Z","iopub.status.idle":"2022-08-12T05:20:32.873906Z","shell.execute_reply.started":"2022-08-12T05:20:32.865878Z","shell.execute_reply":"2022-08-12T05:20:32.872824Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2.0 Preprocessing","metadata":{}},{"cell_type":"markdown","source":"## 2.1 Missing Values as features","metadata":{}},{"cell_type":"code","source":"data = pd.concat([train.drop(['failure'], axis=1), test], axis=0)\nnumerical_cols.remove('failure')","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:13.800922Z","start_time":"2022-08-11T20:57:13.792016Z"},"execution":{"iopub.status.busy":"2022-08-12T05:21:10.111463Z","iopub.execute_input":"2022-08-12T05:21:10.111924Z","iopub.status.idle":"2022-08-12T05:21:10.140919Z","shell.execute_reply.started":"2022-08-12T05:21:10.111887Z","shell.execute_reply":"2022-08-12T05:21:10.139585Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# From https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/342319\n\ndata['measurement_3_na'] = data['measurement_3'].isna().astype(int)\ndata['measurement_4_na'] = data['measurement_4'].isna().astype(int)\ndata['measurement_5_na'] = data['measurement_5'].isna().astype(int)","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:14.139636Z","start_time":"2022-08-11T20:57:14.135295Z"},"execution":{"iopub.status.busy":"2022-08-12T05:21:13.318784Z","iopub.execute_input":"2022-08-12T05:21:13.320042Z","iopub.status.idle":"2022-08-12T05:21:13.331050Z","shell.execute_reply.started":"2022-08-12T05:21:13.319999Z","shell.execute_reply":"2022-08-12T05:21:13.330042Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"data.head()","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:14.389817Z","start_time":"2022-08-11T20:57:14.377026Z"},"execution":{"iopub.status.busy":"2022-08-12T05:21:18.474577Z","iopub.execute_input":"2022-08-12T05:21:18.475006Z","iopub.status.idle":"2022-08-12T05:21:18.504729Z","shell.execute_reply.started":"2022-08-12T05:21:18.474975Z","shell.execute_reply":"2022-08-12T05:21:18.503024Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.2 Imputing Missing Values","metadata":{}},{"cell_type":"code","source":"%%time\n\nimputer = LGBMImputer(n_iter=50)\n\n# Fill missing values on a by product code basis\nfor code in data['product_code'].unique():\n#     imputer = KNNImputer(n_neighbors = 3)\n#     imputer = IterativeImputer(max_iter = 8, random_state = 0, skip_complete = True, n_nearest_features = 12)\n    data.loc[data['product_code'] == code, numerical_cols] = imputer.fit_transform(data.loc[data['product_code']==code,numerical_cols])","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:57:59.545112Z","start_time":"2022-08-11T20:57:49.518932Z"},"execution":{"iopub.status.busy":"2022-08-12T05:21:54.454400Z","iopub.execute_input":"2022-08-12T05:21:54.455154Z","iopub.status.idle":"2022-08-12T05:22:19.704041Z","shell.execute_reply.started":"2022-08-12T05:21:54.455097Z","shell.execute_reply":"2022-08-12T05:22:19.702693Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/code/themikejones/tps-aug-22-votingclassifier/notebook?scriptVersionId=102761965\n\nmeas_gr1_cols = [f\"measurement_{i:d}\" for i in list(range(3, 5)) + list(range(9, 17))]\ndata['meas_gr1_avg'] = np.mean(data[meas_gr1_cols], axis=1)\ndata['meas_gr1_std'] = np.std(data[meas_gr1_cols], axis=1)\n\nmeas_gr2_cols = [f\"measurement_{i:d}\" for i in list(range(5, 9))]\ndata['meas_gr2_avg'] = np.mean(data[meas_gr2_cols], axis=1)\ndata['meas_gr2_std'] = np.std(data[meas_gr2_cols], axis=1)\n\ndata['meas17/meas_gr1_avg'] = data['measurement_17'] / data['meas_gr1_avg']\ndata['meas17/meas_gr2_avg'] = data['measurement_17'] / data['meas_gr2_avg']","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:03.328575Z","start_time":"2022-08-11T20:58:03.303405Z"},"execution":{"iopub.status.busy":"2022-08-12T05:22:27.627595Z","iopub.execute_input":"2022-08-12T05:22:27.628161Z","iopub.status.idle":"2022-08-12T05:22:27.714183Z","shell.execute_reply.started":"2022-08-12T05:22:27.628121Z","shell.execute_reply":"2022-08-12T05:22:27.712985Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"numerical_cols = numerical_cols + ['meas_gr1_avg', 'meas_gr1_std', 'meas_gr2_avg', 'meas_gr2_std', 'meas17/meas_gr1_avg', 'meas17/meas_gr2_avg']","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:03.572551Z","start_time":"2022-08-11T20:58:03.570212Z"},"execution":{"iopub.status.busy":"2022-08-12T05:22:28.776016Z","iopub.execute_input":"2022-08-12T05:22:28.776823Z","iopub.status.idle":"2022-08-12T05:22:28.781279Z","shell.execute_reply.started":"2022-08-12T05:22:28.776780Z","shell.execute_reply":"2022-08-12T05:22:28.780386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.3 Categorical Encoding","metadata":{}},{"cell_type":"code","source":"categorical_cols.remove('product_code')","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:06.598154Z","start_time":"2022-08-11T20:58:06.595836Z"},"execution":{"iopub.status.busy":"2022-08-12T05:22:36.282131Z","iopub.execute_input":"2022-08-12T05:22:36.282597Z","iopub.status.idle":"2022-08-12T05:22:36.288584Z","shell.execute_reply.started":"2022-08-12T05:22:36.282561Z","shell.execute_reply":"2022-08-12T05:22:36.287163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"categorical_cols","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:06.906028Z","start_time":"2022-08-11T20:58:06.903073Z"},"execution":{"iopub.status.busy":"2022-08-12T05:22:41.260645Z","iopub.execute_input":"2022-08-12T05:22:41.261142Z","iopub.status.idle":"2022-08-12T05:22:41.270109Z","shell.execute_reply.started":"2022-08-12T05:22:41.261100Z","shell.execute_reply":"2022-08-12T05:22:41.268615Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"le = LabelEncoder()\n# data['attribute_0'] = le.fit_transform(data['attribute_0']) -- We will apply WoEEncoder later\ndata['attribute_1'] = le.fit_transform(data['attribute_1'])","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:09.987378Z","start_time":"2022-08-11T20:58:09.980104Z"},"execution":{"iopub.status.busy":"2022-08-12T05:23:02.726358Z","iopub.execute_input":"2022-08-12T05:23:02.726942Z","iopub.status.idle":"2022-08-12T05:23:02.753477Z","shell.execute_reply.started":"2022-08-12T05:23:02.726893Z","shell.execute_reply":"2022-08-12T05:23:02.752230Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 2.4 Scaling\n\nI've noticed some progress when scaling 'measurement_17'.","metadata":{"ExecuteTime":{"end_time":"2022-08-08T17:02:37.869921Z","start_time":"2022-08-08T17:02:37.866848Z"}}},{"cell_type":"code","source":"# # Scaling\n# scale_feats = [col for col in train_split.columns if train_split[col].dtypes == 'float64']\n# scaler = StandardScaler()\n# train_scaled = train_split.copy()\n# test_scaled = test_split.copy()\n# train_scaled[scale_feats] = scaler.fit_transform(train_split[scale_feats])\n# test_scaled[scale_feats] = scaler.transform(test_split[scale_feats])","metadata":{"ExecuteTime":{"end_time":"2022-08-11T21:23:25.972976Z","start_time":"2022-08-11T21:23:25.971052Z"},"execution":{"iopub.status.busy":"2022-08-12T05:27:05.878533Z","iopub.execute_input":"2022-08-12T05:27:05.879947Z","iopub.status.idle":"2022-08-12T05:27:05.885061Z","shell.execute_reply.started":"2022-08-12T05:27:05.879884Z","shell.execute_reply":"2022-08-12T05:27:05.883911Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3.0 Feature Engineering","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/code/desalegngeb/tps08-logisticregression-and-some-fe/notebook?scriptVersionId=102685100\ndata['attribute_2*3'] = data['attribute_2'] * data['attribute_3']\ndata['meas17_loading_ratio'] = data['measurement_17'] / data['loading']","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:11.168711Z","start_time":"2022-08-11T20:58:11.164761Z"},"execution":{"iopub.status.busy":"2022-08-12T05:24:27.102978Z","iopub.execute_input":"2022-08-12T05:24:27.104200Z","iopub.status.idle":"2022-08-12T05:24:27.114586Z","shell.execute_reply.started":"2022-08-12T05:24:27.104145Z","shell.execute_reply":"2022-08-12T05:24:27.112871Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense#ROC-curve\n\n# The EDA diagram of measurement 2 shows that the feature is correlated\n# to the target only for values above 10. For this reason, we clip\n# all values below 11.\n\ndata['measurement_2'] = data['measurement_2'].clip(11, None)","metadata":{"ExecuteTime":{"end_time":"2022-08-11T20:58:15.984848Z","start_time":"2022-08-11T20:58:15.981105Z"},"execution":{"iopub.status.busy":"2022-08-12T05:24:36.415406Z","iopub.execute_input":"2022-08-12T05:24:36.415893Z","iopub.status.idle":"2022-08-12T05:24:36.424955Z","shell.execute_reply.started":"2022-08-12T05:24:36.415858Z","shell.execute_reply":"2022-08-12T05:24:36.423548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Re-split\ntrain_split = data.drop('product_code', axis=1).iloc[:train.shape[0],:].copy()\ntest_split = data.drop('product_code', axis=1).iloc[train.shape[0]:,:].copy()","metadata":{"ExecuteTime":{"end_time":"2022-08-11T21:00:43.745659Z","start_time":"2022-08-11T21:00:43.731850Z"},"execution":{"iopub.status.busy":"2022-08-12T05:27:13.824963Z","iopub.execute_input":"2022-08-12T05:27:13.825371Z","iopub.status.idle":"2022-08-12T05:27:13.870739Z","shell.execute_reply.started":"2022-08-12T05:27:13.825338Z","shell.execute_reply":"2022-08-12T05:27:13.869436Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# from https://www.kaggle.com/code/maxsarmento/lb-0-58978-standing-on-the-shoulder-of-giants?scriptVersionId=102785631\n\nwoe_encoder = WoEEncoder(variables=['attribute_0'])\n\nX = train_split.copy()\ny = train['failure']\n\nwoe_encoder.fit(X, y)\n\nX_final = woe_encoder.transform(X)\ntest_final = woe_encoder.transform(test_split)","metadata":{"ExecuteTime":{"end_time":"2022-08-11T21:02:57.270945Z","start_time":"2022-08-11T21:02:57.251795Z"},"execution":{"iopub.status.busy":"2022-08-12T05:27:23.418073Z","iopub.execute_input":"2022-08-12T05:27:23.418454Z","iopub.status.idle":"2022-08-12T05:27:23.474922Z","shell.execute_reply.started":"2022-08-12T05:27:23.418424Z","shell.execute_reply":"2022-08-12T05:27:23.473160Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 3.1 Feature Selection","metadata":{}},{"cell_type":"code","source":"# in progress","metadata":{"ExecuteTime":{"end_time":"2022-08-10T14:35:10.496242Z","start_time":"2022-08-10T14:35:10.494163Z"}},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_final.shape, test_final.shape","metadata":{"ExecuteTime":{"end_time":"2022-08-11T21:24:13.589904Z","start_time":"2022-08-11T21:24:13.586854Z"},"execution":{"iopub.status.busy":"2022-08-12T05:28:22.190985Z","iopub.execute_input":"2022-08-12T05:28:22.191510Z","iopub.status.idle":"2022-08-12T05:28:22.200885Z","shell.execute_reply.started":"2022-08-12T05:28:22.191473Z","shell.execute_reply":"2022-08-12T05:28:22.199509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"X_final.columns","metadata":{"ExecuteTime":{"end_time":"2022-08-11T21:24:22.835613Z","start_time":"2022-08-11T21:24:22.832403Z"},"execution":{"iopub.status.busy":"2022-08-12T05:28:25.292918Z","iopub.execute_input":"2022-08-12T05:28:25.293383Z","iopub.status.idle":"2022-08-12T05:28:25.301046Z","shell.execute_reply.started":"2022-08-12T05:28:25.293347Z","shell.execute_reply":"2022-08-12T05:28:25.300098Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"scaler = StandardScaler()\nX_final[['measurement_17']] = scaler.fit_transform(X_final[['measurement_17']])","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:10:34.697713Z","start_time":"2022-08-12T02:10:34.692016Z"},"execution":{"iopub.status.busy":"2022-08-12T05:28:43.571066Z","iopub.execute_input":"2022-08-12T05:28:43.571523Z","iopub.status.idle":"2022-08-12T05:28:43.583364Z","shell.execute_reply.started":"2022-08-12T05:28:43.571486Z","shell.execute_reply":"2022-08-12T05:28:43.582179Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4.0 Modeling","metadata":{}},{"cell_type":"markdown","source":"**Why using StratifiedGroupKFOLD?**\n\n*by [Robus Mulla](https://www.kaggle.com/robikscube)*\n\n1. Imbalanced datasets -> stratified\n2. Ensure same group for Product Codes -> group\n3. *Shuffle = True*, to make sure that the sorting of our data doesn't play any role into how the splits are made","metadata":{}},{"cell_type":"code","source":"ESTIMATORS = 500\nFOLDS = 5 # must be 5 (product codes)\n\nfeatures_final = [\n                  'loading', \n                  'attribute_0', \n                  'attribute_1',\n#                   'attribute_2',\n                  'measurement_0', \n                  'measurement_1', \n                  'measurement_2', \n#                   'measurement_3',\n#                   'measurement_7',\n#                   'measurement_8',\n#                   'measurement_12',\n                  'measurement_17',\n                  'measurement_3_na', 'measurement_5_na', \n#                   'meas_gr1_avg', \n#                   'meas_gr1_std', \n#                   'meas_gr2_avg', \n#                   'meas_gr2_std',\n                  'attribute_2*3', \n#                   'meas17/meas_gr1_avg', \n                  'meas17/meas_gr2_avg'\n                 ]\n\nn_features = len(features_final)\n\nprint('Using ...')\nprint('')\nprint(f'{n_features:.>5} features.')\nprint('')","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:13:50.172275Z","start_time":"2022-08-12T02:13:50.167030Z"},"execution":{"iopub.status.busy":"2022-08-12T05:37:41.003692Z","iopub.execute_input":"2022-08-12T05:37:41.004124Z","iopub.status.idle":"2022-08-12T05:37:41.015139Z","shell.execute_reply.started":"2022-08-12T05:37:41.004091Z","shell.execute_reply":"2022-08-12T05:37:41.013162Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n\naucs = []\npreds_test = []\nimportances = []\noof_pred = []; oof_tar = []; oof_names = []; oof_folds = [] \nsgkf = StratifiedGroupKFold(n_splits=FOLDS, shuffle=True, random_state=42)\n\nfor fold, (tr_idx, val_idx) in enumerate(sgkf.split(X=train, y=train['failure'], groups=train['product_code'])):\n  \n    X_tr = X_final.iloc[tr_idx][features_final]\n    X_va = X_final.iloc[val_idx][features_final]\n    y_tr = train.loc[tr_idx, 'failure']\n    y_va = train.loc[val_idx, 'failure']\n  \n    model = LogisticRegression(max_iter = 200, C=0.01, penalty='l2', solver='newton-cg')   \n   \n    \n    # Register Feature Importances\n    model.fit(X_tr, y_tr)\n    importances.append(model.coef_.ravel())\n      \n    # Validate model\n    va_preds = model.predict_proba(X_va)[:,1]\n    score = roc_auc_score(y_va, va_preds)\n    print(f\"Fold {fold}: auc = {score:.5f}\") # , C = {model.get_params['C']}, penalty = {model.get_params['penalty']}\n    aucs.append(score)\n    \n    # Test set predictions\n    preds_test.append(model.predict_proba(test_final[features_final])[:,1])\n    \n    \n    # Preprocessing for ensemble\n    oof_pred.append(va_preds)\n    oof_tar.append(y_va.to_numpy())\n    oof_folds.append(np.ones_like(y_va)*fold)\n    oof_names.append(val_idx)\n\n    \n    \nprint(f'\\nAverage auc = {sum(aucs) / len(aucs):.5f}')\nprint('')\npreds = sum(preds_test)/len(preds_test)","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:13:52.649341Z","start_time":"2022-08-12T02:13:51.981675Z"},"execution":{"iopub.status.busy":"2022-08-12T05:46:06.663141Z","iopub.execute_input":"2022-08-12T05:46:06.663648Z","iopub.status.idle":"2022-08-12T05:46:10.247685Z","shell.execute_reply.started":"2022-08-12T05:46:06.663595Z","shell.execute_reply":"2022-08-12T05:46:10.246059Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 4.1 Feature Importances","metadata":{"heading_collapsed":true}},{"cell_type":"code","source":"# Show overall score\nprint(f\"{Fore.GREEN}{Style.BRIGHT}Average auc = {sum(aucs) / len(aucs):.5f}{Style.RESET_ALL}\")\n\n# Show feature importances\nimportance_df = pd.DataFrame(np.array(importances).T, index=features_final)\nimportance_df['mean'] = importance_df.mean(axis=1).abs()\nimportance_df['feature'] = features_final\nimportance_df = importance_df.sort_values('mean', ascending=False).reset_index().head(10)\nplt.figure(figsize=(14, 4))\nplt.barh(importance_df.index, importance_df['mean'], color='lightgreen')\nplt.gca().invert_yaxis()\nplt.yticks(ticks=importance_df.index, labels=importance_df['feature'])\nplt.title('LogisticRegression feature importances')\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:13:12.924344Z","start_time":"2022-08-12T02:13:12.841314Z"},"hidden":true,"scrolled":true,"execution":{"iopub.status.busy":"2022-08-12T05:37:51.204736Z","iopub.execute_input":"2022-08-12T05:37:51.205189Z","iopub.status.idle":"2022-08-12T05:37:51.263501Z","shell.execute_reply.started":"2022-08-12T05:37:51.205150Z","shell.execute_reply":"2022-08-12T05:37:51.262079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(10, 10))\nfpr, tpr, _ = roc_curve(y_va, va_preds)\nplt.plot(fpr, tpr, color='#c00000', lw=3, label=f\"(auc (fold 4) = {roc_auc_score(y_va, va_preds):.5f})\") # curve\nplt.fill_between(fpr, tpr, color='#ffc0c0') # area under the curve\nplt.plot([0, 1], [0, 1], color=\"navy\", lw=1, linestyle=\"--\") # diagonal\nplt.gca().set_aspect('equal')\nplt.xlim([0.0, 1.0])\nplt.ylim([0.0, 1.0])\nplt.xlabel(\"False Positive Rate\")\nplt.ylabel(\"True Positive Rate\")\nplt.title(\"Receiver operating characteristic\")\nplt.legend(loc=\"lower right\")\nplt.show()","metadata":{"ExecuteTime":{"end_time":"2022-08-12T01:00:14.783658Z","start_time":"2022-08-12T01:00:14.705530Z"},"hidden":true,"execution":{"iopub.status.busy":"2022-08-12T05:37:53.112402Z","iopub.execute_input":"2022-08-12T05:37:53.114072Z","iopub.status.idle":"2022-08-12T05:37:53.152367Z","shell.execute_reply.started":"2022-08-12T05:37:53.114009Z","shell.execute_reply":"2022-08-12T05:37:53.150941Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 5.0 Post-Processing\nFor ensemble purposes","metadata":{"heading_collapsed":true}},{"cell_type":"code","source":"# Manual change! Imputing CSV info\n\nMODEL_INFO = 'skgf_lr_01'","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:13:55.291388Z","start_time":"2022-08-12T02:13:55.289260Z"},"execution":{"iopub.status.busy":"2022-08-12T05:36:11.986411Z","iopub.execute_input":"2022-08-12T05:36:11.986911Z","iopub.status.idle":"2022-08-12T05:36:11.993257Z","shell.execute_reply.started":"2022-08-12T05:36:11.986869Z","shell.execute_reply":"2022-08-12T05:36:11.991899Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# COMPUTE OVERALL OOF AUC\noof = np.concatenate(oof_pred); true = np.concatenate(oof_tar);\nnames = np.concatenate(oof_names); folds = np.concatenate(oof_folds)\nauc = roc_auc_score(true,oof)\nprint('Overall OOF AUC = %.5f'%auc)","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:13:55.660176Z","start_time":"2022-08-12T02:13:55.651986Z"},"execution":{"iopub.status.busy":"2022-08-12T05:36:14.427910Z","iopub.execute_input":"2022-08-12T05:36:14.428969Z","iopub.status.idle":"2022-08-12T05:36:14.447739Z","shell.execute_reply.started":"2022-08-12T05:36:14.428923Z","shell.execute_reply":"2022-08-12T05:36:14.446751Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SAVE OOF TO DISK\ndf_oof = pd.DataFrame(dict(\n    id_product = names, target=true, pred = oof, fold=folds))\ndf_oof.to_csv(f'./oof_{MODEL_INFO}.csv',index=False)\ndf_oof.head()","metadata":{"ExecuteTime":{"end_time":"2022-08-12T00:59:50.061658Z","start_time":"2022-08-12T00:59:50.018107Z"},"execution":{"iopub.status.busy":"2022-08-12T05:39:10.406914Z","iopub.execute_input":"2022-08-12T05:39:10.407406Z","iopub.status.idle":"2022-08-12T05:39:10.505363Z","shell.execute_reply.started":"2022-08-12T05:39:10.407369Z","shell.execute_reply":"2022-08-12T05:39:10.503718Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6.0 Neural Network\n\nfrom https://www.kaggle.com/code/sfktrkl/tps-aug-2022/ \n\nThanks, @sfktrlkl","metadata":{}},{"cell_type":"markdown","source":"## 6.1 NN Configuration","metadata":{}},{"cell_type":"code","source":"EPOCHS = 200\nBATCH_SIZE = 512\nACTIVATION = 'swish'\nN_SPLITS = 5\n\ndef load_model():\n    early_stopping = callbacks.EarlyStopping(\n        monitor=\"val_loss\",     # Quantity to be monitored\n        patience=20,                # How many epochs to wait before stopping\n        restore_best_weights=True)\n    \n    reduce_lr = callbacks.ReduceLROnPlateau(\n        monitor='val_loss', \n        factor=0.5,                # Factor by which the learning rate will be reduced\n        patience=5)                # Number of epochs with no improvement\n    \n    model = keras.Sequential([\n        layers.Dense(108, activation=ACTIVATION, input_shape=[X_final[features_final].shape[1]]),      \n        layers.Dense(64, activation=ACTIVATION), \n        layers.Dense(32, activation=ACTIVATION),\n        layers.Dense(1, activation='sigmoid')\n    ])\n\n    model.compile(\n        optimizer='adam',\n        loss='binary_crossentropy',\n        metrics=['AUC'])\n    \n    return model, [early_stopping, reduce_lr]\n","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:58:54.487641Z","start_time":"2022-08-12T02:58:54.483965Z"},"execution":{"iopub.status.busy":"2022-08-12T05:39:45.087310Z","iopub.execute_input":"2022-08-12T05:39:45.087818Z","iopub.status.idle":"2022-08-12T05:39:45.097639Z","shell.execute_reply.started":"2022-08-12T05:39:45.087773Z","shell.execute_reply":"2022-08-12T05:39:45.096300Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.2 Training","metadata":{}},{"cell_type":"code","source":"scores = []\nf_scores = []\ntest_predictions = []\nnn_oof_pred = []; nn_oof_tar = []; nn_oof_names = []; nn_oof_folds = [] \n\ncv = StratifiedKFold(n_splits=N_SPLITS, random_state=42, shuffle=True)\nfor fold, (train_idx, test_idx) in enumerate(cv.split(X=train, y=train['failure'])):\n    train_X, val_X = X_final.iloc[train_idx][features_final], X_final.iloc[test_idx][features_final]\n    train_y, val_y = train.iloc[train_idx]['failure'], train.iloc[test_idx]['failure']\n    \n    model, CALLBACKS = load_model()\n    history = model.fit(\n        train_X, train_y,\n        validation_data=(val_X, val_y),\n        batch_size=BATCH_SIZE,\n        epochs=EPOCHS,\n        callbacks=CALLBACKS,        # Put your callbacks in a list\n        verbose=0)                  # Turn off training log\n\n    predictions = model.predict(val_X)\n    score = roc_auc_score(val_y, predictions)\n    scores.append(score)\n    print(f\"Fold {fold + 1} \\t\\t AUC: {score}\")\n\n    test_predictions.append(model.predict(test_final[features_final]))\n\n    # Saving history to plot at the end\n    hist = pd.DataFrame(history.history)\n    hist['folds'] = fold + 1\n    f_scores = hist if fold == 0 else pd.concat([f_scores, hist], axis=0)\n    \n    \n    # Preprocessing for ensemble\n    nn_oof_pred.append(predictions)\n    nn_oof_tar.append(val_y.to_numpy())\n    nn_oof_folds.append(np.ones_like(val_y)*fold)\n    nn_oof_names.append(test_idx)\n\n    \n    \nprint('Overall AUC: ', np.mean(scores))","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:59:18.862580Z","start_time":"2022-08-12T02:58:56.514889Z"},"execution":{"iopub.status.busy":"2022-08-12T05:40:12.789377Z","iopub.execute_input":"2022-08-12T05:40:12.789951Z","iopub.status.idle":"2022-08-12T05:42:15.040378Z","shell.execute_reply.started":"2022-08-12T05:40:12.789902Z","shell.execute_reply":"2022-08-12T05:42:15.039201Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.3 Outputs","metadata":{"heading_collapsed":true}},{"cell_type":"code","source":"for fold in range(f_scores['folds'].nunique()):\n    fold = fold + 1\n    history_f = f_scores[f_scores['folds'] == fold]\n\n    fig, ax = plt.subplots(1, 2, tight_layout=True, figsize=(14,4))\n    fig.suptitle('Fold : ' + str(fold), fontsize=14)\n        \n    plt.subplot(1,2,1)\n    plt.plot(history_f.loc[:, ['loss', 'val_loss']], label= ['loss', 'val_loss'])\n    plt.legend(fontsize=15)\n    plt.grid()\n    \n    plt.subplot(1,2,2)\n    plt.plot(history_f.loc[:, ['auc', 'val_auc']],label= ['auc', 'val_auc'])\n    plt.legend(fontsize=15)\n    plt.grid()\n    \n    print(\"Validation Loss: {:0.4f}\".format(history_f['val_loss'].min()));","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:40:30.782777Z","start_time":"2022-08-12T02:40:29.958676Z"},"hidden":true,"execution":{"iopub.status.busy":"2022-08-12T05:42:21.090923Z","iopub.execute_input":"2022-08-12T05:42:21.091382Z","iopub.status.idle":"2022-08-12T05:42:21.749230Z","shell.execute_reply.started":"2022-08-12T05:42:21.091346Z","shell.execute_reply":"2022-08-12T05:42:21.748007Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 6.4 Postprocessing","metadata":{}},{"cell_type":"code","source":"# Manual change! Imputing CSV info\n\nNN_INFO = 'nn_oof_01'","metadata":{"ExecuteTime":{"end_time":"2022-08-12T03:00:44.459372Z","start_time":"2022-08-12T03:00:44.456893Z"},"execution":{"iopub.status.busy":"2022-08-12T05:42:27.756236Z","iopub.execute_input":"2022-08-12T05:42:27.757586Z","iopub.status.idle":"2022-08-12T05:42:27.763643Z","shell.execute_reply.started":"2022-08-12T05:42:27.757537Z","shell.execute_reply":"2022-08-12T05:42:27.762255Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# COMPUTE OVERALL OOF AUC\nnn_oof = np.concatenate(nn_oof_pred); nn_true = np.concatenate(nn_oof_tar);\nnn_names = np.concatenate(nn_oof_names); nn_folds = np.concatenate(nn_oof_folds)\nnn_auc = roc_auc_score(nn_true, nn_oof)\nprint('Overall OOF AUC = %.5f'%nn_auc)","metadata":{"ExecuteTime":{"end_time":"2022-08-12T03:01:20.226277Z","start_time":"2022-08-12T03:01:20.217540Z"},"execution":{"iopub.status.busy":"2022-08-12T05:42:30.881126Z","iopub.execute_input":"2022-08-12T05:42:30.882120Z","iopub.status.idle":"2022-08-12T05:42:30.901835Z","shell.execute_reply.started":"2022-08-12T05:42:30.882075Z","shell.execute_reply":"2022-08-12T05:42:30.900495Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"nn_oof.resize(26570)","metadata":{"ExecuteTime":{"end_time":"2022-08-12T04:10:26.366367Z","start_time":"2022-08-12T04:10:26.364172Z"},"execution":{"iopub.status.busy":"2022-08-12T05:42:42.678547Z","iopub.execute_input":"2022-08-12T05:42:42.679880Z","iopub.status.idle":"2022-08-12T05:42:42.686509Z","shell.execute_reply.started":"2022-08-12T05:42:42.679828Z","shell.execute_reply":"2022-08-12T05:42:42.685163Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# SAVE OOF TO DISK\ndf_oof_nn = pd.DataFrame(dict(\n    id_product = nn_names, target=nn_true, pred = nn_oof, fold=nn_folds))\ndf_oof_nn.to_csv(f'./oof_{NN_INFO}.csv',index=False)\ndf_oof_nn.head()","metadata":{"ExecuteTime":{"end_time":"2022-08-12T04:11:09.225143Z","start_time":"2022-08-12T04:11:09.182337Z"},"execution":{"iopub.status.busy":"2022-08-12T05:43:08.230817Z","iopub.execute_input":"2022-08-12T05:43:08.231254Z","iopub.status.idle":"2022-08-12T05:43:08.310273Z","shell.execute_reply.started":"2022-08-12T05:43:08.231219Z","shell.execute_reply":"2022-08-12T05:43:08.308594Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7.0 Submission","metadata":{}},{"cell_type":"markdown","source":"## 7.1 Logistic Classifier Submission","metadata":{}},{"cell_type":"code","source":"submission = pd.DataFrame({'id': test.index,\n                           'failure': preds})\nsubmission.to_csv(f'./sub_{MODEL_INFO}.csv', index=False)\nsubmission","metadata":{"ExecuteTime":{"end_time":"2022-08-12T01:00:20.728491Z","start_time":"2022-08-12T01:00:20.695526Z"},"execution":{"iopub.status.busy":"2022-08-12T05:46:20.400342Z","iopub.execute_input":"2022-08-12T05:46:20.400887Z","iopub.status.idle":"2022-08-12T05:46:20.463740Z","shell.execute_reply.started":"2022-08-12T05:46:20.400845Z","shell.execute_reply":"2022-08-12T05:46:20.462732Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## 7.2 Neural Network Submission","metadata":{}},{"cell_type":"code","source":"sub = pd.read_csv(\"../input/tabular-playground-series-aug-2022/sample_submission.csv\")\nsub['failure'] = np.mean(test_predictions, axis=0)\nsub.to_csv(f'sub_{NN_INFO}.csv', index=False)\nsub","metadata":{"ExecuteTime":{"end_time":"2022-08-12T02:43:42.956339Z","start_time":"2022-08-12T02:43:42.932039Z"},"execution":{"iopub.status.busy":"2022-08-12T05:44:55.720960Z","iopub.execute_input":"2022-08-12T05:44:55.721405Z","iopub.status.idle":"2022-08-12T05:44:55.792796Z","shell.execute_reply.started":"2022-08-12T05:44:55.721371Z","shell.execute_reply":"2022-08-12T05:44:55.791859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}