{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"<div id=\"1\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n   Introduction\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"This notebook is a continuation of my previous notebook, which explained a beginner-level approach to this competition, going over all the necessary steps. It can be accessed using this [link](https://www.kaggle.com/code/reymaster/tps-aug22-xgboost-imputation-feature-engineering).\n\nIn this notebook, we will be going over a technique for handling imbalanced data called Synthetic Minority Over-sampling Technique, or SMOTE for short.","metadata":{}},{"cell_type":"markdown","source":"<div id=\"2\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n   Code from previous notebook\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"I have condensed all my code from my previous notebook into one section. For more detailed explanations of each step, please read [this notebook](https://www.kaggle.com/code/reymaster/tps-aug22-xgboost-imputation-feature-engineering).","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:50.812373Z","iopub.execute_input":"2022-08-02T18:55:50.812783Z","iopub.status.idle":"2022-08-02T18:55:50.817850Z","shell.execute_reply.started":"2022-08-02T18:55:50.812752Z","shell.execute_reply":"2022-08-02T18:55:50.816710Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/tabular-playground-series-aug-2022/train.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:50.823330Z","iopub.execute_input":"2022-08-02T18:55:50.824074Z","iopub.status.idle":"2022-08-02T18:55:51.013440Z","shell.execute_reply.started":"2022-08-02T18:55:50.824028Z","shell.execute_reply":"2022-08-02T18:55:51.012218Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"There are three categorical columns:","metadata":{}},{"cell_type":"code","source":"train_df.select_dtypes(include='object')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.015751Z","iopub.execute_input":"2022-08-02T18:55:51.016232Z","iopub.status.idle":"2022-08-02T18:55:51.046778Z","shell.execute_reply.started":"2022-08-02T18:55:51.016185Z","shell.execute_reply":"2022-08-02T18:55:51.044984Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"None of the categorical columns have any missing values:","metadata":{}},{"cell_type":"code","source":"train_df[['product_code', 'attribute_0', 'attribute_1']].isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.048547Z","iopub.execute_input":"2022-08-02T18:55:51.049227Z","iopub.status.idle":"2022-08-02T18:55:51.065191Z","shell.execute_reply.started":"2022-08-02T18:55:51.049188Z","shell.execute_reply":"2022-08-02T18:55:51.063782Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['product_code'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.067862Z","iopub.execute_input":"2022-08-02T18:55:51.068434Z","iopub.status.idle":"2022-08-02T18:55:51.078858Z","shell.execute_reply.started":"2022-08-02T18:55:51.068398Z","shell.execute_reply":"2022-08-02T18:55:51.077845Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For `product_code`, convert each letter into a number:\n* A = 0\n* B = 1\n* C = 2\n* D = 3\n* E = 4","metadata":{}},{"cell_type":"code","source":"product_code_encoded = []\nfor char in train_df['product_code']:\n    product_code_encoded.append(ord(char) - 65)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.080157Z","iopub.execute_input":"2022-08-02T18:55:51.080533Z","iopub.status.idle":"2022-08-02T18:55:51.096704Z","shell.execute_reply.started":"2022-08-02T18:55:51.080502Z","shell.execute_reply":"2022-08-02T18:55:51.095503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df['product_code'] = product_code_encoded","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.098193Z","iopub.execute_input":"2022-08-02T18:55:51.098568Z","iopub.status.idle":"2022-08-02T18:55:51.115628Z","shell.execute_reply.started":"2022-08-02T18:55:51.098537Z","shell.execute_reply":"2022-08-02T18:55:51.114436Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Use the material number by splitting at the underscore (_) for `attribute_0` and `attribute_1`.","metadata":{}},{"cell_type":"code","source":"train_df['attribute_0'] = train_df[\"attribute_0\"].str.split(\"_\").str[1] # the second value is the material number\ntrain_df['attribute_1'] = train_df[\"attribute_1\"].str.split(\"_\").str[1] # the second value is the material number","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.116899Z","iopub.execute_input":"2022-08-02T18:55:51.117201Z","iopub.status.idle":"2022-08-02T18:55:51.202423Z","shell.execute_reply.started":"2022-08-02T18:55:51.117175Z","shell.execute_reply":"2022-08-02T18:55:51.201440Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Count missing values","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.204473Z","iopub.execute_input":"2022-08-02T18:55:51.204991Z","iopub.status.idle":"2022-08-02T18:55:51.217485Z","shell.execute_reply.started":"2022-08-02T18:55:51.204960Z","shell.execute_reply":"2022-08-02T18:55:51.216437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cols = list(train_df.columns)\ncols.remove('failure')\ncols.remove('id')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.218887Z","iopub.execute_input":"2022-08-02T18:55:51.219472Z","iopub.status.idle":"2022-08-02T18:55:51.224963Z","shell.execute_reply.started":"2022-08-02T18:55:51.219425Z","shell.execute_reply":"2022-08-02T18:55:51.223793Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We can fill in missing values using scikit-learn's `IterativeImputer`. See [official documentation](https://scikit-learn.org/stable/modules/generated/sklearn.impute.IterativeImputer.html) for more details.","metadata":{}},{"cell_type":"markdown","source":"<div class=\"alert alert-block alert-warning\">\n<b>Note:</b> IterativeImputer is currently experimental, and we need to import enable_iterative_imputer to enable it.\n</div>","metadata":{}},{"cell_type":"code","source":"# explicitly enable this experimental feature\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import IterativeImputer","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.228921Z","iopub.execute_input":"2022-08-02T18:55:51.229573Z","iopub.status.idle":"2022-08-02T18:55:51.374654Z","shell.execute_reply.started":"2022-08-02T18:55:51.229527Z","shell.execute_reply":"2022-08-02T18:55:51.373501Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"imputer = IterativeImputer()\ntrain_df[cols] = imputer.fit_transform(train_df[cols])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:55:51.376111Z","iopub.execute_input":"2022-08-02T18:55:51.376498Z","iopub.status.idle":"2022-08-02T18:56:06.065247Z","shell.execute_reply.started":"2022-08-02T18:55:51.376439Z","shell.execute_reply":"2022-08-02T18:56:06.063026Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now, there are no more missing values, and we can check this using the following code.","metadata":{}},{"cell_type":"code","source":"train_df.isna().sum()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:06.069165Z","iopub.execute_input":"2022-08-02T18:56:06.071705Z","iopub.status.idle":"2022-08-02T18:56:06.102511Z","shell.execute_reply.started":"2022-08-02T18:56:06.071640Z","shell.execute_reply":"2022-08-02T18:56:06.101093Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=\"2\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n   Checking For Class Imbalance\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"We can see that the dataset is highly imbalanced using the code below.\n\nFirst, let's plot a histogram.","metadata":{}},{"cell_type":"code","source":"train_df['failure'].value_counts().plot(kind=\"barh\")","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:06.104984Z","iopub.execute_input":"2022-08-02T18:56:06.105994Z","iopub.status.idle":"2022-08-02T18:56:06.331894Z","shell.execute_reply.started":"2022-08-02T18:56:06.105936Z","shell.execute_reply":"2022-08-02T18:56:06.330952Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Class `0` accounts for around 79% of the data, and class `1` accounts for around 21%:","metadata":{}},{"cell_type":"code","source":"train_df['failure'].value_counts() / len(train_df)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:06.333126Z","iopub.execute_input":"2022-08-02T18:56:06.333627Z","iopub.status.idle":"2022-08-02T18:56:06.342215Z","shell.execute_reply.started":"2022-08-02T18:56:06.333596Z","shell.execute_reply":"2022-08-02T18:56:06.341373Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=\"3\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n   SMOTE Explanation and Code\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"One way to solve the problem of imbalanced data (i.e. too few examples of the minority class) is to oversample the examples in the minority class. This can be done by duplicating existing examples from the training data to increase the number of minority samples you have. However, all this does is balance the class distribution; it does not provide any new information for the model to learn from.\n\nSMOTE is used to synthesize new examples from the training data instead of duplicating them. This is a type of data augmentation.","metadata":{}},{"cell_type":"markdown","source":"The imbalanced-learn package offers an implementation of SMOTE.","metadata":{}},{"cell_type":"code","source":"from imblearn.over_sampling import SMOTE\nfrom imblearn.under_sampling import RandomUnderSampler\nfrom imblearn.pipeline import Pipeline","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:06.343465Z","iopub.execute_input":"2022-08-02T18:56:06.343949Z","iopub.status.idle":"2022-08-02T18:56:06.679696Z","shell.execute_reply.started":"2022-08-02T18:56:06.343920Z","shell.execute_reply":"2022-08-02T18:56:06.678579Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=\"4\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n    Find Optimal Sampling Percentages\n</h2>\n</div>","metadata":{}},{"cell_type":"markdown","source":"The original [paper](https://arxiv.org/abs/1106.1813) suggests using random undersampling to reduce the number of samples in the majority class and then using SMOTE to oversample the minority class in order to balance the class distribution.\n\nIn the following code, we will find the optimal percentages to undersample the majority class at.","metadata":{}},{"cell_type":"code","source":"from sklearn.metrics import roc_auc_score\nfrom sklearn.model_selection import train_test_split\nfrom xgboost import XGBClassifier\n\noptimal_percentage = 0\nbest_roc_score = 0\npossible_percentages = [0.4, 0.5, 0.6, 0.7, 0.8, 0.9, 1]\n\nfor percentage in possible_percentages:\n    \n    X_train, X_val, y_train, y_val = train_test_split(train_df[cols], train_df['failure'])\n    \n    #oversample the minority class to the given percentage of the majority class\n    over = SMOTE(sampling_strategy=percentage) \n    \n    #undersample the majority class to the given percentage\n    under = RandomUnderSampler(sampling_strategy=percentage)\n    \n    # Pipeline combines the two steps\n    pipeline = Pipeline([('o', over), ('u', under)])\n    X_train, y_train = pipeline.fit_resample(X_train, y_train)\n    \n    # verbose = 500 produces the model's logs every 500 iterations\n    model = XGBClassifier(n_estimators=1000, learning_rate=0.01, objective = 'binary:logistic', early_stopping_rounds=20)\n    model.fit(X_train, y_train, eval_set=[(X_val, y_val)], verbose=500)\n    \n    roc_score = roc_auc_score(y_val, model.predict_proba(X_val)[:, 1])\n    \n    print(f\"\\nROC Score for {percentage} : {roc_score}\\n\")\n    \n    if roc_score > best_roc_score:\n        best_roc_score = roc_score\n        optimal_percentage = percentage","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:06.681205Z","iopub.execute_input":"2022-08-02T18:56:06.681596Z","iopub.status.idle":"2022-08-02T18:56:13.395517Z","shell.execute_reply.started":"2022-08-02T18:56:06.681563Z","shell.execute_reply":"2022-08-02T18:56:13.393316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"optimal_percentage","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.396582Z","iopub.status.idle":"2022-08-02T18:56:13.397141Z","shell.execute_reply.started":"2022-08-02T18:56:13.396860Z","shell.execute_reply":"2022-08-02T18:56:13.396887Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=\"5\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n    Create Final Model\n</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"X_train, X_val, y_train, y_val = train_test_split(train_df[cols], train_df['failure'])\n\nover = SMOTE(sampling_strategy=optimal_percentage) \nunder = RandomUnderSampler(sampling_strategy=optimal_percentage)\npipeline = Pipeline([('o', over), ('u', under)])\nX_train, y_train = pipeline.fit_resample(X_train, y_train)\n\nmodel = XGBClassifier(n_estimators=1000, learning_rate=0.01, objective = 'binary:logistic', early_stopping_rounds=20)\nmodel.fit(X_train, y_train, eval_set=[(X_val, y_val)], verbose=500)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.398836Z","iopub.status.idle":"2022-08-02T18:56:13.399423Z","shell.execute_reply.started":"2022-08-02T18:56:13.399112Z","shell.execute_reply":"2022-08-02T18:56:13.399138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=\"6\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n    Prepare Testing Data\n</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"test_df = pd.read_csv(\"../input/tabular-playground-series-aug-2022/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.401396Z","iopub.status.idle":"2022-08-02T18:56:13.401945Z","shell.execute_reply.started":"2022-08-02T18:56:13.401682Z","shell.execute_reply":"2022-08-02T18:56:13.401707Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Categorical Variable Encoding","metadata":{}},{"cell_type":"code","source":"product_code_encoded = []\nfor char in test_df['product_code']:\n    product_code_encoded.append(ord(char) - 65)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.403743Z","iopub.status.idle":"2022-08-02T18:56:13.404272Z","shell.execute_reply.started":"2022-08-02T18:56:13.403999Z","shell.execute_reply":"2022-08-02T18:56:13.404025Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df['product_code'] = product_code_encoded","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.405707Z","iopub.status.idle":"2022-08-02T18:56:13.406223Z","shell.execute_reply.started":"2022-08-02T18:56:13.405961Z","shell.execute_reply":"2022-08-02T18:56:13.405986Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_df['attribute_0'] = test_df[\"attribute_0\"].str.split(\"_\").str[1] # the second value is the material number\ntest_df['attribute_1'] = test_df[\"attribute_1\"].str.split(\"_\").str[1] # the second value is the material number","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.407592Z","iopub.status.idle":"2022-08-02T18:56:13.408135Z","shell.execute_reply.started":"2022-08-02T18:56:13.407868Z","shell.execute_reply":"2022-08-02T18:56:13.407894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Imputation","metadata":{}},{"cell_type":"code","source":"test_df[cols] = imputer.transform(test_df[cols])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.410900Z","iopub.status.idle":"2022-08-02T18:56:13.411606Z","shell.execute_reply.started":"2022-08-02T18:56:13.411314Z","shell.execute_reply":"2022-08-02T18:56:13.411341Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div id=\"7\" style=\"color:white;    \n           display:fill;\n           border-radius:5px;\n           background-color:#2ACAEA;\n           font-size:150%;\n           font-family:Calibri;\n           letter-spacing:0.5px;\n           display:flex;\n           justify-content:center;\">\n\n<h2 style=\"padding: 2rem;\n              color:black;\n          text-align:center;\n          margin:0 auto;\n          \">\n    Predictions and Submission\n</h2>\n</div>","metadata":{}},{"cell_type":"code","source":"probabilities = model.predict_proba(test_df[cols])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.413154Z","iopub.status.idle":"2022-08-02T18:56:13.413739Z","shell.execute_reply.started":"2022-08-02T18:56:13.413430Z","shell.execute_reply":"2022-08-02T18:56:13.413473Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"failure_probabilities = [prob[1] for prob in probabilities]","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.415368Z","iopub.status.idle":"2022-08-02T18:56:13.415891Z","shell.execute_reply.started":"2022-08-02T18:56:13.415638Z","shell.execute_reply":"2022-08-02T18:56:13.415662Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission = pd.DataFrame({\n    'id' : test_df['id'],\n    'failure' : failure_probabilities\n})","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.417287Z","iopub.status.idle":"2022-08-02T18:56:13.417815Z","shell.execute_reply.started":"2022-08-02T18:56:13.417546Z","shell.execute_reply":"2022-08-02T18:56:13.417570Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.419048Z","iopub.status.idle":"2022-08-02T18:56:13.419414Z","shell.execute_reply.started":"2022-08-02T18:56:13.419231Z","shell.execute_reply":"2022-08-02T18:56:13.419249Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"submission.to_csv(\"submission.csv\", index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:56:13.421098Z","iopub.status.idle":"2022-08-02T18:56:13.421689Z","shell.execute_reply.started":"2022-08-02T18:56:13.421365Z","shell.execute_reply":"2022-08-02T18:56:13.421394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Please upvote if you found this notebook helpful :)**","metadata":{}}]}