{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# 1. Introduction\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>1.1 Background</b></p>\n</div>\n\nThis months TPS competition is on **product failure prediction**. The dataset represents the results of a large product **testing study** to improve the 'Super Soaker' product. The features consist of a collection of various **attributes** and **measurements**, which we need to use to predict the **probability of failure** for each sample. The competition metric is the Area Under Receiver Operating Characteristic (**AUROC**). \n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>1.2 Libraries</b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"# Core\nimport numpy as np\nimport pandas as pd\nimport seaborn as sns\nsns.set(style='darkgrid', font_scale=1.4)\nimport matplotlib.pyplot as plt\n%matplotlib inline\nfrom itertools import combinations\nimport math\nimport statistics\nfrom scipy import stats\nfrom scipy.stats import pearsonr\nfrom scipy.stats import shapiro\nfrom scipy.stats import chi2\nfrom scipy.stats import poisson\nimport time\nfrom datetime import datetime\nimport matplotlib.dates as mdates\nimport plotly.express as px\nfrom termcolor import colored\nimport warnings\nwarnings.filterwarnings(\"ignore\")\n\n# Sklearn\nimport sklearn\nfrom sklearn.decomposition import PCA\nfrom sklearn.manifold import TSNE\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis as LDA\nfrom sklearn.cluster import KMeans\nfrom sklearn.model_selection import train_test_split, StratifiedKFold, GridSearchCV, TimeSeriesSplit, GroupKFold, cross_validate\nfrom sklearn.preprocessing import StandardScaler, RobustScaler, PowerTransformer, OneHotEncoder, LabelEncoder\nfrom sklearn.impute import SimpleImputer, KNNImputer\nfrom sklearn.experimental import enable_iterative_imputer\nfrom sklearn.impute import IterativeImputer\nfrom sklearn.pipeline import make_pipeline\nfrom sklearn.compose import make_column_transformer\nfrom sklearn.metrics import confusion_matrix, ConfusionMatrixDisplay, accuracy_score, roc_auc_score\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.linear_model import LinearRegression, Ridge\nfrom sklearn.mixture import GaussianMixture, BayesianGaussianMixture\n\n# UMAP\nimport umap\nimport umap.plot\n\n# Models\nfrom sklearn.linear_model import LinearRegression, LogisticRegression\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.svm import SVC\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.ensemble import RandomForestClassifier\nfrom xgboost import XGBClassifier\nfrom lightgbm import LGBMClassifier\nfrom catboost import CatBoostClassifier\nfrom sklearn.naive_bayes import GaussianNB\n\n!git clone https://github.com/analokmaus/kuma_utils.git\nimport sys; sys.path.append(\"kuma_utils/\")\nfrom kuma_utils.preprocessing.imputer import LGBMImputer","metadata":{"_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:13.093034Z","iopub.execute_input":"2022-08-10T13:17:13.093497Z","iopub.status.idle":"2022-08-10T13:17:14.225138Z","shell.execute_reply.started":"2022-08-10T13:17:13.093463Z","shell.execute_reply":"2022-08-10T13:17:14.223843Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 2. Data\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.1 Load data</b></p>\n</div>\n\n* There are **24** features.\n* The **target is binary** (failure/no failure).\n* The **train/test sets are small** with 26570/20775 samples respectively.","metadata":{}},{"cell_type":"code","source":"# Save to df\ntrain = pd.read_csv('../input/tabular-playground-series-aug-2022/train.csv', index_col='id')\ntest = pd.read_csv('../input/tabular-playground-series-aug-2022/test.csv', index_col='id')\n\n# Shape and preview\nprint('Train set shape:', train.shape)\nprint('Test set shape:', test.shape)\ntrain.head(3)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:17:14.227701Z","iopub.execute_input":"2022-08-10T13:17:14.228715Z","iopub.status.idle":"2022-08-10T13:17:14.446556Z","shell.execute_reply.started":"2022-08-10T13:17:14.228669Z","shell.execute_reply":"2022-08-10T13:17:14.445400Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.2 Missing values</b></p>\n</div>\n\n* **54% of samples** contain at least one missing value.\n* **Only continuous features** (but not all of them) have missing values.\n* The number of missing values **steadily increases** with measurement number.\n* The failure rate is **independent** of the number of missing values.","metadata":{}},{"cell_type":"code","source":"# Missing values summary\nmv=pd.DataFrame(train.isna().sum(), columns=['Number_missing (TRAIN)'])\nmv['Percentage_missing (TRAIN)']=np.round(100*mv['Number_missing (TRAIN)']/len(train),2)\nmv['----']=''\nmv['Number_missing (TEST)']=test.isna().sum()\nmv['Percentage_missing (TEST)']=np.round(100*mv['Number_missing (TEST)']/len(test),2)\nmv","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:14.447874Z","iopub.execute_input":"2022-08-10T13:17:14.448231Z","iopub.status.idle":"2022-08-10T13:17:14.487085Z","shell.execute_reply.started":"2022-08-10T13:17:14.448197Z","shell.execute_reply":"2022-08-10T13:17:14.485930Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Heatmap**","metadata":{}},{"cell_type":"code","source":"# Heatmap of missing values\nplt.figure(figsize=(15,8))\nsns.heatmap(train.isna().T, cmap='summer')\nplt.title('Heatmap of missing values')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:17:14.489911Z","iopub.execute_input":"2022-08-10T13:17:14.490301Z","iopub.status.idle":"2022-08-10T13:17:17.154843Z","shell.execute_reply.started":"2022-08-10T13:17:14.490266Z","shell.execute_reply":"2022-08-10T13:17:17.153654Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Distribution of missing values**","metadata":{}},{"cell_type":"code","source":"# Countplot of number of missing values by sample\ntrain['na_count']=train.isna().sum(axis=1)\nplt.figure(figsize=(10,4))\nsns.countplot(data=train, x='na_count', hue='failure')\nplt.title('Number of missing values for each sample')","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.156221Z","iopub.execute_input":"2022-08-10T13:17:17.156578Z","iopub.status.idle":"2022-08-10T13:17:17.441894Z","shell.execute_reply.started":"2022-08-10T13:17:17.156544Z","shell.execute_reply":"2022-08-10T13:17:17.440729Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Fraction of samples with at least one missing value')\nnp.round((train['na_count']>0).sum()/len(train), 3)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.443357Z","iopub.execute_input":"2022-08-10T13:17:17.443712Z","iopub.status.idle":"2022-08-10T13:17:17.455833Z","shell.execute_reply.started":"2022-08-10T13:17:17.443679Z","shell.execute_reply":"2022-08-10T13:17:17.454503Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Number of failures grouped by number of missing values')\ntrain[train['failure']==1]['na_count'].value_counts()","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.457706Z","iopub.execute_input":"2022-08-10T13:17:17.458302Z","iopub.status.idle":"2022-08-10T13:17:17.472093Z","shell.execute_reply.started":"2022-08-10T13:17:17.458256Z","shell.execute_reply":"2022-08-10T13:17:17.471047Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print('Failure rate grouped by number of missing values')\nprint(train[train['failure']==1]['na_count'].value_counts()/train['na_count'].value_counts())\ntrain.drop('na_count', axis=1, inplace=True)","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.473702Z","iopub.execute_input":"2022-08-10T13:17:17.474071Z","iopub.status.idle":"2022-08-10T13:17:17.489374Z","shell.execute_reply.started":"2022-08-10T13:17:17.474039Z","shell.execute_reply":"2022-08-10T13:17:17.488169Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.3 Duplicates</b></p>\n</div>\n\nThere are **no duplicated values**.","metadata":{}},{"cell_type":"code","source":"print(f'Duplicates in train set: {train.duplicated().sum()}, ({np.round(100*train.duplicated().sum()/len(train),1)}%)')\nprint('')\nprint(f'Duplicates in test set: {test.duplicated().sum()}, ({np.round(100*test.duplicated().sum()/len(test),1)}%)')","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.490902Z","iopub.execute_input":"2022-08-10T13:17:17.491257Z","iopub.status.idle":"2022-08-10T13:17:17.612381Z","shell.execute_reply.started":"2022-08-10T13:17:17.491227Z","shell.execute_reply":"2022-08-10T13:17:17.611076Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.4 Data types</b></p>\n</div>\n\nThere are **3 categorical**, **5 discrete** and **16 continuous** features. The target is **binary**.","metadata":{}},{"cell_type":"code","source":"train.dtypes","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.617526Z","iopub.execute_input":"2022-08-10T13:17:17.617902Z","iopub.status.idle":"2022-08-10T13:17:17.628201Z","shell.execute_reply.started":"2022-08-10T13:17:17.617871Z","shell.execute_reply":"2022-08-10T13:17:17.626913Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>2.5 Data cardinality</b></p>\n</div>\n\nThe **cardinality** of categorical/discrete features is defined to be the **number of unique categories/values**. ","metadata":{}},{"cell_type":"code","source":"dis_cols = [col for col in train.columns if (train[col].dtypes == 'object') or (train[col].dtypes == 'int64')]\ntrain[dis_cols].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:17:17.629413Z","iopub.execute_input":"2022-08-10T13:17:17.630207Z","iopub.status.idle":"2022-08-10T13:17:17.652232Z","shell.execute_reply.started":"2022-08-10T13:17:17.630151Z","shell.execute_reply":"2022-08-10T13:17:17.651304Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 3. EDA\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.1 Target distribution</b></p>\n</div>\n\nThe target is **moderately imbalanced**.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(6,6))\ntrain['failure'].value_counts().plot.pie(explode=[0.1,0.1], autopct='%1.1f%%', shadow=True, textprops={'fontsize':16}).set_title(\"Target distribution\")","metadata":{"_kg_hide-input":false,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.653626Z","iopub.execute_input":"2022-08-10T13:17:17.654689Z","iopub.status.idle":"2022-08-10T13:17:17.807304Z","shell.execute_reply.started":"2022-08-10T13:17:17.654656Z","shell.execute_reply":"2022-08-10T13:17:17.805720Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.2 Product code</b></p>\n</div>\n\nThe product codes are **disjoint** between the train and test sets.","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18,4))\nplt.subplot(1,2,1)\nsns.countplot(data=train, x='product_code', hue='failure')\nplt.title('product_code (TRAIN)')\n\nplt.subplot(1,2,2)\nsns.countplot(data=test, x='product_code')\nplt.title('product_code (TEST)')\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:17.809522Z","iopub.execute_input":"2022-08-10T13:17:17.809997Z","iopub.status.idle":"2022-08-10T13:17:18.238612Z","shell.execute_reply.started":"2022-08-10T13:17:17.809954Z","shell.execute_reply":"2022-08-10T13:17:18.237509Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Insight:*\n* We could try to **fill missing values** according to the product codes. \n* We could use **5 fold cross validation**, where the validation set in each fold is one of the product codes (i.e. GroupKFold).","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.3 Attributes</b></p>\n</div>\n\n* The four attributes are **uniquely determined** by the product code. \n* Some test set values **don't appear** in the train set and vice versa.","metadata":{}},{"cell_type":"code","source":"pd.concat([train,test])[['product_code','attribute_0','attribute_1','attribute_2','attribute_3']].drop_duplicates().set_index('product_code')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:17:18.240155Z","iopub.execute_input":"2022-08-10T13:17:18.241102Z","iopub.status.idle":"2022-08-10T13:17:18.287879Z","shell.execute_reply.started":"2022-08-10T13:17:18.241055Z","shell.execute_reply":"2022-08-10T13:17:18.286803Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for i in range(4):\n    plt.figure(figsize=(18,4))\n    plt.subplot(1,2,1)\n    sns.countplot(data=train, x='attribute_'+str(i), hue='failure')\n    plt.title('attribute_'+str(i)+' (TRAIN)')\n\n    plt.subplot(1,2,2)\n    sns.countplot(data=test, x='attribute_'+str(i))\n    plt.title('attribute_'+str(i)+' (TEST)')\n\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:18.289375Z","iopub.execute_input":"2022-08-10T13:17:18.289952Z","iopub.status.idle":"2022-08-10T13:17:19.689961Z","shell.execute_reply.started":"2022-08-10T13:17:18.289917Z","shell.execute_reply":"2022-08-10T13:17:19.688849Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.4 Loading</b></p>\n</div>\n\nThe **loading** feature measures how much **fluid** each product **absorbs** to see whether or not it fails. ","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18,4))\nplt.subplot(1,2,1)\nsns.histplot(data=train, x='loading', hue='failure', kde=True)\nplt.title('loading (TRAIN)')\n\nplt.subplot(1,2,2)\nsns.histplot(data=test, x='loading', kde=True)\nplt.title('loading (TEST)')\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:19.691345Z","iopub.execute_input":"2022-08-10T13:17:19.691772Z","iopub.status.idle":"2022-08-10T13:17:21.063734Z","shell.execute_reply.started":"2022-08-10T13:17:19.691742Z","shell.execute_reply":"2022-08-10T13:17:21.062455Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(18,4))\nplt.subplot(1,2,1)\nsns.violinplot(data=train, x='failure', y='loading', hue='failure', kde=True)\nplt.title('loading (TRAIN)')\n\nplt.subplot(1,2,2)\nsns.violinplot(data=test, y='loading', kde=True)\nplt.title('loading (TEST)')\n\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:21.065085Z","iopub.execute_input":"2022-08-10T13:17:21.065442Z","iopub.status.idle":"2022-08-10T13:17:21.579682Z","shell.execute_reply.started":"2022-08-10T13:17:21.065404Z","shell.execute_reply":"2022-08-10T13:17:21.578580Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.5 Measurements</b></p>\n</div>\n\n* The first 3 measurements are **discrete** and follow roughly **Poisson** distributions.\n* The other 15 measurements are **continuous** and appear to follow **Gaussian** distributions.\n* The features **don't separate** the target well, i.e. there is **low signal** in the data.","metadata":{}},{"cell_type":"code","source":"for i in range(3):\n    plt.figure(figsize=(22,4))\n    plt.subplot(1,2,1)\n    sns.countplot(data=train, x='measurement_'+str(i), hue='failure')\n    plt.title('measurement_'+str(i)+' (TRAIN)')\n\n    plt.subplot(1,2,2)\n    sns.countplot(data=test, x='measurement_'+str(i))\n    plt.title('measurement_'+str(i)+' (TEST)')\n\n    plt.show()\n    \nfor i in range(3,18):\n    plt.figure(figsize=(22,4))\n    plt.subplot(1,2,1)\n    sns.histplot(data=train, x='measurement_'+str(i), hue='failure')\n    plt.title('measurement_'+str(i)+' (TRAIN)')\n\n    plt.subplot(1,2,2)\n    sns.histplot(data=test, x='measurement_'+str(i))\n    plt.title('measurement_'+str(i)+' (TEST)')\n\n    plt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:21.582378Z","iopub.execute_input":"2022-08-10T13:17:21.583091Z","iopub.status.idle":"2022-08-10T13:17:40.086557Z","shell.execute_reply.started":"2022-08-10T13:17:21.583035Z","shell.execute_reply":"2022-08-10T13:17:40.085282Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.6 Train-test distributions</b></p>\n</div>\n\nWe can plot the train and test **distributions on top of each other** to see how closely they match up. This section is shamelessly stolen from AmbrosM; as the saying goes \"copying is the highest form of flattery\" [Reference: AmbrosM](https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense).","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense\nfloat_cols = [col for col in test.columns if test[col].dtypes == 'float64']\nboth = pd.concat([train[test.columns], test])\n_, axs = plt.subplots(4, 4, figsize=(16,16))\nfor f, ax in zip(float_cols, axs.ravel()):\n    mi = min(train[f].min(), test[f].min())\n    ma = max(train[f].max(), test[f].max())\n    bins = np.linspace(mi, ma, 40)\n    ax.hist(train[f], bins=bins, alpha=0.5, density=True, label='train')\n    ax.hist(test[f], bins=bins, alpha=0.5, density=True, label='test')\n    ax.set_xlabel(f)\n    if ax == axs[0, 0]: ax.legend(loc='lower right')\n        \n    ax2 = ax.twinx()\n    total, _ = np.histogram(train[f], bins=bins)\n    failures, _ = np.histogram(train[f][train.failure == 1], bins=bins)\n    with warnings.catch_warnings(): # ignore divide by zero for empty bins\n        warnings.filterwarnings('ignore', category=RuntimeWarning)\n        ax2.scatter((bins[1:] + bins[:-1]) / 2, failures / total,\n                    color='m', s=10, label='failure probability')\n    ax2.set_ylim(0, 0.5)\n    ax2.tick_params(axis='y', colors='m')\n    if ax == axs[0, 0]: ax2.legend(loc='upper right')\nplt.tight_layout(w_pad=1)\nplt.suptitle('Train and test distributions of the continuous features', fontsize=26, y=1.02)\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:40.087891Z","iopub.execute_input":"2022-08-10T13:17:40.088277Z","iopub.status.idle":"2022-08-10T13:17:48.291771Z","shell.execute_reply.started":"2022-08-10T13:17:40.088239Z","shell.execute_reply":"2022-08-10T13:17:48.290631Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"int_cols = [f for f in train.columns if train[f].dtype == int and f != 'failure']\n_, axs = plt.subplots(2, 3, figsize=(18, 10))\nfor f, ax in zip(int_cols, axs.ravel()):\n    temp1 = train.failure.groupby(train[f]).agg(['mean', 'size'])\n    ax.bar(temp1.index, temp1['size'] / len(train), alpha=0.5, label='train')\n    temp2 = test[f].value_counts()\n    ax.bar(temp2.index, temp2 / len(test), alpha=0.5, label='test')\n    ax.set_xlabel(f)\n    ax.set_ylabel('frequency')\n\n    ax2 = ax.twinx()\n    ax2.scatter(temp1.index, temp1['mean'],\n                color='m', label='failure probability')\n    ax2.set_ylim(0, 0.5)\n    ax2.tick_params(axis='y', colors='m')\n    if ax == axs[0, 0]: ax2.legend(loc='upper right')\n\naxs[0, 0].legend()\naxs[1, 2].axis('off')\nplt.tight_layout(w_pad=1)\nplt.suptitle('Train and test distributions of the integer features', fontsize=26, y=1.02)\nplt.show()\ndel temp1, temp2","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:48.293300Z","iopub.execute_input":"2022-08-10T13:17:48.293650Z","iopub.status.idle":"2022-08-10T13:17:50.541833Z","shell.execute_reply.started":"2022-08-10T13:17:48.293616Z","shell.execute_reply":"2022-08-10T13:17:50.540457Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"string_cols = [f for f in train.columns if train[f].dtype == object]\n_, axs = plt.subplots(1, 3, figsize=(20, 5))\nfor f, ax in zip(string_cols, axs.ravel()):\n    temp1 = train[f].value_counts(dropna=False, normalize=True)\n    temp2 = test[f].value_counts(dropna=False, normalize=True)\n    values = sorted(set(temp1.index).union(temp2.index))\n    temp1 = temp1.reindex(values)\n    temp2 = temp2.reindex(values)\n    ax.bar(range(len(values)), temp1, alpha=0.5, label='train')\n    ax.bar(range(len(values)), temp2, alpha=0.5, label='test')\n    ax.set_xlabel(f)\n    ax.set_ylabel('frequency')\n    ax.set_xticks(range(len(values)), values)\n    \n    temp1 = train.failure.groupby(train[f]).agg(['mean', 'size'])\n    temp1 = temp1.reindex(values)\n    ax2 = ax.twinx()\n    ax2.scatter(range(len(values)), temp1['mean'],\n                color='m', label='failure probability')\n    ax2.tick_params(axis='y', colors='m')\n    ax2.set_ylim(0, 0.5)\n    if ax == axs[0]: ax2.legend(loc='lower right')\n\naxs[0].legend()\nplt.suptitle('Train and test distributions of the string features', fontsize=26, y=0.96)\nplt.tight_layout(w_pad=1)\nplt.show()\ndel temp1, temp2","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:17:50.543500Z","iopub.execute_input":"2022-08-10T13:17:50.544389Z","iopub.status.idle":"2022-08-10T13:17:51.592133Z","shell.execute_reply.started":"2022-08-10T13:17:50.544338Z","shell.execute_reply":"2022-08-10T13:17:51.590884Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"*Insight:*\n\n* Most features seem to be **uncorrelated** to the target. **Loading**, **measurement 2** and **measurement 17** appear to be the **most useful**. \n* There is some **train-test drift** present, especially in measurements 0-2, 10-17 and the four attributes.","metadata":{}},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>3.7 Correlations</b></p>\n</div>\n\nPositive correlations are shown in red and negative correlations in blue.","metadata":{}},{"cell_type":"code","source":"# Heatmap of correlations\nplt.figure(figsize=(10,7))\nsns.heatmap(train.corr(), cmap='bwr', vmin=-1, vmax=1)\nplt.title('Correlations')","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:17:51.593900Z","iopub.execute_input":"2022-08-10T13:17:51.594274Z","iopub.status.idle":"2022-08-10T13:17:52.441714Z","shell.execute_reply.started":"2022-08-10T13:17:51.594243Z","shell.execute_reply":"2022-08-10T13:17:52.439208Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 4. Competition metric\n\nThe competition metric is AUROC. To understand it, we first need to look at the ROC curve.\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>4.1 ROC</b></p>\n</div>\n\nThe **Receiver Operating Characteristic** (ROC) curve measures the performance of a binary classifier. It plots the **True Positive Rate** (TPR) against the **False Positive Rate** (FPR) as the classification threshold is varied. \n\nIn a nutshell, the closer the curve is to the **top left corner** the **better** the classifier is. It was originally developed for operators of military radar receivers in 1941, which is where the name comes from. \n\n<center>\n<img src=\"https://vitalflux.com/wp-content/uploads/2020/09/Screenshot-2020-09-01-at-3.44.15-PM.png\" width=\"450\">\n</center>\n\n<br>\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>4.2 AUROC</b></p>\n</div>\n\nThe **Area Under ROC** (AUROC aka AUC) curve is simply the **area under** the ROC curve and it measures how **good** a binary classifier is. \n\n$$\nAUROC = \\int_0^1 ROC (t) \\, dt\n$$\n\nIt condenses an ROC curve down to a **single number**, which makes it easier to compare the two ROC curves at the cost of losing some information. ","metadata":{}},{"cell_type":"markdown","source":"# 5. Preprocessing\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.1 Imputing missing values</b></p>\n</div>\n\nThe **simplest** way to impute missing values is to use the **median** for continuous features and the **mode** for discrete features. This works 'ok' in practice but is usually not optimal.\n\nAlternatively, **KNNImputer** uses the average of the k closest samples to fill missing values. This is a fast algorithm but can be sensitive to the value of k and any outliers in the dataset. \n\nAnother approach is to use **IterativeImputer**, which models each feature with missing values as a function of other features in a round-robin fashion.","metadata":{}},{"cell_type":"code","source":"%%time\n\n# Concatenate df's\ndata = pd.concat([train.drop(['failure'], axis=1), test], axis=0)\n\n# Fill missing values on a by product code basis\nnum_cols = [col for col in train.columns[:-1] if (train[col].dtypes == 'float64') or (train[col].dtypes == 'int64')]\nfor code in data['product_code'].unique():\n    #imputer = KNNImputer(n_neighbors = 3)\n    imputer = IterativeImputer(max_iter = 8, random_state = 0, skip_complete = True, n_nearest_features = 12)\n    #imputer = LGBMImputer(n_iter=50)\n    data.loc[data['product_code']==code,num_cols] = imputer.fit_transform(data.loc[data['product_code']==code,num_cols])","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:17:52.443400Z","iopub.execute_input":"2022-08-10T13:17:52.443834Z","iopub.status.idle":"2022-08-10T13:18:05.868769Z","shell.execute_reply.started":"2022-08-10T13:17:52.443793Z","shell.execute_reply":"2022-08-10T13:18:05.867458Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.2 Categorical encoding</b></p>\n</div>\n\nThe most common approach to dealing with categorical data is to either use **LabelEncoder** (0, 1, 2, ...) or **OneHotEncoder** (elementary vectors).\n\nLabel encoding replaces each category by an integer 0, 1, 2, etc, whilst one hot encoding converts in category into a new column filled with 0's except when samples fall to that category which we denote by 1.\n\n<center>\n<img src=\"https://miro.medium.com/max/1200/0*T5jaa2othYfXZX9W.\" width=\"600\">\n</center>","metadata":{}},{"cell_type":"code","source":"'''\n# Label encoding\ncat_cols = [col for col in train.columns[:-1] if train[col].dtypes == 'object']\nfor col in cat_cols:\n    enc = LabelEncoder()\n    data[col] = enc.fit_transform(data[col])\n'''\n\n# One-hot encoding\nencoded_columns = ['attribute_0', 'attribute_1']\nfor col in encoded_columns:\n    tempdf = pd.get_dummies(data[col], prefix = col)\n    data = pd.merge(left = data, right = tempdf, left_index = True, right_index = True)\ndata = data.drop(encoded_columns, axis = 1)\n\n# Drop one of the binary one-hot columns - cf 'dummy variable trap'\ndata.drop('attribute_0_material_5', axis=1, inplace=True)","metadata":{"_kg_hide-input":false,"_kg_hide-output":false,"execution":{"iopub.status.busy":"2022-08-10T13:18:05.870434Z","iopub.execute_input":"2022-08-10T13:18:05.870895Z","iopub.status.idle":"2022-08-10T13:18:05.942662Z","shell.execute_reply.started":"2022-08-10T13:18:05.870850Z","shell.execute_reply":"2022-08-10T13:18:05.941667Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.3 Feature engineering</b></p>\n</div>\n\n**Better features** make **better models**. ","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/342319\ndata['measurement_3_na'] = data['measurement_3'].isna().astype(int)\ndata['measurement_5_na'] = data['measurement_5'].isna().astype(int)\n\n# From https://www.kaggle.com/code/desalegngeb/tps08-logisticregression-and-some-fe/notebook?scriptVersionId=102685100\ndata['attribute_2*3'] = data['attribute_2'] * data['attribute_3']\n\n# From https://www.kaggle.com/code/heyspaceturtle/feature-selection-is-all-u-need-2\nmeas_gr1_cols = [f\"measurement_{i:d}\" for i in list(range(3, 5)) + list(range(9, 17))]\ndata['meas_gr1_avg'] = np.mean(data[meas_gr1_cols], axis=1)\ndata['meas_gr1_std'] = np.std(data[meas_gr1_cols], axis=1)\nmeas_gr2_cols = [f\"measurement_{i:d}\" for i in list(range(5, 9))]\ndata['meas_gr2_avg'] = np.mean(data[meas_gr2_cols], axis=1)\n\n# Re-split\ntrain_split = data.iloc[:train.shape[0],:].copy()\ntest_split = data.iloc[train.shape[0]:,:].copy()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:18:05.944031Z","iopub.execute_input":"2022-08-10T13:18:05.945021Z","iopub.status.idle":"2022-08-10T13:18:05.999465Z","shell.execute_reply.started":"2022-08-10T13:18:05.944982Z","shell.execute_reply":"2022-08-10T13:18:05.998099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.4 Scaling</b></p>\n</div>\n\n**Scaling** makes the data **more normally distributed** and typically helps classification algorithms **learn parameters faster**. \n\n\n<center>\n<img src=\"https://149695847.v2.pressablecdn.com/wp-content/uploads/2021/09/image-47.png\" width=\"400\">\n</center>\n\nThere are several ways to do this, e.g.\n\n* *StandardScaler*: scales each column independently to have mean 0 and standard deviation 1, by subtracting by the column **mean** and dividing by the column **standard deviation**.\n* *RobustScaler*: does the same as above but uses statistics that are **robust to outliers**, i.e. it subtracts by the **median** and divides by the **interquartile range**. \n* *PowerTransformer*: makes columns more gaussian like by **stabilising variance** and **minising skew**. ","metadata":{}},{"cell_type":"code","source":"# Scaling\nscale_feats = [col for col in train_split.columns if train_split[col].dtypes == 'float64']\nscaler = StandardScaler()\n#scaler = PowerTransformer()\n\ntrain_scaled = train_split.copy()\ntest_scaled = test_split.copy()\ntrain_scaled[scale_feats] = scaler.fit_transform(train_split[scale_feats])\ntest_scaled[scale_feats] = scaler.transform(test_split[scale_feats])","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:18:06.001036Z","iopub.execute_input":"2022-08-10T13:18:06.001411Z","iopub.status.idle":"2022-08-10T13:18:06.036471Z","shell.execute_reply.started":"2022-08-10T13:18:06.001377Z","shell.execute_reply":"2022-08-10T13:18:06.035401Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>5.5 Feature selection</b></p>\n</div>\n\nThe aim of feature selection is to **drop redundant features** to reduce the chance of our models overfitting to noise. There are several ways to do this, e.g. by using metrics like **correlation coeffecient**, **Mutual Information**, **Fisher Information**, etc. This time we'll use Fisher Information, which tells us how much information about an unknown parameter we can get from a sample.","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/code/heyspaceturtle/feature-selection-is-all-u-need-2\ndef FisherScore(bt, y_train, predictors):\n    \"\"\"\n    Verbeke, W., Dejaeger, K., Martens, D., Hur, J., & Baesens, B. (2012). New insights\n    into churn prediction in the telecommunication sector: A profit driven data mining\n    approach. European Journal of Operational Research, 218(1), 211-229.\n    \"\"\"\n    \n    # Get the unique values of dependent variable\n    target_var_val = y_train.unique()\n    \n    # Calculate FisherScore for each predictor\n    predictor_FisherScore = []\n    for v in predictors:\n        fs = np.abs(np.mean(bt.loc[y_train == target_var_val[0], v]) - np.mean(bt.loc[y_train == target_var_val[1], v])) / \\\n             np.sqrt(np.var(bt.loc[y_train == target_var_val[0], v]) + np.var(bt.loc[y_train == target_var_val[1], v]))\n        predictor_FisherScore.append(fs)\n    return predictor_FisherScore\n\n# Calculate Fisher Score for all variables\nfs = FisherScore(train_scaled, train['failure'], train_scaled.drop('product_code', axis=1).columns)\nfs_df = pd.DataFrame({\"predictor\":train_scaled.drop('product_code', axis=1).columns, \"fisherscore\":fs})\nfs_df = fs_df.sort_values('fisherscore', ascending=False).reset_index(drop=True)\nfs_df.head(10)","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:18:06.037895Z","iopub.execute_input":"2022-08-10T13:18:06.038327Z","iopub.status.idle":"2022-08-10T13:18:06.181242Z","shell.execute_reply.started":"2022-08-10T13:18:06.038295Z","shell.execute_reply":"2022-08-10T13:18:06.180015Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Check how AUC changes when adding more variables sorted by fisher score importance\nfs_scores = []\ntop_n_vars = len(fs_df)\nfor i in range(1, top_n_vars+1):\n    top_n_predictors = fs_df['predictor'][:i]\n    clf = LogisticRegression(max_iter = 200, C=0.05, penalty='l1', solver='liblinear')\n    fs_scores.append(cross_validate(clf, train_scaled[top_n_predictors], train['failure'], groups=train_scaled['product_code'],\n                                    scoring='roc_auc', cv=GroupKFold(n_splits=5), verbose=0, n_jobs=-1, return_train_score=True))\n\n# Plot\nplt.plot([s['train_score'].mean() for s in fs_scores], color='blue')\nplt.plot([s['test_score'].mean() for s in fs_scores], color='red')\nplt.xlabel('# vars')\nplt.ylabel('AUC')\nplt.legend(['train', 'test'])\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:18:06.187509Z","iopub.execute_input":"2022-08-10T13:18:06.188481Z","iopub.status.idle":"2022-08-10T13:18:31.865172Z","shell.execute_reply.started":"2022-08-10T13:18:06.188432Z","shell.execute_reply":"2022-08-10T13:18:31.863493Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Select features\nkeep_feats = fs_df['predictor'][:14]\ntrain_final = train_scaled[keep_feats]\ntest_final = test_scaled[keep_feats]","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:18:31.867365Z","iopub.execute_input":"2022-08-10T13:18:31.868216Z","iopub.status.idle":"2022-08-10T13:18:31.878813Z","shell.execute_reply.started":"2022-08-10T13:18:31.868157Z","shell.execute_reply":"2022-08-10T13:18:31.877660Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 6. Modelling\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>6.1 Training</b></p>\n</div>\n\n* We use a **Logistic Regression model** since others have found this to be performing the best on this dataset. \n* We use **5 fold cross validation** grouped by product_code. \n* Hyperparameters are tuned using **grid search**.","metadata":{}},{"cell_type":"code","source":"%%time\n# From https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense\nauc_list = []\ntest_pred_list = []\nimportance_list = []\nkf = GroupKFold(n_splits=5)   # grouped by product_code\nfor fold, (idx_tr, idx_va) in enumerate(kf.split(X=train, y=train['failure'], groups=train['product_code'])):\n    X_tr = train_final.iloc[idx_tr]\n    X_va = train_final.iloc[idx_va]\n    y_tr = train.loc[idx_tr,'failure']\n    y_va = train.loc[idx_va,'failure']\n    \n    # LR with GridSearch\n    #model = LogisticRegression(max_iter=200 , solver='liblinear', random_state=0)\n    #param_grid = {'C': [0.0001, 0.0005, 0.001, 0.005, 0.01, 0.05, 0.1], 'penalty':['l2']}\n    #clf = GridSearchCV(model, param_grid, cv = 5, verbose = 0)\n    \n    clf = LogisticRegression(max_iter = 200, C=0.05, penalty='l1', solver='liblinear')\n    \n    clf.fit(X_tr, y_tr)\n    importance_list.append(clf.coef_.ravel())\n    \n    # Validate model\n    va_preds = clf.predict_proba(X_va)[:,1]\n    score = roc_auc_score(y_va, va_preds)\n    print(f\"Fold {fold}: auc = {score:.5f}\")\n    auc_list.append(score)\n    \n    # Test set predictions\n    test_pred_list.append(clf.predict_proba(test_final)[:,1])\n    \nprint(f'Average auc = {sum(auc_list) / len(auc_list):.5f}')\nprint('')\npreds = sum(test_pred_list)/len(test_pred_list)","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:18:31.880286Z","iopub.execute_input":"2022-08-10T13:18:31.881050Z","iopub.status.idle":"2022-08-10T13:18:32.931386Z","shell.execute_reply.started":"2022-08-10T13:18:31.881014Z","shell.execute_reply":"2022-08-10T13:18:32.927389Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>6.2 Feature importances</b></p>\n</div>\n\n**Feature importance** is a way to analyse which features were the **most useful** for the model. For many industries it is very important that machine learning models are **explainable**. For logistic regression, the importances are are simply given by the **feature weights**. ","metadata":{}},{"cell_type":"code","source":"# From https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense\n# Show feature importances\nimportance_df = pd.DataFrame(np.array(importance_list).T, index=train_final.columns)\nimportance_df['mean'] = importance_df.mean(axis=1).abs()\nimportance_df['feature'] = train_final.columns\nimportance_df = importance_df.sort_values('mean', ascending=False).reset_index().head(10)\nplt.barh(importance_df.index, importance_df['mean'], color='lightgreen')\nplt.gca().invert_yaxis()\nplt.yticks(ticks=importance_df.index, labels=importance_df['feature'])\nplt.title('LogisticRegression feature importances')\nplt.show()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-10T13:18:32.933163Z","iopub.execute_input":"2022-08-10T13:18:32.934175Z","iopub.status.idle":"2022-08-10T13:18:33.268510Z","shell.execute_reply.started":"2022-08-10T13:18:32.934097Z","shell.execute_reply":"2022-08-10T13:18:33.267099Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 7. Submission\n\n<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>7.1 Ensemble</b></p>\n</div>\n\nIn a [discussion post by Lennart Haupts](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/342403) strong evidence was presented that the **public leaderboard** (LB) is made up of data corresponding to **product_code F** only (about 26% of the test set), meaning that the **private LB** will be made up of the other **product_codes G, H, I**.\n\nLet's take a submission that scores highly on the public LB and our one that scores highly on CV. We'll combine them by separately **ranking** the probabilities and **scaling** them to be between 0 and 1. Then we can **concatenate** the two submissions as they are on the same scale. ([reference](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/343202))","metadata":{}},{"cell_type":"code","source":"# Highest scoring submission on LB (atm)\nhigh_sub = pd.read_csv('../input/tps-aug-22-highest-scoring-public-lb-submission/submission.csv')\n\n# My predictions\nsubmission = pd.DataFrame({'id': test.index,'failure': preds})\n\n# Extract indexes\ncode_F = test[test['product_code']=='F'].index\ncode_GHI = test[test['product_code']!='F'].index\n\n# Rank and scale\nhigh_sub.loc[high_sub['id'].isin(code_F),'failure'] = high_sub.loc[high_sub['id'].isin(code_F),'failure'].rank() / high_sub.loc[high_sub['id'].isin(code_F),'failure'].rank().max()\nsubmission.loc[submission['id'].isin(code_GHI),'failure'] = submission.loc[submission['id'].isin(code_GHI),'failure'].rank() / submission.loc[submission['id'].isin(code_GHI),'failure'].rank().max()\n\n# Combine predictions\nsubmission['failure'] = pd.concat([high_sub.loc[high_sub['id'].isin(code_F),'failure'], submission.loc[submission['id'].isin(code_GHI),'failure']])","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:18:33.270324Z","iopub.execute_input":"2022-08-10T13:18:33.270752Z","iopub.status.idle":"2022-08-10T13:18:33.325629Z","shell.execute_reply.started":"2022-08-10T13:18:33.270713Z","shell.execute_reply":"2022-08-10T13:18:33.324344Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>7.2 Save to csv</b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"submission.to_csv('submission.csv', index=False)\nsubmission.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-10T13:18:33.327943Z","iopub.execute_input":"2022-08-10T13:18:33.328373Z","iopub.status.idle":"2022-08-10T13:18:33.395731Z","shell.execute_reply.started":"2022-08-10T13:18:33.328333Z","shell.execute_reply":"2022-08-10T13:18:33.394386Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"color:white;display:fill;\n            background-color:#6b1515;font-size:150%;\n            font-family:Nexa;letter-spacing:0.5px\">\n    <p style=\"padding: 4px;color:white;\"><b>7.3 Probability distribution</b></p>\n</div>","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(10,4))\nsns.histplot(submission['failure'], kde=True)\nplt.title('Predicted probabilities')","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-08-10T13:18:33.397850Z","iopub.execute_input":"2022-08-10T13:18:33.398315Z","iopub.status.idle":"2022-08-10T13:18:33.838227Z","shell.execute_reply.started":"2022-08-10T13:18:33.398278Z","shell.execute_reply":"2022-08-10T13:18:33.836703Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# 8. References\n\n* [TPS AUG - Simple Baseline](https://www.kaggle.com/code/thedevastator/tps-aug-simple-baseline/notebook?scriptVersionId=102469592) by [The Devastator](https://www.kaggle.com/thedevastator).\n* [TPSAUG22 EDA which makes sense](https://www.kaggle.com/code/ambrosm/tpsaug22-eda-which-makes-sense) by [AmbrosM](https://www.kaggle.com/ambrosm).\n* [TPS 08.22 | Logistic Regression](https://www.kaggle.com/code/kostiantynlavronenko/tps-08-22-logistic-regression) by [Kostiantyn Lavronenko](https://www.kaggle.com/kostiantynlavronenko).\n* [TPS08: LogisticRegression, kNN and ExtraTrees](https://www.kaggle.com/code/desalegngeb/tps08-logisticregression-knn-and-extratrees/notebook?scriptVersionId=102603134) by [des.](https://www.kaggle.com/desalegngeb)\n* [The area under the ROC curve. AUROC.](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/341037) by [Marília Prata](https://www.kaggle.com/mpwolke).\n* [feature selection is all u need 2](https://www.kaggle.com/code/heyspaceturtle/feature-selection-is-all-u-need-2) by [Fernando Delgado\n](https://www.kaggle.com/heyspaceturtle).\n* [TPS - Aug 2022](https://www.kaggle.com/code/sfktrkl/tps-aug-2022/notebook) by [Şafak Türkeli](https://www.kaggle.com/sfktrkl).\n* [TPS-Aug22 LB 0.59013](https://www.kaggle.com/code/takanashihumbert/tps-aug22-lb-0-59013) by [Sawaimilert](https://www.kaggle.com/takanashihumbert).","metadata":{}}]}