{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Tabular Playground Series - Aug 2022","metadata":{}},{"cell_type":"markdown","source":"# Introduction","metadata":{}},{"cell_type":"markdown","source":"> The August 2022 edition of the Tabular Playground Series in an opportunity to help the fictional company Keep It Dry improve its main product Super Soaker. The product is used in factories to absorb spills and leaks.\n\n> The company has just completed a large testing study for different product prototypes. Can you use this data to build a model that predicts product failures?","metadata":{}},{"cell_type":"markdown","source":"This notebook contains:\n\n- EDA to give an overview of the data\n- Logistic Regression Baseline Model","metadata":{}},{"cell_type":"markdown","source":"# Part 1 - EDA","metadata":{}},{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nsns.set_style('darkgrid')","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-02T22:01:25.967359Z","iopub.execute_input":"2022-08-02T22:01:25.967752Z","iopub.status.idle":"2022-08-02T22:01:26.463550Z","shell.execute_reply.started":"2022-08-02T22:01:25.967677Z","shell.execute_reply":"2022-08-02T22:01:26.462531Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_df = pd.read_csv(\"../input/tabular-playground-series-aug-2022/train.csv\")\ntest_df = pd.read_csv(\"../input/tabular-playground-series-aug-2022/test.csv\")","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:26.465081Z","iopub.execute_input":"2022-08-02T22:01:26.465416Z","iopub.status.idle":"2022-08-02T22:01:26.636195Z","shell.execute_reply.started":"2022-08-02T22:01:26.465380Z","shell.execute_reply":"2022-08-02T22:01:26.635325Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(train_df.shape)\nprint(test_df.shape)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:26.637055Z","iopub.execute_input":"2022-08-02T22:01:26.637275Z","iopub.status.idle":"2022-08-02T22:01:26.642698Z","shell.execute_reply.started":"2022-08-02T22:01:26.637253Z","shell.execute_reply":"2022-08-02T22:01:26.641634Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n\n- The train and test set are both fairly small, overfitting is likely going to play a part in this competition.","metadata":{}},{"cell_type":"code","source":"display(train_df.head())\ndisplay(test_df.head())","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:26.645169Z","iopub.execute_input":"2022-08-02T22:01:26.645965Z","iopub.status.idle":"2022-08-02T22:01:26.695861Z","shell.execute_reply.started":"2022-08-02T22:01:26.645940Z","shell.execute_reply":"2022-08-02T22:01:26.695216Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Lets try and group some of the different features together so we can refer to them easily later\nmeasurement_cols = [i for i in train_df.columns if \"measurement\" in i]\nmeasurement_int_cols = [i for i in measurement_cols if train_df[i].dtype == int]\nmeasurement_float_cols = [i for i in train_df.columns if train_df[i].dtype == float]\nfloat_cols = [i for i in train_df.columns if train_df[i].dtype == float]\nattribute_cols = [i for i in train_df.columns if \"attribute\" in i]","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:26.696878Z","iopub.execute_input":"2022-08-02T22:01:26.697259Z","iopub.status.idle":"2022-08-02T22:01:26.705905Z","shell.execute_reply.started":"2022-08-02T22:01:26.697234Z","shell.execute_reply":"2022-08-02T22:01:26.704968Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Missing data","metadata":{}},{"cell_type":"markdown","source":"Before we do anything else, lets look at the missing data","metadata":{}},{"cell_type":"code","source":"missing_values = pd.concat([train_df.isna().sum().rename(\"train\"), test_df.isna().sum().rename(\"test\")], axis=1)\n#display(missing_values)\nmissing_values = pd.concat([train_df.isna().sum(), test_df.isna().sum()], axis=0).rename(\"missing values\").reset_index().rename(columns={\"index\":\"column\"})\nmissing_values[\"data\"] = [\"train\"]*len(train_df.columns) + [\"test\"]*len(test_df.columns)\nf,ax = plt.subplots(figsize=(15,15))\nsns.barplot(data = missing_values, y=\"column\", x=\"missing values\", hue=\"data\", orient=\"h\");","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:26.706903Z","iopub.execute_input":"2022-08-02T22:01:26.707166Z","iopub.status.idle":"2022-08-02T22:01:27.217511Z","shell.execute_reply.started":"2022-08-02T22:01:26.707143Z","shell.execute_reply":"2022-08-02T22:01:27.215647Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n- Data is missing for all of the float columns\n- Missing data looks like its been artificially introduced\n- The train and test set have a similar percentage of missing values","metadata":{}},{"cell_type":"markdown","source":"## Target - Failure of the product","metadata":{}},{"cell_type":"code","source":"def val_count_df(df, column_name, sort_by_column_name=False):\n    value_count = df[column_name].value_counts().reset_index().rename(columns={column_name:\"Value Count\",\"index\":column_name}).set_index(column_name)\n    value_count[\"Percentage\"] = df[column_name].value_counts(normalize=True)*100\n    value_count = value_count.reset_index()\n    if sort_by_column_name:\n        value_count = value_count.sort_values(column_name)\n    return value_count","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.219001Z","iopub.execute_input":"2022-08-02T22:01:27.219437Z","iopub.status.idle":"2022-08-02T22:01:27.225724Z","shell.execute_reply.started":"2022-08-02T22:01:27.219401Z","shell.execute_reply":"2022-08-02T22:01:27.224744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def plot_and_display_valuecounts(df, column_name, sort_by_column_name):\n    val_count = val_count_df(df, column_name, sort_by_column_name)\n    display(val_count)\n    \n    val_count.set_index(column_name).plot.pie(y=\"Value Count\", figsize=(5,5), legend=False, ylabel=\"\");\n    \ndef plot_and_display_compare_valuecounts(df1, df2, column_name, sort_by_column_name):\n    val_count_1 = val_count_df(df1, column_name, sort_by_column_name)\n    val_count_1 = val_count_1.rename(columns={\"Value Count\":\"train_value_count\", \"Percentage\":\"train_percentage\"})\n    val_count_2 = val_count_df(df2, column_name, sort_by_column_name)\n    val_count_2 = val_count_2.rename(columns={\"Value Count\":\"test_value_count\", \"Percentage\":\"test_percentage\"})\n    \n    val_count = pd.merge(val_count_1, val_count_2, on=column_name, how=\"outer\")\n    val_count = val_count.fillna(0) # if the data is missing from a column, there is none so we fill with 0's\n    display(val_count)\n    \n    val_count = val_count.drop(columns=[\"train_percentage\", \"test_percentage\"]) #avoid duplicating pie plots\n    val_count.set_index(column_name).plot.pie(figsize=(12,7), legend=False, ylabel=\"\", subplots=True, title=[\"Train\",\"Test\"]);\n    \n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.227121Z","iopub.execute_input":"2022-08-02T22:01:27.227712Z","iopub.status.idle":"2022-08-02T22:01:27.237035Z","shell.execute_reply.started":"2022-08-02T22:01:27.227679Z","shell.execute_reply":"2022-08-02T22:01:27.235889Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plot_and_display_valuecounts(train_df, \"failure\", False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.238224Z","iopub.execute_input":"2022-08-02T22:01:27.238547Z","iopub.status.idle":"2022-08-02T22:01:27.343875Z","shell.execute_reply.started":"2022-08-02T22:01:27.238517Z","shell.execute_reply":"2022-08-02T22:01:27.343138Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n\n- We have imbalanced classes, failure is a lot less common than success (non-failure)","metadata":{}},{"cell_type":"markdown","source":"## Cateogorical columns","metadata":{}},{"cell_type":"markdown","source":"The attribute data and the product code look like cateogorical features, lets take a closer look at the frequency of each possible value that each of these attributes can  take:","metadata":{}},{"cell_type":"markdown","source":"### Product code","metadata":{}},{"cell_type":"code","source":"plot_and_display_compare_valuecounts(train_df, test_df, \"product_code\", True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.347818Z","iopub.execute_input":"2022-08-02T22:01:27.348289Z","iopub.status.idle":"2022-08-02T22:01:27.525840Z","shell.execute_reply.started":"2022-08-02T22:01:27.348257Z","shell.execute_reply":"2022-08-02T22:01:27.525096Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations**\n- The product code is different in train and test, this'll provide an extra challenge for validating our models.\n- There is roughly the same number of occurances of each product code (~5000)","metadata":{}},{"cell_type":"markdown","source":"### attribute_0","metadata":{}},{"cell_type":"code","source":"plot_and_display_compare_valuecounts(train_df, test_df, \"attribute_0\", True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.527132Z","iopub.execute_input":"2022-08-02T22:01:27.527664Z","iopub.status.idle":"2022-08-02T22:01:27.669596Z","shell.execute_reply.started":"2022-08-02T22:01:27.527621Z","shell.execute_reply":"2022-08-02T22:01:27.668702Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n- attribute_0 has a different value distribution in train and test\n- There are no materials just unique to train, or just unique to test","metadata":{}},{"cell_type":"markdown","source":"### attribute_1","metadata":{}},{"cell_type":"code","source":"plot_and_display_compare_valuecounts(train_df, test_df, \"attribute_1\", True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.670710Z","iopub.execute_input":"2022-08-02T22:01:27.671526Z","iopub.status.idle":"2022-08-02T22:01:27.825007Z","shell.execute_reply.started":"2022-08-02T22:01:27.671492Z","shell.execute_reply":"2022-08-02T22:01:27.823937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n- Material_8 is unique to train and material_7 is unique to test\n- The distributins of the materials is different in train and test","metadata":{}},{"cell_type":"markdown","source":"### attribute_2","metadata":{}},{"cell_type":"code","source":"plot_and_display_compare_valuecounts(train_df, test_df, \"attribute_2\", True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.826417Z","iopub.execute_input":"2022-08-02T22:01:27.826720Z","iopub.status.idle":"2022-08-02T22:01:27.973983Z","shell.execute_reply.started":"2022-08-02T22:01:27.826688Z","shell.execute_reply":"2022-08-02T22:01:27.973339Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 7 is unique to test and 5 and 8 are unique to train\n- The value distributins are different in train and test","metadata":{}},{"cell_type":"markdown","source":"### attribute_3","metadata":{}},{"cell_type":"code","source":"plot_and_display_compare_valuecounts(train_df, test_df, \"attribute_3\", True)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:27.975232Z","iopub.execute_input":"2022-08-02T22:01:27.975784Z","iopub.status.idle":"2022-08-02T22:01:28.148365Z","shell.execute_reply.started":"2022-08-02T22:01:27.975744Z","shell.execute_reply":"2022-08-02T22:01:28.147650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"- 7 and 4 are unique to the test data and 6 and 8 are unique to train data\n- The value distributins are different in train and test","metadata":{}},{"cell_type":"markdown","source":"### Product Code and attribute values","metadata":{}},{"cell_type":"markdown","source":"From looking at the attributes we can see that they appear to come in groups of 5000 values, similar to the product codes. It's possible that they are related. takanashihumbert investigated this ([Notebook](https://www.kaggle.com/code/takanashihumbert/interesting-patterns-found-in-category-features)). Each product code only has 1 unique value for each of the four attributes. In otherwords the product code is the key corresponding to specific attribute combinations.","metadata":{}},{"cell_type":"markdown","source":"Heres proof of this:","metadata":{}},{"cell_type":"code","source":"pd.concat([train_df,test_df]).groupby([\"product_code\"])[[\"attribute_0\", \"attribute_1\", \"attribute_2\", \"attribute_3\"]].nunique()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:28.149478Z","iopub.execute_input":"2022-08-02T22:01:28.149887Z","iopub.status.idle":"2022-08-02T22:01:28.201462Z","shell.execute_reply.started":"2022-08-02T22:01:28.149861Z","shell.execute_reply":"2022-08-02T22:01:28.200735Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The attribute combination for each of the product codes are:","metadata":{}},{"cell_type":"code","source":"pd.concat([train_df,test_df]).groupby([\"product_code\"])[[\"attribute_0\", \"attribute_1\", \"attribute_2\", \"attribute_3\"]].first()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:28.202586Z","iopub.execute_input":"2022-08-02T22:01:28.202989Z","iopub.status.idle":"2022-08-02T22:01:28.240617Z","shell.execute_reply.started":"2022-08-02T22:01:28.202963Z","shell.execute_reply":"2022-08-02T22:01:28.239745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distributions","metadata":{}},{"cell_type":"markdown","source":"Taking a look at the distributions of ","metadata":{}},{"cell_type":"markdown","source":"### Float distrubutions","metadata":{}},{"cell_type":"code","source":"plt.subplots(figsize=(25,35))\ntrain_df[\"df\"] = \"train\"\ntest_df[\"df\"] = \"test\"\nfor i, column in enumerate(float_cols):\n    plt.subplot(6,3,i+1)\n    sns.histplot(data=pd.concat([train_df, test_df]).reset_index(drop=True), x=column,hue=\"df\")\n    plt.title(column)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:28.243214Z","iopub.execute_input":"2022-08-02T22:01:28.243605Z","iopub.status.idle":"2022-08-02T22:01:39.008661Z","shell.execute_reply.started":"2022-08-02T22:01:28.243526Z","shell.execute_reply":"2022-08-02T22:01:39.007801Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations:**\n\n\n\n- It looks like the train and test distributions are different for many features (e.g. measurement_11)\n- The measurement float features are all normally distributed\n- The loading feature has a positively skewed distribution","metadata":{}},{"cell_type":"markdown","source":"### Integer distributions","metadata":{}},{"cell_type":"code","source":"plt.subplots(figsize=(25,30))\nfor i, column in enumerate(measurement_int_cols):\n    val_count = pd.concat([train_df, test_df])[[column,\"df\"]].value_counts().rename(\"value_counts\").reset_index()\n    plt.subplot(5,3,i+1)\n    ax = sns.barplot(data = val_count, x=column, y=\"value_counts\", hue=\"df\")\n    ax.set_xlabel(None)\n    plt.title(column)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:39.009855Z","iopub.execute_input":"2022-08-02T22:01:39.010194Z","iopub.status.idle":"2022-08-02T22:01:40.978664Z","shell.execute_reply.started":"2022-08-02T22:01:39.010162Z","shell.execute_reply":"2022-08-02T22:01:40.977765Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations**\n\n- The train and test distributions are different for all three integer measurement features","metadata":{}},{"cell_type":"markdown","source":"## Correlations","metadata":{}},{"cell_type":"code","source":"plt.subplots(figsize=(15,15))\nsns.heatmap(train_df[[\"loading\"] + measurement_cols].corr(),annot=True, cmap=\"RdYlGn\", fmt = '0.2f', vmin=-1, vmax=1, cbar=False);","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:40.979846Z","iopub.execute_input":"2022-08-02T22:01:40.980097Z","iopub.status.idle":"2022-08-02T22:01:42.177994Z","shell.execute_reply.started":"2022-08-02T22:01:40.980074Z","shell.execute_reply":"2022-08-02T22:01:42.176777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Observations**\n\n- There are some correlations between measurements, particularly measurement 17.","metadata":{}},{"cell_type":"markdown","source":"## Relationships with target","metadata":{}},{"cell_type":"markdown","source":"### Mutual Information","metadata":{}},{"cell_type":"markdown","source":"Firsty lets consider Mutual Information:\n\n> The mutual information (MI) between two quantities is a measure of the extent to which knowledge of one quantity reduces uncertainty about the other. If you knew the value of a feature, how much more confident would you be about the target? - [Reference](https://www.kaggle.com/code/ryanholbrook/mutual-information)\n\nMI is similar to the correlation between feature and target only MI can detect non-linear relationships.\n","metadata":{"execution":{"iopub.status.busy":"2022-08-01T23:42:07.773390Z","iopub.execute_input":"2022-08-01T23:42:07.773715Z","iopub.status.idle":"2022-08-01T23:42:07.780314Z","shell.execute_reply.started":"2022-08-01T23:42:07.773689Z","shell.execute_reply":"2022-08-01T23:42:07.779193Z"}}},{"cell_type":"code","source":"X = train_df.drop(columns=[\"id\",\"failure\",\"product_code\",\"df\"]+attribute_cols).dropna()\ny = train_df.loc[X.index, \"failure\"]","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:42.179582Z","iopub.execute_input":"2022-08-02T22:01:42.179911Z","iopub.status.idle":"2022-08-02T22:01:42.191576Z","shell.execute_reply.started":"2022-08-02T22:01:42.179881Z","shell.execute_reply":"2022-08-02T22:01:42.190635Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.feature_selection import mutual_info_classif\n\ndef make_mi_scores(X, y, discrete_features):\n    mi_scores = mutual_info_classif(X, y, discrete_features=discrete_features)\n    mi_scores = pd.Series(mi_scores, name=\"MI Scores\", index=X.columns)\n    mi_scores = mi_scores.sort_values(ascending=False)\n    return mi_scores\n\nmi_scores = make_mi_scores(X, y, X.dtypes == int)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:42.193189Z","iopub.execute_input":"2022-08-02T22:01:42.194167Z","iopub.status.idle":"2022-08-02T22:01:43.012260Z","shell.execute_reply.started":"2022-08-02T22:01:42.194140Z","shell.execute_reply":"2022-08-02T22:01:43.011357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"f,ax = plt.subplots(figsize=(20,10))\nsns.barplot(y=mi_scores.index, x=mi_scores.values, color=\"blue\");","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:43.013526Z","iopub.execute_input":"2022-08-02T22:01:43.013802Z","iopub.status.idle":"2022-08-02T22:01:43.302168Z","shell.execute_reply.started":"2022-08-02T22:01:43.013779Z","shell.execute_reply":"2022-08-02T22:01:43.301311Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**Warning:**\n\nMutual Information can't detect interactions between features. It is a univariate metric so these MI scores (and the plots below) may not be a good predictor of actual feature importance.","metadata":{}},{"cell_type":"markdown","source":"### Visualising float features relationship with target","metadata":{}},{"cell_type":"markdown","source":"Often its more useful to see the features relationship with the target visually:\n\nTo assess the performance of float features we:\n\n1. Sort values by the feature value\n2. Calculate the rolling mean of the target values with a large window size\n3. Plot the feature value against the rolling mean of the target","metadata":{}},{"cell_type":"code","source":"f,ax = plt.subplots(figsize = (20,60))\nfor n, col in enumerate([\"id\"]+measurement_float_cols):\n    temp_df = pd.DataFrame({col: train_df[col].values, \"target\":train_df[\"failure\"]})\n    temp_df = temp_df.sort_values(col).reset_index(drop=True)\n    temp_df[\"rolling_mean\"] = temp_df[\"target\"].rolling(1000, center=True).mean()\n    \n    ax = plt.subplot(10,2,n+1)\n    sns.scatterplot(data=temp_df,x=col,y=\"rolling_mean\",s=3)\n    ax.set_ylim([0.1,0.4])\n    ax.set_ylabel(\"Rolling mean failure\")","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:43.303921Z","iopub.execute_input":"2022-08-02T22:01:43.304690Z","iopub.status.idle":"2022-08-02T22:01:46.360035Z","shell.execute_reply.started":"2022-08-02T22:01:43.304664Z","shell.execute_reply":"2022-08-02T22:01:46.358934Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Visualising integer features relationship with target","metadata":{}},{"cell_type":"markdown","source":"For each feature we plot the mean target value for all categories/integers. We remove any integer values with low value counts for clarity.","metadata":{}},{"cell_type":"code","source":"f,ax = plt.subplots(figsize=(20,20))\nfor i, column in enumerate( [\"product_code\"] + attribute_cols + measurement_int_cols):\n    temp_df = train_df.groupby([column])[\"failure\"].mean()\n    temp_df_2 = train_df[column].value_counts()\n    temp_df = temp_df[temp_df_2 > 50]\n    plt.subplot(6,2,i+1)\n    ax = sns.barplot(x=temp_df.index, y=temp_df.values, color=\"blue\")\n    ax.set_ylim([0.15,0.3])\n    plt.ylabel(\"mean failure\")\n    plt.xlabel(None)\n    plt.title(\"Feature: \" + column)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:46.361421Z","iopub.execute_input":"2022-08-02T22:01:46.361711Z","iopub.status.idle":"2022-08-02T22:01:47.578383Z","shell.execute_reply.started":"2022-08-02T22:01:46.361685Z","shell.execute_reply":"2022-08-02T22:01:47.577731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Part 2 - Baseline Model","metadata":{}},{"cell_type":"markdown","source":"Import relevant modules","metadata":{}},{"cell_type":"code","source":"from sklearn.impute import KNNImputer, SimpleImputer\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.metrics import roc_auc_score, accuracy_score\nfrom sklearn.model_selection import GroupKFold\nfrom sklearn.preprocessing import StandardScaler, PowerTransformer","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:23:09.955430Z","iopub.execute_input":"2022-08-02T22:23:09.955760Z","iopub.status.idle":"2022-08-02T22:23:09.960208Z","shell.execute_reply.started":"2022-08-02T22:23:09.955734Z","shell.execute_reply":"2022-08-02T22:23:09.959508Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Remove the features we no longer need","metadata":{}},{"cell_type":"code","source":"train_df = train_df.drop(columns=[\"id\",\"df\"])\ntest_df = test_df.drop(columns=[\"id\",\"df\"])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:47.591014Z","iopub.execute_input":"2022-08-02T22:01:47.591777Z","iopub.status.idle":"2022-08-02T22:01:47.602320Z","shell.execute_reply.started":"2022-08-02T22:01:47.591750Z","shell.execute_reply":"2022-08-02T22:01:47.601343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Seperate the training data (X) from the labels (y)","metadata":{}},{"cell_type":"code","source":"X = train_df.drop(columns=\"failure\")\ny = train_df[\"failure\"]\n\nX_test = test_df","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:47.603402Z","iopub.execute_input":"2022-08-02T22:01:47.604213Z","iopub.status.idle":"2022-08-02T22:01:47.610588Z","shell.execute_reply.started":"2022-08-02T22:01:47.604179Z","shell.execute_reply":"2022-08-02T22:01:47.609860Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define functions for future use to keep things looking clean/readable inside the cross validation function","metadata":{}},{"cell_type":"code","source":"def _scale(train_data, val_data):\n    \"\"\"scales the data, takes train and validation as input and outputs the scaled train data and scaled validation data\"\"\"\n    #scaler = StandardScaler()\n    scaler = PowerTransformer()\n    \n    scaled_train = scaler.fit_transform(train_data[measurement_cols + [\"loading\"]])\n    scaled_val = scaler.transform(val_data[measurement_cols + [\"loading\"]])\n    \n    #back to dataframe\n    new_train = train_data.copy()\n    new_val = val_data.copy()\n    \n    new_train[measurement_cols + [\"loading\"]] = scaled_train\n    new_val[measurement_cols + [\"loading\"]] = scaled_val\n    \n    assert len(train_data) == len(new_train)\n    assert len(val_data) == len(val_data)\n    \n    return new_train, new_val\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:47.614712Z","iopub.execute_input":"2022-08-02T22:01:47.615385Z","iopub.status.idle":"2022-08-02T22:01:47.621517Z","shell.execute_reply.started":"2022-08-02T22:01:47.615356Z","shell.execute_reply":"2022-08-02T22:01:47.620666Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def _impute(train_data, val_data):\n    \"\"\"imputes the data, takes train and validation as input and outputs the imputed train data and imputed validation data\"\"\"\n    #imputer = KNNImputer(n_neighbors=3)\n    imputer = SimpleImputer(strategy=\"mean\")\n    imputer.fit(train_data[measurement_cols + [\"loading\"]])\n    \n    filled_train = imputer.transform(train_data[measurement_cols + [\"loading\"]])\n    filled_val = imputer.transform(val_data[measurement_cols + [\"loading\"]])\n    \n    #back to dataframe\n    new_train = train_data.copy()\n    new_val = val_data.copy()\n    \n    new_train[measurement_cols + [\"loading\"]] = filled_train\n    new_val[measurement_cols + [\"loading\"]] = filled_val\n    \n    assert len(train_data) == len(new_train)\n    assert len(val_data) == len(val_data)\n    \n    return new_train, new_val\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:47.622774Z","iopub.execute_input":"2022-08-02T22:01:47.623340Z","iopub.status.idle":"2022-08-02T22:01:47.630153Z","shell.execute_reply.started":"2022-08-02T22:01:47.623284Z","shell.execute_reply":"2022-08-02T22:01:47.629595Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def _ohe(train_data, val_data):\n    \"\"\"one-hot-encodes the data, takes train and validation data as input and outputs the \"one-hot-encoded train and validation data\"\"\"\n    \n    new_train = pd.get_dummies(train_data, columns=[\"product_code\",\"attribute_0\", \"attribute_1\", \"attribute_2\", \"attribute_3\"])\n    new_val = pd.get_dummies(val_data, columns=[\"product_code\",\"attribute_0\", \"attribute_1\", \"attribute_2\", \"attribute_3\"])\n    \n    #columns are not currently the same, concat so that they are\n    train_val = pd.concat([new_train, new_val]).fillna(0) #creates some empty columns, fill these with 0's\n    \n    #extract train and val again\n    new_train = train_val.iloc[0:len(train_data)]\n    new_val = train_val.iloc[len(train_data):]\n    \n    assert len(train_data) == len(new_train)\n    assert len(val_data) == len(val_data)\n    \n    return new_train, new_val\n    ","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:01:47.631240Z","iopub.execute_input":"2022-08-02T22:01:47.631756Z","iopub.status.idle":"2022-08-02T22:01:47.641335Z","shell.execute_reply.started":"2022-08-02T22:01:47.631725Z","shell.execute_reply":"2022-08-02T22:01:47.640773Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define the cross-validation function.\n\nWe use groupKFold with the product code as the group. It's important we do this, as, from the EDA we know that the product codes that occur in train do not occur in test. Therefore when we perform validation we want the train set to contain product codes that do not occur in the validation set.","metadata":{}},{"cell_type":"code","source":"def k_fold_cv(model,X,y):\n    kfold = GroupKFold(n_splits=5)\n\n    feature_imp, y_pred_list, y_true_list, roc_list  = [],[],[],[]\n    for fold, (train_index, val_index) in enumerate(kfold.split(X, y, train_df[\"product_code\"])):\n        print(\"===== fold\", fold, \"=====\")\n        X_train = X.loc[train_index]\n        X_val = X.loc[val_index]\n\n        y_train = y.loc[train_index]\n        y_val = y.loc[val_index]\n            \n        #impute\n        X_train, X_val = _impute(X_train, X_val)\n            \n        #scale the data\n        X_train, X_val = _scale(X_train, X_val)\n            \n        #encode categorical variables\n        X_train, X_val = _ohe(X_train, X_val)\n            \n        # fit the model\n        model.fit(X_train,y_train)\n            \n        #make predictions\n        y_pred = model.predict_proba(X_val)[:,1]\n            \n        #save predictions for later\n        y_pred_list = np.append(y_pred_list, y_pred)\n        y_true_list = np.append(y_true_list, y_val)\n        \n        #evaluate performance\n        roc_list.append(roc_auc_score(y_val,y_pred))\n        print(\"roc auc\", roc_auc_score(y_val,y_pred))\n            \n        #feature imporance\n        try:\n            feature_imp.append(model.feature_importances_)\n        except AttributeError: # if model does not have .feature_importances_ attribute\n            pass # returns empty list\n    return feature_imp, y_pred_list, y_true_list, roc_list, X_val, y_val","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:26:04.755198Z","iopub.execute_input":"2022-08-02T22:26:04.755569Z","iopub.status.idle":"2022-08-02T22:26:04.766025Z","shell.execute_reply.started":"2022-08-02T22:26:04.755539Z","shell.execute_reply":"2022-08-02T22:26:04.764734Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We define our model here, you can easily switch it out for your own. I chose the hyperparameters after a little bit of trial and error.","metadata":{}},{"cell_type":"code","source":"model = LogisticRegression(penalty='elasticnet', l1_ratio=0.8, C=0.007, tol = 1e-2, solver='saga', max_iter=3000, random_state=5)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:26:05.583290Z","iopub.execute_input":"2022-08-02T22:26:05.583756Z","iopub.status.idle":"2022-08-02T22:26:05.588737Z","shell.execute_reply.started":"2022-08-02T22:26:05.583730Z","shell.execute_reply":"2022-08-02T22:26:05.587650Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now lets perform our cross validation, I return some variables that might potentially be useful for evaluation","metadata":{}},{"cell_type":"code","source":"%%time\nfeature_imp, y_pred_list, y_true_list, roc_list, X_val, y_val = k_fold_cv(model=model,X=X,y=y)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:26:06.568808Z","iopub.execute_input":"2022-08-02T22:26:06.569181Z","iopub.status.idle":"2022-08-02T22:26:09.083039Z","shell.execute_reply.started":"2022-08-02T22:26:06.569152Z","shell.execute_reply":"2022-08-02T22:26:09.082314Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"print(\"Mean ROC AUC Score:\", np.mean(roc_list))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:26:11.400796Z","iopub.execute_input":"2022-08-02T22:26:11.401139Z","iopub.status.idle":"2022-08-02T22:26:11.407006Z","shell.execute_reply.started":"2022-08-02T22:26:11.401111Z","shell.execute_reply":"2022-08-02T22:26:11.405692Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For reference, if we were to predict 0 (or 1) for every class, the ROC AUC score will be 0.5","metadata":{}},{"cell_type":"code","source":"print(roc_auc_score(y_true_list, np.zeros(len(y_true_list))))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T22:03:04.231949Z","iopub.execute_input":"2022-08-02T22:03:04.232271Z","iopub.status.idle":"2022-08-02T22:03:04.242497Z","shell.execute_reply.started":"2022-08-02T22:03:04.232242Z","shell.execute_reply":"2022-08-02T22:03:04.241719Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"val_preds = pd.DataFrame({\"pred_prob=1\":y_pred_list, \"y_val\":y_true_list})\nf,ax = plt.subplots(figsize=(20,20))\nplt.subplot(2,1,1)\nax = sns.histplot(data=val_preds, x=\"pred_prob=1\", hue=\"y_val\", bins = 100)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T21:07:29.681137Z","iopub.execute_input":"2022-08-02T21:07:29.681856Z","iopub.status.idle":"2022-08-02T21:07:30.742865Z","shell.execute_reply.started":"2022-08-02T21:07:29.681800Z","shell.execute_reply":"2022-08-02T21:07:30.741924Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Inference","metadata":{}},{"cell_type":"markdown","source":"We fit the model multiple times with different random seeds and with different samples from the training data. The results from all the different models are then combined by first ranking all predictions by predicted probability and then taking the mean rank as the final prediction.","metadata":{}},{"cell_type":"code","source":"def inference(X, X_test, iterations):\n    pred_list = []\n    for i in range(iterations):\n        X_train = X.sample(int(0.8*len(X)))\n        y_train = y.loc[X_train.index]\n\n        X_train, X_te = _impute(X_train, X_test)\n\n        #scale the data\n        X_train, X_te = _scale(X_train, X_te)\n\n        #encode categorical variables\n        X_train, X_te = _ohe(X_train, X_te)\n        \n        model = LogisticRegression(penalty='elasticnet', l1_ratio=0.8, C=0.007, tol = 1e-2, solver='saga', max_iter=1000, random_state=i)\n        # fit the model\n        model.fit(X_train,y_train)\n\n        #make predictions\n        y_pred = model.predict_proba(X_te)[:,1]\n        \n        pred_list.append(y_pred)\n    \n    pred_df = pd.DataFrame(pred_list).T\n    pred_df = pred_df.rank()\n    pred_df[\"mean\"] = pred_df.mean(axis=1)\n    \n    return pred_df","metadata":{"execution":{"iopub.status.busy":"2022-08-02T21:07:37.209639Z","iopub.execute_input":"2022-08-02T21:07:37.210055Z","iopub.status.idle":"2022-08-02T21:07:37.219425Z","shell.execute_reply.started":"2022-08-02T21:07:37.210024Z","shell.execute_reply":"2022-08-02T21:07:37.218343Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\npredictions_df = inference(X, X_test, 500)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:50:54.815710Z","iopub.execute_input":"2022-08-01T22:50:54.816357Z","iopub.status.idle":"2022-08-01T22:51:16.869320Z","shell.execute_reply.started":"2022-08-01T22:50:54.816309Z","shell.execute_reply":"2022-08-01T22:51:16.867669Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"predictions_df.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-01T22:50:28.626043Z","iopub.execute_input":"2022-08-01T22:50:28.626442Z","iopub.status.idle":"2022-08-01T22:50:28.658022Z","shell.execute_reply.started":"2022-08-01T22:50:28.626413Z","shell.execute_reply":"2022-08-01T22:50:28.657074Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"We don't have to submit predicted probabilities, we can also sumit ranks. The evaluation metric is ROC AUC. One way of interpreting AUC is: \n\n- **The probability that the model ranks a random positive example (failure == 1) more highly than a random negative example (failure == 0)**\n\nThe absolute values of the predicted probabilities does not matter, it does not matter how much higher the random positive example is than the random negative example, we are only interested in the rankings between them.\n\nIn other words the ROC AUC score is scale invariant, it only measures how well the predictions are ranked.\n\nTherefore we can use the predicted probability ranks rather than the predicted probabilities when calculating the ROC AUC score.\n\nExample:\n","metadata":{}},{"cell_type":"code","source":"pred_df = pd.DataFrame(y_pred_list, columns=[\"pred_prob\"])\npred_df[\"rank\"] = pred_df.rank()\ndisplay(pred_df.head(10))\n\nprint(\"roc auc using prediction probabilities:\", roc_auc_score(y_true_list, pred_df[\"pred_prob\"]))\nprint(\"roc auc using predicted probabilities ranks:\", roc_auc_score(y_true_list, pred_df[\"rank\"]))","metadata":{"execution":{"iopub.status.busy":"2022-08-01T23:25:15.792114Z","iopub.execute_input":"2022-08-01T23:25:15.792485Z","iopub.status.idle":"2022-08-01T23:25:15.829267Z","shell.execute_reply.started":"2022-08-01T23:25:15.792455Z","shell.execute_reply":"2022-08-01T23:25:15.828278Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For a singular model it does not make any difference whether you use ranks or predicted probabilities,\n\nIt may be very slightly better to use ranks rather than probabilities as it allows us to combine multiple sets of predictions together without bias towards one set of predictions. This is particularly true when enembling the predictions from very different models (e.g. logistic regression and gradient boosted decision trees).\n","metadata":{"execution":{"iopub.status.busy":"2022-08-01T23:25:36.088441Z","iopub.execute_input":"2022-08-01T23:25:36.088815Z","iopub.status.idle":"2022-08-01T23:25:36.096495Z","shell.execute_reply.started":"2022-08-01T23:25:36.088785Z","shell.execute_reply":"2022-08-01T23:25:36.095031Z"}}},{"cell_type":"markdown","source":"If you're still interested, I made this comment in one of my old [notebooks](https://www.kaggle.com/code/cabaxiom/tps-may-22-eda-lgbm-model/notebook) where the evaluation metric was also ROCAUC. I'll repeat it here, its a bit rambly and it probably wont make too much difference if you use ranks or predicted probabilities, so you might want to skip it :)","metadata":{}},{"cell_type":"markdown","source":"> @cabaxiom would you be willing to elaborate on this?\n>'It may be better to use ranks rather than probabilities as it allows us to combine multiple sets of predictions together without bias towards one set of predictions.'","metadata":{}},{"cell_type":"markdown","source":"\n\n>...\n>\n>When considering blending the predicted probabilities of a few very different models - say one model gives very confident predicted probabilities (close to 0 or 1) but another returns not so confident probabilities (further away from 0 or 1). Would this be equal to giving the predicted probabilities with more confidence a higher weighting than the other when blending? This sounds like a good thing, but more confident predictions doesn't mean more accurate predictions - especially between different models.\n>\n> How about if one model has a high spread of predicted probabilities (e.g. low density or the median difference between the full list of ordered probabilities is high) and another model's predicted probabilities are in several clusters each with very similar probabilities (e.g. several areas of high density, the median difference between the full list of ordered probabilities is low). Would the model with the higher spread of predicted probabilities have a higher weighting towards the overall ranking when blending?\n>\n> Additionally, not all models have a `predict_proba(X)` method. We could use the `decision_function(X)` method instead, which returns the confidence scores for each sample, which we can also rank. Blending these in their very different raw format e.g. taking the mean is going to cause some bias - it would be much better to take the mean of their rankings.\n>\n> On the other hand, are we loosing useful information when we decide to use rankings rather than probabilities? Do we care how similar the predicted probabilities are? (do we care if one instance is only just ranked higher than another?). When we use rankings, we effectively exaggerate the difference between two neighbouring ordered scores with very similar probabilities - this may be a good thing, as if the ROC AUC score is only focused on ranks it may make more sense to then take the mean of the ranks. But it may be a bad thing because by ranking we lose information about how similar the rankings are.\n>\n> ...","metadata":{}},{"cell_type":"markdown","source":"# Submission","metadata":{}},{"cell_type":"code","source":"sub_df = pd.read_csv(\"../input/tabular-playground-series-aug-2022/sample_submission.csv\")\nsub_df[\"failure\"] = predictions_df[\"mean\"]\nsub_df","metadata":{"execution":{"iopub.status.busy":"2022-08-02T21:07:21.080343Z","iopub.execute_input":"2022-08-02T21:07:21.080857Z","iopub.status.idle":"2022-08-02T21:07:21.126612Z","shell.execute_reply.started":"2022-08-02T21:07:21.080817Z","shell.execute_reply":"2022-08-02T21:07:21.125066Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"sub_df.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-01T21:58:26.357417Z","iopub.execute_input":"2022-08-01T21:58:26.357832Z","iopub.status.idle":"2022-08-01T21:58:26.398059Z","shell.execute_reply.started":"2022-08-01T21:58:26.357800Z","shell.execute_reply":"2022-08-01T21:58:26.397350Z"},"trusted":true},"execution_count":null,"outputs":[]}]}