{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-02T01:46:01.066020Z","iopub.execute_input":"2022-08-02T01:46:01.066852Z","iopub.status.idle":"2022-08-02T01:46:01.598146Z","shell.execute_reply.started":"2022-08-02T01:46:01.066787Z","shell.execute_reply":"2022-08-02T01:46:01.597268Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Introduction\n\nThis month's Tabular Playground Series is a binary classification problem:\n\n> The August 2022 edition of the Tabular Playground Series in an opportunity to help the fictional company Keep It Dry improve its main product Super Soaker. The product is used in factories to absorb spills and leaks.\n> \n> The company has just completed a large testing study for different product prototypes. Can you use this data to build a model that predicts product failures?\n> \n\nLet's first read in the data, look at the first five rows and get a list of the column data formats:","metadata":{}},{"cell_type":"code","source":"sample = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/sample_submission.csv')\ntrain = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/train.csv')\ntest = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/test.csv')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:01.604342Z","iopub.execute_input":"2022-08-02T01:46:01.604689Z","iopub.status.idle":"2022-08-02T01:46:01.788104Z","shell.execute_reply.started":"2022-08-02T01:46:01.604657Z","shell.execute_reply":"2022-08-02T01:46:01.786937Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:01.789954Z","iopub.execute_input":"2022-08-02T01:46:01.790286Z","iopub.status.idle":"2022-08-02T01:46:01.822566Z","shell.execute_reply.started":"2022-08-02T01:46:01.790256Z","shell.execute_reply":"2022-08-02T01:46:01.821465Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()\n\nNrows,Ncols = train.shape","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:01.826509Z","iopub.execute_input":"2022-08-02T01:46:01.826871Z","iopub.status.idle":"2022-08-02T01:46:01.848285Z","shell.execute_reply.started":"2022-08-02T01:46:01.826840Z","shell.execute_reply":"2022-08-02T01:46:01.847441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data consists of some information about different products, we have:\n- `id`: dataset key, unlikely to have any predictive power,\n- `product_code`, a categorical variable, \n- `loading`, a numerical variable, possibly some kind of force or weight measurement,\n- `attribute_0` and `attribute_1`, categorical variables detailing the type of construction material,\n- `attribute_2` and `attribute_3`, categorical (possibly ordinal) variables describing something unknown,\n- `measurement_0` to `measurement_17`, numerical (continuous) measurements of some kind, and\n- `failure`, what we'd like to predict, a binary variable indicating success or failure of the product.\n\nLet's take a look at histograms of the first six features for the training set and the test set:","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize=(24,6))\nfor i in range(6):\n    plt.subplot(1,6,i+1)\n    train[train.columns[i+1]].hist()\n    _=plt.title(f'Training set\\ndistribution of {train.columns[i+1]}')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:01.849706Z","iopub.execute_input":"2022-08-02T01:46:01.850051Z","iopub.status.idle":"2022-08-02T01:46:02.655042Z","shell.execute_reply.started":"2022-08-02T01:46:01.850019Z","shell.execute_reply":"2022-08-02T01:46:02.653320Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The `product_code` feature seems more or less uniformly distributed amongst A, B, C, D, and E. The `loading` feature is a numerical feature with a slight positive skew, but we could possibly get away with assuming this is normally distributed. The `attribute_i` features are categorical variables (possibly ordinal variables), with unbalanced (non-uniform) distributions suggesting that we may need to consider this when splitting the data for cross validation, and is also something to keep in mind when we get to setting up and choosing the model.\n\n","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(24,6))\nfor i in range(6):\n    plt.subplot(1,6,i+1)\n    test[test.columns[i+1]].hist()\n    _=plt.title(f'Test set\\ndistribution of {test.columns[i+1]}')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:02.656593Z","iopub.execute_input":"2022-08-02T01:46:02.657258Z","iopub.status.idle":"2022-08-02T01:46:03.430101Z","shell.execute_reply.started":"2022-08-02T01:46:02.657213Z","shell.execute_reply":"2022-08-02T01:46:03.428825Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The test set consists of new products as evidenced by the values of the `product_code` feature in the test set. For prediction, we can initially drop the product code from the training and test set, although there might be some hidden information in this feature.\n\n`loading` seems to follow a similar distribution in the training and test set. There are different values of `attribute_1`, `attribute_2` and `attribute_3` in the training and test set - we need to consider this in setting up categorical variable encoding for these features (see below).","metadata":{}},{"cell_type":"markdown","source":"# Missing data analysis\n\nLet's turn to missing data. I recently discovered the [missingno](https://github.com/ResidentMario/missingno) package which draws missing data figures very easily. It is already installed in the kaggle environment so there is no need to `pip install` this package, just import it.","metadata":{}},{"cell_type":"code","source":"import missingno as msno\n\nmsno.matrix(train.iloc[np.random.choice(range(Nrows), 250)])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:03.431907Z","iopub.execute_input":"2022-08-02T01:46:03.432407Z","iopub.status.idle":"2022-08-02T01:46:04.107467Z","shell.execute_reply.started":"2022-08-02T01:46:03.432361Z","shell.execute_reply":"2022-08-02T01:46:04.106548Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"It is clear that missing data needs to be dealt with here. Some of the `loading` features are missing, as well as some of the `measurement_i` variables. Next, we'll print out a list of the number of missing items per feature in both the training set and the test set:","metadata":{}},{"cell_type":"code","source":"train.drop(['failure'],axis=1).isna().sum().to_frame().rename(columns={0:'Training set'}).join(test.isna().sum().rename('Test set'))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:04.108785Z","iopub.execute_input":"2022-08-02T01:46:04.109734Z","iopub.status.idle":"2022-08-02T01:46:04.139806Z","shell.execute_reply.started":"2022-08-02T01:46:04.109686Z","shell.execute_reply":"2022-08-02T01:46:04.138725Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The missingness seems similar between both the training and the test set. The `product_code` and the `attribute_i` features are all completely non-missing. `measurement_i`, for $i \\leq 2$ are all non-missing. `measurement_i` for $i \\geq 3$ have some missing values and have progressively increasing proportion of missingness. This suggests that there are a series of measurements, the results of which determine whether or not later tests are carried out. The probabilities of later measurements being missing doesn't appear to be too strongly related to the missingness of earlier measurements, as inferred from `msno.heatmap` (not shown here), however the dendrogram analysis from msno shows a definite pattern:","metadata":{}},{"cell_type":"code","source":"msno.dendrogram(train)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:04.140989Z","iopub.execute_input":"2022-08-02T01:46:04.141323Z","iopub.status.idle":"2022-08-02T01:46:04.671228Z","shell.execute_reply.started":"2022-08-02T01:46:04.141296Z","shell.execute_reply":"2022-08-02T01:46:04.670350Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The interpretation of this figure (from the [missingno readme file](https://github.com/ResidentMario/missingno#readme) is that variables that are more likely to be missing together have join closer to zero on the y-axis. Thus there is a definite relationship between the missingness of `measurement_i` variables as $i$ increases. Missingness might also be related to the value (and not just the presence or absence) of previous measurements. We should probably treat infilling of the measurements carefully and try to understand why measeurements are missing. ","metadata":{}},{"cell_type":"markdown","source":"# Distribution of the measurement variables","metadata":{}},{"cell_type":"markdown","source":"The measurement variables are all nicely normally distributed. It looks like some variables are integers, particularly `measurement_0`, `measurement_1` and `measurement_2`, with the rest continuous variables.","metadata":{}},{"cell_type":"code","source":"_=train[[x for x in train.columns if 'measurement' in x]].hist(figsize=(20,20))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:04.672358Z","iopub.execute_input":"2022-08-02T01:46:04.673347Z","iopub.status.idle":"2022-08-02T01:46:07.112881Z","shell.execute_reply.started":"2022-08-02T01:46:04.673313Z","shell.execute_reply":"2022-08-02T01:46:07.111442Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"for col in [x for x in train.columns if 'measurement' in x]:\n    print(f'Variable {col}: {len(train[col].value_counts())} unique values')\n","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:07.114646Z","iopub.execute_input":"2022-08-02T01:46:07.115195Z","iopub.status.idle":"2022-08-02T01:46:07.149476Z","shell.execute_reply.started":"2022-08-02T01:46:07.115159Z","shell.execute_reply":"2022-08-02T01:46:07.148276Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distribution of target variable (success/failure)\n\nProbability of failure is around 1/5 based on the distribution of the `failure` features in the test set. As with the unbalanced nature of the `attribute` features, we may need to take this into consideration in later modelling steps, particularly when designing any cross-validation schemes.","metadata":{}},{"cell_type":"markdown","source":"","metadata":{}},{"cell_type":"code","source":"train['failure'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:07.151335Z","iopub.execute_input":"2022-08-02T01:46:07.151831Z","iopub.status.idle":"2022-08-02T01:46:07.164045Z","shell.execute_reply.started":"2022-08-02T01:46:07.151789Z","shell.execute_reply":"2022-08-02T01:46:07.163064Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Simple data cleaning pipeline\n\nThe basic steps in data cleaning here are to infill the missing data somehow, and encoding of the categorical variables (`attribute_0` and `attribute_1`). Although we saw that missingness in the measurements is probably related to values and/or missingness in ealier measurements, and possibly other variables. However, for now, let's use a simple imputation method and infill with the median for the measurement variables. \n\nWe can drop `id` as it almost certainly contains no predictive information. `product_code` is not useful at this stage either as there are different products in the training and test dataset. When we get to setting up a cross validation scheme, possibly for a hyperparameter search, it will be good to include the `product_code` variable as a [grouping variable](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/341070). There may also be relationships between the product type and other variables or measurements, which would be good to explore further.\n\nInitially, encode the remaining categorical (string) variables (namely `attribute_0` and `attribute_1`) with `OrdinalEncoder`. Since there are different values of `attribute_1` in the training and test set (see above), we need to include a way to deal with this in the data processing pipeline. Specifically, we set `handle_unknown` parameter to 'use_encoded_value', otherwise the pipeline will fail when applied to the test set. In the first instance, we don't have to apply any preprocessing to `attribute_2` and `attribute_3` as these are numerical (possibly ordinal) variables that will be compatible with a baseline model.","metadata":{}},{"cell_type":"code","source":"missing_cols = list(train.columns[train.isna().sum()>0])\ncategorical_cols = ['attribute_0', 'attribute_1']\nfrom sklearn.compose import ColumnTransformer\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import OrdinalEncoder\npreprocessing = ColumnTransformer([('median_infill', SimpleImputer(strategy='median'), missing_cols),\n                                   ('ordinal_encode', OrdinalEncoder(handle_unknown='use_encoded_value', unknown_value = -1), categorical_cols)],\n                                  remainder='passthrough')\n\n# Test the preprocessing pipeline/transformer:\npreprocessing.fit_transform(train.drop(['id','product_code','failure'], axis=1))\npreprocessing.transform(test.drop(['id','product_code'], axis=1))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:07.167544Z","iopub.execute_input":"2022-08-02T01:46:07.168336Z","iopub.status.idle":"2022-08-02T01:46:07.394726Z","shell.execute_reply.started":"2022-08-02T01:46:07.168300Z","shell.execute_reply":"2022-08-02T01:46:07.393685Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Baseline model - XGBoost regressor\n\nLet's fit a model to the end of the data preprocessing pipeline, and make some predictions as a baseline. XGBoost is a good first start. Since it is a binary classification problem, we can use `XGBRegressor` with binary-logistic objective function. Also, since the [competition evaluation metric](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/overview/evaluation) is area under the ROC curve, set the `eval_metric` to 'auc'. I manually tuned a few of the hyper-parameters for the baseline. Note that XGBoost seems to work best for this data with very shallow trees (max_depth = 2). My initial observations are that deeper trees overfit the training data quite badly. \n","metadata":{}},{"cell_type":"code","source":"from xgboost import XGBRegressor\nfrom sklearn.pipeline import Pipeline\nmodelling_pipeline = Pipeline(steps = (['preprocessing', preprocessing],\n                                       ['xgboost', XGBRegressor(n_estimators = 350,\n                                                                objective = 'binary:logistic',\n                                                                eval_metric = 'auc',\n                                                                eta = 0.2,\n                                                                max_depth = 2,\n                                                                gamma = 1.2,\n                                                                random_state = 200)]))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:07.396277Z","iopub.execute_input":"2022-08-02T01:46:07.396827Z","iopub.status.idle":"2022-08-02T01:46:07.424775Z","shell.execute_reply.started":"2022-08-02T01:46:07.396782Z","shell.execute_reply":"2022-08-02T01:46:07.423582Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To get an estimate of the evaluation score we can expect, split the training data into training/test datasets. For now, let's use `train_test_split`, although in future we may want to use a more sophisticated cross-validation scheme to stratify and group different features, as discussed above.","metadata":{}},{"cell_type":"code","source":"from sklearn.model_selection import train_test_split\n\nXtr, Xte, ytr, yte = train_test_split(train.drop(['id','product_code','failure'], axis=1), train['failure'],\n                                      test_size=0.2, random_state = 123)\nfrom sklearn.metrics import roc_auc_score # Evaluation metric\nmodelling_pipeline.fit(Xtr, ytr)\nprint(f'Estimated score: {roc_auc_score(yte, modelling_pipeline.predict(Xte)):0.3f}')","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:07.425943Z","iopub.execute_input":"2022-08-02T01:46:07.426280Z","iopub.status.idle":"2022-08-02T01:46:11.867910Z","shell.execute_reply.started":"2022-08-02T01:46:07.426249Z","shell.execute_reply":"2022-08-02T01:46:11.866850Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"For submission purposes, use the entire training data set to make predictions:","metadata":{}},{"cell_type":"code","source":"modelling_pipeline.fit(train.drop(['id','product_code','failure'], axis=1), train['failure'])\n\npredictions = modelling_pipeline.predict(test.drop(['id','product_code'], axis=1))\n\nsubmission = pd.DataFrame({'id': test['id'],\n                           'failure': predictions})\n\nsubmission.to_csv('submission.csv', index=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T01:46:11.871949Z","iopub.execute_input":"2022-08-02T01:46:11.872746Z","iopub.status.idle":"2022-08-02T01:46:16.974077Z","shell.execute_reply.started":"2022-08-02T01:46:11.872702Z","shell.execute_reply":"2022-08-02T01:46:16.973198Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Submit to competition\n\nThanks for reading, comments are welcome. 👍","metadata":{}}]}