{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"**Updated**\n\n# Introduction\n\nData drift occurs when there are differences in the distributions of training and test datasets. Maybe there are different classes or instances in the test set, or maybe the test set was collected later, and things have changed since then. Sometimes data drift can be difficult to detect. We saw this in the [February TPS competition](https://www.kaggle.com/code/ambrosm/tpsfeb22-01-eda-which-makes-sense#Comparing-train-and-test) where bacteria had mutated in the test dataset meaning that the same bacteria had slightly different variable distributions in the test set compared to the training set. This makes prediction difficult as the model is fitted on different data.\n\n[Santiago (@svpino)](https://twitter.com/svpino) tweeted this morning about [adversarial validation](https://twitter.com/svpino/status/1554436881734983680) ([unrolled thread](https://threadreaderapp.com/thread/1554436881734983680.html)). Basically, forget about predicting the target variable for now, and instead set up a model to tell if we can predict whether an observation belongs to the training or test dataset. Ideally, if the distribution of variables is the same between training and test set, we won't be able to build a good model to do this. If, however, the training and test sets are distributionally different, a model should be able to tell them apart.\n\nAs an example, I've applied this method to the August 2022 Tabular Playground Series competition: \n\n> The August 2022 edition of the Tabular Playground Series in an opportunity to help the fictional company Keep It Dry improve its main product Super Soaker. The product is used in factories to absorb spills and leaks.\n\n> The company has just completed a large testing study for different product prototypes. Can you use this data to build a model that predicts product failures?\n\nThis dataset is useful to explore the adversarial validation concept as the data has:\n\n- one variable that is mutually exclusive between training and test datasets,\n- a few categorical variables with some overlap, but some values occur exclusively in the training and test sets, and\n- variables with more subtle distributional differences\n\nLet's get started by reading in the data and having a look:","metadata":{}},{"cell_type":"code","source":"import numpy as np \nimport pandas as pd \nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nsample = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/sample_submission.csv')\ntrain = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/train.csv')\ntest = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/test.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:56:55.755011Z","iopub.execute_input":"2022-08-12T13:56:55.756020Z","iopub.status.idle":"2022-08-12T13:56:55.998539Z","shell.execute_reply.started":"2022-08-12T13:56:55.755971Z","shell.execute_reply":"2022-08-12T13:56:55.997079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.head()","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:56:56.001018Z","iopub.execute_input":"2022-08-12T13:56:56.001845Z","iopub.status.idle":"2022-08-12T13:56:56.035572Z","shell.execute_reply.started":"2022-08-12T13:56:56.001776Z","shell.execute_reply":"2022-08-12T13:56:56.034181Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train.info()\n\nNrows,Ncols = train.shape","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:56:56.038003Z","iopub.execute_input":"2022-08-12T13:56:56.038726Z","iopub.status.idle":"2022-08-12T13:56:56.063296Z","shell.execute_reply.started":"2022-08-12T13:56:56.038593Z","shell.execute_reply":"2022-08-12T13:56:56.061856Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data consists of some information about different products, we have:\n- `id`: dataset key, we can forget about this,\n- `product_code`, a categorical variable, \n- `loading`, a numerical variable, possibly some kind of force or weight measurement,\n- `attribute_0` and `attribute_1`, categorical variables detailing the type of construction material,\n- `attribute_2` and `attribute_3`, categorical (possibly ordinal) variables describing something unknown,\n- `measurement_0` to `measurement_17`, numerical (continuous) measurements of some kind, and\n- `failure`, the target variable, a binary variable indicating success or failure of the product. For adversarial validation purposes, we are not interested in this at all, so we can drop this variable.\n\nLet's take a look at histograms of the first six features for the training set and the test set:","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(24,6))\nfor i in range(6):\n    plt.subplot(1,6,i+1)\n    train[train.columns[i+1]].hist()\n    _=plt.title(f'Training set\\ndistribution of {train.columns[i+1]}')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:56:56.065609Z","iopub.execute_input":"2022-08-12T13:56:56.066090Z","iopub.status.idle":"2022-08-12T13:56:56.926024Z","shell.execute_reply.started":"2022-08-12T13:56:56.066051Z","shell.execute_reply":"2022-08-12T13:56:56.924777Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(24,6))\nfor i in range(6):\n    plt.subplot(1,6,i+1)\n    test[test.columns[i+1]].hist()\n    _=plt.title(f'Test set\\ndistribution of {test.columns[i+1]}')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:56:56.927373Z","iopub.execute_input":"2022-08-12T13:56:56.927746Z","iopub.status.idle":"2022-08-12T13:56:58.317726Z","shell.execute_reply.started":"2022-08-12T13:56:56.927714Z","shell.execute_reply":"2022-08-12T13:56:58.316394Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The first thing to note here is that `product_code` is mutually exclusive. The possible values in the training dataset are 'A'-'E', and the test set consists of 'F'-'I'. The training dataset consists of older products that have been tested for failure, and the test dataset consists solely of newer products that have not been tested for failure. The distribution of `loading` looks very similar between the two datasets. The `attribute_0` has the same values in the two datasets, but different proportions. The variables `attribute_1`, `attribute_2` and `attribute_3` have some common values between the training and test datasets, but some differences, e.g. `attribute_1` can take the values 'material_5', 'material_6' or 'material_8' in the training dataset, but 'material_5', 'material_6' or 'material_8' in the test dataset. \n\nThe numerical `measurement` variables look pretty similar between the training and test datasets:","metadata":{}},{"cell_type":"code","source":"_=train[[x for x in train.columns if 'measurement' in x]].hist(figsize=(12,12))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:56:58.320135Z","iopub.execute_input":"2022-08-12T13:56:58.320528Z","iopub.status.idle":"2022-08-12T13:57:00.496725Z","shell.execute_reply.started":"2022-08-12T13:56:58.320493Z","shell.execute_reply":"2022-08-12T13:57:00.495645Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"_=test[[x for x in test.columns if 'measurement' in x]].hist(figsize=(12,12))","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:57:00.498109Z","iopub.execute_input":"2022-08-12T13:57:00.499199Z","iopub.status.idle":"2022-08-12T13:57:02.567678Z","shell.execute_reply.started":"2022-08-12T13:57:00.499031Z","shell.execute_reply":"2022-08-12T13:57:02.566202Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"If you look carefully, you can see a few differences in the distributions, but overall they are very similar. Using adversarial validation we can tease out the differences automatically.\n\n# Data preprocessing pipeline","metadata":{}},{"cell_type":"code","source":"# These modifications to the preprocessing steps allow for variable names to \n# be preserved, which is useful/needed when plotting variable importance\n\nfrom sklearn.impute import SimpleImputer\nfrom sklearn.preprocessing import OrdinalEncoder\nfrom sklearn.compose import ColumnTransformer\nclass SimpleImputerNamed(SimpleImputer):\n    def get_feature_names_out(self):\n        return list(self.feature_names_in_)\nclass OrdinalEncoderNamed(OrdinalEncoder):\n    def get_feature_names_out(self):\n        return list(self.feature_names_in_)\nclass ColumnTransformerNamed(ColumnTransformer):\n    def get_feature_names_out(self):\n        names = []\n        for transformer in self.transformers_:\n            if transformer[0] == 'remainder':\n                if transformer[1] == 'passthrough':\n                    names += list(self.feature_names_in_[transformer[2]])\n                break\n            else:\n                names += transformer[1].get_feature_names_out()\n        return names\n    def fit(self, X, y=None):\n        #print('In fit method')\n        return super().fit(X,y)\n    def transform(self, X):\n        #print('In transform method')\n        transformed = super().transform(X)\n        return pd.DataFrame(transformed, columns= self.get_feature_names_out())\n    def fit_transform(self, X, y=None):\n        #print('In fit_transform method')\n        fit_transformed = super().fit_transform(X,y)\n        return pd.DataFrame(fit_transformed, columns=self.get_feature_names_out())","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-12T13:57:02.569786Z","iopub.execute_input":"2022-08-12T13:57:02.570634Z","iopub.status.idle":"2022-08-12T13:57:02.585668Z","shell.execute_reply.started":"2022-08-12T13:57:02.570579Z","shell.execute_reply":"2022-08-12T13:57:02.584586Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"The data has a [few missing values](https://www.kaggle.com/code/nnjjpp/eda-and-baseline-model-tps-aug-2022), we can impute these with the median. We also need to encode the categorical variables (`product_code`, `attribute_0` and `attribute_1`). These can be encoded with `OrdinalEncoder`. The `attribute` variables are not necessarily ordinal variables, however the XGBoost model we're using (and any tree-based model) will be able to distinguish between different values encoded in this way. If we use a regression-type model, we may have to encode these variables differently (maybe with `OneHotEncoder`).","metadata":{}},{"cell_type":"code","source":"combined_data = pd.concat([train.drop(['id','failure'], axis=1),\n                           test.drop(['id'], axis=1)])\ncombined_data_original = combined_data.copy()\n\nmissing_cols = combined_data.columns[list(combined_data.isna().sum()>0)]\nmissing_cols_original = missing_cols.copy()\n\ncategorical_cols = ['product_code', 'attribute_0', 'attribute_1']\ncategorical_cols_original = categorical_cols.copy()\n\npreprocessing = ColumnTransformerNamed([('median_infill', SimpleImputerNamed(strategy='median'), missing_cols),\n                                        ('ordinal_encode', OrdinalEncoderNamed(), categorical_cols)],\n                                       remainder='passthrough')  \n\n# Test the preprocessing pipeline/transformer:\npreprocessing.fit_transform(combined_data)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-12T13:57:02.587612Z","iopub.execute_input":"2022-08-12T13:57:02.588655Z","iopub.status.idle":"2022-08-12T13:57:02.913781Z","shell.execute_reply.started":"2022-08-12T13:57:02.588580Z","shell.execute_reply":"2022-08-12T13:57:02.912432Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Adversarial validation\n\nThe next steps are to set up a new target variable, which is simply whether or not each row is from the training or the test datset, and to choose a model. Here I use an XGBRegressor with a logistic objective function.","metadata":{}},{"cell_type":"code","source":"in_test_dataset = pd.DataFrame({'test': [0,]*train.shape[0] + [1,]*test.shape[0]})\n\nfrom xgboost import XGBRegressor\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.metrics import roc_auc_score\n\nmod_xgb = Pipeline(steps = (['preprocessing', preprocessing],\n                            ['xgboost', XGBRegressor(objective = 'binary:logistic',\n                                                     eval_metric = 'auc',\n                                                     random_state = 200)]))","metadata":{"execution":{"iopub.status.busy":"2022-08-12T13:57:02.915696Z","iopub.execute_input":"2022-08-12T13:57:02.917105Z","iopub.status.idle":"2022-08-12T13:57:02.939357Z","shell.execute_reply.started":"2022-08-12T13:57:02.917045Z","shell.execute_reply":"2022-08-12T13:57:02.938005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"To assess how different the training and test dataset are, we use the area under the ROC curve. The minimal value of this metric is 0.5, which means that there is no difference between training and test sets. The maximal value is 1.0, which means we can tell 100% whether a row is from the training or the test set based on the variable combinations that the model has chosen.","metadata":{}},{"cell_type":"code","source":"def fit_adversarial_model(X, y, estimator_pipeline, plot=True):\n    from sklearn.model_selection import cross_validate, KFold\n    from sklearn.metrics import make_scorer, roc_auc_score\n    cv_results = cross_validate(estimator = estimator_pipeline,\n                                X = X, \n                                y = y, \n                                scoring = make_scorer(roc_auc_score),\n                                cv= KFold(n_splits=5, random_state = 123, shuffle=True))\n    mean_cv_score = cv_results['test_score'].mean()\n    estimator_pipeline.fit(X, y)\n    print(f'Area under ROC curve (cross-validated): {mean_cv_score:.2f}')\n\n    print('\\nVariable        Importance')\n    print('--------------------------')\n    try:\n        column_heights = estimator_pipeline._final_estimator.feature_importances_\n    except AttributeError: # use absolute value of coefficients if model doesn't support feature importances\n        column_heights = np.abs(estimator_pipeline.steps[-1][1].coef_[0])\n    for x in zip(estimator_pipeline.named_steps['preprocessing'].get_feature_names_out() ,\n                 column_heights):\n        print(f'{x[0]:<15} {x[1]:.2f}')\n        \n    if plot:\n        \n        x1,y1 = zip(*filter(lambda z:z[1] > 0, \n                     zip(estimator_pipeline.named_steps['preprocessing'].get_feature_names_out() ,\n                         column_heights)))\n        sns.barplot(x=list(x1), y=np.array(y1))\n        plt.gca().tick_params(axis='x', rotation=90)\n\nfit_adversarial_model(combined_data, in_test_dataset, estimator_pipeline = mod_xgb, plot=False)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-12T13:57:02.943762Z","iopub.execute_input":"2022-08-12T13:57:02.944296Z","iopub.status.idle":"2022-08-12T13:57:11.470738Z","shell.execute_reply.started":"2022-08-12T13:57:02.944253Z","shell.execute_reply":"2022-08-12T13:57:11.469551Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Since there are no common values of `product_code` in the training and test data (i.e. this variable is mutually exclusive), the adversarial validation model discovered that we can just separate the different classes of `product_code` to tell the training and test datasets apart. Let's drop this variable and see what other variables the model can come up with:","metadata":{}},{"cell_type":"code","source":"try:\n    categorical_cols.remove('product_code')\nexcept ValueError:\n    pass\n\ncombined_data_no_product_code = combined_data.drop(['product_code'], axis=1)\n\nfit_adversarial_model(combined_data_no_product_code, in_test_dataset, estimator_pipeline = mod_xgb)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T13:57:11.475376Z","iopub.execute_input":"2022-08-12T13:57:11.477172Z","iopub.status.idle":"2022-08-12T13:57:26.509258Z","shell.execute_reply.started":"2022-08-12T13:57:11.477117Z","shell.execute_reply":"2022-08-12T13:57:26.507663Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"This is better and more informative. The model has identified that there are different values of the three attribute variables in the training and test set. However, we saw this already from the exploratory data analysis. Let's drop `attribute_1`, `attribute_2` and `attribute_3`. ","metadata":{}},{"cell_type":"code","source":"try:\n    categorical_cols.remove('attribute_1')\nexcept ValueError:\n    pass\ncombined_data2 = combined_data_no_product_code.drop(['attribute_1', 'attribute_2', 'attribute_3'], axis=1)\nfit_adversarial_model(combined_data2, in_test_dataset, estimator_pipeline = mod_xgb)","metadata":{"execution":{"iopub.status.busy":"2022-08-12T13:57:26.510971Z","iopub.execute_input":"2022-08-12T13:57:26.511799Z","iopub.status.idle":"2022-08-12T13:58:01.968262Z","shell.execute_reply.started":"2022-08-12T13:57:26.511760Z","shell.execute_reply":"2022-08-12T13:58:01.966983Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Now the model is picking up the differences in distribution among the values of `attribute_0`, and also differences in the distributions of the `measurement` variables (primarily `measurement_0`).\n\nNote that XGBoost is self-regularising to some extent. As such, it stops after identifying really good splits in the training and test set (like the initial split on `product_code`). To explore the data in more detail with XGBoost we needed to remove these variables. A different model could allow us to identify all variable importances in one go. To demonstrate this, let's carry out the adversarial validation procedure using a logistic regresion. This is regularised by default, but by choosing the `C` parameter large enough it is effectively unregularised, and will estimate coefficients for all variables.","metadata":{}},{"cell_type":"code","source":"from sklearn.linear_model import LogisticRegression\nfrom sklearn.preprocessing import StandardScaler\n\n# Set up the pipeline again as we changed the input data (categorical_cols, etc.)\ncombined_data = combined_data_original.copy()\ncategorical_cols = categorical_cols_original.copy()\nmissing_cols = missing_cols_original.copy()\npreprocessing = ColumnTransformerNamed([('median_infill', SimpleImputerNamed(strategy='median'), missing_cols),\n                                        ('ordinal_encode', OrdinalEncoderNamed(), categorical_cols)],\n                                       remainder='passthrough')  \nmod_logreg = Pipeline(steps = (['preprocessing', preprocessing],\n                               ['scaling', StandardScaler()],\n                               ['logistic_regression', LogisticRegression(penalty='l2',\n                                                                          solver='saga',\n                                                                          class_weight='balanced',\n                                                                          max_iter=2000,\n                                                                          C=10**6,\n                                                                          random_state = 200)]))\n\nfit_adversarial_model(combined_data, np.array(in_test_dataset).ravel(), estimator_pipeline = mod_logreg, plot=True)\n","metadata":{"execution":{"iopub.status.busy":"2022-08-12T14:01:54.519470Z","iopub.execute_input":"2022-08-12T14:01:54.520621Z","iopub.status.idle":"2022-08-12T14:04:27.348610Z","shell.execute_reply.started":"2022-08-12T14:01:54.520571Z","shell.execute_reply":"2022-08-12T14:04:27.347079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Although logistic regression doesn't give us feature importances as XGBoost does, we can use the absolute value of the coefficients after scaling to provide similar information. This looks very similar to the previous analysis with XGBoost, although we got the information in one go.","metadata":{}},{"cell_type":"markdown","source":"# Conclusion\n\nWe can use adversarial validation to identify differences between the distribution of variables between the training and test set. This can help to identify distributional drift between training and test dataset, which will provide information on how to improve our model.\n\nIs data drift a problem? Possibly, and probably yes. Consider the following scenarios:\n1. Values of the predictors have changed, but the relationship with the target, and the model, cover this. For example, we have built a regression model that relates height and weight, and we apply the model to a different population that are taller than the population we trained the model on. The relationship (model) is robust to extrapolation and so provides sensible estimates for data outside of the range of what we have already seen.\n2. Values of the predictors have changed, but the model can't extrapolate. For example, maybe we overfit with a 12-degree polynomial on the height-weight relationship and extrapolations go negative or absurdly large for height measurements outside of what we have seen in the training set.\n3. The relationship between target and predictors has changed. In the [February TPS competition](https://www.kaggle.com/code/ambrosm/tpsfeb22-01-eda-which-makes-sense#Comparing-train-and-test). The genomes of the different bacteria had apparently mutated between training and test dataset, whereas the target (bacteria species) remained the same. Once this kind of drift is identified (which is not necessarily easy to identify!), we can either translate the predictors to encompass the known change, or hack the model to deal with the differences between the training and test dataset. \n4. We have completely new values for the predictor variables. Suppose we have trained a ML model that can distinguish between citrus fruits to a high degree of accuracy, and we then feed it images of bananas. This is essentially what has happened with `product_code`, the values of which are mutually exclusive between the training and test sets. As such, this variable is almost useless for prediction, but is useful for [grouping the training data for cross validation](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/341070).\n\nApart from the obvious differences in `product_code`, we saw differences in all the `attribute` features, and possibly some of the `measurement` features. [Each product takes only one value for each attribute](https://www.kaggle.com/code/nnjjpp/product-composition-tps-aug-2022), and several of the values of the attributes are unseen in the training data. For example, `attribute_2` and `attribute_3` each take value 7 in the test set, but 7 does not appear as a value for these attributes in the training set. Likewise, material_7 occurs as a value of `attribute_1` in the test set but not in the training set, although it does appear in both training and test set as values of `attribute_0`. Have we built models that are sufficiently robust to extrapolate to new combinations of attributes in the test set?\n","metadata":{}},{"cell_type":"markdown","source":"# Further reading\n\nhttps://www.kaggle.com/code/carlmcbrideellis/what-is-adversarial-validation/notebook\n\nhttps://www.kdnuggets.com/2020/02/adversarial-validation-overview.html\n\nhttps://towardsdatascience.com/drift-in-machine-learning-e49df46803a\n\nhttps://medium.com/data-from-the-trenches/a-primer-on-data-drift-18789ef252a6: this describes different types of drift using a more mathematical framework.\n\nSome conversation has arisen around this [in the discussion](https://www.kaggle.com/competitions/tabular-playground-series-aug-2022/discussion/342168).","metadata":{}}]}