{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"import numpy as np\nimport pandas as pd\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\ntrain = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/train.csv')\ntest = pd.read_csv('/kaggle/input/tabular-playground-series-aug-2022/test.csv')","metadata":{"_kg_hide-input":true,"execution":{"iopub.status.busy":"2022-08-13T07:55:41.967624Z","iopub.execute_input":"2022-08-13T07:55:41.968043Z","iopub.status.idle":"2022-08-13T07:55:42.148450Z","shell.execute_reply.started":"2022-08-13T07:55:41.968009Z","shell.execute_reply":"2022-08-13T07:55:42.147126Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Supplementary EDA for Tabular Playground Series August 2022.\n\nThis notebook is just a few plots and a table to help identify differences between products. Hope it is useful!\n\nWe have seen that the training and test sets are split mutually exclusively along different product lines. The test set is a series of new products that haven't been tested for failure, and we need to use the failure probabilities of the products in the training set.","metadata":{}},{"cell_type":"code","source":"print(f\"Training set consists of product codes: {np.unique(train['product_code'])}\")\nprint(f\"Test set consists of product codes: {np.unique(test['product_code'])}\")","metadata":{"execution":{"iopub.status.busy":"2022-08-13T07:55:42.151008Z","iopub.execute_input":"2022-08-13T07:55:42.151723Z","iopub.status.idle":"2022-08-13T07:55:42.184464Z","shell.execute_reply.started":"2022-08-13T07:55:42.151687Z","shell.execute_reply":"2022-08-13T07:55:42.182794Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Each product has only one unique value for each of the four attributes. Using a little panda magic we can see which attributes correspond to which products:","metadata":{}},{"cell_type":"code","source":"test['dataset'] = 'test'\ntrain['dataset'] = 'train'\ntest['failure'] = 0\nboth = pd.concat([train, test])\nx=both.groupby('product_code').agg({var:'value_counts' for var in [colname for colname in train.columns if 'attribute' in colname] + ['dataset']}, sort=False).fillna(0)\nnz = pd.melt(x.reset_index(), id_vars=['product_code', 'level_1'])\nnz2 = nz.loc[nz['value']>0,:]\nnz3 = pd.pivot(nz2,index='product_code',columns='variable',values='level_1')\nnz4 = nz3.reset_index()\nnz4.index = pd.MultiIndex.from_frame(nz4[['dataset','product_code']])\nnz5 = nz4.drop(['product_code', 'dataset'], axis=1)\nnz5","metadata":{"execution":{"iopub.status.busy":"2022-08-13T07:56:02.926912Z","iopub.execute_input":"2022-08-13T07:56:02.927382Z","iopub.status.idle":"2022-08-13T07:56:02.995491Z","shell.execute_reply.started":"2022-08-13T07:56:02.927345Z","shell.execute_reply":"2022-08-13T07:56:02.994588Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distribution of loading stratified by product codes and failure\n\nThe rest of the variables are numerical (a few are integers but most are continuous). The next plots are distributions of the loading and measurement variables for each product/failure combinations. As we don't know what `failure` is for the test set, I have set the value of `failure` equal to 0 for these plots.","metadata":{}},{"cell_type":"code","source":"import seaborn as sns\nplt.figure(figsize=(12,6))\n_=sns.violinplot(data=both, x=\"product_code\", y='loading',\n                   hue='failure', \n                   split=True, inner=\"quart\", linewidth=1)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T07:59:48.544746Z","iopub.execute_input":"2022-08-13T07:59:48.545267Z","iopub.status.idle":"2022-08-13T07:59:49.069623Z","shell.execute_reply.started":"2022-08-13T07:59:48.545213Z","shell.execute_reply":"2022-08-13T07:59:49.067780Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Distribution of measurements stratified by product codes and failure","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(18,12))\nfor i in range(18):\n    ax=plt.subplot(6,3,i+1)\n    sns.violinplot(data=both, x=\"product_code\", y=f\"measurement_{i}\",\n                   hue='failure', \n                   split=True, inner=\"quart\", linewidth=1,ax=ax)\n    if i < 17:\n        ax.legend([],[], frameon=False)","metadata":{"execution":{"iopub.status.busy":"2022-08-13T07:55:42.297478Z","iopub.execute_input":"2022-08-13T07:55:42.298335Z","iopub.status.idle":"2022-08-13T07:55:49.306274Z","shell.execute_reply.started":"2022-08-13T07:55:42.298287Z","shell.execute_reply":"2022-08-13T07:55:49.304930Z"},"trusted":true},"execution_count":null,"outputs":[]}]}