{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Knowing the data before training and testing\n\n## Summary\n\nIt is important to understand the data that we will be using to make any kind of training and predictions.\n\nWe look at the both training and testing data dataframe, then find the missing values in all the columns, make observations, and take decision regarding how to deal with missing values if any.\n\nWe make a separate list of categorical and continuous variables so that we can perform any operations or use them in future\n\nWe plot the disributions of continuous variables and make observation and make assumption and decisions\n\nWe also plot the Frequencies of categorical variables, make observation and make assumptions and decisions\n\n**Work in progress.**\n\nUpvote!\n\nPending\n- finding the correlation between different variables\n- finding correlation between variables and target variable\n- taking decisions regarding onehot encoding other feature preprocessing.\n- Baseline Logistic regression model\n    - finding the feature importance. relative importance\n- Baseline Neural Network\n- Ensemble methods of Machine learning\n- Regularized Fully connected Neural network","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\ndirs = []\n\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        dirs.append(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session\n\nfor path in dirs:\n    print(path)\n    \n# setting to view all columns\npd.set_option('display.max_columns', None)","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","execution":{"iopub.status.busy":"2022-08-02T17:57:45.804862Z","iopub.execute_input":"2022-08-02T17:57:45.805395Z","iopub.status.idle":"2022-08-02T17:57:45.846125Z","shell.execute_reply.started":"2022-08-02T17:57:45.805290Z","shell.execute_reply":"2022-08-02T17:57:45.844182Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# Importing Required Libraries","metadata":{}},{"cell_type":"code","source":"import plotly.express as px\nfrom plotly.subplots import make_subplots\nimport plotly.graph_objects as go\nimport matplotlib.pyplot as plt\nimport seaborn as sbn\n# import tensorflow as tf\n# import keras","metadata":{"execution":{"iopub.status.busy":"2022-08-02T17:57:45.911568Z","iopub.execute_input":"2022-08-02T17:57:45.913023Z","iopub.status.idle":"2022-08-02T17:57:54.678723Z","shell.execute_reply.started":"2022-08-02T17:57:45.912934Z","shell.execute_reply":"2022-08-02T17:57:54.677336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# reading the csv files\n\ntrain_data = pd.read_csv(dirs[1])\ntest_data = pd.read_csv(dirs[2])\nsample_sub = pd.read_csv(dirs[0])","metadata":{"execution":{"iopub.status.busy":"2022-08-02T17:57:54.685537Z","iopub.execute_input":"2022-08-02T17:57:54.687939Z","iopub.status.idle":"2022-08-02T17:57:55.019994Z","shell.execute_reply.started":"2022-08-02T17:57:54.687879Z","shell.execute_reply":"2022-08-02T17:57:55.018869Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# looking at training_data DF\ntrain_data.head(10)","metadata":{"execution":{"iopub.status.busy":"2022-08-02T17:57:55.024126Z","iopub.execute_input":"2022-08-02T17:57:55.024548Z","iopub.status.idle":"2022-08-02T17:57:55.071137Z","shell.execute_reply.started":"2022-08-02T17:57:55.024511Z","shell.execute_reply":"2022-08-02T17:57:55.070222Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data_cols = train_data.columns\ntest_data_cols = test_data.columns\n\n# displaying the name of columns in train_data\nprint(\"Columns in train_data\")\nprint(train_data.columns, '\\n')\n\n# displaying the name of columns in test_data\nprint(\"Columns in test_data\")\nprint(test_data.columns, '\\n')\n\n# displaying which columns are not present in test_data which ARE PRESENT in train_data\nprint(\"Name of Columns of train_data not in test_data\")\nfor col_name in list(train_data.columns):\n    if col_name not in list(test_data_cols):\n        print(f\"'{col_name}'\")","metadata":{"execution":{"iopub.status.busy":"2022-08-02T17:57:55.073465Z","iopub.execute_input":"2022-08-02T17:57:55.073845Z","iopub.status.idle":"2022-08-02T17:57:55.083292Z","shell.execute_reply.started":"2022-08-02T17:57:55.073811Z","shell.execute_reply":"2022-08-02T17:57:55.082003Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation:\n- only target is not present in test_data\n- All other columns are present","metadata":{}},{"cell_type":"code","source":"# checking for null data in both dataset\n\nnull_values_train_data = (train_data.isna().sum()/ train_data.shape[0])*100\nnull_values_test_data = (test_data.isna().sum()/ test_data.shape[0]) * 100\n\n# print(null_values_train_data)\n\nfig = make_subplots(rows=2, cols= 1, subplot_titles=[\"Null values(%) for train DS\", \"Null values(%) for test DS\"])\n\nfig.add_trace(go.Bar(x = null_values_train_data.index, \n                    y =  null_values_train_data.values,\n                    text = null_values_train_data.values,\n                    texttemplate=\"%{text:.2f}%\", \n                    textposition='outside',), \n              row=1, col=1,)\nfig.update_xaxes(row=1,col=1, title='Columns')\nfig.update_yaxes(row=1,col=1, title='Percentage')\n\n\nfig.add_trace(go.Bar(x = null_values_test_data.index,\n                    y = null_values_test_data.values,\n                    text = null_values_test_data.values,\n                    texttemplate=\"%{text:.2f}%\", \n                    textposition='outside'), \n              row=2, col=1,)\nfig.update_xaxes(row=2,col=1, title='Columns')\nfig.update_yaxes(row=2,col=1, title='Percentage')\n\n\nfig.update_layout(autosize=True, height=1000)\nfig.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T17:57:55.084991Z","iopub.execute_input":"2022-08-02T17:57:55.086087Z","iopub.status.idle":"2022-08-02T17:57:55.425104Z","shell.execute_reply.started":"2022-08-02T17:57:55.086041Z","shell.execute_reply":"2022-08-02T17:57:55.423744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation:\n- Both sets have roughly equal amount of null columns\n- No missing values in target variable\n\nDecision regarding dealing with null values:\n- Imputing the missing values by mean and mode\n- alternative possibility we can use Fully Connected Neural Network to predict the missing values.Very lengthy process.","metadata":{}},{"cell_type":"code","source":"print(\"Number of duplicates in train_data = \", train_data.duplicated().sum())\nprint(\"Number of duplicate in test_data = \", test_data.duplicated().sum())","metadata":{"execution":{"iopub.status.busy":"2022-08-02T17:57:55.428000Z","iopub.execute_input":"2022-08-02T17:57:55.428377Z","iopub.status.idle":"2022-08-02T17:57:55.510403Z","shell.execute_reply.started":"2022-08-02T17:57:55.428342Z","shell.execute_reply":"2022-08-02T17:57:55.509464Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# creating a list of columns of cateogrical variables\ncat_var = ['product_code', 'attribute_0', 'attribute_1','attribute_2', 'attribute_3',]\ncont_var = ['loading','measurement_0', 'measurement_1',\n       'measurement_2', 'measurement_3', 'measurement_4', 'measurement_5',\n       'measurement_6', 'measurement_7', 'measurement_8', 'measurement_9',\n       'measurement_10', 'measurement_11', 'measurement_12', 'measurement_13',\n       'measurement_14', 'measurement_15', 'measurement_16', 'measurement_17']","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:21:35.634660Z","iopub.execute_input":"2022-08-02T18:21:35.636242Z","iopub.status.idle":"2022-08-02T18:21:35.643651Z","shell.execute_reply.started":"2022-08-02T18:21:35.636164Z","shell.execute_reply":"2022-08-02T18:21:35.642092Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# helper function to plot distribution in seaborn\ndef plot_dist_sbn(data, title, categorical=False, figsize=(20,18)):\n    if data.shape[1] % 5 == 0:\n        rows = (data.shape[1] // 5)\n    else:\n        rows = (data.shape[1] // 5) + 1\n    f, axs = plt.subplots(rows, 5, figsize=figsize)\n    axs = iter(axs.flatten())\n    col_iter1 = iter(list(data.columns))\n    for row in range(rows):\n        for col in range(5):\n            try: \n                col_name = next(col_iter1)\n            except:\n                break\n            ax = next(axs)\n            if categorical:\n                sbn.countplot(x=data[col_name], ax=ax)\n                ax.set_ylabel(\"Count\")\n            else:\n                sbn.histplot(x=data[col_name], ax=ax, color=\"#00a9e0\")\n                col_std = data[col_name].std()\n                col_mean = data[col_name].mean()\n                ax.set_title(f\"Mean={col_mean:.2f}\\nSTD={col_std:.2f}\")\n                ax.set_ylabel(\"Distribution\")\n            ax.set_xlabel(f\"{col_name}\")\n            \n    plt.suptitle(f\"{title}\", fontsize=20)\n    plt.tight_layout()\n    plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T19:01:19.067541Z","iopub.execute_input":"2022-08-02T19:01:19.068388Z","iopub.status.idle":"2022-08-02T19:01:19.081102Z","shell.execute_reply.started":"2022-08-02T19:01:19.068345Z","shell.execute_reply":"2022-08-02T19:01:19.079770Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"suptitle= \"Distribution of continuous variables in train_data\"\ndisplay(plot_dist_sbn(train_data[cont_var], title=suptitle))\n\nsuptitle= \"Distribution of continuous variables in test_data\"\ndisplay(plot_dist_sbn(test_data[cont_var], title=suptitle))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T19:00:03.421488Z","iopub.execute_input":"2022-08-02T19:00:03.422452Z","iopub.status.idle":"2022-08-02T19:00:17.059806Z","shell.execute_reply.started":"2022-08-02T19:00:03.422400Z","shell.execute_reply":"2022-08-02T19:00:17.058561Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observation:\n- most continuous variables have standard deviation = 1, but mean is not 0\n- columns = loading, measurement_0, measurement_1, measurement_2 have right side tailed data.\n- \n\nDecision:\n- The data needs to be Standardization before training\n- Or perform log on the columns","metadata":{}},{"cell_type":"code","source":"suptitle = \"Frequency of category in categorical variables train_data\"\ndisplay(plot_dist_sbn(train_data[cat_var], title = suptitle, categorical=True, figsize=(18,5)))\n\nsuptitle = \"Frequency of category in categorical variables test_data\"\ndisplay(plot_dist_sbn(test_data[cat_var], title = suptitle, categorical=True, figsize=(18,5)))","metadata":{"execution":{"iopub.status.busy":"2022-08-02T19:01:21.418590Z","iopub.execute_input":"2022-08-02T19:01:21.419042Z","iopub.status.idle":"2022-08-02T19:01:22.704443Z","shell.execute_reply.started":"2022-08-02T19:01:21.419004Z","shell.execute_reply":"2022-08-02T19:01:22.702894Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Observations:\n- the categorical variables have imbalanced data\n- the target variable also has imabalanced labels\n- There are 5 products in train_data whereas we have only 4 product_codes in test_data. If we One-Hot-encode, it will create problem in predictions\n- the cateogries of train_data and test_data have different frequency.\n- The attributes all have different materials in train_data and test_data","metadata":{}},{"cell_type":"code","source":"sbn.countplot(x=train_data['failure'], )\nplt.title(\"Frequency of failure(target) labels\", fontsize=16)\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-08-02T18:49:27.740592Z","iopub.execute_input":"2022-08-02T18:49:27.741045Z","iopub.status.idle":"2022-08-02T18:49:27.917016Z","shell.execute_reply.started":"2022-08-02T18:49:27.741012Z","shell.execute_reply":"2022-08-02T18:49:27.915802Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"","metadata":{},"execution_count":null,"outputs":[]}]}