{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# Explorory Data Analysis on Huge Data with Dask","metadata":{}},{"cell_type":"markdown","source":"First EDA on Default Prediction competition data. Hope, this notebook will be useful for you ;)","metadata":{}},{"cell_type":"markdown","source":"**Warning** In this notebook I explored all train and test data with Dask library. Because there is a very huge dataset, some cells here perform very slow. You are warned ;)","metadata":{}},{"cell_type":"markdown","source":"# Table of contents\n\n1. [Import libraries](#import-libraries)\n1. [Load data](#load-data)\n1. [First look at data](#first-look-at-data)\n1. [Examine shape of given data](#examine-shape-of-given-data)\n    1. [Shape of train data](#shape-of-train-data)\n    1. [Shape of test data](#shape-of-test-data)\n1. [Examine columns in data](#examine-columns-in-data)\n1. [Check types of each column](#check-types-of-each-column)\n1. [Discover NaNs](#discover-nans)\n1. [Build distplot for all numeric features](#build-distplot)","metadata":{}},{"cell_type":"markdown","source":"## Import libraries <a class=\"anchor\" id=\"import-libraries\"></a>","metadata":{}},{"cell_type":"code","source":"import os\n\nimport numpy as np\nimport pandas as pd\nimport dask.dataframe as dd\n\nimport matplotlib.pyplot as plt\nimport seaborn as sns","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:29.952951Z","iopub.execute_input":"2022-05-25T23:43:29.953895Z","iopub.status.idle":"2022-05-25T23:43:31.703120Z","shell.execute_reply.started":"2022-05-25T23:43:29.953843Z","shell.execute_reply":"2022-05-25T23:43:31.702336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"pd.set_option('display.max_columns', 500)\npd.set_option('display.max_rows', 500)\npd.set_option('display.float_format', lambda x: '%.5f' % x)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:31.704702Z","iopub.execute_input":"2022-05-25T23:43:31.705026Z","iopub.status.idle":"2022-05-25T23:43:31.710086Z","shell.execute_reply.started":"2022-05-25T23:43:31.704998Z","shell.execute_reply":"2022-05-25T23:43:31.709280Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Used this option to disable scientific notation","metadata":{}},{"cell_type":"code","source":"PATH_TO_DATA = os.path.join('/kaggle', 'input', 'amex-default-prediction')","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:31.711156Z","iopub.execute_input":"2022-05-25T23:43:31.711703Z","iopub.status.idle":"2022-05-25T23:43:31.725146Z","shell.execute_reply.started":"2022-05-25T23:43:31.711669Z","shell.execute_reply":"2022-05-25T23:43:31.724416Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Load data <a class=\"anchor\" id=\"load-data\"></a>","metadata":{}},{"cell_type":"markdown","source":"[Dask](https://docs.dask.org/en/latest/10-minutes-to-dask.html) allows us to work with data, without loading it fully on RAM. It's API is very similar to Pandas","metadata":{}},{"cell_type":"code","source":"train_data = dd.read_csv(os.path.join(PATH_TO_DATA, 'train_data.csv'))\ntest_data = dd.read_csv(os.path.join(PATH_TO_DATA, 'test_data.csv'))\ntrain_labels = dd.read_csv(os.path.join(PATH_TO_DATA, 'train_labels.csv'))\nsample_submission = dd.read_csv(os.path.join(PATH_TO_DATA, 'sample_submission.csv'))","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:31.728183Z","iopub.execute_input":"2022-05-25T23:43:31.728842Z","iopub.status.idle":"2022-05-25T23:43:32.016234Z","shell.execute_reply.started":"2022-05-25T23:43:31.728803Z","shell.execute_reply":"2022-05-25T23:43:32.015251Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## First look at data  <a class=\"anchor\" id=\"first-look-at-data\"></a>","metadata":{}},{"cell_type":"code","source":"train_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:32.017412Z","iopub.execute_input":"2022-05-25T23:43:32.017719Z","iopub.status.idle":"2022-05-25T23:43:34.385975Z","shell.execute_reply.started":"2022-05-25T23:43:32.017692Z","shell.execute_reply":"2022-05-25T23:43:34.385016Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:34.387060Z","iopub.execute_input":"2022-05-25T23:43:34.387492Z","iopub.status.idle":"2022-05-25T23:43:36.195926Z","shell.execute_reply.started":"2022-05-25T23:43:34.387459Z","shell.execute_reply":"2022-05-25T23:43:36.195010Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.head()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:36.197352Z","iopub.execute_input":"2022-05-25T23:43:36.197688Z","iopub.status.idle":"2022-05-25T23:43:37.042181Z","shell.execute_reply.started":"2022-05-25T23:43:36.197659Z","shell.execute_reply":"2022-05-25T23:43:37.041518Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_data.describe().compute()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:43:37.043332Z","iopub.execute_input":"2022-05-25T23:43:37.043791Z","iopub.status.idle":"2022-05-25T23:47:49.605621Z","shell.execute_reply.started":"2022-05-25T23:43:37.043761Z","shell.execute_reply":"2022-05-25T23:47:49.604336Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.describe().compute()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:47:49.606817Z","iopub.execute_input":"2022-05-25T23:47:49.607140Z","iopub.status.idle":"2022-05-25T23:56:22.243642Z","shell.execute_reply.started":"2022-05-25T23:47:49.607112Z","shell.execute_reply":"2022-05-25T23:56:22.242920Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train_labels.describe().compute()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:56:22.245882Z","iopub.execute_input":"2022-05-25T23:56:22.246451Z","iopub.status.idle":"2022-05-25T23:56:23.531775Z","shell.execute_reply.started":"2022-05-25T23:56:22.246409Z","shell.execute_reply":"2022-05-25T23:56:23.531095Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Examine shape of given data <a class=\"anchor\" id=\"examine-shape-of-given-data\"></a>","metadata":{}},{"cell_type":"markdown","source":"### Shape of train data <a class=\"anchor\" id=\"shape-of-train-data\"></a>","metadata":{}},{"cell_type":"code","source":"train_rows = train_data.size.compute()\ntrain_columns = train_data.shape[1]\n\nprint(f'Shape of train data: {train_rows} {train_columns}')","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:56:23.532905Z","iopub.execute_input":"2022-05-25T23:56:23.533211Z","iopub.status.idle":"2022-05-25T23:58:53.631032Z","shell.execute_reply.started":"2022-05-25T23:56:23.533177Z","shell.execute_reply":"2022-05-25T23:58:53.630030Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### Shape of test data  <a class=\"anchor\" id=\"shape-of-test-data\"></a>","metadata":{}},{"cell_type":"code","source":"test_rows = test_data.size.compute()\ntest_columns = test_data.shape[1]\n\nprint(f'Shape of test data: {test_rows} {test_columns}')","metadata":{"execution":{"iopub.status.busy":"2022-05-25T23:58:53.632901Z","iopub.execute_input":"2022-05-25T23:58:53.633373Z","iopub.status.idle":"2022-05-26T00:03:36.781703Z","shell.execute_reply.started":"2022-05-25T23:58:53.633341Z","shell.execute_reply":"2022-05-26T00:03:36.780374Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Looks pretty huge)","metadata":{}},{"cell_type":"markdown","source":"## Examine columns in data  <a class=\"anchor\" id=\"examine-columns-in-data\"></a>","metadata":{}},{"cell_type":"markdown","source":"### Check is columns in train data same as in test data","metadata":{}},{"cell_type":"code","source":"train_data.columns.tolist() == test_data.columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T00:03:36.784213Z","iopub.execute_input":"2022-05-26T00:03:36.785284Z","iopub.status.idle":"2022-05-26T00:03:36.792170Z","shell.execute_reply.started":"2022-05-26T00:03:36.785234Z","shell.execute_reply":"2022-05-26T00:03:36.791425Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"### List of columns' names","metadata":{}},{"cell_type":"code","source":"train_data.columns.tolist()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T00:03:36.793697Z","iopub.execute_input":"2022-05-26T00:03:36.794439Z","iopub.status.idle":"2022-05-26T00:03:36.810382Z","shell.execute_reply.started":"2022-05-26T00:03:36.794396Z","shell.execute_reply":"2022-05-26T00:03:36.809640Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Check types of each column   <a class=\"anchor\" id=\"check-types-of-each-column\"></a>","metadata":{}},{"cell_type":"code","source":"train_data.dtypes","metadata":{"execution":{"iopub.status.busy":"2022-05-26T00:03:36.811700Z","iopub.execute_input":"2022-05-26T00:03:36.812234Z","iopub.status.idle":"2022-05-26T00:03:36.832778Z","shell.execute_reply.started":"2022-05-26T00:03:36.812200Z","shell.execute_reply":"2022-05-26T00:03:36.831978Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"As we see, nearly all columns are numeric, exclude 'customer_id' ans 'S_2', which represents date.","metadata":{}},{"cell_type":"markdown","source":"## Discover NaNs  <a class=\"anchor\" id=\"discover-nans\"></a>","metadata":{}},{"cell_type":"markdown","source":"As we see here, there are some columns with big number of nans.","metadata":{}},{"cell_type":"code","source":"train_data.isnull().sum(axis = 0).compute()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T00:03:36.834478Z","iopub.execute_input":"2022-05-26T00:03:36.835006Z","iopub.status.idle":"2022-05-26T00:05:56.068624Z","shell.execute_reply.started":"2022-05-26T00:03:36.834975Z","shell.execute_reply":"2022-05-26T00:05:56.067496Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"test_data.isnull().sum(axis = 0).compute()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T00:05:56.070204Z","iopub.execute_input":"2022-05-26T00:05:56.070769Z","iopub.status.idle":"2022-05-26T00:10:57.309849Z","shell.execute_reply.started":"2022-05-26T00:05:56.070735Z","shell.execute_reply":"2022-05-26T00:10:57.308593Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"## Build distplot for all numeric features  <a class=\"anchor\" id=\"build-distplot\"></a>","metadata":{}},{"cell_type":"code","source":"%matplotlib inline\nplt.style.use('seaborn')\ncolumns_to_draw = list(set(train_data.columns.tolist()) - set(['customer_ID', 'S_2']))\nfor current_column in columns_to_draw:\n    if pd.api.types.is_numeric_dtype(train_data[current_column].dtype):\n        fig, ax = plt.subplots(figsize=(10, 4))\n        ax = sns.distplot(train_data[current_column].compute(), kde=False)\n        ax.set_title(current_column)\n        plt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-26T00:10:57.312445Z","iopub.execute_input":"2022-05-26T00:10:57.313241Z","iopub.status.idle":"2022-05-26T05:31:50.778564Z","shell.execute_reply.started":"2022-05-26T00:10:57.313189Z","shell.execute_reply":"2022-05-26T05:31:50.774977Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"**TO BE CONTINUED...**","metadata":{}}]}