{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <b><span style='color:#016FD0;font-size:300%'>1 |</span><span style='color:#016FD0;font-size:300%'> Introduction</span></b>","metadata":{}},{"cell_type":"markdown","source":"Train dataset for [\"American Express - Default Prediction\"][1] competition is visualized for gettting some insight into strategy of EDA, training, etc. The code for visualization with [Matplotlib][2] in the kernel is simple and easy to understand. So it can be helpful for kaggle/machine learning/data science beginners.\n\n(The train dataset for the competition is very large, so it was splitted to approximately equal-sized chunks in the other kernel [\"Chunk of \"AMEX - Default Prediction\" Dataset\"][3] for processing it more efficiently. One of them is used in the kernel.)\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction\n[2]: https://matplotlib.org/\n[3]: https://www.kaggle.com/code/acchiko/chunk-of-amex-default-prediction-dataset","metadata":{}},{"cell_type":"markdown","source":" The kernel was created referring to the following SUPER NICE kernels,\n- [AMEX Default Prediction EDA & LGBM Baseline][3],\n- [Time Series EDA][4],\n- [AMEX EDA which makes sense ⭐️⭐️⭐️⭐️⭐️][5].\n\nIt is recommended to read & upvote those as well.\n\n[3]: https://www.kaggle.com/code/kellibelcher/amex-default-prediction-eda-lgbm-baseline   \n[4]: https://www.kaggle.com/code/cdeotte/time-series-eda/notebook\n[5]: https://www.kaggle.com/code/ambrosm/amex-eda-which-makes-sense ","metadata":{}},{"cell_type":"markdown","source":"<b><div style='color:#9BD4F5;font-size:180%'>NOTE :</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- When you want to save time to read, please check hidden \"Table of Contents\" (in the right side of notebook) first. All of the data, figures, tables, etc. are summarized in it, an you can jump to the item you want to check.</div></b>","metadata":{}},{"cell_type":"markdown","source":"The kernel may have several bugs/wrongs. I am happy to get your comments. Thank you in advance for your kind advice to make the kernel so NICE! and to make me NICE deep learning guy!!","metadata":{}},{"cell_type":"markdown","source":"# <b><span style='color:#016FD0;font-size:300%'>2 |</span><span style='color:#016FD0;font-size:300%'> Load Competition Dataset</span></b>","metadata":{}},{"cell_type":"markdown","source":"Load [\"American Express - Default Prediction\"][1] competition dataset to the kernel and check its contents. But the dataset is very large, so it was splitted to approximately equal-sized chunks in the other kernel [\"Chunk of \"AMEX - Default Prediction\" Dataset\"][3] for processing it more efficiently. Load one of them here.\n\n[1]: https://www.kaggle.com/competitions/amex-default-prediction\n[3]: https://www.kaggle.com/code/acchiko/chunk-of-amex-default-prediction-dataset","metadata":{}},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Required libraries</div></b>","metadata":{}},{"cell_type":"code","source":"# Import libs.\nimport json\nimport glob\nimport pandas as pd\nfrom itertools import islice\nimport gc","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:35.088647Z","iopub.execute_input":"2022-07-17T22:38:35.089553Z","iopub.status.idle":"2022-07-17T22:38:35.120843Z","shell.execute_reply.started":"2022-07-17T22:38:35.089426Z","shell.execute_reply":"2022-07-17T22:38:35.119653Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>List of files</div></b>","metadata":{}},{"cell_type":"code","source":"# Show list of files.\npath_to_dir_chunked_metadata = \"/kaggle/input/chunk-of-amex-default-prediction-dataset\"\n!ls {path_to_dir_chunked_metadata}","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:35.126375Z","iopub.execute_input":"2022-07-17T22:38:35.126752Z","iopub.status.idle":"2022-07-17T22:38:35.922914Z","shell.execute_reply.started":"2022-07-17T22:38:35.126719Z","shell.execute_reply":"2022-07-17T22:38:35.921335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Name of features</div></b>","metadata":{}},{"cell_type":"code","source":"# Define path to json, which grouped name of features are summarized.\npath_to_features_json = f\"{path_to_dir_chunked_metadata}/features.json\"\n!ls {path_to_features_json}","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:35.925863Z","iopub.execute_input":"2022-07-17T22:38:35.927101Z","iopub.status.idle":"2022-07-17T22:38:36.730421Z","shell.execute_reply.started":"2022-07-17T22:38:35.927053Z","shell.execute_reply":"2022-07-17T22:38:36.728316Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Load the json.\nwith open(path_to_features_json) as fin:\n    features = json.load(fin)\n    \nprint(features)","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:36.732332Z","iopub.execute_input":"2022-07-17T22:38:36.732745Z","iopub.status.idle":"2022-07-17T22:38:36.744079Z","shell.execute_reply.started":"2022-07-17T22:38:36.732709Z","shell.execute_reply":"2022-07-17T22:38:36.743127Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Chunked train metadata</div></b>","metadata":{}},{"cell_type":"code","source":"# Define paths to chunked train metadata.\npaths_to_chunked_train_metadata = sorted(glob.glob(f\"{path_to_dir_chunked_metadata}/train_data_chunk_*.parquet\"))\npaths_to_chunked_train_metadata","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:36.746473Z","iopub.execute_input":"2022-07-17T22:38:36.747341Z","iopub.status.idle":"2022-07-17T22:38:36.760715Z","shell.execute_reply.started":"2022-07-17T22:38:36.747302Z","shell.execute_reply":"2022-07-17T22:38:36.759607Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show contents of a chunked train metadata as example.\nchunked_train_metadata = pd.read_parquet(paths_to_chunked_train_metadata[0])\nchunked_train_metadata","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:36.763356Z","iopub.execute_input":"2022-07-17T22:38:36.764032Z","iopub.status.idle":"2022-07-17T22:38:42.436905Z","shell.execute_reply.started":"2022-07-17T22:38:36.763998Z","shell.execute_reply":"2022-07-17T22:38:42.434745Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Memory usage for a chunked train metadata</div></b>","metadata":{}},{"cell_type":"code","source":"# Show memory usage for a chunked train metadata.\nchunked_train_metadata.info(memory_usage=\"deep\")","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:42.438706Z","iopub.execute_input":"2022-07-17T22:38:42.439259Z","iopub.status.idle":"2022-07-17T22:38:42.472636Z","shell.execute_reply.started":"2022-07-17T22:38:42.439208Z","shell.execute_reply":"2022-07-17T22:38:42.471747Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Chunked test metadata</div></b>","metadata":{}},{"cell_type":"code","source":"# Define path to chunked test metadata.\npaths_to_chunked_test_metadata = sorted(glob.glob(f\"{path_to_dir_chunked_metadata}/test_data_chunk_*.parquet\"))\npaths_to_chunked_test_metadata","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:42.474117Z","iopub.execute_input":"2022-07-17T22:38:42.474641Z","iopub.status.idle":"2022-07-17T22:38:42.484028Z","shell.execute_reply.started":"2022-07-17T22:38:42.474612Z","shell.execute_reply":"2022-07-17T22:38:42.482749Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show contents of a chunked test metedata as example.\nchunked_test_metadata = pd.read_parquet(paths_to_chunked_test_metadata[0])\nchunked_test_metadata","metadata":{"execution":{"iopub.status.busy":"2022-07-17T22:38:42.485677Z","iopub.execute_input":"2022-07-17T22:38:42.486335Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Memory usage for chunked test metadata</div></b>","metadata":{}},{"cell_type":"code","source":"# Show memory usage for a chunked test metadata.\nchunked_test_metadata.info(memory_usage=\"deep\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Train labels</div></b>","metadata":{}},{"cell_type":"code","source":"# Define path to train labels.\npath_to_train_labels = f\"{path_to_dir_chunked_metadata}/train_labels.parquet\"\npath_to_train_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show contents of train labels.\ntrain_labels = pd.read_parquet(path_to_train_labels)\ntrain_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Memory usage for train labels</div></b>","metadata":{}},{"cell_type":"code","source":"# Show memory usage for train labels.\ntrain_labels.info(memory_usage=\"deep\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Memory cleaning</div></b>","metadata":{}},{"cell_type":"code","source":"# Clean memory, if it is required.\ndel chunked_train_metadata # It is just a example.\ndel chunked_test_metadata # It is just a example.\ndel train_labels # It is just a example.\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><span style='color:#016FD0;font-size:300%'>3 |</span><span style='color:#016FD0;font-size:300%'> Preview train labels</span></b>","metadata":{}},{"cell_type":"markdown","source":"Preview train labels, which is not chunked.","metadata":{}},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Required libraries</div></b>","metadata":{}},{"cell_type":"code","source":"# Import libs and set configurations for plotting.\nimport numpy as np\nimport matplotlib\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport math\n\nmatplotlib.style.use(\"ggplot\")\npd.set_option(\"display.max_columns\", 500)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Train labels</div></b>","metadata":{}},{"cell_type":"code","source":"# Load train labels.\nlabels = pd.read_parquet(path_to_train_labels)\nlabels.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Number of non-default/default customer IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Count numbers of non-default/default customer IDs.\nnum_customer_IDs = labels.groupby(\"target\")[\"customer_ID\"].count()\nnum_customer_IDs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Define utility function for plotting pie chart.\ndef showPieChart(count_grouped_by_target, labels=[\"Non-default\", \"Default\"], colors=[\"#9BD4F5\", \"#636364\"]):\n    kwargs_pie = dict(autopct=\"%.1f%%\", ylabel=\"\", legend=False)\n    count_grouped_by_target.plot.pie(labels=labels, colors=colors, **kwargs_pie)\n    plt.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show ratio of non-default/default customer IDs.\nshowPieChart(num_customer_IDs)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b><div style='color:#9BD4F5;font-size:180%'>TIPS : </div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- Imbalanced dataset can cause problems as shown in the articles [\"Dealing with Imbalanced Data\"][1] and [\"10 Techniques to deal with Imbalanced Classes in Machine Learning\"][2]. Performance of prediction model can be improved by applying resampling techniques, such as undersamping non-dafault customers, oversampling default customers, generating synthetic samples, etc.</div></b>\n\n[1]: https://towardsdatascience.com/methods-for-dealing-with-imbalanced-data-5b761be45a18\n[2]: https://www.analyticsvidhya.com/blog/2020/07/10-techniques-to-deal-with-class-imbalance-in-machine-learning/#h2_20","metadata":{}},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of non-default/default customer IDs</div></b>\n","metadata":{}},{"cell_type":"code","source":"# Split train labels to approximately equal-sized chunks for plotting distribution of non-default/default customer IDs.\n# The chunked train labels are corresponding to chunked train metadata. But those are used just for plotting distribution.\nnum_chunks = 10\nchunked_customer_ids = np.array_split(ary=labels[\"customer_ID\"].tolist(), indices_or_sections=num_chunks)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show distribution of non-default/default customer IDs for each chunk of train labels.\n# The indices for default customer IDs (i.e. target = \"1\") will be black.\nncols = 5\nnrows = math.ceil(num_chunks / ncols)\nfig, ax = plt.subplots(nrows=nrows, ncols=ncols, figsize=(5*ncols, 12*nrows))\n\nfor chunk_id, (chunked_customer_ids_, ax) in enumerate(zip(chunked_customer_ids, ax.flat)):\n    # Split train labels and replace target to bool for plotting. \"False\" will be black.\n    chunked_labels = labels[labels[\"customer_ID\"].isin(chunked_customer_ids_)]\n    chunked_labels_for_plot = pd.DataFrame(chunked_labels[\"target\"].map({\"0\": True, \"1\": False}).astype(\"bool\"))\n    \n    # Plot distribution (heatmap).\n    sns.heatmap(chunked_labels_for_plot, cbar=False, ax=ax)\n    \n    # Count number of non-default/default customer IDs and show it in title.\n    num_customer_IDs = chunked_labels.groupby(\"target\").count()\n    ax.set_title(f\"Chunk {chunk_id:03d} (#of0/1s: {num_customer_IDs.loc['0'][0]}/{num_customer_IDs.loc['1'][0]})\")\n    \nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Memory cleaning</div></b>","metadata":{}},{"cell_type":"code","source":"# Clean memory, if it is required.\ndel labels\ndel chunked_customer_ids\ndel chunked_labels\ndel chunked_labels_for_plot\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><span style='color:#016FD0;font-size:300%'>4 |</span><span style='color:#016FD0;font-size:300%'> Preview chunked train metadata</span></b>","metadata":{}},{"cell_type":"markdown","source":"Preview chunked train metadata. Configuration for previewing can be changed, if it is required.","metadata":{}},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Configuration</div></b>","metadata":{}},{"cell_type":"code","source":"# Define class for configuration for previewing chunked train metadata.\nclass Config:\n    chunk_ID = 2\n    path_to_metadata = paths_to_chunked_train_metadata[chunk_ID]\n    path_to_labels = path_to_train_labels # Not chunked labels\n    features = features","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Required libraries</div></b>","metadata":{}},{"cell_type":"code","source":"# Import libs.\nimport matplotlib\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nimport seaborn as sns\nimport random\nimport math\n\nmatplotlib.style.use(\"ggplot\")\npd.set_option(\"display.max_columns\", 500)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Chunked train metadata</div></b>","metadata":{}},{"cell_type":"code","source":"# Load a chunked train metadata.\nmetadata = pd.read_parquet(Config.path_to_metadata)\nmetadata.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Unique customer IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Get unique customer IDs in chunked train metadata.\ncustomer_IDs = metadata[\"customer_ID\"].unique()\ncustomer_IDs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Train labels</div></b>","metadata":{}},{"cell_type":"code","source":"# Load train labels.\nlabels = pd.read_parquet(Config.path_to_labels)\nlabels.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Chunked train labels</div></b>","metadata":{}},{"cell_type":"code","source":"# Extract targets corresponding to customer IDs in chunked train metadata.\nchunked_labels = labels[labels[\"customer_ID\"].isin(customer_IDs)]\nchunked_labels","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Number of non-default/default customer IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Count number of non-default/default customer IDs in chunked train labels.\nnum_customer_IDs = chunked_labels.groupby(\"target\")[\"customer_ID\"].count()\nnum_customer_IDs","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show ratio of non-default/default customer IDs in chunked train labels.\nshowPieChart(num_customer_IDs)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Chunked train metadata with labels</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility function for merging train metadata and labels.\ndef mergeLabels(metadata, labels):\n    # Set customer ID as index for extracting target efficiently.\n    indexed_labels = labels.set_index(\"customer_ID\")\n    \n    # Extract target for customer IDs in metadata and add it to metadata.\n    metadata[\"target\"] = [indexed_labels.loc[customer_ID].iloc[0] for customer_ID in metadata[\"customer_ID\"]]\n    metadata[\"target\"] = metadata[\"target\"].astype(\"category\")\n    \n    return metadata","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Merge train metadata and labels. Measure processing time for the discussion in tips below.\nmetadata = mergeLabels(metadata, labels)\nmetadata.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b><div style='color:#9BD4F5;font-size:180%'>TIPS :</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>Q. Which ways for merging labels are fast? There are several ways such as,</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- By indexing labels (above way),</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- By quering target for each customer_ID (the following way).</div></b>\n\n<b><div style='color:#9BD4F5;font-size:120%'>A. The former way is as several times faster than the latter way as shown below.</div></b>","metadata":{}},{"cell_type":"code","source":"# Copy metadata with no \"target\" feature for experiment.\n#metadata_tmp = metadata.drop(columns=\"target\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"%%time\n# Merge train metadata and labels by quering target for each customer ID. Measure processing time for discussion of the tips.\n#metadata_tmp[\"target\"] = [labels[labels[\"customer_ID\"] == customer_ID][\"target\"].iloc[0] for customer_ID in metadata_tmp[\"customer_ID\"]]\n#metadata_tmp.head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Remove copied metadata for experiment.\n#del metadata_tmp\n#gc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Chunked train metadata for non-default/default customer IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Split metadata for non-default/default customer IDs.\ngrouped_metadata = {\n    \"non-default\": metadata[metadata[\"target\"] == \"0\"],\n    \"default\":     metadata[metadata[\"target\"] == \"1\"]\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show splitted metadata for non-default customer IDs.\ngrouped_metadata[\"non-default\"].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show splitted metadata for default customer IDs.\ngrouped_metadata[\"default\"].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Range of date of statements</div></b>","metadata":{}},{"cell_type":"code","source":"# Show range of date of statements for non-default customer IDs.\nfirst_date, last_date = grouped_metadata[\"non-default\"][\"S_2\"].min(), grouped_metadata[\"non-default\"][\"S_2\"].max()\nprint(f\"Date of first statement for non-default customer IDs : {first_date}\")\nprint(f\"Date of last statement for non-default customer IDs : {last_date}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show range of date of statements for default customer IDs.\nfirst_date, last_date = grouped_metadata[\"default\"][\"S_2\"].min(), grouped_metadata[\"default\"][\"S_2\"].max()\nprint(f\"Date of first statement for default customer IDs : {first_date}\")\nprint(f\"Date of last statement for default customer IDs : {last_date}\")","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend of number of statements for non-default/default customers IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Count number of statements by a week for non-default/default customer IDs.\nnum_statements_per_week = {\n    \"non-default\": grouped_metadata[\"non-default\"].groupby([pd.Grouper(key=\"S_2\", freq=\"7d\")])[\"customer_ID\"].count(),\n    \"default\":     grouped_metadata[\"default\"].groupby([pd.Grouper(key=\"S_2\", freq=\"7d\")])[\"customer_ID\"].count()\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show number of statements for non-default customer IDs.\nnum_statements_per_week[\"non-default\"].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show number of statements for default customer IDs.\nnum_statements_per_week[\"default\"].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show trends of number of statements for non-default/default customer IDs.\nfig, ax = plt.subplots(1, 1, figsize=(18, 6))\nkwargs = dict(alpha=0.8, marker=\".\", fillstyle=\"none\")\nnum_statements_per_week[\"non-default\"].plot.line(label=\"Non-default\", color=\"#9BD4F5\", ax=ax, **kwargs)\nnum_statements_per_week[\"default\"].plot.line(label=\"Default\", color=\"#636364\", ax=ax, **kwargs)\nax.set_xlabel(\"\")\nax.set_ylabel(\"Number of statements / week\", fontsize=\"xx-large\")\nax.legend(fontsize=\"large\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histogram of number of statements for non-default/default customers IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Count number of statements by non-default/default customer IDs.\nnum_statements = {\n    \"non-default\": grouped_metadata[\"non-default\"].groupby([\"customer_ID\"])[\"S_2\"].count(),\n    \"default\":     grouped_metadata[\"default\"].groupby([\"customer_ID\"])[\"S_2\"].count()\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show number of statements for non-default customer IDs.\nnum_statements[\"non-default\"].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show number of statements for default customer IDs.\nnum_statements[\"default\"].head()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show histogram of number of statements for non-default/default customer IDs.\nfig, ax = plt.subplots(1, 1, figsize=(18, 6))\nkwargs = dict(alpha=0.5, bins=13)\nnum_statements[\"non-default\"].plot.hist(label=\"Non-default\", color=\"#9BD4F5\", ax=ax, **kwargs)\nnum_statements[\"default\"].plot.hist(label=\"Default\", color=\"#636364\", ax=ax, **kwargs)\nax.set_xlabel(\"Number of statements\", fontsize=\"xx-large\")\nax.set_ylabel(\"Log(Frequency)\", fontsize=\"xx-large\")\nax.set_yscale(\"log\")\nax.legend(fontsize=\"large\")\nplt.show()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of categorical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for calculating several stats.\ndef getStatsForNumericalFeatures(metadata_grouped_by_target, coefs=[1.5, 3.0]):\n    \"\"\"Calculate several standard stats. Argument \"coefs\" is used for finding outliers\n    with interquartile range rule. The following range for finding mild/extreme outliers\n    are calculated, [1Q(25%)-3.0*IQR, 1Q(25%)-1.5*IQR, 3Q(75%)+1.5*IQR, 3Q(75%)+3.0*IQR].\"\"\"\n    # Calculate several standard stats with describe() and convert the result to dict\n    # for making calculation of additional stats easy.\n    stats_grouped_by_target = metadata_grouped_by_target.describe()\n    stats_grouped_by_target.rename(columns={\"count\": \"num_nonnans\"}, level=1, inplace = True)\n    stats = _toDict(stats_grouped_by_target)\n    \n    # Calculate several additional stats.\n    stats[\"num_rows\"] = _fullLike(stats[\"num_nonnans\"], metadata_grouped_by_target.size().to_list(), axis=0)\n    stats[\"num_nans\"] = stats[\"num_rows\"] - stats[\"num_nonnans\"]\n    stats[\"frac_nonnans\"] = stats[\"num_nonnans\"] / stats[\"num_rows\"]\n    stats[\"frac_nans\"] = stats[\"num_nans\"] / stats[\"num_rows\"]\n    \n    stats[\"IQR\"] = stats[\"75%\"] - stats[\"25%\"]\n    stats[\"extreme_lower_range\"] = stats[\"25%\"] - coefs[1] * stats[\"IQR\"]\n    stats[\"mild_lower_range\"] = stats[\"25%\"] - coefs[0] * stats[\"IQR\"]\n    stats[\"mild_upper_range\"] = stats[\"75%\"] + coefs[0] * stats[\"IQR\"]\n    stats[\"extreme_upper_range\"] = stats[\"75%\"] + coefs[1] * stats[\"IQR\"]\n    \n    stats[\"z_min\"] = (stats[\"min\"] - stats[\"mean\"]) / stats[\"std\"]\n    stats[\"z_max\"] = (stats[\"max\"] - stats[\"mean\"]) / stats[\"std\"]\n    stats[\"z_extreme_lower_range\"] = (stats[\"extreme_lower_range\"] - stats[\"mean\"]) / stats[\"std\"]\n    stats[\"z_mild_lower_range\"] = (stats[\"mild_lower_range\"] - stats[\"mean\"]) / stats[\"std\"]\n    stats[\"z_mild_upper_range\"] = (stats[\"mild_upper_range\"] - stats[\"mean\"]) / stats[\"std\"]\n    stats[\"z_extreme_upper_range\"] = (stats[\"extreme_upper_range\"] - stats[\"mean\"]) / stats[\"std\"]\n    \n    # Merge dataframes for all stats and return it.\n    names = [\"frac_nonnans\", \"frac_nans\", \"num_rows\", \"num_nonnans\", \"num_nans\",\n             \"mean\", \"std\", \"IQR\", \"extreme_upper_range\", \"mild_upper_range\", \"max\",\n             \"75%\", \"50%\", \"25%\", \"min\", \"mild_lower_range\", \"extreme_lower_range\"] #,\n             #\"z_extreme_upper_range\", \"z_mild_upper_range\", \"z_max\", \"z_min\",\n             #\"z_mild_lower_range\", \"z_extreme_lower_range\"]\n    return pd.concat([stats[name] for name in names], keys=names)\n\ndef getStatsForCategoricalFeatures(metadata_grouped_by_target):\n    \"\"\"Calculate several standard stats.\"\"\"\n    # Calculate several standard stats with describe() and convert the result to dict\n    # for making calculation of additional stats easy.\n    stats_grouped_by_target = metadata_grouped_by_target.describe()\n    stats_grouped_by_target.rename(columns={\"count\": \"num_nonnans\"}, level=1, inplace = True)\n    stats = _toDict(stats_grouped_by_target)\n    \n    # Calculate several additional stats.\n    stats[\"num_rows\"] = _fullLike(stats[\"num_nonnans\"], metadata_grouped_by_target.size().to_list(), axis=0)\n    stats[\"num_nans\"] = stats[\"num_rows\"] - stats[\"num_nonnans\"]\n    stats[\"frac_nonnans\"] = stats[\"num_nonnans\"] / stats[\"num_rows\"]\n    stats[\"frac_nans\"] = stats[\"num_nans\"] / stats[\"num_rows\"]\n    stats[\"frac_top_category\"] = stats[\"freq\"] / stats[\"num_nonnans\"]\n    \n    # Merge dataframes for all stats and return it.\n    names = [\"frac_nonnans\", \"frac_nans\", \"num_rows\", \"num_nonnans\", \"num_nans\",\n             \"unique\", \"frac_top_category\", \"top\", \"freq\",]\n    return pd.concat([stats[name] for name in names], keys=names)\n\ndef _fullLike(df_ref, fill_values, axis=0):\n    \"\"\"Get a dataframe with the same shape and type as a given dataframe.\"\"\"\n    df = df_ref.copy()\n    if axis == 0 or axis == \"index\":\n        for index_, fill_value in zip(df.index, fill_values):\n            df.loc[index_] = fill_value\n    elif axis == 1 or axis == \"columns\":\n        for column, fill_value in zip(df.columns, fill_values):\n            df[column] = fill_value\n    return df\n\ndef _toDict(stats_grouped_by_target):\n    \"\"\"Create dict. of dataframe for each stat.\"\"\"\n    names = stats_grouped_by_target.columns.get_level_values(level=1).unique().tolist()\n    stats = [stats_grouped_by_target.loc[:, (slice(None), name)].droplevel(level=1, axis=1) for name in names]\n    return dict(zip(names, stats))","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show stats of categorical delinquency features.\nstats = {}\n\nfeature_group = \"categorical_delinquency\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForCategoricalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b><div style='color:#9BD4F5;font-size:180%'>VARIABLES IN THE TABLE : </div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- num_rows : Number of rows</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- num_nonnans : Number of (not missing) values</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- num_nans : Number of missing values</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- frac_nonnans : Fraction of (not missing) values (= num_nonnans / num_rows)</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- frac_nans : Fraction of missing values  (= num_nans / num_rows)</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- frac_top_category : Fraction of top category values  (= freq / num_nonnans)</div></b>","metadata":{}},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of numerical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show stats of numerical delinquency features.\nfeature_group = \"numerical_delinquency\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForNumericalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b><div style='color:#9BD4F5;font-size:180%'>VARIABLES IN THE TABLE : </div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- extreme_upper_range : Upper range for finding extreme outliers (= 3Q(75%) + 3.0\\*IQR)</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- mild_upper_range : Upper range for finding mild outliers (= 3Q(75%) + 1.5\\*IQR)</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- mild_lower_range : Lower range for finding mild outliers (= 1Q(25%) - 1.5\\*IQR)</div></b>\n<b><div style='color:#9BD4F5;font-size:120%'>- extreme_lower_range : Lower range for finding extreme outliers (= 1Q(25%) - 3.0\\*IQR)</div></b>","metadata":{}},{"cell_type":"code","source":"# There are so many features. So it is better to select features based on several conditions.\n# The following is a example for selecting the features, whose fraction of nans in a given range. \nname = \"frac_nans\"\nlower_threshold, upper_threshold = 0.0, 0.9\nconditions = ((stats[feature_group].loc[(name, \"0\"), :] > lower_threshold) & (stats[feature_group].loc[(name, \"0\"), :] < upper_threshold)) & \\\n             ((stats[feature_group].loc[(name, \"1\"), :] > lower_threshold) & (stats[feature_group].loc[(name, \"1\"), :] < upper_threshold))\nselected_features = stats[feature_group].columns[conditions].tolist()\n\n# Show selected features.\nprint(f\"Selceted features ({len(selected_features)}) :\")\nprint(selected_features)\n\n# Stats for selected features can be shown again, if it is required.\n#stats[selected_features]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of numerical spend features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show stats of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForNumericalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of numerical payment features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show stats of numerical payment features.\nfeature_group = \"numerical_payment\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForNumericalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of categorical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show stats of categorical balance features.\nfeature_group = \"categorical_balance\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForCategoricalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of numerical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show stats of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForNumericalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Statistics of numerical risk features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show stats of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Stats of {feature_group} features :\")\nmetadata_grouped_by_target = metadata.groupby(\"target\")[Config.features[feature_group]]\nstats[feature_group] = getStatsForNumericalFeatures(metadata_grouped_by_target)\nstats[feature_group]\n#stats[feature_group].loc[(slice(None), \"0\"), :] # Stats for only target == \"0\" can be shown, if it is reuquired.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Box plots of numerical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility fuctions for plotting boxplot.\n# See https://pandas.pydata.org/docs/reference/api/pandas.DataFrame.boxplot.html for details.\ndef showBoxPlots(metadata, target_features, groupby_feature=\"target\", labels=[\"Non-default\", \"Default\"]):\n    \"\"\"Show multiple boxplot for given target features. It is available for only numerical features.\"\"\"\n    # Set number/size of figures.\n    ncols = 4\n    nrows = math.ceil(len(target_features) / ncols)\n    fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(5*ncols, 4*nrows))\n    \n    # Create boxplot for each feature.\n    for feature, ax in zip(target_features, axes.flat):\n        metadata.boxplot(column=feature, by=groupby_feature, ax=ax)\n        ax.set_title(\"\")\n        ax.set_xticklabels(labels=labels, fontsize=\"x-large\")\n        ax.set_xlabel(\"\")\n        ax.set_ylabel(feature, fontsize=\"xx-large\")\n    \n    # Show all figures.\n    plt.suptitle(\"\")\n    plt.tight_layout()\n    plt.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show box plot of numerical delinquency features.\nfeature_group = \"numerical_delinquency\"\nprint(f\"Box plot of {feature_group} features :\")\nshowBoxPlots(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Box plots of numerical spend features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show box plot of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Box plot of {feature_group} features :\")\nshowBoxPlots(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Box plot of numerical payment features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show box plot of numerical payment features.\nfeature_group = \"numerical_payment\"\nprint(f\"Box plot of {feature_group} features :\")\nshowBoxPlots(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Box plots of numerical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show box plot of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Box plot of {feature_group} features :\")\nshowBoxPlots(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Box plots of numerical risk features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show box plots of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Box plot of {feature_group} features :\")\nshowBoxPlots(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of categorical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility fuctions for plotting histogram.\ndef showHistograms(metadata, target_features, groupby_feature=\"target\", labels=[\"Non-default\", \"Default\"], colors=[\"#9BD4F5\", \"#636364\"], alpha=0.5):\n    \"\"\"Show multiple histograms for given target features. It is available for numerical/categorical features.\n    Categorical features are automatically encoded.\"\"\"\n    # Set number/size of figures.\n    ncols = 4\n    nrows = math.ceil(len(target_features) / ncols)\n    fig, axes = plt.subplots(nrows=nrows, ncols=ncols, figsize=(5*ncols, 5*nrows))\n    \n    # Create histogram for each feature.\n    for feature, ax in zip(target_features, axes.flat):\n        # Set number of bins with Sturges fomula. See \"https://en.wikipedia.org/wiki/Histogram#Sturges'_formula\" for details.\n        bins = int(math.log2(metadata[feature].count()) + 1)\n        \n        # Create histogram. Categorical features are encoded before creating histogram.\n        if metadata[feature].dtype.name is \"category\": # for categorical feature\n            encoded_feature, map_for_decoding = _addEncodedFeature(metadata, feature)\n            _createHistogram(df=metadata, target_feature=encoded_feature, groupby_feature=groupby_feature, labels=labels, colors=colors, alpha=alpha, bins=bins, ax=ax)\n            _decodeYTickLabels(map_for_decoding=map_for_decoding, ax=ax)\n            _removeFeature(metadata=metadata, feature=encoded_feature)\n        else: # for numerical feature\n            _createHistogram(df=metadata, target_feature=feature, groupby_feature=groupby_feature, labels=labels, colors=colors, alpha=alpha, bins=bins, ax=ax)\n\n        # Set font size.\n        ax.legend(fontsize=\"large\")\n        ax.set_ylabel(feature, fontsize=\"xx-large\")\n    \n    # Show all figures.\n    plt.tight_layout()\n    plt.show()\n    \ndef _createHistogram(df, target_feature, groupby_feature, labels, colors, alpha, bins, ax, orientation=\"horizontal\", stacked=False):\n    df_grouped = df.groupby(groupby_feature)\n    groups = df_grouped.groups.keys()\n    df_plot = pd.DataFrame({label : df_grouped.get_group(group)[target_feature] for label, group in zip(labels, groups)})\n    \n    kwargs = dict(stacked=stacked, orientation=orientation, alpha=alpha, bins=bins)\n    df_plot.plot.hist(color=colors, ax=ax, **kwargs)\n    \ndef _addEncodedFeature(metadata, feature):\n    encoded_feature = f\"encoded_{feature}\"\n    metadata[encoded_feature] = metadata[feature].cat.codes\n    map_for_decoding = dict(enumerate(metadata[feature].cat.categories))\n    return encoded_feature, map_for_decoding\n    \ndef _removeFeature(metadata, feature):\n    metadata.drop(columns=feature, inplace=True)\n    \ndef _decodeYTickLabels(map_for_decoding, ax):\n    ax.set_yticks(list(map_for_decoding.keys()))\n    ax.set_yticklabels(list(map_for_decoding.values()))","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show histogram of categorical delinquency features. Frequency with no label shows number of nans.\nfeature_group = \"categorical_delinquency\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of numerical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical delinquency features.\nfeature_group = \"numerical_delinquency\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of numerical spend features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of numerical payment features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical payment features.\nfeature_group = \"numerical_payment\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of categorical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show histogram of categorical balance features. Frequency with no label shows number of nans.\nfeature_group = \"categorical_balance\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of numerical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Histograms of numerical risk features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Histogram of {feature_group} features :\")\nshowHistograms(metadata=metadata, target_features=Config.features[feature_group])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of categorical delinquency features","metadata":{}},{"cell_type":"code","source":"# Define utility function for plotting distribution of valid/missing values.\ndef showValidValueDistributions(metadata, num_chunks=10):\n    \"\"\"Show distributions of valid values for all features in given metadata. The indices for valid values will be black.\"\"\"\n    # Split indices to chunks.\n    indices = metadata.index.tolist()\n    chunked_indices = np.array_split(ary=indices, indices_or_sections=num_chunks)\n    \n    # Set number/size of figures.\n    fig, axes = plt.subplots(nrows=num_chunks, ncols=1, figsize=(min(22, len(metadata.columns)), 9*num_chunks))\n    \n    # Show distributions.\n    for chunked_indices_, ax in zip(chunked_indices, axes.flat):\n        sliced_metadata = metadata[metadata.index.isin(chunked_indices_)]\n        is_valid_values = sliced_metadata.isna()\n        sns.heatmap(is_valid_values, cbar=False, ax=ax)\n        \n    # Show figures.\n    plt.tight_layout()\n    plt.show()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show histogram of categorical delinquency features.\nfeature_group = \"categorical_delinquency\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of numerical delinquency features","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical delinquency features.\nfeature_group = \"numerical_delinquency\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of numerical spend features","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of numerical payment features","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical payment features.\nfeature_group = \"numerical_payment\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of categorical balance features","metadata":{}},{"cell_type":"code","source":"# Show histogram of categorical balance features.\nfeature_group = \"categorical_balance\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of numerical balance features","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Distribution of valid/missing values of numerical risk features","metadata":{}},{"cell_type":"code","source":"# Show histogram of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Distribution of valid/missing values of {feature_group} features :\")\nshowValidValueDistributions(metadata[Config.features[feature_group]])","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><span style='color:#016FD0;font-size:300%'>5 |</span><span style='color:#016FD0;font-size:300%'> Preview samples of chunked train metadata</span></b>","metadata":{}},{"cell_type":"markdown","source":"Preview samples of chunked train metadata (Some figures cannot be drawn at once in the kernel because of limitation of kaggle notebook. So some customer IDs are sampled from chunked train metadata for drawing the figures.). Configuration for previewing can be changed, if it is required.","metadata":{}},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Configuration</div></b>","metadata":{}},{"cell_type":"code","source":"# Define class for configuration of previewing. It takes so long time to plot trends for all customer IDs\n# in chunked train metadata, so some of them are sampled here.\nclass Config:\n    metadata = metadata # Same dataframe as the one for prev. chapter.\n    labels = labels     # Same dataframe as the one for prev. chapter.\n    features = features\n    num_samples = [1000, 1000] # Numbers of samples of non-default/default customers.\n    random_seed = 2 # Random seed for sampling non-default/default customers.","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Required libraries</div></b>","metadata":{}},{"cell_type":"code","source":"# Import libs.\nimport matplotlib\nimport matplotlib.pyplot as plt\nfrom tqdm import tqdm\nimport seaborn as sns\nimport random\nimport math\n\nmatplotlib.style.use(\"ggplot\")\npd.set_option(\"display.max_columns\", 500)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Stratified samples of non-default/default customer IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Sample non-default/default customer IDs in chunked train metadata.\nrandom.seed(Config.random_seed)\nsampled_customer_IDs = {\n    \"non-default\": random.sample(Config.metadata[Config.metadata[\"target\"] == \"0\"][\"customer_ID\"].unique().tolist(), Config.num_samples[0]),\n    \"default\": random.sample(Config.metadata[Config.metadata[\"target\"] == \"1\"][\"customer_ID\"].unique().tolist(), Config.num_samples[1])\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show first 5 samples of non-default customer IDs.\nsampled_customer_IDs[\"non-default\"][:5]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show first 5 samples of default customer IDs.\nsampled_customer_IDs[\"default\"][:5]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Metadata for stratified samples for non-default/default customer IDs</div></b>","metadata":{}},{"cell_type":"code","source":"# Extract metadata correspoinding to sampled customer IDs.\nsampled_metadata_ = {\n    \"non-default\": Config.metadata[Config.metadata[\"customer_ID\"].isin(sampled_customer_IDs[\"non-default\"])],\n    \"default\": Config.metadata[Config.metadata[\"customer_ID\"].isin(sampled_customer_IDs[\"default\"])]\n}","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show sampled metadata for non-default customer IDs.\nsampled_metadata_[\"non-default\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show sampled metadata for default customer IDs.\nsampled_metadata_[\"default\"]","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Merge sampled metadata for non-default/default customer IDs.\nsampled_metadata = pd.concat([sampled_metadata_[\"non-default\"], sampled_metadata_[\"default\"]])\nsampled_metadata","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Memory cleaning</div></b>","metadata":{}},{"cell_type":"code","source":"# Clean memory for metadata/labels, if it is required.\ndel Config.metadata\ndel Config.labels\ngc.collect()","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of categorical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility functions for plotting trend/histogram.\ndef showTrendWithHistogram(metadata, target_feature, groupby_feature=\"target\", labels=[\"Non-default\", \"Default\"], colors=[\"#9BD4F5\", \"#636364\"], alpha=0.5):\n    \"\"\"Show trend with histogram for a given target feature. It is available for numerical/categorical features.\n    Categorical features are automatically encoded.\"\"\"\n    # Set number/size of figures.\n    kwargs = dict(sharey=True, figsize=(20, 6), gridspec_kw={\"width_ratios\": [3, 1]})\n    fig, (ax_line, ax_hist) = plt.subplots(nrows=1, ncols=2, **kwargs)\n    \n    # Set number of bins with Sturges fomula. See \"https://en.wikipedia.org/wiki/Histogram#Sturges'_formula\"\n    # for details. Dummy value is set, if all values are missing.\n    if metadata[feature].count() == 0:\n        message =  f\"All values of {feature} are missing. Check it again!!\"\n        _printErrorMessage(message=message)\n        bins = 1 # dummy\n    else:\n        bins = int(math.log2(metadata[feature].count()) + 1)\n    \n    # Create histogram. Categorical features are encoded before creating histogram.\n    if metadata[feature].dtype.name is \"category\": # for categorical feature\n        encoded_feature, map_for_decoding = _addEncodedFeature(metadata, feature)\n        _createLineChart(df=metadata, target_feature=encoded_feature, groupby_feature=groupby_feature, colors=colors, alpha=alpha, ax=ax_line)\n        _createHistogram(df=metadata, target_feature=encoded_feature, groupby_feature=groupby_feature, labels=labels, colors=colors, alpha=alpha, bins=bins, ax=ax_hist)\n        _decodeYTickLabels(map_for_decoding=map_for_decoding, ax=ax_line)\n        _removeFeature(metadata=metadata, feature=encoded_feature)\n    else: # for numerical feature\n        _createLineChart(df=metadata, target_feature=feature, groupby_feature=groupby_feature, colors=colors, alpha=alpha, ax=ax_line)\n        _createHistogram(df=metadata, target_feature=feature, groupby_feature=groupby_feature, labels=labels, colors=colors, alpha=alpha, bins=bins, ax=ax_hist)\n\n    # Set font size.\n    ax_hist.legend(fontsize=\"large\")\n    ax_line.set_ylabel(feature, fontsize=\"xx-large\")\n    ax_line.set_xlabel(\"\")\n    \n    # Show all figures.\n    plt.tight_layout()\n    plt.show()\n    \ndef _createLineChart(df, target_feature, groupby_feature, colors, alpha, ax):\n    df_grouped = df.groupby(groupby_feature)\n    groups = df_grouped.groups.keys()\n    features = [\"customer_ID\", \"S_2\", target_feature]\n    dfs_plot = [df_grouped.get_group(group)[features] for group in groups]\n    \n    # Plot line chart for each customer IDs.\n    kwargs = dict(marker=\".\", fillstyle=\"none\", linewidth=0.5, legend=False)\n    for df_plot, color in zip(dfs_plot, colors):\n        df_plot.groupby(\"customer_ID\").plot.line(x=\"S_2\", y=target_feature, color=color, alpha=alpha, ax=ax, **kwargs)\n        \ndef _printErrorMessage(message):\n    print()\n    print(f\"========================================================================\")\n    print(f\"   {message}\")\n    print(f\"========================================================================\")\n    print()","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show trend/histogram of categorical delinquency features.\nfeature_group = \"categorical_delinquency\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of numerical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of numerical delinquency features.\nfeature_group = \"numerical_delinquency\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of numerical spend features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of numerical payment features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of numerical payment features.\nfeature_group = \"numerical_payment\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of categorical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of categorical balance features.\nfeature_group = \"categorical_balance\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of numerical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of numerical risk features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Trend/histogram of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of diff. of numerical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Define utility fuctions for plotting trend/histogram of diff. of features.\ndef showDiffTrendWithHistogram(metadata, target_feature, groupby_feature=\"target\", labels=[\"Non-default\", \"Default\"], colors=[\"#9BD4F5\", \"#636364\"], alpha=0.5):\n    \"\"\"Show trend/histogram of diff for a given target feature. It is available for only numerical features.\"\"\"\n    # Set number/size of figures.\n    kwargs = dict(sharey=True, figsize=(20, 6), gridspec_kw={\"width_ratios\": [3, 1]})\n    fig, (ax_line, ax_hist) = plt.subplots(nrows=1, ncols=2, **kwargs)\n    \n    # Set number of bins with Sturges fomula. See \"https://en.wikipedia.org/wiki/Histogram#Sturges'_formula\"\n    # for details. Dummy value is set, if all values are missing.\n    if metadata[feature].count() == 0:\n        message =  f\"All values of {feature} are missing. Check it again!!\"\n        _printErrorMessage(message=message)\n        bins = 1 # dummy\n    else:\n        bins = int(math.log2(metadata[feature].count()) + 1)\n    \n    # Create trend/histogram.\n    diff_feature = _addDiffFeature(metadata=metadata, feature=feature)\n    _createLineChart(df=metadata, target_feature=diff_feature, groupby_feature=groupby_feature, colors=colors, alpha=alpha, ax=ax_line)\n    _createHistogram(df=metadata, target_feature=diff_feature, groupby_feature=groupby_feature, labels=labels, colors=colors, alpha=alpha, bins=bins, ax=ax_hist)\n    _removeFeature(metadata=metadata, feature=diff_feature)\n        \n    # Set font size.\n    ax_hist.legend(fontsize=\"large\")\n    ax_line.set_ylabel(diff_feature, fontsize=\"xx-large\")\n    ax_line.set_xlabel(\"\")\n    \n    # Show all figures.\n    plt.tight_layout()\n    plt.show()\n    \ndef _addDiffFeature(metadata, feature):\n    diff_feature = f\"diff_{feature}\"\n    features = [\"customer_ID\", \"S_2\", feature]\n    metadata[diff_feature] = metadata[features].groupby(\"customer_ID\").diff()[feature]\n    return diff_feature","metadata":{"jupyter":{"source_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"# Show trend/histogram of diff. of numerical delinquency features.\nfeature_group = \"numerical_delinquency\"\nprint(f\"Trend/histogram of diff. of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showDiffTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of diff. of numerical spend features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of diff. of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Trend/histogram of diff. of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showDiffTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of diff. of numerical payment features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of diff. of numerical payment features.\nfeature_group = \"numerical_payment\"\nprint(f\"Trend/histogram of diff. of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showDiffTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of diff. of numerical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of diff. of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Trend/histogram of diff. of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showDiffTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Trend/histogram of diff. of numerical risk features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show trend/histogram of diff. of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Trend/histogram of diff. of {feature_group} features :\")\nfor feature in Config.features[feature_group]:\n    showDiffTrendWithHistogram(metadata=sampled_metadata, target_feature=feature)","metadata":{"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Scatter matrix of numerical delinquency features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show scatter matrix of numerical delinquency features.\n#feature_group = \"numerical_delinquency\"\n#print(f\"Scatter matrix of {feature_group} features :\")\n#sns.pairplot(sampled_metadata, vars=Config.features[feature_group], hue=\"target\", corner=True)\n#plt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Scatter matrix of numerical spend features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show scatter matrix of numerical spend features.\nfeature_group = \"numerical_spend\"\nprint(f\"Scatter matrix of {feature_group} features :\")\nsns.pairplot(sampled_metadata, vars=Config.features[feature_group], hue=\"target\", corner=True)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Scatter matrix of numerical balance features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show scatter matrix of numerical balance features.\nfeature_group = \"numerical_balance\"\nprint(f\"Scatter matrix of {feature_group} features :\")\nsns.pairplot(sampled_metadata, vars=Config.features[feature_group], hue=\"target\", corner=True)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <b><div style='padding:20px;background-color:#636364;color:white;border-radius:5px;font-size:80%'>Scatter matrix of numerical risk features</div></b>","metadata":{}},{"cell_type":"code","source":"# Show scatter matrix of numerical risk features.\nfeature_group = \"numerical_risk\"\nprint(f\"Scatter matrix of {feature_group} features :\")\nsns.pairplot(sampled_metadata, vars=Config.features[feature_group], hue=\"target\", corner=True)\nplt.show()","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<b><div style='padding:20px;background-color:#016FD0;color:white;border-radius:5px;font-size:700%'> Under construction</div></b>","metadata":{}},{"cell_type":"markdown","source":"<b><div style='padding:20px;background-color:#016FD0;color:white;border-radius:5px;font-size:700%'> Thank you for reading !!!</div></b>","metadata":{}}]}