{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"# <h1 align = \"center\"><div style = \"background-color: #6600cc; color:white;font-weight:bold; border-radius: 15px; padding: 20px; margin: 2px;\">💳💳American Express - Default Prediction💳💳</div></h1>\n<img src = \"https://images.unsplash.com/photo-1633522715829-ddb5a54e99eb?ixlib=rb-1.2.1&ixid=MnwxMjA3fDB8MHxwaG90by1wYWdlfHx8fGVufDB8fHx8&auto=format&fit=crop&w=1469&q=80\" width = 100%>\n","metadata":{}},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 15px; padding: 20px; margin: 2px;\"><b><span>📗1 | Story</span></b></div></h2>\n\n <div style=\"background:#1a0033   ;font-family: mono;font-size:15px;color:  #f2f2f2;padding:30px\" >\n \n <p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">   \n <b style=\"font-size:30px\">W</b>hether out at a restaurant or buying tickets to a concert, modern life counts on the convenience of a credit card to make daily purchases. It saves us from carrying large amounts of cash and also can advance a full purchase that can be paid over time. How do card issuers know we’ll pay back what we charge? That’s a complex problem with many existing solutions—and even more potential improvements, to be explored in this competition.\n <p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">'''\nCredit default prediction is central to managing risk in a consumer lending business. Credit default prediction allows lenders to optimize lending decisions, which leads to a better customer experience and sound business economics. Current models exist to help manage risk. But it's possible to create better models that can outperform those currently in use.\n\n <p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">\nIn this competition, you’ll apply your machine learning skills to predict credit default. Specifically, you will leverage an industrial scale data set to build a machine learning model that challenges the current model in production. Training, validation, and testing datasets include time-series behavioral data and anonymized customer profile information. You're free to explore any technique to create the most powerful model, from creating features to using the data in a more organic way within a model.\n     \n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">  '''To best prepare all students, GSU and The Learning Agency Lab have joined forces to encourage data scientists to improve automated writing assessments. This public effort could also encourage higher quality and more accessible automated writing tools. If successful, students will receive more feedback on the argumentative elements of their writing and will apply the skill across many disciplines.\n\n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">'''If successful, you'll help create a better customer experience for cardholders by making it easier to be approved for a credit card. Top solutions could challenge the credit default prediction model used by the world's largest payment card issuer—earning you cash prizes, the opportunity to interview with American Express, and potentially a rewarding new career.\n\n\n\n","metadata":{}},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 15px; padding: 20px; margin: 2px;\"><b><span>💿1 | Data</span></b></div></h2>","metadata":{}},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.1 |Importing Libraries</span></b></div></h2>","metadata":{}},{"cell_type":"code","source":"import numpy as np #linear algebra\nimport pandas as pd #data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\nimport pickle\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.preprocessing import StandardScaler\nfrom sklearn.naive_bayes import GaussianNB\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\n\nfrom sklearn.model_selection import StratifiedKFold\nfrom sklearn.model_selection import GridSearchCV\nfrom sklearn.ensemble import RandomForestClassifier\nfrom sklearn.neighbors import KNeighborsClassifier\nfrom sklearn.tree import DecisionTreeRegressor\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn.pipeline import Pipeline\nfrom sklearn.linear_model import LogisticRegression\nfrom sklearn.linear_model import SGDClassifier\nfrom sklearn.discriminant_analysis import LinearDiscriminantAnalysis\nfrom sklearn.model_selection import train_test_split\nfrom sklearn.ensemble import  BaggingClassifier\nfrom sklearn.neural_network import MLPClassifier\nfrom sklearn.tree import DecisionTreeClassifier\nfrom sklearn import metrics\nfrom sklearn.feature_extraction import DictVectorizer\nimport warnings\n\nfrom matplotlib.font_manager import FontProperties\nwarnings.filterwarnings('ignore')\nfrom sklearn.preprocessing import PowerTransformer\nimport optuna\nfrom lightgbm import LGBMClassifier\nfrom sklearn.ensemble import VotingClassifier\nfrom sklearn import preprocessing\nfrom sklearn.linear_model import LogisticRegression\n\nfrom sklearn.metrics import confusion_matrix\nfrom sklearn.metrics import classification_report\nfrom sklearn.metrics import roc_auc_score\nfrom sklearn.metrics import auc\nfrom sklearn.metrics import precision_score\nfrom sklearn.metrics import recall_score\nfrom sklearn.metrics import accuracy_score\nfrom sklearn.metrics import mean_squared_error\nfrom sklearn.metrics import f1_score\nfrom sklearn.metrics import roc_curve\nfrom sklearn.metrics import plot_roc_curve","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:05:38.841048Z","iopub.execute_input":"2022-06-25T07:05:38.841725Z","iopub.status.idle":"2022-06-25T07:05:42.646431Z","shell.execute_reply.started":"2022-06-25T07:05:38.841646Z","shell.execute_reply":"2022-06-25T07:05:42.645418Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.2 |Data Description</span></b></div></h2>\n\n <div style=\"background:#1a0033   ;font-family: mono;font-size:15px;color:  #f2f2f2;padding:30px\" >\n  <p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">   \n <b style=\"font-size:30px\">T</b>he objective of this competition is to predict the probability that a customer does not pay back their credit card balance amount in the future based on their monthly customer profile. The target binary variable is calculated by observing 18 months performance window after the latest credit card statement, and if the customer does not pay due amount in 120 days after their latest statement date it is considered a default event.\n <p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">\nThe dataset contains aggregated profile features for each customer at each statement date. Features are anonymized and normalized, and fall into the following general categories:\n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">\n    <ul>\n<li><b>D_*</b> = Delinquency variables\n<li><b>S_*</b> = Spend variables\n<li><b>P_*</b> = Payment variables\n<li><b>B_*</b> = Balance variables\n<li><b>R_*</b> = Risk variables\n </ul>   \n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">  with the following features being categorical:\n <p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\"> \n <b>['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']</b>\n\n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">Your task is to predict, for each customer_ID, the probability of a future payment default (target = 1)\n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">Note that the negative class has been subsampled for this dataset at 5%, and thus receives a 20x weighting in the scoring metric.\n\n<p style=\"font-family: mono;font-size:15px;color:  #f2f2f2\">\nFiles\n<ul>\n <li> <b>  train_data.csv </b>- training data with multiple statement dates per customer_ID\n<li><b>train_labels.csv</b> - target label for each customer_ID\n<li><b>test_data.csv </b>- corresponding test data; your objective is to predict the target label for each customer_ID\n<li><b>sample_submission.csv</b> - a sample submission file in the correct format\n    </ul>\n\n","metadata":{}},{"cell_type":"markdown","source":"<body>\n\n<table style=\"width:100%\">\n  <tr>\n    <th style=\" font-size: 20px;padding:20px\", bgcolor='#6600cc'>Feature</th>\n    <th style=\" font-size: 20px\", bgcolor='#6600cc'>Description</th> \n    \n  </tr>\n  <tr>\n      <td style=\" font-size: 17px ;padding:20px \"><b>D_*</b></td>\n      <td style=\"font-size: 17px\">Delinquency Variables</td>\n    </tr>\n      <tr>\n      <td style=\" font-size: 17px;padding:20px\"><b>S_* </b></td>\n      <td style=\"font-size: 17px\">Spend Variables.</td>\n    </tr>\n          <tr>\n      <td style=\" font-size: 17px;padding:20px\"><b>P_* </b></td>\n      <td style=\"font-size: 17px\">Payment Variabels.</td>\n    </tr>\n       <tr>\n      <td style=\" font-size: 17px;padding:20px\"><b>B_*</b></td>\n      <td style=\"font-size: 17px\">Balance Variables. </td>\n    </tr>\n  <tr>\n      <td style=\" font-size: 17px;padding:20px\"><b>R_*</b></td>\n      <td style=\"font-size: 17px\">Risk Variables.</td>\n    </tr>\n\n    \n    \n</table>\n\n</body>\n","metadata":{}},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.3 |Reading Data</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"%%time\ntrain_df = pd.read_feather('../input/parquet-files-amexdefault-prediction/train_data.ftr').join(pd.read_csv('../input/amex-default-prediction/train_labels.csv')['target'])\n#test_df = pd.read_feather('../input/parquet-files-amexdefault-prediction/test_data.ftr')\n\n#sub = pd.read_csv('../input/amex-default-prediction/sample_submission.csv')","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:05:42.649747Z","iopub.execute_input":"2022-06-25T07:05:42.651381Z","iopub.status.idle":"2022-06-25T07:06:05.065796Z","shell.execute_reply.started":"2022-06-25T07:05:42.651336Z","shell.execute_reply":"2022-06-25T07:06:05.064943Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":" <div style=\"background:#1a0033   ;font-family: mono;font-size:15px;color:  white;padding:30px\" > <b>Note</b>:The dataset of this competition has a huge size. It will creeate out of memory error if we feed it. That's why we read the data from @munumbutt's AMEX-Feather-Dataset. In this Feather file, the floating point precision has been reduced from 64 bit to 16 bit. And reading a Feather file is faster than reading a csv file because the Feather file format is binary.\n\n\n    ","metadata":{}},{"cell_type":"code","source":"\ndef overview(df):\n    num_rows=len(df.index)\n    num_col=len(df.columns)\n    fig, ax = plt.subplots()\n      \n    #create values for table\n    lab = ['Number Of Rows', 'Number Of Columns']\n    table_data=[\n    [num_rows,num_col]\n        ]\n    ax.set_title('Samples', \n             fontweight =\"bold\") \n    #create table\n   \n    table = ax.table(cellText=table_data, colLabels=lab,colColours =[\"#6600cc\"] * 10, loc='center')\n    for i in enumerate(lab):\n        table[(0, i[0])].get_text().set_color('white')\n        \n    for (row, col), cell in table.get_celld().items():\n                if (row == 0):\n                    cell.set_text_props(fontproperties=FontProperties(weight='bold',size=25))    \n    #modify table\n    table.set_fontsize(14)\n    table.scale(2,4)\n    ax.axis('off')\n    #display table\n    plt.show()\n    \noverview(train_df) ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:06:05.067200Z","iopub.execute_input":"2022-06-25T07:06:05.067665Z","iopub.status.idle":"2022-06-25T07:06:05.227049Z","shell.execute_reply.started":"2022-06-25T07:06:05.067621Z","shell.execute_reply":"2022-06-25T07:06:05.225876Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"first_five=train_df.head(5)\nlast_five =train_df.tail(5)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:06:05.232619Z","iopub.execute_input":"2022-06-25T07:06:05.233016Z","iopub.status.idle":"2022-06-25T07:06:05.242515Z","shell.execute_reply.started":"2022-06-25T07:06:05.232978Z","shell.execute_reply":"2022-06-25T07:06:05.241505Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from matplotlib import pyplot as plt\nimport pandas as pd\n\nfrom matplotlib.font_manager import FontProperties\n\n# Create some example data\ndff =first_five\ndef data_show(dff,title):\n            fig, ax = plt.subplots()\n            tab2 = ax.table(cellText=dff.values, colLabels=dff.columns, loc='center', cellLoc='center',colColours =[\"#6600cc\"] * len(dff.columns))\n            tab2.auto_set_font_size(False)\n            tab2.set_fontsize(14)\n            for i in enumerate(dff.columns):\n                    tab2[(0, i[0])].get_text().set_color('white')\n\n            for (row, col), cell in tab2.get_celld().items():\n                if (row == 0):\n                    cell.set_text_props(fontproperties=FontProperties(weight='bold',size=25))\n\n            ax.set_title(title)\n            ax.axis(\"off\") \n\n            tab2.auto_set_column_width(col=list(range(len(dff.columns)))) # Provide integer list of columns to adjust\n            tab2.scale(2,8)\n            fig.tight_layout()   \n            plt.show()\n            \ndata_show(first_five,'First Five Rows')            \ndata_show(last_five,'Last Five Rows')            ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T09:46:07.852002Z","iopub.execute_input":"2022-06-25T09:46:07.852388Z","iopub.status.idle":"2022-06-25T09:46:18.530819Z","shell.execute_reply.started":"2022-06-25T09:46:07.852357Z","shell.execute_reply":"2022-06-25T09:46:18.530121Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.31 |First Five Rows</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"display(first_five)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T09:42:13.622532Z","iopub.execute_input":"2022-06-25T09:42:13.623157Z","iopub.status.idle":"2022-06-25T09:42:13.652093Z","shell.execute_reply.started":"2022-06-25T09:42:13.623121Z","shell.execute_reply":"2022-06-25T09:42:13.651195Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.32 |Last Five Rows</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"display(last_five)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T09:53:39.868481Z","iopub.execute_input":"2022-06-25T09:53:39.869015Z","iopub.status.idle":"2022-06-25T09:53:39.930903Z","shell.execute_reply.started":"2022-06-25T09:53:39.868974Z","shell.execute_reply":"2022-06-25T09:53:39.929911Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.33 |General Informations</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def info(df):\n    missing= train_df.isnull().sum()\n    percent_missing = (train_df.isnull().sum() * 100 / len(train_df)).round(2)\n    dtypes=train_df.dtypes\n    data=df\n    data=pd.DataFrame([np.array(list(train_df.columns)).T,np.array(list(missing)).T,np.array(list(percent_missing)).T,np.array(list(dtypes)).T])\n    data = data.transpose()\n    data.columns=['Features','Num of Missing values','percentage Missing','DataType']\n   \n    fig, ax = plt.subplots()\n      \n    #create values for tabl\n\n    #create table\n    ax.set_title(\"General Informations\", fontsize=40, y=27)\n    table = ax.table(cellText=data.values, colLabels=data.columns,colColours =[\"#6600cc\"] * len(data.columns), loc='center')\n    for i in enumerate(data.columns):\n                    table[(0, i[0])].get_text().set_color('white')\n\n    for (row, col), cell in table.get_celld().items():\n                if (row == 0):\n                    cell.set_text_props(fontproperties=FontProperties(weight='bold',size=25))\n\n\n    #modify table\n    table.set_fontsize(14)\n    table.scale(5,5)\n    ax.axis('off')\n    #display table\n  \n    plt.show()\ninfo(train_df)   ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:06:27.540801Z","iopub.execute_input":"2022-06-25T07:06:27.541208Z","iopub.status.idle":"2022-06-25T07:06:44.085177Z","shell.execute_reply.started":"2022-06-25T07:06:27.541173Z","shell.execute_reply":"2022-06-25T07:06:44.084437Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍1.5 |Missing Values</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"\"\"\"What are the counts of missing values in train ?\"\"\"\ndef missing(df):\n        ncounts = (pd.DataFrame([df.isnull().sum() * 100 / len(df)]).T).round(2) \n        ncounts = ncounts.rename(columns={0: \"train_missing\"})\n        margin = 0.60\n        #width = (1.-2.*margin)/len(train_df.columns)\n        ax=ncounts.sort_values('train_missing', ascending=True).plot(\n            kind=\"barh\", figsize=(20, 50), color='#8a3cf6' ,title=\"% of Values Missing\"\n        )\n\n        ax.bar_label(ax.containers[0])\n        plt.tight_layout()\n        plt.show()\n        \nmissing(train_df)        ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:06:44.086417Z","iopub.execute_input":"2022-06-25T07:06:44.086860Z","iopub.status.idle":"2022-06-25T07:06:54.502205Z","shell.execute_reply.started":"2022-06-25T07:06:44.086827Z","shell.execute_reply":"2022-06-25T07:06:54.501302Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import scipy\ndef stats(x,df,dicts):\n    data=df\n    \n    dic_data=dicts\n   # print(f\"Variable: {x}\")\n    if 'variable' not in dic_data.keys():\n        dic_data[\"variable\"]=[str(x)]\n    else :\n        dic_data[\"variable\"].append(str(x))\n    \n    #print(f\"Type of variable: {data[x].dtype}\")\n    if 'Type variable' not in dic_data.keys():\n            \n            dic_data[\"Type variable\"]=[str(data[x].dtype)]\n    else:\n        dic_data[\"Type variable\"].append(str(data[x].dtype))\n        \n    \n   # print(f\"Total observations: {data[x].shape[0]}\")\n    if 'Total Observations' not in dic_data.keys():\n                  dic_data[\"Total Observations\"]=[data[x].shape[0]]\n    else:\n        dic_data[\"Total Observations\"].append(data[x].shape[0])\n        \n    \n    detect_null_val = data[x].isnull().values.any()\n    if detect_null_val:\n       # print(f\"Missing values: {data[x].isnull().sum()} ({(data[x].isnull().sum() / data[x].isnull().shape[0] *100).round(2)}%)\")\n        if 'Missing Values' not in dic_data.keys():\n            \n                dic_data[\"Missing Values\"]=[(data[x].isnull().sum() / data[x].isnull().shape[0] *100).round(2)]\n        else:\n            dic_data[\"Missing Values\"].append((data[x].isnull().sum() / data[x].isnull().shape[0] *100).round(2))    \n    else:\n       # print(f\"Missing values? {data[x].isnull().values.any()}\")\n        if 'Missing Values' not in dic_data.keys():\n            \n                dic_data[\"Missing Values\"]=[data[x].isnull().values.any()  ]\n        else:\n            dic_data[\"Missing Values\"].append(data[x].isnull().values.any())\n            \n  #  print(f\"Unique values: {data[x].nunique()}\")\n    \n    if 'Unique_values' not in dic_data.keys():\n        \n        dic_data[\"Unique_values\"]=[data[x].nunique()]\n    else:\n        dic_data[\"Unique_values\"].append(data[x].nunique())\n        \n    if data[x].dtype != \"O\":\n       # print(f\"Min: {int(data[x].min())}\")\n        if 'Min' not in dic_data.keys():\n                dic_data[\"Min\"]=[int(data[x].min())]\n        else: dic_data[\"Min\"].append(int(data[x].min()))\n            \n      #  print(f\"25%: {int(data[x].quantile(q=[.25]).iloc[-1])}\")\n        if '25%' not in dic_data.keys():\n                dic_data[\"25%\"]=[int(data[x].quantile(q=[.25]).iloc[-1])]\n        else:\n            dic_data[\"25%\"].append(int(data[x].quantile(q=[.25]).iloc[-1]))\n            \n            \n       # print(f\"Median: {int(data[x].median())}\")\n        if 'Median' not in dic_data.keys():\n            \n            dic_data[\"Median\"]=[int(data[x].median())]\n        else:\n            dic_data[\"Median\"].append(int(data[x].median()))\n            \n            \n       # print(f\"75%: {int(data[x].quantile(q=[.75]).iloc[-1])}\")\n        if '75%' not in dic_data.keys():\n            \n            dic_data[\"75%\"]=[int(data[x].quantile(q=[.75]).iloc[-1])]\n        else:\n             dic_data[\"75%\"].append(int(data[x].quantile(q=[.75]).iloc[-1]))\n       # print(f\"Max: {int(data[x].max())}\")\n        if 'Max' not in dic_data.keys():\n            dic_data[\"Max\"]=[int(data[x].max())]\n        else:\n            dic_data[\"Max\"].append(int(data[x].max()))\n       # print(f\"Mean: {data[x].mean()}\")\n        if 'Mean' not in dic_data.keys():\n            \n                dic_data[\"Mean\"]=[data[x].mean()]\n        else:\n            dic_data[\"Mean\"].append(data[x].mean())\n            \n        #print(f\"Std dev: {data[x].std()}\")\n        if 'Std dev' not in dic_data.keys():\n            \n                dic_data[\"Std dev\"]=[data[x].std()]\n                \n        else:\n             dic_data[\"Std dev\"].append(data[x].std())\n                \n            \n        #print(f\"Variance: {data[x].var()}\")\n        if 'Variance' not in dic_data.keys():\n                dic_data[\"Variance\"]=[data[x].var()]\n        else:\n            dic_data[\"Variance\"].append(data[x].var())\n            \n       # print(f\"Skewness: {scipy.stats.skew(data[x])}\")\n        \n        if 'Skewness' not in dic_data.keys():\n            dic_data[\"Skewness\"]=[scipy.stats.skew(data[x])]\n        else:\n            dic_data[\"Skewness\"].append(scipy.stats.skew(data[x]))\n            \n       # print(f\"Kurtosis: {scipy.stats.kurtosis(data[x])}\")\n        if 'Kurtosis' not in dic_data.keys():\n            \n            dic_data[\"Kurtosis\"]=[scipy.stats.kurtosis(data[x])]\n        else:\n            dic_data[\"Kurtosis\"].append(scipy.stats.kurtosis(data[x]))\n            \n       # print(\"\")\n        \n        # Percentiles 1%, 5%, 95% and 99%\n       # print(\"Percentiles 1%, 5%, 95%, 99%\")\n        \n        for x,y  in  zip(['1%', '5%', '95%', '99%'],data[x].quantile(q=[.01, .05, .95, .99])):\n                if f'Percentile {x}' not in dic_data.keys():\n                    dic_data[f'Percentile {x}']=[int(y)]\n                else:\n                    dic_data[f'Percentile {x}'].append(int(y))","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:06:54.503760Z","iopub.execute_input":"2022-06-25T07:06:54.504158Z","iopub.status.idle":"2022-06-25T07:06:54.533169Z","shell.execute_reply.started":"2022-06-25T07:06:54.504119Z","shell.execute_reply":"2022-06-25T07:06:54.532381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"num_col=[]\nfor col in train_df.columns:\n    \n    \n    if train_df[col].dtype =='float16' or train_df[col].dtype =='int16':\n          num_col.append(col)\ndef desc(df,title):\n   \n    fig, ax = plt.subplots()\n      \n    #create values for tabl\n\n    #create table\n    ax.set_title(title, fontsize=40, y=1.01)\n    table = ax.table(cellText=df.values, colLabels=df.columns,colColours =[\"#6600cc\"] * len(df.columns), loc='center')\n    for i in enumerate(range(19)):\n                    table[(0, i[0])].get_text().set_color('white')\n\n    for (row, col), cell in table.get_celld().items():\n                if (row == 0):\n                    cell.set_text_props(fontproperties=FontProperties(weight='bold',size=50))\n    for (row, col), cell in table.get_celld().items():\n                     \n                    cell.set_text_props(fontproperties=FontProperties(size=50))\n\n    #modify table\n    #table.set_fontsize(50)\n    table.scale(10,10)\n    ax.axis('off')\n    #display table\n  \n    plt.show()\n    \nfloat_d={} \n\nfor x in num_col:\n    \n     stats(x,train_df,float_d) \n\n\n    \ndesc(pd.DataFrame(float_d),'Feature Descriptions for Float ')","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:06:54.534555Z","iopub.execute_input":"2022-06-25T07:06:54.535100Z","iopub.status.idle":"2022-06-25T07:15:14.345942Z","shell.execute_reply.started":"2022-06-25T07:06:54.535056Z","shell.execute_reply":"2022-06-25T07:15:14.345228Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 15px; padding: 20px; margin: 2px;\"><b><span>💿2 | EDA</span></b></div></h2>","metadata":{}},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.1 |Target Distribution</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def pie_target(df,col,title):\n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]    \n            fig, ax = plt.subplots(1,2,figsize=(16, 8))\n            fig.suptitle(title, size = 20)\n            labels = list(df[col].value_counts().index)\n            values = df[col].value_counts()\n            ax[0].pie( values,colors=colors[:2],explode=(.05,0),startangle=60, labels=labels,autopct='%1.0f%%', pctdistance=0.6)\n           \n            sns.countplot(x=col, data=df, hue=col,palette=colors[:2], ax=ax[1])\n\n            ax[0].add_artist(plt.Circle((0,0),0.4,fc='white'))\n            plt.show()\n            \npie_target(train_df,'target','target Distrubtion')            ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:15:14.348795Z","iopub.execute_input":"2022-06-25T07:15:14.349297Z","iopub.status.idle":"2022-06-25T07:15:15.211191Z","shell.execute_reply.started":"2022-06-25T07:15:14.349261Z","shell.execute_reply":"2022-06-25T07:15:15.210448Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"cat_var= ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\nobj_col=['customer_ID', 'S_2']\nimport re\nDelinquency_var=[]\nSpend_var=[]\nPayment_var=[]\nBalance_var=[]\nRisk_var=[]\nfor pattern in train_df.columns:\n     if re.findall(\"\\AD\", pattern) and (pattern not in (cat_var+obj_col)):\n            Delinquency_var.append(pattern)\n     elif re.findall(\"\\AS_\", pattern)  and (pattern not in (cat_var+obj_col)) :\n        Spend_var.append(pattern)\n     elif re.findall(\"\\AP_\", pattern)  and (pattern not in (cat_var+obj_col)) :\n        Payment_var.append(pattern)\n     elif re.findall(\"\\AB_\", pattern)  and (pattern not in (cat_var+obj_col)) :\n        Balance_var.append(pattern)\n     elif re.findall(\"\\AR_\", pattern)  and (pattern not in (cat_var+obj_col)):\n        Risk_var.append(pattern)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:15:15.212418Z","iopub.execute_input":"2022-06-25T07:15:15.212846Z","iopub.status.idle":"2022-06-25T07:15:15.223008Z","shell.execute_reply.started":"2022-06-25T07:15:15.212808Z","shell.execute_reply":"2022-06-25T07:15:15.222265Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.2 |Delinquency Variable Distribution</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def  kde_plot(feat,df):\n    fig, ax = plt.subplots(18,5,figsize=(16,54))\n    row=0\n    col=[0,1,2,3,4]*18\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('Count Plot of Delinquency Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Count Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n           # ax[i[0],1].title.set_text(f'Box Plot : {i[1]}')\n            sns.kdeplot(x=i[1], palette=colors, data=df, fill=True, linewidth=2, legend=False, ax= ax[row,col[i[0]]])\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nkde_plot(Delinquency_var,train_df)\n\ndef  box_plot(feat,df):\n    fig, ax = plt.subplots(18,5,figsize=(16,54))\n    row=0\n    col=[0,1,2,3,4]*18\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('Box Delinquency Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Box Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n           # sns.histplot(x=i[1], element=\"step\", kde=True,palette=colors, data=df, ax=ax[i[0],0])\n            sns.boxplot( y=i[1], data=df, ax= ax[row,col[i[0]]] , palette=colors)\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nbox_plot(Delinquency_var,train_df)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:15:15.224424Z","iopub.execute_input":"2022-06-25T07:15:15.224921Z","iopub.status.idle":"2022-06-25T07:36:34.049125Z","shell.execute_reply.started":"2022-06-25T07:15:15.224882Z","shell.execute_reply":"2022-06-25T07:36:34.048441Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"import ctypes\nlibc = ctypes.CDLL(\"libc.so.6\") # clearing cache \nlibc.malloc_trim(0)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:36:34.050216Z","iopub.execute_input":"2022-06-25T07:36:34.051513Z","iopub.status.idle":"2022-06-25T07:36:34.067266Z","shell.execute_reply.started":"2022-06-25T07:36:34.051472Z","shell.execute_reply":"2022-06-25T07:36:34.066381Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.21 |Delinquency Variable Distribution WRT Target</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"\nimport random\nplt.figure(figsize = (16,54))\ndef rel_tar(df,feat_list,target):\n    \n        for i in enumerate(feat_list):\n                colors = [ \"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\", '#967032', '#2734DE'] \n                rand_col = colors[random.sample(range(6), 1)[0]]\n                plt.subplot(18,5,i[0]+1)\n                sns.kdeplot(data = df, x = i[1], hue = target, palette=colors[:2])\n             \n                plt.title (i[1]+ f' vs {target}')\n                plt.xlabel(\" \")\n                plt.ylabel(\" \")\n                if i[1] != 'Age':\n\n                        plt.xlim([-1,1])\n                plt.xticks(rotation = 45)\n                plt.tight_layout()\n                \n                \nrel_tar(train_df,Delinquency_var,'target')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:36:34.068496Z","iopub.execute_input":"2022-06-25T07:36:34.068928Z","iopub.status.idle":"2022-06-25T07:41:10.402832Z","shell.execute_reply.started":"2022-06-25T07:36:34.068890Z","shell.execute_reply":"2022-06-25T07:41:10.402101Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.3 |Payment Variables Distribution</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def  kde_plot(feat,df):\n    fig, ax = plt.subplots(2,3,figsize=(8,8))\n    row=0\n    col=[0,1,2]\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('KDE plot for Payement Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Count Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n           # ax[i[0],1].title.set_text(f'Box Plot : {i[1]}')\n            sns.kdeplot(x=i[1], palette=colors, data=df, fill=True, linewidth=2, legend=False, ax= ax[row,col[i[0]]])\n\n            \n    fig.tight_layout()        \n    plt.show()\n\nkde_plot(Payment_var,train_df)\n\ndef  box_plot(feat,df):\n    fig, ax = plt.subplots(2,3,figsize=(8,8))\n    row=0\n    col=[0,1,2]\n   \n    for i in enumerate(feat):\n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('Box Plot of Payment Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Box Plot : {i[1]}')\n            # labels = list(df[i[1]].value_counts().index)\n            # values = df[i[1]].value_counts()\n            sns.boxplot( y=i[1], data=df, ax= ax[row,col[i[0]]] , palette=colors)\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nbox_plot(Payment_var,train_df)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:41:10.404157Z","iopub.execute_input":"2022-06-25T07:41:10.404750Z","iopub.status.idle":"2022-06-25T07:42:08.937677Z","shell.execute_reply.started":"2022-06-25T07:41:10.404714Z","shell.execute_reply":"2022-06-25T07:42:08.936744Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.31 |Payment Variables Distribution WRT Target</span></b></div></h2>\n","metadata":{"execution":{"iopub.status.busy":"2022-06-24T08:38:24.062433Z","iopub.execute_input":"2022-06-24T08:38:24.063218Z","iopub.status.idle":"2022-06-24T08:38:24.079299Z","shell.execute_reply.started":"2022-06-24T08:38:24.063183Z","shell.execute_reply":"2022-06-24T08:38:24.078481Z"}}},{"cell_type":"code","source":"\nimport random\nplt.figure(figsize = (8,8))\ndef rel_tar(df,feat_list,target):\n    \n        for i in enumerate(feat_list):\n                colors = [ \"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\", '#967032', '#2734DE'] \n                rand_col = colors[random.sample(range(6), 1)[0]]\n                plt.subplot(2,3,i[0]+1)\n                sns.kdeplot(data = df, x = i[1], hue = target, palette=colors[:2])\n             \n                plt.title (i[1]+ f' vs {target}')\n                plt.xlabel(\" \")\n                plt.ylabel(\" \")\n                if i[1] != 'Age':\n\n                        plt.xlim([-1,1])\n                plt.xticks(rotation = 45)\n                plt.tight_layout()\n                \n                \nrel_tar(train_df,Payment_var,'target')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:42:08.939053Z","iopub.execute_input":"2022-06-25T07:42:08.939495Z","iopub.status.idle":"2022-06-25T07:42:18.075699Z","shell.execute_reply.started":"2022-06-25T07:42:08.939455Z","shell.execute_reply":"2022-06-25T07:42:18.074813Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.4 |Risk Variables Distribution</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def  kde_plot(feat,df):\n    fig, ax = plt.subplots(6,5,figsize=(16,16))\n    row=0\n    col=[0,1,2,3,4]*6\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('KDE risk Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Count Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n            sns.kdeplot(x=i[1], palette=colors, data=df, fill=True, linewidth=2, legend=False, ax= ax[row,col[i[0]]])\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n        \n            \n    fig.tight_layout()        \n    plt.show()\n\nkde_plot(Risk_var,train_df)\n\ndef  box_plot(feat,df):\n    fig, ax = plt.subplots(6,5,figsize=(16,16))\n    row=0\n    col=[0,1,2,3,4]*6\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('Box Plot of Risk Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Box Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n           # ax[i[0],1].title.set_text(f'Box Plot : {i[1]}')\n            sns.boxplot( y=i[1], data=df, ax= ax[row,col[i[0]]] , palette=colors)\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nbox_plot(Risk_var,train_df)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:42:18.077037Z","iopub.execute_input":"2022-06-25T07:42:18.077407Z","iopub.status.idle":"2022-06-25T07:50:07.351977Z","shell.execute_reply.started":"2022-06-25T07:42:18.077370Z","shell.execute_reply":"2022-06-25T07:50:07.351254Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.41 |Risk Variables Distribution WRT Target</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize = (32,32))\ndef rel_tar(df,feat_list,target):\n    \n        for i in enumerate(feat_list):\n                colors = [ \"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\", '#967032', '#2734DE'] \n                rand_col = colors[random.sample(range(6), 1)[0]]\n                plt.subplot(6,5,i[0]+1)\n                sns.kdeplot(data = df, x = i[1], hue = target, palette=colors[:2])\n             \n                plt.title (i[1]+ f' vs {target}')\n                plt.xlabel(\" \")\n                plt.ylabel(\" \")\n                if i[1] != 'Age':\n\n                        plt.xlim([-1,1])\n                plt.xticks(rotation = 45)\n                plt.tight_layout()\n                \n                \nrel_tar(train_df,Risk_var,'target')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:50:07.353448Z","iopub.execute_input":"2022-06-25T07:50:07.353834Z","iopub.status.idle":"2022-06-25T07:51:31.745819Z","shell.execute_reply.started":"2022-06-25T07:50:07.353795Z","shell.execute_reply":"2022-06-25T07:51:31.745079Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.5 |Spent Variable Distribution</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def  kde_plot(feat,df):\n    fig, ax = plt.subplots(5,5,figsize=(16,16))\n    row=0\n    col=[0,1,2,3,4]*5\n   \n    for i in enumerate(feat):\n           \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('KDE plot for Spent Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Count Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n           # ax[i[0],1].title.set_text(f'Box Plot : {i[1]}')\n            sns.kdeplot(x=i[1], palette=colors, data=df, fill=True, linewidth=2, legend=False, ax= ax[row,col[i[0]]])\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nkde_plot(Spend_var,train_df)\n\ndef  box_plot(feat,df):\n    fig, ax = plt.subplots(5,5,figsize=(16,16))\n    row=0\n    col=[0,1,2,3,4]*5\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('Box Plot of Spend Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Box Plot : {i[1]}')\n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n           # ax[i[0],1].title.set_text(f'Box Plot : {i[1]}')\n            sns.boxplot( y=i[1], data=df, ax= ax[row,col[i[0]]] , palette=colors)\n            if (i[0]!=0)&(i[0]%5==0):\n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nbox_plot(Spend_var,train_df)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:51:31.747110Z","iopub.execute_input":"2022-06-25T07:51:31.747947Z","iopub.status.idle":"2022-06-25T07:57:47.289577Z","shell.execute_reply.started":"2022-06-25T07:51:31.747911Z","shell.execute_reply":"2022-06-25T07:57:47.288731Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.51 |Spend Variables Distribution WRT Target</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize = (32,32))\ndef rel_tar(df,feat_list,target):\n    \n        for i in enumerate(feat_list):\n                colors = [ \"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\", '#967032', '#2734DE'] \n                rand_col = colors[random.sample(range(6), 1)[0]]\n                plt.subplot(5,5,i[0]+1)\n                sns.kdeplot(data = df, x = i[1], hue = target, palette=colors[:2])\n             \n                plt.title (i[1]+ f' vs {target}')\n                plt.xlabel(\" \")\n                plt.ylabel(\" \")\n                if i[1] != 'Age':\n\n                        plt.xlim([-1,1])\n                plt.xticks(rotation = 45)\n                plt.tight_layout()\n                \n                \nrel_tar(train_df,Spend_var,'target')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:57:47.290740Z","iopub.execute_input":"2022-06-25T07:57:47.291668Z","iopub.status.idle":"2022-06-25T07:58:56.381251Z","shell.execute_reply.started":"2022-06-25T07:57:47.291626Z","shell.execute_reply":"2022-06-25T07:58:56.380525Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.6 |Balance Variables Distribution</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"def  kde_plot(feat,df):\n    fig, ax = plt.subplots(8,5,figsize=(20,20))\n    row=0\n    col=[0,1,2,3,4]*8\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('KDE plot of Balance Variables', size = 29,y=1.01)\n            \n            ax[row,col[i[0]]].title.set_text(f'Count Plot : {i[1]}')\n            sns.kdeplot(x=i[1], palette=colors, data=df, fill=True, linewidth=2, legend=False, ax= ax[row,col[i[0]]])\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n            \n    fig.tight_layout()        \n    plt.show()\n\nkde_plot(Balance_var,train_df)\n\ndef  Box_plot(feat,df):\n    fig, ax = plt.subplots(8,5,figsize=(20,20))\n    row=0\n    col=[0,1,2,3,4]*8\n   \n    for i in enumerate(feat):\n        \n            \n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]\n            fig.suptitle('Box Plot of BAlance  Variables', size = 29,y=1.01)\n            ax[row,col[i[0]]].title.set_text(f'Box Plot : {i[1]}')\n            \n           # labels = list(df[i[1]].value_counts().index)\n           # values = df[i[1]].value_counts()\n            sns.boxplot( y=i[1], data=df, ax= ax[row,col[i[0]]] , palette=colors)\n            if (i[0]!=0)&(i[0]%5==0): \n                       row+=1\n            \n    fig.tight_layout()        \n    plt.show()\n\nBox_plot(Balance_var,train_df)","metadata":{"execution":{"iopub.status.busy":"2022-06-25T07:58:56.382523Z","iopub.execute_input":"2022-06-25T07:58:56.383023Z","iopub.status.idle":"2022-06-25T08:09:42.159117Z","shell.execute_reply.started":"2022-06-25T07:58:56.382982Z","shell.execute_reply":"2022-06-25T08:09:42.158205Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"![](http://)# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.61Balance  Variables Distribution WRT Target</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"\nplt.figure(figsize = (32,32))\ndef rel_tar(df,feat_list,target):\n    \n        for i in enumerate(feat_list):\n                colors = [ \"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\", '#967032', '#2734DE'] \n                rand_col = colors[random.sample(range(6), 1)[0]]\n                plt.subplot(8,5,i[0]+1)\n                sns.kdeplot(data = df, x = i[1], hue = target, palette=colors[:2])\n             \n                plt.title (i[1]+ f' vs {target}')\n                plt.xlabel(\" \")\n                plt.ylabel(\" \")\n                if i[1] != 'Age':\n\n                        plt.xlim([-1,1])\n                plt.xticks(rotation = 45)\n                plt.tight_layout()\n                \n                \nrel_tar(train_df,Balance_var,'target')    ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T08:09:42.160574Z","iopub.execute_input":"2022-06-25T08:09:42.160928Z","iopub.status.idle":"2022-06-25T08:11:48.155379Z","shell.execute_reply.started":"2022-06-25T08:09:42.160896Z","shell.execute_reply":"2022-06-25T08:11:48.154540Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.7 |Categorical Feautres</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"cat_var= ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\nobj_col=['customer_ID', 'S_2']","metadata":{"execution":{"iopub.status.busy":"2022-06-25T08:11:48.156898Z","iopub.execute_input":"2022-06-25T08:11:48.157269Z","iopub.status.idle":"2022-06-25T08:11:48.162072Z","shell.execute_reply.started":"2022-06-25T08:11:48.157221Z","shell.execute_reply":"2022-06-25T08:11:48.161277Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"def pie_target(feat,df):\n    fig, ax = plt.subplots(11,2,figsize=(18,60))\n    for i in enumerate(feat):\n            colors = [\"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\"]    \n            fig.suptitle('Pie chart and Count Plot', size = 29,y=1.01)\n            ax[i[0],0].title.set_text(f'Pie of {i[1]}')\n            labels = list(df[i[1]].value_counts().index)\n            values = df[i[1]].value_counts()\n            plt.figure(facecolor='red') \n            ax[i[0],0].pie( values,colors=colors,startangle=60, labels=labels,autopct='%1.0f%%', pctdistance=0.6)\n           \n            ax[i[0],1].title.set_text(f'Box Plot : {i[1]}')\n        \n            sns.countplot(x=i[1],data=df,palette=colors ,ax=ax[i[0],1])\n            \n            ax[i[0],0].add_artist(plt.Circle((0,0),0.4,fc='white'))\n    \n    fig.tight_layout()        \n    plt.show()\n\n\npie_target(cat_var,train_df)\n    ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T08:32:18.007981Z","iopub.execute_input":"2022-06-25T08:32:18.008510Z","iopub.status.idle":"2022-06-25T08:32:22.722373Z","shell.execute_reply.started":"2022-06-25T08:32:18.008472Z","shell.execute_reply":"2022-06-25T08:32:22.721532Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"# <h2 ><div style = \"background-color: #6600cc; color:white; border-radius: 30px 5px; padding: 20px; margin: 2px;\"><b><span>📍2.71 |Categorical Feautres WRT Target</span></b></div></h2>\n","metadata":{}},{"cell_type":"code","source":"\ncat_var= ['B_30', 'B_38', 'D_114', 'D_116', 'D_117', 'D_120', 'D_126', 'D_63', 'D_64', 'D_66', 'D_68']\nplt.figure(figsize = (18,18))\ndef rel_tar(df,feat_list,target):\n    \n        for i in enumerate(feat_list):\n                colors = [ \"#570990\",\"#e4b6fe\",'#8b22ba', \"#8a3cf6\", '#967032', '#2734DE'] \n                rand_col = colors[random.sample(range(6),1)[0]]\n                plt.subplot(4,4,i[0]+1)\n                sns.countplot(x=i[1], data=df, hue=target,palette=colors)\n                plt.title (i[1]+f' vs {target}')\n                plt.xlabel(\" \")\n                plt.ylabel(\" \")\n                plt.xticks(rotation = 45)\n                plt.tight_layout()\n                \n                \nrel_tar(train_df,cat_var,'target')        ","metadata":{"execution":{"iopub.status.busy":"2022-06-25T08:32:31.726063Z","iopub.execute_input":"2022-06-25T08:32:31.726676Z","iopub.status.idle":"2022-06-25T08:32:37.036130Z","shell.execute_reply.started":"2022-06-25T08:32:31.726640Z","shell.execute_reply":"2022-06-25T08:32:37.035433Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<div style=\"background:#7300e6   ;font-family:'Times';font-size:35px;color:  white\" ><center>&ensp;Thank you</center></div>\n<div style=\"background:#7300e6   ;font-family:'Times';font-size:35px;color:  white\" ><center>&ensp;⚠ WORK IN PROGRESS ⚠\n<br>Please consider upvoting the kernel if you found it useful.</center></div>\n","metadata":{}},{"cell_type":"code","source":"","metadata":{"execution":{"iopub.status.busy":"2022-06-25T09:41:46.907147Z","iopub.execute_input":"2022-06-25T09:41:46.908115Z","iopub.status.idle":"2022-06-25T09:41:47.933037Z","shell.execute_reply.started":"2022-06-25T09:41:46.908066Z","shell.execute_reply":"2022-06-25T09:41:47.931845Z"},"trusted":true},"execution_count":null,"outputs":[]}]}