{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"name":"python","version":"3.10.13","mimetype":"text/x-python","codemirror_mode":{"name":"ipython","version":3},"pygments_lexer":"ipython3","nbconvert_exporter":"python","file_extension":".py"},"kaggle":{"accelerator":"none","dataSources":[{"sourceId":50160,"databundleVersionId":7602123,"sourceType":"competition"}],"dockerImageVersionId":30646,"isInternetEnabled":true,"language":"python","sourceType":"notebook","isGpuEnabled":false}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"markdown","source":"Published on February 06, 2024. By Marília Prata, mpwolke","metadata":{}},{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\nimport matplotlib.pyplot as plt\nimport seaborn as sns\n\nimport plotly.graph_objs as go\nimport plotly.offline as py\nimport plotly.express as px\n\n#Ignore warnings\nimport warnings\nwarnings.filterwarnings('ignore')\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-02-06T03:00:35.774275Z","iopub.execute_input":"2024-02-06T03:00:35.774665Z","iopub.status.idle":"2024-02-06T03:00:39.100701Z","shell.execute_reply.started":"2024-02-06T03:00:35.774632Z","shell.execute_reply":"2024-02-06T03:00:39.099363Z"},"collapsed":true,"jupyter":{"outputs_hidden":true},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Credit Scoring Algorithms: Cracking the Code for a Better Score\n\n![](https://fastercapital.com/i/Credit-Scoring-Algorithms--Cracking-the-Code-for-a-Better-Score--The-Factors-That-Affect-Your-Credit-Score.webp)https://fastercapital.com/content/Credit-Scoring-Algorithms--Cracking-the-Code-for-a-Better-Score.html","metadata":{}},{"cell_type":"markdown","source":"#Artificial intelligence/machine learning based credit scoring model\n\nOn \"Credit scoring models: Techniques and issues - Author: Nazri Engku\n\nJournal of Advanced Research in Business and Management Studies\nJournal homepage:www.akademiabaru.com/arbms.html\nISSN: 2462-1935\n\nARTIFICIAL INTELLIGENCE/ML BASED CREDIT SCORING MODEL \n\n\"Some of the based methods being suggested and explored by researchers are artificial neural networks, genetic algorithms , and artificial immune system.\"\n\n\"The techniques used here are broadly called black boxes in the analytics world because interpreting them is difficult. Banks generally use this type of scoring model for upselling or cross-selling different products of a bank to its customers. These techniques usually outperform the statistical-based credit scoring models but fall behind because of their interpretability issues.\"\n\n\"Logistic regression is superior to other methods in predicting defaults It is easy to explain, very tractable, convenient, most practical and favorable technique in practice. However, it was suggested that neural network is equally superior as its overall predictive ability is comparativelly high. Nevertheless, logit model produces slightly lower type I error rates i.e. error in classifying bad loan as a good loan with an average of score of 16% compared to 17% for the neural networks.\"\n\nSTATISTICAL-BASED CREDIT SCORING MODELS\n\n\"There are various statistical-based credit scoring model that have been introduced such as linear discriminant analysis , decision trees , Markov chain analysis, probit analysis and logistic regression.\"\n\nhttps://www.academia.edu/34562267","metadata":{}},{"cell_type":"markdown","source":"#Base Csv file ","metadata":{}},{"cell_type":"code","source":"df = pd.read_csv('../input/home-credit-credit-risk-model-stability/csv_files/train/train_base.csv')\ndf.tail()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:00:45.796853Z","iopub.execute_input":"2024-02-06T03:00:45.797639Z","iopub.status.idle":"2024-02-06T03:00:47.270895Z","shell.execute_reply.started":"2024-02-06T03:00:45.797591Z","shell.execute_reply":"2024-02-06T03:00:47.269550Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Decision-Making Trial and Evaluation (DEMATEL)\n\n\"Decision-making trial and evaluation laboratory technique (DEMATEL) can be used to identify the causal-effect relations between the candidate criteria to be included in the credit scoring/risk model. Each decision maker is requested to specify the direct influence between any two criteria based on a scale consisting of 0, 1 ,2, 3, and 4 representing \"no influence\", \"low influence\", \"medium influence\", \"high influence\", and \"very high influence\", respectively. then excute steps that include: calculating the directed influenced matrix normalization and produce the total-relation matrix.\"\n\nhttps://www.academia.edu/34562267","metadata":{}},{"cell_type":"markdown","source":"#Distribution of the Target Column","metadata":{}},{"cell_type":"code","source":"df['target'].value_counts()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:14:20.077033Z","iopub.execute_input":"2024-02-06T01:14:20.077485Z","iopub.status.idle":"2024-02-06T01:14:20.102081Z","shell.execute_reply.started":"2024-02-06T01:14:20.077452Z","shell.execute_reply":"2024-02-06T01:14:20.101209Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Determining the weight of Criteria using AHP (Analytic Hierarchy Process)\n\n\"Among the techniques to determine weights are weight-of-evidence and information value technique, ELECTRE and analytic hierarchy process (AHP). Among those techniques, AHP is the most widely used technique. AHP is a technique that simplifies a complex problem by means of hierarchical analysis methodology, which enables subjective judgments among different criteria. One major problem with the AHP process is the consistency of the pairwise comparison matrices. To address this problem a proposed pre-Likert scale-AHP manages to solve the issue.\"\n\n\" The statistical-based techniques are still the methods of choice for bankers. Among the techniques, the most popular one is the logistic regression. However, the variables used must be carefully selected and the weights given to the variables must be carefully determined. Unfortunately, the model developed can only be verified with the availability of previous historical data. If the data are not available, then one suitable method of choice to determine the relevant criteria and the suitable weights for the criteria will be through the combination of DEMATEL and pre-Likert scale AHP whereby the judgments will be executed by experts who are directly involved in performing this credit screening task. \"\n\nhttps://www.academia.edu/34562267","metadata":{}},{"cell_type":"code","source":"df['target'].astype(int).plot.hist();","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:14:27.356616Z","iopub.execute_input":"2024-02-06T01:14:27.357591Z","iopub.status.idle":"2024-02-06T01:14:27.681029Z","shell.execute_reply.started":"2024-02-06T01:14:27.357555Z","shell.execute_reply":"2024-02-06T01:14:27.679630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Trying a cool version bur Don't forget Random. It's almost like math","metadata":{}},{"cell_type":"code","source":"#By Youhan Lee  https://www.kaggle.com/code/youhanlee/extensive-eda-for-application-and-bureau-data\n\nimport random\n\ndef random_color_generator(number_of_colors):\n    color = [\"#\"+''.join([random.choice('0123456789ABCDEF') for j in range(6)])\n                 for i in range(number_of_colors)]\n    return color","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:27:59.167435Z","iopub.execute_input":"2024-02-06T01:27:59.167947Z","iopub.status.idle":"2024-02-06T01:27:59.175728Z","shell.execute_reply.started":"2024-02-06T01:27:59.167910Z","shell.execute_reply":"2024-02-06T01:27:59.174046Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#FICO score\n\nThe FICO score was first introduced in 1989 by FICO, then called Fair, Isaac, and Company. The FICO model is used by the vast majority of banks and credit grantors, and is based on consumer credit files of the three national credit bureaus: Experian, Equifax, and TransUnion. Because a consumer's credit file may contain different information at each of the bureaus, FICO scores can vary depending on which bureau provides the information to FICO to generate the score.\"\n\nhttps://en.wikipedia.org/wiki/Credit_score_in_the_United_States#FICO_score\n\n![](https://www.credit.com/wp-content/uploads/2020/10/Credit-Score-Charts-1024x693.png)Credit.com","metadata":{}},{"cell_type":"code","source":"#By Youhan Lee  https://www.kaggle.com/code/youhanlee/extensive-eda-for-application-and-bureau-data\n\ncnt_srs = df['target'].value_counts()\ntext = ['{:.2f}%'.format(100 * (value / cnt_srs.sum())) for value in cnt_srs.values]\n\ntrace = go.Bar(\n    x = cnt_srs.index,\n    y = (cnt_srs / cnt_srs.sum()) * 100,\n    marker = dict(\n        color = random_color_generator(2),\n        line = dict(color='rgb(8, 48, 107)',\n                   width = 1.5\n                   )\n    ), \n    opacity = 0.7\n)\n\ndata = [trace]\n\nlayout = go.Layout(\n    title = 'Target distribution(%)',\n    margin = dict(\n        l = 100\n    ),\n    xaxis = dict(\n        title = 'Labels (0: repay, 1: not repay)'\n    ),\n    yaxis = dict(\n        title = 'Account(%)'\n    ),\n    width=800,\n    height=500\n)\nannotations = []\nfor i in range(2):\n    annotations.append(dict(\n        x = cnt_srs.index[i],\n        y = ((cnt_srs / cnt_srs.sum()) * 100)[i],\n        text = text[i],\n        font = dict(\n            family = 'Arial',\n            size = 14,\n        ),\n        showarrow = True\n    ))\n    layout['annotations'] = annotations\n\nfig = go.Figure(data=data, layout=layout)\npy.iplot(fig)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:28:04.123631Z","iopub.execute_input":"2024-02-06T01:28:04.124090Z","iopub.status.idle":"2024-02-06T01:28:05.961786Z","shell.execute_reply.started":"2024-02-06T01:28:04.124053Z","shell.execute_reply":"2024-02-06T01:28:05.960828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#What an Imbalanced Class above!","metadata":{}},{"cell_type":"markdown","source":"#No Missing Values on the base file","metadata":{}},{"cell_type":"code","source":"df.isnull().sum()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:17:01.244573Z","iopub.execute_input":"2024-02-06T01:17:01.245075Z","iopub.status.idle":"2024-02-06T01:17:01.424911Z","shell.execute_reply.started":"2024-02-06T01:17:01.245040Z","shell.execute_reply":"2024-02-06T01:17:01.422326Z"},"_kg_hide-output":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Puru Behl https://www.kaggle.com/accountstatus/mt-cars-data-analysis\n\nsns.distplot(df['target'])\nplt.axvline(df['target'].values.mean(), color='red', linestyle='dashed', linewidth=1)\nplt.title('Target Distribution');","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:53:25.575024Z","iopub.execute_input":"2024-02-06T01:53:25.575566Z","iopub.status.idle":"2024-02-06T01:53:32.940856Z","shell.execute_reply.started":"2024-02-06T01:53:25.575520Z","shell.execute_reply":"2024-02-06T01:53:32.939855Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Firat Gonen https://www.kaggle.com/frtgnn/elo-eda-lgbm/notebook \n\nplt.figure(figsize=(10, 6))\nplt.title('Month Distribution')\nsns.despine()\nsns.set_context(\"notebook\", font_scale=1.5, rc={\"lines.linewidth\": 2.5})\n\nsns.distplot(df['MONTH'], hist=True, rug=False,norm_hist=True, color='brown');","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:54:29.140393Z","iopub.execute_input":"2024-02-06T01:54:29.140789Z","iopub.status.idle":"2024-02-06T01:54:36.622021Z","shell.execute_reply.started":"2024-02-06T01:54:29.140759Z","shell.execute_reply":"2024-02-06T01:54:36.620545Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by Puru Behl https://www.kaggle.com/accountstatus/mt-cars-data-analysis\n\nsns.distplot(df['WEEK_NUM'], color = 'g')\nplt.axvline(df['WEEK_NUM'].values.mean(), color='red', linestyle='dashed', linewidth=1)\nplt.title('Week Number Distribution');","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:57:07.505661Z","iopub.execute_input":"2024-02-06T01:57:07.506100Z","iopub.status.idle":"2024-02-06T01:57:14.459103Z","shell.execute_reply.started":"2024-02-06T01:57:07.506068Z","shell.execute_reply":"2024-02-06T01:57:14.458183Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Mustafa Yildirim https://www.kaggle.com/code/homydata/accurate-estimation-of-home-credit-default-risk-by\n\nfig, ax = plt.subplots(figsize=(15,9))\nsns.boxplot(x='target',y='WEEK_NUM',data=df);","metadata":{"execution":{"iopub.status.busy":"2024-02-06T02:10:50.882746Z","iopub.execute_input":"2024-02-06T02:10:50.883194Z","iopub.status.idle":"2024-02-06T02:10:51.314147Z","shell.execute_reply.started":"2024-02-06T02:10:50.883163Z","shell.execute_reply":"2024-02-06T02:10:51.312833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"df.dtypes","metadata":{"execution":{"iopub.status.busy":"2024-02-06T02:07:01.981176Z","iopub.execute_input":"2024-02-06T02:07:01.981659Z","iopub.status.idle":"2024-02-06T02:07:01.994284Z","shell.execute_reply.started":"2024-02-06T02:07:01.981625Z","shell.execute_reply":"2024-02-06T02:07:01.992696Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Furkan Akdag https://www.kaggle.com/code/furkannakdagg/smoking-drinking-prediction-complete-eda-pycaret/notebook\n\n#Pick numerical variables to make the correlation matrix\n\ninterp = df[['case_id', 'MONTH', 'WEEK_NUM', 'target']]\n\n\ndef corr_map(interp, width=8, height=6, annot_kws=15):\n    mtx = np.triu(interp.corr())\n    f, ax = plt.subplots(figsize = (width,height))\n    sns.heatmap(interp.corr(),\n                annot= True,\n                fmt = \".2f\",\n                ax=ax,\n                vmin = -1,\n                vmax = 1,\n                cmap = \"summer\",\n                mask = mtx,\n                linewidth = 0.4,\n                linecolor = \"black\",\n                cbar=False,\n                annot_kws={\"size\": annot_kws})\n    plt.yticks(rotation=0,size=15)\n    plt.xticks(rotation=75,size=15)\n    plt.title('\\nCorrelation of the Credit Risk Base data\\n', size = 20)\n    plt.show();\n\ncorr_map(interp, width=20, height=10, annot_kws=8)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T02:09:06.233938Z","iopub.execute_input":"2024-02-06T02:09:06.234420Z","iopub.status.idle":"2024-02-06T02:09:06.802620Z","shell.execute_reply.started":"2024-02-06T02:09:06.234387Z","shell.execute_reply":"2024-02-06T02:09:06.801127Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"from sklearn.ensemble import RandomForestClassifier\nfrom boruta import BorutaPy","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:01:03.814200Z","iopub.execute_input":"2024-02-06T03:01:03.814691Z","iopub.status.idle":"2024-02-06T03:01:04.277241Z","shell.execute_reply.started":"2024-02-06T03:01:03.814651Z","shell.execute_reply":"2024-02-06T03:01:04.276005Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#All categorical values will be one-hot encoded.","metadata":{}},{"cell_type":"code","source":"#By Robin Smits https://www.kaggle.com/code/rsmits/feature-selection-with-boruta\n\ndf = pd.get_dummies(df, drop_first=True, dummy_na=True)\ndf.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:01:09.484988Z","iopub.execute_input":"2024-02-06T03:01:09.485498Z","iopub.status.idle":"2024-02-06T03:01:10.882117Z","shell.execute_reply.started":"2024-02-06T03:01:09.485455Z","shell.execute_reply":"2024-02-06T03:01:10.881149Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Get all feature names from the dataset","metadata":{}},{"cell_type":"code","source":"features = [f for f in df.columns if f not in ['target','case_id']]\nlen(features)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:01:59.694803Z","iopub.execute_input":"2024-02-06T03:01:59.695201Z","iopub.status.idle":"2024-02-06T03:01:59.702002Z","shell.execute_reply.started":"2024-02-06T03:01:59.695170Z","shell.execute_reply":"2024-02-06T03:01:59.700991Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Replace all missing values with the Mean. Save it for next time","metadata":{}},{"cell_type":"code","source":"df[features] = df[features].fillna(df[features].mean()).clip(-1e9,1e9)","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Get the final dataset X and labels Y","metadata":{}},{"cell_type":"code","source":"#By Robin Smits https://www.kaggle.com/code/rsmits/feature-selection-with-boruta\n\nX = df[features].values\ny = df['target'].values.ravel()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:02:32.459810Z","iopub.execute_input":"2024-02-06T03:02:32.460238Z","iopub.status.idle":"2024-02-06T03:03:17.698784Z","shell.execute_reply.started":"2024-02-06T03:02:32.460206Z","shell.execute_reply":"2024-02-06T03:03:17.697794Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Set up the RandomForrestClassifier as the estimator to use for Boruta. The max_depth of the tree is advised on the Boruta Github page to be between 3 to 7.\n\nclass_weightdict, “balanced” or None\n\n\"If “balanced”, class weights will be given by n_samples / (n_classes * np.bincount(y)). If a dictionary is given, keys are classes and values are corresponding class weights. If None is given, the class weights will be uniform.\"\n\nhttps://scikit-learn.org/stable/modules/generated/sklearn.utils.class_weight.compute_class_weight.html","metadata":{}},{"cell_type":"code","source":"#By Robin Smits https://www.kaggle.com/code/rsmits/feature-selection-with-boruta\n\nrf = RandomForestClassifier(n_jobs=-1, class_weight='balanced', max_depth=5)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:03:36.365409Z","iopub.execute_input":"2024-02-06T03:03:36.366097Z","iopub.status.idle":"2024-02-06T03:03:36.372504Z","shell.execute_reply.started":"2024-02-06T03:03:36.366051Z","shell.execute_reply":"2024-02-06T03:03:36.370879Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#https://github.com/scikit-learn-contrib/boruta_py\n\nfeat_selector = BorutaPy(rf, n_estimators='auto', verbose=2, random_state=1)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:04:19.242610Z","iopub.execute_input":"2024-02-06T03:04:19.243056Z","iopub.status.idle":"2024-02-06T03:04:19.250530Z","shell.execute_reply.started":"2024-02-06T03:04:19.243022Z","shell.execute_reply":"2024-02-06T03:04:19.248799Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Even on Boruta GitHub the next won't run. So I'm not able to fix it. Any clue??","metadata":{}},{"cell_type":"code","source":"#https://github.com/scikit-learn-contrib/boruta_py/blob/master/boruta/examples/Madalon_Data_Set.ipynb\n\n#https://github.com/scikit-learn-contrib/boruta_py\n\n# find all relevant features - 5 features should be selected\nfeat_selector.fit(X, y)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T03:04:43.999431Z","iopub.execute_input":"2024-02-06T03:04:43.999830Z","iopub.status.idle":"2024-02-06T03:06:17.061056Z","shell.execute_reply.started":"2024-02-06T03:04:43.999801Z","shell.execute_reply":"2024-02-06T03:06:17.059130Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#https://github.com/scikit-learn-contrib/boruta_py\n\n# check selected features - first 5 features are selected\nfeat_selector.support_","metadata":{},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#By Robin Smits https://www.kaggle.com/code/rsmits/feature-selection-with-boruta\n\nboruta_feature_selector = BorutaPy(rf, n_estimators='auto', verbose=2, random_state=4242, max_iter = 50, perc = 90)\nboruta_feature_selector.fit(X, y)","metadata":{"_kg_hide-output":true,"execution":{"iopub.status.busy":"2024-02-06T02:48:13.294331Z","iopub.execute_input":"2024-02-06T02:48:13.294731Z","iopub.status.idle":"2024-02-06T02:48:42.140971Z","shell.execute_reply.started":"2024-02-06T02:48:13.294700Z","shell.execute_reply":"2024-02-06T02:48:42.138989Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"After Boruta has run we can transform our dataset.","metadata":{}},{"cell_type":"code","source":"X_filtered = boruta_feature_selector.transform(X)\nX_filtered.shape","metadata":{"execution":{"iopub.status.busy":"2024-02-06T02:30:49.550839Z","iopub.status.idle":"2024-02-06T02:30:49.551904Z","shell.execute_reply.started":"2024-02-06T02:30:49.551603Z","shell.execute_reply":"2024-02-06T02:30:49.551630Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"final_features = list()\nindexes = np.where(boruta_feature_selector.support_ == True)\nfor x in np.nditer(indexes):\n    final_features.append(features[x])\nprint(final_features)","metadata":{"execution":{"iopub.status.busy":"2024-02-06T02:28:14.879951Z","iopub.status.idle":"2024-02-06T02:28:14.880564Z","shell.execute_reply.started":"2024-02-06T02:28:14.880232Z","shell.execute_reply":"2024-02-06T02:28:14.880271Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Checking Feature file csv","metadata":{}},{"cell_type":"code","source":"feat = pd.read_csv('../input/home-credit-credit-risk-model-stability/feature_definitions.csv')\nfeat.tail()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T00:30:03.162450Z","iopub.execute_input":"2024-02-06T00:30:03.162908Z","iopub.status.idle":"2024-02-06T00:30:03.178956Z","shell.execute_reply.started":"2024-02-06T00:30:03.162874Z","shell.execute_reply":"2024-02-06T00:30:03.177750Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = feat['Description'].value_counts()[:20].plot.barh(figsize=(16, 8), color='orange')\nax.set_title('Home Credit Feature Definitions', size=18, color='green')\nax.set_ylabel('Description', size=10)\nax.set_xlabel('Count', size=10);","metadata":{"execution":{"iopub.status.busy":"2024-02-06T00:35:15.604755Z","iopub.execute_input":"2024-02-06T00:35:15.605176Z","iopub.status.idle":"2024-02-06T00:35:16.271884Z","shell.execute_reply.started":"2024-02-06T00:35:15.605146Z","shell.execute_reply":"2024-02-06T00:35:16.270937Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Applicant file","metadata":{}},{"cell_type":"code","source":"appl = pd.read_csv('../input/home-credit-credit-risk-model-stability/csv_files/train/train_applprev_2.csv')\nappl.tail()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T00:42:14.748428Z","iopub.execute_input":"2024-02-06T00:42:14.749097Z","iopub.status.idle":"2024-02-06T00:42:32.461175Z","shell.execute_reply.started":"2024-02-06T00:42:14.749039Z","shell.execute_reply":"2024-02-06T00:42:32.459497Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = appl['conts_type_509L'].value_counts()[:10].plot.barh(figsize=(16, 8), color='green')\nax.set_title('Banking Communication Channels ', size=18, color='orange')\nax.set_ylabel('Applicants communication Channnels', size=10)\nax.set_xlabel('Count', size=10);","metadata":{"execution":{"iopub.status.busy":"2024-02-06T00:48:59.883091Z","iopub.execute_input":"2024-02-06T00:48:59.883508Z","iopub.status.idle":"2024-02-06T00:49:02.541503Z","shell.execute_reply.started":"2024-02-06T00:48:59.883476Z","shell.execute_reply":"2024-02-06T00:49:02.540556Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#One train credit bureau file","metadata":{}},{"cell_type":"code","source":"bur = pd.read_csv('../input/home-credit-credit-risk-model-stability/csv_files/train/train_credit_bureau_b_2.csv')\nbur.tail()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T00:56:42.682459Z","iopub.execute_input":"2024-02-06T00:56:42.682950Z","iopub.status.idle":"2024-02-06T00:56:44.126576Z","shell.execute_reply.started":"2024-02-06T00:56:42.682913Z","shell.execute_reply":"2024-02-06T00:56:44.124925Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Person file","metadata":{}},{"cell_type":"code","source":"per = pd.read_csv('../input/home-credit-credit-risk-model-stability/csv_files/train/train_person_2.csv')\nper.tail()","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:00:02.009969Z","iopub.execute_input":"2024-02-06T01:00:02.010431Z","iopub.status.idle":"2024-02-06T01:00:05.954442Z","shell.execute_reply.started":"2024-02-06T01:00:02.010399Z","shell.execute_reply":"2024-02-06T01:00:05.953111Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = per['addres_role_871L'].value_counts()[:10].plot.barh(figsize=(16, 8), color='blue')\nax.set_title('Applicants Address', size=18, color='red')\nax.set_ylabel('Address Role', size=10)\nax.set_xlabel('Count', size=10);","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:03:23.010053Z","iopub.execute_input":"2024-02-06T01:03:23.010553Z","iopub.status.idle":"2024-02-06T01:03:23.445933Z","shell.execute_reply.started":"2024-02-06T01:03:23.010520Z","shell.execute_reply":"2024-02-06T01:03:23.443358Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"ax = per['relatedpersons_role_762T'].value_counts()[:10].plot.barh(figsize=(16, 8), color='red')\nax.set_title('Applicants Related Persons', size=18, color='blue')\nax.set_ylabel('Related Persons', size=10)\nax.set_xlabel('Count', size=10);","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:05:46.031113Z","iopub.execute_input":"2024-02-06T01:05:46.031551Z","iopub.status.idle":"2024-02-06T01:05:46.479028Z","shell.execute_reply.started":"2024-02-06T01:05:46.031520Z","shell.execute_reply":"2024-02-06T01:05:46.477802Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Closing Arguments\n\nI'm disappointed that I wasn't able to deliver the Boruta Selector\n\n\"Credit scoring algorithms are a vital part of the credit industry, and they play a crucial role in determining creditworthiness. While there are several credit scoring algorithms in use today, each with its unique set of features and methodologies, they all aim to achieve the same end goal: to assess the risk of lending money to an individual. As a borrower, it's essential to understand how credit scoring algorithms work, and what factors they take into account, to ensure that you maintain good creditworthiness.\"\n\nhttps://fastercapital.com/content/Credit-Scoring-Algorithms--Cracking-the-Code-for-a-Better-Score.html","metadata":{"execution":{"iopub.status.busy":"2024-02-06T01:04:26.929787Z","iopub.execute_input":"2024-02-06T01:04:26.930824Z","iopub.status.idle":"2024-02-06T01:04:27.022776Z","shell.execute_reply.started":"2024-02-06T01:04:26.930782Z","shell.execute_reply":"2024-02-06T01:04:27.021354Z"}}},{"cell_type":"markdown","source":"#Acknowledgements:\n\nYouhan Lee  https://www.kaggle.com/code/youhanlee/extensive-eda-for-application-and-bureau-data\n\nFirat Gonen https://www.kaggle.com/frtgnn/elo-eda-lgbm/notebook \n\nPuru Behl https://www.kaggle.com/accountstatus/mt-cars-data-analysis\n\nMustafa Yildirim https://www.kaggle.com/code/homydata/accurate-estimation-of-home-credit-default-risk-by\n\nFurkan Akdag https://www.kaggle.com/code/furkannakdagg/smoking-drinking-prediction-complete-eda-pycaret/notebook","metadata":{}}]}