{"metadata":{"kernelspec":{"language":"python","display_name":"Python 3","name":"python3"},"language_info":{"pygments_lexer":"ipython3","nbconvert_exporter":"python","version":"3.6.4","file_extension":".py","codemirror_mode":{"name":"ipython","version":3},"name":"python","mimetype":"text/x-python"}},"nbformat_minor":4,"nbformat":4,"cells":[{"cell_type":"code","source":"# This Python 3 environment comes with many helpful analytics libraries installed\n# It is defined by the kaggle/python Docker image: https://github.com/kaggle/docker-python\n# For example, here's several helpful packages to load\n\nimport numpy as np # linear algebra\nimport pandas as pd # data processing, CSV file I/O (e.g. pd.read_csv)\n\n# Input data files are available in the read-only \"../input/\" directory\n# For example, running this (by clicking run or pressing Shift+Enter) will list all files under the input directory\n\nimport os\nfor dirname, _, filenames in os.walk('/kaggle/input'):\n    for filename in filenames:\n        print(os.path.join(dirname, filename))\n\n# You can write up to 20GB to the current directory (/kaggle/working/) that gets preserved as output when you create a version using \"Save & Run All\" \n# You can also write temporary files to /kaggle/temp/, but they won't be saved outside of the current session","metadata":{"_uuid":"8f2839f25d086af736a60e9eeb907d3b93b6e0e5","_cell_guid":"b1076dfc-b9ad-4769-8c92-a6c4dae69d19","_kg_hide-input":true,"_kg_hide-output":true,"execution":{"iopub.status.busy":"2022-05-25T21:33:58.719417Z","iopub.execute_input":"2022-05-25T21:33:58.720231Z","iopub.status.idle":"2022-05-25T21:33:58.753056Z","shell.execute_reply.started":"2022-05-25T21:33:58.720112Z","shell.execute_reply":"2022-05-25T21:33:58.752295Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<center style=\"font-family:verdana;\"><h1 style=\"font-size:200%; padding: 10px; background: #1E90FF;\"><b style=\"color:black;\">Normalized Gini Coefficient</b></h1></center>","metadata":{}},{"cell_type":"markdown","source":"#Gini coefficients demystified Definition, by Martin Goldberg March 11, 2013.\n\n![](https://images.slideplayer.com/33/8172270/slides/slide_10.jpg)https://slideplayer.com/slide/8172270/","metadata":{}},{"cell_type":"markdown","source":"\"The Gini coefficient was developed by statistician and sociologist Corrado Gini.\"\n\nhttps://en.wikipedia.org/wiki/Gini_coefficient\n\n\"The Gini coefficient is a statistic which measures the ability of a scorecard or a characteristic to rank order risk. A Gini value of 0% means that the characteristic cannot distinguish good from bad cases.\"\n\n\"A typical credit scorecard has a Gini coefficient of 40-60%. Behaviour scorecards have values of 70-80%. A very powerful characteristic can have a Gini coefficient of 25%.\"\n\n\"To calculate Gini values, assume that one has good and bad accounts rank ordered by score with the score sufficiently finely graded such as that there is only one case per score. The essential notion is that of a “flip”. A flip is a transposition of consecutive good and bad accounts.\"\n\n\"The Gini coefficient is the percentage of flips required to reach the rank ordering from a random assignment of goods and bads by score (i.e. with Gini = 0).\"\n\nhttp://www.rhinorisk.com/Publications/Gini%20Coefficients.pdf","metadata":{}},{"cell_type":"code","source":"#Code by  https://www.kaggle.com/kartushovdanil/ubiquant-market-prediction-eda\n\nfrom pathlib import Path\nimport random\nimport tqdm\n\nfrom argparse import Namespace\nimport random\nimport gc\nimport seaborn as sns\nfrom matplotlib import pyplot as plt\n\n# setting up options\nimport warnings\npd.set_option('display.max_rows', None)\npd.set_option('display.max_columns', None)\nwarnings.filterwarnings('ignore')\nfrom cycler import cycler","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:34:08.029438Z","iopub.execute_input":"2022-05-25T21:34:08.029844Z","iopub.status.idle":"2022-05-25T21:34:08.643139Z","shell.execute_reply.started":"2022-05-25T21:34:08.029812Z","shell.execute_reply":"2022-05-25T21:34:08.642275Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#The Default Rate\n\n\n<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">The Default Rate Formula</span></h1><br>\n\n\"The default rate is the rate of all loans issued by a lender or financial institution that is left unpaid by the borrower and declared to be in default.\"\n\n\"The lending institution will write off the entire value of defaulted loans, removing them from the books altogether. The default rate is important for institutions to reassess their risk from borrowers and is also an important representation of economic conditions.\"\n\nDefault Rate = Number of Defaulted Loans/Total Number of Loans X 100\n\n<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Period Until Default After Last Payment</span></h1><br>\n\nCredit Card: 180 days\n\nMortgage: 30 days\n\nStudent Loan: 270 days\n\n\n<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Routinely Missed Payments</span></h1><br>\n\n\"Lending institutions may implement consequences for borrowers with routinely missed or late payments.\"\n\n\"One strategy a lender may implement is to increase the interest rate on the borrower’s remaining loan after delinquency. The substantially higher interest rate is referred to as the penalty rate. The lender may decide to lower the penalty rate if the borrower successfully makes on-time payments.\"\n\n\"Another strategy allows the lending institution to take hold of personal assets after a defaulted loan. Personal assets may include property, wages, retirement savings, or investments. For example, upon taking ownership of a property, the bank may recover some of its losses on the loan. Through the process of foreclosure, the bank can sell the property.\"\n\nhttps://corporatefinanceinstitute.com/resources/knowledge/credit/default-rate/","metadata":{}},{"cell_type":"code","source":"#Code by Mohsin Hasan https://www.kaggle.com/code/tezdhar/faster-gini-calculation\n\n#The function used in most kernels\ndef gini(actual, pred, cmpcol = 0, sortcol = 1):\n    assert( len(actual) == len(pred) )\n    all = np.asarray(np.c_[ actual, pred, np.arange(len(actual)) ], dtype=np.float)\n    all = all[ np.lexsort((all[:,2], -1*all[:,1])) ]\n    totalLosses = all[:,0].sum()\n    giniSum = all[:,0].cumsum().sum() / totalLosses\n    \n    giniSum -= (len(actual) + 1) / 2.\n    return giniSum / len(actual)\n \ndef gini_normalized(a, p):\n    return gini(a, p) / gini(a, a)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:34:14.391318Z","iopub.execute_input":"2022-05-25T21:34:14.391731Z","iopub.status.idle":"2022-05-25T21:34:14.399853Z","shell.execute_reply.started":"2022-05-25T21:34:14.391698Z","shell.execute_reply":"2022-05-25T21:34:14.398928Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Default Risk Model</span></h1><br>\n\n\n\"Benchmarking Deutsche Bundesbank’s Default Risk Model, the KMV Private Firm Model and Common Financial Ratios for\nGerman Corporations\"\n\nAuthors: Stefan Blochwitz, Thilo Liebig, Mikael Nyberg\n\n\"By comparing Gini curves and Gini coefficients that are determined on the same underlying dataset, the authors assessed the discriminative power of Deutsche Bundesbank’s Default Risk Model, KMV’s Private Firm Model and common financial ratios for\nGerman corporations.\"\n\n\"While the purpose of the Bundesbank Default Risk Model is to decide whether a collateral is eligible for refinancing purposes, the model does this by assessing the creditworthiness of the individual borrowing company. Likewise, the goal of KMV’s Private Firm Model is to determine probabilities of default. However in both cases a best possible discriminative power is desirable.\"\n\n\"In this paper the authors showed that both the statistical model (discriminant analysis) that is the first step in the\nBundesbank’s system and the structural model of KMV (Private Firm Model) provided powerful approaches to credit analysis with similar results. \"\n\n\"When incorporating additional information gained from other sources than the financial statements and\nmarket trends, power of discrimination can further be improved as demonstrated by an expert system that is the second step of the Deutsche Bundesbank’s system.\"\n\n\"The focus of the paper is that of testing the performance of the models not to compare the model approaches in detail. The model construction and features are briefly described rather than exhaustively analysed.\"\n\nhttps://www.bis.org/bcbs/events/oslo/liebigblo.pdf","metadata":{}},{"cell_type":"code","source":"#Code by Mohsin Hasan https://www.kaggle.com/code/tezdhar/faster-gini-calculation\n\na = np.random.randint(0,2,100000)\np = np.random.rand(100000)\nprint(a[10:15], p[10:15])","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:34:20.883364Z","iopub.execute_input":"2022-05-25T21:34:20.884336Z","iopub.status.idle":"2022-05-25T21:34:20.894115Z","shell.execute_reply.started":"2022-05-25T21:34:20.884292Z","shell.execute_reply":"2022-05-25T21:34:20.893298Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Calculating the normalized Gini index</span></h1><br>\n\nThis function calculates the Gini index of a classification rule outputting probabilities. It is a classical metric in the context of Credit Scoring. It is equal to 2 times the AUC (Area Under ROC Curve) minus 1.\n\nhttps://rdrr.io/cran/glmdisc/man/normalizedGini.html\nhttps://search.r-project.org/CRAN/refmans/glmdisc/html/normalizedGini.html","metadata":{}},{"cell_type":"code","source":"%%time\ngini_normalized(a,p)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:34:24.986907Z","iopub.execute_input":"2022-05-25T21:34:24.987302Z","iopub.status.idle":"2022-05-25T21:34:25.031493Z","shell.execute_reply.started":"2022-05-25T21:34:24.987267Z","shell.execute_reply":"2022-05-25T21:34:25.030833Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Is Gini a merely reformulation of AUC?</span></h1><br>\n\n\ngini=2×AUC−1\n\n\"A random prediction will yield a Gini score of 0 as opposed to the AUC which will be 0.5.\"\n\n\"You cannot calculate AUC for a continuous target. However, they also use normalized Gini in regression tasks, like predicting insurance losses.\"\n\nhttps://stats.stackexchange.com/questions/306287/why-use-normalized-gini-score-instead-of-auc-as-evaluation","metadata":{}},{"cell_type":"code","source":"#Code by Mohsin Hasan https://www.kaggle.com/code/tezdhar/faster-gini-calculation\n\n#Remove redundant calls\ndef ginic(actual, pred):\n    actual = np.asarray(actual) #In case, someone passes Series or list\n    n = len(actual)\n    a_s = actual[np.argsort(pred)]\n    a_c = a_s.cumsum()\n    giniSum = a_c.sum() / a_s.sum() - (n + 1) / 2.0\n    return giniSum / n\n \ndef gini_normalizedc(a, p):\n    if p.ndim == 2:#Required for sklearn wrapper\n        p = p[:,1] #If proba array contains proba for both 0 and 1 classes, just pick class 1\n    return ginic(a, p) / ginic(a, a)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:34:29.874635Z","iopub.execute_input":"2022-05-25T21:34:29.875325Z","iopub.status.idle":"2022-05-25T21:34:29.882391Z","shell.execute_reply.started":"2022-05-25T21:34:29.875290Z","shell.execute_reply":"2022-05-25T21:34:29.881361Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">The Normalized Gini Coefficient</span></h1><br>\n\n\n\"The Normalized Gini coefficient is how far away we are with our sorted actual values from a random state measured in number of swaps\"\n\nhttps://theblog.github.io/post/gini-coefficient-intuitive-explanation/#:~:text=The%20Normalized%20Gini%20coefficient%20is,could%20give%20you%20a%20better","metadata":{}},{"cell_type":"code","source":"%%time\ngini_normalizedc(a,p)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:34:35.679163Z","iopub.execute_input":"2022-05-25T21:34:35.679647Z","iopub.status.idle":"2022-05-25T21:34:35.703291Z","shell.execute_reply.started":"2022-05-25T21:34:35.679606Z","shell.execute_reply":"2022-05-25T21:34:35.702357Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Why use Normalized Gini Score instead of AUC as evaluation?</span></h1><br>\n\n\"The Kaggle website used to have this answer: \"There is a maximum achievable area for a \"perfect\" model since not all of the positive examples occur immediately. They use the normalized Gini coefficient by dividing the Gini coefficient of your model by the Gini coefficient of the perfect model.\" but it is not available anymore. webcache.googleusercontent.com/… – \"\n\nBy Sextus Empiricus - Oct 10, 2017 at 1:01\n\nhttps://stats.stackexchange.com/questions/306287/why-use-normalized-gini-score-instead-of-auc-as-evaluation","metadata":{}},{"cell_type":"code","source":"#Code by Mohsin Hasan https://www.kaggle.com/code/tezdhar/faster-gini-calculation\n\n#XGBoost\nfrom sklearn import metrics\ndef gini_xgb(preds, dtrain):\n    labels = dtrain.get_label()\n    gini_score = gini_normalizedc(labels, preds)\n    return [('gini', gini_score)]\n\n#LightGBM\ndef gini_lgb(actuals, preds):\n    return 'gini', gini_normalizedc(actuals, preds), True\n\n#SKlearn\ngini_sklearn = metrics.make_scorer(gini_normalizedc)#Original was (gini_normalizedc, True, True)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:35:00.216110Z","iopub.execute_input":"2022-05-25T21:35:00.216975Z","iopub.status.idle":"2022-05-25T21:35:00.323295Z","shell.execute_reply.started":"2022-05-25T21:35:00.216939Z","shell.execute_reply":"2022-05-25T21:35:00.322500Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"Above: TypeError: make_scorer() takes 1 positional argument but 3 were given. Then I removed both True on the last line of the snippet.","metadata":{}},{"cell_type":"code","source":"#Code by  https://www.kaggle.com/kartushovdanil/ubiquant-market-prediction-eda\n\ndef reduce_mem_usage(df):\n    \"\"\" iterate through all the columns of a dataframe and modify the data type\n        to reduce memory usage.        \n    \"\"\"\n    start_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage of dataframe is {:.2f} MB'.format(start_mem))\n    \n    for col in df.columns:\n        col_type = df[col].dtype\n        \n        if col_type != object:\n            c_min = df[col].min()\n            c_max = df[col].max()\n            if str(col_type)[:3] == 'int':\n                if c_min > np.iinfo(np.int8).min and c_max < np.iinfo(np.int8).max:\n                    df[col] = df[col].astype(np.int8)\n                elif c_min > np.iinfo(np.int16).min and c_max < np.iinfo(np.int16).max:\n                    df[col] = df[col].astype(np.int16)\n                elif c_min > np.iinfo(np.int32).min and c_max < np.iinfo(np.int32).max:\n                    df[col] = df[col].astype(np.int32)\n                elif c_min > np.iinfo(np.int64).min and c_max < np.iinfo(np.int64).max:\n                    df[col] = df[col].astype(np.int64)  \n            else:\n                if c_min > np.finfo(np.float16).min and c_max < np.finfo(np.float16).max:\n                    df[col] = df[col].astype(np.float16)\n                elif c_min > np.finfo(np.float32).min and c_max < np.finfo(np.float32).max:\n                    df[col] = df[col].astype(np.float32)\n                else:\n                    df[col] = df[col].astype(np.float64)\n        else:\n            df[col] = df[col].astype('category')\n\n    end_mem = df.memory_usage().sum() / 1024**2\n    print('Memory usage after optimization is: {:.2f} MB'.format(end_mem))\n    print('Decreased by {:.1f}%'.format(100 * (start_mem - end_mem) / start_mem))\n    \n    return df","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:35:08.762834Z","iopub.execute_input":"2022-05-25T21:35:08.763335Z","iopub.status.idle":"2022-05-25T21:35:08.786272Z","shell.execute_reply.started":"2022-05-25T21:35:08.763292Z","shell.execute_reply":"2022-05-25T21:35:08.785572Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"train = pd.read_csv(\"../input/amex-default-prediction/train_data.csv\",nrows=10000)\ntest =  pd.read_csv(\"../input/amex-default-prediction/test_data.csv\",nrows=10000)\nsub = pd.read_csv('/kaggle/input/amex-default-prediction/sample_submission.csv')\nlabels = pd.read_csv(\"../input/amex-default-prediction/train_labels.csv\", nrows=10000)","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:35:14.430403Z","iopub.execute_input":"2022-05-25T21:35:14.431291Z","iopub.status.idle":"2022-05-25T21:35:17.894992Z","shell.execute_reply.started":"2022-05-25T21:35:14.431254Z","shell.execute_reply":"2022-05-25T21:35:17.893859Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#Let's see some charts. A code without charts has a lack of Soul.","metadata":{}},{"cell_type":"code","source":"inv_ids = random.choices(train['P_2'].unique(), k=3) #Original was ['target'], since train don´t have target feature I chose R_1","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:35:47.340597Z","iopub.execute_input":"2022-05-25T21:35:47.340961Z","iopub.status.idle":"2022-05-25T21:35:47.355810Z","shell.execute_reply.started":"2022-05-25T21:35:47.340932Z","shell.execute_reply":"2022-05-25T21:35:47.354828Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by  https://www.kaggle.com/kartushovdanil/ubiquant-market-prediction-eda\n#https://www.kaggle.com/code/mpwolke/netflix-appetency-charts\n\nplt.rcParams['figure.dpi'] = 600\nfig = plt.figure(figsize=(10, 10), facecolor='#f6f5f5')\ngs = fig.add_gridspec(5, 5)\ngs.update(wspace=0.3, hspace=0.3)\nbackground_color = '#f6f5f5'\nrun_no = 0\n\ncolormap = ['#1DBA94','#1C5ED2', '#FFC300', '#C70039']\nplt.rc('axes', prop_cycle=(cycler('color', colormap)))\n\nfor row in range(0, 5):\n    for col in range(0, 5):\n        locals()[\"ax\"+str(run_no)] = fig.add_subplot(gs[row, col])\n        locals()[\"ax\"+str(run_no)].set_facecolor(background_color)\n        for s in [\"top\",\"right\"]:\n            locals()[\"ax\"+str(run_no)].spines[s].set_visible(False)\n        run_no += 1  \n\n\nfeatures = list(train.columns[7:16]) #column 7 till 16 are uint\n\nrun_no = 0\nfor col in features:\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=train[col], zorder=2, alpha=1, linewidth=1, color='#ffd514')\n    sns.kdeplot(ax=locals()[\"ax\"+str(run_no)], x=train[train['P_2'].isin(inv_ids)][col], hue=train[train['P_2'].isin(inv_ids)]['R_1'],zorder=2, alpha=1, fill=True, color=colormap, linewidth=0.5, legend=False, hue_order=inv_ids.sort(reverse=True))\n    \n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='x', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].grid(which='major', axis='y', zorder=0, color='#EEEEEE', linewidth=0.4)\n    locals()[\"ax\"+str(run_no)].set_ylabel('')\n    locals()[\"ax\"+str(run_no)].set_xlabel(col, fontsize=4, fontweight='bold')\n    locals()[\"ax\"+str(run_no)].tick_params(labelsize=4, width=0.5)\n    locals()[\"ax\"+str(run_no)].xaxis.offsetText.set_fontsize(4)\n    locals()[\"ax\"+str(run_no)].yaxis.offsetText.set_fontsize(4)\n    #locals()[\"ax\"+str(run_no)].get_legend().remove()\n    \n    run_no += 1\n\nplt.show()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:36:31.500870Z","iopub.execute_input":"2022-05-25T21:36:31.501254Z","iopub.status.idle":"2022-05-25T21:36:35.863534Z","shell.execute_reply.started":"2022-05-25T21:36:31.501224Z","shell.execute_reply":"2022-05-25T21:36:35.862511Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by MAYUR DALVI  https://www.kaggle.com/mayurdalvi/tabular-playground-series-simple-and-easy\n\ncols = ['R_'+str(i) for i in range(29)] #Original was range (100) R_29 is the last one","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:39:08.150282Z","iopub.execute_input":"2022-05-25T21:39:08.151253Z","iopub.status.idle":"2022-05-25T21:39:08.155810Z","shell.execute_reply.started":"2022-05-25T21:39:08.151210Z","shell.execute_reply":"2022-05-25T21:39:08.154971Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"#Code by MAYUR DALVI  https://www.kaggle.com/mayurdalvi/tabular-playground-series-simple-and-easy\n#https://www.kaggle.com/code/mpwolke/netflix-appetency-charts\n\n#plot 22 features (R7 - R28) When all are numerical (int/float)\ni = 1\nplt.figure()\nfig, ax = plt.subplots(6,4 ,figsize=(20, 22))\nfor feature in cols[7:180]:\n    plt.subplot(6, 4,i)\n    sns.histplot(train[feature],color=\"blue\", kde=True,bins=100, label='train_'+feature)\n    sns.histplot(test[feature],color=\"olive\", kde=True,bins=100, label='test_'+feature)\n    plt.xlabel(feature, fontsize=9); plt.legend()\n    i += 1\nplt.show();","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:39:24.158995Z","iopub.execute_input":"2022-05-25T21:39:24.159436Z","iopub.status.idle":"2022-05-25T21:39:47.544395Z","shell.execute_reply.started":"2022-05-25T21:39:24.159404Z","shell.execute_reply":"2022-05-25T21:39:47.543463Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize=(6,4))\nsns.catplot(x=\"target\", kind=\"count\",  data=labels,);","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:21:30.940728Z","iopub.execute_input":"2022-05-25T21:21:30.941204Z","iopub.status.idle":"2022-05-25T21:21:31.715933Z","shell.execute_reply.started":"2022-05-25T21:21:30.941166Z","shell.execute_reply":"2022-05-25T21:21:31.714747Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"#That's HUUUUUUUUGE! And took a long time to render an overlapping chart. ","metadata":{}},{"cell_type":"code","source":"plt.figure(figsize=(6,4))\nsns.catplot(x=\"P_2\", kind=\"count\",  data=train);","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:40:46.822072Z","iopub.execute_input":"2022-05-25T21:40:46.823081Z","iopub.status.idle":"2022-05-25T21:43:58.297301Z","shell.execute_reply.started":"2022-05-25T21:40:46.823030Z","shell.execute_reply":"2022-05-25T21:43:58.296522Z"},"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"code","source":"plt.figure(figsize = (12,5))\nax = sns.distplot(train['R_1'], bins=5000)\nplt.xlim(-3,3)\nplt.xlabel(\"Histogram of Risk 1\", size=12)\nplt.show();\ngc.collect()","metadata":{"execution":{"iopub.status.busy":"2022-05-25T21:23:04.584777Z","iopub.execute_input":"2022-05-25T21:23:04.585249Z","iopub.status.idle":"2022-05-25T21:23:16.383493Z","shell.execute_reply.started":"2022-05-25T21:23:04.585212Z","shell.execute_reply":"2022-05-25T21:23:16.382501Z"},"_kg_hide-input":true,"trusted":true},"execution_count":null,"outputs":[]},{"cell_type":"markdown","source":"<h1><span class=\"label label-default\" style=\"background-color:black;border-radius:100px 100px; font-weight: bold; font-family:Garamond; font-size:20px; color:#03e8fc; padding:10px\">Amex Platinum Credit Card </span></h1><br>\n\nThat's all for now with that large Credit Card data.\n\nAmex Platinum credit card: an overview\n\n\"The Amex Platinum is a premium credit card that offers a welcome bonus, luxury perks, and other value-added services. Cardholders benefit from Uber credits, airfare discounts, bonus points, cashback and much more. Unlike typical credit cards, this card allows you to carry a balance for certain charges, but not all.\"\n\n\"There is a catch, though: the Amex Platinum carries a 695 annual fee which is high compared to other options.\"\n\n\"Many people are happy to pay the annual fee because of the incredible perks attached to the Amex Platinum. If you’re calculating whether the cost is worth it, potential applicants must understand whether their lifestyle and financial circumstances are a good fit for the Amex Platinum Card.\"\n\nhttps://www.novacredit.com/resources/who-should-and-who-shouldnt-get-the-amex-platinum-card/#:~:text=The%20Amex%20Platinum%20is%20a,certain%20charges%2C%20but%20not%20all.","metadata":{}},{"cell_type":"markdown","source":"![](https://pics.astrologymemes.com/card-card-first-name-desc-youre-pre-approved-to-apply-for-the-platinum-card%C2%AE-from-58278407.png)astrologymemes.com","metadata":{}},{"cell_type":"markdown","source":"#Acknowledgements\n\nTorch me https://www.kaggle.com/kartushovdanil/ubiquant-market-prediction-eda\n\nMohsin Hasan https://www.kaggle.com/code/tezdhar/faster-gini-calculation","metadata":{}}]}