{
  "id": 333338,
  "title": "Understanding competition metric step by step",
  "url": "/competitions/amex-default-prediction/discussion/333338",
  "author_name": "",
  "post_date": "2022-06-26T02:24:12.645355400Z",
  "votes": 144,
  "comment_count": 26,
  "views": 0,
  "content": "<p>This competition gives us a metric \\(M = 0.5 \\cdot (G + D)\\). It looks simple at a quick glance with the average of \\(G\\) and \\(D\\), but the important part \\(G\\) and \\(D\\) is a little bit complicated to understand. I would like to go through these elements one by one. (I apologize if my English isn't very good.)</p>\n<h2>■ G: Normalized Gini Coefficient</h2>\n<h3>Fundamental Knowledge (ROC curve and AUC)</h3>\n<p>First, in order to understand and calculate the Gini coefficient, we will review the receiver operating characteristic (ROC) curve and the area under the curve (AUC).</p>\n<h4>1. Creating the ROC curve</h4>\n<p>The ROC curve is a curve that can visually capture whether a customer actually defaulted or not when the predicted default rates are ordered. Suppose we have 4 defaulting customers and 8 non-defaulting customers. The following is a rough description of what we are going to do.</p>\n<pre><code>1. Draw a diagram divided vertically into 4 sections and horizontally into 8 sections.\n2. Extend the line directly above if it is a defaulting customer and directly across if it is a non-defaulting customer.\n</code></pre>\n<p>Let us assume that the predicted default rate is calculated as follows: D is the actual defaulting customer and ND is the non-defaulting customer.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Ffdd48f6c2d1d9ef82ae205b25e9cd6fb%2FIMG_0727.jpg?generation=1656141318176769&amp;alt=media\">\n</div>\n<p>Now let's draw a graph based on the example above (↑). Starting from (0, 0), we will draw the ROC curve in the order of increasing the default prediction score (corresponding to the blue line in the image below).</p>\n<pre><code>1. the customer with the highest score (B) is defaulting, so it goes straight up 1/4. \n2. the customer with the next highest score (G) is not defaulting, so it goes straight across 1/8. \n3. the customer with the next highest score (C) is defaulting, so it goes straight up 1/4.\n....\n</code></pre>\n<p>Then we can draw a line that connects the origin (0, 0) and the origin (1, 1) for all customers.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F0f8a41a6fad5681ad85a16528aef1553%2FIMG_CA31680E6CB1-1.jpeg?generation=1656287940943246&amp;alt=media\">\n</div>\n<p>If the model is able to determine whether the customer has defaulted or not completely (perfect model), the ROC curve is drawn like a red line.</p>\n<h4>2. Calculating AUC</h4>\n<p>Let's try to calculate the AUC (the blue area) of the model we created. Since there are 26 rectangles of 1/4 × 1/8 per square, we can find the value of the AUC as follows.</p>\n<p>$$AUC = \\frac{1}{4} \\cdot \\frac{1}{8} \\cdot 26 = \\frac{13}{16}$$</p>\n<p>In this example, we can see that the AUC is \\(\\frac{13}{16}\\).</p>\n<pre><code>- The range of AUC is from 0 to 1.\n- 1 for perfect models and 0.5 for random models.\n</code></pre>\n<h3>The Normalized Gini Coefficient</h3>\n<p>As with the AUC, we will consider the area. A well-known measure of income inequality is the Gini coefficient, which is calculated by doubling the area between the Lorenz curve and the diagonal line. Also, the Gini coefficient calculated using the ROC curve when evaluating a model is defined as \"the area between the ROC curve and the diagonal line\". The normalized Gini coefficient is the ratio of the model created by us to the perfect model.</p>\n<p>$$Gini = A$$<br>\n$$Normalized \\, Gini = \\frac{Gini_{model A}}{Gini_{perfect \\, model}} = \\frac{A}{A+B} = 2A$$</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fa7bfd3b76875b5e06a2b617b7ff240a7%2FIMG_5CA49BBB16DC-1.jpeg?generation=1656136752796766&amp;alt=media\">\n</div>\n<pre><code>- The range of the Gini coefficient is from -0.5 to 0.5, and the range of the normalized Gini coefficient is from -1 to 1. \n- The purpose of the normalized Gini coefficient is to scale the AUC (= 0 for the random model and 1 for the perfect model).\n</code></pre>\n<p>The AUC represents the percentage of the area of the perfect model, and the (normal) Gini coefficient represents the percentage based on the diagonal (random model).</p>\n<p>So far, we have seen how the ROC curve is drawn and how the AUC and normalized Gini coefficient are calculated. It was a little difficult for me to understand the Gini coefficient at first because I had never seen it used in a way not to measure income inequality, but it was interesting to learn that it can also be used to measure the ability to rank in such a way.</p>\n<h2>■ D: The Default Rate Captured at 4%</h2>\n<p>List the predicted default rates of the customers in descending order. Customers whose rank \\(Z\\) (top percentile) is higher than the threshold \\(c\\) (top 4% = 0.04 in this case) are identified as defaults, and customers whose rank is below the threshold are identified as non-defaults.</p>\n<p>To understand the evaluation indicator \\(D\\), the probability of being correctly or incorrectly discriminated, which is determined when we decide the threshold, is important.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F4617d27fecf0273b79f9e65fb25ceada%2F2022-06-27%2014.03.34.png?generation=1656306263171656&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F1eb7b4bf2182090be20c1aa87780f91e%2F2022-06-27%2013.58.31_fpr.png?generation=1656306087392414&amp;alt=media\"></p>\n<p>The FPR represents \"how many non-defaulting customers are incorrectly identified as defaulting,\" and the TNR represents \"how many non-defaulting customers are correctly identified as non-defaulting.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Facfc568279e90ecd426af59096aef760%2FIMG_11CED5E1921A-1.jpeg?generation=1656137055041828&amp;alt=media\">\n</div>\n<ul>\n<li>Horizontal axis: FPR(c) + TNR(c) = 1</li>\n<li>Vertical axis: TPR(c) + FNR(c) = 1</li>\n</ul>\n<p>If we have two models with the same Gini Coefficient, the default rate captured by the 4% (TPR) will change. In the two models below, the value of \\(D\\) in \\(M = 0.5 \\cdot (G + D)\\) is higher in model B, so the final score (\\(M\\)) is also higher in model B (assuming the Gini coefficients \\(G\\) for Model A and Model B in the figure below are the same.)</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F2ef5c997d2de5e9ce5c8f7c56fcd5b16%2FIMG_0D77648FC3E1-1.jpeg?generation=1656137220409055&amp;alt=media\">\n</div>\n<h2>Meaning of combining the two indicators</h2>\n<pre><code>G: Overall correct response rate \nD: Percentage of default customers correctly identified as default customers at the TOP 4% point\n</code></pre>\n<p>\\(G\\) and \\(D\\) each has the above meaning. In other words, in addition to a high overall correct response rate, many defaulting customers must be in the top 4% when sorted in order of predicted default rate (i.e., we should not miss the prediction for customers with high predicted default rates as much as possible).</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fdb7315b0ed2a73cbeb2560d13d8dd7b6%2FIMG_B873FB421C0B-1.jpeg?generation=1656137321111691&amp;alt=media\">\n</div>\n<p>I thought I could just use either \\(G\\) or \\(D\\), but when I understood the logic behind each, mixing both results made a lot of sense. I hope this helps you understand the metric a little better :) </p>\n<p>This post is translated version (JP -&gt; EN). I will post an original Japanese article on note.com after this competition ends.<br>\n-&gt; <a href=\"https://note.com/hanaori/n/n5ced61443058\" target=\"_blank\">https://note.com/hanaori/n/n5ced61443058</a></p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/inversion/amex-competition-metric-python\" target=\"_blank\">https://www.kaggle.com/code/inversion/amex-competition-metric-python</a></li>\n<li><a href=\"https://www.asakura.co.jp/detail.php?book_code=29541\" target=\"_blank\">https://www.asakura.co.jp/detail.php?book_code=29541</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464</a></li>\n</ul>",
  "messages": [
    {
      "id": "1833422",
      "postDate": "06/26/2022 02:24:12",
      "content": "<p>This competition gives us a metric \\(M = 0.5 \\cdot (G + D)\\). It looks simple at a quick glance with the average of \\(G\\) and \\(D\\), but the important part \\(G\\) and \\(D\\) is a little bit complicated to understand. I would like to go through these elements one by one. (I apologize if my English isn't very good.)</p>\n<h2>■ G: Normalized Gini Coefficient</h2>\n<h3>Fundamental Knowledge (ROC curve and AUC)</h3>\n<p>First, in order to understand and calculate the Gini coefficient, we will review the receiver operating characteristic (ROC) curve and the area under the curve (AUC).</p>\n<h4>1. Creating the ROC curve</h4>\n<p>The ROC curve is a curve that can visually capture whether a customer actually defaulted or not when the predicted default rates are ordered. Suppose we have 4 defaulting customers and 8 non-defaulting customers. The following is a rough description of what we are going to do.</p>\n<pre><code>1. Draw a diagram divided vertically into 4 sections and horizontally into 8 sections.\n2. Extend the line directly above if it is a defaulting customer and directly across if it is a non-defaulting customer.\n</code></pre>\n<p>Let us assume that the predicted default rate is calculated as follows: D is the actual defaulting customer and ND is the non-defaulting customer.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Ffdd48f6c2d1d9ef82ae205b25e9cd6fb%2FIMG_0727.jpg?generation=1656141318176769&amp;alt=media\">\n</div>\n<p>Now let's draw a graph based on the example above (↑). Starting from (0, 0), we will draw the ROC curve in the order of increasing the default prediction score (corresponding to the blue line in the image below).</p>\n<pre><code>1. the customer with the highest score (B) is defaulting, so it goes straight up 1/4. \n2. the customer with the next highest score (G) is not defaulting, so it goes straight across 1/8. \n3. the customer with the next highest score (C) is defaulting, so it goes straight up 1/4.\n....\n</code></pre>\n<p>Then we can draw a line that connects the origin (0, 0) and the origin (1, 1) for all customers.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F0f8a41a6fad5681ad85a16528aef1553%2FIMG_CA31680E6CB1-1.jpeg?generation=1656287940943246&amp;alt=media\">\n</div>\n<p>If the model is able to determine whether the customer has defaulted or not completely (perfect model), the ROC curve is drawn like a red line.</p>\n<h4>2. Calculating AUC</h4>\n<p>Let's try to calculate the AUC (the blue area) of the model we created. Since there are 26 rectangles of 1/4 × 1/8 per square, we can find the value of the AUC as follows.</p>\n<p>$$AUC = \\frac{1}{4} \\cdot \\frac{1}{8} \\cdot 26 = \\frac{13}{16}$$</p>\n<p>In this example, we can see that the AUC is \\(\\frac{13}{16}\\).</p>\n<pre><code>- The range of AUC is from 0 to 1.\n- 1 for perfect models and 0.5 for random models.\n</code></pre>\n<h3>The Normalized Gini Coefficient</h3>\n<p>As with the AUC, we will consider the area. A well-known measure of income inequality is the Gini coefficient, which is calculated by doubling the area between the Lorenz curve and the diagonal line. Also, the Gini coefficient calculated using the ROC curve when evaluating a model is defined as \"the area between the ROC curve and the diagonal line\". The normalized Gini coefficient is the ratio of the model created by us to the perfect model.</p>\n<p>$$Gini = A$$<br>\n$$Normalized \\, Gini = \\frac{Gini_{model A}}{Gini_{perfect \\, model}} = \\frac{A}{A+B} = 2A$$</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fa7bfd3b76875b5e06a2b617b7ff240a7%2FIMG_5CA49BBB16DC-1.jpeg?generation=1656136752796766&amp;alt=media\">\n</div>\n<pre><code>- The range of the Gini coefficient is from -0.5 to 0.5, and the range of the normalized Gini coefficient is from -1 to 1. \n- The purpose of the normalized Gini coefficient is to scale the AUC (= 0 for the random model and 1 for the perfect model).\n</code></pre>\n<p>The AUC represents the percentage of the area of the perfect model, and the (normal) Gini coefficient represents the percentage based on the diagonal (random model).</p>\n<p>So far, we have seen how the ROC curve is drawn and how the AUC and normalized Gini coefficient are calculated. It was a little difficult for me to understand the Gini coefficient at first because I had never seen it used in a way not to measure income inequality, but it was interesting to learn that it can also be used to measure the ability to rank in such a way.</p>\n<h2>■ D: The Default Rate Captured at 4%</h2>\n<p>List the predicted default rates of the customers in descending order. Customers whose rank \\(Z\\) (top percentile) is higher than the threshold \\(c\\) (top 4% = 0.04 in this case) are identified as defaults, and customers whose rank is below the threshold are identified as non-defaults.</p>\n<p>To understand the evaluation indicator \\(D\\), the probability of being correctly or incorrectly discriminated, which is determined when we decide the threshold, is important.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F4617d27fecf0273b79f9e65fb25ceada%2F2022-06-27%2014.03.34.png?generation=1656306263171656&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F1eb7b4bf2182090be20c1aa87780f91e%2F2022-06-27%2013.58.31_fpr.png?generation=1656306087392414&amp;alt=media\"></p>\n<p>The FPR represents \"how many non-defaulting customers are incorrectly identified as defaulting,\" and the TNR represents \"how many non-defaulting customers are correctly identified as non-defaulting.</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Facfc568279e90ecd426af59096aef760%2FIMG_11CED5E1921A-1.jpeg?generation=1656137055041828&amp;alt=media\">\n</div>\n<ul>\n<li>Horizontal axis: FPR(c) + TNR(c) = 1</li>\n<li>Vertical axis: TPR(c) + FNR(c) = 1</li>\n</ul>\n<p>If we have two models with the same Gini Coefficient, the default rate captured by the 4% (TPR) will change. In the two models below, the value of \\(D\\) in \\(M = 0.5 \\cdot (G + D)\\) is higher in model B, so the final score (\\(M\\)) is also higher in model B (assuming the Gini coefficients \\(G\\) for Model A and Model B in the figure below are the same.)</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F2ef5c997d2de5e9ce5c8f7c56fcd5b16%2FIMG_0D77648FC3E1-1.jpeg?generation=1656137220409055&amp;alt=media\">\n</div>\n<h2>Meaning of combining the two indicators</h2>\n<pre><code>G: Overall correct response rate \nD: Percentage of default customers correctly identified as default customers at the TOP 4% point\n</code></pre>\n<p>\\(G\\) and \\(D\\) each has the above meaning. In other words, in addition to a high overall correct response rate, many defaulting customers must be in the top 4% when sorted in order of predicted default rate (i.e., we should not miss the prediction for customers with high predicted default rates as much as possible).</p>\n<div>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fdb7315b0ed2a73cbeb2560d13d8dd7b6%2FIMG_B873FB421C0B-1.jpeg?generation=1656137321111691&amp;alt=media\">\n</div>\n<p>I thought I could just use either \\(G\\) or \\(D\\), but when I understood the logic behind each, mixing both results made a lot of sense. I hope this helps you understand the metric a little better :) </p>\n<p>This post is translated version (JP -&gt; EN). I will post an original Japanese article on note.com after this competition ends.<br>\n-&gt; <a href=\"https://note.com/hanaori/n/n5ced61443058\" target=\"_blank\">https://note.com/hanaori/n/n5ced61443058</a></p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://www.kaggle.com/code/inversion/amex-competition-metric-python\" target=\"_blank\">https://www.kaggle.com/code/inversion/amex-competition-metric-python</a></li>\n<li><a href=\"https://www.asakura.co.jp/detail.php?book_code=29541\" target=\"_blank\">https://www.asakura.co.jp/detail.php?book_code=29541</a></li>\n<li><a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\">https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464</a></li>\n</ul>",
      "rawMarkdown": "This competition gives us a metric \\\\(M = 0.5 \\cdot (G + D)\\\\). It looks simple at a quick glance with the average of \\\\(G\\\\) and \\\\(D\\\\), but the important part \\\\(G\\\\) and \\\\(D\\\\) is a little bit complicated to understand. I would like to go through these elements one by one. (I apologize if my English isn't very good.)\n\n## ■ G: Normalized Gini Coefficient\n\n### Fundamental Knowledge (ROC curve and AUC)\n\nFirst, in order to understand and calculate the Gini coefficient, we will review the receiver operating characteristic (ROC) curve and the area under the curve (AUC).\n\n#### 1. Creating the ROC curve\nThe ROC curve is a curve that can visually capture whether a customer actually defaulted or not when the predicted default rates are ordered. Suppose we have 4 defaulting customers and 8 non-defaulting customers. The following is a rough description of what we are going to do.\n\n```\n1. Draw a diagram divided vertically into 4 sections and horizontally into 8 sections.\n2. Extend the line directly above if it is a defaulting customer and directly across if it is a non-defaulting customer.\n```\n\nLet us assume that the predicted default rate is calculated as follows: D is the actual defaulting customer and ND is the non-defaulting customer.\n\n<div style=\"text-align: center;\">\n<img width=800 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Ffdd48f6c2d1d9ef82ae205b25e9cd6fb%2FIMG_0727.jpg?generation=1656141318176769&alt=media\" />\n</div>\n\nNow let's draw a graph based on the example above (↑). Starting from (0, 0), we will draw the ROC curve in the order of increasing the default prediction score (corresponding to the blue line in the image below).\n\n```\n1. the customer with the highest score (B) is defaulting, so it goes straight up 1/4. \n2. the customer with the next highest score (G) is not defaulting, so it goes straight across 1/8. \n3. the customer with the next highest score (C) is defaulting, so it goes straight up 1/4.\n....\n```\nThen we can draw a line that connects the origin (0, 0) and the origin (1, 1) for all customers.\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F0f8a41a6fad5681ad85a16528aef1553%2FIMG_CA31680E6CB1-1.jpeg?generation=1656287940943246&alt=media\" />\n</div>\n\nIf the model is able to determine whether the customer has defaulted or not completely (perfect model), the ROC curve is drawn like a red line.\n\n#### 2. Calculating AUC\nLet's try to calculate the AUC (the blue area) of the model we created. Since there are 26 rectangles of 1/4 × 1/8 per square, we can find the value of the AUC as follows.\n\n$$AUC = \\frac{1}{4} \\cdot \\frac{1}{8} \\cdot 26 = \\frac{13}{16}$$\n\nIn this example, we can see that the AUC is \\\\(\\frac{13}{16}\\\\).\n\n```\n- The range of AUC is from 0 to 1.\n- 1 for perfect models and 0.5 for random models.\n```\n\n### The Normalized Gini Coefficient\n\nAs with the AUC, we will consider the area. A well-known measure of income inequality is the Gini coefficient, which is calculated by doubling the area between the Lorenz curve and the diagonal line. Also, the Gini coefficient calculated using the ROC curve when evaluating a model is defined as \"the area between the ROC curve and the diagonal line\". The normalized Gini coefficient is the ratio of the model created by us to the perfect model.\n\n$$Gini = A$$\n$$Normalized \\, Gini = \\frac{Gini_{model A}}{Gini_{perfect \\, model}} = \\frac{A}{A+B} = 2A$$\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fa7bfd3b76875b5e06a2b617b7ff240a7%2FIMG_5CA49BBB16DC-1.jpeg?generation=1656136752796766&alt=media\" />\n</div>\n\n```\n- The range of the Gini coefficient is from -0.5 to 0.5, and the range of the normalized Gini coefficient is from -1 to 1. \n- The purpose of the normalized Gini coefficient is to scale the AUC (= 0 for the random model and 1 for the perfect model).\n```\nThe AUC represents the percentage of the area of the perfect model, and the (normal) Gini coefficient represents the percentage based on the diagonal (random model).\n\nSo far, we have seen how the ROC curve is drawn and how the AUC and normalized Gini coefficient are calculated. It was a little difficult for me to understand the Gini coefficient at first because I had never seen it used in a way not to measure income inequality, but it was interesting to learn that it can also be used to measure the ability to rank in such a way.\n\n## ■ D: The Default Rate Captured at 4%\n\nList the predicted default rates of the customers in descending order. Customers whose rank \\\\(Z\\\\) (top percentile) is higher than the threshold \\\\(c\\\\) (top 4% = 0.04 in this case) are identified as defaults, and customers whose rank is below the threshold are identified as non-defaults.\n\nTo understand the evaluation indicator \\\\(D\\\\), the probability of being correctly or incorrectly discriminated, which is determined when we decide the threshold, is important.\n\n<img width=700 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F4617d27fecf0273b79f9e65fb25ceada%2F2022-06-27%2014.03.34.png?generation=1656306263171656&alt=media\" />\n<img width=700 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F1eb7b4bf2182090be20c1aa87780f91e%2F2022-06-27%2013.58.31_fpr.png?generation=1656306087392414&alt=media\" />\n\n\nThe FPR represents \"how many non-defaulting customers are incorrectly identified as defaulting,\" and the TNR represents \"how many non-defaulting customers are correctly identified as non-defaulting.\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Facfc568279e90ecd426af59096aef760%2FIMG_11CED5E1921A-1.jpeg?generation=1656137055041828&alt=media\" />\n</div>\n- Horizontal axis: FPR(c) + TNR(c) = 1\n- Vertical axis: TPR(c) + FNR(c) = 1\n\nIf we have two models with the same Gini Coefficient, the default rate captured by the 4% (TPR) will change. In the two models below, the value of \\\\(D\\\\) in \\\\(M = 0.5 \\cdot (G + D)\\\\) is higher in model B, so the final score (\\\\(M\\\\)) is also higher in model B (assuming the Gini coefficients \\\\(G\\\\) for Model A and Model B in the figure below are the same.)\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F2ef5c997d2de5e9ce5c8f7c56fcd5b16%2FIMG_0D77648FC3E1-1.jpeg?generation=1656137220409055&alt=media\" />\n</div>\n\n## Meaning of combining the two indicators\n\n```\nG: Overall correct response rate \nD: Percentage of default customers correctly identified as default customers at the TOP 4% point\n```\n\n\\\\(G\\\\) and \\\\(D\\\\) each has the above meaning. In other words, in addition to a high overall correct response rate, many defaulting customers must be in the top 4% when sorted in order of predicted default rate (i.e., we should not miss the prediction for customers with high predicted default rates as much as possible).\n\n<div style=\"text-align: center;\">\n<img width=800 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fdb7315b0ed2a73cbeb2560d13d8dd7b6%2FIMG_B873FB421C0B-1.jpeg?generation=1656137321111691&alt=media\" />\n</div>\n\nI thought I could just use either \\\\(G\\\\) or \\\\(D\\\\), but when I understood the logic behind each, mixing both results made a lot of sense. I hope this helps you understand the metric a little better :) \n\nThis post is translated version (JP -> EN). I will post an original Japanese article on note.com after this competition ends.\n-> https://note.com/hanaori/n/n5ced61443058\n\n## References\n- https://www.kaggle.com/code/inversion/amex-competition-metric-python\n- https://www.asakura.co.jp/detail.php?book_code=29541\n- https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464",
      "votes": null
    },
    {
      "id": "1834392",
      "postDate": "06/26/2022 22:10:43",
      "content": "<p>Very clearly explained!</p>",
      "rawMarkdown": "Very clearly explained!",
      "votes": null
    },
    {
      "id": "1834442",
      "postDate": "06/26/2022 23:43:15",
      "content": "<p>Thank you for reading 😊</p>",
      "rawMarkdown": "Thank you for reading 😊",
      "votes": null
    },
    {
      "id": "1834511",
      "postDate": "06/27/2022 02:16:54",
      "content": "<p>Thanks for sharing and teaching! </p>",
      "rawMarkdown": "Thanks for sharing and teaching!",
      "votes": null
    },
    {
      "id": "1834551",
      "postDate": "06/27/2022 03:25:43",
      "content": "<p>Thanks for the detailed explanation! I noticed there's a small typo in the FPR formula, it should be <code>false positive (non-default customers with Z &lt; c) / total negatives (all non-default customers)</code> right?</p>",
      "rawMarkdown": "Thanks for the detailed explanation! I noticed there's a small typo in the FPR formula, it should be `false positive (non-default customers with Z < c) / total negatives (all non-default customers)` right?",
      "votes": null
    },
    {
      "id": "1834625",
      "postDate": "06/27/2022 05:11:27",
      "content": "<p>Yes, you are right and I fixed it now. Thank you so much for letting me know!</p>",
      "rawMarkdown": "Yes, you are right and I fixed it now. Thank you so much for letting me know!",
      "votes": null
    },
    {
      "id": "1835232",
      "postDate": "06/27/2022 15:34:09",
      "content": "<p>Great stuff, thanks for your clear explanation 🙌</p>",
      "rawMarkdown": "Great stuff, thanks for your clear explanation 🙌",
      "votes": null
    },
    {
      "id": "1835648",
      "postDate": "06/28/2022 01:07:23",
      "content": "<p><a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> thanks for breaking it down, very clear explanation 👍</p>",
      "rawMarkdown": "hanaori thanks for breaking it down, very clear explanation 👍",
      "votes": null
    },
    {
      "id": "1836011",
      "postDate": "06/28/2022 09:24:47",
      "content": "<p>Great explanation. Thank you for sharing it!</p>",
      "rawMarkdown": "Great explanation. Thank you for sharing it!",
      "votes": null
    },
    {
      "id": "1836039",
      "postDate": "06/28/2022 10:00:35",
      "content": "<p>Very well presented.Thanks for sharing <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> </p>",
      "rawMarkdown": "Very well presented.Thanks for sharing @hanaori",
      "votes": null
    },
    {
      "id": "1836542",
      "postDate": "06/28/2022 19:17:37",
      "content": "<p>This is really helpful.<br>\nThank you so much for sharing it.</p>",
      "rawMarkdown": "This is really helpful.\nThank you so much for sharing it.",
      "votes": null
    },
    {
      "id": "1836887",
      "postDate": "06/29/2022 07:34:59",
      "content": "<p>well explained. but for me its little hard to understand</p>",
      "rawMarkdown": "well explained. but for me its little hard to understand",
      "votes": null
    },
    {
      "id": "1837466",
      "postDate": "06/29/2022 16:06:38",
      "content": "<p>Awesome work!<br>\nThanks!</p>",
      "rawMarkdown": "Awesome work!\nThanks!",
      "votes": null
    },
    {
      "id": "1838117",
      "postDate": "06/30/2022 08:19:06",
      "content": "<p>Thanks for your work, I like the pictures that you explain the principal.</p>",
      "rawMarkdown": "Thanks for your work, I like the pictures that you explain the principal.",
      "votes": null
    },
    {
      "id": "1838182",
      "postDate": "06/30/2022 09:53:55",
      "content": "<p><a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> Kudos to your efforts and a big thanks for bringing about this informative piece. Your article made much of my concepts regarding the metrics clear. Upvoted!</p>",
      "rawMarkdown": "hanaori Kudos to your efforts and a big thanks for bringing about this informative piece. Your article made much of my concepts regarding the metrics clear. Upvoted!",
      "votes": null
    },
    {
      "id": "1839113",
      "postDate": "07/01/2022 06:16:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a>, this is a approachable and clear explanation of the metrics. Your diagrams are beautiful. Thank you!</p>",
      "rawMarkdown": "Hi @hanaori, this is a approachable and clear explanation of the metrics. Your diagrams are beautiful. Thank you!",
      "votes": null
    },
    {
      "id": "1839237",
      "postDate": "07/01/2022 08:16:22",
      "content": "<p>thanks for sharing helped me a lot👍</p>",
      "rawMarkdown": "thanks for sharing helped me a lot👍",
      "votes": null
    },
    {
      "id": "1839523",
      "postDate": "07/01/2022 13:28:39",
      "content": "<p>Thanks a lot for sharing, very helpful info</p>",
      "rawMarkdown": "Thanks a lot for sharing, very helpful info",
      "votes": null
    },
    {
      "id": "1839855",
      "postDate": "07/01/2022 18:29:39",
      "content": "<p>Thank you !<br>\nIt’s very helpful and clearly explained</p>",
      "rawMarkdown": "Thank you !\nIt’s very helpful and clearly explained",
      "votes": null
    },
    {
      "id": "1842099",
      "postDate": "07/03/2022 17:43:05",
      "content": "<p>Haha love your notes! So informative (and cute) :)</p>",
      "rawMarkdown": "Haha love your notes! So informative (and cute) :)",
      "votes": null
    },
    {
      "id": "1843659",
      "postDate": "07/05/2022 04:35:27",
      "content": "<p>great insights</p>",
      "rawMarkdown": "great insights",
      "votes": null
    },
    {
      "id": "1847570",
      "postDate": "07/08/2022 02:04:43",
      "content": "<p>Thank you so much for a patient explanation of the evaluation metric!</p>",
      "rawMarkdown": "Thank you so much for a patient explanation of the evaluation metric!",
      "votes": null
    },
    {
      "id": "1848746",
      "postDate": "07/08/2022 21:26:11",
      "content": "<p>Thank you very much, it is incredibly helpful for this challenge.</p>",
      "rawMarkdown": "Thank you very much, it is incredibly helpful for this challenge.",
      "votes": null
    },
    {
      "id": "1860008",
      "postDate": "07/18/2022 05:53:20",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "1884766",
      "postDate": "08/04/2022 16:25:29",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a>. Very clear explanation!</p>",
      "rawMarkdown": "Thank you @hanaori. Very clear explanation!",
      "votes": null
    },
    {
      "id": "1891021",
      "postDate": "08/09/2022 07:33:45",
      "content": "<p>Great topic, thanks a lot!</p>",
      "rawMarkdown": "Great topic, thanks a lot!",
      "votes": null
    },
    {
      "id": "2011931",
      "postDate": "11/01/2022 00:21:27",
      "content": "<p>Congratulations! Thank you very much for sharing! It help me a lot to understand the topic!</p>",
      "rawMarkdown": "Congratulations! Thank you very much for sharing! It help me a lot to understand the topic!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1834392,
      "author_name": "sharmoul",
      "author_url": "",
      "post_date": "06/26/2022 22:10:43",
      "content": "<p>Very clearly explained!</p>",
      "votes": null,
      "replies": [
        {
          "id": 1834442,
          "author_name": "hanaori",
          "author_url": "",
          "post_date": "06/26/2022 23:43:15",
          "content": "<p>Thank you for reading 😊</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1834511,
      "author_name": "arti1117",
      "author_url": "",
      "post_date": "06/27/2022 02:16:54",
      "content": "<p>Thanks for sharing and teaching! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1834551,
      "author_name": "matthewchau99",
      "author_url": "",
      "post_date": "06/27/2022 03:25:43",
      "content": "<p>Thanks for the detailed explanation! I noticed there's a small typo in the FPR formula, it should be <code>false positive (non-default customers with Z &lt; c) / total negatives (all non-default customers)</code> right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1834625,
          "author_name": "hanaori",
          "author_url": "",
          "post_date": "06/27/2022 05:11:27",
          "content": "<p>Yes, you are right and I fixed it now. Thank you so much for letting me know!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1835232,
      "author_name": "niekvanderzwaag",
      "author_url": "",
      "post_date": "06/27/2022 15:34:09",
      "content": "<p>Great stuff, thanks for your clear explanation 🙌</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1835648,
      "author_name": "ganiesenov",
      "author_url": "",
      "post_date": "06/28/2022 01:07:23",
      "content": "<p><a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> thanks for breaking it down, very clear explanation 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836011,
      "author_name": "mohammadrahmati",
      "author_url": "",
      "post_date": "06/28/2022 09:24:47",
      "content": "<p>Great explanation. Thank you for sharing it!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836039,
      "author_name": "omkargowda",
      "author_url": "",
      "post_date": "06/28/2022 10:00:35",
      "content": "<p>Very well presented.Thanks for sharing <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836542,
      "author_name": "eugenetomashy",
      "author_url": "",
      "post_date": "06/28/2022 19:17:37",
      "content": "<p>This is really helpful.<br>\nThank you so much for sharing it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1836887,
      "author_name": "xodeum",
      "author_url": "",
      "post_date": "06/29/2022 07:34:59",
      "content": "<p>well explained. but for me its little hard to understand</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1837466,
      "author_name": "kanhaiyaisi",
      "author_url": "",
      "post_date": "06/29/2022 16:06:38",
      "content": "<p>Awesome work!<br>\nThanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1838117,
      "author_name": "aidoger",
      "author_url": "",
      "post_date": "06/30/2022 08:19:06",
      "content": "<p>Thanks for your work, I like the pictures that you explain the principal.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1838182,
      "author_name": "sandeepmnair",
      "author_url": "",
      "post_date": "06/30/2022 09:53:55",
      "content": "<p><a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a> Kudos to your efforts and a big thanks for bringing about this informative piece. Your article made much of my concepts regarding the metrics clear. Upvoted!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1839113,
      "author_name": "delwynnlee",
      "author_url": "",
      "post_date": "07/01/2022 06:16:24",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a>, this is a approachable and clear explanation of the metrics. Your diagrams are beautiful. Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1839237,
      "author_name": "sarayaghoobi",
      "author_url": "",
      "post_date": "07/01/2022 08:16:22",
      "content": "<p>thanks for sharing helped me a lot👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1839523,
      "author_name": "alsawymohamed",
      "author_url": "",
      "post_date": "07/01/2022 13:28:39",
      "content": "<p>Thanks a lot for sharing, very helpful info</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1839855,
      "author_name": "houzhangyi",
      "author_url": "",
      "post_date": "07/01/2022 18:29:39",
      "content": "<p>Thank you !<br>\nIt’s very helpful and clearly explained</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1842099,
      "author_name": "simpletonwang",
      "author_url": "",
      "post_date": "07/03/2022 17:43:05",
      "content": "<p>Haha love your notes! So informative (and cute) :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1843659,
      "author_name": "atfandeshmukh",
      "author_url": "",
      "post_date": "07/05/2022 04:35:27",
      "content": "<p>great insights</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1847570,
      "author_name": "chinard1",
      "author_url": "",
      "post_date": "07/08/2022 02:04:43",
      "content": "<p>Thank you so much for a patient explanation of the evaluation metric!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1848746,
      "author_name": "pmassouf",
      "author_url": "",
      "post_date": "07/08/2022 21:26:11",
      "content": "<p>Thank you very much, it is incredibly helpful for this challenge.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1860008,
      "author_name": "rodrigostallsikora",
      "author_url": "",
      "post_date": "07/18/2022 05:53:20",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1884766,
      "author_name": "gderivas",
      "author_url": "",
      "post_date": "08/04/2022 16:25:29",
      "content": "<p>Thank you <a href=\"https://www.kaggle.com/hanaori\" target=\"_blank\">@hanaori</a>. Very clear explanation!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1891021,
      "author_name": "kaggledummie007",
      "author_url": "",
      "post_date": "08/09/2022 07:33:45",
      "content": "<p>Great topic, thanks a lot!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2011931,
      "author_name": "sanechek",
      "author_url": "",
      "post_date": "11/01/2022 00:21:27",
      "content": "<p>Congratulations! Thank you very much for sharing! It help me a lot to understand the topic!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1833422": "This competition gives us a metric \\\\(M = 0.5 \\cdot (G + D)\\\\). It looks simple at a quick glance with the average of \\\\(G\\\\) and \\\\(D\\\\), but the important part \\\\(G\\\\) and \\\\(D\\\\) is a little bit complicated to understand. I would like to go through these elements one by one. (I apologize if my English isn't very good.)\n\n## ■ G: Normalized Gini Coefficient\n\n### Fundamental Knowledge (ROC curve and AUC)\n\nFirst, in order to understand and calculate the Gini coefficient, we will review the receiver operating characteristic (ROC) curve and the area under the curve (AUC).\n\n#### 1. Creating the ROC curve\nThe ROC curve is a curve that can visually capture whether a customer actually defaulted or not when the predicted default rates are ordered. Suppose we have 4 defaulting customers and 8 non-defaulting customers. The following is a rough description of what we are going to do.\n\n```\n1. Draw a diagram divided vertically into 4 sections and horizontally into 8 sections.\n2. Extend the line directly above if it is a defaulting customer and directly across if it is a non-defaulting customer.\n```\n\nLet us assume that the predicted default rate is calculated as follows: D is the actual defaulting customer and ND is the non-defaulting customer.\n\n<div style=\"text-align: center;\">\n<img width=800 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Ffdd48f6c2d1d9ef82ae205b25e9cd6fb%2FIMG_0727.jpg?generation=1656141318176769&alt=media\" />\n</div>\n\nNow let's draw a graph based on the example above (↑). Starting from (0, 0), we will draw the ROC curve in the order of increasing the default prediction score (corresponding to the blue line in the image below).\n\n```\n1. the customer with the highest score (B) is defaulting, so it goes straight up 1/4. \n2. the customer with the next highest score (G) is not defaulting, so it goes straight across 1/8. \n3. the customer with the next highest score (C) is defaulting, so it goes straight up 1/4.\n....\n```\nThen we can draw a line that connects the origin (0, 0) and the origin (1, 1) for all customers.\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F0f8a41a6fad5681ad85a16528aef1553%2FIMG_CA31680E6CB1-1.jpeg?generation=1656287940943246&alt=media\" />\n</div>\n\nIf the model is able to determine whether the customer has defaulted or not completely (perfect model), the ROC curve is drawn like a red line.\n\n#### 2. Calculating AUC\nLet's try to calculate the AUC (the blue area) of the model we created. Since there are 26 rectangles of 1/4 × 1/8 per square, we can find the value of the AUC as follows.\n\n$$AUC = \\frac{1}{4} \\cdot \\frac{1}{8} \\cdot 26 = \\frac{13}{16}$$\n\nIn this example, we can see that the AUC is \\\\(\\frac{13}{16}\\\\).\n\n```\n- The range of AUC is from 0 to 1.\n- 1 for perfect models and 0.5 for random models.\n```\n\n### The Normalized Gini Coefficient\n\nAs with the AUC, we will consider the area. A well-known measure of income inequality is the Gini coefficient, which is calculated by doubling the area between the Lorenz curve and the diagonal line. Also, the Gini coefficient calculated using the ROC curve when evaluating a model is defined as \"the area between the ROC curve and the diagonal line\". The normalized Gini coefficient is the ratio of the model created by us to the perfect model.\n\n$$Gini = A$$\n$$Normalized \\, Gini = \\frac{Gini_{model A}}{Gini_{perfect \\, model}} = \\frac{A}{A+B} = 2A$$\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fa7bfd3b76875b5e06a2b617b7ff240a7%2FIMG_5CA49BBB16DC-1.jpeg?generation=1656136752796766&alt=media\" />\n</div>\n\n```\n- The range of the Gini coefficient is from -0.5 to 0.5, and the range of the normalized Gini coefficient is from -1 to 1. \n- The purpose of the normalized Gini coefficient is to scale the AUC (= 0 for the random model and 1 for the perfect model).\n```\nThe AUC represents the percentage of the area of the perfect model, and the (normal) Gini coefficient represents the percentage based on the diagonal (random model).\n\nSo far, we have seen how the ROC curve is drawn and how the AUC and normalized Gini coefficient are calculated. It was a little difficult for me to understand the Gini coefficient at first because I had never seen it used in a way not to measure income inequality, but it was interesting to learn that it can also be used to measure the ability to rank in such a way.\n\n## ■ D: The Default Rate Captured at 4%\n\nList the predicted default rates of the customers in descending order. Customers whose rank \\\\(Z\\\\) (top percentile) is higher than the threshold \\\\(c\\\\) (top 4% = 0.04 in this case) are identified as defaults, and customers whose rank is below the threshold are identified as non-defaults.\n\nTo understand the evaluation indicator \\\\(D\\\\), the probability of being correctly or incorrectly discriminated, which is determined when we decide the threshold, is important.\n\n<img width=700 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F4617d27fecf0273b79f9e65fb25ceada%2F2022-06-27%2014.03.34.png?generation=1656306263171656&alt=media\" />\n<img width=700 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F1eb7b4bf2182090be20c1aa87780f91e%2F2022-06-27%2013.58.31_fpr.png?generation=1656306087392414&alt=media\" />\n\n\nThe FPR represents \"how many non-defaulting customers are incorrectly identified as defaulting,\" and the TNR represents \"how many non-defaulting customers are correctly identified as non-defaulting.\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Facfc568279e90ecd426af59096aef760%2FIMG_11CED5E1921A-1.jpeg?generation=1656137055041828&alt=media\" />\n</div>\n- Horizontal axis: FPR(c) + TNR(c) = 1\n- Vertical axis: TPR(c) + FNR(c) = 1\n\nIf we have two models with the same Gini Coefficient, the default rate captured by the 4% (TPR) will change. In the two models below, the value of \\\\(D\\\\) in \\\\(M = 0.5 \\cdot (G + D)\\\\) is higher in model B, so the final score (\\\\(M\\\\)) is also higher in model B (assuming the Gini coefficients \\\\(G\\\\) for Model A and Model B in the figure below are the same.)\n\n<div style=\"text-align: center;\">\n<img width=600 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2F2ef5c997d2de5e9ce5c8f7c56fcd5b16%2FIMG_0D77648FC3E1-1.jpeg?generation=1656137220409055&alt=media\" />\n</div>\n\n## Meaning of combining the two indicators\n\n```\nG: Overall correct response rate \nD: Percentage of default customers correctly identified as default customers at the TOP 4% point\n```\n\n\\\\(G\\\\) and \\\\(D\\\\) each has the above meaning. In other words, in addition to a high overall correct response rate, many defaulting customers must be in the top 4% when sorted in order of predicted default rate (i.e., we should not miss the prediction for customers with high predicted default rates as much as possible).\n\n<div style=\"text-align: center;\">\n<img width=800 src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1642598%2Fdb7315b0ed2a73cbeb2560d13d8dd7b6%2FIMG_B873FB421C0B-1.jpeg?generation=1656137321111691&alt=media\" />\n</div>\n\nI thought I could just use either \\\\(G\\\\) or \\\\(D\\\\), but when I understood the logic behind each, mixing both results made a lot of sense. I hope this helps you understand the metric a little better :) \n\nThis post is translated version (JP -> EN). I will post an original Japanese article on note.com after this competition ends.\n-> https://note.com/hanaori/n/n5ced61443058\n\n## References\n- https://www.kaggle.com/code/inversion/amex-competition-metric-python\n- https://www.asakura.co.jp/detail.php?book_code=29541\n- https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464",
    "1834392": "Very clearly explained!",
    "1834442": "Thank you for reading 😊",
    "1834511": "Thanks for sharing and teaching!",
    "1834551": "Thanks for the detailed explanation! I noticed there's a small typo in the FPR formula, it should be `false positive (non-default customers with Z < c) / total negatives (all non-default customers)` right?",
    "1834625": "Yes, you are right and I fixed it now. Thank you so much for letting me know!",
    "1835232": "Great stuff, thanks for your clear explanation 🙌",
    "1835648": "hanaori thanks for breaking it down, very clear explanation 👍",
    "1836011": "Great explanation. Thank you for sharing it!",
    "1836039": "Very well presented.Thanks for sharing @hanaori",
    "1836542": "This is really helpful.\nThank you so much for sharing it.",
    "1836887": "well explained. but for me its little hard to understand",
    "1837466": "Awesome work!\nThanks!",
    "1838117": "Thanks for your work, I like the pictures that you explain the principal.",
    "1838182": "hanaori Kudos to your efforts and a big thanks for bringing about this informative piece. Your article made much of my concepts regarding the metrics clear. Upvoted!",
    "1839113": "Hi @hanaori, this is a approachable and clear explanation of the metrics. Your diagrams are beautiful. Thank you!",
    "1839237": "thanks for sharing helped me a lot👍",
    "1839523": "Thanks a lot for sharing, very helpful info",
    "1839855": "Thank you !\nIt’s very helpful and clearly explained",
    "1842099": "Haha love your notes! So informative (and cute) :)",
    "1843659": "great insights",
    "1847570": "Thank you so much for a patient explanation of the evaluation metric!",
    "1848746": "Thank you very much, it is incredibly helpful for this challenge.",
    "1860008": "Thank you for sharing!",
    "1884766": "Thank you @hanaori. Very clear explanation!",
    "1891021": "Great topic, thanks a lot!",
    "2011931": "Congratulations! Thank you very much for sharing! It help me a lot to understand the topic!"
  },
  "source": "meta"
}