{
  "id": 327464,
  "title": "Graphical explanation of the competition metric",
  "url": "/competitions/amex-default-prediction/discussion/327464",
  "author_name": "AmbrosM",
  "post_date": "2022-05-27T10:33:32.909000",
  "votes": 254,
  "comment_count": 31,
  "views": 0,
  "content": "<p>The competition metric has a simple graphical explanation (assuming that you know the <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#roc-metrics\" target=\"_blank\">ROC (receiver operating characteristic) curve</a>). The following diagram shows the ROC curve of my model in dark red. The area under the curve (AUC) is filled with light red. The green line corresponds to 4 % positive predictions.</p>\n<p><img src=\"https://i.imgur.com/vpJTP8a.png\" alt=\"roc-curve\"></p>\n<p>The competition metric has two components: the normalized Gini coefficient and the default rate captured at 4 %:</p>\n<ul>\n<li>The <em>normalized Gini coefficient</em> is simply a stretched AUC: AUC is the light red area under the curve, which has a value between 0 and 1. The normalized Gini coefficient is equal to 2*AUC-1 and is between -1 and 1. The larger the red area, the better is the score.</li>\n<li>The <em>default rate captured at 4 %</em> is the true positive rate (recall) for a threshold set at 4 % of the total (weighted) sample count. It corresponds to the y coordinate of the intersection between the green line and the red roc curve (marked with a green dot) and is always between 0 and 1. The higher the intersection point, the better is the score.</li>\n</ul>\n<p>The competition metric is the average of these two components. In other words: They want us to simultaneously optimize for a large red area under the curve and a high intersection point with the green line.</p>\n<p>You can estimate the score of my model by looking at the diagram: The red area covers 96 % of the diagram, so the normalized Gini coefficient is 0.92. The green dot has a y coordinate of 0.65 - this is the default rate captured at 4 %. The competition score is the average of both: (0.92 + 0.65) / 2 = 0.78.</p>\n<p>The detailed calculation:</p>\n<pre><code>Area under the curve (AUC):         0.95968\n2*AUC-1:                            0.91936\nNormalized Gini coefficient:        0.91936 (= 2*AUC-1)\n\nPositive samples in validation set:   23765\nNegative samples weighted:          1360340 (unweighted: 68017)\nTotal samples weighted:             1384105\n4 % of Total samples weighted:        55364\nTrue positives at this threshold:     15383\nFalse positives at this threshold:    40040\nDefault rate captured at 4%:        0.64730 (= 15383 / 23765)\n\nCompetition score:                  0.78333 (= (0.91936 + 0.64730) / 2)\n</code></pre>\n<p>The source code for the diagram is <a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\">available here</a>, and more background about the competition metric has been given by <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">here</a> (and congratulations for becoming grandmaster!).</p>",
  "messages": [
    {
      "id": 1802957,
      "postDate": "2022-05-27T10:33:32.910Z",
      "content": "<p>The competition metric has a simple graphical explanation (assuming that you know the <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#roc-metrics\" target=\"_blank\">ROC (receiver operating characteristic) curve</a>). The following diagram shows the ROC curve of my model in dark red. The area under the curve (AUC) is filled with light red. The green line corresponds to 4 % positive predictions.</p>\n<p><img src=\"https://i.imgur.com/vpJTP8a.png\" alt=\"roc-curve\"></p>\n<p>The competition metric has two components: the normalized Gini coefficient and the default rate captured at 4 %:</p>\n<ul>\n<li>The <em>normalized Gini coefficient</em> is simply a stretched AUC: AUC is the light red area under the curve, which has a value between 0 and 1. The normalized Gini coefficient is equal to 2*AUC-1 and is between -1 and 1. The larger the red area, the better is the score.</li>\n<li>The <em>default rate captured at 4 %</em> is the true positive rate (recall) for a threshold set at 4 % of the total (weighted) sample count. It corresponds to the y coordinate of the intersection between the green line and the red roc curve (marked with a green dot) and is always between 0 and 1. The higher the intersection point, the better is the score.</li>\n</ul>\n<p>The competition metric is the average of these two components. In other words: They want us to simultaneously optimize for a large red area under the curve and a high intersection point with the green line.</p>\n<p>You can estimate the score of my model by looking at the diagram: The red area covers 96 % of the diagram, so the normalized Gini coefficient is 0.92. The green dot has a y coordinate of 0.65 - this is the default rate captured at 4 %. The competition score is the average of both: (0.92 + 0.65) / 2 = 0.78.</p>\n<p>The detailed calculation:</p>\n<pre><code>Area under the curve (AUC):         0.95968\n2*AUC-1:                            0.91936\nNormalized Gini coefficient:        0.91936 (= 2*AUC-1)\n\nPositive samples in validation set:   23765\nNegative samples weighted:          1360340 (unweighted: 68017)\nTotal samples weighted:             1384105\n4 % of Total samples weighted:        55364\nTrue positives at this threshold:     15383\nFalse positives at this threshold:    40040\nDefault rate captured at 4%:        0.64730 (= 15383 / 23765)\n\nCompetition score:                  0.78333 (= (0.91936 + 0.64730) / 2)\n</code></pre>\n<p>The source code for the diagram is <a href=\"https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training\" target=\"_blank\">available here</a>, and more background about the competition metric has been given by <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a> <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">here</a> (and congratulations for becoming grandmaster!).</p>",
      "rawMarkdown": "The competition metric has a simple graphical explanation (assuming that you know the [ROC (receiver operating characteristic) curve](https://scikit-learn.org/stable/modules/model_evaluation.html#roc-metrics)). The following diagram shows the ROC curve of my model in dark red. The area under the curve (AUC) is filled with light red. The green line corresponds to 4 % positive predictions.\n\n![roc-curve](https://i.imgur.com/vpJTP8a.png)\n\nThe competition metric has two components: the normalized Gini coefficient and the default rate captured at 4 %:\n- The *normalized Gini coefficient* is simply a stretched AUC: AUC is the light red area under the curve, which has a value between 0 and 1. The normalized Gini coefficient is equal to 2\\*AUC-1 and is between -1 and 1. The larger the red area, the better is the score.\n- The *default rate captured at 4 %* is the true positive rate (recall) for a threshold set at 4 % of the total (weighted) sample count. It corresponds to the y coordinate of the intersection between the green line and the red roc curve (marked with a green dot) and is always between 0 and 1. The higher the intersection point, the better is the score.\n\nThe competition metric is the average of these two components. In other words: They want us to simultaneously optimize for a large red area under the curve and a high intersection point with the green line.\n\nYou can estimate the score of my model by looking at the diagram: The red area covers 96 % of the diagram, so the normalized Gini coefficient is 0.92. The green dot has a y coordinate of 0.65 - this is the default rate captured at 4 %. The competition score is the average of both: (0.92 + 0.65) / 2 = 0.78.\n\nThe detailed calculation:\n```\nArea under the curve (AUC):         0.95968\n2*AUC-1:                            0.91936\nNormalized Gini coefficient:        0.91936 (= 2*AUC-1)\n\nPositive samples in validation set:   23765\nNegative samples weighted:          1360340 (unweighted: 68017)\nTotal samples weighted:             1384105\n4 % of Total samples weighted:        55364\nTrue positives at this threshold:     15383\nFalse positives at this threshold:    40040\nDefault rate captured at 4%:        0.64730 (= 15383 / 23765)\n\nCompetition score:                  0.78333 (= (0.91936 + 0.64730) / 2)\n```\n\nThe source code for the diagram is [available here](https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training), and more background about the competition metric has been given by @mpwolke [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116) (and congratulations for becoming grandmaster!).",
      "votes": 254
    },
    {
      "id": 1803000,
      "postDate": "2022-05-27T11:11:42.467Z",
      "content": "<p>2<em>auc-1 = 2</em>(auc-0.5) corresponds to removing the triangle below the blue dotted line. In other words gini corresponds to the ratio between the area between the red line and the blue dotted line over (that same area + the white area), giving the  whole triangle above the blue dotted line. thanks for clarifying the other metric.</p>",
      "rawMarkdown": "2*auc-1 = 2*(auc-0.5) corresponds to removing the triangle below the blue dotted line. In other words gini corresponds to the ratio between the area between the red line and the blue dotted line over (that same area + the white area), giving the  whole triangle above the blue dotted line. thanks for clarifying the other metric.",
      "votes": 7
    },
    {
      "id": 1803083,
      "postDate": "2022-05-27T12:53:24.200Z",
      "content": "<p>This is the first time I've heard of normalized gini coefficient. Is there any advantages of using it over AUC?</p>",
      "rawMarkdown": "This is the first time I've heard of normalized gini coefficient. Is there any advantages of using it over AUC?",
      "votes": 3,
      "replies": [
        {
          "id": 1803097,
          "postDate": "2022-05-27T13:17:50.050Z",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> I don't see any advantage in this competition. In  <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a>'s <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">discussion</a> they say that the normalized Gini coefficient is more general than AUC and can be used for regression problems as well.</p>",
          "rawMarkdown": "@gunesevitan I don't see any advantage in this competition. In  @mpwolke's [discussion](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116) they say that the normalized Gini coefficient is more general than AUC and can be used for regression problems as well.",
          "votes": 6
        },
        {
          "id": 1805387,
          "postDate": "2022-05-30T05:48:13.353Z",
          "rawMarkdown": "",
          "votes": 7,
          "isDeleted": true
        },
        {
          "id": 1812302,
          "postDate": "2022-06-05T16:39:28.190Z",
          "content": "<p>This is interesting, so basically it makes negative competition metric possible is the only reason?</p>",
          "rawMarkdown": "This is interesting, so basically it makes negative competition metric possible is the only reason?",
          "votes": 1
        },
        {
          "id": 1812833,
          "postDate": "2022-06-06T09:07:47.440Z",
          "content": "<p>AUC also have some interpretation. I don't remember where I saw the proof exactly. But it goes along the line of: AUC is the area (integral) of something (FPR) that can be defined as an integral. You can swap both integral. After some analysis you end up with something that can be interpreted as the ability to distinguish bettween positive and negative score. From Wikipedia: <strong>\"when given one randomly selected positive instance and one randomly selected negative instance, AUC is the probability that the classifier will be able to tell which one is which\"</strong>.  This is the basis for AUC approximation / statistical testing.<br>\nEdit: wikipedia has the proof. </p>",
          "rawMarkdown": "AUC also have some interpretation. I don't remember where I saw the proof exactly. But it goes along the line of: AUC is the area (integral) of something (FPR) that can be defined as an integral. You can swap both integral. After some analysis you end up with something that can be interpreted as the ability to distinguish bettween positive and negative score. From Wikipedia: **\"when given one randomly selected positive instance and one randomly selected negative instance, AUC is the probability that the classifier will be able to tell which one is which\"**.  This is the basis for AUC approximation / statistical testing.\nEdit: wikipedia has the proof. \n",
          "votes": 1
        }
      ]
    },
    {
      "id": 1898616,
      "postDate": "2022-08-14T17:12:54.260Z",
      "content": "<p>So much to learn from this post! <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "So much to learn from this post! @ambrosm ",
      "votes": 1
    },
    {
      "id": 1883759,
      "postDate": "2022-08-04T04:40:41.683Z",
      "content": "<p>The next time I'll first search in the discussion section before to try to understand everything by my own. Great explanation.</p>",
      "rawMarkdown": "The next time I'll first search in the discussion section before to try to understand everything by my own. Great explanation.",
      "votes": 1
    },
    {
      "id": 1869878,
      "postDate": "2022-07-25T05:42:45.167Z",
      "content": "<p>great job!</p>",
      "rawMarkdown": "great job!",
      "votes": 1
    },
    {
      "id": 1865926,
      "postDate": "2022-07-22T07:03:36.983Z",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> Simply 💕💕💕</p>\n<p>Upvoted and Bookmarked ☀️</p>",
      "rawMarkdown": "@ambrosm Simply 💕💕💕\n\nUpvoted and Bookmarked ☀️",
      "votes": 1
    },
    {
      "id": 1858712,
      "postDate": "2022-07-17T06:37:00.950Z",
      "content": "<p>Great job!!</p>",
      "rawMarkdown": "Great job!!",
      "votes": 1
    },
    {
      "id": 1822004,
      "postDate": "2022-06-16T02:15:25.617Z",
      "content": "<p>Wow, you always bring insights to every post. Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>👍🙌</p>",
      "rawMarkdown": "Wow, you always bring insights to every post. Thank you @ambrosm👍🙌",
      "votes": 1
    },
    {
      "id": 1812293,
      "postDate": "2022-06-05T16:34:31.027Z",
      "content": "<p>This is incredible…</p>",
      "rawMarkdown": "This is incredible...",
      "votes": 1
    },
    {
      "id": 1812227,
      "postDate": "2022-06-05T15:39:18.373Z",
      "content": "<p>I really had a hard time understanding the matrix, great explanation thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "I really had a hard time understanding the matrix, great explanation thanks @ambrosm ",
      "votes": 1
    },
    {
      "id": 1806986,
      "postDate": "2022-05-31T16:20:08.613Z",
      "content": "<p>Thanks very much for developing the visualization and providing an example for the key metric within this competition. It really helps to provide insights into how the rank ordering challenge will develop.</p>",
      "rawMarkdown": "Thanks very much for developing the visualization and providing an example for the key metric within this competition. It really helps to provide insights into how the rank ordering challenge will develop.",
      "votes": 1
    },
    {
      "id": 1805331,
      "postDate": "2022-05-30T04:32:58.713Z",
      "content": "<p>Very well illustrated explanation. thanks for sharing 🙌 </p>",
      "rawMarkdown": "Very well illustrated explanation. thanks for sharing 🙌 ",
      "votes": 1
    },
    {
      "id": 1805137,
      "postDate": "2022-05-29T20:05:49.250Z",
      "content": "<p>This is an interesting metric and you make an excellent job explaining and illustrating it!</p>",
      "rawMarkdown": "This is an interesting metric and you make an excellent job explaining and illustrating it!",
      "votes": 1
    },
    {
      "id": 1861730,
      "postDate": "2022-07-19T07:42:35.150Z",
      "content": "<p>Thanks for simplicity explanation.<br>\n<em>The default rate captured at 4 %</em> is sometimes called 「ρTPR(ρ=0.04)」 from an anomaly detection standpoint, but this may not be so famous.</p>",
      "rawMarkdown": "Thanks for simplicity explanation.\n*The default rate captured at 4 %* is sometimes called 「ρTPR(ρ=0.04)」 from an anomaly detection standpoint, but this may not be so famous.",
      "votes": 2
    },
    {
      "id": 1851364,
      "postDate": "2022-07-11T07:32:08Z",
      "content": "<p>Extremely useful and clearly explained!</p>",
      "rawMarkdown": "Extremely useful and clearly explained!",
      "votes": 2
    },
    {
      "id": 1811789,
      "postDate": "2022-06-05T06:20:58.593Z",
      "content": "<p>Wow~ l finally understand the competition metric! thanks a lot👍</p>",
      "rawMarkdown": "Wow~ l finally understand the competition metric! thanks a lot👍",
      "votes": 2
    },
    {
      "id": 1803206,
      "postDate": "2022-05-27T14:58:33.077Z",
      "content": "<p>I hope that what I copied about Normalized Gini is correct. All the sources are there on the topic, since I was trying to learn and understand normalized Gini.</p>\n<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>   It's a honor to me receiving congratulations from one of the most brilliant GM that joined Kaggle on the last months.</p>",
      "rawMarkdown": "I hope that what I copied about Normalized Gini is correct. All the sources are there on the topic, since I was trying to learn and understand normalized Gini.\n\n@ambrosm   It's a honor to me receiving congratulations from one of the most brilliant GM that joined Kaggle on the last months.",
      "votes": 2
    },
    {
      "id": 1899692,
      "postDate": "2022-08-15T12:59:45.390Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1899686,
      "postDate": "2022-08-15T12:54:37.097Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 1898102,
      "postDate": "2022-08-14T09:53:52.863Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 1
    },
    {
      "id": 1896584,
      "postDate": "2022-08-13T00:01:49.370Z",
      "content": "<p>oh… thanks for sharing.</p>",
      "rawMarkdown": "oh... thanks for sharing.",
      "votes": 1
    },
    {
      "id": 1881037,
      "postDate": "2022-08-02T07:52:00.963Z",
      "content": "<p>Thanks for the helpful explanation!</p>",
      "rawMarkdown": "Thanks for the helpful explanation!",
      "votes": 1
    },
    {
      "id": 1877277,
      "postDate": "2022-07-30T13:33:24.353Z",
      "content": "<p>Thanks! Really helpful!</p>",
      "rawMarkdown": "Thanks! Really helpful!",
      "votes": 1
    },
    {
      "id": 1869343,
      "postDate": "2022-07-24T17:07:43.097Z",
      "content": "<p>Great explanation, thanks a lot!</p>",
      "rawMarkdown": "Great explanation, thanks a lot!",
      "votes": 1
    },
    {
      "id": 1864682,
      "postDate": "2022-07-21T08:40:09.840Z",
      "content": "<p>Really useful. Thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "rawMarkdown": "Really useful. Thanks @ambrosm ",
      "votes": 1
    },
    {
      "id": 1864534,
      "postDate": "2022-07-21T05:48:45.690Z",
      "content": "<p>Thank you.<br>\nReally useful this</p>",
      "rawMarkdown": "Thank you.\nReally useful this",
      "votes": 1
    },
    {
      "id": 1864380,
      "postDate": "2022-07-21T02:48:01.297Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing",
      "votes": 1
    }
  ],
  "comments": [
    {
      "id": 1803000,
      "author_name": "Lucas Morin",
      "author_url": "",
      "post_date": "2022-05-27T11:11:42.467000",
      "content": "<p>2<em>auc-1 = 2</em>(auc-0.5) corresponds to removing the triangle below the blue dotted line. In other words gini corresponds to the ratio between the area between the red line and the blue dotted line over (that same area + the white area), giving the  whole triangle above the blue dotted line. thanks for clarifying the other metric.</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 1803083,
      "author_name": "Gunes Evitan",
      "author_url": "",
      "post_date": "2022-05-27T12:53:24.200000",
      "content": "<p>This is the first time I've heard of normalized gini coefficient. Is there any advantages of using it over AUC?</p>",
      "votes": 3,
      "replies": [
        {
          "id": 1803097,
          "author_name": "AmbrosM",
          "author_url": "",
          "post_date": "2022-05-27T13:17:50.050000",
          "content": "<p><a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> I don't see any advantage in this competition. In  <a href=\"https://www.kaggle.com/mpwolke\" target=\"_blank\">@mpwolke</a>'s <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116\" target=\"_blank\">discussion</a> they say that the normalized Gini coefficient is more general than AUC and can be used for regression problems as well.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 1805387,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-05-30T05:48:13.353000",
          "content": "",
          "votes": 7,
          "replies": []
        },
        {
          "id": 1812302,
          "author_name": "Kris Smith",
          "author_url": "",
          "post_date": "2022-06-05T16:39:28.190000",
          "content": "<p>This is interesting, so basically it makes negative competition metric possible is the only reason?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1812833,
          "author_name": "Lucas Morin",
          "author_url": "",
          "post_date": "2022-06-06T09:07:47.440000",
          "content": "<p>AUC also have some interpretation. I don't remember where I saw the proof exactly. But it goes along the line of: AUC is the area (integral) of something (FPR) that can be defined as an integral. You can swap both integral. After some analysis you end up with something that can be interpreted as the ability to distinguish bettween positive and negative score. From Wikipedia: <strong>\"when given one randomly selected positive instance and one randomly selected negative instance, AUC is the probability that the classifier will be able to tell which one is which\"</strong>.  This is the basis for AUC approximation / statistical testing.<br>\nEdit: wikipedia has the proof. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1898616,
      "author_name": "Apurba Pandey",
      "author_url": "",
      "post_date": "2022-08-14T17:12:54.260000",
      "content": "<p>So much to learn from this post! <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1883759,
      "author_name": "Israel Rescalvo",
      "author_url": "",
      "post_date": "2022-08-04T04:40:41.683000",
      "content": "<p>The next time I'll first search in the discussion section before to try to understand everything by my own. Great explanation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1869878,
      "author_name": "Kunlun Zhu",
      "author_url": "",
      "post_date": "2022-07-25T05:42:45.167000",
      "content": "<p>great job!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1865926,
      "author_name": "Muhammad Irfan Azam",
      "author_url": "",
      "post_date": "2022-07-22T07:03:36.983000",
      "content": "<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> Simply 💕💕💕</p>\n<p>Upvoted and Bookmarked ☀️</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1858712,
      "author_name": "Ayush Jaiswal",
      "author_url": "",
      "post_date": "2022-07-17T06:37:00.950000",
      "content": "<p>Great job!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1822004,
      "author_name": "naiborhujosua",
      "author_url": "",
      "post_date": "2022-06-16T02:15:25.617000",
      "content": "<p>Wow, you always bring insights to every post. Thank you <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>👍🙌</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1812293,
      "author_name": "Kris Smith",
      "author_url": "",
      "post_date": "2022-06-05T16:34:31.027000",
      "content": "<p>This is incredible…</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1812227,
      "author_name": "Dhamu",
      "author_url": "",
      "post_date": "2022-06-05T15:39:18.373000",
      "content": "<p>I really had a hard time understanding the matrix, great explanation thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1806986,
      "author_name": "James McNeill",
      "author_url": "",
      "post_date": "2022-05-31T16:20:08.613000",
      "content": "<p>Thanks very much for developing the visualization and providing an example for the key metric within this competition. It really helps to provide insights into how the rank ordering challenge will develop.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1805331,
      "author_name": "Niek van der Zwaag",
      "author_url": "",
      "post_date": "2022-05-30T04:32:58.713000",
      "content": "<p>Very well illustrated explanation. thanks for sharing 🙌 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1805137,
      "author_name": "Adam Wurdits",
      "author_url": "",
      "post_date": "2022-05-29T20:05:49.250000",
      "content": "<p>This is an interesting metric and you make an excellent job explaining and illustrating it!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1861730,
      "author_name": "sqrt4kaido",
      "author_url": "",
      "post_date": "2022-07-19T07:42:35.150000",
      "content": "<p>Thanks for simplicity explanation.<br>\n<em>The default rate captured at 4 %</em> is sometimes called 「ρTPR(ρ=0.04)」 from an anomaly detection standpoint, but this may not be so famous.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1851364,
      "author_name": "Oleg Khudyakov",
      "author_url": "",
      "post_date": "2022-07-11T07:32:08",
      "content": "<p>Extremely useful and clearly explained!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1811789,
      "author_name": "SgangX",
      "author_url": "",
      "post_date": "2022-06-05T06:20:58.593000",
      "content": "<p>Wow~ l finally understand the competition metric! thanks a lot👍</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1803206,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2022-05-27T14:58:33.077000",
      "content": "<p>I hope that what I copied about Normalized Gini is correct. All the sources are there on the topic, since I was trying to learn and understand normalized Gini.</p>\n<p><a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a>   It's a honor to me receiving congratulations from one of the most brilliant GM that joined Kaggle on the last months.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1899692,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-15T12:59:45.390000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1899686,
      "author_name": "Elias",
      "author_url": "",
      "post_date": "2022-08-15T12:54:37.097000",
      "content": "<p>Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1898102,
      "author_name": "",
      "author_url": "",
      "post_date": "2022-08-14T09:53:52.863000",
      "content": "<p>thanks for sharing</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1896584,
      "author_name": "HOSUB JAMES JANG",
      "author_url": "",
      "post_date": "2022-08-13T00:01:49.370000",
      "content": "<p>oh… thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1881037,
      "author_name": "Smruti Sekhar",
      "author_url": "",
      "post_date": "2022-08-02T07:52:00.963000",
      "content": "<p>Thanks for the helpful explanation!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1877277,
      "author_name": "ChuanhaoLi2022",
      "author_url": "",
      "post_date": "2022-07-30T13:33:24.353000",
      "content": "<p>Thanks! Really helpful!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1869343,
      "author_name": "romanchio",
      "author_url": "",
      "post_date": "2022-07-24T17:07:43.097000",
      "content": "<p>Great explanation, thanks a lot!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1864682,
      "author_name": "K Kiran",
      "author_url": "",
      "post_date": "2022-07-21T08:40:09.840000",
      "content": "<p>Really useful. Thanks <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1864534,
      "author_name": "HYEON03",
      "author_url": "",
      "post_date": "2022-07-21T05:48:45.690000",
      "content": "<p>Thank you.<br>\nReally useful this</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1864380,
      "author_name": "ds.wook",
      "author_url": "",
      "post_date": "2022-07-21T02:48:01.297000",
      "content": "<p>thanks for sharing</p>",
      "votes": 1,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1802957": "The competition metric has a simple graphical explanation (assuming that you know the [ROC (receiver operating characteristic) curve](https://scikit-learn.org/stable/modules/model_evaluation.html#roc-metrics)). The following diagram shows the ROC curve of my model in dark red. The area under the curve (AUC) is filled with light red. The green line corresponds to 4 % positive predictions.\n\n![roc-curve](https://i.imgur.com/vpJTP8a.png)\n\nThe competition metric has two components: the normalized Gini coefficient and the default rate captured at 4 %:\n- The *normalized Gini coefficient* is simply a stretched AUC: AUC is the light red area under the curve, which has a value between 0 and 1. The normalized Gini coefficient is equal to 2\\*AUC-1 and is between -1 and 1. The larger the red area, the better is the score.\n- The *default rate captured at 4 %* is the true positive rate (recall) for a threshold set at 4 % of the total (weighted) sample count. It corresponds to the y coordinate of the intersection between the green line and the red roc curve (marked with a green dot) and is always between 0 and 1. The higher the intersection point, the better is the score.\n\nThe competition metric is the average of these two components. In other words: They want us to simultaneously optimize for a large red area under the curve and a high intersection point with the green line.\n\nYou can estimate the score of my model by looking at the diagram: The red area covers 96 % of the diagram, so the normalized Gini coefficient is 0.92. The green dot has a y coordinate of 0.65 - this is the default rate captured at 4 %. The competition score is the average of both: (0.92 + 0.65) / 2 = 0.78.\n\nThe detailed calculation:\n```\nArea under the curve (AUC):         0.95968\n2*AUC-1:                            0.91936\nNormalized Gini coefficient:        0.91936 (= 2*AUC-1)\n\nPositive samples in validation set:   23765\nNegative samples weighted:          1360340 (unweighted: 68017)\nTotal samples weighted:             1384105\n4 % of Total samples weighted:        55364\nTrue positives at this threshold:     15383\nFalse positives at this threshold:    40040\nDefault rate captured at 4%:        0.64730 (= 15383 / 23765)\n\nCompetition score:                  0.78333 (= (0.91936 + 0.64730) / 2)\n```\n\nThe source code for the diagram is [available here](https://www.kaggle.com/code/ambrosm/amex-keras-quickstart-1-training), and more background about the competition metric has been given by @mpwolke [here](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327116) (and congratulations for becoming grandmaster!).",
    "1803000": "2*auc-1 = 2*(auc-0.5) corresponds to removing the triangle below the blue dotted line. In other words gini corresponds to the ratio between the area between the red line and the blue dotted line over (that same area + the white area), giving the  whole triangle above the blue dotted line. thanks for clarifying the other metric.",
    "1803083": "This is the first time I've heard of normalized gini coefficient. Is there any advantages of using it over AUC?",
    "1898616": "So much to learn from this post! @ambrosm ",
    "1883759": "The next time I'll first search in the discussion section before to try to understand everything by my own. Great explanation.",
    "1869878": "great job!",
    "1865926": "@ambrosm Simply 💕💕💕\n\nUpvoted and Bookmarked ☀️",
    "1858712": "Great job!!",
    "1822004": "Wow, you always bring insights to every post. Thank you @ambrosm👍🙌",
    "1812293": "This is incredible...",
    "1812227": "I really had a hard time understanding the matrix, great explanation thanks @ambrosm ",
    "1806986": "Thanks very much for developing the visualization and providing an example for the key metric within this competition. It really helps to provide insights into how the rank ordering challenge will develop.",
    "1805331": "Very well illustrated explanation. thanks for sharing 🙌 ",
    "1805137": "This is an interesting metric and you make an excellent job explaining and illustrating it!",
    "1861730": "Thanks for simplicity explanation.\n*The default rate captured at 4 %* is sometimes called 「ρTPR(ρ=0.04)」 from an anomaly detection standpoint, but this may not be so famous.",
    "1851364": "Extremely useful and clearly explained!",
    "1811789": "Wow~ l finally understand the competition metric! thanks a lot👍",
    "1803206": "I hope that what I copied about Normalized Gini is correct. All the sources are there on the topic, since I was trying to learn and understand normalized Gini.\n\n@ambrosm   It's a honor to me receiving congratulations from one of the most brilliant GM that joined Kaggle on the last months.",
    "1899692": "",
    "1899686": "Thanks for sharing!",
    "1898102": "thanks for sharing",
    "1896584": "oh... thanks for sharing.",
    "1881037": "Thanks for the helpful explanation!",
    "1877277": "Thanks! Really helpful!",
    "1869343": "Great explanation, thanks a lot!",
    "1864682": "Really useful. Thanks @ambrosm ",
    "1864534": "Thank you.\nReally useful this",
    "1864380": "thanks for sharing"
  }
}