{
  "id": 338268,
  "title": "The curse of exotic scoring",
  "url": "/competitions/amex-default-prediction/discussion/338268",
  "author_name": "",
  "post_date": "2022-07-19T21:59:02.649433500Z",
  "votes": 39,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Never liked scoring schemes I couldn't understand, so <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\"><strong>this post</strong></a> helped me a lot. At least I can justify to some degree why they chose this as the competition score.</p>\n<p>The problem is that AmEx score is only loosely related to other quantities we can maximize/minimize during model fitting and early stopping. Here is a long LightGBM run (with DART booster; first fold shown) with each point representing 100 trees.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F647396ef061dce737f7814323ed0ad0a%2Famex-lgbm-dart.png?generation=1658266708075277&amp;alt=media\" alt=\"\"></p>\n<p>The curves are picture-perfect in general, but the iteration with best log-loss (pointed by arrow) is nowhere near the iteration with the best AmEx score.</p>\n<p>Here is a neural network run (Keras with LSTM; first fold shown) with each point representing a single epoch.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2Fadd3d9a24df5d8a7071210a6046cb4ea%2Fkeras-5fold-run-02-v5-epochs.png?generation=1658266881324817&amp;alt=media\" alt=\"\"></p>\n<p>Now, we can do early stopping based on AmEx scores rather than log-loss, but the fact remains that under the hood our models are minimizing a function (log-loss) that is a decent but not great proxy for AmEx scores.</p>\n<p>Without even knowing whether AmEx score is differentiable (presumably not), it is out of question to maximize it directly as a function because it is too expensive to compute. I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable), but even AUC scores do not track perfectly with AmEx scores because AUC/Gini is only part of the calculation.</p>\n<p>Any thoughts on this? I suspect that finding this magical function that can be directly maximized/minimized and tracks better with AmEx scores would probably clinch the top spot, so I am not holding my breath that we will hear about it before the competition ends.</p>",
  "messages": [
    {
      "id": "1862658",
      "postDate": "07/19/2022 21:59:02",
      "content": "<p>Never liked scoring schemes I couldn't understand, so <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\"><strong>this post</strong></a> helped me a lot. At least I can justify to some degree why they chose this as the competition score.</p>\n<p>The problem is that AmEx score is only loosely related to other quantities we can maximize/minimize during model fitting and early stopping. Here is a long LightGBM run (with DART booster; first fold shown) with each point representing 100 trees.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F647396ef061dce737f7814323ed0ad0a%2Famex-lgbm-dart.png?generation=1658266708075277&amp;alt=media\" alt=\"\"></p>\n<p>The curves are picture-perfect in general, but the iteration with best log-loss (pointed by arrow) is nowhere near the iteration with the best AmEx score.</p>\n<p>Here is a neural network run (Keras with LSTM; first fold shown) with each point representing a single epoch.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2Fadd3d9a24df5d8a7071210a6046cb4ea%2Fkeras-5fold-run-02-v5-epochs.png?generation=1658266881324817&amp;alt=media\" alt=\"\"></p>\n<p>Now, we can do early stopping based on AmEx scores rather than log-loss, but the fact remains that under the hood our models are minimizing a function (log-loss) that is a decent but not great proxy for AmEx scores.</p>\n<p>Without even knowing whether AmEx score is differentiable (presumably not), it is out of question to maximize it directly as a function because it is too expensive to compute. I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable), but even AUC scores do not track perfectly with AmEx scores because AUC/Gini is only part of the calculation.</p>\n<p>Any thoughts on this? I suspect that finding this magical function that can be directly maximized/minimized and tracks better with AmEx scores would probably clinch the top spot, so I am not holding my breath that we will hear about it before the competition ends.</p>",
      "rawMarkdown": "Never liked scoring schemes I couldn't understand, so [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464) helped me a lot. At least I can justify to some degree why they chose this as the competition score.\n\nThe problem is that AmEx score is only loosely related to other quantities we can maximize/minimize during model fitting and early stopping. Here is a long LightGBM run (with DART booster; first fold shown) with each point representing 100 trees.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F647396ef061dce737f7814323ed0ad0a%2Famex-lgbm-dart.png?generation=1658266708075277&alt=media)\n\nThe curves are picture-perfect in general, but the iteration with best log-loss (pointed by arrow) is nowhere near the iteration with the best AmEx score.\n\nHere is a neural network run (Keras with LSTM; first fold shown) with each point representing a single epoch.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2Fadd3d9a24df5d8a7071210a6046cb4ea%2Fkeras-5fold-run-02-v5-epochs.png?generation=1658266881324817&alt=media)\n\nNow, we can do early stopping based on AmEx scores rather than log-loss, but the fact remains that under the hood our models are minimizing a function (log-loss) that is a decent but not great proxy for AmEx scores.\n\nWithout even knowing whether AmEx score is differentiable (presumably not), it is out of question to maximize it directly as a function because it is too expensive to compute. I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable), but even AUC scores do not track perfectly with AmEx scores because AUC/Gini is only part of the calculation.\n\nAny thoughts on this? I suspect that finding this magical function that can be directly maximized/minimized and tracks better with AmEx scores would probably clinch the top spot, so I am not holding my breath that we will hear about it before the competition ends.",
      "votes": null
    },
    {
      "id": "1863030",
      "postDate": "07/20/2022 05:52:39",
      "content": "<p>Two months ago, we had a similar discussion during a TPS competition: <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-may-2022/discussion/326116\" target=\"_blank\">Early stopping on loss, auc or accuracy?</a></p>\n<p>And you might want to look at the paper <a href=\"https://arxiv.org/abs/2108.11179\" target=\"_blank\">Recall@k Surrogate Loss with Large Batches and Similarity Mixup</a>, where they show how to modify a recall metric (the d part of our score) to make it differentiable. I didn't apply the ideas of the paper to this competition, though.</p>",
      "rawMarkdown": "Two months ago, we had a similar discussion during a TPS competition: [Early stopping on loss, auc or accuracy?](https://www.kaggle.com/competitions/tabular-playground-series-may-2022/discussion/326116)\n\nAnd you might want to look at the paper [Recall@k Surrogate Loss with Large Batches and Similarity Mixup](https://arxiv.org/abs/2108.11179), where they show how to modify a recall metric (the d part of our score) to make it differentiable. I didn't apply the ideas of the paper to this competition, though.",
      "votes": null
    },
    {
      "id": "1863109",
      "postDate": "07/20/2022 06:47:19",
      "content": "<p>I recall (no pun intended) seeing this paper during my research in the past few days, but I don't know how to re-implement it from PyTorch to Keras/TensorFlow.</p>\n<p>It is uncanny that we are using the same general idea with arrows in plots. But the point still is even when we pick the best iteration by using the appropriate target measure for early stopping, there is no guarantee that our predictor gave us the optimal value simply because it isn't minimizing/maximizing directly for our target measure.</p>",
      "rawMarkdown": "I recall (no pun intended) seeing this paper during my research in the past few days, but I don't know how to re-implement it from PyTorch to Keras/TensorFlow.\n\nIt is uncanny that we are using the same general idea with arrows in plots. But the point still is even when we pick the best iteration by using the appropriate target measure for early stopping, there is no guarantee that our predictor gave us the optimal value simply because it isn't minimizing/maximizing directly for our target measure.",
      "votes": null
    },
    {
      "id": "1863186",
      "postDate": "07/20/2022 07:49:54",
      "content": "<p>I would expect the improvements to the logloss to be cointegrated to the improvements to the Amex metric. So while the improvements are not exactly correlated, in time they head in the same direction. </p>",
      "rawMarkdown": "I would expect the improvements to the logloss to be cointegrated to the improvements to the Amex metric. So while the improvements are not exactly correlated, in time they head in the same direction.",
      "votes": null
    },
    {
      "id": "1863194",
      "postDate": "07/20/2022 07:55:55",
      "content": "<p><strong>TL;DR</strong>: the 4% \"D\" (<a href=\"https://www.kaggle.com/code/carlmcbrideellis/discrimination-threshold-false-positive-negative/notebook\" target=\"_blank\">discrimination threshold</a>) component:</p>\n<ul>\n<li>Great from a business perspective</li>\n<li>Awful from a kaggle LB perspective</li>\n</ul>\n<p>😄</p>",
      "rawMarkdown": "**TL;DR**: the 4% \"D\" ([discrimination threshold](https://www.kaggle.com/code/carlmcbrideellis/discrimination-threshold-false-positive-negative/notebook)) component:\n* Great from a business perspective\n* Awful from a kaggle LB perspective\n\n😄",
      "votes": null
    },
    {
      "id": "1863210",
      "postDate": "07/20/2022 08:07:35",
      "content": "<p>Yes, it was clear why the metric was chosen after I read an explanation in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\"><strong>this post</strong></a>.</p>",
      "rawMarkdown": "Yes, it was clear why the metric was chosen after I read an explanation in [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464).",
      "votes": null
    },
    {
      "id": "1863227",
      "postDate": "07/20/2022 08:17:26",
      "content": "<blockquote>\n  <p>So while the improvements are not exactly correlated, in time they head in the same direction.</p>\n</blockquote>\n<p>From what I can tell, this is not guaranteed at all. In fact, they head in the correct directions early on, but during the late training only luck or an extremely long run without an improvement will give any semblance of guarantee that we found the AmEx minimum.</p>\n<p>Here is an example from neural network training:</p>\n<pre><code>Epoch 29/100\namex: 0.792894 - amex_val: 0.786973\nEpoch 29: amex_val improved from 0.78072 to 0.78697, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 325s - loss: 0.2334 - val_loss: 0.2217 - amex: 0.7929 - amex_val: 0.7870 - 325s/epoch - 113ms/step\n\nEpoch 30/100\namex: 0.798671 - amex_val: 0.789247\nEpoch 30: amex_val improved from 0.78697 to 0.78925, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 324s - loss: 0.2257 - val_loss: 0.2265 - amex: 0.7987 - amex_val: 0.7892 - 324s/epoch - 113ms/step\n\nEpoch 31/100\namex: 0.809482 - amex_val: 0.788469\nEpoch 31: amex_val did not improve from 0.78925\n2869/2869 - 324s - loss: 0.2202 - val_loss: 0.2207 - amex: 0.8095 - amex_val: 0.7885 - 324s/epoch - 113ms/step\n</code></pre>\n<p>From iteration 29 to 30 the validation log-loss gets worse, but validation AmEx score gets better. From iteration 30 to 31 it is the opposite. And this is just a snippet of the log file. If you look at the plots in my original post, best iterations based on one measure or the other are nowhere near each other.</p>\n<p>In real world I would say who cares if we pick an iteration where AmEx score is 0.789247 or 0.788469, because the difference is minor. But we both know that doesn't apply to a competition where hundreds of people have identical score on a 3rd decimal place.</p>",
      "rawMarkdown": "> So while the improvements are not exactly correlated, in time they head in the same direction.\n\nFrom what I can tell, this is not guaranteed at all. In fact, they head in the correct directions early on, but during the late training only luck or an extremely long run without an improvement will give any semblance of guarantee that we found the AmEx minimum.\n\nHere is an example from neural network training:\n\n```\nEpoch 29/100\namex: 0.792894 - amex_val: 0.786973\nEpoch 29: amex_val improved from 0.78072 to 0.78697, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 325s - loss: 0.2334 - val_loss: 0.2217 - amex: 0.7929 - amex_val: 0.7870 - 325s/epoch - 113ms/step\n\nEpoch 30/100\namex: 0.798671 - amex_val: 0.789247\nEpoch 30: amex_val improved from 0.78697 to 0.78925, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 324s - loss: 0.2257 - val_loss: 0.2265 - amex: 0.7987 - amex_val: 0.7892 - 324s/epoch - 113ms/step\n\nEpoch 31/100\namex: 0.809482 - amex_val: 0.788469\nEpoch 31: amex_val did not improve from 0.78925\n2869/2869 - 324s - loss: 0.2202 - val_loss: 0.2207 - amex: 0.8095 - amex_val: 0.7885 - 324s/epoch - 113ms/step\n\n```\n\nFrom iteration 29 to 30 the validation log-loss gets worse, but validation AmEx score gets better. From iteration 30 to 31 it is the opposite. And this is just a snippet of the log file. If you look at the plots in my original post, best iterations based on one measure or the other are nowhere near each other.\n\nIn real world I would say who cares if we pick an iteration where AmEx score is 0.789247 or 0.788469, because the difference is minor. But we both know that doesn't apply to a competition where hundreds of people have identical score on a 3rd decimal place.",
      "votes": null
    },
    {
      "id": "1863240",
      "postDate": "07/20/2022 08:24:15",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>\n<p>…just one of many excellent and insightful posts by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> ! </p>\n<p>One can see in the above-mentioned topic that the ROC curve is almost vertical in the \"4%\" point,  hence the sensitivity (pun intended) in <em>D</em>. If AmEx had gone with just  the <em>G</em> component (basically the AUC) we could then perhaps have a couple more significant figures on the LB, less test data needed, and the medals (particularly the bronze zone) would be far less subject to the vagaries of the metric….</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @tilii7 \n\n...just one of many excellent and insightful posts by @ambrosm ! \n\nOne can see in the above-mentioned topic that the ROC curve is almost vertical in the \"4%\" point,  hence the sensitivity (pun intended) in *D*. If AmEx had gone with just  the *G* component (basically the AUC) we could then perhaps have a couple more significant figures on the LB, less test data needed, and the medals (particularly the bronze zone) would be far less subject to the vagaries of the metric....\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1863255",
      "postDate": "07/20/2022 08:41:30",
      "content": "<p>I believe if there any proxy for AUC, it definitely should be based on ranking / contrastive learning.</p>\n<p>In the end of the day, AUC is a measure of separability of 2 distributions and not a measure of the probability of being 1 or 0.<br>\nIn my tests (not only on this competition) log-loss is better than ranking losses, but If I would be asked to find a proxy for AUC, I would bet on advanced evolution of ranking loss.</p>",
      "rawMarkdown": "I believe if there any proxy for AUC, it definitely should be based on ranking / contrastive learning.\n\nIn the end of the day, AUC is a measure of separability of 2 distributions and not a measure of the probability of being 1 or 0.\nIn my tests (not only on this competition) log-loss is better than ranking losses, but If I would be asked to find a proxy for AUC, I would bet on advanced evolution of ranking loss.",
      "votes": null
    },
    {
      "id": "1863268",
      "postDate": "07/20/2022 08:58:58",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/pavelvod\" target=\"_blank\">@pavelvod</a> </p>\n<blockquote>\n  <p>\"<em>In the end of the day, AUC is a measure of separability of 2 distributions</em>\"</p>\n</blockquote>\n<p>Indeed. I have found didactic plots such as this to help in my understanding:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F889514d11bfd2ca469464b38149f29f9%2Fdiscriminability.jpg?generation=1658307440873561&amp;alt=media\" alt=\"\"></p>\n<p>where the <a href=\"https://en.wikipedia.org/wiki/Sensitivity_index\" target=\"_blank\">discriminability index</a> indicates the ability to discriminate between the signal distribution and the noise distribution. This goes back to the very origin of the  ROC (<em>receiver operating characteristic</em>) curve in comparing radar setups; the more separable the distributions, the greater the area under the curve; a single metric indicating how much better the radar system is in detecting a genuine signal. </p>\n<p>In the context of a classification problem, the greater the AUC, the better the classifier is in separating the two classes. Thus the AUC (area under the curve) is now used as a way comparing the performance of classifiers.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @pavelvod \n\n> \"*In the end of the day, AUC is a measure of separability of 2 distributions*\"\n\nIndeed. I have found didactic plots such as this to help in my understanding:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F889514d11bfd2ca469464b38149f29f9%2Fdiscriminability.jpg?generation=1658307440873561&alt=media)\n\nwhere the [discriminability index](https://en.wikipedia.org/wiki/Sensitivity_index) indicates the ability to discriminate between the signal distribution and the noise distribution. This goes back to the very origin of the  ROC (*receiver operating characteristic*) curve in comparing radar setups; the more separable the distributions, the greater the area under the curve; a single metric indicating how much better the radar system is in detecting a genuine signal. \n\nIn the context of a classification problem, the greater the AUC, the better the classifier is in separating the two classes. Thus the AUC (area under the curve) is now used as a way comparing the performance of classifiers.\n\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1863269",
      "postDate": "07/20/2022 08:59:02",
      "content": "<p>In that example, if you compare 29 to 31 both scores improve, hence what I mean about cointegration. But I guess my “long run” argument doesn’t apply when you hit the end of productive training. </p>",
      "rawMarkdown": "In that example, if you compare 29 to 31 both scores improve, hence what I mean about cointegration. But I guess my “long run” argument doesn’t apply when you hit the end of productive training.",
      "votes": null
    },
    {
      "id": "1863976",
      "postDate": "07/20/2022 16:33:15",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>\n<blockquote>\n  <p>\"<em>I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable)…</em>\"</p>\n</blockquote>\n<p>I have just happened across <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\"><strong>Roc-star</strong></a> \"<em>a differentiable function which is close as possible to AUC.</em>\"</p>\n<ul>\n<li>GitHub: <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\">Roc-star : An objective function for ROC-AUC that actually works.</a></li>\n<li>kaggle notebook: <a href=\"https://www.kaggle.com/code/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe/notebook\" target=\"_blank\">\"Roc-star : An AUC loss function to challenge BxE.\"</a></li>\n</ul>\n<p>where: </p>\n<blockquote>\n  <p>\"<em>It's uber-fast, with speed comparable to BCE (and just as vectorizable for a GPU/MPP). In my tests, it gives higher AUC scores than BCE, is less sensitive to Learning Rate (avoiding the need for a Scheduler in my tests), and eliminates entirely the need for Early Stopping</em>.\"</p>\n</blockquote>\n<p>It is based on the paper <a href=\"https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf\" target=\"_blank\">\"Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @tilii7 \n\n> \"*I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable)...*\"\n\nI have just happened across [**Roc-star**](https://github.com/iridiumblue/roc-star) \"*a differentiable function which is close as possible to AUC.*\"\n\n* GitHub: [Roc-star : An objective function for ROC-AUC that actually works.](https://github.com/iridiumblue/roc-star)\n* kaggle notebook: [\"Roc-star : An AUC loss function to challenge BxE.\"](https://www.kaggle.com/code/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe/notebook)\n\nwhere: \n> \"*It's uber-fast, with speed comparable to BCE (and just as vectorizable for a GPU/MPP). In my tests, it gives higher AUC scores than BCE, is less sensitive to Learning Rate (avoiding the need for a Scheduler in my tests), and eliminates entirely the need for Early Stopping*.\"\n\nIt is based on the paper [\"Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic\"](https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf).\n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1864120",
      "postDate": "07/20/2022 18:38:37",
      "content": "<p>could it be underfitting due to the test split being more than needed?</p>",
      "rawMarkdown": "could it be underfitting due to the test split being more than needed?",
      "votes": null
    },
    {
      "id": "1864164",
      "postDate": "07/20/2022 19:32:37",
      "content": "<p>I'd like to answer the question but I don't understand it. Test data has nothing to do with this discussion as we are discussing the training process. Do you mean validation split instead of test split? And if so, <code>more than needed</code> for what?</p>",
      "rawMarkdown": "I'd like to answer the question but I don't understand it. Test data has nothing to do with this discussion as we are discussing the training process. Do you mean validation split instead of test split? And if so, `more than needed` for what?",
      "votes": null
    },
    {
      "id": "1864166",
      "postDate": "07/20/2022 19:35:27",
      "content": "<p>Thank you for sharing this. It is pretty much a guarantee that I won't be able to re-implement this from PyTorch to Keras, but I will try to make it work so that PyTorch does the calculations for Keras.</p>",
      "rawMarkdown": "Thank you for sharing this. It is pretty much a guarantee that I won't be able to re-implement this from PyTorch to Keras, but I will try to make it work so that PyTorch does the calculations for Keras.",
      "votes": null
    },
    {
      "id": "1864177",
      "postDate": "07/20/2022 19:49:50",
      "content": "<p>Yes ! How much did you split between Validation and training ?</p>",
      "rawMarkdown": "Yes ! How much did you split between Validation and training ?",
      "votes": null
    },
    {
      "id": "1864193",
      "postDate": "07/20/2022 20:11:08",
      "content": "<p>It was a 5-fold training, so 80:20 split. I don't think anything is underfitted here, as this observation has to do with the discrepancy between the actual minimization function (log-loss) and the quantity we need to maximize (AmEx score).</p>",
      "rawMarkdown": "It was a 5-fold training, so 80:20 split. I don't think anything is underfitted here, as this observation has to do with the discrepancy between the actual minimization function (log-loss) and the quantity we need to maximize (AmEx score).",
      "votes": null
    },
    {
      "id": "1865896",
      "postDate": "07/22/2022 06:47:32",
      "content": "<p>Shifting topic only slightly, in my experiment best auc score is very close to…. best logloss. And far from best amex. </p>\n<p>But my hypothesis is that it's irreducible error aka noise. I would suspect (and it might be testable but I don't know how?) that during a long period in which auc marginally improves then marginally gets worse, that the amex score exhibits random walk and regression to mean behaviors. But just a guess. </p>\n<p>A single negative making the \"4%\" stands in for twenty negatives. Watching the metric is just watching a couple negatives go above threshold, oh below, oh above again.</p>\n<p>Looking for a way to better model the competition metric seems like a good area to explore, if successful it could be a big edge. But if out of good ideas id just focus on auc, which fortunately seems pretty highly correlated with logloss. </p>",
      "rawMarkdown": "Shifting topic only slightly, in my experiment best auc score is very close to.... best logloss. And far from best amex. \n\nBut my hypothesis is that it's irreducible error aka noise. I would suspect (and it might be testable but I don't know how?) that during a long period in which auc marginally improves then marginally gets worse, that the amex score exhibits random walk and regression to mean behaviors. But just a guess. \n\nA single negative making the \"4%\" stands in for twenty negatives. Watching the metric is just watching a couple negatives go above threshold, oh below, oh above again.\n\nLooking for a way to better model the competition metric seems like a good area to explore, if successful it could be a big edge. But if out of good ideas id just focus on auc, which fortunately seems pretty highly correlated with logloss.",
      "votes": null
    },
    {
      "id": "1865962",
      "postDate": "07/22/2022 07:26:06",
      "content": "<blockquote>\n  <p>Shifting topic only slightly, in my experiment best auc score is very close to…. best logloss.</p>\n</blockquote>\n<p>Sure, and so is AUC score to the AmEx score. Both of them are solid proxies for AmEx score, because reducing the prediction error of individual points (log-loss) and ensuring their correct ranking (AUC) are both actions that in general should lead to better AmEx scores. Yet when we get close to the optimum, neither one of them seems to track with AmEx as it should.</p>\n<p>Very good point about the effects of negative data points multiplying. They are like Gremlins getting wet.</p>",
      "rawMarkdown": "> Shifting topic only slightly, in my experiment best auc score is very close to…. best logloss.\n\nSure, and so is AUC score to the AmEx score. Both of them are solid proxies for AmEx score, because reducing the prediction error of individual points (log-loss) and ensuring their correct ranking (AUC) are both actions that in general should lead to better AmEx scores. Yet when we get close to the optimum, neither one of them seems to track with AmEx as it should.\n\nVery good point about the effects of negative data points multiplying. They are like Gremlins getting wet.",
      "votes": null
    },
    {
      "id": "1866968",
      "postDate": "07/22/2022 22:33:46",
      "content": "<p>I think all metrics have the same goal, which is to reduce error, maybe some metrics are better than others by relying on proprietary models, but if anyone has a score such as (AUC)  above 95%, he or she will get this Competition for first place.</p>",
      "rawMarkdown": "I think all metrics have the same goal, which is to reduce error, maybe some metrics are better than others by relying on proprietary models, but if anyone has a score such as (AUC)  above 95%, he or she will get this Competition for first place.",
      "votes": null
    },
    {
      "id": "1866993",
      "postDate": "07/22/2022 23:47:53",
      "content": "<blockquote>\n  <p>I think all metrics have the same goal, which is to reduce error</p>\n</blockquote>\n<p>I think you mean that all loss functions try to reduce the error. Metrics only measure the error using various criteria.</p>\n<blockquote>\n  <p>if anyone has a score such as (AUC) above 95%, he or she will get this Competition for first place.</p>\n</blockquote>\n<p>Most of my models have out-of-fold AUC &gt; 0.96, and I am nowhere near the first place.</p>\n<p>Speaking of metrics, not all of them convey the same information, which is why it is best to use multiple metrics in evaluations. Let's say we make a prediction where each target 0 is predicted with probability 0.49 and each target 1 is predicted with probability 0.51. That model would have <code>AUC=1</code> and <code>accuracy=1</code>, but it would have a terrible log-loss score. And rightfully so, because that is a terrible model with all predictions barely on the correct side of the margin. We would have no confidence in those predictions.</p>\n<p>This code will simulate the dataset I just explained:</p>\n<pre><code>import numpy as np\nfrom sklearn.metrics import log_loss, roc_auc_score, accuracy_score\n\ntarget = np.array([1, 1, 1, 0, 0, 0])\nprediction = np.array([0.51, 0.51, 0.51, 0.49, 0.49, 0.49])\n\nprint('Log-loss  %.2f' % log_loss(target, prediction))\nprint('AUC score %.2f' % roc_auc_score(target, prediction))\nprint('Accuracy  %.2f' % accuracy_score(target, prediction &gt; 0.5))\n</code></pre>\n<p>It will print:</p>\n<pre><code>Log-loss  0.67\nAUC score 1.00\nAccuracy  1.00\n</code></pre>",
      "rawMarkdown": "> I think all metrics have the same goal, which is to reduce error\n\nI think you mean that all loss functions try to reduce the error. Metrics only measure the error using various criteria.\n\n> if anyone has a score such as (AUC) above 95%, he or she will get this Competition for first place.\n\nMost of my models have out-of-fold AUC > 0.96, and I am nowhere near the first place.\n\nSpeaking of metrics, not all of them convey the same information, which is why it is best to use multiple metrics in evaluations. Let's say we make a prediction where each target 0 is predicted with probability 0.49 and each target 1 is predicted with probability 0.51. That model would have `AUC=1` and `accuracy=1`, but it would have a terrible log-loss score. And rightfully so, because that is a terrible model with all predictions barely on the correct side of the margin. We would have no confidence in those predictions.\n\nThis code will simulate the dataset I just explained:\n\n```\nimport numpy as np\nfrom sklearn.metrics import log_loss, roc_auc_score, accuracy_score\n\ntarget = np.array([1, 1, 1, 0, 0, 0])\nprediction = np.array([0.51, 0.51, 0.51, 0.49, 0.49, 0.49])\n\nprint('Log-loss  %.2f' % log_loss(target, prediction))\nprint('AUC score %.2f' % roc_auc_score(target, prediction))\nprint('Accuracy  %.2f' % accuracy_score(target, prediction > 0.5))\n\n```\nIt will print:\n\n```\nLog-loss  0.67\nAUC score 1.00\nAccuracy  1.00\n```",
      "votes": null
    },
    {
      "id": "1867325",
      "postDate": "07/23/2022 06:17:23",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a> </p>\n<p>In addition to the excellent answer by <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> I would like to take the liberty of posting this image taken from the notebook <a href=\"https://www.kaggle.com/code/carlmcbrideellis/classification-how-imbalanced-is-imbalanced\" target=\"_blank\">Classification: How imbalanced is \"imbalanced\"?</a>, where various classification <em>metrics</em> are applied to the same data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2Fa2e3e225f3c8ddaaa7f7578f15c61d4f%2Fclassification_metrics.png?generation=1658556303963637&amp;alt=media\" alt=\"\"></p>\n<p>The goal of a metric is to convey information to the end user, and should not be confused with the loss function (or the splitting criteria) used by the estimator to arrive at the model. Although ostensibly similar, they  play very different roles.  Indeed if I am not mistaken, the essence of this topic is the dissonance between the loss function (used to find the solution with the least error) and the AmEx competition metric, which is used to actually score our leaderboard results.  </p>\n<p>All the best,<br>\ncarl</p>",
      "rawMarkdown": "Dear @youneseloiarm \n\nIn addition to the excellent answer by @tilii7 I would like to take the liberty of posting this image taken from the notebook [Classification: How imbalanced is \"imbalanced\"?](https://www.kaggle.com/code/carlmcbrideellis/classification-how-imbalanced-is-imbalanced), where various classification *metrics* are applied to the same data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2Fa2e3e225f3c8ddaaa7f7578f15c61d4f%2Fclassification_metrics.png?generation=1658556303963637&alt=media)\n\nThe goal of a metric is to convey information to the end user, and should not be confused with the loss function (or the splitting criteria) used by the estimator to arrive at the model. Although ostensibly similar, they  play very different roles.  Indeed if I am not mistaken, the essence of this topic is the dissonance between the loss function (used to find the solution with the least error) and the AmEx competition metric, which is used to actually score our leaderboard results.  \n\nAll the best,\ncarl",
      "votes": null
    },
    {
      "id": "1867912",
      "postDate": "07/23/2022 15:39:49",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, Maybe your models have been overfitted, you would do a simulation between all metrics that you have such as Amex's metric, So with that simulation, you probably estimate the approximation between those metrics.</p>",
      "rawMarkdown": "Dear @tilii7, Maybe your models have been overfitted, you would do a simulation between all metrics that you have such as Amex's metric, So with that simulation, you probably estimate the approximation between those metrics.",
      "votes": null
    },
    {
      "id": "1869065",
      "postDate": "07/24/2022 13:04:45",
      "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> thanks for sharing, your posts are very valuable! </p>\n<p>So far I have noticed the same regarding val_loss vs val_AUC, in some runs I monitored also TPR/FPR (that is also interesting to see how changes between epochs) </p>\n<p>On a side note, may I ask if it is possible to share your TF-Keras AmEx metric implementation? </p>",
      "rawMarkdown": "tilii7 thanks for sharing, your posts are very valuable! \n\nSo far I have noticed the same regarding val_loss vs val_AUC, in some runs I monitored also TPR/FPR (that is also interesting to see how changes between epochs) \n\nOn a side note, may I ask if it is possible to share your TF-Keras AmEx metric implementation?",
      "votes": null
    },
    {
      "id": "1869332",
      "postDate": "07/24/2022 17:03:52",
      "content": "<p>Using AmEx metrics from <a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\"><strong>this notebook</strong></a>.</p>",
      "rawMarkdown": "Using AmEx metrics from [**this notebook**](https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations).",
      "votes": null
    },
    {
      "id": "1870601",
      "postDate": "07/25/2022 17:06:34",
      "content": "<p>Thanks for sharing, very insightful</p>",
      "rawMarkdown": "Thanks for sharing, very insightful",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1863030,
      "author_name": "ambrosm",
      "author_url": "",
      "post_date": "07/20/2022 05:52:39",
      "content": "<p>Two months ago, we had a similar discussion during a TPS competition: <a href=\"https://www.kaggle.com/competitions/tabular-playground-series-may-2022/discussion/326116\" target=\"_blank\">Early stopping on loss, auc or accuracy?</a></p>\n<p>And you might want to look at the paper <a href=\"https://arxiv.org/abs/2108.11179\" target=\"_blank\">Recall@k Surrogate Loss with Large Batches and Similarity Mixup</a>, where they show how to modify a recall metric (the d part of our score) to make it differentiable. I didn't apply the ideas of the paper to this competition, though.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1863109,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 06:47:19",
          "content": "<p>I recall (no pun intended) seeing this paper during my research in the past few days, but I don't know how to re-implement it from PyTorch to Keras/TensorFlow.</p>\n<p>It is uncanny that we are using the same general idea with arrows in plots. But the point still is even when we pick the best iteration by using the appropriate target measure for early stopping, there is no guarantee that our predictor gave us the optimal value simply because it isn't minimizing/maximizing directly for our target measure.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1863186,
      "author_name": "burritodan",
      "author_url": "",
      "post_date": "07/20/2022 07:49:54",
      "content": "<p>I would expect the improvements to the logloss to be cointegrated to the improvements to the Amex metric. So while the improvements are not exactly correlated, in time they head in the same direction. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1863227,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 08:17:26",
          "content": "<blockquote>\n  <p>So while the improvements are not exactly correlated, in time they head in the same direction.</p>\n</blockquote>\n<p>From what I can tell, this is not guaranteed at all. In fact, they head in the correct directions early on, but during the late training only luck or an extremely long run without an improvement will give any semblance of guarantee that we found the AmEx minimum.</p>\n<p>Here is an example from neural network training:</p>\n<pre><code>Epoch 29/100\namex: 0.792894 - amex_val: 0.786973\nEpoch 29: amex_val improved from 0.78072 to 0.78697, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 325s - loss: 0.2334 - val_loss: 0.2217 - amex: 0.7929 - amex_val: 0.7870 - 325s/epoch - 113ms/step\n\nEpoch 30/100\namex: 0.798671 - amex_val: 0.789247\nEpoch 30: amex_val improved from 0.78697 to 0.78925, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 324s - loss: 0.2257 - val_loss: 0.2265 - amex: 0.7987 - amex_val: 0.7892 - 324s/epoch - 113ms/step\n\nEpoch 31/100\namex: 0.809482 - amex_val: 0.788469\nEpoch 31: amex_val did not improve from 0.78925\n2869/2869 - 324s - loss: 0.2202 - val_loss: 0.2207 - amex: 0.8095 - amex_val: 0.7885 - 324s/epoch - 113ms/step\n</code></pre>\n<p>From iteration 29 to 30 the validation log-loss gets worse, but validation AmEx score gets better. From iteration 30 to 31 it is the opposite. And this is just a snippet of the log file. If you look at the plots in my original post, best iterations based on one measure or the other are nowhere near each other.</p>\n<p>In real world I would say who cares if we pick an iteration where AmEx score is 0.789247 or 0.788469, because the difference is minor. But we both know that doesn't apply to a competition where hundreds of people have identical score on a 3rd decimal place.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1863269,
          "author_name": "burritodan",
          "author_url": "",
          "post_date": "07/20/2022 08:59:02",
          "content": "<p>In that example, if you compare 29 to 31 both scores improve, hence what I mean about cointegration. But I guess my “long run” argument doesn’t apply when you hit the end of productive training. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1863194,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "07/20/2022 07:55:55",
      "content": "<p><strong>TL;DR</strong>: the 4% \"D\" (<a href=\"https://www.kaggle.com/code/carlmcbrideellis/discrimination-threshold-false-positive-negative/notebook\" target=\"_blank\">discrimination threshold</a>) component:</p>\n<ul>\n<li>Great from a business perspective</li>\n<li>Awful from a kaggle LB perspective</li>\n</ul>\n<p>😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 1863210,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 08:07:35",
          "content": "<p>Yes, it was clear why the metric was chosen after I read an explanation in <a href=\"https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464\" target=\"_blank\"><strong>this post</strong></a>.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1863240,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/20/2022 08:24:15",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>\n<p>…just one of many excellent and insightful posts by <a href=\"https://www.kaggle.com/ambrosm\" target=\"_blank\">@ambrosm</a> ! </p>\n<p>One can see in the above-mentioned topic that the ROC curve is almost vertical in the \"4%\" point,  hence the sensitivity (pun intended) in <em>D</em>. If AmEx had gone with just  the <em>G</em> component (basically the AUC) we could then perhaps have a couple more significant figures on the LB, less test data needed, and the medals (particularly the bronze zone) would be far less subject to the vagaries of the metric….</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1863255,
      "author_name": "pavelvod",
      "author_url": "",
      "post_date": "07/20/2022 08:41:30",
      "content": "<p>I believe if there any proxy for AUC, it definitely should be based on ranking / contrastive learning.</p>\n<p>In the end of the day, AUC is a measure of separability of 2 distributions and not a measure of the probability of being 1 or 0.<br>\nIn my tests (not only on this competition) log-loss is better than ranking losses, but If I would be asked to find a proxy for AUC, I would bet on advanced evolution of ranking loss.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1863268,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/20/2022 08:58:58",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/pavelvod\" target=\"_blank\">@pavelvod</a> </p>\n<blockquote>\n  <p>\"<em>In the end of the day, AUC is a measure of separability of 2 distributions</em>\"</p>\n</blockquote>\n<p>Indeed. I have found didactic plots such as this to help in my understanding:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F889514d11bfd2ca469464b38149f29f9%2Fdiscriminability.jpg?generation=1658307440873561&amp;alt=media\" alt=\"\"></p>\n<p>where the <a href=\"https://en.wikipedia.org/wiki/Sensitivity_index\" target=\"_blank\">discriminability index</a> indicates the ability to discriminate between the signal distribution and the noise distribution. This goes back to the very origin of the  ROC (<em>receiver operating characteristic</em>) curve in comparing radar setups; the more separable the distributions, the greater the area under the curve; a single metric indicating how much better the radar system is in detecting a genuine signal. </p>\n<p>In the context of a classification problem, the greater the AUC, the better the classifier is in separating the two classes. Thus the AUC (area under the curve) is now used as a way comparing the performance of classifiers.</p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1863976,
      "author_name": "carlmcbrideellis",
      "author_url": "",
      "post_date": "07/20/2022 16:33:15",
      "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> </p>\n<blockquote>\n  <p>\"<em>I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable)…</em>\"</p>\n</blockquote>\n<p>I have just happened across <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\"><strong>Roc-star</strong></a> \"<em>a differentiable function which is close as possible to AUC.</em>\"</p>\n<ul>\n<li>GitHub: <a href=\"https://github.com/iridiumblue/roc-star\" target=\"_blank\">Roc-star : An objective function for ROC-AUC that actually works.</a></li>\n<li>kaggle notebook: <a href=\"https://www.kaggle.com/code/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe/notebook\" target=\"_blank\">\"Roc-star : An AUC loss function to challenge BxE.\"</a></li>\n</ul>\n<p>where: </p>\n<blockquote>\n  <p>\"<em>It's uber-fast, with speed comparable to BCE (and just as vectorizable for a GPU/MPP). In my tests, it gives higher AUC scores than BCE, is less sensitive to Learning Rate (avoiding the need for a Scheduler in my tests), and eliminates entirely the need for Early Stopping</em>.\"</p>\n</blockquote>\n<p>It is based on the paper <a href=\"https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf\" target=\"_blank\">\"Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic\"</a>.</p>\n<p>All the best,<br>\ncarl</p>",
      "votes": null,
      "replies": [
        {
          "id": 1864166,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 19:35:27",
          "content": "<p>Thank you for sharing this. It is pretty much a guarantee that I won't be able to re-implement this from PyTorch to Keras, but I will try to make it work so that PyTorch does the calculations for Keras.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1864120,
      "author_name": "muhammadammarjamshed",
      "author_url": "",
      "post_date": "07/20/2022 18:38:37",
      "content": "<p>could it be underfitting due to the test split being more than needed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1864164,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 19:32:37",
          "content": "<p>I'd like to answer the question but I don't understand it. Test data has nothing to do with this discussion as we are discussing the training process. Do you mean validation split instead of test split? And if so, <code>more than needed</code> for what?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1864177,
          "author_name": "muhammadammarjamshed",
          "author_url": "",
          "post_date": "07/20/2022 19:49:50",
          "content": "<p>Yes ! How much did you split between Validation and training ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1864193,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/20/2022 20:11:08",
          "content": "<p>It was a 5-fold training, so 80:20 split. I don't think anything is underfitted here, as this observation has to do with the discrepancy between the actual minimization function (log-loss) and the quantity we need to maximize (AmEx score).</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1865896,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "07/22/2022 06:47:32",
      "content": "<p>Shifting topic only slightly, in my experiment best auc score is very close to…. best logloss. And far from best amex. </p>\n<p>But my hypothesis is that it's irreducible error aka noise. I would suspect (and it might be testable but I don't know how?) that during a long period in which auc marginally improves then marginally gets worse, that the amex score exhibits random walk and regression to mean behaviors. But just a guess. </p>\n<p>A single negative making the \"4%\" stands in for twenty negatives. Watching the metric is just watching a couple negatives go above threshold, oh below, oh above again.</p>\n<p>Looking for a way to better model the competition metric seems like a good area to explore, if successful it could be a big edge. But if out of good ideas id just focus on auc, which fortunately seems pretty highly correlated with logloss. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1865962,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/22/2022 07:26:06",
          "content": "<blockquote>\n  <p>Shifting topic only slightly, in my experiment best auc score is very close to…. best logloss.</p>\n</blockquote>\n<p>Sure, and so is AUC score to the AmEx score. Both of them are solid proxies for AmEx score, because reducing the prediction error of individual points (log-loss) and ensuring their correct ranking (AUC) are both actions that in general should lead to better AmEx scores. Yet when we get close to the optimum, neither one of them seems to track with AmEx as it should.</p>\n<p>Very good point about the effects of negative data points multiplying. They are like Gremlins getting wet.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1866968,
      "author_name": "youneseloiarm",
      "author_url": "",
      "post_date": "07/22/2022 22:33:46",
      "content": "<p>I think all metrics have the same goal, which is to reduce error, maybe some metrics are better than others by relying on proprietary models, but if anyone has a score such as (AUC)  above 95%, he or she will get this Competition for first place.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1866993,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/22/2022 23:47:53",
          "content": "<blockquote>\n  <p>I think all metrics have the same goal, which is to reduce error</p>\n</blockquote>\n<p>I think you mean that all loss functions try to reduce the error. Metrics only measure the error using various criteria.</p>\n<blockquote>\n  <p>if anyone has a score such as (AUC) above 95%, he or she will get this Competition for first place.</p>\n</blockquote>\n<p>Most of my models have out-of-fold AUC &gt; 0.96, and I am nowhere near the first place.</p>\n<p>Speaking of metrics, not all of them convey the same information, which is why it is best to use multiple metrics in evaluations. Let's say we make a prediction where each target 0 is predicted with probability 0.49 and each target 1 is predicted with probability 0.51. That model would have <code>AUC=1</code> and <code>accuracy=1</code>, but it would have a terrible log-loss score. And rightfully so, because that is a terrible model with all predictions barely on the correct side of the margin. We would have no confidence in those predictions.</p>\n<p>This code will simulate the dataset I just explained:</p>\n<pre><code>import numpy as np\nfrom sklearn.metrics import log_loss, roc_auc_score, accuracy_score\n\ntarget = np.array([1, 1, 1, 0, 0, 0])\nprediction = np.array([0.51, 0.51, 0.51, 0.49, 0.49, 0.49])\n\nprint('Log-loss  %.2f' % log_loss(target, prediction))\nprint('AUC score %.2f' % roc_auc_score(target, prediction))\nprint('Accuracy  %.2f' % accuracy_score(target, prediction &gt; 0.5))\n</code></pre>\n<p>It will print:</p>\n<pre><code>Log-loss  0.67\nAUC score 1.00\nAccuracy  1.00\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1867325,
          "author_name": "carlmcbrideellis",
          "author_url": "",
          "post_date": "07/23/2022 06:17:23",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/youneseloiarm\" target=\"_blank\">@youneseloiarm</a> </p>\n<p>In addition to the excellent answer by <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> I would like to take the liberty of posting this image taken from the notebook <a href=\"https://www.kaggle.com/code/carlmcbrideellis/classification-how-imbalanced-is-imbalanced\" target=\"_blank\">Classification: How imbalanced is \"imbalanced\"?</a>, where various classification <em>metrics</em> are applied to the same data:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2Fa2e3e225f3c8ddaaa7f7578f15c61d4f%2Fclassification_metrics.png?generation=1658556303963637&amp;alt=media\" alt=\"\"></p>\n<p>The goal of a metric is to convey information to the end user, and should not be confused with the loss function (or the splitting criteria) used by the estimator to arrive at the model. Although ostensibly similar, they  play very different roles.  Indeed if I am not mistaken, the essence of this topic is the dissonance between the loss function (used to find the solution with the least error) and the AmEx competition metric, which is used to actually score our leaderboard results.  </p>\n<p>All the best,<br>\ncarl</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1867912,
          "author_name": "youneseloiarm",
          "author_url": "",
          "post_date": "07/23/2022 15:39:49",
          "content": "<p>Dear <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, Maybe your models have been overfitted, you would do a simulation between all metrics that you have such as Amex's metric, So with that simulation, you probably estimate the approximation between those metrics.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1869065,
      "author_name": "imeintanis",
      "author_url": "",
      "post_date": "07/24/2022 13:04:45",
      "content": "<p><a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a> thanks for sharing, your posts are very valuable! </p>\n<p>So far I have noticed the same regarding val_loss vs val_AUC, in some runs I monitored also TPR/FPR (that is also interesting to see how changes between epochs) </p>\n<p>On a side note, may I ask if it is possible to share your TF-Keras AmEx metric implementation? </p>",
      "votes": null,
      "replies": [
        {
          "id": 1869332,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "07/24/2022 17:03:52",
          "content": "<p>Using AmEx metrics from <a href=\"https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations\" target=\"_blank\"><strong>this notebook</strong></a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1870601,
      "author_name": "kanakpandit",
      "author_url": "",
      "post_date": "07/25/2022 17:06:34",
      "content": "<p>Thanks for sharing, very insightful</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1862658": "Never liked scoring schemes I couldn't understand, so [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464) helped me a lot. At least I can justify to some degree why they chose this as the competition score.\n\nThe problem is that AmEx score is only loosely related to other quantities we can maximize/minimize during model fitting and early stopping. Here is a long LightGBM run (with DART booster; first fold shown) with each point representing 100 trees.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2F647396ef061dce737f7814323ed0ad0a%2Famex-lgbm-dart.png?generation=1658266708075277&alt=media)\n\nThe curves are picture-perfect in general, but the iteration with best log-loss (pointed by arrow) is nowhere near the iteration with the best AmEx score.\n\nHere is a neural network run (Keras with LSTM; first fold shown) with each point representing a single epoch.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F549055%2Fadd3d9a24df5d8a7071210a6046cb4ea%2Fkeras-5fold-run-02-v5-epochs.png?generation=1658266881324817&alt=media)\n\nNow, we can do early stopping based on AmEx scores rather than log-loss, but the fact remains that under the hood our models are minimizing a function (log-loss) that is a decent but not great proxy for AmEx scores.\n\nWithout even knowing whether AmEx score is differentiable (presumably not), it is out of question to maximize it directly as a function because it is too expensive to compute. I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable), but even AUC scores do not track perfectly with AmEx scores because AUC/Gini is only part of the calculation.\n\nAny thoughts on this? I suspect that finding this magical function that can be directly maximized/minimized and tracks better with AmEx scores would probably clinch the top spot, so I am not holding my breath that we will hear about it before the competition ends.",
    "1863030": "Two months ago, we had a similar discussion during a TPS competition: [Early stopping on loss, auc or accuracy?](https://www.kaggle.com/competitions/tabular-playground-series-may-2022/discussion/326116)\n\nAnd you might want to look at the paper [Recall@k Surrogate Loss with Large Batches and Similarity Mixup](https://arxiv.org/abs/2108.11179), where they show how to modify a recall metric (the d part of our score) to make it differentiable. I didn't apply the ideas of the paper to this competition, though.",
    "1863109": "I recall (no pun intended) seeing this paper during my research in the past few days, but I don't know how to re-implement it from PyTorch to Keras/TensorFlow.\n\nIt is uncanny that we are using the same general idea with arrows in plots. But the point still is even when we pick the best iteration by using the appropriate target measure for early stopping, there is no guarantee that our predictor gave us the optimal value simply because it isn't minimizing/maximizing directly for our target measure.",
    "1863186": "I would expect the improvements to the logloss to be cointegrated to the improvements to the Amex metric. So while the improvements are not exactly correlated, in time they head in the same direction.",
    "1863194": "**TL;DR**: the 4% \"D\" ([discrimination threshold](https://www.kaggle.com/code/carlmcbrideellis/discrimination-threshold-false-positive-negative/notebook)) component:\n* Great from a business perspective\n* Awful from a kaggle LB perspective\n\n😄",
    "1863210": "Yes, it was clear why the metric was chosen after I read an explanation in [**this post**](https://www.kaggle.com/competitions/amex-default-prediction/discussion/327464).",
    "1863227": "> So while the improvements are not exactly correlated, in time they head in the same direction.\n\nFrom what I can tell, this is not guaranteed at all. In fact, they head in the correct directions early on, but during the late training only luck or an extremely long run without an improvement will give any semblance of guarantee that we found the AmEx minimum.\n\nHere is an example from neural network training:\n\n```\nEpoch 29/100\namex: 0.792894 - amex_val: 0.786973\nEpoch 29: amex_val improved from 0.78072 to 0.78697, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 325s - loss: 0.2334 - val_loss: 0.2217 - amex: 0.7929 - amex_val: 0.7870 - 325s/epoch - 113ms/step\n\nEpoch 30/100\namex: 0.798671 - amex_val: 0.789247\nEpoch 30: amex_val improved from 0.78697 to 0.78925, saving model to keras-run-04-v1-fold-01-bag-01.h5\n2869/2869 - 324s - loss: 0.2257 - val_loss: 0.2265 - amex: 0.7987 - amex_val: 0.7892 - 324s/epoch - 113ms/step\n\nEpoch 31/100\namex: 0.809482 - amex_val: 0.788469\nEpoch 31: amex_val did not improve from 0.78925\n2869/2869 - 324s - loss: 0.2202 - val_loss: 0.2207 - amex: 0.8095 - amex_val: 0.7885 - 324s/epoch - 113ms/step\n\n```\n\nFrom iteration 29 to 30 the validation log-loss gets worse, but validation AmEx score gets better. From iteration 30 to 31 it is the opposite. And this is just a snippet of the log file. If you look at the plots in my original post, best iterations based on one measure or the other are nowhere near each other.\n\nIn real world I would say who cares if we pick an iteration where AmEx score is 0.789247 or 0.788469, because the difference is minor. But we both know that doesn't apply to a competition where hundreds of people have identical score on a 3rd decimal place.",
    "1863240": "Dear @tilii7 \n\n...just one of many excellent and insightful posts by @ambrosm ! \n\nOne can see in the above-mentioned topic that the ROC curve is almost vertical in the \"4%\" point,  hence the sensitivity (pun intended) in *D*. If AmEx had gone with just  the *G* component (basically the AUC) we could then perhaps have a couple more significant figures on the LB, less test data needed, and the medals (particularly the bronze zone) would be far less subject to the vagaries of the metric....\n\nAll the best,\ncarl",
    "1863255": "I believe if there any proxy for AUC, it definitely should be based on ranking / contrastive learning.\n\nIn the end of the day, AUC is a measure of separability of 2 distributions and not a measure of the probability of being 1 or 0.\nIn my tests (not only on this competition) log-loss is better than ranking losses, but If I would be asked to find a proxy for AUC, I would bet on advanced evolution of ranking loss.",
    "1863268": "Dear @pavelvod \n\n> \"*In the end of the day, AUC is a measure of separability of 2 distributions*\"\n\nIndeed. I have found didactic plots such as this to help in my understanding:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2F889514d11bfd2ca469464b38149f29f9%2Fdiscriminability.jpg?generation=1658307440873561&alt=media)\n\nwhere the [discriminability index](https://en.wikipedia.org/wiki/Sensitivity_index) indicates the ability to discriminate between the signal distribution and the noise distribution. This goes back to the very origin of the  ROC (*receiver operating characteristic*) curve in comparing radar setups; the more separable the distributions, the greater the area under the curve; a single metric indicating how much better the radar system is in detecting a genuine signal. \n\nIn the context of a classification problem, the greater the AUC, the better the classifier is in separating the two classes. Thus the AUC (area under the curve) is now used as a way comparing the performance of classifiers.\n\n\nAll the best,\ncarl",
    "1863269": "In that example, if you compare 29 to 31 both scores improve, hence what I mean about cointegration. But I guess my “long run” argument doesn’t apply when you hit the end of productive training.",
    "1863976": "Dear @tilii7 \n\n> \"*I spent some time looking for a good AUC score proxy to maximize (AUC is not differentiable)...*\"\n\nI have just happened across [**Roc-star**](https://github.com/iridiumblue/roc-star) \"*a differentiable function which is close as possible to AUC.*\"\n\n* GitHub: [Roc-star : An objective function for ROC-AUC that actually works.](https://github.com/iridiumblue/roc-star)\n* kaggle notebook: [\"Roc-star : An AUC loss function to challenge BxE.\"](https://www.kaggle.com/code/iridiumblue/roc-star-an-auc-loss-function-to-challenge-bxe/notebook)\n\nwhere: \n> \"*It's uber-fast, with speed comparable to BCE (and just as vectorizable for a GPU/MPP). In my tests, it gives higher AUC scores than BCE, is less sensitive to Learning Rate (avoiding the need for a Scheduler in my tests), and eliminates entirely the need for Early Stopping*.\"\n\nIt is based on the paper [\"Optimizing Classifier Performance via an Approximation to the Wilcoxon-Mann-Whitney Statistic\"](https://www.aaai.org/Papers/ICML/2003/ICML03-110.pdf).\n\nAll the best,\ncarl",
    "1864120": "could it be underfitting due to the test split being more than needed?",
    "1864164": "I'd like to answer the question but I don't understand it. Test data has nothing to do with this discussion as we are discussing the training process. Do you mean validation split instead of test split? And if so, `more than needed` for what?",
    "1864166": "Thank you for sharing this. It is pretty much a guarantee that I won't be able to re-implement this from PyTorch to Keras, but I will try to make it work so that PyTorch does the calculations for Keras.",
    "1864177": "Yes ! How much did you split between Validation and training ?",
    "1864193": "It was a 5-fold training, so 80:20 split. I don't think anything is underfitted here, as this observation has to do with the discrepancy between the actual minimization function (log-loss) and the quantity we need to maximize (AmEx score).",
    "1865896": "Shifting topic only slightly, in my experiment best auc score is very close to.... best logloss. And far from best amex. \n\nBut my hypothesis is that it's irreducible error aka noise. I would suspect (and it might be testable but I don't know how?) that during a long period in which auc marginally improves then marginally gets worse, that the amex score exhibits random walk and regression to mean behaviors. But just a guess. \n\nA single negative making the \"4%\" stands in for twenty negatives. Watching the metric is just watching a couple negatives go above threshold, oh below, oh above again.\n\nLooking for a way to better model the competition metric seems like a good area to explore, if successful it could be a big edge. But if out of good ideas id just focus on auc, which fortunately seems pretty highly correlated with logloss.",
    "1865962": "> Shifting topic only slightly, in my experiment best auc score is very close to…. best logloss.\n\nSure, and so is AUC score to the AmEx score. Both of them are solid proxies for AmEx score, because reducing the prediction error of individual points (log-loss) and ensuring their correct ranking (AUC) are both actions that in general should lead to better AmEx scores. Yet when we get close to the optimum, neither one of them seems to track with AmEx as it should.\n\nVery good point about the effects of negative data points multiplying. They are like Gremlins getting wet.",
    "1866968": "I think all metrics have the same goal, which is to reduce error, maybe some metrics are better than others by relying on proprietary models, but if anyone has a score such as (AUC)  above 95%, he or she will get this Competition for first place.",
    "1866993": "> I think all metrics have the same goal, which is to reduce error\n\nI think you mean that all loss functions try to reduce the error. Metrics only measure the error using various criteria.\n\n> if anyone has a score such as (AUC) above 95%, he or she will get this Competition for first place.\n\nMost of my models have out-of-fold AUC > 0.96, and I am nowhere near the first place.\n\nSpeaking of metrics, not all of them convey the same information, which is why it is best to use multiple metrics in evaluations. Let's say we make a prediction where each target 0 is predicted with probability 0.49 and each target 1 is predicted with probability 0.51. That model would have `AUC=1` and `accuracy=1`, but it would have a terrible log-loss score. And rightfully so, because that is a terrible model with all predictions barely on the correct side of the margin. We would have no confidence in those predictions.\n\nThis code will simulate the dataset I just explained:\n\n```\nimport numpy as np\nfrom sklearn.metrics import log_loss, roc_auc_score, accuracy_score\n\ntarget = np.array([1, 1, 1, 0, 0, 0])\nprediction = np.array([0.51, 0.51, 0.51, 0.49, 0.49, 0.49])\n\nprint('Log-loss  %.2f' % log_loss(target, prediction))\nprint('AUC score %.2f' % roc_auc_score(target, prediction))\nprint('Accuracy  %.2f' % accuracy_score(target, prediction > 0.5))\n\n```\nIt will print:\n\n```\nLog-loss  0.67\nAUC score 1.00\nAccuracy  1.00\n```",
    "1867325": "Dear @youneseloiarm \n\nIn addition to the excellent answer by @tilii7 I would like to take the liberty of posting this image taken from the notebook [Classification: How imbalanced is \"imbalanced\"?](https://www.kaggle.com/code/carlmcbrideellis/classification-how-imbalanced-is-imbalanced), where various classification *metrics* are applied to the same data:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4051350%2Fa2e3e225f3c8ddaaa7f7578f15c61d4f%2Fclassification_metrics.png?generation=1658556303963637&alt=media)\n\nThe goal of a metric is to convey information to the end user, and should not be confused with the loss function (or the splitting criteria) used by the estimator to arrive at the model. Although ostensibly similar, they  play very different roles.  Indeed if I am not mistaken, the essence of this topic is the dissonance between the loss function (used to find the solution with the least error) and the AmEx competition metric, which is used to actually score our leaderboard results.  \n\nAll the best,\ncarl",
    "1867912": "Dear @tilii7, Maybe your models have been overfitted, you would do a simulation between all metrics that you have such as Amex's metric, So with that simulation, you probably estimate the approximation between those metrics.",
    "1869065": "tilii7 thanks for sharing, your posts are very valuable! \n\nSo far I have noticed the same regarding val_loss vs val_AUC, in some runs I monitored also TPR/FPR (that is also interesting to see how changes between epochs) \n\nOn a side note, may I ask if it is possible to share your TF-Keras AmEx metric implementation?",
    "1869332": "Using AmEx metrics from [**this notebook**](https://www.kaggle.com/code/rohanrao/amex-competition-metric-implementations).",
    "1870601": "Thanks for sharing, very insightful"
  },
  "source": "meta"
}