{
  "id": 171597,
  "title": "Accuracy 0.98",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/171597",
  "author_name": "",
  "post_date": "2020-08-01T16:38:30.862888700Z",
  "votes": 54,
  "comment_count": 13,
  "views": 0,
  "content": "<p>I am entering this and for first step I am reading the forum.  I see a number of participants asking why accuracy is stuck at 0.98 or so.  There is a very simple explanation for that.</p>\n\n<p>If your model predicts 0 for all images then it is right for all examples with target 0.  Given the proportion of samples with target 0 is 0.9824, your model will have an accuracy of 0.9824.  This is great accuracy, but your model is totally useless. </p>\n\n<p>Even if your model predicts probabilities (which is what you want), using accuracy as a metric assumes predictions are rounded to nearest integer. If your probabilities are well calibrated, then their average is the same as target average, i.e. around 0.017.  Rounding them is most likely to yield 0, and we are back to square one, your accuracy is 0.98xx.</p>\n\n<p>Given you can have a very high accuracy and totally useless model, other metrics are used for imbalanced datasets:  roc-auc as here, or F1 score for instance.  </p>\n\n<p>Conclusion: don't use accuracy, don't even compute it. Use the competition metric for every evaluation you are making.</p>\n\n<p>Sorry if this has been explained already.  It was not in the parts of the forum I read so far.</p>",
  "messages": [
    {
      "id": "954339",
      "postDate": "08/01/2020 16:38:30",
      "content": "<p>I am entering this and for first step I am reading the forum.  I see a number of participants asking why accuracy is stuck at 0.98 or so.  There is a very simple explanation for that.</p>\n\n<p>If your model predicts 0 for all images then it is right for all examples with target 0.  Given the proportion of samples with target 0 is 0.9824, your model will have an accuracy of 0.9824.  This is great accuracy, but your model is totally useless. </p>\n\n<p>Even if your model predicts probabilities (which is what you want), using accuracy as a metric assumes predictions are rounded to nearest integer. If your probabilities are well calibrated, then their average is the same as target average, i.e. around 0.017.  Rounding them is most likely to yield 0, and we are back to square one, your accuracy is 0.98xx.</p>\n\n<p>Given you can have a very high accuracy and totally useless model, other metrics are used for imbalanced datasets:  roc-auc as here, or F1 score for instance.  </p>\n\n<p>Conclusion: don't use accuracy, don't even compute it. Use the competition metric for every evaluation you are making.</p>\n\n<p>Sorry if this has been explained already.  It was not in the parts of the forum I read so far.</p>",
      "rawMarkdown": "I am entering this and for first step I am reading the forum.  I see a number of participants asking why accuracy is stuck at 0.98 or so.  There is a very simple explanation for that.\n\nIf your model predicts 0 for all images then it is right for all examples with target 0.  Given the proportion of samples with target 0 is 0.9824, your model will have an accuracy of 0.9824.  This is great accuracy, but your model is totally useless. \n\nEven if your model predicts probabilities (which is what you want), using accuracy as a metric assumes predictions are rounded to nearest integer. If your probabilities are well calibrated, then their average is the same as target average, i.e. around 0.017.  Rounding them is most likely to yield 0, and we are back to square one, your accuracy is 0.98xx.\n\nGiven you can have a very high accuracy and totally useless model, other metrics are used for imbalanced datasets:  roc-auc as here, or F1 score for instance.  \n\nConclusion: don't use accuracy, don't even compute it. Use the competition metric for every evaluation you are making.\n\nSorry if this has been explained already.  It was not in the parts of the forum I read so far.",
      "votes": null
    },
    {
      "id": "954448",
      "postDate": "08/01/2020 18:26:09",
      "content": "<p>I agree. Using F1 and Recall as my metrics to monitor, as I find ROC-AUC too unstable</p>\n\n<p>cheers\nsid</p>",
      "rawMarkdown": "I agree. Using F1 and Recall as my metrics to monitor, as I find ROC-AUC too unstable\n\ncheers\nsid",
      "votes": null
    },
    {
      "id": "954646",
      "postDate": "08/02/2020 00:29:14",
      "content": "<p>You are so very right. I used aucroc, accuracy, recall and precision and after, well, around 200 iterations on 224X224 images I have the following:</p>\n\n<p>| epoch | train_loss | valid_loss | auroc |accuracy | recall | precision | time|\n| --- | --- |  |  |  |  |  |  |\n| 0 | 0.014131 | 0.256493 | 0.903588 | 0.938095 | 0.483146 | 0.693548 | 01:16|\n| 1 | 0.015516 | 0.239418 | 0.906587 | 0.938095 | 0.561798 | 0.657895| 01:15|\n| 2 | 0.015936 | 0.238434 | 0.911019 | 0.940000 | 0.573034 | 0.671053 | 01:15|\n| 3 | 0.016447 | 0.253368 | 0.895398 | 0.938095 | 0.516854 | 0.676471 | 01:15|\n| 4 | 0.014677 | 0.256090 | 0.911346 | 0.935238 | 0.573034 | 0.629630 | 01:15|\n| 5 | 0.020943 | 0.236019 | 0.912772 | 0.936190 | 0.494382 | 0.666667 | 01:15|\n| 6 | 0.017025 | 0.256317 | 0.904769 | 0.936190 | 0.606742 | 0.627907 | 01:15|\n|  |  |  |  |  |  |  |  |  |</p>\n\n<p>As you can see, it is an overfitted model. However, as CPMP pointed out, accuracy is totally useless. But don't you think aucroc isn't far better? How about Recall and Precision? Could they be used?</p>\n\n<p>Oh and by the way, this model scored a petty ~0.58. </p>",
      "rawMarkdown": "You are so very right. I used aucroc, accuracy, recall and precision and after, well, around 200 iterations on 224X224 images I have the following:\n\n\n\n\n| epoch | train_loss | valid_loss | auroc |accuracy | recall | precision | time|\n| --- | --- |  |  |  |  |  |  |\n| 0 | 0.014131 | 0.256493 | 0.903588 | 0.938095 | 0.483146 | 0.693548 | 01:16|\n| 1 | 0.015516 | 0.239418 | 0.906587 | 0.938095 | 0.561798 | 0.657895| 01:15|\n| 2 | 0.015936 | 0.238434 | 0.911019 | 0.940000 | 0.573034 | 0.671053 | 01:15|\n| 3 | 0.016447 | 0.253368 | 0.895398 | 0.938095 | 0.516854 | 0.676471 | 01:15|\n| 4 | 0.014677 | 0.256090 | 0.911346 | 0.935238 | 0.573034 | 0.629630 | 01:15|\n| 5 | 0.020943 | 0.236019 | 0.912772 | 0.936190 | 0.494382 | 0.666667 | 01:15|\n| 6 | 0.017025 | 0.256317 | 0.904769 | 0.936190 | 0.606742 | 0.627907 | 01:15|\n|  |  |  |  |  |  |  |  |  |\n\nAs you can see, it is an overfitted model. However, as CPMP pointed out, accuracy is totally useless. But don't you think aucroc isn't far better? How about Recall and Precision? Could they be used?\n\nOh and by the way, this model scored a petty ~0.58.",
      "votes": null
    },
    {
      "id": "955093",
      "postDate": "08/02/2020 10:31:35",
      "content": "<p>We are asked to maximize roc-auc.  Why do you bother looking at other metrics?</p>",
      "rawMarkdown": "We are asked to maximize roc-auc.  Why do you bother looking at other metrics?",
      "votes": null
    },
    {
      "id": "955095",
      "postDate": "08/02/2020 10:32:15",
      "content": "<p>How does using F1 aand recall help you maximize roc-auc?</p>",
      "rawMarkdown": "How does using F1 aand recall help you maximize roc-auc?",
      "votes": null
    },
    {
      "id": "955871",
      "postDate": "08/03/2020 03:13:42",
      "content": "<p>By monitor, I mean finding the epoch with best F1 score to save \"best model\".  My Loss fn is still Cross Entropy (using a balanced dataset; under-sample class 0). F1 score is just ROC curve at specific threshold.</p>\n\n<p>My observation is ROC-AUC can be noisy, as seen from simulation results above. This is also mentioned in the ROC wiki (Organizers point to this link) --&gt; <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\">https://en.wikipedia.org/wiki/Receiver_operating_characteristic</a>\n\"The machine learning community most often uses the ROC AUC statistic for model comparison. This practice has been questioned because AUC estimates are quite noisy and suffer from other problems....\nAnother problem with ROC AUC is that reducing the ROC Curve to a single number ignores the fact that it is about the tradeoffs between the different systems or performance points plotted...\"</p>",
      "rawMarkdown": "By monitor, I mean finding the epoch with best F1 score to save \"best model\".  My Loss fn is still Cross Entropy (using a balanced dataset; under-sample class 0). F1 score is just ROC curve at specific threshold.\n\nMy observation is ROC-AUC can be noisy, as seen from simulation results above. This is also mentioned in the ROC wiki (Organizers point to this link) --&gt; https://en.wikipedia.org/wiki/Receiver_operating_characteristic\n\"The machine learning community most often uses the ROC AUC statistic for model comparison. This practice has been questioned because AUC estimates are quite noisy and suffer from other problems....\nAnother problem with ROC AUC is that reducing the ROC Curve to a single number ignores the fact that it is about the tradeoffs between the different systems or performance points plotted...\"",
      "votes": null
    },
    {
      "id": "955971",
      "postDate": "08/03/2020 05:53:54",
      "content": "<p>Well, <a href=\"/sidneyng\">@sidneyng</a> , optimising F1 score (or recall) means that we need to tune yet another hyper-parameter: the threshold. In case of heavily imbalanced data, AUC-ROC is unstable and so is F1 score.</p>\n\n<p>Be sure to tune the threshold well in case you're planning to use F1 score to save the best model.</p>",
      "rawMarkdown": "Well, @sidneyng , optimising F1 score (or recall) means that we need to tune yet another hyper-parameter: the threshold. In case of heavily imbalanced data, AUC-ROC is unstable and so is F1 score.\n\nBe sure to tune the threshold well in case you're planning to use F1 score to save the best model.",
      "votes": null
    },
    {
      "id": "955984",
      "postDate": "08/03/2020 06:09:50",
      "content": "<h1>wow</h1>",
      "rawMarkdown": "# wow",
      "votes": null
    },
    {
      "id": "956147",
      "postDate": "08/03/2020 09:30:22",
      "content": "<p>\"Noisy\": is it not good <em>science</em> pratice to estimate the error of the measure?</p>\n\n<p>If you have small samples, the AUC will have relatively large errors. 1/N+1/M roughly. Another \"noise\" would come from sampling methods. That error can easily be estimated through running two random half splits. Another comes from optimistic/pessimistic ranking. Sort asc &amp; desc will estimate that error.</p>\n\n<p>I find AUC with the associated graph to be informative.</p>\n\n<p>As a tangent: I find it puzzling that Kolmogorov-Smirnov is still used when Cramér-von Mises is available.</p>",
      "rawMarkdown": "\"Noisy\": is it not good *science* pratice to estimate the error of the measure?\n\nIf you have small samples, the AUC will have relatively large errors. 1/N+1/M roughly. Another \"noise\" would come from sampling methods. That error can easily be estimated through running two random half splits. Another comes from optimistic/pessimistic ranking. Sort asc &amp; desc will estimate that error.\n\nI find AUC with the associated graph to be informative.\n\nAs a tangent: I find it puzzling that Kolmogorov-Smirnov is still used when Cramér-von Mises is available.",
      "votes": null
    },
    {
      "id": "958297",
      "postDate": "08/04/2020 23:23:01",
      "content": "<p>Good point.</p>",
      "rawMarkdown": "Good point.",
      "votes": null
    },
    {
      "id": "958551",
      "postDate": "08/05/2020 03:47:03",
      "content": "<p>I still recall my very first kaggle comp. I asked a question on forums about something, maybe my model wasn't training or maybe some leakage, or who knows and I reported on the model's accuracy. Uncle responded very characteristically asking something to the effect of \"Why are you using accuracy when it isn't even the competition metric?\". It was a very good early lesson, one I never forgot.</p>",
      "rawMarkdown": "I still recall my very first kaggle comp. I asked a question on forums about something, maybe my model wasn't training or maybe some leakage, or who knows and I reported on the model's accuracy. Uncle responded very characteristically asking something to the effect of \"Why are you using accuracy when it isn't even the competition metric?\". It was a very good early lesson, one I never forgot.",
      "votes": null
    },
    {
      "id": "959583",
      "postDate": "08/05/2020 18:11:37",
      "content": "<p>Accuracy is usually an illusionary metric and we need to tender a great focus on the precisely overall evaluation metric for the exact kind of the problem we are solving.</p>\n\n<p>Usually, on imbalanced dataset : if considered from just a single point of view of the accuracy metric - we can get the high accuracy as referred with the imbalance but that is totally misleading in itself.</p>",
      "rawMarkdown": "Accuracy is usually an illusionary metric and we need to tender a great focus on the precisely overall evaluation metric for the exact kind of the problem we are solving.\n\nUsually, on imbalanced dataset : if considered from just a single point of view of the accuracy metric - we can get the high accuracy as referred with the imbalance but that is totally misleading in itself.",
      "votes": null
    },
    {
      "id": "959964",
      "postDate": "08/06/2020 04:06:31",
      "content": "<p>Excellent reminder. Thank you.</p>",
      "rawMarkdown": "Excellent reminder. Thank you.",
      "votes": null
    },
    {
      "id": "960072",
      "postDate": "08/06/2020 05:51:23",
      "content": "<p>Auc roc is a ranking metric , there is no thing as a threshold present, so why are you trying to limit everything by strictly lying on either side of 0.5.</p>",
      "rawMarkdown": "Auc roc is a ranking metric , there is no thing as a threshold present, so why are you trying to limit everything by strictly lying on either side of 0.5.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 954448,
      "author_name": "sidneyng",
      "author_url": "",
      "post_date": "08/01/2020 18:26:09",
      "content": "<p>I agree. Using F1 and Recall as my metrics to monitor, as I find ROC-AUC too unstable</p>\n\n<p>cheers\nsid</p>",
      "votes": null,
      "replies": [
        {
          "id": 955095,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/02/2020 10:32:15",
          "content": "<p>How does using F1 aand recall help you maximize roc-auc?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955871,
          "author_name": "sidneyng",
          "author_url": "",
          "post_date": "08/03/2020 03:13:42",
          "content": "<p>By monitor, I mean finding the epoch with best F1 score to save \"best model\".  My Loss fn is still Cross Entropy (using a balanced dataset; under-sample class 0). F1 score is just ROC curve at specific threshold.</p>\n\n<p>My observation is ROC-AUC can be noisy, as seen from simulation results above. This is also mentioned in the ROC wiki (Organizers point to this link) --&gt; <a href=\"https://en.wikipedia.org/wiki/Receiver_operating_characteristic\">https://en.wikipedia.org/wiki/Receiver_operating_characteristic</a>\n\"The machine learning community most often uses the ROC AUC statistic for model comparison. This practice has been questioned because AUC estimates are quite noisy and suffer from other problems....\nAnother problem with ROC AUC is that reducing the ROC Curve to a single number ignores the fact that it is about the tradeoffs between the different systems or performance points plotted...\"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 955971,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "08/03/2020 05:53:54",
          "content": "<p>Well, <a href=\"/sidneyng\">@sidneyng</a> , optimising F1 score (or recall) means that we need to tune yet another hyper-parameter: the threshold. In case of heavily imbalanced data, AUC-ROC is unstable and so is F1 score.</p>\n\n<p>Be sure to tune the threshold well in case you're planning to use F1 score to save the best model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 956147,
          "author_name": "glimmung",
          "author_url": "",
          "post_date": "08/03/2020 09:30:22",
          "content": "<p>\"Noisy\": is it not good <em>science</em> pratice to estimate the error of the measure?</p>\n\n<p>If you have small samples, the AUC will have relatively large errors. 1/N+1/M roughly. Another \"noise\" would come from sampling methods. That error can easily be estimated through running two random half splits. Another comes from optimistic/pessimistic ranking. Sort asc &amp; desc will estimate that error.</p>\n\n<p>I find AUC with the associated graph to be informative.</p>\n\n<p>As a tangent: I find it puzzling that Kolmogorov-Smirnov is still used when Cramér-von Mises is available.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 954646,
      "author_name": "manojacharya",
      "author_url": "",
      "post_date": "08/02/2020 00:29:14",
      "content": "<p>You are so very right. I used aucroc, accuracy, recall and precision and after, well, around 200 iterations on 224X224 images I have the following:</p>\n\n<p>| epoch | train_loss | valid_loss | auroc |accuracy | recall | precision | time|\n| --- | --- |  |  |  |  |  |  |\n| 0 | 0.014131 | 0.256493 | 0.903588 | 0.938095 | 0.483146 | 0.693548 | 01:16|\n| 1 | 0.015516 | 0.239418 | 0.906587 | 0.938095 | 0.561798 | 0.657895| 01:15|\n| 2 | 0.015936 | 0.238434 | 0.911019 | 0.940000 | 0.573034 | 0.671053 | 01:15|\n| 3 | 0.016447 | 0.253368 | 0.895398 | 0.938095 | 0.516854 | 0.676471 | 01:15|\n| 4 | 0.014677 | 0.256090 | 0.911346 | 0.935238 | 0.573034 | 0.629630 | 01:15|\n| 5 | 0.020943 | 0.236019 | 0.912772 | 0.936190 | 0.494382 | 0.666667 | 01:15|\n| 6 | 0.017025 | 0.256317 | 0.904769 | 0.936190 | 0.606742 | 0.627907 | 01:15|\n|  |  |  |  |  |  |  |  |  |</p>\n\n<p>As you can see, it is an overfitted model. However, as CPMP pointed out, accuracy is totally useless. But don't you think aucroc isn't far better? How about Recall and Precision? Could they be used?</p>\n\n<p>Oh and by the way, this model scored a petty ~0.58. </p>",
      "votes": null,
      "replies": [
        {
          "id": 955093,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "08/02/2020 10:31:35",
          "content": "<p>We are asked to maximize roc-auc.  Why do you bother looking at other metrics?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 960072,
          "author_name": "bharatsahu",
          "author_url": "",
          "post_date": "08/06/2020 05:51:23",
          "content": "<p>Auc roc is a ranking metric , there is no thing as a threshold present, so why are you trying to limit everything by strictly lying on either side of 0.5.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 955984,
      "author_name": "mintlemon",
      "author_url": "",
      "post_date": "08/03/2020 06:09:50",
      "content": "<h1>wow</h1>",
      "votes": null,
      "replies": []
    },
    {
      "id": 958297,
      "author_name": "boyangzhang2017",
      "author_url": "",
      "post_date": "08/04/2020 23:23:01",
      "content": "<p>Good point.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 958551,
      "author_name": "authman",
      "author_url": "",
      "post_date": "08/05/2020 03:47:03",
      "content": "<p>I still recall my very first kaggle comp. I asked a question on forums about something, maybe my model wasn't training or maybe some leakage, or who knows and I reported on the model's accuracy. Uncle responded very characteristically asking something to the effect of \"Why are you using accuracy when it isn't even the competition metric?\". It was a very good early lesson, one I never forgot.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959583,
      "author_name": "pksmiles13",
      "author_url": "",
      "post_date": "08/05/2020 18:11:37",
      "content": "<p>Accuracy is usually an illusionary metric and we need to tender a great focus on the precisely overall evaluation metric for the exact kind of the problem we are solving.</p>\n\n<p>Usually, on imbalanced dataset : if considered from just a single point of view of the accuracy metric - we can get the high accuracy as referred with the imbalance but that is totally misleading in itself.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 959964,
      "author_name": "jackisabear",
      "author_url": "",
      "post_date": "08/06/2020 04:06:31",
      "content": "<p>Excellent reminder. Thank you.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "954339": "I am entering this and for first step I am reading the forum.  I see a number of participants asking why accuracy is stuck at 0.98 or so.  There is a very simple explanation for that.\n\nIf your model predicts 0 for all images then it is right for all examples with target 0.  Given the proportion of samples with target 0 is 0.9824, your model will have an accuracy of 0.9824.  This is great accuracy, but your model is totally useless. \n\nEven if your model predicts probabilities (which is what you want), using accuracy as a metric assumes predictions are rounded to nearest integer. If your probabilities are well calibrated, then their average is the same as target average, i.e. around 0.017.  Rounding them is most likely to yield 0, and we are back to square one, your accuracy is 0.98xx.\n\nGiven you can have a very high accuracy and totally useless model, other metrics are used for imbalanced datasets:  roc-auc as here, or F1 score for instance.  \n\nConclusion: don't use accuracy, don't even compute it. Use the competition metric for every evaluation you are making.\n\nSorry if this has been explained already.  It was not in the parts of the forum I read so far.",
    "954448": "I agree. Using F1 and Recall as my metrics to monitor, as I find ROC-AUC too unstable\n\ncheers\nsid",
    "954646": "You are so very right. I used aucroc, accuracy, recall and precision and after, well, around 200 iterations on 224X224 images I have the following:\n\n\n\n\n| epoch | train_loss | valid_loss | auroc |accuracy | recall | precision | time|\n| --- | --- |  |  |  |  |  |  |\n| 0 | 0.014131 | 0.256493 | 0.903588 | 0.938095 | 0.483146 | 0.693548 | 01:16|\n| 1 | 0.015516 | 0.239418 | 0.906587 | 0.938095 | 0.561798 | 0.657895| 01:15|\n| 2 | 0.015936 | 0.238434 | 0.911019 | 0.940000 | 0.573034 | 0.671053 | 01:15|\n| 3 | 0.016447 | 0.253368 | 0.895398 | 0.938095 | 0.516854 | 0.676471 | 01:15|\n| 4 | 0.014677 | 0.256090 | 0.911346 | 0.935238 | 0.573034 | 0.629630 | 01:15|\n| 5 | 0.020943 | 0.236019 | 0.912772 | 0.936190 | 0.494382 | 0.666667 | 01:15|\n| 6 | 0.017025 | 0.256317 | 0.904769 | 0.936190 | 0.606742 | 0.627907 | 01:15|\n|  |  |  |  |  |  |  |  |  |\n\nAs you can see, it is an overfitted model. However, as CPMP pointed out, accuracy is totally useless. But don't you think aucroc isn't far better? How about Recall and Precision? Could they be used?\n\nOh and by the way, this model scored a petty ~0.58.",
    "955093": "We are asked to maximize roc-auc.  Why do you bother looking at other metrics?",
    "955095": "How does using F1 aand recall help you maximize roc-auc?",
    "955871": "By monitor, I mean finding the epoch with best F1 score to save \"best model\".  My Loss fn is still Cross Entropy (using a balanced dataset; under-sample class 0). F1 score is just ROC curve at specific threshold.\n\nMy observation is ROC-AUC can be noisy, as seen from simulation results above. This is also mentioned in the ROC wiki (Organizers point to this link) --&gt; https://en.wikipedia.org/wiki/Receiver_operating_characteristic\n\"The machine learning community most often uses the ROC AUC statistic for model comparison. This practice has been questioned because AUC estimates are quite noisy and suffer from other problems....\nAnother problem with ROC AUC is that reducing the ROC Curve to a single number ignores the fact that it is about the tradeoffs between the different systems or performance points plotted...\"",
    "955971": "Well, @sidneyng , optimising F1 score (or recall) means that we need to tune yet another hyper-parameter: the threshold. In case of heavily imbalanced data, AUC-ROC is unstable and so is F1 score.\n\nBe sure to tune the threshold well in case you're planning to use F1 score to save the best model.",
    "955984": "# wow",
    "956147": "\"Noisy\": is it not good *science* pratice to estimate the error of the measure?\n\nIf you have small samples, the AUC will have relatively large errors. 1/N+1/M roughly. Another \"noise\" would come from sampling methods. That error can easily be estimated through running two random half splits. Another comes from optimistic/pessimistic ranking. Sort asc &amp; desc will estimate that error.\n\nI find AUC with the associated graph to be informative.\n\nAs a tangent: I find it puzzling that Kolmogorov-Smirnov is still used when Cramér-von Mises is available.",
    "958297": "Good point.",
    "958551": "I still recall my very first kaggle comp. I asked a question on forums about something, maybe my model wasn't training or maybe some leakage, or who knows and I reported on the model's accuracy. Uncle responded very characteristically asking something to the effect of \"Why are you using accuracy when it isn't even the competition metric?\". It was a very good early lesson, one I never forgot.",
    "959583": "Accuracy is usually an illusionary metric and we need to tender a great focus on the precisely overall evaluation metric for the exact kind of the problem we are solving.\n\nUsually, on imbalanced dataset : if considered from just a single point of view of the accuracy metric - we can get the high accuracy as referred with the imbalance but that is totally misleading in itself.",
    "959964": "Excellent reminder. Thank you.",
    "960072": "Auc roc is a ranking metric , there is no thing as a threshold present, so why are you trying to limit everything by strictly lying on either side of 0.5."
  },
  "source": "meta"
}