{
  "id": 146989,
  "title": "Discrepancy between accuracy and the competition metric",
  "url": "/competitions/alaska2-image-steganalysis/discussion/146989",
  "author_name": "xhlulu",
  "post_date": "2020-04-29T05:38:45.448000",
  "votes": 12,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hey I just published my new notebook, and I was surprised it got such a high score. Indeed, it was only trained on a small subset (30k/75k), and the accuracy was pretty poor on validation (0.6). But with the competition metric, I got 0.99. </p>\n\n<p>Could there have been an error? </p>",
  "messages": [
    {
      "id": 825576,
      "postDate": "2020-04-29T05:38:45.447Z",
      "content": "<p>Hey I just published my new notebook, and I was surprised it got such a high score. Indeed, it was only trained on a small subset (30k/75k), and the accuracy was pretty poor on validation (0.6). But with the competition metric, I got 0.99. </p>\n\n<p>Could there have been an error? </p>",
      "rawMarkdown": "Hey I just published my new notebook, and I was surprised it got such a high score. Indeed, it was only trained on a small subset (30k/75k), and the accuracy was pretty poor on validation (0.6). But with the competition metric, I got 0.99. \n\nCould there have been an error? ",
      "votes": 12
    },
    {
      "id": 826693,
      "postDate": "2020-04-29T19:39:42.133Z",
      "content": "<p>Hello there,</p>\n\n<p>We have discussed this issue. While it is important for us to focus on low false-positive, it is true that, as it, the metric bring more confusion than pushing towards our goals</p>\n\n<p>Since the competition only started less than 2 days again, we will change the evaluation.</p>\n\n<p>I apologize for the confusion</p>",
      "rawMarkdown": "Hello there,\n\nWe have discussed this issue. While it is important for us to focus on low false-positive, it is true that, as it, the metric bring more confusion than pushing towards our goals\n\nSince the competition only started less than 2 days again, we will change the evaluation.\n\nI apologize for the confusion",
      "votes": 6,
      "replies": [
        {
          "id": 826694,
          "postDate": "2020-04-29T19:40:29.517Z",
          "content": "<p>Thanks for the update. </p>",
          "rawMarkdown": "Thanks for the update. "
        }
      ]
    },
    {
      "id": 826577,
      "postDate": "2020-04-29T18:07:29.930Z",
      "content": "<p>The overall accuracy/AUC do not matter for the competition metrics. The \"perfect\" model for current metrics requires 0 FPR at 40% TPR. Suppose you have such a “perfect” model whose AUC curve consists of two linear lines [(0, 0) -&gt; (0, 0.4), (0, 0.4) -&gt; (1, 1)], the overall AUC is only 0.7 (=0.4 + 0.6/2)! </p>",
      "rawMarkdown": "The overall accuracy/AUC do not matter for the competition metrics. The \"perfect\" model for current metrics requires 0 FPR at 40% TPR. Suppose you have such a “perfect” model whose AUC curve consists of two linear lines [(0, 0) -&gt; (0, 0.4), (0, 0.4) -&gt; (1, 1)], the overall AUC is only 0.7 (=0.4 + 0.6/2)! ",
      "votes": 3
    },
    {
      "id": 826539,
      "postDate": "2020-04-29T17:39:12.973Z",
      "content": "<p>The host's weighted AUC metric seems a bit too generous - it discards everything after 40% TPR!</p>\n\n<p>Even in <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/overview/evaluation\">their own example</a>, they have a model example with 0.804 AUC, and 0.999 weighted AUC. I'm guessing they will have to tweak the scoring function.</p>",
      "rawMarkdown": "The host's weighted AUC metric seems a bit too generous - it discards everything after 40% TPR!\n\nEven in [their own example](https://www.kaggle.com/c/alaska2-image-steganalysis/overview/evaluation), they have a model example with 0.804 AUC, and 0.999 weighted AUC. I'm guessing they will have to tweak the scoring function.",
      "votes": 3
    },
    {
      "id": 825858,
      "postDate": "2020-04-29T09:39:15.587Z",
      "content": "<p>This is the validation set AUC curve of the <a href=\"https://www.kaggle.com/xhlulu/alaska2-efficientnet-on-tpus\">Efficientnet on tpu kernel</a>. \n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1435684%2Ff3986fe49ce0461b332c576e93932b67%2Fcomp%20metric.png?generation=1588152909448200&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>The competition metric ignores the upper part of the AUC curve (everything above TPR &gt; 0.4)</li>\n<li>The slope of the AUC curve at the origin (0, 0) is almost 1, which is extremely beneficial for the comp metric</li>\n</ul>\n\n<p>The metric of 0.9+ seems suspicious nevertheless.</p>",
      "rawMarkdown": "This is the validation set AUC curve of the [Efficientnet on tpu kernel](https://www.kaggle.com/xhlulu/alaska2-efficientnet-on-tpus). \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1435684%2Ff3986fe49ce0461b332c576e93932b67%2Fcomp%20metric.png?generation=1588152909448200&amp;alt=media)\n\n- The competition metric ignores the upper part of the AUC curve (everything above TPR &gt; 0.4)\n- The slope of the AUC curve at the origin (0, 0) is almost 1, which is extremely beneficial for the comp metric\n\n\nThe metric of 0.9+ seems suspicious nevertheless.",
      "votes": 3,
      "replies": [
        {
          "id": 825908,
          "postDate": "2020-04-29T10:18:09.003Z",
          "content": "<p>Yeah, so far this seems suspiciously similar to the Liverpool University competition. I wonder what the expected score range would be for the host.</p>",
          "rawMarkdown": "Yeah, so far this seems suspiciously similar to the Liverpool University competition. I wonder what the expected score range would be for the host."
        },
        {
          "id": 825922,
          "postDate": "2020-04-29T10:28:18.713Z",
          "content": "<p>I created a <a href=\"https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric\">kernel</a> that computes the competition metric. </p>",
          "rawMarkdown": "I created a [kernel](https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric) that computes the competition metric. ",
          "votes": 3
        }
      ]
    },
    {
      "id": 826026,
      "postDate": "2020-04-29T11:52:34.800Z",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> your work is great no doubt. But as you said, it's a bit surprising ofc. This reminds me of the Liverpool University competition, where they had to change the competition metrics. However, the public lb is only for 20% samples, who knows there wouldn't be a shake-up for the rest 80%!!\n<a href=\"/remicogranne\">@remicogranne</a> </p>",
      "rawMarkdown": "@xhlulu your work is great no doubt. But as you said, it's a bit surprising ofc. This reminds me of the Liverpool University competition, where they had to change the competition metrics. However, the public lb is only for 20% samples, who knows there wouldn't be a shake-up for the rest 80%!!\n@remicogranne ",
      "votes": 2
    },
    {
      "id": 825851,
      "postDate": "2020-04-29T09:28:03.127Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 826693,
      "author_name": "Rémi Cogranne",
      "author_url": "",
      "post_date": "2020-04-29T19:39:42.133000",
      "content": "<p>Hello there,</p>\n\n<p>We have discussed this issue. While it is important for us to focus on low false-positive, it is true that, as it, the metric bring more confusion than pushing towards our goals</p>\n\n<p>Since the competition only started less than 2 days again, we will change the evaluation.</p>\n\n<p>I apologize for the confusion</p>",
      "votes": 6,
      "replies": [
        {
          "id": 826694,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-04-29T19:40:29.517000",
          "content": "<p>Thanks for the update. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 826577,
      "author_name": "Xuan Cao",
      "author_url": "",
      "post_date": "2020-04-29T18:07:29.930000",
      "content": "<p>The overall accuracy/AUC do not matter for the competition metrics. The \"perfect\" model for current metrics requires 0 FPR at 40% TPR. Suppose you have such a “perfect” model whose AUC curve consists of two linear lines [(0, 0) -&gt; (0, 0.4), (0, 0.4) -&gt; (1, 1)], the overall AUC is only 0.7 (=0.4 + 0.6/2)! </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 826539,
      "author_name": "anokas",
      "author_url": "",
      "post_date": "2020-04-29T17:39:12.973000",
      "content": "<p>The host's weighted AUC metric seems a bit too generous - it discards everything after 40% TPR!</p>\n\n<p>Even in <a href=\"https://www.kaggle.com/c/alaska2-image-steganalysis/overview/evaluation\">their own example</a>, they have a model example with 0.804 AUC, and 0.999 weighted AUC. I'm guessing they will have to tweak the scoring function.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 825858,
      "author_name": "Max Jeblick",
      "author_url": "",
      "post_date": "2020-04-29T09:39:15.587000",
      "content": "<p>This is the validation set AUC curve of the <a href=\"https://www.kaggle.com/xhlulu/alaska2-efficientnet-on-tpus\">Efficientnet on tpu kernel</a>. \n <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1435684%2Ff3986fe49ce0461b332c576e93932b67%2Fcomp%20metric.png?generation=1588152909448200&amp;alt=media\" alt=\"\"></p>\n\n<ul>\n<li>The competition metric ignores the upper part of the AUC curve (everything above TPR &gt; 0.4)</li>\n<li>The slope of the AUC curve at the origin (0, 0) is almost 1, which is extremely beneficial for the comp metric</li>\n</ul>\n\n<p>The metric of 0.9+ seems suspicious nevertheless.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 825908,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2020-04-29T10:18:09.003000",
          "content": "<p>Yeah, so far this seems suspiciously similar to the Liverpool University competition. I wonder what the expected score range would be for the host.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 825922,
          "author_name": "Max Jeblick",
          "author_url": "",
          "post_date": "2020-04-29T10:28:18.713000",
          "content": "<p>I created a <a href=\"https://www.kaggle.com/maxjeblick/alaska2-efficientnet-on-tpus-competition-metric\">kernel</a> that computes the competition metric. </p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 826026,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-04-29T11:52:34.800000",
      "content": "<p><a href=\"/xhlulu\">@xhlulu</a> your work is great no doubt. But as you said, it's a bit surprising ofc. This reminds me of the Liverpool University competition, where they had to change the competition metrics. However, the public lb is only for 20% samples, who knows there wouldn't be a shake-up for the rest 80%!!\n<a href=\"/remicogranne\">@remicogranne</a> </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 825851,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-29T09:28:03.127000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "825576": "Hey I just published my new notebook, and I was surprised it got such a high score. Indeed, it was only trained on a small subset (30k/75k), and the accuracy was pretty poor on validation (0.6). But with the competition metric, I got 0.99. \n\nCould there have been an error? ",
    "826693": "Hello there,\n\nWe have discussed this issue. While it is important for us to focus on low false-positive, it is true that, as it, the metric bring more confusion than pushing towards our goals\n\nSince the competition only started less than 2 days again, we will change the evaluation.\n\nI apologize for the confusion",
    "826577": "The overall accuracy/AUC do not matter for the competition metrics. The \"perfect\" model for current metrics requires 0 FPR at 40% TPR. Suppose you have such a “perfect” model whose AUC curve consists of two linear lines [(0, 0) -&gt; (0, 0.4), (0, 0.4) -&gt; (1, 1)], the overall AUC is only 0.7 (=0.4 + 0.6/2)! ",
    "826539": "The host's weighted AUC metric seems a bit too generous - it discards everything after 40% TPR!\n\nEven in [their own example](https://www.kaggle.com/c/alaska2-image-steganalysis/overview/evaluation), they have a model example with 0.804 AUC, and 0.999 weighted AUC. I'm guessing they will have to tweak the scoring function.",
    "825858": "This is the validation set AUC curve of the [Efficientnet on tpu kernel](https://www.kaggle.com/xhlulu/alaska2-efficientnet-on-tpus). \n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1435684%2Ff3986fe49ce0461b332c576e93932b67%2Fcomp%20metric.png?generation=1588152909448200&amp;alt=media)\n\n- The competition metric ignores the upper part of the AUC curve (everything above TPR &gt; 0.4)\n- The slope of the AUC curve at the origin (0, 0) is almost 1, which is extremely beneficial for the comp metric\n\n\nThe metric of 0.9+ seems suspicious nevertheless.",
    "826026": "@xhlulu your work is great no doubt. But as you said, it's a bit surprising ofc. This reminds me of the Liverpool University competition, where they had to change the competition metrics. However, the public lb is only for 20% samples, who knows there wouldn't be a shake-up for the rest 80%!!\n@remicogranne ",
    "825851": ""
  }
}