{
  "id": 54521,
  "title": "Is the Leaderboard Difference Meaningful?",
  "url": "/competitions/talkingdata-adtracking-fraud-detection/discussion/54521",
  "author_name": "",
  "post_date": "2018-04-14T09:35:05.555588100Z",
  "votes": 4,
  "comment_count": 6,
  "views": 0,
  "content": "<p>The difference between the top entry and the 50th entry is currently 0.0037.  Is that difference likely to be due to actual improvements in the algorithm or fluke fitting noise in the test set?</p>\n\n<p>I don't really have an intuition for what size difference if score  is actaully meaningful. Does anyone have a feel for this?</p>\n\n<p>Whats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?</p>",
  "messages": [
    {
      "id": "313985",
      "postDate": "04/14/2018 09:35:05",
      "content": "<p>The difference between the top entry and the 50th entry is currently 0.0037.  Is that difference likely to be due to actual improvements in the algorithm or fluke fitting noise in the test set?</p>\n\n<p>I don't really have an intuition for what size difference if score  is actaully meaningful. Does anyone have a feel for this?</p>\n\n<p>Whats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?</p>",
      "rawMarkdown": "The difference between the top entry and the 50th entry is currently 0.0037.  Is that difference likely to be due to actual improvements in the algorithm or fluke fitting noise in the test set?\n\nI don't really have an intuition for what size difference if score  is actaully meaningful. Does anyone have a feel for this?\n\nWhats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?",
      "votes": null
    },
    {
      "id": "313999",
      "postDate": "04/14/2018 10:16:25",
      "content": "<p>Think of it this way.</p>\n\n<p>We have a two solutions with AUC of 0.9785 and AUC of 0.9822. this means that the 50th place solution has still 0.0215 score to improve to reach the perfect model, while the 1st place model has only 0.0178 to improve. So from 50th place point of view, the 1st place has already reached (0.0215 - 0.0178)/0.0215 = 17.2% progress towards the perfect model.  I think such AUC difference has a really significant detection rate increase of True Positives.</p>\n\n<p>Therefore, I would not call such big difference a fluke - there are probably a few key findings 50th place solution has yet to find!</p>",
      "rawMarkdown": "Think of it this way.\n\nWe have a two solutions with AUC of 0.9785 and AUC of 0.9822. this means that the 50th place solution has still 0.0215 score to improve to reach the perfect model, while the 1st place model has only 0.0178 to improve. So from 50th place point of view, the 1st place has already reached (0.0215 - 0.0178)/0.0215 = 17.2% progress towards the perfect model.  I think such AUC difference has a really significant detection rate increase of True Positives.\n\nTherefore, I would not call such big difference a fluke - there are probably a few key findings 50th place solution has yet to find!",
      "votes": null
    },
    {
      "id": "314074",
      "postDate": "04/14/2018 14:47:58",
      "content": "<p>I even think the relative difference is more important, as there cannot be a perfect model here, given duplicate rows can have a different target value.</p>",
      "rawMarkdown": "I even think the relative difference is more important, as there cannot be a perfect model here, given duplicate rows can have a different target value.",
      "votes": null
    },
    {
      "id": "314162",
      "postDate": "04/14/2018 18:47:21",
      "content": "<p>@raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well? any views?\n@CPMP spare me if you have seen this question elsewhere ;)</p>",
      "rawMarkdown": "raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well? any views?\n@CPMP spare me if you have seen this question elsewhere ;)",
      "votes": null
    },
    {
      "id": "314174",
      "postDate": "04/14/2018 19:29:31",
      "content": "<p>They will hold, for sure.</p>",
      "rawMarkdown": "They will hold, for sure.",
      "votes": null
    },
    {
      "id": "314177",
      "postDate": "04/14/2018 19:32:04",
      "content": "<blockquote>\n  <p>Whats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?</p>\n</blockquote>\n\n<p>@raddar covered it well, so I will add a complementary (and complimentary) point of view.</p>\n\n<p>If the data was more balanced and you had a score of 0.975, it wouldn't matter that much if you got to 0.980 because you would already have a very good model for both categories. In a 180-something million dataset with 0.25% of class 1 - which is what they are mostly interested in - this 0.005 improvement can result literally in thousands of class 1 clicks that will be ranked better. That should be very significant.</p>\n\n<blockquote>\n  <p>@raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well?</p>\n</blockquote>\n\n<p>Some models will generalize better than others. It is safe to say that most competitors will have models that score at least 0.96-0.97 simply because of huge data imbalance. But there will be movement in both directions once the private LB is unveiled.</p>",
      "rawMarkdown": "&gt; Whats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?\n\n@raddar covered it well, so I will add a complementary (and complimentary) point of view.\n\nIf the data was more balanced and you had a score of 0.975, it wouldn't matter that much if you got to 0.980 because you would already have a very good model for both categories. In a 180-something million dataset with 0.25% of class 1 - which is what they are mostly interested in - this 0.005 improvement can result literally in thousands of class 1 clicks that will be ranked better. That should be very significant.\n\n&gt; @raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well?\n\nSome models will generalize better than others. It is safe to say that most competitors will have models that score at least 0.96-0.97 simply because of huge data imbalance. But there will be movement in both directions once the private LB is unveiled.",
      "votes": null
    },
    {
      "id": "314178",
      "postDate": "04/14/2018 19:37:16",
      "content": "<p>@raddar @Tilii got my answer. Thanks :)</p>",
      "rawMarkdown": "raddar @Tilii got my answer. Thanks :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 313999,
      "author_name": "raddar",
      "author_url": "",
      "post_date": "04/14/2018 10:16:25",
      "content": "<p>Think of it this way.</p>\n\n<p>We have a two solutions with AUC of 0.9785 and AUC of 0.9822. this means that the 50th place solution has still 0.0215 score to improve to reach the perfect model, while the 1st place model has only 0.0178 to improve. So from 50th place point of view, the 1st place has already reached (0.0215 - 0.0178)/0.0215 = 17.2% progress towards the perfect model.  I think such AUC difference has a really significant detection rate increase of True Positives.</p>\n\n<p>Therefore, I would not call such big difference a fluke - there are probably a few key findings 50th place solution has yet to find!</p>",
      "votes": null,
      "replies": [
        {
          "id": 314074,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "04/14/2018 14:47:58",
          "content": "<p>I even think the relative difference is more important, as there cannot be a perfect model here, given duplicate rows can have a different target value.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314162,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "04/14/2018 18:47:21",
          "content": "<p>@raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well? any views?\n@CPMP spare me if you have seen this question elsewhere ;)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314174,
          "author_name": "raddar",
          "author_url": "",
          "post_date": "04/14/2018 19:29:31",
          "content": "<p>They will hold, for sure.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314177,
          "author_name": "tilii7",
          "author_url": "",
          "post_date": "04/14/2018 19:32:04",
          "content": "<blockquote>\n  <p>Whats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?</p>\n</blockquote>\n\n<p>@raddar covered it well, so I will add a complementary (and complimentary) point of view.</p>\n\n<p>If the data was more balanced and you had a score of 0.975, it wouldn't matter that much if you got to 0.980 because you would already have a very good model for both categories. In a 180-something million dataset with 0.25% of class 1 - which is what they are mostly interested in - this 0.005 improvement can result literally in thousands of class 1 clicks that will be ranked better. That should be very significant.</p>\n\n<blockquote>\n  <p>@raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well?</p>\n</blockquote>\n\n<p>Some models will generalize better than others. It is safe to say that most competitors will have models that score at least 0.96-0.97 simply because of huge data imbalance. But there will be movement in both directions once the private LB is unveiled.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 314178,
          "author_name": "rrqqmm",
          "author_url": "",
          "post_date": "04/14/2018 19:37:16",
          "content": "<p>@raddar @Tilii got my answer. Thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "313985": "The difference between the top entry and the 50th entry is currently 0.0037.  Is that difference likely to be due to actual improvements in the algorithm or fluke fitting noise in the test set?\n\nI don't really have an intuition for what size difference if score  is actaully meaningful. Does anyone have a feel for this?\n\nWhats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?",
    "313999": "Think of it this way.\n\nWe have a two solutions with AUC of 0.9785 and AUC of 0.9822. this means that the 50th place solution has still 0.0215 score to improve to reach the perfect model, while the 1st place model has only 0.0178 to improve. So from 50th place point of view, the 1st place has already reached (0.0215 - 0.0178)/0.0215 = 17.2% progress towards the perfect model.  I think such AUC difference has a really significant detection rate increase of True Positives.\n\nTherefore, I would not call such big difference a fluke - there are probably a few key findings 50th place solution has yet to find!",
    "314074": "I even think the relative difference is more important, as there cannot be a perfect model here, given duplicate rows can have a different target value.",
    "314162": "raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well? any views?\n@CPMP spare me if you have seen this question elsewhere ;)",
    "314174": "They will hold, for sure.",
    "314177": "&gt; Whats the difference to the companies bottom line, if the predictions did actually have a generalisable area-under-the-curve of .0037?\n\n@raddar covered it well, so I will add a complementary (and complimentary) point of view.\n\nIf the data was more balanced and you had a score of 0.975, it wouldn't matter that much if you got to 0.980 because you would already have a very good model for both categories. In a 180-something million dataset with 0.25% of class 1 - which is what they are mostly interested in - this 0.005 improvement can result literally in thousands of class 1 clicks that will be ranked better. That should be very significant.\n\n&gt; @raddar @CPMP would say that these findings are generalized and will hold true for holdout data (private lb) as well?\n\nSome models will generalize better than others. It is safe to say that most competitors will have models that score at least 0.96-0.97 simply because of huge data imbalance. But there will be movement in both directions once the private LB is unveiled.",
    "314178": "raddar @Tilii got my answer. Thanks :)"
  },
  "source": "meta"
}