{
  "id": 22583,
  "title": "Time to talk about LB shaking",
  "url": "/competitions/state-farm-distracted-driver-detection/discussion/22583",
  "author_name": "",
  "post_date": "2016-07-31T06:09:44.207Z",
  "votes": 1,
  "comment_count": 7,
  "views": 1419,
  "content": "<p>Hi, Kagglers,</p>\n\n<p>This is my first time to participate a computer vision competition and I learned a lot here.</p>\n\n<p>In my previous competition experience, we always set a local validation set to prevent overfitting the leaderboard. However, it seems like most public script has cv score differ from the LB score a lot.</p>\n\n<p>Of course, the test set size is big, which could make the result stable.</p>\n\n<p>Considering this two factors, how much we can trust the LB score?</p>\n\n<p>Li</p>",
  "messages": [
    {
      "id": "129558",
      "postDate": "07/31/2016 06:09:44",
      "content": "<p>Hi, Kagglers,</p>\n\n<p>This is my first time to participate a computer vision competition and I learned a lot here.</p>\n\n<p>In my previous competition experience, we always set a local validation set to prevent overfitting the leaderboard. However, it seems like most public script has cv score differ from the LB score a lot.</p>\n\n<p>Of course, the test set size is big, which could make the result stable.</p>\n\n<p>Considering this two factors, how much we can trust the LB score?</p>\n\n<p>Li</p>",
      "rawMarkdown": "Hi, Kagglers,\r\n\r\nThis is my first time to participate a computer vision competition and I learned a lot here.\r\n\r\nIn my previous competition experience, we always set a local validation set to prevent overfitting the leaderboard. However, it seems like most public script has cv score differ from the LB score a lot.\r\n\r\nOf course, the test set size is big, which could make the result stable.\r\n\r\nConsidering this two factors, how much we can trust the LB score?\r\n\r\nLi",
      "votes": null
    },
    {
      "id": "129577",
      "postDate": "07/31/2016 11:31:44",
      "content": "<p>Hi Li Li,</p>\n\n<p>Generally speaking, LB shaking don't occur in computer vision competition.<br>\nHowever, number of drivers is not large in this competition. This may cause LB shaking.</p>",
      "rawMarkdown": "Hi Li Li,\r\n\r\nGenerally speaking, LB shaking don't occur in computer vision competition.<br>\r\nHowever, number of drivers is not large in this competition. This may cause LB shaking.",
      "votes": null
    },
    {
      "id": "129652",
      "postDate": "08/01/2016 09:56:59",
      "content": "<p>If LB split by drivers then there will be large shake up. In other case lower. There are not so many test cases (some is fake and not used for scoring) and split is 31/69. It should give relatively large shake up at the end.</p>",
      "rawMarkdown": "If LB split by drivers then there will be large shake up. In other case lower. There are not so many test cases (some is fake and not used for scoring) and split is 31/69. It should give relatively large shake up at the end.",
      "votes": null
    },
    {
      "id": "129662",
      "postDate": "08/01/2016 12:30:53",
      "content": "<h1>Shakeitoff :)</h1>",
      "rawMarkdown": "#Shakeitoff :)",
      "votes": null
    },
    {
      "id": "129735",
      "postDate": "08/02/2016 01:05:40",
      "content": "<p>0.88 LB to 0.15 Private LB, my hat off :-)</p>\n\n<p>=====================</p>\n\n<p>13  &#8593;474 <br>\nEsper  Team\n0.15209\n58  Mon, 01 Aug 2016 23:17:37 (-0.1h)</p>\n\n<p>=====================</p>",
      "rawMarkdown": "0.88 LB to 0.15 Private LB, my hat off :-)\r\n\r\n=====================\r\n\r\n13\t↑474\t\r\nEsper  Team\r\n0.15209\r\n58\tMon, 01 Aug 2016 23:17:37 (-0.1h)\r\n\r\n=====================",
      "votes": null
    },
    {
      "id": "129749",
      "postDate": "08/02/2016 03:18:57",
      "content": "<p>[quote=Vinh Nguyen;129735]</p>\n\n<p>0.88 LB to 0.15 Private LB, my hat off :-)</p>\n\n<p>=====================</p>\n\n<p>13  &#8593;474 <br>\nEsper  Team\n0.15209\n58  Mon, 01 Aug 2016 23:17:37 (-0.1h)</p>\n\n<p>=====================</p>\n\n<p>[/quote]</p>\n\n<p>Yes, unbelievable!</p>\n\n<p>Anyway, for this &quot;Esper&quot; team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13&#8593;474  jumping is not a result of the data randomness but a result of the competing strategy. </p>\n\n<p>For example, I guess that they have identified some records that belong to the\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13&#8593;474  jump.</p>\n\n<p>This is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^</p>\n\n<p>Anyway, I would be quite curious to hear sth from team &quot;Esper&quot; to confirm they used such a trick (or not)^_^</p>\n\n<p>Best regards,</p>\n\n<p>Shize</p>",
      "rawMarkdown": "[quote=Vinh Nguyen;129735]\r\n\r\n0.88 LB to 0.15 Private LB, my hat off :-)\r\n\r\n=====================\r\n\r\n13\t↑474\t\r\nEsper  Team\r\n0.15209\r\n58\tMon, 01 Aug 2016 23:17:37 (-0.1h)\r\n\r\n=====================\r\n\r\n\r\n[/quote]\r\n\r\nYes, unbelievable!\r\n\r\n Anyway, for this \"Esper\" team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13↑474  jumping is not a result of the data randomness but a result of the competing strategy. \r\n\r\nFor example, I guess that they have identified some records that belong to the\r\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13↑474  jump.\r\n\r\nThis is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^\r\n\r\nAnyway, I would be quite curious to hear sth from team \"Esper\" to confirm they used such a trick (or not)^_^\r\n\r\nBest regards,\r\n\r\nShize",
      "votes": null
    },
    {
      "id": "129793",
      "postDate": "08/02/2016 11:12:08",
      "content": "<p>[quote=Shize Su;129749]</p>\n\n<p>[quote=Vinh Nguyen;129735]</p>\n\n<p>0.88 LB to 0.15 Private LB, my hat off :-)</p>\n\n<p>=====================</p>\n\n<p>13  &#8593;474 <br>\nEsper  Team\n0.15209\n58  Mon, 01 Aug 2016 23:17:37 (-0.1h)</p>\n\n<p>=====================</p>\n\n<p>[/quote]</p>\n\n<p>Yes, unbelievable!</p>\n\n<p>Anyway, for this &quot;Esper&quot; team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13&#8593;474  jumping is not a result of the data randomness but a result of the competing strategy. </p>\n\n<p>For example, I guess that they have identified some records that belong to the\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13&#8593;474  jump.</p>\n\n<p>This is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^</p>\n\n<p>Anyway, I would be quite curious to hear sth from team &quot;Esper&quot; to confirm they used such a trick (or not)^_^</p>\n\n<p>Best regards,</p>\n\n<p>Shize</p>\n\n<p>[/quote]</p>\n\n<p>Score improvement records show this for Esper team:</p>\n\n<p><img src=\"http://i.imgur.com/RGoHUAt.png\" alt=\"enter image description here\" title></p>\n\n<p>This would leave about 40 tries to reverse engineer the private LB in about 1 month + 1 week, and they would have used only 2/3 of available submissions (about 60 available in total).</p>\n\n<p>I can't find the submission scores that are not improvements in the raw data link (<a href=\"https://www.kaggle.com/c/5048/publicleaderboarddata.zip\">https://www.kaggle.com/c/5048/publicleaderboarddata.zip</a>). Once the private raw data is available to download, it might explain more (they used a submission they made the last day, while their best public submissions are 1 month + 1 week ago).</p>",
      "rawMarkdown": "[quote=Shize Su;129749]\r\n\r\n[quote=Vinh Nguyen;129735]\r\n\r\n0.88 LB to 0.15 Private LB, my hat off :-)\r\n\r\n=====================\r\n\r\n13\t↑474\t\r\nEsper  Team\r\n0.15209\r\n58\tMon, 01 Aug 2016 23:17:37 (-0.1h)\r\n\r\n=====================\r\n\r\n\r\n[/quote]\r\n\r\nYes, unbelievable!\r\n\r\n Anyway, for this \"Esper\" team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13↑474  jumping is not a result of the data randomness but a result of the competing strategy. \r\n\r\nFor example, I guess that they have identified some records that belong to the\r\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13↑474  jump.\r\n\r\nThis is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^\r\n\r\nAnyway, I would be quite curious to hear sth from team \"Esper\" to confirm they used such a trick (or not)^_^\r\n\r\nBest regards,\r\n\r\nShize\r\n\r\n\r\n\r\n\r\n[/quote]\r\n\r\nScore improvement records show this for Esper team:\r\n\r\n![enter image description here][1]\r\n\r\nThis would leave about 40 tries to reverse engineer the private LB in about 1 month + 1 week, and they would have used only 2/3 of available submissions (about 60 available in total).\r\n\r\nI can't find the submission scores that are not improvements in the raw data link (https://www.kaggle.com/c/5048/publicleaderboarddata.zip). Once the private raw data is available to download, it might explain more (they used a submission they made the last day, while their best public submissions are 1 month + 1 week ago).\r\n\r\n\r\n  [1]: http://i.imgur.com/RGoHUAt.png",
      "votes": null
    },
    {
      "id": "130660",
      "postDate": "08/09/2016 12:07:27",
      "content": "<p>Once the Meta Kaggle database will be updated we will be able to see what happened exactly with the Esper team: <a href=\"https://www.kaggle.com/laurae2/d/kaggle/meta-kaggle/state-farm-esper-team-performance/code\">https://www.kaggle.com/laurae2/d/kaggle/meta-kaggle/state-farm-esper-team-performance/code</a></p>",
      "rawMarkdown": "Once the Meta Kaggle database will be updated we will be able to see what happened exactly with the Esper team: https://www.kaggle.com/laurae2/d/kaggle/meta-kaggle/state-farm-esper-team-performance/code",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 129577,
      "author_name": "toshik",
      "author_url": "",
      "post_date": "07/31/2016 11:31:44",
      "content": "<p>Hi Li Li,</p>\n\n<p>Generally speaking, LB shaking don't occur in computer vision competition.<br>\nHowever, number of drivers is not large in this competition. This may cause LB shaking.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129652,
      "author_name": "zfturbo",
      "author_url": "",
      "post_date": "08/01/2016 09:56:59",
      "content": "<p>If LB split by drivers then there will be large shake up. In other case lower. There are not so many test cases (some is fake and not used for scoring) and split is 31/69. It should give relatively large shake up at the end.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129662,
      "author_name": "",
      "author_url": "",
      "post_date": "08/01/2016 12:30:53",
      "content": "<h1>Shakeitoff :)</h1>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129735,
      "author_name": "vinhnguyen",
      "author_url": "",
      "post_date": "08/02/2016 01:05:40",
      "content": "<p>0.88 LB to 0.15 Private LB, my hat off :-)</p>\n\n<p>=====================</p>\n\n<p>13  &#8593;474 <br>\nEsper  Team\n0.15209\n58  Mon, 01 Aug 2016 23:17:37 (-0.1h)</p>\n\n<p>=====================</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129749,
      "author_name": "sushize",
      "author_url": "",
      "post_date": "08/02/2016 03:18:57",
      "content": "<p>[quote=Vinh Nguyen;129735]</p>\n\n<p>0.88 LB to 0.15 Private LB, my hat off :-)</p>\n\n<p>=====================</p>\n\n<p>13  &#8593;474 <br>\nEsper  Team\n0.15209\n58  Mon, 01 Aug 2016 23:17:37 (-0.1h)</p>\n\n<p>=====================</p>\n\n<p>[/quote]</p>\n\n<p>Yes, unbelievable!</p>\n\n<p>Anyway, for this &quot;Esper&quot; team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13&#8593;474  jumping is not a result of the data randomness but a result of the competing strategy. </p>\n\n<p>For example, I guess that they have identified some records that belong to the\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13&#8593;474  jump.</p>\n\n<p>This is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^</p>\n\n<p>Anyway, I would be quite curious to hear sth from team &quot;Esper&quot; to confirm they used such a trick (or not)^_^</p>\n\n<p>Best regards,</p>\n\n<p>Shize</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 129793,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "08/02/2016 11:12:08",
      "content": "<p>[quote=Shize Su;129749]</p>\n\n<p>[quote=Vinh Nguyen;129735]</p>\n\n<p>0.88 LB to 0.15 Private LB, my hat off :-)</p>\n\n<p>=====================</p>\n\n<p>13  &#8593;474 <br>\nEsper  Team\n0.15209\n58  Mon, 01 Aug 2016 23:17:37 (-0.1h)</p>\n\n<p>=====================</p>\n\n<p>[/quote]</p>\n\n<p>Yes, unbelievable!</p>\n\n<p>Anyway, for this &quot;Esper&quot; team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13&#8593;474  jumping is not a result of the data randomness but a result of the competing strategy. </p>\n\n<p>For example, I guess that they have identified some records that belong to the\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13&#8593;474  jump.</p>\n\n<p>This is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^</p>\n\n<p>Anyway, I would be quite curious to hear sth from team &quot;Esper&quot; to confirm they used such a trick (or not)^_^</p>\n\n<p>Best regards,</p>\n\n<p>Shize</p>\n\n<p>[/quote]</p>\n\n<p>Score improvement records show this for Esper team:</p>\n\n<p><img src=\"http://i.imgur.com/RGoHUAt.png\" alt=\"enter image description here\" title></p>\n\n<p>This would leave about 40 tries to reverse engineer the private LB in about 1 month + 1 week, and they would have used only 2/3 of available submissions (about 60 available in total).</p>\n\n<p>I can't find the submission scores that are not improvements in the raw data link (<a href=\"https://www.kaggle.com/c/5048/publicleaderboarddata.zip\">https://www.kaggle.com/c/5048/publicleaderboarddata.zip</a>). Once the private raw data is available to download, it might explain more (they used a submission they made the last day, while their best public submissions are 1 month + 1 week ago).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 130660,
      "author_name": "laurae2",
      "author_url": "",
      "post_date": "08/09/2016 12:07:27",
      "content": "<p>Once the Meta Kaggle database will be updated we will be able to see what happened exactly with the Esper team: <a href=\"https://www.kaggle.com/laurae2/d/kaggle/meta-kaggle/state-farm-esper-team-performance/code\">https://www.kaggle.com/laurae2/d/kaggle/meta-kaggle/state-farm-esper-team-performance/code</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "129558": "Hi, Kagglers,\r\n\r\nThis is my first time to participate a computer vision competition and I learned a lot here.\r\n\r\nIn my previous competition experience, we always set a local validation set to prevent overfitting the leaderboard. However, it seems like most public script has cv score differ from the LB score a lot.\r\n\r\nOf course, the test set size is big, which could make the result stable.\r\n\r\nConsidering this two factors, how much we can trust the LB score?\r\n\r\nLi",
    "129577": "Hi Li Li,\r\n\r\nGenerally speaking, LB shaking don't occur in computer vision competition.<br>\r\nHowever, number of drivers is not large in this competition. This may cause LB shaking.",
    "129652": "If LB split by drivers then there will be large shake up. In other case lower. There are not so many test cases (some is fake and not used for scoring) and split is 31/69. It should give relatively large shake up at the end.",
    "129662": "#Shakeitoff :)",
    "129735": "0.88 LB to 0.15 Private LB, my hat off :-)\r\n\r\n=====================\r\n\r\n13\t↑474\t\r\nEsper  Team\r\n0.15209\r\n58\tMon, 01 Aug 2016 23:17:37 (-0.1h)\r\n\r\n=====================",
    "129749": "[quote=Vinh Nguyen;129735]\r\n\r\n0.88 LB to 0.15 Private LB, my hat off :-)\r\n\r\n=====================\r\n\r\n13\t↑474\t\r\nEsper  Team\r\n0.15209\r\n58\tMon, 01 Aug 2016 23:17:37 (-0.1h)\r\n\r\n=====================\r\n\r\n\r\n[/quote]\r\n\r\nYes, unbelievable!\r\n\r\n Anyway, for this \"Esper\" team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13↑474  jumping is not a result of the data randomness but a result of the competing strategy. \r\n\r\nFor example, I guess that they have identified some records that belong to the\r\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13↑474  jump.\r\n\r\nThis is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^\r\n\r\nAnyway, I would be quite curious to hear sth from team \"Esper\" to confirm they used such a trick (or not)^_^\r\n\r\nBest regards,\r\n\r\nShize",
    "129793": "[quote=Shize Su;129749]\r\n\r\n[quote=Vinh Nguyen;129735]\r\n\r\n0.88 LB to 0.15 Private LB, my hat off :-)\r\n\r\n=====================\r\n\r\n13\t↑474\t\r\nEsper  Team\r\n0.15209\r\n58\tMon, 01 Aug 2016 23:17:37 (-0.1h)\r\n\r\n=====================\r\n\r\n\r\n[/quote]\r\n\r\nYes, unbelievable!\r\n\r\n Anyway, for this \"Esper\" team, I have a conjecture that they probably just had disguised their public LB score (as a competing strategy for fun), and such 13↑474  jumping is not a result of the data randomness but a result of the competing strategy. \r\n\r\nFor example, I guess that they have identified some records that belong to the\r\npublic LB from previous submissions, and intentionally made very bad predictions for these identified public LB test records while made regular predictions on the rest, and thus they get such a 13↑474  jump.\r\n\r\nThis is the only explanation I can have. I don't think the data randomness in this competition would be that huge (0.88 Public LB, 0.15 private LB)^_^\r\n\r\nAnyway, I would be quite curious to hear sth from team \"Esper\" to confirm they used such a trick (or not)^_^\r\n\r\nBest regards,\r\n\r\nShize\r\n\r\n\r\n\r\n\r\n[/quote]\r\n\r\nScore improvement records show this for Esper team:\r\n\r\n![enter image description here][1]\r\n\r\nThis would leave about 40 tries to reverse engineer the private LB in about 1 month + 1 week, and they would have used only 2/3 of available submissions (about 60 available in total).\r\n\r\nI can't find the submission scores that are not improvements in the raw data link (https://www.kaggle.com/c/5048/publicleaderboarddata.zip). Once the private raw data is available to download, it might explain more (they used a submission they made the last day, while their best public submissions are 1 month + 1 week ago).\r\n\r\n\r\n  [1]: http://i.imgur.com/RGoHUAt.png",
    "130660": "Once the Meta Kaggle database will be updated we will be able to see what happened exactly with the Esper team: https://www.kaggle.com/laurae2/d/kaggle/meta-kaggle/state-farm-esper-team-performance/code"
  },
  "source": "meta"
}