{
  "id": 19459,
  "title": "Final leaderboard position",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19459",
  "author_name": "CRD",
  "post_date": "2016-03-12T08:30:01.600000",
  "votes": 13,
  "comment_count": 15,
  "views": 2514,
  "content": "<p>Hi,</p>\n\n<p>I was wondering about the final leaderboard position. Will it take into account only people who make a submission on the test dataset or will it consider alll the participants who made a submission during the validation stage?</p>\n\n<p>I am far from the prizes but I could have a chance with top 10% / top 25%. If only test stage is going to be considered, there is a trong self selection so my objective might be further than expected (people with low scores during validation will be less willing to make a submission during test stage).</p>\n\n<p>Could someone clarify this?</p>\n\n<p>Thanks in advance!</p>",
  "messages": [
    {
      "id": 111165,
      "postDate": "2016-03-12T08:30:01.600Z",
      "content": "<p>Hi,</p>\n\n<p>I was wondering about the final leaderboard position. Will it take into account only people who make a submission on the test dataset or will it consider alll the participants who made a submission during the validation stage?</p>\n\n<p>I am far from the prizes but I could have a chance with top 10% / top 25%. If only test stage is going to be considered, there is a trong self selection so my objective might be further than expected (people with low scores during validation will be less willing to make a submission during test stage).</p>\n\n<p>Could someone clarify this?</p>\n\n<p>Thanks in advance!</p>",
      "rawMarkdown": "Hi,\r\n\r\nI was wondering about the final leaderboard position. Will it take into account only people who make a submission on the test dataset or will it consider alll the participants who made a submission during the validation stage?\r\n\r\nI am far from the prizes but I could have a chance with top 10% / top 25%. If only test stage is going to be considered, there is a trong self selection so my objective might be further than expected (people with low scores during validation will be less willing to make a submission during test stage).\r\n\r\nCould someone clarify this?\r\n\r\nThanks in advance!\r\n",
      "votes": 13
    },
    {
      "id": 111811,
      "postDate": "2016-03-16T20:26:37.230Z",
      "content": "<p>Isn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?</p>\n\n<p>I would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.</p>",
      "rawMarkdown": "Isn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?\r\n\r\nI would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\r\n",
      "votes": 9
    },
    {
      "id": 111816,
      "postDate": "2016-03-16T20:54:14.727Z",
      "content": "<p>[quote=Julian de Wit;111811]\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?</p>\n\n<p>I would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\n[/quote]</p>\n\n<p>Exactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?</p>",
      "rawMarkdown": "[quote=Julian de Wit;111811]\r\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?\r\n\r\nI would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\r\n[/quote]\r\n\r\nExactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?",
      "votes": 10
    },
    {
      "id": 112467,
      "postDate": "2016-03-21T12:55:06.110Z",
      "content": "<p>I was not that bad at this competition (around 0.33) but did not upload a new model, because I was occupied with other things, so now I am out of the leaderboard. That is making me a bit sad and it was not that clear to me, that this would happen. In the description it was only written, that they would not be recognised for &quot;prizes&quot; (position 1, 2 and 3), but not that they would not be recognised in the leaderboard. \nThat was a bad communication. </p>\n\n<p>Also the statement of Ryan is very right. </p>",
      "rawMarkdown": "I was not that bad at this competition (around 0.33) but did not upload a new model, because I was occupied with other things, so now I am out of the leaderboard. That is making me a bit sad and it was not that clear to me, that this would happen. In the description it was only written, that they would not be recognised for \"prizes\" (position 1, 2 and 3), but not that they would not be recognised in the leaderboard. \r\nThat was a bad communication. \r\n\r\nAlso the statement of Ryan is very right. ",
      "votes": 8
    },
    {
      "id": 112016,
      "postDate": "2016-03-17T20:02:25.573Z",
      "content": "<p>I feel like people are talking about two different things.</p>\n\n<p>This competition certainly does undervalue a person's final standing in comparison to normal competitions.</p>\n\n<p>Before stage 2 this competition had 700+ teams.  Now it is less than 200 teams.</p>\n\n<p>I am actually surprised it was that high.  Except that people wanted the points, and now position 100 on the final leaderboard is severely undervalued compared to normal Kaggle competitions.</p>",
      "rawMarkdown": "I feel like people are talking about two different things.\r\n\r\nThis competition certainly does undervalue a person's final standing in comparison to normal competitions.\r\n\r\nBefore stage 2 this competition had 700+ teams.  Now it is less than 200 teams.\r\n\r\nI am actually surprised it was that high.  Except that people wanted the points, and now position 100 on the final leaderboard is severely undervalued compared to normal Kaggle competitions.",
      "votes": 8
    },
    {
      "id": 111859,
      "postDate": "2016-03-17T01:01:54.260Z",
      "content": "<p>Not to belabor the point, but there is a certain irony to the two-stage LB setup: The harder the problem, the less impressive your position on the final LB appears to be (since more people drop out).</p>",
      "rawMarkdown": "Not to belabor the point, but there is a certain irony to the two-stage LB setup: The harder the problem, the less impressive your position on the final LB appears to be (since more people drop out).",
      "votes": 7
    },
    {
      "id": 111840,
      "postDate": "2016-03-16T23:18:07.490Z",
      "content": "<p>[quote=inversion;111816]</p>\n\n<p>[quote=Julian de Wit;111811]\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?</p>\n\n<p>I would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\n[/quote]</p>\n\n<p>Exactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?</p>\n\n<p>[/quote]</p>\n\n<p>I could have not expressed it better. It is exactly like if in regular-contests kaggle would not consider people who do not submit during the last week... and Kaggle takes into account even people who only make a single dummy submission.\nI think considering in the final leaderboard everyone who participated in the first stage would make this contest more consistent with other ones in terms of kaggle points. </p>",
      "rawMarkdown": "[quote=inversion;111816]\r\n\r\n[quote=Julian de Wit;111811]\r\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?\r\n\r\nI would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\r\n[/quote]\r\n\r\nExactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?\r\n\r\n[/quote]\r\n\r\nI could have not expressed it better. It is exactly like if in regular-contests kaggle would not consider people who do not submit during the last week... and Kaggle takes into account even people who only make a single dummy submission.\r\nI think considering in the final leaderboard everyone who participated in the first stage would make this contest more consistent with other ones in terms of kaggle points. \r\n",
      "votes": 6
    },
    {
      "id": 111787,
      "postDate": "2016-03-16T19:26:27.887Z",
      "content": "<p>Looks like the leadearboard was calculated only for people who did submit in the second round. </p>\n\n<p>This somewhat doesn't feel right to me due to the very strong self-selction (don't remember the number, but roughly only about 10% of people did submit in the second round) - could admins specify this in advance for future contests ?</p>",
      "rawMarkdown": "Looks like the leadearboard was calculated only for people who did submit in the second round. \r\n\r\nThis somewhat doesn't feel right to me due to the very strong self-selction (don't remember the number, but roughly only about 10% of people did submit in the second round) - could admins specify this in advance for future contests ?",
      "votes": 6
    },
    {
      "id": 112949,
      "postDate": "2016-03-25T12:39:46.780Z",
      "content": "<p>I totally agree with the previous posts too. For us, we believe we had  a consistent and stable public and private LB, and the winning is quite significant instead of noise. Now with this meaningless  new public LB, it gives a impression that the winning has a lot of noise and it included a lot of luck.</p>",
      "rawMarkdown": "I totally agree with the previous posts too. For us, we believe we had  a consistent and stable public and private LB, and the winning is quite significant instead of noise. Now with this meaningless  new public LB, it gives a impression that the winning has a lot of noise and it included a lot of luck.",
      "votes": 3
    },
    {
      "id": 113000,
      "postDate": "2016-03-25T23:57:22.347Z",
      "content": "<p>I would like to join in expressing my disappointment with this issue. I spent many hours on the competition and had a respectable top 10% score on the public leaderboard. I made very few submissions so clearly my success was not due to overfitting. I had very limited time to work on the project during the one week of the 2nd stage so I was not able to reproduce a high scoring submission. </p>\n\n<p>At the very least, it would be nice if the original public leaderboard was made available again so I can brag to my mommy. </p>",
      "rawMarkdown": "I would like to join in expressing my disappointment with this issue. I spent many hours on the competition and had a respectable top 10% score on the public leaderboard. I made very few submissions so clearly my success was not due to overfitting. I had very limited time to work on the project during the one week of the 2nd stage so I was not able to reproduce a high scoring submission. \r\n\r\nAt the very least, it would be nice if the original public leaderboard was made available again so I can brag to my mommy. ",
      "votes": 4
    },
    {
      "id": 113048,
      "postDate": "2016-03-26T19:27:48.240Z",
      "content": "<p>Thanks for the thoughtful comments. </p>\n\n<p>First off, as always, we will not make retrospective changes to how we handle past competitions (including this one). When issues like this come up we use it as an opportunity to evaluate how we might improve in the future.</p>\n\n<p>Internally our debate focused on three issues:  </p>\n\n<ol>\n<li>recognition for those who completed stage one but not stage two </li>\n<li>achievements and how the competition appears on profiles (10th out of 750 looks more impressive than 10th out of 172)  </li>\n<li>how points are handled  </li>\n</ol>\n\n<p><strong>1. recognition for those who completed stage one but not stage two</strong></p>\n\n<p>We need to view the stage one leaderboard as having no weight: if it gets a weight, we incentivize overfitting or hand labeling for stage one.</p>\n\n<p><strong>2. achievements and how the competition appears on profiles</strong></p>\n\n<p>If we did what Julian suggests and add stage one participants to the bottom of the stage two leaderboard, we undermine our rankings by making it very easy for somebody to get an impressive-looking top 25% achievement by finishing 187th out of 750 with a naive submission. </p>\n\n<p><strong>3.   how points are handled</strong>  </p>\n\n<p>The one change we will make in future is the way points are handled. We will add a multiplier to the number of points for a two-stage competition. We have not settled on a formula for doing this yet, but commit to communicating it clearly in the rules of the next two-stage competition.</p>\n\n<p>These are difficult issues but we think this approach strikes the best balance between competing considerations. </p>",
      "rawMarkdown": "Thanks for the thoughtful comments. \r\n\r\nFirst off, as always, we will not make retrospective changes to how we handle past competitions (including this one). When issues like this come up we use it as an opportunity to evaluate how we might improve in the future.\r\n\r\nInternally our debate focused on three issues:  \r\n\r\n 1. recognition for those who completed stage one but not stage two \r\n 2.  achievements and how the competition appears on profiles (10th out of 750 looks more impressive than 10th out of 172)  \r\n 3.   how points are handled  \r\n \r\n**1. recognition for those who completed stage one but not stage two**\r\n\r\nWe need to view the stage one leaderboard as having no weight: if it gets a weight, we incentivize overfitting or hand labeling for stage one.\r\n\r\n**2. achievements and how the competition appears on profiles**\r\n\r\nIf we did what Julian suggests and add stage one participants to the bottom of the stage two leaderboard, we undermine our rankings by making it very easy for somebody to get an impressive-looking top 25% achievement by finishing 187th out of 750 with a naive submission. \r\n\r\n**3.   how points are handled**  \r\n\r\nThe one change we will make in future is the way points are handled. We will add a multiplier to the number of points for a two-stage competition. We have not settled on a formula for doing this yet, but commit to communicating it clearly in the rules of the next two-stage competition.\r\n\r\nThese are difficult issues but we think this approach strikes the best balance between competing considerations. ",
      "votes": 3
    },
    {
      "id": 111889,
      "postDate": "2016-03-17T05:34:10.427Z",
      "content": "<p>The public leaderboard only count 1% of the data, which is extremely low. The scores are so screwed, and it is quite meaningless.</p>\n\n<p>Given that the competition is over, the public leaderboard should at least count more, like 50% as in &quot;Santander Customer Satisfaction&quot;.</p>",
      "rawMarkdown": "The public leaderboard only count 1% of the data, which is extremely low. The scores are so screwed, and it is quite meaningless.\r\n\r\nGiven that the competition is over, the public leaderboard should at least count more, like 50% as in \"Santander Customer Satisfaction\"."
    },
    {
      "id": 111792,
      "postDate": "2016-03-16T19:45:32.717Z",
      "content": "<p>Hey Alchemist,</p>\n\n<p>Sorry to hear that this wasn't clear. This was the intent of this bullet point in the rules:</p>\n\n<ul>\n<li>This is a two-stage competition. The final results of the competition will be determined by the test dataset released for the second stage.</li>\n</ul>\n\n<p>We don't keep two leaderboards because it'd be confusing (also, the first stage was a 100% public leaderboard and could therefore be gamed). We don't have a mixed stage one + stage two leaderboard because there's no right way to order teams.</p>\n\n<p>Two stage competitions are rare on Kaggle, but we'll do our best to message this more clearly in the future. Thanks for the feedback.</p>",
      "rawMarkdown": "Hey Alchemist,\r\n\r\nSorry to hear that this wasn't clear. This was the intent of this bullet point in the rules:\r\n\r\n - This is a two-stage competition. The final results of the competition will be determined by the test dataset released for the second stage.\r\n\r\nWe don't keep two leaderboards because it'd be confusing (also, the first stage was a 100% public leaderboard and could therefore be gamed). We don't have a mixed stage one + stage two leaderboard because there's no right way to order teams.\r\n\r\nTwo stage competitions are rare on Kaggle, but we'll do our best to message this more clearly in the future. Thanks for the feedback.",
      "votes": -5
    },
    {
      "id": 112973,
      "postDate": "2016-03-25T19:21:12.567Z",
      "content": "<p>@yuanfang it saves the result to file after it finishes all cases. I'm traveling in Peru and so did not have time to check bbs </p>",
      "rawMarkdown": "@yuanfang it saves the result to file after it finishes all cases. I'm traveling in Peru and so did not have time to check bbs \r\n"
    },
    {
      "id": 112974,
      "postDate": "2016-03-25T19:27:33.740Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    },
    {
      "id": 112969,
      "postDate": "2016-03-25T18:38:30.237Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 111811,
      "author_name": "Julian de Wit",
      "author_url": "",
      "post_date": "2016-03-16T20:26:37.230000",
      "content": "<p>Isn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?</p>\n\n<p>I would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.</p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 111816,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-16T20:54:14.727000",
      "content": "<p>[quote=Julian de Wit;111811]\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?</p>\n\n<p>I would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\n[/quote]</p>\n\n<p>Exactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?</p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 112467,
      "author_name": "Icedragon",
      "author_url": "",
      "post_date": "2016-03-21T12:55:06.110000",
      "content": "<p>I was not that bad at this competition (around 0.33) but did not upload a new model, because I was occupied with other things, so now I am out of the leaderboard. That is making me a bit sad and it was not that clear to me, that this would happen. In the description it was only written, that they would not be recognised for &quot;prizes&quot; (position 1, 2 and 3), but not that they would not be recognised in the leaderboard. \nThat was a bad communication. </p>\n\n<p>Also the statement of Ryan is very right. </p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 112016,
      "author_name": "Ryan Munion",
      "author_url": "",
      "post_date": "2016-03-17T20:02:25.573000",
      "content": "<p>I feel like people are talking about two different things.</p>\n\n<p>This competition certainly does undervalue a person's final standing in comparison to normal competitions.</p>\n\n<p>Before stage 2 this competition had 700+ teams.  Now it is less than 200 teams.</p>\n\n<p>I am actually surprised it was that high.  Except that people wanted the points, and now position 100 on the final leaderboard is severely undervalued compared to normal Kaggle competitions.</p>",
      "votes": 8,
      "replies": []
    },
    {
      "id": 111859,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-17T01:01:54.260000",
      "content": "<p>Not to belabor the point, but there is a certain irony to the two-stage LB setup: The harder the problem, the less impressive your position on the final LB appears to be (since more people drop out).</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 111840,
      "author_name": "CRD",
      "author_url": "",
      "post_date": "2016-03-16T23:18:07.490000",
      "content": "<p>[quote=inversion;111816]</p>\n\n<p>[quote=Julian de Wit;111811]\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?</p>\n\n<p>I would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\n[/quote]</p>\n\n<p>Exactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?</p>\n\n<p>[/quote]</p>\n\n<p>I could have not expressed it better. It is exactly like if in regular-contests kaggle would not consider people who do not submit during the last week... and Kaggle takes into account even people who only make a single dummy submission.\nI think considering in the final leaderboard everyone who participated in the first stage would make this contest more consistent with other ones in terms of kaggle points. </p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 111787,
      "author_name": "Alchemist",
      "author_url": "",
      "post_date": "2016-03-16T19:26:27.887000",
      "content": "<p>Looks like the leadearboard was calculated only for people who did submit in the second round. </p>\n\n<p>This somewhat doesn't feel right to me due to the very strong self-selction (don't remember the number, but roughly only about 10% of people did submit in the second round) - could admins specify this in advance for future contests ?</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 112949,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-25T12:39:46.780000",
      "content": "<p>I totally agree with the previous posts too. For us, we believe we had  a consistent and stable public and private LB, and the winning is quite significant instead of noise. Now with this meaningless  new public LB, it gives a impression that the winning has a lot of noise and it included a lot of luck.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 113000,
      "author_name": "Ben Alpert",
      "author_url": "",
      "post_date": "2016-03-25T23:57:22.347000",
      "content": "<p>I would like to join in expressing my disappointment with this issue. I spent many hours on the competition and had a respectable top 10% score on the public leaderboard. I made very few submissions so clearly my success was not due to overfitting. I had very limited time to work on the project during the one week of the 2nd stage so I was not able to reproduce a high scoring submission. </p>\n\n<p>At the very least, it would be nice if the original public leaderboard was made available again so I can brag to my mommy. </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 113048,
      "author_name": "Anthony Goldbloom",
      "author_url": "",
      "post_date": "2016-03-26T19:27:48.240000",
      "content": "<p>Thanks for the thoughtful comments. </p>\n\n<p>First off, as always, we will not make retrospective changes to how we handle past competitions (including this one). When issues like this come up we use it as an opportunity to evaluate how we might improve in the future.</p>\n\n<p>Internally our debate focused on three issues:  </p>\n\n<ol>\n<li>recognition for those who completed stage one but not stage two </li>\n<li>achievements and how the competition appears on profiles (10th out of 750 looks more impressive than 10th out of 172)  </li>\n<li>how points are handled  </li>\n</ol>\n\n<p><strong>1. recognition for those who completed stage one but not stage two</strong></p>\n\n<p>We need to view the stage one leaderboard as having no weight: if it gets a weight, we incentivize overfitting or hand labeling for stage one.</p>\n\n<p><strong>2. achievements and how the competition appears on profiles</strong></p>\n\n<p>If we did what Julian suggests and add stage one participants to the bottom of the stage two leaderboard, we undermine our rankings by making it very easy for somebody to get an impressive-looking top 25% achievement by finishing 187th out of 750 with a naive submission. </p>\n\n<p><strong>3.   how points are handled</strong>  </p>\n\n<p>The one change we will make in future is the way points are handled. We will add a multiplier to the number of points for a two-stage competition. We have not settled on a formula for doing this yet, but commit to communicating it clearly in the rules of the next two-stage competition.</p>\n\n<p>These are difficult issues but we think this approach strikes the best balance between competing considerations. </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 111889,
      "author_name": "The Miner",
      "author_url": "",
      "post_date": "2016-03-17T05:34:10.427000",
      "content": "<p>The public leaderboard only count 1% of the data, which is extremely low. The scores are so screwed, and it is quite meaningless.</p>\n\n<p>Given that the competition is over, the public leaderboard should at least count more, like 50% as in &quot;Santander Customer Satisfaction&quot;.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111792,
      "author_name": "Will Cukierski",
      "author_url": "",
      "post_date": "2016-03-16T19:45:32.717000",
      "content": "<p>Hey Alchemist,</p>\n\n<p>Sorry to hear that this wasn't clear. This was the intent of this bullet point in the rules:</p>\n\n<ul>\n<li>This is a two-stage competition. The final results of the competition will be determined by the test dataset released for the second stage.</li>\n</ul>\n\n<p>We don't keep two leaderboards because it'd be confusing (also, the first stage was a 100% public leaderboard and could therefore be gamed). We don't have a mixed stage one + stage two leaderboard because there's no right way to order teams.</p>\n\n<p>Two stage competitions are rare on Kaggle, but we'll do our best to message this more clearly in the future. Thanks for the feedback.</p>",
      "votes": -5,
      "replies": []
    },
    {
      "id": 112973,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-25T19:21:12.567000",
      "content": "<p>@yuanfang it saves the result to file after it finishes all cases. I'm traveling in Peru and so did not have time to check bbs </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 112974,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-25T19:27:33.740000",
      "content": "",
      "votes": -2,
      "replies": []
    },
    {
      "id": 112969,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-25T18:38:30.237000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "111165": "Hi,\r\n\r\nI was wondering about the final leaderboard position. Will it take into account only people who make a submission on the test dataset or will it consider alll the participants who made a submission during the validation stage?\r\n\r\nI am far from the prizes but I could have a chance with top 10% / top 25%. If only test stage is going to be considered, there is a trong self selection so my objective might be further than expected (people with low scores during validation will be less willing to make a submission during test stage).\r\n\r\nCould someone clarify this?\r\n\r\nThanks in advance!\r\n",
    "111811": "Isn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?\r\n\r\nI would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\r\n",
    "111816": "[quote=Julian de Wit;111811]\r\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?\r\n\r\nI would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\r\n[/quote]\r\n\r\nExactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?",
    "112467": "I was not that bad at this competition (around 0.33) but did not upload a new model, because I was occupied with other things, so now I am out of the leaderboard. That is making me a bit sad and it was not that clear to me, that this would happen. In the description it was only written, that they would not be recognised for \"prizes\" (position 1, 2 and 3), but not that they would not be recognised in the leaderboard. \r\nThat was a bad communication. \r\n\r\nAlso the statement of Ryan is very right. ",
    "112016": "I feel like people are talking about two different things.\r\n\r\nThis competition certainly does undervalue a person's final standing in comparison to normal competitions.\r\n\r\nBefore stage 2 this competition had 700+ teams.  Now it is less than 200 teams.\r\n\r\nI am actually surprised it was that high.  Except that people wanted the points, and now position 100 on the final leaderboard is severely undervalued compared to normal Kaggle competitions.",
    "111859": "Not to belabor the point, but there is a certain irony to the two-stage LB setup: The harder the problem, the less impressive your position on the final LB appears to be (since more people drop out).",
    "111840": "[quote=inversion;111816]\r\n\r\n[quote=Julian de Wit;111811]\r\nIsn't it possible to add the pre-existing teams at the bottom of the LB with a score of 1.0 or something like that ?\r\n\r\nI would think that everyone that did the effort to post their model and do a submission deserves their badges and the extra points.\r\n[/quote]\r\n\r\nExactly. The 2-stage approach is effectively the same as stripping a normal-contest leader board of anyone who didn't submit during the last week. Why should those who worked to the end get less recognition because a large number of people stopped trying?\r\n\r\n[/quote]\r\n\r\nI could have not expressed it better. It is exactly like if in regular-contests kaggle would not consider people who do not submit during the last week... and Kaggle takes into account even people who only make a single dummy submission.\r\nI think considering in the final leaderboard everyone who participated in the first stage would make this contest more consistent with other ones in terms of kaggle points. \r\n",
    "111787": "Looks like the leadearboard was calculated only for people who did submit in the second round. \r\n\r\nThis somewhat doesn't feel right to me due to the very strong self-selction (don't remember the number, but roughly only about 10% of people did submit in the second round) - could admins specify this in advance for future contests ?",
    "112949": "I totally agree with the previous posts too. For us, we believe we had  a consistent and stable public and private LB, and the winning is quite significant instead of noise. Now with this meaningless  new public LB, it gives a impression that the winning has a lot of noise and it included a lot of luck.",
    "113000": "I would like to join in expressing my disappointment with this issue. I spent many hours on the competition and had a respectable top 10% score on the public leaderboard. I made very few submissions so clearly my success was not due to overfitting. I had very limited time to work on the project during the one week of the 2nd stage so I was not able to reproduce a high scoring submission. \r\n\r\nAt the very least, it would be nice if the original public leaderboard was made available again so I can brag to my mommy. ",
    "113048": "Thanks for the thoughtful comments. \r\n\r\nFirst off, as always, we will not make retrospective changes to how we handle past competitions (including this one). When issues like this come up we use it as an opportunity to evaluate how we might improve in the future.\r\n\r\nInternally our debate focused on three issues:  \r\n\r\n 1. recognition for those who completed stage one but not stage two \r\n 2.  achievements and how the competition appears on profiles (10th out of 750 looks more impressive than 10th out of 172)  \r\n 3.   how points are handled  \r\n \r\n**1. recognition for those who completed stage one but not stage two**\r\n\r\nWe need to view the stage one leaderboard as having no weight: if it gets a weight, we incentivize overfitting or hand labeling for stage one.\r\n\r\n**2. achievements and how the competition appears on profiles**\r\n\r\nIf we did what Julian suggests and add stage one participants to the bottom of the stage two leaderboard, we undermine our rankings by making it very easy for somebody to get an impressive-looking top 25% achievement by finishing 187th out of 750 with a naive submission. \r\n\r\n**3.   how points are handled**  \r\n\r\nThe one change we will make in future is the way points are handled. We will add a multiplier to the number of points for a two-stage competition. We have not settled on a formula for doing this yet, but commit to communicating it clearly in the rules of the next two-stage competition.\r\n\r\nThese are difficult issues but we think this approach strikes the best balance between competing considerations. ",
    "111889": "The public leaderboard only count 1% of the data, which is extremely low. The scores are so screwed, and it is quite meaningless.\r\n\r\nGiven that the competition is over, the public leaderboard should at least count more, like 50% as in \"Santander Customer Satisfaction\".",
    "111792": "Hey Alchemist,\r\n\r\nSorry to hear that this wasn't clear. This was the intent of this bullet point in the rules:\r\n\r\n - This is a two-stage competition. The final results of the competition will be determined by the test dataset released for the second stage.\r\n\r\nWe don't keep two leaderboards because it'd be confusing (also, the first stage was a 100% public leaderboard and could therefore be gamed). We don't have a mixed stage one + stage two leaderboard because there's no right way to order teams.\r\n\r\nTwo stage competitions are rare on Kaggle, but we'll do our best to message this more clearly in the future. Thanks for the feedback.",
    "112973": "@yuanfang it saves the result to file after it finishes all cases. I'm traveling in Peru and so did not have time to check bbs \r\n",
    "112974": "",
    "112969": ""
  }
}