{
  "id": 19445,
  "title": "Why only 6 public rows in the test set ?",
  "url": "/competitions/second-annual-data-science-bowl/discussion/19445",
  "author_name": "MichelJ",
  "post_date": "2016-03-11T09:00:51.793000",
  "votes": -3,
  "comment_count": 19,
  "views": 2167,
  "content": "<p>In my view, having so few public rows in the test set makes the public leaderboard pretty much useless during this final week.</p>\n\n<p>Witness my current (most likely grossly overestimated) standing on the LB: as I was ranked around 80th just before the upload deadline, I did not upload my model since I intended to keep modifying it anyway. I was then surprised by the ranking I got on the new LB with the same model as before, and that I have been able to move up to 2nd place just by loosely tweaking one coefficient.</p>\n\n<p>Arguably, having a solid validation step makes this unimportant. I guess I need to improve the validation step in my procedure</p>\n\n<p>I think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.</p>",
  "messages": [
    {
      "id": 111101,
      "postDate": "2016-03-11T16:19:07.667Z",
      "content": "<p>I'm glad that we are still at #1 even though it has only 6 (or 5?) cases for the current LB.. :) and it has very little indication about the final standing. I think it is great that the admin only uses 6 cases for the current LB to avoid incentive people to change their code.</p>",
      "rawMarkdown": "I'm glad that we are still at #1 even though it has only 6 (or 5?) cases for the current LB.. :) and it has very little indication about the final standing. I think it is great that the admin only uses 6 cases for the current LB to avoid incentive people to change their code.",
      "votes": 6
    },
    {
      "id": 111120,
      "postDate": "2016-03-11T19:33:55.500Z",
      "content": "<p>[quote=Wei Dong;111115]</p>\n\n<p>I totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.</p>\n\n<p>[/quote]</p>\n\n<p>It's possible in many competitions to hack the LB. (Kaggle does a good job of making it difficult, but it can still happen.)  But I think you are right - this contest has more opportunity to do so, due to the small data set.</p>\n\n<p>I don't fault Kaggle for this. I'd rather have opportunities like this to work on interesting, challenging, and worth while problems - with perhaps ambiguity in final LB positions (except for the money spots, of course), than to have sterile contests that are hack-proof.</p>\n\n<p><strong>Edit:</strong> Further, I don't think it is Kaggle's responsibility to ensure that final LB position can be adequately used as a course grade.</p>",
      "rawMarkdown": "[quote=Wei Dong;111115]\r\n\r\nI totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.\r\n\r\n[/quote]\r\n\r\nIt's possible in many competitions to hack the LB. (Kaggle does a good job of making it difficult, but it can still happen.)  But I think you are right - this contest has more opportunity to do so, due to the small data set.\r\n\r\nI don't fault Kaggle for this. I'd rather have opportunities like this to work on interesting, challenging, and worth while problems - with perhaps ambiguity in final LB positions (except for the money spots, of course), than to have sterile contests that are hack-proof.\r\n\r\n**Edit:** Further, I don't think it is Kaggle's responsibility to ensure that final LB position can be adequately used as a course grade.",
      "votes": 3
    },
    {
      "id": 111087,
      "postDate": "2016-03-11T11:50:26.823Z",
      "content": "<p>[quote=MichelJ;111077]</p>\n\n<p>I think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.</p>\n\n<p>[/quote]</p>\n\n<p>That's exactly why it is only 1% - to minimize any advantage to teams that continue to modify their code after the first deadline.</p>",
      "rawMarkdown": "[quote=MichelJ;111077]\r\n\r\nI think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.\r\n\r\n[/quote]\r\n\r\nThat's exactly why it is only 1% - to minimize any advantage to teams that continue to modify their code after the first deadline.\r\n",
      "votes": 4
    },
    {
      "id": 111122,
      "postDate": "2016-03-11T19:40:23.543Z",
      "content": "<p>It's quite possible to manually label the test set, given the small number of cases. That's why, I presume, the competition was held in a two stage setup in the first place.</p>\n\n<p>You don't need to manually calculate the exact targets for every case. Just doing so for the several cases where your model is most erroneous could provide improvement, as Wei Dong stated.</p>\n\n<p>You could also just label the center of the heart, and do it for the whole test set in a relatively quick manner.</p>\n\n<p>In any case, modifying your submitted model will render your team ineligible for the monetary prize. I agree with Inversion here, that being ranked in top 3 with modified code would likely produce some wry grins from fellow Kagglers.</p>\n\n<p>But still, it's not against the rules, as far as I can tell. This means the final leaderboard standings will not exactly be fair, since those who choose modify their models upon the test set will have a clear advantage. They will not be exposed, unless they make it to a prize eligible position. I agree with Wei Dong and Yuanfang Guan on this.</p>\n\n<p>Personally, I don't have any expectations of ranking high, but I still won't bother modifying my uploaded model.</p>",
      "rawMarkdown": "It's quite possible to manually label the test set, given the small number of cases. That's why, I presume, the competition was held in a two stage setup in the first place.\r\n\r\nYou don't need to manually calculate the exact targets for every case. Just doing so for the several cases where your model is most erroneous could provide improvement, as Wei Dong stated.\r\n\r\nYou could also just label the center of the heart, and do it for the whole test set in a relatively quick manner.\r\n\r\nIn any case, modifying your submitted model will render your team ineligible for the monetary prize. I agree with Inversion here, that being ranked in top 3 with modified code would likely produce some wry grins from fellow Kagglers.\r\n\r\nBut still, it's not against the rules, as far as I can tell. This means the final leaderboard standings will not exactly be fair, since those who choose modify their models upon the test set will have a clear advantage. They will not be exposed, unless they make it to a prize eligible position. I agree with Wei Dong and Yuanfang Guan on this.\r\n\r\nPersonally, I don't have any expectations of ranking high, but I still won't bother modifying my uploaded model.",
      "votes": 3
    },
    {
      "id": 111125,
      "postDate": "2016-03-11T20:20:21.413Z",
      "content": "<p>Inversion,</p>\n\n<p>Good point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.</p>\n\n<p>Barisumog,</p>\n\n<p>Good point about hand labeling the center of the heart.  I did notice that our worst cases are caused by erroneously predicting the center of the heart.   Other than a couple of cases where the predicted center fall into the stomach which happens to be about the size of the heart, most such cases leads to prediction of very big or very small volumes.</p>",
      "rawMarkdown": "Inversion,\r\n\r\nGood point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.\r\n\r\nBarisumog,\r\n\r\nGood point about hand labeling the center of the heart.  I did notice that our worst cases are caused by erroneously predicting the center of the heart.   Other than a couple of cases where the predicted center fall into the stomach which happens to be about the size of the heart, most such cases leads to prediction of very big or very small volumes.",
      "votes": 1
    },
    {
      "id": 111127,
      "postDate": "2016-03-11T20:38:12.823Z",
      "content": "<p>[quote=Wei Dong;111125]</p>\n\n<p>Good point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.</p>\n\n<p>[/quote]</p>\n\n<p>I think it's good to discussion these things openly. It makes for a healthier community. The only time I don't think it is good is when people beat to death their hobby horse. (Which you have not!)</p>",
      "rawMarkdown": "[quote=Wei Dong;111125]\r\n\r\nGood point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.\r\n\r\n[/quote]\r\n\r\nI think it's good to discussion these things openly. It makes for a healthier community. The only time I don't think it is good is when people beat to death their hobby horse. (Which you have not!)\r\n",
      "votes": 2
    },
    {
      "id": 111119,
      "postDate": "2016-03-11T19:27:20.310Z",
      "content": "<p>[quote=Paul Jurczak;111114]</p>\n\n<p>I'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?</p>\n\n<p>[/quote]</p>\n\n<p>Actually, a strict reading of the rules suggests this is only a problem if you're collecting a prize. (I may be mistaken though.)</p>",
      "rawMarkdown": "[quote=Paul Jurczak;111114]\r\n\r\nI'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?\r\n\r\n[/quote]\r\n\r\nActually, a strict reading of the rules suggests this is only a problem if you're collecting a prize. (I may be mistaken though.)",
      "votes": 2
    },
    {
      "id": 111114,
      "postDate": "2016-03-11T18:53:11.653Z",
      "content": "<p>I'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?</p>",
      "rawMarkdown": "I'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?",
      "votes": 2
    },
    {
      "id": 111108,
      "postDate": "2016-03-11T18:32:30.687Z",
      "content": "<p>Wei,</p>\n\n<p>You can't tune your parameters  correctly based on only 6 records let alone beating anyone. In any case,  if you don't deserve to be in the top 10 list forever,you won't be there. I hope you understand that there will be a reshuffinge in the rankings after 3 days.  </p>",
      "rawMarkdown": "Wei,\r\n\r\nYou can't tune your parameters  correctly based on only 6 records let alone beating anyone. In any case,  if you don't deserve to be in the top 10 list forever,you won't be there. I hope you understand that there will be a reshuffinge in the rankings after 3 days.  "
    },
    {
      "id": 111113,
      "postDate": "2016-03-11T18:45:30.023Z",
      "content": "<p>[quote=Wei Dong;111111]</p>\n\n<p>I'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.</p>\n\n<p>[/quote]</p>\n\n<p>To my point above, any team that does this runs the risk of having to decline prize money (and have questionable reputation). </p>\n\n<p>It is far better to keep a good reputation on kaggle, than to place a few places higher.</p>",
      "rawMarkdown": "[quote=Wei Dong;111111]\r\n\r\nI'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.\r\n\r\n[/quote]\r\n\r\nTo my point above, any team that does this runs the risk of having to decline prize money (and have questionable reputation). \r\n\r\nIt is far better to keep a good reputation on kaggle, than to place a few places higher.",
      "votes": 1
    },
    {
      "id": 111112,
      "postDate": "2016-03-11T18:42:40.493Z",
      "content": "<p>[quote=Wei Dong;111107]\nThe question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board? \n[/quote]</p>\n\n<p>If a team has been found to have violated the rules, they will be removed from the LB at a minimum.</p>\n\n<p>If a team is in the money, and declines the prize (because they know their code doesn't follow the rules), I don't think they will get removed from the LB, but, of course, everyone will know what happened, which will result in a reputation problem.</p>",
      "rawMarkdown": "[quote=Wei Dong;111107]\r\nThe question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board? \r\n[/quote]\r\n\r\nIf a team has been found to have violated the rules, they will be removed from the LB at a minimum.\r\n\r\nIf a team is in the money, and declines the prize (because they know their code doesn't follow the rules), I don't think they will get removed from the LB, but, of course, everyone will know what happened, which will result in a reputation problem.",
      "votes": 1
    },
    {
      "id": 111111,
      "postDate": "2016-03-11T18:42:08.827Z",
      "content": "<p>I'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.</p>",
      "rawMarkdown": "I'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.",
      "votes": -1
    },
    {
      "id": 111116,
      "postDate": "2016-03-11T19:03:46.337Z",
      "content": "<p>[quote=Yuanfang Guan;111110]</p>\n\n<p>mohd- we will tune nothing -- i believe the majority of the participants will NOT be that low; and of course we understand the reshuffling in the end.</p>\n\n<p>but do you understand that this is image data, now the gold standard answer is on everyone's hand. one doesn't need to tune against any leaderboard anymore, they can just tune against the gold standard.</p>\n\n<p>[/quote]</p>\n\n<p>When you say that 'gold standard answer is on everyone's hand' are you suggesting that one \ncan manually figure out the results using doctors method.  It will take 10&#215;440 /60=74hrs for all cases And if someone has figured out a way to manually calculate it in less time, I believe the method should be made public after the deadline.</p>",
      "rawMarkdown": "[quote=Yuanfang Guan;111110]\r\n\r\nmohd- we will tune nothing -- i believe the majority of the participants will NOT be that low; and of course we understand the reshuffling in the end.\r\n\r\nbut do you understand that this is image data, now the gold standard answer is on everyone's hand. one doesn't need to tune against any leaderboard anymore, they can just tune against the gold standard.\r\n\r\n[/quote]\r\n\r\nWhen you say that 'gold standard answer is on everyone's hand' are you suggesting that one \r\ncan manually figure out the results using doctors method.  It will take 10×440 /60=74hrs for all cases And if someone has figured out a way to manually calculate it in less time, I believe the method should be made public after the deadline.",
      "votes": -2
    },
    {
      "id": 111115,
      "postDate": "2016-03-11T18:56:16.693Z",
      "content": "<p>Inversion,</p>\n\n<p>I totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.</p>\n\n<p>Paul,</p>\n\n<p>It's not against the rule to keep working on your submissions; the organizer has made this very clear.  But since it's likely that it will no longer be a fair leader board, I don't think such an effort very meaningful either.</p>",
      "rawMarkdown": "Inversion,\r\n\r\nI totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.\r\n\r\nPaul,\r\n\r\nIt's not against the rule to keep working on your submissions; the organizer has made this very clear.  But since it's likely that it will no longer be a fair leader board, I don't think such an effort very meaningful either.",
      "votes": -1
    },
    {
      "id": 111107,
      "postDate": "2016-03-11T18:16:17.650Z",
      "content": "<p>We are now facing the dilemma of whether to take the small chance for money and stick to the honest submissions, or to tune against the test set for the chance of staying in the top 10 list in the front page for ever.  I'm sure most of the original top teams will stick to the honest submissions, which are going to remain the same as they are now, so if we take the second direction, we'll have a big chance of beating them in the final leader board.  We could even be rank 2.  The question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board?  What about the other top 10 teams whose results are not reproducible but are not to be investigated because they are not in the money.</p>\n\n<p>I'm just imagining the ways the interesting rules of this particular competition can be attacked.  My suggestion to the organizer, if they cannot at least make sure the top 10 results are reproducible from the model submissions, is not to maintain a leader board after the competition is done at all.   I think a position in the top 10 list means a lot to someone who's seeking a job.  If there's going to be such a list, I hope it is a fair one.  Our submissions are binary-reproducible and we don't intend to change them no matter what.</p>",
      "rawMarkdown": "We are now facing the dilemma of whether to take the small chance for money and stick to the honest submissions, or to tune against the test set for the chance of staying in the top 10 list in the front page for ever.  I'm sure most of the original top teams will stick to the honest submissions, which are going to remain the same as they are now, so if we take the second direction, we'll have a big chance of beating them in the final leader board.  We could even be rank 2.  The question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board?  What about the other top 10 teams whose results are not reproducible but are not to be investigated because they are not in the money.\r\n\r\nI'm just imagining the ways the interesting rules of this particular competition can be attacked.  My suggestion to the organizer, if they cannot at least make sure the top 10 results are reproducible from the model submissions, is not to maintain a leader board after the competition is done at all.   I think a position in the top 10 list means a lot to someone who's seeking a job.  If there's going to be such a list, I hope it is a fair one.  Our submissions are binary-reproducible and we don't intend to change them no matter what.\r\n",
      "votes": -1
    },
    {
      "id": 111077,
      "postDate": "2016-03-11T09:00:51.793Z",
      "content": "<p>In my view, having so few public rows in the test set makes the public leaderboard pretty much useless during this final week.</p>\n\n<p>Witness my current (most likely grossly overestimated) standing on the LB: as I was ranked around 80th just before the upload deadline, I did not upload my model since I intended to keep modifying it anyway. I was then surprised by the ranking I got on the new LB with the same model as before, and that I have been able to move up to 2nd place just by loosely tweaking one coefficient.</p>\n\n<p>Arguably, having a solid validation step makes this unimportant. I guess I need to improve the validation step in my procedure</p>\n\n<p>I think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.</p>",
      "rawMarkdown": "In my view, having so few public rows in the test set makes the public leaderboard pretty much useless during this final week.\r\n\r\nWitness my current (most likely grossly overestimated) standing on the LB: as I was ranked around 80th just before the upload deadline, I did not upload my model since I intended to keep modifying it anyway. I was then surprised by the ranking I got on the new LB with the same model as before, and that I have been able to move up to 2nd place just by loosely tweaking one coefficient.\r\n\r\nArguably, having a solid validation step makes this unimportant. I guess I need to improve the validation step in my procedure\r\n\r\nI think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.",
      "votes": -3
    },
    {
      "id": 111377,
      "postDate": "2016-03-14T09:01:47.487Z",
      "content": "<p>I just noticed this conversation. It is indeed a choice from the Kaggle admins to allow the teams to compete for the leaderboard only. However it is true that this leads to different problems about the real value of the final leaderboard. Another possible problem is: what happens if some of the teams finishing in the top 3 didn't upload any model and don't provide it later? It is true that they are not eligible for a prize, but wouldn't it be a very big loss with respect to the problem of computing automatically the volume of the left ventricule? \nMaybe allowing further model uploads during the 2nd stage of the competition, for competing for the leaderboard only, would have been a solution?</p>",
      "rawMarkdown": "I just noticed this conversation. It is indeed a choice from the Kaggle admins to allow the teams to compete for the leaderboard only. However it is true that this leads to different problems about the real value of the final leaderboard. Another possible problem is: what happens if some of the teams finishing in the top 3 didn't upload any model and don't provide it later? It is true that they are not eligible for a prize, but wouldn't it be a very big loss with respect to the problem of computing automatically the volume of the left ventricule? \r\nMaybe allowing further model uploads during the 2nd stage of the competition, for competing for the leaderboard only, would have been a solution?\r\n\r\n\r\n",
      "votes": -1
    },
    {
      "id": 111146,
      "postDate": "2016-03-12T00:21:34.523Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 111110,
      "postDate": "2016-03-11T18:40:48.527Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 111102,
      "postDate": "2016-03-11T16:40:33.117Z",
      "rawMarkdown": "",
      "votes": -2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 111101,
      "author_name": "woshialex",
      "author_url": "",
      "post_date": "2016-03-11T16:19:07.667000",
      "content": "<p>I'm glad that we are still at #1 even though it has only 6 (or 5?) cases for the current LB.. :) and it has very little indication about the final standing. I think it is great that the admin only uses 6 cases for the current LB to avoid incentive people to change their code.</p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 111120,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-11T19:33:55.500000",
      "content": "<p>[quote=Wei Dong;111115]</p>\n\n<p>I totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.</p>\n\n<p>[/quote]</p>\n\n<p>It's possible in many competitions to hack the LB. (Kaggle does a good job of making it difficult, but it can still happen.)  But I think you are right - this contest has more opportunity to do so, due to the small data set.</p>\n\n<p>I don't fault Kaggle for this. I'd rather have opportunities like this to work on interesting, challenging, and worth while problems - with perhaps ambiguity in final LB positions (except for the money spots, of course), than to have sterile contests that are hack-proof.</p>\n\n<p><strong>Edit:</strong> Further, I don't think it is Kaggle's responsibility to ensure that final LB position can be adequately used as a course grade.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 111087,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-11T11:50:26.823000",
      "content": "<p>[quote=MichelJ;111077]</p>\n\n<p>I think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.</p>\n\n<p>[/quote]</p>\n\n<p>That's exactly why it is only 1% - to minimize any advantage to teams that continue to modify their code after the first deadline.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 111122,
      "author_name": "barisumog",
      "author_url": "",
      "post_date": "2016-03-11T19:40:23.543000",
      "content": "<p>It's quite possible to manually label the test set, given the small number of cases. That's why, I presume, the competition was held in a two stage setup in the first place.</p>\n\n<p>You don't need to manually calculate the exact targets for every case. Just doing so for the several cases where your model is most erroneous could provide improvement, as Wei Dong stated.</p>\n\n<p>You could also just label the center of the heart, and do it for the whole test set in a relatively quick manner.</p>\n\n<p>In any case, modifying your submitted model will render your team ineligible for the monetary prize. I agree with Inversion here, that being ranked in top 3 with modified code would likely produce some wry grins from fellow Kagglers.</p>\n\n<p>But still, it's not against the rules, as far as I can tell. This means the final leaderboard standings will not exactly be fair, since those who choose modify their models upon the test set will have a clear advantage. They will not be exposed, unless they make it to a prize eligible position. I agree with Wei Dong and Yuanfang Guan on this.</p>\n\n<p>Personally, I don't have any expectations of ranking high, but I still won't bother modifying my uploaded model.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 111125,
      "author_name": "Wei Dong",
      "author_url": "",
      "post_date": "2016-03-11T20:20:21.413000",
      "content": "<p>Inversion,</p>\n\n<p>Good point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.</p>\n\n<p>Barisumog,</p>\n\n<p>Good point about hand labeling the center of the heart.  I did notice that our worst cases are caused by erroneously predicting the center of the heart.   Other than a couple of cases where the predicted center fall into the stomach which happens to be about the size of the heart, most such cases leads to prediction of very big or very small volumes.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111127,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-11T20:38:12.823000",
      "content": "<p>[quote=Wei Dong;111125]</p>\n\n<p>Good point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.</p>\n\n<p>[/quote]</p>\n\n<p>I think it's good to discussion these things openly. It makes for a healthier community. The only time I don't think it is good is when people beat to death their hobby horse. (Which you have not!)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 111119,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-11T19:27:20.310000",
      "content": "<p>[quote=Paul Jurczak;111114]</p>\n\n<p>I'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?</p>\n\n<p>[/quote]</p>\n\n<p>Actually, a strict reading of the rules suggests this is only a problem if you're collecting a prize. (I may be mistaken though.)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 111114,
      "author_name": "Paul Jurczak",
      "author_url": "",
      "post_date": "2016-03-11T18:53:11.653000",
      "content": "<p>I'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 111108,
      "author_name": "Mohammad Shadab Alam",
      "author_url": "",
      "post_date": "2016-03-11T18:32:30.687000",
      "content": "<p>Wei,</p>\n\n<p>You can't tune your parameters  correctly based on only 6 records let alone beating anyone. In any case,  if you don't deserve to be in the top 10 list forever,you won't be there. I hope you understand that there will be a reshuffinge in the rankings after 3 days.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111113,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-11T18:45:30.023000",
      "content": "<p>[quote=Wei Dong;111111]</p>\n\n<p>I'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.</p>\n\n<p>[/quote]</p>\n\n<p>To my point above, any team that does this runs the risk of having to decline prize money (and have questionable reputation). </p>\n\n<p>It is far better to keep a good reputation on kaggle, than to place a few places higher.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111112,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "2016-03-11T18:42:40.493000",
      "content": "<p>[quote=Wei Dong;111107]\nThe question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board? \n[/quote]</p>\n\n<p>If a team has been found to have violated the rules, they will be removed from the LB at a minimum.</p>\n\n<p>If a team is in the money, and declines the prize (because they know their code doesn't follow the rules), I don't think they will get removed from the LB, but, of course, everyone will know what happened, which will result in a reputation problem.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 111111,
      "author_name": "Wei Dong",
      "author_url": "",
      "post_date": "2016-03-11T18:42:08.827000",
      "content": "<p>I'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 111116,
      "author_name": "Mohammad Shadab Alam",
      "author_url": "",
      "post_date": "2016-03-11T19:03:46.337000",
      "content": "<p>[quote=Yuanfang Guan;111110]</p>\n\n<p>mohd- we will tune nothing -- i believe the majority of the participants will NOT be that low; and of course we understand the reshuffling in the end.</p>\n\n<p>but do you understand that this is image data, now the gold standard answer is on everyone's hand. one doesn't need to tune against any leaderboard anymore, they can just tune against the gold standard.</p>\n\n<p>[/quote]</p>\n\n<p>When you say that 'gold standard answer is on everyone's hand' are you suggesting that one \ncan manually figure out the results using doctors method.  It will take 10&#215;440 /60=74hrs for all cases And if someone has figured out a way to manually calculate it in less time, I believe the method should be made public after the deadline.</p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 111115,
      "author_name": "Wei Dong",
      "author_url": "",
      "post_date": "2016-03-11T18:56:16.693000",
      "content": "<p>Inversion,</p>\n\n<p>I totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.</p>\n\n<p>Paul,</p>\n\n<p>It's not against the rule to keep working on your submissions; the organizer has made this very clear.  But since it's likely that it will no longer be a fair leader board, I don't think such an effort very meaningful either.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 111107,
      "author_name": "Wei Dong",
      "author_url": "",
      "post_date": "2016-03-11T18:16:17.650000",
      "content": "<p>We are now facing the dilemma of whether to take the small chance for money and stick to the honest submissions, or to tune against the test set for the chance of staying in the top 10 list in the front page for ever.  I'm sure most of the original top teams will stick to the honest submissions, which are going to remain the same as they are now, so if we take the second direction, we'll have a big chance of beating them in the final leader board.  We could even be rank 2.  The question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board?  What about the other top 10 teams whose results are not reproducible but are not to be investigated because they are not in the money.</p>\n\n<p>I'm just imagining the ways the interesting rules of this particular competition can be attacked.  My suggestion to the organizer, if they cannot at least make sure the top 10 results are reproducible from the model submissions, is not to maintain a leader board after the competition is done at all.   I think a position in the top 10 list means a lot to someone who's seeking a job.  If there's going to be such a list, I hope it is a fair one.  Our submissions are binary-reproducible and we don't intend to change them no matter what.</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 111377,
      "author_name": "Vincent L.",
      "author_url": "",
      "post_date": "2016-03-14T09:01:47.487000",
      "content": "<p>I just noticed this conversation. It is indeed a choice from the Kaggle admins to allow the teams to compete for the leaderboard only. However it is true that this leads to different problems about the real value of the final leaderboard. Another possible problem is: what happens if some of the teams finishing in the top 3 didn't upload any model and don't provide it later? It is true that they are not eligible for a prize, but wouldn't it be a very big loss with respect to the problem of computing automatically the volume of the left ventricule? \nMaybe allowing further model uploads during the 2nd stage of the competition, for competing for the leaderboard only, would have been a solution?</p>",
      "votes": -1,
      "replies": []
    },
    {
      "id": 111146,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-12T00:21:34.523000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 111110,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-11T18:40:48.527000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 111102,
      "author_name": "",
      "author_url": "",
      "post_date": "2016-03-11T16:40:33.117000",
      "content": "",
      "votes": -2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "111101": "I'm glad that we are still at #1 even though it has only 6 (or 5?) cases for the current LB.. :) and it has very little indication about the final standing. I think it is great that the admin only uses 6 cases for the current LB to avoid incentive people to change their code.",
    "111120": "[quote=Wei Dong;111115]\r\n\r\nI totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.\r\n\r\n[/quote]\r\n\r\nIt's possible in many competitions to hack the LB. (Kaggle does a good job of making it difficult, but it can still happen.)  But I think you are right - this contest has more opportunity to do so, due to the small data set.\r\n\r\nI don't fault Kaggle for this. I'd rather have opportunities like this to work on interesting, challenging, and worth while problems - with perhaps ambiguity in final LB positions (except for the money spots, of course), than to have sterile contests that are hack-proof.\r\n\r\n**Edit:** Further, I don't think it is Kaggle's responsibility to ensure that final LB position can be adequately used as a course grade.",
    "111087": "[quote=MichelJ;111077]\r\n\r\nI think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.\r\n\r\n[/quote]\r\n\r\nThat's exactly why it is only 1% - to minimize any advantage to teams that continue to modify their code after the first deadline.\r\n",
    "111122": "It's quite possible to manually label the test set, given the small number of cases. That's why, I presume, the competition was held in a two stage setup in the first place.\r\n\r\nYou don't need to manually calculate the exact targets for every case. Just doing so for the several cases where your model is most erroneous could provide improvement, as Wei Dong stated.\r\n\r\nYou could also just label the center of the heart, and do it for the whole test set in a relatively quick manner.\r\n\r\nIn any case, modifying your submitted model will render your team ineligible for the monetary prize. I agree with Inversion here, that being ranked in top 3 with modified code would likely produce some wry grins from fellow Kagglers.\r\n\r\nBut still, it's not against the rules, as far as I can tell. This means the final leaderboard standings will not exactly be fair, since those who choose modify their models upon the test set will have a clear advantage. They will not be exposed, unless they make it to a prize eligible position. I agree with Wei Dong and Yuanfang Guan on this.\r\n\r\nPersonally, I don't have any expectations of ranking high, but I still won't bother modifying my uploaded model.",
    "111125": "Inversion,\r\n\r\nGood point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.\r\n\r\nBarisumog,\r\n\r\nGood point about hand labeling the center of the heart.  I did notice that our worst cases are caused by erroneously predicting the center of the heart.   Other than a couple of cases where the predicted center fall into the stomach which happens to be about the size of the heart, most such cases leads to prediction of very big or very small volumes.",
    "111127": "[quote=Wei Dong;111125]\r\n\r\nGood point about sterile hack-proof contests.  Now I think I'm worrying too much.  I appreciate that the organizer has made this interesting competition available to the community and apologize if I have ruined the fun of those who want to keep working on the problem.\r\n\r\n[/quote]\r\n\r\nI think it's good to discussion these things openly. It makes for a healthier community. The only time I don't think it is good is when people beat to death their hobby horse. (Which you have not!)\r\n",
    "111119": "[quote=Paul Jurczak;111114]\r\n\r\nI'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?\r\n\r\n[/quote]\r\n\r\nActually, a strict reading of the rules suggests this is only a problem if you're collecting a prize. (I may be mistaken though.)",
    "111114": "I'm a bit confused by this discussion. I didn't submit the models, because I've run out of time to get my competitive solution working. So I'm not in it for the money now, but I would like to keep working for the next few days and make submissions of results from improved algorithms. Is this against the rules?",
    "111108": "Wei,\r\n\r\nYou can't tune your parameters  correctly based on only 6 records let alone beating anyone. In any case,  if you don't deserve to be in the top 10 list forever,you won't be there. I hope you understand that there will be a reshuffinge in the rankings after 3 days.  ",
    "111113": "[quote=Wei Dong;111111]\r\n\r\nI'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.\r\n\r\n[/quote]\r\n\r\nTo my point above, any team that does this runs the risk of having to decline prize money (and have questionable reputation). \r\n\r\nIt is far better to keep a good reputation on kaggle, than to place a few places higher.",
    "111112": "[quote=Wei Dong;111107]\r\nThe question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board? \r\n[/quote]\r\n\r\nIf a team has been found to have violated the rules, they will be removed from the LB at a minimum.\r\n\r\nIf a team is in the money, and declines the prize (because they know their code doesn't follow the rules), I don't think they will get removed from the LB, but, of course, everyone will know what happened, which will result in a reputation problem.",
    "111111": "I'm talking about tuning against the 440 testing cases.  A reasonable error model, which unfortunately happens to be an integral part of this competition so everyone already has one, should be able to find the cases where one's model's likely to make the most mistakes.  I think fixing top 20 of those worst cases will make all the difference.",
    "111116": "[quote=Yuanfang Guan;111110]\r\n\r\nmohd- we will tune nothing -- i believe the majority of the participants will NOT be that low; and of course we understand the reshuffling in the end.\r\n\r\nbut do you understand that this is image data, now the gold standard answer is on everyone's hand. one doesn't need to tune against any leaderboard anymore, they can just tune against the gold standard.\r\n\r\n[/quote]\r\n\r\nWhen you say that 'gold standard answer is on everyone's hand' are you suggesting that one \r\ncan manually figure out the results using doctors method.  It will take 10×440 /60=74hrs for all cases And if someone has figured out a way to manually calculate it in less time, I believe the method should be made public after the deadline.",
    "111115": "Inversion,\r\n\r\nI totally agree with you.  But the scope of this competition has gone beyond monetary interest and reputation on Kaggle alone.   There are academic gradings to be made according to the leader board.  Now that the leader board can be hacked, I think it is beneficial to make this fact known to the public.\r\n\r\nPaul,\r\n\r\nIt's not against the rule to keep working on your submissions; the organizer has made this very clear.  But since it's likely that it will no longer be a fair leader board, I don't think such an effort very meaningful either.",
    "111107": "We are now facing the dilemma of whether to take the small chance for money and stick to the honest submissions, or to tune against the test set for the chance of staying in the top 10 list in the front page for ever.  I'm sure most of the original top teams will stick to the honest submissions, which are going to remain the same as they are now, so if we take the second direction, we'll have a big chance of beating them in the final leader board.  We could even be rank 2.  The question is, if the organizer finds our results non-reproducible, are we to be removed from the leader board?  What about the other top 10 teams whose results are not reproducible but are not to be investigated because they are not in the money.\r\n\r\nI'm just imagining the ways the interesting rules of this particular competition can be attacked.  My suggestion to the organizer, if they cannot at least make sure the top 10 results are reproducible from the model submissions, is not to maintain a leader board after the competition is done at all.   I think a position in the top 10 list means a lot to someone who's seeking a job.  If there's going to be such a list, I hope it is a fair one.  Our submissions are binary-reproducible and we don't intend to change them no matter what.\r\n",
    "111077": "In my view, having so few public rows in the test set makes the public leaderboard pretty much useless during this final week.\r\n\r\nWitness my current (most likely grossly overestimated) standing on the LB: as I was ranked around 80th just before the upload deadline, I did not upload my model since I intended to keep modifying it anyway. I was then surprised by the ranking I got on the new LB with the same model as before, and that I have been able to move up to 2nd place just by loosely tweaking one coefficient.\r\n\r\nArguably, having a solid validation step makes this unimportant. I guess I need to improve the validation step in my procedure\r\n\r\nI think that having at least 10% of the test set used for the public LB would have been more useful for people like me, who want to keep trying things out during the final week.",
    "111377": "I just noticed this conversation. It is indeed a choice from the Kaggle admins to allow the teams to compete for the leaderboard only. However it is true that this leads to different problems about the real value of the final leaderboard. Another possible problem is: what happens if some of the teams finishing in the top 3 didn't upload any model and don't provide it later? It is true that they are not eligible for a prize, but wouldn't it be a very big loss with respect to the problem of computing automatically the volume of the left ventricule? \r\nMaybe allowing further model uploads during the 2nd stage of the competition, for competing for the leaderboard only, would have been a solution?\r\n\r\n\r\n",
    "111146": "",
    "111110": "",
    "111102": ""
  }
}