{
  "id": 34733,
  "title": "Admins -Why release stage1 labels and include stage1 images in stage2?",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/34733",
  "author_name": "",
  "post_date": "2017-06-15T01:07:30.333668900Z",
  "votes": 4,
  "comment_count": 17,
  "views": 0,
  "content": "<p>Am I missing something? stage 1 labels are released yet sample_submission_stg2.csv includes stage 1 images, and it seems the leader-board is still being evaluated on stage 1 images only. This will lead to numerous perfect submissions and a messed up leader-board!</p>",
  "messages": [
    {
      "id": "192857",
      "postDate": "06/15/2017 01:07:30",
      "content": "<p>Am I missing something? stage 1 labels are released yet sample_submission_stg2.csv includes stage 1 images, and it seems the leader-board is still being evaluated on stage 1 images only. This will lead to numerous perfect submissions and a messed up leader-board!</p>",
      "rawMarkdown": "Am I missing something? stage 1 labels are released yet sample_submission_stg2.csv includes stage 1 images, and it seems the leader-board is still being evaluated on stage 1 images only. This will lead to numerous perfect submissions and a messed up leader-board!",
      "votes": null
    },
    {
      "id": "192861",
      "postDate": "06/15/2017 01:37:37",
      "content": "<p>I have the same confusion, anyone can explain ?</p>",
      "rawMarkdown": "I have the same confusion, anyone can explain ?",
      "votes": null
    },
    {
      "id": "192863",
      "postDate": "06/15/2017 01:49:11",
      "content": "<p>Hi, I noted that in the leaderboard there is an illustration</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 13% of the test\n  data. The final results will be based on the other 87%, so the final\n  standings may be different.</p>\n</blockquote>\n\n<p>so I think the the public leaderboard is actually evaluated on stage1 images only (approximately 13% - 512/4019)\nbut the final rank will be evaluated by stage2 images.</p>",
      "rawMarkdown": "Hi, I noted that in the leaderboard there is an illustration\n\n&gt; This leaderboard is calculated with approximately 13% of the test\n&gt; data. The final results will be based on the other 87%, so the final\n&gt; standings may be different.\n\nso I think the the public leaderboard is actually evaluated on stage1 images only (approximately 13% - 512/4019)\nbut the final rank will be evaluated by stage2 images.",
      "votes": null
    },
    {
      "id": "192865",
      "postDate": "06/15/2017 01:54:28",
      "content": "<p>yes, but why are they using stage 1 only in stage 2 leader-board, when they have released stage 1 ground truth? it will only lead to a big mess. In addition, if there are any issues generating stage 2 submissions there will be no lb feedback on this as well.</p>",
      "rawMarkdown": "yes, but why are they using stage 1 only in stage 2 leader-board, when they have released stage 1 ground truth? it will only lead to a big mess. In addition, if there are any issues generating stage 2 submissions there will be no lb feedback on this as well.",
      "votes": null
    },
    {
      "id": "192888",
      "postDate": "06/15/2017 03:15:27",
      "content": "<p>Perhaps they think it have little influence. We are expected not to make any scientific alterations of our codes, which is explained in <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32580\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32580</a></p>\n\n<blockquote>\n  <p>We expect you may need to make some “non scientific” alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</p>\n</blockquote>",
      "rawMarkdown": "Perhaps they think it have little influence. We are expected not to make any scientific alterations of our codes, which is explained in https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32580\n\n&gt; We expect you may need to make some “non scientific” alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.",
      "votes": null
    },
    {
      "id": "192900",
      "postDate": "06/15/2017 03:57:37",
      "content": "<p>But they are not using stage 1 labels for stage 2 public score board, the numbers don't match</p>",
      "rawMarkdown": "But they are not using stage 1 labels for stage 2 public score board, the numbers don't match",
      "votes": null
    },
    {
      "id": "192905",
      "postDate": "06/15/2017 04:15:15",
      "content": "<p>It seemed that way earlier, looks like it just took time to reset the lb, if that is the case that is good news... I still don't understand the reasoning behind including stage1 in stage2 submissions, or including it in stage2 lb calculation. Perhaps maybe to encourage more people to submit.</p>",
      "rawMarkdown": "It seemed that way earlier, looks like it just took time to reset the lb, if that is the case that is good news... I still don't understand the reasoning behind including stage1 in stage2 submissions, or including it in stage2 lb calculation. Perhaps maybe to encourage more people to submit.",
      "votes": null
    },
    {
      "id": "193006",
      "postDate": "06/15/2017 10:45:32",
      "content": "<p>This is absolutely necessary be clear on. <strong>Are the stage1 test image set a part of the stage2 test image set?</strong></p>\n\n<p>In that case the strategy will be to submit the known labels to the stage1 part of the full test set, and add the predictions for the remaining ~3500 images. Those predictions can even be based on a retrain of the model on the stage1 test set images.</p>",
      "rawMarkdown": "This is absolutely necessary be clear on. **Are the stage1 test image set a part of the stage2 test image set?**\n\nIn that case the strategy will be to submit the known labels to the stage1 part of the full test set, and add the predictions for the remaining ~3500 images. Those predictions can even be based on a retrain of the model on the stage1 test set images.",
      "votes": null
    },
    {
      "id": "193008",
      "postDate": "06/15/2017 10:52:21",
      "content": "<p>I think that might be a bad idea because you will simply get a LB loss of 0.0 but you will have no idea how good your model will eventually perform on the private LB.</p>",
      "rawMarkdown": "I think that might be a bad idea because you will simply get a LB loss of 0.0 but you will have no idea how good your model will eventually perform on the private LB.",
      "votes": null
    },
    {
      "id": "193009",
      "postDate": "06/15/2017 10:55:02",
      "content": "<p>I think so too. So essentially what this has done is to create a level playing field versus those who may have already had the ground truth for the stage 1 test set by means of LB mining. But it is STILL unfair since they would have had a mountain of time perfecting their models but others will have only a week to do any work on the same level ground.</p>",
      "rawMarkdown": "I think so too. So essentially what this has done is to create a level playing field versus those who may have already had the ground truth for the stage 1 test set by means of LB mining. But it is STILL unfair since they would have had a mountain of time perfecting their models but others will have only a week to do any work on the same level ground.",
      "votes": null
    },
    {
      "id": "193017",
      "postDate": "06/15/2017 11:21:15",
      "content": "<p>How is the final private leaderborad calculated then? If the stage1 test data is part of the stage2 final private leaderboard, then I will post the true labels that's given. (Unless that's against the rules, of course.)</p>\n\n<p>I'm confused. (as everybody else, I guess)</p>",
      "rawMarkdown": "How is the final private leaderborad calculated then? If the stage1 test data is part of the stage2 final private leaderboard, then I will post the true labels that's given. (Unless that's against the rules, of course.)\n\nI'm confused. (as everybody else, I guess)",
      "votes": null
    },
    {
      "id": "193019",
      "postDate": "06/15/2017 11:26:16",
      "content": "<p>No, my guess is that the private LB will be calculated on the other ~3500 images so submitting the true labels will simply make the public LB score 0 while not helping the private LB score at all. </p>",
      "rawMarkdown": "No, my guess is that the private LB will be calculated on the other ~3500 images so submitting the true labels will simply make the public LB score 0 while not helping the private LB score at all.",
      "votes": null
    },
    {
      "id": "193023",
      "postDate": "06/15/2017 11:35:41",
      "content": "<p>Aha! Now I understand how you're interpreting situation. Thanks.</p>",
      "rawMarkdown": "Aha! Now I understand how you're interpreting situation. Thanks.",
      "votes": null
    },
    {
      "id": "193043",
      "postDate": "06/15/2017 13:13:00",
      "content": "<p>As far as leaderboard mining goes, I would suggest that kaggle start adding a little random noise to the public leaderboards when there are small dataset like this one and mining is effective, so you never get an exact score but rather an estimate of your score, this might slow down leaderboard miners and make the current techniques used in leaderboard mining less effective, and maybe make the leaderboard a little more useful.</p>",
      "rawMarkdown": "As far as leaderboard mining goes, I would suggest that kaggle start adding a little random noise to the public leaderboards when there are small dataset like this one and mining is effective, so you never get an exact score but rather an estimate of your score, this might slow down leaderboard miners and make the current techniques used in leaderboard mining less effective, and maybe make the leaderboard a little more useful.",
      "votes": null
    },
    {
      "id": "193049",
      "postDate": "06/15/2017 13:31:00",
      "content": "<p>no wonder so many 0.0 scores, it is based on stage1 labels.</p>",
      "rawMarkdown": "no wonder so many 0.0 scores, it is based on stage1 labels.",
      "votes": null
    },
    {
      "id": "193052",
      "postDate": "06/15/2017 13:42:27",
      "content": "<p>Yes! @bk0000 I think you got the right idea. It says the private LB will be based on the <em>other</em> 87% of the test data. However, this means that there is no value in the leaderboard in this stage of the competition. You can submit what ever score you like.  It will therefor be a better strategy to include the stage1 test set in the training, and then retrain your model with that. When you then submit you have probably overfitted to the stage1 test, and the LB score will be artificial low.</p>",
      "rawMarkdown": "Yes! @bk0000 I think you got the right idea. It says the private LB will be based on the *other* 87% of the test data. However, this means that there is no value in the leaderboard in this stage of the competition. You can submit what ever score you like.  It will therefor be a better strategy to include the stage1 test set in the training, and then retrain your model with that. When you then submit you have probably overfitted to the stage1 test, and the LB score will be artificial low.",
      "votes": null
    },
    {
      "id": "193056",
      "postDate": "06/15/2017 13:56:29",
      "content": "<p>This way the leaderboard is not reflective of the model in addition to creating this confusion. Instead of revealing the test_stg1 labels, Kaggle could try increasing the precision of the logloss scores that are displayed on a competition based on the total submissions possible over the competition period to prevent LB mining? That way the stg1 labels wouldn't need to be revealed and still would be fair to the competitors.</p>",
      "rawMarkdown": "This way the leaderboard is not reflective of the model in addition to creating this confusion. Instead of revealing the test_stg1 labels, Kaggle could try increasing the precision of the logloss scores that are displayed on a competition based on the total submissions possible over the competition period to prevent LB mining? That way the stg1 labels wouldn't need to be revealed and still would be fair to the competitors.",
      "votes": null
    },
    {
      "id": "193123",
      "postDate": "06/15/2017 17:03:15",
      "content": "<p>For all of the two stage competitions the leader board has been horrible. </p>\n\n<p>This is compounded by the fact that the datasets have been horribly under powered in many of their recent competitions.  Being able to have a 30% larger training set is a big deal.  </p>\n\n<p>They should just release a training set + validation set for first stage and provide labels to both.  Sure the leaderboard will be worthless, but it is already worthless as things stand right now.</p>",
      "rawMarkdown": "For all of the two stage competitions the leader board has been horrible. \n\nThis is compounded by the fact that the datasets have been horribly under powered in many of their recent competitions.  Being able to have a 30% larger training set is a big deal.  \n\nThey should just release a training set + validation set for first stage and provide labels to both.  Sure the leaderboard will be worthless, but it is already worthless as things stand right now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 192861,
      "author_name": "yangyu94",
      "author_url": "",
      "post_date": "06/15/2017 01:37:37",
      "content": "<p>I have the same confusion, anyone can explain ?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192863,
      "author_name": "yangyu94",
      "author_url": "",
      "post_date": "06/15/2017 01:49:11",
      "content": "<p>Hi, I noted that in the leaderboard there is an illustration</p>\n\n<blockquote>\n  <p>This leaderboard is calculated with approximately 13% of the test\n  data. The final results will be based on the other 87%, so the final\n  standings may be different.</p>\n</blockquote>\n\n<p>so I think the the public leaderboard is actually evaluated on stage1 images only (approximately 13% - 512/4019)\nbut the final rank will be evaluated by stage2 images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 193009,
          "author_name": "bk0000",
          "author_url": "",
          "post_date": "06/15/2017 10:55:02",
          "content": "<p>I think so too. So essentially what this has done is to create a level playing field versus those who may have already had the ground truth for the stage 1 test set by means of LB mining. But it is STILL unfair since they would have had a mountain of time perfecting their models but others will have only a week to do any work on the same level ground.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 192865,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "06/15/2017 01:54:28",
      "content": "<p>yes, but why are they using stage 1 only in stage 2 leader-board, when they have released stage 1 ground truth? it will only lead to a big mess. In addition, if there are any issues generating stage 2 submissions there will be no lb feedback on this as well.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192888,
      "author_name": "yangyu94",
      "author_url": "",
      "post_date": "06/15/2017 03:15:27",
      "content": "<p>Perhaps they think it have little influence. We are expected not to make any scientific alterations of our codes, which is explained in <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32580\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32580</a></p>\n\n<blockquote>\n  <p>We expect you may need to make some “non scientific” alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192900,
      "author_name": "kubilai",
      "author_url": "",
      "post_date": "06/15/2017 03:57:37",
      "content": "<p>But they are not using stage 1 labels for stage 2 public score board, the numbers don't match</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192905,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "06/15/2017 04:15:15",
      "content": "<p>It seemed that way earlier, looks like it just took time to reset the lb, if that is the case that is good news... I still don't understand the reasoning behind including stage1 in stage2 submissions, or including it in stage2 lb calculation. Perhaps maybe to encourage more people to submit.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193006,
      "author_name": "oysteijo",
      "author_url": "",
      "post_date": "06/15/2017 10:45:32",
      "content": "<p>This is absolutely necessary be clear on. <strong>Are the stage1 test image set a part of the stage2 test image set?</strong></p>\n\n<p>In that case the strategy will be to submit the known labels to the stage1 part of the full test set, and add the predictions for the remaining ~3500 images. Those predictions can even be based on a retrain of the model on the stage1 test set images.</p>",
      "votes": null,
      "replies": [
        {
          "id": 193008,
          "author_name": "bk0000",
          "author_url": "",
          "post_date": "06/15/2017 10:52:21",
          "content": "<p>I think that might be a bad idea because you will simply get a LB loss of 0.0 but you will have no idea how good your model will eventually perform on the private LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193017,
          "author_name": "oysteijo",
          "author_url": "",
          "post_date": "06/15/2017 11:21:15",
          "content": "<p>How is the final private leaderborad calculated then? If the stage1 test data is part of the stage2 final private leaderboard, then I will post the true labels that's given. (Unless that's against the rules, of course.)</p>\n\n<p>I'm confused. (as everybody else, I guess)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193019,
          "author_name": "bk0000",
          "author_url": "",
          "post_date": "06/15/2017 11:26:16",
          "content": "<p>No, my guess is that the private LB will be calculated on the other ~3500 images so submitting the true labels will simply make the public LB score 0 while not helping the private LB score at all. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193023,
          "author_name": "oysteijo",
          "author_url": "",
          "post_date": "06/15/2017 11:35:41",
          "content": "<p>Aha! Now I understand how you're interpreting situation. Thanks.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193052,
          "author_name": "oysteijo",
          "author_url": "",
          "post_date": "06/15/2017 13:42:27",
          "content": "<p>Yes! @bk0000 I think you got the right idea. It says the private LB will be based on the <em>other</em> 87% of the test data. However, this means that there is no value in the leaderboard in this stage of the competition. You can submit what ever score you like.  It will therefor be a better strategy to include the stage1 test set in the training, and then retrain your model with that. When you then submit you have probably overfitted to the stage1 test, and the LB score will be artificial low.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 193043,
      "author_name": "godaibo",
      "author_url": "",
      "post_date": "06/15/2017 13:13:00",
      "content": "<p>As far as leaderboard mining goes, I would suggest that kaggle start adding a little random noise to the public leaderboards when there are small dataset like this one and mining is effective, so you never get an exact score but rather an estimate of your score, this might slow down leaderboard miners and make the current techniques used in leaderboard mining less effective, and maybe make the leaderboard a little more useful.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193049,
      "author_name": "legion007",
      "author_url": "",
      "post_date": "06/15/2017 13:31:00",
      "content": "<p>no wonder so many 0.0 scores, it is based on stage1 labels.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193056,
      "author_name": "bvineeth007",
      "author_url": "",
      "post_date": "06/15/2017 13:56:29",
      "content": "<p>This way the leaderboard is not reflective of the model in addition to creating this confusion. Instead of revealing the test_stg1 labels, Kaggle could try increasing the precision of the logloss scores that are displayed on a competition based on the total submissions possible over the competition period to prevent LB mining? That way the stg1 labels wouldn't need to be revealed and still would be fair to the competitors.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 193123,
      "author_name": "iamthep",
      "author_url": "",
      "post_date": "06/15/2017 17:03:15",
      "content": "<p>For all of the two stage competitions the leader board has been horrible. </p>\n\n<p>This is compounded by the fact that the datasets have been horribly under powered in many of their recent competitions.  Being able to have a 30% larger training set is a big deal.  </p>\n\n<p>They should just release a training set + validation set for first stage and provide labels to both.  Sure the leaderboard will be worthless, but it is already worthless as things stand right now.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "192857": "Am I missing something? stage 1 labels are released yet sample_submission_stg2.csv includes stage 1 images, and it seems the leader-board is still being evaluated on stage 1 images only. This will lead to numerous perfect submissions and a messed up leader-board!",
    "192861": "I have the same confusion, anyone can explain ?",
    "192863": "Hi, I noted that in the leaderboard there is an illustration\n\n&gt; This leaderboard is calculated with approximately 13% of the test\n&gt; data. The final results will be based on the other 87%, so the final\n&gt; standings may be different.\n\nso I think the the public leaderboard is actually evaluated on stage1 images only (approximately 13% - 512/4019)\nbut the final rank will be evaluated by stage2 images.",
    "192865": "yes, but why are they using stage 1 only in stage 2 leader-board, when they have released stage 1 ground truth? it will only lead to a big mess. In addition, if there are any issues generating stage 2 submissions there will be no lb feedback on this as well.",
    "192888": "Perhaps they think it have little influence. We are expected not to make any scientific alterations of our codes, which is explained in https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/32580\n\n&gt; We expect you may need to make some “non scientific” alterations, such as changes to path names, in order to create your submissions for the second stage. You are allowed to re-train your model (including the stage one data), but your code should not change. You should not be doing any hyper parameter tuning in the second stage. Parameter tuning is permitted as long as it is fully automated.",
    "192900": "But they are not using stage 1 labels for stage 2 public score board, the numbers don't match",
    "192905": "It seemed that way earlier, looks like it just took time to reset the lb, if that is the case that is good news... I still don't understand the reasoning behind including stage1 in stage2 submissions, or including it in stage2 lb calculation. Perhaps maybe to encourage more people to submit.",
    "193006": "This is absolutely necessary be clear on. **Are the stage1 test image set a part of the stage2 test image set?**\n\nIn that case the strategy will be to submit the known labels to the stage1 part of the full test set, and add the predictions for the remaining ~3500 images. Those predictions can even be based on a retrain of the model on the stage1 test set images.",
    "193008": "I think that might be a bad idea because you will simply get a LB loss of 0.0 but you will have no idea how good your model will eventually perform on the private LB.",
    "193009": "I think so too. So essentially what this has done is to create a level playing field versus those who may have already had the ground truth for the stage 1 test set by means of LB mining. But it is STILL unfair since they would have had a mountain of time perfecting their models but others will have only a week to do any work on the same level ground.",
    "193017": "How is the final private leaderborad calculated then? If the stage1 test data is part of the stage2 final private leaderboard, then I will post the true labels that's given. (Unless that's against the rules, of course.)\n\nI'm confused. (as everybody else, I guess)",
    "193019": "No, my guess is that the private LB will be calculated on the other ~3500 images so submitting the true labels will simply make the public LB score 0 while not helping the private LB score at all.",
    "193023": "Aha! Now I understand how you're interpreting situation. Thanks.",
    "193043": "As far as leaderboard mining goes, I would suggest that kaggle start adding a little random noise to the public leaderboards when there are small dataset like this one and mining is effective, so you never get an exact score but rather an estimate of your score, this might slow down leaderboard miners and make the current techniques used in leaderboard mining less effective, and maybe make the leaderboard a little more useful.",
    "193049": "no wonder so many 0.0 scores, it is based on stage1 labels.",
    "193052": "Yes! @bk0000 I think you got the right idea. It says the private LB will be based on the *other* 87% of the test data. However, this means that there is no value in the leaderboard in this stage of the competition. You can submit what ever score you like.  It will therefor be a better strategy to include the stage1 test set in the training, and then retrain your model with that. When you then submit you have probably overfitted to the stage1 test, and the LB score will be artificial low.",
    "193056": "This way the leaderboard is not reflective of the model in addition to creating this confusion. Instead of revealing the test_stg1 labels, Kaggle could try increasing the precision of the logloss scores that are displayed on a competition based on the total submissions possible over the competition period to prevent LB mining? That way the stg1 labels wouldn't need to be revealed and still would be fair to the competitors.",
    "193123": "For all of the two stage competitions the leader board has been horrible. \n\nThis is compounded by the fact that the datasets have been horribly under powered in many of their recent competitions.  Being able to have a 30% larger training set is a big deal.  \n\nThey should just release a training set + validation set for first stage and provide labels to both.  Sure the leaderboard will be worthless, but it is already worthless as things stand right now."
  },
  "source": "meta"
}