{
  "id": 35009,
  "title": "Am I missing something? (stage2 LB)",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/35009",
  "author_name": "",
  "post_date": "2017-06-20T13:27:47.399739700Z",
  "votes": 1,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi all,  Im suddenly having a big  doubt about stage2 rules.  </p>\n\n<p>The reason, I was expecting all teams to have 0 error on the leaderboard, as we already have the labels of stage 1 and -I supposed- we are all expected just to merge them with the rest of the submission... but lots of teams apparently not including those labels, or not yet.</p>\n\n<p>So, am I missing something? Even if it is obvious  maybe the answer to this question will help others that, like me,  may have misunderstood the expected submission.  (There is a group of 0 errors that maybe understood the same). It would be a pity to self destruct so much work because of such a silly (but big) mistake, that we would have still time to correct.</p>\n\n<p>So, <strong>is it merging test labels the correct thing to do, it isn't or it just doesn't matter</strong> cause we will not be evaluated on those observations?</p>",
  "messages": [
    {
      "id": "194417",
      "postDate": "06/20/2017 13:27:47",
      "content": "<p>Hi all,  Im suddenly having a big  doubt about stage2 rules.  </p>\n\n<p>The reason, I was expecting all teams to have 0 error on the leaderboard, as we already have the labels of stage 1 and -I supposed- we are all expected just to merge them with the rest of the submission... but lots of teams apparently not including those labels, or not yet.</p>\n\n<p>So, am I missing something? Even if it is obvious  maybe the answer to this question will help others that, like me,  may have misunderstood the expected submission.  (There is a group of 0 errors that maybe understood the same). It would be a pity to self destruct so much work because of such a silly (but big) mistake, that we would have still time to correct.</p>\n\n<p>So, <strong>is it merging test labels the correct thing to do, it isn't or it just doesn't matter</strong> cause we will not be evaluated on those observations?</p>",
      "rawMarkdown": "Hi all,  Im suddenly having a big  doubt about stage2 rules.  \n\nThe reason, I was expecting all teams to have 0 error on the leaderboard, as we already have the labels of stage 1 and -I supposed- we are all expected just to merge them with the rest of the submission... but lots of teams apparently not including those labels, or not yet.\n\nSo, am I missing something? Even if it is obvious  maybe the answer to this question will help others that, like me,  may have misunderstood the expected submission.  (There is a group of 0 errors that maybe understood the same). It would be a pity to self destruct so much work because of such a silly (but big) mistake, that we would have still time to correct.\n\nSo, **is it merging test labels the correct thing to do, it isn't or it just doesn't matter** cause we will not be evaluated on those observations?",
      "votes": null
    },
    {
      "id": "194419",
      "postDate": "06/20/2017 13:33:17",
      "content": "<p>You can incorporate test images from 1st stage test data but you can use them as validation set or some of the images may be in your validation set when you perform the split, thus not enabling the model to see their labels in training set, in which case you won't have 0 error on LB, as the model will still have to predict them in 'normal' way, not simply using the labels, on which it was trained.</p>",
      "rawMarkdown": "You can incorporate test images from 1st stage test data but you can use them as validation set or some of the images may be in your validation set when you perform the split, thus not enabling the model to see their labels in training set, in which case you won't have 0 error on LB, as the model will still have to predict them in 'normal' way, not simply using the labels, on which it was trained.",
      "votes": null
    },
    {
      "id": "194421",
      "postDate": "06/20/2017 13:45:03",
      "content": "<p>Ouch...  Thank you @Wojtek Rosinski. From your answer I take that we are expected to re-predict stage1test, not just to merge labels... and then of course error can not be 0. It probably will not be 0 even if all test is used to train, unles your train error is 0. </p>\n\n<p>So... good that I can correct that, I wonder if those 60  zero error submissions on the leaderboard are misunderstanding the same I was...  ( I hope they will read this comment cause to be disqualified for such a stupid mistake is a pity for anyone that has worked hard on a comp)</p>\n\n<p>So, again, thank you Wojtek :-)</p>",
      "rawMarkdown": "Ouch...  Thank you @Wojtek Rosinski. From your answer I take that we are expected to re-predict stage1test, not just to merge labels... and then of course error can not be 0. It probably will not be 0 even if all test is used to train, unles your train error is 0. \n\nSo... good that I can correct that, I wonder if those 60  zero error submissions on the leaderboard are misunderstanding the same I was...  ( I hope they will read this comment cause to be disqualified for such a stupid mistake is a pity for anyone that has worked hard on a comp)\n\nSo, again, thank you Wojtek :-)",
      "votes": null
    },
    {
      "id": "194425",
      "postDate": "06/20/2017 14:08:00",
      "content": "<p>We are not expected to predict them again, I think. Whatever we do with now-released labels should, it should not matter for final standings, so you can either predict stage 1 images with your model or just merge labels into submission.</p>\n\n<p>What will mater in the end is your predictions on stage 2. Incorporation of stage 1 may just be used as a sanity check to see if your submission works properly and there's nothing wrong with your model.</p>\n\n<p>At least this is how I understand this situation. If I'm mistaken anywhere, then the admins should clarify this.</p>",
      "rawMarkdown": "We are not expected to predict them again, I think. Whatever we do with now-released labels should, it should not matter for final standings, so you can either predict stage 1 images with your model or just merge labels into submission.\n\nWhat will mater in the end is your predictions on stage 2. Incorporation of stage 1 may just be used as a sanity check to see if your submission works properly and there's nothing wrong with your model.\n\nAt least this is how I understand this situation. If I'm mistaken anywhere, then the admins should clarify this.",
      "votes": null
    },
    {
      "id": "194427",
      "postDate": "06/20/2017 14:14:50",
      "content": "<p>Ah, ok, it all makes sense now. If it just doesnt matter people are doing whatever is more efficient with their submissions, and thus the different LB errors. </p>\n\n<p>So I  guess we dont have to worry about that (and yes, if this is mistaken admins probably should say something seeing this LB )</p>\n\n<p>Thank you again, Wojtek!</p>",
      "rawMarkdown": "Ah, ok, it all makes sense now. If it just doesnt matter people are doing whatever is more efficient with their submissions, and thus the different LB errors. \n\nSo I  guess we dont have to worry about that (and yes, if this is mistaken admins probably should say something seeing this LB )\n\nThank you again, Wojtek!",
      "votes": null
    },
    {
      "id": "194513",
      "postDate": "06/20/2017 22:16:08",
      "content": "<p>I only have 0s because I initially submitted the stage 2 sample submission as a test (merged 512 with correct answers and the rest as-is). All of my other submissions (I will check mark the two I want for the final) have gone through my process for prediction without merges. Hope that helps.</p>\n\n<p>Rodney</p>",
      "rawMarkdown": "I only have 0s because I initially submitted the stage 2 sample submission as a test (merged 512 with correct answers and the rest as-is). All of my other submissions (I will check mark the two I want for the final) have gone through my process for prediction without merges. Hope that helps.\n\nRodney",
      "votes": null
    },
    {
      "id": "194697",
      "postDate": "06/21/2017 12:27:04",
      "content": "<p>Hi @Rodney Thomas, thank you for your input. </p>\n\n<p>In the end after some debate here and in Kagglenoobs we discovered that people are doing it in different ways with stage1 test, including merging, concatenating previous submission or repredicting  and that it is anyway  irrelevant cause those observations will just be ignored in stage2. </p>",
      "rawMarkdown": "Hi @Rodney Thomas, thank you for your input. \n\nIn the end after some debate here and in Kagglenoobs we discovered that people are doing it in different ways with stage1 test, including merging, concatenating previous submission or repredicting  and that it is anyway  irrelevant cause those observations will just be ignored in stage2.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 194419,
      "author_name": "wrosinski",
      "author_url": "",
      "post_date": "06/20/2017 13:33:17",
      "content": "<p>You can incorporate test images from 1st stage test data but you can use them as validation set or some of the images may be in your validation set when you perform the split, thus not enabling the model to see their labels in training set, in which case you won't have 0 error on LB, as the model will still have to predict them in 'normal' way, not simply using the labels, on which it was trained.</p>",
      "votes": null,
      "replies": [
        {
          "id": 194421,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "06/20/2017 13:45:03",
          "content": "<p>Ouch...  Thank you @Wojtek Rosinski. From your answer I take that we are expected to re-predict stage1test, not just to merge labels... and then of course error can not be 0. It probably will not be 0 even if all test is used to train, unles your train error is 0. </p>\n\n<p>So... good that I can correct that, I wonder if those 60  zero error submissions on the leaderboard are misunderstanding the same I was...  ( I hope they will read this comment cause to be disqualified for such a stupid mistake is a pity for anyone that has worked hard on a comp)</p>\n\n<p>So, again, thank you Wojtek :-)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194425,
          "author_name": "wrosinski",
          "author_url": "",
          "post_date": "06/20/2017 14:08:00",
          "content": "<p>We are not expected to predict them again, I think. Whatever we do with now-released labels should, it should not matter for final standings, so you can either predict stage 1 images with your model or just merge labels into submission.</p>\n\n<p>What will mater in the end is your predictions on stage 2. Incorporation of stage 1 may just be used as a sanity check to see if your submission works properly and there's nothing wrong with your model.</p>\n\n<p>At least this is how I understand this situation. If I'm mistaken anywhere, then the admins should clarify this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 194427,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "06/20/2017 14:14:50",
          "content": "<p>Ah, ok, it all makes sense now. If it just doesnt matter people are doing whatever is more efficient with their submissions, and thus the different LB errors. </p>\n\n<p>So I  guess we dont have to worry about that (and yes, if this is mistaken admins probably should say something seeing this LB )</p>\n\n<p>Thank you again, Wojtek!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 194513,
      "author_name": "vasilkor",
      "author_url": "",
      "post_date": "06/20/2017 22:16:08",
      "content": "<p>I only have 0s because I initially submitted the stage 2 sample submission as a test (merged 512 with correct answers and the rest as-is). All of my other submissions (I will check mark the two I want for the final) have gone through my process for prediction without merges. Hope that helps.</p>\n\n<p>Rodney</p>",
      "votes": null,
      "replies": [
        {
          "id": 194697,
          "author_name": "miguelpm",
          "author_url": "",
          "post_date": "06/21/2017 12:27:04",
          "content": "<p>Hi @Rodney Thomas, thank you for your input. </p>\n\n<p>In the end after some debate here and in Kagglenoobs we discovered that people are doing it in different ways with stage1 test, including merging, concatenating previous submission or repredicting  and that it is anyway  irrelevant cause those observations will just be ignored in stage2. </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "194417": "Hi all,  Im suddenly having a big  doubt about stage2 rules.  \n\nThe reason, I was expecting all teams to have 0 error on the leaderboard, as we already have the labels of stage 1 and -I supposed- we are all expected just to merge them with the rest of the submission... but lots of teams apparently not including those labels, or not yet.\n\nSo, am I missing something? Even if it is obvious  maybe the answer to this question will help others that, like me,  may have misunderstood the expected submission.  (There is a group of 0 errors that maybe understood the same). It would be a pity to self destruct so much work because of such a silly (but big) mistake, that we would have still time to correct.\n\nSo, **is it merging test labels the correct thing to do, it isn't or it just doesn't matter** cause we will not be evaluated on those observations?",
    "194419": "You can incorporate test images from 1st stage test data but you can use them as validation set or some of the images may be in your validation set when you perform the split, thus not enabling the model to see their labels in training set, in which case you won't have 0 error on LB, as the model will still have to predict them in 'normal' way, not simply using the labels, on which it was trained.",
    "194421": "Ouch...  Thank you @Wojtek Rosinski. From your answer I take that we are expected to re-predict stage1test, not just to merge labels... and then of course error can not be 0. It probably will not be 0 even if all test is used to train, unles your train error is 0. \n\nSo... good that I can correct that, I wonder if those 60  zero error submissions on the leaderboard are misunderstanding the same I was...  ( I hope they will read this comment cause to be disqualified for such a stupid mistake is a pity for anyone that has worked hard on a comp)\n\nSo, again, thank you Wojtek :-)",
    "194425": "We are not expected to predict them again, I think. Whatever we do with now-released labels should, it should not matter for final standings, so you can either predict stage 1 images with your model or just merge labels into submission.\n\nWhat will mater in the end is your predictions on stage 2. Incorporation of stage 1 may just be used as a sanity check to see if your submission works properly and there's nothing wrong with your model.\n\nAt least this is how I understand this situation. If I'm mistaken anywhere, then the admins should clarify this.",
    "194427": "Ah, ok, it all makes sense now. If it just doesnt matter people are doing whatever is more efficient with their submissions, and thus the different LB errors. \n\nSo I  guess we dont have to worry about that (and yes, if this is mistaken admins probably should say something seeing this LB )\n\nThank you again, Wojtek!",
    "194513": "I only have 0s because I initially submitted the stage 2 sample submission as a test (merged 512 with correct answers and the rest as-is). All of my other submissions (I will check mark the two I want for the final) have gone through my process for prediction without merges. Hope that helps.\n\nRodney",
    "194697": "Hi @Rodney Thomas, thank you for your input. \n\nIn the end after some debate here and in Kagglenoobs we discovered that people are doing it in different ways with stage1 test, including merging, concatenating previous submission or repredicting  and that it is anyway  irrelevant cause those observations will just be ignored in stage2."
  },
  "source": "meta"
}