{
  "id": 34746,
  "title": "How do we get several 0 score submissions for the stage 2?",
  "url": "/competitions/intel-mobileodt-cervical-cancer-screening/discussion/34746",
  "author_name": "",
  "post_date": "2017-06-15T06:45:28.136899Z",
  "votes": 1,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I see first five places have 1-5 submissions and 0 score. How is that possible?</p>",
  "messages": [
    {
      "id": "192941",
      "postDate": "06/15/2017 06:45:28",
      "content": "<p>I see first five places have 1-5 submissions and 0 score. How is that possible?</p>",
      "rawMarkdown": "I see first five places have 1-5 submissions and 0 score. How is that possible?",
      "votes": null
    },
    {
      "id": "192944",
      "postDate": "06/15/2017 06:50:45",
      "content": "<p>yes, seems that the public leaderboard is based on stage1 test data, otherwise it would be really bad :(</p>",
      "rawMarkdown": "yes, seems that the public leaderboard is based on stage1 test data, otherwise it would be really bad :(",
      "votes": null
    },
    {
      "id": "192946",
      "postDate": "06/15/2017 06:57:42",
      "content": "<p>Stage2 public liderboard is completely useless, as stage1 test set labels released.</p>",
      "rawMarkdown": "Stage2 public liderboard is completely useless, as stage1 test set labels released.",
      "votes": null
    },
    {
      "id": "192971",
      "postDate": "06/15/2017 08:43:28",
      "content": "<p>So effectively we are getting no new test information? This means that the advantage of those who mined leaderboard during the first stage is not eliminated. </p>",
      "rawMarkdown": "So effectively we are getting no new test information? This means that the advantage of those who mined leaderboard during the first stage is not eliminated.",
      "votes": null
    },
    {
      "id": "192973",
      "postDate": "06/15/2017 08:46:51",
      "content": "<p>I see the only way this being possible is that the entire Stage 2 Public LB is based on stage1 submission. What I don't get is why would you include Stage1 submissions to start with? Probably to teach the LB miners a lesson by not giving them a clue until the last minute. I like it :D :D</p>",
      "rawMarkdown": "I see the only way this being possible is that the entire Stage 2 Public LB is based on stage1 submission. What I don't get is why would you include Stage1 submissions to start with? Probably to teach the LB miners a lesson by not giving them a clue until the last minute. I like it :D :D",
      "votes": null
    },
    {
      "id": "192985",
      "postDate": "06/15/2017 09:31:29",
      "content": "<p>The test information in stage2 is unusable either way - you are supposed to freeze the pipeline used to generate solutions before the deadline and only run your algorithm once. This can include training on the newly released labels, as long as it's described as a part of your algorithm and the method for training isn't tuned in stage2. At least that's the way I understand it. The multiple submissions are mostly meant for fixing unexpected bugs.</p>",
      "rawMarkdown": "The test information in stage2 is unusable either way - you are supposed to freeze the pipeline used to generate solutions before the deadline and only run your algorithm once. This can include training on the newly released labels, as long as it's described as a part of your algorithm and the method for training isn't tuned in stage2. At least that's the way I understand it. The multiple submissions are mostly meant for fixing unexpected bugs.",
      "votes": null
    },
    {
      "id": "193120",
      "postDate": "06/15/2017 16:57:59",
      "content": "<p>Hi ZadrraS, I'm a little confused, then why admins released stage1 test labels ?</p>",
      "rawMarkdown": "Hi ZadrraS, I'm a little confused, then why admins released stage1 test labels ?",
      "votes": null
    },
    {
      "id": "193121",
      "postDate": "06/15/2017 17:01:03",
      "content": "<p>Hi Sarthak, But kubilai has got different scores from local validation with stage1 test data and stage2 public LB <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/34740\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/34740</a>. So, now I am confused on what to believe </p>",
      "rawMarkdown": "Hi Sarthak, But kubilai has got different scores from local validation with stage1 test data and stage2 public LB https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/34740. So, now I am confused on what to believe",
      "votes": null
    },
    {
      "id": "193135",
      "postDate": "06/15/2017 17:21:30",
      "content": "<p>I was able to validate the stage 1 test data and stage 2 public LB scores. I initially had issues because I realized the the file names in my stage 1 test and the solutions weren't sorted the same. I verified locally as well as replacing the scores from my stage 2 submission with one of my stage 1 submissions and saw the same score. Stage 2 public LB is the 13% of the overall Stage 2.</p>",
      "rawMarkdown": "I was able to validate the stage 1 test data and stage 2 public LB scores. I initially had issues because I realized the the file names in my stage 1 test and the solutions weren't sorted the same. I verified locally as well as replacing the scores from my stage 2 submission with one of my stage 1 submissions and saw the same score. Stage 2 public LB is the 13% of the overall Stage 2.",
      "votes": null
    },
    {
      "id": "193137",
      "postDate": "06/15/2017 17:22:49",
      "content": "<p>@Ravi  I believe they release the labels so that you may train on the additional data(without changing your model parameters) if you'd like to add more information.</p>",
      "rawMarkdown": "Ravi  I believe they release the labels so that you may train on the additional data(without changing your model parameters) if you'd like to add more information.",
      "votes": null
    },
    {
      "id": "193142",
      "postDate": "06/15/2017 17:28:47",
      "content": "<p>Retraining models with more (good) data is always useful, even if you can't tune the details of your algorithm. This isn't a perfect solution, as people who mined the leaderboard early still have a slight advantage, but it's better than nothing. If they allow any sort of tuning in stage2 there isn't any point in doing two stage competitions - the whole point is to see how your (finished) method performs on completely unseen data.</p>",
      "rawMarkdown": "Retraining models with more (good) data is always useful, even if you can't tune the details of your algorithm. This isn't a perfect solution, as people who mined the leaderboard early still have a slight advantage, but it's better than nothing. If they allow any sort of tuning in stage2 there isn't any point in doing two stage competitions - the whole point is to see how your (finished) method performs on completely unseen data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 192944,
      "author_name": "steelrose",
      "author_url": "",
      "post_date": "06/15/2017 06:50:45",
      "content": "<p>yes, seems that the public leaderboard is based on stage1 test data, otherwise it would be really bad :(</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192946,
      "author_name": "uhfband",
      "author_url": "",
      "post_date": "06/15/2017 06:57:42",
      "content": "<p>Stage2 public liderboard is completely useless, as stage1 test set labels released.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 192971,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "06/15/2017 08:43:28",
      "content": "<p>So effectively we are getting no new test information? This means that the advantage of those who mined leaderboard during the first stage is not eliminated. </p>",
      "votes": null,
      "replies": [
        {
          "id": 192985,
          "author_name": "zadrras",
          "author_url": "",
          "post_date": "06/15/2017 09:31:29",
          "content": "<p>The test information in stage2 is unusable either way - you are supposed to freeze the pipeline used to generate solutions before the deadline and only run your algorithm once. This can include training on the newly released labels, as long as it's described as a part of your algorithm and the method for training isn't tuned in stage2. At least that's the way I understand it. The multiple submissions are mostly meant for fixing unexpected bugs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193120,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "06/15/2017 16:57:59",
          "content": "<p>Hi ZadrraS, I'm a little confused, then why admins released stage1 test labels ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193137,
          "author_name": "timothyrimnac",
          "author_url": "",
          "post_date": "06/15/2017 17:22:49",
          "content": "<p>@Ravi  I believe they release the labels so that you may train on the additional data(without changing your model parameters) if you'd like to add more information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193142,
          "author_name": "zadrras",
          "author_url": "",
          "post_date": "06/15/2017 17:28:47",
          "content": "<p>Retraining models with more (good) data is always useful, even if you can't tune the details of your algorithm. This isn't a perfect solution, as people who mined the leaderboard early still have a slight advantage, but it's better than nothing. If they allow any sort of tuning in stage2 there isn't any point in doing two stage competitions - the whole point is to see how your (finished) method performs on completely unseen data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 192973,
      "author_name": "yadavsarthak",
      "author_url": "",
      "post_date": "06/15/2017 08:46:51",
      "content": "<p>I see the only way this being possible is that the entire Stage 2 Public LB is based on stage1 submission. What I don't get is why would you include Stage1 submissions to start with? Probably to teach the LB miners a lesson by not giving them a clue until the last minute. I like it :D :D</p>",
      "votes": null,
      "replies": [
        {
          "id": 193121,
          "author_name": "rteja1113",
          "author_url": "",
          "post_date": "06/15/2017 17:01:03",
          "content": "<p>Hi Sarthak, But kubilai has got different scores from local validation with stage1 test data and stage2 public LB <a href=\"https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/34740\">https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/34740</a>. So, now I am confused on what to believe </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 193135,
          "author_name": "timothyrimnac",
          "author_url": "",
          "post_date": "06/15/2017 17:21:30",
          "content": "<p>I was able to validate the stage 1 test data and stage 2 public LB scores. I initially had issues because I realized the the file names in my stage 1 test and the solutions weren't sorted the same. I verified locally as well as replacing the scores from my stage 2 submission with one of my stage 1 submissions and saw the same score. Stage 2 public LB is the 13% of the overall Stage 2.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "192941": "I see first five places have 1-5 submissions and 0 score. How is that possible?",
    "192944": "yes, seems that the public leaderboard is based on stage1 test data, otherwise it would be really bad :(",
    "192946": "Stage2 public liderboard is completely useless, as stage1 test set labels released.",
    "192971": "So effectively we are getting no new test information? This means that the advantage of those who mined leaderboard during the first stage is not eliminated.",
    "192973": "I see the only way this being possible is that the entire Stage 2 Public LB is based on stage1 submission. What I don't get is why would you include Stage1 submissions to start with? Probably to teach the LB miners a lesson by not giving them a clue until the last minute. I like it :D :D",
    "192985": "The test information in stage2 is unusable either way - you are supposed to freeze the pipeline used to generate solutions before the deadline and only run your algorithm once. This can include training on the newly released labels, as long as it's described as a part of your algorithm and the method for training isn't tuned in stage2. At least that's the way I understand it. The multiple submissions are mostly meant for fixing unexpected bugs.",
    "193120": "Hi ZadrraS, I'm a little confused, then why admins released stage1 test labels ?",
    "193121": "Hi Sarthak, But kubilai has got different scores from local validation with stage1 test data and stage2 public LB https://www.kaggle.com/c/intel-mobileodt-cervical-cancer-screening/discussion/34740. So, now I am confused on what to believe",
    "193135": "I was able to validate the stage 1 test data and stage 2 public LB scores. I initially had issues because I realized the the file names in my stage 1 test and the solutions weren't sorted the same. I verified locally as well as replacing the scores from my stage 2 submission with one of my stage 1 submissions and saw the same score. Stage 2 public LB is the 13% of the overall Stage 2.",
    "193137": "Ravi  I believe they release the labels so that you may train on the additional data(without changing your model parameters) if you'd like to add more information.",
    "193142": "Retraining models with more (good) data is always useful, even if you can't tune the details of your algorithm. This isn't a perfect solution, as people who mined the leaderboard early still have a slight advantage, but it's better than nothing. If they allow any sort of tuning in stage2 there isn't any point in doing two stage competitions - the whole point is to see how your (finished) method performs on completely unseen data."
  },
  "source": "meta"
}