{
  "id": 71208,
  "title": "Why more than one submission in stage 2?",
  "url": "/competitions/inclusive-images-challenge/discussion/71208",
  "author_name": "",
  "post_date": "2018-11-11T12:01:50.660126300Z",
  "votes": 8,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Stage 2 is supposed to about running the model(s) trained in stage 1 on stage 2 data without any modeification.  Why would this require more than one submission, or, a few ones in case of a bug or any other incident?  Yet some have way more submissions...  Maybe they did not read the FAQ...</p>\n\n<p>I bet the final results will generate heated discussions, as was the case for the DSB competition this year. </p>",
  "messages": [
    {
      "id": "419156",
      "postDate": "11/11/2018 12:01:50",
      "content": "<p>Stage 2 is supposed to about running the model(s) trained in stage 1 on stage 2 data without any modeification.  Why would this require more than one submission, or, a few ones in case of a bug or any other incident?  Yet some have way more submissions...  Maybe they did not read the FAQ...</p>\n\n<p>I bet the final results will generate heated discussions, as was the case for the DSB competition this year. </p>",
      "rawMarkdown": "Stage 2 is supposed to about running the model(s) trained in stage 1 on stage 2 data without any modeification.  Why would this require more than one submission, or, a few ones in case of a bug or any other incident?  Yet some have way more submissions...  Maybe they did not read the FAQ...\n\nI bet the final results will generate heated discussions, as was the case for the DSB competition this year.",
      "votes": null
    },
    {
      "id": "419181",
      "postDate": "11/11/2018 12:47:33",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> stage 2 public LB is composed of only 10 images in this competition, so I believe final standings will be closer to stage 1 LB.</p>\n\n<p>I agree that stage 2 don't need so many submissions, 1 per day is enough :-)</p>",
      "rawMarkdown": "cpmpml stage 2 public LB is composed of only 10 images in this competition, so I believe final standings will be closer to stage 1 LB.\n\n I agree that stage 2 don't need so many submissions, 1 per day is enough :-)",
      "votes": null
    },
    {
      "id": "419308",
      "postDate": "11/11/2018 17:23:49",
      "content": "<p>In my opinion, showing the 10 images’ score is more problematic than 5 subs per day. In previous 2 stage competitions like TSA and DSB2018, everyone had 0 score in stage 2, which prevented any LB probing. </p>\n\n<p>In this competition, even though only 10 images’ score is shown, it may still give out information on the stage2 test set, depending on how representative these 10 images are. I noticed some teams improving the LB scores throughout the week, presumably tuning thresholds. I hope the host will verify all teams’ submissions by running the models, especially the gold medalists.</p>",
      "rawMarkdown": "In my opinion, showing the 10 images’ score is more problematic than 5 subs per day. In previous 2 stage competitions like TSA and DSB2018, everyone had 0 score in stage 2, which prevented any LB probing. \n\nIn this competition, even though only 10 images’ score is shown, it may still give out information on the stage2 test set, depending on how representative these 10 images are. I noticed some teams improving the LB scores throughout the week, presumably tuning thresholds. I hope the host will verify all teams’ submissions by running the models, especially the gold medalists.",
      "votes": null
    },
    {
      "id": "419337",
      "postDate": "11/11/2018 18:36:33",
      "content": "<blockquote>\n  <p>presumably tuning thresholds</p>\n</blockquote>\n\n<p>This is forbidden explicitly: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923</a></p>",
      "rawMarkdown": "&gt; presumably tuning thresholds\n\nThis is forbidden explicitly: https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923",
      "votes": null
    },
    {
      "id": "419344",
      "postDate": "11/11/2018 19:00:01",
      "content": "<p>Yes, I'm aware. I'm just worried that Kaggle may not strictly enforce this rule by running every team's model -- based on previous competitions.</p>",
      "rawMarkdown": "Yes, I'm aware. I'm just worried that Kaggle may not strictly enforce this rule by running every team's model -- based on previous competitions.",
      "votes": null
    },
    {
      "id": "419729",
      "postDate": "11/12/2018 13:34:17",
      "content": "<p>@Bo I'm pretty sure that they are trying to pseudo train model on test set stage 2 and tuning threshold, but the models and thresholds should be locked after stage 1. \n<a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923</a>\nI think only dataset of stage 1 is allowed to train model.</p>",
      "rawMarkdown": "Bo I'm pretty sure that they are trying to pseudo train model on test set stage 2 and tuning threshold, but the models and thresholds should be locked after stage 1. \nhttps://www.kaggle.com/c/inclusive-images-challenge/discussion/70923\nI think only dataset of stage 1 is allowed to train model.",
      "votes": null
    },
    {
      "id": "419744",
      "postDate": "11/12/2018 13:55:34",
      "content": "<p>I'm also a bit concerned about this. I made many submissions on the current LB to just see the variance I get with various thresholds and models, just for fun. My personal opinion is that tuning any threshold based on the 10 images is foolish anyways, but despite this the organizers should somehow verify that people don't do any funny tunings after stage-1.</p>",
      "rawMarkdown": "I'm also a bit concerned about this. I made many submissions on the current LB to just see the variance I get with various thresholds and models, just for fun. My personal opinion is that tuning any threshold based on the 10 images is foolish anyways, but despite this the organizers should somehow verify that people don't do any funny tunings after stage-1.",
      "votes": null
    },
    {
      "id": "419751",
      "postDate": "11/12/2018 14:05:58",
      "content": "<p>Model tuning using stage2 testset will have non-trivial advantage. Unfortunately it's unlikely for Kaggle to verify everyone's solution. But one possible improvement could be to give more punishment (more than just remove from LB) to intentional cheaters. Deliberate cheating behavior should result in losing all credits on Kaggle.</p>",
      "rawMarkdown": "Model tuning using stage2 testset will have non-trivial advantage. Unfortunately it's unlikely for Kaggle to verify everyone's solution. But one possible improvement could be to give more punishment (more than just remove from LB) to intentional cheaters. Deliberate cheating behavior should result in losing all credits on Kaggle.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 419181,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "11/11/2018 12:47:33",
      "content": "<p><a href=\"/cpmpml\">@cpmpml</a> stage 2 public LB is composed of only 10 images in this competition, so I believe final standings will be closer to stage 1 LB.</p>\n\n<p>I agree that stage 2 don't need so many submissions, 1 per day is enough :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 419308,
      "author_name": "boliu0",
      "author_url": "",
      "post_date": "11/11/2018 17:23:49",
      "content": "<p>In my opinion, showing the 10 images’ score is more problematic than 5 subs per day. In previous 2 stage competitions like TSA and DSB2018, everyone had 0 score in stage 2, which prevented any LB probing. </p>\n\n<p>In this competition, even though only 10 images’ score is shown, it may still give out information on the stage2 test set, depending on how representative these 10 images are. I noticed some teams improving the LB scores throughout the week, presumably tuning thresholds. I hope the host will verify all teams’ submissions by running the models, especially the gold medalists.</p>",
      "votes": null,
      "replies": [
        {
          "id": 419337,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "11/11/2018 18:36:33",
          "content": "<blockquote>\n  <p>presumably tuning thresholds</p>\n</blockquote>\n\n<p>This is forbidden explicitly: <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419344,
          "author_name": "boliu0",
          "author_url": "",
          "post_date": "11/11/2018 19:00:01",
          "content": "<p>Yes, I'm aware. I'm just worried that Kaggle may not strictly enforce this rule by running every team's model -- based on previous competitions.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419729,
          "author_name": "nguyenbadung",
          "author_url": "",
          "post_date": "11/12/2018 13:34:17",
          "content": "<p>@Bo I'm pretty sure that they are trying to pseudo train model on test set stage 2 and tuning threshold, but the models and thresholds should be locked after stage 1. \n<a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923\">https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923</a>\nI think only dataset of stage 1 is allowed to train model.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419744,
          "author_name": "amrit110",
          "author_url": "",
          "post_date": "11/12/2018 13:55:34",
          "content": "<p>I'm also a bit concerned about this. I made many submissions on the current LB to just see the variance I get with various thresholds and models, just for fun. My personal opinion is that tuning any threshold based on the 10 images is foolish anyways, but despite this the organizers should somehow verify that people don't do any funny tunings after stage-1.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 419751,
          "author_name": "wowfattie",
          "author_url": "",
          "post_date": "11/12/2018 14:05:58",
          "content": "<p>Model tuning using stage2 testset will have non-trivial advantage. Unfortunately it's unlikely for Kaggle to verify everyone's solution. But one possible improvement could be to give more punishment (more than just remove from LB) to intentional cheaters. Deliberate cheating behavior should result in losing all credits on Kaggle.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "419156": "Stage 2 is supposed to about running the model(s) trained in stage 1 on stage 2 data without any modeification.  Why would this require more than one submission, or, a few ones in case of a bug or any other incident?  Yet some have way more submissions...  Maybe they did not read the FAQ...\n\nI bet the final results will generate heated discussions, as was the case for the DSB competition this year.",
    "419181": "cpmpml stage 2 public LB is composed of only 10 images in this competition, so I believe final standings will be closer to stage 1 LB.\n\n I agree that stage 2 don't need so many submissions, 1 per day is enough :-)",
    "419308": "In my opinion, showing the 10 images’ score is more problematic than 5 subs per day. In previous 2 stage competitions like TSA and DSB2018, everyone had 0 score in stage 2, which prevented any LB probing. \n\nIn this competition, even though only 10 images’ score is shown, it may still give out information on the stage2 test set, depending on how representative these 10 images are. I noticed some teams improving the LB scores throughout the week, presumably tuning thresholds. I hope the host will verify all teams’ submissions by running the models, especially the gold medalists.",
    "419337": "&gt; presumably tuning thresholds\n\nThis is forbidden explicitly: https://www.kaggle.com/c/inclusive-images-challenge/discussion/70923",
    "419344": "Yes, I'm aware. I'm just worried that Kaggle may not strictly enforce this rule by running every team's model -- based on previous competitions.",
    "419729": "Bo I'm pretty sure that they are trying to pseudo train model on test set stage 2 and tuning threshold, but the models and thresholds should be locked after stage 1. \nhttps://www.kaggle.com/c/inclusive-images-challenge/discussion/70923\nI think only dataset of stage 1 is allowed to train model.",
    "419744": "I'm also a bit concerned about this. I made many submissions on the current LB to just see the variance I get with various thresholds and models, just for fun. My personal opinion is that tuning any threshold based on the 10 images is foolish anyways, but despite this the organizers should somehow verify that people don't do any funny tunings after stage-1.",
    "419751": "Model tuning using stage2 testset will have non-trivial advantage. Unfortunately it's unlikely for Kaggle to verify everyone's solution. But one possible improvement could be to give more punishment (more than just remove from LB) to intentional cheaters. Deliberate cheating behavior should result in losing all credits on Kaggle."
  },
  "source": "meta"
}