{
  "id": 69761,
  "title": "Misleading Submission Rules",
  "url": "/competitions/rsna-pneumonia-detection-challenge/discussion/69761",
  "author_name": "thesix",
  "post_date": "2018-10-26T22:47:30.505000",
  "votes": 28,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Dear Pneumonia Challenge Organizers,</p>\n\n<p>We believe that the way in which the official rules are written are unclear and, in fact, quite misleading.</p>\n\n<p>In the official \"Rules\" tab of this competition, under the 'MODEL UPLOAD REQUIREMENT' heading, it states that \"each team's Stage 1 submission must include the model uploaded, via Team -&gt; Your Model.\" and \"Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\"</p>\n\n<p>Most people would interpret \"model\" as the model architecture with associated weights. There is absolutely no official mention of a requirement to upload training code.</p>\n\n<p>However, from monitoring the discussion forum, it was mentioned numerous times that we should include whatever files we need to generate the .csv submission file. Logically, this means the model weights and the inference code. Training code is NOT required to generate the .csv submission files if the model weights are uploaded. If the model weights are not uploaded then it is reasonable to upload the training code in addition to inference code, however, this was never clearly stated in the official competition literature and often times repeat training may lead to different model weights even with the exact same code. </p>\n\n<p>From what little we can gather from the recent activity on the discussion forum, it appears that the teams which only uploaded model weights and the inference code during stage 1 are prevented from retraining with the newly available stage 2 labels, placing them at a clear disadvantage. This is unfair. </p>\n\n<p>We hope you can appreciate how this miscommunication on the part of the organizing committee can be very frustrating since many teams have spent significant time and resources on this challenge.</p>\n\n<p>There have already been numerous posts about this on the discussion forum in the past 12 hours, none of which have been addressed by the organizing committee. In fact, unless a team has significant previous experience with two-stage Kaggle competitions, we think it would be impossible to know that submission of training code is mandatory. We wouldn't be surprised if hundreds of teams also didn't submit their training code.</p>\n\n<p>We would very much appreciate a response to our concerns from the organizing committee and also feedback from other teams that may have had a similar experience. </p>\n\n<p>Thank you,</p>\n\n<p>thesix team</p>",
  "messages": [
    {
      "id": 410923,
      "postDate": "2018-10-26T22:47:30.507Z",
      "content": "<p>Dear Pneumonia Challenge Organizers,</p>\n\n<p>We believe that the way in which the official rules are written are unclear and, in fact, quite misleading.</p>\n\n<p>In the official \"Rules\" tab of this competition, under the 'MODEL UPLOAD REQUIREMENT' heading, it states that \"each team's Stage 1 submission must include the model uploaded, via Team -&gt; Your Model.\" and \"Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\"</p>\n\n<p>Most people would interpret \"model\" as the model architecture with associated weights. There is absolutely no official mention of a requirement to upload training code.</p>\n\n<p>However, from monitoring the discussion forum, it was mentioned numerous times that we should include whatever files we need to generate the .csv submission file. Logically, this means the model weights and the inference code. Training code is NOT required to generate the .csv submission files if the model weights are uploaded. If the model weights are not uploaded then it is reasonable to upload the training code in addition to inference code, however, this was never clearly stated in the official competition literature and often times repeat training may lead to different model weights even with the exact same code. </p>\n\n<p>From what little we can gather from the recent activity on the discussion forum, it appears that the teams which only uploaded model weights and the inference code during stage 1 are prevented from retraining with the newly available stage 2 labels, placing them at a clear disadvantage. This is unfair. </p>\n\n<p>We hope you can appreciate how this miscommunication on the part of the organizing committee can be very frustrating since many teams have spent significant time and resources on this challenge.</p>\n\n<p>There have already been numerous posts about this on the discussion forum in the past 12 hours, none of which have been addressed by the organizing committee. In fact, unless a team has significant previous experience with two-stage Kaggle competitions, we think it would be impossible to know that submission of training code is mandatory. We wouldn't be surprised if hundreds of teams also didn't submit their training code.</p>\n\n<p>We would very much appreciate a response to our concerns from the organizing committee and also feedback from other teams that may have had a similar experience. </p>\n\n<p>Thank you,</p>\n\n<p>thesix team</p>",
      "rawMarkdown": "Dear Pneumonia Challenge Organizers,\n\nWe believe that the way in which the official rules are written are unclear and, in fact, quite misleading.\n\nIn the official \"Rules\" tab of this competition, under the 'MODEL UPLOAD REQUIREMENT' heading, it states that \"each team's Stage 1 submission must include the model uploaded, via Team -&gt; Your Model.\" and \"Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\"\n\nMost people would interpret \"model\" as the model architecture with associated weights. There is absolutely no official mention of a requirement to upload training code.\n\nHowever, from monitoring the discussion forum, it was mentioned numerous times that we should include whatever files we need to generate the .csv submission file. Logically, this means the model weights and the inference code. Training code is NOT required to generate the .csv submission files if the model weights are uploaded. If the model weights are not uploaded then it is reasonable to upload the training code in addition to inference code, however, this was never clearly stated in the official competition literature and often times repeat training may lead to different model weights even with the exact same code. \n\nFrom what little we can gather from the recent activity on the discussion forum, it appears that the teams which only uploaded model weights and the inference code during stage 1 are prevented from retraining with the newly available stage 2 labels, placing them at a clear disadvantage. This is unfair. \n\nWe hope you can appreciate how this miscommunication on the part of the organizing committee can be very frustrating since many teams have spent significant time and resources on this challenge.\n\nThere have already been numerous posts about this on the discussion forum in the past 12 hours, none of which have been addressed by the organizing committee. In fact, unless a team has significant previous experience with two-stage Kaggle competitions, we think it would be impossible to know that submission of training code is mandatory. We wouldn't be surprised if hundreds of teams also didn't submit their training code.\n\nWe would very much appreciate a response to our concerns from the organizing committee and also feedback from other teams that may have had a similar experience. \n\n\nThank you,\n\nthesix team",
      "votes": 28
    },
    {
      "id": 412277,
      "postDate": "2018-10-29T22:35:53.300Z",
      "content": "<p>Thanks for sharing your concerns. We'd like to clarify that having uploaded your model weights and inference code as you've described are sufficient to satisfy the model upload requirement. Everyone is eligible to retrain their models on the new stage 2 train set. We haven’t prohibited anyone from retraining on the basis of what they've uploaded at the end of stage 1.</p>\n\n<p>Teams in prize standing can expect the host to scrutinize the model that was uploaded for compliance to ensure that it has not been altered in violation of rules, including inappropriate use of hand annotations or tuning based on having seen the stage 2 dataset. Those prospective winners will be expected to furnish the full extent of their code described in the <a href=\"https://www.kaggle.com/WinningModelDocumentationGuidelines\">Winning Model Guidelines</a>.</p>",
      "rawMarkdown": "Thanks for sharing your concerns. We'd like to clarify that having uploaded your model weights and inference code as you've described are sufficient to satisfy the model upload requirement. Everyone is eligible to retrain their models on the new stage 2 train set. We haven’t prohibited anyone from retraining on the basis of what they've uploaded at the end of stage 1.\n\nTeams in prize standing can expect the host to scrutinize the model that was uploaded for compliance to ensure that it has not been altered in violation of rules, including inappropriate use of hand annotations or tuning based on having seen the stage 2 dataset. Those prospective winners will be expected to furnish the full extent of their code described in the [Winning Model Guidelines](https://www.kaggle.com/WinningModelDocumentationGuidelines).",
      "votes": 1
    },
    {
      "id": 410965,
      "postDate": "2018-10-27T03:11:13.187Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 411297,
          "postDate": "2018-10-27T20:00:24.037Z",
          "content": "<p>I think it would be against the rules, unless it's fully automated and was uploaded as a part of model training and predition script. You can't just re-train you model and for example choose manually the best checkpoint or threshold using the new released labels, it would be a clear rules violation.</p>",
          "rawMarkdown": "I think it would be against the rules, unless it's fully automated and was uploaded as a part of model training and predition script. You can't just re-train you model and for example choose manually the best checkpoint or threshold using the new released labels, it would be a clear rules violation.",
          "votes": 3
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 412277,
      "author_name": "Julia Elliott",
      "author_url": "",
      "post_date": "2018-10-29T22:35:53.300000",
      "content": "<p>Thanks for sharing your concerns. We'd like to clarify that having uploaded your model weights and inference code as you've described are sufficient to satisfy the model upload requirement. Everyone is eligible to retrain their models on the new stage 2 train set. We haven’t prohibited anyone from retraining on the basis of what they've uploaded at the end of stage 1.</p>\n\n<p>Teams in prize standing can expect the host to scrutinize the model that was uploaded for compliance to ensure that it has not been altered in violation of rules, including inappropriate use of hand annotations or tuning based on having seen the stage 2 dataset. Those prospective winners will be expected to furnish the full extent of their code described in the <a href=\"https://www.kaggle.com/WinningModelDocumentationGuidelines\">Winning Model Guidelines</a>.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 410965,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-10-27T03:11:13.187000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 411297,
          "author_name": "Dmytro Poplavskiy",
          "author_url": "",
          "post_date": "2018-10-27T20:00:24.037000",
          "content": "<p>I think it would be against the rules, unless it's fully automated and was uploaded as a part of model training and predition script. You can't just re-train you model and for example choose manually the best checkpoint or threshold using the new released labels, it would be a clear rules violation.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "410923": "Dear Pneumonia Challenge Organizers,\n\nWe believe that the way in which the official rules are written are unclear and, in fact, quite misleading.\n\nIn the official \"Rules\" tab of this competition, under the 'MODEL UPLOAD REQUIREMENT' heading, it states that \"each team's Stage 1 submission must include the model uploaded, via Team -&gt; Your Model.\" and \"Submitters who fail to upload their model by the Stage 1 deadline, or are found not to be in compliance, may be disqualified from Stage 2 and removed from the final leaderboard.\"\n\nMost people would interpret \"model\" as the model architecture with associated weights. There is absolutely no official mention of a requirement to upload training code.\n\nHowever, from monitoring the discussion forum, it was mentioned numerous times that we should include whatever files we need to generate the .csv submission file. Logically, this means the model weights and the inference code. Training code is NOT required to generate the .csv submission files if the model weights are uploaded. If the model weights are not uploaded then it is reasonable to upload the training code in addition to inference code, however, this was never clearly stated in the official competition literature and often times repeat training may lead to different model weights even with the exact same code. \n\nFrom what little we can gather from the recent activity on the discussion forum, it appears that the teams which only uploaded model weights and the inference code during stage 1 are prevented from retraining with the newly available stage 2 labels, placing them at a clear disadvantage. This is unfair. \n\nWe hope you can appreciate how this miscommunication on the part of the organizing committee can be very frustrating since many teams have spent significant time and resources on this challenge.\n\nThere have already been numerous posts about this on the discussion forum in the past 12 hours, none of which have been addressed by the organizing committee. In fact, unless a team has significant previous experience with two-stage Kaggle competitions, we think it would be impossible to know that submission of training code is mandatory. We wouldn't be surprised if hundreds of teams also didn't submit their training code.\n\nWe would very much appreciate a response to our concerns from the organizing committee and also feedback from other teams that may have had a similar experience. \n\n\nThank you,\n\nthesix team",
    "412277": "Thanks for sharing your concerns. We'd like to clarify that having uploaded your model weights and inference code as you've described are sufficient to satisfy the model upload requirement. Everyone is eligible to retrain their models on the new stage 2 train set. We haven’t prohibited anyone from retraining on the basis of what they've uploaded at the end of stage 1.\n\nTeams in prize standing can expect the host to scrutinize the model that was uploaded for compliance to ensure that it has not been altered in violation of rules, including inappropriate use of hand annotations or tuning based on having seen the stage 2 dataset. Those prospective winners will be expected to furnish the full extent of their code described in the [Winning Model Guidelines](https://www.kaggle.com/WinningModelDocumentationGuidelines).",
    "410965": ""
  }
}