{
  "id": 151476,
  "title": "Disqualification versus Rules Violation",
  "url": "/competitions/flower-classification-with-tpus/discussion/151476",
  "author_name": "",
  "post_date": "2020-05-15T16:21:27.125066900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>It is very important for all Kaggle competitions (past, present, future) that we clarify the topic of disqualification versus rules violation with regard to using external data. Currently the same situation that occurred here in Flower Comp is occurring in Tweet Sentiment Comp <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/145363\">here</a>. </p>\n\n<p>Both comps have created their (train and test) data from a public dataset. Both comps have modified the original data and modified the original labels. Can we use the original dataset? In Flower Comp it has been decided \"No\", but what about Tweet Sentiment Comp?</p>\n\n<p>Specifically, in Tweet Sentiment and Flower Comp is it a Kaggle <strong>rules violation</strong> which means that participants in Tweet Sentiment should immediately stop using external data, or is it a Kaggle prize <strong>disqualification</strong> that only applies to Flower Comp and does not affect Tweet Sentiment?</p>\n\n<h1>Rules Violation</h1>\n\n<p>Every competition shares the same rules posted <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/rules\">here</a> accessed by clicking the \"Rules\" tab. The relevant rule is #B7C\n&gt;External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants</p>\n\n<p>And rule #B5\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n\n<p>Do these two rules prevent the use of external data that overlaps test in Tweet Sentiment and Flower Comp?</p>\n\n<h1>Disqualification</h1>\n\n<p>Flower comp is a playground competition that also awards prizes. As such there are additional qualifications to win prizes posted <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/overview/prizes\">here</a>\n&gt;To be eligible for these prizes, your submission must generate predictions from machine learning code fully run on TPUs for both training and inference.</p>\n\n<p>And additionally, Kaggle added another qualification in the last week of the comp <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836394\">here</a>\n&gt; Training on any examples included in the test set will result in disqualification. </p>\n\n<h1>Consequences</h1>\n\n<p>If you are not qualified to win a prize then obviously you do not. If you violate a rule then you will be removed from the leaderboard. So, is using external data that overlaps test a disqualification or rules violation?</p>\n\n<p>Furthermore, how is this checked? Will the solutions from all 848 Flower Comp participants be reviewed? </p>",
  "messages": [
    {
      "id": "849297",
      "postDate": "05/15/2020 16:21:27",
      "content": "<p>It is very important for all Kaggle competitions (past, present, future) that we clarify the topic of disqualification versus rules violation with regard to using external data. Currently the same situation that occurred here in Flower Comp is occurring in Tweet Sentiment Comp <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/145363\">here</a>. </p>\n\n<p>Both comps have created their (train and test) data from a public dataset. Both comps have modified the original data and modified the original labels. Can we use the original dataset? In Flower Comp it has been decided \"No\", but what about Tweet Sentiment Comp?</p>\n\n<p>Specifically, in Tweet Sentiment and Flower Comp is it a Kaggle <strong>rules violation</strong> which means that participants in Tweet Sentiment should immediately stop using external data, or is it a Kaggle prize <strong>disqualification</strong> that only applies to Flower Comp and does not affect Tweet Sentiment?</p>\n\n<h1>Rules Violation</h1>\n\n<p>Every competition shares the same rules posted <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/rules\">here</a> accessed by clicking the \"Rules\" tab. The relevant rule is #B7C\n&gt;External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants</p>\n\n<p>And rule #B5\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.</p>\n\n<p>Do these two rules prevent the use of external data that overlaps test in Tweet Sentiment and Flower Comp?</p>\n\n<h1>Disqualification</h1>\n\n<p>Flower comp is a playground competition that also awards prizes. As such there are additional qualifications to win prizes posted <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/overview/prizes\">here</a>\n&gt;To be eligible for these prizes, your submission must generate predictions from machine learning code fully run on TPUs for both training and inference.</p>\n\n<p>And additionally, Kaggle added another qualification in the last week of the comp <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836394\">here</a>\n&gt; Training on any examples included in the test set will result in disqualification. </p>\n\n<h1>Consequences</h1>\n\n<p>If you are not qualified to win a prize then obviously you do not. If you violate a rule then you will be removed from the leaderboard. So, is using external data that overlaps test a disqualification or rules violation?</p>\n\n<p>Furthermore, how is this checked? Will the solutions from all 848 Flower Comp participants be reviewed? </p>",
      "rawMarkdown": "It is very important for all Kaggle competitions (past, present, future) that we clarify the topic of disqualification versus rules violation with regard to using external data. Currently the same situation that occurred here in Flower Comp is occurring in Tweet Sentiment Comp [here][2]. \n\nBoth comps have created their (train and test) data from a public dataset. Both comps have modified the original data and modified the original labels. Can we use the original dataset? In Flower Comp it has been decided \"No\", but what about Tweet Sentiment Comp?\n\nSpecifically, in Tweet Sentiment and Flower Comp is it a Kaggle **rules violation** which means that participants in Tweet Sentiment should immediately stop using external data, or is it a Kaggle prize **disqualification** that only applies to Flower Comp and does not affect Tweet Sentiment?\n\n# Rules Violation\nEvery competition shares the same rules posted [here][1] accessed by clicking the \"Rules\" tab. The relevant rule is #B7C\n&gt;External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants\n\nAnd rule #B5\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nDo these two rules prevent the use of external data that overlaps test in Tweet Sentiment and Flower Comp?\n\n# Disqualification\nFlower comp is a playground competition that also awards prizes. As such there are additional qualifications to win prizes posted [here][3]\n&gt;To be eligible for these prizes, your submission must generate predictions from machine learning code fully run on TPUs for both training and inference.\n\nAnd additionally, Kaggle added another qualification in the last week of the comp [here][4]\n&gt; Training on any examples included in the test set will result in disqualification. \n\n# Consequences\nIf you are not qualified to win a prize then obviously you do not. If you violate a rule then you will be removed from the leaderboard. So, is using external data that overlaps test a disqualification or rules violation?\n\nFurthermore, how is this checked? Will the solutions from all 848 Flower Comp participants be reviewed? \n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/rules\n[2]: https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/145363\n[3]: https://www.kaggle.com/c/flower-classification-with-tpus/overview/prizes\n[4]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836394",
      "votes": null
    },
    {
      "id": "849505",
      "postDate": "05/15/2020 20:14:05",
      "content": "<p>Maybe you already saw this, but</p>\n\n<p><a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494</a></p>",
      "rawMarkdown": "Maybe you already saw this, but\n\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494",
      "votes": null
    },
    {
      "id": "849507",
      "postDate": "05/15/2020 20:24:11",
      "content": "<p>For playground competitions where there are no points, medals, or monetary prizes, including this one, we are <strong>not</strong> doing leaderboard removals for prohibited external data use. Those who have used prohibited external data (including images which overlap or duplicate the test set), however, will be disqualified from TPU Leaderboard prizes. With the knowledge that countless people have used external data that contains the test set's images, we are now <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494\">soliciting candidates for the TPU Leaderboard prize</a>. I agree we could have been much clearer on this point sooner.</p>\n\n<p>Sorry for any frustration or confusion this has caused. We worked hard to make this competition happen, with the knowledge of the dataset's vulnerabilities, and wished to offer prizes to further reward participants' efforts. But we've learned that this also incentivizes behavior that distracts from the core purpose of the playground competition, which is learning.</p>\n\n<p>The Tweet Sentiment competition has a fully hidden set of test labels, despite its original dataset being public. Any external data used there would <em>not</em> include the test labels or label-adjacent data, unlike this competition. Any hand-labeling of known test samples in that competition would be grounds for disqualification, per the rules. Discussions around that competition should be conducted in those forums.</p>",
      "rawMarkdown": "For playground competitions where there are no points, medals, or monetary prizes, including this one, we are **not** doing leaderboard removals for prohibited external data use. Those who have used prohibited external data (including images which overlap or duplicate the test set), however, will be disqualified from TPU Leaderboard prizes. With the knowledge that countless people have used external data that contains the test set's images, we are now [soliciting candidates for the TPU Leaderboard prize](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494). I agree we could have been much clearer on this point sooner.\n\nSorry for any frustration or confusion this has caused. We worked hard to make this competition happen, with the knowledge of the dataset's vulnerabilities, and wished to offer prizes to further reward participants' efforts. But we've learned that this also incentivizes behavior that distracts from the core purpose of the playground competition, which is learning.\n\nThe Tweet Sentiment competition has a fully hidden set of test labels, despite its original dataset being public. Any external data used there would *not* include the test labels or label-adjacent data, unlike this competition. Any hand-labeling of known test samples in that competition would be grounds for disqualification, per the rules. Discussions around that competition should be conducted in those forums.",
      "votes": null
    },
    {
      "id": "851811",
      "postDate": "05/18/2020 00:10:18",
      "content": "<p>Indeed, I think you should be the best person to get the first place. In the notebook you shared, I learned methods such as GridMask and cutmix that I have never been in contact with before. These are really useful. The current plan for the first place is quite boring to tell the truth. It is simply a model ensemble with additional data. I think he is not qualified to win the first place. Those who use the additional data should not participate in the award。</p>",
      "rawMarkdown": "Indeed, I think you should be the best person to get the first place. In the notebook you shared, I learned methods such as GridMask and cutmix that I have never been in contact with before. These are really useful. The current plan for the first place is quite boring to tell the truth. It is simply a model ensemble with additional data. I think he is not qualified to win the first place. Those who use the additional data should not participate in the award。",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 849505,
      "author_name": "yihdarshieh",
      "author_url": "",
      "post_date": "05/15/2020 20:14:05",
      "content": "<p>Maybe you already saw this, but</p>\n\n<p><a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494\">https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 849507,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "05/15/2020 20:24:11",
      "content": "<p>For playground competitions where there are no points, medals, or monetary prizes, including this one, we are <strong>not</strong> doing leaderboard removals for prohibited external data use. Those who have used prohibited external data (including images which overlap or duplicate the test set), however, will be disqualified from TPU Leaderboard prizes. With the knowledge that countless people have used external data that contains the test set's images, we are now <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494\">soliciting candidates for the TPU Leaderboard prize</a>. I agree we could have been much clearer on this point sooner.</p>\n\n<p>Sorry for any frustration or confusion this has caused. We worked hard to make this competition happen, with the knowledge of the dataset's vulnerabilities, and wished to offer prizes to further reward participants' efforts. But we've learned that this also incentivizes behavior that distracts from the core purpose of the playground competition, which is learning.</p>\n\n<p>The Tweet Sentiment competition has a fully hidden set of test labels, despite its original dataset being public. Any external data used there would <em>not</em> include the test labels or label-adjacent data, unlike this competition. Any hand-labeling of known test samples in that competition would be grounds for disqualification, per the rules. Discussions around that competition should be conducted in those forums.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 851811,
      "author_name": "anyexiezouqu",
      "author_url": "",
      "post_date": "05/18/2020 00:10:18",
      "content": "<p>Indeed, I think you should be the best person to get the first place. In the notebook you shared, I learned methods such as GridMask and cutmix that I have never been in contact with before. These are really useful. The current plan for the first place is quite boring to tell the truth. It is simply a model ensemble with additional data. I think he is not qualified to win the first place. Those who use the additional data should not participate in the award。</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "849297": "It is very important for all Kaggle competitions (past, present, future) that we clarify the topic of disqualification versus rules violation with regard to using external data. Currently the same situation that occurred here in Flower Comp is occurring in Tweet Sentiment Comp [here][2]. \n\nBoth comps have created their (train and test) data from a public dataset. Both comps have modified the original data and modified the original labels. Can we use the original dataset? In Flower Comp it has been decided \"No\", but what about Tweet Sentiment Comp?\n\nSpecifically, in Tweet Sentiment and Flower Comp is it a Kaggle **rules violation** which means that participants in Tweet Sentiment should immediately stop using external data, or is it a Kaggle prize **disqualification** that only applies to Flower Comp and does not affect Tweet Sentiment?\n\n# Rules Violation\nEvery competition shares the same rules posted [here][1] accessed by clicking the \"Rules\" tab. The relevant rule is #B7C\n&gt;External Data. You may use data other than the Competition Data (“External Data”) to develop and test your models and Submissions. However, you will (i) ensure the External Data is available to use by all participants of the competition for purposes of the competition at no cost to the other participants\n\nAnd rule #B5\n&gt; Submissions may not use or incorporate information from hand labeling or human prediction of the validation dataset or test data records.\n\nDo these two rules prevent the use of external data that overlaps test in Tweet Sentiment and Flower Comp?\n\n# Disqualification\nFlower comp is a playground competition that also awards prizes. As such there are additional qualifications to win prizes posted [here][3]\n&gt;To be eligible for these prizes, your submission must generate predictions from machine learning code fully run on TPUs for both training and inference.\n\nAnd additionally, Kaggle added another qualification in the last week of the comp [here][4]\n&gt; Training on any examples included in the test set will result in disqualification. \n\n# Consequences\nIf you are not qualified to win a prize then obviously you do not. If you violate a rule then you will be removed from the leaderboard. So, is using external data that overlaps test a disqualification or rules violation?\n\nFurthermore, how is this checked? Will the solutions from all 848 Flower Comp participants be reviewed? \n\n[1]: https://www.kaggle.com/c/flower-classification-with-tpus/rules\n[2]: https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/145363\n[3]: https://www.kaggle.com/c/flower-classification-with-tpus/overview/prizes\n[4]: https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836394",
    "849505": "Maybe you already saw this, but\n\nhttps://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494",
    "849507": "For playground competitions where there are no points, medals, or monetary prizes, including this one, we are **not** doing leaderboard removals for prohibited external data use. Those who have used prohibited external data (including images which overlap or duplicate the test set), however, will be disqualified from TPU Leaderboard prizes. With the knowledge that countless people have used external data that contains the test set's images, we are now [soliciting candidates for the TPU Leaderboard prize](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/151494). I agree we could have been much clearer on this point sooner.\n\nSorry for any frustration or confusion this has caused. We worked hard to make this competition happen, with the knowledge of the dataset's vulnerabilities, and wished to offer prizes to further reward participants' efforts. But we've learned that this also incentivizes behavior that distracts from the core purpose of the playground competition, which is learning.\n\nThe Tweet Sentiment competition has a fully hidden set of test labels, despite its original dataset being public. Any external data used there would *not* include the test labels or label-adjacent data, unlike this competition. Any hand-labeling of known test samples in that competition would be grounds for disqualification, per the rules. Discussions around that competition should be conducted in those forums.",
    "851811": "Indeed, I think you should be the best person to get the first place. In the notebook you shared, I learned methods such as GridMask and cutmix that I have never been in contact with before. These are really useful. The current plan for the first place is quite boring to tell the truth. It is simply a model ensemble with additional data. I think he is not qualified to win the first place. Those who use the additional data should not participate in the award。"
  },
  "source": "meta"
}