{
  "id": 150057,
  "title": "The first competition I do not want to win :)",
  "url": "/competitions/flower-classification-with-tpus/discussion/150057",
  "author_name": "Victor Paslay",
  "post_date": "2020-05-10T23:10:26.290000",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>So, after reading the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836349\">latest news from Kaggle Team member</a> I started thinking that it would be nice to get the 4th place and do not get disqualified since I am not sure my solution complies with the competition rules.</p>\n\n<p>Most probably all the guys in the top use external datasets and I use them as well. My solution was built also on <em>cleaning</em> the dataset(removing wrong labels) and I spent significant amount of time(at least 4 hours) on making the dataset right. Some of my models have been trained for more than 7-10 hours(multiple runs of the notebook) and now I am too lazy to clean the dataset again to avoid copies of the samples from the test set if there are any and retrain the models. I do not feel like I cheated at all, and I think - woah, I just want to keep the current position after the private LB is revealed(or two place higher :D .</p>\n\n<p>I think that statement \"<strong>It is prohibited to train your models using any mirrors/overlaps/duplication of the competition's test data</strong>\" either vague or can be tricked by simple approaches. \nThe key question is what is a <em>duplication of an image</em>? Exactly the same pixel values? Ok, then I can just change a few pixels or brightness - and I got another image picked from the test set. The algorithm could be the next:\n1. Find the most similar images between the test set and external datasets(hash-values, SURF, HOG or something like this).\n2. Apply labels as in external datasets.\n3. Apply soft augmentations so the images are slightly different.\n4. Add them to your training set and fit your model on this data.\n5. Profit :)</p>\n\n<p>I think the wisest decision could be to <strong>ban external datasets from the beginning</strong> and that would leave the space only to working on modelling/approaches.</p>\n\n<p>Am I frustrated? Slightly. Anyway, in the end only experience matters and I like this competition.</p>",
  "messages": [
    {
      "id": 841646,
      "postDate": "2020-05-10T23:10:26.290Z",
      "content": "<p>So, after reading the <a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836349\">latest news from Kaggle Team member</a> I started thinking that it would be nice to get the 4th place and do not get disqualified since I am not sure my solution complies with the competition rules.</p>\n\n<p>Most probably all the guys in the top use external datasets and I use them as well. My solution was built also on <em>cleaning</em> the dataset(removing wrong labels) and I spent significant amount of time(at least 4 hours) on making the dataset right. Some of my models have been trained for more than 7-10 hours(multiple runs of the notebook) and now I am too lazy to clean the dataset again to avoid copies of the samples from the test set if there are any and retrain the models. I do not feel like I cheated at all, and I think - woah, I just want to keep the current position after the private LB is revealed(or two place higher :D .</p>\n\n<p>I think that statement \"<strong>It is prohibited to train your models using any mirrors/overlaps/duplication of the competition's test data</strong>\" either vague or can be tricked by simple approaches. \nThe key question is what is a <em>duplication of an image</em>? Exactly the same pixel values? Ok, then I can just change a few pixels or brightness - and I got another image picked from the test set. The algorithm could be the next:\n1. Find the most similar images between the test set and external datasets(hash-values, SURF, HOG or something like this).\n2. Apply labels as in external datasets.\n3. Apply soft augmentations so the images are slightly different.\n4. Add them to your training set and fit your model on this data.\n5. Profit :)</p>\n\n<p>I think the wisest decision could be to <strong>ban external datasets from the beginning</strong> and that would leave the space only to working on modelling/approaches.</p>\n\n<p>Am I frustrated? Slightly. Anyway, in the end only experience matters and I like this competition.</p>",
      "rawMarkdown": "So, after reading the [latest news from Kaggle Team member](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836349) I started thinking that it would be nice to get the 4th place and do not get disqualified since I am not sure my solution complies with the competition rules.\n\nMost probably all the guys in the top use external datasets and I use them as well. My solution was built also on *cleaning* the dataset(removing wrong labels) and I spent significant amount of time(at least 4 hours) on making the dataset right. Some of my models have been trained for more than 7-10 hours(multiple runs of the notebook) and now I am too lazy to clean the dataset again to avoid copies of the samples from the test set if there are any and retrain the models. I do not feel like I cheated at all, and I think - woah, I just want to keep the current position after the private LB is revealed(or two place higher :D .\n\nI think that statement \"**It is prohibited to train your models using any mirrors/overlaps/duplication of the competition's test data**\" either vague or can be tricked by simple approaches. \nThe key question is what is a *duplication of an image*? Exactly the same pixel values? Ok, then I can just change a few pixels or brightness - and I got another image picked from the test set. The algorithm could be the next:\n1. Find the most similar images between the test set and external datasets(hash-values, SURF, HOG or something like this).\n2. Apply labels as in external datasets.\n3. Apply soft augmentations so the images are slightly different.\n4. Add them to your training set and fit your model on this data.\n5. Profit :)\n\nI think the wisest decision could be to **ban external datasets from the beginning** and that would leave the space only to working on modelling/approaches.\n\nAm I frustrated? Slightly. Anyway, in the end only experience matters and I like this competition.",
      "votes": 9
    },
    {
      "id": 842827,
      "postDate": "2020-05-11T16:39:05.610Z",
      "content": "<p>Cannot agree with you more. I thought the best way is to ban external dataset at first.</p>",
      "rawMarkdown": "Cannot agree with you more. I thought the best way is to ban external dataset at first.",
      "votes": 2
    },
    {
      "id": 842491,
      "postDate": "2020-05-11T12:29:06.620Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 842827,
      "author_name": "Chung-Hsien Tsai",
      "author_url": "",
      "post_date": "2020-05-11T16:39:05.610000",
      "content": "<p>Cannot agree with you more. I thought the best way is to ban external dataset at first.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 842491,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-11T12:29:06.620000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "841646": "So, after reading the [latest news from Kaggle Team member](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/148329#836349) I started thinking that it would be nice to get the 4th place and do not get disqualified since I am not sure my solution complies with the competition rules.\n\nMost probably all the guys in the top use external datasets and I use them as well. My solution was built also on *cleaning* the dataset(removing wrong labels) and I spent significant amount of time(at least 4 hours) on making the dataset right. Some of my models have been trained for more than 7-10 hours(multiple runs of the notebook) and now I am too lazy to clean the dataset again to avoid copies of the samples from the test set if there are any and retrain the models. I do not feel like I cheated at all, and I think - woah, I just want to keep the current position after the private LB is revealed(or two place higher :D .\n\nI think that statement \"**It is prohibited to train your models using any mirrors/overlaps/duplication of the competition's test data**\" either vague or can be tricked by simple approaches. \nThe key question is what is a *duplication of an image*? Exactly the same pixel values? Ok, then I can just change a few pixels or brightness - and I got another image picked from the test set. The algorithm could be the next:\n1. Find the most similar images between the test set and external datasets(hash-values, SURF, HOG or something like this).\n2. Apply labels as in external datasets.\n3. Apply soft augmentations so the images are slightly different.\n4. Add them to your training set and fit your model on this data.\n5. Profit :)\n\nI think the wisest decision could be to **ban external datasets from the beginning** and that would leave the space only to working on modelling/approaches.\n\nAm I frustrated? Slightly. Anyway, in the end only experience matters and I like this competition.",
    "842827": "Cannot agree with you more. I thought the best way is to ban external dataset at first.",
    "842491": ""
  }
}