{
  "id": 134701,
  "title": "External Data Should be Disallowed",
  "url": "/competitions/flower-classification-with-tpus/discussion/134701",
  "author_name": "",
  "post_date": "2020-03-09T21:43:20.889236300Z",
  "votes": 18,
  "comment_count": 6,
  "views": 0,
  "content": "<p>My current LB of 0.96730 only uses the training data. However i discovered today that if you use external datasets, you can most likely achieve a perfect solution of LB 100%.</p>\n\n<p>This is a playground competition similar to Titanic and MNIST competition. People are supposed to avoid using public answers from the internet. The problem with allowing external data, is that it is difficult to determine if the external data are the answers or not because the images can be cropped, resized, rotated etc. </p>\n\n<p>Below are some examples using deep learning image retrieval from the internet. You can see that the test images were created by modifying the below images on the left. Therefore including these external images even though they are different from test images would be unfair.</p>\n\n<p><strong>How do we know whether it is fair or not to include a certain external image?</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F034c6d1e9952ac439a747f8e36d64f60%2FScreen%20Shot%202020-03-09%20at%202.33.29%20PM.png?generation=1583789680653875&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc1c219c4b77133d153f557ccae1247cd%2FScreen%20Shot%202020-03-09%20at%202.37.48%20PM.png?generation=1583789911852633&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8bb5a35eea211db3e3b4c4b107bcf75e%2FScreen%20Shot%202020-03-09%20at%202.37.56%20PM.png?generation=1583789921218906&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5c8a311a4d4205623ac6ca41d2f0b256%2FScreen%20Shot%202020-03-09%20at%202.38.16%20PM.png?generation=1583789940811594&amp;alt=media\" alt=\"\"></p>\n\n<h1>Update</h1>\n\n<p>Kaggle has responded below saying that we can use all public datasets without restriction. In the interest of fairness, I have made the Oxford Flowers data available to everyone as tfrecords <a href=\"https://www.kaggle.com/cdeotte/oxford-flowers-tfrecords\">here</a>. </p>",
  "messages": [
    {
      "id": "767597",
      "postDate": "03/09/2020 21:43:20",
      "content": "<p>My current LB of 0.96730 only uses the training data. However i discovered today that if you use external datasets, you can most likely achieve a perfect solution of LB 100%.</p>\n\n<p>This is a playground competition similar to Titanic and MNIST competition. People are supposed to avoid using public answers from the internet. The problem with allowing external data, is that it is difficult to determine if the external data are the answers or not because the images can be cropped, resized, rotated etc. </p>\n\n<p>Below are some examples using deep learning image retrieval from the internet. You can see that the test images were created by modifying the below images on the left. Therefore including these external images even though they are different from test images would be unfair.</p>\n\n<p><strong>How do we know whether it is fair or not to include a certain external image?</strong></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F034c6d1e9952ac439a747f8e36d64f60%2FScreen%20Shot%202020-03-09%20at%202.33.29%20PM.png?generation=1583789680653875&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc1c219c4b77133d153f557ccae1247cd%2FScreen%20Shot%202020-03-09%20at%202.37.48%20PM.png?generation=1583789911852633&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8bb5a35eea211db3e3b4c4b107bcf75e%2FScreen%20Shot%202020-03-09%20at%202.37.56%20PM.png?generation=1583789921218906&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5c8a311a4d4205623ac6ca41d2f0b256%2FScreen%20Shot%202020-03-09%20at%202.38.16%20PM.png?generation=1583789940811594&amp;alt=media\" alt=\"\"></p>\n\n<h1>Update</h1>\n\n<p>Kaggle has responded below saying that we can use all public datasets without restriction. In the interest of fairness, I have made the Oxford Flowers data available to everyone as tfrecords <a href=\"https://www.kaggle.com/cdeotte/oxford-flowers-tfrecords\">here</a>. </p>",
      "rawMarkdown": "My current LB of 0.96730 only uses the training data. However i discovered today that if you use external datasets, you can most likely achieve a perfect solution of LB 100%.\n\nThis is a playground competition similar to Titanic and MNIST competition. People are supposed to avoid using public answers from the internet. The problem with allowing external data, is that it is difficult to determine if the external data are the answers or not because the images can be cropped, resized, rotated etc. \n\nBelow are some examples using deep learning image retrieval from the internet. You can see that the test images were created by modifying the below images on the left. Therefore including these external images even though they are different from test images would be unfair.\n\n**How do we know whether it is fair or not to include a certain external image?**\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F034c6d1e9952ac439a747f8e36d64f60%2FScreen%20Shot%202020-03-09%20at%202.33.29%20PM.png?generation=1583789680653875&amp;alt=media)\n  \n  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc1c219c4b77133d153f557ccae1247cd%2FScreen%20Shot%202020-03-09%20at%202.37.48%20PM.png?generation=1583789911852633&amp;alt=media)\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8bb5a35eea211db3e3b4c4b107bcf75e%2FScreen%20Shot%202020-03-09%20at%202.37.56%20PM.png?generation=1583789921218906&amp;alt=media)\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5c8a311a4d4205623ac6ca41d2f0b256%2FScreen%20Shot%202020-03-09%20at%202.38.16%20PM.png?generation=1583789940811594&amp;alt=media)\n\n# Update\nKaggle has responded below saying that we can use all public datasets without restriction. In the interest of fairness, I have made the Oxford Flowers data available to everyone as tfrecords [here][1]. \n\n[1]: https://www.kaggle.com/cdeotte/oxford-flowers-tfrecords",
      "votes": null
    },
    {
      "id": "767604",
      "postDate": "03/09/2020 21:58:12",
      "content": "<p>Here are two examples of very close looking images but not the same image. So it would be fair to include these external images. Are we required to test every external image against the test set to make sure it is fair versus unfair?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb8e78ac1ac549085a5b546c3acce5754%2FScreen%20Shot%202020-03-09%20at%202.56.30%20PM.png?generation=1583791052794182&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F729fea020a5596024573182a2d4b26ed%2FScreen%20Shot%202020-03-09%20at%202.56.40%20PM.png?generation=1583791064960097&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Here are two examples of very close looking images but not the same image. So it would be fair to include these external images. Are we required to test every external image against the test set to make sure it is fair versus unfair?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb8e78ac1ac549085a5b546c3acce5754%2FScreen%20Shot%202020-03-09%20at%202.56.30%20PM.png?generation=1583791052794182&amp;alt=media)\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F729fea020a5596024573182a2d4b26ed%2FScreen%20Shot%202020-03-09%20at%202.56.40%20PM.png?generation=1583791064960097&amp;alt=media)",
      "votes": null
    },
    {
      "id": "768052",
      "postDate": "03/10/2020 12:02:44",
      "content": "<p>According to the rules, it seems that it is possible to use a public data set, but only to make it public one week before the competition deadline.<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130276\">offical argument</a>\nHowever, if all the data sets are used for training, it is really meaningless.</p>",
      "rawMarkdown": "According to the rules, it seems that it is possible to use a public data set, but only to make it public one week before the competition deadline.[offical argument](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130276)\nHowever, if all the data sets are used for training, it is really meaningless.",
      "votes": null
    },
    {
      "id": "768403",
      "postDate": "03/10/2020 18:46:56",
      "content": "<p>Yes, this dataset was assembled / cleaned up / reformatted from \"five different public datasets\" as the description states. The list is in the \"Rules\" section: \"collectively sourced and sampled from 5 datasets: <a href=\"http://image-net.org/index\">ImageNet</a>, <a href=\"https://www.robots.ox.ac.uk/~vgg/data/flowers/102/\">Oxford 102 Category Flowers</a>, <a href=\"https://www.tensorflow.org/datasets/catalog/tf_flowers\">TF Flowers</a>, <a href=\"https://storage.googleapis.com/openimages/web/index.html\">Open Images</a>, and <a href=\"https://www.inaturalist.org/\">iNaturalist</a>.\"</p>\n\n<p>These are all public datasets, although we did spend a considerable amount of time making sure the result was nice and clean. The goal was to have a new good quality dataset to play with TPUs.</p>",
      "rawMarkdown": "Yes, this dataset was assembled / cleaned up / reformatted from \"five different public datasets\" as the description states. The list is in the \"Rules\" section: \"collectively sourced and sampled from 5 datasets: [ImageNet](http://image-net.org/index), [Oxford 102 Category Flowers](https://www.robots.ox.ac.uk/~vgg/data/flowers/102/), [TF Flowers](https://www.tensorflow.org/datasets/catalog/tf_flowers), [Open Images](https://storage.googleapis.com/openimages/web/index.html), and [iNaturalist](https://www.inaturalist.org/).\"\n\nThese are all public datasets, although we did spend a considerable amount of time making sure the result was nice and clean. The goal was to have a new good quality dataset to play with TPUs.",
      "votes": null
    },
    {
      "id": "768507",
      "postDate": "03/10/2020 22:09:09",
      "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> for clarifying. And thanks for building a nice flower dataset. When this competition says \"external datasets allowed\", does it really mean \"external datasets excluding those 5\"? The rules seem confusing.</p>",
      "rawMarkdown": "Thanks @mgornergoogle for clarifying. And thanks for building a nice flower dataset. When this competition says \"external datasets allowed\", does it really mean \"external datasets excluding those 5\"? The rules seem confusing.",
      "votes": null
    },
    {
      "id": "768702",
      "postDate": "03/11/2020 05:20:01",
      "content": "<p>In fact, I tried to create one external dataset (I labeled about 133 images) through searching the internet.  But the training result   was not very good. After adding the external dataset, my cv was lower and lb score increased only 0.0002.  Hence I gave up to use external dataset for training.  </p>",
      "rawMarkdown": "In fact, I tried to create one external dataset (I labeled about 133 images) through searching the internet.  But the training result   was not very good. After adding the external dataset, my cv was lower and lb score increased only 0.0002.  Hence I gave up to use external dataset for training.",
      "votes": null
    },
    {
      "id": "769297",
      "postDate": "03/11/2020 18:35:05",
      "content": "<p>Thanks for raising this concern. We are not prohibiting use of any of the 5 datasets used to source the training/test sets. However, the rules do specify that hand-labeling of the test set and incorporating hand-labeled or pre-labeled versions of the test set in your training is prohibited. We recognize the similarities between flowers contained within those datasets and the test set. But ultimately, this is a playground competition and the prizes are centered around demonstrating use of TPUs. </p>\n\n<p>Worth noting also that we will release test set labels at the end of the competition to become more useful for research and learning in the future.</p>",
      "rawMarkdown": "Thanks for raising this concern. We are not prohibiting use of any of the 5 datasets used to source the training/test sets. However, the rules do specify that hand-labeling of the test set and incorporating hand-labeled or pre-labeled versions of the test set in your training is prohibited. We recognize the similarities between flowers contained within those datasets and the test set. But ultimately, this is a playground competition and the prizes are centered around demonstrating use of TPUs. \n\nWorth noting also that we will release test set labels at the end of the competition to become more useful for research and learning in the future.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 767604,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "03/09/2020 21:58:12",
      "content": "<p>Here are two examples of very close looking images but not the same image. So it would be fair to include these external images. Are we required to test every external image against the test set to make sure it is fair versus unfair?</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb8e78ac1ac549085a5b546c3acce5754%2FScreen%20Shot%202020-03-09%20at%202.56.30%20PM.png?generation=1583791052794182&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F729fea020a5596024573182a2d4b26ed%2FScreen%20Shot%202020-03-09%20at%202.56.40%20PM.png?generation=1583791064960097&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 768052,
      "author_name": "anyexiezouqu",
      "author_url": "",
      "post_date": "03/10/2020 12:02:44",
      "content": "<p>According to the rules, it seems that it is possible to use a public data set, but only to make it public one week before the competition deadline.<a href=\"https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130276\">offical argument</a>\nHowever, if all the data sets are used for training, it is really meaningless.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 768403,
      "author_name": "mgorner",
      "author_url": "",
      "post_date": "03/10/2020 18:46:56",
      "content": "<p>Yes, this dataset was assembled / cleaned up / reformatted from \"five different public datasets\" as the description states. The list is in the \"Rules\" section: \"collectively sourced and sampled from 5 datasets: <a href=\"http://image-net.org/index\">ImageNet</a>, <a href=\"https://www.robots.ox.ac.uk/~vgg/data/flowers/102/\">Oxford 102 Category Flowers</a>, <a href=\"https://www.tensorflow.org/datasets/catalog/tf_flowers\">TF Flowers</a>, <a href=\"https://storage.googleapis.com/openimages/web/index.html\">Open Images</a>, and <a href=\"https://www.inaturalist.org/\">iNaturalist</a>.\"</p>\n\n<p>These are all public datasets, although we did spend a considerable amount of time making sure the result was nice and clean. The goal was to have a new good quality dataset to play with TPUs.</p>",
      "votes": null,
      "replies": [
        {
          "id": 768507,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "03/10/2020 22:09:09",
          "content": "<p>Thanks <a href=\"/mgornergoogle\">@mgornergoogle</a> for clarifying. And thanks for building a nice flower dataset. When this competition says \"external datasets allowed\", does it really mean \"external datasets excluding those 5\"? The rules seem confusing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 769297,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "03/11/2020 18:35:05",
          "content": "<p>Thanks for raising this concern. We are not prohibiting use of any of the 5 datasets used to source the training/test sets. However, the rules do specify that hand-labeling of the test set and incorporating hand-labeled or pre-labeled versions of the test set in your training is prohibited. We recognize the similarities between flowers contained within those datasets and the test set. But ultimately, this is a playground competition and the prizes are centered around demonstrating use of TPUs. </p>\n\n<p>Worth noting also that we will release test set labels at the end of the competition to become more useful for research and learning in the future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 768702,
      "author_name": "qinhui1999",
      "author_url": "",
      "post_date": "03/11/2020 05:20:01",
      "content": "<p>In fact, I tried to create one external dataset (I labeled about 133 images) through searching the internet.  But the training result   was not very good. After adding the external dataset, my cv was lower and lb score increased only 0.0002.  Hence I gave up to use external dataset for training.  </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "767597": "My current LB of 0.96730 only uses the training data. However i discovered today that if you use external datasets, you can most likely achieve a perfect solution of LB 100%.\n\nThis is a playground competition similar to Titanic and MNIST competition. People are supposed to avoid using public answers from the internet. The problem with allowing external data, is that it is difficult to determine if the external data are the answers or not because the images can be cropped, resized, rotated etc. \n\nBelow are some examples using deep learning image retrieval from the internet. You can see that the test images were created by modifying the below images on the left. Therefore including these external images even though they are different from test images would be unfair.\n\n**How do we know whether it is fair or not to include a certain external image?**\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F034c6d1e9952ac439a747f8e36d64f60%2FScreen%20Shot%202020-03-09%20at%202.33.29%20PM.png?generation=1583789680653875&amp;alt=media)\n  \n  ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fc1c219c4b77133d153f557ccae1247cd%2FScreen%20Shot%202020-03-09%20at%202.37.48%20PM.png?generation=1583789911852633&amp;alt=media)\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F8bb5a35eea211db3e3b4c4b107bcf75e%2FScreen%20Shot%202020-03-09%20at%202.37.56%20PM.png?generation=1583789921218906&amp;alt=media)\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F5c8a311a4d4205623ac6ca41d2f0b256%2FScreen%20Shot%202020-03-09%20at%202.38.16%20PM.png?generation=1583789940811594&amp;alt=media)\n\n# Update\nKaggle has responded below saying that we can use all public datasets without restriction. In the interest of fairness, I have made the Oxford Flowers data available to everyone as tfrecords [here][1]. \n\n[1]: https://www.kaggle.com/cdeotte/oxford-flowers-tfrecords",
    "767604": "Here are two examples of very close looking images but not the same image. So it would be fair to include these external images. Are we required to test every external image against the test set to make sure it is fair versus unfair?\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2Fb8e78ac1ac549085a5b546c3acce5754%2FScreen%20Shot%202020-03-09%20at%202.56.30%20PM.png?generation=1583791052794182&amp;alt=media)\n  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1723677%2F729fea020a5596024573182a2d4b26ed%2FScreen%20Shot%202020-03-09%20at%202.56.40%20PM.png?generation=1583791064960097&amp;alt=media)",
    "768052": "According to the rules, it seems that it is possible to use a public data set, but only to make it public one week before the competition deadline.[offical argument](https://www.kaggle.com/c/flower-classification-with-tpus/discussion/130276)\nHowever, if all the data sets are used for training, it is really meaningless.",
    "768403": "Yes, this dataset was assembled / cleaned up / reformatted from \"five different public datasets\" as the description states. The list is in the \"Rules\" section: \"collectively sourced and sampled from 5 datasets: [ImageNet](http://image-net.org/index), [Oxford 102 Category Flowers](https://www.robots.ox.ac.uk/~vgg/data/flowers/102/), [TF Flowers](https://www.tensorflow.org/datasets/catalog/tf_flowers), [Open Images](https://storage.googleapis.com/openimages/web/index.html), and [iNaturalist](https://www.inaturalist.org/).\"\n\nThese are all public datasets, although we did spend a considerable amount of time making sure the result was nice and clean. The goal was to have a new good quality dataset to play with TPUs.",
    "768507": "Thanks @mgornergoogle for clarifying. And thanks for building a nice flower dataset. When this competition says \"external datasets allowed\", does it really mean \"external datasets excluding those 5\"? The rules seem confusing.",
    "768702": "In fact, I tried to create one external dataset (I labeled about 133 images) through searching the internet.  But the training result   was not very good. After adding the external dataset, my cv was lower and lb score increased only 0.0002.  Hence I gave up to use external dataset for training.",
    "769297": "Thanks for raising this concern. We are not prohibiting use of any of the 5 datasets used to source the training/test sets. However, the rules do specify that hand-labeling of the test set and incorporating hand-labeled or pre-labeled versions of the test set in your training is prohibited. We recognize the similarities between flowers contained within those datasets and the test set. But ultimately, this is a playground competition and the prizes are centered around demonstrating use of TPUs. \n\nWorth noting also that we will release test set labels at the end of the competition to become more useful for research and learning in the future."
  },
  "source": "meta"
}