{
  "id": 65713,
  "title": "What is the best validate approach?",
  "url": "/competitions/inclusive-images-challenge/discussion/65713",
  "author_name": "",
  "post_date": "2018-09-14T04:05:45.073062800Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Since the distributions of test set and training set are clearly not same, how can we set the validation set? Is there a way to split the training set by location?</p>",
  "messages": [
    {
      "id": "386975",
      "postDate": "09/14/2018 04:05:45",
      "content": "<p>Since the distributions of test set and training set are clearly not same, how can we set the validation set? Is there a way to split the training set by location?</p>",
      "rawMarkdown": "Since the distributions of test set and training set are clearly not same, how can we set the validation set? Is there a way to split the training set by location?",
      "votes": null
    },
    {
      "id": "387504",
      "postDate": "09/15/2018 04:39:56",
      "content": "<p>I am not sure the usage of location metadata of images is allowed or not.</p>",
      "rawMarkdown": "I am not sure the usage of location metadata of images is allowed or not.",
      "votes": null
    },
    {
      "id": "389162",
      "postDate": "09/18/2018 09:24:30",
      "content": "<p>From another <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65069\">thread</a>:</p>\n\n<pre><code>Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.\n\nMetadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.\n\nIf you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the \"Full dataset\" which is larger than the \"Bounding box subset\"].\"\n</code></pre>",
      "rawMarkdown": "From another [thread](https://www.kaggle.com/c/inclusive-images-challenge/discussion/65069):\n\n    Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.\n\n    Metadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.\n\n    If you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the \"Full dataset\" which is larger than the \"Bounding box subset\"].\"\n\n\n  [1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/65069",
      "votes": null
    },
    {
      "id": "389279",
      "postDate": "09/18/2018 13:18:20",
      "content": "<p>Thanks @Krisztian!</p>",
      "rawMarkdown": "Thanks @Krisztian!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 387504,
      "author_name": "ryanzhang",
      "author_url": "",
      "post_date": "09/15/2018 04:39:56",
      "content": "<p>I am not sure the usage of location metadata of images is allowed or not.</p>",
      "votes": null,
      "replies": [
        {
          "id": 389162,
          "author_name": "kk1694",
          "author_url": "",
          "post_date": "09/18/2018 09:24:30",
          "content": "<p>From another <a href=\"https://www.kaggle.com/c/inclusive-images-challenge/discussion/65069\">thread</a>:</p>\n\n<pre><code>Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.\n\nMetadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.\n\nIf you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the \"Full dataset\" which is larger than the \"Bounding box subset\"].\"\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 389279,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "09/18/2018 13:18:20",
          "content": "<p>Thanks @Krisztian!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "386975": "Since the distributions of test set and training set are clearly not same, how can we set the validation set? Is there a way to split the training set by location?",
    "387504": "I am not sure the usage of location metadata of images is allowed or not.",
    "389162": "From another [thread](https://www.kaggle.com/c/inclusive-images-challenge/discussion/65069):\n\n    Metadata that you can find in the image files that we have pointed to (e.g., any exif headers, etc) is fine to use for development as long as your model does not rely on it for inference. At inference time, no metadata can be used other than the image itself.\n\n    Metadata that cannot be found in the image files themselves (e.g., using the attribution files in the challenge dataset or by following links online to find out more about the training images) may not be used at all. That would be considered external data used for development. The attribution files are provided to ensure that the people who donated the images are properly attributed, but they are not intended to be used in the competition at all.\n\n    If you'd like to download the images a different way (e.g., through the Flickr URLs) that's fine as long as you ensure that you are using the proper subset of images training images [note that the TSV files with urls on the CVDF site contain urls for the \"Full dataset\" which is larger than the \"Bounding box subset\"].\"\n\n\n  [1]: https://www.kaggle.com/c/inclusive-images-challenge/discussion/65069",
    "389279": "Thanks @Krisztian!"
  },
  "source": "meta"
}