{
  "id": 60887,
  "title": "Images appear in both train and test",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/60887",
  "author_name": "",
  "post_date": "2018-07-11T15:11:16.784914400Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I downloaded all the images from CVDF, around  July 7 to 9. I got my train and test set via <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">commands instructed by the website</a> :</p>\n\n<p>for train, I got 1743042 images via:  </p>\n\n<pre><code>  gsutil -m rsync -r gs://open-images-dataset/train  [some dir]\n</code></pre>\n\n<p>and for test, I got 99999 images via:  </p>\n\n<pre><code>  gsutil -m rsync -r gs://open-images-dataset/challenge2018 [some dir]\n</code></pre>\n\n<p>I found there are 2 images appearing both in train and test set :   </p>\n\n<pre><code> 3e5c900366f1ee45.jpg &amp; 93fc92807fae6417.jpg\n</code></pre>\n\n<p>Challenge-2018-train-annotations-bbox.csv  also contains the boxes for these two files.  </p>\n\n<hr>\n\n<p><br>\nAlso a minor issue,  <em>challenge-2018-train-annotations-bbox.csv</em> provided <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">here</a> has only \nboxes for 1674979 images (out of 1743042), and  <em>challenge-2018-train-annotations-human-imagelabels.csv</em> has 1717554 image ids, aslo not covering all the image in train set. While <a href=\"https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\">Figure Eight</a> has <em>train-annotations-bbox.csv</em> and <em>train-images-boxable.csv</em>, which do have the complete image ids. </p>",
  "messages": [
    {
      "id": "355378",
      "postDate": "07/11/2018 15:11:16",
      "content": "<p>I downloaded all the images from CVDF, around  July 7 to 9. I got my train and test set via <a href=\"https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\">commands instructed by the website</a> :</p>\n\n<p>for train, I got 1743042 images via:  </p>\n\n<pre><code>  gsutil -m rsync -r gs://open-images-dataset/train  [some dir]\n</code></pre>\n\n<p>and for test, I got 99999 images via:  </p>\n\n<pre><code>  gsutil -m rsync -r gs://open-images-dataset/challenge2018 [some dir]\n</code></pre>\n\n<p>I found there are 2 images appearing both in train and test set :   </p>\n\n<pre><code> 3e5c900366f1ee45.jpg &amp; 93fc92807fae6417.jpg\n</code></pre>\n\n<p>Challenge-2018-train-annotations-bbox.csv  also contains the boxes for these two files.  </p>\n\n<hr>\n\n<p><br>\nAlso a minor issue,  <em>challenge-2018-train-annotations-bbox.csv</em> provided <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">here</a> has only \nboxes for 1674979 images (out of 1743042), and  <em>challenge-2018-train-annotations-human-imagelabels.csv</em> has 1717554 image ids, aslo not covering all the image in train set. While <a href=\"https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\">Figure Eight</a> has <em>train-annotations-bbox.csv</em> and <em>train-images-boxable.csv</em>, which do have the complete image ids. </p>",
      "rawMarkdown": "I downloaded all the images from CVDF, around  July 7 to 9. I got my train and test set via [commands instructed by the website][1] :\n\nfor train, I got 1743042 images via:  \n\n      gsutil -m rsync -r gs://open-images-dataset/train  [some dir]\n\nand for test, I got 99999 images via:  \n\n      gsutil -m rsync -r gs://open-images-dataset/challenge2018 [some dir]\n\nI found there are 2 images appearing both in train and test set :   \n\n     3e5c900366f1ee45.jpg &amp; 93fc92807fae6417.jpg\n\nChallenge-2018-train-annotations-bbox.csv  also contains the boxes for these two files.  \n\n\n------  \n <br>\nAlso a minor issue,  *challenge-2018-train-annotations-bbox.csv* provided [here][2] has only \nboxes for 1674979 images (out of 1743042), and  *challenge-2018-train-annotations-human-imagelabels.csv* has 1717554 image ids, aslo not covering all the image in train set. While [Figure Eight][3] has *train-annotations-bbox.csv* and *train-images-boxable.csv*, which do have the complete image ids. \n  \n  [1]: https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\n  [2]: https://storage.googleapis.com/openimages/web/challenge.html\n  [3]: https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/",
      "votes": null
    },
    {
      "id": "355421",
      "postDate": "07/11/2018 16:47:50",
      "content": "<p>I believe the test set needs to be downloaded from Kaggle.</p>",
      "rawMarkdown": "I believe the test set needs to be downloaded from Kaggle.",
      "votes": null
    },
    {
      "id": "355482",
      "postDate": "07/11/2018 19:17:17",
      "content": "<p>I've checked, image ids are same for test set on CVDF and  on Kaggle.</p>",
      "rawMarkdown": "I've checked, image ids are same for test set on CVDF and  on Kaggle.",
      "votes": null
    },
    {
      "id": "356580",
      "postDate": "07/13/2018 22:45:12",
      "content": "<p>Thanks for raising the issue. Indeed a very small fraction of images is included in both test and train sets, and it does not influence the result of evaluation significantly. To make the results clean though those images are ignored in evaluation.</p>",
      "rawMarkdown": "Thanks for raising the issue. Indeed a very small fraction of images is included in both test and train sets, and it does not influence the result of evaluation significantly. To make the results clean though those images are ignored in evaluation.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 355421,
      "author_name": "radek1",
      "author_url": "",
      "post_date": "07/11/2018 16:47:50",
      "content": "<p>I believe the test set needs to be downloaded from Kaggle.</p>",
      "votes": null,
      "replies": [
        {
          "id": 355482,
          "author_name": "mrchou",
          "author_url": "",
          "post_date": "07/11/2018 19:17:17",
          "content": "<p>I've checked, image ids are same for test set on CVDF and  on Kaggle.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 356580,
      "author_name": "akuznetsa",
      "author_url": "",
      "post_date": "07/13/2018 22:45:12",
      "content": "<p>Thanks for raising the issue. Indeed a very small fraction of images is included in both test and train sets, and it does not influence the result of evaluation significantly. To make the results clean though those images are ignored in evaluation.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "355378": "I downloaded all the images from CVDF, around  July 7 to 9. I got my train and test set via [commands instructed by the website][1] :\n\nfor train, I got 1743042 images via:  \n\n      gsutil -m rsync -r gs://open-images-dataset/train  [some dir]\n\nand for test, I got 99999 images via:  \n\n      gsutil -m rsync -r gs://open-images-dataset/challenge2018 [some dir]\n\nI found there are 2 images appearing both in train and test set :   \n\n     3e5c900366f1ee45.jpg &amp; 93fc92807fae6417.jpg\n\nChallenge-2018-train-annotations-bbox.csv  also contains the boxes for these two files.  \n\n\n------  \n <br>\nAlso a minor issue,  *challenge-2018-train-annotations-bbox.csv* provided [here][2] has only \nboxes for 1674979 images (out of 1743042), and  *challenge-2018-train-annotations-human-imagelabels.csv* has 1717554 image ids, aslo not covering all the image in train set. While [Figure Eight][3] has *train-annotations-bbox.csv* and *train-images-boxable.csv*, which do have the complete image ids. \n  \n  [1]: https://github.com/cvdfoundation/open-images-dataset#download-images-with-bounding-boxes-annotations\n  [2]: https://storage.googleapis.com/openimages/web/challenge.html\n  [3]: https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/",
    "355421": "I believe the test set needs to be downloaded from Kaggle.",
    "355482": "I've checked, image ids are same for test set on CVDF and  on Kaggle.",
    "356580": "Thanks for raising the issue. Indeed a very small fraction of images is included in both test and train sets, and it does not influence the result of evaluation significantly. To make the results clean though those images are ignored in evaluation."
  },
  "source": "meta"
}