{
  "id": 61895,
  "title": "Fewer images annotated than in training set?",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/61895",
  "author_name": "",
  "post_date": "2018-07-24T22:20:02.296903500Z",
  "votes": null,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Sorry if this has been addressed before.</p>\n\n<p>But there seems to be approx 68K fewer images annotated than in the train set.</p>\n\n<p>If I do something like:\n\"cat challenge-2018-train-annotations-bbox.csv | cut -d',' -f1 | sort | uniq | wc -l\"\nGives me: 1674980</p>\n\n<p>and: \"ls train/ | wc -l\"\nGives me: 1743042 (this seems correct from the challenge description)</p>\n\n<p>And so: 1743042-1674980  = 68062</p>\n\n<p>I know that the suggested validation set is mixed in with the train set, but even this does not add up as val set is 100K.</p>\n\n<p>Any clues?</p>",
  "messages": [
    {
      "id": "361657",
      "postDate": "07/24/2018 22:20:02",
      "content": "<p>Sorry if this has been addressed before.</p>\n\n<p>But there seems to be approx 68K fewer images annotated than in the train set.</p>\n\n<p>If I do something like:\n\"cat challenge-2018-train-annotations-bbox.csv | cut -d',' -f1 | sort | uniq | wc -l\"\nGives me: 1674980</p>\n\n<p>and: \"ls train/ | wc -l\"\nGives me: 1743042 (this seems correct from the challenge description)</p>\n\n<p>And so: 1743042-1674980  = 68062</p>\n\n<p>I know that the suggested validation set is mixed in with the train set, but even this does not add up as val set is 100K.</p>\n\n<p>Any clues?</p>",
      "rawMarkdown": "Sorry if this has been addressed before.\n\nBut there seems to be approx 68K fewer images annotated than in the train set.\n\nIf I do something like:\n\"cat challenge-2018-train-annotations-bbox.csv | cut -d',' -f1 | sort | uniq | wc -l\"\nGives me: 1674980\n\nand: \"ls train/ | wc -l\"\nGives me: 1743042 (this seems correct from the challenge description)\n\nAnd so: 1743042-1674980  = 68062\n\nI know that the suggested validation set is mixed in with the train set, but even this does not add up as val set is 100K.\n\nAny clues?",
      "votes": null
    },
    {
      "id": "361862",
      "postDate": "07/25/2018 08:01:04",
      "content": "<p>If an image does not appear in challenge-2018-train-annotations-bbox.csv, then it means it does not have any object from the 500 categories of the challenge annotated (remember that Open Images V4 has 600 categories, but the challenge \"only\" 500 of those). </p>\n\n<p>It sounds about right that 68k images (4%) are those images that only had instances of these 100 discarded categories.</p>",
      "rawMarkdown": "If an image does not appear in challenge-2018-train-annotations-bbox.csv, then it means it does not have any object from the 500 categories of the challenge annotated (remember that Open Images V4 has 600 categories, but the challenge \"only\" 500 of those). \n\nIt sounds about right that 68k images (4%) are those images that only had instances of these 100 discarded categories.",
      "votes": null
    },
    {
      "id": "361904",
      "postDate": "07/25/2018 09:33:13",
      "content": "<p>Thanks for this information, this should be explicitly mentioned somewhere, it's' not self explanatory.</p>",
      "rawMarkdown": "Thanks for this information, this should be explicitly mentioned somewhere, it's' not self explanatory.",
      "votes": null
    },
    {
      "id": "361908",
      "postDate": "07/25/2018 09:37:37",
      "content": "<p>From the <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">challenge website</a>:</p>\n\n<blockquote>\n  <p>All images and annotations are a subset of Open Images V4 training set, restricted to the 500 object classes of the challenge. We provide bounding box annotations and image-level annotations (both positive and negative).</p>\n</blockquote>",
      "rawMarkdown": "From the [challenge website][1]:\n\n&gt; All images and annotations are a subset of Open Images V4 training set, restricted to the 500 object classes of the challenge. We provide bounding box annotations and image-level annotations (both positive and negative).\n\n\n  [1]: https://storage.googleapis.com/openimages/web/challenge.html",
      "votes": null
    },
    {
      "id": "361988",
      "postDate": "07/25/2018 12:55:27",
      "content": "<p>Ah! Thanks for the explanation. It did not occur to me.</p>",
      "rawMarkdown": "Ah! Thanks for the explanation. It did not occur to me.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 361862,
      "author_name": "jponttusset",
      "author_url": "",
      "post_date": "07/25/2018 08:01:04",
      "content": "<p>If an image does not appear in challenge-2018-train-annotations-bbox.csv, then it means it does not have any object from the 500 categories of the challenge annotated (remember that Open Images V4 has 600 categories, but the challenge \"only\" 500 of those). </p>\n\n<p>It sounds about right that 68k images (4%) are those images that only had instances of these 100 discarded categories.</p>",
      "votes": null,
      "replies": [
        {
          "id": 361904,
          "author_name": "rishabhiitbhu",
          "author_url": "",
          "post_date": "07/25/2018 09:33:13",
          "content": "<p>Thanks for this information, this should be explicitly mentioned somewhere, it's' not self explanatory.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361908,
          "author_name": "jponttusset",
          "author_url": "",
          "post_date": "07/25/2018 09:37:37",
          "content": "<p>From the <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">challenge website</a>:</p>\n\n<blockquote>\n  <p>All images and annotations are a subset of Open Images V4 training set, restricted to the 500 object classes of the challenge. We provide bounding box annotations and image-level annotations (both positive and negative).</p>\n</blockquote>",
          "votes": null,
          "replies": []
        },
        {
          "id": 361988,
          "author_name": "ggopalan",
          "author_url": "",
          "post_date": "07/25/2018 12:55:27",
          "content": "<p>Ah! Thanks for the explanation. It did not occur to me.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "361657": "Sorry if this has been addressed before.\n\nBut there seems to be approx 68K fewer images annotated than in the train set.\n\nIf I do something like:\n\"cat challenge-2018-train-annotations-bbox.csv | cut -d',' -f1 | sort | uniq | wc -l\"\nGives me: 1674980\n\nand: \"ls train/ | wc -l\"\nGives me: 1743042 (this seems correct from the challenge description)\n\nAnd so: 1743042-1674980  = 68062\n\nI know that the suggested validation set is mixed in with the train set, but even this does not add up as val set is 100K.\n\nAny clues?",
    "361862": "If an image does not appear in challenge-2018-train-annotations-bbox.csv, then it means it does not have any object from the 500 categories of the challenge annotated (remember that Open Images V4 has 600 categories, but the challenge \"only\" 500 of those). \n\nIt sounds about right that 68k images (4%) are those images that only had instances of these 100 discarded categories.",
    "361904": "Thanks for this information, this should be explicitly mentioned somewhere, it's' not self explanatory.",
    "361908": "From the [challenge website][1]:\n\n&gt; All images and annotations are a subset of Open Images V4 training set, restricted to the 500 object classes of the challenge. We provide bounding box annotations and image-level annotations (both positive and negative).\n\n\n  [1]: https://storage.googleapis.com/openimages/web/challenge.html",
    "361988": "Ah! Thanks for the explanation. It did not occur to me."
  },
  "source": "meta"
}