{
  "id": 61146,
  "title": "Where \"500 classes\" come from? (dataset has 599 labels)",
  "url": "/competitions/google-ai-open-images-object-detection-track/discussion/61146",
  "author_name": "lyakaap",
  "post_date": "2018-07-14T20:25:30.392000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Evaluation page says that,\n<code>\nThe final mAP is computed as the average AP over the 500 classes. The participants will be ranked on this final metric.\n</code>\nbut I found that the dataset has 599 unique labels.</p>\n\n<p>First, I thought \"500 classes\" means the number of leaf-most classes, but the number of leaf-most classes is 525.</p>\n\n<p>I read this page <a href=\"https://storage.googleapis.com/openimages/web/object_detection_metric.html\">https://storage.googleapis.com/openimages/web/object_detection_metric.html</a> , but I'm not still clear.\nWhere 500 classes come from?</p>",
  "messages": [
    {
      "id": 356940,
      "postDate": "2018-07-14T20:25:30.393Z",
      "content": "<p>Evaluation page says that,\n<code>\nThe final mAP is computed as the average AP over the 500 classes. The participants will be ranked on this final metric.\n</code>\nbut I found that the dataset has 599 unique labels.</p>\n\n<p>First, I thought \"500 classes\" means the number of leaf-most classes, but the number of leaf-most classes is 525.</p>\n\n<p>I read this page <a href=\"https://storage.googleapis.com/openimages/web/object_detection_metric.html\">https://storage.googleapis.com/openimages/web/object_detection_metric.html</a> , but I'm not still clear.\nWhere 500 classes come from?</p>",
      "rawMarkdown": "Evaluation page says that,\n```\nThe final mAP is computed as the average AP over the 500 classes. The participants will be ranked on this final metric.\n```\nbut I found that the dataset has 599 unique labels.\n\nFirst, I thought \"500 classes\" means the number of leaf-most classes, but the number of leaf-most classes is 525.\n\nI read this page https://storage.googleapis.com/openimages/web/object_detection_metric.html , but I'm not still clear.\nWhere 500 classes come from?",
      "votes": 6
    },
    {
      "id": 357024,
      "postDate": "2018-07-15T05:19:45.837Z",
      "content": "<p>Yes they are subset of All 599 unique classes</p>",
      "rawMarkdown": "Yes they are subset of All 599 unique classes",
      "replies": [
        {
          "id": 357077,
          "postDate": "2018-07-15T07:54:08.367Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 366353,
      "postDate": "2018-08-04T21:49:07.510Z",
      "content": "<p>Can someone check my understanding of what data is supposed to be used for this competition?</p>\n\n<p>There's the V4 data set downloads here\n<a href=\"https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\">https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/</a>\nThere's also a bunch of csv files here\n<a href=\"https://storage.googleapis.com/openimages/web/download.html\">https://storage.googleapis.com/openimages/web/download.html</a></p>\n\n<p>With training/validation sets, annotations, and bounding box data. Except these bounding boxes aren't the correct ones to use?</p>\n\n<p>There's the information here:\n<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a>\nWhich has different bbox data and suggests using a subset of the training data as validation instead of the validation set. </p>\n\n<p>So is the correct thing to do to download all the training images, ignore the validation images, and use the csv files provided in the second link?</p>",
      "rawMarkdown": "Can someone check my understanding of what data is supposed to be used for this competition?\n\nThere's the V4 data set downloads here\nhttps://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\nThere's also a bunch of csv files here\nhttps://storage.googleapis.com/openimages/web/download.html\n\nWith training/validation sets, annotations, and bounding box data. Except these bounding boxes aren't the correct ones to use?\n\nThere's the information here:\nhttps://storage.googleapis.com/openimages/web/challenge.html\nWhich has different bbox data and suggests using a subset of the training data as validation instead of the validation set. \n\nSo is the correct thing to do to download all the training images, ignore the validation images, and use the csv files provided in the second link?",
      "replies": [
        {
          "id": 366642,
          "postDate": "2018-08-06T06:15:40.277Z",
          "content": "<p>Dear Karl,</p>\n\n<p>You can find all data to download linked from the challenge website:\n<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a></p>\n\n<p>V4 contains a superset of categories (600 instead of 500) and you are right in that the recommended validation set for the challenge is a subset of the V4 train set, instead of the V4 validation set.</p>\n\n<blockquote>\n  <p>So is the correct thing to do to download all the training images,\n  ignore the validation images, and use the csv files provided in the\n  second link?</p>\n</blockquote>\n\n<p>Yes, except using the download from your \"third\" link:\n<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a></p>\n\n<p>Best,</p>",
          "rawMarkdown": "Dear Karl,\n\nYou can find all data to download linked from the challenge website:\nhttps://storage.googleapis.com/openimages/web/challenge.html\n\nV4 contains a superset of categories (600 instead of 500) and you are right in that the recommended validation set for the challenge is a subset of the V4 train set, instead of the V4 validation set.\n\n&gt; So is the correct thing to do to download all the training images,\n&gt; ignore the validation images, and use the csv files provided in the\n&gt; second link?\n\nYes, except using the download from your \"third\" link:\nhttps://storage.googleapis.com/openimages/web/challenge.html\n\nBest,"
        }
      ]
    },
    {
      "id": 356993,
      "postDate": "2018-07-15T03:14:47.287Z",
      "content": "<p>I think it is according to this page here: <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a></p>\n\n<p>If you scroll down to where it says \"classes\" you will find this file: <a href=\"https://storage.googleapis.com/openimages/challenge_2018/challenge-2018-class-descriptions-500.csv\">https://storage.googleapis.com/openimages/challenge_2018/challenge-2018-class-descriptions-500.csv</a></p>\n\n<p>Which will have 500 classes. I do not know whether these are only the leaf nodes or not. I think these are probably a subset of the full 599 unique classes.</p>",
      "rawMarkdown": "I think it is according to this page here: https://storage.googleapis.com/openimages/web/challenge.html\n\nIf you scroll down to where it says \"classes\" you will find this file: https://storage.googleapis.com/openimages/challenge_2018/challenge-2018-class-descriptions-500.csv\n\nWhich will have 500 classes. I do not know whether these are only the leaf nodes or not. I think these are probably a subset of the full 599 unique classes.",
      "votes": 5,
      "isDeleted": true,
      "replies": [
        {
          "id": 357076,
          "postDate": "2018-07-15T07:53:37.973Z",
          "content": "<p><code>\nThe Challenge is based on Open Images V4. The Object Detection track covers 500 classes out of the 600 annotated with bounding boxes in Open Images V4. We removed some very broad classes (e.g. \"clothing\") and some infrequent ones (e.g. \"paper cutter\").\n</code>\nI overlooked this (<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a>). Now I know the 500 classes. Thanks a lot!</p>",
          "rawMarkdown": "```\nThe Challenge is based on Open Images V4. The Object Detection track covers 500 classes out of the 600 annotated with bounding boxes in Open Images V4. We removed some very broad classes (e.g. \"clothing\") and some infrequent ones (e.g. \"paper cutter\").\n```\nI overlooked this (https://storage.googleapis.com/openimages/web/challenge.html). Now I know the 500 classes. Thanks a lot!",
          "votes": 2
        },
        {
          "id": 357584,
          "postDate": "2018-07-16T11:09:34.580Z",
          "content": "<p>thank you</p>",
          "rawMarkdown": "thank you"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 357024,
      "author_name": "dineshbarri",
      "author_url": "",
      "post_date": "2018-07-15T05:19:45.837000",
      "content": "<p>Yes they are subset of All 599 unique classes</p>",
      "votes": 0,
      "replies": [
        {
          "id": 357077,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2018-07-15T07:54:08.367000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 366353,
      "author_name": "Karl Heyer",
      "author_url": "",
      "post_date": "2018-08-04T21:49:07.510000",
      "content": "<p>Can someone check my understanding of what data is supposed to be used for this competition?</p>\n\n<p>There's the V4 data set downloads here\n<a href=\"https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\">https://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/</a>\nThere's also a bunch of csv files here\n<a href=\"https://storage.googleapis.com/openimages/web/download.html\">https://storage.googleapis.com/openimages/web/download.html</a></p>\n\n<p>With training/validation sets, annotations, and bounding box data. Except these bounding boxes aren't the correct ones to use?</p>\n\n<p>There's the information here:\n<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a>\nWhich has different bbox data and suggests using a subset of the training data as validation instead of the validation set. </p>\n\n<p>So is the correct thing to do to download all the training images, ignore the validation images, and use the csv files provided in the second link?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 366642,
          "author_name": "Jordi Pont-Tuset",
          "author_url": "",
          "post_date": "2018-08-06T06:15:40.277000",
          "content": "<p>Dear Karl,</p>\n\n<p>You can find all data to download linked from the challenge website:\n<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a></p>\n\n<p>V4 contains a superset of categories (600 instead of 500) and you are right in that the recommended validation set for the challenge is a subset of the V4 train set, instead of the V4 validation set.</p>\n\n<blockquote>\n  <p>So is the correct thing to do to download all the training images,\n  ignore the validation images, and use the csv files provided in the\n  second link?</p>\n</blockquote>\n\n<p>Yes, except using the download from your \"third\" link:\n<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a></p>\n\n<p>Best,</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 356993,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-07-15T03:14:47.287000",
      "content": "<p>I think it is according to this page here: <a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a></p>\n\n<p>If you scroll down to where it says \"classes\" you will find this file: <a href=\"https://storage.googleapis.com/openimages/challenge_2018/challenge-2018-class-descriptions-500.csv\">https://storage.googleapis.com/openimages/challenge_2018/challenge-2018-class-descriptions-500.csv</a></p>\n\n<p>Which will have 500 classes. I do not know whether these are only the leaf nodes or not. I think these are probably a subset of the full 599 unique classes.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 357076,
          "author_name": "lyakaap",
          "author_url": "",
          "post_date": "2018-07-15T07:53:37.973000",
          "content": "<p><code>\nThe Challenge is based on Open Images V4. The Object Detection track covers 500 classes out of the 600 annotated with bounding boxes in Open Images V4. We removed some very broad classes (e.g. \"clothing\") and some infrequent ones (e.g. \"paper cutter\").\n</code>\nI overlooked this (<a href=\"https://storage.googleapis.com/openimages/web/challenge.html\">https://storage.googleapis.com/openimages/web/challenge.html</a>). Now I know the 500 classes. Thanks a lot!</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 357584,
          "author_name": "earhian",
          "author_url": "",
          "post_date": "2018-07-16T11:09:34.580000",
          "content": "<p>thank you</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "356940": "Evaluation page says that,\n```\nThe final mAP is computed as the average AP over the 500 classes. The participants will be ranked on this final metric.\n```\nbut I found that the dataset has 599 unique labels.\n\nFirst, I thought \"500 classes\" means the number of leaf-most classes, but the number of leaf-most classes is 525.\n\nI read this page https://storage.googleapis.com/openimages/web/object_detection_metric.html , but I'm not still clear.\nWhere 500 classes come from?",
    "357024": "Yes they are subset of All 599 unique classes",
    "366353": "Can someone check my understanding of what data is supposed to be used for this competition?\n\nThere's the V4 data set downloads here\nhttps://www.figure-eight.com/dataset/open-images-annotated-with-bounding-boxes/\nThere's also a bunch of csv files here\nhttps://storage.googleapis.com/openimages/web/download.html\n\nWith training/validation sets, annotations, and bounding box data. Except these bounding boxes aren't the correct ones to use?\n\nThere's the information here:\nhttps://storage.googleapis.com/openimages/web/challenge.html\nWhich has different bbox data and suggests using a subset of the training data as validation instead of the validation set. \n\nSo is the correct thing to do to download all the training images, ignore the validation images, and use the csv files provided in the second link?",
    "356993": "I think it is according to this page here: https://storage.googleapis.com/openimages/web/challenge.html\n\nIf you scroll down to where it says \"classes\" you will find this file: https://storage.googleapis.com/openimages/challenge_2018/challenge-2018-class-descriptions-500.csv\n\nWhich will have 500 classes. I do not know whether these are only the leaf nodes or not. I think these are probably a subset of the full 599 unique classes."
  }
}