{
  "id": 105478,
  "title": "Open Images instance segmentation challenge metrics with Tensorflow/models",
  "url": "/competitions/open-images-2019-instance-segmentation/discussion/105478",
  "author_name": "",
  "post_date": "2019-08-23T12:15:12.799623800Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone,\nThe Open Images challenge official metrics are implemented within <code>tensorflow/models</code> as described <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md\">here</a>. The section for instance segmentation explains that masks have to be converted to some <code>.csv</code> format and COCO RLE encoding, and that \"the util to make the transformation will be released soon\". I've been waiting for a while, so I decided to implement it myself...</p>\n\n<p>First of all, let me point out that - to my understanding - the column named <code>GroupOf</code> should be called <code>IsGroupOf</code>. Even after that, I am still getting a <code>KeyError: 'LabelName'</code> in\n<code>File \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary\n    data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),</code>\nAnyone managed to run the segmentation metrics successfully? What tweaks did you need?</p>",
  "messages": [
    {
      "id": "606311",
      "postDate": "08/23/2019 12:15:12",
      "content": "<p>Hi everyone,\nThe Open Images challenge official metrics are implemented within <code>tensorflow/models</code> as described <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md\">here</a>. The section for instance segmentation explains that masks have to be converted to some <code>.csv</code> format and COCO RLE encoding, and that \"the util to make the transformation will be released soon\". I've been waiting for a while, so I decided to implement it myself...</p>\n\n<p>First of all, let me point out that - to my understanding - the column named <code>GroupOf</code> should be called <code>IsGroupOf</code>. Even after that, I am still getting a <code>KeyError: 'LabelName'</code> in\n<code>File \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary\n    data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),</code>\nAnyone managed to run the segmentation metrics successfully? What tweaks did you need?</p>",
      "rawMarkdown": "Hi everyone,\nThe Open Images challenge official metrics are implemented within `tensorflow/models` as described [here](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md). The section for instance segmentation explains that masks have to be converted to some `.csv` format and COCO RLE encoding, and that \"the util to make the transformation will be released soon\". I've been waiting for a while, so I decided to implement it myself...\n\nFirst of all, let me point out that - to my understanding - the column named `GroupOf` should be called `IsGroupOf`. Even after that, I am still getting a `KeyError: 'LabelName'` in\n```File \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary\n    data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),```\nAnyone managed to run the segmentation metrics successfully? What tweaks did you need?",
      "votes": null
    },
    {
      "id": "609820",
      "postDate": "08/28/2019 06:14:56",
      "content": "<p>Any update? I have the same problem,thanks!</p>",
      "rawMarkdown": "Any update? I have the same problem,thanks!",
      "votes": null
    },
    {
      "id": "609940",
      "postDate": "08/28/2019 08:52:25",
      "content": "<p>Sorry no update so far... The community seems to be more interested in flares about private sharing U.U\nThanks for sharing that you're also having the same issue.</p>",
      "rawMarkdown": "Sorry no update so far... The community seems to be more interested in flares about private sharing U.U\nThanks for sharing that you're also having the same issue.",
      "votes": null
    },
    {
      "id": "614913",
      "postDate": "09/01/2019 08:31:40",
      "content": "<p>You are correct that the column named <code>GroupOf</code> should be called <code>IsGroupOf</code>. As for the<code>KeyError: 'LabelName'</code>, it's because Kaggle's prediction format is different. Instead of separate columns for label name, confidence score, etc., Kaggle wants one column called \"PredictionString\". Assuming you made predictions in this format, you have to modify the evaluation code. You can modify the lambda function to parse the prediction string and get the fields you want.</p>\n\n<p><a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py#L181\">https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py#L181</a></p>",
      "rawMarkdown": "You are correct that the column named `GroupOf` should be called `IsGroupOf`. As for the`KeyError: 'LabelName'`, it's because Kaggle's prediction format is different. Instead of separate columns for label name, confidence score, etc., Kaggle wants one column called \"PredictionString\". Assuming you made predictions in this format, you have to modify the evaluation code. You can modify the lambda function to parse the prediction string and get the fields you want.\n\nhttps://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py#L181",
      "votes": null
    },
    {
      "id": "615645",
      "postDate": "09/02/2019 07:50:41",
      "content": "<p>Thank you very much for your reply!\nSo if I understand correctly, for instance segmentation we have:</p>\n\n<ul>\n<li><p>the groundtruth annotation format, which is described in <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track\"><code>challenge_evaluation.md</code></a>, which uses columns\n<code>ImageID, LabelName, ImageWidth, ImageHeight, XMin, YMin, XMax, YMax, IsGroupOf, Mask</code></p></li>\n<li><p>the predictions Kaggle submission format, described in the <a href=\"https://www.kaggle.com/c/open-images-2019-instance-segmentation/overview/evaluation\">Kaggle evaluation page</a>, with columns\n<code>ImageID, ImageWidth, ImageHeight, PredictionString</code></p></li>\n<li><p>a prediction <code>Tensorflow/models</code> format, which is not described but can be evinced from <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation.py\"><code>oid_challenge_evaluation.py</code></a> and <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py\"><code>oid_challenge_evaluation_utils.py</code></a>, with columns\n<code>ImageID, ImageWidth, ImageHeight, LabelName, Score, Mask</code></p></li>\n</ul>\n\n<p>I will test this ASAP. Thanks again!</p>",
      "rawMarkdown": "Thank you very much for your reply!\nSo if I understand correctly, for instance segmentation we have:\n\n- the groundtruth annotation format, which is described in [`challenge_evaluation.md`](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track), which uses columns\n```ImageID, LabelName, ImageWidth, ImageHeight, XMin, YMin, XMax, YMax, IsGroupOf, Mask```\n\n- the predictions Kaggle submission format, described in the [Kaggle evaluation page](https://www.kaggle.com/c/open-images-2019-instance-segmentation/overview/evaluation), with columns\n```ImageID, ImageWidth, ImageHeight, PredictionString```\n\n- a prediction `Tensorflow/models` format, which is not described but can be evinced from [`oid_challenge_evaluation.py`](https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation.py) and [`oid_challenge_evaluation_utils.py`](https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py), with columns\n```ImageID, ImageWidth, ImageHeight, LabelName, Score, Mask```\n\nI will test this ASAP. Thanks again!",
      "votes": null
    },
    {
      "id": "616005",
      "postDate": "09/02/2019 15:34:40",
      "content": "<p>I was finally able to run the metrics!\nSome of the nitty-gritty details I had to figure out that might be useful to other participants:\n- before encoding them to COCO RLE, I had to resize the masks for ground-truth annotations to match height and width of the corresponding image (otherwise I get a nice seg fault);\n- after generating a compressed representation of masks as described in the official <a href=\"https://gist.github.com/pculliton/209398a2a52867580c6103e25e55d93c\">gist</a>, if <code>pandas</code> is used to generate the <code>.csv</code> file, masks should be decoded to strings using <code>.decode('utf-8')</code> (otherwise they get printed out as byte-strings and generate errors when uncompressed);\n- the evaluation script is able to deal with different levels of <code>zlib</code> compression because the first bytes contain the information about compression level (not trivial for someone with as little experience as myself).</p>",
      "rawMarkdown": "I was finally able to run the metrics!\nSome of the nitty-gritty details I had to figure out that might be useful to other participants:\n- before encoding them to COCO RLE, I had to resize the masks for ground-truth annotations to match height and width of the corresponding image (otherwise I get a nice seg fault);\n- after generating a compressed representation of masks as described in the official [gist](https://gist.github.com/pculliton/209398a2a52867580c6103e25e55d93c), if `pandas` is used to generate the `.csv` file, masks should be decoded to strings using `.decode('utf-8')` (otherwise they get printed out as byte-strings and generate errors when uncompressed);\n- the evaluation script is able to deal with different levels of `zlib` compression because the first bytes contain the information about compression level (not trivial for someone with as little experience as myself).",
      "votes": null
    },
    {
      "id": "616193",
      "postDate": "09/02/2019 20:26:00",
      "content": "<p>Glad to hear!</p>",
      "rawMarkdown": "Glad to hear!",
      "votes": null
    },
    {
      "id": "627641",
      "postDate": "09/16/2019 07:53:07",
      "content": "<p>Hi, what about the speed of the evaluation scripts? Mine is too slow.</p>",
      "rawMarkdown": "Hi, what about the speed of the evaluation scripts? Mine is too slow.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 609820,
      "author_name": "summery2015",
      "author_url": "",
      "post_date": "08/28/2019 06:14:56",
      "content": "<p>Any update? I have the same problem,thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 609940,
          "author_name": "lntsmn",
          "author_url": "",
          "post_date": "08/28/2019 08:52:25",
          "content": "<p>Sorry no update so far... The community seems to be more interested in flares about private sharing U.U\nThanks for sharing that you're also having the same issue.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 614913,
      "author_name": "dfan97",
      "author_url": "",
      "post_date": "09/01/2019 08:31:40",
      "content": "<p>You are correct that the column named <code>GroupOf</code> should be called <code>IsGroupOf</code>. As for the<code>KeyError: 'LabelName'</code>, it's because Kaggle's prediction format is different. Instead of separate columns for label name, confidence score, etc., Kaggle wants one column called \"PredictionString\". Assuming you made predictions in this format, you have to modify the evaluation code. You can modify the lambda function to parse the prediction string and get the fields you want.</p>\n\n<p><a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py#L181\">https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py#L181</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 615645,
          "author_name": "lntsmn",
          "author_url": "",
          "post_date": "09/02/2019 07:50:41",
          "content": "<p>Thank you very much for your reply!\nSo if I understand correctly, for instance segmentation we have:</p>\n\n<ul>\n<li><p>the groundtruth annotation format, which is described in <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track\"><code>challenge_evaluation.md</code></a>, which uses columns\n<code>ImageID, LabelName, ImageWidth, ImageHeight, XMin, YMin, XMax, YMax, IsGroupOf, Mask</code></p></li>\n<li><p>the predictions Kaggle submission format, described in the <a href=\"https://www.kaggle.com/c/open-images-2019-instance-segmentation/overview/evaluation\">Kaggle evaluation page</a>, with columns\n<code>ImageID, ImageWidth, ImageHeight, PredictionString</code></p></li>\n<li><p>a prediction <code>Tensorflow/models</code> format, which is not described but can be evinced from <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation.py\"><code>oid_challenge_evaluation.py</code></a> and <a href=\"https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py\"><code>oid_challenge_evaluation_utils.py</code></a>, with columns\n<code>ImageID, ImageWidth, ImageHeight, LabelName, Score, Mask</code></p></li>\n</ul>\n\n<p>I will test this ASAP. Thanks again!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 616005,
          "author_name": "lntsmn",
          "author_url": "",
          "post_date": "09/02/2019 15:34:40",
          "content": "<p>I was finally able to run the metrics!\nSome of the nitty-gritty details I had to figure out that might be useful to other participants:\n- before encoding them to COCO RLE, I had to resize the masks for ground-truth annotations to match height and width of the corresponding image (otherwise I get a nice seg fault);\n- after generating a compressed representation of masks as described in the official <a href=\"https://gist.github.com/pculliton/209398a2a52867580c6103e25e55d93c\">gist</a>, if <code>pandas</code> is used to generate the <code>.csv</code> file, masks should be decoded to strings using <code>.decode('utf-8')</code> (otherwise they get printed out as byte-strings and generate errors when uncompressed);\n- the evaluation script is able to deal with different levels of <code>zlib</code> compression because the first bytes contain the information about compression level (not trivial for someone with as little experience as myself).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 616193,
          "author_name": "dfan97",
          "author_url": "",
          "post_date": "09/02/2019 20:26:00",
          "content": "<p>Glad to hear!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 627641,
      "author_name": "wangchuanyuan",
      "author_url": "",
      "post_date": "09/16/2019 07:53:07",
      "content": "<p>Hi, what about the speed of the evaluation scripts? Mine is too slow.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "606311": "Hi everyone,\nThe Open Images challenge official metrics are implemented within `tensorflow/models` as described [here](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md). The section for instance segmentation explains that masks have to be converted to some `.csv` format and COCO RLE encoding, and that \"the util to make the transformation will be released soon\". I've been waiting for a while, so I decided to implement it myself...\n\nFirst of all, let me point out that - to my understanding - the column named `GroupOf` should be called `IsGroupOf`. Even after that, I am still getting a `KeyError: 'LabelName'` in\n```File \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary\n    data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),```\nAnyone managed to run the segmentation metrics successfully? What tweaks did you need?",
    "609820": "Any update? I have the same problem,thanks!",
    "609940": "Sorry no update so far... The community seems to be more interested in flares about private sharing U.U\nThanks for sharing that you're also having the same issue.",
    "614913": "You are correct that the column named `GroupOf` should be called `IsGroupOf`. As for the`KeyError: 'LabelName'`, it's because Kaggle's prediction format is different. Instead of separate columns for label name, confidence score, etc., Kaggle wants one column called \"PredictionString\". Assuming you made predictions in this format, you have to modify the evaluation code. You can modify the lambda function to parse the prediction string and get the fields you want.\n\nhttps://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py#L181",
    "615645": "Thank you very much for your reply!\nSo if I understand correctly, for instance segmentation we have:\n\n- the groundtruth annotation format, which is described in [`challenge_evaluation.md`](https://github.com/tensorflow/models/blob/master/research/object_detection/g3doc/challenge_evaluation.md#instance-segmentation-track), which uses columns\n```ImageID, LabelName, ImageWidth, ImageHeight, XMin, YMin, XMax, YMax, IsGroupOf, Mask```\n\n- the predictions Kaggle submission format, described in the [Kaggle evaluation page](https://www.kaggle.com/c/open-images-2019-instance-segmentation/overview/evaluation), with columns\n```ImageID, ImageWidth, ImageHeight, PredictionString```\n\n- a prediction `Tensorflow/models` format, which is not described but can be evinced from [`oid_challenge_evaluation.py`](https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation.py) and [`oid_challenge_evaluation_utils.py`](https://github.com/tensorflow/models/blob/master/research/object_detection/metrics/oid_challenge_evaluation_utils.py), with columns\n```ImageID, ImageWidth, ImageHeight, LabelName, Score, Mask```\n\nI will test this ASAP. Thanks again!",
    "616005": "I was finally able to run the metrics!\nSome of the nitty-gritty details I had to figure out that might be useful to other participants:\n- before encoding them to COCO RLE, I had to resize the masks for ground-truth annotations to match height and width of the corresponding image (otherwise I get a nice seg fault);\n- after generating a compressed representation of masks as described in the official [gist](https://gist.github.com/pculliton/209398a2a52867580c6103e25e55d93c), if `pandas` is used to generate the `.csv` file, masks should be decoded to strings using `.decode('utf-8')` (otherwise they get printed out as byte-strings and generate errors when uncompressed);\n- the evaluation script is able to deal with different levels of `zlib` compression because the first bytes contain the information about compression level (not trivial for someone with as little experience as myself).",
    "616193": "Glad to hear!",
    "627641": "Hi, what about the speed of the evaluation scripts? Mine is too slow."
  },
  "source": "meta"
}