{
  "id": 100595,
  "title": "Unusually low public LB score",
  "url": "/competitions/open-images-2019-instance-segmentation/discussion/100595",
  "author_name": "",
  "post_date": "2019-07-19T12:10:53.650118300Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi. We are having a problem of getting unusually low test score with our submission. We have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nHowever, the same model only achieves 0.0001mAP on the public leaderboard.\nWe have tried to figure out the reason but no luck.\nHere is a script that writes the submission CSV file.\nWe call <code>write_instance_segmentation_to_csv</code> to write CSV file.\nCould someone (possibly the organizers) help us?</p>\n\n<p>Also, we noticed that the sizes of masks in <code>sample_truncated_submission.csv</code> are different from the sizes of the corresponding images. Is this related?</p>\n\n<p>```\nimport base64\nimport numpy as np\nfrom pycocotools import _mask as coco_mask\nimport zlib</p>\n\n<p>def _encode_binary_m(m):\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n    assert m.dtype == np.bool\n    assert len(m.shape) == 2</p>\n\n<pre><code># convert input mask to expected COCO API input\nm_to_encode = m.reshape(m.shape[0], m.shape[1], 1)\nm_to_encode = m_to_encode.astype(np.uint8)\nm_to_encode = np.asfortranarray(m_to_encode)\n\n# RLE encode mask\nencoded_m = coco_mask.encode(m_to_encode)[0][\"counts\"]\n\n# compress and base64 encoding\nbinary_str = zlib.compress(encoded_m, zlib.Z_BEST_COMPRESSION)\nbase64_str = base64.b64encode(binary_str)\nreturn base64_str.decode('utf-8')\n</code></pre>\n\n<p>def write_instance_segmentation_to_csv(\n        masks, labels, scores, img_ids, label_names, csv_path):\n    \"\"\"Write submission CSV</p>\n\n<pre><code>Args:\n    masks (iterable of ndarray): Iterable of arrays with shape (R, H, W).\n        There are N arrays (N corresponds to the number of images).\n        R is the number of instances for each image.\n    labels (iterabale of ndarray): Iterable of arrays with shape (R,).\n    scores (iterabale of ndarray): Iterable of arrays with shape (R,).\n    img_ids (iterabale of strings)\n    label_names (list): maps integer index to label name (e.g, /m/01g317).\n    csv_path (str)\n\n\"\"\"\nlines = ['ImageID,ImageWidth,ImageHeight,PredictionString']\n\nfor mask, label, score, img_id in zip(\n        masks, labels, scores, img_ids):\n    assert len(mask) == len(label) == len(score)\n    _, H, W = mask.shape\n    line = '{},{},{},'.format(img_id, W, H)\n    for m, lbl, sc in zip(mask, label, score):\n        encoded_m = _encode_binary_m(m)\n        lbl_id = label_names[lbl]\n        line += '{} {:.6f} {} '.format(lbl_id, sc, encoded_m)\n    lines.append(line)\n\nwith open(csv_path, 'w') as fw:\n    for line in lines:\n        fw.write('{}\\n'.format(line))\n</code></pre>\n\n<p>```</p>\n\n<p>Also, here is one of the rows in the submission file created from val set and the visualization of the prediction.</p>\n\n<p><code>\n0009bad4d8539bb4,1024,681,/m/0cmf2 0.9336771965026855 eNqFUkkOwjAM/JLtVhWlR4TgQGIJJDhy4QL/fwBxnMXpojZq4ozXjP3+vobx94HpdLweztNldMjEwOQRgOVPksoMEUU5F5+YbGFVN0PCESXGDJWdMeQpjuizI5nswSpsrIs4GEVJsFaH0Op4+YgNrLw4/1pLkljriek9esrpWvt0txqTT6rjWinrs2wIDa9rHqYySrzKMNhT+S6U23syKe6w0s/1zu8ol3hp5m7MzUFLXoXKMgu1961eOOtdHGjOLbH2aYa88W6iCI4t+7mDaQRkALpgQY64u/W3zpG/Px/DHyVMu5M= /m/0cmf2 0.37426382303237915 eNp1ket2ojAQgF9pgtX10m739KhduWSsvcglICRQcWvr+//bGUBBa+Gc5JtvJgkTPtblcHjnLnMYfUboHAyM9xGC+/DPQCjAxzeFgliuFVo7A56vsFe2sxso7L8bcMJ6tiOFg8LAIlZo20Ar+gmTKLt0s1HoLFpykCvHzLYtBfEoZXY85mGm0F3Y0iprdtBLL8kWXkKkycGJhKeIcs7KkKggAhHQl2yZLJ/OL4nwhsjaMQ3eDMAH069XA7gnkuMX6vOLafJMnR6YblfU6zImunuibp+Yfi8NPD7H6HhMf1fs7g+6craXAsl5hRmQnXGlzOBLw/QlRtfVbB9eGXP41PBnTTi3rT37BAOKZk5vzwuqCB5zziikw2f9iVNwJuJo3p/QFm1061RRiKtrUcGHBUjfMO0bsdWwkypGXxgNpYxjDITWsJVJjKHINLzLTZcikWooZHpOG1674bXkSjh7BB6HroTaCrTkZNbzLCkAKBJK0p+kDQ0YaRQqkTCRi9llkv59KnJzthmK04TQ8GmujqaNa1X7ZmyGdnlNlzGhhd2DrjycqXrsllwvP1oqv7yVH4qvl7W2kxfHhqF763zPx7dONZdCt9/zGg/tVUFjzmzHffenjLToHU1N6g/+AyYCLnE= /m/0k5j 0.054499007761478424 eNqVUtsOgjAM/aV2i4BfgIRAfRBffDLRRBP//9l1XccuqHELrDvrLef09VwaPN/B9rBcyD5usB+boR3avjt0424yZGbkDwCQ/8Sf3L3NOKmVrwr0gKLxBNR7+RQt9QGtV9/EE9mT9JTWjbRNKLh3QMo2h6RpoykvRVWfxCWVhOsG5QhiKQoNeG9FpT3atteiCnH/LkEaDJAUN06krI2KvoT4+CaUKulylmynEZvylFrU9z/Xh/AfMNKX0Qjkog7yhgjKnJmtk2qVLxcmDD7pDMVqOnGcJ/jIkEUdiRV0Ks08DHY6XY/NG9Icvsc= /m/083wq 0.09189055114984512 eNpLism1NglNNTDyM/Q39DcAAUMDCG3gbwilgWJ+cSn5FgACpAr8\n</code></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F413121%2F8c7be97f4de98c7b6f1047d5e5e0decb%2Faa.png?generation=1563538150709091&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "579929",
      "postDate": "07/19/2019 12:10:53",
      "content": "<p>Hi. We are having a problem of getting unusually low test score with our submission. We have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nHowever, the same model only achieves 0.0001mAP on the public leaderboard.\nWe have tried to figure out the reason but no luck.\nHere is a script that writes the submission CSV file.\nWe call <code>write_instance_segmentation_to_csv</code> to write CSV file.\nCould someone (possibly the organizers) help us?</p>\n\n<p>Also, we noticed that the sizes of masks in <code>sample_truncated_submission.csv</code> are different from the sizes of the corresponding images. Is this related?</p>\n\n<p>```\nimport base64\nimport numpy as np\nfrom pycocotools import _mask as coco_mask\nimport zlib</p>\n\n<p>def _encode_binary_m(m):\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n    assert m.dtype == np.bool\n    assert len(m.shape) == 2</p>\n\n<pre><code># convert input mask to expected COCO API input\nm_to_encode = m.reshape(m.shape[0], m.shape[1], 1)\nm_to_encode = m_to_encode.astype(np.uint8)\nm_to_encode = np.asfortranarray(m_to_encode)\n\n# RLE encode mask\nencoded_m = coco_mask.encode(m_to_encode)[0][\"counts\"]\n\n# compress and base64 encoding\nbinary_str = zlib.compress(encoded_m, zlib.Z_BEST_COMPRESSION)\nbase64_str = base64.b64encode(binary_str)\nreturn base64_str.decode('utf-8')\n</code></pre>\n\n<p>def write_instance_segmentation_to_csv(\n        masks, labels, scores, img_ids, label_names, csv_path):\n    \"\"\"Write submission CSV</p>\n\n<pre><code>Args:\n    masks (iterable of ndarray): Iterable of arrays with shape (R, H, W).\n        There are N arrays (N corresponds to the number of images).\n        R is the number of instances for each image.\n    labels (iterabale of ndarray): Iterable of arrays with shape (R,).\n    scores (iterabale of ndarray): Iterable of arrays with shape (R,).\n    img_ids (iterabale of strings)\n    label_names (list): maps integer index to label name (e.g, /m/01g317).\n    csv_path (str)\n\n\"\"\"\nlines = ['ImageID,ImageWidth,ImageHeight,PredictionString']\n\nfor mask, label, score, img_id in zip(\n        masks, labels, scores, img_ids):\n    assert len(mask) == len(label) == len(score)\n    _, H, W = mask.shape\n    line = '{},{},{},'.format(img_id, W, H)\n    for m, lbl, sc in zip(mask, label, score):\n        encoded_m = _encode_binary_m(m)\n        lbl_id = label_names[lbl]\n        line += '{} {:.6f} {} '.format(lbl_id, sc, encoded_m)\n    lines.append(line)\n\nwith open(csv_path, 'w') as fw:\n    for line in lines:\n        fw.write('{}\\n'.format(line))\n</code></pre>\n\n<p>```</p>\n\n<p>Also, here is one of the rows in the submission file created from val set and the visualization of the prediction.</p>\n\n<p><code>\n0009bad4d8539bb4,1024,681,/m/0cmf2 0.9336771965026855 eNqFUkkOwjAM/JLtVhWlR4TgQGIJJDhy4QL/fwBxnMXpojZq4ozXjP3+vobx94HpdLweztNldMjEwOQRgOVPksoMEUU5F5+YbGFVN0PCESXGDJWdMeQpjuizI5nswSpsrIs4GEVJsFaH0Op4+YgNrLw4/1pLkljriek9esrpWvt0txqTT6rjWinrs2wIDa9rHqYySrzKMNhT+S6U23syKe6w0s/1zu8ol3hp5m7MzUFLXoXKMgu1961eOOtdHGjOLbH2aYa88W6iCI4t+7mDaQRkALpgQY64u/W3zpG/Px/DHyVMu5M= /m/0cmf2 0.37426382303237915 eNp1ket2ojAQgF9pgtX10m739KhduWSsvcglICRQcWvr+//bGUBBa+Gc5JtvJgkTPtblcHjnLnMYfUboHAyM9xGC+/DPQCjAxzeFgliuFVo7A56vsFe2sxso7L8bcMJ6tiOFg8LAIlZo20Ar+gmTKLt0s1HoLFpykCvHzLYtBfEoZXY85mGm0F3Y0iprdtBLL8kWXkKkycGJhKeIcs7KkKggAhHQl2yZLJ/OL4nwhsjaMQ3eDMAH069XA7gnkuMX6vOLafJMnR6YblfU6zImunuibp+Yfi8NPD7H6HhMf1fs7g+6craXAsl5hRmQnXGlzOBLw/QlRtfVbB9eGXP41PBnTTi3rT37BAOKZk5vzwuqCB5zziikw2f9iVNwJuJo3p/QFm1061RRiKtrUcGHBUjfMO0bsdWwkypGXxgNpYxjDITWsJVJjKHINLzLTZcikWooZHpOG1674bXkSjh7BB6HroTaCrTkZNbzLCkAKBJK0p+kDQ0YaRQqkTCRi9llkv59KnJzthmK04TQ8GmujqaNa1X7ZmyGdnlNlzGhhd2DrjycqXrsllwvP1oqv7yVH4qvl7W2kxfHhqF763zPx7dONZdCt9/zGg/tVUFjzmzHffenjLToHU1N6g/+AyYCLnE= /m/0k5j 0.054499007761478424 eNqVUtsOgjAM/aV2i4BfgIRAfRBffDLRRBP//9l1XccuqHELrDvrLef09VwaPN/B9rBcyD5usB+boR3avjt0424yZGbkDwCQ/8Sf3L3NOKmVrwr0gKLxBNR7+RQt9QGtV9/EE9mT9JTWjbRNKLh3QMo2h6RpoykvRVWfxCWVhOsG5QhiKQoNeG9FpT3atteiCnH/LkEaDJAUN06krI2KvoT4+CaUKulylmynEZvylFrU9z/Xh/AfMNKX0Qjkog7yhgjKnJmtk2qVLxcmDD7pDMVqOnGcJ/jIkEUdiRV0Ks08DHY6XY/NG9Icvsc= /m/083wq 0.09189055114984512 eNpLism1NglNNTDyM/Q39DcAAUMDCG3gbwilgWJ+cSn5FgACpAr8\n</code></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F413121%2F8c7be97f4de98c7b6f1047d5e5e0decb%2Faa.png?generation=1563538150709091&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi. We are having a problem of getting unusually low test score with our submission. We have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nHowever, the same model only achieves 0.0001mAP on the public leaderboard.\nWe have tried to figure out the reason but no luck.\nHere is a script that writes the submission CSV file.\nWe call `write_instance_segmentation_to_csv` to write CSV file.\nCould someone (possibly the organizers) help us?\n\nAlso, we noticed that the sizes of masks in `sample_truncated_submission.csv` are different from the sizes of the corresponding images. Is this related?\n\n```\nimport base64\nimport numpy as np\nfrom pycocotools import _mask as coco_mask\nimport zlib\n\n\ndef _encode_binary_m(m):\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n    assert m.dtype == np.bool\n    assert len(m.shape) == 2\n\n    # convert input mask to expected COCO API input\n    m_to_encode = m.reshape(m.shape[0], m.shape[1], 1)\n    m_to_encode = m_to_encode.astype(np.uint8)\n    m_to_encode = np.asfortranarray(m_to_encode)\n\n    # RLE encode mask\n    encoded_m = coco_mask.encode(m_to_encode)[0][\"counts\"]\n\n    # compress and base64 encoding\n    binary_str = zlib.compress(encoded_m, zlib.Z_BEST_COMPRESSION)\n    base64_str = base64.b64encode(binary_str)\n    return base64_str.decode('utf-8')\n\n\ndef write_instance_segmentation_to_csv(\n        masks, labels, scores, img_ids, label_names, csv_path):\n    \"\"\"Write submission CSV\n\n    Args:\n        masks (iterable of ndarray): Iterable of arrays with shape (R, H, W).\n            There are N arrays (N corresponds to the number of images).\n            R is the number of instances for each image.\n        labels (iterabale of ndarray): Iterable of arrays with shape (R,).\n        scores (iterabale of ndarray): Iterable of arrays with shape (R,).\n        img_ids (iterabale of strings)\n        label_names (list): maps integer index to label name (e.g, /m/01g317).\n        csv_path (str)\n\n    \"\"\"\n    lines = ['ImageID,ImageWidth,ImageHeight,PredictionString']\n\n    for mask, label, score, img_id in zip(\n            masks, labels, scores, img_ids):\n        assert len(mask) == len(label) == len(score)\n        _, H, W = mask.shape\n        line = '{},{},{},'.format(img_id, W, H)\n        for m, lbl, sc in zip(mask, label, score):\n            encoded_m = _encode_binary_m(m)\n            lbl_id = label_names[lbl]\n            line += '{} {:.6f} {} '.format(lbl_id, sc, encoded_m)\n        lines.append(line)\n\n    with open(csv_path, 'w') as fw:\n        for line in lines:\n            fw.write('{}\\n'.format(line))\n\n```\n\nAlso, here is one of the rows in the submission file created from val set and the visualization of the prediction.\n\n```\n0009bad4d8539bb4,1024,681,/m/0cmf2 0.9336771965026855 eNqFUkkOwjAM/JLtVhWlR4TgQGIJJDhy4QL/fwBxnMXpojZq4ozXjP3+vobx94HpdLweztNldMjEwOQRgOVPksoMEUU5F5+YbGFVN0PCESXGDJWdMeQpjuizI5nswSpsrIs4GEVJsFaH0Op4+YgNrLw4/1pLkljriek9esrpWvt0txqTT6rjWinrs2wIDa9rHqYySrzKMNhT+S6U23syKe6w0s/1zu8ol3hp5m7MzUFLXoXKMgu1961eOOtdHGjOLbH2aYa88W6iCI4t+7mDaQRkALpgQY64u/W3zpG/Px/DHyVMu5M= /m/0cmf2 0.37426382303237915 eNp1ket2ojAQgF9pgtX10m739KhduWSsvcglICRQcWvr+//bGUBBa+Gc5JtvJgkTPtblcHjnLnMYfUboHAyM9xGC+/DPQCjAxzeFgliuFVo7A56vsFe2sxso7L8bcMJ6tiOFg8LAIlZo20Ar+gmTKLt0s1HoLFpykCvHzLYtBfEoZXY85mGm0F3Y0iprdtBLL8kWXkKkycGJhKeIcs7KkKggAhHQl2yZLJ/OL4nwhsjaMQ3eDMAH069XA7gnkuMX6vOLafJMnR6YblfU6zImunuibp+Yfi8NPD7H6HhMf1fs7g+6craXAsl5hRmQnXGlzOBLw/QlRtfVbB9eGXP41PBnTTi3rT37BAOKZk5vzwuqCB5zziikw2f9iVNwJuJo3p/QFm1061RRiKtrUcGHBUjfMO0bsdWwkypGXxgNpYxjDITWsJVJjKHINLzLTZcikWooZHpOG1674bXkSjh7BB6HroTaCrTkZNbzLCkAKBJK0p+kDQ0YaRQqkTCRi9llkv59KnJzthmK04TQ8GmujqaNa1X7ZmyGdnlNlzGhhd2DrjycqXrsllwvP1oqv7yVH4qvl7W2kxfHhqF763zPx7dONZdCt9/zGg/tVUFjzmzHffenjLToHU1N6g/+AyYCLnE= /m/0k5j 0.054499007761478424 eNqVUtsOgjAM/aV2i4BfgIRAfRBffDLRRBP//9l1XccuqHELrDvrLef09VwaPN/B9rBcyD5usB+boR3avjt0424yZGbkDwCQ/8Sf3L3NOKmVrwr0gKLxBNR7+RQt9QGtV9/EE9mT9JTWjbRNKLh3QMo2h6RpoykvRVWfxCWVhOsG5QhiKQoNeG9FpT3atteiCnH/LkEaDJAUN06krI2KvoT4+CaUKulylmynEZvylFrU9z/Xh/AfMNKX0Qjkog7yhgjKnJmtk2qVLxcmDD7pDMVqOnGcJ/jIkEUdiRV0Ks08DHY6XY/NG9Icvsc= /m/083wq 0.09189055114984512 eNpLism1NglNNTDyM/Q39DcAAUMDCG3gbwilgWJ+cSn5FgACpAr8\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F413121%2F8c7be97f4de98c7b6f1047d5e5e0decb%2Faa.png?generation=1563538150709091&amp;alt=media)",
      "votes": null
    },
    {
      "id": "579970",
      "postDate": "07/19/2019 13:35:11",
      "content": "<p>Thanks for reporting these issues, we will look into them. \nDue to time-zones differences, expect a status update early next week.</p>",
      "rawMarkdown": "Thanks for reporting these issues, we will look into them. \nDue to time-zones differences, expect a status update early next week.",
      "votes": null
    },
    {
      "id": "580011",
      "postDate": "07/19/2019 14:40:23",
      "content": "<p>Quick update:\nWe have inspected the data and we can confirm we have found an issue.\nWe have established a plan to fix the problem, and we should be able to deploy it early next week.</p>\n\n<p>We will update this thread as things evolve. Thank you for your patience.</p>",
      "rawMarkdown": "Quick update:\nWe have inspected the data and we can confirm we have found an issue.\nWe have established a plan to fix the problem, and we should be able to deploy it early next week.\n\nWe will update this thread as things evolve. Thank you for your patience.",
      "votes": null
    },
    {
      "id": "580257",
      "postDate": "07/19/2019 22:21:40",
      "content": "<p>Thank you for checking the data. I will wait to hear from you.</p>",
      "rawMarkdown": "Thank you for checking the data. I will wait to hear from you.",
      "votes": null
    },
    {
      "id": "580285",
      "postDate": "07/20/2019 00:41:50",
      "content": "<blockquote>\n  <p>I will wait to hear from you.</p>\n</blockquote>\n\n<p>In the mean time you can keep working towards preparing submission files where the detection binary masks have the same resolution as the corresponding test set images.</p>\n\n<p>The current plan of action will lead to an update on the leaderboard scores (early next week), without change on the submission format for the participants.</p>",
      "rawMarkdown": "&gt; I will wait to hear from you.\n\nIn the mean time you can keep working towards preparing submission files where the detection binary masks have the same resolution as the corresponding test set images.\n\nThe current plan of action will lead to an update on the leaderboard scores (early next week), without change on the submission format for the participants.",
      "votes": null
    },
    {
      "id": "582108",
      "postDate": "07/22/2019 18:25:12",
      "content": "<p>Update:</p>\n\n<p>The issue has now been resolved and the <a href=\"https://www.kaggle.com/c/open-images-2019-instance-segmentation/leaderboard\">leaderboard</a> updated.\nNow 6 team have a mAP &gt; 10%, and the values observed seem closer to the expected range of results quality.</p>\n\n<p>Thanks again for having reported this issue.\nLooking forward to see how the teams will push upwards these results !</p>",
      "rawMarkdown": "Update:\n\nThe issue has now been resolved and the [leaderboard](https://www.kaggle.com/c/open-images-2019-instance-segmentation/leaderboard) updated.\nNow 6 team have a mAP &gt; 10%, and the values observed seem closer to the expected range of results quality.\n\nThanks again for having reported this issue.\nLooking forward to see how the teams will push upwards these results !",
      "votes": null
    },
    {
      "id": "582225",
      "postDate": "07/22/2019 23:25:20",
      "content": "<p>Thank you for fixing the evaluation. The evaluation score looks good!</p>",
      "rawMarkdown": "Thank you for fixing the evaluation. The evaluation score looks good!",
      "votes": null
    },
    {
      "id": "609842",
      "postDate": "08/28/2019 06:49:00",
      "content": "<p>hi, \nWe have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nCan you tell me how to use the using official evaluation script on local validation?</p>\n\n<p>I can get some information from:\nThe download consists of a set of .zip archives containing binary .png masks. Those should be transformed into a single CSV file in the format:</p>\n\n<p>ImageID,LabelName,ImageWidth,ImageHeight,XMin,YMin,XMax,YMax,GroupOf,Mask where Mask is MS COCO RLE encoding of a binary mask stored in .png file.</p>\n\n<p>NOTE: the util to make the transformation will be released soon.</p>\n\n<p>however: i don't know how to produce the CSV file, So I try to implement it myself…, but some error will appear.</p>\n\n<p>eg:\nFirst of all, let me point out that - to my understanding - the column named GroupOf should be called IsGroupOf. Even after that, I am still getting a KeyError: 'LabelName' in\nFile \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),</p>\n\n<p>thanks so much!</p>",
      "rawMarkdown": "hi, \nWe have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nCan you tell me how to use the using official evaluation script on local validation?\n\nI can get some information from:\nThe download consists of a set of .zip archives containing binary .png masks. Those should be transformed into a single CSV file in the format:\n\nImageID,LabelName,ImageWidth,ImageHeight,XMin,YMin,XMax,YMax,GroupOf,Mask where Mask is MS COCO RLE encoding of a binary mask stored in .png file.\n\nNOTE: the util to make the transformation will be released soon.\n\nhowever: i don't know how to produce the CSV file, So I try to implement it myself…, but some error will appear.\n\neg:\nFirst of all, let me point out that - to my understanding - the column named GroupOf should be called IsGroupOf. Even after that, I am still getting a KeyError: 'LabelName' in\nFile \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),\n\nthanks so much!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 579970,
      "author_name": "benenson",
      "author_url": "",
      "post_date": "07/19/2019 13:35:11",
      "content": "<p>Thanks for reporting these issues, we will look into them. \nDue to time-zones differences, expect a status update early next week.</p>",
      "votes": null,
      "replies": [
        {
          "id": 580011,
          "author_name": "benenson",
          "author_url": "",
          "post_date": "07/19/2019 14:40:23",
          "content": "<p>Quick update:\nWe have inspected the data and we can confirm we have found an issue.\nWe have established a plan to fix the problem, and we should be able to deploy it early next week.</p>\n\n<p>We will update this thread as things evolve. Thank you for your patience.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 580257,
          "author_name": "yuyu2172",
          "author_url": "",
          "post_date": "07/19/2019 22:21:40",
          "content": "<p>Thank you for checking the data. I will wait to hear from you.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 580285,
          "author_name": "benenson",
          "author_url": "",
          "post_date": "07/20/2019 00:41:50",
          "content": "<blockquote>\n  <p>I will wait to hear from you.</p>\n</blockquote>\n\n<p>In the mean time you can keep working towards preparing submission files where the detection binary masks have the same resolution as the corresponding test set images.</p>\n\n<p>The current plan of action will lead to an update on the leaderboard scores (early next week), without change on the submission format for the participants.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 582108,
          "author_name": "benenson",
          "author_url": "",
          "post_date": "07/22/2019 18:25:12",
          "content": "<p>Update:</p>\n\n<p>The issue has now been resolved and the <a href=\"https://www.kaggle.com/c/open-images-2019-instance-segmentation/leaderboard\">leaderboard</a> updated.\nNow 6 team have a mAP &gt; 10%, and the values observed seem closer to the expected range of results quality.</p>\n\n<p>Thanks again for having reported this issue.\nLooking forward to see how the teams will push upwards these results !</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 582225,
          "author_name": "yuyu2172",
          "author_url": "",
          "post_date": "07/22/2019 23:25:20",
          "content": "<p>Thank you for fixing the evaluation. The evaluation score looks good!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 609842,
      "author_name": "summery2015",
      "author_url": "",
      "post_date": "08/28/2019 06:49:00",
      "content": "<p>hi, \nWe have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nCan you tell me how to use the using official evaluation script on local validation?</p>\n\n<p>I can get some information from:\nThe download consists of a set of .zip archives containing binary .png masks. Those should be transformed into a single CSV file in the format:</p>\n\n<p>ImageID,LabelName,ImageWidth,ImageHeight,XMin,YMin,XMax,YMax,GroupOf,Mask where Mask is MS COCO RLE encoding of a binary mask stored in .png file.</p>\n\n<p>NOTE: the util to make the transformation will be released soon.</p>\n\n<p>however: i don't know how to produce the CSV file, So I try to implement it myself…, but some error will appear.</p>\n\n<p>eg:\nFirst of all, let me point out that - to my understanding - the column named GroupOf should be called IsGroupOf. Even after that, I am still getting a KeyError: 'LabelName' in\nFile \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),</p>\n\n<p>thanks so much!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "579929": "Hi. We are having a problem of getting unusually low test score with our submission. We have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nHowever, the same model only achieves 0.0001mAP on the public leaderboard.\nWe have tried to figure out the reason but no luck.\nHere is a script that writes the submission CSV file.\nWe call `write_instance_segmentation_to_csv` to write CSV file.\nCould someone (possibly the organizers) help us?\n\nAlso, we noticed that the sizes of masks in `sample_truncated_submission.csv` are different from the sizes of the corresponding images. Is this related?\n\n```\nimport base64\nimport numpy as np\nfrom pycocotools import _mask as coco_mask\nimport zlib\n\n\ndef _encode_binary_m(m):\n    \"\"\"Converts a binary mask into OID challenge encoding ascii text.\"\"\"\n    assert m.dtype == np.bool\n    assert len(m.shape) == 2\n\n    # convert input mask to expected COCO API input\n    m_to_encode = m.reshape(m.shape[0], m.shape[1], 1)\n    m_to_encode = m_to_encode.astype(np.uint8)\n    m_to_encode = np.asfortranarray(m_to_encode)\n\n    # RLE encode mask\n    encoded_m = coco_mask.encode(m_to_encode)[0][\"counts\"]\n\n    # compress and base64 encoding\n    binary_str = zlib.compress(encoded_m, zlib.Z_BEST_COMPRESSION)\n    base64_str = base64.b64encode(binary_str)\n    return base64_str.decode('utf-8')\n\n\ndef write_instance_segmentation_to_csv(\n        masks, labels, scores, img_ids, label_names, csv_path):\n    \"\"\"Write submission CSV\n\n    Args:\n        masks (iterable of ndarray): Iterable of arrays with shape (R, H, W).\n            There are N arrays (N corresponds to the number of images).\n            R is the number of instances for each image.\n        labels (iterabale of ndarray): Iterable of arrays with shape (R,).\n        scores (iterabale of ndarray): Iterable of arrays with shape (R,).\n        img_ids (iterabale of strings)\n        label_names (list): maps integer index to label name (e.g, /m/01g317).\n        csv_path (str)\n\n    \"\"\"\n    lines = ['ImageID,ImageWidth,ImageHeight,PredictionString']\n\n    for mask, label, score, img_id in zip(\n            masks, labels, scores, img_ids):\n        assert len(mask) == len(label) == len(score)\n        _, H, W = mask.shape\n        line = '{},{},{},'.format(img_id, W, H)\n        for m, lbl, sc in zip(mask, label, score):\n            encoded_m = _encode_binary_m(m)\n            lbl_id = label_names[lbl]\n            line += '{} {:.6f} {} '.format(lbl_id, sc, encoded_m)\n        lines.append(line)\n\n    with open(csv_path, 'w') as fw:\n        for line in lines:\n            fw.write('{}\\n'.format(line))\n\n```\n\nAlso, here is one of the rows in the submission file created from val set and the visualization of the prediction.\n\n```\n0009bad4d8539bb4,1024,681,/m/0cmf2 0.9336771965026855 eNqFUkkOwjAM/JLtVhWlR4TgQGIJJDhy4QL/fwBxnMXpojZq4ozXjP3+vobx94HpdLweztNldMjEwOQRgOVPksoMEUU5F5+YbGFVN0PCESXGDJWdMeQpjuizI5nswSpsrIs4GEVJsFaH0Op4+YgNrLw4/1pLkljriek9esrpWvt0txqTT6rjWinrs2wIDa9rHqYySrzKMNhT+S6U23syKe6w0s/1zu8ol3hp5m7MzUFLXoXKMgu1961eOOtdHGjOLbH2aYa88W6iCI4t+7mDaQRkALpgQY64u/W3zpG/Px/DHyVMu5M= /m/0cmf2 0.37426382303237915 eNp1ket2ojAQgF9pgtX10m739KhduWSsvcglICRQcWvr+//bGUBBa+Gc5JtvJgkTPtblcHjnLnMYfUboHAyM9xGC+/DPQCjAxzeFgliuFVo7A56vsFe2sxso7L8bcMJ6tiOFg8LAIlZo20Ar+gmTKLt0s1HoLFpykCvHzLYtBfEoZXY85mGm0F3Y0iprdtBLL8kWXkKkycGJhKeIcs7KkKggAhHQl2yZLJ/OL4nwhsjaMQ3eDMAH069XA7gnkuMX6vOLafJMnR6YblfU6zImunuibp+Yfi8NPD7H6HhMf1fs7g+6craXAsl5hRmQnXGlzOBLw/QlRtfVbB9eGXP41PBnTTi3rT37BAOKZk5vzwuqCB5zziikw2f9iVNwJuJo3p/QFm1061RRiKtrUcGHBUjfMO0bsdWwkypGXxgNpYxjDITWsJVJjKHINLzLTZcikWooZHpOG1674bXkSjh7BB6HroTaCrTkZNbzLCkAKBJK0p+kDQ0YaRQqkTCRi9llkv59KnJzthmK04TQ8GmujqaNa1X7ZmyGdnlNlzGhhd2DrjycqXrsllwvP1oqv7yVH4qvl7W2kxfHhqF763zPx7dONZdCt9/zGg/tVUFjzmzHffenjLToHU1N6g/+AyYCLnE= /m/0k5j 0.054499007761478424 eNqVUtsOgjAM/aV2i4BfgIRAfRBffDLRRBP//9l1XccuqHELrDvrLef09VwaPN/B9rBcyD5usB+boR3avjt0424yZGbkDwCQ/8Sf3L3NOKmVrwr0gKLxBNR7+RQt9QGtV9/EE9mT9JTWjbRNKLh3QMo2h6RpoykvRVWfxCWVhOsG5QhiKQoNeG9FpT3atteiCnH/LkEaDJAUN06krI2KvoT4+CaUKulylmynEZvylFrU9z/Xh/AfMNKX0Qjkog7yhgjKnJmtk2qVLxcmDD7pDMVqOnGcJ/jIkEUdiRV0Ks08DHY6XY/NG9Icvsc= /m/083wq 0.09189055114984512 eNpLism1NglNNTDyM/Q39DcAAUMDCG3gbwilgWJ+cSn5FgACpAr8\n```\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F413121%2F8c7be97f4de98c7b6f1047d5e5e0decb%2Faa.png?generation=1563538150709091&amp;alt=media)",
    "579970": "Thanks for reporting these issues, we will look into them. \nDue to time-zones differences, expect a status update early next week.",
    "580011": "Quick update:\nWe have inspected the data and we can confirm we have found an issue.\nWe have established a plan to fix the problem, and we should be able to deploy it early next week.\n\nWe will update this thread as things evolve. Thank you for your patience.",
    "580257": "Thank you for checking the data. I will wait to hear from you.",
    "580285": "&gt; I will wait to hear from you.\n\nIn the mean time you can keep working towards preparing submission files where the detection binary masks have the same resolution as the corresponding test set images.\n\nThe current plan of action will lead to an update on the leaderboard scores (early next week), without change on the submission format for the participants.",
    "582108": "Update:\n\nThe issue has now been resolved and the [leaderboard](https://www.kaggle.com/c/open-images-2019-instance-segmentation/leaderboard) updated.\nNow 6 team have a mAP &gt; 10%, and the values observed seem closer to the expected range of results quality.\n\nThanks again for having reported this issue.\nLooking forward to see how the teams will push upwards these results !",
    "582225": "Thank you for fixing the evaluation. The evaluation score looks good!",
    "609842": "hi, \nWe have prepared a decently working model that achieves 0.30mAP on the validation set evaluated using official evaluation script.\nCan you tell me how to use the using official evaluation script on local validation?\n\nI can get some information from:\nThe download consists of a set of .zip archives containing binary .png masks. Those should be transformed into a single CSV file in the format:\n\nImageID,LabelName,ImageWidth,ImageHeight,XMin,YMin,XMax,YMax,GroupOf,Mask where Mask is MS COCO RLE encoding of a binary mask stored in .png file.\n\nNOTE: the util to make the transformation will be released soon.\n\nhowever: i don't know how to produce the CSV file, So I try to implement it myself…, but some error will appear.\n\neg:\nFirst of all, let me point out that - to my understanding - the column named GroupOf should be called IsGroupOf. Even after that, I am still getting a KeyError: 'LabelName' in\nFile \"/content/models/research/object_detection/metrics/oid_challenge_evaluation_utils.py\", line 182, in build_predictions_dictionary data['LabelName'].map(lambda x: class_label_map[x]).as_matrix(),\n\nthanks so much!"
  },
  "source": "meta"
}