{
  "id": 53776,
  "title": "Welcome",
  "url": "/competitions/cvpr-2018-autonomous-driving/discussion/53776",
  "author_name": "",
  "post_date": "2018-04-05T00:10:25.270530Z",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Welcome to the Video Segmentation Challenge from 2018 CVPR Workshop on Autonomous Driving (<a href=\"http://www.wad.ai/\">www.wad.ai</a>). We are excited for to host this competition, which provides a challenging data set of 60K street-view images with per-pixel semantic annotation, courtesy of Baidu's <a href=\"http://apolloscape.auto/\">ApolloScape</a> project.  An even bigger set containing 140K images is available there.  </p>\n\n<p>In this challenge, you are asked to segment out each instance of different objects in the test video.\nThe video frames exhibit challenging cases of large lighting variations, heavy occlusions, and busy background. To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. </p>\n\n<p>Please feel free to ask your questions in this thread.\nGood luck!</p>",
  "messages": [
    {
      "id": "309278",
      "postDate": "04/05/2018 00:10:25",
      "content": "<p>Welcome to the Video Segmentation Challenge from 2018 CVPR Workshop on Autonomous Driving (<a href=\"http://www.wad.ai/\">www.wad.ai</a>). We are excited for to host this competition, which provides a challenging data set of 60K street-view images with per-pixel semantic annotation, courtesy of Baidu's <a href=\"http://apolloscape.auto/\">ApolloScape</a> project.  An even bigger set containing 140K images is available there.  </p>\n\n<p>In this challenge, you are asked to segment out each instance of different objects in the test video.\nThe video frames exhibit challenging cases of large lighting variations, heavy occlusions, and busy background. To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. </p>\n\n<p>Please feel free to ask your questions in this thread.\nGood luck!</p>",
      "rawMarkdown": "Welcome to the Video Segmentation Challenge from 2018 CVPR Workshop on Autonomous Driving ([www.wad.ai][1]). We are excited for to host this competition, which provides a challenging data set of 60K street-view images with per-pixel semantic annotation, courtesy of Baidu's [ApolloScape][2] project.  An even bigger set containing 140K images is available there.  \n\nIn this challenge, you are asked to segment out each instance of different objects in the test video.\nThe video frames exhibit challenging cases of large lighting variations, heavy occlusions, and busy background. To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. \n\nPlease feel free to ask your questions in this thread.\nGood luck!\n\n\n  [1]: http://www.wad.ai\n  [2]: http://apolloscape.auto",
      "votes": null
    },
    {
      "id": "310173",
      "postDate": "04/06/2018 18:10:22",
      "content": "<p>&gt; We are excited for to host this competition, which provides a challenging data set of 800K street-view images with per-pixel semantic annotation, courtesy of Baidu's ApolloScape project.</p>\n\n<p>It looks like there are only 39k images in train labels. I still haven't finished downloading full train data but it looks like there is also less then 800k images. Are other 761k images hosted elsewhere or am I missing something obvious?</p>",
      "rawMarkdown": "&gt; We are excited for to host this competition, which provides a challenging data set of 800K street-view images with per-pixel semantic annotation, courtesy of Baidu's ApolloScape project.\n\nIt looks like there are only 39k images in train labels. I still haven't finished downloading full train data but it looks like there is also less then 800k images. Are other 761k images hosted elsewhere or am I missing something obvious?",
      "votes": null
    },
    {
      "id": "310471",
      "postDate": "04/07/2018 16:45:44",
      "content": "<p>is there any requirement to develop a real-time algorithm?</p>",
      "rawMarkdown": "is there any requirement to develop a real-time algorithm?",
      "votes": null
    },
    {
      "id": "310479",
      "postDate": "04/07/2018 17:05:37",
      "content": "<p>Sorry， I made a typo in the initial introduciton.  We have uploaded to Kaggle about 60K images, we have a richer set at Apolloscape.auto that has 140K iamges. </p>",
      "rawMarkdown": "Sorry， I made a typo in the initial introduciton.  We have uploaded to Kaggle about 60K images, we have a richer set at Apolloscape.auto that has 140K iamges.",
      "votes": null
    },
    {
      "id": "310480",
      "postDate": "04/07/2018 17:05:50",
      "content": "<p>No, the emphasis is on accuracy</p>",
      "rawMarkdown": "No, the emphasis is on accuracy",
      "votes": null
    },
    {
      "id": "310500",
      "postDate": "04/07/2018 18:13:08",
      "content": "<p>Hi Ruigang Yang, sorry for double posting, but could you please have a look at <a href=\"https://www.kaggle.com/c/cvpr-2018-autonomous-driving/discussion/53845\">https://www.kaggle.com/c/cvpr-2018-autonomous-driving/discussion/53845</a> - I suspect that train labels were uploaded in a wrong format.</p>",
      "rawMarkdown": "Hi Ruigang Yang, sorry for double posting, but could you please have a look at https://www.kaggle.com/c/cvpr-2018-autonomous-driving/discussion/53845 - I suspect that train labels were uploaded in a wrong format.",
      "votes": null
    },
    {
      "id": "312542",
      "postDate": "04/12/2018 00:56:34",
      "content": "<p>We have submitted some results before, but there are some problems. The shape of our predicted mask is (H, W), and the first time we  use numpy.ndarry.flatten directly, then get the Length Encoding format and submit it successfully. After that, we realize that the pixels should be numbered from top to bottom, then left to right. So we transpose the array (numpy.ndarry.T) before use numpy.ndarry.flatten and submit it again but get the errors as: \nEvaluation Exception: Sequence contains no elements\nWe have no idea what happens as we just transpose the array. \nBy the way, why does the leaderboard display our score with others is zero?</p>",
      "rawMarkdown": "We have submitted some results before, but there are some problems. The shape of our predicted mask is (H, W), and the first time we  use numpy.ndarry.flatten directly, then get the Length Encoding format and submit it successfully. After that, we realize that the pixels should be numbered from top to bottom, then left to right. So we transpose the array (numpy.ndarry.T) before use numpy.ndarry.flatten and submit it again but get the errors as: \nEvaluation Exception: Sequence contains no elements\nWe have no idea what happens as we just transpose the array. \nBy the way, why does the leaderboard display our score with others is zero?",
      "votes": null
    },
    {
      "id": "312974",
      "postDate": "04/12/2018 15:51:49",
      "content": "<p>&gt; To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. </p>\n\n<p>Test set files are named as <code>0aa38f3841f7aa778c0e2cdd6a29dc1a.jpg</code>, does this mean that frame order is not explicitly provided for the test set, so if we want to use information from neighbor video frames, we'll have to recover them ourselves? Same question for left/right camera?</p>",
      "rawMarkdown": "&gt; To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. \n\nTest set files are named as `0aa38f3841f7aa778c0e2cdd6a29dc1a.jpg`, does this mean that frame order is not explicitly provided for the test set, so if we want to use information from neighbor video frames, we'll have to recover them ourselves? Same question for left/right camera?",
      "votes": null
    },
    {
      "id": "313272",
      "postDate": "04/13/2018 03:10:29",
      "content": "<p>Hi, Konstantin, I uploaded a zip file in the Data session that contains video lists and a mapping from md5 to the original timestamp for the testing set. Please check if it meets your needs.</p>",
      "rawMarkdown": "Hi, Konstantin, I uploaded a zip file in the Data session that contains video lists and a mapping from md5 to the original timestamp for the testing set. Please check if it meets your needs.",
      "votes": null
    },
    {
      "id": "313383",
      "postDate": "04/13/2018 07:31:09",
      "content": "<p>Hi Xinyu, this is exactly what I wanted, thanks! I checked test ids, there are 10 that are missing from the mapping files:</p>\n\n<pre><code>['30226074f61ab422d3f0ddbb284fa852',\n '9df73c4dd7d6efab6e39c9e91f4fe3ae',\n '12eeb997abcf94235b65f663d2cff198',\n 'abffd3db151bca65c1a72281c98592f2',\n 'b0a81467ffa2b260fa84fdda78d245a9',\n '0d5920509771a8d99dcfc265ce046df6',\n 'b0a539102aba1dee708d14edaf499bc0',\n 'e714994348b85004d09a447b2e0d5fc0',\n 'fa23f779c6c842dc2ad6445b41650707',\n 'c745d4d27662374bebc0ab0d1cd31dd6']\n</code></pre>\n\n<p>This is also visible in <code>wc -l</code>: it gives 1907 lines for mappings, while I see 1917 test files.</p>\n\n<p>It would be awesome if you could provide mappings for these 10 files as well.</p>",
      "rawMarkdown": "Hi Xinyu, this is exactly what I wanted, thanks! I checked test ids, there are 10 that are missing from the mapping files:\n\n    ['30226074f61ab422d3f0ddbb284fa852',\n     '9df73c4dd7d6efab6e39c9e91f4fe3ae',\n     '12eeb997abcf94235b65f663d2cff198',\n     'abffd3db151bca65c1a72281c98592f2',\n     'b0a81467ffa2b260fa84fdda78d245a9',\n     '0d5920509771a8d99dcfc265ce046df6',\n     'b0a539102aba1dee708d14edaf499bc0',\n     'e714994348b85004d09a447b2e0d5fc0',\n     'fa23f779c6c842dc2ad6445b41650707',\n     'c745d4d27662374bebc0ab0d1cd31dd6']\n\nThis is also visible in `wc -l`: it gives 1907 lines for mappings, while I see 1917 test files.\n\nIt would be awesome if you could provide mappings for these 10 files as well.",
      "votes": null
    },
    {
      "id": "314335",
      "postDate": "04/15/2018 09:21:08",
      "content": "<p>Hi Xinyu Huang, I'd like to clarify how evaluation metric is calculated - I notice that submissions must contain the confidence field, but it seems unused by the evaluation, is this correct? This is the relevant bit I believe:</p>\n\n<blockquote>\n  <p>If there are multiple predicted instances matched to a ground truth instance, the predicted instance with the largest IoU is considered as the true positive, and remaining predicted instances are false positives. </p>\n</blockquote>\n\n<p>This means that matches are selected using only IoU, not confidence (as also often done for instance segmentation), right?</p>",
      "rawMarkdown": "Hi Xinyu Huang, I'd like to clarify how evaluation metric is calculated - I notice that submissions must contain the confidence field, but it seems unused by the evaluation, is this correct? This is the relevant bit I believe:\n\n&gt; If there are multiple predicted instances matched to a ground truth instance, the predicted instance with the largest IoU is considered as the true positive, and remaining predicted instances are false positives. \n\nThis means that matches are selected using only IoU, not confidence (as also often done for instance segmentation), right?",
      "votes": null
    },
    {
      "id": "321898",
      "postDate": "05/02/2018 05:37:14",
      "content": "<p>Thanks for pointing out. We have updated the description to \"If there are multiple predicted instances matched to a ground truth instance, the predicted instance that is larger than the IoU threshold and has the largest confidence is considered as the true positive, and remaining predicted instances are false positives. \"</p>",
      "rawMarkdown": "Thanks for pointing out. We have updated the description to \"If there are multiple predicted instances matched to a ground truth instance, the predicted instance that is larger than the IoU threshold and has the largest confidence is considered as the true positive, and remaining predicted instances are false positives. \"",
      "votes": null
    },
    {
      "id": "321903",
      "postDate": "05/02/2018 05:52:19",
      "content": "<p>Unless you allow overlapping predictions, a matching IoU &gt; 0.5 means 1) the prediction cannot match to any other ground-truth with IoU &gt; 0.5 and 2) the ground truth region cannot match to any other prediction with IoU &gt; 0.5.   So in a practical sense (excluding extremely rare cases when 0.5+05=1), confidence score is useless here.</p>",
      "rawMarkdown": "Unless you allow overlapping predictions, a matching IoU &gt; 0.5 means 1) the prediction cannot match to any other ground-truth with IoU &gt; 0.5 and 2) the ground truth region cannot match to any other prediction with IoU &gt; 0.5.   So in a practical sense (excluding extremely rare cases when 0.5+05=1), confidence score is useless here.",
      "votes": null
    },
    {
      "id": "324043",
      "postDate": "05/07/2018 02:37:43",
      "content": "<p>In our evaluation, a predicted result will not be matched to more than one ground truth instance. A predicted result with IoU larger than a threshold is either matched or not-matched. This is based on the confidence scores.</p>",
      "rawMarkdown": "In our evaluation, a predicted result will not be matched to more than one ground truth instance. A predicted result with IoU larger than a threshold is either matched or not-matched. This is based on the confidence scores.",
      "votes": null
    },
    {
      "id": "334776",
      "postDate": "05/28/2018 12:06:20",
      "content": "<p>Hi! Any reason why the bottom car is annotated in 170927_064423787_Camera_5 but not in the rest of images? Thanks!</p>",
      "rawMarkdown": "Hi! Any reason why the bottom car is annotated in 170927_064423787_Camera_5 but not in the rest of images? Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 310173,
      "author_name": "ceperaang",
      "author_url": "",
      "post_date": "04/06/2018 18:10:22",
      "content": "<p>&gt; We are excited for to host this competition, which provides a challenging data set of 800K street-view images with per-pixel semantic annotation, courtesy of Baidu's ApolloScape project.</p>\n\n<p>It looks like there are only 39k images in train labels. I still haven't finished downloading full train data but it looks like there is also less then 800k images. Are other 761k images hosted elsewhere or am I missing something obvious?</p>",
      "votes": null,
      "replies": [
        {
          "id": 310479,
          "author_name": "ruigangyang",
          "author_url": "",
          "post_date": "04/07/2018 17:05:37",
          "content": "<p>Sorry， I made a typo in the initial introduciton.  We have uploaded to Kaggle about 60K images, we have a richer set at Apolloscape.auto that has 140K iamges. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310471,
      "author_name": "erruoru",
      "author_url": "",
      "post_date": "04/07/2018 16:45:44",
      "content": "<p>is there any requirement to develop a real-time algorithm?</p>",
      "votes": null,
      "replies": [
        {
          "id": 310480,
          "author_name": "ruigangyang",
          "author_url": "",
          "post_date": "04/07/2018 17:05:50",
          "content": "<p>No, the emphasis is on accuracy</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 310500,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "04/07/2018 18:13:08",
      "content": "<p>Hi Ruigang Yang, sorry for double posting, but could you please have a look at <a href=\"https://www.kaggle.com/c/cvpr-2018-autonomous-driving/discussion/53845\">https://www.kaggle.com/c/cvpr-2018-autonomous-driving/discussion/53845</a> - I suspect that train labels were uploaded in a wrong format.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 312542,
      "author_name": "tzt202",
      "author_url": "",
      "post_date": "04/12/2018 00:56:34",
      "content": "<p>We have submitted some results before, but there are some problems. The shape of our predicted mask is (H, W), and the first time we  use numpy.ndarry.flatten directly, then get the Length Encoding format and submit it successfully. After that, we realize that the pixels should be numbered from top to bottom, then left to right. So we transpose the array (numpy.ndarry.T) before use numpy.ndarry.flatten and submit it again but get the errors as: \nEvaluation Exception: Sequence contains no elements\nWe have no idea what happens as we just transpose the array. \nBy the way, why does the leaderboard display our score with others is zero?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 312974,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "04/12/2018 15:51:49",
      "content": "<p>&gt; To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. </p>\n\n<p>Test set files are named as <code>0aa38f3841f7aa778c0e2cdd6a29dc1a.jpg</code>, does this mean that frame order is not explicitly provided for the test set, so if we want to use information from neighbor video frames, we'll have to recover them ourselves? Same question for left/right camera?</p>",
      "votes": null,
      "replies": [
        {
          "id": 313272,
          "author_name": "huangxinyu01",
          "author_url": "",
          "post_date": "04/13/2018 03:10:29",
          "content": "<p>Hi, Konstantin, I uploaded a zip file in the Data session that contains video lists and a mapping from md5 to the original timestamp for the testing set. Please check if it meets your needs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 313383,
          "author_name": "lopuhin",
          "author_url": "",
          "post_date": "04/13/2018 07:31:09",
          "content": "<p>Hi Xinyu, this is exactly what I wanted, thanks! I checked test ids, there are 10 that are missing from the mapping files:</p>\n\n<pre><code>['30226074f61ab422d3f0ddbb284fa852',\n '9df73c4dd7d6efab6e39c9e91f4fe3ae',\n '12eeb997abcf94235b65f663d2cff198',\n 'abffd3db151bca65c1a72281c98592f2',\n 'b0a81467ffa2b260fa84fdda78d245a9',\n '0d5920509771a8d99dcfc265ce046df6',\n 'b0a539102aba1dee708d14edaf499bc0',\n 'e714994348b85004d09a447b2e0d5fc0',\n 'fa23f779c6c842dc2ad6445b41650707',\n 'c745d4d27662374bebc0ab0d1cd31dd6']\n</code></pre>\n\n<p>This is also visible in <code>wc -l</code>: it gives 1907 lines for mappings, while I see 1917 test files.</p>\n\n<p>It would be awesome if you could provide mappings for these 10 files as well.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 314335,
      "author_name": "lopuhin",
      "author_url": "",
      "post_date": "04/15/2018 09:21:08",
      "content": "<p>Hi Xinyu Huang, I'd like to clarify how evaluation metric is calculated - I notice that submissions must contain the confidence field, but it seems unused by the evaluation, is this correct? This is the relevant bit I believe:</p>\n\n<blockquote>\n  <p>If there are multiple predicted instances matched to a ground truth instance, the predicted instance with the largest IoU is considered as the true positive, and remaining predicted instances are false positives. </p>\n</blockquote>\n\n<p>This means that matches are selected using only IoU, not confidence (as also often done for instance segmentation), right?</p>",
      "votes": null,
      "replies": [
        {
          "id": 321898,
          "author_name": "huangxinyu01",
          "author_url": "",
          "post_date": "05/02/2018 05:37:14",
          "content": "<p>Thanks for pointing out. We have updated the description to \"If there are multiple predicted instances matched to a ground truth instance, the predicted instance that is larger than the IoU threshold and has the largest confidence is considered as the true positive, and remaining predicted instances are false positives. \"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 321903,
          "author_name": "aaalgo",
          "author_url": "",
          "post_date": "05/02/2018 05:52:19",
          "content": "<p>Unless you allow overlapping predictions, a matching IoU &gt; 0.5 means 1) the prediction cannot match to any other ground-truth with IoU &gt; 0.5 and 2) the ground truth region cannot match to any other prediction with IoU &gt; 0.5.   So in a practical sense (excluding extremely rare cases when 0.5+05=1), confidence score is useless here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 324043,
          "author_name": "huangxinyu01",
          "author_url": "",
          "post_date": "05/07/2018 02:37:43",
          "content": "<p>In our evaluation, a predicted result will not be matched to more than one ground truth instance. A predicted result with IoU larger than a threshold is either matched or not-matched. This is based on the confidence scores.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 334776,
      "author_name": "bertocasta",
      "author_url": "",
      "post_date": "05/28/2018 12:06:20",
      "content": "<p>Hi! Any reason why the bottom car is annotated in 170927_064423787_Camera_5 but not in the rest of images? Thanks!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "309278": "Welcome to the Video Segmentation Challenge from 2018 CVPR Workshop on Autonomous Driving ([www.wad.ai][1]). We are excited for to host this competition, which provides a challenging data set of 60K street-view images with per-pixel semantic annotation, courtesy of Baidu's [ApolloScape][2] project.  An even bigger set containing 140K images is available there.  \n\nIn this challenge, you are asked to segment out each instance of different objects in the test video.\nThe video frames exhibit challenging cases of large lighting variations, heavy occlusions, and busy background. To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. \n\nPlease feel free to ask your questions in this thread.\nGood luck!\n\n\n  [1]: http://www.wad.ai\n  [2]: http://apolloscape.auto",
    "310173": "&gt; We are excited for to host this competition, which provides a challenging data set of 800K street-view images with per-pixel semantic annotation, courtesy of Baidu's ApolloScape project.\n\nIt looks like there are only 39k images in train labels. I still haven't finished downloading full train data but it looks like there is also less then 800k images. Are other 761k images hosted elsewhere or am I missing something obvious?",
    "310471": "is there any requirement to develop a real-time algorithm?",
    "310479": "Sorry， I made a typo in the initial introduciton.  We have uploaded to Kaggle about 60K images, we have a richer set at Apolloscape.auto that has 140K iamges.",
    "310480": "No, the emphasis is on accuracy",
    "310500": "Hi Ruigang Yang, sorry for double posting, but could you please have a look at https://www.kaggle.com/c/cvpr-2018-autonomous-driving/discussion/53845 - I suspect that train labels were uploaded in a wrong format.",
    "312542": "We have submitted some results before, but there are some problems. The shape of our predicted mask is (H, W), and the first time we  use numpy.ndarry.flatten directly, then get the Length Encoding format and submit it successfully. After that, we realize that the pixels should be numbered from top to bottom, then left to right. So we transpose the array (numpy.ndarry.T) before use numpy.ndarry.flatten and submit it again but get the errors as: \nEvaluation Exception: Sequence contains no elements\nWe have no idea what happens as we just transpose the array. \nBy the way, why does the leaderboard display our score with others is zero?",
    "312974": "&gt; To get the best results, you are encouraged to explore the continuously annotated video frames that are only available in this data set. \n\nTest set files are named as `0aa38f3841f7aa778c0e2cdd6a29dc1a.jpg`, does this mean that frame order is not explicitly provided for the test set, so if we want to use information from neighbor video frames, we'll have to recover them ourselves? Same question for left/right camera?",
    "313272": "Hi, Konstantin, I uploaded a zip file in the Data session that contains video lists and a mapping from md5 to the original timestamp for the testing set. Please check if it meets your needs.",
    "313383": "Hi Xinyu, this is exactly what I wanted, thanks! I checked test ids, there are 10 that are missing from the mapping files:\n\n    ['30226074f61ab422d3f0ddbb284fa852',\n     '9df73c4dd7d6efab6e39c9e91f4fe3ae',\n     '12eeb997abcf94235b65f663d2cff198',\n     'abffd3db151bca65c1a72281c98592f2',\n     'b0a81467ffa2b260fa84fdda78d245a9',\n     '0d5920509771a8d99dcfc265ce046df6',\n     'b0a539102aba1dee708d14edaf499bc0',\n     'e714994348b85004d09a447b2e0d5fc0',\n     'fa23f779c6c842dc2ad6445b41650707',\n     'c745d4d27662374bebc0ab0d1cd31dd6']\n\nThis is also visible in `wc -l`: it gives 1907 lines for mappings, while I see 1917 test files.\n\nIt would be awesome if you could provide mappings for these 10 files as well.",
    "314335": "Hi Xinyu Huang, I'd like to clarify how evaluation metric is calculated - I notice that submissions must contain the confidence field, but it seems unused by the evaluation, is this correct? This is the relevant bit I believe:\n\n&gt; If there are multiple predicted instances matched to a ground truth instance, the predicted instance with the largest IoU is considered as the true positive, and remaining predicted instances are false positives. \n\nThis means that matches are selected using only IoU, not confidence (as also often done for instance segmentation), right?",
    "321898": "Thanks for pointing out. We have updated the description to \"If there are multiple predicted instances matched to a ground truth instance, the predicted instance that is larger than the IoU threshold and has the largest confidence is considered as the true positive, and remaining predicted instances are false positives. \"",
    "321903": "Unless you allow overlapping predictions, a matching IoU &gt; 0.5 means 1) the prediction cannot match to any other ground-truth with IoU &gt; 0.5 and 2) the ground truth region cannot match to any other prediction with IoU &gt; 0.5.   So in a practical sense (excluding extremely rare cases when 0.5+05=1), confidence score is useless here.",
    "324043": "In our evaluation, a predicted result will not be matched to more than one ground truth instance. A predicted result with IoU larger than a threshold is either matched or not-matched. This is based on the confidence scores.",
    "334776": "Hi! Any reason why the bottom car is annotated in 170927_064423787_Camera_5 but not in the rest of images? Thanks!"
  },
  "source": "meta"
}