{
  "id": 110207,
  "title": "Leaderboard Update",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/110207",
  "author_name": "Phil Culliton",
  "post_date": "2019-09-25T21:28:07.654000",
  "votes": 21,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all!</p>\n\n<p>First, thanks to <a href=\"/joonl04\">@joonl04</a> and everyone in their kernel thread and elsewhere that spotted and reported the problem.</p>\n\n<p>There was a bug in the evaluation metric, specifically in the code we used for polygon intersection. It was returning bad IoU values due to a scaling factor error on my part that didn't show up in our normal pre-launch tests, but showed up really well in the predictions created by <a href=\"/joonl04\">@joonl04</a>'s kernel! The bug has been fixed, and the leaderboard has been rescored.</p>\n\n<p>I'm sorry for not catching this in the first place! The translation to C# for this metric was complex, and while the host team and I did our best to make sure our metrics matched exactly, I missed this case in testing. It's entirely on me.</p>\n\n<p>Please let me know if anything seems odd or off - I believe we've tracked down all of the corner cases here, but I'm very open to continued feedback. I'd really like this competition to be fun and fruitful for everyone!</p>\n\n<p>You'll notice that scores went down significantly - this is expected. This is a tough problem and I've been really impressed with the submissions I've seen so far - but the submissions that caught the problem definitely won't score well any longer.</p>\n\n<p>Once again, I want to thank everyone that reported the issue, and for everyone's patience as we fixed the problem. I really appreciate it. Feel free to ask any questions you might have!</p>",
  "messages": [
    {
      "id": 634117,
      "postDate": "2019-09-25T21:28:07.653Z",
      "content": "<p>Hi all!</p>\n\n<p>First, thanks to <a href=\"/joonl04\">@joonl04</a> and everyone in their kernel thread and elsewhere that spotted and reported the problem.</p>\n\n<p>There was a bug in the evaluation metric, specifically in the code we used for polygon intersection. It was returning bad IoU values due to a scaling factor error on my part that didn't show up in our normal pre-launch tests, but showed up really well in the predictions created by <a href=\"/joonl04\">@joonl04</a>'s kernel! The bug has been fixed, and the leaderboard has been rescored.</p>\n\n<p>I'm sorry for not catching this in the first place! The translation to C# for this metric was complex, and while the host team and I did our best to make sure our metrics matched exactly, I missed this case in testing. It's entirely on me.</p>\n\n<p>Please let me know if anything seems odd or off - I believe we've tracked down all of the corner cases here, but I'm very open to continued feedback. I'd really like this competition to be fun and fruitful for everyone!</p>\n\n<p>You'll notice that scores went down significantly - this is expected. This is a tough problem and I've been really impressed with the submissions I've seen so far - but the submissions that caught the problem definitely won't score well any longer.</p>\n\n<p>Once again, I want to thank everyone that reported the issue, and for everyone's patience as we fixed the problem. I really appreciate it. Feel free to ask any questions you might have!</p>",
      "rawMarkdown": "Hi all!\n\nFirst, thanks to @joonl04 and everyone in their kernel thread and elsewhere that spotted and reported the problem.\n\nThere was a bug in the evaluation metric, specifically in the code we used for polygon intersection. It was returning bad IoU values due to a scaling factor error on my part that didn't show up in our normal pre-launch tests, but showed up really well in the predictions created by @joonl04's kernel! The bug has been fixed, and the leaderboard has been rescored.\n\nI'm sorry for not catching this in the first place! The translation to C# for this metric was complex, and while the host team and I did our best to make sure our metrics matched exactly, I missed this case in testing. It's entirely on me.\n\nPlease let me know if anything seems odd or off - I believe we've tracked down all of the corner cases here, but I'm very open to continued feedback. I'd really like this competition to be fun and fruitful for everyone!\n\nYou'll notice that scores went down significantly - this is expected. This is a tough problem and I've been really impressed with the submissions I've seen so far - but the submissions that caught the problem definitely won't score well any longer.\n\nOnce again, I want to thank everyone that reported the issue, and for everyone's patience as we fixed the problem. I really appreciate it. Feel free to ask any questions you might have!\n",
      "votes": 21
    },
    {
      "id": 634151,
      "postDate": "2019-09-26T00:05:41.190Z",
      "content": "<p>Thanks (and to anyone who participated in resolving this issue) <a href=\"/philculliton\">@philculliton</a> !</p>",
      "rawMarkdown": "Thanks (and to anyone who participated in resolving this issue) @philculliton !",
      "votes": 3
    },
    {
      "id": 640445,
      "postDate": "2019-10-04T03:11:26.423Z",
      "content": "<p>Serhii Hrynko, you can use evaluation code  from GitHub <a href=\"https://github.com/lyft/nuscenes-devkit\">https://github.com/lyft/nuscenes-devkit</a> to calculate scores  for different thresholdes (0.5...0.95) and then take average of the scores.</p>",
      "rawMarkdown": "Serhii Hrynko, you can use evaluation code  from GitHub https://github.com/lyft/nuscenes-devkit to calculate scores  for different thresholdes (0.5...0.95) and then take average of the scores.",
      "votes": 2,
      "replies": [
        {
          "id": 640517,
          "postDate": "2019-10-04T04:39:54Z",
          "content": "<p>Thanks a lot!</p>",
          "rawMarkdown": "Thanks a lot!"
        }
      ]
    },
    {
      "id": 640234,
      "postDate": "2019-10-04T00:14:14.640Z",
      "content": "<p>Anybody emulated score calculation to use with training annotations so far? </p>",
      "rawMarkdown": "Anybody emulated score calculation to use with training annotations so far? "
    },
    {
      "id": 634164,
      "postDate": "2019-09-26T01:09:01.523Z",
      "content": "<p>Many thanks! And could you please also update the evaluation code at github?</p>",
      "rawMarkdown": "Many thanks! And could you please also update the evaluation code at github?",
      "replies": [
        {
          "id": 634175,
          "postDate": "2019-09-26T01:49:06.273Z",
          "content": "<p>Hi! The Python code in GitHub is still the same. Only Kaggle's back end code (in C#) has changed, to better match the Python code.</p>",
          "rawMarkdown": "Hi! The Python code in GitHub is still the same. Only Kaggle's back end code (in C#) has changed, to better match the Python code.",
          "votes": 3
        },
        {
          "id": 634182,
          "postDate": "2019-09-26T02:00:20.593Z",
          "content": "<p>OK. Thanks for your explanation!</p>",
          "rawMarkdown": "OK. Thanks for your explanation!"
        },
        {
          "id": 659330,
          "postDate": "2019-10-27T13:14:27.183Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a>  . I may be wrong  :) but I think C# can run native code(un managed) as well . may be directly calling the snippet may help if rewrite of metric needs to be avoided</p>",
          "rawMarkdown": "@philculliton  . I may be wrong  :) but I think C# can run native code(un managed) as well . may be directly calling the snippet may help if rewrite of metric needs to be avoided"
        }
      ]
    },
    {
      "id": 661530,
      "postDate": "2019-10-30T12:38:45.433Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 634151,
      "author_name": "Brian Lee",
      "author_url": "",
      "post_date": "2019-09-26T00:05:41.190000",
      "content": "<p>Thanks (and to anyone who participated in resolving this issue) <a href=\"/philculliton\">@philculliton</a> !</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 640445,
      "author_name": "Rustem Iskuzhin",
      "author_url": "",
      "post_date": "2019-10-04T03:11:26.423000",
      "content": "<p>Serhii Hrynko, you can use evaluation code  from GitHub <a href=\"https://github.com/lyft/nuscenes-devkit\">https://github.com/lyft/nuscenes-devkit</a> to calculate scores  for different thresholdes (0.5...0.95) and then take average of the scores.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 640517,
          "author_name": "Serhii Hrynko",
          "author_url": "",
          "post_date": "2019-10-04T04:39:54",
          "content": "<p>Thanks a lot!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 640234,
      "author_name": "Serhii Hrynko",
      "author_url": "",
      "post_date": "2019-10-04T00:14:14.640000",
      "content": "<p>Anybody emulated score calculation to use with training annotations so far? </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 634164,
      "author_name": "twang",
      "author_url": "",
      "post_date": "2019-09-26T01:09:01.523000",
      "content": "<p>Many thanks! And could you please also update the evaluation code at github?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 634175,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2019-09-26T01:49:06.273000",
          "content": "<p>Hi! The Python code in GitHub is still the same. Only Kaggle's back end code (in C#) has changed, to better match the Python code.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 634182,
          "author_name": "twang",
          "author_url": "",
          "post_date": "2019-09-26T02:00:20.593000",
          "content": "<p>OK. Thanks for your explanation!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 659330,
          "author_name": "Kishore M",
          "author_url": "",
          "post_date": "2019-10-27T13:14:27.183000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a>  . I may be wrong  :) but I think C# can run native code(un managed) as well . may be directly calling the snippet may help if rewrite of metric needs to be avoided</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 661530,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-10-30T12:38:45.433000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "634117": "Hi all!\n\nFirst, thanks to @joonl04 and everyone in their kernel thread and elsewhere that spotted and reported the problem.\n\nThere was a bug in the evaluation metric, specifically in the code we used for polygon intersection. It was returning bad IoU values due to a scaling factor error on my part that didn't show up in our normal pre-launch tests, but showed up really well in the predictions created by @joonl04's kernel! The bug has been fixed, and the leaderboard has been rescored.\n\nI'm sorry for not catching this in the first place! The translation to C# for this metric was complex, and while the host team and I did our best to make sure our metrics matched exactly, I missed this case in testing. It's entirely on me.\n\nPlease let me know if anything seems odd or off - I believe we've tracked down all of the corner cases here, but I'm very open to continued feedback. I'd really like this competition to be fun and fruitful for everyone!\n\nYou'll notice that scores went down significantly - this is expected. This is a tough problem and I've been really impressed with the submissions I've seen so far - but the submissions that caught the problem definitely won't score well any longer.\n\nOnce again, I want to thank everyone that reported the issue, and for everyone's patience as we fixed the problem. I really appreciate it. Feel free to ask any questions you might have!\n",
    "634151": "Thanks (and to anyone who participated in resolving this issue) @philculliton !",
    "640445": "Serhii Hrynko, you can use evaluation code  from GitHub https://github.com/lyft/nuscenes-devkit to calculate scores  for different thresholdes (0.5...0.95) and then take average of the scores.",
    "640234": "Anybody emulated score calculation to use with training annotations so far? ",
    "634164": "Many thanks! And could you please also update the evaluation code at github?",
    "661530": ""
  }
}