{
  "id": 115376,
  "title": "PointRCNN trouble, pls halp :(",
  "url": "/competitions/3d-object-detection-for-autonomous-vehicles/discussion/115376",
  "author_name": "Igor Kotenkov",
  "post_date": "2019-11-02T12:18:24.475000",
  "votes": 17,
  "comment_count": 50,
  "views": 0,
  "content": "<p>Hello all participants! \nI'm trying to solve 3D object detection task on pointclouds with PointRCNN network. I use official implementation from here: <a href=\"https://github.com/sshaoshuai/PointRCNN\">https://github.com/sshaoshuai/PointRCNN</a> .\nWhen i run inference on pretrained model (checkpoint link on repo) i got normal predictions (only shown one of predictions, all other also well detected):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8f8db26537d153e61d3be91338ef3ca3%2Fphoto_2019-10-27_11-05-34.jpg?generation=1572695211708053&amp;alt=media\" alt=\"\">\nThen i try train network from scratch - firstly RPN and then RCNN. I used 5000 samples from LYFT dataset, with 6 camera it's a 30000 samples - approx. x8 from original KITTI train data.\nRPN was trained on 50 epochs (intead 200 from repo, but dataset is larger), RCNN has 10 epochs (instead of 70). </p>\n\n<p>The day before yesterday i run inference on test set and got this terrible PROBLEM:\npredictions have well translation, but absolutely disoriented (have bad YAW parameter):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa354920533067b35152c0f38bf581ce7%2FScreenshot_1.png?generation=1572695601960018&amp;alt=media\" alt=\"\"></p>\n\n<p>Then i use debugger and take step-by-step in-depth look, where this YAW was generated - maybe the network gives the right YAWs, but I have an error in postprocessing (at the same time, it’s strange that the pretrained checkpoint works fine). But i find that end writted YAW parameter is the same that generated as ROI in network, before NMS, limitations and other operations (\nafter the decode bbox from bins, that described in article, the angle is slightly different, but this is understandable - after all, this is the result of several proposals nearby, i.e. 2.41 YAW become 2.47)\nSo, I almost ran out of ideas, what else can I check and what to do next. I'm almost desperate :)\nThe last hypothesis is that the error is somewhere in the data loader, and the angle is literally white noise, so the neural network cannot learn it and outputs ~random~ +-, so i can check this today. </p>\n\n<p>Maybe someone already retrained this network from scratch? and was able to avoid this problem? or fix it? or do you have any suggestions what else to check?</p>\n\n<p>It's a shame to spend so much time and have no result ;( <br>\nI added classes, since the network initially learns only for cars, and they even work, but there is no point in this without correct YAW (green associated with class \"bicycle\"):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa8ee160c5f893f60ac0e057b8f6eeedd%2FScreenshot_2.png?generation=1572696349587432&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F2438ff8b2466cc94138bab1e1e6b0c00%2Fphoto_2019-11-01_00-18-22.jpg?generation=1572696398633333&amp;alt=media\" alt=\"\"></p>\n\n<p>so another idea is network is underfitted in relation to the YAW prediction, idk</p>\n\n<p>PS: yes, as you may have noticed - I'm an native speaker, so I apologize (:kekeke:)</p>",
  "messages": [
    {
      "id": 663609,
      "postDate": "2019-11-02T12:18:24.477Z",
      "content": "<p>Hello all participants! \nI'm trying to solve 3D object detection task on pointclouds with PointRCNN network. I use official implementation from here: <a href=\"https://github.com/sshaoshuai/PointRCNN\">https://github.com/sshaoshuai/PointRCNN</a> .\nWhen i run inference on pretrained model (checkpoint link on repo) i got normal predictions (only shown one of predictions, all other also well detected):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8f8db26537d153e61d3be91338ef3ca3%2Fphoto_2019-10-27_11-05-34.jpg?generation=1572695211708053&amp;alt=media\" alt=\"\">\nThen i try train network from scratch - firstly RPN and then RCNN. I used 5000 samples from LYFT dataset, with 6 camera it's a 30000 samples - approx. x8 from original KITTI train data.\nRPN was trained on 50 epochs (intead 200 from repo, but dataset is larger), RCNN has 10 epochs (instead of 70). </p>\n\n<p>The day before yesterday i run inference on test set and got this terrible PROBLEM:\npredictions have well translation, but absolutely disoriented (have bad YAW parameter):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa354920533067b35152c0f38bf581ce7%2FScreenshot_1.png?generation=1572695601960018&amp;alt=media\" alt=\"\"></p>\n\n<p>Then i use debugger and take step-by-step in-depth look, where this YAW was generated - maybe the network gives the right YAWs, but I have an error in postprocessing (at the same time, it’s strange that the pretrained checkpoint works fine). But i find that end writted YAW parameter is the same that generated as ROI in network, before NMS, limitations and other operations (\nafter the decode bbox from bins, that described in article, the angle is slightly different, but this is understandable - after all, this is the result of several proposals nearby, i.e. 2.41 YAW become 2.47)\nSo, I almost ran out of ideas, what else can I check and what to do next. I'm almost desperate :)\nThe last hypothesis is that the error is somewhere in the data loader, and the angle is literally white noise, so the neural network cannot learn it and outputs ~random~ +-, so i can check this today. </p>\n\n<p>Maybe someone already retrained this network from scratch? and was able to avoid this problem? or fix it? or do you have any suggestions what else to check?</p>\n\n<p>It's a shame to spend so much time and have no result ;( <br>\nI added classes, since the network initially learns only for cars, and they even work, but there is no point in this without correct YAW (green associated with class \"bicycle\"):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa8ee160c5f893f60ac0e057b8f6eeedd%2FScreenshot_2.png?generation=1572696349587432&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F2438ff8b2466cc94138bab1e1e6b0c00%2Fphoto_2019-11-01_00-18-22.jpg?generation=1572696398633333&amp;alt=media\" alt=\"\"></p>\n\n<p>so another idea is network is underfitted in relation to the YAW prediction, idk</p>\n\n<p>PS: yes, as you may have noticed - I'm an native speaker, so I apologize (:kekeke:)</p>",
      "rawMarkdown": "Hello all participants! \nI'm trying to solve 3D object detection task on pointclouds with PointRCNN network. I use official implementation from here: https://github.com/sshaoshuai/PointRCNN .\nWhen i run inference on pretrained model (checkpoint link on repo) i got normal predictions (only shown one of predictions, all other also well detected):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8f8db26537d153e61d3be91338ef3ca3%2Fphoto_2019-10-27_11-05-34.jpg?generation=1572695211708053&amp;alt=media)\nThen i try train network from scratch - firstly RPN and then RCNN. I used 5000 samples from LYFT dataset, with 6 camera it's a 30000 samples - approx. x8 from original KITTI train data.\nRPN was trained on 50 epochs (intead 200 from repo, but dataset is larger), RCNN has 10 epochs (instead of 70). \n\nThe day before yesterday i run inference on test set and got this terrible PROBLEM:\npredictions have well translation, but absolutely disoriented (have bad YAW parameter):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa354920533067b35152c0f38bf581ce7%2FScreenshot_1.png?generation=1572695601960018&amp;alt=media)\n\nThen i use debugger and take step-by-step in-depth look, where this YAW was generated - maybe the network gives the right YAWs, but I have an error in postprocessing (at the same time, it’s strange that the pretrained checkpoint works fine). But i find that end writted YAW parameter is the same that generated as ROI in network, before NMS, limitations and other operations (\nafter the decode bbox from bins, that described in article, the angle is slightly different, but this is understandable - after all, this is the result of several proposals nearby, i.e. 2.41 YAW become 2.47)\nSo, I almost ran out of ideas, what else can I check and what to do next. I'm almost desperate :)\nThe last hypothesis is that the error is somewhere in the data loader, and the angle is literally white noise, so the neural network cannot learn it and outputs ~random~ +-, so i can check this today. \n\n\nMaybe someone already retrained this network from scratch? and was able to avoid this problem? or fix it? or do you have any suggestions what else to check?\n\n\nIt's a shame to spend so much time and have no result ;(  \nI added classes, since the network initially learns only for cars, and they even work, but there is no point in this without correct YAW (green associated with class \"bicycle\"):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa8ee160c5f893f60ac0e057b8f6eeedd%2FScreenshot_2.png?generation=1572696349587432&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F2438ff8b2466cc94138bab1e1e6b0c00%2Fphoto_2019-11-01_00-18-22.jpg?generation=1572696398633333&amp;alt=media)\n\nso another idea is network is underfitted in relation to the YAW prediction, idk\n\nPS: yes, as you may have noticed - I'm an native speaker, so I apologize (:kekeke:)",
      "votes": 17
    },
    {
      "id": 664010,
      "postDate": "2019-11-03T03:04:22.500Z",
      "content": "<p>From Lyft to PointRCNN: \npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground) </p>\n\n<p>I think this is the correct conversion. Hope it helps. But you need to double check it. Be very careful!\nKeep going. You are actually doing very well.</p>",
      "rawMarkdown": "From Lyft to PointRCNN: \npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground) \n\nI think this is the correct conversion. Hope it helps. But you need to double check it. Be very careful!\nKeep going. You are actually doing very well.",
      "votes": 12,
      "replies": [
        {
          "id": 664504,
          "postDate": "2019-11-03T18:55:08.787Z",
          "content": "<p>I'm curious, at what point do these changes need to be applied? Do you train the PointRCNN as is and then convert the predictions? Or do you have to apply transformations to the data before you can train PointRCNN?</p>",
          "rawMarkdown": "I'm curious, at what point do these changes need to be applied? Do you train the PointRCNN as is and then convert the predictions? Or do you have to apply transformations to the data before you can train PointRCNN?"
        },
        {
          "id": 664540,
          "postDate": "2019-11-03T20:42:03.973Z",
          "content": "<p>Zhang, thanks for reply and good words! \nI'm debugging dataloader step by step and check augmentations in PointRCNN. I found that authors use alpha when rotate boxes. But alpha after converting from LYFT was constant -10 xD \nI change this part (recalc alpha from yaw) and run retrain network. Tomorrow see what i'm training - \nI am full of hopes and expectations.</p>\n\n<p>How you check that yours converting is right? Do you render submission? I'm render .csv using that kernel: <a href=\"https://www.kaggle.com/rishabhiitbhu/visualizing-predictions\">https://www.kaggle.com/rishabhiitbhu/visualizing-predictions</a> and try to check posiitons, classes and rotations.</p>\n\n<p>Good luck in the Competition!</p>",
          "rawMarkdown": "Zhang, thanks for reply and good words! \nI'm debugging dataloader step by step and check augmentations in PointRCNN. I found that authors use alpha when rotate boxes. But alpha after converting from LYFT was constant -10 xD \nI change this part (recalc alpha from yaw) and run retrain network. Tomorrow see what i'm training - \nI am full of hopes and expectations.\n\nHow you check that yours converting is right? Do you render submission? I'm render .csv using that kernel: https://www.kaggle.com/rishabhiitbhu/visualizing-predictions and try to check posiitons, classes and rotations.\n\nGood luck in the Competition!",
          "votes": 1
        },
        {
          "id": 664602,
          "postDate": "2019-11-04T00:26:24.630Z",
          "content": "<p>One way to check whether your conversion is correct is to check whether the same points are in the same box. Say you have point and box in lyft and convert it to PointRCNN as point-cnn and box-cnn. Now you can use points_in_box in lyft utils to find which points are inside the box. You can also use a similar function in PointRCNN to find which point-cnn are inside box-cnn. The indices of the two findings should be the same.</p>",
          "rawMarkdown": "One way to check whether your conversion is correct is to check whether the same points are in the same box. Say you have point and box in lyft and convert it to PointRCNN as point-cnn and box-cnn. Now you can use points_in_box in lyft utils to find which points are inside the box. You can also use a similar function in PointRCNN to find which point-cnn are inside box-cnn. The indices of the two findings should be the same.",
          "votes": 1
        },
        {
          "id": 664603,
          "postDate": "2019-11-04T00:28:32.367Z",
          "content": "<p>Generally, you need to convert points and box from lyft toPointRCNN first and do the training and prediction, then you need to convert the prediction back from PointRCNN to lyft</p>",
          "rawMarkdown": "Generally, you need to convert points and box from lyft toPointRCNN first and do the training and prediction, then you need to convert the prediction back from PointRCNN to lyft",
          "votes": 1
        },
        {
          "id": 664889,
          "postDate": "2019-11-04T11:05:48.210Z",
          "content": "<p>Today i  see retrained PointRCNN network with totation angle bug fix. So they looks much better (on second scrren left big bbox also classifiend as truck, wow):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Ff269bb111604680306e4db6695c5b81a%2Fphoto_2019-11-04_05-11-56.jpg?generation=1572865402713654&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8a5a1fe7522c579bb10b8a281ac1aac9%2Fphoto_2019-11-04_04-02-34.jpg?generation=1572865378923645&amp;alt=media\" alt=\"\">\nSo i submit and got 0.007 on lb (only front camera used at this time).\nZhang, how do you rate this score for PointRCNN? \nI have a feeling that somewhere there is a mistake in the conversion, since everything is cool on the visualization, and the cars are detected well.</p>",
          "rawMarkdown": "Today i  see retrained PointRCNN network with totation angle bug fix. So they looks much better (on second scrren left big bbox also classifiend as truck, wow):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Ff269bb111604680306e4db6695c5b81a%2Fphoto_2019-11-04_05-11-56.jpg?generation=1572865402713654&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8a5a1fe7522c579bb10b8a281ac1aac9%2Fphoto_2019-11-04_04-02-34.jpg?generation=1572865378923645&amp;alt=media)\nSo i submit and got 0.007 on lb (only front camera used at this time).\nZhang, how do you rate this score for PointRCNN? \nI have a feeling that somewhere there is a mistake in the conversion, since everything is cool on the visualization, and the cars are detected well.",
          "votes": 2
        },
        {
          "id": 664900,
          "postDate": "2019-11-04T11:20:52.180Z",
          "content": "<p>what do you mean by only front camera used? it is strange, the score is too low while the boxes look fine. Do you convert the box back to lyft format correctly? also rember to transform the boxes back to world coordinates.</p>",
          "rawMarkdown": "what do you mean by only front camera used? it is strange, the score is too low while the boxes look fine. Do you convert the box back to lyft format correctly? also rember to transform the boxes back to world coordinates.",
          "votes": 3
        },
        {
          "id": 664907,
          "postDate": "2019-11-04T11:39:43.330Z",
          "content": "<p>I use front camera only for time, just check and debugging - after got a working pipeline i'm add another cameras and increase score (dream...).\nYes, i'm converting to world coordinates. \nFor example,  can you check one bbox of one token in your submission?\nTOKEN: <code>da0671a11e837c86c131f04ded4fc91941b936f70fdead745259361f1f06f250</code>\ncoordinates : 2231.559051779429 987.5928335963141 -17.873463805634827\nwlh: 2.946 11.6001 3.5798 \nyaw: 2.5836576429783755\nclass: bus\ndoes it look like what you wrote in .csv submission?\nIt's that bus (or not bus, but i'm talking only for checking coordinates and yaw):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F27eac95b6cbcc58231e8b7d3101279e4%2FScreenshot_2.png?generation=1572867506015369&amp;alt=media\" alt=\"\"></p>\n\n<p>think competition host will not ban us for one of thousands rpedictions public share)</p>",
          "rawMarkdown": "I use front camera only for time, just check and debugging - after got a working pipeline i'm add another cameras and increase score (dream...).\nYes, i'm converting to world coordinates. \nFor example,  can you check one bbox of one token in your submission?\nTOKEN: `da0671a11e837c86c131f04ded4fc91941b936f70fdead745259361f1f06f250`\ncoordinates : 2231.559051779429 987.5928335963141 -17.873463805634827\nwlh: 2.946 11.6001 3.5798 \nyaw: 2.5836576429783755\nclass: bus\ndoes it look like what you wrote in .csv submission?\nIt's that bus (or not bus, but i'm talking only for checking coordinates and yaw):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F27eac95b6cbcc58231e8b7d3101279e4%2FScreenshot_2.png?generation=1572867506015369&amp;alt=media)\n\nthink competition host will not ban us for one of thousands rpedictions public share)"
        },
        {
          "id": 664924,
          "postDate": "2019-11-04T12:12:01.350Z",
          "content": "<p>Still have no idea how you only use the front camera. But I previously made a mistake to clip the detected boxes: only keep the boxes with x&gt;0 (I forget the exact number, but somehow like this). After the correction, I see  the score increases more than tenfold . If your implementation(like only using the front camera) faces the same problem, your detection should actually be good.</p>",
          "rawMarkdown": "Still have no idea how you only use the front camera. But I previously made a mistake to clip the detected boxes: only keep the boxes with x&gt;0 (I forget the exact number, but somehow like this). After the correction, I see  the score increases more than tenfold . If your implementation(like only using the front camera) faces the same problem, your detection should actually be good.",
          "votes": 1
        },
        {
          "id": 664947,
          "postDate": "2019-11-04T12:57:10.383Z",
          "content": "<p>In PointRCNN implemetation i used all boxes outside camera FOV are cuted (on train and inference stages both). I don't fix that, and currently train on 6 cameras, but inference test set for 1 camera. They clip boxes by x/z axis too, and i know that. So if score on 1 camera is 0.007, 6 cameras maximum gain is 0.042, it's lower than unet on BEV :) </p>\n\n<p>what do you think, what score PointRCNN can schive on class Car on ths task?</p>",
          "rawMarkdown": "In PointRCNN implemetation i used all boxes outside camera FOV are cuted (on train and inference stages both). I don't fix that, and currently train on 6 cameras, but inference test set for 1 camera. They clip boxes by x/z axis too, and i know that. So if score on 1 camera is 0.007, 6 cameras maximum gain is 0.042, it's lower than unet on BEV :) \n\n\nwhat do you think, what score PointRCNN can schive on class Car on ths task?"
        },
        {
          "id": 664953,
          "postDate": "2019-11-04T13:10:54.850Z",
          "content": "<p>I think the increase will be much more than that. As I mentioned, I once wrongly removed around half or more of the detected boxes and the score drops from 0.086 to 0.006. If you fix all the issues, the score should be well above 0.1 (including all classes, no idea what you will get if you use car only, maybe 0.05?) </p>",
          "rawMarkdown": "I think the increase will be much more than that. As I mentioned, I once wrongly removed around half or more of the detected boxes and the score drops from 0.086 to 0.006. If you fix all the issues, the score should be well above 0.1 (including all classes, no idea what you will get if you use car only, maybe 0.05?) ",
          "votes": 4
        },
        {
          "id": 665149,
          "postDate": "2019-11-04T17:16:12.703Z",
          "content": "<p>Zhang, is 0.086 score of pure multiclass PointRCNN? </p>",
          "rawMarkdown": "Zhang, is 0.086 score of pure multiclass PointRCNN? "
        },
        {
          "id": 665290,
          "postDate": "2019-11-04T21:15:46.030Z",
          "content": "<blockquote>\n  <p>I think the increase will be much more than that. As I mentioned, I once wrongly removed around half or more of the detected boxes and the score drops from 0.086 to 0.006. If you fix all the issues, the score should be well above 0.1 (including all classes, no idea what you will get if you use car only, maybe 0.05?)</p>\n</blockquote>\n\n<p><a href=\"/nywenjing\">@nywenjing</a> How many epochs do you think it will take to reach that?</p>",
          "rawMarkdown": "&gt; I think the increase will be much more than that. As I mentioned, I once wrongly removed around half or more of the detected boxes and the score drops from 0.086 to 0.006. If you fix all the issues, the score should be well above 0.1 (including all classes, no idea what you will get if you use car only, maybe 0.05?)\n\n@nywenjing How many epochs do you think it will take to reach that?"
        },
        {
          "id": 665404,
          "postDate": "2019-11-05T01:17:25.870Z",
          "content": "<p>Actually, I did not perform a full submission with PointRCNN. I only gave a tried and finished the data conversion and train several epochs for RPN. The score is obtained with another implementation. But  I think the mistake might be similar. Here are my first submission scores:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F185a9ff84628ff68656c93ae761cc1cd%2FScreenshot%20from%202019-11-05%2008-53-11.png?generation=1572915721940258&amp;alt=media\" alt=\"\"></p>\n\n<p>The first score is almost 0 since I wrongly transform the box back to world coordinates.</p>\n\n<p>The second is 0.006 with corrected transformation to world coordinates. It is still very low. I was frustrating then. I can believe it could be so bad. So I give a try with the baseline.</p>\n\n<p>The third with 0.036 is my try of baseline then I found the yaw angle problem and correct it. I thought now I might be right. </p>\n\n<p>But the forth is still 0.006. I was very sad at that point. But when I try to do prediction on the training set, I found my predicted boxes are actually good but seems only half are present (since the other parts are wrongly removed). Then I did the correction. Just as <a href=\"/stalkermustang\">@stalkermustang</a> , I only thought It might give me 0.012 score, which is still well below the baseline. So I did not give much expectation. </p>\n\n<p>To my surprise, the corrected score was increase more than tenfold to 0.089 at the fifth submission. So <a href=\"/stalkermustang\">@stalkermustang</a>, if you face the same issue, I think you are actually very near to the right one. Keep going!</p>\n\n<p><a href=\"/usmannkhan\">@usmannkhan</a> , I have not run the pipeline of PointRCNN fully. But I guess the RPN may need 60 epochs. The author use 200 epochs for kitti. Lyft have around 4 times more data so we might need 40 epochs.  Considering there are more classes in Lyft, I will give another 20 epochs.</p>",
          "rawMarkdown": "Actually, I did not perform a full submission with PointRCNN. I only gave a tried and finished the data conversion and train several epochs for RPN. The score is obtained with another implementation. But  I think the mistake might be similar. Here are my first submission scores:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F185a9ff84628ff68656c93ae761cc1cd%2FScreenshot%20from%202019-11-05%2008-53-11.png?generation=1572915721940258&amp;alt=media)\n\nThe first score is almost 0 since I wrongly transform the box back to world coordinates.\n\nThe second is 0.006 with corrected transformation to world coordinates. It is still very low. I was frustrating then. I can believe it could be so bad. So I give a try with the baseline.\n\nThe third with 0.036 is my try of baseline then I found the yaw angle problem and correct it. I thought now I might be right. \n\nBut the forth is still 0.006. I was very sad at that point. But when I try to do prediction on the training set, I found my predicted boxes are actually good but seems only half are present (since the other parts are wrongly removed). Then I did the correction. Just as @stalkermustang , I only thought It might give me 0.012 score, which is still well below the baseline. So I did not give much expectation. \n\nTo my surprise, the corrected score was increase more than tenfold to 0.089 at the fifth submission. So @stalkermustang, if you face the same issue, I think you are actually very near to the right one. Keep going!\n\n@usmannkhan , I have not run the pipeline of PointRCNN fully. But I guess the RPN may need 60 epochs. The author use 200 epochs for kitti. Lyft have around 4 times more data so we might need 40 epochs.  Considering there are more classes in Lyft, I will give another 20 epochs.",
          "votes": 6
        },
        {
          "id": 665568,
          "postDate": "2019-11-05T05:47:51.607Z",
          "content": "<p><a href=\"/nywenjing\">@nywenjing</a> when you say that you have to convert Lyft to PointRCNN here\n<code>\nFrom Lyft to PointRCNN:\npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground)\n</code>\nwhat exactly do you mean? I mean I converted Lyft dataset to KITTI using <a href=\"/stalkermustang\">@stalkermustang</a> kernel. Do I have apply the above tranformation while loading the KITTI dataset in the PointRCNN codebase or I am doing something wrong?\nHelp would be appreciated.</p>",
          "rawMarkdown": "@nywenjing when you say that you have to convert Lyft to PointRCNN here\n```\nFrom Lyft to PointRCNN:\npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground)\n```\nwhat exactly do you mean? I mean I converted Lyft dataset to KITTI using @stalkermustang kernel. Do I have apply the above tranformation while loading the KITTI dataset in the PointRCNN codebase or I am doing something wrong?\nHelp would be appreciated."
        },
        {
          "id": 665574,
          "postDate": "2019-11-05T06:09:55.050Z",
          "content": "<p>I did not look into <a href=\"/stalkermustang\">@stalkermustang</a>  kernel. It seems too complex for me. I only need point and have nothing to do with the Camera. If you want to train PointRCNN on Lyft, you need to first convert the point cloud and gt boxes to its corresponding format, which is what my code do. If you can ensure <a href=\"/stalkermustang\">@stalkermustang</a>  kernel code does the correct conversion, then you can just do the conversion and let it run. The conversion code is only applied to Lyft dataset.</p>",
          "rawMarkdown": "I did not look into @stalkermustang  kernel. It seems too complex for me. I only need point and have nothing to do with the Camera. If you want to train PointRCNN on Lyft, you need to first convert the point cloud and gt boxes to its corresponding format, which is what my code do. If you can ensure @stalkermustang  kernel code does the correct conversion, then you can just do the conversion and let it run. The conversion code is only applied to Lyft dataset.",
          "votes": 1
        },
        {
          "id": 665605,
          "postDate": "2019-11-05T07:04:15.730Z",
          "content": "<p>Maybe this is the score, that can't be achived with PointRCNN, because you use another RCNN part? Ypu write only about RPN :( So this score is score of 2-3-4 well detected classes..?</p>",
          "rawMarkdown": "Maybe this is the score, that can't be achived with PointRCNN, because you use another RCNN part? Ypu write only about RPN :( So this score is score of 2-3-4 well detected classes..?"
        }
      ]
    },
    {
      "id": 664004,
      "postDate": "2019-11-03T02:50:30.183Z",
      "content": "<p>From my perspective of view, it might be the coordinate system problem. Lyft  uses z up and size wlh but pointrcnn uses -y up and size hwl. You need to be very careful to convert the boxes and point clouds between lyft and pointrcnn. You should also be careful of the sign of yaw angle. Some use clockwise as positive while some use anti clockwise as positive. I have not checked it. You need to double check it. I feel your prediction is correct up to a sign and/or an offset of the yaw angle.</p>",
      "rawMarkdown": "From my perspective of view, it might be the coordinate system problem. Lyft  uses z up and size wlh but pointrcnn uses -y up and size hwl. You need to be very careful to convert the boxes and point clouds between lyft and pointrcnn. You should also be careful of the sign of yaw angle. Some use clockwise as positive while some use anti clockwise as positive. I have not checked it. You need to double check it. I feel your prediction is correct up to a sign and/or an offset of the yaw angle.",
      "votes": 5
    },
    {
      "id": 665245,
      "postDate": "2019-11-04T19:36:08.377Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fc8466e24b2bad268fa4cff18f6fcea7e%2FScreen%20Shot%202019-11-04%20at%202.22.42%20PM.png?generation=1572895519112648&amp;alt=media\" alt=\"\"></p>\n\n<p>I am seeing a similar problem with my predictions. I am using PointPillars based on this code <a href=\"https://github.com/traveller59/second.pytorch\">https://github.com/traveller59/second.pytorch</a> .</p>\n\n<p>The ground truth in orange looks ok but the predictions in blue all look off by a similar angle. When creating the predictions strings I use the same transformation on to get yaw </p>\n\n<p><code>yaw = 2*np.arccos(box3ds[i].rotation[0])</code></p>\n\n<p>and this is how I create the boxes:</p>\n\n<p>```\n    def _box_translate(self, translation, x):\n        # box.center += x\n        # In Box3d center is called translation\n        translation += x\n        return translation</p>\n\n<pre><code>def _box_rotate(self, t, r, quaternion):\n    # box.center = np.dot(quaternion.rotation_matrix, box.center)\n    # In Box3d center is called translation\n    t = np.dot(quaternion.rotation_matrix, t)\n    # box.orientation = quaternion * box.orientation\n    # In Box3d orientation is called rotation, I think\n    r = quaternion * r\n    # box.velocity = np.dot(quaternion.rotation_matrix,box.velocity)\n    return t, r\n\ndef _get_lyft_pred3d_boxes(self, detection, token2info):\n    pred_box3ds = []\n    box3d = detection[\"box3d_lidar\"].detach().cpu().numpy()\n    scores = detection[\"scores\"].detach().cpu().numpy()\n    labels = detection[\"label_preds\"].detach().cpu().numpy()\n    sample_token = detection[\"metadata\"][\"token\"]\n\n    box3d[:, 6] = -box3d[:, 6] - np.pi / 2\n    for i in range(box3d.shape[0]):\n\n        # Rotate about the z axis by the z value and create a new Quaternion\n        quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n        translation = box3d[i, :3]\n        size = box3d[i, 3:6]\n\n        # Say everything is a car until we figure out the issues with the classes\n        model_class_names =  ['car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car']\n        class_name = model_class_names[labels[i]]\n\n        trans_info = token2info[detection[\"metadata\"][\"token\"]]\n\n        ###### Move box to ego vehicle coord system #######\n        translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['lidar2ego_rotation']))\n        translation = self._box_translate(translation, np.array(trans_info['lidar2ego_translation']))\n\n        ####### Move box to global coord system ########\n        translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['ego2global_rotation']))\n        translation = self._box_translate(translation, np.array(trans_info['ego2global_translation']))\n\n        box = Box3D(\n            sample_token=sample_token,\n            translation=list(translation),\n            size=list(size),\n            rotation=list(quat),\n            name=class_name,\n            score=scores[i]\n        )\n        pred_box3ds.append(box)\n\n    return pred_box3ds\n</code></pre>\n\n<p>```</p>\n\n<p>I think the predictions of the model are ok but I'm going wrong somewhere applying the rotations in post processing.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fc8466e24b2bad268fa4cff18f6fcea7e%2FScreen%20Shot%202019-11-04%20at%202.22.42%20PM.png?generation=1572895519112648&amp;alt=media)\n\nI am seeing a similar problem with my predictions. I am using PointPillars based on this code https://github.com/traveller59/second.pytorch .\n\nThe ground truth in orange looks ok but the predictions in blue all look off by a similar angle. When creating the predictions strings I use the same transformation on to get yaw \n\n`yaw = 2*np.arccos(box3ds[i].rotation[0])`\n\nand this is how I create the boxes:\n\n```\n    def _box_translate(self, translation, x):\n        # box.center += x\n        # In Box3d center is called translation\n        translation += x\n        return translation\n\n    def _box_rotate(self, t, r, quaternion):\n        # box.center = np.dot(quaternion.rotation_matrix, box.center)\n        # In Box3d center is called translation\n        t = np.dot(quaternion.rotation_matrix, t)\n        # box.orientation = quaternion * box.orientation\n        # In Box3d orientation is called rotation, I think\n        r = quaternion * r\n        # box.velocity = np.dot(quaternion.rotation_matrix,box.velocity)\n        return t, r\n    \n    def _get_lyft_pred3d_boxes(self, detection, token2info):\n        pred_box3ds = []\n        box3d = detection[\"box3d_lidar\"].detach().cpu().numpy()\n        scores = detection[\"scores\"].detach().cpu().numpy()\n        labels = detection[\"label_preds\"].detach().cpu().numpy()\n        sample_token = detection[\"metadata\"][\"token\"]\n        \n        box3d[:, 6] = -box3d[:, 6] - np.pi / 2\n        for i in range(box3d.shape[0]):\n            \n            # Rotate about the z axis by the z value and create a new Quaternion\n            quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n            translation = box3d[i, :3]\n            size = box3d[i, 3:6]\n           \n            # Say everything is a car until we figure out the issues with the classes\n            model_class_names =  ['car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car']\n            class_name = model_class_names[labels[i]]\n            \n            trans_info = token2info[detection[\"metadata\"][\"token\"]]\n            \n            ###### Move box to ego vehicle coord system #######\n            translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['lidar2ego_rotation']))\n            translation = self._box_translate(translation, np.array(trans_info['lidar2ego_translation']))\n            \n            ####### Move box to global coord system ########\n            translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['ego2global_rotation']))\n            translation = self._box_translate(translation, np.array(trans_info['ego2global_translation']))\n            \n            box = Box3D(\n                sample_token=sample_token,\n                translation=list(translation),\n                size=list(size),\n                rotation=list(quat),\n                name=class_name,\n                score=scores[i]\n            )\n            pred_box3ds.append(box)\n            \n        return pred_box3ds\n```\n\nI think the predictions of the model are ok but I'm going wrong somewhere applying the rotations in post processing.",
      "votes": 1,
      "replies": [
        {
          "id": 665260,
          "postDate": "2019-11-04T20:10:24.923Z",
          "content": "<p>your code snippet looks fine to me, maybe there's a bug in plotting the boxes?</p>",
          "rawMarkdown": "your code snippet looks fine to me, maybe there's a bug in plotting the boxes?",
          "votes": 2
        },
        {
          "id": 665278,
          "postDate": "2019-11-04T20:51:11.007Z",
          "content": "<p>Thanks for taking a look <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a>, I am using your code from here <a href=\"https://www.kaggle.com/rishabhiitbhu/visualizing-predictions\">https://www.kaggle.com/rishabhiitbhu/visualizing-predictions</a> for visualization and I didn't change anything so pretty sure the problem is not there, but I will check again in case I maybe did change something and don't remember. </p>\n\n<p>Do you think this piece of code from the other inference kernels is needed for the rotation?\n<code>\n        v = (sample_boxes[i,0] - sample_boxes[i,1]) # an edge vector\n        v /= np.linalg.norm(v) # normalized edge vector\n        r = R.from_dcm([\n            [v[0], -v[1], 0],\n            [v[1],  v[0], 0],\n            [   0,     0, 1],\n        ])  # a rotation matrix for rotation along z axis. \n        quat = r.as_quat()\n        # XYZW -&amp;gt; WXYZ order of elements\n        quat = quat[[3,0,1,2]]\n</code></p>\n\n<p>or is the following line enough in my case?\n<code>\nquat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n</code></p>",
          "rawMarkdown": "Thanks for taking a look @rishabhiitbhu, I am using your code from here https://www.kaggle.com/rishabhiitbhu/visualizing-predictions for visualization and I didn't change anything so pretty sure the problem is not there, but I will check again in case I maybe did change something and don't remember. \n\nDo you think this piece of code from the other inference kernels is needed for the rotation?\n```\n        v = (sample_boxes[i,0] - sample_boxes[i,1]) # an edge vector\n        v /= np.linalg.norm(v) # normalized edge vector\n        r = R.from_dcm([\n            [v[0], -v[1], 0],\n            [v[1],  v[0], 0],\n            [   0,     0, 1],\n        ])  # a rotation matrix for rotation along z axis. \n        quat = r.as_quat()\n        # XYZW -&gt; WXYZ order of elements\n        quat = quat[[3,0,1,2]]\n```\n\nor is the following line enough in my case?\n```\nquat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n```\n",
          "votes": 1
        },
        {
          "id": 665283,
          "postDate": "2019-11-04T20:56:07.093Z",
          "content": "<p>I see you do <code>box3d[:, 6] = -box3d[:, 6] - np.pi / 2</code> which is specific for second.pytorch code and then <code>quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])</code> both are correct.</p>",
          "rawMarkdown": "I see you do `box3d[:, 6] = -box3d[:, 6] - np.pi / 2` which is specific for second.pytorch code and then `quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])` both are correct.",
          "votes": 2
        },
        {
          "id": 665296,
          "postDate": "2019-11-04T21:23:22.557Z",
          "content": "<p>Thank you again, this helps narrow down the problem for me.</p>\n\n<p>Maybe it is because I am doing the ego and global transformations before creating the boxes rather than after like how it is done in second? </p>",
          "rawMarkdown": "Thank you again, this helps narrow down the problem for me.\n\nMaybe it is because I am doing the ego and global transformations before creating the boxes rather than after like how it is done in second? ",
          "votes": 1
        },
        {
          "id": 665325,
          "postDate": "2019-11-04T21:57:45.353Z",
          "content": "<p>In second, you get the predictions, get yaw in lyft format (<code>box3d[:, 6] = -box3d[:, 6] - np.pi / 2</code>), get the boxes in ego vehicle's frame, then get the boxes in global frame (first rotation, then translation, in both cases) That's exactly what you have done, I'm not able find any bug in there.</p>",
          "rawMarkdown": "In second, you get the predictions, get yaw in lyft format (`box3d[:, 6] = -box3d[:, 6] - np.pi / 2`), get the boxes in ego vehicle's frame, then get the boxes in global frame (first rotation, then translation, in both cases) That's exactly what you have done, I'm not able find any bug in there.",
          "votes": 1
        },
        {
          "id": 665327,
          "postDate": "2019-11-04T22:00:06.443Z",
          "content": "<p>here's how I do this:\n```\nfrom nuscenes.eval.detection.utils import quaternion_yaw\n** import other stuff **</p>\n\n<p>def to_glb(box, info):\n    # lidar -&gt; ego -&gt; global\n    # info should belong to exact same element in <code>gt</code> dict\n    box.rotate(Quaternion(info['lidar2ego_rotation']))\n    box.translate(np.array(info['lidar2ego_translation']))</p>\n\n<pre><code>box.rotate(Quaternion(info['ego2global_rotation']))\nbox.translate(np.array(info['ego2global_translation']))\nreturn box\n</code></pre>\n\n<p>def get_pred_glb(pred, sample_token, form='str'):\n    boxes_lidar = pred[\"box3d_lidar\"]\n    boxes_class = pred[\"label_preds\"]\n    scores = pred['scores']\n    preds_classes = [classes[x] for x in boxes_class]\n    box_centers = boxes_lidar[:, :3]\n    box_yaws = boxes_lidar[:, -1]\n    box_wlh = boxes_lidar[:, 3:6]\n    info = token2info[sample_token] # a <code>sample</code> token\n    boxes = []\n    pred_str = ''\n    for idx in range(len(boxes_lidar)):\n        translation = box_centers[idx]\n        yaw = - box_yaws[idx] - pi/2 # second to lyft format\n        size = box_wlh[idx]\n        name = preds_classes[idx]\n        detection_score = scores[idx]\n        quat = Quaternion(scalar=np.cos(yaw / 2), vector=[0, 0, np.sin(yaw / 2)])\n        box = Box(\n            center=box_centers[idx],\n            size=size,\n            orientation=quat,\n            score=detection_score,\n            name=name,\n            token=sample_token\n        )\n        #return box\n        box = to_glb(box, info)\n        if form=='str':\n            pred =  str(box.score) + ' ' + str(box.center[0])  + ' '  + \\\n                    str(box.center[1]) + ' '  + str(box.center[2]) + ' '  + \\\n                    str(box.wlh[0]) + ' ' \\\n                    + str(box.wlh[1]) + ' '  + str(box.wlh[2]) + ' ' + str(quaternion_yaw(box.orientation)) + ' ' \\\n                    + str(name) + ' ' \n            pred_str += pred\n        else:\n            boxes.append(box)\n    if form=='str':\n        return pred_str.strip()\n    else:\n        return boxes\n```</p>",
          "rawMarkdown": "here's how I do this:\n```\nfrom nuscenes.eval.detection.utils import quaternion_yaw\n** import other stuff **\n\ndef to_glb(box, info):\n    # lidar -&gt; ego -&gt; global\n    # info should belong to exact same element in `gt` dict\n    box.rotate(Quaternion(info['lidar2ego_rotation']))\n    box.translate(np.array(info['lidar2ego_translation']))\n\n    box.rotate(Quaternion(info['ego2global_rotation']))\n    box.translate(np.array(info['ego2global_translation']))\n    return box\n\n\ndef get_pred_glb(pred, sample_token, form='str'):\n    boxes_lidar = pred[\"box3d_lidar\"]\n    boxes_class = pred[\"label_preds\"]\n    scores = pred['scores']\n    preds_classes = [classes[x] for x in boxes_class]\n    box_centers = boxes_lidar[:, :3]\n    box_yaws = boxes_lidar[:, -1]\n    box_wlh = boxes_lidar[:, 3:6]\n    info = token2info[sample_token] # a `sample` token\n    boxes = []\n    pred_str = ''\n    for idx in range(len(boxes_lidar)):\n        translation = box_centers[idx]\n        yaw = - box_yaws[idx] - pi/2 # second to lyft format\n        size = box_wlh[idx]\n        name = preds_classes[idx]\n        detection_score = scores[idx]\n        quat = Quaternion(scalar=np.cos(yaw / 2), vector=[0, 0, np.sin(yaw / 2)])\n        box = Box(\n            center=box_centers[idx],\n            size=size,\n            orientation=quat,\n            score=detection_score,\n            name=name,\n            token=sample_token\n        )\n        #return box\n        box = to_glb(box, info)\n        if form=='str':\n            pred =  str(box.score) + ' ' + str(box.center[0])  + ' '  + \\\n                    str(box.center[1]) + ' '  + str(box.center[2]) + ' '  + \\\n                    str(box.wlh[0]) + ' ' \\\n                    + str(box.wlh[1]) + ' '  + str(box.wlh[2]) + ' ' + str(quaternion_yaw(box.orientation)) + ' ' \\\n                    + str(name) + ' ' \n            pred_str += pred\n        else:\n            boxes.append(box)\n    if form=='str':\n        return pred_str.strip()\n    else:\n        return boxes\n```",
          "votes": 5
        },
        {
          "id": 665333,
          "postDate": "2019-11-04T22:12:11.897Z",
          "content": "<p>The main difference I see is you are using Box like in the original second code, I am using Box3d but I will try using Box instead and updating my prediction string creation to follow your code and see if that fixes it. Box also has the rotate and translate methods so that will be more convenient and less error prone. </p>\n\n<p>Thank you so much for sharing this, I'll try update my code and see if that fixes the issue. </p>",
          "rawMarkdown": "The main difference I see is you are using Box like in the original second code, I am using Box3d but I will try using Box instead and updating my prediction string creation to follow your code and see if that fixes it. Box also has the rotate and translate methods so that will be more convenient and less error prone. \n\nThank you so much for sharing this, I'll try update my code and see if that fixes the issue. ",
          "votes": 2
        },
        {
          "id": 665412,
          "postDate": "2019-11-05T01:28:18.370Z",
          "content": "<p>The way I use to check transformation is to see whether the points and boxes are consistently transformed. Say you have point_o and box_o and transformed to point_t and box_t. What I do is find the indices of point_o inside box_o and also the indices of point_t inside box_t. The two indices should be exactly same (up to numerical limit).   </p>",
          "rawMarkdown": "The way I use to check transformation is to see whether the points and boxes are consistently transformed. Say you have point_o and box_o and transformed to point_t and box_t. What I do is find the indices of point_o inside box_o and also the indices of point_t inside box_t. The two indices should be exactly same (up to numerical limit).   ",
          "votes": 1
        },
        {
          "id": 665483,
          "postDate": "2019-11-05T03:38:08.897Z",
          "content": "<p>Success! The boxes line up much better now. Thanks <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> </p>\n\n<p>There is still quite a few FP and FN but I'm not training on all the data yet so hopefully I can improve results and fixes some other issues with classes before the competition ends! 🙏 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fb73532e8c4dfd85ded9b927eaa09862b%2FScreen%20Shot%202019-11-04%20at%2010.30.34%20PM.png?generation=1572924676224488&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fb7827335b1105cca49fe9dad657ce3f3%2Fyaw_fixed_3d.png?generation=1572924602327411&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "Success! The boxes line up much better now. Thanks @rishabhiitbhu \n\nThere is still quite a few FP and FN but I'm not training on all the data yet so hopefully I can improve results and fixes some other issues with classes before the competition ends! 🙏 \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fb73532e8c4dfd85ded9b927eaa09862b%2FScreen%20Shot%202019-11-04%20at%2010.30.34%20PM.png?generation=1572924676224488&amp;alt=media)\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fb7827335b1105cca49fe9dad657ce3f3%2Fyaw_fixed_3d.png?generation=1572924602327411&amp;alt=media)\n",
          "votes": 3
        },
        {
          "id": 665496,
          "postDate": "2019-11-05T04:05:52.563Z",
          "content": "<p>Thanks <a href=\"/nywenjing\">@nywenjing</a> , I'll have to try that technique!</p>",
          "rawMarkdown": "Thanks @nywenjing , I'll have to try that technique!"
        },
        {
          "id": 665556,
          "postDate": "2019-11-05T05:30:39.357Z",
          "content": "<p>Good Job!\nSo, what was the bug then 🤔</p>",
          "rawMarkdown": "Good Job!\nSo, what was the bug then 🤔",
          "votes": 2
        },
        {
          "id": 666236,
          "postDate": "2019-11-05T22:58:08.100Z",
          "content": "<p>Thank you <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> .\nI'm not exactly sure what it was yet, I think it was the other the transformations were applied. </p>",
          "rawMarkdown": "Thank you @rishabhiitbhu .\nI'm not exactly sure what it was yet, I think it was the other the transformations were applied. "
        }
      ]
    },
    {
      "id": 664507,
      "postDate": "2019-11-03T19:05:14.857Z",
      "content": "<p>You might've been affected by this issue in the KITTI converter: <a href=\"https://github.com/lyft/nuscenes-devkit/pull/65\">https://github.com/lyft/nuscenes-devkit/pull/65</a></p>",
      "rawMarkdown": "You might've been affected by this issue in the KITTI converter: https://github.com/lyft/nuscenes-devkit/pull/65"
    },
    {
      "id": 688294,
      "postDate": "2019-12-05T12:38:00.120Z",
      "content": "<p>I believe it is because of the different definition of yaw. This makes some bugs in the rotation transformation. </p>",
      "rawMarkdown": "I believe it is because of the different definition of yaw. This makes some bugs in the rotation transformation. "
    },
    {
      "id": 670432,
      "postDate": "2019-11-11T13:04:49.960Z",
      "content": "<p>note that it seems the test set doesn’t include the side lidars.  so maybe don’t use them for training; the test clouds aren’t as dense :(</p>",
      "rawMarkdown": "note that it seems the test set doesn’t include the side lidars.  so maybe don’t use them for training; the test clouds aren’t as dense :("
    },
    {
      "id": 665138,
      "postDate": "2019-11-04T17:09:02.250Z",
      "content": "<p>Hi, I'm also trying PointRCNN but failed to add support for multiclass.... I have tried to use eval_rcnn to test the rpn part for the official pretrain model, but the result is very bad, rpn iou avg is 0 and bbox recall is very low. How about your result? I'm trying to debug... I have used the latest KITTI converter.😩 Thanks!</p>",
      "rawMarkdown": "Hi, I'm also trying PointRCNN but failed to add support for multiclass.... I have tried to use eval_rcnn to test the rpn part for the official pretrain model, but the result is very bad, rpn iou avg is 0 and bbox recall is very low. How about your result? I'm trying to debug... I have used the latest KITTI converter.😩 Thanks!",
      "replies": [
        {
          "id": 665254,
          "postDate": "2019-11-04T19:51:31.290Z",
          "content": "<p>It's a converter bug: alpha parameter in label file is set to default -10, and this param used in PointRCNN to augment data. This make inputs YAW of network noisy (i mean random).\nYou can off augmentations or refactor code to fix this (in code i used alpha calculcations was implemented, but not used)</p>\n\n<p>GL :)</p>",
          "rawMarkdown": "It's a converter bug: alpha parameter in label file is set to default -10, and this param used in PointRCNN to augment data. This make inputs YAW of network noisy (i mean random).\nYou can off augmentations or refactor code to fix this (in code i used alpha calculcations was implemented, but not used)\n\nGL :)",
          "votes": 1
        },
        {
          "id": 665335,
          "postDate": "2019-11-04T22:18:47.513Z",
          "content": "<p>Thanks! I just used the render_kitti function in nuscenes-devkit to render the data I converted, and they're all wrong.. I found that I haven't use the correct codes for converting.  If I used the correct codes for converting at least right now the visualization is good, do I still need to modified the alpha parameter?</p>",
          "rawMarkdown": "Thanks! I just used the render_kitti function in nuscenes-devkit to render the data I converted, and they're all wrong.. I found that I haven't use the correct codes for converting.  If I used the correct codes for converting at least right now the visualization is good, do I still need to modified the alpha parameter?"
        },
        {
          "id": 665364,
          "postDate": "2019-11-05T00:09:46.057Z",
          "content": "<p>Hmmmm, as i know this rendering tool doesn't use alpha parameter to visualization :/ and previosluy, when bug was unfounded, i render train and test both and get correct images :/ </p>",
          "rawMarkdown": "Hmmmm, as i know this rendering tool doesn't use alpha parameter to visualization :/ and previosluy, when bug was unfounded, i render train and test both and get correct images :/ \n",
          "votes": 1
        },
        {
          "id": 665475,
          "postDate": "2019-11-05T03:27:35.400Z",
          "content": "<p>I see, thanks!!</p>",
          "rawMarkdown": "I see, thanks!!"
        },
        {
          "id": 665531,
          "postDate": "2019-11-05T04:59:18.277Z",
          "content": "<p><a href=\"/jionie\">@jionie</a> what do you mean by correct codes? I am using the kernel by <a href=\"/stalkermustang\">@stalkermustang</a> to convert to KITTi format, did you use the same?</p>",
          "rawMarkdown": "@jionie what do you mean by correct codes? I am using the kernel by @stalkermustang to convert to KITTi format, did you use the same?"
        },
        {
          "id": 665548,
          "postDate": "2019-11-05T05:25:27.080Z",
          "content": "<p><a href=\"/stalkermustang\">@stalkermustang</a> does your conversion code for Lyft to KITTi is enough for PointRCNN or do we have to make any other changes?</p>",
          "rawMarkdown": "@stalkermustang does your conversion code for Lyft to KITTi is enough for PointRCNN or do we have to make any other changes?"
        },
        {
          "id": 665554,
          "postDate": "2019-11-05T05:29:17.500Z",
          "content": "<p>I used codes from here, <a href=\"https://github.com/lyft/nuscenes-devkit/pull/65\">https://github.com/lyft/nuscenes-devkit/pull/65</a></p>",
          "rawMarkdown": "I used codes from here, https://github.com/lyft/nuscenes-devkit/pull/65"
        },
        {
          "id": 665572,
          "postDate": "2019-11-05T06:04:55.190Z",
          "content": "<p><a href=\"/jionie\">@jionie</a> so you didn't use this kernel <a href=\"https://www.kaggle.com/stalkermustang/converting-lyft-dataset-to-kitty-format\">https://www.kaggle.com/stalkermustang/converting-lyft-dataset-to-kitty-format</a> ?</p>",
          "rawMarkdown": "@jionie so you didn't use this kernel https://www.kaggle.com/stalkermustang/converting-lyft-dataset-to-kitty-format ?"
        },
        {
          "id": 665594,
          "postDate": "2019-11-05T06:42:45.673Z",
          "content": "<p>you need fix alpha. It can be done in converter on pointrcnn directly.</p>",
          "rawMarkdown": "you need fix alpha. It can be done in converter on pointrcnn directly."
        },
        {
          "id": 665595,
          "postDate": "2019-11-05T06:45:58.960Z",
          "content": "<p>i think fix is need after <a href=\"https://github.com/lyft/nuscenes-devkit/commit/80fc3edf88fbf3cfa779eb7915e338c6a343acfd#diff-3f1352c0f518bcc0109188caba927807R591\">https://github.com/lyft/nuscenes-devkit/commit/80fc3edf88fbf3cfa779eb7915e338c6a343acfd#diff-3f1352c0f518bcc0109188caba927807R591</a> changes, so i dont update my tool and convert still looks good on visualization. Don't submit my currentlu baseline, rpn still training. Well...see later what's correct :)</p>",
          "rawMarkdown": "i think fix is need after https://github.com/lyft/nuscenes-devkit/commit/80fc3edf88fbf3cfa779eb7915e338c6a343acfd#diff-3f1352c0f518bcc0109188caba927807R591 changes, so i dont update my tool and convert still looks good on visualization. Don't submit my currentlu baseline, rpn still training. Well...see later what's correct :)"
        },
        {
          "id": 665611,
          "postDate": "2019-11-05T07:15:09.750Z",
          "content": "<p>I had a similar situation, during the RPN phase (50 epoch), the recall was very low, no more than 0.40. \n The analysis of the output of the network showed a good segmentation effect, but the pred_reg was somewhat bad, the predicted center points were scattered everywhere (but some of them could capture the position of GT), increasing training time doesn't help much.</p>\n\n<p>PS: I didn't use any augmentations.</p>",
          "rawMarkdown": "I had a similar situation, during the RPN phase (50 epoch), the recall was very low, no more than 0.40. \n The analysis of the output of the network showed a good segmentation effect, but the pred_reg was somewhat bad, the predicted center points were scattered everywhere (but some of them could capture the position of GT), increasing training time doesn't help much.\n\nPS: I didn't use any augmentations."
        },
        {
          "id": 665616,
          "postDate": "2019-11-05T07:22:40.887Z",
          "content": "<p><a href=\"/stalkermustang\">@stalkermustang</a> so you are saying conversion to KITTI from the new fixes of lyft_dataset_sdk does not need the fix of <code>alpha</code> ? Also where do you fix the <code>alpha</code> in KITTI POINTRCNN dataset?</p>",
          "rawMarkdown": "@stalkermustang so you are saying conversion to KITTI from the new fixes of lyft_dataset_sdk does not need the fix of `alpha` ? Also where do you fix the `alpha` in KITTI POINTRCNN dataset?"
        },
        {
          "id": 666057,
          "postDate": "2019-11-05T17:13:17.630Z",
          "content": "<p>i didn't say that. You need fix alpha additionaly. I'm manually calc it in dataloader.</p>",
          "rawMarkdown": "i didn't say that. You need fix alpha additionaly. I'm manually calc it in dataloader."
        }
      ]
    },
    {
      "id": 670174,
      "postDate": "2019-11-11T05:57:00.367Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 665644,
      "postDate": "2019-11-05T08:02:53.487Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 664010,
      "author_name": "Wenjing Zhang",
      "author_url": "",
      "post_date": "2019-11-03T03:04:22.500000",
      "content": "<p>From Lyft to PointRCNN: \npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground) </p>\n\n<p>I think this is the correct conversion. Hope it helps. But you need to double check it. Be very careful!\nKeep going. You are actually doing very well.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 664504,
          "author_name": "William Horton",
          "author_url": "",
          "post_date": "2019-11-03T18:55:08.787000",
          "content": "<p>I'm curious, at what point do these changes need to be applied? Do you train the PointRCNN as is and then convert the predictions? Or do you have to apply transformations to the data before you can train PointRCNN?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 664540,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-03T20:42:03.973000",
          "content": "<p>Zhang, thanks for reply and good words! \nI'm debugging dataloader step by step and check augmentations in PointRCNN. I found that authors use alpha when rotate boxes. But alpha after converting from LYFT was constant -10 xD \nI change this part (recalc alpha from yaw) and run retrain network. Tomorrow see what i'm training - \nI am full of hopes and expectations.</p>\n\n<p>How you check that yours converting is right? Do you render submission? I'm render .csv using that kernel: <a href=\"https://www.kaggle.com/rishabhiitbhu/visualizing-predictions\">https://www.kaggle.com/rishabhiitbhu/visualizing-predictions</a> and try to check posiitons, classes and rotations.</p>\n\n<p>Good luck in the Competition!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 664602,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-04T00:26:24.630000",
          "content": "<p>One way to check whether your conversion is correct is to check whether the same points are in the same box. Say you have point and box in lyft and convert it to PointRCNN as point-cnn and box-cnn. Now you can use points_in_box in lyft utils to find which points are inside the box. You can also use a similar function in PointRCNN to find which point-cnn are inside box-cnn. The indices of the two findings should be the same.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 664603,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-04T00:28:32.367000",
          "content": "<p>Generally, you need to convert points and box from lyft toPointRCNN first and do the training and prediction, then you need to convert the prediction back from PointRCNN to lyft</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 664889,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-04T11:05:48.210000",
          "content": "<p>Today i  see retrained PointRCNN network with totation angle bug fix. So they looks much better (on second scrren left big bbox also classifiend as truck, wow):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Ff269bb111604680306e4db6695c5b81a%2Fphoto_2019-11-04_05-11-56.jpg?generation=1572865402713654&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8a5a1fe7522c579bb10b8a281ac1aac9%2Fphoto_2019-11-04_04-02-34.jpg?generation=1572865378923645&amp;alt=media\" alt=\"\">\nSo i submit and got 0.007 on lb (only front camera used at this time).\nZhang, how do you rate this score for PointRCNN? \nI have a feeling that somewhere there is a mistake in the conversion, since everything is cool on the visualization, and the cars are detected well.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 664900,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-04T11:20:52.180000",
          "content": "<p>what do you mean by only front camera used? it is strange, the score is too low while the boxes look fine. Do you convert the box back to lyft format correctly? also rember to transform the boxes back to world coordinates.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 664907,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-04T11:39:43.330000",
          "content": "<p>I use front camera only for time, just check and debugging - after got a working pipeline i'm add another cameras and increase score (dream...).\nYes, i'm converting to world coordinates. \nFor example,  can you check one bbox of one token in your submission?\nTOKEN: <code>da0671a11e837c86c131f04ded4fc91941b936f70fdead745259361f1f06f250</code>\ncoordinates : 2231.559051779429 987.5928335963141 -17.873463805634827\nwlh: 2.946 11.6001 3.5798 \nyaw: 2.5836576429783755\nclass: bus\ndoes it look like what you wrote in .csv submission?\nIt's that bus (or not bus, but i'm talking only for checking coordinates and yaw):\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F27eac95b6cbcc58231e8b7d3101279e4%2FScreenshot_2.png?generation=1572867506015369&amp;alt=media\" alt=\"\"></p>\n\n<p>think competition host will not ban us for one of thousands rpedictions public share)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 664924,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-04T12:12:01.350000",
          "content": "<p>Still have no idea how you only use the front camera. But I previously made a mistake to clip the detected boxes: only keep the boxes with x&gt;0 (I forget the exact number, but somehow like this). After the correction, I see  the score increases more than tenfold . If your implementation(like only using the front camera) faces the same problem, your detection should actually be good.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 664947,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-04T12:57:10.383000",
          "content": "<p>In PointRCNN implemetation i used all boxes outside camera FOV are cuted (on train and inference stages both). I don't fix that, and currently train on 6 cameras, but inference test set for 1 camera. They clip boxes by x/z axis too, and i know that. So if score on 1 camera is 0.007, 6 cameras maximum gain is 0.042, it's lower than unet on BEV :) </p>\n\n<p>what do you think, what score PointRCNN can schive on class Car on ths task?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 664953,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-04T13:10:54.850000",
          "content": "<p>I think the increase will be much more than that. As I mentioned, I once wrongly removed around half or more of the detected boxes and the score drops from 0.086 to 0.006. If you fix all the issues, the score should be well above 0.1 (including all classes, no idea what you will get if you use car only, maybe 0.05?) </p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 665149,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-04T17:16:12.703000",
          "content": "<p>Zhang, is 0.086 score of pure multiclass PointRCNN? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665290,
          "author_name": "Usmann Khan",
          "author_url": "",
          "post_date": "2019-11-04T21:15:46.030000",
          "content": "<blockquote>\n  <p>I think the increase will be much more than that. As I mentioned, I once wrongly removed around half or more of the detected boxes and the score drops from 0.086 to 0.006. If you fix all the issues, the score should be well above 0.1 (including all classes, no idea what you will get if you use car only, maybe 0.05?)</p>\n</blockquote>\n\n<p><a href=\"/nywenjing\">@nywenjing</a> How many epochs do you think it will take to reach that?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665404,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-05T01:17:25.870000",
          "content": "<p>Actually, I did not perform a full submission with PointRCNN. I only gave a tried and finished the data conversion and train several epochs for RPN. The score is obtained with another implementation. But  I think the mistake might be similar. Here are my first submission scores:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3701925%2F185a9ff84628ff68656c93ae761cc1cd%2FScreenshot%20from%202019-11-05%2008-53-11.png?generation=1572915721940258&amp;alt=media\" alt=\"\"></p>\n\n<p>The first score is almost 0 since I wrongly transform the box back to world coordinates.</p>\n\n<p>The second is 0.006 with corrected transformation to world coordinates. It is still very low. I was frustrating then. I can believe it could be so bad. So I give a try with the baseline.</p>\n\n<p>The third with 0.036 is my try of baseline then I found the yaw angle problem and correct it. I thought now I might be right. </p>\n\n<p>But the forth is still 0.006. I was very sad at that point. But when I try to do prediction on the training set, I found my predicted boxes are actually good but seems only half are present (since the other parts are wrongly removed). Then I did the correction. Just as <a href=\"/stalkermustang\">@stalkermustang</a> , I only thought It might give me 0.012 score, which is still well below the baseline. So I did not give much expectation. </p>\n\n<p>To my surprise, the corrected score was increase more than tenfold to 0.089 at the fifth submission. So <a href=\"/stalkermustang\">@stalkermustang</a>, if you face the same issue, I think you are actually very near to the right one. Keep going!</p>\n\n<p><a href=\"/usmannkhan\">@usmannkhan</a> , I have not run the pipeline of PointRCNN fully. But I guess the RPN may need 60 epochs. The author use 200 epochs for kitti. Lyft have around 4 times more data so we might need 40 epochs.  Considering there are more classes in Lyft, I will give another 20 epochs.</p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 665568,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-11-05T05:47:51.607000",
          "content": "<p><a href=\"/nywenjing\">@nywenjing</a> when you say that you have to convert Lyft to PointRCNN here\n<code>\nFrom Lyft to PointRCNN:\npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground)\n</code>\nwhat exactly do you mean? I mean I converted Lyft dataset to KITTI using <a href=\"/stalkermustang\">@stalkermustang</a> kernel. Do I have apply the above tranformation while loading the KITTI dataset in the PointRCNN codebase or I am doing something wrong?\nHelp would be appreciated.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665574,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-05T06:09:55.050000",
          "content": "<p>I did not look into <a href=\"/stalkermustang\">@stalkermustang</a>  kernel. It seems too complex for me. I only need point and have nothing to do with the Camera. If you want to train PointRCNN on Lyft, you need to first convert the point cloud and gt boxes to its corresponding format, which is what my code do. If you can ensure <a href=\"/stalkermustang\">@stalkermustang</a>  kernel code does the correct conversion, then you can just do the conversion and let it run. The conversion code is only applied to Lyft dataset.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665605,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-05T07:04:15.730000",
          "content": "<p>Maybe this is the score, that can't be achived with PointRCNN, because you use another RCNN part? Ypu write only about RPN :( So this score is score of 2-3-4 well detected classes..?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 664004,
      "author_name": "Wenjing Zhang",
      "author_url": "",
      "post_date": "2019-11-03T02:50:30.183000",
      "content": "<p>From my perspective of view, it might be the coordinate system problem. Lyft  uses z up and size wlh but pointrcnn uses -y up and size hwl. You need to be very careful to convert the boxes and point clouds between lyft and pointrcnn. You should also be careful of the sign of yaw angle. Some use clockwise as positive while some use anti clockwise as positive. I have not checked it. You need to double check it. I feel your prediction is correct up to a sign and/or an offset of the yaw angle.</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 665245,
      "author_name": "Jack Vial",
      "author_url": "",
      "post_date": "2019-11-04T19:36:08.377000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fc8466e24b2bad268fa4cff18f6fcea7e%2FScreen%20Shot%202019-11-04%20at%202.22.42%20PM.png?generation=1572895519112648&amp;alt=media\" alt=\"\"></p>\n\n<p>I am seeing a similar problem with my predictions. I am using PointPillars based on this code <a href=\"https://github.com/traveller59/second.pytorch\">https://github.com/traveller59/second.pytorch</a> .</p>\n\n<p>The ground truth in orange looks ok but the predictions in blue all look off by a similar angle. When creating the predictions strings I use the same transformation on to get yaw </p>\n\n<p><code>yaw = 2*np.arccos(box3ds[i].rotation[0])</code></p>\n\n<p>and this is how I create the boxes:</p>\n\n<p>```\n    def _box_translate(self, translation, x):\n        # box.center += x\n        # In Box3d center is called translation\n        translation += x\n        return translation</p>\n\n<pre><code>def _box_rotate(self, t, r, quaternion):\n    # box.center = np.dot(quaternion.rotation_matrix, box.center)\n    # In Box3d center is called translation\n    t = np.dot(quaternion.rotation_matrix, t)\n    # box.orientation = quaternion * box.orientation\n    # In Box3d orientation is called rotation, I think\n    r = quaternion * r\n    # box.velocity = np.dot(quaternion.rotation_matrix,box.velocity)\n    return t, r\n\ndef _get_lyft_pred3d_boxes(self, detection, token2info):\n    pred_box3ds = []\n    box3d = detection[\"box3d_lidar\"].detach().cpu().numpy()\n    scores = detection[\"scores\"].detach().cpu().numpy()\n    labels = detection[\"label_preds\"].detach().cpu().numpy()\n    sample_token = detection[\"metadata\"][\"token\"]\n\n    box3d[:, 6] = -box3d[:, 6] - np.pi / 2\n    for i in range(box3d.shape[0]):\n\n        # Rotate about the z axis by the z value and create a new Quaternion\n        quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n        translation = box3d[i, :3]\n        size = box3d[i, 3:6]\n\n        # Say everything is a car until we figure out the issues with the classes\n        model_class_names =  ['car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car']\n        class_name = model_class_names[labels[i]]\n\n        trans_info = token2info[detection[\"metadata\"][\"token\"]]\n\n        ###### Move box to ego vehicle coord system #######\n        translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['lidar2ego_rotation']))\n        translation = self._box_translate(translation, np.array(trans_info['lidar2ego_translation']))\n\n        ####### Move box to global coord system ########\n        translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['ego2global_rotation']))\n        translation = self._box_translate(translation, np.array(trans_info['ego2global_translation']))\n\n        box = Box3D(\n            sample_token=sample_token,\n            translation=list(translation),\n            size=list(size),\n            rotation=list(quat),\n            name=class_name,\n            score=scores[i]\n        )\n        pred_box3ds.append(box)\n\n    return pred_box3ds\n</code></pre>\n\n<p>```</p>\n\n<p>I think the predictions of the model are ok but I'm going wrong somewhere applying the rotations in post processing.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 665260,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-11-04T20:10:24.923000",
          "content": "<p>your code snippet looks fine to me, maybe there's a bug in plotting the boxes?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665278,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-11-04T20:51:11.007000",
          "content": "<p>Thanks for taking a look <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a>, I am using your code from here <a href=\"https://www.kaggle.com/rishabhiitbhu/visualizing-predictions\">https://www.kaggle.com/rishabhiitbhu/visualizing-predictions</a> for visualization and I didn't change anything so pretty sure the problem is not there, but I will check again in case I maybe did change something and don't remember. </p>\n\n<p>Do you think this piece of code from the other inference kernels is needed for the rotation?\n<code>\n        v = (sample_boxes[i,0] - sample_boxes[i,1]) # an edge vector\n        v /= np.linalg.norm(v) # normalized edge vector\n        r = R.from_dcm([\n            [v[0], -v[1], 0],\n            [v[1],  v[0], 0],\n            [   0,     0, 1],\n        ])  # a rotation matrix for rotation along z axis. \n        quat = r.as_quat()\n        # XYZW -&amp;gt; WXYZ order of elements\n        quat = quat[[3,0,1,2]]\n</code></p>\n\n<p>or is the following line enough in my case?\n<code>\nquat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665283,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-11-04T20:56:07.093000",
          "content": "<p>I see you do <code>box3d[:, 6] = -box3d[:, 6] - np.pi / 2</code> which is specific for second.pytorch code and then <code>quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])</code> both are correct.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665296,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-11-04T21:23:22.557000",
          "content": "<p>Thank you again, this helps narrow down the problem for me.</p>\n\n<p>Maybe it is because I am doing the ego and global transformations before creating the boxes rather than after like how it is done in second? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665325,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-11-04T21:57:45.353000",
          "content": "<p>In second, you get the predictions, get yaw in lyft format (<code>box3d[:, 6] = -box3d[:, 6] - np.pi / 2</code>), get the boxes in ego vehicle's frame, then get the boxes in global frame (first rotation, then translation, in both cases) That's exactly what you have done, I'm not able find any bug in there.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665327,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-11-04T22:00:06.443000",
          "content": "<p>here's how I do this:\n```\nfrom nuscenes.eval.detection.utils import quaternion_yaw\n** import other stuff **</p>\n\n<p>def to_glb(box, info):\n    # lidar -&gt; ego -&gt; global\n    # info should belong to exact same element in <code>gt</code> dict\n    box.rotate(Quaternion(info['lidar2ego_rotation']))\n    box.translate(np.array(info['lidar2ego_translation']))</p>\n\n<pre><code>box.rotate(Quaternion(info['ego2global_rotation']))\nbox.translate(np.array(info['ego2global_translation']))\nreturn box\n</code></pre>\n\n<p>def get_pred_glb(pred, sample_token, form='str'):\n    boxes_lidar = pred[\"box3d_lidar\"]\n    boxes_class = pred[\"label_preds\"]\n    scores = pred['scores']\n    preds_classes = [classes[x] for x in boxes_class]\n    box_centers = boxes_lidar[:, :3]\n    box_yaws = boxes_lidar[:, -1]\n    box_wlh = boxes_lidar[:, 3:6]\n    info = token2info[sample_token] # a <code>sample</code> token\n    boxes = []\n    pred_str = ''\n    for idx in range(len(boxes_lidar)):\n        translation = box_centers[idx]\n        yaw = - box_yaws[idx] - pi/2 # second to lyft format\n        size = box_wlh[idx]\n        name = preds_classes[idx]\n        detection_score = scores[idx]\n        quat = Quaternion(scalar=np.cos(yaw / 2), vector=[0, 0, np.sin(yaw / 2)])\n        box = Box(\n            center=box_centers[idx],\n            size=size,\n            orientation=quat,\n            score=detection_score,\n            name=name,\n            token=sample_token\n        )\n        #return box\n        box = to_glb(box, info)\n        if form=='str':\n            pred =  str(box.score) + ' ' + str(box.center[0])  + ' '  + \\\n                    str(box.center[1]) + ' '  + str(box.center[2]) + ' '  + \\\n                    str(box.wlh[0]) + ' ' \\\n                    + str(box.wlh[1]) + ' '  + str(box.wlh[2]) + ' ' + str(quaternion_yaw(box.orientation)) + ' ' \\\n                    + str(name) + ' ' \n            pred_str += pred\n        else:\n            boxes.append(box)\n    if form=='str':\n        return pred_str.strip()\n    else:\n        return boxes\n```</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 665333,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-11-04T22:12:11.897000",
          "content": "<p>The main difference I see is you are using Box like in the original second code, I am using Box3d but I will try using Box instead and updating my prediction string creation to follow your code and see if that fixes it. Box also has the rotate and translate methods so that will be more convenient and less error prone. </p>\n\n<p>Thank you so much for sharing this, I'll try update my code and see if that fixes the issue. </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 665412,
          "author_name": "Wenjing Zhang",
          "author_url": "",
          "post_date": "2019-11-05T01:28:18.370000",
          "content": "<p>The way I use to check transformation is to see whether the points and boxes are consistently transformed. Say you have point_o and box_o and transformed to point_t and box_t. What I do is find the indices of point_o inside box_o and also the indices of point_t inside box_t. The two indices should be exactly same (up to numerical limit).   </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665483,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-11-05T03:38:08.897000",
          "content": "<p>Success! The boxes line up much better now. Thanks <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> </p>\n\n<p>There is still quite a few FP and FN but I'm not training on all the data yet so hopefully I can improve results and fixes some other issues with classes before the competition ends! 🙏 </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fb73532e8c4dfd85ded9b927eaa09862b%2FScreen%20Shot%202019-11-04%20at%2010.30.34%20PM.png?generation=1572924676224488&amp;alt=media\" alt=\"\"></p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fb7827335b1105cca49fe9dad657ce3f3%2Fyaw_fixed_3d.png?generation=1572924602327411&amp;alt=media\" alt=\"\"></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 665496,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-11-05T04:05:52.563000",
          "content": "<p>Thanks <a href=\"/nywenjing\">@nywenjing</a> , I'll have to try that technique!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665556,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-11-05T05:30:39.357000",
          "content": "<p>Good Job!\nSo, what was the bug then 🤔</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 666236,
          "author_name": "Jack Vial",
          "author_url": "",
          "post_date": "2019-11-05T22:58:08.100000",
          "content": "<p>Thank you <a href=\"/rishabhiitbhu\">@rishabhiitbhu</a> .\nI'm not exactly sure what it was yet, I think it was the other the transformations were applied. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 664507,
      "author_name": "William Horton",
      "author_url": "",
      "post_date": "2019-11-03T19:05:14.857000",
      "content": "<p>You might've been affected by this issue in the KITTI converter: <a href=\"https://github.com/lyft/nuscenes-devkit/pull/65\">https://github.com/lyft/nuscenes-devkit/pull/65</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 688294,
      "author_name": "Xianzhong",
      "author_url": "",
      "post_date": "2019-12-05T12:38:00.120000",
      "content": "<p>I believe it is because of the different definition of yaw. This makes some bugs in the rotation transformation. </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 670432,
      "author_name": "oarph",
      "author_url": "",
      "post_date": "2019-11-11T13:04:49.960000",
      "content": "<p>note that it seems the test set doesn’t include the side lidars.  so maybe don’t use them for training; the test clouds aren’t as dense :(</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665138,
      "author_name": "jionie",
      "author_url": "",
      "post_date": "2019-11-04T17:09:02.250000",
      "content": "<p>Hi, I'm also trying PointRCNN but failed to add support for multiclass.... I have tried to use eval_rcnn to test the rpn part for the official pretrain model, but the result is very bad, rpn iou avg is 0 and bbox recall is very low. How about your result? I'm trying to debug... I have used the latest KITTI converter.😩 Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 665254,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-04T19:51:31.290000",
          "content": "<p>It's a converter bug: alpha parameter in label file is set to default -10, and this param used in PointRCNN to augment data. This make inputs YAW of network noisy (i mean random).\nYou can off augmentations or refactor code to fix this (in code i used alpha calculcations was implemented, but not used)</p>\n\n<p>GL :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665335,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-11-04T22:18:47.513000",
          "content": "<p>Thanks! I just used the render_kitti function in nuscenes-devkit to render the data I converted, and they're all wrong.. I found that I haven't use the correct codes for converting.  If I used the correct codes for converting at least right now the visualization is good, do I still need to modified the alpha parameter?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665364,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-05T00:09:46.057000",
          "content": "<p>Hmmmm, as i know this rendering tool doesn't use alpha parameter to visualization :/ and previosluy, when bug was unfounded, i render train and test both and get correct images :/ </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 665475,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-11-05T03:27:35.400000",
          "content": "<p>I see, thanks!!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665531,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-11-05T04:59:18.277000",
          "content": "<p><a href=\"/jionie\">@jionie</a> what do you mean by correct codes? I am using the kernel by <a href=\"/stalkermustang\">@stalkermustang</a> to convert to KITTi format, did you use the same?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665548,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-11-05T05:25:27.080000",
          "content": "<p><a href=\"/stalkermustang\">@stalkermustang</a> does your conversion code for Lyft to KITTi is enough for PointRCNN or do we have to make any other changes?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665554,
          "author_name": "jionie",
          "author_url": "",
          "post_date": "2019-11-05T05:29:17.500000",
          "content": "<p>I used codes from here, <a href=\"https://github.com/lyft/nuscenes-devkit/pull/65\">https://github.com/lyft/nuscenes-devkit/pull/65</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665572,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-11-05T06:04:55.190000",
          "content": "<p><a href=\"/jionie\">@jionie</a> so you didn't use this kernel <a href=\"https://www.kaggle.com/stalkermustang/converting-lyft-dataset-to-kitty-format\">https://www.kaggle.com/stalkermustang/converting-lyft-dataset-to-kitty-format</a> ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665594,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-05T06:42:45.673000",
          "content": "<p>you need fix alpha. It can be done in converter on pointrcnn directly.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665595,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-05T06:45:58.960000",
          "content": "<p>i think fix is need after <a href=\"https://github.com/lyft/nuscenes-devkit/commit/80fc3edf88fbf3cfa779eb7915e338c6a343acfd#diff-3f1352c0f518bcc0109188caba927807R591\">https://github.com/lyft/nuscenes-devkit/commit/80fc3edf88fbf3cfa779eb7915e338c6a343acfd#diff-3f1352c0f518bcc0109188caba927807R591</a> changes, so i dont update my tool and convert still looks good on visualization. Don't submit my currentlu baseline, rpn still training. Well...see later what's correct :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665611,
          "author_name": "gakki",
          "author_url": "",
          "post_date": "2019-11-05T07:15:09.750000",
          "content": "<p>I had a similar situation, during the RPN phase (50 epoch), the recall was very low, no more than 0.40. \n The analysis of the output of the network showed a good segmentation effect, but the pred_reg was somewhat bad, the predicted center points were scattered everywhere (but some of them could capture the position of GT), increasing training time doesn't help much.</p>\n\n<p>PS: I didn't use any augmentations.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 665616,
          "author_name": "Ram Ramrakhya",
          "author_url": "",
          "post_date": "2019-11-05T07:22:40.887000",
          "content": "<p><a href=\"/stalkermustang\">@stalkermustang</a> so you are saying conversion to KITTI from the new fixes of lyft_dataset_sdk does not need the fix of <code>alpha</code> ? Also where do you fix the <code>alpha</code> in KITTI POINTRCNN dataset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 666057,
          "author_name": "Igor Kotenkov",
          "author_url": "",
          "post_date": "2019-11-05T17:13:17.630000",
          "content": "<p>i didn't say that. You need fix alpha additionaly. I'm manually calc it in dataloader.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 670174,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-11T05:57:00.367000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 665644,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-05T08:02:53.487000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "663609": "Hello all participants! \nI'm trying to solve 3D object detection task on pointclouds with PointRCNN network. I use official implementation from here: https://github.com/sshaoshuai/PointRCNN .\nWhen i run inference on pretrained model (checkpoint link on repo) i got normal predictions (only shown one of predictions, all other also well detected):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F8f8db26537d153e61d3be91338ef3ca3%2Fphoto_2019-10-27_11-05-34.jpg?generation=1572695211708053&amp;alt=media)\nThen i try train network from scratch - firstly RPN and then RCNN. I used 5000 samples from LYFT dataset, with 6 camera it's a 30000 samples - approx. x8 from original KITTI train data.\nRPN was trained on 50 epochs (intead 200 from repo, but dataset is larger), RCNN has 10 epochs (instead of 70). \n\nThe day before yesterday i run inference on test set and got this terrible PROBLEM:\npredictions have well translation, but absolutely disoriented (have bad YAW parameter):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa354920533067b35152c0f38bf581ce7%2FScreenshot_1.png?generation=1572695601960018&amp;alt=media)\n\nThen i use debugger and take step-by-step in-depth look, where this YAW was generated - maybe the network gives the right YAWs, but I have an error in postprocessing (at the same time, it’s strange that the pretrained checkpoint works fine). But i find that end writted YAW parameter is the same that generated as ROI in network, before NMS, limitations and other operations (\nafter the decode bbox from bins, that described in article, the angle is slightly different, but this is understandable - after all, this is the result of several proposals nearby, i.e. 2.41 YAW become 2.47)\nSo, I almost ran out of ideas, what else can I check and what to do next. I'm almost desperate :)\nThe last hypothesis is that the error is somewhere in the data loader, and the angle is literally white noise, so the neural network cannot learn it and outputs ~random~ +-, so i can check this today. \n\n\nMaybe someone already retrained this network from scratch? and was able to avoid this problem? or fix it? or do you have any suggestions what else to check?\n\n\nIt's a shame to spend so much time and have no result ;(  \nI added classes, since the network initially learns only for cars, and they even work, but there is no point in this without correct YAW (green associated with class \"bicycle\"):\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2Fa8ee160c5f893f60ac0e057b8f6eeedd%2FScreenshot_2.png?generation=1572696349587432&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1002569%2F2438ff8b2466cc94138bab1e1e6b0c00%2Fphoto_2019-11-01_00-18-22.jpg?generation=1572696398633333&amp;alt=media)\n\nso another idea is network is underfitted in relation to the YAW prediction, idk\n\nPS: yes, as you may have noticed - I'm an native speaker, so I apologize (:kekeke:)",
    "664010": "From Lyft to PointRCNN: \npoint cloud [x,y,z] -&gt;[x,-z,y]\nboxes: [x,y,z,w,l,h,yaw] -&gt;[x,-z + 0.5 * h,y,h,l,w,-yaw - np.pi /2 ] ## box center at PointRCNN is actually at the bottom (ground) \n\nI think this is the correct conversion. Hope it helps. But you need to double check it. Be very careful!\nKeep going. You are actually doing very well.",
    "664004": "From my perspective of view, it might be the coordinate system problem. Lyft  uses z up and size wlh but pointrcnn uses -y up and size hwl. You need to be very careful to convert the boxes and point clouds between lyft and pointrcnn. You should also be careful of the sign of yaw angle. Some use clockwise as positive while some use anti clockwise as positive. I have not checked it. You need to double check it. I feel your prediction is correct up to a sign and/or an offset of the yaw angle.",
    "665245": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1005864%2Fc8466e24b2bad268fa4cff18f6fcea7e%2FScreen%20Shot%202019-11-04%20at%202.22.42%20PM.png?generation=1572895519112648&amp;alt=media)\n\nI am seeing a similar problem with my predictions. I am using PointPillars based on this code https://github.com/traveller59/second.pytorch .\n\nThe ground truth in orange looks ok but the predictions in blue all look off by a similar angle. When creating the predictions strings I use the same transformation on to get yaw \n\n`yaw = 2*np.arccos(box3ds[i].rotation[0])`\n\nand this is how I create the boxes:\n\n```\n    def _box_translate(self, translation, x):\n        # box.center += x\n        # In Box3d center is called translation\n        translation += x\n        return translation\n\n    def _box_rotate(self, t, r, quaternion):\n        # box.center = np.dot(quaternion.rotation_matrix, box.center)\n        # In Box3d center is called translation\n        t = np.dot(quaternion.rotation_matrix, t)\n        # box.orientation = quaternion * box.orientation\n        # In Box3d orientation is called rotation, I think\n        r = quaternion * r\n        # box.velocity = np.dot(quaternion.rotation_matrix,box.velocity)\n        return t, r\n    \n    def _get_lyft_pred3d_boxes(self, detection, token2info):\n        pred_box3ds = []\n        box3d = detection[\"box3d_lidar\"].detach().cpu().numpy()\n        scores = detection[\"scores\"].detach().cpu().numpy()\n        labels = detection[\"label_preds\"].detach().cpu().numpy()\n        sample_token = detection[\"metadata\"][\"token\"]\n        \n        box3d[:, 6] = -box3d[:, 6] - np.pi / 2\n        for i in range(box3d.shape[0]):\n            \n            # Rotate about the z axis by the z value and create a new Quaternion\n            quat = pyquaternion.Quaternion(axis=[0, 0, 1], radians=box3d[i, 6])\n            translation = box3d[i, :3]\n            size = box3d[i, 3:6]\n           \n            # Say everything is a car until we figure out the issues with the classes\n            model_class_names =  ['car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car', 'car']\n            class_name = model_class_names[labels[i]]\n            \n            trans_info = token2info[detection[\"metadata\"][\"token\"]]\n            \n            ###### Move box to ego vehicle coord system #######\n            translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['lidar2ego_rotation']))\n            translation = self._box_translate(translation, np.array(trans_info['lidar2ego_translation']))\n            \n            ####### Move box to global coord system ########\n            translation, quat = self._box_rotate(translation, quat, pyquaternion.Quaternion(trans_info['ego2global_rotation']))\n            translation = self._box_translate(translation, np.array(trans_info['ego2global_translation']))\n            \n            box = Box3D(\n                sample_token=sample_token,\n                translation=list(translation),\n                size=list(size),\n                rotation=list(quat),\n                name=class_name,\n                score=scores[i]\n            )\n            pred_box3ds.append(box)\n            \n        return pred_box3ds\n```\n\nI think the predictions of the model are ok but I'm going wrong somewhere applying the rotations in post processing.",
    "664507": "You might've been affected by this issue in the KITTI converter: https://github.com/lyft/nuscenes-devkit/pull/65",
    "688294": "I believe it is because of the different definition of yaw. This makes some bugs in the rotation transformation. ",
    "670432": "note that it seems the test set doesn’t include the side lidars.  so maybe don’t use them for training; the test clouds aren’t as dense :(",
    "665138": "Hi, I'm also trying PointRCNN but failed to add support for multiclass.... I have tried to use eval_rcnn to test the rpn part for the official pretrain model, but the result is very bad, rpn iou avg is 0 and bbox recall is very low. How about your result? I'm trying to debug... I have used the latest KITTI converter.😩 Thanks!",
    "670174": "",
    "665644": ""
  }
}