{
  "id": 123261,
  "title": "The Other 6DVNET",
  "url": "/competitions/pku-autonomous-driving/discussion/123261",
  "author_name": "",
  "post_date": "2019-12-26T08:20:53.682594800Z",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I've seen a few posts from people who have produced some bad results using 6DVNET  (<a href=\"https://github.com/stevenwudi/6DVNET\">https://github.com/stevenwudi/6DVNET</a>), however I haven't anything yet about the related ApolloScape repo (<a href=\"https://github.com/stevenwudi/ApolloScape_InstanceSeg\">https://github.com/stevenwudi/ApolloScape_InstanceSeg</a>). I'm just going to throw a few questions out in hopes that someone (or maybe even <a href=\"/stevenwudi\">@stevenwudi</a> himself) has some answers.</p>\n\n<p>1) Has anyone produced good results using this? My LB scores using this are 0.003-0.004. I've even tried overfitting to the training set an things still never look that great when plotted. I've attached a few images of results on training set images.\n2) What is the expected order of inputs? According to the Data page the data we are provided is (yaw, pitch, roll, x, y, z). Can this be fed to the model in this same order or is it expecting some other order? I've tried changing the orders of yaw, pitch, and roll (I'm assuming x,y,z are in the correct place and correct order), but my results don't seem to change a whole lot.</p>",
  "messages": [
    {
      "id": "703471",
      "postDate": "12/26/2019 08:20:53",
      "content": "<p>I've seen a few posts from people who have produced some bad results using 6DVNET  (<a href=\"https://github.com/stevenwudi/6DVNET\">https://github.com/stevenwudi/6DVNET</a>), however I haven't anything yet about the related ApolloScape repo (<a href=\"https://github.com/stevenwudi/ApolloScape_InstanceSeg\">https://github.com/stevenwudi/ApolloScape_InstanceSeg</a>). I'm just going to throw a few questions out in hopes that someone (or maybe even <a href=\"/stevenwudi\">@stevenwudi</a> himself) has some answers.</p>\n\n<p>1) Has anyone produced good results using this? My LB scores using this are 0.003-0.004. I've even tried overfitting to the training set an things still never look that great when plotted. I've attached a few images of results on training set images.\n2) What is the expected order of inputs? According to the Data page the data we are provided is (yaw, pitch, roll, x, y, z). Can this be fed to the model in this same order or is it expecting some other order? I've tried changing the orders of yaw, pitch, and roll (I'm assuming x,y,z are in the correct place and correct order), but my results don't seem to change a whole lot.</p>",
      "rawMarkdown": "I've seen a few posts from people who have produced some bad results using 6DVNET  (https://github.com/stevenwudi/6DVNET), however I haven't anything yet about the related ApolloScape repo (https://github.com/stevenwudi/ApolloScape_InstanceSeg). I'm just going to throw a few questions out in hopes that someone (or maybe even @stevenwudi himself) has some answers.\n\n1) Has anyone produced good results using this? My LB scores using this are 0.003-0.004. I've even tried overfitting to the training set an things still never look that great when plotted. I've attached a few images of results on training set images.\n2) What is the expected order of inputs? According to the Data page the data we are provided is (yaw, pitch, roll, x, y, z). Can this be fed to the model in this same order or is it expecting some other order? I've tried changing the orders of yaw, pitch, and roll (I'm assuming x,y,z are in the correct place and correct order), but my results don't seem to change a whole lot.",
      "votes": null
    },
    {
      "id": "703912",
      "postDate": "12/26/2019 19:36:51",
      "content": "<p>1) I got 0.018 using config <code>e2e_3d_car_101_FPN_triple_head_non_local_weighted_homoscedastic</code>,\nset \n<code>TRAIN.SCALES=(800,1200)</code>\nand \n<code>SCORE_THRESH_FOR_TRUTH_DETECTION=0.9</code>\n2) for me this transformation\n <code>if abs(pose[1]) &amp;gt; math.pi / 2: pose[0] = -pose[0]</code>\nworks, you don't need to change input angle order, \nthe cars are still not perfectly aligned, and I've not solved that.</p>\n\n<p>P.S. Now I see the projected 3d-bboxes are not aligned with predicted 2d-bboxes, \nthere is <code>LOSS_3D_2D</code> there which may help...</p>",
      "rawMarkdown": "1) I got 0.018 using config `e2e_3d_car_101_FPN_triple_head_non_local_weighted_homoscedastic`,\nset \n`TRAIN.SCALES=(800,1200)`\nand \n`SCORE_THRESH_FOR_TRUTH_DETECTION=0.9`\n2) for me this transformation\n `if abs(pose[1]) &gt; math.pi / 2: pose[0] = -pose[0]`\nworks, you don't need to change input angle order, \nthe cars are still not perfectly aligned, and I've not solved that.\n\nP.S. Now I see the projected 3d-bboxes are not aligned with predicted 2d-bboxes, \nthere is `LOSS_3D_2D` there which may help...",
      "votes": null
    },
    {
      "id": "704047",
      "postDate": "12/27/2019 01:37:55",
      "content": "<p>Hey <a href=\"/brandenkmurray\">@brandenkmurray</a>  Thank you for your interests of my work. Concerning that I am currently also a participant of this challenge. I am afraid I cannot provide the code for this challenge. But I can point out some pits and falls:\n(1) Eular angle is very tricky for this challenge, the order is not defined in an orthodox way: the right way to generate mesh according to the given euler angle (EA) is (-EA[1], -EA[0], -EA2[2])。 More over, double check the code for transform it into quaternion because we change the axis -- I believe my original code base script should be altered...\n(2) x, y, z are in the correct place (it's simply a regression in that implementation)</p>\n\n<p>Good luck (to both of us)</p>",
      "rawMarkdown": "Hey @brandenkmurray  Thank you for your interests of my work. Concerning that I am currently also a participant of this challenge. I am afraid I cannot provide the code for this challenge. But I can point out some pits and falls:\n(1) Eular angle is very tricky for this challenge, the order is not defined in an orthodox way: the right way to generate mesh according to the given euler angle (EA) is (-EA[1], -EA[0], -EA2[2])。 More over, double check the code for transform it into quaternion because we change the axis -- I believe my original code base script should be altered...\n(2) x, y, z are in the correct place (it's simply a regression in that implementation)\n\nGood luck (to both of us)",
      "votes": null
    },
    {
      "id": "704286",
      "postDate": "12/27/2019 09:15:54",
      "content": "<p>Hi Steven:</p>\n\n<p>If you don't mind, could you explain why you set the q[3] = 0 at the last if-condition in quaternion-upper-hemispher function(in your 6D-VNet github  repo)?</p>\n\n<p>Since in your article, you said when a, b and c equal to 0, then only d equal to 1 is allowed. But in the repo, you set the d to 0 when a, b and c equal to 0.</p>",
      "rawMarkdown": "Hi Steven:\n\n\nIf you don't mind, could you explain why you set the q[3] = 0 at the last if-condition in quaternion-upper-hemispher function(in your 6D-VNet github  repo)?\n\nSince in your article, you said when a, b and c equal to 0, then only d equal to 1 is allowed. But in the repo, you set the d to 0 when a, b and c equal to 0.",
      "votes": null
    },
    {
      "id": "704424",
      "postDate": "12/27/2019 13:14:12",
      "content": "<p><a href=\"/stevenwudi\">@stevenwudi</a> \nthanks for your input few q.\n1) i see that while generating the labels people generate 2d coordinates using x,y,z and simply assiging the 6d to the coordinates ,my ask how we should take into account the Y,P,R details here </p>\n\n<p>2) which loss can help overcome the rotational errors. \n3) in popular posenet it seems it calculates loss for quaternion. Is there a way to convert the y,p,r into quarternions </p>",
      "rawMarkdown": "stevenwudi \nthanks for your input few q.\n1) i see that while generating the labels people generate 2d coordinates using x,y,z and simply assiging the 6d to the coordinates ,my ask how we should take into account the Y,P,R details here \n\n2) which loss can help overcome the rotational errors. \n3) in popular posenet it seems it calculates loss for quaternion. Is there a way to convert the y,p,r into quarternions",
      "votes": null
    },
    {
      "id": "704431",
      "postDate": "12/27/2019 13:19:34",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> I believe answers of all your questions can be found in Steven's paper.</p>\n\n<p><a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf\">http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf</a></p>",
      "rawMarkdown": "jaideepvalani I believe answers of all your questions can be found in Steven's paper.\n\nhttp://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf",
      "votes": null
    },
    {
      "id": "705617",
      "postDate": "12/29/2019 06:49:34",
      "content": "<p>Thanks <a href=\"/stevenwudi\">@stevenwudi</a>, I appreciate your response. I'll dig into these after the new year.</p>",
      "rawMarkdown": "Thanks @stevenwudi, I appreciate your response. I'll dig into these after the new year.",
      "votes": null
    },
    {
      "id": "705645",
      "postDate": "12/29/2019 08:20:59",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  You are correct： when a=b=c=0, d=1, it was a typo in my script... \nBut in fact, this never happens in the dataset, so it does not influence the output. But thank you for pointing my error!</p>",
      "rawMarkdown": "xiejialun  You are correct： when a=b=c=0, d=1, it was a typo in my script... \nBut in fact, this never happens in the dataset, so it does not influence the output. But thank you for pointing my error!",
      "votes": null
    },
    {
      "id": "705746",
      "postDate": "12/29/2019 12:05:41",
      "content": "<p>It's okay! Thanks a lot for the great work. Good luck in the competition!</p>",
      "rawMarkdown": "It's okay! Thanks a lot for the great work. Good luck in the competition!",
      "votes": null
    },
    {
      "id": "705779",
      "postDate": "12/29/2019 13:15:09",
      "content": "<p><a href=\"/stevenwudi\">@stevenwudi</a> \n<a href=\"/xiejialun\">@xiejialun</a> \n are you regressing X,y,z through a separate loss function than y,p,r ?\nis there a need for any rotational loss function. If yes would appreciateif you point to one or the approach for same. </p>\n\n<p>For the reason cited here\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/123284\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/123284</a>  </p>",
      "rawMarkdown": "stevenwudi \n@xiejialun \n are you regressing X,y,z through a separate loss function than y,p,r ?\nis there a need for any rotational loss function. If yes would appreciateif you point to one or the approach for same. \n\nFor the reason cited here\nhttps://www.kaggle.com/c/pku-autonomous-driving/discussion/123284",
      "votes": null
    },
    {
      "id": "706140",
      "postDate": "12/30/2019 00:55:58",
      "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \nIt might not be the right answer, but I would answer as much as I can,\nI think the rotation loss is for yaw, pitch and row(in quaternion form) , and of course you can separate these terms with specific loss, which you can observe in Steven's paper. For X, Y and Z, I don't know whether there is a good way to encode them for regression. So I am not able to answer that. </p>\n\n<p>Maybe you should take some time to read and understand the papers and useful articles. It would be very helpful.</p>",
      "rawMarkdown": "jaideepvalani \nIt might not be the right answer, but I would answer as much as I can,\nI think the rotation loss is for yaw, pitch and row(in quaternion form) , and of course you can separate these terms with specific loss, which you can observe in Steven's paper. For X, Y and Z, I don't know whether there is a good way to encode them for regression. So I am not able to answer that. \n\nMaybe you should take some time to read and understand the papers and useful articles. It would be very helpful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 703912,
      "author_name": "kruntuid",
      "author_url": "",
      "post_date": "12/26/2019 19:36:51",
      "content": "<p>1) I got 0.018 using config <code>e2e_3d_car_101_FPN_triple_head_non_local_weighted_homoscedastic</code>,\nset \n<code>TRAIN.SCALES=(800,1200)</code>\nand \n<code>SCORE_THRESH_FOR_TRUTH_DETECTION=0.9</code>\n2) for me this transformation\n <code>if abs(pose[1]) &amp;gt; math.pi / 2: pose[0] = -pose[0]</code>\nworks, you don't need to change input angle order, \nthe cars are still not perfectly aligned, and I've not solved that.</p>\n\n<p>P.S. Now I see the projected 3d-bboxes are not aligned with predicted 2d-bboxes, \nthere is <code>LOSS_3D_2D</code> there which may help...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 704047,
      "author_name": "stevenwudi",
      "author_url": "",
      "post_date": "12/27/2019 01:37:55",
      "content": "<p>Hey <a href=\"/brandenkmurray\">@brandenkmurray</a>  Thank you for your interests of my work. Concerning that I am currently also a participant of this challenge. I am afraid I cannot provide the code for this challenge. But I can point out some pits and falls:\n(1) Eular angle is very tricky for this challenge, the order is not defined in an orthodox way: the right way to generate mesh according to the given euler angle (EA) is (-EA[1], -EA[0], -EA2[2])。 More over, double check the code for transform it into quaternion because we change the axis -- I believe my original code base script should be altered...\n(2) x, y, z are in the correct place (it's simply a regression in that implementation)</p>\n\n<p>Good luck (to both of us)</p>",
      "votes": null,
      "replies": [
        {
          "id": 704286,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/27/2019 09:15:54",
          "content": "<p>Hi Steven:</p>\n\n<p>If you don't mind, could you explain why you set the q[3] = 0 at the last if-condition in quaternion-upper-hemispher function(in your 6D-VNet github  repo)?</p>\n\n<p>Since in your article, you said when a, b and c equal to 0, then only d equal to 1 is allowed. But in the repo, you set the d to 0 when a, b and c equal to 0.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704424,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/27/2019 13:14:12",
          "content": "<p><a href=\"/stevenwudi\">@stevenwudi</a> \nthanks for your input few q.\n1) i see that while generating the labels people generate 2d coordinates using x,y,z and simply assiging the 6d to the coordinates ,my ask how we should take into account the Y,P,R details here </p>\n\n<p>2) which loss can help overcome the rotational errors. \n3) in popular posenet it seems it calculates loss for quaternion. Is there a way to convert the y,p,r into quarternions </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 704431,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/27/2019 13:19:34",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> I believe answers of all your questions can be found in Steven's paper.</p>\n\n<p><a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf\">http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705617,
          "author_name": "brandenkmurray",
          "author_url": "",
          "post_date": "12/29/2019 06:49:34",
          "content": "<p>Thanks <a href=\"/stevenwudi\">@stevenwudi</a>, I appreciate your response. I'll dig into these after the new year.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 705779,
          "author_name": "jaideepvalani",
          "author_url": "",
          "post_date": "12/29/2019 13:15:09",
          "content": "<p><a href=\"/stevenwudi\">@stevenwudi</a> \n<a href=\"/xiejialun\">@xiejialun</a> \n are you regressing X,y,z through a separate loss function than y,p,r ?\nis there a need for any rotational loss function. If yes would appreciateif you point to one or the approach for same. </p>\n\n<p>For the reason cited here\n<a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/123284\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/123284</a>  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706140,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/30/2019 00:55:58",
          "content": "<p><a href=\"/jaideepvalani\">@jaideepvalani</a> \nIt might not be the right answer, but I would answer as much as I can,\nI think the rotation loss is for yaw, pitch and row(in quaternion form) , and of course you can separate these terms with specific loss, which you can observe in Steven's paper. For X, Y and Z, I don't know whether there is a good way to encode them for regression. So I am not able to answer that. </p>\n\n<p>Maybe you should take some time to read and understand the papers and useful articles. It would be very helpful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 705645,
      "author_name": "stevenwudi",
      "author_url": "",
      "post_date": "12/29/2019 08:20:59",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a>  You are correct： when a=b=c=0, d=1, it was a typo in my script... \nBut in fact, this never happens in the dataset, so it does not influence the output. But thank you for pointing my error!</p>",
      "votes": null,
      "replies": [
        {
          "id": 705746,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "12/29/2019 12:05:41",
          "content": "<p>It's okay! Thanks a lot for the great work. Good luck in the competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "703471": "I've seen a few posts from people who have produced some bad results using 6DVNET  (https://github.com/stevenwudi/6DVNET), however I haven't anything yet about the related ApolloScape repo (https://github.com/stevenwudi/ApolloScape_InstanceSeg). I'm just going to throw a few questions out in hopes that someone (or maybe even @stevenwudi himself) has some answers.\n\n1) Has anyone produced good results using this? My LB scores using this are 0.003-0.004. I've even tried overfitting to the training set an things still never look that great when plotted. I've attached a few images of results on training set images.\n2) What is the expected order of inputs? According to the Data page the data we are provided is (yaw, pitch, roll, x, y, z). Can this be fed to the model in this same order or is it expecting some other order? I've tried changing the orders of yaw, pitch, and roll (I'm assuming x,y,z are in the correct place and correct order), but my results don't seem to change a whole lot.",
    "703912": "1) I got 0.018 using config `e2e_3d_car_101_FPN_triple_head_non_local_weighted_homoscedastic`,\nset \n`TRAIN.SCALES=(800,1200)`\nand \n`SCORE_THRESH_FOR_TRUTH_DETECTION=0.9`\n2) for me this transformation\n `if abs(pose[1]) &gt; math.pi / 2: pose[0] = -pose[0]`\nworks, you don't need to change input angle order, \nthe cars are still not perfectly aligned, and I've not solved that.\n\nP.S. Now I see the projected 3d-bboxes are not aligned with predicted 2d-bboxes, \nthere is `LOSS_3D_2D` there which may help...",
    "704047": "Hey @brandenkmurray  Thank you for your interests of my work. Concerning that I am currently also a participant of this challenge. I am afraid I cannot provide the code for this challenge. But I can point out some pits and falls:\n(1) Eular angle is very tricky for this challenge, the order is not defined in an orthodox way: the right way to generate mesh according to the given euler angle (EA) is (-EA[1], -EA[0], -EA2[2])。 More over, double check the code for transform it into quaternion because we change the axis -- I believe my original code base script should be altered...\n(2) x, y, z are in the correct place (it's simply a regression in that implementation)\n\nGood luck (to both of us)",
    "704286": "Hi Steven:\n\n\nIf you don't mind, could you explain why you set the q[3] = 0 at the last if-condition in quaternion-upper-hemispher function(in your 6D-VNet github  repo)?\n\nSince in your article, you said when a, b and c equal to 0, then only d equal to 1 is allowed. But in the repo, you set the d to 0 when a, b and c equal to 0.",
    "704424": "stevenwudi \nthanks for your input few q.\n1) i see that while generating the labels people generate 2d coordinates using x,y,z and simply assiging the 6d to the coordinates ,my ask how we should take into account the Y,P,R details here \n\n2) which loss can help overcome the rotational errors. \n3) in popular posenet it seems it calculates loss for quaternion. Is there a way to convert the y,p,r into quarternions",
    "704431": "jaideepvalani I believe answers of all your questions can be found in Steven's paper.\n\nhttp://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf",
    "705617": "Thanks @stevenwudi, I appreciate your response. I'll dig into these after the new year.",
    "705645": "xiejialun  You are correct： when a=b=c=0, d=1, it was a typo in my script... \nBut in fact, this never happens in the dataset, so it does not influence the output. But thank you for pointing my error!",
    "705746": "It's okay! Thanks a lot for the great work. Good luck in the competition!",
    "705779": "stevenwudi \n@xiejialun \n are you regressing X,y,z through a separate loss function than y,p,r ?\nis there a need for any rotational loss function. If yes would appreciateif you point to one or the approach for same. \n\nFor the reason cited here\nhttps://www.kaggle.com/c/pku-autonomous-driving/discussion/123284",
    "706140": "jaideepvalani \nIt might not be the right answer, but I would answer as much as I can,\nI think the rotation loss is for yaw, pitch and row(in quaternion form) , and of course you can separate these terms with specific loss, which you can observe in Steven's paper. For X, Y and Z, I don't know whether there is a good way to encode them for regression. So I am not able to answer that. \n\nMaybe you should take some time to read and understand the papers and useful articles. It would be very helpful."
  },
  "source": "meta"
}