{
  "id": 119173,
  "title": "How does models such as Centernet predict poses?",
  "url": "/competitions/pku-autonomous-driving/discussion/119173",
  "author_name": "",
  "post_date": "2019-11-27T01:56:55.377935300Z",
  "votes": 2,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Hi!</p>\n\n<p>I am confused about how do models like Centernet predict car poses information since they are object detection models. Thanks!</p>",
  "messages": [
    {
      "id": "682142",
      "postDate": "11/27/2019 01:56:55",
      "content": "<p>Hi!</p>\n\n<p>I am confused about how do models like Centernet predict car poses information since they are object detection models. Thanks!</p>",
      "rawMarkdown": "Hi!\n\nI am confused about how do models like Centernet predict car poses information since they are object detection models. Thanks!",
      "votes": null
    },
    {
      "id": "682686",
      "postDate": "11/27/2019 18:47:27",
      "content": "<p>For each pixel you are regressing to a yaw,pitch,roll,x,y, and z value. The detection part is used to mask those values to actual cars in the picture. \nAfter that, you may do some fusion using the predicted x,y img coords and the regressed x,y,z coords.</p>",
      "rawMarkdown": "For each pixel you are regressing to a yaw,pitch,roll,x,y, and z value. The detection part is used to mask those values to actual cars in the picture. \nAfter that, you may do some fusion using the predicted x,y img coords and the regressed x,y,z coords.",
      "votes": null
    },
    {
      "id": "683007",
      "postDate": "11/28/2019 02:24:12",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "688890",
      "postDate": "12/06/2019 08:10:24",
      "content": "<p>Hi, I have some confusions, could you please give me some suggestions? @ llu000</p>\n\n<p><strong>Orientation is a single scalar by default. However, it can\nbe hard to regress to. We follow Mousavian et al. [38] and\nrepresent the orientation as two bins with in-bin regression.\nSpecifically, the orientation is encoded using 8 scalars, with\n4 scalars for each bin. For one bin, two scalars are used\nfor softmax classification and the rest two scalar regress to\nan angle within each bin.</strong></p>",
      "rawMarkdown": "Hi, I have some confusions, could you please give me some suggestions? @ llu000\n\n**Orientation is a single scalar by default. However, it can\nbe hard to regress to. We follow Mousavian et al. [38] and\nrepresent the orientation as two bins with in-bin regression.\nSpecifically, the orientation is encoded using 8 scalars, with\n4 scalars for each bin. For one bin, two scalars are used\nfor softmax classification and the rest two scalar regress to\nan angle within each bin.**",
      "votes": null
    },
    {
      "id": "689615",
      "postDate": "12/07/2019 06:59:37",
      "content": "<p>Hi, could you explain a little about how to do the fusion? Like adding a neural network that take predicted mask and 6d coordinates as input?</p>",
      "rawMarkdown": "Hi, could you explain a little about how to do the fusion? Like adding a neural network that take predicted mask and 6d coordinates as input?",
      "votes": null
    },
    {
      "id": "693125",
      "postDate": "12/12/2019 04:03:38",
      "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> <br>\nI haven't read centernet's paper carefully, but based on this passage, I think it's similar to pointrcnn: <br>\nIn pointrcnn, for Orientation, 2π is divided into several bin, the Angle is determined by classification, and the offset of the Angle is also predicted. <br>\nIt's just a reference.</p>",
      "rawMarkdown": "diegojohnson   \nI haven't read centernet's paper carefully, but based on this passage, I think it's similar to pointrcnn:  \nIn pointrcnn, for Orientation, 2π is divided into several bin, the Angle is determined by classification, and the offset of the Angle is also predicted.  \nIt's just a reference.",
      "votes": null
    },
    {
      "id": "693129",
      "postDate": "12/12/2019 04:12:19",
      "content": "<p>Good job, thanks👍 <a href=\"/sdeagggg\">@sdeagggg</a> </p>",
      "rawMarkdown": "Good job, thanks👍 @sdeagggg",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 682686,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/27/2019 18:47:27",
      "content": "<p>For each pixel you are regressing to a yaw,pitch,roll,x,y, and z value. The detection part is used to mask those values to actual cars in the picture. \nAfter that, you may do some fusion using the predicted x,y img coords and the regressed x,y,z coords.</p>",
      "votes": null,
      "replies": [
        {
          "id": 683007,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "11/28/2019 02:24:12",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 688890,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/06/2019 08:10:24",
          "content": "<p>Hi, I have some confusions, could you please give me some suggestions? @ llu000</p>\n\n<p><strong>Orientation is a single scalar by default. However, it can\nbe hard to regress to. We follow Mousavian et al. [38] and\nrepresent the orientation as two bins with in-bin regression.\nSpecifically, the orientation is encoded using 8 scalars, with\n4 scalars for each bin. For one bin, two scalars are used\nfor softmax classification and the rest two scalar regress to\nan angle within each bin.</strong></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 689615,
          "author_name": "yihaoshao",
          "author_url": "",
          "post_date": "12/07/2019 06:59:37",
          "content": "<p>Hi, could you explain a little about how to do the fusion? Like adding a neural network that take predicted mask and 6d coordinates as input?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 693125,
          "author_name": "sdeagggg",
          "author_url": "",
          "post_date": "12/12/2019 04:03:38",
          "content": "<p><a href=\"/diegojohnson\">@diegojohnson</a> <br>\nI haven't read centernet's paper carefully, but based on this passage, I think it's similar to pointrcnn: <br>\nIn pointrcnn, for Orientation, 2π is divided into several bin, the Angle is determined by classification, and the offset of the Angle is also predicted. <br>\nIt's just a reference.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 693129,
          "author_name": "diegojohnson",
          "author_url": "",
          "post_date": "12/12/2019 04:12:19",
          "content": "<p>Good job, thanks👍 <a href=\"/sdeagggg\">@sdeagggg</a> </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "682142": "Hi!\n\nI am confused about how do models like Centernet predict car poses information since they are object detection models. Thanks!",
    "682686": "For each pixel you are regressing to a yaw,pitch,roll,x,y, and z value. The detection part is used to mask those values to actual cars in the picture. \nAfter that, you may do some fusion using the predicted x,y img coords and the regressed x,y,z coords.",
    "683007": "Thanks!",
    "688890": "Hi, I have some confusions, could you please give me some suggestions? @ llu000\n\n**Orientation is a single scalar by default. However, it can\nbe hard to regress to. We follow Mousavian et al. [38] and\nrepresent the orientation as two bins with in-bin regression.\nSpecifically, the orientation is encoded using 8 scalars, with\n4 scalars for each bin. For one bin, two scalars are used\nfor softmax classification and the rest two scalar regress to\nan angle within each bin.**",
    "689615": "Hi, could you explain a little about how to do the fusion? Like adding a neural network that take predicted mask and 6d coordinates as input?",
    "693125": "diegojohnson   \nI haven't read centernet's paper carefully, but based on this passage, I think it's similar to pointrcnn:  \nIn pointrcnn, for Orientation, 2π is divided into several bin, the Angle is determined by classification, and the offset of the Angle is also predicted.  \nIt's just a reference.",
    "693129": "Good job, thanks👍 @sdeagggg"
  },
  "source": "meta"
}