{
  "id": 118304,
  "title": "OFT Net - Paper + PyTorch Code",
  "url": "/competitions/pku-autonomous-driving/discussion/118304",
  "author_name": "",
  "post_date": "2019-11-20T16:01:49.183184900Z",
  "votes": 26,
  "comment_count": 9,
  "views": 0,
  "content": "<h3>Orthographic Feature Transform for Monocular 3D Object Detection</h3>\n\n<p>Paper: <a href=\"https://arxiv.org/pdf/1811.08188.pdf\">https://arxiv.org/pdf/1811.08188.pdf</a>\nPyTorch Code (kitti support): <a href=\"https://github.com/tom-roddick/oft\">https://github.com/tom-roddick/oft</a></p>\n\n<p>Currently working on adapting the above code to work with the competition dataset, will leave an update here if I have any success. </p>",
  "messages": [
    {
      "id": "677810",
      "postDate": "11/20/2019 16:01:49",
      "content": "<h3>Orthographic Feature Transform for Monocular 3D Object Detection</h3>\n\n<p>Paper: <a href=\"https://arxiv.org/pdf/1811.08188.pdf\">https://arxiv.org/pdf/1811.08188.pdf</a>\nPyTorch Code (kitti support): <a href=\"https://github.com/tom-roddick/oft\">https://github.com/tom-roddick/oft</a></p>\n\n<p>Currently working on adapting the above code to work with the competition dataset, will leave an update here if I have any success. </p>",
      "rawMarkdown": "### Orthographic Feature Transform for Monocular 3D Object Detection\n\nPaper: https://arxiv.org/pdf/1811.08188.pdf\nPyTorch Code (kitti support): https://github.com/tom-roddick/oft\n\nCurrently working on adapting the above code to work with the competition dataset, will leave an update here if I have any success.",
      "votes": null
    },
    {
      "id": "677834",
      "postDate": "11/20/2019 16:32:24",
      "content": "<p>Great\nThanks for Sharing!! <a href=\"/jackvial\">@jackvial</a> </p>",
      "rawMarkdown": "Great\nThanks for Sharing!! @jackvial",
      "votes": null
    },
    {
      "id": "677864",
      "postDate": "11/20/2019 17:18:36",
      "content": "<p><a href=\"/jackvial\">@jackvial</a>  Excellent coincidence . I was about to post the same . I have adopted it partially except for the orthographic projection part . I will post the Kernel in couple of hours , once it finishes training . It's just a base version with some mistake but could be made to work . You too please let us know .</p>",
      "rawMarkdown": "jackvial  Excellent coincidence . I was about to post the same . I have adopted it partially except for the orthographic projection part . I will post the Kernel in couple of hours , once it finishes training . It's just a base version with some mistake but could be made to work . You too please let us know .",
      "votes": null
    },
    {
      "id": "677869",
      "postDate": "11/20/2019 17:25:40",
      "content": "<p>Looking forward to seeing your Kernel. What is your reason for removing the orthographic step?</p>",
      "rawMarkdown": "Looking forward to seeing your Kernel. What is your reason for removing the orthographic step?",
      "votes": null
    },
    {
      "id": "677870",
      "postDate": "11/20/2019 17:29:45",
      "content": "<p>I am very new to this , I was looking more of adopting some Resnet backbone with the Centernet kernel . I didn't understand what is Orthographic projection yet . So I have planned to use only the backbone with Centernet public kernel and then try the OFTNET altogether . </p>\n\n<p>Another reason of not chosing OFTNET altogether is because I saw it had quite poor result in KITTI car dataset as published on their site.  </p>\n\n<p>As a beginner I am being conservative that's all . </p>\n\n<p>One question , any reason why Seconds will not work here ? You had made it customized and made it work for Lyft correct ? </p>",
      "rawMarkdown": "I am very new to this , I was looking more of adopting some Resnet backbone with the Centernet kernel . I didn't understand what is Orthographic projection yet . So I have planned to use only the backbone with Centernet public kernel and then try the OFTNET altogether . \n\nAnother reason of not chosing OFTNET altogether is because I saw it had quite poor result in KITTI car dataset as published on their site.  \n\nAs a beginner I am being conservative that's all . \n\nOne question , any reason why Seconds will not work here ? You had made it customized and made it work for Lyft correct ?",
      "votes": null
    },
    {
      "id": "677878",
      "postDate": "11/20/2019 17:44:21",
      "content": "<p>Resent backbone + CenterNet sounds like a good approach too. Orthographic projection in the case of OFT net basically means the 2d images are projected into 3d space to help overcome the depth perception issue. </p>\n\n<p>3d Object detection from only 2d images is a more challenging tasks so that might explain the low score on kitti compared to other methods using multiple cameras + lidar but I will have to review it.</p>\n\n<p>SECOND uses only lidar data so can't be used for this competition unfortunately :(</p>",
      "rawMarkdown": "Resent backbone + CenterNet sounds like a good approach too. Orthographic projection in the case of OFT net basically means the 2d images are projected into 3d space to help overcome the depth perception issue. \n\n3d Object detection from only 2d images is a more challenging tasks so that might explain the low score on kitti compared to other methods using multiple cameras + lidar but I will have to review it.\n\nSECOND uses only lidar data so can't be used for this competition unfortunately :(",
      "votes": null
    },
    {
      "id": "678114",
      "postDate": "11/21/2019 02:06:48",
      "content": "<p><a href=\"/jackvial\">@jackvial</a> <br>\nHere is my Kernel.</p>\n\n<p><a href=\"https://www.kaggle.com/phoenix9032/center-resnet-trial\">https://www.kaggle.com/phoenix9032/center-resnet-trial</a></p>",
      "rawMarkdown": "jackvial  \nHere is my Kernel.\n\nhttps://www.kaggle.com/phoenix9032/center-resnet-trial",
      "votes": null
    },
    {
      "id": "678292",
      "postDate": "11/21/2019 07:41:31",
      "content": "<p>Thanks for sharing this!!! this is very helpful....</p>",
      "rawMarkdown": "Thanks for sharing this!!! this is very helpful....",
      "votes": null
    },
    {
      "id": "678473",
      "postDate": "11/21/2019 12:49:54",
      "content": "<p>Welcome, please upvote if you find it useful.</p>",
      "rawMarkdown": "Welcome, please upvote if you find it useful.",
      "votes": null
    },
    {
      "id": "678474",
      "postDate": "11/21/2019 12:50:03",
      "content": "<p>Welcome, please upvote if you find it useful.</p>",
      "rawMarkdown": "Welcome, please upvote if you find it useful.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 677834,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "11/20/2019 16:32:24",
      "content": "<p>Great\nThanks for Sharing!! <a href=\"/jackvial\">@jackvial</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 678474,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "11/21/2019 12:50:03",
          "content": "<p>Welcome, please upvote if you find it useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 677864,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/20/2019 17:18:36",
      "content": "<p><a href=\"/jackvial\">@jackvial</a>  Excellent coincidence . I was about to post the same . I have adopted it partially except for the orthographic projection part . I will post the Kernel in couple of hours , once it finishes training . It's just a base version with some mistake but could be made to work . You too please let us know .</p>",
      "votes": null,
      "replies": [
        {
          "id": 677869,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "11/20/2019 17:25:40",
          "content": "<p>Looking forward to seeing your Kernel. What is your reason for removing the orthographic step?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677870,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/20/2019 17:29:45",
          "content": "<p>I am very new to this , I was looking more of adopting some Resnet backbone with the Centernet kernel . I didn't understand what is Orthographic projection yet . So I have planned to use only the backbone with Centernet public kernel and then try the OFTNET altogether . </p>\n\n<p>Another reason of not chosing OFTNET altogether is because I saw it had quite poor result in KITTI car dataset as published on their site.  </p>\n\n<p>As a beginner I am being conservative that's all . </p>\n\n<p>One question , any reason why Seconds will not work here ? You had made it customized and made it work for Lyft correct ? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 677878,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "11/20/2019 17:44:21",
          "content": "<p>Resent backbone + CenterNet sounds like a good approach too. Orthographic projection in the case of OFT net basically means the 2d images are projected into 3d space to help overcome the depth perception issue. </p>\n\n<p>3d Object detection from only 2d images is a more challenging tasks so that might explain the low score on kitti compared to other methods using multiple cameras + lidar but I will have to review it.</p>\n\n<p>SECOND uses only lidar data so can't be used for this competition unfortunately :(</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 678114,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/21/2019 02:06:48",
          "content": "<p><a href=\"/jackvial\">@jackvial</a> <br>\nHere is my Kernel.</p>\n\n<p><a href=\"https://www.kaggle.com/phoenix9032/center-resnet-trial\">https://www.kaggle.com/phoenix9032/center-resnet-trial</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 678292,
      "author_name": "nanditab35",
      "author_url": "",
      "post_date": "11/21/2019 07:41:31",
      "content": "<p>Thanks for sharing this!!! this is very helpful....</p>",
      "votes": null,
      "replies": [
        {
          "id": 678473,
          "author_name": "jackvial",
          "author_url": "",
          "post_date": "11/21/2019 12:49:54",
          "content": "<p>Welcome, please upvote if you find it useful.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "677810": "### Orthographic Feature Transform for Monocular 3D Object Detection\n\nPaper: https://arxiv.org/pdf/1811.08188.pdf\nPyTorch Code (kitti support): https://github.com/tom-roddick/oft\n\nCurrently working on adapting the above code to work with the competition dataset, will leave an update here if I have any success.",
    "677834": "Great\nThanks for Sharing!! @jackvial",
    "677864": "jackvial  Excellent coincidence . I was about to post the same . I have adopted it partially except for the orthographic projection part . I will post the Kernel in couple of hours , once it finishes training . It's just a base version with some mistake but could be made to work . You too please let us know .",
    "677869": "Looking forward to seeing your Kernel. What is your reason for removing the orthographic step?",
    "677870": "I am very new to this , I was looking more of adopting some Resnet backbone with the Centernet kernel . I didn't understand what is Orthographic projection yet . So I have planned to use only the backbone with Centernet public kernel and then try the OFTNET altogether . \n\nAnother reason of not chosing OFTNET altogether is because I saw it had quite poor result in KITTI car dataset as published on their site.  \n\nAs a beginner I am being conservative that's all . \n\nOne question , any reason why Seconds will not work here ? You had made it customized and made it work for Lyft correct ?",
    "677878": "Resent backbone + CenterNet sounds like a good approach too. Orthographic projection in the case of OFT net basically means the 2d images are projected into 3d space to help overcome the depth perception issue. \n\n3d Object detection from only 2d images is a more challenging tasks so that might explain the low score on kitti compared to other methods using multiple cameras + lidar but I will have to review it.\n\nSECOND uses only lidar data so can't be used for this competition unfortunately :(",
    "678114": "jackvial  \nHere is my Kernel.\n\nhttps://www.kaggle.com/phoenix9032/center-resnet-trial",
    "678292": "Thanks for sharing this!!! this is very helpful....",
    "678473": "Welcome, please upvote if you find it useful.",
    "678474": "Welcome, please upvote if you find it useful."
  },
  "source": "meta"
}