{
  "id": 118163,
  "title": "How to extend Mask RCNN for this problem?",
  "url": "/competitions/pku-autonomous-driving/discussion/118163",
  "author_name": "",
  "post_date": "2019-11-19T20:04:32.387635900Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "",
  "messages": [
    {
      "id": "677108",
      "postDate": "11/19/2019 20:04:32",
      "content": "",
      "rawMarkdown": "",
      "votes": null
    },
    {
      "id": "677183",
      "postDate": "11/19/2019 22:26:47",
      "content": "<p>first step: \nstart trying yourself, we can only help if you have a codebase.</p>",
      "rawMarkdown": "first step: \nstart trying yourself, we can only help if you have a codebase.",
      "votes": null
    },
    {
      "id": "678359",
      "postDate": "11/21/2019 09:25:18",
      "content": "<p>Hello <a href=\"/abhayakumar\">@abhayakumar</a>  . What I understand from the public Kernel a bit is that :  Your CNN model is creating some area of interest. i.e it is understanding what is a car feature  in the whole images consisting of road , building , people , railing etc . Then those  intermediate features are extracted and the pattern and pose of those features are extrapolated to understand a car pose , distance etc . Based on training data . \nI have not worked on Mask RCNN  but what i understand is that Mask RCNN provides output as mask , class , 2D bounding boxes etc . If you can extract those mask related features out just like its done for efficientnet and resnet in public kernels , then you can use that information to find the pose co-ordinates . </p>\n\n<p>Something like that is already done in this paper \n<a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf\">http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf</a></p>",
      "rawMarkdown": "Hello @abhayakumar  . What I understand from the public Kernel a bit is that :  Your CNN model is creating some area of interest. i.e it is understanding what is a car feature  in the whole images consisting of road , building , people , railing etc . Then those  intermediate features are extracted and the pattern and pose of those features are extrapolated to understand a car pose , distance etc . Based on training data . \nI have not worked on Mask RCNN  but what i understand is that Mask RCNN provides output as mask , class , 2D bounding boxes etc . If you can extract those mask related features out just like its done for efficientnet and resnet in public kernels , then you can use that information to find the pose co-ordinates . \n\nSomething like that is already done in this paper \nhttp://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 677183,
      "author_name": "ilu000",
      "author_url": "",
      "post_date": "11/19/2019 22:26:47",
      "content": "<p>first step: \nstart trying yourself, we can only help if you have a codebase.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 678359,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/21/2019 09:25:18",
      "content": "<p>Hello <a href=\"/abhayakumar\">@abhayakumar</a>  . What I understand from the public Kernel a bit is that :  Your CNN model is creating some area of interest. i.e it is understanding what is a car feature  in the whole images consisting of road , building , people , railing etc . Then those  intermediate features are extracted and the pattern and pose of those features are extrapolated to understand a car pose , distance etc . Based on training data . \nI have not worked on Mask RCNN  but what i understand is that Mask RCNN provides output as mask , class , 2D bounding boxes etc . If you can extract those mask related features out just like its done for efficientnet and resnet in public kernels , then you can use that information to find the pose co-ordinates . </p>\n\n<p>Something like that is already done in this paper \n<a href=\"http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf\">http://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf</a></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "677108": "",
    "677183": "first step: \nstart trying yourself, we can only help if you have a codebase.",
    "678359": "Hello @abhayakumar  . What I understand from the public Kernel a bit is that :  Your CNN model is creating some area of interest. i.e it is understanding what is a car feature  in the whole images consisting of road , building , people , railing etc . Then those  intermediate features are extracted and the pattern and pose of those features are extrapolated to understand a car pose , distance etc . Based on training data . \nI have not worked on Mask RCNN  but what i understand is that Mask RCNN provides output as mask , class , 2D bounding boxes etc . If you can extract those mask related features out just like its done for efficientnet and resnet in public kernels , then you can use that information to find the pose co-ordinates . \n\nSomething like that is already done in this paper \nhttp://openaccess.thecvf.com/content_CVPRW_2019/papers/Autonomous%20Driving/Wu_6D-VNet_End-to-End_6-DoF_Vehicle_Pose_Estimation_From_Monocular_RGB_Images_CVPRW_2019_paper.pdf"
  },
  "source": "meta"
}