{
  "id": 118440,
  "title": "Starting from Basic - Part -II - CAR POSE Model Explanation",
  "url": "/competitions/pku-autonomous-driving/discussion/118440",
  "author_name": "",
  "post_date": "2019-11-21T13:29:55.667118700Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>DISCLAIMER : I Wanted to continue under the same topic rather than flooding the forum . But , chose against it as  a) I dont know the word limit for a topic and b) this is a little different than the last topic .\nGENERAL DISCLAIMER : very novice in this area , so if I am wrong , please feel free to correct me . I will be very happy . \nHere are some more things that I found out during my study in past couple of days .  Here is a site which explains it very nicely , so I will quote from there . \n<a href=\"https://www.fritz.ai/pose-estimation/\">https://www.fritz.ai/pose-estimation/</a></p>\n\n<p>*<em>WHAT ARE  BOTTOM-UP NETWORK\n*</em>\n With a bottom-up approach, the model detects every instance of a particular keypoint in a given image and then attempts to assemble groups of keypoints into skeletons for distinct objects.</p>\n\n<p>Here is a  heatmap prediction from my Effiecientnet-Centernet model . Some of the car heatmap is marked </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fd4182d6c3bad8e36c50003f6bba07e20%2FHeatmap.png?generation=1574341719184422&amp;alt=media\" alt=\"\"></p>\n\n<p>Here are the postprocesing  on the heatmap to get the pose estimation .</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F595a05d0ccadef20bdedbc6b0c8a70cc%2Fboxes.png?generation=1574341765761785&amp;alt=media\" alt=\"\"></p>\n\n<p>This is a portion of the code is used to extract pose from the heatmap .</p>\n\n<p><code>\n    output = model(torch.tensor(img[None]).to(device)).data.cpu().numpy()\n    coords_pred = extract_coords(output[0])\n    coords_true = extract_coords(np.concatenate([mask[None], regr], 0))\n</code>\nHOW DOES IT WORK ?</p>\n\n<p>&gt; It simply starts with  an encoder that accepts an image as input and extracts features using a series of narrowing convolution blocks. What comes after the encoder depends on the method of pose estimation.</p>\n\n<p>The most conceptually simple method uses a regressor to output final predictions of each keypoint location. The resulting model accepts an image as input and outputs X, Y, and potentially Z coordinates for each keypoint you’re trying to predict.  In our case , it is X,Y,Z ,yaw,pitch, roll . </p>\n\n<p>A slightly more complicated approach uses an encoder-decoder architecture. Instead of estimating keypoint coordinates directly, the encoder is fed into a decoder, which creates heatmaps representing the likelihood that a keypoint is found in a given region of an image.</p>\n\n<p>During post-processing, the exact coordinates of a keypoint are found by selecting heatmap locations with the highest keypoint likelihood. In the case of multi-pose estimation, a heatmap may contain multiple areas of high keypoint likelihood (e.g. multiple right hands in an image). In these cases, additional post-processing is required to assign each area to a specific object instance.</p>\n\n<p>In our case  Efficientnetb0 is the encoder and then we have a small decoder which takes the features from efficientnetb0 and gives output of the heatmap as shown in the above picture .\nEncoder : \n<code>\n    def __init__(self, n_classes):\n        super(MyUNet, self).__init__()\n        self.base_model =EfficientNet.from_pretrained('efficientnet-b0')\n</code>\nDecoder:\n<code>\n        x = self.up1(feats, x4)\n        x = self.up2(x, x3)\n        x = self.outc(x)\n        return x\n</code>\nNow there are some literature and codebase to start with BOTTOM-UP Network  : </p>\n\n<p>Ofcourse below two public  Kernels : \n<a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>\n<a href=\"https://www.kaggle.com/phoenix9032/center-resnet-trial\">https://www.kaggle.com/phoenix9032/center-resnet-trial</a></p>\n\n<p>OFTNET : \n<a href=\"https://arxiv.org/abs/1811.08188\">https://arxiv.org/abs/1811.08188</a>\n<a href=\"https://github.com/tom-roddick/oft\">https://github.com/tom-roddick/oft</a></p>\n\n<p>CENTERNET : \n<a href=\"https://arxiv.org/abs/1904.07850\">https://arxiv.org/abs/1904.07850</a>\n<a href=\"https://github.com/xingyizhou/CenterNet\">https://github.com/xingyizhou/CenterNet</a></p>\n\n<p>PVNET:\n<a href=\"https://arxiv.org/pdf/1812.11788.pdf\">https://arxiv.org/pdf/1812.11788.pdf</a>\n<a href=\"https://github.com/zju3dv/pvnet\">https://github.com/zju3dv/pvnet</a></p>\n\n<p>Once I get familiar with Mask-RCNN and the corresponding pose detection frameworks e.g. 6D-VN , I will try to post a topic on TOP-DOWN NETWORK</p>",
  "messages": [
    {
      "id": "678505",
      "postDate": "11/21/2019 13:29:55",
      "content": "<p>DISCLAIMER : I Wanted to continue under the same topic rather than flooding the forum . But , chose against it as  a) I dont know the word limit for a topic and b) this is a little different than the last topic .\nGENERAL DISCLAIMER : very novice in this area , so if I am wrong , please feel free to correct me . I will be very happy . \nHere are some more things that I found out during my study in past couple of days .  Here is a site which explains it very nicely , so I will quote from there . \n<a href=\"https://www.fritz.ai/pose-estimation/\">https://www.fritz.ai/pose-estimation/</a></p>\n\n<p>*<em>WHAT ARE  BOTTOM-UP NETWORK\n*</em>\n With a bottom-up approach, the model detects every instance of a particular keypoint in a given image and then attempts to assemble groups of keypoints into skeletons for distinct objects.</p>\n\n<p>Here is a  heatmap prediction from my Effiecientnet-Centernet model . Some of the car heatmap is marked </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fd4182d6c3bad8e36c50003f6bba07e20%2FHeatmap.png?generation=1574341719184422&amp;alt=media\" alt=\"\"></p>\n\n<p>Here are the postprocesing  on the heatmap to get the pose estimation .</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F595a05d0ccadef20bdedbc6b0c8a70cc%2Fboxes.png?generation=1574341765761785&amp;alt=media\" alt=\"\"></p>\n\n<p>This is a portion of the code is used to extract pose from the heatmap .</p>\n\n<p><code>\n    output = model(torch.tensor(img[None]).to(device)).data.cpu().numpy()\n    coords_pred = extract_coords(output[0])\n    coords_true = extract_coords(np.concatenate([mask[None], regr], 0))\n</code>\nHOW DOES IT WORK ?</p>\n\n<p>&gt; It simply starts with  an encoder that accepts an image as input and extracts features using a series of narrowing convolution blocks. What comes after the encoder depends on the method of pose estimation.</p>\n\n<p>The most conceptually simple method uses a regressor to output final predictions of each keypoint location. The resulting model accepts an image as input and outputs X, Y, and potentially Z coordinates for each keypoint you’re trying to predict.  In our case , it is X,Y,Z ,yaw,pitch, roll . </p>\n\n<p>A slightly more complicated approach uses an encoder-decoder architecture. Instead of estimating keypoint coordinates directly, the encoder is fed into a decoder, which creates heatmaps representing the likelihood that a keypoint is found in a given region of an image.</p>\n\n<p>During post-processing, the exact coordinates of a keypoint are found by selecting heatmap locations with the highest keypoint likelihood. In the case of multi-pose estimation, a heatmap may contain multiple areas of high keypoint likelihood (e.g. multiple right hands in an image). In these cases, additional post-processing is required to assign each area to a specific object instance.</p>\n\n<p>In our case  Efficientnetb0 is the encoder and then we have a small decoder which takes the features from efficientnetb0 and gives output of the heatmap as shown in the above picture .\nEncoder : \n<code>\n    def __init__(self, n_classes):\n        super(MyUNet, self).__init__()\n        self.base_model =EfficientNet.from_pretrained('efficientnet-b0')\n</code>\nDecoder:\n<code>\n        x = self.up1(feats, x4)\n        x = self.up2(x, x3)\n        x = self.outc(x)\n        return x\n</code>\nNow there are some literature and codebase to start with BOTTOM-UP Network  : </p>\n\n<p>Ofcourse below two public  Kernels : \n<a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">https://www.kaggle.com/hocop1/centernet-baseline</a>\n<a href=\"https://www.kaggle.com/phoenix9032/center-resnet-trial\">https://www.kaggle.com/phoenix9032/center-resnet-trial</a></p>\n\n<p>OFTNET : \n<a href=\"https://arxiv.org/abs/1811.08188\">https://arxiv.org/abs/1811.08188</a>\n<a href=\"https://github.com/tom-roddick/oft\">https://github.com/tom-roddick/oft</a></p>\n\n<p>CENTERNET : \n<a href=\"https://arxiv.org/abs/1904.07850\">https://arxiv.org/abs/1904.07850</a>\n<a href=\"https://github.com/xingyizhou/CenterNet\">https://github.com/xingyizhou/CenterNet</a></p>\n\n<p>PVNET:\n<a href=\"https://arxiv.org/pdf/1812.11788.pdf\">https://arxiv.org/pdf/1812.11788.pdf</a>\n<a href=\"https://github.com/zju3dv/pvnet\">https://github.com/zju3dv/pvnet</a></p>\n\n<p>Once I get familiar with Mask-RCNN and the corresponding pose detection frameworks e.g. 6D-VN , I will try to post a topic on TOP-DOWN NETWORK</p>",
      "rawMarkdown": "DISCLAIMER : I Wanted to continue under the same topic rather than flooding the forum . But , chose against it as  a) I dont know the word limit for a topic and b) this is a little different than the last topic .\nGENERAL DISCLAIMER : very novice in this area , so if I am wrong , please feel free to correct me . I will be very happy . \nHere are some more things that I found out during my study in past couple of days .  Here is a site which explains it very nicely , so I will quote from there . \nhttps://www.fritz.ai/pose-estimation/\n\n**WHAT ARE  BOTTOM-UP NETWORK\n**\n With a bottom-up approach, the model detects every instance of a particular keypoint in a given image and then attempts to assemble groups of keypoints into skeletons for distinct objects.\n\nHere is a  heatmap prediction from my Effiecientnet-Centernet model . Some of the car heatmap is marked \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fd4182d6c3bad8e36c50003f6bba07e20%2FHeatmap.png?generation=1574341719184422&amp;alt=media)\n\nHere are the postprocesing  on the heatmap to get the pose estimation .\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F595a05d0ccadef20bdedbc6b0c8a70cc%2Fboxes.png?generation=1574341765761785&amp;alt=media)\n\nThis is a portion of the code is used to extract pose from the heatmap .\n\n```\n    output = model(torch.tensor(img[None]).to(device)).data.cpu().numpy()\n    coords_pred = extract_coords(output[0])\n    coords_true = extract_coords(np.concatenate([mask[None], regr], 0))\n```\nHOW DOES IT WORK ?\n\n&gt; It simply starts with  an encoder that accepts an image as input and extracts features using a series of narrowing convolution blocks. What comes after the encoder depends on the method of pose estimation.\n\nThe most conceptually simple method uses a regressor to output final predictions of each keypoint location. The resulting model accepts an image as input and outputs X, Y, and potentially Z coordinates for each keypoint you’re trying to predict.  In our case , it is X,Y,Z ,yaw,pitch, roll . \n\nA slightly more complicated approach uses an encoder-decoder architecture. Instead of estimating keypoint coordinates directly, the encoder is fed into a decoder, which creates heatmaps representing the likelihood that a keypoint is found in a given region of an image.\n\nDuring post-processing, the exact coordinates of a keypoint are found by selecting heatmap locations with the highest keypoint likelihood. In the case of multi-pose estimation, a heatmap may contain multiple areas of high keypoint likelihood (e.g. multiple right hands in an image). In these cases, additional post-processing is required to assign each area to a specific object instance.\n\nIn our case  Efficientnetb0 is the encoder and then we have a small decoder which takes the features from efficientnetb0 and gives output of the heatmap as shown in the above picture .\nEncoder : \n```\n    def __init__(self, n_classes):\n        super(MyUNet, self).__init__()\n        self.base_model =EfficientNet.from_pretrained('efficientnet-b0')\n```\nDecoder:\n```\n        x = self.up1(feats, x4)\n        x = self.up2(x, x3)\n        x = self.outc(x)\n        return x\n```\nNow there are some literature and codebase to start with BOTTOM-UP Network  : \n\nOfcourse below two public  Kernels : \nhttps://www.kaggle.com/hocop1/centernet-baseline\nhttps://www.kaggle.com/phoenix9032/center-resnet-trial\n\nOFTNET : \nhttps://arxiv.org/abs/1811.08188\nhttps://github.com/tom-roddick/oft\n\nCENTERNET : \nhttps://arxiv.org/abs/1904.07850\nhttps://github.com/xingyizhou/CenterNet\n\nPVNET:\nhttps://arxiv.org/pdf/1812.11788.pdf\nhttps://github.com/zju3dv/pvnet\n\nOnce I get familiar with Mask-RCNN and the corresponding pose detection frameworks e.g. 6D-VN , I will try to post a topic on TOP-DOWN NETWORK",
      "votes": null
    },
    {
      "id": "685946",
      "postDate": "12/02/2019 14:53:24",
      "content": "<p>good job!\nthis competition is really difficult for me😭 </p>",
      "rawMarkdown": "good job!\nthis competition is really difficult for me😭",
      "votes": null
    },
    {
      "id": "692419",
      "postDate": "12/11/2019 08:42:27",
      "content": "<p>Hi buddy, I have read PVNET, the PVNET method requires the segmentation of vehicles, do you have any method to get the segmentation information？Officials did not provide the Mask for cars of interest.☹️ </p>",
      "rawMarkdown": "Hi buddy, I have read PVNET, the PVNET method requires the segmentation of vehicles, do you have any method to get the segmentation information？Officials did not provide the Mask for cars of interest.☹️",
      "votes": null
    },
    {
      "id": "750790",
      "postDate": "02/19/2020 17:34:24",
      "content": "<p>nice</p>",
      "rawMarkdown": "nice",
      "votes": null
    },
    {
      "id": "752202",
      "postDate": "02/20/2020 19:39:18",
      "content": "<p>great</p>",
      "rawMarkdown": "great",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 685946,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/02/2019 14:53:24",
      "content": "<p>good job!\nthis competition is really difficult for me😭 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 692419,
      "author_name": "kiruto",
      "author_url": "",
      "post_date": "12/11/2019 08:42:27",
      "content": "<p>Hi buddy, I have read PVNET, the PVNET method requires the segmentation of vehicles, do you have any method to get the segmentation information？Officials did not provide the Mask for cars of interest.☹️ </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 750790,
      "author_name": "",
      "author_url": "",
      "post_date": "02/19/2020 17:34:24",
      "content": "<p>nice</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 752202,
      "author_name": "",
      "author_url": "",
      "post_date": "02/20/2020 19:39:18",
      "content": "<p>great</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "678505": "DISCLAIMER : I Wanted to continue under the same topic rather than flooding the forum . But , chose against it as  a) I dont know the word limit for a topic and b) this is a little different than the last topic .\nGENERAL DISCLAIMER : very novice in this area , so if I am wrong , please feel free to correct me . I will be very happy . \nHere are some more things that I found out during my study in past couple of days .  Here is a site which explains it very nicely , so I will quote from there . \nhttps://www.fritz.ai/pose-estimation/\n\n**WHAT ARE  BOTTOM-UP NETWORK\n**\n With a bottom-up approach, the model detects every instance of a particular keypoint in a given image and then attempts to assemble groups of keypoints into skeletons for distinct objects.\n\nHere is a  heatmap prediction from my Effiecientnet-Centernet model . Some of the car heatmap is marked \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2Fd4182d6c3bad8e36c50003f6bba07e20%2FHeatmap.png?generation=1574341719184422&amp;alt=media)\n\nHere are the postprocesing  on the heatmap to get the pose estimation .\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F595a05d0ccadef20bdedbc6b0c8a70cc%2Fboxes.png?generation=1574341765761785&amp;alt=media)\n\nThis is a portion of the code is used to extract pose from the heatmap .\n\n```\n    output = model(torch.tensor(img[None]).to(device)).data.cpu().numpy()\n    coords_pred = extract_coords(output[0])\n    coords_true = extract_coords(np.concatenate([mask[None], regr], 0))\n```\nHOW DOES IT WORK ?\n\n&gt; It simply starts with  an encoder that accepts an image as input and extracts features using a series of narrowing convolution blocks. What comes after the encoder depends on the method of pose estimation.\n\nThe most conceptually simple method uses a regressor to output final predictions of each keypoint location. The resulting model accepts an image as input and outputs X, Y, and potentially Z coordinates for each keypoint you’re trying to predict.  In our case , it is X,Y,Z ,yaw,pitch, roll . \n\nA slightly more complicated approach uses an encoder-decoder architecture. Instead of estimating keypoint coordinates directly, the encoder is fed into a decoder, which creates heatmaps representing the likelihood that a keypoint is found in a given region of an image.\n\nDuring post-processing, the exact coordinates of a keypoint are found by selecting heatmap locations with the highest keypoint likelihood. In the case of multi-pose estimation, a heatmap may contain multiple areas of high keypoint likelihood (e.g. multiple right hands in an image). In these cases, additional post-processing is required to assign each area to a specific object instance.\n\nIn our case  Efficientnetb0 is the encoder and then we have a small decoder which takes the features from efficientnetb0 and gives output of the heatmap as shown in the above picture .\nEncoder : \n```\n    def __init__(self, n_classes):\n        super(MyUNet, self).__init__()\n        self.base_model =EfficientNet.from_pretrained('efficientnet-b0')\n```\nDecoder:\n```\n        x = self.up1(feats, x4)\n        x = self.up2(x, x3)\n        x = self.outc(x)\n        return x\n```\nNow there are some literature and codebase to start with BOTTOM-UP Network  : \n\nOfcourse below two public  Kernels : \nhttps://www.kaggle.com/hocop1/centernet-baseline\nhttps://www.kaggle.com/phoenix9032/center-resnet-trial\n\nOFTNET : \nhttps://arxiv.org/abs/1811.08188\nhttps://github.com/tom-roddick/oft\n\nCENTERNET : \nhttps://arxiv.org/abs/1904.07850\nhttps://github.com/xingyizhou/CenterNet\n\nPVNET:\nhttps://arxiv.org/pdf/1812.11788.pdf\nhttps://github.com/zju3dv/pvnet\n\nOnce I get familiar with Mask-RCNN and the corresponding pose detection frameworks e.g. 6D-VN , I will try to post a topic on TOP-DOWN NETWORK",
    "685946": "good job!\nthis competition is really difficult for me😭",
    "692419": "Hi buddy, I have read PVNET, the PVNET method requires the segmentation of vehicles, do you have any method to get the segmentation information？Officials did not provide the Mask for cars of interest.☹️",
    "750790": "nice",
    "752202": "great"
  },
  "source": "meta"
}