{
  "id": 127065,
  "title": "(part of) 5th place solution",
  "url": "/competitions/pku-autonomous-driving/discussion/127065",
  "author_name": "4ui_iurz1",
  "post_date": "2020-01-22T03:43:16.262000",
  "votes": 30,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Congrats to all the competitors!<br>\nThanks the whole team <a href=\"https://www.kaggle.com/erniechiew\" target=\"_blank\">@erniechiew</a> <a href=\"https://www.kaggle.com/css919\" target=\"_blank\">@css919</a> for the great collaboration!</p>\n<p>For the other part, please see <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127145\" target=\"_blank\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/127145</a>.</p>\n<p>code:<br>\n<a href=\"https://github.com/4uiiurz1/kaggle-pku-autonomous-driving\" target=\"_blank\">https://github.com/4uiiurz1/kaggle-pku-autonomous-driving</a></p>\n<p>My approach is based on <a href=\"https://github.com/xingyizhou/CenterNet\" target=\"_blank\">CenterNet</a>.</p>\n<h3>Heads</h3>\n<ul>\n<li>heatmap[1]</li>\n<li>xy offset[2]</li>\n<li>z (depth)[1]</li>\n<li>pose[6]: cos(yaw), sin(yaw) cos(pitch), sin(pitch), cos(roll), sin(roll)</li>\n<li>wh[2]: It's not used for prediction, but PublicLB was improved by learning this as an auxiliary task.</li>\n</ul>\n<p>Heatmap's loss is Focal Loss, and the others are L1Loss. The weight of wh loss is 0.05. Mask regions of mask images are ignored when calculating loss.</p>\n<h3>Network Architecture</h3>\n<ul>\n<li><a href=\"https://github.com/Cadene/pretrained-models.pytorch\" target=\"_blank\">ResNet18 (pretrained ImageNet)</a> + FPN (channels: 256-&gt;128-&gt;64)</li>\n<li><a href=\"https://github.com/xingyizhou/CenterNet/blob/master/readme/MODEL_ZOO.md\" target=\"_blank\">DLA34 (pretrained KITTI 3DOP)</a> + FPN (channels: 256-&gt;256-&gt;256)</li>\n<li>Input size: 2560 x 2048 (2560 x 1024)</li>\n<li>Output size: 640 x 512 (640 x 256)</li>\n</ul>\n<p>Increasing the input size is very effective, mAP was improved dramatically.<br>\nI tried deeper networks (ResNet34, 50) but not worked.</p>\n<h3>Augmentation</h3>\n<ul>\n<li>HFlip (p=0.5): Flip images horizontally and <code>yaw *= -1, roll *= -1</code>.</li>\n<li>RandomShift (p=0.5, limit=0.1): Shift images and positions (x, y).</li>\n<li>RandomScale (p=0.5, limit=0.1): Scale images and positions (x, y, z).</li>\n<li>RandomHueSaturationValue (p=0.5, hue_limit=20)</li>\n<li>RandomBrightness (p=0.5, limit=0.2)</li>\n<li>RandomContrast (p=0.5, limit=0.2)</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>Optimizer: RAdam</li>\n<li>LR scheduler: CosineAnnealingLR (lr=1e-3 -&gt; 1e-5)</li>\n<li>50epochs</li>\n<li>5-folds cv</li>\n<li>Batch size: 4</li>\n</ul>\n<h3>Post Processing</h3>\n<ul>\n<li>Remove mask regions from predictions by multiplying heatmap by masks.</li>\n<li>NMS (distance threshold: 0.1): I'm not sure how effective this is…</li>\n<li>Find duplicate images with imagehash and ensemble them. PublicLB was slighly improved.</li>\n<li>Score threshold: 0.3 (for val mAP: 0.1)</li>\n</ul>\n<h3>Ensemble</h3>\n<p>Ensemble each fold models and two models (ResNet18, DLA34) by averaging the raw output maps.</p>\n<h3>Score Summary</h3>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>val mAP (tito's script)</th>\n<th>PublicLB</th>\n<th>PrivateLB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet18 + FPN</td>\n<td>0.257224305900695</td>\n<td>0.118</td>\n<td>0.109</td>\n</tr>\n<tr>\n<td>DLA34 + FPN</td>\n<td>0.2681900383192367</td>\n<td>0.118</td>\n<td>0.112</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>0.27363538075234406</td>\n<td>0.121</td>\n<td>0.115</td>\n</tr>\n</tbody>\n</table>\n<h3>What Didn't Work</h3>\n<ul>\n<li>Pseudo labeling</li>\n<li>TTA (hflip)</li>\n<li>Weight Standardization</li>\n<li>Group Normalization</li>\n<li>Deformable Convolution V2</li>\n<li>Quaternion + L1Loss</li>\n<li>Very large input size (3360 x 2688)</li>\n<li>Eigen's depth prediction method used in CenterNet paper (<code>z = 1 / sigmoid(output) − 1</code>)</li>\n</ul>",
  "messages": [
    {
      "id": 725396,
      "postDate": "2020-01-22T03:43:16.263Z",
      "content": "<p>Congrats to all the competitors!<br>\nThanks the whole team <a href=\"https://www.kaggle.com/erniechiew\" target=\"_blank\">@erniechiew</a> <a href=\"https://www.kaggle.com/css919\" target=\"_blank\">@css919</a> for the great collaboration!</p>\n<p>For the other part, please see <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127145\" target=\"_blank\">https://www.kaggle.com/c/pku-autonomous-driving/discussion/127145</a>.</p>\n<p>code:<br>\n<a href=\"https://github.com/4uiiurz1/kaggle-pku-autonomous-driving\" target=\"_blank\">https://github.com/4uiiurz1/kaggle-pku-autonomous-driving</a></p>\n<p>My approach is based on <a href=\"https://github.com/xingyizhou/CenterNet\" target=\"_blank\">CenterNet</a>.</p>\n<h3>Heads</h3>\n<ul>\n<li>heatmap[1]</li>\n<li>xy offset[2]</li>\n<li>z (depth)[1]</li>\n<li>pose[6]: cos(yaw), sin(yaw) cos(pitch), sin(pitch), cos(roll), sin(roll)</li>\n<li>wh[2]: It's not used for prediction, but PublicLB was improved by learning this as an auxiliary task.</li>\n</ul>\n<p>Heatmap's loss is Focal Loss, and the others are L1Loss. The weight of wh loss is 0.05. Mask regions of mask images are ignored when calculating loss.</p>\n<h3>Network Architecture</h3>\n<ul>\n<li><a href=\"https://github.com/Cadene/pretrained-models.pytorch\" target=\"_blank\">ResNet18 (pretrained ImageNet)</a> + FPN (channels: 256-&gt;128-&gt;64)</li>\n<li><a href=\"https://github.com/xingyizhou/CenterNet/blob/master/readme/MODEL_ZOO.md\" target=\"_blank\">DLA34 (pretrained KITTI 3DOP)</a> + FPN (channels: 256-&gt;256-&gt;256)</li>\n<li>Input size: 2560 x 2048 (2560 x 1024)</li>\n<li>Output size: 640 x 512 (640 x 256)</li>\n</ul>\n<p>Increasing the input size is very effective, mAP was improved dramatically.<br>\nI tried deeper networks (ResNet34, 50) but not worked.</p>\n<h3>Augmentation</h3>\n<ul>\n<li>HFlip (p=0.5): Flip images horizontally and <code>yaw *= -1, roll *= -1</code>.</li>\n<li>RandomShift (p=0.5, limit=0.1): Shift images and positions (x, y).</li>\n<li>RandomScale (p=0.5, limit=0.1): Scale images and positions (x, y, z).</li>\n<li>RandomHueSaturationValue (p=0.5, hue_limit=20)</li>\n<li>RandomBrightness (p=0.5, limit=0.2)</li>\n<li>RandomContrast (p=0.5, limit=0.2)</li>\n</ul>\n<h3>Training</h3>\n<ul>\n<li>Optimizer: RAdam</li>\n<li>LR scheduler: CosineAnnealingLR (lr=1e-3 -&gt; 1e-5)</li>\n<li>50epochs</li>\n<li>5-folds cv</li>\n<li>Batch size: 4</li>\n</ul>\n<h3>Post Processing</h3>\n<ul>\n<li>Remove mask regions from predictions by multiplying heatmap by masks.</li>\n<li>NMS (distance threshold: 0.1): I'm not sure how effective this is…</li>\n<li>Find duplicate images with imagehash and ensemble them. PublicLB was slighly improved.</li>\n<li>Score threshold: 0.3 (for val mAP: 0.1)</li>\n</ul>\n<h3>Ensemble</h3>\n<p>Ensemble each fold models and two models (ResNet18, DLA34) by averaging the raw output maps.</p>\n<h3>Score Summary</h3>\n<table>\n<thead>\n<tr>\n<th>model</th>\n<th>val mAP (tito's script)</th>\n<th>PublicLB</th>\n<th>PrivateLB</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>ResNet18 + FPN</td>\n<td>0.257224305900695</td>\n<td>0.118</td>\n<td>0.109</td>\n</tr>\n<tr>\n<td>DLA34 + FPN</td>\n<td>0.2681900383192367</td>\n<td>0.118</td>\n<td>0.112</td>\n</tr>\n<tr>\n<td>Ensemble</td>\n<td>0.27363538075234406</td>\n<td>0.121</td>\n<td>0.115</td>\n</tr>\n</tbody>\n</table>\n<h3>What Didn't Work</h3>\n<ul>\n<li>Pseudo labeling</li>\n<li>TTA (hflip)</li>\n<li>Weight Standardization</li>\n<li>Group Normalization</li>\n<li>Deformable Convolution V2</li>\n<li>Quaternion + L1Loss</li>\n<li>Very large input size (3360 x 2688)</li>\n<li>Eigen's depth prediction method used in CenterNet paper (<code>z = 1 / sigmoid(output) − 1</code>)</li>\n</ul>",
      "rawMarkdown": "Congrats to all the competitors!\nThanks the whole team @erniechiew @css919 for the great collaboration!\n\nFor the other part, please see https://www.kaggle.com/c/pku-autonomous-driving/discussion/127145.\n\ncode:\nhttps://github.com/4uiiurz1/kaggle-pku-autonomous-driving\n\nMy approach is based on [CenterNet](https://github.com/xingyizhou/CenterNet).\n\n### Heads\n- heatmap[1]\n- xy offset[2]\n- z (depth)[1]\n- pose[6]: cos(yaw), sin(yaw) cos(pitch), sin(pitch), cos(roll), sin(roll)\n- wh[2]: It's not used for prediction, but PublicLB was improved by learning this as an auxiliary task.\n\nHeatmap's loss is Focal Loss, and the others are L1Loss. The weight of wh loss is 0.05. Mask regions of mask images are ignored when calculating loss.\n\n### Network Architecture\n- [ResNet18 (pretrained ImageNet)](https://github.com/Cadene/pretrained-models.pytorch) + FPN (channels: 256-&gt;128-&gt;64)\n- [DLA34 (pretrained KITTI 3DOP)](https://github.com/xingyizhou/CenterNet/blob/master/readme/MODEL_ZOO.md) + FPN (channels: 256-&gt;256-&gt;256)\n- Input size: 2560 x 2048 (2560 x 1024)\n- Output size: 640 x 512 (640 x 256)\n\nIncreasing the input size is very effective, mAP was improved dramatically.\nI tried deeper networks (ResNet34, 50) but not worked.\n\n### Augmentation\n- HFlip (p=0.5): Flip images horizontally and `yaw *= -1, roll *= -1`.\n- RandomShift (p=0.5, limit=0.1): Shift images and positions (x, y).\n- RandomScale (p=0.5, limit=0.1): Scale images and positions (x, y, z).\n- RandomHueSaturationValue (p=0.5, hue_limit=20)\n- RandomBrightness (p=0.5, limit=0.2)\n- RandomContrast (p=0.5, limit=0.2)\n\n### Training\n- Optimizer: RAdam\n- LR scheduler: CosineAnnealingLR (lr=1e-3 -&gt; 1e-5)\n- 50epochs\n- 5-folds cv\n- Batch size: 4\n\n### Post Processing\n- Remove mask regions from predictions by multiplying heatmap by masks.\n- NMS (distance threshold: 0.1): I'm not sure how effective this is...\n- Find duplicate images with imagehash and ensemble them. PublicLB was slighly improved.\n- Score threshold: 0.3 (for val mAP: 0.1)\n\n### Ensemble\nEnsemble each fold models and two models (ResNet18, DLA34) by averaging the raw output maps.\n\n### Score Summary\n| model          | val mAP (tito's script) | PublicLB   | PrivateLB |\n|----------------|:-----------------------:|:----------:|:----------|\n| ResNet18 + FPN | 0.257224305900695       | 0.118      | 0.109     |\n| DLA34 + FPN    | 0.2681900383192367      | 0.118      | 0.112     |\n| Ensemble       | 0.27363538075234406     | 0.121      | 0.115     |\n\n### What Didn't Work\n- Pseudo labeling\n- TTA (hflip)\n- Weight Standardization\n- Group Normalization\n- Deformable Convolution V2\n- Quaternion + L1Loss\n- Very large input size (3360 x 2688)\n- Eigen's depth prediction method used in CenterNet paper (`z = 1 / sigmoid(output) − 1`)\n",
      "votes": 30
    },
    {
      "id": 725408,
      "postDate": "2020-01-22T04:03:34.133Z",
      "content": "<p>Thanks you for the great collaboration too! Looking forward to working together again!</p>",
      "rawMarkdown": "Thanks you for the great collaboration too! Looking forward to working together again!",
      "votes": 3,
      "replies": [
        {
          "id": 725679,
          "postDate": "2020-01-22T11:23:12.647Z",
          "content": "<p><a href=\"/css919\">@css919</a> Let's win the prize next time!</p>",
          "rawMarkdown": "@css919 Let's win the prize next time!",
          "votes": 3
        }
      ]
    },
    {
      "id": 732533,
      "postDate": "2020-01-29T23:20:15.543Z",
      "content": "<p>Congrats, thanks for sharing!</p>",
      "rawMarkdown": "Congrats, thanks for sharing!",
      "replies": [
        {
          "id": 733856,
          "postDate": "2020-01-31T15:41:47.230Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        }
      ]
    },
    {
      "id": 726171,
      "postDate": "2020-01-22T22:28:06.043Z",
      "content": "<p>I used a similar approach based on CenterNet with stacked hour glass as base network, probably should have used some pretrained network. Reading through this post, I can also see other details which could have helped my solution. Thanks for sharing :). \nP.S. I too found that Quaternion + L1 loss and eigen depth prediction method didn't work.</p>",
      "rawMarkdown": "I used a similar approach based on CenterNet with stacked hour glass as base network, probably should have used some pretrained network. Reading through this post, I can also see other details which could have helped my solution. Thanks for sharing :). \nP.S. I too found that Quaternion + L1 loss and eigen depth prediction method didn't work.",
      "replies": [
        {
          "id": 728560,
          "postDate": "2020-01-24T22:35:50.177Z",
          "content": "<p><a href=\"/shanmukhamanoj11\">@shanmukhamanoj11</a> You're welcome!</p>",
          "rawMarkdown": "@shanmukhamanoj11 You're welcome!"
        }
      ]
    },
    {
      "id": 725607,
      "postDate": "2020-01-22T09:32:52.837Z",
      "content": "<p>Thank you for sharing！Looking forward to a whole solution</p>",
      "rawMarkdown": "Thank you for sharing！Looking forward to a whole solution",
      "replies": [
        {
          "id": 728555,
          "postDate": "2020-01-24T22:23:44.830Z",
          "content": "<p><a href=\"/miaorain\">@miaorain</a> You're welcome!</p>",
          "rawMarkdown": "@miaorain You're welcome!"
        }
      ]
    },
    {
      "id": 725431,
      "postDate": "2020-01-22T05:03:58.837Z",
      "content": "<p>Thank you! So what made a large difference is FPN?</p>",
      "rawMarkdown": "Thank you! So what made a large difference is FPN?",
      "replies": [
        {
          "id": 728559,
          "postDate": "2020-01-24T22:35:26.870Z",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> You're welcome!\nThe following is maybe what improved my score.\n- large input size\n- RAdam (made training stable)\n- width &amp; height prediction \n- KITTI pretrained backbone</p>",
          "rawMarkdown": "@tonychenxyz You're welcome!\nThe following is maybe what improved my score.\n- large input size\n- RAdam (made training stable)\n- width &amp; height prediction \n- KITTI pretrained backbone"
        }
      ]
    },
    {
      "id": 733379,
      "postDate": "2020-01-31T03:44:20.580Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 737773,
      "postDate": "2020-02-05T19:00:20.657Z",
      "content": "<p>Thanks for sharing! learned a lot! </p>",
      "rawMarkdown": "Thanks for sharing! learned a lot! "
    }
  ],
  "comments": [
    {
      "id": 725408,
      "author_name": "ShinSiang",
      "author_url": "",
      "post_date": "2020-01-22T04:03:34.133000",
      "content": "<p>Thanks you for the great collaboration too! Looking forward to working together again!</p>",
      "votes": 3,
      "replies": [
        {
          "id": 725679,
          "author_name": "4ui_iurz1",
          "author_url": "",
          "post_date": "2020-01-22T11:23:12.647000",
          "content": "<p><a href=\"/css919\">@css919</a> Let's win the prize next time!</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 732533,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-29T23:20:15.543000",
      "content": "<p>Congrats, thanks for sharing!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 733856,
          "author_name": "4ui_iurz1",
          "author_url": "",
          "post_date": "2020-01-31T15:41:47.230000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 726171,
      "author_name": "Manoj",
      "author_url": "",
      "post_date": "2020-01-22T22:28:06.043000",
      "content": "<p>I used a similar approach based on CenterNet with stacked hour glass as base network, probably should have used some pretrained network. Reading through this post, I can also see other details which could have helped my solution. Thanks for sharing :). \nP.S. I too found that Quaternion + L1 loss and eigen depth prediction method didn't work.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728560,
          "author_name": "4ui_iurz1",
          "author_url": "",
          "post_date": "2020-01-24T22:35:50.177000",
          "content": "<p><a href=\"/shanmukhamanoj11\">@shanmukhamanoj11</a> You're welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 725607,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-22T09:32:52.837000",
      "content": "<p>Thank you for sharing！Looking forward to a whole solution</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728555,
          "author_name": "4ui_iurz1",
          "author_url": "",
          "post_date": "2020-01-24T22:23:44.830000",
          "content": "<p><a href=\"/miaorain\">@miaorain</a> You're welcome!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 725431,
      "author_name": "Gold Retriever",
      "author_url": "",
      "post_date": "2020-01-22T05:03:58.837000",
      "content": "<p>Thank you! So what made a large difference is FPN?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728559,
          "author_name": "4ui_iurz1",
          "author_url": "",
          "post_date": "2020-01-24T22:35:26.870000",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> You're welcome!\nThe following is maybe what improved my score.\n- large input size\n- RAdam (made training stable)\n- width &amp; height prediction \n- KITTI pretrained backbone</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 733379,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-31T03:44:20.580000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 737773,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-02-05T19:00:20.657000",
      "content": "<p>Thanks for sharing! learned a lot! </p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "725396": "Congrats to all the competitors!\nThanks the whole team @erniechiew @css919 for the great collaboration!\n\nFor the other part, please see https://www.kaggle.com/c/pku-autonomous-driving/discussion/127145.\n\ncode:\nhttps://github.com/4uiiurz1/kaggle-pku-autonomous-driving\n\nMy approach is based on [CenterNet](https://github.com/xingyizhou/CenterNet).\n\n### Heads\n- heatmap[1]\n- xy offset[2]\n- z (depth)[1]\n- pose[6]: cos(yaw), sin(yaw) cos(pitch), sin(pitch), cos(roll), sin(roll)\n- wh[2]: It's not used for prediction, but PublicLB was improved by learning this as an auxiliary task.\n\nHeatmap's loss is Focal Loss, and the others are L1Loss. The weight of wh loss is 0.05. Mask regions of mask images are ignored when calculating loss.\n\n### Network Architecture\n- [ResNet18 (pretrained ImageNet)](https://github.com/Cadene/pretrained-models.pytorch) + FPN (channels: 256-&gt;128-&gt;64)\n- [DLA34 (pretrained KITTI 3DOP)](https://github.com/xingyizhou/CenterNet/blob/master/readme/MODEL_ZOO.md) + FPN (channels: 256-&gt;256-&gt;256)\n- Input size: 2560 x 2048 (2560 x 1024)\n- Output size: 640 x 512 (640 x 256)\n\nIncreasing the input size is very effective, mAP was improved dramatically.\nI tried deeper networks (ResNet34, 50) but not worked.\n\n### Augmentation\n- HFlip (p=0.5): Flip images horizontally and `yaw *= -1, roll *= -1`.\n- RandomShift (p=0.5, limit=0.1): Shift images and positions (x, y).\n- RandomScale (p=0.5, limit=0.1): Scale images and positions (x, y, z).\n- RandomHueSaturationValue (p=0.5, hue_limit=20)\n- RandomBrightness (p=0.5, limit=0.2)\n- RandomContrast (p=0.5, limit=0.2)\n\n### Training\n- Optimizer: RAdam\n- LR scheduler: CosineAnnealingLR (lr=1e-3 -&gt; 1e-5)\n- 50epochs\n- 5-folds cv\n- Batch size: 4\n\n### Post Processing\n- Remove mask regions from predictions by multiplying heatmap by masks.\n- NMS (distance threshold: 0.1): I'm not sure how effective this is...\n- Find duplicate images with imagehash and ensemble them. PublicLB was slighly improved.\n- Score threshold: 0.3 (for val mAP: 0.1)\n\n### Ensemble\nEnsemble each fold models and two models (ResNet18, DLA34) by averaging the raw output maps.\n\n### Score Summary\n| model          | val mAP (tito's script) | PublicLB   | PrivateLB |\n|----------------|:-----------------------:|:----------:|:----------|\n| ResNet18 + FPN | 0.257224305900695       | 0.118      | 0.109     |\n| DLA34 + FPN    | 0.2681900383192367      | 0.118      | 0.112     |\n| Ensemble       | 0.27363538075234406     | 0.121      | 0.115     |\n\n### What Didn't Work\n- Pseudo labeling\n- TTA (hflip)\n- Weight Standardization\n- Group Normalization\n- Deformable Convolution V2\n- Quaternion + L1Loss\n- Very large input size (3360 x 2688)\n- Eigen's depth prediction method used in CenterNet paper (`z = 1 / sigmoid(output) − 1`)\n",
    "725408": "Thanks you for the great collaboration too! Looking forward to working together again!",
    "732533": "Congrats, thanks for sharing!",
    "726171": "I used a similar approach based on CenterNet with stacked hour glass as base network, probably should have used some pretrained network. Reading through this post, I can also see other details which could have helped my solution. Thanks for sharing :). \nP.S. I too found that Quaternion + L1 loss and eigen depth prediction method didn't work.",
    "725607": "Thank you for sharing！Looking forward to a whole solution",
    "725431": "Thank you! So what made a large difference is FPN?",
    "733379": "",
    "737773": "Thanks for sharing! learned a lot! "
  }
}