{
  "id": 127546,
  "title": "37th place brief writeup",
  "url": "/competitions/pku-autonomous-driving/discussion/127546",
  "author_name": "yama",
  "post_date": "2020-01-24T16:13:39.117000",
  "votes": 10,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Thanks to the host and useful kernels and discussions. \nI learned a lot during this competition.</p>\n\n<p>My code is based on <a href=\"https://www.kaggle.com/phoenix9032/center-resnet-starter\">Center-Resnet Starter</a>.\nMy public LB was 0.094 and private LB was 0.086.</p>\n\n<h2>Model</h2>\n\n<ul>\n<li>model architecture is the same as Center-Resnet Starter Kernel except below</li>\n<li>feed mask image and x,y position into model</li>\n<li>predict heatmap with focal loss following CenterNet paper</li>\n<li>change regression target from (x,y,z) to (u-diff, v-diff, z) following CenterNet paper</li>\n<li>regress log(z) instead of z ∵ depth affects by multiplication and log(z) distribution is more balanced than z distribution</li>\n</ul>\n\n<h2>Data Augumentation</h2>\n\n<ul>\n<li>(x/z, y/z) position jittering</li>\n<li>slight gauss noise</li>\n<li>slight randomContrastBrightness</li>\n</ul>\n\n<h2>Preprocessing / Post Processing</h2>\n\n<ul>\n<li>masking prediction using given masks</li>\n<li>restore color distorted test images.\nfor each image and each channel, stretch [0, '95 percentile value'] to [0, 255]\n<ul><li>probably no effect on LB, pointed out by <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127060\">this discussion</a></li></ul></li>\n</ul>\n\n<h2>Others</h2>\n\n<ul>\n<li>replacing confidence by Y-position has no effect. At this point, I doubted the evaluation metric.</li>\n<li>remove corruputed 5 train images</li>\n<li>adaptive heatmap threshold to predict at least one car per an image. discarded it since LB does not change</li>\n<li>increase epochs and change scheduling to ReduceLROnPlateau</li>\n<li>2x weights to regression targets to balance two types of losses. It improved LB</li>\n<li>add (x,y) position info as head input and add two 1x1 convs to head. \nbetter localCV and private LB (my final sub score + 0.002)\nI discarded it since public LB was bad (my final sub score - 0.007)</li>\n</ul>\n\n<h2>What did not work for me</h2>\n\n<ul>\n<li>smaller input size (w,h = 1536,512)</li>\n<li>larger input size + grad accumulation (accumlation_step=2)</li>\n<li>deformable convolution V2 (maybe because of my poor modeling skill)</li>\n<li>bins with in-bins regression for pitch (bins=4) following CenterNet paper</li>\n<li>predict pitch from camera view following CenterNet paper</li>\n</ul>\n\n<h2>What I should have tried</h2>\n\n<ul>\n<li>improve predictions of large cars. Below may be relevant:\n9th solution : difference models for cars at difference positions\n5th solution : FPN network</li>\n<li>ensemble</li>\n<li>other backbones (DLA34 or resnet34 or efficientnet-b0)</li>\n<li>change the each-car distance threshold</li>\n<li>flip augmentation</li>\n<li>use pretrained model (and use mask info for loss calculation)</li>\n</ul>\n\n<hr>\n\n<p>code is <a href=\"https://github.com/lisosia/kaggle-pku-autonomous-driving\">here</a></p>",
  "messages": [
    {
      "id": 728333,
      "postDate": "2020-01-24T16:13:39.117Z",
      "content": "<p>Thanks to the host and useful kernels and discussions. \nI learned a lot during this competition.</p>\n\n<p>My code is based on <a href=\"https://www.kaggle.com/phoenix9032/center-resnet-starter\">Center-Resnet Starter</a>.\nMy public LB was 0.094 and private LB was 0.086.</p>\n\n<h2>Model</h2>\n\n<ul>\n<li>model architecture is the same as Center-Resnet Starter Kernel except below</li>\n<li>feed mask image and x,y position into model</li>\n<li>predict heatmap with focal loss following CenterNet paper</li>\n<li>change regression target from (x,y,z) to (u-diff, v-diff, z) following CenterNet paper</li>\n<li>regress log(z) instead of z ∵ depth affects by multiplication and log(z) distribution is more balanced than z distribution</li>\n</ul>\n\n<h2>Data Augumentation</h2>\n\n<ul>\n<li>(x/z, y/z) position jittering</li>\n<li>slight gauss noise</li>\n<li>slight randomContrastBrightness</li>\n</ul>\n\n<h2>Preprocessing / Post Processing</h2>\n\n<ul>\n<li>masking prediction using given masks</li>\n<li>restore color distorted test images.\nfor each image and each channel, stretch [0, '95 percentile value'] to [0, 255]\n<ul><li>probably no effect on LB, pointed out by <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127060\">this discussion</a></li></ul></li>\n</ul>\n\n<h2>Others</h2>\n\n<ul>\n<li>replacing confidence by Y-position has no effect. At this point, I doubted the evaluation metric.</li>\n<li>remove corruputed 5 train images</li>\n<li>adaptive heatmap threshold to predict at least one car per an image. discarded it since LB does not change</li>\n<li>increase epochs and change scheduling to ReduceLROnPlateau</li>\n<li>2x weights to regression targets to balance two types of losses. It improved LB</li>\n<li>add (x,y) position info as head input and add two 1x1 convs to head. \nbetter localCV and private LB (my final sub score + 0.002)\nI discarded it since public LB was bad (my final sub score - 0.007)</li>\n</ul>\n\n<h2>What did not work for me</h2>\n\n<ul>\n<li>smaller input size (w,h = 1536,512)</li>\n<li>larger input size + grad accumulation (accumlation_step=2)</li>\n<li>deformable convolution V2 (maybe because of my poor modeling skill)</li>\n<li>bins with in-bins regression for pitch (bins=4) following CenterNet paper</li>\n<li>predict pitch from camera view following CenterNet paper</li>\n</ul>\n\n<h2>What I should have tried</h2>\n\n<ul>\n<li>improve predictions of large cars. Below may be relevant:\n9th solution : difference models for cars at difference positions\n5th solution : FPN network</li>\n<li>ensemble</li>\n<li>other backbones (DLA34 or resnet34 or efficientnet-b0)</li>\n<li>change the each-car distance threshold</li>\n<li>flip augmentation</li>\n<li>use pretrained model (and use mask info for loss calculation)</li>\n</ul>\n\n<hr>\n\n<p>code is <a href=\"https://github.com/lisosia/kaggle-pku-autonomous-driving\">here</a></p>",
      "rawMarkdown": "Thanks to the host and useful kernels and discussions. \nI learned a lot during this competition.\n\nMy code is based on [Center-Resnet Starter](https://www.kaggle.com/phoenix9032/center-resnet-starter).\nMy public LB was 0.094 and private LB was 0.086.\n\n## Model\n- model architecture is the same as Center-Resnet Starter Kernel except below\n- feed mask image and x,y position into model\n- predict heatmap with focal loss following CenterNet paper\n- change regression target from (x,y,z) to (u-diff, v-diff, z) following CenterNet paper\n- regress log(z) instead of z ∵ depth affects by multiplication and log(z) distribution is more balanced than z distribution\n\n## Data Augumentation\n- (x/z, y/z) position jittering\n- slight gauss noise\n- slight randomContrastBrightness\n\n## Preprocessing / Post Processing\n- masking prediction using given masks\n- restore color distorted test images.\n  for each image and each channel, stretch [0, '95 percentile value'] to [0, 255]\n  - probably no effect on LB, pointed out by [this discussion](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127060)\n  \n## Others\n- replacing confidence by Y-position has no effect. At this point, I doubted the evaluation metric.\n- remove corruputed 5 train images\n- adaptive heatmap threshold to predict at least one car per an image. discarded it since LB does not change\n- increase epochs and change scheduling to ReduceLROnPlateau\n- 2x weights to regression targets to balance two types of losses. It improved LB\n- add (x,y) position info as head input and add two 1x1 convs to head. \n   better localCV and private LB (my final sub score + 0.002)\n   I discarded it since public LB was bad (my final sub score - 0.007)\n\n## What did not work for me\n- smaller input size (w,h = 1536,512)\n- larger input size + grad accumulation (accumlation_step=2)\n- deformable convolution V2 (maybe because of my poor modeling skill)\n- bins with in-bins regression for pitch (bins=4) following CenterNet paper\n- predict pitch from camera view following CenterNet paper\n\n## What I should have tried\n- improve predictions of large cars. Below may be relevant:\n   9th solution : difference models for cars at difference positions\n   5th solution : FPN network\n- ensemble\n- other backbones (DLA34 or resnet34 or efficientnet-b0)\n- change the each-car distance threshold\n- flip augmentation\n- use pretrained model (and use mask info for loss calculation)\n\n-------------------------\ncode is [here](https://github.com/lisosia/kaggle-pku-autonomous-driving)",
      "votes": 10
    },
    {
      "id": 728647,
      "postDate": "2020-01-25T02:36:59.170Z",
      "content": "<p>Good job, I also want to follow original centernet paper, but I'm just lack of programming skill.😭 </p>",
      "rawMarkdown": "Good job, I also want to follow original centernet paper, but I'm just lack of programming skill.😭 "
    },
    {
      "id": 728428,
      "postDate": "2020-01-24T17:59:23.850Z",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 "
    },
    {
      "id": 728356,
      "postDate": "2020-01-24T16:32:44.830Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 728647,
      "author_name": "DiegoJohnson",
      "author_url": "",
      "post_date": "2020-01-25T02:36:59.170000",
      "content": "<p>Good job, I also want to follow original centernet paper, but I'm just lack of programming skill.😭 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728428,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-01-24T17:59:23.850000",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728356,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T16:32:44.830000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "728333": "Thanks to the host and useful kernels and discussions. \nI learned a lot during this competition.\n\nMy code is based on [Center-Resnet Starter](https://www.kaggle.com/phoenix9032/center-resnet-starter).\nMy public LB was 0.094 and private LB was 0.086.\n\n## Model\n- model architecture is the same as Center-Resnet Starter Kernel except below\n- feed mask image and x,y position into model\n- predict heatmap with focal loss following CenterNet paper\n- change regression target from (x,y,z) to (u-diff, v-diff, z) following CenterNet paper\n- regress log(z) instead of z ∵ depth affects by multiplication and log(z) distribution is more balanced than z distribution\n\n## Data Augumentation\n- (x/z, y/z) position jittering\n- slight gauss noise\n- slight randomContrastBrightness\n\n## Preprocessing / Post Processing\n- masking prediction using given masks\n- restore color distorted test images.\n  for each image and each channel, stretch [0, '95 percentile value'] to [0, 255]\n  - probably no effect on LB, pointed out by [this discussion](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127060)\n  \n## Others\n- replacing confidence by Y-position has no effect. At this point, I doubted the evaluation metric.\n- remove corruputed 5 train images\n- adaptive heatmap threshold to predict at least one car per an image. discarded it since LB does not change\n- increase epochs and change scheduling to ReduceLROnPlateau\n- 2x weights to regression targets to balance two types of losses. It improved LB\n- add (x,y) position info as head input and add two 1x1 convs to head. \n   better localCV and private LB (my final sub score + 0.002)\n   I discarded it since public LB was bad (my final sub score - 0.007)\n\n## What did not work for me\n- smaller input size (w,h = 1536,512)\n- larger input size + grad accumulation (accumlation_step=2)\n- deformable convolution V2 (maybe because of my poor modeling skill)\n- bins with in-bins regression for pitch (bins=4) following CenterNet paper\n- predict pitch from camera view following CenterNet paper\n\n## What I should have tried\n- improve predictions of large cars. Below may be relevant:\n   9th solution : difference models for cars at difference positions\n   5th solution : FPN network\n- ensemble\n- other backbones (DLA34 or resnet34 or efficientnet-b0)\n- change the each-car distance threshold\n- flip augmentation\n- use pretrained model (and use mask info for loss calculation)\n\n-------------------------\ncode is [here](https://github.com/lisosia/kaggle-pku-autonomous-driving)",
    "728647": "Good job, I also want to follow original centernet paper, but I'm just lack of programming skill.😭 ",
    "728428": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 ",
    "728356": ""
  }
}