{
  "id": 127056,
  "title": "(part of) 7th place solution with code",
  "url": "/competitions/pku-autonomous-driving/discussion/127056",
  "author_name": "",
  "post_date": "2020-01-22T02:47:02.081755100Z",
  "votes": 15,
  "comment_count": 7,
  "views": 0,
  "content": "<p>At first I'd like to express my appreciation for my teammate @phalanx @hesene @lanjunyelan, you guys are awesome! And congrats for those who finished with prize or gold zone:) I also would like to thank to kaggle and host for hosting such a nice competition!</p>\n\n<p>I wrote a brief summary of my solution and opened up the code in <a href=\"https://github.com/bamps53/kaggle-autonomous-driving2019\">my github</a>.</p>\n\n<p>Our entire solution is in <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034\">this thread</a>, please check it out. </p>\n\n<hr>\n\n<p>I started with <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">hocop1's really great kernel</a>, so basic part of my code is based on his.\nMy main contributions are below:\n - Change the CenterNet implementation to <a href=\"https://github.com/see--/kuzushiji-recognition\">see---'s code for kuzushiji competition</a>\n - Change backbone to resnet152\n - change the depth representation to 1 / (sigmoid(z) - 1\n - remove x, y, yaw(actually pitch), roll prediction\n - remove optimize xyz, just convert rcz to xyz\n - increase image size to (832, 2080) or (960, 2400) or (1088, 2720)\n - change the distance threshold to 4\n - change the confidence threshold to 0.1\n - ensemble 5 folds\n - remove false positive by using masks in 'test_masks/'</p>\n\n<p>I tried 3 scale ensemble, but it might be not so different from just 5 fold ensemble of image_size=(1088,2720). <br>\n My single fold model achieve public0.097/private0.091 and 5 fold ensemble got public0.117/private0.105. <br>\n With ensmebling Phalanx's model and Jhui's model, we got public0.129/private0.119.  </p>\n\n<hr>\n\n<p>code is available <a href=\"https://github.com/bamps53/kaggle-autonomous-driving2019\">here</a>.</p>",
  "messages": [
    {
      "id": "725356",
      "postDate": "01/22/2020 02:47:02",
      "content": "<p>At first I'd like to express my appreciation for my teammate @phalanx @hesene @lanjunyelan, you guys are awesome! And congrats for those who finished with prize or gold zone:) I also would like to thank to kaggle and host for hosting such a nice competition!</p>\n\n<p>I wrote a brief summary of my solution and opened up the code in <a href=\"https://github.com/bamps53/kaggle-autonomous-driving2019\">my github</a>.</p>\n\n<p>Our entire solution is in <a href=\"https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034\">this thread</a>, please check it out. </p>\n\n<hr>\n\n<p>I started with <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">hocop1's really great kernel</a>, so basic part of my code is based on his.\nMy main contributions are below:\n - Change the CenterNet implementation to <a href=\"https://github.com/see--/kuzushiji-recognition\">see---'s code for kuzushiji competition</a>\n - Change backbone to resnet152\n - change the depth representation to 1 / (sigmoid(z) - 1\n - remove x, y, yaw(actually pitch), roll prediction\n - remove optimize xyz, just convert rcz to xyz\n - increase image size to (832, 2080) or (960, 2400) or (1088, 2720)\n - change the distance threshold to 4\n - change the confidence threshold to 0.1\n - ensemble 5 folds\n - remove false positive by using masks in 'test_masks/'</p>\n\n<p>I tried 3 scale ensemble, but it might be not so different from just 5 fold ensemble of image_size=(1088,2720). <br>\n My single fold model achieve public0.097/private0.091 and 5 fold ensemble got public0.117/private0.105. <br>\n With ensmebling Phalanx's model and Jhui's model, we got public0.129/private0.119.  </p>\n\n<hr>\n\n<p>code is available <a href=\"https://github.com/bamps53/kaggle-autonomous-driving2019\">here</a>.</p>",
      "rawMarkdown": "At first I'd like to express my appreciation for my teammate @phalanx @hesene @lanjunyelan, you guys are awesome! And congrats for those who finished with prize or gold zone:) I also would like to thank to kaggle and host for hosting such a nice competition!\n\nI wrote a brief summary of my solution and opened up the code in [my github](https://github.com/bamps53/kaggle-autonomous-driving2019).\n\nOur entire solution is in [this thread](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034), please check it out. \n\n---\n\nI started with [hocop1's really great kernel](https://www.kaggle.com/hocop1/centernet-baseline), so basic part of my code is based on his.\nMy main contributions are below:\n - Change the CenterNet implementation to [see---'s code for kuzushiji competition](https://github.com/see--/kuzushiji-recognition)\n - Change backbone to resnet152\n - change the depth representation to 1 / (sigmoid(z) - 1\n - remove x, y, yaw(actually pitch), roll prediction\n - remove optimize xyz, just convert rcz to xyz\n - increase image size to (832, 2080) or (960, 2400) or (1088, 2720)\n - change the distance threshold to 4\n - change the confidence threshold to 0.1\n - ensemble 5 folds\n - remove false positive by using masks in 'test_masks/'\n \n\nI tried 3 scale ensemble, but it might be not so different from just 5 fold ensemble of image_size=(1088,2720).  \n My single fold model achieve public0.097/private0.091 and 5 fold ensemble got public0.117/private0.105.  \n With ensmebling Phalanx's model and Jhui's model, we got public0.129/private0.119.  \n\n---\ncode is available [here](https://github.com/bamps53/kaggle-autonomous-driving2019).",
      "votes": null
    },
    {
      "id": "725402",
      "postDate": "01/22/2020 03:52:22",
      "content": "<p>Awesome solution,I learned many from it.Thanks for sharing codes.I can't understand why you remove optimize xyz, just convert rcz to xyz.</p>",
      "rawMarkdown": "Awesome solution,I learned many from it.Thanks for sharing codes.I can't understand why you remove optimize xyz, just convert rcz to xyz.",
      "votes": null
    },
    {
      "id": "725616",
      "postDate": "01/22/2020 09:52:35",
      "content": "<p>Congrats dude...\n1) Could throw some light on  what is it trying to do..</p>\n\n<p>```\ndef gather_embeddings(self, embeddings, centers):\n        gathered_embeddings = []\n        for sample_index in range(len(centers)):\n            center_mask = centers[sample_index, :, 0] != -1\n            if center_mask.sum().item() == 0:\n                continue\n            per_sample_centers = centers[sample_index][center_mask]\n            emb = embeddings[sample_index][:, per_sample_centers[:, 1],\n                                           per_sample_centers[:, 0]].transpose(0, 1)\n            gathered_embeddings.append(emb)\n        gathered_embeddings = torch.cat(gathered_embeddings, 0)</p>\n\n<pre><code>    return gathered_embeddings\n</code></pre>\n\n<p>```</p>\n\n<p>2)\" depth representation to 1 / (sigmoid(z) - 1 \"  what is motivation for this in Loss function where i think u are using this.</p>",
      "rawMarkdown": "Congrats dude...\n1) Could throw some light on  what is it trying to do..\n\n \n```\ndef gather_embeddings(self, embeddings, centers):\n        gathered_embeddings = []\n        for sample_index in range(len(centers)):\n            center_mask = centers[sample_index, :, 0] != -1\n            if center_mask.sum().item() == 0:\n                continue\n            per_sample_centers = centers[sample_index][center_mask]\n            emb = embeddings[sample_index][:, per_sample_centers[:, 1],\n                                           per_sample_centers[:, 0]].transpose(0, 1)\n            gathered_embeddings.append(emb)\n        gathered_embeddings = torch.cat(gathered_embeddings, 0)\n\n        return gathered_embeddings\n\n```\n \n\n2)\" depth representation to 1 / (sigmoid(z) - 1 \"  what is motivation for this in Loss function where i think u are using this.",
      "votes": null
    },
    {
      "id": "725850",
      "postDate": "01/22/2020 14:46:02",
      "content": "<p>Hi <a href=\"/ynhuhu\">@ynhuhu</a>, I tried both and score was almost same or slightly better if you just convert. Somehow x,y predictions were not so accurate and were just noise for me. I also tried to optimize distance in world coordinate directly, but just converting was best. And also optimize xyz was so slow.</p>",
      "rawMarkdown": "Hi @ynhuhu, I tried both and score was almost same or slightly better if you just convert. Somehow x,y predictions were not so accurate and were just noise for me. I also tried to optimize distance in world coordinate directly, but just converting was best. And also optimize xyz was so slow.",
      "votes": null
    },
    {
      "id": "725859",
      "postDate": "01/22/2020 14:52:12",
      "content": "<p>Hi <a href=\"/jaideepvalani\">@jaideepvalani</a>,\n1) I just copied code from see---'s implementation and I haven't used this function for my solution. But what this function does is collect prediction for the point at where the ground truth is.\n2) I just followed the way CenterNet original paper or other 3d monocular prediciton paper use. I guess it helps model to represent wide range of numerical value,but some participants report it didn't work for him, so it could be not necessary.</p>",
      "rawMarkdown": "Hi @jaideepvalani,\n1) I just copied code from see---'s implementation and I haven't used this function for my solution. But what this function does is collect prediction for the point at where the ground truth is.\n2) I just followed the way CenterNet original paper or other 3d monocular prediciton paper use. I guess it helps model to represent wide range of numerical value,but some participants report it didn't work for him, so it could be not necessary.",
      "votes": null
    },
    {
      "id": "727498",
      "postDate": "01/23/2020 18:47:40",
      "content": "<p><a href=\"/bamps53\">@bamps53</a> thanks for the explanations. could you tell how long was the training (1 fold)? Thanks!</p>",
      "rawMarkdown": "bamps53 thanks for the explanations. could you tell how long was the training (1 fold)? Thanks!",
      "votes": null
    },
    {
      "id": "727689",
      "postDate": "01/24/2020 01:05:55",
      "content": "<p>For largest image size=(1088, 2720), it took 1 and half days to train 1 fold with 2080Ti....</p>",
      "rawMarkdown": "For largest image size=(1088, 2720), it took 1 and half days to train 1 fold with 2080Ti....",
      "votes": null
    },
    {
      "id": "732538",
      "postDate": "01/29/2020 23:26:55",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 725402,
      "author_name": "ynhuhu",
      "author_url": "",
      "post_date": "01/22/2020 03:52:22",
      "content": "<p>Awesome solution,I learned many from it.Thanks for sharing codes.I can't understand why you remove optimize xyz, just convert rcz to xyz.</p>",
      "votes": null,
      "replies": [
        {
          "id": 725850,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "01/22/2020 14:46:02",
          "content": "<p>Hi <a href=\"/ynhuhu\">@ynhuhu</a>, I tried both and score was almost same or slightly better if you just convert. Somehow x,y predictions were not so accurate and were just noise for me. I also tried to optimize distance in world coordinate directly, but just converting was best. And also optimize xyz was so slow.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 725616,
      "author_name": "jaideepvalani",
      "author_url": "",
      "post_date": "01/22/2020 09:52:35",
      "content": "<p>Congrats dude...\n1) Could throw some light on  what is it trying to do..</p>\n\n<p>```\ndef gather_embeddings(self, embeddings, centers):\n        gathered_embeddings = []\n        for sample_index in range(len(centers)):\n            center_mask = centers[sample_index, :, 0] != -1\n            if center_mask.sum().item() == 0:\n                continue\n            per_sample_centers = centers[sample_index][center_mask]\n            emb = embeddings[sample_index][:, per_sample_centers[:, 1],\n                                           per_sample_centers[:, 0]].transpose(0, 1)\n            gathered_embeddings.append(emb)\n        gathered_embeddings = torch.cat(gathered_embeddings, 0)</p>\n\n<pre><code>    return gathered_embeddings\n</code></pre>\n\n<p>```</p>\n\n<p>2)\" depth representation to 1 / (sigmoid(z) - 1 \"  what is motivation for this in Loss function where i think u are using this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 725859,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "01/22/2020 14:52:12",
          "content": "<p>Hi <a href=\"/jaideepvalani\">@jaideepvalani</a>,\n1) I just copied code from see---'s implementation and I haven't used this function for my solution. But what this function does is collect prediction for the point at where the ground truth is.\n2) I just followed the way CenterNet original paper or other 3d monocular prediciton paper use. I guess it helps model to represent wide range of numerical value,but some participants report it didn't work for him, so it could be not necessary.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 727498,
      "author_name": "jesucristo",
      "author_url": "",
      "post_date": "01/23/2020 18:47:40",
      "content": "<p><a href=\"/bamps53\">@bamps53</a> thanks for the explanations. could you tell how long was the training (1 fold)? Thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 727689,
          "author_name": "bamps53",
          "author_url": "",
          "post_date": "01/24/2020 01:05:55",
          "content": "<p>For largest image size=(1088, 2720), it took 1 and half days to train 1 fold with 2080Ti....</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 732538,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/29/2020 23:26:55",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "725356": "At first I'd like to express my appreciation for my teammate @phalanx @hesene @lanjunyelan, you guys are awesome! And congrats for those who finished with prize or gold zone:) I also would like to thank to kaggle and host for hosting such a nice competition!\n\nI wrote a brief summary of my solution and opened up the code in [my github](https://github.com/bamps53/kaggle-autonomous-driving2019).\n\nOur entire solution is in [this thread](https://www.kaggle.com/c/pku-autonomous-driving/discussion/127034), please check it out. \n\n---\n\nI started with [hocop1's really great kernel](https://www.kaggle.com/hocop1/centernet-baseline), so basic part of my code is based on his.\nMy main contributions are below:\n - Change the CenterNet implementation to [see---'s code for kuzushiji competition](https://github.com/see--/kuzushiji-recognition)\n - Change backbone to resnet152\n - change the depth representation to 1 / (sigmoid(z) - 1\n - remove x, y, yaw(actually pitch), roll prediction\n - remove optimize xyz, just convert rcz to xyz\n - increase image size to (832, 2080) or (960, 2400) or (1088, 2720)\n - change the distance threshold to 4\n - change the confidence threshold to 0.1\n - ensemble 5 folds\n - remove false positive by using masks in 'test_masks/'\n \n\nI tried 3 scale ensemble, but it might be not so different from just 5 fold ensemble of image_size=(1088,2720).  \n My single fold model achieve public0.097/private0.091 and 5 fold ensemble got public0.117/private0.105.  \n With ensmebling Phalanx's model and Jhui's model, we got public0.129/private0.119.  \n\n---\ncode is available [here](https://github.com/bamps53/kaggle-autonomous-driving2019).",
    "725402": "Awesome solution,I learned many from it.Thanks for sharing codes.I can't understand why you remove optimize xyz, just convert rcz to xyz.",
    "725616": "Congrats dude...\n1) Could throw some light on  what is it trying to do..\n\n \n```\ndef gather_embeddings(self, embeddings, centers):\n        gathered_embeddings = []\n        for sample_index in range(len(centers)):\n            center_mask = centers[sample_index, :, 0] != -1\n            if center_mask.sum().item() == 0:\n                continue\n            per_sample_centers = centers[sample_index][center_mask]\n            emb = embeddings[sample_index][:, per_sample_centers[:, 1],\n                                           per_sample_centers[:, 0]].transpose(0, 1)\n            gathered_embeddings.append(emb)\n        gathered_embeddings = torch.cat(gathered_embeddings, 0)\n\n        return gathered_embeddings\n\n```\n \n\n2)\" depth representation to 1 / (sigmoid(z) - 1 \"  what is motivation for this in Loss function where i think u are using this.",
    "725850": "Hi @ynhuhu, I tried both and score was almost same or slightly better if you just convert. Somehow x,y predictions were not so accurate and were just noise for me. I also tried to optimize distance in world coordinate directly, but just converting was best. And also optimize xyz was so slow.",
    "725859": "Hi @jaideepvalani,\n1) I just copied code from see---'s implementation and I haven't used this function for my solution. But what this function does is collect prediction for the point at where the ground truth is.\n2) I just followed the way CenterNet original paper or other 3d monocular prediciton paper use. I guess it helps model to represent wide range of numerical value,but some participants report it didn't work for him, so it could be not necessary.",
    "727498": "bamps53 thanks for the explanations. could you tell how long was the training (1 fold)? Thanks!",
    "727689": "For largest image size=(1088, 2720), it took 1 and half days to train 1 fold with 2080Ti....",
    "732538": "Congrats and thanks for sharing!"
  },
  "source": "meta"
}