{
  "id": 127048,
  "title": "What Makes the Differences?",
  "url": "/competitions/pku-autonomous-driving/discussion/127048",
  "author_name": "",
  "post_date": "2020-01-22T01:22:08.470516600Z",
  "votes": 4,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Congrats to everyone who gets awards.</p>\n\n<p>Missed the bronze by a bit but learned a lot.</p>\n\n<p>I just don't know how my results are so difference from higher ranking people although I pretty much did the similar things.</p>\n\n<p>My final solution consists of ensemble of 3 centernet-hourglass104 models with different random seeds and augmentation rate. I did horizontal flip, brightness, gaussian noise, color shift, and random contrast.</p>\n\n<p>I tried focal loss with different gamma parameters, other models such as dla, resnet, densenet etc., and I tried psuedolabeling, but none of them worked.</p>\n\n<p>I would like to hear what you think.</p>",
  "messages": [
    {
      "id": "725294",
      "postDate": "01/22/2020 01:22:08",
      "content": "<p>Congrats to everyone who gets awards.</p>\n\n<p>Missed the bronze by a bit but learned a lot.</p>\n\n<p>I just don't know how my results are so difference from higher ranking people although I pretty much did the similar things.</p>\n\n<p>My final solution consists of ensemble of 3 centernet-hourglass104 models with different random seeds and augmentation rate. I did horizontal flip, brightness, gaussian noise, color shift, and random contrast.</p>\n\n<p>I tried focal loss with different gamma parameters, other models such as dla, resnet, densenet etc., and I tried psuedolabeling, but none of them worked.</p>\n\n<p>I would like to hear what you think.</p>",
      "rawMarkdown": "Congrats to everyone who gets awards.\n\nMissed the bronze by a bit but learned a lot.\n\nI just don't know how my results are so difference from higher ranking people although I pretty much did the similar things.\n\nMy final solution consists of ensemble of 3 centernet-hourglass104 models with different random seeds and augmentation rate. I did horizontal flip, brightness, gaussian noise, color shift, and random contrast.\n\nI tried focal loss with different gamma parameters, other models such as dla, resnet, densenet etc., and I tried psuedolabeling, but none of them worked.\n\nI would like to hear what you think.",
      "votes": null
    },
    {
      "id": "725298",
      "postDate": "01/22/2020 01:28:21",
      "content": "<p>me too, i just splite three part of train dataset. one for train 81% train, 9%val,10%test. and at test   dataset, use public metric, my kernel can get 0.158, but at LB i just get 0.06, so I'm very confuse . </p>",
      "rawMarkdown": "me too, i just splite three part of train dataset. one for train 81% train, 9%val,10%test. and at test   dataset, use public metric, my kernel can get 0.158, but at LB i just get 0.06, so I'm very confuse .",
      "votes": null
    },
    {
      "id": "725299",
      "postDate": "01/22/2020 01:29:56",
      "content": "<p>the local CV metric might be different from official one, so I guess that difference is normal.</p>",
      "rawMarkdown": "the local CV metric might be different from official one, so I guess that difference is normal.",
      "votes": null
    },
    {
      "id": "725306",
      "postDate": "01/22/2020 01:39:43",
      "content": "<p>for now, I think maybe submission format  has some different or metric has bugs. I will wait some guyes publish theirs kernel make a compare.\nI find some publish kernel used train-dataset cal score just get 0.04, and their LB do not shake up so mutch. \nI use train-dataset by publish metric can get almost 0.22.</p>",
      "rawMarkdown": "for now, I think maybe submission format  has some different or metric has bugs. I will wait some guyes publish theirs kernel make a compare.\nI find some publish kernel used train-dataset cal score just get 0.04, and their LB do not shake up so mutch. \nI use train-dataset by publish metric can get almost 0.22.",
      "votes": null
    },
    {
      "id": "725334",
      "postDate": "01/22/2020 02:06:40",
      "content": "<p>Same here. I have to say that those little details matter and they can only be obtained through past experiences and careful experiments. We've still got a long way to go.</p>",
      "rawMarkdown": "Same here. I have to say that those little details matter and they can only be obtained through past experiences and careful experiments. We've still got a long way to go.",
      "votes": null
    },
    {
      "id": "725341",
      "postDate": "01/22/2020 02:23:34",
      "content": "<p>And the unclear metric for this competition is also a bummer, so I could not do many effective experiments.</p>",
      "rawMarkdown": "And the unclear metric for this competition is also a bummer, so I could not do many effective experiments.",
      "votes": null
    },
    {
      "id": "725433",
      "postDate": "01/22/2020 05:13:46",
      "content": "<p>And one more thing: the mAP metric is not closely related to the hm loss or regression loss. :)</p>",
      "rawMarkdown": "And one more thing: the mAP metric is not closely related to the hm loss or regression loss. :)",
      "votes": null
    },
    {
      "id": "725488",
      "postDate": "01/22/2020 06:50:22",
      "content": "<p>My final solution is based on <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">hocop1's kernel</a>.\nI adopted the same model architecture as original and my best solution didn't use ensemble.\nHere are the main differences from the original kernel.</p>\n\n<ul>\n<li>100 epoch learning</li>\n<li>Focal loss for mask whose parameters are default</li>\n<li>Sigmoid threshold: 0.2</li>\n<li>Input resolution for cnn model: 1024x320 -&gt; 2048x640</li>\n<li>Data augmentation\n<ul><li>Pixel level augmentation like yours</li>\n<li>Random scale and crop</li></ul></li>\n<li>Postprocess\n<ul><li>Modify <code>optimize_xy</code> not to use <code>xzy_slope</code></li>\n<li>Drop false positives by using test_masks</li></ul></li>\n</ul>\n\n<p>From the above, I think preprocess and postprocess have huge impacts for bronze-level score.</p>\n\n<p>I hope this will help.</p>",
      "rawMarkdown": "My final solution is based on [hocop1's kernel](https://www.kaggle.com/hocop1/centernet-baseline).\nI adopted the same model architecture as original and my best solution didn't use ensemble.\nHere are the main differences from the original kernel.\n\n- 100 epoch learning\n- Focal loss for mask whose parameters are default\n- Sigmoid threshold: 0.2\n- Input resolution for cnn model: 1024x320 -&gt; 2048x640\n- Data augmentation\n  - Pixel level augmentation like yours\n  - Random scale and crop\n- Postprocess\n  - Modify `optimize_xy` not to use `xzy_slope`\n  - Drop false positives by using test_masks\n\nFrom the above, I think preprocess and postprocess have huge impacts for bronze-level score.\n\nI hope this will help.",
      "votes": null
    },
    {
      "id": "725493",
      "postDate": "01/22/2020 06:58:00",
      "content": "<p>Thank you very much! Could you specify the focal loss parameters used? (especially sigma for Gaussian heatmap since it was not specified in paper). Also, it is really impressive that you did 100 epoch training because my dev loss or mAP never improved after about 10-15 epochs of training. <a href=\"/kmizunoster\">@kmizunoster</a> </p>",
      "rawMarkdown": "Thank you very much! Could you specify the focal loss parameters used? (especially sigma for Gaussian heatmap since it was not specified in paper). Also, it is really impressive that you did 100 epoch training because my dev loss or mAP never improved after about 10-15 epochs of training. @kmizunoster",
      "votes": null
    },
    {
      "id": "725537",
      "postDate": "01/22/2020 08:11:52",
      "content": "<p>Hi <a href=\"/kmizunoster\">@kmizunoster</a> does the image size 2048x640 lead to the 100 epochs training? I've tried a larger size but the improvement was really slow with little computation resources so I finally gave it up ...</p>",
      "rawMarkdown": "Hi @kmizunoster does the image size 2048x640 lead to the 100 epochs training? I've tried a larger size but the improvement was really slow with little computation resources so I finally gave it up ...",
      "votes": null
    },
    {
      "id": "725566",
      "postDate": "01/22/2020 08:53:04",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> \nI used <a href=\"https://github.com/xingyizhou/CenterNet/blob/8ef87b433529ac8f8bd4f95707f6bc05052c55e9/src/lib/models/losses.py#L42\">CenterNet's official implementation of FocalLoss</a> as it is.\nIn my case (test_size=0.2), my dev loss is saturated around 35 epoch, but a model after 100 epoch training got better LB score than one chosen by best dev loss.</p>\n\n<p><a href=\"/syoya1997\">@syoya1997</a> \nI'm not sure that 100 epochs training is required for the image size 2048x640,\nbut longer epochs training provided better results regerdless of image size in my experience on this competition.</p>",
      "rawMarkdown": "tonychenxyz \nI used [CenterNet's official implementation of FocalLoss](https://github.com/xingyizhou/CenterNet/blob/8ef87b433529ac8f8bd4f95707f6bc05052c55e9/src/lib/models/losses.py#L42) as it is.\nIn my case (test_size=0.2), my dev loss is saturated around 35 epoch, but a model after 100 epoch training got better LB score than one chosen by best dev loss.\n\n@syoya1997 \nI'm not sure that 100 epochs training is required for the image size 2048x640,\nbut longer epochs training provided better results regerdless of image size in my experience on this competition.",
      "votes": null
    },
    {
      "id": "725572",
      "postDate": "01/22/2020 08:58:59",
      "content": "<p>Thanks. I usually trained until my dev loss not decreasing any more. Wondering why you choose to train for this long? 100 epochs take a lot of time.</p>",
      "rawMarkdown": "Thanks. I usually trained until my dev loss not decreasing any more. Wondering why you choose to train for this long? 100 epochs take a lot of time.",
      "votes": null
    },
    {
      "id": "727804",
      "postDate": "01/24/2020 04:27:51",
      "content": "<p>I think image size, focal loss, longer training scheduling makes a big difference in getting a bronze score.\nI made a kernel to check if we can get LB 0.07 with the baseline model.\n<a href=\"https://www.kaggle.com/kyoshioka47/centernet-baseline-lb-pb-0-07\">https://www.kaggle.com/kyoshioka47/centernet-baseline-lb-pb-0-07</a>\nThe trained model gets LB 0.07, PB 0.069.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F347059bdaab492b0507096cd77689ea3%2Faaa.JPG?generation=1579840252559961&amp;alt=media\" alt=\"\"></p>\n\n<p>To get a silver level score, we need to implement a better model (centernet + FPN) to reach LB 0.09~.\nI will open that kernel when the training is done.</p>",
      "rawMarkdown": "I think image size, focal loss, longer training scheduling makes a big difference in getting a bronze score.\nI made a kernel to check if we can get LB 0.07 with the baseline model.\nhttps://www.kaggle.com/kyoshioka47/centernet-baseline-lb-pb-0-07\nThe trained model gets LB 0.07, PB 0.069.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F347059bdaab492b0507096cd77689ea3%2Faaa.JPG?generation=1579840252559961&amp;alt=media)\n\n\nTo get a silver level score, we need to implement a better model (centernet + FPN) to reach LB 0.09~.\nI will open that kernel when the training is done.",
      "votes": null
    },
    {
      "id": "728541",
      "postDate": "01/24/2020 21:58:41",
      "content": "<p>In my case implementing centernet repo's gaussian heatmap method (umich_gaussian) and applying maxpool as nms helped the model improve a lot, and I changed the encoder to resnet18 (simplest one that seemed to work well) as I saw severe overfitting and numerical issues with other models. I'm surprised hourglass worked for some, given how large it is. If you are looking to learn more I suggest to refer to higher ranked approaches that are most similar to your pipeline and improve the scores until you get to a similar score! Kaggle allows late submission for such purposes I guess. I'm also up to share more of my approach if you need details.</p>",
      "rawMarkdown": "In my case implementing centernet repo's gaussian heatmap method (umich_gaussian) and applying maxpool as nms helped the model improve a lot, and I changed the encoder to resnet18 (simplest one that seemed to work well) as I saw severe overfitting and numerical issues with other models. I'm surprised hourglass worked for some, given how large it is. If you are looking to learn more I suggest to refer to higher ranked approaches that are most similar to your pipeline and improve the scores until you get to a similar score! Kaggle allows late submission for such purposes I guess. I'm also up to share more of my approach if you need details.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 725298,
      "author_name": "pp2file",
      "author_url": "",
      "post_date": "01/22/2020 01:28:21",
      "content": "<p>me too, i just splite three part of train dataset. one for train 81% train, 9%val,10%test. and at test   dataset, use public metric, my kernel can get 0.158, but at LB i just get 0.06, so I'm very confuse . </p>",
      "votes": null,
      "replies": [
        {
          "id": 725299,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "01/22/2020 01:29:56",
          "content": "<p>the local CV metric might be different from official one, so I guess that difference is normal.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725306,
          "author_name": "pp2file",
          "author_url": "",
          "post_date": "01/22/2020 01:39:43",
          "content": "<p>for now, I think maybe submission format  has some different or metric has bugs. I will wait some guyes publish theirs kernel make a compare.\nI find some publish kernel used train-dataset cal score just get 0.04, and their LB do not shake up so mutch. \nI use train-dataset by publish metric can get almost 0.22.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 725334,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "01/22/2020 02:06:40",
      "content": "<p>Same here. I have to say that those little details matter and they can only be obtained through past experiences and careful experiments. We've still got a long way to go.</p>",
      "votes": null,
      "replies": [
        {
          "id": 725341,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "01/22/2020 02:23:34",
          "content": "<p>And the unclear metric for this competition is also a bummer, so I could not do many effective experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725433,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/22/2020 05:13:46",
          "content": "<p>And one more thing: the mAP metric is not closely related to the hm loss or regression loss. :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 725488,
      "author_name": "kmizunoster",
      "author_url": "",
      "post_date": "01/22/2020 06:50:22",
      "content": "<p>My final solution is based on <a href=\"https://www.kaggle.com/hocop1/centernet-baseline\">hocop1's kernel</a>.\nI adopted the same model architecture as original and my best solution didn't use ensemble.\nHere are the main differences from the original kernel.</p>\n\n<ul>\n<li>100 epoch learning</li>\n<li>Focal loss for mask whose parameters are default</li>\n<li>Sigmoid threshold: 0.2</li>\n<li>Input resolution for cnn model: 1024x320 -&gt; 2048x640</li>\n<li>Data augmentation\n<ul><li>Pixel level augmentation like yours</li>\n<li>Random scale and crop</li></ul></li>\n<li>Postprocess\n<ul><li>Modify <code>optimize_xy</code> not to use <code>xzy_slope</code></li>\n<li>Drop false positives by using test_masks</li></ul></li>\n</ul>\n\n<p>From the above, I think preprocess and postprocess have huge impacts for bronze-level score.</p>\n\n<p>I hope this will help.</p>",
      "votes": null,
      "replies": [
        {
          "id": 725493,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "01/22/2020 06:58:00",
          "content": "<p>Thank you very much! Could you specify the focal loss parameters used? (especially sigma for Gaussian heatmap since it was not specified in paper). Also, it is really impressive that you did 100 epoch training because my dev loss or mAP never improved after about 10-15 epochs of training. <a href=\"/kmizunoster\">@kmizunoster</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725537,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/22/2020 08:11:52",
          "content": "<p>Hi <a href=\"/kmizunoster\">@kmizunoster</a> does the image size 2048x640 lead to the 100 epochs training? I've tried a larger size but the improvement was really slow with little computation resources so I finally gave it up ...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725566,
          "author_name": "kmizunoster",
          "author_url": "",
          "post_date": "01/22/2020 08:53:04",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> \nI used <a href=\"https://github.com/xingyizhou/CenterNet/blob/8ef87b433529ac8f8bd4f95707f6bc05052c55e9/src/lib/models/losses.py#L42\">CenterNet's official implementation of FocalLoss</a> as it is.\nIn my case (test_size=0.2), my dev loss is saturated around 35 epoch, but a model after 100 epoch training got better LB score than one chosen by best dev loss.</p>\n\n<p><a href=\"/syoya1997\">@syoya1997</a> \nI'm not sure that 100 epochs training is required for the image size 2048x640,\nbut longer epochs training provided better results regerdless of image size in my experience on this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 725572,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "01/22/2020 08:58:59",
          "content": "<p>Thanks. I usually trained until my dev loss not decreasing any more. Wondering why you choose to train for this long? 100 epochs take a lot of time.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 727804,
      "author_name": "kyoshioka47",
      "author_url": "",
      "post_date": "01/24/2020 04:27:51",
      "content": "<p>I think image size, focal loss, longer training scheduling makes a big difference in getting a bronze score.\nI made a kernel to check if we can get LB 0.07 with the baseline model.\n<a href=\"https://www.kaggle.com/kyoshioka47/centernet-baseline-lb-pb-0-07\">https://www.kaggle.com/kyoshioka47/centernet-baseline-lb-pb-0-07</a>\nThe trained model gets LB 0.07, PB 0.069.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F347059bdaab492b0507096cd77689ea3%2Faaa.JPG?generation=1579840252559961&amp;alt=media\" alt=\"\"></p>\n\n<p>To get a silver level score, we need to implement a better model (centernet + FPN) to reach LB 0.09~.\nI will open that kernel when the training is done.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 728541,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "01/24/2020 21:58:41",
      "content": "<p>In my case implementing centernet repo's gaussian heatmap method (umich_gaussian) and applying maxpool as nms helped the model improve a lot, and I changed the encoder to resnet18 (simplest one that seemed to work well) as I saw severe overfitting and numerical issues with other models. I'm surprised hourglass worked for some, given how large it is. If you are looking to learn more I suggest to refer to higher ranked approaches that are most similar to your pipeline and improve the scores until you get to a similar score! Kaggle allows late submission for such purposes I guess. I'm also up to share more of my approach if you need details.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "725294": "Congrats to everyone who gets awards.\n\nMissed the bronze by a bit but learned a lot.\n\nI just don't know how my results are so difference from higher ranking people although I pretty much did the similar things.\n\nMy final solution consists of ensemble of 3 centernet-hourglass104 models with different random seeds and augmentation rate. I did horizontal flip, brightness, gaussian noise, color shift, and random contrast.\n\nI tried focal loss with different gamma parameters, other models such as dla, resnet, densenet etc., and I tried psuedolabeling, but none of them worked.\n\nI would like to hear what you think.",
    "725298": "me too, i just splite three part of train dataset. one for train 81% train, 9%val,10%test. and at test   dataset, use public metric, my kernel can get 0.158, but at LB i just get 0.06, so I'm very confuse .",
    "725299": "the local CV metric might be different from official one, so I guess that difference is normal.",
    "725306": "for now, I think maybe submission format  has some different or metric has bugs. I will wait some guyes publish theirs kernel make a compare.\nI find some publish kernel used train-dataset cal score just get 0.04, and their LB do not shake up so mutch. \nI use train-dataset by publish metric can get almost 0.22.",
    "725334": "Same here. I have to say that those little details matter and they can only be obtained through past experiences and careful experiments. We've still got a long way to go.",
    "725341": "And the unclear metric for this competition is also a bummer, so I could not do many effective experiments.",
    "725433": "And one more thing: the mAP metric is not closely related to the hm loss or regression loss. :)",
    "725488": "My final solution is based on [hocop1's kernel](https://www.kaggle.com/hocop1/centernet-baseline).\nI adopted the same model architecture as original and my best solution didn't use ensemble.\nHere are the main differences from the original kernel.\n\n- 100 epoch learning\n- Focal loss for mask whose parameters are default\n- Sigmoid threshold: 0.2\n- Input resolution for cnn model: 1024x320 -&gt; 2048x640\n- Data augmentation\n  - Pixel level augmentation like yours\n  - Random scale and crop\n- Postprocess\n  - Modify `optimize_xy` not to use `xzy_slope`\n  - Drop false positives by using test_masks\n\nFrom the above, I think preprocess and postprocess have huge impacts for bronze-level score.\n\nI hope this will help.",
    "725493": "Thank you very much! Could you specify the focal loss parameters used? (especially sigma for Gaussian heatmap since it was not specified in paper). Also, it is really impressive that you did 100 epoch training because my dev loss or mAP never improved after about 10-15 epochs of training. @kmizunoster",
    "725537": "Hi @kmizunoster does the image size 2048x640 lead to the 100 epochs training? I've tried a larger size but the improvement was really slow with little computation resources so I finally gave it up ...",
    "725566": "tonychenxyz \nI used [CenterNet's official implementation of FocalLoss](https://github.com/xingyizhou/CenterNet/blob/8ef87b433529ac8f8bd4f95707f6bc05052c55e9/src/lib/models/losses.py#L42) as it is.\nIn my case (test_size=0.2), my dev loss is saturated around 35 epoch, but a model after 100 epoch training got better LB score than one chosen by best dev loss.\n\n@syoya1997 \nI'm not sure that 100 epochs training is required for the image size 2048x640,\nbut longer epochs training provided better results regerdless of image size in my experience on this competition.",
    "725572": "Thanks. I usually trained until my dev loss not decreasing any more. Wondering why you choose to train for this long? 100 epochs take a lot of time.",
    "727804": "I think image size, focal loss, longer training scheduling makes a big difference in getting a bronze score.\nI made a kernel to check if we can get LB 0.07 with the baseline model.\nhttps://www.kaggle.com/kyoshioka47/centernet-baseline-lb-pb-0-07\nThe trained model gets LB 0.07, PB 0.069.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3508221%2F347059bdaab492b0507096cd77689ea3%2Faaa.JPG?generation=1579840252559961&amp;alt=media)\n\n\nTo get a silver level score, we need to implement a better model (centernet + FPN) to reach LB 0.09~.\nI will open that kernel when the training is done.",
    "728541": "In my case implementing centernet repo's gaussian heatmap method (umich_gaussian) and applying maxpool as nms helped the model improve a lot, and I changed the encoder to resnet18 (simplest one that seemed to work well) as I saw severe overfitting and numerical issues with other models. I'm surprised hourglass worked for some, given how large it is. If you are looking to learn more I suggest to refer to higher ranked approaches that are most similar to your pipeline and improve the scores until you get to a similar score! Kaggle allows late submission for such purposes I guess. I'm also up to share more of my approach if you need details."
  },
  "source": "meta"
}