{
  "id": 82393,
  "title": "31st place solution + source code",
  "url": "/competitions/humpback-whale-identification/writeups/4-people-31st-place-solution-source-code",
  "author_name": "",
  "post_date": "2019-03-01T06:36:44.249543900Z",
  "votes": 25,
  "comment_count": 13,
  "views": 0,
  "content": "<p>Source code: <a href=\"https://github.com/suicao/Siamese-Whale-Identification\">https://github.com/suicao/Siamese-Whale-Identification</a> . I'll try to update this repo later but it should be simple enough to follow for now.</p>\n\n<p>Our approach was simple. We took the amazing solution by <a href=\"/martinpiotte\">@martinpiotte</a> and added a few twists:</p>\n\n<ul>\n<li>Using RGB instead of grayscale images.</li>\n<li>Changing the feature extraction CNN with Imagenet trained models. My teammate <a href=\"/iafoss\">@iafoss</a> was able to achieve 0.937 single model with a DenseNet121 encoder. At the last few weeks he also noticed a severe bug in my code where I froze the branch model instead of the feature extractor in the first few epochs, which helped boosting the score.</li>\n<li>Simply training on bigger image size worked, but we didn't have the resources needed to try anything bigger than 512x512.</li>\n<li>Adding TTA made the results worse, we haven't got time to investigate this just yet.</li>\n</ul>\n\n<p>That's it, at the end I made an ensemble of a few high scoring models trained by my teammates using <a href=\"https://www.kaggle.com/matthewa313/ensembling-algorithm-for-average-precision-metric\">this method</a>.  Big thanks to <a href=\"/matthewa313\">@matthewa313</a> .</p>\n\n<p>I didn't even have any accesses to GPUs for the majority of this competition, and one of my other teammate couldn't compete either due to hardware problems, so this is quite unfortunate for us. </p>\n\n<p>Anyway I'm happy with the final standings, congrats everyone, I wish you do <em>whale</em> in the future competitions.</p>",
  "messages": [
    {
      "id": "481201",
      "postDate": "03/01/2019 06:36:44",
      "content": "<p>Source code: <a href=\"https://github.com/suicao/Siamese-Whale-Identification\">https://github.com/suicao/Siamese-Whale-Identification</a> . I'll try to update this repo later but it should be simple enough to follow for now.</p>\n\n<p>Our approach was simple. We took the amazing solution by <a href=\"/martinpiotte\">@martinpiotte</a> and added a few twists:</p>\n\n<ul>\n<li>Using RGB instead of grayscale images.</li>\n<li>Changing the feature extraction CNN with Imagenet trained models. My teammate <a href=\"/iafoss\">@iafoss</a> was able to achieve 0.937 single model with a DenseNet121 encoder. At the last few weeks he also noticed a severe bug in my code where I froze the branch model instead of the feature extractor in the first few epochs, which helped boosting the score.</li>\n<li>Simply training on bigger image size worked, but we didn't have the resources needed to try anything bigger than 512x512.</li>\n<li>Adding TTA made the results worse, we haven't got time to investigate this just yet.</li>\n</ul>\n\n<p>That's it, at the end I made an ensemble of a few high scoring models trained by my teammates using <a href=\"https://www.kaggle.com/matthewa313/ensembling-algorithm-for-average-precision-metric\">this method</a>.  Big thanks to <a href=\"/matthewa313\">@matthewa313</a> .</p>\n\n<p>I didn't even have any accesses to GPUs for the majority of this competition, and one of my other teammate couldn't compete either due to hardware problems, so this is quite unfortunate for us. </p>\n\n<p>Anyway I'm happy with the final standings, congrats everyone, I wish you do <em>whale</em> in the future competitions.</p>",
      "rawMarkdown": "Source code: https://github.com/suicao/Siamese-Whale-Identification . I'll try to update this repo later but it should be simple enough to follow for now.\n\nOur approach was simple. We took the amazing solution by @martinpiotte and added a few twists:\n\n- Using RGB instead of grayscale images.\n- Changing the feature extraction CNN with Imagenet trained models. My teammate @iafoss was able to achieve 0.937 single model with a DenseNet121 encoder. At the last few weeks he also noticed a severe bug in my code where I froze the branch model instead of the feature extractor in the first few epochs, which helped boosting the score.\n- Simply training on bigger image size worked, but we didn't have the resources needed to try anything bigger than 512x512.\n- Adding TTA made the results worse, we haven't got time to investigate this just yet.\n\nThat's it, at the end I made an ensemble of a few high scoring models trained by my teammates using [this method][1].  Big thanks to @matthewa313 .\n\nI didn't even have any accesses to GPUs for the majority of this competition, and one of my other teammate couldn't compete either due to hardware problems, so this is quite unfortunate for us. \n\nAnyway I'm happy with the final standings, congrats everyone, I wish you do *whale* in the future competitions.\n\n\n  [1]: https://www.kaggle.com/matthewa313/ensembling-algorithm-for-average-precision-metric",
      "votes": null
    },
    {
      "id": "481206",
      "postDate": "03/01/2019 06:46:59",
      "content": "<p>Congratulations and thanks for sharing the code <a href=\"/suicaokhoailang\">@suicaokhoailang</a>...</p>",
      "rawMarkdown": "Congratulations and thanks for sharing the code @suicaokhoailang...",
      "votes": null
    },
    {
      "id": "481209",
      "postDate": "03/01/2019 06:51:49",
      "content": "<p>Congrats Khoi ^^</p>",
      "rawMarkdown": "Congrats Khoi ^^",
      "votes": null
    },
    {
      "id": "481218",
      "postDate": "03/01/2019 07:07:04",
      "content": "<p>Congrats. Thank for sharing !</p>",
      "rawMarkdown": "Congrats. Thank for sharing !",
      "votes": null
    },
    {
      "id": "481225",
      "postDate": "03/01/2019 07:17:28",
      "content": "<p>Congratulations! Thank you for sharing your solution. </p>",
      "rawMarkdown": "Congratulations! Thank you for sharing your solution.",
      "votes": null
    },
    {
      "id": "481227",
      "postDate": "03/01/2019 07:21:39",
      "content": "<p>awesome result with little access to GPUs</p>",
      "rawMarkdown": "awesome result with little access to GPUs",
      "votes": null
    },
    {
      "id": "481263",
      "postDate": "03/01/2019 07:48:34",
      "content": "<p>One thing to add, with this code I also tried DenseNet169 with concatenative pooling and gradient accumulation and was able to reach ~0.930 after 80 epochs on 384x384 + 30 epochs on 512x512. DenseNet121 after training for the same number of epochs gave only ~0.910.\nTraining for longer boosted the score of DenseNet121 model to 0.938 on public LB. Unfortunately, we didn't have time to run DenseNet169 model for more epochs, but I expect it could reach ~0.950 single model prediction.</p>",
      "rawMarkdown": "One thing to add, with this code I also tried DenseNet169 with concatenative pooling and gradient accumulation and was able to reach ~0.930 after 80 epochs on 384x384 + 30 epochs on 512x512. DenseNet121 after training for the same number of epochs gave only ~0.910.\nTraining for longer boosted the score of DenseNet121 model to 0.938 on public LB. Unfortunately, we didn't have time to run DenseNet169 model for more epochs, but I expect it could reach ~0.950 single model prediction.",
      "votes": null
    },
    {
      "id": "481307",
      "postDate": "03/01/2019 08:27:54",
      "content": "<p>Congrats <a href=\"/suicaokhoailang\">@suicaokhoailang</a> and team. Thanks for sharing.</p>",
      "rawMarkdown": "Congrats @suicaokhoailang and team. Thanks for sharing.",
      "votes": null
    },
    {
      "id": "481574",
      "postDate": "03/01/2019 15:18:19",
      "content": "<p>One thing to add too\nFirst, thanks for Khoi Iafoss and Ocean’s efforts. And thanks for Alex and Lu Yang’s advice.\nWe are focus on Siamese network, when the competition was going to over, I tried to use the classification methods of SE-ResNeXt101 and 152 for 150 epochs, the new_whale didn’t be the part of trainset. But because of the mistake in my code, the loss wasn’t decrease after 10 epoch, so I only get 0.3xx on LB. I will go and see the first place’s se-resnext solution, and also check my code.\nAll in all, if we have more time and more submissions chances, we can do better.</p>",
      "rawMarkdown": "One thing to add too\nFirst, thanks for Khoi Iafoss and Ocean’s efforts. And thanks for Alex and Lu Yang’s advice.\nWe are focus on Siamese network, when the competition was going to over, I tried to use the classification methods of SE-ResNeXt101 and 152 for 150 epochs, the new_whale didn’t be the part of trainset. But because of the mistake in my code, the loss wasn’t decrease after 10 epoch, so I only get 0.3xx on LB. I will go and see the first place’s se-resnext solution, and also check my code.\nAll in all, if we have more time and more submissions chances, we can do better.",
      "votes": null
    },
    {
      "id": "481826",
      "postDate": "03/01/2019 22:17:41",
      "content": "<p>Nice approach, that Siamese architecture are pretty interesting. congrats for the top31</p>",
      "rawMarkdown": "Nice approach, that Siamese architecture are pretty interesting. congrats for the top31",
      "votes": null
    },
    {
      "id": "481942",
      "postDate": "03/02/2019 04:43:54",
      "content": "<p>TTA is useful。Using fliplr, but flipud is useless,Since I tried the TTA at the last time, it is not clear why the fliplr is useful, and it may be related to the image characteristics of the whale.</p>",
      "rawMarkdown": "TTA is useful。Using fliplr, but flipud is useless,Since I tried the TTA at the last time, it is not clear why the fliplr is useful, and it may be related to the image characteristics of the whale.",
      "votes": null
    },
    {
      "id": "481945",
      "postDate": "03/02/2019 04:51:25",
      "content": "<p>We tried TTA for two times, but it didn't work well, our score droped instead</p>",
      "rawMarkdown": "We tried TTA for two times, but it didn't work well, our score droped instead",
      "votes": null
    },
    {
      "id": "482664",
      "postDate": "03/03/2019 13:10:48",
      "content": "<p>Congratulations <a href=\"/iafoss\">@iafoss</a> and to your team mates. You deserve it after all that hard work!\nYour trials that you are referring is with the fastai code or the keras code (Martin's approach)?\nI am still curious why your fastai code could not get to &gt;0.84</p>",
      "rawMarkdown": "Congratulations @iafoss and to your team mates. You deserve it after all that hard work!\nYour trials that you are referring is with the fastai code or the keras code (Martin's approach)?\nI am still curious why your fastai code could not get to &gt;0.84",
      "votes": null
    },
    {
      "id": "482789",
      "postDate": "03/03/2019 16:42:40",
      "content": "<p>I think the problem is that I didn't use LAP there. At later stage of training, probably, optimization of individual pairs becomes important. Without LAP,  I expect switching from batch all to batch hard could be useful, but in this case batches should contain a lot of images, ~100+, and the size of batches should increase with training (similar to decrease of the noise amplitude in LAP). So batch hard approach is not very useful for large image resolution. Hard negative sampling could be another alternative, but likely it should not be combined with batch all loss. In this case at every n epochs N hardest negative samples is selected for each image, and during training the model randomly select one of those N images. Similar with noise, N should decrease with training. Another thing is the speed of nose reduction in LAP which should be carefully adjusted: the training with LAP may be quite slow if noise is not reduced enough, while if it is reduced too much, the model would not converge. So LAP is quite interesting approach, but should be studied in more details and compared more carefully with traditional hard negative mining.</p>",
      "rawMarkdown": "I think the problem is that I didn't use LAP there. At later stage of training, probably, optimization of individual pairs becomes important. Without LAP,  I expect switching from batch all to batch hard could be useful, but in this case batches should contain a lot of images, ~100+, and the size of batches should increase with training (similar to decrease of the noise amplitude in LAP). So batch hard approach is not very useful for large image resolution. Hard negative sampling could be another alternative, but likely it should not be combined with batch all loss. In this case at every n epochs N hardest negative samples is selected for each image, and during training the model randomly select one of those N images. Similar with noise, N should decrease with training. Another thing is the speed of nose reduction in LAP which should be carefully adjusted: the training with LAP may be quite slow if noise is not reduced enough, while if it is reduced too much, the model would not converge. So LAP is quite interesting approach, but should be studied in more details and compared more carefully with traditional hard negative mining.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 481206,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "03/01/2019 06:46:59",
      "content": "<p>Congratulations and thanks for sharing the code <a href=\"/suicaokhoailang\">@suicaokhoailang</a>...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481209,
      "author_name": "truocpham",
      "author_url": "",
      "post_date": "03/01/2019 06:51:49",
      "content": "<p>Congrats Khoi ^^</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481218,
      "author_name": "backaggle",
      "author_url": "",
      "post_date": "03/01/2019 07:07:04",
      "content": "<p>Congrats. Thank for sharing !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481225,
      "author_name": "devilears",
      "author_url": "",
      "post_date": "03/01/2019 07:17:28",
      "content": "<p>Congratulations! Thank you for sharing your solution. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481227,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "03/01/2019 07:21:39",
      "content": "<p>awesome result with little access to GPUs</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481263,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "03/01/2019 07:48:34",
      "content": "<p>One thing to add, with this code I also tried DenseNet169 with concatenative pooling and gradient accumulation and was able to reach ~0.930 after 80 epochs on 384x384 + 30 epochs on 512x512. DenseNet121 after training for the same number of epochs gave only ~0.910.\nTraining for longer boosted the score of DenseNet121 model to 0.938 on public LB. Unfortunately, we didn't have time to run DenseNet169 model for more epochs, but I expect it could reach ~0.950 single model prediction.</p>",
      "votes": null,
      "replies": [
        {
          "id": 482664,
          "author_name": "hwasiti",
          "author_url": "",
          "post_date": "03/03/2019 13:10:48",
          "content": "<p>Congratulations <a href=\"/iafoss\">@iafoss</a> and to your team mates. You deserve it after all that hard work!\nYour trials that you are referring is with the fastai code or the keras code (Martin's approach)?\nI am still curious why your fastai code could not get to &gt;0.84</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 482789,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "03/03/2019 16:42:40",
          "content": "<p>I think the problem is that I didn't use LAP there. At later stage of training, probably, optimization of individual pairs becomes important. Without LAP,  I expect switching from batch all to batch hard could be useful, but in this case batches should contain a lot of images, ~100+, and the size of batches should increase with training (similar to decrease of the noise amplitude in LAP). So batch hard approach is not very useful for large image resolution. Hard negative sampling could be another alternative, but likely it should not be combined with batch all loss. In this case at every n epochs N hardest negative samples is selected for each image, and during training the model randomly select one of those N images. Similar with noise, N should decrease with training. Another thing is the speed of nose reduction in LAP which should be carefully adjusted: the training with LAP may be quite slow if noise is not reduced enough, while if it is reduced too much, the model would not converge. So LAP is quite interesting approach, but should be studied in more details and compared more carefully with traditional hard negative mining.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 481307,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "03/01/2019 08:27:54",
      "content": "<p>Congrats <a href=\"/suicaokhoailang\">@suicaokhoailang</a> and team. Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481574,
      "author_name": "",
      "author_url": "",
      "post_date": "03/01/2019 15:18:19",
      "content": "<p>One thing to add too\nFirst, thanks for Khoi Iafoss and Ocean’s efforts. And thanks for Alex and Lu Yang’s advice.\nWe are focus on Siamese network, when the competition was going to over, I tried to use the classification methods of SE-ResNeXt101 and 152 for 150 epochs, the new_whale didn’t be the part of trainset. But because of the mistake in my code, the loss wasn’t decrease after 10 epoch, so I only get 0.3xx on LB. I will go and see the first place’s se-resnext solution, and also check my code.\nAll in all, if we have more time and more submissions chances, we can do better.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481826,
      "author_name": "titericz",
      "author_url": "",
      "post_date": "03/01/2019 22:17:41",
      "content": "<p>Nice approach, that Siamese architecture are pretty interesting. congrats for the top31</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 481942,
      "author_name": "zhangeng",
      "author_url": "",
      "post_date": "03/02/2019 04:43:54",
      "content": "<p>TTA is useful。Using fliplr, but flipud is useless,Since I tried the TTA at the last time, it is not clear why the fliplr is useful, and it may be related to the image characteristics of the whale.</p>",
      "votes": null,
      "replies": [
        {
          "id": 481945,
          "author_name": "",
          "author_url": "",
          "post_date": "03/02/2019 04:51:25",
          "content": "<p>We tried TTA for two times, but it didn't work well, our score droped instead</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "481201": "Source code: https://github.com/suicao/Siamese-Whale-Identification . I'll try to update this repo later but it should be simple enough to follow for now.\n\nOur approach was simple. We took the amazing solution by @martinpiotte and added a few twists:\n\n- Using RGB instead of grayscale images.\n- Changing the feature extraction CNN with Imagenet trained models. My teammate @iafoss was able to achieve 0.937 single model with a DenseNet121 encoder. At the last few weeks he also noticed a severe bug in my code where I froze the branch model instead of the feature extractor in the first few epochs, which helped boosting the score.\n- Simply training on bigger image size worked, but we didn't have the resources needed to try anything bigger than 512x512.\n- Adding TTA made the results worse, we haven't got time to investigate this just yet.\n\nThat's it, at the end I made an ensemble of a few high scoring models trained by my teammates using [this method][1].  Big thanks to @matthewa313 .\n\nI didn't even have any accesses to GPUs for the majority of this competition, and one of my other teammate couldn't compete either due to hardware problems, so this is quite unfortunate for us. \n\nAnyway I'm happy with the final standings, congrats everyone, I wish you do *whale* in the future competitions.\n\n\n  [1]: https://www.kaggle.com/matthewa313/ensembling-algorithm-for-average-precision-metric",
    "481206": "Congratulations and thanks for sharing the code @suicaokhoailang...",
    "481209": "Congrats Khoi ^^",
    "481218": "Congrats. Thank for sharing !",
    "481225": "Congratulations! Thank you for sharing your solution.",
    "481227": "awesome result with little access to GPUs",
    "481263": "One thing to add, with this code I also tried DenseNet169 with concatenative pooling and gradient accumulation and was able to reach ~0.930 after 80 epochs on 384x384 + 30 epochs on 512x512. DenseNet121 after training for the same number of epochs gave only ~0.910.\nTraining for longer boosted the score of DenseNet121 model to 0.938 on public LB. Unfortunately, we didn't have time to run DenseNet169 model for more epochs, but I expect it could reach ~0.950 single model prediction.",
    "481307": "Congrats @suicaokhoailang and team. Thanks for sharing.",
    "481574": "One thing to add too\nFirst, thanks for Khoi Iafoss and Ocean’s efforts. And thanks for Alex and Lu Yang’s advice.\nWe are focus on Siamese network, when the competition was going to over, I tried to use the classification methods of SE-ResNeXt101 and 152 for 150 epochs, the new_whale didn’t be the part of trainset. But because of the mistake in my code, the loss wasn’t decrease after 10 epoch, so I only get 0.3xx on LB. I will go and see the first place’s se-resnext solution, and also check my code.\nAll in all, if we have more time and more submissions chances, we can do better.",
    "481826": "Nice approach, that Siamese architecture are pretty interesting. congrats for the top31",
    "481942": "TTA is useful。Using fliplr, but flipud is useless,Since I tried the TTA at the last time, it is not clear why the fliplr is useful, and it may be related to the image characteristics of the whale.",
    "481945": "We tried TTA for two times, but it didn't work well, our score droped instead",
    "482664": "Congratulations @iafoss and to your team mates. You deserve it after all that hard work!\nYour trials that you are referring is with the fastai code or the keras code (Martin's approach)?\nI am still curious why your fastai code could not get to &gt;0.84",
    "482789": "I think the problem is that I didn't use LAP there. At later stage of training, probably, optimization of individual pairs becomes important. Without LAP,  I expect switching from batch all to batch hard could be useful, but in this case batches should contain a lot of images, ~100+, and the size of batches should increase with training (similar to decrease of the noise amplitude in LAP). So batch hard approach is not very useful for large image resolution. Hard negative sampling could be another alternative, but likely it should not be combined with batch all loss. In this case at every n epochs N hardest negative samples is selected for each image, and during training the model randomly select one of those N images. Similar with noise, N should decrease with training. Another thing is the speed of nose reduction in LAP which should be carefully adjusted: the training with LAP may be quite slow if noise is not reduced enough, while if it is reduced too much, the model would not converge. So LAP is quite interesting approach, but should be studied in more details and compared more carefully with traditional hard negative mining."
  },
  "source": "meta"
}