{
  "id": 77300,
  "title": "part of 4th place solution: GAPNet & dual loss ResNet",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/77300",
  "author_name": "Dieter",
  "post_date": "2019-01-11T08:45:10.832000",
  "votes": 81,
  "comment_count": 24,
  "views": 0,
  "content": "<p>First a big thanks to all my team-members. They all made this competition an awesome experience. I also want to thank @brian, @heng and @lafoss for their fruitful input. </p>\n\n<p>Since major ingredient of our solution was diversity and in the end we used around 25 different models, I want to give a separate few notes on two of my main contributions. </p>\n\n<p>I want to split up my notes into strategy, architecture as I see each equally important.</p>\n\n<p><strong>Strategy:</strong></p>\n\n<p><a href=\"/tunguz\">@tunguz</a> <a href=\"/sasrdw\">@sasrdw</a> and myself teamed-up quite early which enabled us to work in different directions right from the beginning. We always had diversity of our models in mind. So I concentrated on models that seem a bit different. Diversity to my team-members was also the main reason I sticked to keras, although in my opinion pytorch would have been more suitable for this competition due to its flexibility. We also used different cross validation schemes for the sake of diversity.</p>\n\n<p>After some trouble in the beginning for getting the cross-validation right, I started exploring different architectures as posted by <a href=\"/hengck23\">@hengck23</a> . I found it quite efficient to only use 256x256 RGB images in the beginning because it allows for high iteration of different ideas.</p>\n\n<p><strong>Architectures:</strong></p>\n\n<p><em>GAPNet</em></p>\n\n<p>Immediately,  reading the GAPNet paper, I had the idea to change the illustrated architecture to use a pretrained backbone instead. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/a95f150c7153a17538b074def2255e21/GAP.png\" alt=\"original GAPNet\"></p>\n\n<p>I think one important advantage of the GAPNet architecture is its ability for multiscale. So I tried different backbones and ended up with ResNet18, which also enabled to use a batchsize of 32 on a GTX1080Ti. I also saw minor improvements adding SE-Blocks before the Average Pooling layers with nearly no computational cost, so I added those. I saw no improvement in using RGBY images. I used a weighted bce and f1 loss and a cosine annealing lr schedule and only trained for 20 epochs. After applying our thresholding method to the predictions GAPNet trained on 512x512 RGB images also using the HPA external data a single 5-fold model was able to achieve 0.602 Public LB. I also experimented with different internal/external data proportions, RGBY and 512cropping from 1024 images so I had 4 5-fold GAPNet models which I could ensemble resulted in LB  0.609</p>\n\n<p><em>Dual Loss ResNet</em></p>\n\n<p>Following another post from <a href=\"/hengck23\">@hengck23</a> I implemented a ResNet34 with a dual loss:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png\" alt=\"enter image description here\"></p>\n\n<p>Additionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".\nThe additional supervised attention added quite some benefit to regularization, and the computational cost was bearable. I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618</p>\n\n<p>I guess <a href=\"/tunguz\">@tunguz</a> will write an overall summary where he explains how my models were then incorporated into our overall ensemble.</p>",
  "messages": [
    {
      "id": 454209,
      "postDate": "2019-01-11T08:45:10.833Z",
      "content": "<p>First a big thanks to all my team-members. They all made this competition an awesome experience. I also want to thank @brian, @heng and @lafoss for their fruitful input. </p>\n\n<p>Since major ingredient of our solution was diversity and in the end we used around 25 different models, I want to give a separate few notes on two of my main contributions. </p>\n\n<p>I want to split up my notes into strategy, architecture as I see each equally important.</p>\n\n<p><strong>Strategy:</strong></p>\n\n<p><a href=\"/tunguz\">@tunguz</a> <a href=\"/sasrdw\">@sasrdw</a> and myself teamed-up quite early which enabled us to work in different directions right from the beginning. We always had diversity of our models in mind. So I concentrated on models that seem a bit different. Diversity to my team-members was also the main reason I sticked to keras, although in my opinion pytorch would have been more suitable for this competition due to its flexibility. We also used different cross validation schemes for the sake of diversity.</p>\n\n<p>After some trouble in the beginning for getting the cross-validation right, I started exploring different architectures as posted by <a href=\"/hengck23\">@hengck23</a> . I found it quite efficient to only use 256x256 RGB images in the beginning because it allows for high iteration of different ideas.</p>\n\n<p><strong>Architectures:</strong></p>\n\n<p><em>GAPNet</em></p>\n\n<p>Immediately,  reading the GAPNet paper, I had the idea to change the illustrated architecture to use a pretrained backbone instead. </p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/a95f150c7153a17538b074def2255e21/GAP.png\" alt=\"original GAPNet\"></p>\n\n<p>I think one important advantage of the GAPNet architecture is its ability for multiscale. So I tried different backbones and ended up with ResNet18, which also enabled to use a batchsize of 32 on a GTX1080Ti. I also saw minor improvements adding SE-Blocks before the Average Pooling layers with nearly no computational cost, so I added those. I saw no improvement in using RGBY images. I used a weighted bce and f1 loss and a cosine annealing lr schedule and only trained for 20 epochs. After applying our thresholding method to the predictions GAPNet trained on 512x512 RGB images also using the HPA external data a single 5-fold model was able to achieve 0.602 Public LB. I also experimented with different internal/external data proportions, RGBY and 512cropping from 1024 images so I had 4 5-fold GAPNet models which I could ensemble resulted in LB  0.609</p>\n\n<p><em>Dual Loss ResNet</em></p>\n\n<p>Following another post from <a href=\"/hengck23\">@hengck23</a> I implemented a ResNet34 with a dual loss:</p>\n\n<p><img src=\"https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png\" alt=\"enter image description here\"></p>\n\n<p>Additionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".\nThe additional supervised attention added quite some benefit to regularization, and the computational cost was bearable. I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618</p>\n\n<p>I guess <a href=\"/tunguz\">@tunguz</a> will write an overall summary where he explains how my models were then incorporated into our overall ensemble.</p>",
      "rawMarkdown": "First a big thanks to all my team-members. They all made this competition an awesome experience. I also want to thank @brian, @heng and @lafoss for their fruitful input. \n\nSince major ingredient of our solution was diversity and in the end we used around 25 different models, I want to give a separate few notes on two of my main contributions. \n\nI want to split up my notes into strategy, architecture as I see each equally important.\n\n**Strategy:**\n\n@tunguz @sasrdw and myself teamed-up quite early which enabled us to work in different directions right from the beginning. We always had diversity of our models in mind. So I concentrated on models that seem a bit different. Diversity to my team-members was also the main reason I sticked to keras, although in my opinion pytorch would have been more suitable for this competition due to its flexibility. We also used different cross validation schemes for the sake of diversity.\n\nAfter some trouble in the beginning for getting the cross-validation right, I started exploring different architectures as posted by @hengck23 . I found it quite efficient to only use 256x256 RGB images in the beginning because it allows for high iteration of different ideas.\n\n**Architectures:**\n\n*GAPNet*\n\nImmediately,  reading the GAPNet paper, I had the idea to change the illustrated architecture to use a pretrained backbone instead. \n\n![original GAPNet][1]\n\nI think one important advantage of the GAPNet architecture is its ability for multiscale. So I tried different backbones and ended up with ResNet18, which also enabled to use a batchsize of 32 on a GTX1080Ti. I also saw minor improvements adding SE-Blocks before the Average Pooling layers with nearly no computational cost, so I added those. I saw no improvement in using RGBY images. I used a weighted bce and f1 loss and a cosine annealing lr schedule and only trained for 20 epochs. After applying our thresholding method to the predictions GAPNet trained on 512x512 RGB images also using the HPA external data a single 5-fold model was able to achieve 0.602 Public LB. I also experimented with different internal/external data proportions, RGBY and 512cropping from 1024 images so I had 4 5-fold GAPNet models which I could ensemble resulted in LB  0.609\n\n*Dual Loss ResNet*\n\nFollowing another post from @hengck23 I implemented a ResNet34 with a dual loss:\n\n![enter image description here][2]\n\nAdditionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".\nThe additional supervised attention added quite some benefit to regularization, and the computational cost was bearable. I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618\n\nI guess @tunguz will write an overall summary where he explains how my models were then incorporated into our overall ensemble.\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/a95f150c7153a17538b074def2255e21/GAP.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png",
      "votes": 81
    },
    {
      "id": 454299,
      "postDate": "2019-01-11T11:39:58.637Z",
      "content": "<p>congratulations! are you going to share the code ?</p>",
      "rawMarkdown": "congratulations! are you going to share the code ?\n",
      "votes": 1
    },
    {
      "id": 454284,
      "postDate": "2019-01-11T11:11:27.993Z",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> - congratulations on your result! I basically did an almost identical thing re. butchering a resnet18 into a GAPNet to use the pre-trained weights. Which layers did you attach the GAP blocks to (this is the bit that slows things down)?</p>\n\n<p>Also, when you say you used different CV schemes - how did you combine overall predictions and validate - some form of voting/mean function? </p>",
      "rawMarkdown": "@christofhenkel - congratulations on your result! I basically did an almost identical thing re. butchering a resnet18 into a GAPNet to use the pre-trained weights. Which layers did you attach the GAP blocks to (this is the bit that slows things down)?\n\nAlso, when you say you used different CV schemes - how did you combine overall predictions and validate - some form of voting/mean function? ",
      "votes": 1,
      "replies": [
        {
          "id": 454297,
          "postDate": "2019-01-11T11:33:21.073Z",
          "content": "<p>I attached the GAP blocks to the last layers with 16,32,64,128,256 filters.</p>\n\n<p>Even if you use different cross validation schemes (say 5-fold vs 4-fold) you can still stack models using the oof predictions. Nevertheless we used a voting scheme as very last ensembling layer as we did not have oof predictions for all our models and it worked quite well</p>",
          "rawMarkdown": "I attached the GAP blocks to the last layers with 16,32,64,128,256 filters.\n\nEven if you use different cross validation schemes (say 5-fold vs 4-fold) you can still stack models using the oof predictions. Nevertheless we used a voting scheme as very last ensembling layer as we did not have oof predictions for all our models and it worked quite well",
          "votes": 2
        },
        {
          "id": 454312,
          "postDate": "2019-01-11T12:20:38.177Z",
          "content": "<p>Thanks, and sorry, I wasn't clear. Different CV schemes could imply using different images for the validation sets (obviously usually they'll be some overlap) i.e. one kaggle data, one kaggle + hpa, one just hpa etc. Guess I'm really asking if you all used the same set of images overall and did you exclude any due to concerns around LB leakage, similarity, uncertain labels etc...</p>",
          "rawMarkdown": "Thanks, and sorry, I wasn't clear. Different CV schemes could imply using different images for the validation sets (obviously usually they'll be some overlap) i.e. one kaggle data, one kaggle + hpa, one just hpa etc. Guess I'm really asking if you all used the same set of images overall and did you exclude any due to concerns around LB leakage, similarity, uncertain labels etc..."
        },
        {
          "id": 454451,
          "postDate": "2019-01-11T16:48:43.633Z",
          "content": "<p>What was the CV scheme that you guys used? Was it just a random split?</p>",
          "rawMarkdown": "What was the CV scheme that you guys used? Was it just a random split?"
        },
        {
          "id": 454484,
          "postDate": "2019-01-11T17:52:19.027Z",
          "content": "<p>I mainly used stratified multilabel split as posted in one of the discussions. But we (me included) also used a 4-fold cv based on clusters of images with similar statistics. For the last day we also ran last minute 2-fold models. </p>",
          "rawMarkdown": "I mainly used stratified multilabel split as posted in one of the discussions. But we (me included) also used a 4-fold cv based on clusters of images with similar statistics. For the last day we also ran last minute 2-fold models. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 456006,
      "postDate": "2019-01-15T00:09:51.427Z",
      "content": "<p>Congratulations! Always fascinating to read top solutions for Kaggle competitions. </p>\n\n<p>Is this the GAPNet paper that you mentioned?</p>\n\n<p>Link (\"Human-level Protein Localization with Convolutional Neural Networks\"):\n<a href=\"https://openreview.net/forum?id=ryl5khRcKm\">https://openreview.net/forum?id=ryl5khRcKm</a></p>",
      "rawMarkdown": "Congratulations! Always fascinating to read top solutions for Kaggle competitions. \n\nIs this the GAPNet paper that you mentioned?\n\nLink (\"Human-level Protein Localization with Convolutional Neural Networks\"):\nhttps://openreview.net/forum?id=ryl5khRcKm",
      "replies": [
        {
          "id": 456101,
          "postDate": "2019-01-15T06:36:21.387Z",
          "content": "<p>thanks, yes thats the paper</p>",
          "rawMarkdown": "thanks, yes thats the paper",
          "votes": 1
        },
        {
          "id": 456222,
          "postDate": "2019-01-15T11:26:02.457Z",
          "content": "<p>👍</p>",
          "rawMarkdown": "👍"
        }
      ]
    },
    {
      "id": 455602,
      "postDate": "2019-01-14T08:35:54.683Z",
      "content": "<p>Thanks for sharing! \n\"I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618\"\nI saw lb score is 0.65331 ? </p>",
      "rawMarkdown": "Thanks for sharing! \n\"I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618\"\nI saw lb score is 0.65331 ? ",
      "replies": [
        {
          "id": 455617,
          "postDate": "2019-01-14T08:56:19.807Z",
          "content": "<p>Here I only address a (small) part of our solution. Since our whole solution consists of many models and tricks we are not finished writing the complete summary yet</p>",
          "rawMarkdown": "Here I only address a (small) part of our solution. Since our whole solution consists of many models and tricks we are not finished writing the complete summary yet"
        }
      ]
    },
    {
      "id": 454690,
      "postDate": "2019-01-12T02:47:51.587Z",
      "content": "<p>Very Great work! </p>",
      "rawMarkdown": "Very Great work! "
    },
    {
      "id": 454428,
      "postDate": "2019-01-11T16:18:37.320Z",
      "content": "<p>Hi <a href=\"/christofhenkel\">@christofhenkel</a> and thanks for sharing your ideas. I was wondering how you implemented the autoencoder. Could you explain it a little bit further?</p>",
      "rawMarkdown": "Hi @christofhenkel and thanks for sharing your ideas. I was wondering how you implemented the autoencoder. Could you explain it a little bit further?",
      "replies": [
        {
          "id": 454486,
          "postDate": "2019-01-11T17:58:16Z",
          "content": "<p>Congratulations to your result Sven :D</p>\n\n<p>You mean for the dual loss ResNet ? I actually did not implement an auto encoder (I wanted but it was really computational expensive and a quick test showed, at least in my case, that it was not promising). Additionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".</p>",
          "rawMarkdown": "Congratulations to your result Sven :D\n\nYou mean for the dual loss ResNet ? I actually did not implement an auto encoder (I wanted but it was really computational expensive and a quick test showed, at least in my case, that it was not promising). Additionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".",
          "votes": 1
        },
        {
          "id": 455465,
          "postDate": "2019-01-14T00:50:02.880Z",
          "content": "<p>Thanks, man!</p>",
          "rawMarkdown": "Thanks, man!"
        }
      ]
    },
    {
      "id": 454410,
      "postDate": "2019-01-11T15:44:25.940Z",
      "content": "<p>Excellent work! Truly inspiring!</p>",
      "rawMarkdown": "Excellent work! Truly inspiring!"
    },
    {
      "id": 454345,
      "postDate": "2019-01-11T13:33:28.943Z",
      "content": "<p>congrats to you and your team, your tip on Gapnet tweak is indeed very helpful..</p>",
      "rawMarkdown": "congrats to you and your team, your tip on Gapnet tweak is indeed very helpful.."
    },
    {
      "id": 454225,
      "postDate": "2019-01-11T09:15:41.560Z",
      "content": "<p>Congratulations! Wonderful work! </p>",
      "rawMarkdown": "Congratulations! Wonderful work! "
    },
    {
      "id": 454224,
      "postDate": "2019-01-11T09:15:16.280Z",
      "content": "<p>Thank you for the insight! Excellent work.</p>",
      "rawMarkdown": "Thank you for the insight! Excellent work."
    },
    {
      "id": 454214,
      "postDate": "2019-01-11T08:56:39.733Z",
      "content": "<p>Great job. congratulations!</p>",
      "rawMarkdown": "Great job. congratulations!"
    },
    {
      "id": 909766,
      "postDate": "2020-06-30T20:07:35.860Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 454257,
      "postDate": "2019-01-11T10:10:55.190Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 454406,
      "postDate": "2019-01-11T15:39:47.743Z",
      "content": "<p>Great work! Thanks for sharing!</p>",
      "rawMarkdown": "Great work! Thanks for sharing!"
    },
    {
      "id": 454217,
      "postDate": "2019-01-11T09:00:37.500Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!"
    }
  ],
  "comments": [
    {
      "id": 454299,
      "author_name": "soywu",
      "author_url": "",
      "post_date": "2019-01-11T11:39:58.637000",
      "content": "<p>congratulations! are you going to share the code ?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 454284,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2019-01-11T11:11:27.993000",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> - congratulations on your result! I basically did an almost identical thing re. butchering a resnet18 into a GAPNet to use the pre-trained weights. Which layers did you attach the GAP blocks to (this is the bit that slows things down)?</p>\n\n<p>Also, when you say you used different CV schemes - how did you combine overall predictions and validate - some form of voting/mean function? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 454297,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-01-11T11:33:21.073000",
          "content": "<p>I attached the GAP blocks to the last layers with 16,32,64,128,256 filters.</p>\n\n<p>Even if you use different cross validation schemes (say 5-fold vs 4-fold) you can still stack models using the oof predictions. Nevertheless we used a voting scheme as very last ensembling layer as we did not have oof predictions for all our models and it worked quite well</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 454312,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2019-01-11T12:20:38.177000",
          "content": "<p>Thanks, and sorry, I wasn't clear. Different CV schemes could imply using different images for the validation sets (obviously usually they'll be some overlap) i.e. one kaggle data, one kaggle + hpa, one just hpa etc. Guess I'm really asking if you all used the same set of images overall and did you exclude any due to concerns around LB leakage, similarity, uncertain labels etc...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 454451,
          "author_name": "Kevin Lu",
          "author_url": "",
          "post_date": "2019-01-11T16:48:43.633000",
          "content": "<p>What was the CV scheme that you guys used? Was it just a random split?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 454484,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-01-11T17:52:19.027000",
          "content": "<p>I mainly used stratified multilabel split as posted in one of the discussions. But we (me included) also used a 4-fold cv based on clusters of images with similar statistics. For the last day we also ran last minute 2-fold models. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 456006,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2019-01-15T00:09:51.427000",
      "content": "<p>Congratulations! Always fascinating to read top solutions for Kaggle competitions. </p>\n\n<p>Is this the GAPNet paper that you mentioned?</p>\n\n<p>Link (\"Human-level Protein Localization with Convolutional Neural Networks\"):\n<a href=\"https://openreview.net/forum?id=ryl5khRcKm\">https://openreview.net/forum?id=ryl5khRcKm</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 456101,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-01-15T06:36:21.387000",
          "content": "<p>thanks, yes thats the paper</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 456222,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-01-15T11:26:02.457000",
          "content": "<p>👍</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 455602,
      "author_name": "zhengjie",
      "author_url": "",
      "post_date": "2019-01-14T08:35:54.683000",
      "content": "<p>Thanks for sharing! \n\"I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618\"\nI saw lb score is 0.65331 ? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 455617,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-01-14T08:56:19.807000",
          "content": "<p>Here I only address a (small) part of our solution. Since our whole solution consists of many models and tricks we are not finished writing the complete summary yet</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454690,
      "author_name": "LbyG",
      "author_url": "",
      "post_date": "2019-01-12T02:47:51.587000",
      "content": "<p>Very Great work! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454428,
      "author_name": "Sven",
      "author_url": "",
      "post_date": "2019-01-11T16:18:37.320000",
      "content": "<p>Hi <a href=\"/christofhenkel\">@christofhenkel</a> and thanks for sharing your ideas. I was wondering how you implemented the autoencoder. Could you explain it a little bit further?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 454486,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2019-01-11T17:58:16",
          "content": "<p>Congratulations to your result Sven :D</p>\n\n<p>You mean for the dual loss ResNet ? I actually did not implement an auto encoder (I wanted but it was really computational expensive and a quick test showed, at least in my case, that it was not promising). Additionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 455465,
          "author_name": "Sven",
          "author_url": "",
          "post_date": "2019-01-14T00:50:02.880000",
          "content": "<p>Thanks, man!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 454410,
      "author_name": "Ashish Lal",
      "author_url": "",
      "post_date": "2019-01-11T15:44:25.940000",
      "content": "<p>Excellent work! Truly inspiring!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454345,
      "author_name": "Vishy",
      "author_url": "",
      "post_date": "2019-01-11T13:33:28.943000",
      "content": "<p>congrats to you and your team, your tip on Gapnet tweak is indeed very helpful..</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454225,
      "author_name": "zhangboshen",
      "author_url": "",
      "post_date": "2019-01-11T09:15:41.560000",
      "content": "<p>Congratulations! Wonderful work! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454224,
      "author_name": "interneuron",
      "author_url": "",
      "post_date": "2019-01-11T09:15:16.280000",
      "content": "<p>Thank you for the insight! Excellent work.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454214,
      "author_name": "xftts",
      "author_url": "",
      "post_date": "2019-01-11T08:56:39.733000",
      "content": "<p>Great job. congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 909766,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-06-30T20:07:35.860000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 454257,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-11T10:10:55.190000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454406,
      "author_name": "FlYM",
      "author_url": "",
      "post_date": "2019-01-11T15:39:47.743000",
      "content": "<p>Great work! Thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 454217,
      "author_name": "Femi",
      "author_url": "",
      "post_date": "2019-01-11T09:00:37.500000",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "454209": "First a big thanks to all my team-members. They all made this competition an awesome experience. I also want to thank @brian, @heng and @lafoss for their fruitful input. \n\nSince major ingredient of our solution was diversity and in the end we used around 25 different models, I want to give a separate few notes on two of my main contributions. \n\nI want to split up my notes into strategy, architecture as I see each equally important.\n\n**Strategy:**\n\n@tunguz @sasrdw and myself teamed-up quite early which enabled us to work in different directions right from the beginning. We always had diversity of our models in mind. So I concentrated on models that seem a bit different. Diversity to my team-members was also the main reason I sticked to keras, although in my opinion pytorch would have been more suitable for this competition due to its flexibility. We also used different cross validation schemes for the sake of diversity.\n\nAfter some trouble in the beginning for getting the cross-validation right, I started exploring different architectures as posted by @hengck23 . I found it quite efficient to only use 256x256 RGB images in the beginning because it allows for high iteration of different ideas.\n\n**Architectures:**\n\n*GAPNet*\n\nImmediately,  reading the GAPNet paper, I had the idea to change the illustrated architecture to use a pretrained backbone instead. \n\n![original GAPNet][1]\n\nI think one important advantage of the GAPNet architecture is its ability for multiscale. So I tried different backbones and ended up with ResNet18, which also enabled to use a batchsize of 32 on a GTX1080Ti. I also saw minor improvements adding SE-Blocks before the Average Pooling layers with nearly no computational cost, so I added those. I saw no improvement in using RGBY images. I used a weighted bce and f1 loss and a cosine annealing lr schedule and only trained for 20 epochs. After applying our thresholding method to the predictions GAPNet trained on 512x512 RGB images also using the HPA external data a single 5-fold model was able to achieve 0.602 Public LB. I also experimented with different internal/external data proportions, RGBY and 512cropping from 1024 images so I had 4 5-fold GAPNet models which I could ensemble resulted in LB  0.609\n\n*Dual Loss ResNet*\n\nFollowing another post from @hengck23 I implemented a ResNet34 with a dual loss:\n\n![enter image description here][2]\n\nAdditionally to the \"normal\" classification loss I used the output of the last 32x32x128 layer within ResNet34 did Conv2D to 32x32x28 and then used a downsampling of the green channel with the according labels as ground truth mask to have a segmentation loss. This segmentation loss works like a regularizer that ensures that that activations of the 32x32x128 layer are \"nice\".\nThe additional supervised attention added quite some benefit to regularization, and the computational cost was bearable. I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618\n\nI guess @tunguz will write an overall summary where he explains how my models were then incorporated into our overall ensemble.\n\n  [1]: https://storage.googleapis.com/kaggle-forum-message-attachments/inbox/113660/a95f150c7153a17538b074def2255e21/GAP.png\n  [2]: https://storage.googleapis.com/kaggle-forum-message-attachments/415080/10599/attention%20is%20what%20you%20needq.png",
    "454299": "congratulations! are you going to share the code ?\n",
    "454284": "@christofhenkel - congratulations on your result! I basically did an almost identical thing re. butchering a resnet18 into a GAPNet to use the pre-trained weights. Which layers did you attach the GAP blocks to (this is the bit that slows things down)?\n\nAlso, when you say you used different CV schemes - how did you combine overall predictions and validate - some form of voting/mean function? ",
    "456006": "Congratulations! Always fascinating to read top solutions for Kaggle competitions. \n\nIs this the GAPNet paper that you mentioned?\n\nLink (\"Human-level Protein Localization with Convolutional Neural Networks\"):\nhttps://openreview.net/forum?id=ryl5khRcKm",
    "455602": "Thanks for sharing! \n\"I trained 2 variants 5-fold each scoring LB 0.603 and added them to the GAPNet ensemble -&gt; LB 0.618\"\nI saw lb score is 0.65331 ? ",
    "454690": "Very Great work! ",
    "454428": "Hi @christofhenkel and thanks for sharing your ideas. I was wondering how you implemented the autoencoder. Could you explain it a little bit further?",
    "454410": "Excellent work! Truly inspiring!",
    "454345": "congrats to you and your team, your tip on Gapnet tweak is indeed very helpful..",
    "454225": "Congratulations! Wonderful work! ",
    "454224": "Thank you for the insight! Excellent work.",
    "454214": "Great job. congratulations!",
    "909766": "",
    "454257": "",
    "454406": "Great work! Thanks for sharing!",
    "454217": "Congratulations and thanks for sharing!"
  }
}