{
  "id": 117420,
  "title": "how many networks or diversities in the ensemble? ",
  "url": "/competitions/understanding_cloud_organization/discussion/117420",
  "author_name": "",
  "post_date": "2019-11-15T12:01:41.389187600Z",
  "votes": 9,
  "comment_count": 40,
  "views": 0,
  "content": "<p>After two weeks struggling with training a good single model, I give up at .663\nThere are someone can achieve more than .667, cannot wait seeing their solution after it ends.\nNow it's time to ensemble, personally I have three networks for ensemble, two are still training different folds, till now two networks are &gt;.660, one is .657, first ensemble result in .668.</p>\n\n<p>How about networks in your ensemble, or other kinds, like different image sizes?</p>",
  "messages": [
    {
      "id": "673725",
      "postDate": "11/15/2019 12:01:41",
      "content": "<p>After two weeks struggling with training a good single model, I give up at .663\nThere are someone can achieve more than .667, cannot wait seeing their solution after it ends.\nNow it's time to ensemble, personally I have three networks for ensemble, two are still training different folds, till now two networks are &gt;.660, one is .657, first ensemble result in .668.</p>\n\n<p>How about networks in your ensemble, or other kinds, like different image sizes?</p>",
      "rawMarkdown": "After two weeks struggling with training a good single model, I give up at .663\nThere are someone can achieve more than .667, cannot wait seeing their solution after it ends.\nNow it's time to ensemble, personally I have three networks for ensemble, two are still training different folds, till now two networks are &gt;.660, one is .657, first ensemble result in .668.\n\nHow about networks in your ensemble, or other kinds, like different image sizes?",
      "votes": null
    },
    {
      "id": "673789",
      "postDate": "11/15/2019 13:39:39",
      "content": "<p>My single model best only got 0.659... I gave up much earlier.. about one week ago. My current best LB was getting by ensemble 4 models with classifier. 1-&gt;0.659, 2 -&gt;0.660, 3-&gt;0.664, 4-&gt;0.668\nBut it seems to reach to the upper bound, since all of my single model just not that good. 😂 </p>",
      "rawMarkdown": "My single model best only got 0.659... I gave up much earlier.. about one week ago. My current best LB was getting by ensemble 4 models with classifier. 1-&gt;0.659, 2 -&gt;0.660, 3-&gt;0.664, 4-&gt;0.668\nBut it seems to reach to the upper bound, since all of my single model just not that good. 😂",
      "votes": null
    },
    {
      "id": "673808",
      "postDate": "11/15/2019 14:12:47",
      "content": "<p>try to go beyond fpn, unet, deeplab. there are many new architecture like jpu, pspnet, texture encoding, object context encoding. </p>\n\n<p>i currently about 8 network of different segmentation head, loss and input size to get LB 0.672. each network is in the range of 0.65 to 0.66</p>",
      "rawMarkdown": "try to go beyond fpn, unet, deeplab. there are many new architecture like jpu, pspnet, texture encoding, object context encoding. \n\ni currently about 8 network of different segmentation head, loss and input size to get LB 0.672. each network is in the range of 0.65 to 0.66",
      "votes": null
    },
    {
      "id": "673823",
      "postDate": "11/15/2019 14:26:22",
      "content": "<p>It seems most of the people will ensemble at least 3 models. Does this means the shake up will be less possible?</p>",
      "rawMarkdown": "It seems most of the people will ensemble at least 3 models. Does this means the shake up will be less possible?",
      "votes": null
    },
    {
      "id": "673825",
      "postDate": "11/15/2019 14:30:00",
      "content": "<p>yes. ensemble always give better results (public and private).</p>\n\n<p>resnet34 seems sufficient. it takes about 1 to 2 hr to train one segmentation model. \nso i think most people will have more than 3 models.  </p>",
      "rawMarkdown": "yes. ensemble always give better results (public and private).\n\nresnet34 seems sufficient. it takes about 1 to 2 hr to train one segmentation model. \nso i think most people will have more than 3 models.",
      "votes": null
    },
    {
      "id": "673836",
      "postDate": "11/15/2019 14:41:34",
      "content": "<p>True, my best single model was trained with seresnet18 backbone.</p>",
      "rawMarkdown": "True, my best single model was trained with seresnet18 backbone.",
      "votes": null
    },
    {
      "id": "673838",
      "postDate": "11/15/2019 14:42:37",
      "content": "<p>Thanks Heng, I'll definitely try. </p>",
      "rawMarkdown": "Thanks Heng, I'll definitely try.",
      "votes": null
    },
    {
      "id": "673841",
      "postDate": "11/15/2019 14:44:16",
      "content": "<p>That's OK, it's wise decision for this competition.</p>",
      "rawMarkdown": "That's OK, it's wise decision for this competition.",
      "votes": null
    },
    {
      "id": "673871",
      "postDate": "11/15/2019 15:29:07",
      "content": "<p><a href=\"/xiejialun\">@xiejialun</a> haha, my best is deeplabv3+, which was always my best performance network from previous competitions...</p>",
      "rawMarkdown": "xiejialun haha, my best is deeplabv3+, which was always my best performance network from previous competitions...",
      "votes": null
    },
    {
      "id": "673981",
      "postDate": "11/15/2019 18:28:13",
      "content": "<p>Hey <a href=\"/xiejialun\">@xiejialun</a> how are you ensembling your models? As for me LB is not consistent with different ensembles</p>",
      "rawMarkdown": "Hey @xiejialun how are you ensembling your models? As for me LB is not consistent with different ensembles",
      "votes": null
    },
    {
      "id": "674126",
      "postDate": "11/15/2019 23:09:07",
      "content": "<p><a href=\"/ubamba98\">@ubamba98</a> I'm using average the raw prediction for ensemble. I did train these model with slightly different process/loss on purpose to make them can help each other's performance. I don't sure this will help or I just get lucky😂. But the number of model is quite consistent with LB. </p>",
      "rawMarkdown": "ubamba98 I'm using average the raw prediction for ensemble. I did train these model with slightly different process/loss on purpose to make them can help each other's performance. I don't sure this will help or I just get lucky😂. But the number of model is quite consistent with LB.",
      "votes": null
    },
    {
      "id": "674193",
      "postDate": "11/16/2019 01:48:37",
      "content": "<p>Thanks I will try that in future.</p>",
      "rawMarkdown": "Thanks I will try that in future.",
      "votes": null
    },
    {
      "id": "674199",
      "postDate": "11/16/2019 01:54:17",
      "content": "<p>I'm also struggling to find a good model for ensemble. 😂 </p>",
      "rawMarkdown": "I'm also struggling to find a good model for ensemble. 😂",
      "votes": null
    },
    {
      "id": "674202",
      "postDate": "11/16/2019 01:59:46",
      "content": "<p>it is difficult. sometimes we use stacking or learned another level-2 network. \nyou can also see this: <a href=\"https://ai.googleblog.com/2018/10/introducing-adanet-fast-and-flexible.html\">https://ai.googleblog.com/2018/10/introducing-adanet-fast-and-flexible.html</a></p>\n\n<p>\"Building an ensemble of neural networks has several challenges: What are the best subnetwork architectures to consider? Is it best to reuse the same architectures or encourage diversity? While complex subnetworks with more parameters will tend to perform better on the training set, they may not generalize to unseen data due to their greater complexity. \"</p>",
      "rawMarkdown": "it is difficult. sometimes we use stacking or learned another level-2 network. \nyou can also see this: https://ai.googleblog.com/2018/10/introducing-adanet-fast-and-flexible.html\n\n\"Building an ensemble of neural networks has several challenges: What are the best subnetwork architectures to consider? Is it best to reuse the same architectures or encourage diversity? While complex subnetworks with more parameters will tend to perform better on the training set, they may not generalize to unseen data due to their greater complexity. \"",
      "votes": null
    },
    {
      "id": "674211",
      "postDate": "11/16/2019 02:37:09",
      "content": "<p>Single fpn 6640 + single unet 6612 only got 6647  : (</p>",
      "rawMarkdown": "Single fpn 6640 + single unet 6612 only got 6647  : (",
      "votes": null
    },
    {
      "id": "674212",
      "postDate": "11/16/2019 02:37:59",
      "content": "<p>You're right, it's difficult.\nSo I haven't improved my LB score in days. 😭 </p>",
      "rawMarkdown": "You're right, it's difficult.\nSo I haven't improved my LB score in days. 😭",
      "votes": null
    },
    {
      "id": "674224",
      "postDate": "11/16/2019 03:09:12",
      "content": "<p>Sometimes it’s fun to see the 1st place make a joke 😄 😂 </p>",
      "rawMarkdown": "Sometimes it’s fun to see the 1st place make a joke 😄 😂",
      "votes": null
    },
    {
      "id": "674227",
      "postDate": "11/16/2019 03:34:27",
      "content": "<p>\"Single fpn 6640 + single unet 6612 only got 6647 : (\"</p>\n\n<p>my suggestion is:</p>\n\n<ol>\n<li><p>breakdown the results of single model and ensemble, e.g.\nsingle model false positive =xx,  ensemble model false positive =yy</p></li>\n<li><p>if yy is better than xx, than we know that the effects of ensemble is reduction of fp.</p></li>\n<li><p>then we can think of ways to further improve this effect or to address problems that are solved by ensemble </p></li>\n</ol>",
      "rawMarkdown": "\"Single fpn 6640 + single unet 6612 only got 6647 : (\"\n\nmy suggestion is:\n\n1. breakdown the results of single model and ensemble, e.g.\nsingle model false positive =xx,  ensemble model false positive =yy\n\n\n2. if yy is better than xx, than we know that the effects of ensemble is reduction of fp.\n\n3. then we can think of ways to further improve this effect or to address problems that are solved by ensemble",
      "votes": null
    },
    {
      "id": "674228",
      "postDate": "11/16/2019 03:35:21",
      "content": "<p>\"So I haven't improved my LB score in days\"</p>\n\n<p>lesson from steel competition is that  preventing shakeup is also important </p>",
      "rawMarkdown": "\"So I haven't improved my LB score in days\"\n\nlesson from steel competition is that  preventing shakeup is also important",
      "votes": null
    },
    {
      "id": "674236",
      "postDate": "11/16/2019 03:59:29",
      "content": "<p>Yes... It's a sad experience for us.\nAnd when I found that I got a score [private: 0.90828, public: 0.92286], when I just set threshold to [0.7, 0.7, 0.7, 0.7], I cried... 😿 </p>\n\n<p>So I will only use high thresholds in this competition...</p>",
      "rawMarkdown": "Yes... It's a sad experience for us.\nAnd when I found that I got a score [private: 0.90828, public: 0.92286], when I just set threshold to [0.7, 0.7, 0.7, 0.7], I cried... 😿 \n\nSo I will only use high thresholds in this competition...",
      "votes": null
    },
    {
      "id": "674243",
      "postDate": "11/16/2019 04:16:16",
      "content": "<p>my experiment shows higher threshold works better but we still have to be cautious.</p>\n\n<p>i haven't think of a way yet but we should think of a way to use both the public and private data to confirm. (unlike the steel competition, this time we get to see the private data set as well)</p>\n\n<p>some possible solutions\n1. visual inspection of test data\n2. measure the similarity of test data \n3. think of a way to verify label consistency in test data. e.g. we make prediction on test. then we use pseudo label only to train and verify on train set. then we apply mixup augmentation pseudo label and see the effects on error of the same train set, etc ...</p>\n\n<p>since we are only interested in image level label for fp rejection, we have a lot of options actually. it is a problem of verifying (or robustness, uncertainty estimation) the classification (not segmentation) of test set</p>",
      "rawMarkdown": "my experiment shows higher threshold works better but we still have to be cautious.\n\ni haven't think of a way yet but we should think of a way to use both the public and private data to confirm. (unlike the steel competition, this time we get to see the private data set as well)\n\nsome possible solutions\n1. visual inspection of test data\n2. measure the similarity of test data \n3. think of a way to verify label consistency in test data. e.g. we make prediction on test. then we use pseudo label only to train and verify on train set. then we apply mixup augmentation pseudo label and see the effects on error of the same train set, etc ...\n\nsince we are only interested in image level label for fp rejection, we have a lot of options actually. it is a problem of verifying (or robustness, uncertainty estimation) the classification (not segmentation) of test set",
      "votes": null
    },
    {
      "id": "674260",
      "postDate": "11/16/2019 05:33:37",
      "content": "<p>Thanks for your suggestions.\nHeng, I respect your sharing and I'm always rooting for you.</p>",
      "rawMarkdown": "Thanks for your suggestions.\nHeng, I respect your sharing and I'm always rooting for you.",
      "votes": null
    },
    {
      "id": "674448",
      "postDate": "11/16/2019 14:24:32",
      "content": "<p>single image size, 8x folds Unet, 3 more folds Unet with different encoders. Contribution of the different encoder is negligble I think.</p>",
      "rawMarkdown": "single image size, 8x folds Unet, 3 more folds Unet with different encoders. Contribution of the different encoder is negligble I think.",
      "votes": null
    },
    {
      "id": "674452",
      "postDate": "11/16/2019 14:33:40",
      "content": "<p>I can only gain very small benefit on multi-fold training, but different encoder did improve my score. This is quite weird haha.</p>\n\n<p>Btw, can't wait to see your solution! single model with 0.667..that's insane!</p>",
      "rawMarkdown": "I can only gain very small benefit on multi-fold training, but different encoder did improve my score. This is quite weird haha.\n\nBtw, can't wait to see your solution! single model with 0.667..that's insane!",
      "votes": null
    },
    {
      "id": "674518",
      "postDate": "11/16/2019 16:14:18",
      "content": "<p>another is diversity in loss, e.g</p>\n\n<p>```</p>\n\n<p>class Net(nn.Module):\n    ....\n    def forward(self, x):\n        feature =  self.block(x)\n       ....\n       .....\n       label_probability1 = self.one_way_to_predict( ....)\n       ....\n       .....\n       label_probability2 = self.another_way_to_predict( .... different feature used ....)\n       .....\n       label_probability3 = self.yet_another_way_to_predict( .... yet different feature used ....)</p>\n\n<p>```</p>\n\n<p>for inference, average the 3 label prediction (since image label is important, we only create multi classifiers).</p>\n\n<p>this is like using shared features to learn 3 classifiers. and then average them like ensemble.</p>\n\n<p>for back propagation, use:</p>\n\n<p>```\nloss = a*loss(label_probability1 ) +  b*loss(label_probability2 ) +  c*loss(label_probability3 ) + ...</p>\n\n<p>a,b,c are random variable with certain range e.g. a+b+c = 1 and a in range 0.3333 +/- noise, etc ....\n```</p>\n\n<p>we inject the weighing loss noise to make it more robust</p>",
      "rawMarkdown": "another is diversity in loss, e.g\n\n```\n\nclass Net(nn.Module):\n    ....\n    def forward(self, x):\n        feature =  self.block(x)\n       ....\n       .....\n       label_probability1 = self.one_way_to_predict( ....)\n       ....\n       .....\n       label_probability2 = self.another_way_to_predict( .... different feature used ....)\n       .....\n       label_probability3 = self.yet_another_way_to_predict( .... yet different feature used ....)\n\n\n```\n\nfor inference, average the 3 label prediction (since image label is important, we only create multi classifiers).\n\nthis is like using shared features to learn 3 classifiers. and then average them like ensemble.\n\nfor back propagation, use:\n\n```\nloss = a*loss(label_probability1 ) +  b*loss(label_probability2 ) +  c*loss(label_probability3 ) + ...\n\na,b,c are random variable with certain range e.g. a+b+c = 1 and a in range 0.3333 +/- noise, etc ....\n```\n\nwe inject the weighing loss noise to make it more robust",
      "votes": null
    },
    {
      "id": "674780",
      "postDate": "11/17/2019 04:15:00",
      "content": "<p>Got .6671 network, now I suspect some kaggler got even higher single model performance...</p>",
      "rawMarkdown": "Got .6671 network, now I suspect some kaggler got even higher single model performance...",
      "votes": null
    },
    {
      "id": "674790",
      "postDate": "11/17/2019 04:41:52",
      "content": "<p>definitely yes.</p>\n\n<p>the ensemble results just prove that there is that much discriminative information that can extracted from the train set.</p>\n\n<p>it is whether we can design good learning method to extract the information for a single  model</p>\n\n<p>if you read through the papers of imagenet, you will find that the ensemble of many vgg-16 is close to  the performance of single resnet50</p>",
      "rawMarkdown": "definitely yes.\n\nthe ensemble results just prove that there is that much discriminative information that can extracted from the train set.\n\nit is whether we can design good learning method to extract the information for a single  model\n\nif you read through the papers of imagenet, you will find that the ensemble of many vgg-16 is close to  the performance of single resnet50",
      "votes": null
    },
    {
      "id": "674842",
      "postDate": "11/17/2019 06:24:25",
      "content": "<p>single fpn with or without classifier?</p>",
      "rawMarkdown": "single fpn with or without classifier?",
      "votes": null
    },
    {
      "id": "674869",
      "postDate": "11/17/2019 07:10:03",
      "content": "<p><a href=\"/niuddd\">@niuddd</a> amazing work! is it K-fold or single fold? </p>",
      "rawMarkdown": "niuddd amazing work! is it K-fold or single fold?",
      "votes": null
    },
    {
      "id": "675374",
      "postDate": "11/18/2019 02:19:08",
      "content": "<p>I got 0.6686 from a single network with 4 folds. However, looks like it doesn't matter; at least for the Public LB.  Because while ensembling, the result becomes funnier.  For example, the ensemble of lower scored networks giving higher public LB compared to the ensemble with the score of the best-scored network. I am not sure what is causing that.</p>",
      "rawMarkdown": "I got 0.6686 from a single network with 4 folds. However, looks like it doesn't matter; at least for the Public LB.  Because while ensembling, the result becomes funnier.  For example, the ensemble of lower scored networks giving higher public LB compared to the ensemble with the score of the best-scored network. I am not sure what is causing that.",
      "votes": null
    },
    {
      "id": "675375",
      "postDate": "11/18/2019 02:21:57",
      "content": "<p><a href=\"/mykttu\">@mykttu</a> true I am also experiencing the same effect I don't think private LB is a good generalization.</p>",
      "rawMarkdown": "mykttu true I am also experiencing the same effect I don't think private LB is a good generalization.",
      "votes": null
    },
    {
      "id": "675415",
      "postDate": "11/18/2019 03:53:38",
      "content": "<p>I made a mistake, 6640 model is unet resnet34. So it is a reasonable ensemble score.</p>",
      "rawMarkdown": "I made a mistake, 6640 model is unet resnet34. So it is a reasonable ensemble score.",
      "votes": null
    },
    {
      "id": "675442",
      "postDate": "11/18/2019 05:00:35",
      "content": "<p>True. Going beyond .670, I realized there's sub-competition of tuning ensembles.</p>",
      "rawMarkdown": "True. Going beyond .670, I realized there's sub-competition of tuning ensembles.",
      "votes": null
    },
    {
      "id": "675446",
      "postDate": "11/18/2019 05:08:23",
      "content": "<p>Another question is would you select two best scoring (public LB) submissions, or would you prefer two that have most confidence with? Personally, I would choose at least one to be a big ensemble (of as many models as I have), which statistically has lower variance (though lower LB) to fight against the possible shake up. What about you?</p>",
      "rawMarkdown": "Another question is would you select two best scoring (public LB) submissions, or would you prefer two that have most confidence with? Personally, I would choose at least one to be a big ensemble (of as many models as I have), which statistically has lower variance (though lower LB) to fight against the possible shake up. What about you?",
      "votes": null
    },
    {
      "id": "675455",
      "postDate": "11/18/2019 05:14:42",
      "content": "<p>I think submission should be selected such that you can minimize the false positives. I will select my submission based on it not the lb score.</p>",
      "rawMarkdown": "I think submission should be selected such that you can minimize the false positives. I will select my submission based on it not the lb score.",
      "votes": null
    },
    {
      "id": "675463",
      "postDate": "11/18/2019 05:23:41",
      "content": "<p>Same as you, one with best public LB, one with big ensemble.\nGood luck</p>",
      "rawMarkdown": "Same as you, one with best public LB, one with big ensemble.\nGood luck",
      "votes": null
    },
    {
      "id": "675482",
      "postDate": "11/18/2019 05:49:19",
      "content": "<p>5 different input size model with same head and decoder got me from 0.65 to 0.666</p>",
      "rawMarkdown": "5 different input size model with same head and decoder got me from 0.65 to 0.666",
      "votes": null
    },
    {
      "id": "675544",
      "postDate": "11/18/2019 07:53:00",
      "content": "<p>5 different architectures each 5 fold.\nEnsemble Strategy in 5 folds -&gt; sigmoid(mean(probabilities))\nEnsemble Strategy to merge binary masks -&gt; max voting\nBest single architecture performance -&gt; 0.6658\nLB score -&gt; 0.6674</p>",
      "rawMarkdown": "5 different architectures each 5 fold.\nEnsemble Strategy in 5 folds -&gt; sigmoid(mean(probabilities))\nEnsemble Strategy to merge binary masks -&gt; max voting\nBest single architecture performance -&gt; 0.6658\nLB score -&gt; 0.6674",
      "votes": null
    },
    {
      "id": "675551",
      "postDate": "11/18/2019 08:05:17",
      "content": "<p>What is the voting threshold did you use I am experiencing a huge difference in LB score with different thresh in voting. Is that the case with you too?</p>",
      "rawMarkdown": "What is the voting threshold did you use I am experiencing a huge difference in LB score with different thresh in voting. Is that the case with you too?",
      "votes": null
    },
    {
      "id": "675587",
      "postDate": "11/18/2019 09:12:08",
      "content": "<p>I have experience such situation. Shakeup is around 0.001. Currently I found out higher theshold has better LB score.</p>",
      "rawMarkdown": "I have experience such situation. Shakeup is around 0.001. Currently I found out higher theshold has better LB score.",
      "votes": null
    },
    {
      "id": "675594",
      "postDate": "11/18/2019 09:26:45",
      "content": "<p>Ya same. Its really hard to select final submission.</p>",
      "rawMarkdown": "Ya same. Its really hard to select final submission.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 673789,
      "author_name": "xiejialun",
      "author_url": "",
      "post_date": "11/15/2019 13:39:39",
      "content": "<p>My single model best only got 0.659... I gave up much earlier.. about one week ago. My current best LB was getting by ensemble 4 models with classifier. 1-&gt;0.659, 2 -&gt;0.660, 3-&gt;0.664, 4-&gt;0.668\nBut it seems to reach to the upper bound, since all of my single model just not that good. 😂 </p>",
      "votes": null,
      "replies": [
        {
          "id": 673841,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "11/15/2019 14:44:16",
          "content": "<p>That's OK, it's wise decision for this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673981,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "11/15/2019 18:28:13",
          "content": "<p>Hey <a href=\"/xiejialun\">@xiejialun</a> how are you ensembling your models? As for me LB is not consistent with different ensembles</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674126,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/15/2019 23:09:07",
          "content": "<p><a href=\"/ubamba98\">@ubamba98</a> I'm using average the raw prediction for ensemble. I did train these model with slightly different process/loss on purpose to make them can help each other's performance. I don't sure this will help or I just get lucky😂. But the number of model is quite consistent with LB. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674193,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "11/16/2019 01:48:37",
          "content": "<p>Thanks I will try that in future.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 673808,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "11/15/2019 14:12:47",
      "content": "<p>try to go beyond fpn, unet, deeplab. there are many new architecture like jpu, pspnet, texture encoding, object context encoding. </p>\n\n<p>i currently about 8 network of different segmentation head, loss and input size to get LB 0.672. each network is in the range of 0.65 to 0.66</p>",
      "votes": null,
      "replies": [
        {
          "id": 673823,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/15/2019 14:26:22",
          "content": "<p>It seems most of the people will ensemble at least 3 models. Does this means the shake up will be less possible?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673825,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/15/2019 14:30:00",
          "content": "<p>yes. ensemble always give better results (public and private).</p>\n\n<p>resnet34 seems sufficient. it takes about 1 to 2 hr to train one segmentation model. \nso i think most people will have more than 3 models.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673836,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/15/2019 14:41:34",
          "content": "<p>True, my best single model was trained with seresnet18 backbone.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673838,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "11/15/2019 14:42:37",
          "content": "<p>Thanks Heng, I'll definitely try. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673871,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "11/15/2019 15:29:07",
          "content": "<p><a href=\"/xiejialun\">@xiejialun</a> haha, my best is deeplabv3+, which was always my best performance network from previous competitions...</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 674199,
      "author_name": "limerobot",
      "author_url": "",
      "post_date": "11/16/2019 01:54:17",
      "content": "<p>I'm also struggling to find a good model for ensemble. 😂 </p>",
      "votes": null,
      "replies": [
        {
          "id": 674202,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/16/2019 01:59:46",
          "content": "<p>it is difficult. sometimes we use stacking or learned another level-2 network. \nyou can also see this: <a href=\"https://ai.googleblog.com/2018/10/introducing-adanet-fast-and-flexible.html\">https://ai.googleblog.com/2018/10/introducing-adanet-fast-and-flexible.html</a></p>\n\n<p>\"Building an ensemble of neural networks has several challenges: What are the best subnetwork architectures to consider? Is it best to reuse the same architectures or encourage diversity? While complex subnetworks with more parameters will tend to perform better on the training set, they may not generalize to unseen data due to their greater complexity. \"</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674212,
          "author_name": "limerobot",
          "author_url": "",
          "post_date": "11/16/2019 02:37:59",
          "content": "<p>You're right, it's difficult.\nSo I haven't improved my LB score in days. 😭 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674224,
          "author_name": "ratthachat",
          "author_url": "",
          "post_date": "11/16/2019 03:09:12",
          "content": "<p>Sometimes it’s fun to see the 1st place make a joke 😄 😂 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674228,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/16/2019 03:35:21",
          "content": "<p>\"So I haven't improved my LB score in days\"</p>\n\n<p>lesson from steel competition is that  preventing shakeup is also important </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674236,
          "author_name": "limerobot",
          "author_url": "",
          "post_date": "11/16/2019 03:59:29",
          "content": "<p>Yes... It's a sad experience for us.\nAnd when I found that I got a score [private: 0.90828, public: 0.92286], when I just set threshold to [0.7, 0.7, 0.7, 0.7], I cried... 😿 </p>\n\n<p>So I will only use high thresholds in this competition...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674243,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/16/2019 04:16:16",
          "content": "<p>my experiment shows higher threshold works better but we still have to be cautious.</p>\n\n<p>i haven't think of a way yet but we should think of a way to use both the public and private data to confirm. (unlike the steel competition, this time we get to see the private data set as well)</p>\n\n<p>some possible solutions\n1. visual inspection of test data\n2. measure the similarity of test data \n3. think of a way to verify label consistency in test data. e.g. we make prediction on test. then we use pseudo label only to train and verify on train set. then we apply mixup augmentation pseudo label and see the effects on error of the same train set, etc ...</p>\n\n<p>since we are only interested in image level label for fp rejection, we have a lot of options actually. it is a problem of verifying (or robustness, uncertainty estimation) the classification (not segmentation) of test set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674260,
          "author_name": "limerobot",
          "author_url": "",
          "post_date": "11/16/2019 05:33:37",
          "content": "<p>Thanks for your suggestions.\nHeng, I respect your sharing and I'm always rooting for you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 674211,
      "author_name": "qiaoshiji",
      "author_url": "",
      "post_date": "11/16/2019 02:37:09",
      "content": "<p>Single fpn 6640 + single unet 6612 only got 6647  : (</p>",
      "votes": null,
      "replies": [
        {
          "id": 674227,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/16/2019 03:34:27",
          "content": "<p>\"Single fpn 6640 + single unet 6612 only got 6647 : (\"</p>\n\n<p>my suggestion is:</p>\n\n<ol>\n<li><p>breakdown the results of single model and ensemble, e.g.\nsingle model false positive =xx,  ensemble model false positive =yy</p></li>\n<li><p>if yy is better than xx, than we know that the effects of ensemble is reduction of fp.</p></li>\n<li><p>then we can think of ways to further improve this effect or to address problems that are solved by ensemble </p></li>\n</ol>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674842,
          "author_name": "ekan825",
          "author_url": "",
          "post_date": "11/17/2019 06:24:25",
          "content": "<p>single fpn with or without classifier?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675415,
          "author_name": "qiaoshiji",
          "author_url": "",
          "post_date": "11/18/2019 03:53:38",
          "content": "<p>I made a mistake, 6640 model is unet resnet34. So it is a reasonable ensemble score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 674448,
      "author_name": "tugstugi",
      "author_url": "",
      "post_date": "11/16/2019 14:24:32",
      "content": "<p>single image size, 8x folds Unet, 3 more folds Unet with different encoders. Contribution of the different encoder is negligble I think.</p>",
      "votes": null,
      "replies": [
        {
          "id": 674452,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/16/2019 14:33:40",
          "content": "<p>I can only gain very small benefit on multi-fold training, but different encoder did improve my score. This is quite weird haha.</p>\n\n<p>Btw, can't wait to see your solution! single model with 0.667..that's insane!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674518,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/16/2019 16:14:18",
          "content": "<p>another is diversity in loss, e.g</p>\n\n<p>```</p>\n\n<p>class Net(nn.Module):\n    ....\n    def forward(self, x):\n        feature =  self.block(x)\n       ....\n       .....\n       label_probability1 = self.one_way_to_predict( ....)\n       ....\n       .....\n       label_probability2 = self.another_way_to_predict( .... different feature used ....)\n       .....\n       label_probability3 = self.yet_another_way_to_predict( .... yet different feature used ....)</p>\n\n<p>```</p>\n\n<p>for inference, average the 3 label prediction (since image label is important, we only create multi classifiers).</p>\n\n<p>this is like using shared features to learn 3 classifiers. and then average them like ensemble.</p>\n\n<p>for back propagation, use:</p>\n\n<p>```\nloss = a*loss(label_probability1 ) +  b*loss(label_probability2 ) +  c*loss(label_probability3 ) + ...</p>\n\n<p>a,b,c are random variable with certain range e.g. a+b+c = 1 and a in range 0.3333 +/- noise, etc ....\n```</p>\n\n<p>we inject the weighing loss noise to make it more robust</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 674780,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "11/17/2019 04:15:00",
      "content": "<p>Got .6671 network, now I suspect some kaggler got even higher single model performance...</p>",
      "votes": null,
      "replies": [
        {
          "id": 674790,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "11/17/2019 04:41:52",
          "content": "<p>definitely yes.</p>\n\n<p>the ensemble results just prove that there is that much discriminative information that can extracted from the train set.</p>\n\n<p>it is whether we can design good learning method to extract the information for a single  model</p>\n\n<p>if you read through the papers of imagenet, you will find that the ensemble of many vgg-16 is close to  the performance of single resnet50</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 674869,
          "author_name": "adish333",
          "author_url": "",
          "post_date": "11/17/2019 07:10:03",
          "content": "<p><a href=\"/niuddd\">@niuddd</a> amazing work! is it K-fold or single fold? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675374,
          "author_name": "mykttu",
          "author_url": "",
          "post_date": "11/18/2019 02:19:08",
          "content": "<p>I got 0.6686 from a single network with 4 folds. However, looks like it doesn't matter; at least for the Public LB.  Because while ensembling, the result becomes funnier.  For example, the ensemble of lower scored networks giving higher public LB compared to the ensemble with the score of the best-scored network. I am not sure what is causing that.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675375,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "11/18/2019 02:21:57",
          "content": "<p><a href=\"/mykttu\">@mykttu</a> true I am also experiencing the same effect I don't think private LB is a good generalization.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675442,
          "author_name": "niuddd",
          "author_url": "",
          "post_date": "11/18/2019 05:00:35",
          "content": "<p>True. Going beyond .670, I realized there's sub-competition of tuning ensembles.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 675446,
      "author_name": "niuddd",
      "author_url": "",
      "post_date": "11/18/2019 05:08:23",
      "content": "<p>Another question is would you select two best scoring (public LB) submissions, or would you prefer two that have most confidence with? Personally, I would choose at least one to be a big ensemble (of as many models as I have), which statistically has lower variance (though lower LB) to fight against the possible shake up. What about you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 675455,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "11/18/2019 05:14:42",
          "content": "<p>I think submission should be selected such that you can minimize the false positives. I will select my submission based on it not the lb score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675463,
          "author_name": "xiejialun",
          "author_url": "",
          "post_date": "11/18/2019 05:23:41",
          "content": "<p>Same as you, one with best public LB, one with big ensemble.\nGood luck</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 675482,
      "author_name": "markson14",
      "author_url": "",
      "post_date": "11/18/2019 05:49:19",
      "content": "<p>5 different input size model with same head and decoder got me from 0.65 to 0.666</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 675544,
      "author_name": "adish333",
      "author_url": "",
      "post_date": "11/18/2019 07:53:00",
      "content": "<p>5 different architectures each 5 fold.\nEnsemble Strategy in 5 folds -&gt; sigmoid(mean(probabilities))\nEnsemble Strategy to merge binary masks -&gt; max voting\nBest single architecture performance -&gt; 0.6658\nLB score -&gt; 0.6674</p>",
      "votes": null,
      "replies": [
        {
          "id": 675551,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "11/18/2019 08:05:17",
          "content": "<p>What is the voting threshold did you use I am experiencing a huge difference in LB score with different thresh in voting. Is that the case with you too?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675587,
          "author_name": "markson14",
          "author_url": "",
          "post_date": "11/18/2019 09:12:08",
          "content": "<p>I have experience such situation. Shakeup is around 0.001. Currently I found out higher theshold has better LB score.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675594,
          "author_name": "ubamba98",
          "author_url": "",
          "post_date": "11/18/2019 09:26:45",
          "content": "<p>Ya same. Its really hard to select final submission.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "673725": "After two weeks struggling with training a good single model, I give up at .663\nThere are someone can achieve more than .667, cannot wait seeing their solution after it ends.\nNow it's time to ensemble, personally I have three networks for ensemble, two are still training different folds, till now two networks are &gt;.660, one is .657, first ensemble result in .668.\n\nHow about networks in your ensemble, or other kinds, like different image sizes?",
    "673789": "My single model best only got 0.659... I gave up much earlier.. about one week ago. My current best LB was getting by ensemble 4 models with classifier. 1-&gt;0.659, 2 -&gt;0.660, 3-&gt;0.664, 4-&gt;0.668\nBut it seems to reach to the upper bound, since all of my single model just not that good. 😂",
    "673808": "try to go beyond fpn, unet, deeplab. there are many new architecture like jpu, pspnet, texture encoding, object context encoding. \n\ni currently about 8 network of different segmentation head, loss and input size to get LB 0.672. each network is in the range of 0.65 to 0.66",
    "673823": "It seems most of the people will ensemble at least 3 models. Does this means the shake up will be less possible?",
    "673825": "yes. ensemble always give better results (public and private).\n\nresnet34 seems sufficient. it takes about 1 to 2 hr to train one segmentation model. \nso i think most people will have more than 3 models.",
    "673836": "True, my best single model was trained with seresnet18 backbone.",
    "673838": "Thanks Heng, I'll definitely try.",
    "673841": "That's OK, it's wise decision for this competition.",
    "673871": "xiejialun haha, my best is deeplabv3+, which was always my best performance network from previous competitions...",
    "673981": "Hey @xiejialun how are you ensembling your models? As for me LB is not consistent with different ensembles",
    "674126": "ubamba98 I'm using average the raw prediction for ensemble. I did train these model with slightly different process/loss on purpose to make them can help each other's performance. I don't sure this will help or I just get lucky😂. But the number of model is quite consistent with LB.",
    "674193": "Thanks I will try that in future.",
    "674199": "I'm also struggling to find a good model for ensemble. 😂",
    "674202": "it is difficult. sometimes we use stacking or learned another level-2 network. \nyou can also see this: https://ai.googleblog.com/2018/10/introducing-adanet-fast-and-flexible.html\n\n\"Building an ensemble of neural networks has several challenges: What are the best subnetwork architectures to consider? Is it best to reuse the same architectures or encourage diversity? While complex subnetworks with more parameters will tend to perform better on the training set, they may not generalize to unseen data due to their greater complexity. \"",
    "674211": "Single fpn 6640 + single unet 6612 only got 6647  : (",
    "674212": "You're right, it's difficult.\nSo I haven't improved my LB score in days. 😭",
    "674224": "Sometimes it’s fun to see the 1st place make a joke 😄 😂",
    "674227": "\"Single fpn 6640 + single unet 6612 only got 6647 : (\"\n\nmy suggestion is:\n\n1. breakdown the results of single model and ensemble, e.g.\nsingle model false positive =xx,  ensemble model false positive =yy\n\n\n2. if yy is better than xx, than we know that the effects of ensemble is reduction of fp.\n\n3. then we can think of ways to further improve this effect or to address problems that are solved by ensemble",
    "674228": "\"So I haven't improved my LB score in days\"\n\nlesson from steel competition is that  preventing shakeup is also important",
    "674236": "Yes... It's a sad experience for us.\nAnd when I found that I got a score [private: 0.90828, public: 0.92286], when I just set threshold to [0.7, 0.7, 0.7, 0.7], I cried... 😿 \n\nSo I will only use high thresholds in this competition...",
    "674243": "my experiment shows higher threshold works better but we still have to be cautious.\n\ni haven't think of a way yet but we should think of a way to use both the public and private data to confirm. (unlike the steel competition, this time we get to see the private data set as well)\n\nsome possible solutions\n1. visual inspection of test data\n2. measure the similarity of test data \n3. think of a way to verify label consistency in test data. e.g. we make prediction on test. then we use pseudo label only to train and verify on train set. then we apply mixup augmentation pseudo label and see the effects on error of the same train set, etc ...\n\nsince we are only interested in image level label for fp rejection, we have a lot of options actually. it is a problem of verifying (or robustness, uncertainty estimation) the classification (not segmentation) of test set",
    "674260": "Thanks for your suggestions.\nHeng, I respect your sharing and I'm always rooting for you.",
    "674448": "single image size, 8x folds Unet, 3 more folds Unet with different encoders. Contribution of the different encoder is negligble I think.",
    "674452": "I can only gain very small benefit on multi-fold training, but different encoder did improve my score. This is quite weird haha.\n\nBtw, can't wait to see your solution! single model with 0.667..that's insane!",
    "674518": "another is diversity in loss, e.g\n\n```\n\nclass Net(nn.Module):\n    ....\n    def forward(self, x):\n        feature =  self.block(x)\n       ....\n       .....\n       label_probability1 = self.one_way_to_predict( ....)\n       ....\n       .....\n       label_probability2 = self.another_way_to_predict( .... different feature used ....)\n       .....\n       label_probability3 = self.yet_another_way_to_predict( .... yet different feature used ....)\n\n\n```\n\nfor inference, average the 3 label prediction (since image label is important, we only create multi classifiers).\n\nthis is like using shared features to learn 3 classifiers. and then average them like ensemble.\n\nfor back propagation, use:\n\n```\nloss = a*loss(label_probability1 ) +  b*loss(label_probability2 ) +  c*loss(label_probability3 ) + ...\n\na,b,c are random variable with certain range e.g. a+b+c = 1 and a in range 0.3333 +/- noise, etc ....\n```\n\nwe inject the weighing loss noise to make it more robust",
    "674780": "Got .6671 network, now I suspect some kaggler got even higher single model performance...",
    "674790": "definitely yes.\n\nthe ensemble results just prove that there is that much discriminative information that can extracted from the train set.\n\nit is whether we can design good learning method to extract the information for a single  model\n\nif you read through the papers of imagenet, you will find that the ensemble of many vgg-16 is close to  the performance of single resnet50",
    "674842": "single fpn with or without classifier?",
    "674869": "niuddd amazing work! is it K-fold or single fold?",
    "675374": "I got 0.6686 from a single network with 4 folds. However, looks like it doesn't matter; at least for the Public LB.  Because while ensembling, the result becomes funnier.  For example, the ensemble of lower scored networks giving higher public LB compared to the ensemble with the score of the best-scored network. I am not sure what is causing that.",
    "675375": "mykttu true I am also experiencing the same effect I don't think private LB is a good generalization.",
    "675415": "I made a mistake, 6640 model is unet resnet34. So it is a reasonable ensemble score.",
    "675442": "True. Going beyond .670, I realized there's sub-competition of tuning ensembles.",
    "675446": "Another question is would you select two best scoring (public LB) submissions, or would you prefer two that have most confidence with? Personally, I would choose at least one to be a big ensemble (of as many models as I have), which statistically has lower variance (though lower LB) to fight against the possible shake up. What about you?",
    "675455": "I think submission should be selected such that you can minimize the false positives. I will select my submission based on it not the lb score.",
    "675463": "Same as you, one with best public LB, one with big ensemble.\nGood luck",
    "675482": "5 different input size model with same head and decoder got me from 0.65 to 0.666",
    "675544": "5 different architectures each 5 fold.\nEnsemble Strategy in 5 folds -&gt; sigmoid(mean(probabilities))\nEnsemble Strategy to merge binary masks -&gt; max voting\nBest single architecture performance -&gt; 0.6658\nLB score -&gt; 0.6674",
    "675551": "What is the voting threshold did you use I am experiencing a huge difference in LB score with different thresh in voting. Is that the case with you too?",
    "675587": "I have experience such situation. Shakeup is around 0.001. Currently I found out higher theshold has better LB score.",
    "675594": "Ya same. Its really hard to select final submission."
  },
  "source": "meta"
}