{
  "id": 110303,
  "title": "Some thoughts on this competition",
  "url": "/competitions/understanding_cloud_organization/discussion/110303",
  "author_name": "Andrey Lukyanenko",
  "post_date": "2019-09-26T18:26:39.032000",
  "votes": 49,
  "comment_count": 38,
  "views": 0,
  "content": "<p>I want to share some thoughts and experience about this competition after several weeks of taking part in it. They are subjective and could be not entirely correct in other experiments:\n- My approach is almost the same as in my public kernel;\n- I still use <code>BCEDiceLoss</code>. <code>BCEJaccardLoss</code> or simple <code>BCE</code> or simple <code>Dice</code> work worse for me;\n- Higher score on validation doesn't always mean higher score on leaderboard. Maybe the problem is in my validation split, maybe test distribution is different;\n- Using deeper backbones helps, but only little. Tried <code>resnet152</code> and <code>se_resnext50_32x4d</code>;\n- I'm experimenting with augmentations. It seems some of them either cause overfitting or worsen the model;\n- Catalyst is a very convenient framework and makes many things easier;\n- Using classifier to predict which masks should have masks boosts score!\n- I suppose some better/different post-processing is necessary...\n<img src=\"https://i.imgur.com/fhmwcUw.png\" alt=\"\">\n- RAdam optimizer is better than simple Adam;</p>\n\n<p>My loss and dice curves:\n<img src=\"https://i.imgur.com/rtsm4Th.png\" alt=\"\"></p>\n\n<p>Some of the things which I plan to try in future:\n- different losses (like correct multiclass dice);\n- use different model architectures;\n- take some ideas from severstal competition;</p>",
  "messages": [
    {
      "id": 634769,
      "postDate": "2019-09-26T18:26:39.033Z",
      "content": "<p>I want to share some thoughts and experience about this competition after several weeks of taking part in it. They are subjective and could be not entirely correct in other experiments:\n- My approach is almost the same as in my public kernel;\n- I still use <code>BCEDiceLoss</code>. <code>BCEJaccardLoss</code> or simple <code>BCE</code> or simple <code>Dice</code> work worse for me;\n- Higher score on validation doesn't always mean higher score on leaderboard. Maybe the problem is in my validation split, maybe test distribution is different;\n- Using deeper backbones helps, but only little. Tried <code>resnet152</code> and <code>se_resnext50_32x4d</code>;\n- I'm experimenting with augmentations. It seems some of them either cause overfitting or worsen the model;\n- Catalyst is a very convenient framework and makes many things easier;\n- Using classifier to predict which masks should have masks boosts score!\n- I suppose some better/different post-processing is necessary...\n<img src=\"https://i.imgur.com/fhmwcUw.png\" alt=\"\">\n- RAdam optimizer is better than simple Adam;</p>\n\n<p>My loss and dice curves:\n<img src=\"https://i.imgur.com/rtsm4Th.png\" alt=\"\"></p>\n\n<p>Some of the things which I plan to try in future:\n- different losses (like correct multiclass dice);\n- use different model architectures;\n- take some ideas from severstal competition;</p>",
      "rawMarkdown": "I want to share some thoughts and experience about this competition after several weeks of taking part in it. They are subjective and could be not entirely correct in other experiments:\n- My approach is almost the same as in my public kernel;\n- I still use `BCEDiceLoss`. `BCEJaccardLoss` or simple `BCE` or simple `Dice` work worse for me;\n- Higher score on validation doesn't always mean higher score on leaderboard. Maybe the problem is in my validation split, maybe test distribution is different;\n- Using deeper backbones helps, but only little. Tried `resnet152` and `se_resnext50_32x4d`;\n- I'm experimenting with augmentations. It seems some of them either cause overfitting or worsen the model;\n- Catalyst is a very convenient framework and makes many things easier;\n- Using classifier to predict which masks should have masks boosts score!\n- I suppose some better/different post-processing is necessary...\n![](https://i.imgur.com/fhmwcUw.png)\n- RAdam optimizer is better than simple Adam;\n\nMy loss and dice curves:\n![](https://i.imgur.com/rtsm4Th.png)\n\nSome of the things which I plan to try in future:\n- different losses (like correct multiclass dice);\n- use different model architectures;\n- take some ideas from severstal competition;",
      "votes": 49
    },
    {
      "id": 635262,
      "postDate": "2019-09-27T10:19:28.063Z",
      "content": "<p>Did you apply RAdam from here: <a href=\"https://github.com/LiyuanLucasLiu/RAdam\">https://github.com/LiyuanLucasLiu/RAdam</a> ? And did you use the same parameters as you did in your Adam optimizer?  </p>\n\n<p>Oh, and thank you for sharing all this information and your kernel! :)</p>",
      "rawMarkdown": "Did you apply RAdam from here: [https://github.com/LiyuanLucasLiu/RAdam](https://github.com/LiyuanLucasLiu/RAdam) ? And did you use the same parameters as you did in your Adam optimizer?  \n\nOh, and thank you for sharing all this information and your kernel! :)",
      "votes": 3,
      "replies": [
        {
          "id": 635283,
          "postDate": "2019-09-27T10:34:09.640Z",
          "content": "<p>Yes, I used this implementation of RAdam with the same parameters.</p>",
          "rawMarkdown": "Yes, I used this implementation of RAdam with the same parameters.",
          "votes": 4
        },
        {
          "id": 635286,
          "postDate": "2019-09-27T10:36:23.320Z",
          "content": "<p>Okay, thank you.</p>",
          "rawMarkdown": "Okay, thank you.",
          "votes": 1
        },
        {
          "id": 637899,
          "postDate": "2019-10-01T10:59:33.063Z",
          "content": "<p><a href=\"/artgor\">@artgor</a> Andrew, for some reason when I plugged in RAdam implementation with the same parameters into your kernel LB score got worse. Are there any tricks for using it?</p>",
          "rawMarkdown": "@artgor Andrew, for some reason when I plugged in RAdam implementation with the same parameters into your kernel LB score got worse. Are there any tricks for using it?",
          "votes": 1
        },
        {
          "id": 637929,
          "postDate": "2019-10-01T11:31:01.630Z",
          "content": "<p>I'm not sure... when I train locally, I usually decrease the value of <code>factor</code> in reduceonplateu and train longer. Maybe this was the reason.\nAnd I think I decreased decoder lr to 1e-3 (I used the same in Adam and Radam locally).</p>",
          "rawMarkdown": "I'm not sure... when I train locally, I usually decrease the value of `factor` in reduceonplateu and train longer. Maybe this was the reason.\nAnd I think I decreased decoder lr to 1e-3 (I used the same in Adam and Radam locally).",
          "votes": 2
        }
      ]
    },
    {
      "id": 635183,
      "postDate": "2019-09-27T08:35:28.240Z",
      "content": "<p><a href=\"/artgor\">@artgor</a>, I really liked the topic, please share your future findings too!</p>",
      "rawMarkdown": "@artgor, I really liked the topic, please share your future findings too!",
      "votes": 1
    },
    {
      "id": 634805,
      "postDate": "2019-09-26T19:25:22.807Z",
      "content": "<p>If you don't mind sharing, is classifying before masking already in your current working pipeline? I tried doing a classifier and then have the thresholds searched above and below that binary classification threshold and it improved local validation but always decreases lb. Not sure if you might be doing something less prone to overfitting. </p>",
      "rawMarkdown": "If you don't mind sharing, is classifying before masking already in your current working pipeline? I tried doing a classifier and then have the thresholds searched above and below that binary classification threshold and it improved local validation but always decreases lb. Not sure if you might be doing something less prone to overfitting. ",
      "votes": 1,
      "replies": [
        {
          "id": 634807,
          "postDate": "2019-09-26T19:34:16.377Z",
          "content": "<p>I tried classifier for the first time today and use a very basic solution now:\n- take my submission with <code>resnet152</code> and lb <code>0.648</code>;\n- take submission from this kernel with classifier: <a href=\"https://www.kaggle.com/samusram/cloud-classifier-for-post-processing\">https://www.kaggle.com/samusram/cloud-classifier-for-post-processing</a>;\n- replace my predictions with empty masks where prediction is empty in the classifier-submission. This way I reduce the number of false positives;\n- submit and get <code>0.657</code>\nSo I apply classifier even after applying postprocessing (like in my kernel)</p>\n\n<p>Now I'll try training my own classifier. Maybe it could be efficiently combined with segmentation model (as far as I know, in severstal it is better to keep them separate).</p>",
          "rawMarkdown": "I tried classifier for the first time today and use a very basic solution now:\n- take my submission with `resnet152` and lb `0.648`;\n- take submission from this kernel with classifier: https://www.kaggle.com/samusram/cloud-classifier-for-post-processing;\n- replace my predictions with empty masks where prediction is empty in the classifier-submission. This way I reduce the number of false positives;\n- submit and get `0.657`\nSo I apply classifier even after applying postprocessing (like in my kernel)\n\nNow I'll try training my own classifier. Maybe it could be efficiently combined with segmentation model (as far as I know, in severstal it is better to keep them separate).",
          "votes": 4
        },
        {
          "id": 634913,
          "postDate": "2019-09-27T00:08:46.977Z",
          "content": "<p>I toyed with various configurations in the pneumothorax challenge and didnt have any stellar return from it. Reducing false positives is very important here, but I found that the classifiers I trained had very similar false positives as my segmentation mask predictions. Might have to revisit it though. </p>",
          "rawMarkdown": "I toyed with various configurations in the pneumothorax challenge and didnt have any stellar return from it. Reducing false positives is very important here, but I found that the classifiers I trained had very similar false positives as my segmentation mask predictions. Might have to revisit it though. ",
          "votes": 1
        }
      ]
    },
    {
      "id": 636698,
      "postDate": "2019-09-30T04:13:21.777Z",
      "content": "<p>Interesting, thanks for sharing! I also have to agree on the points you have made. My latest boost was from using the largest encoder I could handle, yet barely provided a boost (~0.002?), not too sure what is going on there. To make things wore, it saw larger reduction in validation than and that of rise in LB, which I do not like to trust. I think it just got lucky. However, my recent 5 fold average score saw somewhat similar values with the public LB, so I hope to continue working on using it.</p>\n\n<p>Also... the fact that the nature of this competition is different from a typical segmentation task (more like bounding box labels) adds another layer of confusion and unknown.  :(</p>",
      "rawMarkdown": "Interesting, thanks for sharing! I also have to agree on the points you have made. My latest boost was from using the largest encoder I could handle, yet barely provided a boost (~0.002?), not too sure what is going on there. To make things wore, it saw larger reduction in validation than and that of rise in LB, which I do not like to trust. I think it just got lucky. However, my recent 5 fold average score saw somewhat similar values with the public LB, so I hope to continue working on using it.\n\nAlso... the fact that the nature of this competition is different from a typical segmentation task (more like bounding box labels) adds another layer of confusion and unknown.  :(",
      "votes": 2,
      "replies": [
        {
          "id": 636727,
          "postDate": "2019-09-30T05:30:56.400Z",
          "content": "<p>How did you calculate 5 fold average score in this case? I'm not able to understand how ensembling works in segmentation task! Also, what do you mean by a largest encoder, highest number of parameters or maximum number of layers?  </p>",
          "rawMarkdown": "How did you calculate 5 fold average score in this case? I'm not able to understand how ensembling works in segmentation task! Also, what do you mean by a largest encoder, highest number of parameters or maximum number of layers?  "
        },
        {
          "id": 636728,
          "postDate": "2019-09-30T05:33:34.587Z",
          "content": "<p>There are various ways of combining segmentation models. For example you can simply average the raw predictions and find best threshold after this.</p>\n\n<p>As for encoders - there are many encoders. One of the smallest is resnet18. Resnet152 is obviously bigger. And there are others, like senet154.</p>",
          "rawMarkdown": "There are various ways of combining segmentation models. For example you can simply average the raw predictions and find best threshold after this.\n\nAs for encoders - there are many encoders. One of the smallest is resnet18. Resnet152 is obviously bigger. And there are others, like senet154.",
          "votes": 1
        },
        {
          "id": 636766,
          "postDate": "2019-09-30T07:05:51.830Z",
          "content": "<p>So this week I’ve raised my local cross validation score from .650-.655 to .683+ with model not post process changes. But my LB score remains .665-.670 before and after. </p>\n\n<p>All is good with val set, etc, no leakage, no overfitting signs. I can only imagine that single fold model strength is now as good as a 5-fold model. Unlikely. Or a bug. The hunt goes on...\nEDIT: Reason/Bug found :)</p>",
          "rawMarkdown": "So this week I’ve raised my local cross validation score from .650-.655 to .683+ with model not post process changes. But my LB score remains .665-.670 before and after. \n\nAll is good with val set, etc, no leakage, no overfitting signs. I can only imagine that single fold model strength is now as good as a 5-fold model. Unlikely. Or a bug. The hunt goes on...\nEDIT: Reason/Bug found :)",
          "votes": 1
        },
        {
          "id": 636796,
          "postDate": "2019-09-30T07:58:34.977Z",
          "content": "<p><a href=\"/artgor\">@artgor</a>  I've experimented with different encoders like resnet34, resnet50, Densenet201, seresnet34 etc with Unet and FPN architecture (keeping all other parameters same) but they all are almost similar in terms of raw_predictions, in fact, my best raw prediction is using resent34. Am I missing something?</p>",
          "rawMarkdown": "@artgor  I've experimented with different encoders like resnet34, resnet50, Densenet201, seresnet34 etc with Unet and FPN architecture (keeping all other parameters same) but they all are almost similar in terms of raw_predictions, in fact, my best raw prediction is using resent34. Am I missing something?"
        },
        {
          "id": 637675,
          "postDate": "2019-10-01T07:36:54.293Z",
          "content": "<p>Fyi I am using SEnet154, but after some CV (I just do 5 fold to get val/LB score, and also get an average just for further comparison) I am also thinking a single OOF is quite consistent... struggling with finding a good way to verify improvement tbh</p>",
          "rawMarkdown": "Fyi I am using SEnet154, but after some CV (I just do 5 fold to get val/LB score, and also get an average just for further comparison) I am also thinking a single OOF is quite consistent... struggling with finding a good way to verify improvement tbh",
          "votes": 2
        },
        {
          "id": 637687,
          "postDate": "2019-10-01T07:46:10.133Z",
          "content": "<p>I've been exclusively using a single fold and have had basically perfect scaling from Val to lb. My fear at this point though is we aren't trying to segment clouds better, we're trying to match the extremely noisy annotations better. Hence no real gain occurring from bigger networks or higher res images or different models. </p>",
          "rawMarkdown": "I've been exclusively using a single fold and have had basically perfect scaling from Val to lb. My fear at this point though is we aren't trying to segment clouds better, we're trying to match the extremely noisy annotations better. Hence no real gain occurring from bigger networks or higher res images or different models. ",
          "votes": 1
        },
        {
          "id": 637689,
          "postDate": "2019-10-01T07:47:17.860Z",
          "content": "<p>Not sure if I'm going to continue on this competition because of that. </p>",
          "rawMarkdown": "Not sure if I'm going to continue on this competition because of that. ",
          "votes": 1
        },
        {
          "id": 637761,
          "postDate": "2019-10-01T08:37:47.750Z",
          "content": "<p>you are right <a href=\"/ryches\">@ryches</a> </p>",
          "rawMarkdown": "you are right @ryches "
        },
        {
          "id": 638320,
          "postDate": "2019-10-01T19:05:47.817Z",
          "content": "<p>I have seen improvement with image resolution. Specifically, I went from 128*192 to 256*384 and improved LB (and CV) by roughly 0.03. I'm trying an even higher resolution now (taking awhile!). Perhaps it's because my score is overall in a lower range? Or maybe fluke?</p>",
          "rawMarkdown": "I have seen improvement with image resolution. Specifically, I went from 128*192 to 256*384 and improved LB (and CV) by roughly 0.03. I'm trying an even higher resolution now (taking awhile!). Perhaps it's because my score is overall in a lower range? Or maybe fluke?"
        },
        {
          "id": 642546,
          "postDate": "2019-10-06T09:17:33.740Z",
          "rawMarkdown": "",
          "votes": -2
        }
      ]
    },
    {
      "id": 635948,
      "postDate": "2019-09-28T14:00:22.073Z",
      "content": "<p>One question for the people with more experience in the matter, how do you usually evaluate a specific augmentation method?</p>\n\n<p>You use only that method and see how it performs? Keep the previous ones and add the new? Sometimes I feel hard to evaluate this because of all the randomness. </p>",
      "rawMarkdown": "One question for the people with more experience in the matter, how do you usually evaluate a specific augmentation method?\n\nYou use only that method and see how it performs? Keep the previous ones and add the new? Sometimes I feel hard to evaluate this because of all the randomness. ",
      "votes": 2,
      "replies": [
        {
          "id": 643016,
          "postDate": "2019-10-07T00:25:56.207Z",
          "content": "<p>I found this interesting I haven't tried it yet. <a href=\"https://www.youtube.com/watch?v=c-oth0OWv3Y\">Data Distribution Search: Deep Reinforcement Learning To Improvise Input Datasets</a></p>",
          "rawMarkdown": "I found this interesting I haven't tried it yet. [Data Distribution Search: Deep Reinforcement Learning To Improvise Input Datasets](https://www.youtube.com/watch?v=c-oth0OWv3Y)",
          "votes": 1
        },
        {
          "id": 643017,
          "postDate": "2019-10-07T00:37:04.530Z",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> I try to experiment +-1 if it's new aug, if its the values I modify one or more at once. Current OOF val works well enough, but if you are concerned about stochasticity I'm sure Kfold provides good enough evaluation (compare Dice, etc).</p>",
          "rawMarkdown": "@dimitreoliveira I try to experiment +-1 if it's new aug, if its the values I modify one or more at once. Current OOF val works well enough, but if you are concerned about stochasticity I'm sure Kfold provides good enough evaluation (compare Dice, etc).",
          "votes": 3
        },
        {
          "id": 643032,
          "postDate": "2019-10-07T01:43:13.527Z",
          "content": "<p>Thanks <a href=\"/vivekwisdom\">@vivekwisdom</a> , it seems very interesting, I would also love to see an implementation, once I saw a library used by Google that also used reinforcement learning to improve augmentations, but never saw an open implementation of it.</p>",
          "rawMarkdown": "Thanks @vivekwisdom , it seems very interesting, I would also love to see an implementation, once I saw a library used by Google that also used reinforcement learning to improve augmentations, but never saw an open implementation of it."
        },
        {
          "id": 643033,
          "postDate": "2019-10-07T01:46:40.700Z",
          "content": "<p>Thanks <a href=\"/joonl04\">@joonl04</a> , I'm also trying +-1 on new augs, I agree that K-Fold should give more confidence, unfortunately with the new GPU quota too many experimentations are harder 😄 </p>",
          "rawMarkdown": "Thanks @joonl04 , I'm also trying +-1 on new augs, I agree that K-Fold should give more confidence, unfortunately with the new GPU quota too many experimentations are harder 😄 "
        },
        {
          "id": 643232,
          "postDate": "2019-10-07T09:42:17.327Z",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> I guess you're referring to this paper <a href=\"https://arxiv.org/pdf/1805.09501.pdf\">AutoAugment: Learning Augmentation Strategies from Data</a>, they have <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">EfficientNet release for TPU with AutoAugment</a>.</p>",
          "rawMarkdown": "@dimitreoliveira I guess you're referring to this paper [AutoAugment: Learning Augmentation Strategies from Data](https://arxiv.org/pdf/1805.09501.pdf), they have [EfficientNet release for TPU with AutoAugment](https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet).",
          "votes": 1
        },
        {
          "id": 643282,
          "postDate": "2019-10-07T11:19:39.247Z",
          "content": "<p>Yes <a href=\"/vivekwisdom\">@vivekwisdom</a> , that's the one I was thinking, do you think we can run this on Google colab?</p>",
          "rawMarkdown": "Yes @vivekwisdom , that's the one I was thinking, do you think we can run this on Google colab?",
          "votes": 1
        },
        {
          "id": 643328,
          "postDate": "2019-10-07T12:11:51.267Z",
          "content": "<p>I don't think so. According to the paper, they needed 15000 GPU hours for imagenet. Unless we want to run for months. You can use pre-trained weights from Google though, link in repo readme.</p>",
          "rawMarkdown": "I don't think so. According to the paper, they needed 15000 GPU hours for imagenet. Unless we want to run for months. You can use pre-trained weights from Google though, link in repo readme.",
          "votes": 1
        }
      ]
    },
    {
      "id": 635280,
      "postDate": "2019-09-27T10:31:58.323Z",
      "content": "<p>My experience on most segmentation tasks is the segmenter with threshold is as good a classifier as any separate classifier. Though you can train a good set of classifiers early on to maybe get a headstart. It can also help if you have an inference time limit but we are OK here. I am just using a segmentation model.</p>",
      "rawMarkdown": "My experience on most segmentation tasks is the segmenter with threshold is as good a classifier as any separate classifier. Though you can train a good set of classifiers early on to maybe get a headstart. It can also help if you have an inference time limit but we are OK here. I am just using a segmentation model.",
      "votes": 2,
      "replies": [
        {
          "id": 635284,
          "postDate": "2019-09-27T10:35:15.087Z",
          "content": "<p>I suppose you are correct. Maybe I need better preprocessing or a different loss. Also I didn't exclude bad images yet - maybe this is the problem.</p>",
          "rawMarkdown": "I suppose you are correct. Maybe I need better preprocessing or a different loss. Also I didn't exclude bad images yet - maybe this is the problem.\n",
          "votes": 1
        },
        {
          "id": 638476,
          "postDate": "2019-10-02T00:21:03.430Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 634777,
      "postDate": "2019-09-26T18:40:51.290Z",
      "content": "<p>hi sir,thank you for this post,you are absolutely right,i will say \"taking some ideas from severstal competitions is great idea\" but i can't find any competition related to this one,here we are predicting masks on small scale and training on large size dataset,i played a lot with your kernel and i observe the loss function doesn't converge and i don't care the public lb,our all models are overfitted,even if we get medal,i am sure our models overfitted, only thing i am planing to do next is \"spending a lot of time on the dataset to find some leaks and problems\",your help is highly appreciated</p>\n\n<p>i am planning to implement <a href=\"https://www.kaggle.com/rishabhiitbhu/unet-starter-kernel-pytorch-lb-0-88\">this kernel</a> in this competition? do you recommend it or something else?</p>\n\n<p>thank you in advance</p>",
      "rawMarkdown": "hi sir,thank you for this post,you are absolutely right,i will say \"taking some ideas from severstal competitions is great idea\" but i can't find any competition related to this one,here we are predicting masks on small scale and training on large size dataset,i played a lot with your kernel and i observe the loss function doesn't converge and i don't care the public lb,our all models are overfitted,even if we get medal,i am sure our models overfitted, only thing i am planing to do next is \"spending a lot of time on the dataset to find some leaks and problems\",your help is highly appreciated\n\ni am planning to implement [this kernel](https://www.kaggle.com/rishabhiitbhu/unet-starter-kernel-pytorch-lb-0-88) in this competition? do you recommend it or something else?\n\n\nthank you in advance",
      "votes": 2,
      "replies": [
        {
          "id": 634795,
          "postDate": "2019-09-26T19:07:09.343Z",
          "content": "<p>Why do you think there is an overfit?</p>\n\n<p>I think that kernel is okay.</p>",
          "rawMarkdown": "Why do you think there is an overfit?\n\nI think that kernel is okay.",
          "votes": 1
        },
        {
          "id": 634804,
          "postDate": "2019-09-26T19:23:46.087Z",
          "content": "<p>I don't perceive it to be overfitting either. I have almost 1:1 matching between local validation and lb. Always within +-.003. That's pretty close considering I get similarity across runs. </p>",
          "rawMarkdown": "I don't perceive it to be overfitting either. I have almost 1:1 matching between local validation and lb. Always within +-.003. That's pretty close considering I get similarity across runs. "
        },
        {
          "id": 635170,
          "postDate": "2019-09-27T08:19:40.727Z",
          "content": "<p><a href=\"/artgor\">@artgor</a>  <a href=\"/ryches\">@ryches</a>  perhaps you guys are right but please plot the learning curve,from the attached picture below,we can see gap and pure inconsistency between train and validation which should describe overfitting(this is what i learnt from andrew ng's class),i feel the model is learning from noise and should not generalize,sorry if i am making mistakes,just shared my thought..!!! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F1e561067a36a6f007377a9374e47e487%2Flc.PNG?generation=1569572249738691&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "@artgor  @ryches  perhaps you guys are right but please plot the learning curve,from the attached picture below,we can see gap and pure inconsistency between train and validation which should describe overfitting(this is what i learnt from andrew ng's class),i feel the model is learning from noise and should not generalize,sorry if i am making mistakes,just shared my thought..!!! ![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F1e561067a36a6f007377a9374e47e487%2Flc.PNG?generation=1569572249738691&amp;alt=media)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 635572,
      "postDate": "2019-09-27T21:01:31.727Z",
      "content": "<p>Thanks for sharing those thoughts about your work!\nI'm not sure I fully understand, what do you mean by using a classifier to determine which masks should have masks boost score?</p>",
      "rawMarkdown": "Thanks for sharing those thoughts about your work!\nI'm not sure I fully understand, what do you mean by using a classifier to determine which masks should have masks boost score?",
      "replies": [
        {
          "id": 635719,
          "postDate": "2019-09-28T04:53:11.540Z",
          "content": "<p>He is saying to train a model that doesn't necessarily try to pin point where the clouds are but rather just makes the determination whether a cloud of the certain type was present or not in the image. You can then use this classifier as some additional postprocessing on top of the model you train for locating the clouds. In this case andrew is using it in order to remove any masks that had classifications below a certain threshold from the classifying cloud/no-cloud model</p>",
          "rawMarkdown": "He is saying to train a model that doesn't necessarily try to pin point where the clouds are but rather just makes the determination whether a cloud of the certain type was present or not in the image. You can then use this classifier as some additional postprocessing on top of the model you train for locating the clouds. In this case andrew is using it in order to remove any masks that had classifications below a certain threshold from the classifying cloud/no-cloud model",
          "votes": 4
        },
        {
          "id": 635780,
          "postDate": "2019-09-28T07:38:11.200Z",
          "content": "<p>Makes sense, thanks for the info!</p>",
          "rawMarkdown": "Makes sense, thanks for the info!"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 635262,
      "author_name": "Christoffer G",
      "author_url": "",
      "post_date": "2019-09-27T10:19:28.063000",
      "content": "<p>Did you apply RAdam from here: <a href=\"https://github.com/LiyuanLucasLiu/RAdam\">https://github.com/LiyuanLucasLiu/RAdam</a> ? And did you use the same parameters as you did in your Adam optimizer?  </p>\n\n<p>Oh, and thank you for sharing all this information and your kernel! :)</p>",
      "votes": 3,
      "replies": [
        {
          "id": 635283,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-27T10:34:09.640000",
          "content": "<p>Yes, I used this implementation of RAdam with the same parameters.</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 635286,
          "author_name": "Christoffer G",
          "author_url": "",
          "post_date": "2019-09-27T10:36:23.320000",
          "content": "<p>Okay, thank you.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 637899,
          "author_name": "Anna Novikova",
          "author_url": "",
          "post_date": "2019-10-01T10:59:33.063000",
          "content": "<p><a href=\"/artgor\">@artgor</a> Andrew, for some reason when I plugged in RAdam implementation with the same parameters into your kernel LB score got worse. Are there any tricks for using it?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 637929,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-10-01T11:31:01.630000",
          "content": "<p>I'm not sure... when I train locally, I usually decrease the value of <code>factor</code> in reduceonplateu and train longer. Maybe this was the reason.\nAnd I think I decreased decoder lr to 1e-3 (I used the same in Adam and Radam locally).</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 635183,
      "author_name": "Brun",
      "author_url": "",
      "post_date": "2019-09-27T08:35:28.240000",
      "content": "<p><a href=\"/artgor\">@artgor</a>, I really liked the topic, please share your future findings too!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 634805,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-09-26T19:25:22.807000",
      "content": "<p>If you don't mind sharing, is classifying before masking already in your current working pipeline? I tried doing a classifier and then have the thresholds searched above and below that binary classification threshold and it improved local validation but always decreases lb. Not sure if you might be doing something less prone to overfitting. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 634807,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-26T19:34:16.377000",
          "content": "<p>I tried classifier for the first time today and use a very basic solution now:\n- take my submission with <code>resnet152</code> and lb <code>0.648</code>;\n- take submission from this kernel with classifier: <a href=\"https://www.kaggle.com/samusram/cloud-classifier-for-post-processing\">https://www.kaggle.com/samusram/cloud-classifier-for-post-processing</a>;\n- replace my predictions with empty masks where prediction is empty in the classifier-submission. This way I reduce the number of false positives;\n- submit and get <code>0.657</code>\nSo I apply classifier even after applying postprocessing (like in my kernel)</p>\n\n<p>Now I'll try training my own classifier. Maybe it could be efficiently combined with segmentation model (as far as I know, in severstal it is better to keep them separate).</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 634913,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-27T00:08:46.977000",
          "content": "<p>I toyed with various configurations in the pneumothorax challenge and didnt have any stellar return from it. Reducing false positives is very important here, but I found that the classifiers I trained had very similar false positives as my segmentation mask predictions. Might have to revisit it though. </p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 636698,
      "author_name": "Brian Lee",
      "author_url": "",
      "post_date": "2019-09-30T04:13:21.777000",
      "content": "<p>Interesting, thanks for sharing! I also have to agree on the points you have made. My latest boost was from using the largest encoder I could handle, yet barely provided a boost (~0.002?), not too sure what is going on there. To make things wore, it saw larger reduction in validation than and that of rise in LB, which I do not like to trust. I think it just got lucky. However, my recent 5 fold average score saw somewhat similar values with the public LB, so I hope to continue working on using it.</p>\n\n<p>Also... the fact that the nature of this competition is different from a typical segmentation task (more like bounding box labels) adds another layer of confusion and unknown.  :(</p>",
      "votes": 2,
      "replies": [
        {
          "id": 636727,
          "author_name": "0DD1",
          "author_url": "",
          "post_date": "2019-09-30T05:30:56.400000",
          "content": "<p>How did you calculate 5 fold average score in this case? I'm not able to understand how ensembling works in segmentation task! Also, what do you mean by a largest encoder, highest number of parameters or maximum number of layers?  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 636728,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-30T05:33:34.587000",
          "content": "<p>There are various ways of combining segmentation models. For example you can simply average the raw predictions and find best threshold after this.</p>\n\n<p>As for encoders - there are many encoders. One of the smallest is resnet18. Resnet152 is obviously bigger. And there are others, like senet154.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636766,
          "author_name": "robga",
          "author_url": "",
          "post_date": "2019-09-30T07:05:51.830000",
          "content": "<p>So this week I’ve raised my local cross validation score from .650-.655 to .683+ with model not post process changes. But my LB score remains .665-.670 before and after. </p>\n\n<p>All is good with val set, etc, no leakage, no overfitting signs. I can only imagine that single fold model strength is now as good as a 5-fold model. Unlikely. Or a bug. The hunt goes on...\nEDIT: Reason/Bug found :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 636796,
          "author_name": "0DD1",
          "author_url": "",
          "post_date": "2019-09-30T07:58:34.977000",
          "content": "<p><a href=\"/artgor\">@artgor</a>  I've experimented with different encoders like resnet34, resnet50, Densenet201, seresnet34 etc with Unet and FPN architecture (keeping all other parameters same) but they all are almost similar in terms of raw_predictions, in fact, my best raw prediction is using resent34. Am I missing something?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 637675,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2019-10-01T07:36:54.293000",
          "content": "<p>Fyi I am using SEnet154, but after some CV (I just do 5 fold to get val/LB score, and also get an average just for further comparison) I am also thinking a single OOF is quite consistent... struggling with finding a good way to verify improvement tbh</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 637687,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-10-01T07:46:10.133000",
          "content": "<p>I've been exclusively using a single fold and have had basically perfect scaling from Val to lb. My fear at this point though is we aren't trying to segment clouds better, we're trying to match the extremely noisy annotations better. Hence no real gain occurring from bigger networks or higher res images or different models. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 637689,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-10-01T07:47:17.860000",
          "content": "<p>Not sure if I'm going to continue on this competition because of that. </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 637761,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-10-01T08:37:47.750000",
          "content": "<p>you are right <a href=\"/ryches\">@ryches</a> </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 638320,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2019-10-01T19:05:47.817000",
          "content": "<p>I have seen improvement with image resolution. Specifically, I went from 128*192 to 256*384 and improved LB (and CV) by roughly 0.03. I'm trying an even higher resolution now (taking awhile!). Perhaps it's because my score is overall in a lower range? Or maybe fluke?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 642546,
          "author_name": "Anurag Trivedi",
          "author_url": "",
          "post_date": "2019-10-06T09:17:33.740000",
          "content": "",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 635948,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2019-09-28T14:00:22.073000",
      "content": "<p>One question for the people with more experience in the matter, how do you usually evaluate a specific augmentation method?</p>\n\n<p>You use only that method and see how it performs? Keep the previous ones and add the new? Sometimes I feel hard to evaluate this because of all the randomness. </p>",
      "votes": 2,
      "replies": [
        {
          "id": 643016,
          "author_name": "sapientiae",
          "author_url": "",
          "post_date": "2019-10-07T00:25:56.207000",
          "content": "<p>I found this interesting I haven't tried it yet. <a href=\"https://www.youtube.com/watch?v=c-oth0OWv3Y\">Data Distribution Search: Deep Reinforcement Learning To Improvise Input Datasets</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 643017,
          "author_name": "Brian Lee",
          "author_url": "",
          "post_date": "2019-10-07T00:37:04.530000",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> I try to experiment +-1 if it's new aug, if its the values I modify one or more at once. Current OOF val works well enough, but if you are concerned about stochasticity I'm sure Kfold provides good enough evaluation (compare Dice, etc).</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 643032,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-10-07T01:43:13.527000",
          "content": "<p>Thanks <a href=\"/vivekwisdom\">@vivekwisdom</a> , it seems very interesting, I would also love to see an implementation, once I saw a library used by Google that also used reinforcement learning to improve augmentations, but never saw an open implementation of it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643033,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-10-07T01:46:40.700000",
          "content": "<p>Thanks <a href=\"/joonl04\">@joonl04</a> , I'm also trying +-1 on new augs, I agree that K-Fold should give more confidence, unfortunately with the new GPU quota too many experimentations are harder 😄 </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 643232,
          "author_name": "sapientiae",
          "author_url": "",
          "post_date": "2019-10-07T09:42:17.327000",
          "content": "<p><a href=\"/dimitreoliveira\">@dimitreoliveira</a> I guess you're referring to this paper <a href=\"https://arxiv.org/pdf/1805.09501.pdf\">AutoAugment: Learning Augmentation Strategies from Data</a>, they have <a href=\"https://github.com/tensorflow/tpu/tree/master/models/official/efficientnet\">EfficientNet release for TPU with AutoAugment</a>.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 643282,
          "author_name": "DimitreOliveira",
          "author_url": "",
          "post_date": "2019-10-07T11:19:39.247000",
          "content": "<p>Yes <a href=\"/vivekwisdom\">@vivekwisdom</a> , that's the one I was thinking, do you think we can run this on Google colab?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 643328,
          "author_name": "sapientiae",
          "author_url": "",
          "post_date": "2019-10-07T12:11:51.267000",
          "content": "<p>I don't think so. According to the paper, they needed 15000 GPU hours for imagenet. Unless we want to run for months. You can use pre-trained weights from Google though, link in repo readme.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 635280,
      "author_name": "robga",
      "author_url": "",
      "post_date": "2019-09-27T10:31:58.323000",
      "content": "<p>My experience on most segmentation tasks is the segmenter with threshold is as good a classifier as any separate classifier. Though you can train a good set of classifiers early on to maybe get a headstart. It can also help if you have an inference time limit but we are OK here. I am just using a segmentation model.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 635284,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-27T10:35:15.087000",
          "content": "<p>I suppose you are correct. Maybe I need better preprocessing or a different loss. Also I didn't exclude bad images yet - maybe this is the problem.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 638476,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-10-02T00:21:03.430000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 634777,
      "author_name": "Mobassir",
      "author_url": "",
      "post_date": "2019-09-26T18:40:51.290000",
      "content": "<p>hi sir,thank you for this post,you are absolutely right,i will say \"taking some ideas from severstal competitions is great idea\" but i can't find any competition related to this one,here we are predicting masks on small scale and training on large size dataset,i played a lot with your kernel and i observe the loss function doesn't converge and i don't care the public lb,our all models are overfitted,even if we get medal,i am sure our models overfitted, only thing i am planing to do next is \"spending a lot of time on the dataset to find some leaks and problems\",your help is highly appreciated</p>\n\n<p>i am planning to implement <a href=\"https://www.kaggle.com/rishabhiitbhu/unet-starter-kernel-pytorch-lb-0-88\">this kernel</a> in this competition? do you recommend it or something else?</p>\n\n<p>thank you in advance</p>",
      "votes": 2,
      "replies": [
        {
          "id": 634795,
          "author_name": "Andrey Lukyanenko",
          "author_url": "",
          "post_date": "2019-09-26T19:07:09.343000",
          "content": "<p>Why do you think there is an overfit?</p>\n\n<p>I think that kernel is okay.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 634804,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-26T19:23:46.087000",
          "content": "<p>I don't perceive it to be overfitting either. I have almost 1:1 matching between local validation and lb. Always within +-.003. That's pretty close considering I get similarity across runs. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 635170,
          "author_name": "Mobassir",
          "author_url": "",
          "post_date": "2019-09-27T08:19:40.727000",
          "content": "<p><a href=\"/artgor\">@artgor</a>  <a href=\"/ryches\">@ryches</a>  perhaps you guys are right but please plot the learning curve,from the attached picture below,we can see gap and pure inconsistency between train and validation which should describe overfitting(this is what i learnt from andrew ng's class),i feel the model is learning from noise and should not generalize,sorry if i am making mistakes,just shared my thought..!!! <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2034058%2F1e561067a36a6f007377a9374e47e487%2Flc.PNG?generation=1569572249738691&amp;alt=media\" alt=\"\"></p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 635572,
      "author_name": "Maxime Lenormand",
      "author_url": "",
      "post_date": "2019-09-27T21:01:31.727000",
      "content": "<p>Thanks for sharing those thoughts about your work!\nI'm not sure I fully understand, what do you mean by using a classifier to determine which masks should have masks boost score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 635719,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "2019-09-28T04:53:11.540000",
          "content": "<p>He is saying to train a model that doesn't necessarily try to pin point where the clouds are but rather just makes the determination whether a cloud of the certain type was present or not in the image. You can then use this classifier as some additional postprocessing on top of the model you train for locating the clouds. In this case andrew is using it in order to remove any masks that had classifications below a certain threshold from the classifying cloud/no-cloud model</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 635780,
          "author_name": "Maxime Lenormand",
          "author_url": "",
          "post_date": "2019-09-28T07:38:11.200000",
          "content": "<p>Makes sense, thanks for the info!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "634769": "I want to share some thoughts and experience about this competition after several weeks of taking part in it. They are subjective and could be not entirely correct in other experiments:\n- My approach is almost the same as in my public kernel;\n- I still use `BCEDiceLoss`. `BCEJaccardLoss` or simple `BCE` or simple `Dice` work worse for me;\n- Higher score on validation doesn't always mean higher score on leaderboard. Maybe the problem is in my validation split, maybe test distribution is different;\n- Using deeper backbones helps, but only little. Tried `resnet152` and `se_resnext50_32x4d`;\n- I'm experimenting with augmentations. It seems some of them either cause overfitting or worsen the model;\n- Catalyst is a very convenient framework and makes many things easier;\n- Using classifier to predict which masks should have masks boosts score!\n- I suppose some better/different post-processing is necessary...\n![](https://i.imgur.com/fhmwcUw.png)\n- RAdam optimizer is better than simple Adam;\n\nMy loss and dice curves:\n![](https://i.imgur.com/rtsm4Th.png)\n\nSome of the things which I plan to try in future:\n- different losses (like correct multiclass dice);\n- use different model architectures;\n- take some ideas from severstal competition;",
    "635262": "Did you apply RAdam from here: [https://github.com/LiyuanLucasLiu/RAdam](https://github.com/LiyuanLucasLiu/RAdam) ? And did you use the same parameters as you did in your Adam optimizer?  \n\nOh, and thank you for sharing all this information and your kernel! :)",
    "635183": "@artgor, I really liked the topic, please share your future findings too!",
    "634805": "If you don't mind sharing, is classifying before masking already in your current working pipeline? I tried doing a classifier and then have the thresholds searched above and below that binary classification threshold and it improved local validation but always decreases lb. Not sure if you might be doing something less prone to overfitting. ",
    "636698": "Interesting, thanks for sharing! I also have to agree on the points you have made. My latest boost was from using the largest encoder I could handle, yet barely provided a boost (~0.002?), not too sure what is going on there. To make things wore, it saw larger reduction in validation than and that of rise in LB, which I do not like to trust. I think it just got lucky. However, my recent 5 fold average score saw somewhat similar values with the public LB, so I hope to continue working on using it.\n\nAlso... the fact that the nature of this competition is different from a typical segmentation task (more like bounding box labels) adds another layer of confusion and unknown.  :(",
    "635948": "One question for the people with more experience in the matter, how do you usually evaluate a specific augmentation method?\n\nYou use only that method and see how it performs? Keep the previous ones and add the new? Sometimes I feel hard to evaluate this because of all the randomness. ",
    "635280": "My experience on most segmentation tasks is the segmenter with threshold is as good a classifier as any separate classifier. Though you can train a good set of classifiers early on to maybe get a headstart. It can also help if you have an inference time limit but we are OK here. I am just using a segmentation model.",
    "634777": "hi sir,thank you for this post,you are absolutely right,i will say \"taking some ideas from severstal competitions is great idea\" but i can't find any competition related to this one,here we are predicting masks on small scale and training on large size dataset,i played a lot with your kernel and i observe the loss function doesn't converge and i don't care the public lb,our all models are overfitted,even if we get medal,i am sure our models overfitted, only thing i am planing to do next is \"spending a lot of time on the dataset to find some leaks and problems\",your help is highly appreciated\n\ni am planning to implement [this kernel](https://www.kaggle.com/rishabhiitbhu/unet-starter-kernel-pytorch-lb-0-88) in this competition? do you recommend it or something else?\n\n\nthank you in advance",
    "635572": "Thanks for sharing those thoughts about your work!\nI'm not sure I fully understand, what do you mean by using a classifier to determine which masks should have masks boost score?"
  }
}