{
  "id": 130311,
  "title": "7 things that did not worked",
  "url": "/competitions/bengaliai-cv19/discussion/130311",
  "author_name": "",
  "post_date": "2020-02-13T11:19:36.797431300Z",
  "votes": 62,
  "comment_count": 52,
  "views": 0,
  "content": "<p>Among other things that worked that I will shared it after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower</p>\n\n<ol>\n<li>Replacing Adam with RAdam (almost identical  behavior)</li>\n<li>Cutout combined with Cutmix/Mixup for the same images (the image becomes too noisy and the model is having a hard time understanding)</li>\n<li>Wide resnet50 architecture (for me se_resnet101 was better)</li>\n<li>Small image size inputs (64x64)</li>\n<li>Vertical and horizontal flipping (images are not symmetrical in x or y axis)</li>\n<li>OneCycleLearning  (for me ReduceOnPlateau worked better)</li>\n<li>It may be surprising but for me cross-entropy did better than focal loss</li>\n</ol>",
  "messages": [
    {
      "id": "744993",
      "postDate": "02/13/2020 11:19:36",
      "content": "<p>Among other things that worked that I will shared it after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower</p>\n\n<ol>\n<li>Replacing Adam with RAdam (almost identical  behavior)</li>\n<li>Cutout combined with Cutmix/Mixup for the same images (the image becomes too noisy and the model is having a hard time understanding)</li>\n<li>Wide resnet50 architecture (for me se_resnet101 was better)</li>\n<li>Small image size inputs (64x64)</li>\n<li>Vertical and horizontal flipping (images are not symmetrical in x or y axis)</li>\n<li>OneCycleLearning  (for me ReduceOnPlateau worked better)</li>\n<li>It may be surprising but for me cross-entropy did better than focal loss</li>\n</ol>",
      "rawMarkdown": "Among other things that worked that I will shared it after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower\n\n1. Replacing Adam with RAdam (almost identical  behavior)\n2. Cutout combined with Cutmix/Mixup for the same images (the image becomes too noisy and the model is having a hard time understanding)\n3. Wide resnet50 architecture (for me se_resnet101 was better)\n4. Small image size inputs (64x64)\n5. Vertical and horizontal flipping (images are not symmetrical in x or y axis)\n6. OneCycleLearning  (for me ReduceOnPlateau worked better)\n7. It may be surprising but for me cross-entropy did better than focal loss",
      "votes": null
    },
    {
      "id": "745100",
      "postDate": "02/13/2020 13:34:33",
      "content": "<p>Thanks for sharing!</p>\n\n<p>1) I haven't tried RAdam, have tried AdamW which seems like giving me similar result on validation set as Adam, but training set metric was lower (so maybe could generalize better, but didnt see much change in LB)</p>\n\n<p>3) Having a hard time to train models bigger than resnet50, e.g. these se-resnext50 etc, the training time is so long, do you mind sharing your input size, batch size and epochs? a 50 epochs on resnet34 with 224x224 input took me 2hrs to train </p>\n\n<p>4) that didnt work for me either, although there was a post saying how he/she got high score with just 64x64, so far 224 seems like giving me best result (except it took time to train)</p>\n\n<p>6) I kept using OneCycleLR due to benefits mentioned in the super convergency paper, will try ReduceOnPlateau soon then, if you dont mind sharing again, do you monitor the epoch/iteration loss or the metric and patience you use?</p>\n\n<p>7) Same here, focal loss was clearly worse for the first half way training, although it starts converging but didnt get as good as a just weighted CE</p>",
      "rawMarkdown": "Thanks for sharing!\n\n1) I haven't tried RAdam, have tried AdamW which seems like giving me similar result on validation set as Adam, but training set metric was lower (so maybe could generalize better, but didnt see much change in LB)\n\n3) Having a hard time to train models bigger than resnet50, e.g. these se-resnext50 etc, the training time is so long, do you mind sharing your input size, batch size and epochs? a 50 epochs on resnet34 with 224x224 input took me 2hrs to train \n\n4) that didnt work for me either, although there was a post saying how he/she got high score with just 64x64, so far 224 seems like giving me best result (except it took time to train)\n\n6) I kept using OneCycleLR due to benefits mentioned in the super convergency paper, will try ReduceOnPlateau soon then, if you dont mind sharing again, do you monitor the epoch/iteration loss or the metric and patience you use?\n\n7) Same here, focal loss was clearly worse for the first half way training, although it starts converging but didnt get as good as a just weighted CE",
      "votes": null
    },
    {
      "id": "745105",
      "postDate": "02/13/2020 13:44:29",
      "content": "<p>Good feedback <a href=\"/samshipengs\">@samshipengs</a> \nI use input size 128x128 , batch size of 64 and 120 epochs (last 15-20 are not very useful so around 100 will be enough). I will try higher resolution, 164x164 and 224x224. A epoch now takes about 25 mins on a 2080Ti\nI only print the learning rate at each epoch to see when is decreasing, now at every 5 epochs without improvement, I multiply the learning rate (which starts at 0.001) with 0.8</p>",
      "rawMarkdown": "Good feedback @samshipengs \nI use input size 128x128 , batch size of 64 and 120 epochs (last 15-20 are not very useful so around 100 will be enough). I will try higher resolution, 164x164 and 224x224. A epoch now takes about 25 mins on a 2080Ti\nI only print the learning rate at each epoch to see when is decreasing, now at every 5 epochs without improvement, I multiply the learning rate (which starts at 0.001) with 0.8",
      "votes": null
    },
    {
      "id": "745130",
      "postDate": "02/13/2020 14:30:10",
      "content": "<p>I've had similar results in my experiments:\n- WRN never worked for me, same with smaller image size.\n- Along with Cutout+Mixup/Cutmix not working well, I tried Augmix+Mixup/Cutmix and had a similar experience with the model not being able to learn from the messy images. Not sure how people are using Augmix.\n- I have not tried Focal loss but I have tried OHEM loss a few times still unsure on its performance. Have you tried OHEM loss?</p>",
      "rawMarkdown": "I've had similar results in my experiments:\n- WRN never worked for me, same with smaller image size.\n- Along with Cutout+Mixup/Cutmix not working well, I tried Augmix+Mixup/Cutmix and had a similar experience with the model not being able to learn from the messy images. Not sure how people are using Augmix.\n- I have not tried Focal loss but I have tried OHEM loss a few times still unsure on its performance. Have you tried OHEM loss?",
      "votes": null
    },
    {
      "id": "745131",
      "postDate": "02/13/2020 14:32:22",
      "content": "<p>Hi <a href=\"/greatgamedota\">@greatgamedota</a> \nNo, I did not try OHEM loss, it is on my to do list.</p>",
      "rawMarkdown": "Hi @greatgamedota \nNo, I did not try OHEM loss, it is on my to do list.",
      "votes": null
    },
    {
      "id": "745250",
      "postDate": "02/13/2020 16:40:37",
      "content": "<p>same here.\n1. most \"cool\" optimizers can't beat old trusty adamw.\n2. can't get gridmask work better with cutmix\n7. focal loss &lt; ce</p>",
      "rawMarkdown": "same here.\n1. most \"cool\" optimizers can't beat old trusty adamw.\n2. can't get gridmask work better with cutmix\n7. focal loss &lt; ce",
      "votes": null
    },
    {
      "id": "745404",
      "postDate": "02/13/2020 19:34:05",
      "content": "<p>Thanks for creating this topic. Its always helpful for the community to know what failed = )</p>\n\n<p><code>Point 2</code>\nDo you use a fix probability for cutout and the number of cutouts ?</p>\n\n<p>After some optimization for cutout I found  up to 1-10 cutouts per image works the best, with probability of image getting cutout 30%. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fb3e37be93354fc3d75a6f9ae0330a46a%2FScreen%20Shot%202020-02-13%20at%202.32.09%20PM.png?generation=1581622382327634&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see from the image... its very soft cutout almost like scratch... For me anything more than this doesn't improve the performance . </p>\n\n<p><code>Point 1</code>\nAgree with you regarding optimizers, They kind of perform very similar </p>\n\n<p><code>Point 8</code>\nVery small difference if you train on 3 channel or 1 channel</p>\n\n<p><code>Point 9</code>\nNot a big difference if you normalize image using <code>imagnet</code> or <code>dataset</code> set</p>\n\n<p><code>Point 10</code>\nTraining longer improves the performance for <code>mixup</code> and <code>cutmix</code>\nmodel 1 trained for <code>100 epoch</code> - CV - 0.9841688871383667.\nmodel 2 (same as model 1) trained for <code>200 epoch</code> - CV - 0.9896742701530457.</p>\n\n<p>Not sure what i am doing here wrong... I saw some people could achieve CV of<code>0.99</code> + only using <code>150</code> epoch:</p>\n\n<p>```\nccchang\nCV: 0.9937\nLB: 0.9838</p>\n\n<p>Interesting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%,\nguess I need totally different way to break through 99%</p>\n\n<p>pheadrus 150 epochs\nipythonx My CV : 0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant) = 0.9937\nthanatoz Yes I combined augmentation methods and Cutmix is one of them\n<code>``\nor Gary only in</code>80 epoch`</p>\n\n<p>```\nmodel: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix\nCV : 0.994\nLB: 0.985\n80 epochs</p>\n\n<p>```</p>",
      "rawMarkdown": "Thanks for creating this topic. Its always helpful for the community to know what failed = )\n\n`Point 2`\nDo you use a fix probability for cutout and the number of cutouts ?\n\nAfter some optimization for cutout I found  up to 1-10 cutouts per image works the best, with probability of image getting cutout 30%. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fb3e37be93354fc3d75a6f9ae0330a46a%2FScreen%20Shot%202020-02-13%20at%202.32.09%20PM.png?generation=1581622382327634&amp;alt=media)\n\nAs you can see from the image... its very soft cutout almost like scratch... For me anything more than this doesn't improve the performance . \n\n`Point 1`\nAgree with you regarding optimizers, They kind of perform very similar \n\n`Point 8`\nVery small difference if you train on 3 channel or 1 channel\n\n`Point 9`\nNot a big difference if you normalize image using `imagnet` or `dataset` set\n\n`Point 10`\nTraining longer improves the performance for `mixup` and `cutmix`\nmodel 1 trained for `100 epoch` - CV - 0.9841688871383667.\nmodel 2 (same as model 1) trained for `200 epoch` - CV - 0.9896742701530457.\n\nNot sure what i am doing here wrong... I saw some people could achieve CV of` 0.99` + only using `150` epoch:\n\n```\nccchang\nCV: 0.9937\nLB: 0.9838\n\nInteresting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%,\nguess I need totally different way to break through 99%\n\npheadrus 150 epochs\nipythonx My CV : 0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant) = 0.9937\nthanatoz Yes I combined augmentation methods and Cutmix is one of them\n```\nor Gary only in `80 epoch`\n\n```\nmodel: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix\nCV : 0.994\nLB: 0.985\n80 epochs\n\n\n```",
      "votes": null
    },
    {
      "id": "745469",
      "postDate": "02/13/2020 21:03:44",
      "content": "<p>Good to know</p>",
      "rawMarkdown": "Good to know",
      "votes": null
    },
    {
      "id": "745484",
      "postDate": "02/13/2020 21:23:22",
      "content": "<p>What cutout probability are you using?</p>",
      "rawMarkdown": "What cutout probability are you using?",
      "votes": null
    },
    {
      "id": "745485",
      "postDate": "02/13/2020 21:26:30",
      "content": "<p>That's impressive CV and LB! I am very frustrated now because my CV is almost 0.99 after about 120 epochs with lb is only around 0.966. Do you know what might be the problem? I am using stratified shuffle split with 80/20 ratio. Thanks! <a href=\"/drhabib\">@drhabib</a> </p>",
      "rawMarkdown": "That's impressive CV and LB! I am very frustrated now because my CV is almost 0.99 after about 120 epochs with lb is only around 0.966. Do you know what might be the problem? I am using stratified shuffle split with 80/20 ratio. Thanks! @drhabib",
      "votes": null
    },
    {
      "id": "745487",
      "postDate": "02/13/2020 21:28:17",
      "content": "<p>Adding a point about loss: OHEM for me did not work.</p>",
      "rawMarkdown": "Adding a point about loss: OHEM for me did not work.",
      "votes": null
    },
    {
      "id": "745488",
      "postDate": "02/13/2020 21:29:34",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> I have tried OHEM. It worked worse than CE. I was wondering what are your best cv and lb?</p>",
      "rawMarkdown": "greatgamedota I have tried OHEM. It worked worse than CE. I was wondering what are your best cv and lb?",
      "votes": null
    },
    {
      "id": "745489",
      "postDate": "02/13/2020 21:30:00",
      "content": "<p>Interesting gap.  Not sure what could be the problem. Maybe you have leak? Maybe you can use random split and re run the model with the clean code..... Maybe other people will have better advice. </p>",
      "rawMarkdown": "Interesting gap.  Not sure what could be the problem. Maybe you have leak? Maybe you can use random split and re run the model with the clean code..... Maybe other people will have better advice.",
      "votes": null
    },
    {
      "id": "745494",
      "postDate": "02/13/2020 21:34:54",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Thanks for your reply! I tried pure shuffle split too but there was no difference. This situation has been the same for me with different losses, model structures, and augmentations. I used this notebook as my baseline: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a>\nMy cv/lb difference ranges from 0.0134 to 0.0224, and it seems that the more epochs I train, the larger the gap. I checked my codes for generating validation set and calculating validation score, but everything seemed all right.</p>",
      "rawMarkdown": "drhabib Thanks for your reply! I tried pure shuffle split too but there was no difference. This situation has been the same for me with different losses, model structures, and augmentations. I used this notebook as my baseline: https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\nMy cv/lb difference ranges from 0.0134 to 0.0224, and it seems that the more epochs I train, the larger the gap. I checked my codes for generating validation set and calculating validation score, but everything seemed all right.",
      "votes": null
    },
    {
      "id": "745495",
      "postDate": "02/13/2020 21:35:15",
      "content": "<p>That gap is not good. I was able to lower the gap with harder augmentation. Should be around 1%. I've never gotten my CV to .98 yet.</p>",
      "rawMarkdown": "That gap is not good. I was able to lower the gap with harder augmentation. Should be around 1%. I've never gotten my CV to .98 yet.",
      "votes": null
    },
    {
      "id": "745497",
      "postDate": "02/13/2020 21:36:29",
      "content": "<p>I guess check how validation score is computed, no leak if you are doing normalization with given dataset, and the submission kernel setup is exactly same as your training? (i had one submission where the prediction is run on different preprocessing method, and it gave me worse result than anticipated)</p>",
      "rawMarkdown": "I guess check how validation score is computed, no leak if you are doing normalization with given dataset, and the submission kernel setup is exactly same as your training? (i had one submission where the prediction is run on different preprocessing method, and it gave me worse result than anticipated)",
      "votes": null
    },
    {
      "id": "745499",
      "postDate": "02/13/2020 21:37:33",
      "content": "<p>My current score isn't with OHEM but my previous score (LB .9661) used decaying OHEM. (Though I was able to reproduce without it)</p>",
      "rawMarkdown": "My current score isn't with OHEM but my previous score (LB .9661) used decaying OHEM. (Though I was able to reproduce without it)",
      "votes": null
    },
    {
      "id": "745509",
      "postDate": "02/13/2020 21:49:19",
      "content": "<p><a href=\"/samshipengs\">@samshipengs</a> Thanks! I think there shouldn't be a problem with prediction kernel. I checked it and found no problem, and I used the same one written by the author of my training baseline and didn't change anything.</p>",
      "rawMarkdown": "samshipengs Thanks! I think there shouldn't be a problem with prediction kernel. I checked it and found no problem, and I used the same one written by the author of my training baseline and didn't change anything.",
      "votes": null
    },
    {
      "id": "745510",
      "postDate": "02/13/2020 21:49:46",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Thanks! I'll try more augmentation.</p>",
      "rawMarkdown": "greatgamedota Thanks! I'll try more augmentation.",
      "votes": null
    },
    {
      "id": "745564",
      "postDate": "02/13/2020 23:49:35",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> , thank you for sharing, I really am learning a lot from your sharing and what others have shared, but so far you really have taught me a lot especially the CNN Tails post.</p>\n\n<p>First of all, I want to point out something silly: doesn't the 2nd image you showed look like a bulldog's face? You can see the eye, the ears, the mouth..</p>\n\n<p>The serious question: Can you comment on <code>\"Very small difference if you train on 3 channel or 1 channel\"</code>. You mentioned that you can change the 1 grayscale channel into 3 RGB channels by using \"cloning\" in a previous post. I don't really know what that means. Currently, our code does this:</p>\n\n<p><code>\nself.conv0 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=1, bias=True)\n</code></p>\n\n<p>This makes me uncomfortable, because the model is learning Conv2D weights to change my grayscale image into RGB using a 3x3 window, which is altering those grayscale images. So I think cloning may fare better? (e.g. <code>input = torch.stack([input, input, input)])</code></p>\n\n<p>If you can comment any opinions, and also please clarify what you mean by \"training on 1 channel\" - which I deem is not feasible, because the <code>se-resnext50</code> is expecting RGB, unless you are referring to the technique I spoke about here, or you are modifying the first part of <code>se-resnext50</code>to accept 1 channel instead of 3 channels, in which case you cannot use pretrained weights from the model and would have to basically reformat the entire architecture to work with 1 channel instead of the expected 3.</p>",
      "rawMarkdown": "drhabib , thank you for sharing, I really am learning a lot from your sharing and what others have shared, but so far you really have taught me a lot especially the CNN Tails post.\n\nFirst of all, I want to point out something silly: doesn't the 2nd image you showed look like a bulldog's face? You can see the eye, the ears, the mouth..\n\nThe serious question: Can you comment on `\"Very small difference if you train on 3 channel or 1 channel\"`. You mentioned that you can change the 1 grayscale channel into 3 RGB channels by using \"cloning\" in a previous post. I don't really know what that means. Currently, our code does this:\n\n```\nself.conv0 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=1, bias=True)\n```\n\nThis makes me uncomfortable, because the model is learning Conv2D weights to change my grayscale image into RGB using a 3x3 window, which is altering those grayscale images. So I think cloning may fare better? (e.g. `input = torch.stack([input, input, input)])`\n\nIf you can comment any opinions, and also please clarify what you mean by \"training on 1 channel\" - which I deem is not feasible, because the `se-resnext50` is expecting RGB, unless you are referring to the technique I spoke about here, or you are modifying the first part of `se-resnext50`to accept 1 channel instead of 3 channels, in which case you cannot use pretrained weights from the model and would have to basically reformat the entire architecture to work with 1 channel instead of the expected 3.",
      "votes": null
    },
    {
      "id": "745589",
      "postDate": "02/14/2020 01:06:49",
      "content": "<p>Hey Corey! \nThanks for the kind words!! I learned from you a lot as well =) </p>\n\n<p>Ok we have <code>1 channel</code> image and we want to use pertained model  on <code>imagenet</code> (e.g <code>se_resnext50_32x4d</code>).  Since all the current pertained models are trained to accept <code>3 channel</code> image (<code>RGB</code>), we can solve our problem in 3 ways. </p>\n\n<p><code>----------------------------------------------------</code>\n<code>Option 1: Convert 1 channel image two 3 channel</code>\nThis one is pretty easy as you suggested we can clone our image in to 3 channels and problem is solved.</p>\n\n<p><code>----------------------------------------------------</code>\n<code>Option 2: Replace first Conv2D to accept 1 channel.</code>\nMost CNN's start with <code>Conv2D</code> and they accept <code>3 channel</code>input. If we want to feed our network with <code>1 channel</code> image we have to modify <strong>ONLY</strong> this layer. Below is the example how I will do This in Pytorch. </p>\n\n<p><code>arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F046131f8b12d7a29fc35fbb2e9963a29%2FScreen%20Shot%202020-02-13%20at%207.31.26%20PM.png?generation=1581640446846621&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see our first <code>Conv2d</code> accepts <code>3 channel</code> image and outputs <code>64 channel</code>. In order to make our model accept 1 channel image <code>ONLY</code> thing we have to do is create new Conv2D which accept <code>1 channel</code>and  outputs <code>64 channel</code>.  </p>\n\n<p>And this how It looks:\n```</p>\n\n<h1>loading our model</h1>\n\n<p>arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)</p>\n\n<h1>converting to list</h1>\n\n<p>arch = list(arch.children())</p>\n\n<h1>replacing first Conv2D</h1>\n\n<p>arch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)</p>\n\n<h1>converting back to sequential</h1>\n\n<p>arch = nn.Sequential(*arch)\n```</p>\n\n<p>Thats it. Now disadvantage of this method is that we are initiating new <code>Conv2d</code> layer in the beginning of our network for which  we have to retrain weights from scratch. In some cases this can lead to slow convergence or weird spikes during training. In option 3 I show how we can solve this problem.</p>\n\n<p><code>----------------------------------------------------</code>\n<code>Option 3: Replace first Conv2d to accept 1 channel, with pre trained weights.</code>\nInstead of initiating our <code>Conv2d</code> with new weights we can reuse imagenet weights. \nIf you remember our first <code>Conv2d</code>  accept 3 channels and outputs 64 channels. We can  do following: </p>\n\n<p>1) save the weights for the first <code>Covn2d</code> and average all the channels in to 1. \n2) create new <code>Conv2d</code> which accepts <code>1 channel</code>\n3) substitute newly <code>Conv2d</code> weights with averaged imagenet weights that we created early </p>\n\n<p>Here how it looks in the code.</p>\n\n<p>```</p>\n\n<h1>loading our model</h1>\n\n<p>arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)</p>\n\n<h1>converting to list</h1>\n\n<p>arch = list(arch.children())</p>\n\n<h1>saving the weights of the forst conv in w</h1>\n\n<p>w = arch[0][0].weight</p>\n\n<h1>creating new Conv2d to accept 1 channel</h1>\n\n<p>arch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)</p>\n\n<h1>substituting weights of newly created Conv2d with w from but we have to take mean</h1>\n\n<h1>to go from  3 channel to 1</h1>\n\n<p>arch[0][0].weight = nn.Parameter(torch.mean(w, dim=1, keepdim=True))\narch = nn.Sequential(*arch)\n```</p>\n\n<p>Thats it. This method resolve all the problems. We can train 1 channel images, We are still taking advantage of the pertained weights =) </p>\n\n<p>When I mentioned that I trained 1 channel image I was referring to  <code>option 3.</code> </p>\n\n<p>I hope this answered all your question. Let me know if something is not clear. </p>\n\n<p>P.S you can imagine that if you have <code>6 channel</code> image we can do the same trick.  We preserve the 3 channel weights, create new <code>Conv2d</code> witch accepts 6 channel and just substitute it with stacked version of the saved weights. This is how it will look like </p>\n\n<p><code>\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\narch = list(arch.children())\nw = arch[0][0].weight\narch[0][0] = nn.Conv2d(6, 64, kernel_size=7, stride=2, padding=2, bias=False)\narch[0][0].weight = nn.Parameter(torch.stack([w,w] dim=1,))\narch = nn.Sequential(*arch)\n</code></p>",
      "rawMarkdown": "Hey Corey! \nThanks for the kind words!! I learned from you a lot as well =) \n\nOk we have `1 channel` image and we want to use pertained model  on `imagenet` (e.g `se_resnext50_32x4d`).  Since all the current pertained models are trained to accept `3 channel ` image (`RGB`), we can solve our problem in 3 ways. \n\n`----------------------------------------------------`\n`Option 1: Convert 1 channel image two 3 channel `\nThis one is pretty easy as you suggested we can clone our image in to 3 channels and problem is solved.\n\n\n`----------------------------------------------------`\n`Option 2: Replace first Conv2D to accept 1 channel.`\nMost CNN's start with `Conv2D` and they accept `3 channel `input. If we want to feed our network with `1 channel` image we have to modify **ONLY** this layer. Below is the example how I will do This in Pytorch. \n\n`arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F046131f8b12d7a29fc35fbb2e9963a29%2FScreen%20Shot%202020-02-13%20at%207.31.26%20PM.png?generation=1581640446846621&amp;alt=media)\n\nAs you can see our first `Conv2d` accepts `3 channel` image and outputs `64 channel `. In order to make our model accept 1 channel image `ONLY` thing we have to do is create new Conv2D which accept `1 channel `and  outputs `64 channel `.  \n\nAnd this how It looks:\n```\n#loading our model\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\n#converting to list\narch = list(arch.children())\n#replacing first Conv2D\narch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)\n#converting back to sequential\narch = nn.Sequential(*arch)\n```\n\nThats it. Now disadvantage of this method is that we are initiating new `Conv2d` layer in the beginning of our network for which  we have to retrain weights from scratch. In some cases this can lead to slow convergence or weird spikes during training. In option 3 I show how we can solve this problem.\n\n\n`----------------------------------------------------`\n`Option 3: Replace first Conv2d to accept 1 channel, with pre trained weights. `\nInstead of initiating our `Conv2d` with new weights we can reuse imagenet weights. \nIf you remember our first `Conv2d`  accept 3 channels and outputs 64 channels. We can  do following: \n\n1) save the weights for the first `Covn2d` and average all the channels in to 1. \n2) create new `Conv2d` which accepts `1 channel`\n3) substitute newly `Conv2d` weights with averaged imagenet weights that we created early \n\nHere how it looks in the code.\n\n```\n#loading our model\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\n#converting to list\narch = list(arch.children())\n#saving the weights of the forst conv in w\nw = arch[0][0].weight\n#creating new Conv2d to accept 1 channel \narch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)\n#substituting weights of newly created Conv2d with w from but we have to take mean\n#to go from  3 channel to 1\narch[0][0].weight = nn.Parameter(torch.mean(w, dim=1, keepdim=True))\narch = nn.Sequential(*arch)\n```\n\nThats it. This method resolve all the problems. We can train 1 channel images, We are still taking advantage of the pertained weights =) \n\nWhen I mentioned that I trained 1 channel image I was referring to  `option 3.` \n\nI hope this answered all your question. Let me know if something is not clear. \n\nP.S you can imagine that if you have `6 channel ` image we can do the same trick.  We preserve the 3 channel weights, create new `Conv2d` witch accepts 6 channel and just substitute it with stacked version of the saved weights. This is how it will look like \n\n```\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\narch = list(arch.children())\nw = arch[0][0].weight\narch[0][0] = nn.Conv2d(6, 64, kernel_size=7, stride=2, padding=2, bias=False)\narch[0][0].weight = nn.Parameter(torch.stack([w,w] dim=1,))\narch = nn.Sequential(*arch)\n```",
      "votes": null
    },
    {
      "id": "745597",
      "postDate": "02/14/2020 01:20:47",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Hi, Did you use ReduceOnPlatau or OneCycleLearning , In my experiment, I set the epoch to 120, but when the epoch equals 70 ~ 80, the learning rate is already very small, and the CV score is not much improved. </p>",
      "rawMarkdown": "drhabib Hi, Did you use ReduceOnPlatau or OneCycleLearning , In my experiment, I set the epoch to 120, but when the epoch equals 70 ~ 80, the learning rate is already very small, and the CV score is not much improved.",
      "votes": null
    },
    {
      "id": "745640",
      "postDate": "02/14/2020 03:12:34",
      "content": "<p>Hi DrHB, thank you for a fantastic reply, and also thank you for your verbosity and also your code snippets; it really is helpful to someone like me who is a total noob in CNN.</p>\n\n<p>Honestly though, if you think about it, Option 3 isn't very attractive in my opinion. Averaging the weights I think can yield bad outcomes that have to be fixed and basically in my eyes make option 2 and option 3 equivalent. Here's how I think about it. Let's imagine we have weights <code>[w1, w2, w3, ..., w10]</code>. Well, on the Red color spectrum, <code>w1</code> might refer to characteristic <code>c1</code> and <code>w2</code> might refer to characteristic <code>c2</code>. However in the green spectrum, you might have <code>w1</code> refer to characteristic <code>c2</code> and <code>w2</code> refer to characteristic <code>c1</code>. I feel like averaging in this case does not really seem to yield something better than Option 2, however it feels more \"elegant\"...</p>\n\n<p>What is more attractive, in my opinion, is to use the concatenation of <code>red weights, green weights, blue weights</code> and use the concatenation of that information! I think that is more powerful. So it sounds like I would prefer Option 1.</p>\n\n<p>However, I don't really want the colors to \"interact\" with each other. For example the se-resnext model may have learned \"if red==blue, then DO THIS else DO THAT\". Unfortunately this type of interaction you just have to suck up and live with, and hope your model can correct some of these weights I guess.</p>\n\n<p>I am going to code up Option 1 on our team because despite the flaw I mentioned, I still think it would be more powerful than options 2 &amp; 3. Furthermore, I think I remember you posted that using 3 channels instead of 1 channel gave a tiny boost, so hopefully this intuition is correct</p>\n\n<p>Thank you again for your contributions, and good luck in this competition</p>",
      "rawMarkdown": "Hi DrHB, thank you for a fantastic reply, and also thank you for your verbosity and also your code snippets; it really is helpful to someone like me who is a total noob in CNN.\n\nHonestly though, if you think about it, Option 3 isn't very attractive in my opinion. Averaging the weights I think can yield bad outcomes that have to be fixed and basically in my eyes make option 2 and option 3 equivalent. Here's how I think about it. Let's imagine we have weights `[w1, w2, w3, ..., w10]`. Well, on the Red color spectrum, `w1` might refer to characteristic `c1` and `w2` might refer to characteristic `c2`. However in the green spectrum, you might have `w1` refer to characteristic `c2` and `w2` refer to characteristic `c1`. I feel like averaging in this case does not really seem to yield something better than Option 2, however it feels more \"elegant\"...\n\nWhat is more attractive, in my opinion, is to use the concatenation of `red weights, green weights, blue weights` and use the concatenation of that information! I think that is more powerful. So it sounds like I would prefer Option 1.\n\nHowever, I don't really want the colors to \"interact\" with each other. For example the se-resnext model may have learned \"if red==blue, then DO THIS else DO THAT\". Unfortunately this type of interaction you just have to suck up and live with, and hope your model can correct some of these weights I guess.\n\nI am going to code up Option 1 on our team because despite the flaw I mentioned, I still think it would be more powerful than options 2 &amp; 3. Furthermore, I think I remember you posted that using 3 channels instead of 1 channel gave a tiny boost, so hopefully this intuition is correct\n\nThank you again for your contributions, and good luck in this competition",
      "votes": null
    },
    {
      "id": "745686",
      "postDate": "02/14/2020 04:48:42",
      "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> I have tested both option 1 and option 2. They do not make much difference. Also, I was wondering if I could join your team. I have competed in numerous CV competitions before, and I have a 2080 Ti and a lot of time to do testing. Thanks!</p>",
      "rawMarkdown": "returnofsputnik I have tested both option 1 and option 2. They do not make much difference. Also, I was wondering if I could join your team. I have competed in numerous CV competitions before, and I have a 2080 Ti and a lot of time to do testing. Thanks!",
      "votes": null
    },
    {
      "id": "745804",
      "postDate": "02/14/2020 08:13:52",
      "content": "<p>Thank you! This is my list of \"What didn't work for me\":\n1. OHEM loss\n2. Class balanced loss\n3. Balanced data sampler\n4. AugMix - low, medium and high severity\n5. GridDropout</p>",
      "rawMarkdown": "Thank you! This is my list of \"What didn't work for me\":\n1. OHEM loss\n2. Class balanced loss\n3. Balanced data sampler\n4. AugMix - low, medium and high severity\n5. GridDropout",
      "votes": null
    },
    {
      "id": "745862",
      "postDate": "02/14/2020 09:48:07",
      "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> , glad to see you here\nI used a probability of 25% for cutout and max 8 cutouts. But this was combined with either cutmix or mixup. I usually train for around 120 epochs, it takes a pretty long time and to be honest I never tried to leave it more epochs, maybe I will have a try to see if it helps</p>",
      "rawMarkdown": "Hi @drhabib , glad to see you here\nI used a probability of 25% for cutout and max 8 cutouts. But this was combined with either cutmix or mixup. I usually train for around 120 epochs, it takes a pretty long time and to be honest I never tried to leave it more epochs, maybe I will have a try to see if it helps",
      "votes": null
    },
    {
      "id": "745865",
      "postDate": "02/14/2020 09:48:59",
      "content": "<p>Good feedback <a href=\"/dukhovnik\">@dukhovnik</a> \nDo you use mixup/cutmix both or just one of them ?</p>",
      "rawMarkdown": "Good feedback @dukhovnik \nDo you use mixup/cutmix both or just one of them ?",
      "votes": null
    },
    {
      "id": "745919",
      "postDate": "02/14/2020 11:37:18",
      "content": "<p>random select [cutmix, mixup, gridmask] doesn't work for me. For me, gridmask works best.</p>",
      "rawMarkdown": "random select [cutmix, mixup, gridmask] doesn't work for me. For me, gridmask works best.",
      "votes": null
    },
    {
      "id": "745927",
      "postDate": "02/14/2020 11:54:37",
      "content": "<p>I am in progress with tuning them right now :)</p>",
      "rawMarkdown": "I am in progress with tuning them right now :)",
      "votes": null
    },
    {
      "id": "746007",
      "postDate": "02/14/2020 13:25:47",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> (Gold Retriever) Thank you for your offer, please go to my profile and click 'Contact User' so we can handle that chat privately. I notified my teammate as well.</p>",
      "rawMarkdown": "tonychenxyz (Gold Retriever) Thank you for your offer, please go to my profile and click 'Contact User' so we can handle that chat privately. I notified my teammate as well.",
      "votes": null
    },
    {
      "id": "746615",
      "postDate": "02/15/2020 09:18:08",
      "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> hi Corey, I think using  Op1 is the same as using Op2, as in this case, Op1 is using three same gray img, and the result of conv0, is the average of the orginal weight conv the img, which is Op3.</p>",
      "rawMarkdown": "returnofsputnik hi Corey, I think using  Op1 is the same as using Op2, as in this case, Op1 is using three same gray img, and the result of conv0, is the average of the orginal weight conv the img, which is Op3.",
      "votes": null
    },
    {
      "id": "746793",
      "postDate": "02/15/2020 14:58:09",
      "content": "<p><a href=\"/liang23333\">@liang23333</a> With cutmix and mixup I found that the model has to be trained for much longer (40 epochs -&gt; &gt;100 epochs). Are you observing the same effect with gridmask?</p>",
      "rawMarkdown": "liang23333 With cutmix and mixup I found that the model has to be trained for much longer (40 epochs -&gt; &gt;100 epochs). Are you observing the same effect with gridmask?",
      "votes": null
    },
    {
      "id": "746798",
      "postDate": "02/15/2020 15:02:03",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Nice to see you here! I wonder what batch size you are using? For big batch sizes, same number of epochs correspond to much fewer steps, which might explain the relatively slow training. However small batch sizes have to correspond to smaller LR, so I am still in the process of tuning them.</p>\n\n<p>Also, if you don't mind disclosing, do you find Onecycle or ReduceLRonPlateau more handy ?</p>",
      "rawMarkdown": "drhabib Nice to see you here! I wonder what batch size you are using? For big batch sizes, same number of epochs correspond to much fewer steps, which might explain the relatively slow training. However small batch sizes have to correspond to smaller LR, so I am still in the process of tuning them.\n\nAlso, if you don't mind disclosing, do you find Onecycle or ReduceLRonPlateau more handy ?",
      "votes": null
    },
    {
      "id": "746831",
      "postDate": "02/15/2020 15:52:31",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Hey, i'm also having this problem, i can reach 0.982 cv but then i get 0.966lb. I'm using cutmix, and rotate as augments, maybe other augmentations will help? Only thing that help a little bit was adding weights in the loss function to handle imbalanced classes, then i get lower (0.974) cv but about 0.9685 lb! What are your thoughts? :)</p>",
      "rawMarkdown": "tonychenxyz Hey, i'm also having this problem, i can reach 0.982 cv but then i get 0.966lb. I'm using cutmix, and rotate as augments, maybe other augmentations will help? Only thing that help a little bit was adding weights in the loss function to handle imbalanced classes, then i get lower (0.974) cv but about 0.9685 lb! What are your thoughts? :)",
      "votes": null
    },
    {
      "id": "746872",
      "postDate": "02/15/2020 17:12:54",
      "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> How many epochs are you training? Could you share what specfically you added in loss function? Thanks! I'm now trying bigger image size, which doesn't seem to work. I'm going to add more augmentation to see if anything improves. Another thing I noticed that is that the gap increases as epoch number increases.</p>",
      "rawMarkdown": "yannmajewski How many epochs are you training? Could you share what specfically you added in loss function? Thanks! I'm now trying bigger image size, which doesn't seem to work. I'm going to add more augmentation to see if anything improves. Another thing I noticed that is that the gap increases as epoch number increases.",
      "votes": null
    },
    {
      "id": "746899",
      "postDate": "02/15/2020 17:46:09",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> im currently training for 100 epochs to test things out and see what works and what doesn't. And for the weights in the loss functions you have to calculate (1/nb of occurrences) of each class, so it will give you a tensor of the same size as the nb of classes. Then add this to the weight param in the loss function!</p>\n\n<p>Keep me updated on your tests! i'll do the same :)</p>",
      "rawMarkdown": "tonychenxyz im currently training for 100 epochs to test things out and see what works and what doesn't. And for the weights in the loss functions you have to calculate (1/nb of occurrences) of each class, so it will give you a tensor of the same size as the nb of classes. Then add this to the weight param in the loss function!\n\nKeep me updated on your tests! i'll do the same :)",
      "votes": null
    },
    {
      "id": "746985",
      "postDate": "02/15/2020 20:20:56",
      "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Cool! Thanks!</p>",
      "rawMarkdown": "yannmajewski Cool! Thanks!",
      "votes": null
    },
    {
      "id": "747384",
      "postDate": "02/16/2020 11:14:35",
      "content": "<p>Thank you for sharing your list of “things that did not work”. Here is my version of list.\n1.  Replacing Adam with RAdam did not work for me. I did not try Adamw.\n2.  Cutmix, mixup and cutout did not work. Simple augmentation (such as shift and rotation) would be fine.\n3.  It is surprising that ResNet50 performed worse than VGG19 on my validation set. In the meanwhile, SEResNeXt101 performed better than VGG19.\n4.  OneCycleLearning performed worse than ReduceLROnPlateau.\n5.  CrossEntropy performed better than focal loss.</p>",
      "rawMarkdown": "Thank you for sharing your list of “things that did not work”. Here is my version of list.\n1.\tReplacing Adam with RAdam did not work for me. I did not try Adamw.\n2.\tCutmix, mixup and cutout did not work. Simple augmentation (such as shift and rotation) would be fine.\n3.\tIt is surprising that ResNet50 performed worse than VGG19 on my validation set. In the meanwhile, SEResNeXt101 performed better than VGG19.\n4.\tOneCycleLearning performed worse than ReduceLROnPlateau.\n5.\tCrossEntropy performed better than focal loss.",
      "votes": null
    },
    {
      "id": "747865",
      "postDate": "02/17/2020 00:20:22",
      "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Quick question, are you using a pretrained model?</p>",
      "rawMarkdown": "tonychenxyz Quick question, are you using a pretrained model?",
      "votes": null
    },
    {
      "id": "748271",
      "postDate": "02/17/2020 10:38:30",
      "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a>, i also use seresnext50, and also have the same question with yours. I train 300 epoch, the CV is 0.9906, and i also focus that somebody get 0.99+ CV with little epoch, such as <a href=\"/garybios\">@garybios</a> result, that's so strange, and i haven't found why it appear. </p>",
      "rawMarkdown": "Hi @drhabib, i also use seresnext50, and also have the same question with yours. I train 300 epoch, the CV is 0.9906, and i also focus that somebody get 0.99+ CV with little epoch, such as @garybios result, that's so strange, and i haven't found why it appear.",
      "votes": null
    },
    {
      "id": "748277",
      "postDate": "02/17/2020 10:46:12",
      "content": "<p>For me:</p>\n\n<blockquote>\n  <p>Cutmix gives a more boost than random select cutmix and mixup</p>\n  \n  <p>Radam converge a little quickly than Adam</p>\n  \n  <p>Adding GridDistortion RandomGamma, OpticalDistortion, GaussianBlur augment not  effect the result</p>\n  \n  <p>More epochs may give better result</p>\n</blockquote>",
      "rawMarkdown": "For me:\n&gt; Cutmix gives a more boost than random select cutmix and mixup\n\n&gt; Radam converge a little quickly than Adam\n\n&gt; Adding GridDistortion RandomGamma, OpticalDistortion, GaussianBlur augment not  effect the result\n\n&gt; More epochs may give better result",
      "votes": null
    },
    {
      "id": "748288",
      "postDate": "02/17/2020 10:57:05",
      "content": "<p>In my experiments including all seperate augmentations,  i use the best recall in validation dataset. And my optimizer is adam, learning rate is 1e-4. Which optimizer and lr do you choose?</p>",
      "rawMarkdown": "In my experiments including all seperate augmentations,  i use the best recall in validation dataset. And my optimizer is adam, learning rate is 1e-4. Which optimizer and lr do you choose?",
      "votes": null
    },
    {
      "id": "748521",
      "postDate": "02/17/2020 16:12:00",
      "content": "<p>What lr scheduler do you use for 300 epochs ? <a href=\"/cswwp347724\">@cswwp347724</a> </p>",
      "rawMarkdown": "What lr scheduler do you use for 300 epochs ? @cswwp347724",
      "votes": null
    },
    {
      "id": "748822",
      "postDate": "02/18/2020 03:04:44",
      "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Yes. I'm using seresnext50 pretrained with imagenet.</p>",
      "rawMarkdown": "yannmajewski Yes. I'm using seresnext50 pretrained with imagenet.",
      "votes": null
    },
    {
      "id": "748851",
      "postDate": "02/18/2020 03:48:17",
      "content": "<p>Quick update on my OHEM loss experiments: It performs worse than plain CE. It converges worse and causes spiking in CV as training progresses.</p>",
      "rawMarkdown": "Quick update on my OHEM loss experiments: It performs worse than plain CE. It converges worse and causes spiking in CV as training progresses.",
      "votes": null
    },
    {
      "id": "749120",
      "postDate": "02/18/2020 10:19:51",
      "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  I used OneCycleLearning, but now it seems that ReduceOnPlateau may converge quickly</p>",
      "rawMarkdown": "greatgamedota  I used OneCycleLearning, but now it seems that ReduceOnPlateau may converge quickly",
      "votes": null
    },
    {
      "id": "749203",
      "postDate": "02/18/2020 12:40:36",
      "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> I realize that batch size has a large effect on training time. If it is convenient, what batch size are you using to train 300 epochs? It must be taking the hell of a long time :)</p>",
      "rawMarkdown": "cswwp347724 I realize that batch size has a large effect on training time. If it is convenient, what batch size are you using to train 300 epochs? It must be taking the hell of a long time :)",
      "votes": null
    },
    {
      "id": "749871",
      "postDate": "02/19/2020 00:39:33",
      "content": "<p><a href=\"/vladvdv\">@vladvdv</a> I wonder how is \"epochs without improvement\" defined? I have observed cases where validation metric and loss are both increasing :( It's very baffling</p>",
      "rawMarkdown": "vladvdv I wonder how is \"epochs without improvement\" defined? I have observed cases where validation metric and loss are both increasing :( It's very baffling",
      "votes": null
    },
    {
      "id": "750325",
      "postDate": "02/19/2020 09:49:48",
      "content": "<p>those are impressive cv scores, I find the loss weighting of .5, .25, .25 between root, vowel and consonant not working for me. the vowel and consonant will become very good (.98+ easily in 30 epochs or so, but the root just gets stuck around .96)  struggling...</p>",
      "rawMarkdown": "those are impressive cv scores, I find the loss weighting of .5, .25, .25 between root, vowel and consonant not working for me. the vowel and consonant will become very good (.98+ easily in 30 epochs or so, but the root just gets stuck around .96)  struggling...",
      "votes": null
    },
    {
      "id": "752398",
      "postDate": "02/21/2020 01:50:03",
      "content": "<p>Did you use different weights for CE loss?</p>",
      "rawMarkdown": "Did you use different weights for CE loss?",
      "votes": null
    },
    {
      "id": "752583",
      "postDate": "02/21/2020 07:22:14",
      "content": "<p>my batch size is 860</p>",
      "rawMarkdown": "my batch size is 860",
      "votes": null
    },
    {
      "id": "753064",
      "postDate": "02/21/2020 17:13:18",
      "content": "<ol>\n<li>We've tried mixup, cutmix, and 50% cutmix + 50% mixup, and 50% cutmix + 50% mixup works the best.</li>\n<li>cropping image doesn't work for us. In fact, we found that cropping image dropped the accuracy by about 0.005. I don't quite understand the reason behind it though. </li>\n</ol>",
      "rawMarkdown": "1. We've tried mixup, cutmix, and 50% cutmix + 50% mixup, and 50% cutmix + 50% mixup works the best.\n2. cropping image doesn't work for us. In fact, we found that cropping image dropped the accuracy by about 0.005. I don't quite understand the reason behind it though.",
      "votes": null
    },
    {
      "id": "755354",
      "postDate": "02/24/2020 17:37:51",
      "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Thanks for you tips, I forget this point, maybe it's crucial, and i'm doing this experiment</p>",
      "rawMarkdown": "roguekk007 Thanks for you tips, I forget this point, maybe it's crucial, and i'm doing this experiment",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 745100,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "02/13/2020 13:34:33",
      "content": "<p>Thanks for sharing!</p>\n\n<p>1) I haven't tried RAdam, have tried AdamW which seems like giving me similar result on validation set as Adam, but training set metric was lower (so maybe could generalize better, but didnt see much change in LB)</p>\n\n<p>3) Having a hard time to train models bigger than resnet50, e.g. these se-resnext50 etc, the training time is so long, do you mind sharing your input size, batch size and epochs? a 50 epochs on resnet34 with 224x224 input took me 2hrs to train </p>\n\n<p>4) that didnt work for me either, although there was a post saying how he/she got high score with just 64x64, so far 224 seems like giving me best result (except it took time to train)</p>\n\n<p>6) I kept using OneCycleLR due to benefits mentioned in the super convergency paper, will try ReduceOnPlateau soon then, if you dont mind sharing again, do you monitor the epoch/iteration loss or the metric and patience you use?</p>\n\n<p>7) Same here, focal loss was clearly worse for the first half way training, although it starts converging but didnt get as good as a just weighted CE</p>",
      "votes": null,
      "replies": [
        {
          "id": 745105,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/13/2020 13:44:29",
          "content": "<p>Good feedback <a href=\"/samshipengs\">@samshipengs</a> \nI use input size 128x128 , batch size of 64 and 120 epochs (last 15-20 are not very useful so around 100 will be enough). I will try higher resolution, 164x164 and 224x224. A epoch now takes about 25 mins on a 2080Ti\nI only print the learning rate at each epoch to see when is decreasing, now at every 5 epochs without improvement, I multiply the learning rate (which starts at 0.001) with 0.8</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749871,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/19/2020 00:39:33",
          "content": "<p><a href=\"/vladvdv\">@vladvdv</a> I wonder how is \"epochs without improvement\" defined? I have observed cases where validation metric and loss are both increasing :( It's very baffling</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745130,
      "author_name": "greatgamedota",
      "author_url": "",
      "post_date": "02/13/2020 14:30:10",
      "content": "<p>I've had similar results in my experiments:\n- WRN never worked for me, same with smaller image size.\n- Along with Cutout+Mixup/Cutmix not working well, I tried Augmix+Mixup/Cutmix and had a similar experience with the model not being able to learn from the messy images. Not sure how people are using Augmix.\n- I have not tried Focal loss but I have tried OHEM loss a few times still unsure on its performance. Have you tried OHEM loss?</p>",
      "votes": null,
      "replies": [
        {
          "id": 745131,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/13/2020 14:32:22",
          "content": "<p>Hi <a href=\"/greatgamedota\">@greatgamedota</a> \nNo, I did not try OHEM loss, it is on my to do list.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745488,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/13/2020 21:29:34",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> I have tried OHEM. It worked worse than CE. I was wondering what are your best cv and lb?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745499,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/13/2020 21:37:33",
          "content": "<p>My current score isn't with OHEM but my previous score (LB .9661) used decaying OHEM. (Though I was able to reproduce without it)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748851,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/18/2020 03:48:17",
          "content": "<p>Quick update on my OHEM loss experiments: It performs worse than plain CE. It converges worse and causes spiking in CV as training progresses.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745250,
      "author_name": "moewie94",
      "author_url": "",
      "post_date": "02/13/2020 16:40:37",
      "content": "<p>same here.\n1. most \"cool\" optimizers can't beat old trusty adamw.\n2. can't get gridmask work better with cutmix\n7. focal loss &lt; ce</p>",
      "votes": null,
      "replies": [
        {
          "id": 745469,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/13/2020 21:03:44",
          "content": "<p>Good to know</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745404,
      "author_name": "drhabib",
      "author_url": "",
      "post_date": "02/13/2020 19:34:05",
      "content": "<p>Thanks for creating this topic. Its always helpful for the community to know what failed = )</p>\n\n<p><code>Point 2</code>\nDo you use a fix probability for cutout and the number of cutouts ?</p>\n\n<p>After some optimization for cutout I found  up to 1-10 cutouts per image works the best, with probability of image getting cutout 30%. </p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fb3e37be93354fc3d75a6f9ae0330a46a%2FScreen%20Shot%202020-02-13%20at%202.32.09%20PM.png?generation=1581622382327634&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see from the image... its very soft cutout almost like scratch... For me anything more than this doesn't improve the performance . </p>\n\n<p><code>Point 1</code>\nAgree with you regarding optimizers, They kind of perform very similar </p>\n\n<p><code>Point 8</code>\nVery small difference if you train on 3 channel or 1 channel</p>\n\n<p><code>Point 9</code>\nNot a big difference if you normalize image using <code>imagnet</code> or <code>dataset</code> set</p>\n\n<p><code>Point 10</code>\nTraining longer improves the performance for <code>mixup</code> and <code>cutmix</code>\nmodel 1 trained for <code>100 epoch</code> - CV - 0.9841688871383667.\nmodel 2 (same as model 1) trained for <code>200 epoch</code> - CV - 0.9896742701530457.</p>\n\n<p>Not sure what i am doing here wrong... I saw some people could achieve CV of<code>0.99</code> + only using <code>150</code> epoch:</p>\n\n<p>```\nccchang\nCV: 0.9937\nLB: 0.9838</p>\n\n<p>Interesting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%,\nguess I need totally different way to break through 99%</p>\n\n<p>pheadrus 150 epochs\nipythonx My CV : 0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant) = 0.9937\nthanatoz Yes I combined augmentation methods and Cutmix is one of them\n<code>``\nor Gary only in</code>80 epoch`</p>\n\n<p>```\nmodel: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix\nCV : 0.994\nLB: 0.985\n80 epochs</p>\n\n<p>```</p>",
      "votes": null,
      "replies": [
        {
          "id": 745485,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/13/2020 21:26:30",
          "content": "<p>That's impressive CV and LB! I am very frustrated now because my CV is almost 0.99 after about 120 epochs with lb is only around 0.966. Do you know what might be the problem? I am using stratified shuffle split with 80/20 ratio. Thanks! <a href=\"/drhabib\">@drhabib</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745489,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/13/2020 21:30:00",
          "content": "<p>Interesting gap.  Not sure what could be the problem. Maybe you have leak? Maybe you can use random split and re run the model with the clean code..... Maybe other people will have better advice. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745494,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/13/2020 21:34:54",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Thanks for your reply! I tried pure shuffle split too but there was no difference. This situation has been the same for me with different losses, model structures, and augmentations. I used this notebook as my baseline: <a href=\"https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\">https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch</a>\nMy cv/lb difference ranges from 0.0134 to 0.0224, and it seems that the more epochs I train, the larger the gap. I checked my codes for generating validation set and calculating validation score, but everything seemed all right.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745495,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/13/2020 21:35:15",
          "content": "<p>That gap is not good. I was able to lower the gap with harder augmentation. Should be around 1%. I've never gotten my CV to .98 yet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745497,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "02/13/2020 21:36:29",
          "content": "<p>I guess check how validation score is computed, no leak if you are doing normalization with given dataset, and the submission kernel setup is exactly same as your training? (i had one submission where the prediction is run on different preprocessing method, and it gave me worse result than anticipated)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745509,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/13/2020 21:49:19",
          "content": "<p><a href=\"/samshipengs\">@samshipengs</a> Thanks! I think there shouldn't be a problem with prediction kernel. I checked it and found no problem, and I used the same one written by the author of my training baseline and didn't change anything.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745510,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/13/2020 21:49:46",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a> Thanks! I'll try more augmentation.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745564,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "02/13/2020 23:49:35",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> , thank you for sharing, I really am learning a lot from your sharing and what others have shared, but so far you really have taught me a lot especially the CNN Tails post.</p>\n\n<p>First of all, I want to point out something silly: doesn't the 2nd image you showed look like a bulldog's face? You can see the eye, the ears, the mouth..</p>\n\n<p>The serious question: Can you comment on <code>\"Very small difference if you train on 3 channel or 1 channel\"</code>. You mentioned that you can change the 1 grayscale channel into 3 RGB channels by using \"cloning\" in a previous post. I don't really know what that means. Currently, our code does this:</p>\n\n<p><code>\nself.conv0 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=1, bias=True)\n</code></p>\n\n<p>This makes me uncomfortable, because the model is learning Conv2D weights to change my grayscale image into RGB using a 3x3 window, which is altering those grayscale images. So I think cloning may fare better? (e.g. <code>input = torch.stack([input, input, input)])</code></p>\n\n<p>If you can comment any opinions, and also please clarify what you mean by \"training on 1 channel\" - which I deem is not feasible, because the <code>se-resnext50</code> is expecting RGB, unless you are referring to the technique I spoke about here, or you are modifying the first part of <code>se-resnext50</code>to accept 1 channel instead of 3 channels, in which case you cannot use pretrained weights from the model and would have to basically reformat the entire architecture to work with 1 channel instead of the expected 3.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745589,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "02/14/2020 01:06:49",
          "content": "<p>Hey Corey! \nThanks for the kind words!! I learned from you a lot as well =) </p>\n\n<p>Ok we have <code>1 channel</code> image and we want to use pertained model  on <code>imagenet</code> (e.g <code>se_resnext50_32x4d</code>).  Since all the current pertained models are trained to accept <code>3 channel</code> image (<code>RGB</code>), we can solve our problem in 3 ways. </p>\n\n<p><code>----------------------------------------------------</code>\n<code>Option 1: Convert 1 channel image two 3 channel</code>\nThis one is pretty easy as you suggested we can clone our image in to 3 channels and problem is solved.</p>\n\n<p><code>----------------------------------------------------</code>\n<code>Option 2: Replace first Conv2D to accept 1 channel.</code>\nMost CNN's start with <code>Conv2D</code> and they accept <code>3 channel</code>input. If we want to feed our network with <code>1 channel</code> image we have to modify <strong>ONLY</strong> this layer. Below is the example how I will do This in Pytorch. </p>\n\n<p><code>arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)</code>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F046131f8b12d7a29fc35fbb2e9963a29%2FScreen%20Shot%202020-02-13%20at%207.31.26%20PM.png?generation=1581640446846621&amp;alt=media\" alt=\"\"></p>\n\n<p>As you can see our first <code>Conv2d</code> accepts <code>3 channel</code> image and outputs <code>64 channel</code>. In order to make our model accept 1 channel image <code>ONLY</code> thing we have to do is create new Conv2D which accept <code>1 channel</code>and  outputs <code>64 channel</code>.  </p>\n\n<p>And this how It looks:\n```</p>\n\n<h1>loading our model</h1>\n\n<p>arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)</p>\n\n<h1>converting to list</h1>\n\n<p>arch = list(arch.children())</p>\n\n<h1>replacing first Conv2D</h1>\n\n<p>arch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)</p>\n\n<h1>converting back to sequential</h1>\n\n<p>arch = nn.Sequential(*arch)\n```</p>\n\n<p>Thats it. Now disadvantage of this method is that we are initiating new <code>Conv2d</code> layer in the beginning of our network for which  we have to retrain weights from scratch. In some cases this can lead to slow convergence or weird spikes during training. In option 3 I show how we can solve this problem.</p>\n\n<p><code>----------------------------------------------------</code>\n<code>Option 3: Replace first Conv2d to accept 1 channel, with pre trained weights.</code>\nInstead of initiating our <code>Conv2d</code> with new weights we can reuse imagenet weights. \nIf you remember our first <code>Conv2d</code>  accept 3 channels and outputs 64 channels. We can  do following: </p>\n\n<p>1) save the weights for the first <code>Covn2d</code> and average all the channels in to 1. \n2) create new <code>Conv2d</code> which accepts <code>1 channel</code>\n3) substitute newly <code>Conv2d</code> weights with averaged imagenet weights that we created early </p>\n\n<p>Here how it looks in the code.</p>\n\n<p>```</p>\n\n<h1>loading our model</h1>\n\n<p>arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)</p>\n\n<h1>converting to list</h1>\n\n<p>arch = list(arch.children())</p>\n\n<h1>saving the weights of the forst conv in w</h1>\n\n<p>w = arch[0][0].weight</p>\n\n<h1>creating new Conv2d to accept 1 channel</h1>\n\n<p>arch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)</p>\n\n<h1>substituting weights of newly created Conv2d with w from but we have to take mean</h1>\n\n<h1>to go from  3 channel to 1</h1>\n\n<p>arch[0][0].weight = nn.Parameter(torch.mean(w, dim=1, keepdim=True))\narch = nn.Sequential(*arch)\n```</p>\n\n<p>Thats it. This method resolve all the problems. We can train 1 channel images, We are still taking advantage of the pertained weights =) </p>\n\n<p>When I mentioned that I trained 1 channel image I was referring to  <code>option 3.</code> </p>\n\n<p>I hope this answered all your question. Let me know if something is not clear. </p>\n\n<p>P.S you can imagine that if you have <code>6 channel</code> image we can do the same trick.  We preserve the 3 channel weights, create new <code>Conv2d</code> witch accepts 6 channel and just substitute it with stacked version of the saved weights. This is how it will look like </p>\n\n<p><code>\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\narch = list(arch.children())\nw = arch[0][0].weight\narch[0][0] = nn.Conv2d(6, 64, kernel_size=7, stride=2, padding=2, bias=False)\narch[0][0].weight = nn.Parameter(torch.stack([w,w] dim=1,))\narch = nn.Sequential(*arch)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745597,
          "author_name": "hesene",
          "author_url": "",
          "post_date": "02/14/2020 01:20:47",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Hi, Did you use ReduceOnPlatau or OneCycleLearning , In my experiment, I set the epoch to 120, but when the epoch equals 70 ~ 80, the learning rate is already very small, and the CV score is not much improved. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745640,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "02/14/2020 03:12:34",
          "content": "<p>Hi DrHB, thank you for a fantastic reply, and also thank you for your verbosity and also your code snippets; it really is helpful to someone like me who is a total noob in CNN.</p>\n\n<p>Honestly though, if you think about it, Option 3 isn't very attractive in my opinion. Averaging the weights I think can yield bad outcomes that have to be fixed and basically in my eyes make option 2 and option 3 equivalent. Here's how I think about it. Let's imagine we have weights <code>[w1, w2, w3, ..., w10]</code>. Well, on the Red color spectrum, <code>w1</code> might refer to characteristic <code>c1</code> and <code>w2</code> might refer to characteristic <code>c2</code>. However in the green spectrum, you might have <code>w1</code> refer to characteristic <code>c2</code> and <code>w2</code> refer to characteristic <code>c1</code>. I feel like averaging in this case does not really seem to yield something better than Option 2, however it feels more \"elegant\"...</p>\n\n<p>What is more attractive, in my opinion, is to use the concatenation of <code>red weights, green weights, blue weights</code> and use the concatenation of that information! I think that is more powerful. So it sounds like I would prefer Option 1.</p>\n\n<p>However, I don't really want the colors to \"interact\" with each other. For example the se-resnext model may have learned \"if red==blue, then DO THIS else DO THAT\". Unfortunately this type of interaction you just have to suck up and live with, and hope your model can correct some of these weights I guess.</p>\n\n<p>I am going to code up Option 1 on our team because despite the flaw I mentioned, I still think it would be more powerful than options 2 &amp; 3. Furthermore, I think I remember you posted that using 3 channels instead of 1 channel gave a tiny boost, so hopefully this intuition is correct</p>\n\n<p>Thank you again for your contributions, and good luck in this competition</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745686,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/14/2020 04:48:42",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> I have tested both option 1 and option 2. They do not make much difference. Also, I was wondering if I could join your team. I have competed in numerous CV competitions before, and I have a 2080 Ti and a lot of time to do testing. Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746007,
          "author_name": "returnofsputnik",
          "author_url": "",
          "post_date": "02/14/2020 13:25:47",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> (Gold Retriever) Thank you for your offer, please go to my profile and click 'Contact User' so we can handle that chat privately. I notified my teammate as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746615,
          "author_name": "jiangjiangzuijiang",
          "author_url": "",
          "post_date": "02/15/2020 09:18:08",
          "content": "<p><a href=\"/returnofsputnik\">@returnofsputnik</a> hi Corey, I think using  Op1 is the same as using Op2, as in this case, Op1 is using three same gray img, and the result of conv0, is the average of the orginal weight conv the img, which is Op3.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746798,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/15/2020 15:02:03",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Nice to see you here! I wonder what batch size you are using? For big batch sizes, same number of epochs correspond to much fewer steps, which might explain the relatively slow training. However small batch sizes have to correspond to smaller LR, so I am still in the process of tuning them.</p>\n\n<p>Also, if you don't mind disclosing, do you find Onecycle or ReduceLRonPlateau more handy ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746831,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/15/2020 15:52:31",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Hey, i'm also having this problem, i can reach 0.982 cv but then i get 0.966lb. I'm using cutmix, and rotate as augments, maybe other augmentations will help? Only thing that help a little bit was adding weights in the loss function to handle imbalanced classes, then i get lower (0.974) cv but about 0.9685 lb! What are your thoughts? :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746872,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/15/2020 17:12:54",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> How many epochs are you training? Could you share what specfically you added in loss function? Thanks! I'm now trying bigger image size, which doesn't seem to work. I'm going to add more augmentation to see if anything improves. Another thing I noticed that is that the gap increases as epoch number increases.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746899,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/15/2020 17:46:09",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> im currently training for 100 epochs to test things out and see what works and what doesn't. And for the weights in the loss functions you have to calculate (1/nb of occurrences) of each class, so it will give you a tensor of the same size as the nb of classes. Then add this to the weight param in the loss function!</p>\n\n<p>Keep me updated on your tests! i'll do the same :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 746985,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/15/2020 20:20:56",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Cool! Thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 747865,
          "author_name": "yannmajewski",
          "author_url": "",
          "post_date": "02/17/2020 00:20:22",
          "content": "<p><a href=\"/tonychenxyz\">@tonychenxyz</a> Quick question, are you using a pretrained model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748271,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/17/2020 10:38:30",
          "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a>, i also use seresnext50, and also have the same question with yours. I train 300 epoch, the CV is 0.9906, and i also focus that somebody get 0.99+ CV with little epoch, such as <a href=\"/garybios\">@garybios</a> result, that's so strange, and i haven't found why it appear. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748521,
          "author_name": "greatgamedota",
          "author_url": "",
          "post_date": "02/17/2020 16:12:00",
          "content": "<p>What lr scheduler do you use for 300 epochs ? <a href=\"/cswwp347724\">@cswwp347724</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748822,
          "author_name": "tonychenxyz",
          "author_url": "",
          "post_date": "02/18/2020 03:04:44",
          "content": "<p><a href=\"/yannmajewski\">@yannmajewski</a> Yes. I'm using seresnext50 pretrained with imagenet.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749120,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/18/2020 10:19:51",
          "content": "<p><a href=\"/greatgamedota\">@greatgamedota</a>  I used OneCycleLearning, but now it seems that ReduceOnPlateau may converge quickly</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 749203,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/18/2020 12:40:36",
          "content": "<p><a href=\"/cswwp347724\">@cswwp347724</a> I realize that batch size has a large effect on training time. If it is convenient, what batch size are you using to train 300 epochs? It must be taking the hell of a long time :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 750325,
          "author_name": "yl1202",
          "author_url": "",
          "post_date": "02/19/2020 09:49:48",
          "content": "<p>those are impressive cv scores, I find the loss weighting of .5, .25, .25 between root, vowel and consonant not working for me. the vowel and consonant will become very good (.98+ easily in 30 epochs or so, but the root just gets stuck around .96)  struggling...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 752583,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/21/2020 07:22:14",
          "content": "<p>my batch size is 860</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 755354,
          "author_name": "cswwp347724",
          "author_url": "",
          "post_date": "02/24/2020 17:37:51",
          "content": "<p><a href=\"/roguekk007\">@roguekk007</a> Thanks for you tips, I forget this point, maybe it's crucial, and i'm doing this experiment</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745484,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/13/2020 21:23:22",
      "content": "<p>What cutout probability are you using?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745487,
      "author_name": "tonychenxyz",
      "author_url": "",
      "post_date": "02/13/2020 21:28:17",
      "content": "<p>Adding a point about loss: OHEM for me did not work.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745804,
      "author_name": "dukhovnik",
      "author_url": "",
      "post_date": "02/14/2020 08:13:52",
      "content": "<p>Thank you! This is my list of \"What didn't work for me\":\n1. OHEM loss\n2. Class balanced loss\n3. Balanced data sampler\n4. AugMix - low, medium and high severity\n5. GridDropout</p>",
      "votes": null,
      "replies": [
        {
          "id": 745865,
          "author_name": "vladvdv",
          "author_url": "",
          "post_date": "02/14/2020 09:48:59",
          "content": "<p>Good feedback <a href=\"/dukhovnik\">@dukhovnik</a> \nDo you use mixup/cutmix both or just one of them ?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 745927,
          "author_name": "dukhovnik",
          "author_url": "",
          "post_date": "02/14/2020 11:54:37",
          "content": "<p>I am in progress with tuning them right now :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 745862,
      "author_name": "vladvdv",
      "author_url": "",
      "post_date": "02/14/2020 09:48:07",
      "content": "<p>Hi <a href=\"/drhabib\">@drhabib</a> , glad to see you here\nI used a probability of 25% for cutout and max 8 cutouts. But this was combined with either cutmix or mixup. I usually train for around 120 epochs, it takes a pretty long time and to be honest I never tried to leave it more epochs, maybe I will have a try to see if it helps</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 745919,
      "author_name": "liang23333",
      "author_url": "",
      "post_date": "02/14/2020 11:37:18",
      "content": "<p>random select [cutmix, mixup, gridmask] doesn't work for me. For me, gridmask works best.</p>",
      "votes": null,
      "replies": [
        {
          "id": 746793,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "02/15/2020 14:58:09",
          "content": "<p><a href=\"/liang23333\">@liang23333</a> With cutmix and mixup I found that the model has to be trained for much longer (40 epochs -&gt; &gt;100 epochs). Are you observing the same effect with gridmask?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 748288,
          "author_name": "liang23333",
          "author_url": "",
          "post_date": "02/17/2020 10:57:05",
          "content": "<p>In my experiments including all seperate augmentations,  i use the best recall in validation dataset. And my optimizer is adam, learning rate is 1e-4. Which optimizer and lr do you choose?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 747384,
      "author_name": "andrewsher",
      "author_url": "",
      "post_date": "02/16/2020 11:14:35",
      "content": "<p>Thank you for sharing your list of “things that did not work”. Here is my version of list.\n1.  Replacing Adam with RAdam did not work for me. I did not try Adamw.\n2.  Cutmix, mixup and cutout did not work. Simple augmentation (such as shift and rotation) would be fine.\n3.  It is surprising that ResNet50 performed worse than VGG19 on my validation set. In the meanwhile, SEResNeXt101 performed better than VGG19.\n4.  OneCycleLearning performed worse than ReduceLROnPlateau.\n5.  CrossEntropy performed better than focal loss.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 748277,
      "author_name": "cswwp347724",
      "author_url": "",
      "post_date": "02/17/2020 10:46:12",
      "content": "<p>For me:</p>\n\n<blockquote>\n  <p>Cutmix gives a more boost than random select cutmix and mixup</p>\n  \n  <p>Radam converge a little quickly than Adam</p>\n  \n  <p>Adding GridDistortion RandomGamma, OpticalDistortion, GaussianBlur augment not  effect the result</p>\n  \n  <p>More epochs may give better result</p>\n</blockquote>",
      "votes": null,
      "replies": []
    },
    {
      "id": 752398,
      "author_name": "mnikita",
      "author_url": "",
      "post_date": "02/21/2020 01:50:03",
      "content": "<p>Did you use different weights for CE loss?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 753064,
      "author_name": "axiostpc",
      "author_url": "",
      "post_date": "02/21/2020 17:13:18",
      "content": "<ol>\n<li>We've tried mixup, cutmix, and 50% cutmix + 50% mixup, and 50% cutmix + 50% mixup works the best.</li>\n<li>cropping image doesn't work for us. In fact, we found that cropping image dropped the accuracy by about 0.005. I don't quite understand the reason behind it though. </li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "744993": "Among other things that worked that I will shared it after the competition ends, here are 7 things that for me did not work, they either give the same cv/public leaderboard result as without them or lower\n\n1. Replacing Adam with RAdam (almost identical  behavior)\n2. Cutout combined with Cutmix/Mixup for the same images (the image becomes too noisy and the model is having a hard time understanding)\n3. Wide resnet50 architecture (for me se_resnet101 was better)\n4. Small image size inputs (64x64)\n5. Vertical and horizontal flipping (images are not symmetrical in x or y axis)\n6. OneCycleLearning  (for me ReduceOnPlateau worked better)\n7. It may be surprising but for me cross-entropy did better than focal loss",
    "745100": "Thanks for sharing!\n\n1) I haven't tried RAdam, have tried AdamW which seems like giving me similar result on validation set as Adam, but training set metric was lower (so maybe could generalize better, but didnt see much change in LB)\n\n3) Having a hard time to train models bigger than resnet50, e.g. these se-resnext50 etc, the training time is so long, do you mind sharing your input size, batch size and epochs? a 50 epochs on resnet34 with 224x224 input took me 2hrs to train \n\n4) that didnt work for me either, although there was a post saying how he/she got high score with just 64x64, so far 224 seems like giving me best result (except it took time to train)\n\n6) I kept using OneCycleLR due to benefits mentioned in the super convergency paper, will try ReduceOnPlateau soon then, if you dont mind sharing again, do you monitor the epoch/iteration loss or the metric and patience you use?\n\n7) Same here, focal loss was clearly worse for the first half way training, although it starts converging but didnt get as good as a just weighted CE",
    "745105": "Good feedback @samshipengs \nI use input size 128x128 , batch size of 64 and 120 epochs (last 15-20 are not very useful so around 100 will be enough). I will try higher resolution, 164x164 and 224x224. A epoch now takes about 25 mins on a 2080Ti\nI only print the learning rate at each epoch to see when is decreasing, now at every 5 epochs without improvement, I multiply the learning rate (which starts at 0.001) with 0.8",
    "745130": "I've had similar results in my experiments:\n- WRN never worked for me, same with smaller image size.\n- Along with Cutout+Mixup/Cutmix not working well, I tried Augmix+Mixup/Cutmix and had a similar experience with the model not being able to learn from the messy images. Not sure how people are using Augmix.\n- I have not tried Focal loss but I have tried OHEM loss a few times still unsure on its performance. Have you tried OHEM loss?",
    "745131": "Hi @greatgamedota \nNo, I did not try OHEM loss, it is on my to do list.",
    "745250": "same here.\n1. most \"cool\" optimizers can't beat old trusty adamw.\n2. can't get gridmask work better with cutmix\n7. focal loss &lt; ce",
    "745404": "Thanks for creating this topic. Its always helpful for the community to know what failed = )\n\n`Point 2`\nDo you use a fix probability for cutout and the number of cutouts ?\n\nAfter some optimization for cutout I found  up to 1-10 cutouts per image works the best, with probability of image getting cutout 30%. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2Fb3e37be93354fc3d75a6f9ae0330a46a%2FScreen%20Shot%202020-02-13%20at%202.32.09%20PM.png?generation=1581622382327634&amp;alt=media)\n\nAs you can see from the image... its very soft cutout almost like scratch... For me anything more than this doesn't improve the performance . \n\n`Point 1`\nAgree with you regarding optimizers, They kind of perform very similar \n\n`Point 8`\nVery small difference if you train on 3 channel or 1 channel\n\n`Point 9`\nNot a big difference if you normalize image using `imagnet` or `dataset` set\n\n`Point 10`\nTraining longer improves the performance for `mixup` and `cutmix`\nmodel 1 trained for `100 epoch` - CV - 0.9841688871383667.\nmodel 2 (same as model 1) trained for `200 epoch` - CV - 0.9896742701530457.\n\nNot sure what i am doing here wrong... I saw some people could achieve CV of` 0.99` + only using `150` epoch:\n\n```\nccchang\nCV: 0.9937\nLB: 0.9838\n\nInteresting that no matter how I changed model structure, augmentation or image size,\nLB scores are always equal to my CV scores minus about 1~1.3%,\nguess I need totally different way to break through 99%\n\npheadrus 150 epochs\nipythonx My CV : 0.5 * 0.99107(root) + 0.25 * 0.99648(vowel) + 0.25 * 0.99641(consonant) = 0.9937\nthanatoz Yes I combined augmentation methods and Cutmix is one of them\n```\nor Gary only in `80 epoch`\n\n```\nmodel: se-resnext50\nimg_size: 3x137x236\naugmentation: rotate, cutmix\nCV : 0.994\nLB: 0.985\n80 epochs\n\n\n```",
    "745469": "Good to know",
    "745484": "What cutout probability are you using?",
    "745485": "That's impressive CV and LB! I am very frustrated now because my CV is almost 0.99 after about 120 epochs with lb is only around 0.966. Do you know what might be the problem? I am using stratified shuffle split with 80/20 ratio. Thanks! @drhabib",
    "745487": "Adding a point about loss: OHEM for me did not work.",
    "745488": "greatgamedota I have tried OHEM. It worked worse than CE. I was wondering what are your best cv and lb?",
    "745489": "Interesting gap.  Not sure what could be the problem. Maybe you have leak? Maybe you can use random split and re run the model with the clean code..... Maybe other people will have better advice.",
    "745494": "drhabib Thanks for your reply! I tried pure shuffle split too but there was no difference. This situation has been the same for me with different losses, model structures, and augmentations. I used this notebook as my baseline: https://www.kaggle.com/corochann/bengali-seresnext-training-with-pytorch\nMy cv/lb difference ranges from 0.0134 to 0.0224, and it seems that the more epochs I train, the larger the gap. I checked my codes for generating validation set and calculating validation score, but everything seemed all right.",
    "745495": "That gap is not good. I was able to lower the gap with harder augmentation. Should be around 1%. I've never gotten my CV to .98 yet.",
    "745497": "I guess check how validation score is computed, no leak if you are doing normalization with given dataset, and the submission kernel setup is exactly same as your training? (i had one submission where the prediction is run on different preprocessing method, and it gave me worse result than anticipated)",
    "745499": "My current score isn't with OHEM but my previous score (LB .9661) used decaying OHEM. (Though I was able to reproduce without it)",
    "745509": "samshipengs Thanks! I think there shouldn't be a problem with prediction kernel. I checked it and found no problem, and I used the same one written by the author of my training baseline and didn't change anything.",
    "745510": "greatgamedota Thanks! I'll try more augmentation.",
    "745564": "drhabib , thank you for sharing, I really am learning a lot from your sharing and what others have shared, but so far you really have taught me a lot especially the CNN Tails post.\n\nFirst of all, I want to point out something silly: doesn't the 2nd image you showed look like a bulldog's face? You can see the eye, the ears, the mouth..\n\nThe serious question: Can you comment on `\"Very small difference if you train on 3 channel or 1 channel\"`. You mentioned that you can change the 1 grayscale channel into 3 RGB channels by using \"cloning\" in a previous post. I don't really know what that means. Currently, our code does this:\n\n```\nself.conv0 = nn.Conv2d(1, 3, kernel_size=3, stride=1, padding=1, bias=True)\n```\n\nThis makes me uncomfortable, because the model is learning Conv2D weights to change my grayscale image into RGB using a 3x3 window, which is altering those grayscale images. So I think cloning may fare better? (e.g. `input = torch.stack([input, input, input)])`\n\nIf you can comment any opinions, and also please clarify what you mean by \"training on 1 channel\" - which I deem is not feasible, because the `se-resnext50` is expecting RGB, unless you are referring to the technique I spoke about here, or you are modifying the first part of `se-resnext50`to accept 1 channel instead of 3 channels, in which case you cannot use pretrained weights from the model and would have to basically reformat the entire architecture to work with 1 channel instead of the expected 3.",
    "745589": "Hey Corey! \nThanks for the kind words!! I learned from you a lot as well =) \n\nOk we have `1 channel` image and we want to use pertained model  on `imagenet` (e.g `se_resnext50_32x4d`).  Since all the current pertained models are trained to accept `3 channel ` image (`RGB`), we can solve our problem in 3 ways. \n\n`----------------------------------------------------`\n`Option 1: Convert 1 channel image two 3 channel `\nThis one is pretty easy as you suggested we can clone our image in to 3 channels and problem is solved.\n\n\n`----------------------------------------------------`\n`Option 2: Replace first Conv2D to accept 1 channel.`\nMost CNN's start with `Conv2D` and they accept `3 channel `input. If we want to feed our network with `1 channel` image we have to modify **ONLY** this layer. Below is the example how I will do This in Pytorch. \n\n`arch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)`\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F991320%2F046131f8b12d7a29fc35fbb2e9963a29%2FScreen%20Shot%202020-02-13%20at%207.31.26%20PM.png?generation=1581640446846621&amp;alt=media)\n\nAs you can see our first `Conv2d` accepts `3 channel` image and outputs `64 channel `. In order to make our model accept 1 channel image `ONLY` thing we have to do is create new Conv2D which accept `1 channel `and  outputs `64 channel `.  \n\nAnd this how It looks:\n```\n#loading our model\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\n#converting to list\narch = list(arch.children())\n#replacing first Conv2D\narch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)\n#converting back to sequential\narch = nn.Sequential(*arch)\n```\n\nThats it. Now disadvantage of this method is that we are initiating new `Conv2d` layer in the beginning of our network for which  we have to retrain weights from scratch. In some cases this can lead to slow convergence or weird spikes during training. In option 3 I show how we can solve this problem.\n\n\n`----------------------------------------------------`\n`Option 3: Replace first Conv2d to accept 1 channel, with pre trained weights. `\nInstead of initiating our `Conv2d` with new weights we can reuse imagenet weights. \nIf you remember our first `Conv2d`  accept 3 channels and outputs 64 channels. We can  do following: \n\n1) save the weights for the first `Covn2d` and average all the channels in to 1. \n2) create new `Conv2d` which accepts `1 channel`\n3) substitute newly `Conv2d` weights with averaged imagenet weights that we created early \n\nHere how it looks in the code.\n\n```\n#loading our model\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\n#converting to list\narch = list(arch.children())\n#saving the weights of the forst conv in w\nw = arch[0][0].weight\n#creating new Conv2d to accept 1 channel \narch[0][0] = nn.Conv2d(1, 64, kernel_size=7, stride=2, padding=3, bias=False)\n#substituting weights of newly created Conv2d with w from but we have to take mean\n#to go from  3 channel to 1\narch[0][0].weight = nn.Parameter(torch.mean(w, dim=1, keepdim=True))\narch = nn.Sequential(*arch)\n```\n\nThats it. This method resolve all the problems. We can train 1 channel images, We are still taking advantage of the pertained weights =) \n\nWhen I mentioned that I trained 1 channel image I was referring to  `option 3.` \n\nI hope this answered all your question. Let me know if something is not clear. \n\nP.S you can imagine that if you have `6 channel ` image we can do the same trick.  We preserve the 3 channel weights, create new `Conv2d` witch accepts 6 channel and just substitute it with stacked version of the saved weights. This is how it will look like \n\n```\narch = pretrainedmodels.se_resnext50_32x4d(num_classes=1000)\narch = list(arch.children())\nw = arch[0][0].weight\narch[0][0] = nn.Conv2d(6, 64, kernel_size=7, stride=2, padding=2, bias=False)\narch[0][0].weight = nn.Parameter(torch.stack([w,w] dim=1,))\narch = nn.Sequential(*arch)\n```",
    "745597": "drhabib Hi, Did you use ReduceOnPlatau or OneCycleLearning , In my experiment, I set the epoch to 120, but when the epoch equals 70 ~ 80, the learning rate is already very small, and the CV score is not much improved.",
    "745640": "Hi DrHB, thank you for a fantastic reply, and also thank you for your verbosity and also your code snippets; it really is helpful to someone like me who is a total noob in CNN.\n\nHonestly though, if you think about it, Option 3 isn't very attractive in my opinion. Averaging the weights I think can yield bad outcomes that have to be fixed and basically in my eyes make option 2 and option 3 equivalent. Here's how I think about it. Let's imagine we have weights `[w1, w2, w3, ..., w10]`. Well, on the Red color spectrum, `w1` might refer to characteristic `c1` and `w2` might refer to characteristic `c2`. However in the green spectrum, you might have `w1` refer to characteristic `c2` and `w2` refer to characteristic `c1`. I feel like averaging in this case does not really seem to yield something better than Option 2, however it feels more \"elegant\"...\n\nWhat is more attractive, in my opinion, is to use the concatenation of `red weights, green weights, blue weights` and use the concatenation of that information! I think that is more powerful. So it sounds like I would prefer Option 1.\n\nHowever, I don't really want the colors to \"interact\" with each other. For example the se-resnext model may have learned \"if red==blue, then DO THIS else DO THAT\". Unfortunately this type of interaction you just have to suck up and live with, and hope your model can correct some of these weights I guess.\n\nI am going to code up Option 1 on our team because despite the flaw I mentioned, I still think it would be more powerful than options 2 &amp; 3. Furthermore, I think I remember you posted that using 3 channels instead of 1 channel gave a tiny boost, so hopefully this intuition is correct\n\nThank you again for your contributions, and good luck in this competition",
    "745686": "returnofsputnik I have tested both option 1 and option 2. They do not make much difference. Also, I was wondering if I could join your team. I have competed in numerous CV competitions before, and I have a 2080 Ti and a lot of time to do testing. Thanks!",
    "745804": "Thank you! This is my list of \"What didn't work for me\":\n1. OHEM loss\n2. Class balanced loss\n3. Balanced data sampler\n4. AugMix - low, medium and high severity\n5. GridDropout",
    "745862": "Hi @drhabib , glad to see you here\nI used a probability of 25% for cutout and max 8 cutouts. But this was combined with either cutmix or mixup. I usually train for around 120 epochs, it takes a pretty long time and to be honest I never tried to leave it more epochs, maybe I will have a try to see if it helps",
    "745865": "Good feedback @dukhovnik \nDo you use mixup/cutmix both or just one of them ?",
    "745919": "random select [cutmix, mixup, gridmask] doesn't work for me. For me, gridmask works best.",
    "745927": "I am in progress with tuning them right now :)",
    "746007": "tonychenxyz (Gold Retriever) Thank you for your offer, please go to my profile and click 'Contact User' so we can handle that chat privately. I notified my teammate as well.",
    "746615": "returnofsputnik hi Corey, I think using  Op1 is the same as using Op2, as in this case, Op1 is using three same gray img, and the result of conv0, is the average of the orginal weight conv the img, which is Op3.",
    "746793": "liang23333 With cutmix and mixup I found that the model has to be trained for much longer (40 epochs -&gt; &gt;100 epochs). Are you observing the same effect with gridmask?",
    "746798": "drhabib Nice to see you here! I wonder what batch size you are using? For big batch sizes, same number of epochs correspond to much fewer steps, which might explain the relatively slow training. However small batch sizes have to correspond to smaller LR, so I am still in the process of tuning them.\n\nAlso, if you don't mind disclosing, do you find Onecycle or ReduceLRonPlateau more handy ?",
    "746831": "tonychenxyz Hey, i'm also having this problem, i can reach 0.982 cv but then i get 0.966lb. I'm using cutmix, and rotate as augments, maybe other augmentations will help? Only thing that help a little bit was adding weights in the loss function to handle imbalanced classes, then i get lower (0.974) cv but about 0.9685 lb! What are your thoughts? :)",
    "746872": "yannmajewski How many epochs are you training? Could you share what specfically you added in loss function? Thanks! I'm now trying bigger image size, which doesn't seem to work. I'm going to add more augmentation to see if anything improves. Another thing I noticed that is that the gap increases as epoch number increases.",
    "746899": "tonychenxyz im currently training for 100 epochs to test things out and see what works and what doesn't. And for the weights in the loss functions you have to calculate (1/nb of occurrences) of each class, so it will give you a tensor of the same size as the nb of classes. Then add this to the weight param in the loss function!\n\nKeep me updated on your tests! i'll do the same :)",
    "746985": "yannmajewski Cool! Thanks!",
    "747384": "Thank you for sharing your list of “things that did not work”. Here is my version of list.\n1.\tReplacing Adam with RAdam did not work for me. I did not try Adamw.\n2.\tCutmix, mixup and cutout did not work. Simple augmentation (such as shift and rotation) would be fine.\n3.\tIt is surprising that ResNet50 performed worse than VGG19 on my validation set. In the meanwhile, SEResNeXt101 performed better than VGG19.\n4.\tOneCycleLearning performed worse than ReduceLROnPlateau.\n5.\tCrossEntropy performed better than focal loss.",
    "747865": "tonychenxyz Quick question, are you using a pretrained model?",
    "748271": "Hi @drhabib, i also use seresnext50, and also have the same question with yours. I train 300 epoch, the CV is 0.9906, and i also focus that somebody get 0.99+ CV with little epoch, such as @garybios result, that's so strange, and i haven't found why it appear.",
    "748277": "For me:\n&gt; Cutmix gives a more boost than random select cutmix and mixup\n\n&gt; Radam converge a little quickly than Adam\n\n&gt; Adding GridDistortion RandomGamma, OpticalDistortion, GaussianBlur augment not  effect the result\n\n&gt; More epochs may give better result",
    "748288": "In my experiments including all seperate augmentations,  i use the best recall in validation dataset. And my optimizer is adam, learning rate is 1e-4. Which optimizer and lr do you choose?",
    "748521": "What lr scheduler do you use for 300 epochs ? @cswwp347724",
    "748822": "yannmajewski Yes. I'm using seresnext50 pretrained with imagenet.",
    "748851": "Quick update on my OHEM loss experiments: It performs worse than plain CE. It converges worse and causes spiking in CV as training progresses.",
    "749120": "greatgamedota  I used OneCycleLearning, but now it seems that ReduceOnPlateau may converge quickly",
    "749203": "cswwp347724 I realize that batch size has a large effect on training time. If it is convenient, what batch size are you using to train 300 epochs? It must be taking the hell of a long time :)",
    "749871": "vladvdv I wonder how is \"epochs without improvement\" defined? I have observed cases where validation metric and loss are both increasing :( It's very baffling",
    "750325": "those are impressive cv scores, I find the loss weighting of .5, .25, .25 between root, vowel and consonant not working for me. the vowel and consonant will become very good (.98+ easily in 30 epochs or so, but the root just gets stuck around .96)  struggling...",
    "752398": "Did you use different weights for CE loss?",
    "752583": "my batch size is 860",
    "753064": "1. We've tried mixup, cutmix, and 50% cutmix + 50% mixup, and 50% cutmix + 50% mixup works the best.\n2. cropping image doesn't work for us. In fact, we found that cropping image dropped the accuracy by about 0.005. I don't quite understand the reason behind it though.",
    "755354": "roguekk007 Thanks for you tips, I forget this point, maybe it's crucial, and i'm doing this experiment"
  },
  "source": "meta"
}