{
  "id": 93580,
  "title": "Many attempts, but no success",
  "url": "/competitions/imet-2019-fgvc6/discussion/93580",
  "author_name": "",
  "post_date": "2019-05-28T11:39:12.920673100Z",
  "votes": 4,
  "comment_count": 10,
  "views": 0,
  "content": "<p>I started this competition with the intention of learn fastai, pytorch and deep metric learning. \nI finally learnt about the two first and about Kaggle Kernels, and it was nice!</p>\n\n<p>For all my experiments I used fastai 1.0, thanks a lot to @miguelobpinto for the starter code</p>\n\n<p>I tried the next architectures:</p>\n\n<pre><code>- resnet50, resnet152\n\n- densenet121\n\n- GAPnets:  gap-resnet18, gap-resnet50, gap-se-resnext101\n\n- se-resnext101\n</code></pre>\n\n<p>I trained it:</p>\n\n<pre><code>- stratified K-fold (similar to Konstantin Lopuhin)\n\n- manual decreasing LR on plateau\n\n- using the always highest batch size that fits in memory: I think that having 1000 classes, a high batch size must be important\n\n- mixed FP32/FP16 \n\n- steadily increasing image size on each epoch (for se-resnexts and GAPnets), and decreasing the batch size to make it fit in memory\n\n- LabelSmoothingCrossEntropy + mixup  or using 0.5*loss_beta + 0.5*loss_focal\n\n- Augmentation.  I used default fastai get_trainsforms plus flip_vert/dihedral_affine. Also, tried to use an augmentation lower than the default, and increase all the parameters during the training.  I think augmentation was very important, and I had to invest more time here\n\n- increased default L2/WD, added dropout\n\n- oversample less represented classes\n</code></pre>\n\n<p>Doing it, I wasn't able to surpass the 0.56 CV score without applying any threshold.  Which is to low.</p>\n\n<p>Any tip or advice will be welcomed.</p>\n\n<p>I'm looking forward to read about your solutions after the competition!</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": "538303",
      "postDate": "05/28/2019 11:39:12",
      "content": "<p>I started this competition with the intention of learn fastai, pytorch and deep metric learning. \nI finally learnt about the two first and about Kaggle Kernels, and it was nice!</p>\n\n<p>For all my experiments I used fastai 1.0, thanks a lot to @miguelobpinto for the starter code</p>\n\n<p>I tried the next architectures:</p>\n\n<pre><code>- resnet50, resnet152\n\n- densenet121\n\n- GAPnets:  gap-resnet18, gap-resnet50, gap-se-resnext101\n\n- se-resnext101\n</code></pre>\n\n<p>I trained it:</p>\n\n<pre><code>- stratified K-fold (similar to Konstantin Lopuhin)\n\n- manual decreasing LR on plateau\n\n- using the always highest batch size that fits in memory: I think that having 1000 classes, a high batch size must be important\n\n- mixed FP32/FP16 \n\n- steadily increasing image size on each epoch (for se-resnexts and GAPnets), and decreasing the batch size to make it fit in memory\n\n- LabelSmoothingCrossEntropy + mixup  or using 0.5*loss_beta + 0.5*loss_focal\n\n- Augmentation.  I used default fastai get_trainsforms plus flip_vert/dihedral_affine. Also, tried to use an augmentation lower than the default, and increase all the parameters during the training.  I think augmentation was very important, and I had to invest more time here\n\n- increased default L2/WD, added dropout\n\n- oversample less represented classes\n</code></pre>\n\n<p>Doing it, I wasn't able to surpass the 0.56 CV score without applying any threshold.  Which is to low.</p>\n\n<p>Any tip or advice will be welcomed.</p>\n\n<p>I'm looking forward to read about your solutions after the competition!</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "I started this competition with the intention of learn fastai, pytorch and deep metric learning. \nI finally learnt about the two first and about Kaggle Kernels, and it was nice!\n\nFor all my experiments I used fastai 1.0, thanks a lot to @miguelobpinto for the starter code\n\nI tried the next architectures:\n\n\t- resnet50, resnet152\n\n\t- densenet121\n\n\t- GAPnets:  gap-resnet18, gap-resnet50, gap-se-resnext101\n\n\t- se-resnext101\n\nI trained it:\n\n\t- stratified K-fold (similar to Konstantin Lopuhin)\n\n\t- manual decreasing LR on plateau\n\n\t- using the always highest batch size that fits in memory: I think that having 1000 classes, a high batch size must be important\n\n\t- mixed FP32/FP16 \n\n\t- steadily increasing image size on each epoch (for se-resnexts and GAPnets), and decreasing the batch size to make it fit in memory\n\n\t- LabelSmoothingCrossEntropy + mixup  or using 0.5*loss_beta + 0.5*loss_focal\n\n\t- Augmentation.  I used default fastai get_trainsforms plus flip_vert/dihedral_affine. Also, tried to use an augmentation lower than the default, and increase all the parameters during the training.\tI think augmentation was very important, and I had to invest more time here\n\n\t- increased default L2/WD, added dropout\n\n\t- oversample less represented classes\n\nDoing it, I wasn't able to surpass the 0.56 CV score without applying any threshold.  Which is to low.\n\nAny tip or advice will be welcomed.\n\nI'm looking forward to read about your solutions after the competition!\n\nCheers",
      "votes": null
    },
    {
      "id": "538304",
      "postDate": "05/28/2019 11:40:58",
      "content": "<p>What was your score after a global threshold search?</p>",
      "rawMarkdown": "What was your score after a global threshold search?",
      "votes": null
    },
    {
      "id": "538307",
      "postDate": "05/28/2019 11:46:23",
      "content": "<p>0.585 after thresholding</p>",
      "rawMarkdown": "0.585 after thresholding",
      "votes": null
    },
    {
      "id": "538310",
      "postDate": "05/28/2019 11:49:00",
      "content": "<p>Depending on your scale/crop configuration, TTA might also boost your score significantly.</p>",
      "rawMarkdown": "Depending on your scale/crop configuration, TTA might also boost your score significantly.",
      "votes": null
    },
    {
      "id": "538315",
      "postDate": "05/28/2019 12:03:46",
      "content": "<p>how significantly?</p>",
      "rawMarkdown": "how significantly?",
      "votes": null
    },
    {
      "id": "538322",
      "postDate": "05/28/2019 12:21:46",
      "content": "<p>for up to 0.03 cv but as I said it depends on how you train</p>",
      "rawMarkdown": "for up to 0.03 cv but as I said it depends on how you train",
      "votes": null
    },
    {
      "id": "538770",
      "postDate": "05/29/2019 05:11:21",
      "content": "<p>For mixed precision training, if you're using SGDM it's fine. However, with Adam, the result may suffer from numerical instability. So if you have enough computational resources, I'd recommend you still use FP32 mode and it's also easier for you to debug your code because it's really hard to find out numerical problems.</p>\n\n<p>As for batch size, it's not the larger the better thing. A large batch size can lead to less variance in gradients and this means that it's less possible to escape from local minima or saddle points. Moreover, large batch size means less parameter updates and worse convergence rate.</p>",
      "rawMarkdown": "For mixed precision training, if you're using SGDM it's fine. However, with Adam, the result may suffer from numerical instability. So if you have enough computational resources, I'd recommend you still use FP32 mode and it's also easier for you to debug your code because it's really hard to find out numerical problems.\n\nAs for batch size, it's not the larger the better thing. A large batch size can lead to less variance in gradients and this means that it's less possible to escape from local minima or saddle points. Moreover, large batch size means less parameter updates and worse convergence rate.",
      "votes": null
    },
    {
      "id": "538810",
      "postDate": "05/29/2019 06:30:00",
      "content": "<p>thanks a lot <a href=\"/yaroshevskiy\">@yaroshevskiy</a>  !!!</p>\n\n<p>I've to read fastai <a href=\"https://docs.fast.ai/vision.transform.html#get_transforms\">get_transforms()</a> twice.  </p>\n\n<p>I remember it uses some randomness regarding ....</p>\n\n<p>And I remember bestfitting used \"<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#523162\">no other augmentation, randomly cropping from 768x768 then resize to 512x512 was enough,we should use external data</a>\" for Human Protein Atlas Image Classification</p>\n\n<p>So I'd like to give it a try with fixed size cropping and resizing. </p>\n\n<p>Did you use fixed size cropping / scale?</p>",
      "rawMarkdown": "thanks a lot @yaroshevskiy  !!!\n\nI've to read fastai [get_transforms()](https://docs.fast.ai/vision.transform.html#get_transforms) twice.  \n\nI remember it uses some randomness regarding ....\n\nAnd I remember bestfitting used \"[no other augmentation, randomly cropping from 768x768 then resize to 512x512 was enough,we should use external data](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#523162)\" for Human Protein Atlas Image Classification\n\nSo I'd like to give it a try with fixed size cropping and resizing. \n\nDid you use fixed size cropping / scale?",
      "votes": null
    },
    {
      "id": "538820",
      "postDate": "05/29/2019 06:42:08",
      "content": "<p>Thanks a lot for your advice <a href=\"/syoya1997\">@syoya1997</a>  !!!</p>\n\n<p>I tried FP16 for Resnet-152 with success but for Se-ResNext101 it got quickly a 0.45 score, but after that, I started having NaNs for the metric.\nI'll try to move to SGDM, or to stop using mixed FP32/FP16.</p>\n\n<p>Also, I'll try to find a better lower batch-size as you recommended.</p>\n\n<p>In addition, do you think it is a good idea to change LR and batch size at the same time during the training?</p>\n\n<p>Even more, changing the size of the input image at the same time of the LR and batch size?</p>",
      "rawMarkdown": "Thanks a lot for your advice @syoya1997  !!!\n\nI tried FP16 for Resnet-152 with success but for Se-ResNext101 it got quickly a 0.45 score, but after that, I started having NaNs for the metric.\nI'll try to move to SGDM, or to stop using mixed FP32/FP16.\n\nAlso, I'll try to find a better lower batch-size as you recommended.\n\nIn addition, do you think it is a good idea to change LR and batch size at the same time during the training?\n\nEven more, changing the size of the input image at the same time of the LR and batch size?",
      "votes": null
    },
    {
      "id": "538828",
      "postDate": "05/29/2019 06:51:10",
      "content": "<p>I didn't read much paper about this part. But for batch size and learning rate, maybe you can refer to this one: <a href=\"https://arxiv.org/abs/1609.04836\">On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima</a>. And I also find one paper containing multiple tricks for training classification models: <a href=\"https://arxiv.org/abs/1812.01187\">Bag of Tricks for Image Classification with Convolutional Neural Networks</a>. Hope this can help.</p>",
      "rawMarkdown": "I didn't read much paper about this part. But for batch size and learning rate, maybe you can refer to this one: [On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima](https://arxiv.org/abs/1609.04836). And I also find one paper containing multiple tricks for training classification models: [Bag of Tricks for Image Classification with Convolutional Neural Networks](https://arxiv.org/abs/1812.01187). Hope this can help.",
      "votes": null
    },
    {
      "id": "539076",
      "postDate": "05/29/2019 13:36:02",
      "content": "<p>I do random resize of original image to f.e. 256-512 size and then randomly crop 224. I do same for my inference.</p>",
      "rawMarkdown": "I do random resize of original image to f.e. 256-512 size and then randomly crop 224. I do same for my inference.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 538304,
      "author_name": "yaroshevskiy",
      "author_url": "",
      "post_date": "05/28/2019 11:40:58",
      "content": "<p>What was your score after a global threshold search?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 538307,
      "author_name": "virilo",
      "author_url": "",
      "post_date": "05/28/2019 11:46:23",
      "content": "<p>0.585 after thresholding</p>",
      "votes": null,
      "replies": [
        {
          "id": 538310,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "05/28/2019 11:49:00",
          "content": "<p>Depending on your scale/crop configuration, TTA might also boost your score significantly.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538315,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "05/28/2019 12:03:46",
          "content": "<p>how significantly?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538322,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "05/28/2019 12:21:46",
          "content": "<p>for up to 0.03 cv but as I said it depends on how you train</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538810,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "05/29/2019 06:30:00",
          "content": "<p>thanks a lot <a href=\"/yaroshevskiy\">@yaroshevskiy</a>  !!!</p>\n\n<p>I've to read fastai <a href=\"https://docs.fast.ai/vision.transform.html#get_transforms\">get_transforms()</a> twice.  </p>\n\n<p>I remember it uses some randomness regarding ....</p>\n\n<p>And I remember bestfitting used \"<a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#523162\">no other augmentation, randomly cropping from 768x768 then resize to 512x512 was enough,we should use external data</a>\" for Human Protein Atlas Image Classification</p>\n\n<p>So I'd like to give it a try with fixed size cropping and resizing. </p>\n\n<p>Did you use fixed size cropping / scale?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 539076,
          "author_name": "yaroshevskiy",
          "author_url": "",
          "post_date": "05/29/2019 13:36:02",
          "content": "<p>I do random resize of original image to f.e. 256-512 size and then randomly crop 224. I do same for my inference.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 538770,
      "author_name": "syoya1997",
      "author_url": "",
      "post_date": "05/29/2019 05:11:21",
      "content": "<p>For mixed precision training, if you're using SGDM it's fine. However, with Adam, the result may suffer from numerical instability. So if you have enough computational resources, I'd recommend you still use FP32 mode and it's also easier for you to debug your code because it's really hard to find out numerical problems.</p>\n\n<p>As for batch size, it's not the larger the better thing. A large batch size can lead to less variance in gradients and this means that it's less possible to escape from local minima or saddle points. Moreover, large batch size means less parameter updates and worse convergence rate.</p>",
      "votes": null,
      "replies": [
        {
          "id": 538820,
          "author_name": "virilo",
          "author_url": "",
          "post_date": "05/29/2019 06:42:08",
          "content": "<p>Thanks a lot for your advice <a href=\"/syoya1997\">@syoya1997</a>  !!!</p>\n\n<p>I tried FP16 for Resnet-152 with success but for Se-ResNext101 it got quickly a 0.45 score, but after that, I started having NaNs for the metric.\nI'll try to move to SGDM, or to stop using mixed FP32/FP16.</p>\n\n<p>Also, I'll try to find a better lower batch-size as you recommended.</p>\n\n<p>In addition, do you think it is a good idea to change LR and batch size at the same time during the training?</p>\n\n<p>Even more, changing the size of the input image at the same time of the LR and batch size?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 538828,
          "author_name": "syoya1997",
          "author_url": "",
          "post_date": "05/29/2019 06:51:10",
          "content": "<p>I didn't read much paper about this part. But for batch size and learning rate, maybe you can refer to this one: <a href=\"https://arxiv.org/abs/1609.04836\">On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima</a>. And I also find one paper containing multiple tricks for training classification models: <a href=\"https://arxiv.org/abs/1812.01187\">Bag of Tricks for Image Classification with Convolutional Neural Networks</a>. Hope this can help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "538303": "I started this competition with the intention of learn fastai, pytorch and deep metric learning. \nI finally learnt about the two first and about Kaggle Kernels, and it was nice!\n\nFor all my experiments I used fastai 1.0, thanks a lot to @miguelobpinto for the starter code\n\nI tried the next architectures:\n\n\t- resnet50, resnet152\n\n\t- densenet121\n\n\t- GAPnets:  gap-resnet18, gap-resnet50, gap-se-resnext101\n\n\t- se-resnext101\n\nI trained it:\n\n\t- stratified K-fold (similar to Konstantin Lopuhin)\n\n\t- manual decreasing LR on plateau\n\n\t- using the always highest batch size that fits in memory: I think that having 1000 classes, a high batch size must be important\n\n\t- mixed FP32/FP16 \n\n\t- steadily increasing image size on each epoch (for se-resnexts and GAPnets), and decreasing the batch size to make it fit in memory\n\n\t- LabelSmoothingCrossEntropy + mixup  or using 0.5*loss_beta + 0.5*loss_focal\n\n\t- Augmentation.  I used default fastai get_trainsforms plus flip_vert/dihedral_affine. Also, tried to use an augmentation lower than the default, and increase all the parameters during the training.\tI think augmentation was very important, and I had to invest more time here\n\n\t- increased default L2/WD, added dropout\n\n\t- oversample less represented classes\n\nDoing it, I wasn't able to surpass the 0.56 CV score without applying any threshold.  Which is to low.\n\nAny tip or advice will be welcomed.\n\nI'm looking forward to read about your solutions after the competition!\n\nCheers",
    "538304": "What was your score after a global threshold search?",
    "538307": "0.585 after thresholding",
    "538310": "Depending on your scale/crop configuration, TTA might also boost your score significantly.",
    "538315": "how significantly?",
    "538322": "for up to 0.03 cv but as I said it depends on how you train",
    "538770": "For mixed precision training, if you're using SGDM it's fine. However, with Adam, the result may suffer from numerical instability. So if you have enough computational resources, I'd recommend you still use FP32 mode and it's also easier for you to debug your code because it's really hard to find out numerical problems.\n\nAs for batch size, it's not the larger the better thing. A large batch size can lead to less variance in gradients and this means that it's less possible to escape from local minima or saddle points. Moreover, large batch size means less parameter updates and worse convergence rate.",
    "538810": "thanks a lot @yaroshevskiy  !!!\n\nI've to read fastai [get_transforms()](https://docs.fast.ai/vision.transform.html#get_transforms) twice.  \n\nI remember it uses some randomness regarding ....\n\nAnd I remember bestfitting used \"[no other augmentation, randomly cropping from 768x768 then resize to 512x512 was enough,we should use external data](https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/78109#523162)\" for Human Protein Atlas Image Classification\n\nSo I'd like to give it a try with fixed size cropping and resizing. \n\nDid you use fixed size cropping / scale?",
    "538820": "Thanks a lot for your advice @syoya1997  !!!\n\nI tried FP16 for Resnet-152 with success but for Se-ResNext101 it got quickly a 0.45 score, but after that, I started having NaNs for the metric.\nI'll try to move to SGDM, or to stop using mixed FP32/FP16.\n\nAlso, I'll try to find a better lower batch-size as you recommended.\n\nIn addition, do you think it is a good idea to change LR and batch size at the same time during the training?\n\nEven more, changing the size of the input image at the same time of the LR and batch size?",
    "538828": "I didn't read much paper about this part. But for batch size and learning rate, maybe you can refer to this one: [On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima](https://arxiv.org/abs/1609.04836). And I also find one paper containing multiple tricks for training classification models: [Bag of Tricks for Image Classification with Convolutional Neural Networks](https://arxiv.org/abs/1812.01187). Hope this can help.",
    "539076": "I do random resize of original image to f.e. 256-512 size and then randomly crop 224. I do same for my inference."
  },
  "source": "meta"
}