{
  "id": 129790,
  "title": "se_resnext50_32x4d  vs se_resnext101_32x4d",
  "url": "/competitions/bengaliai-cv19/discussion/129790",
  "author_name": "Vlad Vaduva",
  "post_date": "2020-02-10T19:11:03.488000",
  "votes": 27,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi guys,</p>\n\n<p>I want to share a little of my experiences with se_resnext50 and se_resnext101.\nI worked in parallel with both architectures  and made identical tests to see which gives better in identical situations</p>\n\n<p><strong>First test</strong>\n*<em>arhitecture</em>*: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n<strong>image size</strong>: 96x96\n<strong>augmentation</strong>: scale(0.85,1.15) , rotation(5 degress), translation(4), sheer(10 degres) + small gausian noise\n<strong>Test results</strong>: se_resnext101(0.9620 for a fold) &gt; se_resnext50(0.9611 for a fold)</p>\n\n<p><strong>Second test</strong>\n*<em>arhitecture</em>*: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n<strong>image size</strong>: 96x96\n<strong>augmentation</strong>: mixup/cutmix\n<strong>Test results</strong>: se_resnext101(0.9655 for a fold) &gt; se_resnext50(0.9441 for a fold)</p>\n\n<p>So, from this tests I wanted to see if:\na) is it se_resnext101 is too big for our dataset and leads to overfit\nb) is it se_resnext50 too small for the complexity needed for our problem\nc) is really mixup/cutmix a better solution than let's say normal augmenting (affine transforms+noises)</p>\n\n<p>Conclusions\na) se_resnext101 seems to not be a too big architecture for our dataset considering that we augment the data and did not lead to overfiting \nb) In my trials se_resnext50 seemed to be to small (underfit), I will probably try to add extra layers or neurons after the se_resnext50 output to increase the complexity and analyze the overfit/underfit ratio\nc) it seems that in this case, yes is the answer, mixup/cutmix really gives better results than classical augmenting methods. Also, what is interesting is that when I combined mixup/cutmix with the affine transforms the results went down, probably the images are way too different and the algorithm is having a hard time understanding what is what</p>\n\n<p>To try:\na) Increase the se_resnext50 based model complexity by adding extra layers or neurons\nb) Combine mixup/cutmix with some noise transforms\nc) Try augMix, I understand that it give good results to some of you guys\nd) Increase image size to 128x128</p>\n\n<p>Cheers</p>",
  "messages": [
    {
      "id": 741540,
      "postDate": "2020-02-10T19:11:03.490Z",
      "content": "<p>Hi guys,</p>\n\n<p>I want to share a little of my experiences with se_resnext50 and se_resnext101.\nI worked in parallel with both architectures  and made identical tests to see which gives better in identical situations</p>\n\n<p><strong>First test</strong>\n*<em>arhitecture</em>*: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n<strong>image size</strong>: 96x96\n<strong>augmentation</strong>: scale(0.85,1.15) , rotation(5 degress), translation(4), sheer(10 degres) + small gausian noise\n<strong>Test results</strong>: se_resnext101(0.9620 for a fold) &gt; se_resnext50(0.9611 for a fold)</p>\n\n<p><strong>Second test</strong>\n*<em>arhitecture</em>*: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n<strong>image size</strong>: 96x96\n<strong>augmentation</strong>: mixup/cutmix\n<strong>Test results</strong>: se_resnext101(0.9655 for a fold) &gt; se_resnext50(0.9441 for a fold)</p>\n\n<p>So, from this tests I wanted to see if:\na) is it se_resnext101 is too big for our dataset and leads to overfit\nb) is it se_resnext50 too small for the complexity needed for our problem\nc) is really mixup/cutmix a better solution than let's say normal augmenting (affine transforms+noises)</p>\n\n<p>Conclusions\na) se_resnext101 seems to not be a too big architecture for our dataset considering that we augment the data and did not lead to overfiting \nb) In my trials se_resnext50 seemed to be to small (underfit), I will probably try to add extra layers or neurons after the se_resnext50 output to increase the complexity and analyze the overfit/underfit ratio\nc) it seems that in this case, yes is the answer, mixup/cutmix really gives better results than classical augmenting methods. Also, what is interesting is that when I combined mixup/cutmix with the affine transforms the results went down, probably the images are way too different and the algorithm is having a hard time understanding what is what</p>\n\n<p>To try:\na) Increase the se_resnext50 based model complexity by adding extra layers or neurons\nb) Combine mixup/cutmix with some noise transforms\nc) Try augMix, I understand that it give good results to some of you guys\nd) Increase image size to 128x128</p>\n\n<p>Cheers</p>",
      "rawMarkdown": "Hi guys,\n\nI want to share a little of my experiences with se_resnext50 and se_resnext101.\nI worked in parallel with both architectures  and made identical tests to see which gives better in identical situations\n\n**First test**\n**arhitecture**: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n**image size**: 96x96\n**augmentation**: scale(0.85,1.15) , rotation(5 degress), translation(4), sheer(10 degres) + small gausian noise\n**Test results**: se_resnext101(0.9620 for a fold) &gt; se_resnext50(0.9611 for a fold)\n\n**Second test**\n**arhitecture**: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n**image size**: 96x96\n**augmentation**: mixup/cutmix\n**Test results**: se_resnext101(0.9655 for a fold) &gt; se_resnext50(0.9441 for a fold)\n\nSo, from this tests I wanted to see if:\na) is it se_resnext101 is too big for our dataset and leads to overfit\nb) is it se_resnext50 too small for the complexity needed for our problem\nc) is really mixup/cutmix a better solution than let's say normal augmenting (affine transforms+noises)\n\nConclusions\na) se_resnext101 seems to not be a too big architecture for our dataset considering that we augment the data and did not lead to overfiting \nb) In my trials se_resnext50 seemed to be to small (underfit), I will probably try to add extra layers or neurons after the se_resnext50 output to increase the complexity and analyze the overfit/underfit ratio\nc) it seems that in this case, yes is the answer, mixup/cutmix really gives better results than classical augmenting methods. Also, what is interesting is that when I combined mixup/cutmix with the affine transforms the results went down, probably the images are way too different and the algorithm is having a hard time understanding what is what\n\nTo try:\na) Increase the se_resnext50 based model complexity by adding extra layers or neurons\nb) Combine mixup/cutmix with some noise transforms\nc) Try augMix, I understand that it give good results to some of you guys\nd) Increase image size to 128x128\n\n\nCheers",
      "votes": 27
    },
    {
      "id": 750736,
      "postDate": "2020-02-19T16:42:54.770Z",
      "content": "<p>Interesting, I replaced my seresnext50 with 101 and trained with exact same parameters and it increased my score quite a bit even when I didn't hit the a record PB CV. LB: .9683 CV: .9769</p>",
      "rawMarkdown": "Interesting, I replaced my seresnext50 with 101 and trained with exact same parameters and it increased my score quite a bit even when I didn't hit the a record PB CV. LB: .9683 CV: .9769",
      "votes": 1
    },
    {
      "id": 743098,
      "postDate": "2020-02-11T19:37:12.697Z",
      "content": "<p>I wanted to try se-resnext but as I was using Keras, It was difficult for me to build up the architecture. I wonder if anyone tried already. </p>",
      "rawMarkdown": "I wanted to try se-resnext but as I was using Keras, It was difficult for me to build up the architecture. I wonder if anyone tried already. ",
      "votes": 1,
      "replies": [
        {
          "id": 743130,
          "postDate": "2020-02-11T20:14:47.540Z",
          "content": "<p><a href=\"https://github.com/osmr/imgclsmob/tree/master/tensorflow2\">Here</a> you can find tensorflow2 implementation of se-resnext models (and many others)</p>",
          "rawMarkdown": "[Here](https://github.com/osmr/imgclsmob/tree/master/tensorflow2) you can find tensorflow2 implementation of se-resnext models (and many others)",
          "votes": 1
        },
        {
          "id": 743146,
          "postDate": "2020-02-11T20:43:06.117Z",
          "content": "<p>Have you tried? </p>",
          "rawMarkdown": "Have you tried? "
        },
        {
          "id": 743671,
          "postDate": "2020-02-12T07:16:21.167Z",
          "content": "<p>yes, I did. here is one of my experiments:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4033611%2F23d77818da60f745a33da1208c7cd400%2Fs3d7.png?generation=1581491739245854&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "yes, I did. here is one of my experiments:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4033611%2F23d77818da60f745a33da1208c7cd400%2Fs3d7.png?generation=1581491739245854&amp;alt=media)\n",
          "votes": 6
        }
      ]
    },
    {
      "id": 741792,
      "postDate": "2020-02-11T00:27:17.753Z",
      "content": "<p>In second test, how many epochs did you train your models?</p>",
      "rawMarkdown": "In second test, how many epochs did you train your models?",
      "votes": 1,
      "replies": [
        {
          "id": 742199,
          "postDate": "2020-02-11T06:31:32.883Z",
          "content": "<p>100 epochs, mixup/cutmix requires more time than usual augmentation for reaching it's peak</p>",
          "rawMarkdown": "100 epochs, mixup/cutmix requires more time than usual augmentation for reaching it's peak"
        },
        {
          "id": 742355,
          "postDate": "2020-02-11T08:55:43.480Z",
          "content": "<p>How long is an epoch taking for training 101?</p>",
          "rawMarkdown": "How long is an epoch taking for training 101?"
        },
        {
          "id": 742395,
          "postDate": "2020-02-11T09:23:44.037Z",
          "content": "<p>On a 1070Ti around 25 mins on image size 96x96 and on a 2080Ti around 22 mins on image size 128x128</p>",
          "rawMarkdown": "On a 1070Ti around 25 mins on image size 96x96 and on a 2080Ti around 22 mins on image size 128x128",
          "votes": 3
        }
      ]
    },
    {
      "id": 743119,
      "postDate": "2020-02-11T20:05:06.763Z",
      "content": "<p>Have you tried wide resnet50? \nI got better result on wide resnet than se_resnet50 in same situations.</p>",
      "rawMarkdown": "Have you tried wide resnet50? \nI got better result on wide resnet than se_resnet50 in same situations.",
      "replies": [
        {
          "id": 743148,
          "postDate": "2020-02-11T20:45:52.907Z",
          "content": "<p>No, I did not but I will try it. It is difficult to tune a lot of things, training takes a lot of time and you have to be careful what you choose to tune and with what values, it is not enough time to try a lot of things (unless for the guys with access to several servers 😄 )</p>",
          "rawMarkdown": "No, I did not but I will try it. It is difficult to tune a lot of things, training takes a lot of time and you have to be careful what you choose to tune and with what values, it is not enough time to try a lot of things (unless for the guys with access to several servers 😄 )"
        },
        {
          "id": 743351,
          "postDate": "2020-02-12T02:16:03.453Z",
          "content": "<p>WRN 50 performs worse for me, do you need to train it longer for it to converge?</p>",
          "rawMarkdown": "WRN 50 performs worse for me, do you need to train it longer for it to converge?"
        },
        {
          "id": 743633,
          "postDate": "2020-02-12T06:41:03.210Z",
          "content": "<p>I tested WRN50 to choose best baseline model.Don't know whether it degrades after any sort of augmentations.\nConvergence time was almost same as se_resnext50.</p>",
          "rawMarkdown": "I tested WRN50 to choose best baseline model.Don't know whether it degrades after any sort of augmentations.\nConvergence time was almost same as se_resnext50."
        }
      ]
    },
    {
      "id": 742706,
      "postDate": "2020-02-11T13:54:31.943Z",
      "content": "<p>Do you use different learning rates for your pretrained mode and for classifier?</p>",
      "rawMarkdown": "Do you use different learning rates for your pretrained mode and for classifier?",
      "replies": [
        {
          "id": 742785,
          "postDate": "2020-02-11T14:38:58.613Z",
          "content": "<p>No, I used the same learning rate with ReduceOnPlateu. Did you get better results with different learning rates ?</p>",
          "rawMarkdown": "No, I used the same learning rate with ReduceOnPlateu. Did you get better results with different learning rates ?"
        },
        {
          "id": 742847,
          "postDate": "2020-02-11T15:33:27.923Z",
          "content": "<p>Can you please share LR. We not training for as many epochs, stop after like 20 or 25. So wondering if LR is too high.</p>",
          "rawMarkdown": "Can you please share LR. We not training for as many epochs, stop after like 20 or 25. So wondering if LR is too high."
        },
        {
          "id": 742861,
          "postDate": "2020-02-11T15:49:29.993Z",
          "content": "<p>Sure <a href=\"/returnofsputnik\">@returnofsputnik</a>  . It is 0.001</p>",
          "rawMarkdown": "Sure @returnofsputnik  . It is 0.001",
          "votes": 1
        },
        {
          "id": 742876,
          "postDate": "2020-02-11T16:07:56.653Z",
          "content": "<p>What optimizer are you using ?</p>",
          "rawMarkdown": "What optimizer are you using ?"
        },
        {
          "id": 743195,
          "postDate": "2020-02-11T21:52:30.190Z",
          "content": "<p>sgd, and thank you for answering.</p>",
          "rawMarkdown": "sgd, and thank you for answering."
        }
      ]
    },
    {
      "id": 749294,
      "postDate": "2020-02-18T14:46:08.783Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 746981,
      "postDate": "2020-02-15T20:10:22.573Z",
      "content": "<p>Thanks for sharing, very informative!</p>",
      "rawMarkdown": "Thanks for sharing, very informative!"
    }
  ],
  "comments": [
    {
      "id": 750736,
      "author_name": "GreatGameDota",
      "author_url": "",
      "post_date": "2020-02-19T16:42:54.770000",
      "content": "<p>Interesting, I replaced my seresnext50 with 101 and trained with exact same parameters and it increased my score quite a bit even when I didn't hit the a record PB CV. LB: .9683 CV: .9769</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 743098,
      "author_name": "Innat",
      "author_url": "",
      "post_date": "2020-02-11T19:37:12.697000",
      "content": "<p>I wanted to try se-resnext but as I was using Keras, It was difficult for me to build up the architecture. I wonder if anyone tried already. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 743130,
          "author_name": "Andrey Zotov",
          "author_url": "",
          "post_date": "2020-02-11T20:14:47.540000",
          "content": "<p><a href=\"https://github.com/osmr/imgclsmob/tree/master/tensorflow2\">Here</a> you can find tensorflow2 implementation of se-resnext models (and many others)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 743146,
          "author_name": "Innat",
          "author_url": "",
          "post_date": "2020-02-11T20:43:06.117000",
          "content": "<p>Have you tried? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743671,
          "author_name": "Andrey Zotov",
          "author_url": "",
          "post_date": "2020-02-12T07:16:21.167000",
          "content": "<p>yes, I did. here is one of my experiments:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F4033611%2F23d77818da60f745a33da1208c7cd400%2Fs3d7.png?generation=1581491739245854&amp;alt=media\" alt=\"\"></p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 741792,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2020-02-11T00:27:17.753000",
      "content": "<p>In second test, how many epochs did you train your models?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 742199,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-11T06:31:32.883000",
          "content": "<p>100 epochs, mixup/cutmix requires more time than usual augmentation for reaching it's peak</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742355,
          "author_name": "timetraveller",
          "author_url": "",
          "post_date": "2020-02-11T08:55:43.480000",
          "content": "<p>How long is an epoch taking for training 101?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742395,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-11T09:23:44.037000",
          "content": "<p>On a 1070Ti around 25 mins on image size 96x96 and on a 2080Ti around 22 mins on image size 128x128</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 743119,
      "author_name": "Md.Abrar Istiak Akib",
      "author_url": "",
      "post_date": "2020-02-11T20:05:06.763000",
      "content": "<p>Have you tried wide resnet50? \nI got better result on wide resnet than se_resnet50 in same situations.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 743148,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-11T20:45:52.907000",
          "content": "<p>No, I did not but I will try it. It is difficult to tune a lot of things, training takes a lot of time and you have to be careful what you choose to tune and with what values, it is not enough time to try a lot of things (unless for the guys with access to several servers 😄 )</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743351,
          "author_name": "GreatGameDota",
          "author_url": "",
          "post_date": "2020-02-12T02:16:03.453000",
          "content": "<p>WRN 50 performs worse for me, do you need to train it longer for it to converge?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743633,
          "author_name": "Md.Abrar Istiak Akib",
          "author_url": "",
          "post_date": "2020-02-12T06:41:03.210000",
          "content": "<p>I tested WRN50 to choose best baseline model.Don't know whether it degrades after any sort of augmentations.\nConvergence time was almost same as se_resnext50.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 742706,
      "author_name": "Volodymyr",
      "author_url": "",
      "post_date": "2020-02-11T13:54:31.943000",
      "content": "<p>Do you use different learning rates for your pretrained mode and for classifier?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 742785,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-11T14:38:58.613000",
          "content": "<p>No, I used the same learning rate with ReduceOnPlateu. Did you get better results with different learning rates ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742847,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-02-11T15:33:27.923000",
          "content": "<p>Can you please share LR. We not training for as many epochs, stop after like 20 or 25. So wondering if LR is too high.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 742861,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-11T15:49:29.993000",
          "content": "<p>Sure <a href=\"/returnofsputnik\">@returnofsputnik</a>  . It is 0.001</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 742876,
          "author_name": "Vlad Vaduva",
          "author_url": "",
          "post_date": "2020-02-11T16:07:56.653000",
          "content": "<p>What optimizer are you using ?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 743195,
          "author_name": "CoreyJamesLevinson",
          "author_url": "",
          "post_date": "2020-02-11T21:52:30.190000",
          "content": "<p>sgd, and thank you for answering.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 749294,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-18T14:46:08.783000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 746981,
      "author_name": "Helen",
      "author_url": "",
      "post_date": "2020-02-15T20:10:22.573000",
      "content": "<p>Thanks for sharing, very informative!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "741540": "Hi guys,\n\nI want to share a little of my experiences with se_resnext50 and se_resnext101.\nI worked in parallel with both architectures  and made identical tests to see which gives better in identical situations\n\n**First test**\n**arhitecture**: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n**image size**: 96x96\n**augmentation**: scale(0.85,1.15) , rotation(5 degress), translation(4), sheer(10 degres) + small gausian noise\n**Test results**: se_resnext101(0.9620 for a fold) &gt; se_resnext50(0.9611 for a fold)\n\n**Second test**\n**arhitecture**: 1 conv layer + se_resnext101/ se_resnext50+2 dense layers\n**image size**: 96x96\n**augmentation**: mixup/cutmix\n**Test results**: se_resnext101(0.9655 for a fold) &gt; se_resnext50(0.9441 for a fold)\n\nSo, from this tests I wanted to see if:\na) is it se_resnext101 is too big for our dataset and leads to overfit\nb) is it se_resnext50 too small for the complexity needed for our problem\nc) is really mixup/cutmix a better solution than let's say normal augmenting (affine transforms+noises)\n\nConclusions\na) se_resnext101 seems to not be a too big architecture for our dataset considering that we augment the data and did not lead to overfiting \nb) In my trials se_resnext50 seemed to be to small (underfit), I will probably try to add extra layers or neurons after the se_resnext50 output to increase the complexity and analyze the overfit/underfit ratio\nc) it seems that in this case, yes is the answer, mixup/cutmix really gives better results than classical augmenting methods. Also, what is interesting is that when I combined mixup/cutmix with the affine transforms the results went down, probably the images are way too different and the algorithm is having a hard time understanding what is what\n\nTo try:\na) Increase the se_resnext50 based model complexity by adding extra layers or neurons\nb) Combine mixup/cutmix with some noise transforms\nc) Try augMix, I understand that it give good results to some of you guys\nd) Increase image size to 128x128\n\n\nCheers",
    "750736": "Interesting, I replaced my seresnext50 with 101 and trained with exact same parameters and it increased my score quite a bit even when I didn't hit the a record PB CV. LB: .9683 CV: .9769",
    "743098": "I wanted to try se-resnext but as I was using Keras, It was difficult for me to build up the architecture. I wonder if anyone tried already. ",
    "741792": "In second test, how many epochs did you train your models?",
    "743119": "Have you tried wide resnet50? \nI got better result on wide resnet than se_resnet50 in same situations.",
    "742706": "Do you use different learning rates for your pretrained mode and for classifier?",
    "749294": "",
    "746981": "Thanks for sharing, very informative!"
  }
}