{
  "id": 124249,
  "title": "Shake-Shake ",
  "url": "/competitions/bengaliai-cv19/discussion/124249",
  "author_name": "",
  "post_date": "2020-01-02T21:44:09.162047Z",
  "votes": 9,
  "comment_count": 6,
  "views": 0,
  "content": "<p>Why  should we do augmentations only for image where we can do augmentations for features as well to mix things up ! . Based on this idea only shake-shake regularization has been proposed . \nShake-Shake gave SOTA performance on Kuzushiji handwritten characters . Therefore the motivation for using it for Bengali . \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F612614a264786e69efbba4462984f015%2FSOTA.PNG?generation=1578001385829597&amp;alt=media\" alt=\"\"></p>\n\n<p>Here is the discussion where Heng Mentioned it  : \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757</a></p>\n\n<p>Here is the paper: \n<a href=\"https://arxiv.org/abs/1705.07485\">https://arxiv.org/abs/1705.07485</a></p>\n\n<p>Here is my adaptation of the code  :\n<a href=\"https://www.kaggle.com/phoenix9032/shake-shake-regularization-starter-code\">https://www.kaggle.com/phoenix9032/shake-shake-regularization-starter-code</a></p>\n\n<p>I dont have any GPU left , so I have executed it only for 1 epoch . You can change the model_config and run it as much as you want . I have also switched the normal image augmentation off to save around 10 minutes per epoch . Anyone adopting it , can make it ON.</p>\n\n<p>Do let me know if it does any good .</p>",
  "messages": [
    {
      "id": "708953",
      "postDate": "01/02/2020 21:44:09",
      "content": "<p>Why  should we do augmentations only for image where we can do augmentations for features as well to mix things up ! . Based on this idea only shake-shake regularization has been proposed . \nShake-Shake gave SOTA performance on Kuzushiji handwritten characters . Therefore the motivation for using it for Bengali . \n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F612614a264786e69efbba4462984f015%2FSOTA.PNG?generation=1578001385829597&amp;alt=media\" alt=\"\"></p>\n\n<p>Here is the discussion where Heng Mentioned it  : \n<a href=\"https://www.kaggle.com/c/bengaliai-cv19/discussion/123757\">https://www.kaggle.com/c/bengaliai-cv19/discussion/123757</a></p>\n\n<p>Here is the paper: \n<a href=\"https://arxiv.org/abs/1705.07485\">https://arxiv.org/abs/1705.07485</a></p>\n\n<p>Here is my adaptation of the code  :\n<a href=\"https://www.kaggle.com/phoenix9032/shake-shake-regularization-starter-code\">https://www.kaggle.com/phoenix9032/shake-shake-regularization-starter-code</a></p>\n\n<p>I dont have any GPU left , so I have executed it only for 1 epoch . You can change the model_config and run it as much as you want . I have also switched the normal image augmentation off to save around 10 minutes per epoch . Anyone adopting it , can make it ON.</p>\n\n<p>Do let me know if it does any good .</p>",
      "rawMarkdown": "Why  should we do augmentations only for image where we can do augmentations for features as well to mix things up ! . Based on this idea only shake-shake regularization has been proposed . \nShake-Shake gave SOTA performance on Kuzushiji handwritten characters . Therefore the motivation for using it for Bengali . \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F612614a264786e69efbba4462984f015%2FSOTA.PNG?generation=1578001385829597&amp;alt=media)\n\n\nHere is the discussion where Heng Mentioned it  : \nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123757\n\nHere is the paper: \nhttps://arxiv.org/abs/1705.07485\n\nHere is my adaptation of the code  :\nhttps://www.kaggle.com/phoenix9032/shake-shake-regularization-starter-code\n\nI dont have any GPU left , so I have executed it only for 1 epoch . You can change the model_config and run it as much as you want . I have also switched the normal image augmentation off to save around 10 minutes per epoch . Anyone adopting it , can make it ON.\n\nDo let me know if it does any good .",
      "votes": null
    },
    {
      "id": "709410",
      "postDate": "01/03/2020 13:34:22",
      "content": "<p>I have already tried the shake shake regularization from this following pytorch implementation:\n<a href=\"https://github.com/hysts/pytorch_shake_shake\">https://github.com/hysts/pytorch_shake_shake</a>\nhowever, from my experience, the convergence rate was way too slow compared to the other sota architectures (like effnet, resnet, densenet etc.) I ran for 25 epochs only, the validation recall was around 78% so I dropped this approach. however I strongly believe that I must have missed something important that has caused such poor performance. Looking forward to seeing the results of your experiments! </p>",
      "rawMarkdown": "I have already tried the shake shake regularization from this following pytorch implementation:\nhttps://github.com/hysts/pytorch_shake_shake\nhowever, from my experience, the convergence rate was way too slow compared to the other sota architectures (like effnet, resnet, densenet etc.) I ran for 25 epochs only, the validation recall was around 78% so I dropped this approach. however I strongly believe that I must have missed something important that has caused such poor performance. Looking forward to seeing the results of your experiments!",
      "votes": null
    },
    {
      "id": "709489",
      "postDate": "01/03/2020 15:21:59",
      "content": "<p>you have to train for upto 150- 200 epochs according to repo. Which make sense because we are training CNN from scratch =) Good luck!</p>",
      "rawMarkdown": "you have to train for upto 150- 200 epochs according to repo. Which make sense because we are training CNN from scratch =) Good luck!",
      "votes": null
    },
    {
      "id": "709501",
      "postDate": "01/03/2020 15:33:19",
      "content": "<p>Yes , if we see the results , it looks like the loss flatlines for long time and then at the end suddenly decreases . Normally  good old SGD helps in long training ..I think!</p>",
      "rawMarkdown": "Yes , if we see the results , it looks like the loss flatlines for long time and then at the end suddenly decreases . Normally  good old SGD helps in long training ..I think!",
      "votes": null
    },
    {
      "id": "709524",
      "postDate": "01/03/2020 16:04:43",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> Yes, it seems that might be the root cause for the poor performance in my case. But for other deeper and bigger nets, especially SE-Resnext50 or Resnet50, training from scratch with the exact same setup and data, required only 10-15 epochs to converge! That is why I kind of felt disheartened with the shake-shake. I think I need to be more patient from now on! :( </p>",
      "rawMarkdown": "drhabib Yes, it seems that might be the root cause for the poor performance in my case. But for other deeper and bigger nets, especially SE-Resnext50 or Resnet50, training from scratch with the exact same setup and data, required only 10-15 epochs to converge! That is why I kind of felt disheartened with the shake-shake. I think I need to be more patient from now on! :(",
      "votes": null
    },
    {
      "id": "709563",
      "postDate": "01/03/2020 17:06:28",
      "content": "<p>Thanks. I tried the model with augmentation. About 35 mins per epoch. Validation recall .90 + after about 7 epochs.</p>",
      "rawMarkdown": "Thanks. I tried the model with augmentation. About 35 mins per epoch. Validation recall .90 + after about 7 epochs.",
      "votes": null
    },
    {
      "id": "724650",
      "postDate": "01/21/2020 10:45:27",
      "content": "<p>Thank you very much sir been looking for this :)</p>",
      "rawMarkdown": "Thank you very much sir been looking for this :)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 709410,
      "author_name": "udaykamal",
      "author_url": "",
      "post_date": "01/03/2020 13:34:22",
      "content": "<p>I have already tried the shake shake regularization from this following pytorch implementation:\n<a href=\"https://github.com/hysts/pytorch_shake_shake\">https://github.com/hysts/pytorch_shake_shake</a>\nhowever, from my experience, the convergence rate was way too slow compared to the other sota architectures (like effnet, resnet, densenet etc.) I ran for 25 epochs only, the validation recall was around 78% so I dropped this approach. however I strongly believe that I must have missed something important that has caused such poor performance. Looking forward to seeing the results of your experiments! </p>",
      "votes": null,
      "replies": [
        {
          "id": 709489,
          "author_name": "drhabib",
          "author_url": "",
          "post_date": "01/03/2020 15:21:59",
          "content": "<p>you have to train for upto 150- 200 epochs according to repo. Which make sense because we are training CNN from scratch =) Good luck!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709501,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "01/03/2020 15:33:19",
          "content": "<p>Yes , if we see the results , it looks like the loss flatlines for long time and then at the end suddenly decreases . Normally  good old SGD helps in long training ..I think!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 709524,
          "author_name": "udaykamal",
          "author_url": "",
          "post_date": "01/03/2020 16:04:43",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> Yes, it seems that might be the root cause for the poor performance in my case. But for other deeper and bigger nets, especially SE-Resnext50 or Resnet50, training from scratch with the exact same setup and data, required only 10-15 epochs to converge! That is why I kind of felt disheartened with the shake-shake. I think I need to be more patient from now on! :( </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 709563,
      "author_name": "shayekh",
      "author_url": "",
      "post_date": "01/03/2020 17:06:28",
      "content": "<p>Thanks. I tried the model with augmentation. About 35 mins per epoch. Validation recall .90 + after about 7 epochs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 724650,
      "author_name": "roguekk007",
      "author_url": "",
      "post_date": "01/21/2020 10:45:27",
      "content": "<p>Thank you very much sir been looking for this :)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "708953": "Why  should we do augmentations only for image where we can do augmentations for features as well to mix things up ! . Based on this idea only shake-shake regularization has been proposed . \nShake-Shake gave SOTA performance on Kuzushiji handwritten characters . Therefore the motivation for using it for Bengali . \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2234817%2F612614a264786e69efbba4462984f015%2FSOTA.PNG?generation=1578001385829597&amp;alt=media)\n\n\nHere is the discussion where Heng Mentioned it  : \nhttps://www.kaggle.com/c/bengaliai-cv19/discussion/123757\n\nHere is the paper: \nhttps://arxiv.org/abs/1705.07485\n\nHere is my adaptation of the code  :\nhttps://www.kaggle.com/phoenix9032/shake-shake-regularization-starter-code\n\nI dont have any GPU left , so I have executed it only for 1 epoch . You can change the model_config and run it as much as you want . I have also switched the normal image augmentation off to save around 10 minutes per epoch . Anyone adopting it , can make it ON.\n\nDo let me know if it does any good .",
    "709410": "I have already tried the shake shake regularization from this following pytorch implementation:\nhttps://github.com/hysts/pytorch_shake_shake\nhowever, from my experience, the convergence rate was way too slow compared to the other sota architectures (like effnet, resnet, densenet etc.) I ran for 25 epochs only, the validation recall was around 78% so I dropped this approach. however I strongly believe that I must have missed something important that has caused such poor performance. Looking forward to seeing the results of your experiments!",
    "709489": "you have to train for upto 150- 200 epochs according to repo. Which make sense because we are training CNN from scratch =) Good luck!",
    "709501": "Yes , if we see the results , it looks like the loss flatlines for long time and then at the end suddenly decreases . Normally  good old SGD helps in long training ..I think!",
    "709524": "drhabib Yes, it seems that might be the root cause for the poor performance in my case. But for other deeper and bigger nets, especially SE-Resnext50 or Resnet50, training from scratch with the exact same setup and data, required only 10-15 epochs to converge! That is why I kind of felt disheartened with the shake-shake. I think I need to be more patient from now on! :(",
    "709563": "Thanks. I tried the model with augmentation. About 35 mins per epoch. Validation recall .90 + after about 7 epochs.",
    "724650": "Thank you very much sir been looking for this :)"
  },
  "source": "meta"
}