{
  "id": 117325,
  "title": "Predicting with different epochs weights might give better outcome!",
  "url": "/competitions/understanding_cloud_organization/discussion/117325",
  "author_name": "",
  "post_date": "2019-11-14T17:19:29.869137400Z",
  "votes": 10,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Generally, we tend to use the best weights from each fold epochs for generating predictions. I observed that instead of going for best or near best it might be a good idea to try one from earlier epochs. For example, use one with CV 0.61 from epoch 25 and one with 0.60 from epoch 15 performing better compared to two weights with 0.610 and 0.605 from epochs 25 and 27. </p>\n\n<p>Also, you can try more weight from the same fold. Like instead of 2 we can use 4.  This generalizing the public lb better. </p>\n\n<p>Any different thoughts or suggestions will be highly appreciated. </p>",
  "messages": [
    {
      "id": "673218",
      "postDate": "11/14/2019 17:19:29",
      "content": "<p>Generally, we tend to use the best weights from each fold epochs for generating predictions. I observed that instead of going for best or near best it might be a good idea to try one from earlier epochs. For example, use one with CV 0.61 from epoch 25 and one with 0.60 from epoch 15 performing better compared to two weights with 0.610 and 0.605 from epochs 25 and 27. </p>\n\n<p>Also, you can try more weight from the same fold. Like instead of 2 we can use 4.  This generalizing the public lb better. </p>\n\n<p>Any different thoughts or suggestions will be highly appreciated. </p>",
      "rawMarkdown": "Generally, we tend to use the best weights from each fold epochs for generating predictions. I observed that instead of going for best or near best it might be a good idea to try one from earlier epochs. For example, use one with CV 0.61 from epoch 25 and one with 0.60 from epoch 15 performing better compared to two weights with 0.610 and 0.605 from epochs 25 and 27. \n\nAlso, you can try more weight from the same fold. Like instead of 2 we can use 4.  This generalizing the public lb better. \n\nAny different thoughts or suggestions will be highly appreciated.",
      "votes": null
    },
    {
      "id": "673497",
      "postDate": "11/15/2019 04:13:48",
      "content": "<p>Nice suggestions. There is a lot of literature on these ideas. Ideally (when training an NN) we would like to find the global minimum of the loss function but instead we find a local minimum. So you can bounce around your training and find many local minimums and then ensemble them all together. There's a nice blog about this <a href=\"https://towardsdatascience.com/https-medium-com-reina-wang-tw-stochastic-gradient-descent-with-restarts-5f511975163\">here</a>. There are also nice example Kaggle notebooks from past competitions showing this.</p>",
      "rawMarkdown": "Nice suggestions. There is a lot of literature on these ideas. Ideally (when training an NN) we would like to find the global minimum of the loss function but instead we find a local minimum. So you can bounce around your training and find many local minimums and then ensemble them all together. There's a nice blog about this [here][1]. There are also nice example Kaggle notebooks from past competitions showing this.\n\n[1]: https://towardsdatascience.com/https-medium-com-reina-wang-tw-stochastic-gradient-descent-with-restarts-5f511975163",
      "votes": null
    },
    {
      "id": "673551",
      "postDate": "11/15/2019 06:22:33",
      "content": "<p>I tried your method but unfortunately my result get worse. sad :(</p>",
      "rawMarkdown": "I tried your method but unfortunately my result get worse. sad :(",
      "votes": null
    },
    {
      "id": "673573",
      "postDate": "11/15/2019 07:17:31",
      "content": "<p>I tried several times and almost every time I got better results. Try to get the weights with less loss too. A balance between less loss and kaggle dice would be better in my opinion. </p>",
      "rawMarkdown": "I tried several times and almost every time I got better results. Try to get the weights with less loss too. A balance between less loss and kaggle dice would be better in my opinion.",
      "votes": null
    },
    {
      "id": "673645",
      "postDate": "11/15/2019 09:10:41",
      "content": "<p>i got you. Thanks</p>",
      "rawMarkdown": "i got you. Thanks",
      "votes": null
    },
    {
      "id": "674001",
      "postDate": "11/15/2019 19:23:46",
      "content": "<p>Thanks for the comment. I am going to look at the blog. </p>",
      "rawMarkdown": "Thanks for the comment. I am going to look at the blog.",
      "votes": null
    },
    {
      "id": "675864",
      "postDate": "11/18/2019 17:16:26",
      "content": "<p>Is this similar to snapshot ensemble ?  Like when we use a varying learning rate cycle at the end of one cycle we save weight and then ensemble all such weights ?  I have never tried it . Probably I will try once with the late submission . For now submission limit is over . </p>",
      "rawMarkdown": "Is this similar to snapshot ensemble ?  Like when we use a varying learning rate cycle at the end of one cycle we save weight and then ensemble all such weights ?  I have never tried it . Probably I will try once with the late submission . For now submission limit is over .",
      "votes": null
    },
    {
      "id": "675875",
      "postDate": "11/18/2019 17:46:32",
      "content": "<p>Yes, that is what I'm referring to. Here's another <a href=\"https://machinelearningmastery.com/snapshot-ensemble-deep-learning-neural-network/\">blog</a> about snapshot ensemble.</p>",
      "rawMarkdown": "Yes, that is what I'm referring to. Here's another [blog][1] about snapshot ensemble.\n\n[1]: https://machinelearningmastery.com/snapshot-ensemble-deep-learning-neural-network/",
      "votes": null
    },
    {
      "id": "676001",
      "postDate": "11/18/2019 22:09:29",
      "content": "<p>apart from snapshot ensembling, <a href=\"https://arxiv.org/abs/1803.05407\">SWA</a> (Stocastic  Weights Averaging) or <a href=\"https://arxiv.org/abs/1703.01780\">EMA</a> (Exponential Moving Average) of models also perform better sometimes. However, adjusting the learning rate, choice of the optimizer,  and choice of epochs to perform SWA or EMA makes it a more time-consuming process of optimization. Snapshot/SWA/EMA performed almost similarly for me, except for the fact that snapshot is much easier/less time consuming to do.</p>",
      "rawMarkdown": "apart from snapshot ensembling, [SWA](https://arxiv.org/abs/1803.05407) (Stocastic  Weights Averaging) or [EMA](https://arxiv.org/abs/1703.01780) (Exponential Moving Average) of models also perform better sometimes. However, adjusting the learning rate, choice of the optimizer,  and choice of epochs to perform SWA or EMA makes it a more time-consuming process of optimization. Snapshot/SWA/EMA performed almost similarly for me, except for the fact that snapshot is much easier/less time consuming to do.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 673497,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/15/2019 04:13:48",
      "content": "<p>Nice suggestions. There is a lot of literature on these ideas. Ideally (when training an NN) we would like to find the global minimum of the loss function but instead we find a local minimum. So you can bounce around your training and find many local minimums and then ensemble them all together. There's a nice blog about this <a href=\"https://towardsdatascience.com/https-medium-com-reina-wang-tw-stochastic-gradient-descent-with-restarts-5f511975163\">here</a>. There are also nice example Kaggle notebooks from past competitions showing this.</p>",
      "votes": null,
      "replies": [
        {
          "id": 674001,
          "author_name": "mykttu",
          "author_url": "",
          "post_date": "11/15/2019 19:23:46",
          "content": "<p>Thanks for the comment. I am going to look at the blog. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675864,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "11/18/2019 17:16:26",
          "content": "<p>Is this similar to snapshot ensemble ?  Like when we use a varying learning rate cycle at the end of one cycle we save weight and then ensemble all such weights ?  I have never tried it . Probably I will try once with the late submission . For now submission limit is over . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 675875,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "11/18/2019 17:46:32",
          "content": "<p>Yes, that is what I'm referring to. Here's another <a href=\"https://machinelearningmastery.com/snapshot-ensemble-deep-learning-neural-network/\">blog</a> about snapshot ensemble.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 676001,
          "author_name": "sgalib",
          "author_url": "",
          "post_date": "11/18/2019 22:09:29",
          "content": "<p>apart from snapshot ensembling, <a href=\"https://arxiv.org/abs/1803.05407\">SWA</a> (Stocastic  Weights Averaging) or <a href=\"https://arxiv.org/abs/1703.01780\">EMA</a> (Exponential Moving Average) of models also perform better sometimes. However, adjusting the learning rate, choice of the optimizer,  and choice of epochs to perform SWA or EMA makes it a more time-consuming process of optimization. Snapshot/SWA/EMA performed almost similarly for me, except for the fact that snapshot is much easier/less time consuming to do.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 673551,
      "author_name": "jiahecao",
      "author_url": "",
      "post_date": "11/15/2019 06:22:33",
      "content": "<p>I tried your method but unfortunately my result get worse. sad :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 673573,
          "author_name": "mykttu",
          "author_url": "",
          "post_date": "11/15/2019 07:17:31",
          "content": "<p>I tried several times and almost every time I got better results. Try to get the weights with less loss too. A balance between less loss and kaggle dice would be better in my opinion. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 673645,
          "author_name": "jiahecao",
          "author_url": "",
          "post_date": "11/15/2019 09:10:41",
          "content": "<p>i got you. Thanks</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "673218": "Generally, we tend to use the best weights from each fold epochs for generating predictions. I observed that instead of going for best or near best it might be a good idea to try one from earlier epochs. For example, use one with CV 0.61 from epoch 25 and one with 0.60 from epoch 15 performing better compared to two weights with 0.610 and 0.605 from epochs 25 and 27. \n\nAlso, you can try more weight from the same fold. Like instead of 2 we can use 4.  This generalizing the public lb better. \n\nAny different thoughts or suggestions will be highly appreciated.",
    "673497": "Nice suggestions. There is a lot of literature on these ideas. Ideally (when training an NN) we would like to find the global minimum of the loss function but instead we find a local minimum. So you can bounce around your training and find many local minimums and then ensemble them all together. There's a nice blog about this [here][1]. There are also nice example Kaggle notebooks from past competitions showing this.\n\n[1]: https://towardsdatascience.com/https-medium-com-reina-wang-tw-stochastic-gradient-descent-with-restarts-5f511975163",
    "673551": "I tried your method but unfortunately my result get worse. sad :(",
    "673573": "I tried several times and almost every time I got better results. Try to get the weights with less loss too. A balance between less loss and kaggle dice would be better in my opinion.",
    "673645": "i got you. Thanks",
    "674001": "Thanks for the comment. I am going to look at the blog.",
    "675864": "Is this similar to snapshot ensemble ?  Like when we use a varying learning rate cycle at the end of one cycle we save weight and then ensemble all such weights ?  I have never tried it . Probably I will try once with the late submission . For now submission limit is over .",
    "675875": "Yes, that is what I'm referring to. Here's another [blog][1] about snapshot ensemble.\n\n[1]: https://machinelearningmastery.com/snapshot-ensemble-deep-learning-neural-network/",
    "676001": "apart from snapshot ensembling, [SWA](https://arxiv.org/abs/1803.05407) (Stocastic  Weights Averaging) or [EMA](https://arxiv.org/abs/1703.01780) (Exponential Moving Average) of models also perform better sometimes. However, adjusting the learning rate, choice of the optimizer,  and choice of epochs to perform SWA or EMA makes it a more time-consuming process of optimization. Snapshot/SWA/EMA performed almost similarly for me, except for the fact that snapshot is much easier/less time consuming to do."
  },
  "source": "meta"
}