{
  "id": 173718,
  "title": "Training in parts vs together!",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/173718",
  "author_name": "",
  "post_date": "2020-08-10T11:50:29.681099100Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hey, first of all, Good luck for the competition.</p>\n<p>What I am experiencing is a <code>Resource Exhausted Error</code> when training the Triple Stratified notebook provided by Chris when trained on Image of size 768 and model b6 or b7, on other image sizes and models, it works just fine. I get this error only after first fold, first fold runs nicely, after that, this error is produced.</p>\n<p>Now, I have tried training it in parts, i.e. 3 folds in 3 sessions and then a new session for prediction, while doing it, I noticed that my score is very different in both the cases, for example, with image size 512, <code>[0.945, b6]</code> in case of together while <code>[0.939, b6]</code> in case of parts.</p>\n<p>I have two questions.</p>\n<ul>\n<li><p>Is there any other way to train with this setup without getting a Resource Exhausted Error?</p></li>\n<li><p>What might be causing this much fluctuation in the score?</p></li>\n</ul>",
  "messages": [
    {
      "id": "965115",
      "postDate": "08/10/2020 11:50:29",
      "content": "<p>Hey, first of all, Good luck for the competition.</p>\n<p>What I am experiencing is a <code>Resource Exhausted Error</code> when training the Triple Stratified notebook provided by Chris when trained on Image of size 768 and model b6 or b7, on other image sizes and models, it works just fine. I get this error only after first fold, first fold runs nicely, after that, this error is produced.</p>\n<p>Now, I have tried training it in parts, i.e. 3 folds in 3 sessions and then a new session for prediction, while doing it, I noticed that my score is very different in both the cases, for example, with image size 512, <code>[0.945, b6]</code> in case of together while <code>[0.939, b6]</code> in case of parts.</p>\n<p>I have two questions.</p>\n<ul>\n<li><p>Is there any other way to train with this setup without getting a Resource Exhausted Error?</p></li>\n<li><p>What might be causing this much fluctuation in the score?</p></li>\n</ul>",
      "rawMarkdown": "Hey, first of all, Good luck for the competition.\n\nWhat I am experiencing is a `Resource Exhausted Error` when training the Triple Stratified notebook provided by Chris when trained on Image of size 768 and model b6 or b7, on other image sizes and models, it works just fine. I get this error only after first fold, first fold runs nicely, after that, this error is produced.\n\nNow, I have tried training it in parts, i.e. 3 folds in 3 sessions and then a new session for prediction, while doing it, I noticed that my score is very different in both the cases, for example, with image size 512, `[0.945, b6]` in case of together while `[0.939, b6]` in case of parts.\n\nI have two questions.\n\n- Is there any other way to train with this setup without getting a Resource Exhausted Error?\n\n- What might be causing this much fluctuation in the score?",
      "votes": null
    },
    {
      "id": "965124",
      "postDate": "08/10/2020 11:59:59",
      "content": "<p>My guess is you are using a smaller batch size for a larger input resolution which can certain affect how the model learns. Larger batch size can certainly improve convergence score as the model can learn faster with a larger learning rate. Interesting <a href=\"https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e\">article on batch size</a></p>",
      "rawMarkdown": "My guess is you are using a smaller batch size for a larger input resolution which can certain affect how the model learns. Larger batch size can certainly improve convergence score as the model can learn faster with a larger learning rate. Interesting [article on batch size](https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e)",
      "votes": null
    },
    {
      "id": "965131",
      "postDate": "08/10/2020 12:03:40",
      "content": "<p>I'm keeping everything same, batch size is same in both setups, even then it is giving this much difference is score. </p>\n<p>IMO, 3 folds are producing 3 different models and then prediction is done, in both the cases, I don't what is making the score fall down this much.</p>\n<p>I have high hopes for 768, but when tried doing it in parts, it is giving me only 0.93, and then I found this absurd thing happening.</p>\n<p>P.S.: I'm sorry if I misunderstood, you are talking about the error I guess? If yes, I will try increasing the batch size then, thanks anyway! :)</p>\n<hr>\n<p>Update: I tried higher batch size, it gave me error even before computing a single epoch. :/</p>",
      "rawMarkdown": "I'm keeping everything same, batch size is same in both setups, even then it is giving this much difference is score. \n\nIMO, 3 folds are producing 3 different models and then prediction is done, in both the cases, I don't what is making the score fall down this much.\n\nI have high hopes for 768, but when tried doing it in parts, it is giving me only 0.93, and then I found this absurd thing happening.\n\nP.S.: I'm sorry if I misunderstood, you are talking about the error I guess? If yes, I will try increasing the batch size then, thanks anyway! :)\n\n***********************************\n\nUpdate: I tried higher batch size, it gave me error even before computing a single epoch. :/",
      "votes": null
    },
    {
      "id": "968679",
      "postDate": "08/13/2020 07:17:13",
      "content": "<p>You are doing something wrong.  It's hard to say what without seeing your code.  I assume you are re-using the same data structures each fold?  </p>\n<p>I'll take a stab at what you may be doing.  Are you instantiating a completely new model each fold?  And then leaving the old one around?  You should just be re-using the same variables so it would be over written.  If you can share the code we can probably see right away whats wrong.</p>",
      "rawMarkdown": "You are doing something wrong.  It's hard to say what without seeing your code.  I assume you are re-using the same data structures each fold?  \n\nI'll take a stab at what you may be doing.  Are you instantiating a completely new model each fold?  And then leaving the old one around?  You should just be re-using the same variables so it would be over written.  If you can share the code we can probably see right away whats wrong.",
      "votes": null
    },
    {
      "id": "968840",
      "postDate": "08/13/2020 09:15:49",
      "content": "<p>Although I am using Chris's notebook and I ignored this thought when it did came to my mind, I will surely check it again. Thanks!</p>",
      "rawMarkdown": "Although I am using Chris's notebook and I ignored this thought when it did came to my mind, I will surely check it again. Thanks!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 965124,
      "author_name": "teeyee314",
      "author_url": "",
      "post_date": "08/10/2020 11:59:59",
      "content": "<p>My guess is you are using a smaller batch size for a larger input resolution which can certain affect how the model learns. Larger batch size can certainly improve convergence score as the model can learn faster with a larger learning rate. Interesting <a href=\"https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e\">article on batch size</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 965131,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/10/2020 12:03:40",
          "content": "<p>I'm keeping everything same, batch size is same in both setups, even then it is giving this much difference is score. </p>\n<p>IMO, 3 folds are producing 3 different models and then prediction is done, in both the cases, I don't what is making the score fall down this much.</p>\n<p>I have high hopes for 768, but when tried doing it in parts, it is giving me only 0.93, and then I found this absurd thing happening.</p>\n<p>P.S.: I'm sorry if I misunderstood, you are talking about the error I guess? If yes, I will try increasing the batch size then, thanks anyway! :)</p>\n<hr>\n<p>Update: I tried higher batch size, it gave me error even before computing a single epoch. :/</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 968679,
          "author_name": "brianfeeny",
          "author_url": "",
          "post_date": "08/13/2020 07:17:13",
          "content": "<p>You are doing something wrong.  It's hard to say what without seeing your code.  I assume you are re-using the same data structures each fold?  </p>\n<p>I'll take a stab at what you may be doing.  Are you instantiating a completely new model each fold?  And then leaving the old one around?  You should just be re-using the same variables so it would be over written.  If you can share the code we can probably see right away whats wrong.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 968840,
          "author_name": "sarques",
          "author_url": "",
          "post_date": "08/13/2020 09:15:49",
          "content": "<p>Although I am using Chris's notebook and I ignored this thought when it did came to my mind, I will surely check it again. Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "965115": "Hey, first of all, Good luck for the competition.\n\nWhat I am experiencing is a `Resource Exhausted Error` when training the Triple Stratified notebook provided by Chris when trained on Image of size 768 and model b6 or b7, on other image sizes and models, it works just fine. I get this error only after first fold, first fold runs nicely, after that, this error is produced.\n\nNow, I have tried training it in parts, i.e. 3 folds in 3 sessions and then a new session for prediction, while doing it, I noticed that my score is very different in both the cases, for example, with image size 512, `[0.945, b6]` in case of together while `[0.939, b6]` in case of parts.\n\nI have two questions.\n\n- Is there any other way to train with this setup without getting a Resource Exhausted Error?\n\n- What might be causing this much fluctuation in the score?",
    "965124": "My guess is you are using a smaller batch size for a larger input resolution which can certain affect how the model learns. Larger batch size can certainly improve convergence score as the model can learn faster with a larger learning rate. Interesting [article on batch size](https://medium.com/mini-distill/effect-of-batch-size-on-training-dynamics-21c14f7a716e)",
    "965131": "I'm keeping everything same, batch size is same in both setups, even then it is giving this much difference is score. \n\nIMO, 3 folds are producing 3 different models and then prediction is done, in both the cases, I don't what is making the score fall down this much.\n\nI have high hopes for 768, but when tried doing it in parts, it is giving me only 0.93, and then I found this absurd thing happening.\n\nP.S.: I'm sorry if I misunderstood, you are talking about the error I guess? If yes, I will try increasing the batch size then, thanks anyway! :)\n\n***********************************\n\nUpdate: I tried higher batch size, it gave me error even before computing a single epoch. :/",
    "968679": "You are doing something wrong.  It's hard to say what without seeing your code.  I assume you are re-using the same data structures each fold?  \n\nI'll take a stab at what you may be doing.  Are you instantiating a completely new model each fold?  And then leaving the old one around?  You should just be re-using the same variables so it would be over written.  If you can share the code we can probably see right away whats wrong.",
    "968840": "Although I am using Chris's notebook and I ignored this thought when it did came to my mind, I will surely check it again. Thanks!"
  },
  "source": "meta"
}