{
  "id": 129064,
  "title": "Fast.ai training issue using multiple notebooks? (100 epochs)",
  "url": "/competitions/bengaliai-cv19/discussion/129064",
  "author_name": "",
  "post_date": "2020-02-05T06:49:51.750332600Z",
  "votes": null,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I am trying to train a model in fast.ai for 100 epochs , however since Kaggle only allows 9 hour runtimes , I have to train it in seperate notebooks for around ~35 epochs each. So we load the optimizer state , model weights , and start the training in a new notebook , however when the new training starts (36th epoch ) , the recall is reduced by about ~2% , and loss has increased. This makes it worthless training for 100 epochs. Even after trying different things , the issue still remains. Does anyone have an idea as to how to fix this?</p>\n\n<p>Training code snippets -</p>\n\n<p>First notebook (0-35 epochs):\n<code>learn = Learner(data, model, loss_func=Loss_combine(),opt_func=Over9000,\n        metrics=[Metric_grapheme(),Metric_vowel(),Metric_consonant(),Metric_tot()])\nlogger = CSVLogger(learn,f'log{fold}')\nlearn.clip_grad = 1.0\nlearn.split([model.head1])\nlearn.unfreeze()</code></p>\n\n<p><code>cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100)</code>\n<code>learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])</code></p>\n\n<p>Second notebook ( 36-70 epochs):</p>\n\n<p><code>cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100, jump_epochs=35)</code>\n<code>learn.model.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_model_34.pth'))</code>\n<code>learn.opt.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_optimizer_34.pth'))</code>\n<code>learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])</code></p>\n\n<p>Epoch callback saves the model weights and optimizer state at the 35th epoch.</p>\n\n<p>Last epoch of first notebook(total recall is 0.97)-</p>\n\n<p>33  0.989179    0.148242    0.958493    0.983685    0.980485    0.970289    </p>\n\n<p>First epoch of 2nd notebook(notice the total recall has fallen to 0.962)-</p>\n\n<p>0   1.128961    0.167398    0.949028    0.980022    0.970367    0.962112</p>",
  "messages": [
    {
      "id": "737302",
      "postDate": "02/05/2020 06:49:51",
      "content": "<p>I am trying to train a model in fast.ai for 100 epochs , however since Kaggle only allows 9 hour runtimes , I have to train it in seperate notebooks for around ~35 epochs each. So we load the optimizer state , model weights , and start the training in a new notebook , however when the new training starts (36th epoch ) , the recall is reduced by about ~2% , and loss has increased. This makes it worthless training for 100 epochs. Even after trying different things , the issue still remains. Does anyone have an idea as to how to fix this?</p>\n\n<p>Training code snippets -</p>\n\n<p>First notebook (0-35 epochs):\n<code>learn = Learner(data, model, loss_func=Loss_combine(),opt_func=Over9000,\n        metrics=[Metric_grapheme(),Metric_vowel(),Metric_consonant(),Metric_tot()])\nlogger = CSVLogger(learn,f'log{fold}')\nlearn.clip_grad = 1.0\nlearn.split([model.head1])\nlearn.unfreeze()</code></p>\n\n<p><code>cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100)</code>\n<code>learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])</code></p>\n\n<p>Second notebook ( 36-70 epochs):</p>\n\n<p><code>cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100, jump_epochs=35)</code>\n<code>learn.model.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_model_34.pth'))</code>\n<code>learn.opt.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_optimizer_34.pth'))</code>\n<code>learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])</code></p>\n\n<p>Epoch callback saves the model weights and optimizer state at the 35th epoch.</p>\n\n<p>Last epoch of first notebook(total recall is 0.97)-</p>\n\n<p>33  0.989179    0.148242    0.958493    0.983685    0.980485    0.970289    </p>\n\n<p>First epoch of 2nd notebook(notice the total recall has fallen to 0.962)-</p>\n\n<p>0   1.128961    0.167398    0.949028    0.980022    0.970367    0.962112</p>",
      "rawMarkdown": "I am trying to train a model in fast.ai for 100 epochs , however since Kaggle only allows 9 hour runtimes , I have to train it in seperate notebooks for around ~35 epochs each. So we load the optimizer state , model weights , and start the training in a new notebook , however when the new training starts (36th epoch ) , the recall is reduced by about ~2% , and loss has increased. This makes it worthless training for 100 epochs. Even after trying different things , the issue still remains. Does anyone have an idea as to how to fix this?\n\nTraining code snippets -\n\nFirst notebook (0-35 epochs):\n`learn = Learner(data, model, loss_func=Loss_combine(),opt_func=Over9000,\n        metrics=[Metric_grapheme(),Metric_vowel(),Metric_consonant(),Metric_tot()])\nlogger = CSVLogger(learn,f'log{fold}')\nlearn.clip_grad = 1.0\nlearn.split([model.head1])\nlearn.unfreeze()`\n\n`cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100)`\n`learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])`\n\n\nSecond notebook ( 36-70 epochs):\n\n`cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100, jump_epochs=35)`\n`learn.model.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_model_34.pth'))`\n`learn.opt.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_optimizer_34.pth'))`\n`learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])`\n\nEpoch callback saves the model weights and optimizer state at the 35th epoch.\n\n\nLast epoch of first notebook(total recall is 0.97)-\n\n33\t0.989179\t0.148242\t0.958493\t0.983685\t0.980485\t0.970289\t\n\n\nFirst epoch of 2nd notebook(notice the total recall has fallen to 0.962)-\n\n0\t1.128961\t0.167398\t0.949028\t0.980022\t0.970367\t0.962112",
      "votes": null
    },
    {
      "id": "737867",
      "postDate": "02/05/2020 21:07:20",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> can u pls help?</p>",
      "rawMarkdown": "drhabib can u pls help?",
      "votes": null
    },
    {
      "id": "741083",
      "postDate": "02/10/2020 06:46:52",
      "content": "<p>Did you still get improve after the 2nd notebook comparing with the 1st notebook?</p>",
      "rawMarkdown": "Did you still get improve after the 2nd notebook comparing with the 1st notebook?",
      "votes": null
    },
    {
      "id": "741186",
      "postDate": "02/10/2020 09:58:12",
      "content": "<p>Yes , but it was not worth it anyways. First notebook after 35 epochs was ~97.01 % CV , 2nd notebook started 36th epoch from 96.2% CV , that's very less,  and after 70 epochs it was 97.3% CV. If it properly worked , and started 36th epoch at 97.01 CV , we could have reached ~98% CV after 70-100 epochs. This issue still remains.</p>",
      "rawMarkdown": "Yes , but it was not worth it anyways. First notebook after 35 epochs was ~97.01 % CV , 2nd notebook started 36th epoch from 96.2% CV , that's very less,  and after 70 epochs it was 97.3% CV. If it properly worked , and started 36th epoch at 97.01 CV , we could have reached ~98% CV after 70-100 epochs. This issue still remains.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 737867,
      "author_name": "p4rallax",
      "author_url": "",
      "post_date": "02/05/2020 21:07:20",
      "content": "<p><a href=\"/drhabib\">@drhabib</a> can u pls help?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 741083,
      "author_name": "kasim0226",
      "author_url": "",
      "post_date": "02/10/2020 06:46:52",
      "content": "<p>Did you still get improve after the 2nd notebook comparing with the 1st notebook?</p>",
      "votes": null,
      "replies": [
        {
          "id": 741186,
          "author_name": "p4rallax",
          "author_url": "",
          "post_date": "02/10/2020 09:58:12",
          "content": "<p>Yes , but it was not worth it anyways. First notebook after 35 epochs was ~97.01 % CV , 2nd notebook started 36th epoch from 96.2% CV , that's very less,  and after 70 epochs it was 97.3% CV. If it properly worked , and started 36th epoch at 97.01 CV , we could have reached ~98% CV after 70-100 epochs. This issue still remains.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "737302": "I am trying to train a model in fast.ai for 100 epochs , however since Kaggle only allows 9 hour runtimes , I have to train it in seperate notebooks for around ~35 epochs each. So we load the optimizer state , model weights , and start the training in a new notebook , however when the new training starts (36th epoch ) , the recall is reduced by about ~2% , and loss has increased. This makes it worthless training for 100 epochs. Even after trying different things , the issue still remains. Does anyone have an idea as to how to fix this?\n\nTraining code snippets -\n\nFirst notebook (0-35 epochs):\n`learn = Learner(data, model, loss_func=Loss_combine(),opt_func=Over9000,\n        metrics=[Metric_grapheme(),Metric_vowel(),Metric_consonant(),Metric_tot()])\nlogger = CSVLogger(learn,f'log{fold}')\nlearn.clip_grad = 1.0\nlearn.split([model.head1])\nlearn.unfreeze()`\n\n`cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100)`\n`learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])`\n\n\nSecond notebook ( 36-70 epochs):\n\n`cb = OneCycleScheduler(learn, lr_max=max_lr, pct_start=0.0, div_factor=100, jump_epochs=35)`\n`learn.model.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_model_34.pth'))`\n`learn.opt.load_state_dict(torch.load('/kaggle/input/fast-ai-se-resnext50-mixup-cutmix/latest_optimizer_34.pth'))`\n`learn.fit(100, max_lr,wd=[1e-3,0.1e-1], callbacks=[logger, SaveModelCallback(learn,monitor='metric_tot',\n    mode='max',name=f'model_{fold}'),Choice(learn), cb, EpochCallback(learn)])`\n\nEpoch callback saves the model weights and optimizer state at the 35th epoch.\n\n\nLast epoch of first notebook(total recall is 0.97)-\n\n33\t0.989179\t0.148242\t0.958493\t0.983685\t0.980485\t0.970289\t\n\n\nFirst epoch of 2nd notebook(notice the total recall has fallen to 0.962)-\n\n0\t1.128961\t0.167398\t0.949028\t0.980022\t0.970367\t0.962112",
    "737867": "drhabib can u pls help?",
    "741083": "Did you still get improve after the 2nd notebook comparing with the 1st notebook?",
    "741186": "Yes , but it was not worth it anyways. First notebook after 35 epochs was ~97.01 % CV , 2nd notebook started 36th epoch from 96.2% CV , that's very less,  and after 70 epochs it was 97.3% CV. If it properly worked , and started 36th epoch at 97.01 CV , we could have reached ~98% CV after 70-100 epochs. This issue still remains."
  },
  "source": "meta"
}