{
  "id": 222795,
  "title": "Any Problem with my Callbacks? Training on 60th Epoch, No Convergence Yet",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222795",
  "author_name": "samuel adeshina",
  "post_date": "2021-03-01T06:29:07.885000",
  "votes": 0,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I used ModelCheckpoint(monitor='val_auc', mode='max' ), EarlyStopping(mode='max') and ReduceLROnPlateau(monitor='val_auc', mode='max' ) as my callbacks. I also set epoch as 100. My questions are:</p>\n<p>(1) do these settings make sense ( I am thinking of setting mode = 'min' for EarlyStopping but I feel that it will not be in tune with the settings of the other callbacks)?</p>\n<p>(2) My training is now on the 60th epoch and the results do not show any sign of convergence yet despite using three callbacks. Is this a problem? Or should I increase the number of epochs (my guess is this will take more than 100 epochs for convergence to be reached)? </p>\n<p>Thanks</p>",
  "messages": [
    {
      "id": 1221773,
      "postDate": "2021-03-01T09:06:15.117Z",
      "content": "<p>How many epochs you should be training for is very much a function of your model (e.g. some models notoriously train more slowly), whether you are using transfer learning (with transfer learning you should need fewer epochs), what augmentations you use (some very heavy augmentation can need more epochs) and your learning rate schedule (e.g. too low a learning rate can slow down training a lot, some clever schedules like the one cycle policy can speed it up a lot etc.). So, there's not really any generally applicable answers here (but, 60 to 100 seems a lot to me).</p>\n<p>You'd definitely want to maximize your validation set ROC AuC.</p>",
      "rawMarkdown": "How many epochs you should be training for is very much a function of your model (e.g. some models notoriously train more slowly), whether you are using transfer learning (with transfer learning you should need fewer epochs), what augmentations you use (some very heavy augmentation can need more epochs) and your learning rate schedule (e.g. too low a learning rate can slow down training a lot, some clever schedules like the one cycle policy can speed it up a lot etc.). So, there's not really any generally applicable answers here (but, 60 to 100 seems a lot to me).\n\nYou'd definitely want to maximize your validation set ROC AuC.",
      "votes": 1,
      "replies": [
        {
          "id": 1222184,
          "postDate": "2021-03-01T16:01:40.023Z",
          "content": "<p>Thanks a lot. I think I need to increase the learning rate. <br>\nI don't understand what you mean by the '<strong>one cycle policy</strong>'. I will be grateful if you can shed more light on it.</p>\n<p>Thanks.</p>",
          "rawMarkdown": "Thanks a lot. I think I need to increase the learning rate. \nI don't understand what you mean by the '**one cycle policy**'. I will be grateful if you can shed more light on it.\n\nThanks."
        },
        {
          "id": 1222351,
          "postDate": "2021-03-01T18:04:14.850Z",
          "content": "<p>For the one-cycle policy, see e.g. the <a href=\"https://docs.fast.ai/callback.schedule.html#Learner.fit_one_cycle\" target=\"_blank\">fastai documentation</a> - or the book by Jeremy Howard and Sylvain Gugger, where <a href=\"https://github.com/fastai/fastbook/blob/master/05_pet_breeds.ipynb\" target=\"_blank\">Chapter 5</a> discusses this.</p>",
          "rawMarkdown": "For the one-cycle policy, see e.g. the [fastai documentation](https://docs.fast.ai/callback.schedule.html#Learner.fit_one_cycle) - or the book by Jeremy Howard and Sylvain Gugger, where [Chapter 5](https://github.com/fastai/fastbook/blob/master/05_pet_breeds.ipynb) discusses this."
        },
        {
          "id": 1222460,
          "postDate": "2021-03-01T20:09:28.433Z",
          "content": "<p>Thanks. Will do that. </p>\n<p>Just one more question if you don't mind: is patience = 3 okay? Or should I reduce it to 2? </p>\n<p>Thanks</p>",
          "rawMarkdown": "Thanks. Will do that. \n\nJust one more question if you don't mind: is patience = 3 okay? Or should I reduce it to 2? \n\nThanks"
        },
        {
          "id": 1222555,
          "postDate": "2021-03-01T22:29:31.907Z",
          "content": "<p>That really depends on how much the validation set score fluctuates randomly for you vs. your total number of epochs. There more noisy the signal, the more patience you want, the less noisy, the closer you can get to 2 or 3. I.e. you don't want to stop training too early just because of random noise, if you are still slowly improving, on the other hand if you only train 10 epochs, then waiting for 5 epochs is probably a bit pointless. If you really need to train for 50 to 100 epochs, then a patience of 3 seems lowish to me. But, that's just me guessing.</p>",
          "rawMarkdown": "That really depends on how much the validation set score fluctuates randomly for you vs. your total number of epochs. There more noisy the signal, the more patience you want, the less noisy, the closer you can get to 2 or 3. I.e. you don't want to stop training too early just because of random noise, if you are still slowly improving, on the other hand if you only train 10 epochs, then waiting for 5 epochs is probably a bit pointless. If you really need to train for 50 to 100 epochs, then a patience of 3 seems lowish to me. But, that's just me guessing."
        },
        {
          "id": 1224420,
          "postDate": "2021-03-02T17:52:53.250Z",
          "content": "<p>I get it. Thanks a lot.</p>",
          "rawMarkdown": "I get it. Thanks a lot."
        }
      ]
    },
    {
      "id": 1221769,
      "postDate": "2021-03-01T09:02:51.933Z",
      "content": "<p>Hello!</p>\n<p>There's a mistake in your <code>EarlyStopping</code> callback. As you do not provide any arguments except for the <code>mode='max'</code>, according to <strong><a href=\"https://keras.io/api/callbacks/early_stopping/\" target=\"_blank\">keras docs</a></strong> the <code>monitor</code> argument sets by default to <code>val_loss</code>. So what you are literally doing is training until reaching the highest <code>val_loss</code>, which I think, is not what you want :)</p>\n<p>To fix that, consider one of the following options of the <code>EarlyStopping</code> instance:</p>\n<ul>\n<li><code>EarlyStopping(monitor='val_auc', mode='max')</code></li>\n<li><code>EarlyStopping()</code> (which is equal to <code>EarlyStopping(monitor='val_loss', mode='min')</code>)</li>\n</ul>",
      "rawMarkdown": "Hello!\n\nThere's a mistake in your `EarlyStopping` callback. As you do not provide any arguments except for the `mode='max'`, according to **[keras docs](https://keras.io/api/callbacks/early_stopping/)** the `monitor` argument sets by default to `val_loss`. So what you are literally doing is training until reaching the highest `val_loss`, which I think, is not what you want :)\n\nTo fix that, consider one of the following options of the `EarlyStopping` instance:\n* `EarlyStopping(monitor='val_auc', mode='max')`\n* `EarlyStopping()` (which is equal to `EarlyStopping(monitor='val_loss', mode='min')`)",
      "votes": 2,
      "replies": [
        {
          "id": 1222180,
          "postDate": "2021-03-01T15:58:15.153Z",
          "content": "<p>I am sorry I did not provide that information. I actually used EarlyStopping(monitor='val_auc', mode='max')</p>\n<p>Thanks a lot</p>",
          "rawMarkdown": "I am sorry I did not provide that information. I actually used EarlyStopping(monitor='val_auc', mode='max')\n\nThanks a lot"
        }
      ]
    },
    {
      "id": 1221629,
      "postDate": "2021-03-01T06:29:07.887Z",
      "content": "<p>I used ModelCheckpoint(monitor='val_auc', mode='max' ), EarlyStopping(mode='max') and ReduceLROnPlateau(monitor='val_auc', mode='max' ) as my callbacks. I also set epoch as 100. My questions are:</p>\n<p>(1) do these settings make sense ( I am thinking of setting mode = 'min' for EarlyStopping but I feel that it will not be in tune with the settings of the other callbacks)?</p>\n<p>(2) My training is now on the 60th epoch and the results do not show any sign of convergence yet despite using three callbacks. Is this a problem? Or should I increase the number of epochs (my guess is this will take more than 100 epochs for convergence to be reached)? </p>\n<p>Thanks</p>",
      "rawMarkdown": "I used ModelCheckpoint(monitor='val_auc', mode='max' ), EarlyStopping(mode='max') and ReduceLROnPlateau(monitor='val_auc', mode='max' ) as my callbacks. I also set epoch as 100. My questions are:\n \n(1) do these settings make sense ( I am thinking of setting mode = 'min' for EarlyStopping but I feel that it will not be in tune with the settings of the other callbacks)?\n\n(2) My training is now on the 60th epoch and the results do not show any sign of convergence yet despite using three callbacks. Is this a problem? Or should I increase the number of epochs (my guess is this will take more than 100 epochs for convergence to be reached)? \n\nThanks"
    },
    {
      "id": 1222376,
      "postDate": "2021-03-01T18:23:18.750Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1221773,
      "author_name": "Björn",
      "author_url": "",
      "post_date": "2021-03-01T09:06:15.117000",
      "content": "<p>How many epochs you should be training for is very much a function of your model (e.g. some models notoriously train more slowly), whether you are using transfer learning (with transfer learning you should need fewer epochs), what augmentations you use (some very heavy augmentation can need more epochs) and your learning rate schedule (e.g. too low a learning rate can slow down training a lot, some clever schedules like the one cycle policy can speed it up a lot etc.). So, there's not really any generally applicable answers here (but, 60 to 100 seems a lot to me).</p>\n<p>You'd definitely want to maximize your validation set ROC AuC.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1222184,
          "author_name": "samuel adeshina",
          "author_url": "",
          "post_date": "2021-03-01T16:01:40.023000",
          "content": "<p>Thanks a lot. I think I need to increase the learning rate. <br>\nI don't understand what you mean by the '<strong>one cycle policy</strong>'. I will be grateful if you can shed more light on it.</p>\n<p>Thanks.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222351,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-03-01T18:04:14.850000",
          "content": "<p>For the one-cycle policy, see e.g. the <a href=\"https://docs.fast.ai/callback.schedule.html#Learner.fit_one_cycle\" target=\"_blank\">fastai documentation</a> - or the book by Jeremy Howard and Sylvain Gugger, where <a href=\"https://github.com/fastai/fastbook/blob/master/05_pet_breeds.ipynb\" target=\"_blank\">Chapter 5</a> discusses this.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222460,
          "author_name": "samuel adeshina",
          "author_url": "",
          "post_date": "2021-03-01T20:09:28.433000",
          "content": "<p>Thanks. Will do that. </p>\n<p>Just one more question if you don't mind: is patience = 3 okay? Or should I reduce it to 2? </p>\n<p>Thanks</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1222555,
          "author_name": "Björn",
          "author_url": "",
          "post_date": "2021-03-01T22:29:31.907000",
          "content": "<p>That really depends on how much the validation set score fluctuates randomly for you vs. your total number of epochs. There more noisy the signal, the more patience you want, the less noisy, the closer you can get to 2 or 3. I.e. you don't want to stop training too early just because of random noise, if you are still slowly improving, on the other hand if you only train 10 epochs, then waiting for 5 epochs is probably a bit pointless. If you really need to train for 50 to 100 epochs, then a patience of 3 seems lowish to me. But, that's just me guessing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1224420,
          "author_name": "samuel adeshina",
          "author_url": "",
          "post_date": "2021-03-02T17:52:53.250000",
          "content": "<p>I get it. Thanks a lot.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1221769,
      "author_name": "Nikita Kuzmenkov",
      "author_url": "",
      "post_date": "2021-03-01T09:02:51.933000",
      "content": "<p>Hello!</p>\n<p>There's a mistake in your <code>EarlyStopping</code> callback. As you do not provide any arguments except for the <code>mode='max'</code>, according to <strong><a href=\"https://keras.io/api/callbacks/early_stopping/\" target=\"_blank\">keras docs</a></strong> the <code>monitor</code> argument sets by default to <code>val_loss</code>. So what you are literally doing is training until reaching the highest <code>val_loss</code>, which I think, is not what you want :)</p>\n<p>To fix that, consider one of the following options of the <code>EarlyStopping</code> instance:</p>\n<ul>\n<li><code>EarlyStopping(monitor='val_auc', mode='max')</code></li>\n<li><code>EarlyStopping()</code> (which is equal to <code>EarlyStopping(monitor='val_loss', mode='min')</code>)</li>\n</ul>",
      "votes": 2,
      "replies": [
        {
          "id": 1222180,
          "author_name": "samuel adeshina",
          "author_url": "",
          "post_date": "2021-03-01T15:58:15.153000",
          "content": "<p>I am sorry I did not provide that information. I actually used EarlyStopping(monitor='val_auc', mode='max')</p>\n<p>Thanks a lot</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1222376,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-03-01T18:23:18.750000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1221773": "How many epochs you should be training for is very much a function of your model (e.g. some models notoriously train more slowly), whether you are using transfer learning (with transfer learning you should need fewer epochs), what augmentations you use (some very heavy augmentation can need more epochs) and your learning rate schedule (e.g. too low a learning rate can slow down training a lot, some clever schedules like the one cycle policy can speed it up a lot etc.). So, there's not really any generally applicable answers here (but, 60 to 100 seems a lot to me).\n\nYou'd definitely want to maximize your validation set ROC AuC.",
    "1221769": "Hello!\n\nThere's a mistake in your `EarlyStopping` callback. As you do not provide any arguments except for the `mode='max'`, according to **[keras docs](https://keras.io/api/callbacks/early_stopping/)** the `monitor` argument sets by default to `val_loss`. So what you are literally doing is training until reaching the highest `val_loss`, which I think, is not what you want :)\n\nTo fix that, consider one of the following options of the `EarlyStopping` instance:\n* `EarlyStopping(monitor='val_auc', mode='max')`\n* `EarlyStopping()` (which is equal to `EarlyStopping(monitor='val_loss', mode='min')`)",
    "1221629": "I used ModelCheckpoint(monitor='val_auc', mode='max' ), EarlyStopping(mode='max') and ReduceLROnPlateau(monitor='val_auc', mode='max' ) as my callbacks. I also set epoch as 100. My questions are:\n \n(1) do these settings make sense ( I am thinking of setting mode = 'min' for EarlyStopping but I feel that it will not be in tune with the settings of the other callbacks)?\n\n(2) My training is now on the 60th epoch and the results do not show any sign of convergence yet despite using three callbacks. Is this a problem? Or should I increase the number of epochs (my guess is this will take more than 100 epochs for convergence to be reached)? \n\nThanks",
    "1222376": ""
  }
}