{
  "id": 215874,
  "title": "Can we believe in early stopping? especially for this competition",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/215874",
  "author_name": "",
  "post_date": "2021-01-31T15:20:16.240860800Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>If i use early stopping with small valid loss , am getting less score than without early stopping(esp. in this competition)</p>\n<p>any suggestion?</p>",
  "messages": [
    {
      "id": "1179456",
      "postDate": "01/31/2021 15:20:16",
      "content": "<p>If i use early stopping with small valid loss , am getting less score than without early stopping(esp. in this competition)</p>\n<p>any suggestion?</p>",
      "rawMarkdown": "If i use early stopping with small valid loss , am getting less score than without early stopping(esp. in this competition)\n\nany suggestion?",
      "votes": null
    },
    {
      "id": "1179502",
      "postDate": "01/31/2021 15:41:50",
      "content": "<p>Hello! </p>\n<p>If by \"score\" you mean LB (and not CV) your \"not early stopped\" models are likely overfitting on private test data. </p>\n<p>I also suggest monitoring validation accuracy rather than validation loss (sure, they are highly correlated, but accuracy is the parameter of interest, not the loss). Furthermore, if the application of early stopping results in a poorer CV score, consider increasing <code>patience</code> and make sure that your best weights are actually being restored (e.g. if you are using Keras <code>EarlyStopping</code> callback, the best weights are restored only if the training process finishes before the last epoch, even if <code>restore_best_weights</code> argument is explicitly set to <code>True</code>).</p>",
      "rawMarkdown": "Hello! \n\nIf by \"score\" you mean LB (and not CV) your \"not early stopped\" models are likely overfitting on private test data. \n\nI also suggest monitoring validation accuracy rather than validation loss (sure, they are highly correlated, but accuracy is the parameter of interest, not the loss). Furthermore, if the application of early stopping results in a poorer CV score, consider increasing `patience` and make sure that your best weights are actually being restored (e.g. if you are using Keras `EarlyStopping` callback, the best weights are restored only if the training process finishes before the last epoch, even if `restore_best_weights` argument is explicitly set to `True`).",
      "votes": null
    },
    {
      "id": "1179524",
      "postDate": "01/31/2021 15:53:10",
      "content": "<p>Sounds great! <a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> Increasing the patience did not help me, monitoring accuracy would be d the right choice as u said</p>",
      "rawMarkdown": "Sounds great! @nickuzmenkov Increasing the patience did not help me, monitoring accuracy would be d the right choice as u said",
      "votes": null
    },
    {
      "id": "1179794",
      "postDate": "01/31/2021 20:34:42",
      "content": "<p>I'm actually a bit apprehensive about using validation accuracy for early stopping. The validation set with 5-fold CV (and 7- or 10-fold only gets worse) is pretty small and just pushing a tiny number of cases across the threshold affects accuracy a lot, but could be just chance fluctuation. The worry is that that may not generalize to the test set. Better CV accuracy resulting in worse public LB performance could be a hint in that direction, but to be fair, that's again a rather noisy signal. One idea to alleviate that concern could be to define your own metric that is more continuous, but aligns closely with accuracy (e.g. giving some credit for close misses, only giving full score if you are a decent but ahead of the next most likely class).</p>",
      "rawMarkdown": "I'm actually a bit apprehensive about using validation accuracy for early stopping. The validation set with 5-fold CV (and 7- or 10-fold only gets worse) is pretty small and just pushing a tiny number of cases across the threshold affects accuracy a lot, but could be just chance fluctuation. The worry is that that may not generalize to the test set. Better CV accuracy resulting in worse public LB performance could be a hint in that direction, but to be fair, that's again a rather noisy signal. One idea to alleviate that concern could be to define your own metric that is more continuous, but aligns closely with accuracy (e.g. giving some credit for close misses, only giving full score if you are a decent but ahead of the next most likely class).",
      "votes": null
    },
    {
      "id": "1180266",
      "postDate": "02/01/2021 07:37:48",
      "content": "<p>Thats great <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>, thanks🙌 </p>",
      "rawMarkdown": "Thats great @bjoernholzhauer, thanks🙌",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1179502,
      "author_name": "nickuzmenkov",
      "author_url": "",
      "post_date": "01/31/2021 15:41:50",
      "content": "<p>Hello! </p>\n<p>If by \"score\" you mean LB (and not CV) your \"not early stopped\" models are likely overfitting on private test data. </p>\n<p>I also suggest monitoring validation accuracy rather than validation loss (sure, they are highly correlated, but accuracy is the parameter of interest, not the loss). Furthermore, if the application of early stopping results in a poorer CV score, consider increasing <code>patience</code> and make sure that your best weights are actually being restored (e.g. if you are using Keras <code>EarlyStopping</code> callback, the best weights are restored only if the training process finishes before the last epoch, even if <code>restore_best_weights</code> argument is explicitly set to <code>True</code>).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1179524,
          "author_name": "santhoshkumarv",
          "author_url": "",
          "post_date": "01/31/2021 15:53:10",
          "content": "<p>Sounds great! <a href=\"https://www.kaggle.com/nickuzmenkov\" target=\"_blank\">@nickuzmenkov</a> Increasing the patience did not help me, monitoring accuracy would be d the right choice as u said</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1179794,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "01/31/2021 20:34:42",
      "content": "<p>I'm actually a bit apprehensive about using validation accuracy for early stopping. The validation set with 5-fold CV (and 7- or 10-fold only gets worse) is pretty small and just pushing a tiny number of cases across the threshold affects accuracy a lot, but could be just chance fluctuation. The worry is that that may not generalize to the test set. Better CV accuracy resulting in worse public LB performance could be a hint in that direction, but to be fair, that's again a rather noisy signal. One idea to alleviate that concern could be to define your own metric that is more continuous, but aligns closely with accuracy (e.g. giving some credit for close misses, only giving full score if you are a decent but ahead of the next most likely class).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1180266,
          "author_name": "santhoshkumarv",
          "author_url": "",
          "post_date": "02/01/2021 07:37:48",
          "content": "<p>Thats great <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a>, thanks🙌 </p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1179456": "If i use early stopping with small valid loss , am getting less score than without early stopping(esp. in this competition)\n\nany suggestion?",
    "1179502": "Hello! \n\nIf by \"score\" you mean LB (and not CV) your \"not early stopped\" models are likely overfitting on private test data. \n\nI also suggest monitoring validation accuracy rather than validation loss (sure, they are highly correlated, but accuracy is the parameter of interest, not the loss). Furthermore, if the application of early stopping results in a poorer CV score, consider increasing `patience` and make sure that your best weights are actually being restored (e.g. if you are using Keras `EarlyStopping` callback, the best weights are restored only if the training process finishes before the last epoch, even if `restore_best_weights` argument is explicitly set to `True`).",
    "1179524": "Sounds great! @nickuzmenkov Increasing the patience did not help me, monitoring accuracy would be d the right choice as u said",
    "1179794": "I'm actually a bit apprehensive about using validation accuracy for early stopping. The validation set with 5-fold CV (and 7- or 10-fold only gets worse) is pretty small and just pushing a tiny number of cases across the threshold affects accuracy a lot, but could be just chance fluctuation. The worry is that that may not generalize to the test set. Better CV accuracy resulting in worse public LB performance could be a hint in that direction, but to be fair, that's again a rather noisy signal. One idea to alleviate that concern could be to define your own metric that is more continuous, but aligns closely with accuracy (e.g. giving some credit for close misses, only giving full score if you are a decent but ahead of the next most likely class).",
    "1180266": "Thats great @bjoernholzhauer, thanks🙌"
  },
  "source": "meta"
}