{
  "id": 222845,
  "title": "Model checkpoints & early stopping: AUC vs valid loss",
  "url": "/competitions/ranzcr-clip-catheter-line-classification/discussion/222845",
  "author_name": "",
  "post_date": "2021-03-01T11:41:07.221566300Z",
  "votes": 7,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi guys, this has been bugging me for a while and I'm not sure what the answer is.</p>\n<p>What do you use to determine when to early stop &amp; checkpoint your models?</p>\n<ol>\n<li>Validation loss (minimise)</li>\n<li>AUC (maximise)</li>\n</ol>\n<p>Sometimes you can get a strange situation where valid loss increases (i.e. overfitting) but AUC is also increasing (noise maybe??). If you stop/checkpoint using valid loss, you could potentially be leaving a few AUC points on the table, but minimising the risk of overfitting and vice versa if you use AUC. What are your thoughts?</p>",
  "messages": [
    {
      "id": "1221897",
      "postDate": "03/01/2021 11:41:07",
      "content": "<p>Hi guys, this has been bugging me for a while and I'm not sure what the answer is.</p>\n<p>What do you use to determine when to early stop &amp; checkpoint your models?</p>\n<ol>\n<li>Validation loss (minimise)</li>\n<li>AUC (maximise)</li>\n</ol>\n<p>Sometimes you can get a strange situation where valid loss increases (i.e. overfitting) but AUC is also increasing (noise maybe??). If you stop/checkpoint using valid loss, you could potentially be leaving a few AUC points on the table, but minimising the risk of overfitting and vice versa if you use AUC. What are your thoughts?</p>",
      "rawMarkdown": "Hi guys, this has been bugging me for a while and I'm not sure what the answer is.\n\nWhat do you use to determine when to early stop & checkpoint your models?\n1. Validation loss (minimise)\n2. AUC (maximise)\n\nSometimes you can get a strange situation where valid loss increases (i.e. overfitting) but AUC is also increasing (noise maybe??). If you stop/checkpoint using valid loss, you could potentially be leaving a few AUC points on the table, but minimising the risk of overfitting and vice versa if you use AUC. What are your thoughts?",
      "votes": null
    },
    {
      "id": "1221916",
      "postDate": "03/01/2021 12:01:56",
      "content": "<p>If you've doubt, why not use both and average their predictions during inference ? </p>\n<p>In Pytorch it's easy to add in the training process, in TF you can create custom callback if you use <code>model.fit</code>  or add it like in Pytorch if you use custom training loop. </p>",
      "rawMarkdown": "If you've doubt, why not use both and average their predictions during inference ? \n\nIn Pytorch it's easy to add in the training process, in TF you can create custom callback if you use `model.fit`  or add it like in Pytorch if you use custom training loop.",
      "votes": null
    },
    {
      "id": "1222001",
      "postDate": "03/01/2021 13:12:37",
      "content": "<p>or use SWA and avoid minima from a single epoch!?</p>",
      "rawMarkdown": "or use SWA and avoid minima from a single epoch!?",
      "votes": null
    },
    {
      "id": "1222023",
      "postDate": "03/01/2021 13:33:29",
      "content": "<p>I asked a similar question before, I chose val loss. Then according to my observation, if val loss and val auc rise at the same time, the model is usually not very good and the LB score is not high.</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145</a></p>",
      "rawMarkdown": "I asked a similar question before, I chose val loss. Then according to my observation, if val loss and val auc rise at the same time, the model is usually not very good and the LB score is not high.\n\n[https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145)",
      "votes": null
    },
    {
      "id": "1222048",
      "postDate": "03/01/2021 14:01:50",
      "content": "<p>SWA/EMA is  good alternative</p>",
      "rawMarkdown": "SWA/EMA is  good alternative",
      "votes": null
    },
    {
      "id": "1222537",
      "postDate": "03/01/2021 21:50:48",
      "content": "<p>It is a bit of a difficult question because auc is rather noisy in comparison. ETT - Abnormal only has a handful of positive samples so a single poorly classified sample can cause large variation, but at the same time loss does not perfectly correlate with auc. it is possible to improve loss sometimes simply by calibrating predictions up or down, but gaining nothing in terms of auc. I am tending to focus more on loss because it is much smoother and does generally map to auc.  </p>",
      "rawMarkdown": "It is a bit of a difficult question because auc is rather noisy in comparison. ETT - Abnormal only has a handful of positive samples so a single poorly classified sample can cause large variation, but at the same time loss does not perfectly correlate with auc. it is possible to improve loss sometimes simply by calibrating predictions up or down, but gaining nothing in terms of auc. I am tending to focus more on loss because it is much smoother and does generally map to auc.",
      "votes": null
    },
    {
      "id": "1222538",
      "postDate": "03/01/2021 21:50:54",
      "content": "<p>Also tried averaging checkpoints for the best auc and best loss and found no real performance gain there</p>",
      "rawMarkdown": "Also tried averaging checkpoints for the best auc and best loss and found no real performance gain there",
      "votes": null
    },
    {
      "id": "1224440",
      "postDate": "03/02/2021 18:14:15",
      "content": "<p>Yes, I had the same question too so I did an experiment by submitting lowest val loss and highest val auc (same fold but different epoch). lowest val loss always give me better LB result, so i'm not sure which one to trust (People always say trust your CV).</p>\n<p>I also noticed in <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> notebooks. He saved both best loss and best accuracy epochs but in the training he picked best loss epoch for each stage. </p>",
      "rawMarkdown": "Yes, I had the same question too so I did an experiment by submitting lowest val loss and highest val auc (same fold but different epoch). lowest val loss always give me better LB result, so i'm not sure which one to trust (People always say trust your CV).\n\nI also noticed in @yasufuminakama notebooks. He saved both best loss and best accuracy epochs but in the training he picked best loss epoch for each stage.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1221916,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "03/01/2021 12:01:56",
      "content": "<p>If you've doubt, why not use both and average their predictions during inference ? </p>\n<p>In Pytorch it's easy to add in the training process, in TF you can create custom callback if you use <code>model.fit</code>  or add it like in Pytorch if you use custom training loop. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1222001,
          "author_name": "valanm",
          "author_url": "",
          "post_date": "03/01/2021 13:12:37",
          "content": "<p>or use SWA and avoid minima from a single epoch!?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1222048,
          "author_name": "serigne",
          "author_url": "",
          "post_date": "03/01/2021 14:01:50",
          "content": "<p>SWA/EMA is  good alternative</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1222023,
      "author_name": "h053473666",
      "author_url": "",
      "post_date": "03/01/2021 13:33:29",
      "content": "<p>I asked a similar question before, I chose val loss. Then according to my observation, if val loss and val auc rise at the same time, the model is usually not very good and the LB score is not high.</p>\n<p><a href=\"https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145\" target=\"_blank\">https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1222537,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "03/01/2021 21:50:48",
      "content": "<p>It is a bit of a difficult question because auc is rather noisy in comparison. ETT - Abnormal only has a handful of positive samples so a single poorly classified sample can cause large variation, but at the same time loss does not perfectly correlate with auc. it is possible to improve loss sometimes simply by calibrating predictions up or down, but gaining nothing in terms of auc. I am tending to focus more on loss because it is much smoother and does generally map to auc.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 1222538,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "03/01/2021 21:50:54",
          "content": "<p>Also tried averaging checkpoints for the best auc and best loss and found no real performance gain there</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1224440,
      "author_name": "tom88jerry",
      "author_url": "",
      "post_date": "03/02/2021 18:14:15",
      "content": "<p>Yes, I had the same question too so I did an experiment by submitting lowest val loss and highest val auc (same fold but different epoch). lowest val loss always give me better LB result, so i'm not sure which one to trust (People always say trust your CV).</p>\n<p>I also noticed in <a href=\"https://www.kaggle.com/yasufuminakama\" target=\"_blank\">@yasufuminakama</a> notebooks. He saved both best loss and best accuracy epochs but in the training he picked best loss epoch for each stage. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1221897": "Hi guys, this has been bugging me for a while and I'm not sure what the answer is.\n\nWhat do you use to determine when to early stop & checkpoint your models?\n1. Validation loss (minimise)\n2. AUC (maximise)\n\nSometimes you can get a strange situation where valid loss increases (i.e. overfitting) but AUC is also increasing (noise maybe??). If you stop/checkpoint using valid loss, you could potentially be leaving a few AUC points on the table, but minimising the risk of overfitting and vice versa if you use AUC. What are your thoughts?",
    "1221916": "If you've doubt, why not use both and average their predictions during inference ? \n\nIn Pytorch it's easy to add in the training process, in TF you can create custom callback if you use `model.fit`  or add it like in Pytorch if you use custom training loop.",
    "1222001": "or use SWA and avoid minima from a single epoch!?",
    "1222023": "I asked a similar question before, I chose val loss. Then according to my observation, if val loss and val auc rise at the same time, the model is usually not very good and the LB score is not high.\n\n[https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145](https://www.kaggle.com/c/ranzcr-clip-catheter-line-classification/discussion/217145)",
    "1222048": "SWA/EMA is  good alternative",
    "1222537": "It is a bit of a difficult question because auc is rather noisy in comparison. ETT - Abnormal only has a handful of positive samples so a single poorly classified sample can cause large variation, but at the same time loss does not perfectly correlate with auc. it is possible to improve loss sometimes simply by calibrating predictions up or down, but gaining nothing in terms of auc. I am tending to focus more on loss because it is much smoother and does generally map to auc.",
    "1222538": "Also tried averaging checkpoints for the best auc and best loss and found no real performance gain there",
    "1224440": "Yes, I had the same question too so I did an experiment by submitting lowest val loss and highest val auc (same fold but different epoch). lowest val loss always give me better LB result, so i'm not sure which one to trust (People always say trust your CV).\n\nI also noticed in @yasufuminakama notebooks. He saved both best loss and best accuracy epochs but in the training he picked best loss epoch for each stage."
  },
  "source": "meta"
}