{
  "id": 239684,
  "title": "Using validation loss or validation AUC for early stopping",
  "url": "/competitions/seti-breakthrough-listen/discussion/239684",
  "author_name": "James Howard",
  "post_date": "2021-05-17T10:16:47.921000",
  "votes": 16,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Obviously there are issues with using normal augmentations in this competition, so our models overfit quickly and we often have to stop early. Assuming you save your best epoch as you train, how do you select best?</p>\n<p>It's an interesting question, I think.</p>\n<p>Whilst obviously your loss and AUC will be correlated, your BCEloss may be dominated by a few examples that are very 'wrong' in confidence, whilst they will contribute much less to AUC. If BCEloss didn't have an epsilon, you could technically get infinite loss from a single example predicted wrongly with 100% confidence.</p>\n<p>Furthermore, one could argue the absolute 'confidence' with which you make a prediction is irrelevant in this task, it is just the relative magnitude between needles and hay.</p>\n<p>However, early stopping using AUC also has its limitations - it's even easier to overfit to AUC (I strongly suspect) as you lose your granularity of overall model performance when it's only the few cases at the borders that are make any movement, with big swings epoch to epoch (although maybe this is a good thing?)</p>\n<p>EDIT: I suspected that fold selection based on validation loss would be superior, but in my (very limited) experimentation, that's not the case. I have a few runs where CV validation loss using the same architecture was much better, but a worse CV AUC, and indeed, they perform worse on the LB. I'm surprised.</p>",
  "messages": [
    {
      "id": 1311323,
      "postDate": "2021-05-17T10:16:47.920Z",
      "content": "<p>Obviously there are issues with using normal augmentations in this competition, so our models overfit quickly and we often have to stop early. Assuming you save your best epoch as you train, how do you select best?</p>\n<p>It's an interesting question, I think.</p>\n<p>Whilst obviously your loss and AUC will be correlated, your BCEloss may be dominated by a few examples that are very 'wrong' in confidence, whilst they will contribute much less to AUC. If BCEloss didn't have an epsilon, you could technically get infinite loss from a single example predicted wrongly with 100% confidence.</p>\n<p>Furthermore, one could argue the absolute 'confidence' with which you make a prediction is irrelevant in this task, it is just the relative magnitude between needles and hay.</p>\n<p>However, early stopping using AUC also has its limitations - it's even easier to overfit to AUC (I strongly suspect) as you lose your granularity of overall model performance when it's only the few cases at the borders that are make any movement, with big swings epoch to epoch (although maybe this is a good thing?)</p>\n<p>EDIT: I suspected that fold selection based on validation loss would be superior, but in my (very limited) experimentation, that's not the case. I have a few runs where CV validation loss using the same architecture was much better, but a worse CV AUC, and indeed, they perform worse on the LB. I'm surprised.</p>",
      "rawMarkdown": "Obviously there are issues with using normal augmentations in this competition, so our models overfit quickly and we often have to stop early. Assuming you save your best epoch as you train, how do you select best?\n\nIt's an interesting question, I think.\n\nWhilst obviously your loss and AUC will be correlated, your BCEloss may be dominated by a few examples that are very 'wrong' in confidence, whilst they will contribute much less to AUC. If BCEloss didn't have an epsilon, you could technically get infinite loss from a single example predicted wrongly with 100% confidence.\n\nFurthermore, one could argue the absolute 'confidence' with which you make a prediction is irrelevant in this task, it is just the relative magnitude between needles and hay.\n\nHowever, early stopping using AUC also has its limitations - it's even easier to overfit to AUC (I strongly suspect) as you lose your granularity of overall model performance when it's only the few cases at the borders that are make any movement, with big swings epoch to epoch (although maybe this is a good thing?)\n\nEDIT: I suspected that fold selection based on validation loss would be superior, but in my (very limited) experimentation, that's not the case. I have a few runs where CV validation loss using the same architecture was much better, but a worse CV AUC, and indeed, they perform worse on the LB. I'm surprised.",
      "votes": 16
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "1311323": "Obviously there are issues with using normal augmentations in this competition, so our models overfit quickly and we often have to stop early. Assuming you save your best epoch as you train, how do you select best?\n\nIt's an interesting question, I think.\n\nWhilst obviously your loss and AUC will be correlated, your BCEloss may be dominated by a few examples that are very 'wrong' in confidence, whilst they will contribute much less to AUC. If BCEloss didn't have an epsilon, you could technically get infinite loss from a single example predicted wrongly with 100% confidence.\n\nFurthermore, one could argue the absolute 'confidence' with which you make a prediction is irrelevant in this task, it is just the relative magnitude between needles and hay.\n\nHowever, early stopping using AUC also has its limitations - it's even easier to overfit to AUC (I strongly suspect) as you lose your granularity of overall model performance when it's only the few cases at the borders that are make any movement, with big swings epoch to epoch (although maybe this is a good thing?)\n\nEDIT: I suspected that fold selection based on validation loss would be superior, but in my (very limited) experimentation, that's not the case. I have a few runs where CV validation loss using the same architecture was much better, but a worse CV AUC, and indeed, they perform worse on the LB. I'm surprised."
  }
}