{
  "id": 170379,
  "title": "What should we use as a monitor? val_auc or val_loss?",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/170379",
  "author_name": "",
  "post_date": "2020-07-27T12:40:25.988432200Z",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi! I wonder whether I should use val_auc as a monitor. In other words, which measurement should I use when I pick the best model in each fold? My concern is that if we use val_auc to select the best model in each fold, we may overfit to the fold because val_auc only takes the order of prediction into account. I'm doing some experiments now and will report the result when it's ready. I really want to hear your opinion.</p>",
  "messages": [
    {
      "id": "947702",
      "postDate": "07/27/2020 12:40:25",
      "content": "<p>Hi! I wonder whether I should use val_auc as a monitor. In other words, which measurement should I use when I pick the best model in each fold? My concern is that if we use val_auc to select the best model in each fold, we may overfit to the fold because val_auc only takes the order of prediction into account. I'm doing some experiments now and will report the result when it's ready. I really want to hear your opinion.</p>",
      "rawMarkdown": "Hi! I wonder whether I should use val_auc as a monitor. In other words, which measurement should I use when I pick the best model in each fold? My concern is that if we use val_auc to select the best model in each fold, we may overfit to the fold because val_auc only takes the order of prediction into account. I'm doing some experiments now and will report the result when it's ready. I really want to hear your opinion.",
      "votes": null
    },
    {
      "id": "947708",
      "postDate": "07/27/2020 12:43:49",
      "content": "<p>It is up to you to pick. In my experiments, I use <code>val_auc</code> to measure the performance but you can use <code>val_loss</code>. It just depends on what kind of model you want to save. The one with the highest val auc or the one with the lowest val loss. Based on that choice you can pick one of them to monitor. </p>",
      "rawMarkdown": "It is up to you to pick. In my experiments, I use `val_auc` to measure the performance but you can use `val_loss`. It just depends on what kind of model you want to save. The one with the highest val auc or the one with the lowest val loss. Based on that choice you can pick one of them to monitor.",
      "votes": null
    },
    {
      "id": "947789",
      "postDate": "07/27/2020 13:39:13",
      "content": "<p>But <code>val_auc</code> is the one we should look for, as val_auc closely resembles to the selected metric, which we are trying to optimise for?</p>",
      "rawMarkdown": "But `val_auc` is the one we should look for, as val_auc closely resembles to the selected metric, which we are trying to optimise for?",
      "votes": null
    },
    {
      "id": "947798",
      "postDate": "07/27/2020 13:41:59",
      "content": "<p>Indeed. We should look for <code>val_auc</code>. But, in the end it is an individual's choice. </p>",
      "rawMarkdown": "Indeed. We should look for `val_auc`. But, in the end it is an individual's choice.",
      "votes": null
    },
    {
      "id": "948000",
      "postDate": "07/27/2020 15:49:21",
      "content": "<p>This is a good question and one which I have debated with myself. The typical answer is 'personal preference' which didn't sit well with me for some time as my instinct was that the answer should be more scientific than this.</p>\n\n<p>However, now having run many experiments across different models and image sizes using cross validation, in actual fact using either gives very similar results in terms of assessing how long to train models for optimum performance with the validation data. Some folds may diverge a little but they average out.</p>",
      "rawMarkdown": "This is a good question and one which I have debated with myself. The typical answer is 'personal preference' which didn't sit well with me for some time as my instinct was that the answer should be more scientific than this.\n\nHowever, now having run many experiments across different models and image sizes using cross validation, in actual fact using either gives very similar results in terms of assessing how long to train models for optimum performance with the validation data. Some folds may diverge a little but they average out.",
      "votes": null
    },
    {
      "id": "948025",
      "postDate": "07/27/2020 16:07:03",
      "content": "<p>This is valid question. My experience shows that score is very noisy, people often choose model based on noisy score and they turn their brain off, the result is always big shakeup after competition end.</p>",
      "rawMarkdown": "This is valid question. My experience shows that score is very noisy, people often choose model based on noisy score and they turn their brain off, the result is always big shakeup after competition end.",
      "votes": null
    },
    {
      "id": "949496",
      "postDate": "07/28/2020 17:09:39",
      "content": "<p>I tried both as early stopping criterion val-auc and val-loss, my results were better (by 5%) when I used oof val-auc. With the same settings I tried val-loss and the results were worse. The model my ES algorithm choose was just not good enough in this case.</p>",
      "rawMarkdown": "I tried both as early stopping criterion val-auc and val-loss, my results were better (by 5%) when I used oof val-auc. With the same settings I tried val-loss and the results were worse. The model my ES algorithm choose was just not good enough in this case.",
      "votes": null
    },
    {
      "id": "949555",
      "postDate": "07/28/2020 17:56:29",
      "content": "<p>I would say, it depends. It is hard to find a loss which will directly optimize for ROC AUC. Most popular losses like Cross Entropy Loss or Focal Loss are proxies to optimize classification performance of the model. There are more esoteric losses discussed <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611\">here</a> which are trying to minimize miss-classification for both classes and more close to the ROC AUC task (as a task of ranking samples correctly). Unfortunately, they are not plug-and-play type of losses and need be implemented with a great care.\nAnswering the question: if using common losses, I think it would not hurt to monitor ROC-AUC score ;)\nCheers!</p>",
      "rawMarkdown": "I would say, it depends. It is hard to find a loss which will directly optimize for ROC AUC. Most popular losses like Cross Entropy Loss or Focal Loss are proxies to optimize classification performance of the model. There are more esoteric losses discussed [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611) which are trying to minimize miss-classification for both classes and more close to the ROC AUC task (as a task of ranking samples correctly). Unfortunately, they are not plug-and-play type of losses and need be implemented with a great care.\nAnswering the question: if using common losses, I think it would not hurt to monitor ROC-AUC score ;)\nCheers!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 947708,
      "author_name": "urvishp80",
      "author_url": "",
      "post_date": "07/27/2020 12:43:49",
      "content": "<p>It is up to you to pick. In my experiments, I use <code>val_auc</code> to measure the performance but you can use <code>val_loss</code>. It just depends on what kind of model you want to save. The one with the highest val auc or the one with the lowest val loss. Based on that choice you can pick one of them to monitor. </p>",
      "votes": null,
      "replies": [
        {
          "id": 947789,
          "author_name": "kurianbenoy",
          "author_url": "",
          "post_date": "07/27/2020 13:39:13",
          "content": "<p>But <code>val_auc</code> is the one we should look for, as val_auc closely resembles to the selected metric, which we are trying to optimise for?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 947798,
          "author_name": "urvishp80",
          "author_url": "",
          "post_date": "07/27/2020 13:41:59",
          "content": "<p>Indeed. We should look for <code>val_auc</code>. But, in the end it is an individual's choice. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 948000,
      "author_name": "jsyphil",
      "author_url": "",
      "post_date": "07/27/2020 15:49:21",
      "content": "<p>This is a good question and one which I have debated with myself. The typical answer is 'personal preference' which didn't sit well with me for some time as my instinct was that the answer should be more scientific than this.</p>\n\n<p>However, now having run many experiments across different models and image sizes using cross validation, in actual fact using either gives very similar results in terms of assessing how long to train models for optimum performance with the validation data. Some folds may diverge a little but they average out.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 948025,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "07/27/2020 16:07:03",
      "content": "<p>This is valid question. My experience shows that score is very noisy, people often choose model based on noisy score and they turn their brain off, the result is always big shakeup after competition end.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 949496,
      "author_name": "aliabdin1",
      "author_url": "",
      "post_date": "07/28/2020 17:09:39",
      "content": "<p>I tried both as early stopping criterion val-auc and val-loss, my results were better (by 5%) when I used oof val-auc. With the same settings I tried val-loss and the results were worse. The model my ES algorithm choose was just not good enough in this case.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 949555,
      "author_name": "ademyanchuk",
      "author_url": "",
      "post_date": "07/28/2020 17:56:29",
      "content": "<p>I would say, it depends. It is hard to find a loss which will directly optimize for ROC AUC. Most popular losses like Cross Entropy Loss or Focal Loss are proxies to optimize classification performance of the model. There are more esoteric losses discussed <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611\">here</a> which are trying to minimize miss-classification for both classes and more close to the ROC AUC task (as a task of ranking samples correctly). Unfortunately, they are not plug-and-play type of losses and need be implemented with a great care.\nAnswering the question: if using common losses, I think it would not hurt to monitor ROC-AUC score ;)\nCheers!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "947702": "Hi! I wonder whether I should use val_auc as a monitor. In other words, which measurement should I use when I pick the best model in each fold? My concern is that if we use val_auc to select the best model in each fold, we may overfit to the fold because val_auc only takes the order of prediction into account. I'm doing some experiments now and will report the result when it's ready. I really want to hear your opinion.",
    "947708": "It is up to you to pick. In my experiments, I use `val_auc` to measure the performance but you can use `val_loss`. It just depends on what kind of model you want to save. The one with the highest val auc or the one with the lowest val loss. Based on that choice you can pick one of them to monitor.",
    "947789": "But `val_auc` is the one we should look for, as val_auc closely resembles to the selected metric, which we are trying to optimise for?",
    "947798": "Indeed. We should look for `val_auc`. But, in the end it is an individual's choice.",
    "948000": "This is a good question and one which I have debated with myself. The typical answer is 'personal preference' which didn't sit well with me for some time as my instinct was that the answer should be more scientific than this.\n\nHowever, now having run many experiments across different models and image sizes using cross validation, in actual fact using either gives very similar results in terms of assessing how long to train models for optimum performance with the validation data. Some folds may diverge a little but they average out.",
    "948025": "This is valid question. My experience shows that score is very noisy, people often choose model based on noisy score and they turn their brain off, the result is always big shakeup after competition end.",
    "949496": "I tried both as early stopping criterion val-auc and val-loss, my results were better (by 5%) when I used oof val-auc. With the same settings I tried val-loss and the results were worse. The model my ES algorithm choose was just not good enough in this case.",
    "949555": "I would say, it depends. It is hard to find a loss which will directly optimize for ROC AUC. Most popular losses like Cross Entropy Loss or Focal Loss are proxies to optimize classification performance of the model. There are more esoteric losses discussed [here](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611) which are trying to minimize miss-classification for both classes and more close to the ROC AUC task (as a task of ranking samples correctly). Unfortunately, they are not plug-and-play type of losses and need be implemented with a great care.\nAnswering the question: if using common losses, I think it would not hurt to monitor ROC-AUC score ;)\nCheers!"
  },
  "source": "meta"
}