{
  "id": 164291,
  "title": "Which metric to use and workflow for model tuning",
  "url": "/competitions/siim-isic-melanoma-classification/discussion/164291",
  "author_name": "",
  "post_date": "2020-07-05T14:37:54.197654200Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Hello,</p>\n\n<p>I am new to Kaggle. In ML problems in general I am use to split my data between train and test, then pick a metric that is relevant for the problem and follow the evolution of that metric to see if the changes I am doing are making the model better or worse.</p>\n\n<p>In this case I tried with AUC and F1 score (over the test set, which in my case was 10% randomly sampled, also I included some melanoma images from 2019 dataset as to overcome the class imbalance problem, I also tried data augmentation with same results), and in both cases although I managed to improve those metrics in my test set when I am submitting my results the LB score is getting worse.</p>\n\n<p>For example, I started with a relatively simple model with a CNN, and achieved 0.835, then I improved my best AUC score and tried submitting that and in the LB I got a score of ~0.75. Then I tried F1 score, on the test set got a result of 0.99 for benign class and 0.92 for malignant class, which I thought it was great, but then again, I submit my predictions and got a lower score.  I generally pick the model that resulted with the best validation metric from all my epochs (using in Keras ModelCheckpoint), but seems this is a bit nonsense in this case, as the validation metrics I am picking don't have anything to do with the LB end results.</p>\n\n<p>What will be the optimal workflow in this case?</p>\n\n<p>Thank you</p>",
  "messages": [
    {
      "id": "916306",
      "postDate": "07/05/2020 14:37:54",
      "content": "<p>Hello,</p>\n\n<p>I am new to Kaggle. In ML problems in general I am use to split my data between train and test, then pick a metric that is relevant for the problem and follow the evolution of that metric to see if the changes I am doing are making the model better or worse.</p>\n\n<p>In this case I tried with AUC and F1 score (over the test set, which in my case was 10% randomly sampled, also I included some melanoma images from 2019 dataset as to overcome the class imbalance problem, I also tried data augmentation with same results), and in both cases although I managed to improve those metrics in my test set when I am submitting my results the LB score is getting worse.</p>\n\n<p>For example, I started with a relatively simple model with a CNN, and achieved 0.835, then I improved my best AUC score and tried submitting that and in the LB I got a score of ~0.75. Then I tried F1 score, on the test set got a result of 0.99 for benign class and 0.92 for malignant class, which I thought it was great, but then again, I submit my predictions and got a lower score.  I generally pick the model that resulted with the best validation metric from all my epochs (using in Keras ModelCheckpoint), but seems this is a bit nonsense in this case, as the validation metrics I am picking don't have anything to do with the LB end results.</p>\n\n<p>What will be the optimal workflow in this case?</p>\n\n<p>Thank you</p>",
      "rawMarkdown": "Hello,\n\nI am new to Kaggle. In ML problems in general I am use to split my data between train and test, then pick a metric that is relevant for the problem and follow the evolution of that metric to see if the changes I am doing are making the model better or worse.\n\nIn this case I tried with AUC and F1 score (over the test set, which in my case was 10% randomly sampled, also I included some melanoma images from 2019 dataset as to overcome the class imbalance problem, I also tried data augmentation with same results), and in both cases although I managed to improve those metrics in my test set when I am submitting my results the LB score is getting worse.\n\nFor example, I started with a relatively simple model with a CNN, and achieved 0.835, then I improved my best AUC score and tried submitting that and in the LB I got a score of ~0.75. Then I tried F1 score, on the test set got a result of 0.99 for benign class and 0.92 for malignant class, which I thought it was great, but then again, I submit my predictions and got a lower score.  I generally pick the model that resulted with the best validation metric from all my epochs (using in Keras ModelCheckpoint), but seems this is a bit nonsense in this case, as the validation metrics I am picking don't have anything to do with the LB end results.\n\nWhat will be the optimal workflow in this case?\n\nThank you",
      "votes": null
    },
    {
      "id": "916330",
      "postDate": "07/05/2020 15:00:51",
      "content": "<p>This topic has been discussed already.</p>\n\n<p>See <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880870\">Soft margin focal loss</a>\nAlso see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611\">ROC-Star</a></p>",
      "rawMarkdown": "This topic has been discussed already.\n\nSee [Soft margin focal loss](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880870)\nAlso see [ROC-Star](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611)",
      "votes": null
    },
    {
      "id": "916350",
      "postDate": "07/05/2020 15:14:22",
      "content": "<p>Thank you very much for your answer, so you suggest to use the metric mentioned in those posts? Might you have any opinion on the early stopping for this problem?</p>",
      "rawMarkdown": "Thank you very much for your answer, so you suggest to use the metric mentioned in those posts? Might you have any opinion on the early stopping for this problem?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 916330,
      "author_name": "sirishks",
      "author_url": "",
      "post_date": "07/05/2020 15:00:51",
      "content": "<p>This topic has been discussed already.</p>\n\n<p>See <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880870\">Soft margin focal loss</a>\nAlso see <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611\">ROC-Star</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 916350,
          "author_name": "maiskovich",
          "author_url": "",
          "post_date": "07/05/2020 15:14:22",
          "content": "<p>Thank you very much for your answer, so you suggest to use the metric mentioned in those posts? Might you have any opinion on the early stopping for this problem?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "916306": "Hello,\n\nI am new to Kaggle. In ML problems in general I am use to split my data between train and test, then pick a metric that is relevant for the problem and follow the evolution of that metric to see if the changes I am doing are making the model better or worse.\n\nIn this case I tried with AUC and F1 score (over the test set, which in my case was 10% randomly sampled, also I included some melanoma images from 2019 dataset as to overcome the class imbalance problem, I also tried data augmentation with same results), and in both cases although I managed to improve those metrics in my test set when I am submitting my results the LB score is getting worse.\n\nFor example, I started with a relatively simple model with a CNN, and achieved 0.835, then I improved my best AUC score and tried submitting that and in the LB I got a score of ~0.75. Then I tried F1 score, on the test set got a result of 0.99 for benign class and 0.92 for malignant class, which I thought it was great, but then again, I submit my predictions and got a lower score.  I generally pick the model that resulted with the best validation metric from all my epochs (using in Keras ModelCheckpoint), but seems this is a bit nonsense in this case, as the validation metrics I am picking don't have anything to do with the LB end results.\n\nWhat will be the optimal workflow in this case?\n\nThank you",
    "916330": "This topic has been discussed already.\n\nSee [Soft margin focal loss](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/155201#880870)\nAlso see [ROC-Star](https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/160611)",
    "916350": "Thank you very much for your answer, so you suggest to use the metric mentioned in those posts? Might you have any opinion on the early stopping for this problem?"
  },
  "source": "meta"
}