{
  "id": 211532,
  "title": "3 tips to optimize CV and LB relation",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/211532",
  "author_name": "",
  "post_date": "2021-01-15T14:51:22.775019500Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<ol>\n<li><p>Try different seeds when hyperparameter tuning. I've found that models can have significantly different performance with the same hyperparameters with different seeds. While ultimately, it shouldn't matter if you perform 5-fold CV, this is helpful for searching for a validation split that gives you results consistent with the LB (while you're still experimenting).</p></li>\n<li><p>Regularize your models well. At some point in your training, especially when the model goes above 90% training accuracy, you will start to see the validation loss decrease along with your validation accuracy. This means that your model is becoming surer of its predictions, but it is getting worse results. Worse still, your training accuracy continues to climb while your validation accuracy continues to decrease. Clearly, it has overfit on the training set and possibly picked up noise present in the training set. Regularization helps prevent overfitting if applied correctly.</p></li>\n<li><p>Clean the data. This is the tricky part, and for me the least rewarding. I used out-of-fold accuracy as the benchmark to determine noisy labels. You have to experiment to determine the appropriate confidence level for your model to clean the data. Too high, and there is little impact overall. Too low, and you end up removing too much data. Very time-consuming, but helped me climb 20 positions on the LB. </p></li>\n</ol>",
  "messages": [
    {
      "id": "1154293",
      "postDate": "01/15/2021 14:51:22",
      "content": "<ol>\n<li><p>Try different seeds when hyperparameter tuning. I've found that models can have significantly different performance with the same hyperparameters with different seeds. While ultimately, it shouldn't matter if you perform 5-fold CV, this is helpful for searching for a validation split that gives you results consistent with the LB (while you're still experimenting).</p></li>\n<li><p>Regularize your models well. At some point in your training, especially when the model goes above 90% training accuracy, you will start to see the validation loss decrease along with your validation accuracy. This means that your model is becoming surer of its predictions, but it is getting worse results. Worse still, your training accuracy continues to climb while your validation accuracy continues to decrease. Clearly, it has overfit on the training set and possibly picked up noise present in the training set. Regularization helps prevent overfitting if applied correctly.</p></li>\n<li><p>Clean the data. This is the tricky part, and for me the least rewarding. I used out-of-fold accuracy as the benchmark to determine noisy labels. You have to experiment to determine the appropriate confidence level for your model to clean the data. Too high, and there is little impact overall. Too low, and you end up removing too much data. Very time-consuming, but helped me climb 20 positions on the LB. </p></li>\n</ol>",
      "rawMarkdown": "1. Try different seeds when hyperparameter tuning. I've found that models can have significantly different performance with the same hyperparameters with different seeds. While ultimately, it shouldn't matter if you perform 5-fold CV, this is helpful for searching for a validation split that gives you results consistent with the LB (while you're still experimenting).\n\n2. Regularize your models well. At some point in your training, especially when the model goes above 90% training accuracy, you will start to see the validation loss decrease along with your validation accuracy. This means that your model is becoming surer of its predictions, but it is getting worse results. Worse still, your training accuracy continues to climb while your validation accuracy continues to decrease. Clearly, it has overfit on the training set and possibly picked up noise present in the training set. Regularization helps prevent overfitting if applied correctly.\n\n3. Clean the data. This is the tricky part, and for me the least rewarding. I used out-of-fold accuracy as the benchmark to determine noisy labels. You have to experiment to determine the appropriate confidence level for your model to clean the data. Too high, and there is little impact overall. Too low, and you end up removing too much data. Very time-consuming, but helped me climb 20 positions on the LB.",
      "votes": null
    },
    {
      "id": "1154691",
      "postDate": "01/15/2021 19:31:23",
      "content": "<p>Have you tried implementing the cleanlab from another discussion thread from this competition? I had some good results for item 3, but have not figured out 100% how to deploy it as inference yet.</p>",
      "rawMarkdown": "Have you tried implementing the cleanlab from another discussion thread from this competition? I had some good results for item 3, but have not figured out 100% how to deploy it as inference yet.",
      "votes": null
    },
    {
      "id": "1154837",
      "postDate": "01/16/2021 01:27:10",
      "content": "<p>I've tried it briefly, but wasn't able to make it work with Keras implementation of pre-trained models. So I can't help there, sorry!</p>",
      "rawMarkdown": "I've tried it briefly, but wasn't able to make it work with Keras implementation of pre-trained models. So I can't help there, sorry!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1154691,
      "author_name": "capiru",
      "author_url": "",
      "post_date": "01/15/2021 19:31:23",
      "content": "<p>Have you tried implementing the cleanlab from another discussion thread from this competition? I had some good results for item 3, but have not figured out 100% how to deploy it as inference yet.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1154837,
          "author_name": "junyingsg",
          "author_url": "",
          "post_date": "01/16/2021 01:27:10",
          "content": "<p>I've tried it briefly, but wasn't able to make it work with Keras implementation of pre-trained models. So I can't help there, sorry!</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1154293": "1. Try different seeds when hyperparameter tuning. I've found that models can have significantly different performance with the same hyperparameters with different seeds. While ultimately, it shouldn't matter if you perform 5-fold CV, this is helpful for searching for a validation split that gives you results consistent with the LB (while you're still experimenting).\n\n2. Regularize your models well. At some point in your training, especially when the model goes above 90% training accuracy, you will start to see the validation loss decrease along with your validation accuracy. This means that your model is becoming surer of its predictions, but it is getting worse results. Worse still, your training accuracy continues to climb while your validation accuracy continues to decrease. Clearly, it has overfit on the training set and possibly picked up noise present in the training set. Regularization helps prevent overfitting if applied correctly.\n\n3. Clean the data. This is the tricky part, and for me the least rewarding. I used out-of-fold accuracy as the benchmark to determine noisy labels. You have to experiment to determine the appropriate confidence level for your model to clean the data. Too high, and there is little impact overall. Too low, and you end up removing too much data. Very time-consuming, but helped me climb 20 positions on the LB.",
    "1154691": "Have you tried implementing the cleanlab from another discussion thread from this competition? I had some good results for item 3, but have not figured out 100% how to deploy it as inference yet.",
    "1154837": "I've tried it briefly, but wasn't able to make it work with Keras implementation of pre-trained models. So I can't help there, sorry!"
  },
  "source": "meta"
}