{
  "id": 200952,
  "title": "How to properly pick the test fold in k-fold validation?",
  "url": "/competitions/cassava-leaf-disease-classification/discussion/200952",
  "author_name": "",
  "post_date": "2020-12-02T14:00:14.839596300Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Image you do a 5-fold startegy to test your models </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3245800%2F2574edf8c66c54472a2e3fe5230c8ee4%2F1fXzJ.png?generation=1606917533195392&amp;alt=media\" alt=\"\"></p>\n<p>Wich of the 5 experiments do you pick? do you keep that test set fixed during all the competition?<br>\nI guess for the first question is the one getting  a better socre in test set right? But if private LB has a different distribtion , does that become a problem?</p>",
  "messages": [
    {
      "id": "1099642",
      "postDate": "12/02/2020 14:00:14",
      "content": "<p>Image you do a 5-fold startegy to test your models </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3245800%2F2574edf8c66c54472a2e3fe5230c8ee4%2F1fXzJ.png?generation=1606917533195392&amp;alt=media\" alt=\"\"></p>\n<p>Wich of the 5 experiments do you pick? do you keep that test set fixed during all the competition?<br>\nI guess for the first question is the one getting  a better socre in test set right? But if private LB has a different distribtion , does that become a problem?</p>",
      "rawMarkdown": "Image you do a 5-fold startegy to test your models \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3245800%2F2574edf8c66c54472a2e3fe5230c8ee4%2F1fXzJ.png?generation=1606917533195392&alt=media)\n\nWich of the 5 experiments do you pick? do you keep that test set fixed during all the competition?\nI guess for the first question is the one getting  a better socre in test set right? But if private LB has a different distribtion , does that become a problem?",
      "votes": null
    },
    {
      "id": "1099681",
      "postDate": "12/02/2020 14:29:45",
      "content": "<p>Best move is to inference all five experiments (keep probabilities for each class). Than average the probabilities for each class and take the max class for your prediction.</p>\n<p>I think the 9 hour limit in committed GPU notebooks is sufficient for running all 5 inferences.</p>\n<p>-Rich</p>",
      "rawMarkdown": "Best move is to inference all five experiments (keep probabilities for each class). Than average the probabilities for each class and take the max class for your prediction.\n\nI think the 9 hour limit in committed GPU notebooks is sufficient for running all 5 inferences.\n\n-Rich",
      "votes": null
    },
    {
      "id": "1099710",
      "postDate": "12/02/2020 14:55:38",
      "content": "<p>Yes but then you dont have a CV to check if something is better than previous models, trusting the public LB is not the best option right?</p>",
      "rawMarkdown": "Yes but then you dont have a CV to check if something is better than previous models, trusting the public LB is not the best option right?",
      "votes": null
    },
    {
      "id": "1101039",
      "postDate": "12/03/2020 15:30:07",
      "content": "<p>IMO:</p>\n<ol>\n<li><p>cross-validation is used to look at whether your model can generalize well on every split test set data, then you just need to average the metrics of all the experiments. Then, you do not need to pick one of the experiments.</p></li>\n<li><p>Yes. therefore, your model should be robust under different distribution and unseen data.</p></li>\n</ol>",
      "rawMarkdown": "IMO:\n1. cross-validation is used to look at whether your model can generalize well on every split test set data, then you just need to average the metrics of all the experiments. Then, you do not need to pick one of the experiments.\n\n2. Yes. therefore, your model should be robust under different distribution and unseen data.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1099681,
      "author_name": "richardepstein",
      "author_url": "",
      "post_date": "12/02/2020 14:29:45",
      "content": "<p>Best move is to inference all five experiments (keep probabilities for each class). Than average the probabilities for each class and take the max class for your prediction.</p>\n<p>I think the 9 hour limit in committed GPU notebooks is sufficient for running all 5 inferences.</p>\n<p>-Rich</p>",
      "votes": null,
      "replies": [
        {
          "id": 1099710,
          "author_name": "marcelosanchezortega",
          "author_url": "",
          "post_date": "12/02/2020 14:55:38",
          "content": "<p>Yes but then you dont have a CV to check if something is better than previous models, trusting the public LB is not the best option right?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1101039,
      "author_name": "andreaschandra",
      "author_url": "",
      "post_date": "12/03/2020 15:30:07",
      "content": "<p>IMO:</p>\n<ol>\n<li><p>cross-validation is used to look at whether your model can generalize well on every split test set data, then you just need to average the metrics of all the experiments. Then, you do not need to pick one of the experiments.</p></li>\n<li><p>Yes. therefore, your model should be robust under different distribution and unseen data.</p></li>\n</ol>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1099642": "Image you do a 5-fold startegy to test your models \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F3245800%2F2574edf8c66c54472a2e3fe5230c8ee4%2F1fXzJ.png?generation=1606917533195392&alt=media)\n\nWich of the 5 experiments do you pick? do you keep that test set fixed during all the competition?\nI guess for the first question is the one getting  a better socre in test set right? But if private LB has a different distribtion , does that become a problem?",
    "1099681": "Best move is to inference all five experiments (keep probabilities for each class). Than average the probabilities for each class and take the max class for your prediction.\n\nI think the 9 hour limit in committed GPU notebooks is sufficient for running all 5 inferences.\n\n-Rich",
    "1099710": "Yes but then you dont have a CV to check if something is better than previous models, trusting the public LB is not the best option right?",
    "1101039": "IMO:\n1. cross-validation is used to look at whether your model can generalize well on every split test set data, then you just need to average the metrics of all the experiments. Then, you do not need to pick one of the experiments.\n\n2. Yes. therefore, your model should be robust under different distribution and unseen data."
  },
  "source": "meta"
}