{
  "id": 418733,
  "title": "Did anyone understand what ratio to split to train and validation or any other stratergy ?",
  "url": "/competitions/hubmap-hacking-the-human-vasculature/discussion/418733",
  "author_name": "",
  "post_date": "2023-06-22T09:51:08.335711Z",
  "votes": 2,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Since the test data is full of dataset 1 that is expert verified,</p>\n<p>and the public test set is just 28 percent of the whole test set.</p>\n<p>What would be the reasonable approach to corelate CV and LB??</p>",
  "messages": [
    {
      "id": "2312959",
      "postDate": "06/22/2023 09:51:08",
      "content": "<p>Since the test data is full of dataset 1 that is expert verified,</p>\n<p>and the public test set is just 28 percent of the whole test set.</p>\n<p>What would be the reasonable approach to corelate CV and LB??</p>",
      "rawMarkdown": "Since the test data is full of dataset 1 that is expert verified,\n\nand the public test set is just 28 percent of the whole test set.\n\nWhat would be the reasonable approach to corelate CV and LB??",
      "votes": null
    },
    {
      "id": "2321625",
      "postDate": "06/28/2023 17:50:23",
      "content": "<p>I think it is really difficult to know the really the function between the CV and LB. </p>\n<p>Now I didn't split dataset 1 or 2, I just split all data with some ratio. </p>\n<p>Usually, we use 0.8:0.2 ratio for training and test, or 0.7:0.3</p>\n<p>But to utilize all dataset, don't forget retrain your model with all data. </p>\n<p>if you want to track the performance during training, K-fold is a GOOD CHOICE</p>\n<p>and this <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038#2279515\" target=\"_blank\">discussion </a> and this <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2315708\" target=\"_blank\">one </a> maybe help a little.</p>",
      "rawMarkdown": "I think it is really difficult to know the really the function between the CV and LB. \n\nNow I didn't split dataset 1 or 2, I just split all data with some ratio. \n\nUsually, we use 0.8:0.2 ratio for training and test, or 0.7:0.3\n\nBut to utilize all dataset, don't forget retrain your model with all data. \n\nif you want to track the performance during training, K-fold is a GOOD CHOICE\n\nand this [discussion ](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038#2279515) and this [one ](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2315708) maybe help a little.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2321625,
      "author_name": "chg0901",
      "author_url": "",
      "post_date": "06/28/2023 17:50:23",
      "content": "<p>I think it is really difficult to know the really the function between the CV and LB. </p>\n<p>Now I didn't split dataset 1 or 2, I just split all data with some ratio. </p>\n<p>Usually, we use 0.8:0.2 ratio for training and test, or 0.7:0.3</p>\n<p>But to utilize all dataset, don't forget retrain your model with all data. </p>\n<p>if you want to track the performance during training, K-fold is a GOOD CHOICE</p>\n<p>and this <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038#2279515\" target=\"_blank\">discussion </a> and this <a href=\"https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2315708\" target=\"_blank\">one </a> maybe help a little.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2312959": "Since the test data is full of dataset 1 that is expert verified,\n\nand the public test set is just 28 percent of the whole test set.\n\nWhat would be the reasonable approach to corelate CV and LB??",
    "2321625": "I think it is really difficult to know the really the function between the CV and LB. \n\nNow I didn't split dataset 1 or 2, I just split all data with some ratio. \n\nUsually, we use 0.8:0.2 ratio for training and test, or 0.7:0.3\n\nBut to utilize all dataset, don't forget retrain your model with all data. \n\nif you want to track the performance during training, K-fold is a GOOD CHOICE\n\nand this [discussion ](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/413038#2279515) and this [one ](https://www.kaggle.com/competitions/hubmap-hacking-the-human-vasculature/discussion/419143#2315708) maybe help a little."
  },
  "source": "meta"
}