{
  "id": 549003,
  "title": "How to design a good Validation?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/549003",
  "author_name": "",
  "post_date": "2024-11-30T07:10:57.954149100Z",
  "votes": 1,
  "comment_count": 3,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> <br>\nIn actual fact, trusting the public LB is not advisable because it is quite often misleading. </p>\n<p>I have some considerations.</p>\n<ol>\n<li>Are the training samples and testing samples of same kind? I mean is there a sampling bias?</li>\n<li>If I try to get more samples from <a href=\"https://cryoetdataportal.czscience.com/\" target=\"_blank\">Cryo ET Data portal</a> and <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544740\" target=\"_blank\">previous literature</a> then will that lead to sampling bias?</li>\n<li>How does overfitting really happen specifically to this type of data and approaches? The focus of the competition is clearly to develop a generalizable machine learning solution.</li>\n<li>What choice of validation scheme is advisable? Holdout system, Kfold CV, Bootstrap</li>\n</ol>\n<p>As suggested by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, using scanning window for processing and modelling is quite useful and normalization in important.</p>",
  "messages": [
    {
      "id": "3058958",
      "postDate": "11/30/2024 07:10:57",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> <br>\nIn actual fact, trusting the public LB is not advisable because it is quite often misleading. </p>\n<p>I have some considerations.</p>\n<ol>\n<li>Are the training samples and testing samples of same kind? I mean is there a sampling bias?</li>\n<li>If I try to get more samples from <a href=\"https://cryoetdataportal.czscience.com/\" target=\"_blank\">Cryo ET Data portal</a> and <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544740\" target=\"_blank\">previous literature</a> then will that lead to sampling bias?</li>\n<li>How does overfitting really happen specifically to this type of data and approaches? The focus of the competition is clearly to develop a generalizable machine learning solution.</li>\n<li>What choice of validation scheme is advisable? Holdout system, Kfold CV, Bootstrap</li>\n</ol>\n<p>As suggested by <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, using scanning window for processing and modelling is quite useful and normalization in important.</p>",
      "rawMarkdown": "kharrington \nIn actual fact, trusting the public LB is not advisable because it is quite often misleading. \n\nI have some considerations.\n1. Are the training samples and testing samples of same kind? I mean is there a sampling bias?\n2. If I try to get more samples from [Cryo ET Data portal](https://cryoetdataportal.czscience.com/) and [previous literature](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544740) then will that lead to sampling bias?\n3. How does overfitting really happen specifically to this type of data and approaches? The focus of the competition is clearly to develop a generalizable machine learning solution.\n4. What choice of validation scheme is advisable? Holdout system, Kfold CV, Bootstrap\n\nAs suggested by @hengck23, using scanning window for processing and modelling is quite useful and normalization in important.",
      "votes": null
    },
    {
      "id": "3059694",
      "postDate": "12/01/2024 01:52:59",
      "content": "<p>Of the 7 training sets, I'm using 1-5 for training, 1 for validation, and 1 as a holdout at least initially while I'm exploring different approaches.  Eventually, I'll probably turn that into 7-fold cv, but what I have right now is good enough to run meaningful tests.</p>",
      "rawMarkdown": "Of the 7 training sets, I'm using 1-5 for training, 1 for validation, and 1 as a holdout at least initially while I'm exploring different approaches.  Eventually, I'll probably turn that into 7-fold cv, but what I have right now is good enough to run meaningful tests.",
      "votes": null
    },
    {
      "id": "3060130",
      "postDate": "12/01/2024 12:20:04",
      "content": "<p>Please imagine ( and try it if u have time ). Download cifar10 images. Select 80% as train. 2% as validation 18% as test. They represent your local cv and hidden test respectively. Try a simple cnn like resenet34. Do experiments, does cv correlates to lb here for the cifar experiment here? If no how to improve </p>",
      "rawMarkdown": "Please imagine ( and try it if u have time ). Download cifar10 images. Select 80% as train. 2% as validation 18% as test. They represent your local cv and hidden test respectively. Try a simple cnn like resenet34. Do experiments, does cv correlates to lb here for the cifar experiment here? If no how to improve",
      "votes": null
    },
    {
      "id": "3063090",
      "postDate": "12/04/2024 06:33:12",
      "content": "<p>Thank you!  I will give it a try.</p>",
      "rawMarkdown": "Thank you!  I will give it a try.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3059694,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "12/01/2024 01:52:59",
      "content": "<p>Of the 7 training sets, I'm using 1-5 for training, 1 for validation, and 1 as a holdout at least initially while I'm exploring different approaches.  Eventually, I'll probably turn that into 7-fold cv, but what I have right now is good enough to run meaningful tests.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3060130,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "12/01/2024 12:20:04",
          "content": "<p>Please imagine ( and try it if u have time ). Download cifar10 images. Select 80% as train. 2% as validation 18% as test. They represent your local cv and hidden test respectively. Try a simple cnn like resenet34. Do experiments, does cv correlates to lb here for the cifar experiment here? If no how to improve </p>",
          "votes": null,
          "replies": [
            {
              "id": 3063090,
              "author_name": "davidlist",
              "author_url": "",
              "post_date": "12/04/2024 06:33:12",
              "content": "<p>Thank you!  I will give it a try.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3058958": "kharrington \nIn actual fact, trusting the public LB is not advisable because it is quite often misleading. \n\nI have some considerations.\n1. Are the training samples and testing samples of same kind? I mean is there a sampling bias?\n2. If I try to get more samples from [Cryo ET Data portal](https://cryoetdataportal.czscience.com/) and [previous literature](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544740) then will that lead to sampling bias?\n3. How does overfitting really happen specifically to this type of data and approaches? The focus of the competition is clearly to develop a generalizable machine learning solution.\n4. What choice of validation scheme is advisable? Holdout system, Kfold CV, Bootstrap\n\nAs suggested by @hengck23, using scanning window for processing and modelling is quite useful and normalization in important.",
    "3059694": "Of the 7 training sets, I'm using 1-5 for training, 1 for validation, and 1 as a holdout at least initially while I'm exploring different approaches.  Eventually, I'll probably turn that into 7-fold cv, but what I have right now is good enough to run meaningful tests.",
    "3060130": "Please imagine ( and try it if u have time ). Download cifar10 images. Select 80% as train. 2% as validation 18% as test. They represent your local cv and hidden test respectively. Try a simple cnn like resenet34. Do experiments, does cv correlates to lb here for the cifar experiment here? If no how to improve",
    "3063090": "Thank you!  I will give it a try."
  },
  "source": "meta"
}