{
  "id": 264087,
  "title": "Possible Solutions To CV *NOT* Correlating LB.",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/264087",
  "author_name": "",
  "post_date": "2021-08-11T04:11:22.701708900Z",
  "votes": 13,
  "comment_count": 4,
  "views": 0,
  "content": "<p>First I would like to point out that the LB public dataset is too small to consider and the private dataset is not big enough to give us a stable score so even after following CV the private LB can be expected to have a shakeup.</p>\n<ol>\n<li><p>Training in 5-10 Folds can give you a way more stable CV. </p></li>\n<li><p>It is also important to notice that the models overfits quite quickly on train dataset and methods like heavy augmentations, high dropout can be pretty effective.</p></li>\n<li><p>The main goal should be to improve generalization, but look out for possible \"model-specific\" bias in your trainings. Sometimes these bias can pay off but they are not worth the risk in my opinion…</p></li>\n<li><p>Do try out different loss functions, mixing up metrics or multiple losses to provide a better way to generalize can also be a major factor.</p></li>\n<li><p>Too big of an image size can especially lead the model to be biased so be aware of that while trying out big models or big image size.</p></li>\n</ol>\n<p>Please feel free to share your thoughts, I will edit the post to add more points from comments with credit ofc.</p>",
  "messages": [
    {
      "id": "1465467",
      "postDate": "08/11/2021 04:11:22",
      "content": "<p>First I would like to point out that the LB public dataset is too small to consider and the private dataset is not big enough to give us a stable score so even after following CV the private LB can be expected to have a shakeup.</p>\n<ol>\n<li><p>Training in 5-10 Folds can give you a way more stable CV. </p></li>\n<li><p>It is also important to notice that the models overfits quite quickly on train dataset and methods like heavy augmentations, high dropout can be pretty effective.</p></li>\n<li><p>The main goal should be to improve generalization, but look out for possible \"model-specific\" bias in your trainings. Sometimes these bias can pay off but they are not worth the risk in my opinion…</p></li>\n<li><p>Do try out different loss functions, mixing up metrics or multiple losses to provide a better way to generalize can also be a major factor.</p></li>\n<li><p>Too big of an image size can especially lead the model to be biased so be aware of that while trying out big models or big image size.</p></li>\n</ol>\n<p>Please feel free to share your thoughts, I will edit the post to add more points from comments with credit ofc.</p>",
      "rawMarkdown": "First I would like to point out that the LB public dataset is too small to consider and the private dataset is not big enough to give us a stable score so even after following CV the private LB can be expected to have a shakeup.\n\n1. Training in 5-10 Folds can give you a way more stable CV. \n\n2. It is also important to notice that the models overfits quite quickly on train dataset and methods like heavy augmentations, high dropout can be pretty effective.\n\n3. The main goal should be to improve generalization, but look out for possible \"model-specific\" bias in your trainings. Sometimes these bias can pay off but they are not worth the risk in my opinion...\n\n4. Do try out different loss functions, mixing up metrics or multiple losses to provide a better way to generalize can also be a major factor.\n\n5. Too big of an image size can especially lead the model to be biased so be aware of that while trying out big models or big image size.\n\nPlease feel free to share your thoughts, I will edit the post to add more points from comments with credit ofc.",
      "votes": null
    },
    {
      "id": "1466079",
      "postDate": "08/11/2021 09:59:56",
      "content": "<p>couldn't understand your 5th point</p>",
      "rawMarkdown": "couldn't understand your 5th point",
      "votes": null
    },
    {
      "id": "1466355",
      "postDate": "08/11/2021 12:30:46",
      "content": "<p>I have also been struggling with my model overfitting on training data and for some reason reluctant to use high dropouts. <br>\nWould like to know how many epochs on average are people training when using high dropout.</p>",
      "rawMarkdown": "I have also been struggling with my model overfitting on training data and for some reason reluctant to use high dropouts. \nWould like to know how many epochs on average are people training when using high dropout.",
      "votes": null
    },
    {
      "id": "1467128",
      "postDate": "08/11/2021 19:35:32",
      "content": "<p>When the model has too many parameters, it will give up on generalization rather quickly because you only need a fraction of those to make a bias… Similarly if the image size is too big you can see the model not being able to understand the image (try training in incredibly massive image sizes and you will see an immense performance difference in pretty much every dataset and every model)</p>",
      "rawMarkdown": "When the model has too many parameters, it will give up on generalization rather quickly because you only need a fraction of those to make a bias... Similarly if the image size is too big you can see the model not being able to understand the image (try training in incredibly massive image sizes and you will see an immense performance difference in pretty much every dataset and every model)",
      "votes": null
    },
    {
      "id": "1479276",
      "postDate": "08/18/2021 11:31:04",
      "content": "<p>I think there is no need to be afraid of this problem, i am not relying on my public lb anymore as it is totally out of its mind . Rather i am trusting on my cv more and yes just a little on lb. Some participants have also wrote code for generating random submission values and has aced a good score , so there is nothing to worry . Just hope for a good private lb and trust on your cv like 98%.</p>",
      "rawMarkdown": "I think there is no need to be afraid of this problem, i am not relying on my public lb anymore as it is totally out of its mind . Rather i am trusting on my cv more and yes just a little on lb. Some participants have also wrote code for generating random submission values and has aced a good score , so there is nothing to worry . Just hope for a good private lb and trust on your cv like 98%.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1466079,
      "author_name": "mrinath",
      "author_url": "",
      "post_date": "08/11/2021 09:59:56",
      "content": "<p>couldn't understand your 5th point</p>",
      "votes": null,
      "replies": [
        {
          "id": 1467128,
          "author_name": "harshitsheoran",
          "author_url": "",
          "post_date": "08/11/2021 19:35:32",
          "content": "<p>When the model has too many parameters, it will give up on generalization rather quickly because you only need a fraction of those to make a bias… Similarly if the image size is too big you can see the model not being able to understand the image (try training in incredibly massive image sizes and you will see an immense performance difference in pretty much every dataset and every model)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1466355,
      "author_name": "aryamansharma47",
      "author_url": "",
      "post_date": "08/11/2021 12:30:46",
      "content": "<p>I have also been struggling with my model overfitting on training data and for some reason reluctant to use high dropouts. <br>\nWould like to know how many epochs on average are people training when using high dropout.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1479276,
      "author_name": "swaralipibose",
      "author_url": "",
      "post_date": "08/18/2021 11:31:04",
      "content": "<p>I think there is no need to be afraid of this problem, i am not relying on my public lb anymore as it is totally out of its mind . Rather i am trusting on my cv more and yes just a little on lb. Some participants have also wrote code for generating random submission values and has aced a good score , so there is nothing to worry . Just hope for a good private lb and trust on your cv like 98%.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1465467": "First I would like to point out that the LB public dataset is too small to consider and the private dataset is not big enough to give us a stable score so even after following CV the private LB can be expected to have a shakeup.\n\n1. Training in 5-10 Folds can give you a way more stable CV. \n\n2. It is also important to notice that the models overfits quite quickly on train dataset and methods like heavy augmentations, high dropout can be pretty effective.\n\n3. The main goal should be to improve generalization, but look out for possible \"model-specific\" bias in your trainings. Sometimes these bias can pay off but they are not worth the risk in my opinion...\n\n4. Do try out different loss functions, mixing up metrics or multiple losses to provide a better way to generalize can also be a major factor.\n\n5. Too big of an image size can especially lead the model to be biased so be aware of that while trying out big models or big image size.\n\nPlease feel free to share your thoughts, I will edit the post to add more points from comments with credit ofc.",
    "1466079": "couldn't understand your 5th point",
    "1466355": "I have also been struggling with my model overfitting on training data and for some reason reluctant to use high dropouts. \nWould like to know how many epochs on average are people training when using high dropout.",
    "1467128": "When the model has too many parameters, it will give up on generalization rather quickly because you only need a fraction of those to make a bias... Similarly if the image size is too big you can see the model not being able to understand the image (try training in incredibly massive image sizes and you will see an immense performance difference in pretty much every dataset and every model)",
    "1479276": "I think there is no need to be afraid of this problem, i am not relying on my public lb anymore as it is totally out of its mind . Rather i am trusting on my cv more and yes just a little on lb. Some participants have also wrote code for generating random submission values and has aced a good score , so there is nothing to worry . Just hope for a good private lb and trust on your cv like 98%."
  },
  "source": "meta"
}