{
  "id": 255352,
  "title": "[CV vs LB]",
  "url": "/competitions/rsna-miccai-brain-tumor-radiogenomic-classification/discussion/255352",
  "author_name": "",
  "post_date": "2021-07-27T07:21:29.491694Z",
  "votes": 27,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I am really thinking if in this competition should we believe the LB?  For me, my models with lower CVs are performing better in the LB.<br>\nCurrently for my best model CV: 0.69 and LB 0.656 StratifiedKFold 5 folds<br>\nWhat do you guys think, is your CV correlating with LB?</p>",
  "messages": [
    {
      "id": "1401289",
      "postDate": "07/27/2021 07:21:29",
      "content": "<p>I am really thinking if in this competition should we believe the LB?  For me, my models with lower CVs are performing better in the LB.<br>\nCurrently for my best model CV: 0.69 and LB 0.656 StratifiedKFold 5 folds<br>\nWhat do you guys think, is your CV correlating with LB?</p>",
      "rawMarkdown": "I am really thinking if in this competition should we believe the LB?  For me, my models with lower CVs are performing better in the LB.\nCurrently for my best model CV: 0.69 and LB 0.656 StratifiedKFold 5 folds\nWhat do you guys think, is your CV correlating with LB?",
      "votes": null
    },
    {
      "id": "1401370",
      "postDate": "07/27/2021 09:11:19",
      "content": "<p>I am in a same situation. I'm also confused, too. </p>",
      "rawMarkdown": "I am in a same situation. I'm also confused, too.",
      "votes": null
    },
    {
      "id": "1401493",
      "postDate": "07/27/2021 11:16:29",
      "content": "<p>If you don't mind can you share your CV Score</p>",
      "rawMarkdown": "If you don't mind can you share your CV Score",
      "votes": null
    },
    {
      "id": "1401505",
      "postDate": "07/27/2021 11:26:46",
      "content": "<p>I'm also confused. My CV and LB are completely unrelated, and both have huge randomness(0.5~0.7).</p>",
      "rawMarkdown": "I'm also confused. My CV and LB are completely unrelated, and both have huge randomness(0.5~0.7).",
      "votes": null
    },
    {
      "id": "1401625",
      "postDate": "07/27/2021 12:52:12",
      "content": "<p>Do you think test data is very much different than the train data?</p>",
      "rawMarkdown": "Do you think test data is very much different than the train data?",
      "votes": null
    },
    {
      "id": "1401872",
      "postDate": "07/27/2021 15:54:16",
      "content": "<p>The public test data has only 87 samples. So that, the public LB score is unstable. A score of about 0.6 is almost by chance. I check it in <a href=\"https://www.kaggle.com/osciiart/public-lb-simulation?scriptVersionId=69168956\" target=\"_blank\">my notebook</a>.</p>",
      "rawMarkdown": "The public test data has only 87 samples. So that, the public LB score is unstable. A score of about 0.6 is almost by chance. I check it in [my notebook](https://www.kaggle.com/osciiart/public-lb-simulation?scriptVersionId=69168956).",
      "votes": null
    },
    {
      "id": "1401898",
      "postDate": "07/27/2021 16:09:00",
      "content": "<p>Completely agree, limited sample size makes auc score unstable and not a good indicator of model performance.  I won't be surprised to see a big shakeup and the end of this competition </p>",
      "rawMarkdown": "Completely agree, limited sample size makes auc score unstable and not a good indicator of model performance.  I won't be surprised to see a big shakeup and the end of this competition",
      "votes": null
    },
    {
      "id": "1402062",
      "postDate": "07/27/2021 19:47:42",
      "content": "<p>No strong correlation for me. I picked models with the best validation AUC from training across all scan types (3D simple model) and it did not correlate. My 3rd worst scan type model scored highest on the public LB (0.65) (no kfold).</p>\n<p>See <strong><a href=\"https://www.kaggle.com/dschettler8845/eda-3d-baseline-rsna-glioma-radiogenomics#3d_numpy\" target=\"_blank\">this notebook</a></strong> I just created for more details of the training and model structures.</p>\n<p>In general a <strong>CV AUC score of high 0.6xx up to low to mid 0.7xx</strong> made no difference on the public LB. That being said I did notice that certain scan types are more consistent and score higher on LB (no random shakeup). I want to rerun training a few times and benchmark the public LB scores to see how much is just chance and how much is consistent.</p>\n<hr>\n<p>I'm pretty certain a diverse approach combined with a strong local CV dataset will be imperative in this competition. Looking forward to see how things develop going forward…</p>",
      "rawMarkdown": "No strong correlation for me. I picked models with the best validation AUC from training across all scan types (3D simple model) and it did not correlate. My 3rd worst scan type model scored highest on the public LB (0.65) (no kfold).\n\nSee **[this notebook](https://www.kaggle.com/dschettler8845/eda-3d-baseline-rsna-glioma-radiogenomics#3d_numpy)** I just created for more details of the training and model structures.\n\nIn general a **CV AUC score of high 0.6xx up to low to mid 0.7xx** made no difference on the public LB. That being said I did notice that certain scan types are more consistent and score higher on LB (no random shakeup). I want to rerun training a few times and benchmark the public LB scores to see how much is just chance and how much is consistent.\n\n---\n\nI'm pretty certain a diverse approach combined with a strong local CV dataset will be imperative in this competition. Looking forward to see how things develop going forward...",
      "votes": null
    },
    {
      "id": "1402817",
      "postDate": "07/28/2021 15:03:56",
      "content": "<p>As discussed below, the number of data points in this competition is extremely small, so I don't think it is possible to get much correlation. And since the number of private data is only about 5 times as large as the number of public data, so even the private rankings are quite unstable and the competition will have a large shake.</p>\n<p>And, depending on how much expertise there is among the participants, I think the public LB will eventually become completely unrelated to the private rankings (I mean useless public LB). Since the number of data points is only 87, hand labeling will not be that difficult, and I think the AUC will become close to 1.0.</p>\n<p>Therefore, I think it is better to trust the local CV without worrying too much about the public score or ranking.</p>",
      "rawMarkdown": "As discussed below, the number of data points in this competition is extremely small, so I don't think it is possible to get much correlation. And since the number of private data is only about 5 times as large as the number of public data, so even the private rankings are quite unstable and the competition will have a large shake.\n\nAnd, depending on how much expertise there is among the participants, I think the public LB will eventually become completely unrelated to the private rankings (I mean useless public LB). Since the number of data points is only 87, hand labeling will not be that difficult, and I think the AUC will become close to 1.0.\n\nTherefore, I think it is better to trust the local CV without worrying too much about the public score or ranking.",
      "votes": null
    },
    {
      "id": "1403098",
      "postDate": "07/28/2021 19:25:44",
      "content": "<p>Makes sense but taking the nature of the problem into consideration the best approach to achieving best results on evaluating the quality of models and how they would perform in real time would be to use a higher public ratio between 50% - 70%.  else what if after the competition cv fails , public LB fails, private LB fails and we couldn't help improve the usage of MRI in accurately predicting brain tumors. To me that would be a loss. Though I do understand that as enthusiastic data scientists its up to us to find ways to solve this problem irrespective of obstacles. I hope the competition host can reconsider </p>",
      "rawMarkdown": "Makes sense but taking the nature of the problem into consideration the best approach to achieving best results on evaluating the quality of models and how they would perform in real time would be to use a higher public ratio between 50% - 70%.  else what if after the competition cv fails , public LB fails, private LB fails and we couldn't help improve the usage of MRI in accurately predicting brain tumors. To me that would be a loss. Though I do understand that as enthusiastic data scientists its up to us to find ways to solve this problem irrespective of obstacles. I hope the competition host can reconsider",
      "votes": null
    },
    {
      "id": "1408707",
      "postDate": "08/02/2021 15:12:36",
      "content": "<p>no not at all in my case a model which works with 3d cnn , and gets auc of 0.58CV but in the competition 0.604 while other with 0.5 CV but with 0.612 LB  ,0.55 CV and 0.405 LB. Efficient Net creates a huge difference , if you ask me . Efficient Net Models are just amazing here .</p>",
      "rawMarkdown": "no not at all in my case a model which works with 3d cnn , and gets auc of 0.58CV but in the competition 0.604 while other with 0.5 CV but with 0.612 LB  ,0.55 CV and 0.405 LB. Efficient Net creates a huge difference , if you ask me . Efficient Net Models are just amazing here .",
      "votes": null
    },
    {
      "id": "1464618",
      "postDate": "08/10/2021 16:27:48",
      "content": "<p>Any more thoughts anyone about how to address this CV/LB inconsistency?</p>\n<p>Thanks,</p>",
      "rawMarkdown": "Any more thoughts anyone about how to address this CV/LB inconsistency?\n\nThanks,",
      "votes": null
    },
    {
      "id": "1464749",
      "postDate": "08/10/2021 17:33:04",
      "content": "<p>I think what <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> said is completely true, believing in your CV is important in this competition.</p>",
      "rawMarkdown": "I think what @maxwell110 said is completely true, believing in your CV is important in this competition.",
      "votes": null
    },
    {
      "id": "1466469",
      "postDate": "08/11/2021 13:33:14",
      "content": "<p>I'm curious if in this case, it would be better to monitor the loss? My loss seems to hover around ~0.7 assume it's not learning anything.  </p>",
      "rawMarkdown": "I'm curious if in this case, it would be better to monitor the loss? My loss seems to hover around ~0.7 assume it's not learning anything.",
      "votes": null
    },
    {
      "id": "1467892",
      "postDate": "08/12/2021 07:07:59",
      "content": "<p>An idea, train a <a href=\"https://github.com/qubvel/segmentation_models\" target=\"_blank\">segmentation model</a> (which also gives classification score) with previous BraTS data set and use this data set for validation.  Note, in the validation and also inference time, just don't take into account the predicted mast but the predicted class score. </p>\n<p>Or, adopt a <strong>Supervised Contrastive Learning</strong> approach. In the first phase, train a segmentation model as an encoder to learn the vector representation on the previous BraTS data set, and later in the second phase, use the competition data set just to train the top classifier by freezing the segmentation encoder. </p>",
      "rawMarkdown": "An idea, train a [segmentation model](https://github.com/qubvel/segmentation_models) (which also gives classification score) with previous BraTS data set and use this data set for validation.  Note, in the validation and also inference time, just don't take into account the predicted mast but the predicted class score. \n\nOr, adopt a **Supervised Contrastive Learning** approach. In the first phase, train a segmentation model as an encoder to learn the vector representation on the previous BraTS data set, and later in the second phase, use the competition data set just to train the top classifier by freezing the segmentation encoder.",
      "votes": null
    },
    {
      "id": "1467981",
      "postDate": "08/12/2021 07:55:02",
      "content": "<p>was wondering if there is a way to create a mask for segmentation?</p>",
      "rawMarkdown": "was wondering if there is a way to create a mask for segmentation?",
      "votes": null
    },
    {
      "id": "1475658",
      "postDate": "08/16/2021 19:36:08",
      "content": "<p>Stratified KFold (5 folds)<br>\n2D-CNN<br>\nCV / LB<br>\n0.603 / 0.455<br>\n0.620 / 0.583<br>\n0.654 / 0.568<br>\n0.658 / 0.593</p>",
      "rawMarkdown": "Stratified KFold (5 folds)\n2D-CNN\nCV / LB\n0.603 / 0.455\n0.620 / 0.583\n0.654 / 0.568\n0.658 / 0.593",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1401370,
      "author_name": "relurelu",
      "author_url": "",
      "post_date": "07/27/2021 09:11:19",
      "content": "<p>I am in a same situation. I'm also confused, too. </p>",
      "votes": null,
      "replies": [
        {
          "id": 1401493,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "07/27/2021 11:16:29",
          "content": "<p>If you don't mind can you share your CV Score</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1401505,
      "author_name": "zzy990106",
      "author_url": "",
      "post_date": "07/27/2021 11:26:46",
      "content": "<p>I'm also confused. My CV and LB are completely unrelated, and both have huge randomness(0.5~0.7).</p>",
      "votes": null,
      "replies": [
        {
          "id": 1401625,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "07/27/2021 12:52:12",
          "content": "<p>Do you think test data is very much different than the train data?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1401872,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "07/27/2021 15:54:16",
      "content": "<p>The public test data has only 87 samples. So that, the public LB score is unstable. A score of about 0.6 is almost by chance. I check it in <a href=\"https://www.kaggle.com/osciiart/public-lb-simulation?scriptVersionId=69168956\" target=\"_blank\">my notebook</a>.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1401898,
          "author_name": "tivfrvqhs5",
          "author_url": "",
          "post_date": "07/27/2021 16:09:00",
          "content": "<p>Completely agree, limited sample size makes auc score unstable and not a good indicator of model performance.  I won't be surprised to see a big shakeup and the end of this competition </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1402062,
      "author_name": "dschettler8845",
      "author_url": "",
      "post_date": "07/27/2021 19:47:42",
      "content": "<p>No strong correlation for me. I picked models with the best validation AUC from training across all scan types (3D simple model) and it did not correlate. My 3rd worst scan type model scored highest on the public LB (0.65) (no kfold).</p>\n<p>See <strong><a href=\"https://www.kaggle.com/dschettler8845/eda-3d-baseline-rsna-glioma-radiogenomics#3d_numpy\" target=\"_blank\">this notebook</a></strong> I just created for more details of the training and model structures.</p>\n<p>In general a <strong>CV AUC score of high 0.6xx up to low to mid 0.7xx</strong> made no difference on the public LB. That being said I did notice that certain scan types are more consistent and score higher on LB (no random shakeup). I want to rerun training a few times and benchmark the public LB scores to see how much is just chance and how much is consistent.</p>\n<hr>\n<p>I'm pretty certain a diverse approach combined with a strong local CV dataset will be imperative in this competition. Looking forward to see how things develop going forward…</p>",
      "votes": null,
      "replies": [
        {
          "id": 1466469,
          "author_name": "marshath",
          "author_url": "",
          "post_date": "08/11/2021 13:33:14",
          "content": "<p>I'm curious if in this case, it would be better to monitor the loss? My loss seems to hover around ~0.7 assume it's not learning anything.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1402817,
      "author_name": "maxwell110",
      "author_url": "",
      "post_date": "07/28/2021 15:03:56",
      "content": "<p>As discussed below, the number of data points in this competition is extremely small, so I don't think it is possible to get much correlation. And since the number of private data is only about 5 times as large as the number of public data, so even the private rankings are quite unstable and the competition will have a large shake.</p>\n<p>And, depending on how much expertise there is among the participants, I think the public LB will eventually become completely unrelated to the private rankings (I mean useless public LB). Since the number of data points is only 87, hand labeling will not be that difficult, and I think the AUC will become close to 1.0.</p>\n<p>Therefore, I think it is better to trust the local CV without worrying too much about the public score or ranking.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1403098,
          "author_name": "troublem1",
          "author_url": "",
          "post_date": "07/28/2021 19:25:44",
          "content": "<p>Makes sense but taking the nature of the problem into consideration the best approach to achieving best results on evaluating the quality of models and how they would perform in real time would be to use a higher public ratio between 50% - 70%.  else what if after the competition cv fails , public LB fails, private LB fails and we couldn't help improve the usage of MRI in accurately predicting brain tumors. To me that would be a loss. Though I do understand that as enthusiastic data scientists its up to us to find ways to solve this problem irrespective of obstacles. I hope the competition host can reconsider </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1408707,
      "author_name": "swaralipibose",
      "author_url": "",
      "post_date": "08/02/2021 15:12:36",
      "content": "<p>no not at all in my case a model which works with 3d cnn , and gets auc of 0.58CV but in the competition 0.604 while other with 0.5 CV but with 0.612 LB  ,0.55 CV and 0.405 LB. Efficient Net creates a huge difference , if you ask me . Efficient Net Models are just amazing here .</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1464618,
      "author_name": "bopengiowa",
      "author_url": "",
      "post_date": "08/10/2021 16:27:48",
      "content": "<p>Any more thoughts anyone about how to address this CV/LB inconsistency?</p>\n<p>Thanks,</p>",
      "votes": null,
      "replies": [
        {
          "id": 1464749,
          "author_name": "mrinath",
          "author_url": "",
          "post_date": "08/10/2021 17:33:04",
          "content": "<p>I think what <a href=\"https://www.kaggle.com/maxwell110\" target=\"_blank\">@maxwell110</a> said is completely true, believing in your CV is important in this competition.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1467892,
          "author_name": "ipythonx",
          "author_url": "",
          "post_date": "08/12/2021 07:07:59",
          "content": "<p>An idea, train a <a href=\"https://github.com/qubvel/segmentation_models\" target=\"_blank\">segmentation model</a> (which also gives classification score) with previous BraTS data set and use this data set for validation.  Note, in the validation and also inference time, just don't take into account the predicted mast but the predicted class score. </p>\n<p>Or, adopt a <strong>Supervised Contrastive Learning</strong> approach. In the first phase, train a segmentation model as an encoder to learn the vector representation on the previous BraTS data set, and later in the second phase, use the competition data set just to train the top classifier by freezing the segmentation encoder. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1467981,
          "author_name": "marshath",
          "author_url": "",
          "post_date": "08/12/2021 07:55:02",
          "content": "<p>was wondering if there is a way to create a mask for segmentation?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1475658,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "08/16/2021 19:36:08",
      "content": "<p>Stratified KFold (5 folds)<br>\n2D-CNN<br>\nCV / LB<br>\n0.603 / 0.455<br>\n0.620 / 0.583<br>\n0.654 / 0.568<br>\n0.658 / 0.593</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1401289": "I am really thinking if in this competition should we believe the LB?  For me, my models with lower CVs are performing better in the LB.\nCurrently for my best model CV: 0.69 and LB 0.656 StratifiedKFold 5 folds\nWhat do you guys think, is your CV correlating with LB?",
    "1401370": "I am in a same situation. I'm also confused, too.",
    "1401493": "If you don't mind can you share your CV Score",
    "1401505": "I'm also confused. My CV and LB are completely unrelated, and both have huge randomness(0.5~0.7).",
    "1401625": "Do you think test data is very much different than the train data?",
    "1401872": "The public test data has only 87 samples. So that, the public LB score is unstable. A score of about 0.6 is almost by chance. I check it in [my notebook](https://www.kaggle.com/osciiart/public-lb-simulation?scriptVersionId=69168956).",
    "1401898": "Completely agree, limited sample size makes auc score unstable and not a good indicator of model performance.  I won't be surprised to see a big shakeup and the end of this competition",
    "1402062": "No strong correlation for me. I picked models with the best validation AUC from training across all scan types (3D simple model) and it did not correlate. My 3rd worst scan type model scored highest on the public LB (0.65) (no kfold).\n\nSee **[this notebook](https://www.kaggle.com/dschettler8845/eda-3d-baseline-rsna-glioma-radiogenomics#3d_numpy)** I just created for more details of the training and model structures.\n\nIn general a **CV AUC score of high 0.6xx up to low to mid 0.7xx** made no difference on the public LB. That being said I did notice that certain scan types are more consistent and score higher on LB (no random shakeup). I want to rerun training a few times and benchmark the public LB scores to see how much is just chance and how much is consistent.\n\n---\n\nI'm pretty certain a diverse approach combined with a strong local CV dataset will be imperative in this competition. Looking forward to see how things develop going forward...",
    "1402817": "As discussed below, the number of data points in this competition is extremely small, so I don't think it is possible to get much correlation. And since the number of private data is only about 5 times as large as the number of public data, so even the private rankings are quite unstable and the competition will have a large shake.\n\nAnd, depending on how much expertise there is among the participants, I think the public LB will eventually become completely unrelated to the private rankings (I mean useless public LB). Since the number of data points is only 87, hand labeling will not be that difficult, and I think the AUC will become close to 1.0.\n\nTherefore, I think it is better to trust the local CV without worrying too much about the public score or ranking.",
    "1403098": "Makes sense but taking the nature of the problem into consideration the best approach to achieving best results on evaluating the quality of models and how they would perform in real time would be to use a higher public ratio between 50% - 70%.  else what if after the competition cv fails , public LB fails, private LB fails and we couldn't help improve the usage of MRI in accurately predicting brain tumors. To me that would be a loss. Though I do understand that as enthusiastic data scientists its up to us to find ways to solve this problem irrespective of obstacles. I hope the competition host can reconsider",
    "1408707": "no not at all in my case a model which works with 3d cnn , and gets auc of 0.58CV but in the competition 0.604 while other with 0.5 CV but with 0.612 LB  ,0.55 CV and 0.405 LB. Efficient Net creates a huge difference , if you ask me . Efficient Net Models are just amazing here .",
    "1464618": "Any more thoughts anyone about how to address this CV/LB inconsistency?\n\nThanks,",
    "1464749": "I think what @maxwell110 said is completely true, believing in your CV is important in this competition.",
    "1466469": "I'm curious if in this case, it would be better to monitor the loss? My loss seems to hover around ~0.7 assume it's not learning anything.",
    "1467892": "An idea, train a [segmentation model](https://github.com/qubvel/segmentation_models) (which also gives classification score) with previous BraTS data set and use this data set for validation.  Note, in the validation and also inference time, just don't take into account the predicted mast but the predicted class score. \n\nOr, adopt a **Supervised Contrastive Learning** approach. In the first phase, train a segmentation model as an encoder to learn the vector representation on the previous BraTS data set, and later in the second phase, use the competition data set just to train the top classifier by freezing the segmentation encoder.",
    "1467981": "was wondering if there is a way to create a mask for segmentation?",
    "1475658": "Stratified KFold (5 folds)\n2D-CNN\nCV / LB\n0.603 / 0.455\n0.620 / 0.583\n0.654 / 0.568\n0.658 / 0.593"
  },
  "source": "meta"
}