{
  "id": 114938,
  "title": "the gap between CV and LB",
  "url": "/competitions/understanding_cloud_organization/discussion/114938",
  "author_name": "",
  "post_date": "2019-10-30T06:41:30.986721100Z",
  "votes": 5,
  "comment_count": 24,
  "views": 0,
  "content": "<p>I fork this kernel@<a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">Segmentation in PyTorch using convenient tools</a> to sart this competion, I had got pretty good CV score, But LB score didn't improve much.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fcc99499a8a4c45f73c1b7c669050c341%2Fb60f72f6455e81b07f15c4c9f21657c.png?generation=1572416904412103&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fec2d04b17a0670b2170600fd431f2611%2F81627d0c1747f3a4700bc22987a878f.png?generation=1572417084234463&amp;alt=media\" alt=\"\">\nIt seems that these thresholds lead to overfitting.\nSo what's the gap between CV and LB in you model?😳 </p>",
  "messages": [
    {
      "id": "661333",
      "postDate": "10/30/2019 06:41:30",
      "content": "<p>I fork this kernel@<a href=\"https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools\">Segmentation in PyTorch using convenient tools</a> to sart this competion, I had got pretty good CV score, But LB score didn't improve much.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fcc99499a8a4c45f73c1b7c669050c341%2Fb60f72f6455e81b07f15c4c9f21657c.png?generation=1572416904412103&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fec2d04b17a0670b2170600fd431f2611%2F81627d0c1747f3a4700bc22987a878f.png?generation=1572417084234463&amp;alt=media\" alt=\"\">\nIt seems that these thresholds lead to overfitting.\nSo what's the gap between CV and LB in you model?😳 </p>",
      "rawMarkdown": "I fork this kernel@[Segmentation in PyTorch using convenient tools](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) to sart this competion, I had got pretty good CV score, But LB score didn't improve much.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fcc99499a8a4c45f73c1b7c669050c341%2Fb60f72f6455e81b07f15c4c9f21657c.png?generation=1572416904412103&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fec2d04b17a0670b2170600fd431f2611%2F81627d0c1747f3a4700bc22987a878f.png?generation=1572417084234463&amp;alt=media)\nIt seems that these thresholds lead to overfitting.\nSo what's the gap between CV and LB in you model?😳",
      "votes": null
    },
    {
      "id": "661375",
      "postDate": "10/30/2019 08:03:07",
      "content": "<p>If you are using multifold .can you do one thing (Not sure if it's common practice ) .\nWhat I do is , I use Trained model from Fold1 to do prediction of Fold2 Valid_Loader ...</p>",
      "rawMarkdown": "If you are using multifold .can you do one thing (Not sure if it's common practice ) .\nWhat I do is , I use Trained model from Fold1 to do prediction of Fold2 Valid_Loader ...",
      "votes": null
    },
    {
      "id": "661376",
      "postDate": "10/30/2019 08:04:25",
      "content": "<p>..also in hindsight , it's pretty great CV score ..since LB is only a small percentage of Test Set , probably you should trust your CV eh ?</p>",
      "rawMarkdown": "..also in hindsight , it's pretty great CV score ..since LB is only a small percentage of Test Set , probably you should trust your CV eh ?",
      "votes": null
    },
    {
      "id": "661400",
      "postDate": "10/30/2019 08:47:34",
      "content": "<p>I haven't used multifold， maybe it's the time to use it.</p>",
      "rawMarkdown": "I haven't used multifold， maybe it's the time to use it.",
      "votes": null
    },
    {
      "id": "661410",
      "postDate": "10/30/2019 09:02:39",
      "content": "<p>I removed some  small area by the threshold and min_size, What I'm worried about is that using these too early will make me over fit the validation data, so i will also try without any post-processing 🙏 .</p>",
      "rawMarkdown": "I removed some  small area by the threshold and min_size, What I'm worried about is that using these too early will make me over fit the validation data, so i will also try without any post-processing 🙏 .",
      "votes": null
    },
    {
      "id": "661466",
      "postDate": "10/30/2019 10:45:58",
      "content": "<p>Is this using oof data for k folds? It seems your CV is wrong.\nI have a pretty consistent correlation between CV/LB with 0.65x/0.66x</p>",
      "rawMarkdown": "Is this using oof data for k folds? It seems your CV is wrong.\nI have a pretty consistent correlation between CV/LB with 0.65x/0.66x",
      "votes": null
    },
    {
      "id": "661472",
      "postDate": "10/30/2019 10:58:17",
      "content": "<p>Without Post Processing</p>\n\n<p>```\n0\n    threshold   size      dice\n5        0.00  30000  0.628982 \n11       0.05  30000  0.627289\n17       0.10  30000  0.625166\n23       0.15  30000  0.624697\n29       0.20  30000  0.622637\n1\n    threshold   size      dice\n10       0.05  10000  0.743060 \n16       0.10  10000  0.742804\n4        0.00  10000  0.741858\n22       0.15  10000  0.741653\n34       0.25  10000  0.739891\n2\n    threshold   size      dice\n22       0.15  10000  0.614364 \n16       0.10  10000  0.613529\n10       0.05  10000  0.613313\n28       0.20  10000  0.612224\n5        0.00  30000  0.611968\n3\n    threshold  size      dice\n15       0.10  5000  0.601419 \n21       0.15  5000  0.601251\n9        0.05  5000  0.601139\n3        0.00  5000  0.600664\n27       0.20  5000  0.597595</p>\n\n<p>```\nWith ConvexHull and Removing small mask : </p>\n\n<p><code>\n0\n    threshold   size      dice\n23       0.15  30000  0.635331\n29       0.20  30000  0.634952\n35       0.25  30000  0.633626\n41       0.30  30000  0.631460\n11       0.05  30000  0.630965\n1\n    threshold   size      dice\n10       0.05  10000  0.745643\n58       0.45  10000  0.745479\n4        0.00  10000  0.744895\n52       0.40  10000  0.744391\n22       0.15  10000  0.744311\n2\n     threshold   size      dice\n5         0.00  30000  0.623816\n11        0.05  30000  0.621720\n23        0.15  30000  0.619979\n17        0.10  30000  0.618410\n118       0.95  10000  0.617806\n3\n    threshold   size      dice\n21       0.15   5000  0.605091\n4        0.00  10000  0.604875\n10       0.05  10000  0.604359\n16       0.10  10000  0.603477\n9        0.05   5000  0.602316\n</code></p>",
      "rawMarkdown": "Without Post Processing\n\n```\n0\n    threshold   size      dice\n5        0.00  30000  0.628982 \n11       0.05  30000  0.627289\n17       0.10  30000  0.625166\n23       0.15  30000  0.624697\n29       0.20  30000  0.622637\n1\n    threshold   size      dice\n10       0.05  10000  0.743060 \n16       0.10  10000  0.742804\n4        0.00  10000  0.741858\n22       0.15  10000  0.741653\n34       0.25  10000  0.739891\n2\n    threshold   size      dice\n22       0.15  10000  0.614364 \n16       0.10  10000  0.613529\n10       0.05  10000  0.613313\n28       0.20  10000  0.612224\n5        0.00  30000  0.611968\n3\n    threshold  size      dice\n15       0.10  5000  0.601419 \n21       0.15  5000  0.601251\n9        0.05  5000  0.601139\n3        0.00  5000  0.600664\n27       0.20  5000  0.597595\n\n```\nWith ConvexHull and Removing small mask : \n\n```\n0\n    threshold   size      dice\n23       0.15  30000  0.635331\n29       0.20  30000  0.634952\n35       0.25  30000  0.633626\n41       0.30  30000  0.631460\n11       0.05  30000  0.630965\n1\n    threshold   size      dice\n10       0.05  10000  0.745643\n58       0.45  10000  0.745479\n4        0.00  10000  0.744895\n52       0.40  10000  0.744391\n22       0.15  10000  0.744311\n2\n     threshold   size      dice\n5         0.00  30000  0.623816\n11        0.05  30000  0.621720\n23        0.15  30000  0.619979\n17        0.10  30000  0.618410\n118       0.95  10000  0.617806\n3\n    threshold   size      dice\n21       0.15   5000  0.605091\n4        0.00  10000  0.604875\n10       0.05  10000  0.604359\n16       0.10  10000  0.603477\n9        0.05   5000  0.602316\n```",
      "votes": null
    },
    {
      "id": "661481",
      "postDate": "10/30/2019 11:13:15",
      "content": "<p>I didn't use k folds, i use the *train_test_split* to get train and val data. So the CV score may not be convincing.</p>",
      "rawMarkdown": "I didn't use k folds, i use the *train_test_split* to get train and val data. So the CV score may not be convincing.",
      "votes": null
    },
    {
      "id": "661535",
      "postDate": "10/30/2019 12:50:41",
      "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> can you explain what you mean with 'without post processing'? Isn't threshold and minsize search a kind of post processing to remove small masks?</p>",
      "rawMarkdown": "phoenix9032 can you explain what you mean with 'without post processing'? Isn't threshold and minsize search a kind of post processing to remove small masks?",
      "votes": null
    },
    {
      "id": "661541",
      "postDate": "10/30/2019 12:53:16",
      "content": "<p>Agreed . I have used that , i should have been more specific . I have not used convex-hull , when I said without post processing </p>",
      "rawMarkdown": "Agreed . I have used that , i should have been more specific . I have not used convex-hull , when I said without post processing",
      "votes": null
    },
    {
      "id": "661782",
      "postDate": "10/30/2019 17:42:51",
      "content": "<p>My thought is maybe you are splitting based on image rather than on the imageid. I.e. you may be showing it examples in the train and the test from the same image but not the same class</p>",
      "rawMarkdown": "My thought is maybe you are splitting based on image rather than on the imageid. I.e. you may be showing it examples in the train and the test from the same image but not the same class",
      "votes": null
    },
    {
      "id": "662304",
      "postDate": "10/31/2019 11:56:32",
      "content": "<p>My CV and LB have been agreeing well so far. My first model uses bounding boxes and has CV 0.582 and LB 0.611 posted <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a>. My second model uses segmentation and has CV 0.612 and LB 0.628 posted <a href=\"https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\">here</a></p>",
      "rawMarkdown": "My CV and LB have been agreeing well so far. My first model uses bounding boxes and has CV 0.582 and LB 0.611 posted [here][1]. My second model uses segmentation and has CV 0.612 and LB 0.628 posted [here][2]\n\n[1]: https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\n[2]: https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60",
      "votes": null
    },
    {
      "id": "662311",
      "postDate": "10/31/2019 12:12:49",
      "content": "<p>As usual , the code is very neat .  When you say random crop . Am i doing the similar thing as you here ?  I am essentially using a cv2.resize() option to resize my train and valid images . </p>\n\n<p><code>\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.ShiftScaleRotate(\n            scale_limit=0.5,\n            rotate_limit=0,\n            shift_limit=0.1,\n            p=0.5,\n            border_mode=0\n        ),\n        albu.GridDistortion(p=0.5),\n        albu.Resize(320, 640),\n        albu.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225)),\n    ]\n    return albu.Compose(train_transform)\n</code></p>",
      "rawMarkdown": "As usual , the code is very neat .  When you say random crop . Am i doing the similar thing as you here ?  I am essentially using a cv2.resize() option to resize my train and valid images . \n\n```\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.ShiftScaleRotate(\n            scale_limit=0.5,\n            rotate_limit=0,\n            shift_limit=0.1,\n            p=0.5,\n            border_mode=0\n        ),\n        albu.GridDistortion(p=0.5),\n        albu.Resize(320, 640),\n        albu.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225)),\n    ]\n    return albu.Compose(train_transform)\n```",
      "votes": null
    },
    {
      "id": "662333",
      "postDate": "10/31/2019 12:42:39",
      "content": "<p>No. That is not the same. I'm essentially using <code>albu.RandomCrop(320,640)</code> which is different than performing stuff and then using <code>albu.Resize(320,640)</code>. Your network is more often \"seeing\" the global image whereas my network is more often \"seeing\" a localized region within images.</p>",
      "rawMarkdown": "No. That is not the same. I'm essentially using `albu.RandomCrop(320,640)` which is different than performing stuff and then using `albu.Resize(320,640)`. Your network is more often \"seeing\" the global image whereas my network is more often \"seeing\" a localized region within images.",
      "votes": null
    },
    {
      "id": "662358",
      "postDate": "10/31/2019 13:25:11",
      "content": "<p>why do all people report LB score +0.01 bigger than CV? Because of post processing? Isn't that a sign for public LB overfitting?</p>",
      "rawMarkdown": "why do all people report LB score +0.01 bigger than CV? Because of post processing? Isn't that a sign for public LB overfitting?",
      "votes": null
    },
    {
      "id": "662364",
      "postDate": "10/31/2019 13:33:37",
      "content": "<p>No. My CV scores include the result of post processing the OOF. </p>\n\n<p>This difference in CV and LB is a natural occurrence with small training sets. Your OOF are the result of one model's predictions (the one in-fold) and your test predictions are the result of ensembling three model predictions (all folds). </p>",
      "rawMarkdown": "No. My CV scores include the result of post processing the OOF. \n\nThis difference in CV and LB is a natural occurrence with small training sets. Your OOF are the result of one model's predictions (the one in-fold) and your test predictions are the result of ensembling three model predictions (all folds).",
      "votes": null
    },
    {
      "id": "662370",
      "postDate": "10/31/2019 13:37:30",
      "content": "<p>For me CV is calculated after post processing .   With stringent thresholding ,choices of overfitting does remain . Here there is no private dataset . So one can take different samples from test set  and try to visually validate if predictions are way off or not . \nIn my opinion  ,We don't know which 25% of test data is taken to show the score . If you are lucky ,all your good predictions are falling in the remaining 75% and you will get a surprise after result . </p>",
      "rawMarkdown": "For me CV is calculated after post processing .   With stringent thresholding ,choices of overfitting does remain . Here there is no private dataset . So one can take different samples from test set  and try to visually validate if predictions are way off or not . \nIn my opinion  ,We don't know which 25% of test data is taken to show the score . If you are lucky ,all your good predictions are falling in the remaining 75% and you will get a surprise after result .",
      "votes": null
    },
    {
      "id": "662423",
      "postDate": "10/31/2019 14:35:29",
      "content": "<p>I believe the organisers said the train, public LB, and private LB were random splits. </p>\n\n<p>If you are playing around with thresholds, or doing post processing, you should consider using a separate holdout set as the validation, not the public LB. Take 20% of the labelled data and hold it back as a validation of the thresholding. I started the challenge tuning separate thresholds for classes but now am just using the same global threshold I chose on day 1, which is the same threshold I used on Airbus. </p>\n\n<p>Overfitting? Select your final models based on local CV not public LB. Select model(s) where your val score is best. </p>",
      "rawMarkdown": "I believe the organisers said the train, public LB, and private LB were random splits. \n\nIf you are playing around with thresholds, or doing post processing, you should consider using a separate holdout set as the validation, not the public LB. Take 20% of the labelled data and hold it back as a validation of the thresholding. I started the challenge tuning separate thresholds for classes but now am just using the same global threshold I chose on day 1, which is the same threshold I used on Airbus. \n\nOverfitting? Select your final models based on local CV not public LB. Select model(s) where your val score is best.",
      "votes": null
    },
    {
      "id": "662433",
      "postDate": "10/31/2019 14:52:07",
      "content": "<p>thank you. 😄 </p>",
      "rawMarkdown": "thank you. 😄",
      "votes": null
    },
    {
      "id": "662452",
      "postDate": "10/31/2019 15:20:17",
      "content": "<p>One example of the difference is that if you randomly scale Sugar to twice as large then it looks like Gravel. And if you randomly scale Gravel to half it's size, then it looks like Sugar. My crops are randomly cutting rectangles out of the original image which is different than randomly resizing the original image. </p>",
      "rawMarkdown": "One example of the difference is that if you randomly scale Sugar to twice as large then it looks like Gravel. And if you randomly scale Gravel to half it's size, then it looks like Sugar. My crops are randomly cutting rectangles out of the original image which is different than randomly resizing the original image.",
      "votes": null
    },
    {
      "id": "662460",
      "postDate": "10/31/2019 15:25:05",
      "content": "<p>My CV score is calculated after post processing for oof data of 6 folds. That means about 5500 images. I'm also stratifying by number of masks in images, but i'm not sure if it helped.</p>",
      "rawMarkdown": "My CV score is calculated after post processing for oof data of 6 folds. That means about 5500 images. I'm also stratifying by number of masks in images, but i'm not sure if it helped.",
      "votes": null
    },
    {
      "id": "662477",
      "postDate": "10/31/2019 15:48:16",
      "content": "<p>If we are talking about Global Context , then I thought PSP should be able to address it better as promised . I tried it few days back , didn't do extremely well . May be I should try again and confirm.</p>",
      "rawMarkdown": "If we are talking about Global Context , then I thought PSP should be able to address it better as promised . I tried it few days back , didn't do extremely well . May be I should try again and confirm.",
      "votes": null
    },
    {
      "id": "663032",
      "postDate": "11/01/2019 11:10:36",
      "content": "<p>UPDATE: My model three has CV 0.632 and LB 0.644. There is a nice correlation between CV and LB.</p>",
      "rawMarkdown": "UPDATE: My model three has CV 0.632 and LB 0.644. There is a nice correlation between CV and LB.",
      "votes": null
    },
    {
      "id": "663054",
      "postDate": "11/01/2019 12:00:28",
      "content": "<p>The public kernel also shows a best dice score by threshold and size .  Can you please post what was the value of that ? I have surprisingly seen that the final value it shows is lot closer to sugar dice than flower dice . I just want to confirm that it's not a bug from my side .</p>",
      "rawMarkdown": "The public kernel also shows a best dice score by threshold and size .  Can you please post what was the value of that ? I have surprisingly seen that the final value it shows is lot closer to sugar dice than flower dice . I just want to confirm that it's not a bug from my side .",
      "votes": null
    },
    {
      "id": "1023908",
      "postDate": "09/23/2020 14:23:09",
      "content": "<p>hi i am new to the kaggle, can you explain the significant this table? i am trying to find the documents in support of this, but unable to find out? can any one help me to understand the table of \"threshold vs Size vs Dice\" please? </p>",
      "rawMarkdown": "hi i am new to the kaggle, can you explain the significant this table? i am trying to find the documents in support of this, but unable to find out? can any one help me to understand the table of \"threshold vs Size vs Dice\" please?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1023908,
      "author_name": "harimrai",
      "author_url": "",
      "post_date": "09/23/2020 14:23:09",
      "content": "<p>hi i am new to the kaggle, can you explain the significant this table? i am trying to find the documents in support of this, but unable to find out? can any one help me to understand the table of \"threshold vs Size vs Dice\" please? </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 661375,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "10/30/2019 08:03:07",
      "content": "<p>If you are using multifold .can you do one thing (Not sure if it's common practice ) .\nWhat I do is , I use Trained model from Fold1 to do prediction of Fold2 Valid_Loader ...</p>",
      "votes": null,
      "replies": [
        {
          "id": 661400,
          "author_name": "dandingclam",
          "author_url": "",
          "post_date": "10/30/2019 08:47:34",
          "content": "<p>I haven't used multifold， maybe it's the time to use it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 661376,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "10/30/2019 08:04:25",
      "content": "<p>..also in hindsight , it's pretty great CV score ..since LB is only a small percentage of Test Set , probably you should trust your CV eh ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 661410,
          "author_name": "dandingclam",
          "author_url": "",
          "post_date": "10/30/2019 09:02:39",
          "content": "<p>I removed some  small area by the threshold and min_size, What I'm worried about is that using these too early will make me over fit the validation data, so i will also try without any post-processing 🙏 .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 661472,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "10/30/2019 10:58:17",
          "content": "<p>Without Post Processing</p>\n\n<p>```\n0\n    threshold   size      dice\n5        0.00  30000  0.628982 \n11       0.05  30000  0.627289\n17       0.10  30000  0.625166\n23       0.15  30000  0.624697\n29       0.20  30000  0.622637\n1\n    threshold   size      dice\n10       0.05  10000  0.743060 \n16       0.10  10000  0.742804\n4        0.00  10000  0.741858\n22       0.15  10000  0.741653\n34       0.25  10000  0.739891\n2\n    threshold   size      dice\n22       0.15  10000  0.614364 \n16       0.10  10000  0.613529\n10       0.05  10000  0.613313\n28       0.20  10000  0.612224\n5        0.00  30000  0.611968\n3\n    threshold  size      dice\n15       0.10  5000  0.601419 \n21       0.15  5000  0.601251\n9        0.05  5000  0.601139\n3        0.00  5000  0.600664\n27       0.20  5000  0.597595</p>\n\n<p>```\nWith ConvexHull and Removing small mask : </p>\n\n<p><code>\n0\n    threshold   size      dice\n23       0.15  30000  0.635331\n29       0.20  30000  0.634952\n35       0.25  30000  0.633626\n41       0.30  30000  0.631460\n11       0.05  30000  0.630965\n1\n    threshold   size      dice\n10       0.05  10000  0.745643\n58       0.45  10000  0.745479\n4        0.00  10000  0.744895\n52       0.40  10000  0.744391\n22       0.15  10000  0.744311\n2\n     threshold   size      dice\n5         0.00  30000  0.623816\n11        0.05  30000  0.621720\n23        0.15  30000  0.619979\n17        0.10  30000  0.618410\n118       0.95  10000  0.617806\n3\n    threshold   size      dice\n21       0.15   5000  0.605091\n4        0.00  10000  0.604875\n10       0.05  10000  0.604359\n16       0.10  10000  0.603477\n9        0.05   5000  0.602316\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 661535,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "10/30/2019 12:50:41",
          "content": "<p><a href=\"/phoenix9032\">@phoenix9032</a> can you explain what you mean with 'without post processing'? Isn't threshold and minsize search a kind of post processing to remove small masks?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 661541,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "10/30/2019 12:53:16",
          "content": "<p>Agreed . I have used that , i should have been more specific . I have not used convex-hull , when I said without post processing </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 661466,
      "author_name": "igormunizims",
      "author_url": "",
      "post_date": "10/30/2019 10:45:58",
      "content": "<p>Is this using oof data for k folds? It seems your CV is wrong.\nI have a pretty consistent correlation between CV/LB with 0.65x/0.66x</p>",
      "votes": null,
      "replies": [
        {
          "id": 661481,
          "author_name": "dandingclam",
          "author_url": "",
          "post_date": "10/30/2019 11:13:15",
          "content": "<p>I didn't use k folds, i use the *train_test_split* to get train and val data. So the CV score may not be convincing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662358,
          "author_name": "tugstugi",
          "author_url": "",
          "post_date": "10/31/2019 13:25:11",
          "content": "<p>why do all people report LB score +0.01 bigger than CV? Because of post processing? Isn't that a sign for public LB overfitting?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662364,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 13:33:37",
          "content": "<p>No. My CV scores include the result of post processing the OOF. </p>\n\n<p>This difference in CV and LB is a natural occurrence with small training sets. Your OOF are the result of one model's predictions (the one in-fold) and your test predictions are the result of ensembling three model predictions (all folds). </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662370,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "10/31/2019 13:37:30",
          "content": "<p>For me CV is calculated after post processing .   With stringent thresholding ,choices of overfitting does remain . Here there is no private dataset . So one can take different samples from test set  and try to visually validate if predictions are way off or not . \nIn my opinion  ,We don't know which 25% of test data is taken to show the score . If you are lucky ,all your good predictions are falling in the remaining 75% and you will get a surprise after result . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662423,
          "author_name": "robga",
          "author_url": "",
          "post_date": "10/31/2019 14:35:29",
          "content": "<p>I believe the organisers said the train, public LB, and private LB were random splits. </p>\n\n<p>If you are playing around with thresholds, or doing post processing, you should consider using a separate holdout set as the validation, not the public LB. Take 20% of the labelled data and hold it back as a validation of the thresholding. I started the challenge tuning separate thresholds for classes but now am just using the same global threshold I chose on day 1, which is the same threshold I used on Airbus. </p>\n\n<p>Overfitting? Select your final models based on local CV not public LB. Select model(s) where your val score is best. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662433,
          "author_name": "dandingclam",
          "author_url": "",
          "post_date": "10/31/2019 14:52:07",
          "content": "<p>thank you. 😄 </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662460,
          "author_name": "igormunizims",
          "author_url": "",
          "post_date": "10/31/2019 15:25:05",
          "content": "<p>My CV score is calculated after post processing for oof data of 6 folds. That means about 5500 images. I'm also stratifying by number of masks in images, but i'm not sure if it helped.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 661782,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "10/30/2019 17:42:51",
      "content": "<p>My thought is maybe you are splitting based on image rather than on the imageid. I.e. you may be showing it examples in the train and the test from the same image but not the same class</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 662304,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "10/31/2019 11:56:32",
      "content": "<p>My CV and LB have been agreeing well so far. My first model uses bounding boxes and has CV 0.582 and LB 0.611 posted <a href=\"https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\">here</a>. My second model uses segmentation and has CV 0.612 and LB 0.628 posted <a href=\"https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60\">here</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 662311,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "10/31/2019 12:12:49",
          "content": "<p>As usual , the code is very neat .  When you say random crop . Am i doing the similar thing as you here ?  I am essentially using a cv2.resize() option to resize my train and valid images . </p>\n\n<p><code>\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.ShiftScaleRotate(\n            scale_limit=0.5,\n            rotate_limit=0,\n            shift_limit=0.1,\n            p=0.5,\n            border_mode=0\n        ),\n        albu.GridDistortion(p=0.5),\n        albu.Resize(320, 640),\n        albu.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225)),\n    ]\n    return albu.Compose(train_transform)\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662333,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 12:42:39",
          "content": "<p>No. That is not the same. I'm essentially using <code>albu.RandomCrop(320,640)</code> which is different than performing stuff and then using <code>albu.Resize(320,640)</code>. Your network is more often \"seeing\" the global image whereas my network is more often \"seeing\" a localized region within images.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662452,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "10/31/2019 15:20:17",
          "content": "<p>One example of the difference is that if you randomly scale Sugar to twice as large then it looks like Gravel. And if you randomly scale Gravel to half it's size, then it looks like Sugar. My crops are randomly cutting rectangles out of the original image which is different than randomly resizing the original image. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 662477,
          "author_name": "phoenix9032",
          "author_url": "",
          "post_date": "10/31/2019 15:48:16",
          "content": "<p>If we are talking about Global Context , then I thought PSP should be able to address it better as promised . I tried it few days back , didn't do extremely well . May be I should try again and confirm.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 663032,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "11/01/2019 11:10:36",
      "content": "<p>UPDATE: My model three has CV 0.632 and LB 0.644. There is a nice correlation between CV and LB.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 663054,
      "author_name": "phoenix9032",
      "author_url": "",
      "post_date": "11/01/2019 12:00:28",
      "content": "<p>The public kernel also shows a best dice score by threshold and size .  Can you please post what was the value of that ? I have surprisingly seen that the final value it shows is lot closer to sugar dice than flower dice . I just want to confirm that it's not a bug from my side .</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "661333": "I fork this kernel@[Segmentation in PyTorch using convenient tools](https://www.kaggle.com/artgor/segmentation-in-pytorch-using-convenient-tools) to sart this competion, I had got pretty good CV score, But LB score didn't improve much.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fcc99499a8a4c45f73c1b7c669050c341%2Fb60f72f6455e81b07f15c4c9f21657c.png?generation=1572416904412103&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3930531%2Fec2d04b17a0670b2170600fd431f2611%2F81627d0c1747f3a4700bc22987a878f.png?generation=1572417084234463&amp;alt=media)\nIt seems that these thresholds lead to overfitting.\nSo what's the gap between CV and LB in you model?😳",
    "661375": "If you are using multifold .can you do one thing (Not sure if it's common practice ) .\nWhat I do is , I use Trained model from Fold1 to do prediction of Fold2 Valid_Loader ...",
    "661376": "..also in hindsight , it's pretty great CV score ..since LB is only a small percentage of Test Set , probably you should trust your CV eh ?",
    "661400": "I haven't used multifold， maybe it's the time to use it.",
    "661410": "I removed some  small area by the threshold and min_size, What I'm worried about is that using these too early will make me over fit the validation data, so i will also try without any post-processing 🙏 .",
    "661466": "Is this using oof data for k folds? It seems your CV is wrong.\nI have a pretty consistent correlation between CV/LB with 0.65x/0.66x",
    "661472": "Without Post Processing\n\n```\n0\n    threshold   size      dice\n5        0.00  30000  0.628982 \n11       0.05  30000  0.627289\n17       0.10  30000  0.625166\n23       0.15  30000  0.624697\n29       0.20  30000  0.622637\n1\n    threshold   size      dice\n10       0.05  10000  0.743060 \n16       0.10  10000  0.742804\n4        0.00  10000  0.741858\n22       0.15  10000  0.741653\n34       0.25  10000  0.739891\n2\n    threshold   size      dice\n22       0.15  10000  0.614364 \n16       0.10  10000  0.613529\n10       0.05  10000  0.613313\n28       0.20  10000  0.612224\n5        0.00  30000  0.611968\n3\n    threshold  size      dice\n15       0.10  5000  0.601419 \n21       0.15  5000  0.601251\n9        0.05  5000  0.601139\n3        0.00  5000  0.600664\n27       0.20  5000  0.597595\n\n```\nWith ConvexHull and Removing small mask : \n\n```\n0\n    threshold   size      dice\n23       0.15  30000  0.635331\n29       0.20  30000  0.634952\n35       0.25  30000  0.633626\n41       0.30  30000  0.631460\n11       0.05  30000  0.630965\n1\n    threshold   size      dice\n10       0.05  10000  0.745643\n58       0.45  10000  0.745479\n4        0.00  10000  0.744895\n52       0.40  10000  0.744391\n22       0.15  10000  0.744311\n2\n     threshold   size      dice\n5         0.00  30000  0.623816\n11        0.05  30000  0.621720\n23        0.15  30000  0.619979\n17        0.10  30000  0.618410\n118       0.95  10000  0.617806\n3\n    threshold   size      dice\n21       0.15   5000  0.605091\n4        0.00  10000  0.604875\n10       0.05  10000  0.604359\n16       0.10  10000  0.603477\n9        0.05   5000  0.602316\n```",
    "661481": "I didn't use k folds, i use the *train_test_split* to get train and val data. So the CV score may not be convincing.",
    "661535": "phoenix9032 can you explain what you mean with 'without post processing'? Isn't threshold and minsize search a kind of post processing to remove small masks?",
    "661541": "Agreed . I have used that , i should have been more specific . I have not used convex-hull , when I said without post processing",
    "661782": "My thought is maybe you are splitting based on image rather than on the imageid. I.e. you may be showing it examples in the train and the test from the same image but not the same class",
    "662304": "My CV and LB have been agreeing well so far. My first model uses bounding boxes and has CV 0.582 and LB 0.611 posted [here][1]. My second model uses segmentation and has CV 0.612 and LB 0.628 posted [here][2]\n\n[1]: https://www.kaggle.com/cdeotte/cloud-bounding-boxes-cv-0-58\n[2]: https://www.kaggle.com/cdeotte/train-with-crops-cv-0-60",
    "662311": "As usual , the code is very neat .  When you say random crop . Am i doing the similar thing as you here ?  I am essentially using a cv2.resize() option to resize my train and valid images . \n\n```\ndef get_training_augmentation():\n    train_transform = [\n        albu.HorizontalFlip(p=0.5),\n        albu.ShiftScaleRotate(\n            scale_limit=0.5,\n            rotate_limit=0,\n            shift_limit=0.1,\n            p=0.5,\n            border_mode=0\n        ),\n        albu.GridDistortion(p=0.5),\n        albu.Resize(320, 640),\n        albu.Normalize(mean=(0.485, 0.456, 0.406), std=(0.229, 0.224, 0.225)),\n    ]\n    return albu.Compose(train_transform)\n```",
    "662333": "No. That is not the same. I'm essentially using `albu.RandomCrop(320,640)` which is different than performing stuff and then using `albu.Resize(320,640)`. Your network is more often \"seeing\" the global image whereas my network is more often \"seeing\" a localized region within images.",
    "662358": "why do all people report LB score +0.01 bigger than CV? Because of post processing? Isn't that a sign for public LB overfitting?",
    "662364": "No. My CV scores include the result of post processing the OOF. \n\nThis difference in CV and LB is a natural occurrence with small training sets. Your OOF are the result of one model's predictions (the one in-fold) and your test predictions are the result of ensembling three model predictions (all folds).",
    "662370": "For me CV is calculated after post processing .   With stringent thresholding ,choices of overfitting does remain . Here there is no private dataset . So one can take different samples from test set  and try to visually validate if predictions are way off or not . \nIn my opinion  ,We don't know which 25% of test data is taken to show the score . If you are lucky ,all your good predictions are falling in the remaining 75% and you will get a surprise after result .",
    "662423": "I believe the organisers said the train, public LB, and private LB were random splits. \n\nIf you are playing around with thresholds, or doing post processing, you should consider using a separate holdout set as the validation, not the public LB. Take 20% of the labelled data and hold it back as a validation of the thresholding. I started the challenge tuning separate thresholds for classes but now am just using the same global threshold I chose on day 1, which is the same threshold I used on Airbus. \n\nOverfitting? Select your final models based on local CV not public LB. Select model(s) where your val score is best.",
    "662433": "thank you. 😄",
    "662452": "One example of the difference is that if you randomly scale Sugar to twice as large then it looks like Gravel. And if you randomly scale Gravel to half it's size, then it looks like Sugar. My crops are randomly cutting rectangles out of the original image which is different than randomly resizing the original image.",
    "662460": "My CV score is calculated after post processing for oof data of 6 folds. That means about 5500 images. I'm also stratifying by number of masks in images, but i'm not sure if it helped.",
    "662477": "If we are talking about Global Context , then I thought PSP should be able to address it better as promised . I tried it few days back , didn't do extremely well . May be I should try again and confirm.",
    "663032": "UPDATE: My model three has CV 0.632 and LB 0.644. There is a nice correlation between CV and LB.",
    "663054": "The public kernel also shows a best dice score by threshold and size .  Can you please post what was the value of that ? I have surprisingly seen that the final value it shows is lot closer to sugar dice than flower dice . I just want to confirm that it's not a bug from my side .",
    "1023908": "hi i am new to the kaggle, can you explain the significant this table? i am trying to find the documents in support of this, but unable to find out? can any one help me to understand the table of \"threshold vs Size vs Dice\" please?"
  },
  "source": "meta"
}