{
  "id": 231795,
  "title": "big gap between CV and LB, please help",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/231795",
  "author_name": "",
  "post_date": "2021-04-10T11:27:02.625333Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I have big gap between metrics on my validation dataset and private dataset:<br>\nCV: 0.91<br>\nLB: 0.85<br>\nI don't know why? What the reason? Please help.<br>\nMy ideas:</p>\n<ul>\n<li>glomeruli of another type more frequent in the private dataset</li>\n<li>pale or blured images in the private dataset</li>\n<li>some images in the private dataset not in RGB colorsapce</li>\n<li>problems with image reading in the private dataset</li>\n</ul>",
  "messages": [
    {
      "id": "1269269",
      "postDate": "04/10/2021 11:27:02",
      "content": "<p>I have big gap between metrics on my validation dataset and private dataset:<br>\nCV: 0.91<br>\nLB: 0.85<br>\nI don't know why? What the reason? Please help.<br>\nMy ideas:</p>\n<ul>\n<li>glomeruli of another type more frequent in the private dataset</li>\n<li>pale or blured images in the private dataset</li>\n<li>some images in the private dataset not in RGB colorsapce</li>\n<li>problems with image reading in the private dataset</li>\n</ul>",
      "rawMarkdown": "I have big gap between metrics on my validation dataset and private dataset:\nCV: 0.91\nLB: 0.85\nI don't know why? What the reason? Please help.\nMy ideas:\n- glomeruli of another type more frequent in the private dataset\n- pale or blured images in the private dataset\n- some images in the private dataset not in RGB colorsapce\n- problems with image reading in the private dataset",
      "votes": null
    },
    {
      "id": "1269455",
      "postDate": "04/10/2021 14:46:11",
      "content": "<p>are you using any augementations ? if not then model is probably overfitting , I trained effnetb7 ns encoder with and without augs , without augs :- cv :- 0.93x , lb :- 0.89x , with low augs :- cv :- 0.92x  , lb :- 0.91x , initial learning rate :- 1e-4 ,if i started with 1e-3 the model was converging very fast to 0.9x validation dice score within 3-4 epochs but after that my model was improving inconsistently , training started with 1e-4 my model was able to converge to 0.9x in 7-8 epochs and the validation score after every epochs was constantly increasing , although i trained for 40 epochs , training longer /  using psuedolabels / ensembling different model architectures / post processing  may be the key to success in this comp . also i found out that using large batch sizes improves my model . </p>",
      "rawMarkdown": "are you using any augementations ? if not then model is probably overfitting , I trained effnetb7 ns encoder with and without augs , without augs :- cv :- 0.93x , lb :- 0.89x , with low augs :- cv :- 0.92x  , lb :- 0.91x , initial learning rate :- 1e-4 ,if i started with 1e-3 the model was converging very fast to 0.9x validation dice score within 3-4 epochs but after that my model was improving inconsistently , training started with 1e-4 my model was able to converge to 0.9x in 7-8 epochs and the validation score after every epochs was constantly increasing , although i trained for 40 epochs , training longer /  using psuedolabels / ensembling different model architectures / post processing  may be the key to success in this comp . also i found out that using large batch sizes improves my model .",
      "votes": null
    },
    {
      "id": "1269978",
      "postDate": "04/11/2021 06:32:09",
      "content": "<p>Thank you for the answer. I do use the augmentation: random rotation, scale, contrast, gamma, hue shift. Should I add a blur? Anyway my gap is match bigger than your: 0.06 vs 0.01. How do you read images with rasterio crop by crop?</p>",
      "rawMarkdown": "Thank you for the answer. I do use the augmentation: random rotation, scale, contrast, gamma, hue shift. Should I add a blur? Anyway my gap is match bigger than your: 0.06 vs 0.01. How do you read images with rasterio crop by crop?",
      "votes": null
    },
    {
      "id": "1270057",
      "postDate": "04/11/2021 08:29:20",
      "content": "<p>I just increased my augmentation and my LB score decreased to 0.811</p>",
      "rawMarkdown": "I just increased my augmentation and my LB score decreased to 0.811",
      "votes": null
    },
    {
      "id": "1270074",
      "postDate": "04/11/2021 08:59:10",
      "content": "<p>my update :- single model effnetb7ns 3 fold cv : 0.9213 lb :- 0.920 , btw i am using the trick that <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> has shared here <a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a> to generate images . </p>",
      "rawMarkdown": "my update :- single model effnetb7ns 3 fold cv : 0.9213 lb :- 0.920 , btw i am using the trick that @iafoss has shared here https://www.kaggle.com/iafoss/256x256-images to generate images .",
      "votes": null
    },
    {
      "id": "1270136",
      "postDate": "04/11/2021 10:30:08",
      "content": "<p>I don't have folds, so my CV is not pricise. I see in the 256x256 code:<br>\nhsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)<br>\nIs rasterio return BGR?</p>",
      "rawMarkdown": "I don't have folds, so my CV is not pricise. I see in the 256x256 code:\nhsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\nIs rasterio return BGR?",
      "votes": null
    },
    {
      "id": "1270424",
      "postDate": "04/11/2021 16:39:20",
      "content": "<p>I think it returns RGB, but BGR2HSV and RGB2HSV give the same results</p>",
      "rawMarkdown": "I think it returns RGB, but BGR2HSV and RGB2HSV give the same results",
      "votes": null
    },
    {
      "id": "1271209",
      "postDate": "04/12/2021 12:00:48",
      "content": "<p>You can try augmentations. We used more than 20 types of methods to augment the datasets. And you can add dropout layer in your model.😄</p>",
      "rawMarkdown": "You can try augmentations. We used more than 20 types of methods to augment the datasets. And you can add dropout layer in your model.😄",
      "votes": null
    },
    {
      "id": "1286909",
      "postDate": "04/28/2021 14:20:13",
      "content": "<p>I am using the segmentation_models library popular here. But how to add dropout layer to backbone of the Unet model?</p>",
      "rawMarkdown": "I am using the segmentation_models library popular here. But how to add dropout layer to backbone of the Unet model?",
      "votes": null
    },
    {
      "id": "1287431",
      "postDate": "04/29/2021 04:42:45",
      "content": "<p>Usually add dropout layer behind the backbone just the CNN</p>",
      "rawMarkdown": "Usually add dropout layer behind the backbone just the CNN",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1269455,
      "author_name": "trooperog",
      "author_url": "",
      "post_date": "04/10/2021 14:46:11",
      "content": "<p>are you using any augementations ? if not then model is probably overfitting , I trained effnetb7 ns encoder with and without augs , without augs :- cv :- 0.93x , lb :- 0.89x , with low augs :- cv :- 0.92x  , lb :- 0.91x , initial learning rate :- 1e-4 ,if i started with 1e-3 the model was converging very fast to 0.9x validation dice score within 3-4 epochs but after that my model was improving inconsistently , training started with 1e-4 my model was able to converge to 0.9x in 7-8 epochs and the validation score after every epochs was constantly increasing , although i trained for 40 epochs , training longer /  using psuedolabels / ensembling different model architectures / post processing  may be the key to success in this comp . also i found out that using large batch sizes improves my model . </p>",
      "votes": null,
      "replies": [
        {
          "id": 1269978,
          "author_name": "vedenev",
          "author_url": "",
          "post_date": "04/11/2021 06:32:09",
          "content": "<p>Thank you for the answer. I do use the augmentation: random rotation, scale, contrast, gamma, hue shift. Should I add a blur? Anyway my gap is match bigger than your: 0.06 vs 0.01. How do you read images with rasterio crop by crop?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270057,
          "author_name": "vedenev",
          "author_url": "",
          "post_date": "04/11/2021 08:29:20",
          "content": "<p>I just increased my augmentation and my LB score decreased to 0.811</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270074,
          "author_name": "trooperog",
          "author_url": "",
          "post_date": "04/11/2021 08:59:10",
          "content": "<p>my update :- single model effnetb7ns 3 fold cv : 0.9213 lb :- 0.920 , btw i am using the trick that <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> has shared here <a href=\"https://www.kaggle.com/iafoss/256x256-images\" target=\"_blank\">https://www.kaggle.com/iafoss/256x256-images</a> to generate images . </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270136,
          "author_name": "vedenev",
          "author_url": "",
          "post_date": "04/11/2021 10:30:08",
          "content": "<p>I don't have folds, so my CV is not pricise. I see in the 256x256 code:<br>\nhsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)<br>\nIs rasterio return BGR?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1270424,
          "author_name": "vkendo",
          "author_url": "",
          "post_date": "04/11/2021 16:39:20",
          "content": "<p>I think it returns RGB, but BGR2HSV and RGB2HSV give the same results</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1271209,
      "author_name": "jasonhuangcn",
      "author_url": "",
      "post_date": "04/12/2021 12:00:48",
      "content": "<p>You can try augmentations. We used more than 20 types of methods to augment the datasets. And you can add dropout layer in your model.😄</p>",
      "votes": null,
      "replies": [
        {
          "id": 1286909,
          "author_name": "vsharkeev",
          "author_url": "",
          "post_date": "04/28/2021 14:20:13",
          "content": "<p>I am using the segmentation_models library popular here. But how to add dropout layer to backbone of the Unet model?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1287431,
          "author_name": "jasonhuangcn",
          "author_url": "",
          "post_date": "04/29/2021 04:42:45",
          "content": "<p>Usually add dropout layer behind the backbone just the CNN</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1269269": "I have big gap between metrics on my validation dataset and private dataset:\nCV: 0.91\nLB: 0.85\nI don't know why? What the reason? Please help.\nMy ideas:\n- glomeruli of another type more frequent in the private dataset\n- pale or blured images in the private dataset\n- some images in the private dataset not in RGB colorsapce\n- problems with image reading in the private dataset",
    "1269455": "are you using any augementations ? if not then model is probably overfitting , I trained effnetb7 ns encoder with and without augs , without augs :- cv :- 0.93x , lb :- 0.89x , with low augs :- cv :- 0.92x  , lb :- 0.91x , initial learning rate :- 1e-4 ,if i started with 1e-3 the model was converging very fast to 0.9x validation dice score within 3-4 epochs but after that my model was improving inconsistently , training started with 1e-4 my model was able to converge to 0.9x in 7-8 epochs and the validation score after every epochs was constantly increasing , although i trained for 40 epochs , training longer /  using psuedolabels / ensembling different model architectures / post processing  may be the key to success in this comp . also i found out that using large batch sizes improves my model .",
    "1269978": "Thank you for the answer. I do use the augmentation: random rotation, scale, contrast, gamma, hue shift. Should I add a blur? Anyway my gap is match bigger than your: 0.06 vs 0.01. How do you read images with rasterio crop by crop?",
    "1270057": "I just increased my augmentation and my LB score decreased to 0.811",
    "1270074": "my update :- single model effnetb7ns 3 fold cv : 0.9213 lb :- 0.920 , btw i am using the trick that @iafoss has shared here https://www.kaggle.com/iafoss/256x256-images to generate images .",
    "1270136": "I don't have folds, so my CV is not pricise. I see in the 256x256 code:\nhsv = cv2.cvtColor(img, cv2.COLOR_BGR2HSV)\nIs rasterio return BGR?",
    "1270424": "I think it returns RGB, but BGR2HSV and RGB2HSV give the same results",
    "1271209": "You can try augmentations. We used more than 20 types of methods to augment the datasets. And you can add dropout layer in your model.😄",
    "1286909": "I am using the segmentation_models library popular here. But how to add dropout layer to backbone of the Unet model?",
    "1287431": "Usually add dropout layer behind the backbone just the CNN"
  },
  "source": "meta"
}