{
  "id": 105606,
  "title": "Correlation LB and CV scores",
  "url": "/competitions/recursion-cellular-image-classification/discussion/105606",
  "author_name": "",
  "post_date": "2019-08-24T16:00:06.371100300Z",
  "votes": 7,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Nobody seems to talk about correlation between LB and CV scores anymore.</p>\n\n<p>I'm using DenseNet121 (PyTorch) pretrained on ImageNet with Adam and no schedulers. Simple classification using cross-entropy. I'm not doing CV but I just split into training and test sets using Scikit-learn's <code>train_test_split</code>: I tried to use either the <code>sirna</code> or <code>experiment</code> fields for stratification but it does not seem to matter that much. I'm now facing the issue where my test accuracy is much higher than the LB score. Below you'll find some relevant figures: I augmented the dataset with random horizontal flipping, resize to 256 and random crop to 224 (with Albumentation). I also use both sites for prediction (averaging).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F80a173ddedb794addc774177c1b081a4%2Fkaggle.png?generation=1566662146285345&amp;alt=media\" alt=\"\"></p>\n\n<p>The corresponding LB score is 0.197. Is anybody facing similar issues? Now that the LB scores are much higher than a few weeks ago, can somebody report some example scores between CV and LB? I'm currently not doing any regularization since I read some papers saying that augmentation should suffice in general.</p>",
  "messages": [
    {
      "id": "607104",
      "postDate": "08/24/2019 16:00:06",
      "content": "<p>Nobody seems to talk about correlation between LB and CV scores anymore.</p>\n\n<p>I'm using DenseNet121 (PyTorch) pretrained on ImageNet with Adam and no schedulers. Simple classification using cross-entropy. I'm not doing CV but I just split into training and test sets using Scikit-learn's <code>train_test_split</code>: I tried to use either the <code>sirna</code> or <code>experiment</code> fields for stratification but it does not seem to matter that much. I'm now facing the issue where my test accuracy is much higher than the LB score. Below you'll find some relevant figures: I augmented the dataset with random horizontal flipping, resize to 256 and random crop to 224 (with Albumentation). I also use both sites for prediction (averaging).</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F80a173ddedb794addc774177c1b081a4%2Fkaggle.png?generation=1566662146285345&amp;alt=media\" alt=\"\"></p>\n\n<p>The corresponding LB score is 0.197. Is anybody facing similar issues? Now that the LB scores are much higher than a few weeks ago, can somebody report some example scores between CV and LB? I'm currently not doing any regularization since I read some papers saying that augmentation should suffice in general.</p>",
      "rawMarkdown": "Nobody seems to talk about correlation between LB and CV scores anymore.\n\nI'm using DenseNet121 (PyTorch) pretrained on ImageNet with Adam and no schedulers. Simple classification using cross-entropy. I'm not doing CV but I just split into training and test sets using Scikit-learn's `train_test_split`: I tried to use either the `sirna` or `experiment` fields for stratification but it does not seem to matter that much. I'm now facing the issue where my test accuracy is much higher than the LB score. Below you'll find some relevant figures: I augmented the dataset with random horizontal flipping, resize to 256 and random crop to 224 (with Albumentation). I also use both sites for prediction (averaging).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F80a173ddedb794addc774177c1b081a4%2Fkaggle.png?generation=1566662146285345&amp;alt=media)\n\nThe corresponding LB score is 0.197. Is anybody facing similar issues? Now that the LB scores are much higher than a few weeks ago, can somebody report some example scores between CV and LB? I'm currently not doing any regularization since I read some papers saying that augmentation should suffice in general.",
      "votes": null
    },
    {
      "id": "607122",
      "postDate": "08/24/2019 16:26:29",
      "content": "<p>One way to explain your CV vs LB divergence is the difference between experiments. You can see in pixel_stats.csv that the experiments are quite different in brightness. If you don't account for that and train your model on one mode of brightness it may perform poorly on the other mode of brightness.</p>",
      "rawMarkdown": "One way to explain your CV vs LB divergence is the difference between experiments. You can see in pixel_stats.csv that the experiments are quite different in brightness. If you don't account for that and train your model on one mode of brightness it may perform poorly on the other mode of brightness.",
      "votes": null
    },
    {
      "id": "607124",
      "postDate": "08/24/2019 16:31:49",
      "content": "<p>Interesting. But if I stratify by experiment, shouldn't the model be able to handle it?</p>",
      "rawMarkdown": "Interesting. But if I stratify by experiment, shouldn't the model be able to handle it?",
      "votes": null
    },
    {
      "id": "607136",
      "postDate": "08/24/2019 16:42:44",
      "content": "<p>Not if your validation set shares any experiments with the training set, it will just learn stuff about those specific experiments. That is why your val act is high. </p>\n\n<p>The test set explicitly comes from experiments (batches) that are not in the training set. The goal of this competition is to figure out a way to identify and control for arbitrary batch, really instrumentation, induced effects to be able to cleanly see effects caused by the biological manipulation alone - in this case treatment with a bunch of different sirnas. </p>\n\n<p>To make validation make sense you must move entire experiments out of training and into valid. </p>",
      "rawMarkdown": "Not if your validation set shares any experiments with the training set, it will just learn stuff about those specific experiments. That is why your val act is high. \n\nThe test set explicitly comes from experiments (batches) that are not in the training set. The goal of this competition is to figure out a way to identify and control for arbitrary batch, really instrumentation, induced effects to be able to cleanly see effects caused by the biological manipulation alone - in this case treatment with a bunch of different sirnas. \n\nTo make validation make sense you must move entire experiments out of training and into valid.",
      "votes": null
    },
    {
      "id": "607138",
      "postDate": "08/24/2019 16:49:15",
      "content": "<p>Okay, that makes sense now. If I simply do that kind of split, my test accuracy will decrease.\nSo I guess normalizing the images in the proper way is a big part of this competition?</p>",
      "rawMarkdown": "Okay, that makes sense now. If I simply do that kind of split, my test accuracy will decrease.\nSo I guess normalizing the images in the proper way is a big part of this competition?",
      "votes": null
    },
    {
      "id": "607297",
      "postDate": "08/25/2019 01:32:27",
      "content": "<p>Test acc might decrease, but the cv lb difference should shrink. Really the only way to tell if your model is learning what it needs to for this competition is to gage its performance on totally heldout experiments for validation. Else it’s only learning stuff specific to the batches on which it was trained and will fail to classify the same treatment in new batches. </p>\n\n<p>Normalizing is, in a sense, what this competition is about. But if you could normalize away the batch effects prior to classification you would have already won. </p>",
      "rawMarkdown": "Test acc might decrease, but the cv lb difference should shrink. Really the only way to tell if your model is learning what it needs to for this competition is to gage its performance on totally heldout experiments for validation. Else it’s only learning stuff specific to the batches on which it was trained and will fail to classify the same treatment in new batches. \n\nNormalizing is, in a sense, what this competition is about. But if you could normalize away the batch effects prior to classification you would have already won.",
      "votes": null
    },
    {
      "id": "607795",
      "postDate": "08/25/2019 22:09:33",
      "content": "<p>In case you missed it - <a href=\"/lopuhin\">@lopuhin</a>'s thread discusses CV/LB correlation <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98116#latest-579835\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98116#latest-579835</a> . </p>",
      "rawMarkdown": "In case you missed it - @lopuhin's thread discusses CV/LB correlation https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98116#latest-579835 .",
      "votes": null
    },
    {
      "id": "608708",
      "postDate": "08/27/2019 05:08:37",
      "content": "<p><a href=\"/interneuron\">@interneuron</a> <a href=\"/zaharch\">@zaharch</a> You were right. I kept one cell line out for validation and I realized that my network is learning absolutely nothing! :)</p>",
      "rawMarkdown": "interneuron @zaharch You were right. I kept one cell line out for validation and I realized that my network is learning absolutely nothing! :)",
      "votes": null
    },
    {
      "id": "611068",
      "postDate": "08/29/2019 04:46:59",
      "content": "<p>Thanks for this useful thread. I have just got going on this competition using the TPU quota and experiencing a similar problem. First few runs, my validation accuracy was around 60% but my LB score was 12%. Now I will work on normalizing for brightness as suggested by <a href=\"/zaharch\">@zaharch</a> and <a href=\"/interneuron\">@interneuron</a> </p>",
      "rawMarkdown": "Thanks for this useful thread. I have just got going on this competition using the TPU quota and experiencing a similar problem. First few runs, my validation accuracy was around 60% but my LB score was 12%. Now I will work on normalizing for brightness as suggested by @zaharch and @interneuron",
      "votes": null
    },
    {
      "id": "611862",
      "postDate": "08/29/2019 13:21:55",
      "content": "<p>I'm still looking for a proper way to normalize the images. I did not find anything relevant to high-resolution microscopy images, though.</p>",
      "rawMarkdown": "I'm still looking for a proper way to normalize the images. I did not find anything relevant to high-resolution microscopy images, though.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 607122,
      "author_name": "zaharch",
      "author_url": "",
      "post_date": "08/24/2019 16:26:29",
      "content": "<p>One way to explain your CV vs LB divergence is the difference between experiments. You can see in pixel_stats.csv that the experiments are quite different in brightness. If you don't account for that and train your model on one mode of brightness it may perform poorly on the other mode of brightness.</p>",
      "votes": null,
      "replies": [
        {
          "id": 607124,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/24/2019 16:31:49",
          "content": "<p>Interesting. But if I stratify by experiment, shouldn't the model be able to handle it?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607136,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "08/24/2019 16:42:44",
          "content": "<p>Not if your validation set shares any experiments with the training set, it will just learn stuff about those specific experiments. That is why your val act is high. </p>\n\n<p>The test set explicitly comes from experiments (batches) that are not in the training set. The goal of this competition is to figure out a way to identify and control for arbitrary batch, really instrumentation, induced effects to be able to cleanly see effects caused by the biological manipulation alone - in this case treatment with a bunch of different sirnas. </p>\n\n<p>To make validation make sense you must move entire experiments out of training and into valid. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607138,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/24/2019 16:49:15",
          "content": "<p>Okay, that makes sense now. If I simply do that kind of split, my test accuracy will decrease.\nSo I guess normalizing the images in the proper way is a big part of this competition?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 607297,
          "author_name": "interneuron",
          "author_url": "",
          "post_date": "08/25/2019 01:32:27",
          "content": "<p>Test acc might decrease, but the cv lb difference should shrink. Really the only way to tell if your model is learning what it needs to for this competition is to gage its performance on totally heldout experiments for validation. Else it’s only learning stuff specific to the batches on which it was trained and will fail to classify the same treatment in new batches. </p>\n\n<p>Normalizing is, in a sense, what this competition is about. But if you could normalize away the batch effects prior to classification you would have already won. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 608708,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/27/2019 05:08:37",
          "content": "<p><a href=\"/interneuron\">@interneuron</a> <a href=\"/zaharch\">@zaharch</a> You were right. I kept one cell line out for validation and I realized that my network is learning absolutely nothing! :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 607795,
      "author_name": "michelml",
      "author_url": "",
      "post_date": "08/25/2019 22:09:33",
      "content": "<p>In case you missed it - <a href=\"/lopuhin\">@lopuhin</a>'s thread discusses CV/LB correlation <a href=\"https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98116#latest-579835\">https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98116#latest-579835</a> . </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 611068,
      "author_name": "kenkrige",
      "author_url": "",
      "post_date": "08/29/2019 04:46:59",
      "content": "<p>Thanks for this useful thread. I have just got going on this competition using the TPU quota and experiencing a similar problem. First few runs, my validation accuracy was around 60% but my LB score was 12%. Now I will work on normalizing for brightness as suggested by <a href=\"/zaharch\">@zaharch</a> and <a href=\"/interneuron\">@interneuron</a> </p>",
      "votes": null,
      "replies": [
        {
          "id": 611862,
          "author_name": "lorenzofabbri92",
          "author_url": "",
          "post_date": "08/29/2019 13:21:55",
          "content": "<p>I'm still looking for a proper way to normalize the images. I did not find anything relevant to high-resolution microscopy images, though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "607104": "Nobody seems to talk about correlation between LB and CV scores anymore.\n\nI'm using DenseNet121 (PyTorch) pretrained on ImageNet with Adam and no schedulers. Simple classification using cross-entropy. I'm not doing CV but I just split into training and test sets using Scikit-learn's `train_test_split`: I tried to use either the `sirna` or `experiment` fields for stratification but it does not seem to matter that much. I'm now facing the issue where my test accuracy is much higher than the LB score. Below you'll find some relevant figures: I augmented the dataset with random horizontal flipping, resize to 256 and random crop to 224 (with Albumentation). I also use both sites for prediction (averaging).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2325848%2F80a173ddedb794addc774177c1b081a4%2Fkaggle.png?generation=1566662146285345&amp;alt=media)\n\nThe corresponding LB score is 0.197. Is anybody facing similar issues? Now that the LB scores are much higher than a few weeks ago, can somebody report some example scores between CV and LB? I'm currently not doing any regularization since I read some papers saying that augmentation should suffice in general.",
    "607122": "One way to explain your CV vs LB divergence is the difference between experiments. You can see in pixel_stats.csv that the experiments are quite different in brightness. If you don't account for that and train your model on one mode of brightness it may perform poorly on the other mode of brightness.",
    "607124": "Interesting. But if I stratify by experiment, shouldn't the model be able to handle it?",
    "607136": "Not if your validation set shares any experiments with the training set, it will just learn stuff about those specific experiments. That is why your val act is high. \n\nThe test set explicitly comes from experiments (batches) that are not in the training set. The goal of this competition is to figure out a way to identify and control for arbitrary batch, really instrumentation, induced effects to be able to cleanly see effects caused by the biological manipulation alone - in this case treatment with a bunch of different sirnas. \n\nTo make validation make sense you must move entire experiments out of training and into valid.",
    "607138": "Okay, that makes sense now. If I simply do that kind of split, my test accuracy will decrease.\nSo I guess normalizing the images in the proper way is a big part of this competition?",
    "607297": "Test acc might decrease, but the cv lb difference should shrink. Really the only way to tell if your model is learning what it needs to for this competition is to gage its performance on totally heldout experiments for validation. Else it’s only learning stuff specific to the batches on which it was trained and will fail to classify the same treatment in new batches. \n\nNormalizing is, in a sense, what this competition is about. But if you could normalize away the batch effects prior to classification you would have already won.",
    "607795": "In case you missed it - @lopuhin's thread discusses CV/LB correlation https://www.kaggle.com/c/recursion-cellular-image-classification/discussion/98116#latest-579835 .",
    "608708": "interneuron @zaharch You were right. I kept one cell line out for validation and I realized that my network is learning absolutely nothing! :)",
    "611068": "Thanks for this useful thread. I have just got going on this competition using the TPU quota and experiencing a similar problem. First few runs, my validation accuracy was around 60% but my LB score was 12%. Now I will work on normalizing for brightness as suggested by @zaharch and @interneuron",
    "611862": "I'm still looking for a proper way to normalize the images. I did not find anything relevant to high-resolution microscopy images, though."
  },
  "source": "meta"
}