{
  "id": 100815,
  "title": "Discrepancies between CV/LB",
  "url": "/competitions/aptos2019-blindness-detection/discussion/100815",
  "author_name": "Tom Aindow",
  "post_date": "2019-07-21T11:15:52.963000",
  "votes": 67,
  "comment_count": 75,
  "views": 0,
  "content": "<p>I've seen a few posts now where people are seeing cv 0.80+ but submitting and receiving abysmal LB scores (sometimes even negative). For example:</p>\n\n<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100680\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100680</a>\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100583\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100583</a>\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98194\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98194</a></p>\n\n<p>I believe a major part of this issue is down to the fact that image size and black space cropping are strongly indicative of diagnosis in the training data, but not in the public test data. </p>\n\n<p>I've put together a kernel showing how by <strong>using just meta features you can score 0.70+ validation kappa locally, but get a negative LB score</strong>:</p>\n\n<p><a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on\">Be careful what you train on ...</a></p>\n\n<p>My suggestion if you are having this problem is to focus on pre-processing and ensuring that the leaky image size and cropping information is not influencing your model. Pre-trained models will always help since we start from a set of robust features.</p>\n\n<p>I think we need to be very careful about how we handle this pre-processing, since it is going to have a large effect on local CV scores. \"Trust local CV\" is always the mantra, but sometimes it can lead you astray ...</p>",
  "messages": [
    {
      "id": 581073,
      "postDate": "2019-07-21T11:15:52.963Z",
      "content": "<p>I've seen a few posts now where people are seeing cv 0.80+ but submitting and receiving abysmal LB scores (sometimes even negative). For example:</p>\n\n<p><a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100680\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100680</a>\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100583\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100583</a>\n<a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98194\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98194</a></p>\n\n<p>I believe a major part of this issue is down to the fact that image size and black space cropping are strongly indicative of diagnosis in the training data, but not in the public test data. </p>\n\n<p>I've put together a kernel showing how by <strong>using just meta features you can score 0.70+ validation kappa locally, but get a negative LB score</strong>:</p>\n\n<p><a href=\"https://www.kaggle.com/taindow/be-careful-what-you-train-on\">Be careful what you train on ...</a></p>\n\n<p>My suggestion if you are having this problem is to focus on pre-processing and ensuring that the leaky image size and cropping information is not influencing your model. Pre-trained models will always help since we start from a set of robust features.</p>\n\n<p>I think we need to be very careful about how we handle this pre-processing, since it is going to have a large effect on local CV scores. \"Trust local CV\" is always the mantra, but sometimes it can lead you astray ...</p>",
      "rawMarkdown": "I've seen a few posts now where people are seeing cv 0.80+ but submitting and receiving abysmal LB scores (sometimes even negative). For example:\n\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100680\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100583\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98194\n\nI believe a major part of this issue is down to the fact that image size and black space cropping are strongly indicative of diagnosis in the training data, but not in the public test data. \n\nI've put together a kernel showing how by **using just meta features you can score 0.70+ validation kappa locally, but get a negative LB score**:\n\n[Be careful what you train on ...](https://www.kaggle.com/taindow/be-careful-what-you-train-on)\n\nMy suggestion if you are having this problem is to focus on pre-processing and ensuring that the leaky image size and cropping information is not influencing your model. Pre-trained models will always help since we start from a set of robust features.\n\nI think we need to be very careful about how we handle this pre-processing, since it is going to have a large effect on local CV scores. \"Trust local CV\" is always the mantra, but sometimes it can lead you astray ...",
      "votes": 66
    },
    {
      "id": 581136,
      "postDate": "2019-07-21T13:49:09.043Z",
      "content": "<p>Great Post thank you for doing this! </p>\n\n<p>I have also noticed this was an issue so one way I overcome this was using a lot of Augmentations. Here is the results of some of my experiments. </p>\n\n<p>```\nsetup: \nRegression Problem\nIMG_SIZE: 224 \nMODEL: B3\nno TTA:\n5 fold CV (same across experiments) </p>\n\n<p>EDIT:\npertained on old competition data (train)(full not cropped). \noptimizer: ADAM\nmonitor: Valid Loss\nvalid_dataset: 2019\ntraining_phase_1: Load Image Net weights, freeze up to last layer train for 5 epoch \ntraining_phase_2: Unfreeze whole model and train for 15 epoch. </p>\n\n<p>for the Experiments below. \nI load weights which were trained above unfreeze the model and train for 5 epochs, monitoring  valid loss \n```</p>\n\n<p>EXP_1:\n```\nno AUGMENT:\nCV: 0.93 LB: 0.78</p>\n\n<p><code>\nEXP_2:\n</code>\nAUGMENT = [FLIPS]\nCV: 0.934, LB: 0.783\n```</p>\n\n<p>EXP_3:\n```\nAUGMENT = [rotation(360)]\nCV: 0.92 LB: 0.79</p>\n\n<p><code>\nEXP4:\n</code>\nAUGMENT = [zoom_to_center_area(1. 3)]\nCV: 0.90 LB: 0.793\n```</p>\n\n<p>EXP5: \n```\nAUGMENT = [FLIP, ZOOM_TO_CENTER, ROTATE(360), CONTRAST]\nCV: 0.89 LB: 0.809 </p>\n\n<p>```</p>",
      "rawMarkdown": "Great Post thank you for doing this! \n\nI have also noticed this was an issue so one way I overcome this was using a lot of Augmentations. Here is the results of some of my experiments. \n\n```\nsetup: \nRegression Problem\nIMG_SIZE: 224 \nMODEL: B3\nno TTA:\n5 fold CV (same across experiments) \n\nEDIT:\npertained on old competition data (train)(full not cropped). \noptimizer: ADAM\nmonitor: Valid Loss\nvalid_dataset: 2019\ntraining_phase_1: Load Image Net weights, freeze up to last layer train for 5 epoch \ntraining_phase_2: Unfreeze whole model and train for 15 epoch. \n\nfor the Experiments below. \nI load weights which were trained above unfreeze the model and train for 5 epochs, monitoring  valid loss \n```\n\nEXP_1:\n```\nno AUGMENT:\nCV: 0.93 LB: 0.78\n\n```\nEXP_2:\n```\nAUGMENT = [FLIPS]\nCV: 0.934, LB: 0.783\n```\n\n\nEXP_3:\n```\nAUGMENT = [rotation(360)]\nCV: 0.92 LB: 0.79\n\n```\nEXP4:\n```\nAUGMENT = [zoom_to_center_area(1. 3)]\nCV: 0.90 LB: 0.793\n```\n\nEXP5: \n```\nAUGMENT = [FLIP, ZOOM_TO_CENTER, ROTATE(360), CONTRAST]\nCV: 0.89 LB: 0.809 \n\n```\n\n",
      "votes": 56,
      "replies": [
        {
          "id": 581165,
          "postDate": "2019-07-21T14:42:54.043Z",
          "content": "<p>Only 2019 data? It's so weird how good of a score you can get with a simple model.</p>",
          "rawMarkdown": "Only 2019 data? It's so weird how good of a score you can get with a simple model.",
          "votes": 1
        },
        {
          "id": 581179,
          "postDate": "2019-07-21T14:59:31.650Z",
          "content": "<p>Thanks for brining this up. I made edits to my post with proper description of my setup </p>",
          "rawMarkdown": "Thanks for brining this up. I made edits to my post with proper description of my setup ",
          "votes": 1
        },
        {
          "id": 581210,
          "postDate": "2019-07-21T15:34:16.417Z",
          "content": "<p>Nice post, I imagine the zoom to center will mitigate the issue a lot :)</p>",
          "rawMarkdown": "Nice post, I imagine the zoom to center will mitigate the issue a lot :)",
          "votes": 1
        },
        {
          "id": 581236,
          "postDate": "2019-07-21T16:20:12.760Z",
          "content": "<p>Thanks a lot <a href=\"/drhabib\">@drhabib</a>, really helpful. Just one quick follow-up question: Are you fitting for the whole 20 epochs in pretraining or taking best epoch based on 2019 loss?</p>",
          "rawMarkdown": "Thanks a lot @drhabib, really helpful. Just one quick follow-up question: Are you fitting for the whole 20 epochs in pretraining or taking best epoch based on 2019 loss?",
          "votes": 1
        },
        {
          "id": 581305,
          "postDate": "2019-07-21T18:27:58.597Z",
          "content": "<p>After 5+ 15 I am taking the best epoch based on 2019 loss.</p>",
          "rawMarkdown": "After 5+ 15 I am taking the best epoch based on 2019 loss.",
          "votes": 2
        },
        {
          "id": 581365,
          "postDate": "2019-07-21T20:32:16.297Z",
          "content": "<p>how did you balance the classes? did you just add old data until each of the classes had the same amount of examples?</p>",
          "rawMarkdown": "how did you balance the classes? did you just add old data until each of the classes had the same amount of examples?",
          "votes": 2
        },
        {
          "id": 581421,
          "postDate": "2019-07-21T23:59:45.520Z",
          "content": "<p>Thanks for sharing! <br>\nWhat is your training data?  </p>\n\n<p>You wrote that your valid data is 2019 and pretrained is 2015. <br>\nWhich is correct? training is 2019 or 2015 + 2019.  </p>",
          "rawMarkdown": "Thanks for sharing!  \nWhat is your training data?  \n\nYou wrote that your valid data is 2019 and pretrained is 2015.  \nWhich is correct? training is 2019 or 2015 + 2019.  ",
          "votes": 1
        },
        {
          "id": 581427,
          "postDate": "2019-07-22T00:17:21.577Z",
          "content": "<p>Hi, \nI train on all 2015 and use 2019 as validation and save the weights with best validation loss. Hope its clear now  </p>",
          "rawMarkdown": "Hi, \nI train on all 2015 and use 2019 as validation and save the weights with best validation loss. Hope its clear now  ",
          "votes": 9
        },
        {
          "id": 581444,
          "postDate": "2019-07-22T00:32:30.063Z",
          "content": "<p>Wait you don't ever train on the 2019 competition data?</p>",
          "rawMarkdown": "Wait you don't ever train on the 2019 competition data?"
        },
        {
          "id": 581447,
          "postDate": "2019-07-22T00:41:42.093Z",
          "content": "<p>I pretrain my initial model on old data. Afterwards I do standard training on new data (<strong>with the weights trained from old data</strong>) with CV splits.  But just as fun I was able to get LB 0.719-0.730 with the model trained just on 2015 data =) </p>",
          "rawMarkdown": "I pretrain my initial model on old data. Afterwards I do standard training on new data (**with the weights trained from old data**) with CV splits.  But just as fun I was able to get LB 0.719-0.730 with the model trained just on 2015 data =) ",
          "votes": 3
        },
        {
          "id": 581449,
          "postDate": "2019-07-22T00:45:52.920Z",
          "content": "<p>OK if I understand correctly:\n1. Train on old dataset with valid. from current dataset\n2. Train on current dataset (weights from old) and valid. from current dataset </p>",
          "rawMarkdown": "OK if I understand correctly:\n1. Train on old dataset with valid. from current dataset\n2. Train on current dataset (weights from old) and valid. from current dataset ",
          "votes": 10
        },
        {
          "id": 581453,
          "postDate": "2019-07-22T00:50:37.660Z",
          "content": "<p>Yes =) </p>",
          "rawMarkdown": "Yes =) ",
          "votes": 3
        },
        {
          "id": 581462,
          "postDate": "2019-07-22T01:12:31.580Z",
          "content": "<p>another question, why did you switch from b5 to b3? for speed/performance?</p>",
          "rawMarkdown": "another question, why did you switch from b5 to b3? for speed/performance?"
        },
        {
          "id": 581464,
          "postDate": "2019-07-22T01:18:31.543Z",
          "content": "<p>No particular reason, just wanted to understand better efficientnet also speed was advantage. After playing around with efficientnet it seems like they are a bit harder to train... </p>",
          "rawMarkdown": "No particular reason, just wanted to understand better efficientnet also speed was advantage. After playing around with efficientnet it seems like they are a bit harder to train... ",
          "votes": 3
        },
        {
          "id": 581471,
          "postDate": "2019-07-22T01:43:07.093Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> in my experiments, i found that CV score is inversely proportional to public LB =))). \nCV: 92 -&gt;&gt;&gt; LB: 81.0\nCV: 95 -&gt;&gt;&gt; LB: 79.9\nCV: 97.5 -&gt;&gt;&gt; lB: 78.2\nSo, don't trust LB :v</p>",
          "rawMarkdown": "@drhabib in my experiments, i found that CV score is inversely proportional to public LB =))). \nCV: 92 -&gt;&gt;&gt; LB: 81.0\nCV: 95 -&gt;&gt;&gt; LB: 79.9\nCV: 97.5 -&gt;&gt;&gt; lB: 78.2\nSo, don't trust LB :v",
          "votes": 4
        },
        {
          "id": 581477,
          "postDate": "2019-07-22T01:46:59.720Z",
          "content": "<p>hahah =) I am with you there  =) I am definitely not going to choose my highest LB as a submission =) </p>",
          "rawMarkdown": "hahah =) I am with you there  =) I am definitely not going to choose my highest LB as a submission =) ",
          "votes": 2
        },
        {
          "id": 581485,
          "postDate": "2019-07-22T02:03:49.830Z",
          "content": "<p>I am not sure what to trust though. For all we know the LB is very similar to the private LB, and we should trust the LB over our CV. </p>",
          "rawMarkdown": "I am not sure what to trust though. For all we know the LB is very similar to the private LB, and we should trust the LB over our CV. "
        },
        {
          "id": 581852,
          "postDate": "2019-07-22T13:23:08.597Z",
          "content": "<p>Can anyone please share their knowledge about,\nWhat's the difference between b0, b1, ..., b5?\nAre they in increasing order of complexity?\n Thanks you.</p>",
          "rawMarkdown": "Can anyone please share their knowledge about,\nWhat's the difference between b0, b1, ..., b5?\nAre they in increasing order of complexity?\n Thanks you."
        },
        {
          "id": 581853,
          "postDate": "2019-07-22T13:23:48.683Z",
          "content": "<p>yes, more parameters</p>",
          "rawMarkdown": "yes, more parameters",
          "votes": 1
        },
        {
          "id": 582926,
          "postDate": "2019-07-23T19:35:38.403Z",
          "content": "<p><a href=\"/prashantkikani\">@prashantkikani</a> yes. EfficientNet B1 to B7 are obtained by scaling up the EfficientNet B0 architecture. You can read a short summary <a href=\"https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html\">here</a></p>",
          "rawMarkdown": "@prashantkikani yes. EfficientNet B1 to B7 are obtained by scaling up the EfficientNet B0 architecture. You can read a short summary [here](https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html)",
          "votes": 3
        },
        {
          "id": 583414,
          "postDate": "2019-07-24T13:02:16.517Z",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> why do you think private LB distribution is more similar to public LB? I think the private distribution can be more similar to the training data distribution (more labels 0). The reason that makes me believe it, is that the old dataset, which has a similar amount of images compared to private data, also have more labels 0.</p>",
          "rawMarkdown": "@tanlikesmath why do you think private LB distribution is more similar to public LB? I think the private distribution can be more similar to the training data distribution (more labels 0). The reason that makes me believe it, is that the old dataset, which has a similar amount of images compared to private data, also have more labels 0.",
          "votes": 1
        },
        {
          "id": 583423,
          "postDate": "2019-07-24T13:15:13.393Z",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> The old dataset you are using is this <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a>?\nThanks for sharing!</p>",
          "rawMarkdown": "@drhabib The old dataset you are using is this https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized?\nThanks for sharing!",
          "votes": 1
        },
        {
          "id": 583425,
          "postDate": "2019-07-24T13:18:13.997Z",
          "content": "<p>yes =) </p>",
          "rawMarkdown": "yes =) ",
          "votes": 1
        },
        {
          "id": 583953,
          "postDate": "2019-07-25T08:38:19.887Z",
          "content": "<p>do you use some preprocess like <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping</a>?</p>",
          "rawMarkdown": "do you use some preprocess like https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping?"
        },
        {
          "id": 585876,
          "postDate": "2019-07-28T08:30:52.267Z",
          "content": "<p>Hi, DrHB, Is this CV score processed by optimized kappa, you choose a constant threshold?</p>",
          "rawMarkdown": "Hi, DrHB, Is this CV score processed by optimized kappa, you choose a constant threshold?"
        },
        {
          "id": 586126,
          "postDate": "2019-07-28T16:45:28.020Z",
          "content": "<p>no Optimized Kappa. Just standard thresholds [0.5, 1.5, 2.5, 3.5] =)</p>",
          "rawMarkdown": "no Optimized Kappa. Just standard thresholds [0.5, 1.5, 2.5, 3.5] =)",
          "votes": 4
        },
        {
          "id": 586917,
          "postDate": "2019-07-29T21:51:15.653Z",
          "content": "<p>Thanks for sharing :)</p>",
          "rawMarkdown": "Thanks for sharing :)"
        },
        {
          "id": 588927,
          "postDate": "2019-07-31T08:29:28.297Z",
          "content": "<p>Hi, DrHB. Did you just train 5 epochs(the lr is?) on 2019 datas with the pretrained model on 2015 data? </p>",
          "rawMarkdown": "Hi, DrHB. Did you just train 5 epochs(the lr is?) on 2019 datas with the pretrained model on 2015 data? ",
          "votes": 1
        },
        {
          "id": 589050,
          "postDate": "2019-07-31T11:46:08.373Z",
          "content": "<p>yes! and my lr (learning rate) was 1e-3/2. </p>",
          "rawMarkdown": "yes! and my lr (learning rate) was 1e-3/2. "
        },
        {
          "id": 589095,
          "postDate": "2019-07-31T12:53:22.510Z",
          "content": "<p>Thanks for your sharing.</p>",
          "rawMarkdown": "Thanks for your sharing."
        },
        {
          "id": 589103,
          "postDate": "2019-07-31T13:08:38.493Z",
          "content": "<p>you are welcome</p>",
          "rawMarkdown": "you are welcome"
        },
        {
          "id": 589248,
          "postDate": "2019-07-31T17:04:05.333Z",
          "content": "<p>Hi <a href=\"/dathudeptrai\">@dathudeptrai</a>, as you said in your experiments: \"CV: 92 -&gt;&gt;&gt; LB: 81.0\".\nMay I ask more about this experiment: Did you solve the problem with a regression model? Is CV a kappa score? If so, you set the thresholds as [0.5, 1.5, 2.5, 3.5] to calculate CV kappa or you use optimized thresholds?</p>",
          "rawMarkdown": "Hi @dathudeptrai, as you said in your experiments: \"CV: 92 -&gt;&gt;&gt; LB: 81.0\".\nMay I ask more about this experiment: Did you solve the problem with a regression model? Is CV a kappa score? If so, you set the thresholds as [0.5, 1.5, 2.5, 3.5] to calculate CV kappa or you use optimized thresholds?"
        },
        {
          "id": 590060,
          "postDate": "2019-08-01T18:48:36.020Z",
          "content": "<p>Interesting HB, I notice that you specified that you use the full image for pretraining on old competition data. Do you do any augments on that image, or do you just train on what's given?</p>",
          "rawMarkdown": "Interesting HB, I notice that you specified that you use the full image for pretraining on old competition data. Do you do any augments on that image, or do you just train on what's given?"
        },
        {
          "id": 590653,
          "postDate": "2019-08-02T12:49:10.617Z",
          "content": "<p>Thanks for sharing your setup! It beautifully clears up the discussion of \"to augment or not to augment\".</p>",
          "rawMarkdown": "Thanks for sharing your setup! It beautifully clears up the discussion of \"to augment or not to augment\"."
        },
        {
          "id": 596837,
          "postDate": "2019-08-11T11:43:05.413Z",
          "content": "<p>Thanks for sharing! This thread is really insightful.\n&gt;Train on old dataset with valid. from current dataset\n&gt;Train on current dataset (weights from old) and valid. from current dataset</p>\n\n<p><a href=\"/drhabib\">@drhabib</a> \nI'm curious about how many epochs on the latter process. Thanks!</p>",
          "rawMarkdown": "Thanks for sharing! This thread is really insightful.\n&gt;Train on old dataset with valid. from current dataset\n&gt;Train on current dataset (weights from old) and valid. from current dataset\n\n@drhabib \nI'm curious about how many epochs on the latter process. Thanks!"
        },
        {
          "id": 598120,
          "postDate": "2019-08-13T06:50:24.670Z",
          "content": "<p>Just curious, for the EfficientNet, where did you cut? I did some simple test for B0\n1. no cut\n2. cut at m._conv_head, so layer_groups now is 2\n3. cut at m._conv_head and m._blocks[8], so layer_groups now is 3</p>\n\n<p>for 3, the reason that cut at m._blocks[8]. Just observed the channel size is changed from 80-&gt;112 after two 80-&gt;80 blocks, and it is right in the middle of the m._blocks</p>\n\n<p>Exp 1 and 3 was trained for 5 epochs (without freezing), Exp 2 freezed last layer for 5 epochs, then unfreeze for another 8 epochs.  All with same argumentation and all with pretrained weights, and all used discriminative lr except for 1 the layer groups is 1</p>\n\n<p>The results are 1 &gt;= 3 &gt;&gt; 2.</p>",
          "rawMarkdown": "Just curious, for the EfficientNet, where did you cut? I did some simple test for B0\n1. no cut\n2. cut at m._conv_head, so layer_groups now is 2\n3. cut at m._conv_head and m._blocks[8], so layer_groups now is 3\n\nfor 3, the reason that cut at m._blocks[8]. Just observed the channel size is changed from 80-&gt;112 after two 80-&gt;80 blocks, and it is right in the middle of the m._blocks\n\nExp 1 and 3 was trained for 5 epochs (without freezing), Exp 2 freezed last layer for 5 epochs, then unfreeze for another 8 epochs.  All with same argumentation and all with pretrained weights, and all used discriminative lr except for 1 the layer groups is 1\n\nThe results are 1 &gt;= 3 &gt;&gt; 2.\n\n"
        },
        {
          "id": 607744,
          "postDate": "2019-08-25T20:00:41.687Z",
          "content": "<p>can I do this in kaggle kernel (the old data training part)??</p>",
          "rawMarkdown": "can I do this in kaggle kernel (the old data training part)??"
        },
        {
          "id": 607755,
          "postDate": "2019-08-25T20:36:20.593Z",
          "content": "<p>yes. </p>",
          "rawMarkdown": "yes. "
        },
        {
          "id": 607959,
          "postDate": "2019-08-26T06:37:08.620Z",
          "content": "<p>What does CV mean?</p>",
          "rawMarkdown": "What does CV mean?"
        },
        {
          "id": 607994,
          "postDate": "2019-08-26T07:24:36.303Z",
          "content": "<p>It takes 45 min for 1 epoch of 35000 images. Am I doing anything wrong?</p>",
          "rawMarkdown": "It takes 45 min for 1 epoch of 35000 images. Am I doing anything wrong?"
        },
        {
          "id": 608031,
          "postDate": "2019-08-26T08:29:39.117Z",
          "content": "<p><a href=\"/apthagowda\">@apthagowda</a> It might be because of your preprocessing method which takes a lot of time.. Mostly the culprit will be resizing images. If that is the case, download the dataset and resize the images and upload it to kaggle. Using already resized images will reduce the epoch time. In my case an epoch takes around 10 minutes using resized dataset. Hope it helps :)</p>",
          "rawMarkdown": "@apthagowda It might be because of your preprocessing method which takes a lot of time.. Mostly the culprit will be resizing images. If that is the case, download the dataset and resize the images and upload it to kaggle. Using already resized images will reduce the epoch time. In my case an epoch takes around 10 minutes using resized dataset. Hope it helps :)",
          "votes": 1
        },
        {
          "id": 608081,
          "postDate": "2019-08-26T09:58:41.677Z",
          "rawMarkdown": ""
        }
      ]
    },
    {
      "id": 582273,
      "postDate": "2019-07-23T02:02:57.277Z",
      "content": "<p>To me, I found a different evidence to <a href=\"/drhabib\">@drhabib</a> as my CV/LB is quite correlated;</p>\n\n<p>CV &gt;&gt; LB\n.92x &gt;&gt; 0.77x\n.93x &gt;&gt; 0.79x\n.94x &gt;&gt; 0.80x</p>\n\n<p>Perhaps because I use heavy preprocessing / augmentation from the start. (I shared my six augmentations in my 2nd kernel)\nAlso there still a large gap even though those preprocessing/augmentations.</p>\n\n<p>IMO, “trust your CV” is true if we are able to set a correct validation setting (which I still not be able to do yet :)</p>",
      "rawMarkdown": "To me, I found a different evidence to @drhabib as my CV/LB is quite correlated;\n\nCV &gt;&gt; LB\n.92x &gt;&gt; 0.77x\n.93x &gt;&gt; 0.79x\n.94x &gt;&gt; 0.80x\n\nPerhaps because I use heavy preprocessing / augmentation from the start. (I shared my six augmentations in my 2nd kernel)\nAlso there still a large gap even though those preprocessing/augmentations.\n\nIMO, “trust your CV” is true if we are able to set a correct validation setting (which I still not be able to do yet :)",
      "votes": 7,
      "replies": [
        {
          "id": 585801,
          "postDate": "2019-07-28T05:08:38.590Z",
          "content": "<p>HI, Is this CV score processed by optimized kappa?</p>",
          "rawMarkdown": "HI, Is this CV score processed by optimized kappa?"
        }
      ]
    },
    {
      "id": 588223,
      "postDate": "2019-07-30T09:59:11.910Z",
      "content": "<p><a href=\"/taindow\">@taindow</a> the \"CV\" you mentioned is \"Cross Validation\"? \nso far as i known, K-fold Cross Validation will generate K models, so the CV value here is the average of those K models?  just like Jason Huang said in <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98609#latest-568539\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98609#latest-568539</a>\n\"\nMake CV: \nSplit the data to K folds and train K models by using each fold as the validation dataset. Then you can get CV score by computing the average validation score of each fold.</p>\n\n<p>Ensemble the result from different models:\nLoad each model and give out their corresponding predictions on the test dataset, then you can ensemble them.\n\"\nam i right?</p>",
      "rawMarkdown": "@taindow the \"CV\" you mentioned is \"Cross Validation\"? \nso far as i known, K-fold Cross Validation will generate K models, so the CV value here is the average of those K models?  just like Jason Huang said in https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98609#latest-568539\n\"\nMake CV: \nSplit the data to K folds and train K models by using each fold as the validation dataset. Then you can get CV score by computing the average validation score of each fold.\n\nEnsemble the result from different models:\nLoad each model and give out their corresponding predictions on the test dataset, then you can ensemble them.\n\"\nam i right?",
      "votes": 5,
      "replies": [
        {
          "id": 588229,
          "postDate": "2019-07-30T10:04:43.720Z",
          "content": "<p>Yes</p>",
          "rawMarkdown": "Yes",
          "votes": 1
        },
        {
          "id": 588517,
          "postDate": "2019-07-30T17:37:21.070Z",
          "content": "<p>This is correct :)</p>",
          "rawMarkdown": "This is correct :)",
          "votes": 1
        }
      ]
    },
    {
      "id": 581400,
      "postDate": "2019-07-21T22:09:36.973Z",
      "content": "<p>Another big problem is we have no idea what the private test distribution looks like and with these concerns about CV/LB discrepancies it is possible there may be a shake-up and it will be hard to prevent this.</p>\n\n<p>I have discussed this over <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493\">here</a></p>",
      "rawMarkdown": "Another big problem is we have no idea what the private test distribution looks like and with these concerns about CV/LB discrepancies it is possible there may be a shake-up and it will be hard to prevent this.\n\nI have discussed this over [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493)",
      "votes": 3
    },
    {
      "id": 581148,
      "postDate": "2019-07-21T14:11:14.287Z",
      "content": "<p>Oh  my  god , I  have  just  check  the  test  dataset ,  it's  quite   different  from  train  dataset ,  there  is  nearly  very  less  black   space   here ,   it   is   a  very  huge   gap   between  train_data  and     test_data </p>",
      "rawMarkdown": "Oh  my  god , I  have  just  check  the  test  dataset ,  it's  quite   different  from  train  dataset ,  there  is  nearly  very  less  black   space   here ,   it   is   a  very  huge   gap   between  train_data  and     test_data ",
      "votes": 3
    },
    {
      "id": 581089,
      "postDate": "2019-07-21T11:41:48.663Z",
      "content": "<p>That's a great finding. I've seen my LB improve after center cropping the retina. It's better to avoid the black region altogether, otherwise model is very prone to learn this weird correlation and fit itself on it.</p>",
      "rawMarkdown": "That's a great finding. I've seen my LB improve after center cropping the retina. It's better to avoid the black region altogether, otherwise model is very prone to learn this weird correlation and fit itself on it.",
      "votes": 4,
      "replies": [
        {
          "id": 582124,
          "postDate": "2019-07-22T19:17:00.767Z",
          "content": "<p>I also notice LB improvement with cropping, but even then local CV is not correlating. Are you seeing a good correlation? </p>",
          "rawMarkdown": "I also notice LB improvement with cropping, but even then local CV is not correlating. Are you seeing a good correlation? "
        },
        {
          "id": 582315,
          "postDate": "2019-07-23T03:45:50.300Z",
          "content": "<p>Tbh, No.\nI expect a big shake-up after the private test results are announced. After doing ~60 submissions so far I can literally look at the unique counts of predictions and say which submission is gonna perform good on the LB. Which is definitely not a good approach, I don't trust local qwk anymore, As of now I'm monitoring a lot of other classification metrics, I'm trying to find a correlation between them and the LB, will share if I find something useful there.</p>",
          "rawMarkdown": "Tbh, No.\nI expect a big shake-up after the private test results are announced. After doing ~60 submissions so far I can literally look at the unique counts of predictions and say which submission is gonna perform good on the LB. Which is definitely not a good approach, I don't trust local qwk anymore, As of now I'm monitoring a lot of other classification metrics, I'm trying to find a correlation between them and the LB, will share if I find something useful there.",
          "votes": 3
        }
      ]
    },
    {
      "id": 590270,
      "postDate": "2019-08-02T01:27:58.433Z",
      "content": "<p>Is this necessary for preprocessing? \nWhen I use preprocessing and removed the black pixels, I got a lower LB.</p>",
      "rawMarkdown": "Is this necessary for preprocessing? \nWhen I use preprocessing and removed the black pixels, I got a lower LB.",
      "votes": 2,
      "replies": [
        {
          "id": 590488,
          "postDate": "2019-08-02T08:40:30.750Z",
          "content": "<p>I got a higher LB score when I forgot to preprocess the test images in my inference kernel like I did when training.\nAlso resizing the images with OpenCV and PIL gave different scores</p>",
          "rawMarkdown": "I got a higher LB score when I forgot to preprocess the test images in my inference kernel like I did when training.\nAlso resizing the images with OpenCV and PIL gave different scores"
        },
        {
          "id": 590494,
          "postDate": "2019-08-02T08:43:53.190Z",
          "content": "<p>Different resizing on inference or also on train? On train it can be just the seed effect.</p>",
          "rawMarkdown": "Different resizing on inference or also on train? On train it can be just the seed effect."
        },
        {
          "id": 590521,
          "postDate": "2019-08-02T09:35:11.213Z",
          "content": "<p>@<a href=\"https://www.kaggle.com/suicaokhoailang\">Khoi Nguyen</a>What a wonderful coincidence..</p>",
          "rawMarkdown": "@[Khoi Nguyen](https://www.kaggle.com/suicaokhoailang)What a wonderful coincidence.."
        },
        {
          "id": 590522,
          "postDate": "2019-08-02T09:36:17.447Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\">@Psi</a>I think I did the same preprocessing on train and test.</p>",
          "rawMarkdown": "[@Psi](https://www.kaggle.com/philippsinger)I think I did the same preprocessing on train and test."
        },
        {
          "id": 590533,
          "postDate": "2019-08-02T09:54:24.390Z",
          "content": "<p>Then I am not surprised. Resizing in cv2 vs Pil cannot have any impact.</p>",
          "rawMarkdown": "Then I am not surprised. Resizing in cv2 vs Pil cannot have any impact."
        },
        {
          "id": 590540,
          "postDate": "2019-08-02T10:02:40.357Z",
          "content": "<p>Psi, is that you current LB score through preprocessing?</p>",
          "rawMarkdown": "Psi, is that you current LB score through preprocessing?"
        },
        {
          "id": 590560,
          "postDate": "2019-08-02T10:42:10.383Z",
          "content": "<p>What do you mean?</p>",
          "rawMarkdown": "What do you mean?"
        },
        {
          "id": 590564,
          "postDate": "2019-08-02T10:47:25.500Z",
          "content": "<p>sorry, did you use ben preprocessing to get the current LB score?</p>",
          "rawMarkdown": "sorry, did you use ben preprocessing to get the current LB score?"
        },
        {
          "id": 590587,
          "postDate": "2019-08-02T11:19:04.587Z",
          "content": "<p>Combination of few things.</p>",
          "rawMarkdown": "Combination of few things."
        },
        {
          "id": 590603,
          "postDate": "2019-08-02T11:43:46.460Z",
          "content": "<p>yep, thank you.</p>",
          "rawMarkdown": "yep, thank you."
        },
        {
          "id": 596946,
          "postDate": "2019-08-11T15:02:55.663Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 607942,
      "postDate": "2019-08-26T05:32:31.970Z",
      "content": "<p>Thanks for this. I'm wondering, when using a pre-trained model, should I freeze any layers, or just unfreeze all?</p>",
      "rawMarkdown": "Thanks for this. I'm wondering, when using a pre-trained model, should I freeze any layers, or just unfreeze all?"
    },
    {
      "id": 590661,
      "postDate": "2019-08-02T12:59:50.800Z",
      "content": "<p>My theory: Pre-processing allows models to extract features much easier therefore learn much faster and models don't have to be too deep (I guess we need less layers to do feature extracting after pre-processing?)</p>",
      "rawMarkdown": "My theory: Pre-processing allows models to extract features much easier therefore learn much faster and models don't have to be too deep (I guess we need less layers to do feature extracting after pre-processing?)"
    },
    {
      "id": 590213,
      "postDate": "2019-08-01T21:55:10.180Z",
      "content": "<p>Ah ok! Thank you for clearing that up! Time to dive deeper into pre-processing!</p>",
      "rawMarkdown": "Ah ok! Thank you for clearing that up! Time to dive deeper into pre-processing!"
    },
    {
      "id": 584263,
      "postDate": "2019-07-25T16:25:42.773Z",
      "content": "<p>I really think, we should focus on our model's interpretability. If it picks correct features, it would be in quite good agreement with doctors. In real world scenario (after deployment), model will get single data point to predict on, so there will nothing be like \"Test Distribution\". And this is the reason, why I liked your kernel so much. Thanks.</p>",
      "rawMarkdown": "I really think, we should focus on our model's interpretability. If it picks correct features, it would be in quite good agreement with doctors. In real world scenario (after deployment), model will get single data point to predict on, so there will nothing be like \"Test Distribution\". And this is the reason, why I liked your kernel so much. Thanks.",
      "replies": [
        {
          "id": 584269,
          "postDate": "2019-07-25T16:35:07.013Z",
          "content": "<p>there** is** \"test distribution\" in real world scenario. and it is likely close to train set. Consider previous competition which also had train set  with similar distribution. </p>",
          "rawMarkdown": "there** is** \"test distribution\" in real world scenario. and it is likely close to train set. Consider previous competition which also had train set  with similar distribution. "
        }
      ]
    },
    {
      "id": 582932,
      "postDate": "2019-07-23T19:55:08.367Z",
      "content": "<p>Great post!</p>",
      "rawMarkdown": "Great post!"
    },
    {
      "id": 581950,
      "postDate": "2019-07-22T15:24:30.060Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 581856,
      "postDate": "2019-07-22T13:25:18.823Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 582668,
      "postDate": "2019-07-23T12:26:39.770Z",
      "content": "<p>Thanks for the post</p>",
      "rawMarkdown": "Thanks for the post"
    },
    {
      "id": 582005,
      "postDate": "2019-07-22T16:44:20.847Z",
      "content": "<p>Thanks for the post</p>",
      "rawMarkdown": "Thanks for the post"
    }
  ],
  "comments": [
    {
      "id": 581136,
      "author_name": "DrHB",
      "author_url": "",
      "post_date": "2019-07-21T13:49:09.043000",
      "content": "<p>Great Post thank you for doing this! </p>\n\n<p>I have also noticed this was an issue so one way I overcome this was using a lot of Augmentations. Here is the results of some of my experiments. </p>\n\n<p>```\nsetup: \nRegression Problem\nIMG_SIZE: 224 \nMODEL: B3\nno TTA:\n5 fold CV (same across experiments) </p>\n\n<p>EDIT:\npertained on old competition data (train)(full not cropped). \noptimizer: ADAM\nmonitor: Valid Loss\nvalid_dataset: 2019\ntraining_phase_1: Load Image Net weights, freeze up to last layer train for 5 epoch \ntraining_phase_2: Unfreeze whole model and train for 15 epoch. </p>\n\n<p>for the Experiments below. \nI load weights which were trained above unfreeze the model and train for 5 epochs, monitoring  valid loss \n```</p>\n\n<p>EXP_1:\n```\nno AUGMENT:\nCV: 0.93 LB: 0.78</p>\n\n<p><code>\nEXP_2:\n</code>\nAUGMENT = [FLIPS]\nCV: 0.934, LB: 0.783\n```</p>\n\n<p>EXP_3:\n```\nAUGMENT = [rotation(360)]\nCV: 0.92 LB: 0.79</p>\n\n<p><code>\nEXP4:\n</code>\nAUGMENT = [zoom_to_center_area(1. 3)]\nCV: 0.90 LB: 0.793\n```</p>\n\n<p>EXP5: \n```\nAUGMENT = [FLIP, ZOOM_TO_CENTER, ROTATE(360), CONTRAST]\nCV: 0.89 LB: 0.809 </p>\n\n<p>```</p>",
      "votes": 56,
      "replies": [
        {
          "id": 581165,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-21T14:42:54.043000",
          "content": "<p>Only 2019 data? It's so weird how good of a score you can get with a simple model.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581179,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-21T14:59:31.650000",
          "content": "<p>Thanks for brining this up. I made edits to my post with proper description of my setup </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581210,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-21T15:34:16.417000",
          "content": "<p>Nice post, I imagine the zoom to center will mitigate the issue a lot :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581236,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-21T16:20:12.760000",
          "content": "<p>Thanks a lot <a href=\"/drhabib\">@drhabib</a>, really helpful. Just one quick follow-up question: Are you fitting for the whole 20 epochs in pretraining or taking best epoch based on 2019 loss?</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581305,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-21T18:27:58.597000",
          "content": "<p>After 5+ 15 I am taking the best epoch based on 2019 loss.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 581365,
          "author_name": "sh",
          "author_url": "",
          "post_date": "2019-07-21T20:32:16.297000",
          "content": "<p>how did you balance the classes? did you just add old data until each of the classes had the same amount of examples?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 581421,
          "author_name": "tereka",
          "author_url": "",
          "post_date": "2019-07-21T23:59:45.520000",
          "content": "<p>Thanks for sharing! <br>\nWhat is your training data?  </p>\n\n<p>You wrote that your valid data is 2019 and pretrained is 2015. <br>\nWhich is correct? training is 2019 or 2015 + 2019.  </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 581427,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-22T00:17:21.577000",
          "content": "<p>Hi, \nI train on all 2015 and use 2019 as validation and save the weights with best validation loss. Hope its clear now  </p>",
          "votes": 9,
          "replies": []
        },
        {
          "id": 581444,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-07-22T00:32:30.063000",
          "content": "<p>Wait you don't ever train on the 2019 competition data?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 581447,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-22T00:41:42.093000",
          "content": "<p>I pretrain my initial model on old data. Afterwards I do standard training on new data (<strong>with the weights trained from old data</strong>) with CV splits.  But just as fun I was able to get LB 0.719-0.730 with the model trained just on 2015 data =) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 581449,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-07-22T00:45:52.920000",
          "content": "<p>OK if I understand correctly:\n1. Train on old dataset with valid. from current dataset\n2. Train on current dataset (weights from old) and valid. from current dataset </p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 581453,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-22T00:50:37.660000",
          "content": "<p>Yes =) </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 581462,
          "author_name": "sh",
          "author_url": "",
          "post_date": "2019-07-22T01:12:31.580000",
          "content": "<p>another question, why did you switch from b5 to b3? for speed/performance?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 581464,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-22T01:18:31.543000",
          "content": "<p>No particular reason, just wanted to understand better efficientnet also speed was advantage. After playing around with efficientnet it seems like they are a bit harder to train... </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 581471,
          "author_name": "Nguyen Quan Anh Minh",
          "author_url": "",
          "post_date": "2019-07-22T01:43:07.093000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> in my experiments, i found that CV score is inversely proportional to public LB =))). \nCV: 92 -&gt;&gt;&gt; LB: 81.0\nCV: 95 -&gt;&gt;&gt; LB: 79.9\nCV: 97.5 -&gt;&gt;&gt; lB: 78.2\nSo, don't trust LB :v</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 581477,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-22T01:46:59.720000",
          "content": "<p>hahah =) I am with you there  =) I am definitely not going to choose my highest LB as a submission =) </p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 581485,
          "author_name": "ilovescience",
          "author_url": "",
          "post_date": "2019-07-22T02:03:49.830000",
          "content": "<p>I am not sure what to trust though. For all we know the LB is very similar to the private LB, and we should trust the LB over our CV. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 581852,
          "author_name": "Prashant Kikani",
          "author_url": "",
          "post_date": "2019-07-22T13:23:08.597000",
          "content": "<p>Can anyone please share their knowledge about,\nWhat's the difference between b0, b1, ..., b5?\nAre they in increasing order of complexity?\n Thanks you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 581853,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-07-22T13:23:48.683000",
          "content": "<p>yes, more parameters</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 582926,
          "author_name": "Nikita Kozodoi",
          "author_url": "",
          "post_date": "2019-07-23T19:35:38.403000",
          "content": "<p><a href=\"/prashantkikani\">@prashantkikani</a> yes. EfficientNet B1 to B7 are obtained by scaling up the EfficientNet B0 architecture. You can read a short summary <a href=\"https://ai.googleblog.com/2019/05/efficientnet-improving-accuracy-and.html\">here</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 583414,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-07-24T13:02:16.517000",
          "content": "<p><a href=\"/tanlikesmath\">@tanlikesmath</a> why do you think private LB distribution is more similar to public LB? I think the private distribution can be more similar to the training data distribution (more labels 0). The reason that makes me believe it, is that the old dataset, which has a similar amount of images compared to private data, also have more labels 0.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 583423,
          "author_name": "Carlos Prades K.",
          "author_url": "",
          "post_date": "2019-07-24T13:15:13.393000",
          "content": "<p><a href=\"/drhabib\">@drhabib</a> The old dataset you are using is this <a href=\"https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized\">https://www.kaggle.com/tanlikesmath/diabetic-retinopathy-resized</a>?\nThanks for sharing!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 583425,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-24T13:18:13.997000",
          "content": "<p>yes =) </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 583953,
          "author_name": "kevin",
          "author_url": "",
          "post_date": "2019-07-25T08:38:19.887000",
          "content": "<p>do you use some preprocess like <a href=\"https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping\">https://www.kaggle.com/ratthachat/aptos-updatedv14-preprocessing-ben-s-cropping</a>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 585876,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-07-28T08:30:52.267000",
          "content": "<p>Hi, DrHB, Is this CV score processed by optimized kappa, you choose a constant threshold?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 586126,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-28T16:45:28.020000",
          "content": "<p>no Optimized Kappa. Just standard thresholds [0.5, 1.5, 2.5, 3.5] =)</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 586917,
          "author_name": "Tahsin Mostafiz",
          "author_url": "",
          "post_date": "2019-07-29T21:51:15.653000",
          "content": "<p>Thanks for sharing :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 588927,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-07-31T08:29:28.297000",
          "content": "<p>Hi, DrHB. Did you just train 5 epochs(the lr is?) on 2019 datas with the pretrained model on 2015 data? </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 589050,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-31T11:46:08.373000",
          "content": "<p>yes! and my lr (learning rate) was 1e-3/2. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 589095,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-07-31T12:53:22.510000",
          "content": "<p>Thanks for your sharing.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 589103,
          "author_name": "DrHB",
          "author_url": "",
          "post_date": "2019-07-31T13:08:38.493000",
          "content": "<p>you are welcome</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 589248,
          "author_name": "Pi",
          "author_url": "",
          "post_date": "2019-07-31T17:04:05.333000",
          "content": "<p>Hi <a href=\"/dathudeptrai\">@dathudeptrai</a>, as you said in your experiments: \"CV: 92 -&gt;&gt;&gt; LB: 81.0\".\nMay I ask more about this experiment: Did you solve the problem with a regression model? Is CV a kappa score? If so, you set the thresholds as [0.5, 1.5, 2.5, 3.5] to calculate CV kappa or you use optimized thresholds?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590060,
          "author_name": "Drei",
          "author_url": "",
          "post_date": "2019-08-01T18:48:36.020000",
          "content": "<p>Interesting HB, I notice that you specified that you use the full image for pretraining on old competition data. Do you do any augments on that image, or do you just train on what's given?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590653,
          "author_name": "Carlo",
          "author_url": "",
          "post_date": "2019-08-02T12:49:10.617000",
          "content": "<p>Thanks for sharing your setup! It beautifully clears up the discussion of \"to augment or not to augment\".</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 596837,
          "author_name": "PGiN",
          "author_url": "",
          "post_date": "2019-08-11T11:43:05.413000",
          "content": "<p>Thanks for sharing! This thread is really insightful.\n&gt;Train on old dataset with valid. from current dataset\n&gt;Train on current dataset (weights from old) and valid. from current dataset</p>\n\n<p><a href=\"/drhabib\">@drhabib</a> \nI'm curious about how many epochs on the latter process. Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 598120,
          "author_name": "Hao He",
          "author_url": "",
          "post_date": "2019-08-13T06:50:24.670000",
          "content": "<p>Just curious, for the EfficientNet, where did you cut? I did some simple test for B0\n1. no cut\n2. cut at m._conv_head, so layer_groups now is 2\n3. cut at m._conv_head and m._blocks[8], so layer_groups now is 3</p>\n\n<p>for 3, the reason that cut at m._blocks[8]. Just observed the channel size is changed from 80-&gt;112 after two 80-&gt;80 blocks, and it is right in the middle of the m._blocks</p>\n\n<p>Exp 1 and 3 was trained for 5 epochs (without freezing), Exp 2 freezed last layer for 5 epochs, then unfreeze for another 8 epochs.  All with same argumentation and all with pretrained weights, and all used discriminative lr except for 1 the layer groups is 1</p>\n\n<p>The results are 1 &gt;= 3 &gt;&gt; 2.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 607744,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2019-08-25T20:00:41.687000",
          "content": "<p>can I do this in kaggle kernel (the old data training part)??</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 607755,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-08-25T20:36:20.593000",
          "content": "<p>yes. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 607959,
          "author_name": "Yangfan",
          "author_url": "",
          "post_date": "2019-08-26T06:37:08.620000",
          "content": "<p>What does CV mean?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 607994,
          "author_name": "Aptha K S",
          "author_url": "",
          "post_date": "2019-08-26T07:24:36.303000",
          "content": "<p>It takes 45 min for 1 epoch of 35000 images. Am I doing anything wrong?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 608031,
          "author_name": "Jayasooryan K V",
          "author_url": "",
          "post_date": "2019-08-26T08:29:39.117000",
          "content": "<p><a href=\"/apthagowda\">@apthagowda</a> It might be because of your preprocessing method which takes a lot of time.. Mostly the culprit will be resizing images. If that is the case, download the dataset and resize the images and upload it to kaggle. Using already resized images will reduce the epoch time. In my case an epoch takes around 10 minutes using resized dataset. Hope it helps :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 608081,
          "author_name": "Yangfan",
          "author_url": "",
          "post_date": "2019-08-26T09:58:41.677000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 582273,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-07-23T02:02:57.277000",
      "content": "<p>To me, I found a different evidence to <a href=\"/drhabib\">@drhabib</a> as my CV/LB is quite correlated;</p>\n\n<p>CV &gt;&gt; LB\n.92x &gt;&gt; 0.77x\n.93x &gt;&gt; 0.79x\n.94x &gt;&gt; 0.80x</p>\n\n<p>Perhaps because I use heavy preprocessing / augmentation from the start. (I shared my six augmentations in my 2nd kernel)\nAlso there still a large gap even though those preprocessing/augmentations.</p>\n\n<p>IMO, “trust your CV” is true if we are able to set a correct validation setting (which I still not be able to do yet :)</p>",
      "votes": 7,
      "replies": [
        {
          "id": 585801,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-07-28T05:08:38.590000",
          "content": "<p>HI, Is this CV score processed by optimized kappa?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 588223,
      "author_name": "Frank Jing",
      "author_url": "",
      "post_date": "2019-07-30T09:59:11.910000",
      "content": "<p><a href=\"/taindow\">@taindow</a> the \"CV\" you mentioned is \"Cross Validation\"? \nso far as i known, K-fold Cross Validation will generate K models, so the CV value here is the average of those K models?  just like Jason Huang said in <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98609#latest-568539\">https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98609#latest-568539</a>\n\"\nMake CV: \nSplit the data to K folds and train K models by using each fold as the validation dataset. Then you can get CV score by computing the average validation score of each fold.</p>\n\n<p>Ensemble the result from different models:\nLoad each model and give out their corresponding predictions on the test dataset, then you can ensemble them.\n\"\nam i right?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 588229,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-07-30T10:04:43.720000",
          "content": "<p>Yes</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 588517,
          "author_name": "Tom Aindow",
          "author_url": "",
          "post_date": "2019-07-30T17:37:21.070000",
          "content": "<p>This is correct :)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 581400,
      "author_name": "ilovescience",
      "author_url": "",
      "post_date": "2019-07-21T22:09:36.973000",
      "content": "<p>Another big problem is we have no idea what the private test distribution looks like and with these concerns about CV/LB discrepancies it is possible there may be a shake-up and it will be hard to prevent this.</p>\n\n<p>I have discussed this over <a href=\"https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493\">here</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 581148,
      "author_name": "哈尔的移动城堡",
      "author_url": "",
      "post_date": "2019-07-21T14:11:14.287000",
      "content": "<p>Oh  my  god , I  have  just  check  the  test  dataset ,  it's  quite   different  from  train  dataset ,  there  is  nearly  very  less  black   space   here ,   it   is   a  very  huge   gap   between  train_data  and     test_data </p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 581089,
      "author_name": "Rishabh Agrahari",
      "author_url": "",
      "post_date": "2019-07-21T11:41:48.663000",
      "content": "<p>That's a great finding. I've seen my LB improve after center cropping the retina. It's better to avoid the black region altogether, otherwise model is very prone to learn this weird correlation and fit itself on it.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 582124,
          "author_name": "Kartik Godawat",
          "author_url": "",
          "post_date": "2019-07-22T19:17:00.767000",
          "content": "<p>I also notice LB improvement with cropping, but even then local CV is not correlating. Are you seeing a good correlation? </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 582315,
          "author_name": "Rishabh Agrahari",
          "author_url": "",
          "post_date": "2019-07-23T03:45:50.300000",
          "content": "<p>Tbh, No.\nI expect a big shake-up after the private test results are announced. After doing ~60 submissions so far I can literally look at the unique counts of predictions and say which submission is gonna perform good on the LB. Which is definitely not a good approach, I don't trust local qwk anymore, As of now I'm monitoring a lot of other classification metrics, I'm trying to find a correlation between them and the LB, will share if I find something useful there.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 590270,
      "author_name": "Gary",
      "author_url": "",
      "post_date": "2019-08-02T01:27:58.433000",
      "content": "<p>Is this necessary for preprocessing? \nWhen I use preprocessing and removed the black pixels, I got a lower LB.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 590488,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2019-08-02T08:40:30.750000",
          "content": "<p>I got a higher LB score when I forgot to preprocess the test images in my inference kernel like I did when training.\nAlso resizing the images with OpenCV and PIL gave different scores</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590494,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-08-02T08:43:53.190000",
          "content": "<p>Different resizing on inference or also on train? On train it can be just the seed effect.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590521,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-02T09:35:11.213000",
          "content": "<p>@<a href=\"https://www.kaggle.com/suicaokhoailang\">Khoi Nguyen</a>What a wonderful coincidence..</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590522,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-02T09:36:17.447000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\">@Psi</a>I think I did the same preprocessing on train and test.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590533,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-08-02T09:54:24.390000",
          "content": "<p>Then I am not surprised. Resizing in cv2 vs Pil cannot have any impact.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590540,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-02T10:02:40.357000",
          "content": "<p>Psi, is that you current LB score through preprocessing?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590560,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-08-02T10:42:10.383000",
          "content": "<p>What do you mean?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590564,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-02T10:47:25.500000",
          "content": "<p>sorry, did you use ben preprocessing to get the current LB score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590587,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-08-02T11:19:04.587000",
          "content": "<p>Combination of few things.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 590603,
          "author_name": "Gary",
          "author_url": "",
          "post_date": "2019-08-02T11:43:46.460000",
          "content": "<p>yep, thank you.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 596946,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-08-11T15:02:55.663000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 607942,
      "author_name": "Josh Myers",
      "author_url": "",
      "post_date": "2019-08-26T05:32:31.970000",
      "content": "<p>Thanks for this. I'm wondering, when using a pre-trained model, should I freeze any layers, or just unfreeze all?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 590661,
      "author_name": "Quan",
      "author_url": "",
      "post_date": "2019-08-02T12:59:50.800000",
      "content": "<p>My theory: Pre-processing allows models to extract features much easier therefore learn much faster and models don't have to be too deep (I guess we need less layers to do feature extracting after pre-processing?)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 590213,
      "author_name": "Carlo",
      "author_url": "",
      "post_date": "2019-08-01T21:55:10.180000",
      "content": "<p>Ah ok! Thank you for clearing that up! Time to dive deeper into pre-processing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 584263,
      "author_name": "Ashwani Pandey",
      "author_url": "",
      "post_date": "2019-07-25T16:25:42.773000",
      "content": "<p>I really think, we should focus on our model's interpretability. If it picks correct features, it would be in quite good agreement with doctors. In real world scenario (after deployment), model will get single data point to predict on, so there will nothing be like \"Test Distribution\". And this is the reason, why I liked your kernel so much. Thanks.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 584269,
          "author_name": "Miroslav Valan",
          "author_url": "",
          "post_date": "2019-07-25T16:35:07.013000",
          "content": "<p>there** is** \"test distribution\" in real world scenario. and it is likely close to train set. Consider previous competition which also had train set  with similar distribution. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 582932,
      "author_name": "AndrioC",
      "author_url": "",
      "post_date": "2019-07-23T19:55:08.367000",
      "content": "<p>Great post!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 581950,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-22T15:24:30.060000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 581856,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-22T13:25:18.823000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 582668,
      "author_name": "saltboy",
      "author_url": "",
      "post_date": "2019-07-23T12:26:39.770000",
      "content": "<p>Thanks for the post</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 582005,
      "author_name": "Guillaume LG",
      "author_url": "",
      "post_date": "2019-07-22T16:44:20.847000",
      "content": "<p>Thanks for the post</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "581073": "I've seen a few posts now where people are seeing cv 0.80+ but submitting and receiving abysmal LB scores (sometimes even negative). For example:\n\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100680\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/100583\nhttps://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98194\n\nI believe a major part of this issue is down to the fact that image size and black space cropping are strongly indicative of diagnosis in the training data, but not in the public test data. \n\nI've put together a kernel showing how by **using just meta features you can score 0.70+ validation kappa locally, but get a negative LB score**:\n\n[Be careful what you train on ...](https://www.kaggle.com/taindow/be-careful-what-you-train-on)\n\nMy suggestion if you are having this problem is to focus on pre-processing and ensuring that the leaky image size and cropping information is not influencing your model. Pre-trained models will always help since we start from a set of robust features.\n\nI think we need to be very careful about how we handle this pre-processing, since it is going to have a large effect on local CV scores. \"Trust local CV\" is always the mantra, but sometimes it can lead you astray ...",
    "581136": "Great Post thank you for doing this! \n\nI have also noticed this was an issue so one way I overcome this was using a lot of Augmentations. Here is the results of some of my experiments. \n\n```\nsetup: \nRegression Problem\nIMG_SIZE: 224 \nMODEL: B3\nno TTA:\n5 fold CV (same across experiments) \n\nEDIT:\npertained on old competition data (train)(full not cropped). \noptimizer: ADAM\nmonitor: Valid Loss\nvalid_dataset: 2019\ntraining_phase_1: Load Image Net weights, freeze up to last layer train for 5 epoch \ntraining_phase_2: Unfreeze whole model and train for 15 epoch. \n\nfor the Experiments below. \nI load weights which were trained above unfreeze the model and train for 5 epochs, monitoring  valid loss \n```\n\nEXP_1:\n```\nno AUGMENT:\nCV: 0.93 LB: 0.78\n\n```\nEXP_2:\n```\nAUGMENT = [FLIPS]\nCV: 0.934, LB: 0.783\n```\n\n\nEXP_3:\n```\nAUGMENT = [rotation(360)]\nCV: 0.92 LB: 0.79\n\n```\nEXP4:\n```\nAUGMENT = [zoom_to_center_area(1. 3)]\nCV: 0.90 LB: 0.793\n```\n\nEXP5: \n```\nAUGMENT = [FLIP, ZOOM_TO_CENTER, ROTATE(360), CONTRAST]\nCV: 0.89 LB: 0.809 \n\n```\n\n",
    "582273": "To me, I found a different evidence to @drhabib as my CV/LB is quite correlated;\n\nCV &gt;&gt; LB\n.92x &gt;&gt; 0.77x\n.93x &gt;&gt; 0.79x\n.94x &gt;&gt; 0.80x\n\nPerhaps because I use heavy preprocessing / augmentation from the start. (I shared my six augmentations in my 2nd kernel)\nAlso there still a large gap even though those preprocessing/augmentations.\n\nIMO, “trust your CV” is true if we are able to set a correct validation setting (which I still not be able to do yet :)",
    "588223": "@taindow the \"CV\" you mentioned is \"Cross Validation\"? \nso far as i known, K-fold Cross Validation will generate K models, so the CV value here is the average of those K models?  just like Jason Huang said in https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98609#latest-568539\n\"\nMake CV: \nSplit the data to K folds and train K models by using each fold as the validation dataset. Then you can get CV score by computing the average validation score of each fold.\n\nEnsemble the result from different models:\nLoad each model and give out their corresponding predictions on the test dataset, then you can ensemble them.\n\"\nam i right?",
    "581400": "Another big problem is we have no idea what the private test distribution looks like and with these concerns about CV/LB discrepancies it is possible there may be a shake-up and it will be hard to prevent this.\n\nI have discussed this over [here](https://www.kaggle.com/c/aptos2019-blindness-detection/discussion/98493)",
    "581148": "Oh  my  god , I  have  just  check  the  test  dataset ,  it's  quite   different  from  train  dataset ,  there  is  nearly  very  less  black   space   here ,   it   is   a  very  huge   gap   between  train_data  and     test_data ",
    "581089": "That's a great finding. I've seen my LB improve after center cropping the retina. It's better to avoid the black region altogether, otherwise model is very prone to learn this weird correlation and fit itself on it.",
    "590270": "Is this necessary for preprocessing? \nWhen I use preprocessing and removed the black pixels, I got a lower LB.",
    "607942": "Thanks for this. I'm wondering, when using a pre-trained model, should I freeze any layers, or just unfreeze all?",
    "590661": "My theory: Pre-processing allows models to extract features much easier therefore learn much faster and models don't have to be too deep (I guess we need less layers to do feature extracting after pre-processing?)",
    "590213": "Ah ok! Thank you for clearing that up! Time to dive deeper into pre-processing!",
    "584263": "I really think, we should focus on our model's interpretability. If it picks correct features, it would be in quite good agreement with doctors. In real world scenario (after deployment), model will get single data point to predict on, so there will nothing be like \"Test Distribution\". And this is the reason, why I liked your kernel so much. Thanks.",
    "582932": "Great post!",
    "581950": "",
    "581856": "",
    "582668": "Thanks for the post",
    "582005": "Thanks for the post"
  }
}