{
  "id": 69462,
  "title": "Has anyone noticed a ~35% drop from local validation score to LB score?",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/69462",
  "author_name": "",
  "post_date": "2018-10-24T00:14:15.102600800Z",
  "votes": 14,
  "comment_count": 29,
  "views": 0,
  "content": "<p>All my submissions with local validation around 0.630, their LB scores are usually only around 0.400.</p>\n\n<p>I have confirmed that</p>\n\n<ul>\n<li>my training process did not peek into the validation dataset, there's no cheating</li>\n<li>my macro f1 score formula is correct, and is calculated on the entire validation dataset (as opposed to on separate batches and then averaged). I tried sklearn's <code>f1_score(y_true, y_pred, average='macro')</code>, and also the one posted <a href=\"https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras\">here</a></li>\n<li>the order of ids in my submission is the same as that of the <code>sample_submission.csv</code></li>\n<li>my validation dataset has enough number of samples (11,000+)</li>\n</ul>\n\n<p>Seems like <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">iafoss's kernel also suffers a drop from ~0.540 to ~0.432</a></p>",
  "messages": [
    {
      "id": "409198",
      "postDate": "10/24/2018 00:14:15",
      "content": "<p>All my submissions with local validation around 0.630, their LB scores are usually only around 0.400.</p>\n\n<p>I have confirmed that</p>\n\n<ul>\n<li>my training process did not peek into the validation dataset, there's no cheating</li>\n<li>my macro f1 score formula is correct, and is calculated on the entire validation dataset (as opposed to on separate batches and then averaged). I tried sklearn's <code>f1_score(y_true, y_pred, average='macro')</code>, and also the one posted <a href=\"https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras\">here</a></li>\n<li>the order of ids in my submission is the same as that of the <code>sample_submission.csv</code></li>\n<li>my validation dataset has enough number of samples (11,000+)</li>\n</ul>\n\n<p>Seems like <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">iafoss's kernel also suffers a drop from ~0.540 to ~0.432</a></p>",
      "rawMarkdown": "All my submissions with local validation around 0.630, their LB scores are usually only around 0.400.\n\nI have confirmed that\n\n  * my training process did not peek into the validation dataset, there's no cheating\n  * my macro f1 score formula is correct, and is calculated on the entire validation dataset (as opposed to on separate batches and then averaged). I tried sklearn's `f1_score(y_true, y_pred, average='macro')`, and also the one posted [here](https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras)\n  * the order of ids in my submission is the same as that of the `sample_submission.csv`\n  * my validation dataset has enough number of samples (11,000+)\n\nSeems like [iafoss's kernel also suffers a drop from ~0.540 to ~0.432](https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai)",
      "votes": null
    },
    {
      "id": "409286",
      "postDate": "10/24/2018 03:59:11",
      "content": "<p>I just set up a very first and simple baseline model - and get the same drop (about 40 percent). My only explanation is that the distribution of classes is different in the test set. If you look at </p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678#404591\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678#404591</a></p>\n\n<p>the class distribution for the test set has been worked out. And there are differences.</p>",
      "rawMarkdown": "I just set up a very first and simple baseline model - and get the same drop (about 40 percent). My only explanation is that the distribution of classes is different in the test set. If you look at \n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678#404591\n\nthe class distribution for the test set has been worked out. And there are differences.",
      "votes": null
    },
    {
      "id": "409301",
      "postDate": "10/24/2018 04:49:42",
      "content": "<p>Actually even pixel statistics is different for train and test. In the first version of my kernel I checked it and got the following values for mean and standard deviation in each channel:</p>\n\n<ul>\n<li>train: (array([0.08069, 0.05258, 0.05487, 0.08282]), array([0.13704,0.10145, 0.15313, 0.13814])) </li>\n<li>test: (array([0.05913, 0.0454 , 0.04066, 0.05928]), array([0.11734, 0.09503, 0.129 , 0.11528]))</li>\n</ul>",
      "rawMarkdown": "Actually even pixel statistics is different for train and test. In the first version of my kernel I checked it and got the following values for mean and standard deviation in each channel:\n\n - train: (array([0.08069, 0.05258, 0.05487, 0.08282]), array([0.13704,0.10145, 0.15313, 0.13814])) \n - test: (array([0.05913, 0.0454 , 0.04066, 0.05928]), array([0.11734, 0.09503, 0.129 , 0.11528]))",
      "votes": null
    },
    {
      "id": "409304",
      "postDate": "10/24/2018 04:55:39",
      "content": "<p>What are your ideas to handle this gap?</p>",
      "rawMarkdown": "What are your ideas to handle this gap?",
      "votes": null
    },
    {
      "id": "409313",
      "postDate": "10/24/2018 05:31:11",
      "content": "<p>My LB follows this pattern since the beginning:\nLB = 0.64 * CV +0.04\nSo yes, I can confirm similar drop, too.</p>",
      "rawMarkdown": "My LB follows this pattern since the beginning:\nLB = 0.64 * CV +0.04\nSo yes, I can confirm similar drop, too.",
      "votes": null
    },
    {
      "id": "409353",
      "postDate": "10/24/2018 06:54:53",
      "content": "<p>Your pattern is so interesting. It is true for me too.</p>",
      "rawMarkdown": "Your pattern is so interesting. It is true for me too.",
      "votes": null
    },
    {
      "id": "409606",
      "postDate": "10/24/2018 15:09:45",
      "content": "<p>Hi lafoss, \njust out of interest: How did you calculate these values? I am asking because I did the same a couple of days ago, on the original 512x512 images and got slightly different results, i.e.:\nMeans for train image data (originals)</p>\n\n<p>Red average:    0.080441904331346\nGreen average:  0.05262986230955176\nBlue average:   0.05474700710311806\nYellow average: 0.08270895676048498</p>\n\n<p>Means for test image data (originals)</p>\n\n<p>Red average:    0.05908022413399168\nGreen average:  0.04532851916280794\nBlue average:   0.040652325092460015\nYellow average: 0.05923425759572161</p>\n\n<p>Did you resize the images before checking the means? \nAs I say, just out of interest, \ncheers and thanks, \nWolfgang</p>",
      "rawMarkdown": "Hi lafoss, \njust out of interest: How did you calculate these values? I am asking because I did the same a couple of days ago, on the original 512x512 images and got slightly different results, i.e.:\nMeans for train image data (originals)\n\nRed average:    0.080441904331346\nGreen average:  0.05262986230955176\nBlue average:   0.05474700710311806\nYellow average: 0.08270895676048498\n\nMeans for test image data (originals)\n\nRed average:    0.05908022413399168\nGreen average:  0.04532851916280794\nBlue average:   0.040652325092460015\nYellow average: 0.05923425759572161\n\nDid you resize the images before checking the means? \nAs I say, just out of interest, \ncheers and thanks, \nWolfgang",
      "votes": null
    },
    {
      "id": "409610",
      "postDate": "10/24/2018 15:26:40",
      "content": "<p>Yes, it is in the first version of <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-448-public-lb\">the kernel</a>. I resized images to 256 and converted the channels to [0.0, 1.0] range before calculating the statistics . </p>\n\n<p>I found a problem there, my bad. My mean values must be multiplied by 16, and standard deviation should be recalculated as well. But after this correction, your results are the same as my ones. I will add corrected values to my post, thank you for pointing it out.</p>",
      "rawMarkdown": "Yes, it is in the first version of [the kernel][1]. I resized images to 256 and converted the channels to [0.0, 1.0] range before calculating the statistics . \n\nI found a problem there, my bad. My mean values must be multiplied by 16, and standard deviation should be recalculated as well. But after this correction, your results are the same as my ones. I will add corrected values to my post, thank you for pointing it out.\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-448-public-lb",
      "votes": null
    },
    {
      "id": "410278",
      "postDate": "10/25/2018 18:16:47",
      "content": "<p>+1 - I have noticed score drop too</p>",
      "rawMarkdown": "1 - I have noticed score drop too",
      "votes": null
    },
    {
      "id": "410471",
      "postDate": "10/26/2018 05:27:56",
      "content": "<p>I see a drop when I use sklearn f1_score over my entire validation set. One interesting point is that my F1 per batch is much closer to the public LB.</p>",
      "rawMarkdown": "I see a drop when I use sklearn f1_score over my entire validation set. One interesting point is that my F1 per batch is much closer to the public LB.",
      "votes": null
    },
    {
      "id": "410555",
      "postDate": "10/26/2018 08:52:09",
      "content": "<p>Well within your batch some classes may not be present and the f1 score is then (depending on the implementation) 0. Thus your batch f1 score is a pessimistic estimate</p>",
      "rawMarkdown": "Well within your batch some classes may not be present and the f1 score is then (depending on the implementation) 0. Thus your batch f1 score is a pessimistic estimate",
      "votes": null
    },
    {
      "id": "410755",
      "postDate": "10/26/2018 15:53:27",
      "content": "<p>That makes sense. It does go up/down with the batch size. I still see the same overall score drop. One model that locally was 0.69 with sklearn macro came out to 0.453 on the public LB.</p>",
      "rawMarkdown": "That makes sense. It does go up/down with the batch size. I still see the same overall score drop. One model that locally was 0.69 with sklearn macro came out to 0.453 on the public LB.",
      "votes": null
    },
    {
      "id": "410761",
      "postDate": "10/26/2018 16:12:04",
      "content": "<p>Lucky you my 0.73 was a 0.346 on the LB. I will stop using that leaderboard now and just use the training data -&gt; Let's see where this ends</p>",
      "rawMarkdown": "Lucky you my 0.73 was a 0.346 on the LB. I will stop using that leaderboard now and just use the training data -&gt; Let's see where this ends",
      "votes": null
    },
    {
      "id": "410776",
      "postDate": "10/26/2018 16:30:26",
      "content": "<p>Are you searching for an optimal threshold? I picked up a bit just from doing that better. Based on everyone having the same issue. I have a feeling that the public vs private LB was split with some imbalance.</p>",
      "rawMarkdown": "Are you searching for an optimal threshold? I picked up a bit just from doing that better. Based on everyone having the same issue. I have a feeling that the public vs private LB was split with some imbalance.",
      "votes": null
    },
    {
      "id": "410777",
      "postDate": "10/26/2018 16:38:49",
      "content": "<p>Any threshold you optimize on the public LB will not generalize to the held out test set - so why bother? I might play with thresholds based on my cross-validation on the training data but only at the very end of the submission phase</p>",
      "rawMarkdown": "Any threshold you optimize on the public LB will not generalize to the held out test set - so why bother? I might play with thresholds based on my cross-validation on the training data but only at the very end of the submission phase",
      "votes": null
    },
    {
      "id": "410778",
      "postDate": "10/26/2018 16:40:16",
      "content": "<p>Here is what my metrics are on the most recent epoch that finished. This model is showing good progress, running 512x512x3 RGB with a batch size of 30. Still at least a day or two left of training, it is the one scoring 0.453. All training metrics are looking nice for this one so far, I've got my fingers crossed.</p>\n\n<pre>Epoch 5/2000\n4154/4154 [==============================] - 1460s 351ms/step - loss: 3.3850 - weighted_binary_crossentropy: 0.3285 - acc: 0.6499 - f1: 0.4167 - val_loss: 3.2685 - val_weighted_binary_crossentropy: 0.2365 - val_acc: 0.5959 - val_f1: 0.4327\n\nEpoch 00005: val_loss improved from 3.26920 to 3.26846, saving model to hpa_sub_16.h5\nval_f1: 0.697597 — val_precision: 0.639676 — val_recall 0.839397\n</pre>",
      "rawMarkdown": "Here is what my metrics are on the most recent epoch that finished. This model is showing good progress, running 512x512x3 RGB with a batch size of 30. Still at least a day or two left of training, it is the one scoring 0.453. All training metrics are looking nice for this one so far, I've got my fingers crossed.\n\n<pre>Epoch 5/2000\n4154/4154 [==============================] - 1460s 351ms/step - loss: 3.3850 - weighted_binary_crossentropy: 0.3285 - acc: 0.6499 - f1: 0.4167 - val_loss: 3.2685 - val_weighted_binary_crossentropy: 0.2365 - val_acc: 0.5959 - val_f1: 0.4327\n\nEpoch 00005: val_loss improved from 3.26920 to 3.26846, saving model to hpa_sub_16.h5\nval_f1: 0.697597 — val_precision: 0.639676 — val_recall 0.839397\n</pre>",
      "votes": null
    },
    {
      "id": "410897",
      "postDate": "10/26/2018 21:01:38",
      "content": "<p>Fingers crossed and the best of luck!\nIs that some kind of expanded pretrained model? Or do you train from scratch?</p>",
      "rawMarkdown": "Fingers crossed and the best of luck!\nIs that some kind of expanded pretrained model? Or do you train from scratch?",
      "votes": null
    },
    {
      "id": "410901",
      "postDate": "10/26/2018 21:10:02",
      "content": "<p>Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.</p>\n\n<p>4 hours later, the latest sklearn macro scores are:</p>\n\n<pre>val_f1: 0.712903 — val_precision: 0.663954 — val_recall 0.841232\n</pre>",
      "rawMarkdown": "Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.\n\n4 hours later, the latest sklearn macro scores are:\n<pre>val_f1: 0.712903 — val_precision: 0.663954 — val_recall 0.841232\n</pre>",
      "votes": null
    },
    {
      "id": "411192",
      "postDate": "10/27/2018 15:38:39",
      "content": "<p>Not bad! Sounds pretty good =)\nJust out of curiosity (you dont need to answer if you dont want to) but how are you training with a dense output if the labels are just one per image?</p>",
      "rawMarkdown": "Not bad! Sounds pretty good =)\nJust out of curiosity (you dont need to answer if you dont want to) but how are you training with a dense output if the labels are just one per image?",
      "votes": null
    },
    {
      "id": "411340",
      "postDate": "10/27/2018 22:07:46",
      "content": "<p>The last dense layer has 28 units with sigmoid activation.  I use a global average pooling layer to flatten the last conv block before the dense layers.</p>",
      "rawMarkdown": "The last dense layer has 28 units with sigmoid activation.  I use a global average pooling layer to flatten the last conv block before the dense layers.",
      "votes": null
    },
    {
      "id": "411460",
      "postDate": "10/28/2018 07:53:54",
      "content": "<p>Oh alright thats what you meant by dense. For some reason I was thinking about a spatial output (as in segmentation) and I was wondering how that would be possible</p>",
      "rawMarkdown": "Oh alright thats what you meant by dense. For some reason I was thinking about a spatial output (as in segmentation) and I was wondering how that would be possible",
      "votes": null
    },
    {
      "id": "411549",
      "postDate": "10/28/2018 12:50:50",
      "content": "<p>how do you know that metrics used by public Kaggle kernels is same as metrics used in competition?</p>",
      "rawMarkdown": "how do you know that metrics used by public Kaggle kernels is same as metrics used in competition?",
      "votes": null
    },
    {
      "id": "411776",
      "postDate": "10/29/2018 01:33:40",
      "content": "<p>So far the discrepancy got worse as the score went up. </p>",
      "rawMarkdown": "So far the discrepancy got worse as the score went up.",
      "votes": null
    },
    {
      "id": "411928",
      "postDate": "10/29/2018 08:21:28",
      "content": "<p>Macro F1 score is pretty unambiguous: <a href=\"https://datascience.stackexchange.com/questions/15989/micro-average-vs-macro-average-performance-in-a-multiclass-classification-settin/16001\">https://datascience.stackexchange.com/questions/15989/micro-average-vs-macro-average-performance-in-a-multiclass-classification-settin/16001</a>\nCompute the F1 score for each class, then average over the classes</p>",
      "rawMarkdown": "Macro F1 score is pretty unambiguous: https://datascience.stackexchange.com/questions/15989/micro-average-vs-macro-average-performance-in-a-multiclass-classification-settin/16001\nCompute the F1 score for each class, then average over the classes",
      "votes": null
    },
    {
      "id": "412083",
      "postDate": "10/29/2018 14:48:34",
      "content": "<p>Macro F1 calculates metrics for each label, and find their unweighted mean. (7 missing classes in LB) / (28 total classes) = 0.25, and if the organizer is interpreting 0/0 as 0, this explains 0.25 of the LB drop. The other 0.1 could be that public LB has more hard examples. This made me suspicious of if the public/private LB split is truly random. It is possible that private dataset has more balanced classes.</p>",
      "rawMarkdown": "Macro F1 calculates metrics for each label, and find their unweighted mean. (7 missing classes in LB) / (28 total classes) = 0.25, and if the organizer is interpreting 0/0 as 0, this explains 0.25 of the LB drop. The other 0.1 could be that public LB has more hard examples. This made me suspicious of if the public/private LB split is truly random. It is possible that private dataset has more balanced classes.",
      "votes": null
    },
    {
      "id": "412103",
      "postDate": "10/29/2018 15:10:40",
      "content": "<p>I fully agree and for my part will run all optimizations on a cross-validation on the training data rather than concentrating on the leaderboard only. This could go really well, or be a very bad idea. But I believe it is the right thing to do here</p>",
      "rawMarkdown": "I fully agree and for my part will run all optimizations on a cross-validation on the training data rather than concentrating on the leaderboard only. This could go really well, or be a very bad idea. But I believe it is the right thing to do here",
      "votes": null
    },
    {
      "id": "412863",
      "postDate": "10/30/2018 23:09:06",
      "content": "<blockquote>\n  <p>Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.</p>\n</blockquote>\n\n<p>Did you try cosine-annealing? </p>",
      "rawMarkdown": "&gt; Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.\n\nDid you try cosine-annealing?",
      "votes": null
    },
    {
      "id": "412909",
      "postDate": "10/31/2018 01:32:52",
      "content": "<p>Yes, using SGD with cosine annealing schedule. Also used Adadelta to start training, Padam for mid training, and SGD at the end. Then I freeze parts of the model and train the other layers. My current leading model is 2.3M params. Performs great locally, but public LB is 45% lower.</p>",
      "rawMarkdown": "Yes, using SGD with cosine annealing schedule. Also used Adadelta to start training, Padam for mid training, and SGD at the end. Then I freeze parts of the model and train the other layers. My current leading model is 2.3M params. Performs great locally, but public LB is 45% lower.",
      "votes": null
    },
    {
      "id": "437197",
      "postDate": "12/11/2018 14:28:51",
      "content": "<p>Hello, how did you fix the gap between local lb and public lb?</p>",
      "rawMarkdown": "Hello, how did you fix the gap between local lb and public lb?",
      "votes": null
    },
    {
      "id": "437286",
      "postDate": "12/11/2018 17:15:07",
      "content": "<p>Hi Alexander Liao,\nCan you explain how you determined that there are 7 missing classes in LB?\nThanks.\nAh, I see that it was done by LB probing, which however is unreliable because of score truncation to 3 digits.</p>",
      "rawMarkdown": "Hi Alexander Liao,\nCan you explain how you determined that there are 7 missing classes in LB?\nThanks.\nAh, I see that it was done by LB probing, which however is unreliable because of score truncation to 3 digits.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 409286,
      "author_name": "wreuter",
      "author_url": "",
      "post_date": "10/24/2018 03:59:11",
      "content": "<p>I just set up a very first and simple baseline model - and get the same drop (about 40 percent). My only explanation is that the distribution of classes is different in the test set. If you look at </p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678#404591\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678#404591</a></p>\n\n<p>the class distribution for the test set has been worked out. And there are differences.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 409301,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "10/24/2018 04:49:42",
      "content": "<p>Actually even pixel statistics is different for train and test. In the first version of my kernel I checked it and got the following values for mean and standard deviation in each channel:</p>\n\n<ul>\n<li>train: (array([0.08069, 0.05258, 0.05487, 0.08282]), array([0.13704,0.10145, 0.15313, 0.13814])) </li>\n<li>test: (array([0.05913, 0.0454 , 0.04066, 0.05928]), array([0.11734, 0.09503, 0.129 , 0.11528]))</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 409606,
          "author_name": "wreuter",
          "author_url": "",
          "post_date": "10/24/2018 15:09:45",
          "content": "<p>Hi lafoss, \njust out of interest: How did you calculate these values? I am asking because I did the same a couple of days ago, on the original 512x512 images and got slightly different results, i.e.:\nMeans for train image data (originals)</p>\n\n<p>Red average:    0.080441904331346\nGreen average:  0.05262986230955176\nBlue average:   0.05474700710311806\nYellow average: 0.08270895676048498</p>\n\n<p>Means for test image data (originals)</p>\n\n<p>Red average:    0.05908022413399168\nGreen average:  0.04532851916280794\nBlue average:   0.040652325092460015\nYellow average: 0.05923425759572161</p>\n\n<p>Did you resize the images before checking the means? \nAs I say, just out of interest, \ncheers and thanks, \nWolfgang</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 409610,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/24/2018 15:26:40",
          "content": "<p>Yes, it is in the first version of <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-448-public-lb\">the kernel</a>. I resized images to 256 and converted the channels to [0.0, 1.0] range before calculating the statistics . </p>\n\n<p>I found a problem there, my bad. My mean values must be multiplied by 16, and standard deviation should be recalculated as well. But after this correction, your results are the same as my ones. I will add corrected values to my post, thank you for pointing it out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409304,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "10/24/2018 04:55:39",
      "content": "<p>What are your ideas to handle this gap?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 409313,
      "author_name": "rejpalcz",
      "author_url": "",
      "post_date": "10/24/2018 05:31:11",
      "content": "<p>My LB follows this pattern since the beginning:\nLB = 0.64 * CV +0.04\nSo yes, I can confirm similar drop, too.</p>",
      "votes": null,
      "replies": [
        {
          "id": 409353,
          "author_name": "backaggle",
          "author_url": "",
          "post_date": "10/24/2018 06:54:53",
          "content": "<p>Your pattern is so interesting. It is true for me too.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410278,
      "author_name": "pinullmezon",
      "author_url": "",
      "post_date": "10/25/2018 18:16:47",
      "content": "<p>+1 - I have noticed score drop too</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 410471,
      "author_name": "ldm314",
      "author_url": "",
      "post_date": "10/26/2018 05:27:56",
      "content": "<p>I see a drop when I use sklearn f1_score over my entire validation set. One interesting point is that my F1 per batch is much closer to the public LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 410555,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/26/2018 08:52:09",
          "content": "<p>Well within your batch some classes may not be present and the f1 score is then (depending on the implementation) 0. Thus your batch f1 score is a pessimistic estimate</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410755,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/26/2018 15:53:27",
          "content": "<p>That makes sense. It does go up/down with the batch size. I still see the same overall score drop. One model that locally was 0.69 with sklearn macro came out to 0.453 on the public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410761,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/26/2018 16:12:04",
          "content": "<p>Lucky you my 0.73 was a 0.346 on the LB. I will stop using that leaderboard now and just use the training data -&gt; Let's see where this ends</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410776,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/26/2018 16:30:26",
          "content": "<p>Are you searching for an optimal threshold? I picked up a bit just from doing that better. Based on everyone having the same issue. I have a feeling that the public vs private LB was split with some imbalance.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410777,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/26/2018 16:38:49",
          "content": "<p>Any threshold you optimize on the public LB will not generalize to the held out test set - so why bother? I might play with thresholds based on my cross-validation on the training data but only at the very end of the submission phase</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410778,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/26/2018 16:40:16",
          "content": "<p>Here is what my metrics are on the most recent epoch that finished. This model is showing good progress, running 512x512x3 RGB with a batch size of 30. Still at least a day or two left of training, it is the one scoring 0.453. All training metrics are looking nice for this one so far, I've got my fingers crossed.</p>\n\n<pre>Epoch 5/2000\n4154/4154 [==============================] - 1460s 351ms/step - loss: 3.3850 - weighted_binary_crossentropy: 0.3285 - acc: 0.6499 - f1: 0.4167 - val_loss: 3.2685 - val_weighted_binary_crossentropy: 0.2365 - val_acc: 0.5959 - val_f1: 0.4327\n\nEpoch 00005: val_loss improved from 3.26920 to 3.26846, saving model to hpa_sub_16.h5\nval_f1: 0.697597 — val_precision: 0.639676 — val_recall 0.839397\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410897,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/26/2018 21:01:38",
          "content": "<p>Fingers crossed and the best of luck!\nIs that some kind of expanded pretrained model? Or do you train from scratch?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 410901,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/26/2018 21:10:02",
          "content": "<p>Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.</p>\n\n<p>4 hours later, the latest sklearn macro scores are:</p>\n\n<pre>val_f1: 0.712903 — val_precision: 0.663954 — val_recall 0.841232\n</pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411192,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/27/2018 15:38:39",
          "content": "<p>Not bad! Sounds pretty good =)\nJust out of curiosity (you dont need to answer if you dont want to) but how are you training with a dense output if the labels are just one per image?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411340,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/27/2018 22:07:46",
          "content": "<p>The last dense layer has 28 units with sigmoid activation.  I use a global average pooling layer to flatten the last conv block before the dense layers.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411460,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/28/2018 07:53:54",
          "content": "<p>Oh alright thats what you meant by dense. For some reason I was thinking about a spatial output (as in segmentation) and I was wondering how that would be possible</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 411776,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/29/2018 01:33:40",
          "content": "<p>So far the discrepancy got worse as the score went up. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 412863,
          "author_name": "kokecacao",
          "author_url": "",
          "post_date": "10/30/2018 23:09:06",
          "content": "<blockquote>\n  <p>Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.</p>\n</blockquote>\n\n<p>Did you try cosine-annealing? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 412909,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "10/31/2018 01:32:52",
          "content": "<p>Yes, using SGD with cosine annealing schedule. Also used Adadelta to start training, Padam for mid training, and SGD at the end. Then I freeze parts of the model and train the other layers. My current leading model is 2.3M params. Performs great locally, but public LB is 45% lower.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 411549,
      "author_name": "jacekpoplawski",
      "author_url": "",
      "post_date": "10/28/2018 12:50:50",
      "content": "<p>how do you know that metrics used by public Kaggle kernels is same as metrics used in competition?</p>",
      "votes": null,
      "replies": [
        {
          "id": 411928,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/29/2018 08:21:28",
          "content": "<p>Macro F1 score is pretty unambiguous: <a href=\"https://datascience.stackexchange.com/questions/15989/micro-average-vs-macro-average-performance-in-a-multiclass-classification-settin/16001\">https://datascience.stackexchange.com/questions/15989/micro-average-vs-macro-average-performance-in-a-multiclass-classification-settin/16001</a>\nCompute the F1 score for each class, then average over the classes</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 412083,
      "author_name": "alexanderliao",
      "author_url": "",
      "post_date": "10/29/2018 14:48:34",
      "content": "<p>Macro F1 calculates metrics for each label, and find their unweighted mean. (7 missing classes in LB) / (28 total classes) = 0.25, and if the organizer is interpreting 0/0 as 0, this explains 0.25 of the LB drop. The other 0.1 could be that public LB has more hard examples. This made me suspicious of if the public/private LB split is truly random. It is possible that private dataset has more balanced classes.</p>",
      "votes": null,
      "replies": [
        {
          "id": 412103,
          "author_name": "fabianisensee",
          "author_url": "",
          "post_date": "10/29/2018 15:10:40",
          "content": "<p>I fully agree and for my part will run all optimizations on a cross-validation on the training data rather than concentrating on the leaderboard only. This could go really well, or be a very bad idea. But I believe it is the right thing to do here</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437286,
          "author_name": "dslate",
          "author_url": "",
          "post_date": "12/11/2018 17:15:07",
          "content": "<p>Hi Alexander Liao,\nCan you explain how you determined that there are 7 missing classes in LB?\nThanks.\nAh, I see that it was done by LB probing, which however is unreliable because of score truncation to 3 digits.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437197,
      "author_name": "lianandrew",
      "author_url": "",
      "post_date": "12/11/2018 14:28:51",
      "content": "<p>Hello, how did you fix the gap between local lb and public lb?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "409198": "All my submissions with local validation around 0.630, their LB scores are usually only around 0.400.\n\nI have confirmed that\n\n  * my training process did not peek into the validation dataset, there's no cheating\n  * my macro f1 score formula is correct, and is calculated on the entire validation dataset (as opposed to on separate batches and then averaged). I tried sklearn's `f1_score(y_true, y_pred, average='macro')`, and also the one posted [here](https://www.kaggle.com/guglielmocamporese/macro-f1-score-keras)\n  * the order of ids in my submission is the same as that of the `sample_submission.csv`\n  * my validation dataset has enough number of samples (11,000+)\n\nSeems like [iafoss's kernel also suffers a drop from ~0.540 to ~0.432](https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai)",
    "409286": "I just set up a very first and simple baseline model - and get the same drop (about 40 percent). My only explanation is that the distribution of classes is different in the test set. If you look at \n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/68678#404591\n\nthe class distribution for the test set has been worked out. And there are differences.",
    "409301": "Actually even pixel statistics is different for train and test. In the first version of my kernel I checked it and got the following values for mean and standard deviation in each channel:\n\n - train: (array([0.08069, 0.05258, 0.05487, 0.08282]), array([0.13704,0.10145, 0.15313, 0.13814])) \n - test: (array([0.05913, 0.0454 , 0.04066, 0.05928]), array([0.11734, 0.09503, 0.129 , 0.11528]))",
    "409304": "What are your ideas to handle this gap?",
    "409313": "My LB follows this pattern since the beginning:\nLB = 0.64 * CV +0.04\nSo yes, I can confirm similar drop, too.",
    "409353": "Your pattern is so interesting. It is true for me too.",
    "409606": "Hi lafoss, \njust out of interest: How did you calculate these values? I am asking because I did the same a couple of days ago, on the original 512x512 images and got slightly different results, i.e.:\nMeans for train image data (originals)\n\nRed average:    0.080441904331346\nGreen average:  0.05262986230955176\nBlue average:   0.05474700710311806\nYellow average: 0.08270895676048498\n\nMeans for test image data (originals)\n\nRed average:    0.05908022413399168\nGreen average:  0.04532851916280794\nBlue average:   0.040652325092460015\nYellow average: 0.05923425759572161\n\nDid you resize the images before checking the means? \nAs I say, just out of interest, \ncheers and thanks, \nWolfgang",
    "409610": "Yes, it is in the first version of [the kernel][1]. I resized images to 256 and converted the channels to [0.0, 1.0] range before calculating the statistics . \n\nI found a problem there, my bad. My mean values must be multiplied by 16, and standard deviation should be recalculated as well. But after this correction, your results are the same as my ones. I will add corrected values to my post, thank you for pointing it out.\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-448-public-lb",
    "410278": "1 - I have noticed score drop too",
    "410471": "I see a drop when I use sklearn f1_score over my entire validation set. One interesting point is that my F1 per batch is much closer to the public LB.",
    "410555": "Well within your batch some classes may not be present and the f1 score is then (depending on the implementation) 0. Thus your batch f1 score is a pessimistic estimate",
    "410755": "That makes sense. It does go up/down with the batch size. I still see the same overall score drop. One model that locally was 0.69 with sklearn macro came out to 0.453 on the public LB.",
    "410761": "Lucky you my 0.73 was a 0.346 on the LB. I will stop using that leaderboard now and just use the training data -&gt; Let's see where this ends",
    "410776": "Are you searching for an optimal threshold? I picked up a bit just from doing that better. Based on everyone having the same issue. I have a feeling that the public vs private LB was split with some imbalance.",
    "410777": "Any threshold you optimize on the public LB will not generalize to the held out test set - so why bother? I might play with thresholds based on my cross-validation on the training data but only at the very end of the submission phase",
    "410778": "Here is what my metrics are on the most recent epoch that finished. This model is showing good progress, running 512x512x3 RGB with a batch size of 30. Still at least a day or two left of training, it is the one scoring 0.453. All training metrics are looking nice for this one so far, I've got my fingers crossed.\n\n<pre>Epoch 5/2000\n4154/4154 [==============================] - 1460s 351ms/step - loss: 3.3850 - weighted_binary_crossentropy: 0.3285 - acc: 0.6499 - f1: 0.4167 - val_loss: 3.2685 - val_weighted_binary_crossentropy: 0.2365 - val_acc: 0.5959 - val_f1: 0.4327\n\nEpoch 00005: val_loss improved from 3.26920 to 3.26846, saving model to hpa_sub_16.h5\nval_f1: 0.697597 — val_precision: 0.639676 — val_recall 0.839397\n</pre>",
    "410897": "Fingers crossed and the best of luck!\nIs that some kind of expanded pretrained model? Or do you train from scratch?",
    "410901": "Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.\n\n4 hours later, the latest sklearn macro scores are:\n<pre>val_f1: 0.712903 — val_precision: 0.663954 — val_recall 0.841232\n</pre>",
    "411192": "Not bad! Sounds pretty good =)\nJust out of curiosity (you dont need to answer if you dont want to) but how are you training with a dense output if the labels are just one per image?",
    "411340": "The last dense layer has 28 units with sigmoid activation.  I use a global average pooling layer to flatten the last conv block before the dense layers.",
    "411460": "Oh alright thats what you meant by dense. For some reason I was thinking about a spatial output (as in segmentation) and I was wondering how that would be possible",
    "411549": "how do you know that metrics used by public Kaggle kernels is same as metrics used in competition?",
    "411776": "So far the discrepancy got worse as the score went up.",
    "411928": "Macro F1 score is pretty unambiguous: https://datascience.stackexchange.com/questions/15989/micro-average-vs-macro-average-performance-in-a-multiclass-classification-settin/16001\nCompute the F1 score for each class, then average over the classes",
    "412083": "Macro F1 calculates metrics for each label, and find their unweighted mean. (7 missing classes in LB) / (28 total classes) = 0.25, and if the organizer is interpreting 0/0 as 0, this explains 0.25 of the LB drop. The other 0.1 could be that public LB has more hard examples. This made me suspicious of if the public/private LB split is truly random. It is possible that private dataset has more balanced classes.",
    "412103": "I fully agree and for my part will run all optimizations on a cross-validation on the training data rather than concentrating on the leaderboard only. This could go really well, or be a very bad idea. But I believe it is the right thing to do here",
    "412863": "&gt; Made the model and have trained from scratch. It is basically a resnet with dense categorical output. Still tuning the optimizer and learning rate. Recently set the rate to something much more aggressive and instead of blowing up like I thought it would, it continued to converge quickly. 3.5M parameters total.\n\nDid you try cosine-annealing?",
    "412909": "Yes, using SGD with cosine annealing schedule. Also used Adadelta to start training, Padam for mid training, and SGD at the end. Then I freeze parts of the model and train the other layers. My current leading model is 2.3M params. Performs great locally, but public LB is 45% lower.",
    "437197": "Hello, how did you fix the gap between local lb and public lb?",
    "437286": "Hi Alexander Liao,\nCan you explain how you determined that there are 7 missing classes in LB?\nThanks.\nAh, I see that it was done by LB probing, which however is unreliable because of score truncation to 3 digits."
  },
  "source": "meta"
}