{
  "id": 68678,
  "title": "LB probing",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/68678",
  "author_name": "",
  "post_date": "2018-10-16T03:46:18.232245800Z",
  "votes": 40,
  "comment_count": 21,
  "views": 0,
  "content": "<p>I submitted files containing the same prediction for all images and got following results (class, score, fraction in public LB). The fraction of ground truth predictions of a particular class in the public LB if \"i\" is set for all images can be calculated based on macro averaging of the score between 28 classes: <code>F_i = score_i*28</code>; <code>p_i = F_i/(2 - F_i)</code>.</p>\n\n<pre><code> - 0 -&amp;gt; 0.019 -&amp;gt; 0.36239782\n - 1 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 2 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 3 -&amp;gt; 0.004 -&amp;gt; 0.059322034\n - 4 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 5 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 6 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 7 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 8 -&amp;gt; 0 -&amp;gt; 0\n - 9 -&amp;gt; 0 -&amp;gt; 0\n - 10 -&amp;gt; 0 -&amp;gt; 0\n - 11 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 12 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 13 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n - 14 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 15 -&amp;gt; 0 -&amp;gt; 0\n - 16 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n - 17 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n - 18 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n - 19 -&amp;gt; 0.004 -&amp;gt;  0.059322034\n - 20 -&amp;gt; 0 -&amp;gt; 0\n - 21 -&amp;gt; 0.008 -&amp;gt; 0.126126126\n - 22 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n - 23 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 24 -&amp;gt; 0 -&amp;gt; 0\n - 25 -&amp;gt; 0.013 -&amp;gt; 0.222493888\n - 26 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n - 27 -&amp;gt; 0 -&amp;gt; 0\n</code></pre>",
  "messages": [
    {
      "id": "404591",
      "postDate": "10/16/2018 03:46:18",
      "content": "<p>I submitted files containing the same prediction for all images and got following results (class, score, fraction in public LB). The fraction of ground truth predictions of a particular class in the public LB if \"i\" is set for all images can be calculated based on macro averaging of the score between 28 classes: <code>F_i = score_i*28</code>; <code>p_i = F_i/(2 - F_i)</code>.</p>\n\n<pre><code> - 0 -&amp;gt; 0.019 -&amp;gt; 0.36239782\n - 1 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 2 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 3 -&amp;gt; 0.004 -&amp;gt; 0.059322034\n - 4 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 5 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 6 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 7 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 8 -&amp;gt; 0 -&amp;gt; 0\n - 9 -&amp;gt; 0 -&amp;gt; 0\n - 10 -&amp;gt; 0 -&amp;gt; 0\n - 11 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 12 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 13 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n - 14 -&amp;gt; 0.003 -&amp;gt; 0.043841336\n - 15 -&amp;gt; 0 -&amp;gt; 0\n - 16 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n - 17 -&amp;gt; 0.001 -&amp;gt; 0.014198783\n - 18 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n - 19 -&amp;gt; 0.004 -&amp;gt;  0.059322034\n - 20 -&amp;gt; 0 -&amp;gt; 0\n - 21 -&amp;gt; 0.008 -&amp;gt; 0.126126126\n - 22 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n - 23 -&amp;gt; 0.005 -&amp;gt; 0.075268817\n - 24 -&amp;gt; 0 -&amp;gt; 0\n - 25 -&amp;gt; 0.013 -&amp;gt; 0.222493888\n - 26 -&amp;gt; 0.002 -&amp;gt; 0.028806584\n - 27 -&amp;gt; 0 -&amp;gt; 0\n</code></pre>",
      "rawMarkdown": "I submitted files containing the same prediction for all images and got following results (class, score, fraction in public LB). The fraction of ground truth predictions of a particular class in the public LB if \"i\" is set for all images can be calculated based on macro averaging of the score between 28 classes: `F_i = score_i*28`; `p_i = F_i/(2 - F_i)`.\n\n     - 0 -&gt; 0.019 -&gt; 0.36239782\n     - 1 -&gt; 0.003 -&gt; 0.043841336\n     - 2 -&gt; 0.005 -&gt; 0.075268817\n     - 3 -&gt; 0.004 -&gt; 0.059322034\n     - 4 -&gt; 0.005 -&gt; 0.075268817\n     - 5 -&gt; 0.005 -&gt; 0.075268817\n     - 6 -&gt; 0.003 -&gt; 0.043841336\n     - 7 -&gt; 0.005 -&gt; 0.075268817\n     - 8 -&gt; 0 -&gt; 0\n     - 9 -&gt; 0 -&gt; 0\n     - 10 -&gt; 0 -&gt; 0\n     - 11 -&gt; 0.003 -&gt; 0.043841336\n     - 12 -&gt; 0.003 -&gt; 0.043841336\n     - 13 -&gt; 0.001 -&gt; 0.014198783\n     - 14 -&gt; 0.003 -&gt; 0.043841336\n     - 15 -&gt; 0 -&gt; 0\n     - 16 -&gt; 0.001 -&gt; 0.014198783\n     - 17 -&gt; 0.001 -&gt; 0.014198783\n     - 18 -&gt; 0.002 -&gt; 0.028806584\n     - 19 -&gt; 0.004 -&gt;  0.059322034\n     - 20 -&gt; 0 -&gt; 0\n     - 21 -&gt; 0.008 -&gt; 0.126126126\n     - 22 -&gt; 0.002 -&gt; 0.028806584\n     - 23 -&gt; 0.005 -&gt; 0.075268817\n     - 24 -&gt; 0 -&gt; 0\n     - 25 -&gt; 0.013 -&gt; 0.222493888\n     - 26 -&gt; 0.002 -&gt; 0.028806584\n     - 27 -&gt; 0 -&gt; 0",
      "votes": null
    },
    {
      "id": "404664",
      "postDate": "10/16/2018 07:05:47",
      "content": "<p>Nice work! Just note, that 0 doesn't have to mean exactly 0. Due to rounding to 3 digits, we only know that the score is less than 0.000499, which means frequency &lt;0.007</p>",
      "rawMarkdown": "Nice work! Just note, that 0 doesn't have to mean exactly 0. Due to rounding to 3 digits, we only know that the score is less than 0.000499, which means frequency &lt;0.007",
      "votes": null
    },
    {
      "id": "404909",
      "postDate": "10/16/2018 14:53:21",
      "content": "<p>You are right, those fractions are calculated up to the rounding errors. Also, there can be no examples of a particular class in the public LB, while they are present in a private one.</p>",
      "rawMarkdown": "You are right, those fractions are calculated up to the rounding errors. Also, there can be no examples of a particular class in the public LB, while they are present in a private one.",
      "votes": null
    },
    {
      "id": "405999",
      "postDate": "10/18/2018 13:29:58",
      "content": "<p>Hi,\ni checked the class frequency for the training (and my random validation set) and the probabilities you posted, and it seemes they are very similar.\nJust a FYI incase you were wondering if your method for sampling a  validation set are different from the training set up.\nCheers,\nGunther</p>",
      "rawMarkdown": "Hi,\ni checked the class frequency for the training (and my random validation set) and the probabilities you posted, and it seemes they are very similar.\nJust a FYI incase you were wondering if your method for sampling a  validation set are different from the training set up.\nCheers,\nGunther",
      "votes": null
    },
    {
      "id": "406052",
      "postDate": "10/18/2018 15:06:58",
      "content": "<p>Nice analysis, thank you.</p>",
      "rawMarkdown": "Nice analysis, thank you.",
      "votes": null
    },
    {
      "id": "409282",
      "postDate": "10/24/2018 03:12:54",
      "content": "<p>Hi, I think this is very clever. But does it really help? As far as I understand it, the final evaluation is done on a different test set, which we can't check out beforehand. Or did I get something wrong there? Thanks for your answers, best, Wolfgang</p>",
      "rawMarkdown": "Hi, I think this is very clever. But does it really help? As far as I understand it, the final evaluation is done on a different test set, which we can't check out beforehand. Or did I get something wrong there? Thanks for your answers, best, Wolfgang",
      "votes": null
    },
    {
      "id": "409298",
      "postDate": "10/24/2018 04:39:51",
      "content": "<p>Yes, the thing that we may hope is that private test dataset is closer to public test dataset rather than to train one. Train data is quite off, even if u just check pixel statistics. Also it is confirmed by necessity of using  <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">different thresholds for validation and test</a> and <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462\">huge difference between validation and test score</a>.</p>",
      "rawMarkdown": "Yes, the thing that we may hope is that private test dataset is closer to public test dataset rather than to train one. Train data is quite off, even if u just check pixel statistics. Also it is confirmed by necessity of using  [different thresholds for validation and test][1] and [huge difference between validation and test score][2].\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462",
      "votes": null
    },
    {
      "id": "410557",
      "postDate": "10/26/2018 08:54:07",
      "content": "<p>Dangerous territory here. You make assumptions about all the test cases by looking at those 29% of the 10k test images the LB is actually based on. That introduces a lot of noise in the estimated class frequencies.</p>",
      "rawMarkdown": "Dangerous territory here. You make assumptions about all the test cases by looking at those 29% of the 10k test images the LB is actually based on. That introduces a lot of noise in the estimated class frequencies.",
      "votes": null
    },
    {
      "id": "410722",
      "postDate": "10/26/2018 14:41:15",
      "content": "<p>It is right, but it is also helpful to have information about LB statistics. @Gunther made a comparison of public LB statistics and train set, so looks that they are quite consistent, but some classes may have different frequencies.</p>",
      "rawMarkdown": "It is right, but it is also helpful to have information about LB statistics. @Gunther made a comparison of public LB statistics and train set, so looks that they are quite consistent, but some classes may have different frequencies.",
      "votes": null
    },
    {
      "id": "423440",
      "postDate": "11/18/2018 08:38:19",
      "content": "<p>And a frequency of &lt;0.007 for the ~ 3 400 images in the public LB gives that if there is less than 24 occurrences of a specific class in the public LB then you may not see it with probing.</p>",
      "rawMarkdown": "And a frequency of &lt;0.007 for the ~ 3 400 images in the public LB gives that if there is less than 24 occurrences of a specific class in the public LB then you may not see it with probing.",
      "votes": null
    },
    {
      "id": "423586",
      "postDate": "11/18/2018 16:15:09",
      "content": "<p>Actually right now I think that the discrepancy between public LB and val is resulted by missing those 7 rare classes. Therefore:</p>\n\n<pre><code>public LB ~ val - 0.25\n</code></pre>",
      "rawMarkdown": "Actually right now I think that the discrepancy between public LB and val is resulted by missing those 7 rare classes. Therefore:\n\n    public LB ~ val - 0.25",
      "votes": null
    },
    {
      "id": "423714",
      "postDate": "11/18/2018 23:03:15",
      "content": "<p>I disagree that 9 and 10 are missing classes. These are all public LB scores. I took a submission that scored .527 and here are some test results:</p>\n\n<ul>\n<li>Change any rows predicted as  '9 10 26' to '26' resulted in a .525</li>\n<li>Change any rows predicted as  '9 10 26' to '9 10' resulted in a .531</li>\n</ul>\n\n<p>To me this shows that 9 and 10 are indeed used in the public LB.</p>",
      "rawMarkdown": "I disagree that 9 and 10 are missing classes. These are all public LB scores. I took a submission that scored .527 and here are some test results:\n\n* Change any rows predicted as  '9 10 26' to '26' resulted in a .525\n* Change any rows predicted as  '9 10 26' to '9 10' resulted in a .531\n\nTo me this shows that 9 and 10 are indeed used in the public LB.",
      "votes": null
    },
    {
      "id": "423730",
      "postDate": "11/19/2018 00:19:12",
      "content": "<p>Thank you @Brian for letting me know. Do I understand correctly that if you drop all 9 and 10 classes in the submission your score goes down? did you try to do similar testing dropping only one class each time?</p>",
      "rawMarkdown": "Thank you @Brian for letting me know. Do I understand correctly that if you drop all 9 and 10 classes in the submission your score goes down? did you try to do similar testing dropping only one class each time?",
      "votes": null
    },
    {
      "id": "423877",
      "postDate": "11/19/2018 07:19:13",
      "content": "<p>I only tested dropping 9's and 10's when there was also a 26. I left in values that were only 9 and 10. I also added 10 anywhere there is a 9. Didn't try any other classes, I was doing some targeted changes after prediction based on the correlation between classes and some clear errors the model was making.</p>",
      "rawMarkdown": "I only tested dropping 9's and 10's when there was also a 26. I left in values that were only 9 and 10. I also added 10 anywhere there is a 9. Didn't try any other classes, I was doing some targeted changes after prediction based on the correlation between classes and some clear errors the model was making.",
      "votes": null
    },
    {
      "id": "425314",
      "postDate": "11/21/2018 12:41:18",
      "content": "<p>In the training set, the label which has the smallest occurence probability is Rods&amp;rings, 0.000357. For this label, the probability that the label doesn't occur in the test set is 0.03. That is to say, this is impossible from aspect of hypothesis test. As for other labels,  the probability that labels doesn't occur in the test set is lower than 4.5*10-5.</p>",
      "rawMarkdown": "In the training set, the label which has the smallest occurence probability is Rods&amp;rings, 0.000357. For this label, the probability that the label doesn't occur in the test set is 0.03. That is to say, this is impossible from aspect of hypothesis test. As for other labels,  the probability that labels doesn't occur in the test set is lower than 4.5*10-5.",
      "votes": null
    },
    {
      "id": "425498",
      "postDate": "11/21/2018 17:14:17",
      "content": "<p>Initially I also thought that the fraction of these classes is low. However, if some of them are not present in public LB, it can explain the discrepancy between public LB score and val. Also, when you do hypothesis testing, you should remember that it is competition, not just randomly selected data: you never know which trick organizers use to confuse participants.</p>",
      "rawMarkdown": "Initially I also thought that the fraction of these classes is low. However, if some of them are not present in public LB, it can explain the discrepancy between public LB score and val. Also, when you do hypothesis testing, you should remember that it is competition, not just randomly selected data: you never know which trick organizers use to confuse participants.",
      "votes": null
    },
    {
      "id": "447561",
      "postDate": "12/30/2018 04:18:42",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> can you share the result on the LB after the leak removed?</p>",
      "rawMarkdown": "iafoss can you share the result on the LB after the leak removed?",
      "votes": null
    },
    {
      "id": "447565",
      "postDate": "12/30/2018 04:31:26",
      "content": "<p>It is already updated, I think in the new LB only one value has changed, but I do not remember which one.</p>",
      "rawMarkdown": "It is already updated, I think in the new LB only one value has changed, but I do not remember which one.",
      "votes": null
    },
    {
      "id": "447570",
      "postDate": "12/30/2018 04:59:58",
      "content": "<p>I think that the leak was not \"removed\". On the contrary, it is now mostly in the public lb. So, using the leak gives you excellent lb score but this will have nothing to do with the private lb scores. Also, probably now the class distribution of the public lb and private lb are different so the probing might not help the final score.\n@lafoss, just out of curiosity, what is the math behind the lb probing? I understand you probe with a submission with one class for all ids but i couldn't figure out what and how to deduct anything from this. </p>",
      "rawMarkdown": "I think that the leak was not \"removed\". On the contrary, it is now mostly in the public lb. So, using the leak gives you excellent lb score but this will have nothing to do with the private lb scores. Also, probably now the class distribution of the public lb and private lb are different so the probing might not help the final score.\n@lafoss, just out of curiosity, what is the math behind the lb probing? I understand you probe with a submission with one class for all ids but i couldn't figure out what and how to deduct anything from this.",
      "votes": null
    },
    {
      "id": "447583",
      "postDate": "12/30/2018 05:35:27",
      "content": "<p>Everything is coming from <strong>macro</strong> averaging. If you submit a file with only one label for all images, you will get <code>TP = f*N</code>, <code>FP = (1-f)*N</code>, <code>FN = 0</code>, where f - is the fraction of the particular label in the dataset, and <code>F1 = 2*TP/(2*TP+FP+FN)</code>. Also, you know that the F1 score for all other labels is 0, so the real F1 score for a particular class is <code>public LB*28</code>. That is all that you need to calculate f.</p>",
      "rawMarkdown": "Everything is coming from **macro** averaging. If you submit a file with only one label for all images, you will get `TP = f*N`, `FP = (1-f)*N`, `FN = 0`, where f - is the fraction of the particular label in the dataset, and `F1 = 2*TP/(2*TP+FP+FN)`. Also, you know that the F1 score for all other labels is 0, so the real F1 score for a particular class is `public LB*28`. That is all that you need to calculate f.",
      "votes": null
    },
    {
      "id": "447627",
      "postDate": "12/30/2018 07:37:04",
      "content": "<p>Thank you! Much appreciated </p>",
      "rawMarkdown": "Thank you! Much appreciated",
      "votes": null
    },
    {
      "id": "447842",
      "postDate": "12/30/2018 17:55:47",
      "content": "<p>@Iafoss Thank you! I've got it.</p>",
      "rawMarkdown": "Iafoss Thank you! I've got it.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 404664,
      "author_name": "rejpalcz",
      "author_url": "",
      "post_date": "10/16/2018 07:05:47",
      "content": "<p>Nice work! Just note, that 0 doesn't have to mean exactly 0. Due to rounding to 3 digits, we only know that the score is less than 0.000499, which means frequency &lt;0.007</p>",
      "votes": null,
      "replies": [
        {
          "id": 404909,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/16/2018 14:53:21",
          "content": "<p>You are right, those fractions are calculated up to the rounding errors. Also, there can be no examples of a particular class in the public LB, while they are present in a private one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423440,
          "author_name": "manuscrits",
          "author_url": "",
          "post_date": "11/18/2018 08:38:19",
          "content": "<p>And a frequency of &lt;0.007 for the ~ 3 400 images in the public LB gives that if there is less than 24 occurrences of a specific class in the public LB then you may not see it with probing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423586,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/18/2018 16:15:09",
          "content": "<p>Actually right now I think that the discrepancy between public LB and val is resulted by missing those 7 rare classes. Therefore:</p>\n\n<pre><code>public LB ~ val - 0.25\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423714,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/18/2018 23:03:15",
          "content": "<p>I disagree that 9 and 10 are missing classes. These are all public LB scores. I took a submission that scored .527 and here are some test results:</p>\n\n<ul>\n<li>Change any rows predicted as  '9 10 26' to '26' resulted in a .525</li>\n<li>Change any rows predicted as  '9 10 26' to '9 10' resulted in a .531</li>\n</ul>\n\n<p>To me this shows that 9 and 10 are indeed used in the public LB.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423730,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/19/2018 00:19:12",
          "content": "<p>Thank you @Brian for letting me know. Do I understand correctly that if you drop all 9 and 10 classes in the submission your score goes down? did you try to do similar testing dropping only one class each time?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 423877,
          "author_name": "ldm314",
          "author_url": "",
          "post_date": "11/19/2018 07:19:13",
          "content": "<p>I only tested dropping 9's and 10's when there was also a 26. I left in values that were only 9 and 10. I also added 10 anywhere there is a 9. Didn't try any other classes, I was doing some targeted changes after prediction based on the correlation between classes and some clear errors the model was making.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 405999,
      "author_name": "guntherthepenguin",
      "author_url": "",
      "post_date": "10/18/2018 13:29:58",
      "content": "<p>Hi,\ni checked the class frequency for the training (and my random validation set) and the probabilities you posted, and it seemes they are very similar.\nJust a FYI incase you were wondering if your method for sampling a  validation set are different from the training set up.\nCheers,\nGunther</p>",
      "votes": null,
      "replies": [
        {
          "id": 406052,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/18/2018 15:06:58",
          "content": "<p>Nice analysis, thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 409282,
      "author_name": "wreuter",
      "author_url": "",
      "post_date": "10/24/2018 03:12:54",
      "content": "<p>Hi, I think this is very clever. But does it really help? As far as I understand it, the final evaluation is done on a different test set, which we can't check out beforehand. Or did I get something wrong there? Thanks for your answers, best, Wolfgang</p>",
      "votes": null,
      "replies": [
        {
          "id": 409298,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/24/2018 04:39:51",
          "content": "<p>Yes, the thing that we may hope is that private test dataset is closer to public test dataset rather than to train one. Train data is quite off, even if u just check pixel statistics. Also it is confirmed by necessity of using  <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\">different thresholds for validation and test</a> and <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462\">huge difference between validation and test score</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 410557,
      "author_name": "fabianisensee",
      "author_url": "",
      "post_date": "10/26/2018 08:54:07",
      "content": "<p>Dangerous territory here. You make assumptions about all the test cases by looking at those 29% of the 10k test images the LB is actually based on. That introduces a lot of noise in the estimated class frequencies.</p>",
      "votes": null,
      "replies": [
        {
          "id": 410722,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "10/26/2018 14:41:15",
          "content": "<p>It is right, but it is also helpful to have information about LB statistics. @Gunther made a comparison of public LB statistics and train set, so looks that they are quite consistent, but some classes may have different frequencies.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 425314,
      "author_name": "zhengzhuangqiang",
      "author_url": "",
      "post_date": "11/21/2018 12:41:18",
      "content": "<p>In the training set, the label which has the smallest occurence probability is Rods&amp;rings, 0.000357. For this label, the probability that the label doesn't occur in the test set is 0.03. That is to say, this is impossible from aspect of hypothesis test. As for other labels,  the probability that labels doesn't occur in the test set is lower than 4.5*10-5.</p>",
      "votes": null,
      "replies": [
        {
          "id": 425498,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "11/21/2018 17:14:17",
          "content": "<p>Initially I also thought that the fraction of these classes is low. However, if some of them are not present in public LB, it can explain the discrepancy between public LB score and val. Also, when you do hypothesis testing, you should remember that it is competition, not just randomly selected data: you never know which trick organizers use to confuse participants.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447561,
      "author_name": "osciiart",
      "author_url": "",
      "post_date": "12/30/2018 04:18:42",
      "content": "<p><a href=\"/iafoss\">@iafoss</a> can you share the result on the LB after the leak removed?</p>",
      "votes": null,
      "replies": [
        {
          "id": 447565,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/30/2018 04:31:26",
          "content": "<p>It is already updated, I think in the new LB only one value has changed, but I do not remember which one.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447570,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/30/2018 04:59:58",
          "content": "<p>I think that the leak was not \"removed\". On the contrary, it is now mostly in the public lb. So, using the leak gives you excellent lb score but this will have nothing to do with the private lb scores. Also, probably now the class distribution of the public lb and private lb are different so the probing might not help the final score.\n@lafoss, just out of curiosity, what is the math behind the lb probing? I understand you probe with a submission with one class for all ids but i couldn't figure out what and how to deduct anything from this. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447583,
          "author_name": "iafoss",
          "author_url": "",
          "post_date": "12/30/2018 05:35:27",
          "content": "<p>Everything is coming from <strong>macro</strong> averaging. If you submit a file with only one label for all images, you will get <code>TP = f*N</code>, <code>FP = (1-f)*N</code>, <code>FN = 0</code>, where f - is the fraction of the particular label in the dataset, and <code>F1 = 2*TP/(2*TP+FP+FN)</code>. Also, you know that the F1 score for all other labels is 0, so the real F1 score for a particular class is <code>public LB*28</code>. That is all that you need to calculate f.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447627,
          "author_name": "moshel",
          "author_url": "",
          "post_date": "12/30/2018 07:37:04",
          "content": "<p>Thank you! Much appreciated </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447842,
          "author_name": "osciiart",
          "author_url": "",
          "post_date": "12/30/2018 17:55:47",
          "content": "<p>@Iafoss Thank you! I've got it.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "404591": "I submitted files containing the same prediction for all images and got following results (class, score, fraction in public LB). The fraction of ground truth predictions of a particular class in the public LB if \"i\" is set for all images can be calculated based on macro averaging of the score between 28 classes: `F_i = score_i*28`; `p_i = F_i/(2 - F_i)`.\n\n     - 0 -&gt; 0.019 -&gt; 0.36239782\n     - 1 -&gt; 0.003 -&gt; 0.043841336\n     - 2 -&gt; 0.005 -&gt; 0.075268817\n     - 3 -&gt; 0.004 -&gt; 0.059322034\n     - 4 -&gt; 0.005 -&gt; 0.075268817\n     - 5 -&gt; 0.005 -&gt; 0.075268817\n     - 6 -&gt; 0.003 -&gt; 0.043841336\n     - 7 -&gt; 0.005 -&gt; 0.075268817\n     - 8 -&gt; 0 -&gt; 0\n     - 9 -&gt; 0 -&gt; 0\n     - 10 -&gt; 0 -&gt; 0\n     - 11 -&gt; 0.003 -&gt; 0.043841336\n     - 12 -&gt; 0.003 -&gt; 0.043841336\n     - 13 -&gt; 0.001 -&gt; 0.014198783\n     - 14 -&gt; 0.003 -&gt; 0.043841336\n     - 15 -&gt; 0 -&gt; 0\n     - 16 -&gt; 0.001 -&gt; 0.014198783\n     - 17 -&gt; 0.001 -&gt; 0.014198783\n     - 18 -&gt; 0.002 -&gt; 0.028806584\n     - 19 -&gt; 0.004 -&gt;  0.059322034\n     - 20 -&gt; 0 -&gt; 0\n     - 21 -&gt; 0.008 -&gt; 0.126126126\n     - 22 -&gt; 0.002 -&gt; 0.028806584\n     - 23 -&gt; 0.005 -&gt; 0.075268817\n     - 24 -&gt; 0 -&gt; 0\n     - 25 -&gt; 0.013 -&gt; 0.222493888\n     - 26 -&gt; 0.002 -&gt; 0.028806584\n     - 27 -&gt; 0 -&gt; 0",
    "404664": "Nice work! Just note, that 0 doesn't have to mean exactly 0. Due to rounding to 3 digits, we only know that the score is less than 0.000499, which means frequency &lt;0.007",
    "404909": "You are right, those fractions are calculated up to the rounding errors. Also, there can be no examples of a particular class in the public LB, while they are present in a private one.",
    "405999": "Hi,\ni checked the class frequency for the training (and my random validation set) and the probabilities you posted, and it seemes they are very similar.\nJust a FYI incase you were wondering if your method for sampling a  validation set are different from the training set up.\nCheers,\nGunther",
    "406052": "Nice analysis, thank you.",
    "409282": "Hi, I think this is very clever. But does it really help? As far as I understand it, the final evaluation is done on a different test set, which we can't check out beforehand. Or did I get something wrong there? Thanks for your answers, best, Wolfgang",
    "409298": "Yes, the thing that we may hope is that private test dataset is closer to public test dataset rather than to train one. Train data is quite off, even if u just check pixel statistics. Also it is confirmed by necessity of using  [different thresholds for validation and test][1] and [huge difference between validation and test score][2].\n\n\n  [1]: https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-fast-ai\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69462",
    "410557": "Dangerous territory here. You make assumptions about all the test cases by looking at those 29% of the 10k test images the LB is actually based on. That introduces a lot of noise in the estimated class frequencies.",
    "410722": "It is right, but it is also helpful to have information about LB statistics. @Gunther made a comparison of public LB statistics and train set, so looks that they are quite consistent, but some classes may have different frequencies.",
    "423440": "And a frequency of &lt;0.007 for the ~ 3 400 images in the public LB gives that if there is less than 24 occurrences of a specific class in the public LB then you may not see it with probing.",
    "423586": "Actually right now I think that the discrepancy between public LB and val is resulted by missing those 7 rare classes. Therefore:\n\n    public LB ~ val - 0.25",
    "423714": "I disagree that 9 and 10 are missing classes. These are all public LB scores. I took a submission that scored .527 and here are some test results:\n\n* Change any rows predicted as  '9 10 26' to '26' resulted in a .525\n* Change any rows predicted as  '9 10 26' to '9 10' resulted in a .531\n\nTo me this shows that 9 and 10 are indeed used in the public LB.",
    "423730": "Thank you @Brian for letting me know. Do I understand correctly that if you drop all 9 and 10 classes in the submission your score goes down? did you try to do similar testing dropping only one class each time?",
    "423877": "I only tested dropping 9's and 10's when there was also a 26. I left in values that were only 9 and 10. I also added 10 anywhere there is a 9. Didn't try any other classes, I was doing some targeted changes after prediction based on the correlation between classes and some clear errors the model was making.",
    "425314": "In the training set, the label which has the smallest occurence probability is Rods&amp;rings, 0.000357. For this label, the probability that the label doesn't occur in the test set is 0.03. That is to say, this is impossible from aspect of hypothesis test. As for other labels,  the probability that labels doesn't occur in the test set is lower than 4.5*10-5.",
    "425498": "Initially I also thought that the fraction of these classes is low. However, if some of them are not present in public LB, it can explain the discrepancy between public LB score and val. Also, when you do hypothesis testing, you should remember that it is competition, not just randomly selected data: you never know which trick organizers use to confuse participants.",
    "447561": "iafoss can you share the result on the LB after the leak removed?",
    "447565": "It is already updated, I think in the new LB only one value has changed, but I do not remember which one.",
    "447570": "I think that the leak was not \"removed\". On the contrary, it is now mostly in the public lb. So, using the leak gives you excellent lb score but this will have nothing to do with the private lb scores. Also, probably now the class distribution of the public lb and private lb are different so the probing might not help the final score.\n@lafoss, just out of curiosity, what is the math behind the lb probing? I understand you probe with a submission with one class for all ids but i couldn't figure out what and how to deduct anything from this.",
    "447583": "Everything is coming from **macro** averaging. If you submit a file with only one label for all images, you will get `TP = f*N`, `FP = (1-f)*N`, `FN = 0`, where f - is the fraction of the particular label in the dataset, and `F1 = 2*TP/(2*TP+FP+FN)`. Also, you know that the F1 score for all other labels is 0, so the real F1 score for a particular class is `public LB*28`. That is all that you need to calculate f.",
    "447627": "Thank you! Much appreciated",
    "447842": "Iafoss Thank you! I've got it."
  },
  "source": "meta"
}