{
  "id": 76336,
  "title": "Do you trust HPA? (and some ramblings)",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/76336",
  "author_name": "",
  "post_date": "2019-01-01T13:25:50.478981900Z",
  "votes": 17,
  "comment_count": 10,
  "views": 0,
  "content": "<p><strong>Disclaimer</strong>: I've gotten around to exploring this too late and a little half-heartedly so I may merely be pointing out some parts of a jigsaw that others have solved.</p>\n\n<p><strong>TLDR</strong></p>\n\n<p>I can't help but think this is a competition with huge potential for CV errors. My present hypothesis (which is likely only partially true) is that if you've picked up the HPA data and simply mixed it into the train data and run k-fold over it then you might be in for a shock on the private LB.</p>\n\n<p><strong>Using HPA</strong></p>\n\n<p>First up, if we start from the position of the organisers who are experienced and have had people build models on their data before it bothers me that one can (potentially) get such huge score improvements from using the HPA data. </p>\n\n<p>Bluntly, why would you set-up a competition with massive class imbalance and use a metric that punishes this heavily on a limited data-set when all the time in your broom cupboard you have 2x the data and you don't make it openly available?</p>\n\n<p><strong>In other words:</strong> construct a competition with massive class imbalance and use macro f1 metric =&gt; predicting rare classes is the key to the competition =&gt; kagglers are not going to ignore potentially 2x the data for these classes </p>\n\n<p>Training observations:</p>\n\n<ul>\n<li>Models with train data only (5fold): local CV 0.72 - 0.74 =&gt; LB 0.47 - 0.5</li>\n<li>Models trained on train + HPA data and validation sets on train data only: unstable, didn't actually submit but this was my preferred way to incorporate HPA initially to try to respect CV integrity.</li>\n<li>Mix train + HPA and run 5 fold and train for 12 epochs: train loss stable, validation loss wild. Local CV f1 0.6, LB 0.55, LB with leak file 0.582, LB removing all leak entries 0.532</li>\n</ul>\n\n<p>A few options:</p>\n\n<ol>\n<li><strong>The organisers knew all along HPA was crucial but chose not to include it.</strong> e.g. to not overwhelm kagglers with an even more enormous dataset. If this is the case I think it's poor competition set-up as huge amounts of time has been spent by hundreds of people trying to understand this extra data - it would be much fairer and efficient to include it as part of the competition. In this scenario they've basically given us data that isn't sufficient for the problem at hand. </li>\n<li><strong>The organisers know HPA is helpful but can't speak for the accuracy of some of it.</strong> Perhaps more likely than option 1 but this ambiguity must be detrimental to the competition given all the energy spent on figuring this out. A simple statement clarifying this fact would have been helpful.</li>\n<li><strong>The HPA is a red herring.</strong> It's possible (but pretty unlikely) that the HPA data is completely misleading and the public LB looks very different to the private LB. The lack of stability from train to public LB could be that the organisers didn't have enough data they were happy with to construct a bigger overall test set and so put the better data (i.e. matching train closely) into the private LB and the public LB set is a little unrepresentative (this could be a justifiable thing to do on the basis that the 'trust your CV' mantra will pay-off on private LB though causes a lot of red flags for kagglers).</li>\n<li><strong>Is HPA closer to train than train is to test?</strong> It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because...</li>\n</ol>\n\n<p><strong>Drip drop - mind the leak</strong></p>\n\n<p>An obvious concern is simply leakage. There are known similar images in train and HPA and known leakage to the public LB from HPA. It's possible to get a decent local CV by not taking care when constructing your validation sets (e.g. having rare similar images in both validation and train) and then when you submit you score well due to public LB leakage. In both cases you may be committing an error that may cause a (possibly severe) private LB fall.</p>\n\n<p><strong>Safely using HPA</strong></p>\n\n<p>A 'safe' way to use HPA (which I am going to do over the next week) in my view is something like: run a quick k-fold on train + HPA or on train only and predict on HPA. Throw away any HPA data that looks mislabelled. Throw away any similar images. Throw away any known leak images (and blank these out when submitting to LB). Use BCE as a loss function with some sampling strategy.</p>\n\n<p>It's almost certain the above 'safe' use of the HPA is too aggressive in the data it will throw away but might give one more peace of mind about leakage.</p>\n\n<p><em>Open question</em>: does HPA look more like test than train does test?</p>\n\n<p><strong>A note on using focal loss</strong></p>\n\n<p>Training with focal loss and HPA data is unstable. If you are experiencing large spikes in validation loss but still getting okay f1 my guess is that it's due to bad labels in the HPA data. Focal loss will focus on examples that are hard to classify and if these are mislabelled you can get wild spikes in validation loss as the model works to get these bad training labels right (i.e. probably ends up memorising them) and obviously this causes poor validation performance. If you still want to use focal loss you can add a line (in PyTorch) into the function to clip the logits, e.g.:</p>\n\n<pre><code>preds = preds.clamp(min=-15, max=15)   # need to try varying the best clamping\n</code></pre>\n\n<p>Though this seems to slow training down quite a bit it does bring stability to the validation loss.</p>\n\n<p>Keen to hear thoughts,</p>\n\n<p>Mark</p>\n\n<p>P. S. It's also pretty likely the current public LB scores are not to be trusted too much as it's unclear who has used the leak information/probed LB. </p>",
  "messages": [
    {
      "id": "448550",
      "postDate": "01/01/2019 13:25:50",
      "content": "<p><strong>Disclaimer</strong>: I've gotten around to exploring this too late and a little half-heartedly so I may merely be pointing out some parts of a jigsaw that others have solved.</p>\n\n<p><strong>TLDR</strong></p>\n\n<p>I can't help but think this is a competition with huge potential for CV errors. My present hypothesis (which is likely only partially true) is that if you've picked up the HPA data and simply mixed it into the train data and run k-fold over it then you might be in for a shock on the private LB.</p>\n\n<p><strong>Using HPA</strong></p>\n\n<p>First up, if we start from the position of the organisers who are experienced and have had people build models on their data before it bothers me that one can (potentially) get such huge score improvements from using the HPA data. </p>\n\n<p>Bluntly, why would you set-up a competition with massive class imbalance and use a metric that punishes this heavily on a limited data-set when all the time in your broom cupboard you have 2x the data and you don't make it openly available?</p>\n\n<p><strong>In other words:</strong> construct a competition with massive class imbalance and use macro f1 metric =&gt; predicting rare classes is the key to the competition =&gt; kagglers are not going to ignore potentially 2x the data for these classes </p>\n\n<p>Training observations:</p>\n\n<ul>\n<li>Models with train data only (5fold): local CV 0.72 - 0.74 =&gt; LB 0.47 - 0.5</li>\n<li>Models trained on train + HPA data and validation sets on train data only: unstable, didn't actually submit but this was my preferred way to incorporate HPA initially to try to respect CV integrity.</li>\n<li>Mix train + HPA and run 5 fold and train for 12 epochs: train loss stable, validation loss wild. Local CV f1 0.6, LB 0.55, LB with leak file 0.582, LB removing all leak entries 0.532</li>\n</ul>\n\n<p>A few options:</p>\n\n<ol>\n<li><strong>The organisers knew all along HPA was crucial but chose not to include it.</strong> e.g. to not overwhelm kagglers with an even more enormous dataset. If this is the case I think it's poor competition set-up as huge amounts of time has been spent by hundreds of people trying to understand this extra data - it would be much fairer and efficient to include it as part of the competition. In this scenario they've basically given us data that isn't sufficient for the problem at hand. </li>\n<li><strong>The organisers know HPA is helpful but can't speak for the accuracy of some of it.</strong> Perhaps more likely than option 1 but this ambiguity must be detrimental to the competition given all the energy spent on figuring this out. A simple statement clarifying this fact would have been helpful.</li>\n<li><strong>The HPA is a red herring.</strong> It's possible (but pretty unlikely) that the HPA data is completely misleading and the public LB looks very different to the private LB. The lack of stability from train to public LB could be that the organisers didn't have enough data they were happy with to construct a bigger overall test set and so put the better data (i.e. matching train closely) into the private LB and the public LB set is a little unrepresentative (this could be a justifiable thing to do on the basis that the 'trust your CV' mantra will pay-off on private LB though causes a lot of red flags for kagglers).</li>\n<li><strong>Is HPA closer to train than train is to test?</strong> It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because...</li>\n</ol>\n\n<p><strong>Drip drop - mind the leak</strong></p>\n\n<p>An obvious concern is simply leakage. There are known similar images in train and HPA and known leakage to the public LB from HPA. It's possible to get a decent local CV by not taking care when constructing your validation sets (e.g. having rare similar images in both validation and train) and then when you submit you score well due to public LB leakage. In both cases you may be committing an error that may cause a (possibly severe) private LB fall.</p>\n\n<p><strong>Safely using HPA</strong></p>\n\n<p>A 'safe' way to use HPA (which I am going to do over the next week) in my view is something like: run a quick k-fold on train + HPA or on train only and predict on HPA. Throw away any HPA data that looks mislabelled. Throw away any similar images. Throw away any known leak images (and blank these out when submitting to LB). Use BCE as a loss function with some sampling strategy.</p>\n\n<p>It's almost certain the above 'safe' use of the HPA is too aggressive in the data it will throw away but might give one more peace of mind about leakage.</p>\n\n<p><em>Open question</em>: does HPA look more like test than train does test?</p>\n\n<p><strong>A note on using focal loss</strong></p>\n\n<p>Training with focal loss and HPA data is unstable. If you are experiencing large spikes in validation loss but still getting okay f1 my guess is that it's due to bad labels in the HPA data. Focal loss will focus on examples that are hard to classify and if these are mislabelled you can get wild spikes in validation loss as the model works to get these bad training labels right (i.e. probably ends up memorising them) and obviously this causes poor validation performance. If you still want to use focal loss you can add a line (in PyTorch) into the function to clip the logits, e.g.:</p>\n\n<pre><code>preds = preds.clamp(min=-15, max=15)   # need to try varying the best clamping\n</code></pre>\n\n<p>Though this seems to slow training down quite a bit it does bring stability to the validation loss.</p>\n\n<p>Keen to hear thoughts,</p>\n\n<p>Mark</p>\n\n<p>P. S. It's also pretty likely the current public LB scores are not to be trusted too much as it's unclear who has used the leak information/probed LB. </p>",
      "rawMarkdown": "**Disclaimer**: I've gotten around to exploring this too late and a little half-heartedly so I may merely be pointing out some parts of a jigsaw that others have solved.\n\n**TLDR**\n\nI can't help but think this is a competition with huge potential for CV errors. My present hypothesis (which is likely only partially true) is that if you've picked up the HPA data and simply mixed it into the train data and run k-fold over it then you might be in for a shock on the private LB.\n\n**Using HPA**\n\nFirst up, if we start from the position of the organisers who are experienced and have had people build models on their data before it bothers me that one can (potentially) get such huge score improvements from using the HPA data. \n\nBluntly, why would you set-up a competition with massive class imbalance and use a metric that punishes this heavily on a limited data-set when all the time in your broom cupboard you have 2x the data and you don't make it openly available?\n\n**In other words:** construct a competition with massive class imbalance and use macro f1 metric =&gt; predicting rare classes is the key to the competition =&gt; kagglers are not going to ignore potentially 2x the data for these classes \n\nTraining observations:\n\n - Models with train data only (5fold): local CV 0.72 - 0.74 =&gt; LB 0.47 - 0.5\n - Models trained on train + HPA data and validation sets on train data only: unstable, didn't actually submit but this was my preferred way to incorporate HPA initially to try to respect CV integrity.\n - Mix train + HPA and run 5 fold and train for 12 epochs: train loss stable, validation loss wild. Local CV f1 0.6, LB 0.55, LB with leak file 0.582, LB removing all leak entries 0.532\n\nA few options:\n\n 1. **The organisers knew all along HPA was crucial but chose not to include it.** e.g. to not overwhelm kagglers with an even more enormous dataset. If this is the case I think it's poor competition set-up as huge amounts of time has been spent by hundreds of people trying to understand this extra data - it would be much fairer and efficient to include it as part of the competition. In this scenario they've basically given us data that isn't sufficient for the problem at hand. \n 2. **The organisers know HPA is helpful but can't speak for the accuracy of some of it.** Perhaps more likely than option 1 but this ambiguity must be detrimental to the competition given all the energy spent on figuring this out. A simple statement clarifying this fact would have been helpful.\n 3. **The HPA is a red herring.** It's possible (but pretty unlikely) that the HPA data is completely misleading and the public LB looks very different to the private LB. The lack of stability from train to public LB could be that the organisers didn't have enough data they were happy with to construct a bigger overall test set and so put the better data (i.e. matching train closely) into the private LB and the public LB set is a little unrepresentative (this could be a justifiable thing to do on the basis that the 'trust your CV' mantra will pay-off on private LB though causes a lot of red flags for kagglers).\n 4. **Is HPA closer to train than train is to test?** It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because...\n\n\n**Drip drop - mind the leak**\n\nAn obvious concern is simply leakage. There are known similar images in train and HPA and known leakage to the public LB from HPA. It's possible to get a decent local CV by not taking care when constructing your validation sets (e.g. having rare similar images in both validation and train) and then when you submit you score well due to public LB leakage. In both cases you may be committing an error that may cause a (possibly severe) private LB fall.\n\n**Safely using HPA**\n\nA 'safe' way to use HPA (which I am going to do over the next week) in my view is something like: run a quick k-fold on train + HPA or on train only and predict on HPA. Throw away any HPA data that looks mislabelled. Throw away any similar images. Throw away any known leak images (and blank these out when submitting to LB). Use BCE as a loss function with some sampling strategy.\n\nIt's almost certain the above 'safe' use of the HPA is too aggressive in the data it will throw away but might give one more peace of mind about leakage.\n\n*Open question*: does HPA look more like test than train does test?\n\n**A note on using focal loss**\n\nTraining with focal loss and HPA data is unstable. If you are experiencing large spikes in validation loss but still getting okay f1 my guess is that it's due to bad labels in the HPA data. Focal loss will focus on examples that are hard to classify and if these are mislabelled you can get wild spikes in validation loss as the model works to get these bad training labels right (i.e. probably ends up memorising them) and obviously this causes poor validation performance. If you still want to use focal loss you can add a line (in PyTorch) into the function to clip the logits, e.g.:\n\n    preds = preds.clamp(min=-15, max=15)   # need to try varying the best clamping\n\nThough this seems to slow training down quite a bit it does bring stability to the validation loss.\n\nKeen to hear thoughts,\n\nMark\n\nP. S. It's also pretty likely the current public LB scores are not to be trusted too much as it's unclear who has used the leak information/probed LB.",
      "votes": null
    },
    {
      "id": "448557",
      "postDate": "01/01/2019 13:42:03",
      "content": "<p>\"Is HPA closer to train than train is to test? It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because…\"</p>\n\n<p>this can be proven by the experiments:</p>\n\n<ol>\n<li><p>train=kaggle, valid =hpa,  test=public LB</p></li>\n<li><p>train=hpa, valid =kaggle,  test=public LB</p></li>\n</ol>\n\n<p>you can measure and compare the gap of validation and test public LB</p>",
      "rawMarkdown": "\"Is HPA closer to train than train is to test? It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because…\"\n\n\nthis can be proven by the experiments:\n\n1. train=kaggle, valid =hpa,  test=public LB\n\n2. train=hpa, valid =kaggle,  test=public LB\n\nyou can measure and compare the gap of validation and test public LB",
      "votes": null
    },
    {
      "id": "448566",
      "postDate": "01/01/2019 14:10:04",
      "content": "<p>Thanks Heng. Have you done this (I probably won't have time now)? </p>\n\n<p>If you deduce HPA is closer to public LB than kaggle data this raises big concerns about the kaggle data/competition set-up.</p>",
      "rawMarkdown": "Thanks Heng. Have you done this (I probably won't have time now)? \n\nIf you deduce HPA is closer to public LB than kaggle data this raises big concerns about the kaggle data/competition set-up.",
      "votes": null
    },
    {
      "id": "448570",
      "postDate": "01/01/2019 14:25:00",
      "content": "<p>I've done step 2: train=hpa, valid=kaggle (f1 score: 0.54), test=pulic LB (score=0.433)</p>",
      "rawMarkdown": "I've done step 2: train=hpa, valid=kaggle (f1 score: 0.54), test=pulic LB (score=0.433)",
      "votes": null
    },
    {
      "id": "448601",
      "postDate": "01/01/2019 16:04:12",
      "content": "<p>do note that there are several set of external data</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>\n\n<p>　RGB_wodpl 75,040 sample</p>\n\n<p>　RGBY_wodpl 74,606 sample</p>\n\n<p>　RGBY withoutUncertain_wodpl:71,437 sample</p>\n\n<p>performances may be different for different set</p>",
      "rawMarkdown": "do note that there are several set of external data\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\n\n　RGB_wodpl 75,040 sample\n\n　RGBY_wodpl 74,606 sample\n\n　RGBY withoutUncertain_wodpl:71,437 sample\n\nperformances may be different for different set",
      "votes": null
    },
    {
      "id": "448753",
      "postDate": "01/02/2019 02:32:42",
      "content": "<p>my result: train=hpa, valid=kaggle (f1 score: 0.0.584), test=pulic LB (score=0.473)</p>",
      "rawMarkdown": "my result: train=hpa, valid=kaggle (f1 score: 0.0.584), test=pulic LB (score=0.473)",
      "votes": null
    },
    {
      "id": "449208",
      "postDate": "01/02/2019 19:08:57",
      "content": "<p>I did training of a number of models on the extended dataset (mix) and didn't notice any unstable behavior of val loss, such as spikes. I use almost the same setup as in my public kernel <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#</a> . The val loss relatively smoothly goes down (excluding small increase at cycles with higher lr during lr annealing). The only problem I saw is in training of SE model using extended dataset, but it was not about the spikes of val.</p>\n\n<p>This discussion points to a valid concern, and we do not know until the end of the competition if it is true or not. I can suggest another explanation of val drop when extended dataset is used. I would think that the presented images are just crops of even larger images done for entire cells (or their larger parts). So there can be a bunch of images with the same labels taken for a similar location within a cell. Not surprising that if we train our model on part of them, we can get quite high val on another part, but <strong>it is just a data leak and overfitting</strong>. If we go to another dataset with images from different cells, of course the score drops a lot, like for the public LB, and higher val score can result lower public LB. If the external data is collected from different cells, the val score on mixed dataset would be lower, but the model generalizes better. </p>",
      "rawMarkdown": "I did training of a number of models on the extended dataset (mix) and didn't notice any unstable behavior of val loss, such as spikes. I use almost the same setup as in my public kernel https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb# . The val loss relatively smoothly goes down (excluding small increase at cycles with higher lr during lr annealing). The only problem I saw is in training of SE model using extended dataset, but it was not about the spikes of val.\n\nThis discussion points to a valid concern, and we do not know until the end of the competition if it is true or not. I can suggest another explanation of val drop when extended dataset is used. I would think that the presented images are just crops of even larger images done for entire cells (or their larger parts). So there can be a bunch of images with the same labels taken for a similar location within a cell. Not surprising that if we train our model on part of them, we can get quite high val on another part, but **it is just a data leak and overfitting**. If we go to another dataset with images from different cells, of course the score drops a lot, like for the public LB, and higher val score can result lower public LB. If the external data is collected from different cells, the val score on mixed dataset would be lower, but the model generalizes better.",
      "votes": null
    },
    {
      "id": "449529",
      "postDate": "01/03/2019 09:55:05",
      "content": "<p>My brain tells me to apply Occam's razor: this is just a not really well set-up competition, I mean allowing 2048x2048 was already a bad choice since it might turn this competition into an arms race (did it?).</p>\n\n<p>But my heart keeps telling me something is off, they are probably trolling us and we will drop to hell with the private set. </p>\n\n<p>Btw, I hope there will be more kernels-only comps. </p>",
      "rawMarkdown": "My brain tells me to apply Occam's razor: this is just a not really well set-up competition, I mean allowing 2048x2048 was already a bad choice since it might turn this competition into an arms race (did it?).\n\nBut my heart keeps telling me something is off, they are probably trolling us and we will drop to hell with the private set. \n\nBtw, I hope there will be more kernels-only comps.",
      "votes": null
    },
    {
      "id": "450384",
      "postDate": "01/04/2019 20:38:15",
      "content": "<p>@Khoi Nguyen I agree with every single word you have written :)</p>",
      "rawMarkdown": "Khoi Nguyen I agree with every single word you have written :)",
      "votes": null
    },
    {
      "id": "450407",
      "postDate": "01/04/2019 21:51:59",
      "content": "<p>I hope you are right, because if this competitions resumes in \"I got more images than you\" is going to be deceiving.</p>",
      "rawMarkdown": "I hope you are right, because if this competitions resumes in \"I got more images than you\" is going to be deceiving.",
      "votes": null
    },
    {
      "id": "450629",
      "postDate": "01/05/2019 12:32:22",
      "content": "<p><a href=\"/stecasasso\">@stecasasso</a> I almost forgot this. Thank you for sending me a merging invitation earlier on, unfortunately I decided to team up with my friends from the beginning (and Kaggle not sending a notification for such requests doesn't help either). So maybe see you next time ;)</p>",
      "rawMarkdown": "stecasasso I almost forgot this. Thank you for sending me a merging invitation earlier on, unfortunately I decided to team up with my friends from the beginning (and Kaggle not sending a notification for such requests doesn't help either). So maybe see you next time ;)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 448557,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/01/2019 13:42:03",
      "content": "<p>\"Is HPA closer to train than train is to test? It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because…\"</p>\n\n<p>this can be proven by the experiments:</p>\n\n<ol>\n<li><p>train=kaggle, valid =hpa,  test=public LB</p></li>\n<li><p>train=hpa, valid =kaggle,  test=public LB</p></li>\n</ol>\n\n<p>you can measure and compare the gap of validation and test public LB</p>",
      "votes": null,
      "replies": [
        {
          "id": 448566,
          "author_name": "maw501",
          "author_url": "",
          "post_date": "01/01/2019 14:10:04",
          "content": "<p>Thanks Heng. Have you done this (I probably won't have time now)? </p>\n\n<p>If you deduce HPA is closer to public LB than kaggle data this raises big concerns about the kaggle data/competition set-up.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448570,
          "author_name": "mathormad",
          "author_url": "",
          "post_date": "01/01/2019 14:25:00",
          "content": "<p>I've done step 2: train=hpa, valid=kaggle (f1 score: 0.54), test=pulic LB (score=0.433)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448753,
          "author_name": "action",
          "author_url": "",
          "post_date": "01/02/2019 02:32:42",
          "content": "<p>my result: train=hpa, valid=kaggle (f1 score: 0.0.584), test=pulic LB (score=0.473)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 448601,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/01/2019 16:04:12",
      "content": "<p>do note that there are several set of external data</p>\n\n<p><a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a></p>\n\n<p>　RGB_wodpl 75,040 sample</p>\n\n<p>　RGBY_wodpl 74,606 sample</p>\n\n<p>　RGBY withoutUncertain_wodpl:71,437 sample</p>\n\n<p>performances may be different for different set</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449208,
      "author_name": "iafoss",
      "author_url": "",
      "post_date": "01/02/2019 19:08:57",
      "content": "<p>I did training of a number of models on the extended dataset (mix) and didn't notice any unstable behavior of val loss, such as spikes. I use almost the same setup as in my public kernel <a href=\"https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#\">https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb#</a> . The val loss relatively smoothly goes down (excluding small increase at cycles with higher lr during lr annealing). The only problem I saw is in training of SE model using extended dataset, but it was not about the spikes of val.</p>\n\n<p>This discussion points to a valid concern, and we do not know until the end of the competition if it is true or not. I can suggest another explanation of val drop when extended dataset is used. I would think that the presented images are just crops of even larger images done for entire cells (or their larger parts). So there can be a bunch of images with the same labels taken for a similar location within a cell. Not surprising that if we train our model on part of them, we can get quite high val on another part, but <strong>it is just a data leak and overfitting</strong>. If we go to another dataset with images from different cells, of course the score drops a lot, like for the public LB, and higher val score can result lower public LB. If the external data is collected from different cells, the val score on mixed dataset would be lower, but the model generalizes better. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 449529,
      "author_name": "suicaokhoailang",
      "author_url": "",
      "post_date": "01/03/2019 09:55:05",
      "content": "<p>My brain tells me to apply Occam's razor: this is just a not really well set-up competition, I mean allowing 2048x2048 was already a bad choice since it might turn this competition into an arms race (did it?).</p>\n\n<p>But my heart keeps telling me something is off, they are probably trolling us and we will drop to hell with the private set. </p>\n\n<p>Btw, I hope there will be more kernels-only comps. </p>",
      "votes": null,
      "replies": [
        {
          "id": 450384,
          "author_name": "stecasasso",
          "author_url": "",
          "post_date": "01/04/2019 20:38:15",
          "content": "<p>@Khoi Nguyen I agree with every single word you have written :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450407,
          "author_name": "tcapelle",
          "author_url": "",
          "post_date": "01/04/2019 21:51:59",
          "content": "<p>I hope you are right, because if this competitions resumes in \"I got more images than you\" is going to be deceiving.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 450629,
          "author_name": "suicaokhoailang",
          "author_url": "",
          "post_date": "01/05/2019 12:32:22",
          "content": "<p><a href=\"/stecasasso\">@stecasasso</a> I almost forgot this. Thank you for sending me a merging invitation earlier on, unfortunately I decided to team up with my friends from the beginning (and Kaggle not sending a notification for such requests doesn't help either). So maybe see you next time ;)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "448550": "**Disclaimer**: I've gotten around to exploring this too late and a little half-heartedly so I may merely be pointing out some parts of a jigsaw that others have solved.\n\n**TLDR**\n\nI can't help but think this is a competition with huge potential for CV errors. My present hypothesis (which is likely only partially true) is that if you've picked up the HPA data and simply mixed it into the train data and run k-fold over it then you might be in for a shock on the private LB.\n\n**Using HPA**\n\nFirst up, if we start from the position of the organisers who are experienced and have had people build models on their data before it bothers me that one can (potentially) get such huge score improvements from using the HPA data. \n\nBluntly, why would you set-up a competition with massive class imbalance and use a metric that punishes this heavily on a limited data-set when all the time in your broom cupboard you have 2x the data and you don't make it openly available?\n\n**In other words:** construct a competition with massive class imbalance and use macro f1 metric =&gt; predicting rare classes is the key to the competition =&gt; kagglers are not going to ignore potentially 2x the data for these classes \n\nTraining observations:\n\n - Models with train data only (5fold): local CV 0.72 - 0.74 =&gt; LB 0.47 - 0.5\n - Models trained on train + HPA data and validation sets on train data only: unstable, didn't actually submit but this was my preferred way to incorporate HPA initially to try to respect CV integrity.\n - Mix train + HPA and run 5 fold and train for 12 epochs: train loss stable, validation loss wild. Local CV f1 0.6, LB 0.55, LB with leak file 0.582, LB removing all leak entries 0.532\n\nA few options:\n\n 1. **The organisers knew all along HPA was crucial but chose not to include it.** e.g. to not overwhelm kagglers with an even more enormous dataset. If this is the case I think it's poor competition set-up as huge amounts of time has been spent by hundreds of people trying to understand this extra data - it would be much fairer and efficient to include it as part of the competition. In this scenario they've basically given us data that isn't sufficient for the problem at hand. \n 2. **The organisers know HPA is helpful but can't speak for the accuracy of some of it.** Perhaps more likely than option 1 but this ambiguity must be detrimental to the competition given all the energy spent on figuring this out. A simple statement clarifying this fact would have been helpful.\n 3. **The HPA is a red herring.** It's possible (but pretty unlikely) that the HPA data is completely misleading and the public LB looks very different to the private LB. The lack of stability from train to public LB could be that the organisers didn't have enough data they were happy with to construct a bigger overall test set and so put the better data (i.e. matching train closely) into the private LB and the public LB set is a little unrepresentative (this could be a justifiable thing to do on the basis that the 'trust your CV' mantra will pay-off on private LB though causes a lot of red flags for kagglers).\n 4. **Is HPA closer to train than train is to test?** It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because...\n\n\n**Drip drop - mind the leak**\n\nAn obvious concern is simply leakage. There are known similar images in train and HPA and known leakage to the public LB from HPA. It's possible to get a decent local CV by not taking care when constructing your validation sets (e.g. having rare similar images in both validation and train) and then when you submit you score well due to public LB leakage. In both cases you may be committing an error that may cause a (possibly severe) private LB fall.\n\n**Safely using HPA**\n\nA 'safe' way to use HPA (which I am going to do over the next week) in my view is something like: run a quick k-fold on train + HPA or on train only and predict on HPA. Throw away any HPA data that looks mislabelled. Throw away any similar images. Throw away any known leak images (and blank these out when submitting to LB). Use BCE as a loss function with some sampling strategy.\n\nIt's almost certain the above 'safe' use of the HPA is too aggressive in the data it will throw away but might give one more peace of mind about leakage.\n\n*Open question*: does HPA look more like test than train does test?\n\n**A note on using focal loss**\n\nTraining with focal loss and HPA data is unstable. If you are experiencing large spikes in validation loss but still getting okay f1 my guess is that it's due to bad labels in the HPA data. Focal loss will focus on examples that are hard to classify and if these are mislabelled you can get wild spikes in validation loss as the model works to get these bad training labels right (i.e. probably ends up memorising them) and obviously this causes poor validation performance. If you still want to use focal loss you can add a line (in PyTorch) into the function to clip the logits, e.g.:\n\n    preds = preds.clamp(min=-15, max=15)   # need to try varying the best clamping\n\nThough this seems to slow training down quite a bit it does bring stability to the validation loss.\n\nKeen to hear thoughts,\n\nMark\n\nP. S. It's also pretty likely the current public LB scores are not to be trusted too much as it's unclear who has used the leak information/probed LB.",
    "448557": "\"Is HPA closer to train than train is to test? It seems that adding HPA to training worsens local CV, in fact, destroys it if we keep our original validation sets with just train data. Adding HPA to the CV can give relatively poor local CV results but much stronger public LB - this just feels completely wrong and could be because…\"\n\n\nthis can be proven by the experiments:\n\n1. train=kaggle, valid =hpa,  test=public LB\n\n2. train=hpa, valid =kaggle,  test=public LB\n\nyou can measure and compare the gap of validation and test public LB",
    "448566": "Thanks Heng. Have you done this (I probably won't have time now)? \n\nIf you deduce HPA is closer to public LB than kaggle data this raises big concerns about the kaggle data/competition set-up.",
    "448570": "I've done step 2: train=hpa, valid=kaggle (f1 score: 0.54), test=pulic LB (score=0.433)",
    "448601": "do note that there are several set of external data\n\nhttps://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\n\n　RGB_wodpl 75,040 sample\n\n　RGBY_wodpl 74,606 sample\n\n　RGBY withoutUncertain_wodpl:71,437 sample\n\nperformances may be different for different set",
    "448753": "my result: train=hpa, valid=kaggle (f1 score: 0.0.584), test=pulic LB (score=0.473)",
    "449208": "I did training of a number of models on the extended dataset (mix) and didn't notice any unstable behavior of val loss, such as spikes. I use almost the same setup as in my public kernel https://www.kaggle.com/iafoss/pretrained-resnet34-with-rgby-0-460-public-lb# . The val loss relatively smoothly goes down (excluding small increase at cycles with higher lr during lr annealing). The only problem I saw is in training of SE model using extended dataset, but it was not about the spikes of val.\n\nThis discussion points to a valid concern, and we do not know until the end of the competition if it is true or not. I can suggest another explanation of val drop when extended dataset is used. I would think that the presented images are just crops of even larger images done for entire cells (or their larger parts). So there can be a bunch of images with the same labels taken for a similar location within a cell. Not surprising that if we train our model on part of them, we can get quite high val on another part, but **it is just a data leak and overfitting**. If we go to another dataset with images from different cells, of course the score drops a lot, like for the public LB, and higher val score can result lower public LB. If the external data is collected from different cells, the val score on mixed dataset would be lower, but the model generalizes better.",
    "449529": "My brain tells me to apply Occam's razor: this is just a not really well set-up competition, I mean allowing 2048x2048 was already a bad choice since it might turn this competition into an arms race (did it?).\n\nBut my heart keeps telling me something is off, they are probably trolling us and we will drop to hell with the private set. \n\nBtw, I hope there will be more kernels-only comps.",
    "450384": "Khoi Nguyen I agree with every single word you have written :)",
    "450407": "I hope you are right, because if this competitions resumes in \"I got more images than you\" is going to be deceiving.",
    "450629": "stecasasso I almost forgot this. Thank you for sending me a merging invitation earlier on, unfortunately I decided to team up with my friends from the beginning (and Kaggle not sending a notification for such requests doesn't help either). So maybe see you next time ;)"
  },
  "source": "meta"
}