{
  "id": 73395,
  "title": "Let's talk about the HPA data leakage",
  "url": "/competitions/human-protein-atlas-image-classification/discussion/73395",
  "author_name": "Mark Worrall",
  "post_date": "2018-12-02T17:28:20.183000",
  "votes": 52,
  "comment_count": 83,
  "views": 0,
  "content": "<p>Hi all,</p>\n\n<p>I realise this might not make me the most popular person on this forum but it's something I think needs to be addressed and kaggle (e.g <a href=\"/philculliton\">@philculliton</a>) should provide an answer.</p>\n\n<p>Tucked away in a discussion <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206\">here</a> is the revelation that the (difficult to access) HPA data has been used to boost <a href=\"/tomomimoriyama\">@tomomimoriyama</a> (plus I imagine others) to the top of the LB. In all fairness to the few people discussing this topic they have acknowledged it's uncertain as to whether it's allowed or not but I for one would like this explicit and clarified. </p>\n\n<p><strong>Use of the HPA data</strong>: I'm only going on what little is mentioned but it appears the HPA data is not being used in the sense I understand external data should to be used (e.g. word embeddings) in order to supplement a problem but instead people have basically found the same images in the HPA data as appear in the train and test sets. i.e. they've found some of the answers. I cannot imagine this is what the competition organisers intend when they challenge us to automate <em>'biomedical image analysis to accelerate the understanding of human cells and disease'.</em></p>\n\n<p>Further, from what I can gather, the HPA data is a pain to access and as pointed out by <a href=\"/stecasasso\">@stecasasso</a> (Chase the Trane) this sort of external data has been banned in previous competitions. I can't force the kaggle or the organisers to ban it in this case but I can safely say that if a prerequisite for tackling image competitions with huge data is to have to parse even more hard to reach huge data to simply uncover test images it's definitely not what I had in mind when I decided to start this competition and use machine learning for good. </p>\n\n<p>A final point: if I've gotten any of the above mixed up I am happy to amend my post for inaccuracies or clarify misunderstanding. I'd also like to thank those using the HPA data for being honest about finding similar images in it, particularly <a href=\"/ldm314\">@ldm314</a> (Brian) and <a href=\"/tomomimoriyama\">@tomomimoriyama</a>.</p>\n\n<p>Cheers,</p>\n\n<p>Mark</p>\n\n<p>P. S. also mentioned <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">here</a></p>",
  "messages": [
    {
      "id": 431652,
      "postDate": "2018-12-02T17:28:20.183Z",
      "content": "<p>Hi all,</p>\n\n<p>I realise this might not make me the most popular person on this forum but it's something I think needs to be addressed and kaggle (e.g <a href=\"/philculliton\">@philculliton</a>) should provide an answer.</p>\n\n<p>Tucked away in a discussion <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206\">here</a> is the revelation that the (difficult to access) HPA data has been used to boost <a href=\"/tomomimoriyama\">@tomomimoriyama</a> (plus I imagine others) to the top of the LB. In all fairness to the few people discussing this topic they have acknowledged it's uncertain as to whether it's allowed or not but I for one would like this explicit and clarified. </p>\n\n<p><strong>Use of the HPA data</strong>: I'm only going on what little is mentioned but it appears the HPA data is not being used in the sense I understand external data should to be used (e.g. word embeddings) in order to supplement a problem but instead people have basically found the same images in the HPA data as appear in the train and test sets. i.e. they've found some of the answers. I cannot imagine this is what the competition organisers intend when they challenge us to automate <em>'biomedical image analysis to accelerate the understanding of human cells and disease'.</em></p>\n\n<p>Further, from what I can gather, the HPA data is a pain to access and as pointed out by <a href=\"/stecasasso\">@stecasasso</a> (Chase the Trane) this sort of external data has been banned in previous competitions. I can't force the kaggle or the organisers to ban it in this case but I can safely say that if a prerequisite for tackling image competitions with huge data is to have to parse even more hard to reach huge data to simply uncover test images it's definitely not what I had in mind when I decided to start this competition and use machine learning for good. </p>\n\n<p>A final point: if I've gotten any of the above mixed up I am happy to amend my post for inaccuracies or clarify misunderstanding. I'd also like to thank those using the HPA data for being honest about finding similar images in it, particularly <a href=\"/ldm314\">@ldm314</a> (Brian) and <a href=\"/tomomimoriyama\">@tomomimoriyama</a>.</p>\n\n<p>Cheers,</p>\n\n<p>Mark</p>\n\n<p>P. S. also mentioned <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">here</a></p>",
      "rawMarkdown": "Hi all,\n\nI realise this might not make me the most popular person on this forum but it's something I think needs to be addressed and kaggle (e.g @philculliton) should provide an answer.\n\nTucked away in a discussion [here][1] is the revelation that the (difficult to access) HPA data has been used to boost @tomomimoriyama (plus I imagine others) to the top of the LB. In all fairness to the few people discussing this topic they have acknowledged it's uncertain as to whether it's allowed or not but I for one would like this explicit and clarified. \n\n**Use of the HPA data**: I'm only going on what little is mentioned but it appears the HPA data is not being used in the sense I understand external data should to be used (e.g. word embeddings) in order to supplement a problem but instead people have basically found the same images in the HPA data as appear in the train and test sets. i.e. they've found some of the answers. I cannot imagine this is what the competition organisers intend when they challenge us to automate *'biomedical image analysis to accelerate the understanding of human cells and disease'.*\n\nFurther, from what I can gather, the HPA data is a pain to access and as pointed out by @stecasasso (Chase the Trane) this sort of external data has been banned in previous competitions. I can't force the kaggle or the organisers to ban it in this case but I can safely say that if a prerequisite for tackling image competitions with huge data is to have to parse even more hard to reach huge data to simply uncover test images it's definitely not what I had in mind when I decided to start this competition and use machine learning for good. \n\nA final point: if I've gotten any of the above mixed up I am happy to amend my post for inaccuracies or clarify misunderstanding. I'd also like to thank those using the HPA data for being honest about finding similar images in it, particularly @ldm314 (Brian) and @tomomimoriyama.\n\nCheers,\n\nMark\n\nP. S. also mentioned [here][2]\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534",
      "votes": 52
    },
    {
      "id": 433361,
      "postDate": "2018-12-05T02:09:09.363Z",
      "content": "<p>It's a mistake to use competition participants' findings (the leaks) for test set update. Just ban the external data.</p>",
      "rawMarkdown": "It's a mistake to use competition participants' findings (the leaks) for test set update. Just ban the external data.",
      "votes": 21,
      "replies": [
        {
          "id": 433971,
          "postDate": "2018-12-05T17:54:39.610Z",
          "content": "<p>Lots of people have already spent lots of time on the external data, so it would be unfair to ban it now.</p>",
          "rawMarkdown": "Lots of people have already spent lots of time on the external data, so it would be unfair to ban it now.",
          "votes": -2
        }
      ]
    },
    {
      "id": 433019,
      "postDate": "2018-12-04T15:12:24.667Z",
      "content": "<p>Hi all - thanks for bringing this to my attention. The samples in question will no longer be included in scoring. I'll be making that change (and rescoring / resetting the leaderboard) today.</p>\n\n<p>Mark, <a href=\"/ldm314\">@ldm314</a> and <a href=\"/tomomimoriyama\">@tomomimoriyama</a>, I'll be in touch with you via DM. I appreciate your efforts on this.</p>",
      "rawMarkdown": "Hi all - thanks for bringing this to my attention. The samples in question will no longer be included in scoring. I'll be making that change (and rescoring / resetting the leaderboard) today.\n\nMark, @ldm314 and @tomomimoriyama, I'll be in touch with you via DM. I appreciate your efforts on this.",
      "votes": 12,
      "replies": [
        {
          "id": 433021,
          "postDate": "2018-12-04T15:16:38.683Z",
          "content": "<p>Thank you!</p>\n\n<p>Could you please provide a full list of leaked test IDs as well as their classes after rescoring?</p>",
          "rawMarkdown": "Thank you!\n\nCould you please provide a full list of leaked test IDs as well as their classes after rescoring?",
          "votes": 2
        },
        {
          "id": 433022,
          "postDate": "2018-12-04T15:17:48.313Z",
          "content": "<p>Thank you Phil!</p>",
          "rawMarkdown": "Thank you Phil!"
        },
        {
          "id": 433031,
          "postDate": "2018-12-04T15:32:10.070Z",
          "content": "<p>Thank you!</p>",
          "rawMarkdown": "Thank you!"
        },
        {
          "id": 433062,
          "postDate": "2018-12-04T16:08:07.513Z",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a>,</p>\n\n<p>First thanks for your quick reaction on this.</p>\n\n<p>Is the aim to simply remove those samples ? Or replace them with images that are not part of the public data ? </p>\n\n<p>I'm afraid that removing them will make the distribution even more imbalanced and increase the instability of the f1 score.</p>\n\n<p>For instance for class 27, we have only 11 samples in the training set for 31072 images. By removing those 259 samples you'd remove 5 samples with class 27 from 11702 samples. So it's probably a very important % of this class, if not the whole class.</p>",
          "rawMarkdown": "Hi @philculliton,\n\nFirst thanks for your quick reaction on this.\n\nIs the aim to simply remove those samples ? Or replace them with images that are not part of the public data ? \n\nI'm afraid that removing them will make the distribution even more imbalanced and increase the instability of the f1 score.\n\nFor instance for class 27, we have only 11 samples in the training set for 31072 images. By removing those 259 samples you'd remove 5 samples with class 27 from 11702 samples. So it's probably a very important % of this class, if not the whole class.",
          "votes": 2
        },
        {
          "id": 433091,
          "postDate": "2018-12-04T17:22:17.183Z",
          "content": "<p>Thank you Phil.</p>",
          "rawMarkdown": "Thank you Phil."
        },
        {
          "id": 433186,
          "postDate": "2018-12-04T20:13:34.483Z",
          "content": "<p>Thanks!</p>",
          "rawMarkdown": "Thanks!"
        },
        {
          "id": 433348,
          "postDate": "2018-12-05T01:48:00.893Z",
          "content": "<p>Looks like the leaderboard has not been reset yet - is that true? I guess the comments here mean that new submissions ARE being reset?</p>",
          "rawMarkdown": "Looks like the leaderboard has not been reset yet - is that true? I guess the comments here mean that new submissions ARE being reset?"
        },
        {
          "id": 433372,
          "postDate": "2018-12-05T02:20:06.457Z",
          "content": "<p>I just submitted one of my earlier submissions as a test and got the same score as before, so new submissions have not been reset as yet.</p>\n\n<p>Incidentally my latest model which got 0.595 on this un-rescored LB uses the external HPA data but does not do anything special with the offending images. It will be interesting to see the difference once the leaderboard has been re-scored. I would have assumed that re-scoring will cause people that made special use of those images would see their scores drop, but based on Brian's comment below, maybe that will not always be the case!</p>",
          "rawMarkdown": "I just submitted one of my earlier submissions as a test and got the same score as before, so new submissions have not been reset as yet.\n\nIncidentally my latest model which got 0.595 on this un-rescored LB uses the external HPA data but does not do anything special with the offending images. It will be interesting to see the difference once the leaderboard has been re-scored. I would have assumed that re-scoring will cause people that made special use of those images would see their scores drop, but based on Brian's comment below, maybe that will not always be the case!"
        },
        {
          "id": 433394,
          "postDate": "2018-12-05T02:58:16.230Z",
          "content": "<p>looks ike something's happening to the LB now. In progress.</p>",
          "rawMarkdown": "looks ike something's happening to the LB now. In progress."
        },
        {
          "id": 433481,
          "postDate": "2018-12-05T05:53:13.037Z",
          "content": "<p>Thanks Phil for your prompt attention to this problem.\nAre you also going to delete the leaked test images from test.zip and test_full_size.7z?</p>",
          "rawMarkdown": "Thanks Phil for your prompt attention to this problem.\nAre you also going to delete the leaked test images from test.zip and test_full_size.7z?"
        },
        {
          "id": 433981,
          "postDate": "2018-12-05T18:10:51.790Z",
          "content": "<p>Hi <a href=\"/dslate\">@dslate</a> - </p>\n\n<p>Good question!  I won't be removing the leaked test images.  All current competitors have already downloaded them, so this would primarily negatively impact later competitors.  Removing them from scoring mitigates the effects of the leak just as readily, without asking later competitors to go without information / images that everyone else had.</p>",
          "rawMarkdown": "Hi @dslate - \n\nGood question!  I won't be removing the leaked test images.  All current competitors have already downloaded them, so this would primarily negatively impact later competitors.  Removing them from scoring mitigates the effects of the leak just as readily, without asking later competitors to go without information / images that everyone else had."
        }
      ]
    },
    {
      "id": 433234,
      "postDate": "2018-12-04T21:48:41.850Z",
      "content": "<p>Here is the 126 I've been using, this time with the image IDS. </p>",
      "rawMarkdown": "Here is the 126 I've been using, this time with the image IDS. ",
      "votes": 9
    },
    {
      "id": 431818,
      "postDate": "2018-12-03T00:14:07.397Z",
      "content": "<p>My proposal to even out the playing field is to take all of the v18 HPA and include it here with the training data. Remove the leaked images from the test set. I found and use ~125 matches and <a href=\"/tomomimoriyama\">@tomomimoriyama</a> found over twice that. With the test set being 11k images, removing the overlap would leave plenty to test with.</p>",
      "rawMarkdown": "My proposal to even out the playing field is to take all of the v18 HPA and include it here with the training data. Remove the leaked images from the test set. I found and use ~125 matches and @tomomimoriyama found over twice that. With the test set being 11k images, removing the overlap would leave plenty to test with.",
      "votes": 7,
      "replies": [
        {
          "id": 431976,
          "postDate": "2018-12-03T08:03:11.823Z",
          "content": "<p>Yeah, or just announce they aren't part of the private LB in which case using for public LB will bias your results. This of course relies on them being able to reliably identify the true extent of the overlap/leak.</p>",
          "rawMarkdown": "Yeah, or just announce they aren't part of the private LB in which case using for public LB will bias your results. This of course relies on them being able to reliably identify the true extent of the overlap/leak.",
          "votes": 1
        },
        {
          "id": 432107,
          "postDate": "2018-12-03T12:24:11.930Z",
          "content": "<p>I agree with Brian proposal concerning making HPA data available.   </p>\n\n<p>For the leakage, I don't think the organizers will take any action on that, because the leakage is not severe and the vast majority of the test set is not affected. This expectation is based on the few competitions \"with leak\" that I took part in. For example, in Santander the leakage was much more evident and the organizers did not do anything about it. However, someone shared part of the leakage in kernels, but the full set of leaky rows of the test set was never released and only the top of the LB managed to exploit it 100%. </p>",
          "rawMarkdown": "I agree with Brian proposal concerning making HPA data available.   \n\nFor the leakage, I don't think the organizers will take any action on that, because the leakage is not severe and the vast majority of the test set is not affected. This expectation is based on the few competitions \"with leak\" that I took part in. For example, in Santander the leakage was much more evident and the organizers did not do anything about it. However, someone shared part of the leakage in kernels, but the full set of leaky rows of the test set was never released and only the top of the LB managed to exploit it 100%. "
        },
        {
          "id": 432140,
          "postDate": "2018-12-03T13:04:45.983Z",
          "content": "<p>Jumping from 0.528 to 0.588 because of the leak is severe in my view. </p>\n\n<p>I also don't think adding the HPA data to the training data really levels the playing field - it favours even more those with massive compute power.</p>",
          "rawMarkdown": "Jumping from 0.528 to 0.588 because of the leak is severe in my view. \n\nI also don't think adding the HPA data to the training data really levels the playing field - it favours even more those with massive compute power."
        },
        {
          "id": 432251,
          "postDate": "2018-12-03T16:01:10.893Z",
          "content": "<p>Regarding the jump in the LB: I may agree with you (what I wrote before is what I think it's the reasoning of the organizers, not necessarily mine...) .   </p>\n\n<p>For the second point, some personal considerations: <br>\n - Kaggle is biased towards competitors with large computing power and there's little or nothing we can do about it . <br>\n - Making the additional data available to everyone is the only thing that can be done to level the field again. External data has been announced (by Brian) on the public thread <strong>a month ago</strong> and no one from the Kaggle team has replied that it was prohibited. If they banned it now, it would be very unfair towards people who have built their models and spent their submissions using that data (N.B.: it's not me, I am not using HPA data so far). <br>\n - More data instances is, technically speaking, more \"democratic\" than higher resolution data. You could train a model in several steps on different Kaggle kernels, by re-loading the weights from the previous step. On the other hand, higher resolution images require gpu with larger memory, which are very expensive and clearly not available to everyone.</p>\n\n<p>Now, my opinion on what happened in this competition: <br>\n- it was a mistake of the organizers to make higher resolution images available, as it increases the \"technology gap\" <strong>a lot</strong>; <br>\n- the HPA dataset should be made available by someone who owns it - better if in agreement with the organizers, so we are sure that really <strong>all</strong> the available external data is included; <br>\n- we should not blame on people who discovered and used the leak: this is just fine according to the rules, as they announced publicly the dataset they were using</p>",
          "rawMarkdown": "Regarding the jump in the LB: I may agree with you (what I wrote before is what I think it's the reasoning of the organizers, not necessarily mine...) .   \n\nFor the second point, some personal considerations:   \n - Kaggle is biased towards competitors with large computing power and there's little or nothing we can do about it .   \n - Making the additional data available to everyone is the only thing that can be done to level the field again. External data has been announced (by Brian) on the public thread **a month ago** and no one from the Kaggle team has replied that it was prohibited. If they banned it now, it would be very unfair towards people who have built their models and spent their submissions using that data (N.B.: it's not me, I am not using HPA data so far).    \n - More data instances is, technically speaking, more \"democratic\" than higher resolution data. You could train a model in several steps on different Kaggle kernels, by re-loading the weights from the previous step. On the other hand, higher resolution images require gpu with larger memory, which are very expensive and clearly not available to everyone.\n\nNow, my opinion on what happened in this competition:   \n- it was a mistake of the organizers to make higher resolution images available, as it increases the \"technology gap\" **a lot**;    \n- the HPA dataset should be made available by someone who owns it - better if in agreement with the organizers, so we are sure that really **all** the available external data is included;    \n- we should not blame on people who discovered and used the leak: this is just fine according to the rules, as they announced publicly the dataset they were using",
          "votes": 2
        },
        {
          "id": 432269,
          "postDate": "2018-12-03T16:33:11.217Z",
          "content": "<p>Hi Chase the Trane (<a href=\"/stecasasso\">@stecasasso</a>),</p>\n\n<p>I agree with nearly all of what you said but think you are slightly conflating two points.</p>\n\n<p>I don't think it's an issue for people to use the HPA data (and don't want it banned) - we just need to make sure that the HPA data doesn't contain any of the answers to the live competition. </p>\n\n<p>Mark</p>",
          "rawMarkdown": "Hi Chase the Trane (@stecasasso),\n\nI agree with nearly all of what you said but think you are slightly conflating two points.\n\nI don't think it's an issue for people to use the HPA data (and don't want it banned) - we just need to make sure that the HPA data doesn't contain any of the answers to the live competition. \n\nMark",
          "votes": 2
        },
        {
          "id": 432517,
          "postDate": "2018-12-04T01:39:40.273Z",
          "content": "<p>Hi Mark Worrall,\nI agree with you that people should be able to use the HPA data but that data shouldn't contain any leaked labels for the competition test set.  As far as I can see, the only way to accomplish this would be to find all the leaks and remove their images from the test set, both public and private.  This has the downside that every submission will have to be rescored and the LB recomputed, but with over a month to go until the contest ends this may be the best solution.  The organizers would have to decide to do this pretty soon, however.</p>",
          "rawMarkdown": "Hi Mark Worrall,\nI agree with you that people should be able to use the HPA data but that data shouldn't contain any leaked labels for the competition test set.  As far as I can see, the only way to accomplish this would be to find all the leaks and remove their images from the test set, both public and private.  This has the downside that every submission will have to be rescored and the LB recomputed, but with over a month to go until the contest ends this may be the best solution.  The organizers would have to decide to do this pretty soon, however.",
          "votes": 5
        },
        {
          "id": 432810,
          "postDate": "2018-12-04T10:40:37.930Z",
          "content": "<p>Even if no other action is taken, we need an <strong>urgent examination of the private test set</strong> to find out just how contaminated it is. 126 images is only &gt;1% of the public test. For all we know the percentage is much higher in the unreleased data. I would really like some reassurance that contamination is really only the minor issue that it appears to be at present time.</p>",
          "rawMarkdown": "Even if no other action is taken, we need an **urgent examination of the private test set** to find out just how contaminated it is. 126 images is only &gt;1% of the public test. For all we know the percentage is much higher in the unreleased data. I would really like some reassurance that contamination is really only the minor issue that it appears to be at present time."
        }
      ]
    },
    {
      "id": 433830,
      "postDate": "2018-12-05T14:35:58.963Z",
      "content": "<p>Hey all -</p>\n\n<p>I updated and reset the LB last night.  Here's what changed:</p>\n\n<p>1) The majority of the impacted samples are no longer included in scoring.</p>\n\n<p>2) A small number of samples with very rare labels were not removed from scoring, but were shifted to the public leaderboard.</p>\n\n<p>Both of these moves will affect public LB scores, especially for less common labels.  If you're concerned about your score changing after this fix was implemented, the fix did impact everyone's scores - whether they were using the leak or not.  Any shift on the leaderboard, whether ignoring samples or moving them between public and private LBs, will affect scores. I do understand the confusion expressed about scores changing: I sincerely apologize and hope this clears things up!</p>\n\n<p>Thanks again to everyone for your feedback and help!  It's much appreciated.</p>",
      "rawMarkdown": "Hey all -\n\nI updated and reset the LB last night.  Here's what changed:\n\n1) The majority of the impacted samples are no longer included in scoring.\n\n2) A small number of samples with very rare labels were not removed from scoring, but were shifted to the public leaderboard.\n\nBoth of these moves will affect public LB scores, especially for less common labels.  If you're concerned about your score changing after this fix was implemented, the fix did impact everyone's scores - whether they were using the leak or not.  Any shift on the leaderboard, whether ignoring samples or moving them between public and private LBs, will affect scores. I do understand the confusion expressed about scores changing: I sincerely apologize and hope this clears things up!\n\nThanks again to everyone for your feedback and help!  It's much appreciated.",
      "votes": 7,
      "replies": [
        {
          "id": 433899,
          "postDate": "2018-12-05T16:21:31.900Z",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Thanks for the info. I think this should be a separate topic and pinned to the top of discussion board, as not everyone is subscribed to the whole forum.</p>",
          "rawMarkdown": "@philculliton Thanks for the info. I think this should be a separate topic and pinned to the top of discussion board, as not everyone is subscribed to the whole forum.",
          "votes": 5
        },
        {
          "id": 433917,
          "postDate": "2018-12-05T16:45:29.167Z",
          "content": "<p>I guess that rare classes that were already on the public LB were left there? And more shifted from the private? It’s hard to see how that would lower public scores of those using the leak. It might generate a massive shakeup at the end, though. Wouldnt it have been better to just remove leaks entirely, even from rare classes? And if some classes were rendered 1/10000 or something, remove the classes.</p>",
          "rawMarkdown": "I guess that rare classes that were already on the public LB were left there? And more shifted from the private? It’s hard to see how that would lower public scores of those using the leak. It might generate a massive shakeup at the end, though. Wouldnt it have been better to just remove leaks entirely, even from rare classes? And if some classes were rendered 1/10000 or something, remove the classes."
        },
        {
          "id": 433920,
          "postDate": "2018-12-05T16:49:30.530Z",
          "content": "<p>Yeah, so using the file can cover for a model missing some of the rare classes as these are kept in public LB...but this means private LB will likely be very different.</p>",
          "rawMarkdown": "Yeah, so using the file can cover for a model missing some of the rare classes as these are kept in public LB...but this means private LB will likely be very different."
        },
        {
          "id": 433946,
          "postDate": "2018-12-05T17:20:24.307Z",
          "content": "<p>There going to be a nice shakeup in this competition :)</p>",
          "rawMarkdown": "There going to be a nice shakeup in this competition :)",
          "votes": 1
        },
        {
          "id": 433972,
          "postDate": "2018-12-05T17:58:50.057Z",
          "content": "<p>Thank you for the update.\nCould you please confirm that there is no leakage in the private set?\nCould you also publish a list of leaked images with their classes to even the odds?</p>",
          "rawMarkdown": "Thank you for the update.\nCould you please confirm that there is no leakage in the private set?\nCould you also publish a list of leaked images with their classes to even the odds?"
        },
        {
          "id": 433975,
          "postDate": "2018-12-05T18:03:34.657Z",
          "content": "<p>Hi Dmytro -</p>\n\n<p>I can confirm that there is no leakage in the private set.  Re: posting a list of the leaked images... I'm considering it.  Because some of the images aren't simply being ignored, that makes the impact of releasing the list more complex.  I may release the list of the ignored images.</p>",
          "rawMarkdown": "Hi Dmytro -\n\nI can confirm that there is no leakage in the private set.  Re: posting a list of the leaked images... I'm considering it.  Because some of the images aren't simply being ignored, that makes the impact of releasing the list more complex.  I may release the list of the ignored images.",
          "votes": 2
        },
        {
          "id": 433986,
          "postDate": "2018-12-05T18:21:10.297Z",
          "content": "<p>Here's some random chatter from someone without much experience in DL or Kaggle:</p>\n\n<p>I'll incorporate the leak to get the 0.05 public LB asap, so  can brag to my wife. But the question is now: how do I avoid the shakeup?</p>\n\n<ol>\n<li>the leak advantage will go away (democratically, I guess)</li>\n<li>the private LB will have a relative dearth of rare classes.</li>\n</ol>\n\n<p>Point 2 might not be a huge killer because we will be in the prediction stage so the model will have been trained with non-decimated rare classes. However, the leak may reduce the impetus to deal with rare classes, thus deteriorating the effectiveness of the winning models.</p>\n\n<p>So the strategy I guess will be to incorporate the leak (altering submission file only) for bragging rights in the next month (and to see exactly where you are in the competition) and develop your model just like nothing happened.</p>\n\n<p>Or you could copy all leak data (rare classes) into your personal training set to help you better train those classes. It won't hurt the public LB as you are replacing the predictions of those samples anyway.</p>\n\n<p>Regarding point 1, for completeness, I wonder how many of the leaks have been found. It appears that the organizers are just using the results contributed on these threads - not doing any checks. Perhaps there are many more leaks with very small differences (say by different resampling/antiallias, different versions of photo(?)). And it's not easy to compare a hundred thousand huge images - hashing has been used - how accurate is that in this context? Is there a better way that runs in finite time?</p>\n\n<p>Definitely a footnote in the Nature paper.</p>",
          "rawMarkdown": "Here's some random chatter from someone without much experience in DL or Kaggle:\n\nI'll incorporate the leak to get the 0.05 public LB asap, so  can brag to my wife. But the question is now: how do I avoid the shakeup?\n\n  1. the leak advantage will go away (democratically, I guess)\n  2. the private LB will have a relative dearth of rare classes.\n\nPoint 2 might not be a huge killer because we will be in the prediction stage so the model will have been trained with non-decimated rare classes. However, the leak may reduce the impetus to deal with rare classes, thus deteriorating the effectiveness of the winning models.\n\nSo the strategy I guess will be to incorporate the leak (altering submission file only) for bragging rights in the next month (and to see exactly where you are in the competition) and develop your model just like nothing happened.\n\nOr you could copy all leak data (rare classes) into your personal training set to help you better train those classes. It won't hurt the public LB as you are replacing the predictions of those samples anyway.\n\nRegarding point 1, for completeness, I wonder how many of the leaks have been found. It appears that the organizers are just using the results contributed on these threads - not doing any checks. Perhaps there are many more leaks with very small differences (say by different resampling/antiallias, different versions of photo(?)). And it's not easy to compare a hundred thousand huge images - hashing has been used - how accurate is that in this context? Is there a better way that runs in finite time?\n\nDefinitely a footnote in the Nature paper.",
          "votes": 2
        },
        {
          "id": 434012,
          "postDate": "2018-12-05T19:08:29.897Z",
          "content": "<p>Hi pete -</p>\n\n<p>Thanks for the feedback - I appreciate the thought everyone is putting into this!</p>\n\n<p>Re: your point 2.  The rare classes in question still exist in the private LB.  Samples were shifted between the private and public LB to ensure this.  Anyone who feels a reduced impetus to deal with rare classes will definitely take a hit on the private LB.</p>\n\n<p>Re: the leaks... we absolutely did our own checks, and did <em>not</em> use the results contributed on the forums.  The forum results were useful because they indicated that people had identified a problem.  We know the full extent of the potential leakage based on our own work.</p>\n\n<p>On a related note, I should mention that hashing is <em>not</em> 100% accurate in this context, and that some of the \"leaked\" images are not, in fact, leaks (and further, some of the labels assigned them are incorrect).  The leak appears to be so effective because some of the real, actual leaks include rare classes that have a high impact on LB score.</p>\n\n<p>The people who found the leak did excellent work, but it's very difficult to get completely right.  I'd recommend not using the leak in any way.</p>",
          "rawMarkdown": "Hi pete -\n\nThanks for the feedback - I appreciate the thought everyone is putting into this!\n\nRe: your point 2.  The rare classes in question still exist in the private LB.  Samples were shifted between the private and public LB to ensure this.  Anyone who feels a reduced impetus to deal with rare classes will definitely take a hit on the private LB.\n\nRe: the leaks... we absolutely did our own checks, and did *not* use the results contributed on the forums.  The forum results were useful because they indicated that people had identified a problem.  We know the full extent of the potential leakage based on our own work.\n\nOn a related note, I should mention that hashing is *not* 100% accurate in this context, and that some of the \"leaked\" images are not, in fact, leaks (and further, some of the labels assigned them are incorrect).  The leak appears to be so effective because some of the real, actual leaks include rare classes that have a high impact on LB score.\n\nThe people who found the leak did excellent work, but it's very difficult to get completely right.  I'd recommend not using the leak in any way.",
          "votes": 5
        },
        {
          "id": 434100,
          "postDate": "2018-12-05T22:10:15.150Z",
          "content": "<p>Phil,</p>\n\n<p>Thanks very much for your comments. My apologies for assuming that you did not look carefully into this - looks like you definitely did. While I'm not fully on board with all of your suggestions, they are very helpful and I appreciate what you have done with this, and your policy of keeping us up to date.</p>\n\n<p>Also to those who found the leak and did great work calling attention to it and helped others understand, thank you also.  I am super impressed with how fast you understood that complicated website (I still don't really understand it). You improved the competition immeasurably.</p>",
          "rawMarkdown": "Phil,\n\nThanks very much for your comments. My apologies for assuming that you did not look carefully into this - looks like you definitely did. While I'm not fully on board with all of your suggestions, they are very helpful and I appreciate what you have done with this, and your policy of keeping us up to date.\n\nAlso to those who found the leak and did great work calling attention to it and helped others understand, thank you also.  I am super impressed with how fast you understood that complicated website (I still don't really understand it). You improved the competition immeasurably.",
          "votes": 2
        },
        {
          "id": 434463,
          "postDate": "2018-12-06T12:43:31.860Z",
          "content": "<p>Hi Phil - thanks for the info.  Do we need to download the test set again if samples were shifted between private and public LB?</p>",
          "rawMarkdown": "Hi Phil - thanks for the info.  Do we need to download the test set again if samples were shifted between private and public LB?"
        },
        {
          "id": 434472,
          "postDate": "2018-12-06T13:08:17.340Z",
          "content": "<p>@Fionnán Alt: I guess no... \ntest set = public LB + private LB</p>",
          "rawMarkdown": "@Fionnán Alt: I guess no... \ntest set = public LB + private LB"
        },
        {
          "id": 434473,
          "postDate": "2018-12-06T13:08:18.310Z",
          "content": "<p>Hi Fionnán,\nGood question! You don't need to download anything new - the shift happened entirely on our end, so you shouldn't have to make any changes.</p>",
          "rawMarkdown": "Hi Fionnán,\nGood question! You don't need to download anything new - the shift happened entirely on our end, so you shouldn't have to make any changes."
        },
        {
          "id": 434501,
          "postDate": "2018-12-06T14:08:53.157Z",
          "content": "<p>Thanks Phil.</p>",
          "rawMarkdown": "Thanks Phil."
        }
      ]
    },
    {
      "id": 432969,
      "postDate": "2018-12-04T14:26:38.920Z",
      "content": "<p>Test matches found by Tomomi Moriyama with labels from HPA.</p>",
      "rawMarkdown": "Test matches found by Tomomi Moriyama with labels from HPA.",
      "votes": 5,
      "replies": [
        {
          "id": 432980,
          "postDate": "2018-12-04T14:41:25.410Z",
          "content": "<p>Thank you Alexander - I hadn't been through all those files.</p>\n\n<p>P. S. that just took me from 0.482 -&gt; 0.527 ... wonder how many others in the top 100 are using it already. You might also wish to make that csv more public (i.e. kernel).</p>",
          "rawMarkdown": "Thank you Alexander - I hadn't been through all those files.\n\nP. S. that just took me from 0.482 -&gt; 0.527 ... wonder how many others in the top 100 are using it already. You might also wish to make that csv more public (i.e. kernel)."
        },
        {
          "id": 433045,
          "postDate": "2018-12-04T15:49:52.583Z",
          "content": "<p>It's funny, using the leak our 0.537 submission using oversampling only got to 0.544, while the 0.507 submission which clearly overfitted to the major classes got to 0.556. I'm glad they decided to take action.</p>",
          "rawMarkdown": "It's funny, using the leak our 0.537 submission using oversampling only got to 0.544, while the 0.507 submission which clearly overfitted to the major classes got to 0.556. I'm glad they decided to take action."
        },
        {
          "id": 433190,
          "postDate": "2018-12-04T20:20:06.293Z",
          "content": "<p>Amusingly enough if I apply this file, my score goes down from .531 to .528...</p>",
          "rawMarkdown": "Amusingly enough if I apply this file, my score goes down from .531 to .528...",
          "votes": 1
        },
        {
          "id": 433202,
          "postDate": "2018-12-04T20:39:31.683Z",
          "content": "<p><a href=\"/ldm314\">@ldm314</a> do you think <a href=\"/tomomimoriyama\">@tomomimoriyama</a> is trying to miss lead us? :D</p>",
          "rawMarkdown": " @ldm314 do you think @tomomimoriyama is trying to miss lead us? :D",
          "votes": -1
        },
        {
          "id": 433231,
          "postDate": "2018-12-04T21:47:13.377Z",
          "content": "<p>I don't think so. I came up with a similar list independently after <a href=\"/tomomimoriyama\">@tomomimoriyama</a>  demonstrated that the yellow channels were available. I ended up with the same score drop vs the 126 list I came up with from only RGB.</p>",
          "rawMarkdown": "I don't think so. I came up with a similar list independently after @tomomimoriyama  demonstrated that the yellow channels were available. I ended up with the same score drop vs the 126 list I came up with from only RGB."
        },
        {
          "id": 433461,
          "postDate": "2018-12-05T05:18:30.957Z",
          "content": "<p>It is very upset that I don't use leak but all my scores were dropped :(</p>",
          "rawMarkdown": "It is very upset that I don't use leak but all my scores were dropped :("
        }
      ]
    },
    {
      "id": 433563,
      "postDate": "2018-12-05T07:24:09.127Z",
      "content": "<p>I don't think the LB has been updated properly (or something weird is going on) as my score has gone from:</p>\n\n<p>0.482 no leaked data\n0.527 using leaked data\n0.528 post LB update?</p>",
      "rawMarkdown": "I don't think the LB has been updated properly (or something weird is going on) as my score has gone from:\n\n0.482 no leaked data\n0.527 using leaked data\n0.528 post LB update?",
      "votes": 6,
      "replies": [
        {
          "id": 433575,
          "postDate": "2018-12-05T07:39:32.640Z",
          "content": "<p>Yes! I suppose the same! My score without leak has dropped!</p>",
          "rawMarkdown": "Yes! I suppose the same! My score without leak has dropped!"
        },
        {
          "id": 433576,
          "postDate": "2018-12-05T07:40:53.367Z",
          "content": "<p>Definitely something weird going on</p>\n\n<p>0.525 no leak\n0.561 leak \n0.561 after update</p>",
          "rawMarkdown": "Definitely something weird going on\n\n0.525 no leak\n0.561 leak \n0.561 after update",
          "votes": 3
        },
        {
          "id": 433600,
          "postDate": "2018-12-05T08:05:46.760Z",
          "content": "<p>Drop the samples, they said\nIt would be fair, they said</p>\n\n<p>P.S. no offense :)</p>",
          "rawMarkdown": "Drop the samples, they said\nIt would be fair, they said\n\nP.S. no offense :)",
          "votes": 5
        },
        {
          "id": 433603,
          "postDate": "2018-12-05T08:08:26.280Z",
          "content": "<p>Haha - no offence taken. I don't think it's clear what's gone on at the moment but the ids from the csv you provided are still being scored in the public LB somehow.</p>",
          "rawMarkdown": "Haha - no offence taken. I don't think it's clear what's gone on at the moment but the ids from the csv you provided are still being scored in the public LB somehow."
        },
        {
          "id": 433610,
          "postDate": "2018-12-05T08:13:51.300Z",
          "content": "<p>Same! Our score without leak has decreased from 0.506 to 0.491 .</p>",
          "rawMarkdown": "Same! Our score without leak has decreased from 0.506 to 0.491 ."
        },
        {
          "id": 433648,
          "postDate": "2018-12-05T09:43:38.643Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 433653,
          "postDate": "2018-12-05T09:49:56.027Z",
          "content": "<p>I did not use the leaked labels, and my highest LB score has dropped from 0.504 to 0.499.  But that doesn't necessarily indicate any scoring error by the organizers.  It could be that even without the leaks, my model did a good job of predicting the affected images, so without them my score dropped a bit.\nThe LB still looks dubious at the high end, however, since those scores seem to have been hardly affected by the \"reset\" at all.</p>",
          "rawMarkdown": "I did not use the leaked labels, and my highest LB score has dropped from 0.504 to 0.499.  But that doesn't necessarily indicate any scoring error by the organizers.  It could be that even without the leaks, my model did a good job of predicting the affected images, so without them my score dropped a bit.\nThe LB still looks dubious at the high end, however, since those scores seem to have been hardly affected by the \"reset\" at all."
        },
        {
          "id": 433697,
          "postDate": "2018-12-05T11:14:43.027Z",
          "content": "<p>my score also decreased from 0.501 to 0.495.</p>",
          "rawMarkdown": "my score also decreased from 0.501 to 0.495."
        },
        {
          "id": 433756,
          "postDate": "2018-12-05T12:50:26.927Z",
          "content": "<p>0.570 -&gt; 0.553\nWas I using too many leaks? :P</p>",
          "rawMarkdown": "0.570 -&gt; 0.553\nWas I using too many leaks? :P",
          "votes": 1
        },
        {
          "id": 433963,
          "postDate": "2018-12-05T17:43:55.623Z",
          "content": "<p>haha, I think so.\nmy score dropped from 0.498 to 0.494,\nusing the leaks, it goes to 0.549</p>",
          "rawMarkdown": "haha, I think so.\nmy score dropped from 0.498 to 0.494,\nusing the leaks, it goes to 0.549",
          "votes": 1
        }
      ]
    },
    {
      "id": 434133,
      "postDate": "2018-12-06T00:06:11.560Z",
      "content": "<p>After the leaderboard changes, in general, models submitted with the 126 samples I have posted went down in score. Models using the list from <a href=\"/tomomimoriyama\">@tomomimoriyama</a> went up in score. What meaning this has, if any, I won't speculate on as some of these posted answers are confirmed to be inaccurate.</p>\n\n<p>Thanks to <a href=\"/philculliton\">@philculliton</a> for cleaning thing up and evening out the field.</p>",
      "rawMarkdown": "After the leaderboard changes, in general, models submitted with the 126 samples I have posted went down in score. Models using the list from @tomomimoriyama went up in score. What meaning this has, if any, I won't speculate on as some of these posted answers are confirmed to be inaccurate.\n\nThanks to @philculliton for cleaning thing up and evening out the field.\n\n",
      "votes": 3
    },
    {
      "id": 436658,
      "postDate": "2018-12-10T17:56:22.613Z",
      "content": "<p>since this challenge is permited to use the external datasets,  can anyone provide a complete HPA dataset to download including full HPA image and train.csv? don't think we should cost much time exploring external datasets.  </p>",
      "rawMarkdown": "since this challenge is permited to use the external datasets,  can anyone provide a complete HPA dataset to download including full HPA image and train.csv? don't think we should cost much time exploring external datasets.  ",
      "votes": 4,
      "replies": [
        {
          "id": 442604,
          "postDate": "2018-12-20T07:29:07.140Z",
          "rawMarkdown": "",
          "votes": 3,
          "isDeleted": true
        },
        {
          "id": 442840,
          "postDate": "2018-12-20T15:39:13.173Z",
          "content": "<p>Hi Jarvis, Thank you for these useful information! Are we allowed to use HPA datasets in training a model? Does it not have any implication such as data leakage? Sorry, I didn't check all HPA datasets. Thanks </p>",
          "rawMarkdown": "Hi Jarvis, Thank you for these useful information! Are we allowed to use HPA datasets in training a model? Does it not have any implication such as data leakage? Sorry, I didn't check all HPA datasets. Thanks ",
          "votes": 9
        }
      ]
    },
    {
      "id": 443202,
      "postDate": "2018-12-21T07:55:34.103Z",
      "content": "<p>Hello guys,\nCould anybody clarify, does it make sense to include the leak in the submission now or not?\nThank you!</p>",
      "rawMarkdown": "Hello guys,\nCould anybody clarify, does it make sense to include the leak in the submission now or not?\nThank you!",
      "votes": 1,
      "replies": [
        {
          "id": 443225,
          "postDate": "2018-12-21T08:41:46.163Z",
          "content": "<p>The advantage provided by the leak will probably vanish in the private leaderboard but it doesnt hurt to include it if only to see where you stand. The hpa data themselves give an extra edge in addition to any leak advantage. I dont see any reason why this extra edge wont remain in the private leaderboard.</p>",
          "rawMarkdown": "The advantage provided by the leak will probably vanish in the private leaderboard but it doesnt hurt to include it if only to see where you stand. The hpa data themselves give an extra edge in addition to any leak advantage. I dont see any reason why this extra edge wont remain in the private leaderboard."
        }
      ]
    },
    {
      "id": 433691,
      "postDate": "2018-12-05T10:58:30.890Z",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> After LB rescoring, my leak only submission (no training) scored 0.085, now it is scoring 0.062 (about two rare classes with F1 close to 1.00), so definitely there are still leakage in public LB. Don't know about private LB though.</p>",
      "rawMarkdown": "@philculliton After LB rescoring, my leak only submission (no training) scored 0.085, now it is scoring 0.062 (about two rare classes with F1 close to 1.00), so definitely there are still leakage in public LB. Don't know about private LB though.",
      "votes": 2
    },
    {
      "id": 432185,
      "postDate": "2018-12-03T14:25:42.633Z",
      "content": "<p>I concur.</p>\n\n<p>If they make this data official and give us an easy zipped download of the HPA in a smaller image size (512p), that would be the best approach. If HPA can afford to have people trawling their servers for tens of thousands of 2048p images from individual urls they can afford to give a direct link to a compressed archive. They must have known that the HPA would be the obvious first point of call for anybody looking for external data.</p>\n\n<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a> and <a href=\"/ldm314\">@ldm314</a> (Brian) have kindly provided access to parsing scripts that allow download of the larger images and label them accordingly. Having to download at mostly 2048px and resize is still a pain. <a href=\"/artemtprv\">@artemtprv</a> provided a web-scraping kernel to access 'preview' images from the HPA website that are 800x800, but these appear to be either cropped (or at high magnification), and there are far fewer than in the XML urls.</p>\n\n<p>I am guessing the organisers pre-reduced the training data for the competition to 512px to allow easy download and training for amateur enthusiasts and typical users. Some people don't even have the connection or hdd space to acquire the larger original image set, let alone access to memory and GPUs capable of training models with them properly. Having to download a full size external dataset (even if resized later) to make it onto the top half of the leaderboard is a little extreme.</p>\n\n<p>There's still a good amount of time left to fix this so please address it. Even if you just exclude the leaked images, that leaves the option for kind community members to upload the data somewhere independent at a reduced image size with accompanied csv files.</p>",
      "rawMarkdown": "I concur.\n\nIf they make this data official and give us an easy zipped download of the HPA in a smaller image size (512p), that would be the best approach. If HPA can afford to have people trawling their servers for tens of thousands of 2048p images from individual urls they can afford to give a direct link to a compressed archive. They must have known that the HPA would be the obvious first point of call for anybody looking for external data.\n\n@tomomimoriyama and @ldm314 (Brian) have kindly provided access to parsing scripts that allow download of the larger images and label them accordingly. Having to download at mostly 2048px and resize is still a pain. @artemtprv provided a web-scraping kernel to access 'preview' images from the HPA website that are 800x800, but these appear to be either cropped (or at high magnification), and there are far fewer than in the XML urls.\n\nI am guessing the organisers pre-reduced the training data for the competition to 512px to allow easy download and training for amateur enthusiasts and typical users. Some people don't even have the connection or hdd space to acquire the larger original image set, let alone access to memory and GPUs capable of training models with them properly. Having to download a full size external dataset (even if resized later) to make it onto the top half of the leaderboard is a little extreme.\n\nThere's still a good amount of time left to fix this so please address it. Even if you just exclude the leaked images, that leaves the option for kind community members to upload the data somewhere independent at a reduced image size with accompanied csv files.",
      "votes": 2
    },
    {
      "id": 432072,
      "postDate": "2018-12-03T11:14:54.117Z",
      "content": "<p>Cannot agree more with you. The competition needs to focus on methods and skills instead of \"cheating\" to get a high score.</p>",
      "rawMarkdown": "Cannot agree more with you. The competition needs to focus on methods and skills instead of \"cheating\" to get a high score.",
      "votes": 2
    },
    {
      "id": 432160,
      "postDate": "2018-12-03T13:46:03.287Z",
      "content": "<p>Wait, what? There is a leak?</p>",
      "rawMarkdown": "Wait, what? There is a leak?"
    },
    {
      "id": 433965,
      "postDate": "2018-12-05T17:48:11.510Z",
      "content": "<p>Hmm... based on the scores that people are posting it looks like most of the leaks are still being scored.</p>",
      "rawMarkdown": "Hmm... based on the scores that people are posting it looks like most of the leaks are still being scored.",
      "replies": [
        {
          "id": 433974,
          "postDate": "2018-12-05T18:00:19.807Z",
          "content": "<p>Hi Robert -</p>\n\n<p>We've changed the public LB exposure for two rare classes, which is having a noticeable effect on public LB scores.  An overwhelming majority of the leaked samples are no longer being scored at all.</p>",
          "rawMarkdown": "Hi Robert -\n\nWe've changed the public LB exposure for two rare classes, which is having a noticeable effect on public LB scores.  An overwhelming majority of the leaked samples are no longer being scored at all."
        },
        {
          "id": 435873,
          "postDate": "2018-12-09T00:00:58.267Z",
          "content": "<p>Is there any mistake here? Before you adjusted, my score was 0.529. After the adjustment, it was 0.513. I didn't use data leak before, but when I use data leak,now my score is 0.555. </p>",
          "rawMarkdown": "Is there any mistake here? Before you adjusted, my score was 0.529. After the adjustment, it was 0.513. I didn't use data leak before, but when I use data leak,now my score is 0.555. "
        },
        {
          "id": 435874,
          "postDate": "2018-12-09T00:05:24.707Z",
          "content": "<p>It sounds like reasonable behaviour. I would say it's OK.</p>",
          "rawMarkdown": "It sounds like reasonable behaviour. I would say it's OK."
        },
        {
          "id": 435877,
          "postDate": "2018-12-09T00:12:00.187Z",
          "content": "<p>Hi,pete.\nI don't think it's a problem either. I'm mainly concerned about whether there will be a second adjustment, because the adjustment still has a great impact on the score. This may cause some trouble for me to adjust the model. </p>",
          "rawMarkdown": "Hi,pete.\nI don't think it's a problem either. I'm mainly concerned about whether there will be a second adjustment, because the adjustment still has a great impact on the score. This may cause some trouble for me to adjust the model. "
        },
        {
          "id": 439847,
          "postDate": "2018-12-16T14:09:50.683Z",
          "content": "<p>Yes, I'm agree with you. This will cause trouble to adjust the model.</p>",
          "rawMarkdown": "Yes, I'm agree with you. This will cause trouble to adjust the model."
        }
      ]
    },
    {
      "id": 433469,
      "postDate": "2018-12-05T05:30:00.543Z",
      "content": "<p>Leaderboard have been reseted, but leak is applied to public LB.<br>\n* no leak submission: public LB. 0.496<br>\n* no leak submission + leak: public LB. 0.502<br>\nThe effective of leak is small, but public LB is improved.</p>",
      "rawMarkdown": "Leaderboard have been reseted, but leak is applied to public LB.<br>\n* no leak submission: public LB. 0.496<br>\n* no leak submission + leak: public LB. 0.502<br>\nThe effective of leak is small, but public LB is improved.",
      "replies": [
        {
          "id": 433472,
          "postDate": "2018-12-05T05:39:05.337Z",
          "content": "<p>From my perspective they didn't fix anything much.</p>\n\n<ul>\n<li><p>Before reset + leak: 0.556</p></li>\n<li><p>After reset + leak: 0.555</p></li>\n</ul>",
          "rawMarkdown": "From my perspective they didn't fix anything much.\n\n- Before reset + leak: 0.556\n\n- After reset + leak: 0.555",
          "votes": 1
        },
        {
          "id": 433477,
          "postDate": "2018-12-05T05:47:22.053Z",
          "content": "<p>Thanks Khoi. I think so too.<br>\nIt is strange that the scores of many participants have not changed.</p>",
          "rawMarkdown": "Thanks Khoi. I think so too.<br>\nIt is strange that the scores of many participants have not changed.",
          "votes": 1
        }
      ]
    },
    {
      "id": 432876,
      "postDate": "2018-12-04T12:28:23.320Z",
      "content": "<p>I want to point out that some classes are very rare in train and test set, so removing them from test set will reduce them even more and score will be affected a lot by correct/incorrect predictions for that classes. IMHO it will cause a big shakeup later.</p>",
      "rawMarkdown": "I want to point out that some classes are very rare in train and test set, so removing them from test set will reduce them even more and score will be affected a lot by correct/incorrect predictions for that classes. IMHO it will cause a big shakeup later.",
      "replies": [
        {
          "id": 432892,
          "postDate": "2018-12-04T12:43:51.427Z",
          "content": "<p>Well this is only the case if the HPA data contains the answers to some of the rare classes - that's the only case where we need to remove from test. The flip side of this is that the leak is even more serious as the impact is large.</p>\n\n<p>To be honest, I'm not sure why we are worrying about LB stability when it's just been revealed part of the answers are externally available - this compromises the integrity of the competition which has a metric defined precisely to try to encourage rare class predictions.</p>",
          "rawMarkdown": "Well this is only the case if the HPA data contains the answers to some of the rare classes - that's the only case where we need to remove from test. The flip side of this is that the leak is even more serious as the impact is large.\n\nTo be honest, I'm not sure why we are worrying about LB stability when it's just been revealed part of the answers are externally available - this compromises the integrity of the competition which has a metric defined precisely to try to encourage rare class predictions."
        },
        {
          "id": 432908,
          "postDate": "2018-12-04T13:07:35.870Z",
          "content": "<p>\"Leaked\" data have been published already here <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534</a> (you can download images like described here <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a> ) and available to everyone. So it should be +- ok if everyone use same labels for that samples, as it doesn't change overall rating. So yes, I'm worrying about stability more.</p>",
          "rawMarkdown": "\"Leaked\" data have been published already here https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534 (you can download images like described here https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984 ) and available to everyone. So it should be +- ok if everyone use same labels for that samples, as it doesn't change overall rating. So yes, I'm worrying about stability more.",
          "votes": 1
        },
        {
          "id": 432915,
          "postDate": "2018-12-04T13:12:59.543Z",
          "content": "<p>The leaked test examples have a relative majority of class 16 (Cytokinetic bridge), which is a pretty rare class overall, so it is a bit worrying. <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#423736\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#423736</a></p>",
          "rawMarkdown": "The leaked test examples have a relative majority of class 16 (Cytokinetic bridge), which is a pretty rare class overall, so it is a bit worrying. https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#423736"
        },
        {
          "id": 432919,
          "postDate": "2018-12-04T13:20:49.757Z",
          "content": "<p>Part of the outcome of this competition is meant to be a Nature Methods paper - hardly worth much if the metric is biased as some people found the answers online.</p>\n<p>This is exacerbated as the extent isn't confirmed officially (it's being left to drag out) though I agree if the organisers can confirm the extent and either give a csv with those answers in or exclude them I think that is fine.</p>\n<p><strong>Just to be clear:</strong> in my view asking people to parse and download massive amounts of extra data just to plug a leak is a weak fix. You can worry all you like about stability but if the integrity is compromised it undermines the whole competition.</p>\n<p>Mark</p>\n<p>P. S. Alexander Kiselev: if you truly want to help out why not provide a csv with all the images from the test set and their labels from HPA if you have them? We can all get back on with the problem at hand then.</p>",
          "rawMarkdown": "Part of the outcome of this competition is meant to be a Nature Methods paper - hardly worth much if the metric is biased as some people found the answers online.\n\nThis is exacerbated as the extent isn't confirmed officially (it's being left to drag out) though I agree if the organisers can confirm the extent and either give a csv with those answers in or exclude them I think that is fine.\n\n**Just to be clear:** in my view asking people to parse and download massive amounts of extra data just to plug a leak is a weak fix. You can worry all you like about stability but if the integrity is compromised it undermines the whole competition.\n\nMark\n\nP. S. Alexander Kiselev: if you truly want to help out why not provide a csv with all the images from the test set and their labels from HPA if you have them? We can all get back on with the problem at hand then."
        },
        {
          "id": 432968,
          "postDate": "2018-12-04T14:24:38.563Z",
          "content": "<p>That nice csv file have been already sincerely provided by Tomomi Moriyama (TestEtraMatchingUnder_259_R14_G12_B10.csv from <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534</a> ) all you need is just join it with another csv... Ok let me do it.</p>",
          "rawMarkdown": "That nice csv file have been already sincerely provided by Tomomi Moriyama (TestEtraMatchingUnder_259_R14_G12_B10.csv from https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534 ) all you need is just join it with another csv... Ok let me do it."
        },
        {
          "id": 433039,
          "postDate": "2018-12-04T15:42:35.720Z",
          "content": "<p>Hi Alexander - thanks for bringing this up!  The host and I are aware of the presence / impact of rare classes here.  I'll be doing what I can to mitigate the impact - there will be some shifts made outside of the affected samples.</p>",
          "rawMarkdown": "Hi Alexander - thanks for bringing this up!  The host and I are aware of the presence / impact of rare classes here.  I'll be doing what I can to mitigate the impact - there will be some shifts made outside of the affected samples."
        },
        {
          "id": 433601,
          "postDate": "2018-12-05T08:06:14.010Z",
          "content": "<p>I guess you have done these shifts because I have not trained yet with the new data leak and there's been a huge jump with my LB.</p>",
          "rawMarkdown": "I guess you have done these shifts because I have not trained yet with the new data leak and there's been a huge jump with my LB.",
          "votes": 1
        }
      ]
    },
    {
      "id": 432301,
      "postDate": "2018-12-03T17:44:01.497Z",
      "content": "<p>This is so annoying, another leak of the solutions to the test set.  So it looks like getting a high score is about finding the leaked test set data rather than building a good model.\nI agree with some of the suggestions for making this extra data public from someone who already has it.  The organizers could then limit external data to this extra data.  It would also need to be done soon so that we have time to train on the new data.</p>",
      "rawMarkdown": "This is so annoying, another leak of the solutions to the test set.  So it looks like getting a high score is about finding the leaked test set data rather than building a good model.\nI agree with some of the suggestions for making this extra data public from someone who already has it.  The organizers could then limit external data to this extra data.  It would also need to be done soon so that we have time to train on the new data."
    }
  ],
  "comments": [
    {
      "id": 433361,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2018-12-05T02:09:09.363000",
      "content": "<p>It's a mistake to use competition participants' findings (the leaks) for test set update. Just ban the external data.</p>",
      "votes": 21,
      "replies": [
        {
          "id": 433971,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2018-12-05T17:54:39.610000",
          "content": "<p>Lots of people have already spent lots of time on the external data, so it would be unfair to ban it now.</p>",
          "votes": -2,
          "replies": []
        }
      ]
    },
    {
      "id": 433019,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2018-12-04T15:12:24.667000",
      "content": "<p>Hi all - thanks for bringing this to my attention. The samples in question will no longer be included in scoring. I'll be making that change (and rescoring / resetting the leaderboard) today.</p>\n\n<p>Mark, <a href=\"/ldm314\">@ldm314</a> and <a href=\"/tomomimoriyama\">@tomomimoriyama</a>, I'll be in touch with you via DM. I appreciate your efforts on this.</p>",
      "votes": 12,
      "replies": [
        {
          "id": 433021,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2018-12-04T15:16:38.683000",
          "content": "<p>Thank you!</p>\n\n<p>Could you please provide a full list of leaked test IDs as well as their classes after rescoring?</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433022,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-12-04T15:17:48.313000",
          "content": "<p>Thank you Phil!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433031,
          "author_name": "TomomiMoriyama",
          "author_url": "",
          "post_date": "2018-12-04T15:32:10.070000",
          "content": "<p>Thank you!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433062,
          "author_name": "Antoine",
          "author_url": "",
          "post_date": "2018-12-04T16:08:07.513000",
          "content": "<p>Hi <a href=\"/philculliton\">@philculliton</a>,</p>\n\n<p>First thanks for your quick reaction on this.</p>\n\n<p>Is the aim to simply remove those samples ? Or replace them with images that are not part of the public data ? </p>\n\n<p>I'm afraid that removing them will make the distribution even more imbalanced and increase the instability of the f1 score.</p>\n\n<p>For instance for class 27, we have only 11 samples in the training set for 31072 images. By removing those 259 samples you'd remove 5 samples with class 27 from 11702 samples. So it's probably a very important % of this class, if not the whole class.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433091,
          "author_name": "Robert",
          "author_url": "",
          "post_date": "2018-12-04T17:22:17.183000",
          "content": "<p>Thank you Phil.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433186,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-04T20:13:34.483000",
          "content": "<p>Thanks!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433348,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-05T01:48:00.893000",
          "content": "<p>Looks like the leaderboard has not been reset yet - is that true? I guess the comments here mean that new submissions ARE being reset?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433372,
          "author_name": "Tim H",
          "author_url": "",
          "post_date": "2018-12-05T02:20:06.457000",
          "content": "<p>I just submitted one of my earlier submissions as a test and got the same score as before, so new submissions have not been reset as yet.</p>\n\n<p>Incidentally my latest model which got 0.595 on this un-rescored LB uses the external HPA data but does not do anything special with the offending images. It will be interesting to see the difference once the leaderboard has been re-scored. I would have assumed that re-scoring will cause people that made special use of those images would see their scores drop, but based on Brian's comment below, maybe that will not always be the case!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433394,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-05T02:58:16.230000",
          "content": "<p>looks ike something's happening to the LB now. In progress.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433481,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2018-12-05T05:53:13.037000",
          "content": "<p>Thanks Phil for your prompt attention to this problem.\nAre you also going to delete the leaked test images from test.zip and test_full_size.7z?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433981,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-12-05T18:10:51.790000",
          "content": "<p>Hi <a href=\"/dslate\">@dslate</a> - </p>\n\n<p>Good question!  I won't be removing the leaked test images.  All current competitors have already downloaded them, so this would primarily negatively impact later competitors.  Removing them from scoring mitigates the effects of the leak just as readily, without asking later competitors to go without information / images that everyone else had.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433234,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-12-04T21:48:41.850000",
      "content": "<p>Here is the 126 I've been using, this time with the image IDS. </p>",
      "votes": 9,
      "replies": []
    },
    {
      "id": 431818,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-12-03T00:14:07.397000",
      "content": "<p>My proposal to even out the playing field is to take all of the v18 HPA and include it here with the training data. Remove the leaked images from the test set. I found and use ~125 matches and <a href=\"/tomomimoriyama\">@tomomimoriyama</a> found over twice that. With the test set being 11k images, removing the overlap would leave plenty to test with.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 431976,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-03T08:03:11.823000",
          "content": "<p>Yeah, or just announce they aren't part of the private LB in which case using for public LB will bias your results. This of course relies on them being able to reliably identify the true extent of the overlap/leak.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 432107,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-12-03T12:24:11.930000",
          "content": "<p>I agree with Brian proposal concerning making HPA data available.   </p>\n\n<p>For the leakage, I don't think the organizers will take any action on that, because the leakage is not severe and the vast majority of the test set is not affected. This expectation is based on the few competitions \"with leak\" that I took part in. For example, in Santander the leakage was much more evident and the organizers did not do anything about it. However, someone shared part of the leakage in kernels, but the full set of leaky rows of the test set was never released and only the top of the LB managed to exploit it 100%. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432140,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-03T13:04:45.983000",
          "content": "<p>Jumping from 0.528 to 0.588 because of the leak is severe in my view. </p>\n\n<p>I also don't think adding the HPA data to the training data really levels the playing field - it favours even more those with massive compute power.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432251,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-12-03T16:01:10.893000",
          "content": "<p>Regarding the jump in the LB: I may agree with you (what I wrote before is what I think it's the reasoning of the organizers, not necessarily mine...) .   </p>\n\n<p>For the second point, some personal considerations: <br>\n - Kaggle is biased towards competitors with large computing power and there's little or nothing we can do about it . <br>\n - Making the additional data available to everyone is the only thing that can be done to level the field again. External data has been announced (by Brian) on the public thread <strong>a month ago</strong> and no one from the Kaggle team has replied that it was prohibited. If they banned it now, it would be very unfair towards people who have built their models and spent their submissions using that data (N.B.: it's not me, I am not using HPA data so far). <br>\n - More data instances is, technically speaking, more \"democratic\" than higher resolution data. You could train a model in several steps on different Kaggle kernels, by re-loading the weights from the previous step. On the other hand, higher resolution images require gpu with larger memory, which are very expensive and clearly not available to everyone.</p>\n\n<p>Now, my opinion on what happened in this competition: <br>\n- it was a mistake of the organizers to make higher resolution images available, as it increases the \"technology gap\" <strong>a lot</strong>; <br>\n- the HPA dataset should be made available by someone who owns it - better if in agreement with the organizers, so we are sure that really <strong>all</strong> the available external data is included; <br>\n- we should not blame on people who discovered and used the leak: this is just fine according to the rules, as they announced publicly the dataset they were using</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 432269,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-03T16:33:11.217000",
          "content": "<p>Hi Chase the Trane (<a href=\"/stecasasso\">@stecasasso</a>),</p>\n\n<p>I agree with nearly all of what you said but think you are slightly conflating two points.</p>\n\n<p>I don't think it's an issue for people to use the HPA data (and don't want it banned) - we just need to make sure that the HPA data doesn't contain any of the answers to the live competition. </p>\n\n<p>Mark</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 432517,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2018-12-04T01:39:40.273000",
          "content": "<p>Hi Mark Worrall,\nI agree with you that people should be able to use the HPA data but that data shouldn't contain any leaked labels for the competition test set.  As far as I can see, the only way to accomplish this would be to find all the leaks and remove their images from the test set, both public and private.  This has the downside that every submission will have to be rescored and the LB recomputed, but with over a month to go until the contest ends this may be the best solution.  The organizers would have to decide to do this pretty soon, however.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 432810,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-12-04T10:40:37.930000",
          "content": "<p>Even if no other action is taken, we need an <strong>urgent examination of the private test set</strong> to find out just how contaminated it is. 126 images is only &gt;1% of the public test. For all we know the percentage is much higher in the unreleased data. I would really like some reassurance that contamination is really only the minor issue that it appears to be at present time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433830,
      "author_name": "Phil Culliton",
      "author_url": "",
      "post_date": "2018-12-05T14:35:58.963000",
      "content": "<p>Hey all -</p>\n\n<p>I updated and reset the LB last night.  Here's what changed:</p>\n\n<p>1) The majority of the impacted samples are no longer included in scoring.</p>\n\n<p>2) A small number of samples with very rare labels were not removed from scoring, but were shifted to the public leaderboard.</p>\n\n<p>Both of these moves will affect public LB scores, especially for less common labels.  If you're concerned about your score changing after this fix was implemented, the fix did impact everyone's scores - whether they were using the leak or not.  Any shift on the leaderboard, whether ignoring samples or moving them between public and private LBs, will affect scores. I do understand the confusion expressed about scores changing: I sincerely apologize and hope this clears things up!</p>\n\n<p>Thanks again to everyone for your feedback and help!  It's much appreciated.</p>",
      "votes": 7,
      "replies": [
        {
          "id": 433899,
          "author_name": "Tilii",
          "author_url": "",
          "post_date": "2018-12-05T16:21:31.900000",
          "content": "<p><a href=\"/philculliton\">@philculliton</a> Thanks for the info. I think this should be a separate topic and pinned to the top of discussion board, as not everyone is subscribed to the whole forum.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 433917,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-05T16:45:29.167000",
          "content": "<p>I guess that rare classes that were already on the public LB were left there? And more shifted from the private? It’s hard to see how that would lower public scores of those using the leak. It might generate a massive shakeup at the end, though. Wouldnt it have been better to just remove leaks entirely, even from rare classes? And if some classes were rendered 1/10000 or something, remove the classes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433920,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-05T16:49:30.530000",
          "content": "<p>Yeah, so using the file can cover for a model missing some of the rare classes as these are kept in public LB...but this means private LB will likely be very different.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433946,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-12-05T17:20:24.307000",
          "content": "<p>There going to be a nice shakeup in this competition :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433972,
          "author_name": "Dmytro Panchenko",
          "author_url": "",
          "post_date": "2018-12-05T17:58:50.057000",
          "content": "<p>Thank you for the update.\nCould you please confirm that there is no leakage in the private set?\nCould you also publish a list of leaked images with their classes to even the odds?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433975,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-12-05T18:03:34.657000",
          "content": "<p>Hi Dmytro -</p>\n\n<p>I can confirm that there is no leakage in the private set.  Re: posting a list of the leaked images... I'm considering it.  Because some of the images aren't simply being ignored, that makes the impact of releasing the list more complex.  I may release the list of the ignored images.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 433986,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-05T18:21:10.297000",
          "content": "<p>Here's some random chatter from someone without much experience in DL or Kaggle:</p>\n\n<p>I'll incorporate the leak to get the 0.05 public LB asap, so  can brag to my wife. But the question is now: how do I avoid the shakeup?</p>\n\n<ol>\n<li>the leak advantage will go away (democratically, I guess)</li>\n<li>the private LB will have a relative dearth of rare classes.</li>\n</ol>\n\n<p>Point 2 might not be a huge killer because we will be in the prediction stage so the model will have been trained with non-decimated rare classes. However, the leak may reduce the impetus to deal with rare classes, thus deteriorating the effectiveness of the winning models.</p>\n\n<p>So the strategy I guess will be to incorporate the leak (altering submission file only) for bragging rights in the next month (and to see exactly where you are in the competition) and develop your model just like nothing happened.</p>\n\n<p>Or you could copy all leak data (rare classes) into your personal training set to help you better train those classes. It won't hurt the public LB as you are replacing the predictions of those samples anyway.</p>\n\n<p>Regarding point 1, for completeness, I wonder how many of the leaks have been found. It appears that the organizers are just using the results contributed on these threads - not doing any checks. Perhaps there are many more leaks with very small differences (say by different resampling/antiallias, different versions of photo(?)). And it's not easy to compare a hundred thousand huge images - hashing has been used - how accurate is that in this context? Is there a better way that runs in finite time?</p>\n\n<p>Definitely a footnote in the Nature paper.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 434012,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-12-05T19:08:29.897000",
          "content": "<p>Hi pete -</p>\n\n<p>Thanks for the feedback - I appreciate the thought everyone is putting into this!</p>\n\n<p>Re: your point 2.  The rare classes in question still exist in the private LB.  Samples were shifted between the private and public LB to ensure this.  Anyone who feels a reduced impetus to deal with rare classes will definitely take a hit on the private LB.</p>\n\n<p>Re: the leaks... we absolutely did our own checks, and did <em>not</em> use the results contributed on the forums.  The forum results were useful because they indicated that people had identified a problem.  We know the full extent of the potential leakage based on our own work.</p>\n\n<p>On a related note, I should mention that hashing is <em>not</em> 100% accurate in this context, and that some of the \"leaked\" images are not, in fact, leaks (and further, some of the labels assigned them are incorrect).  The leak appears to be so effective because some of the real, actual leaks include rare classes that have a high impact on LB score.</p>\n\n<p>The people who found the leak did excellent work, but it's very difficult to get completely right.  I'd recommend not using the leak in any way.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 434100,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-05T22:10:15.150000",
          "content": "<p>Phil,</p>\n\n<p>Thanks very much for your comments. My apologies for assuming that you did not look carefully into this - looks like you definitely did. While I'm not fully on board with all of your suggestions, they are very helpful and I appreciate what you have done with this, and your policy of keeping us up to date.</p>\n\n<p>Also to those who found the leak and did great work calling attention to it and helped others understand, thank you also.  I am super impressed with how fast you understood that complicated website (I still don't really understand it). You improved the competition immeasurably.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 434463,
          "author_name": "Fionnán Alt",
          "author_url": "",
          "post_date": "2018-12-06T12:43:31.860000",
          "content": "<p>Hi Phil - thanks for the info.  Do we need to download the test set again if samples were shifted between private and public LB?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434472,
          "author_name": "bluetrain",
          "author_url": "",
          "post_date": "2018-12-06T13:08:17.340000",
          "content": "<p>@Fionnán Alt: I guess no... \ntest set = public LB + private LB</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434473,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-12-06T13:08:18.310000",
          "content": "<p>Hi Fionnán,\nGood question! You don't need to download anything new - the shift happened entirely on our end, so you shouldn't have to make any changes.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 434501,
          "author_name": "Fionnán Alt",
          "author_url": "",
          "post_date": "2018-12-06T14:08:53.157000",
          "content": "<p>Thanks Phil.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 432969,
      "author_name": "Aleksandr Kiselev",
      "author_url": "",
      "post_date": "2018-12-04T14:26:38.920000",
      "content": "<p>Test matches found by Tomomi Moriyama with labels from HPA.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 432980,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-04T14:41:25.410000",
          "content": "<p>Thank you Alexander - I hadn't been through all those files.</p>\n\n<p>P. S. that just took me from 0.482 -&gt; 0.527 ... wonder how many others in the top 100 are using it already. You might also wish to make that csv more public (i.e. kernel).</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433045,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2018-12-04T15:49:52.583000",
          "content": "<p>It's funny, using the leak our 0.537 submission using oversampling only got to 0.544, while the 0.507 submission which clearly overfitted to the major classes got to 0.556. I'm glad they decided to take action.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433190,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-04T20:20:06.293000",
          "content": "<p>Amusingly enough if I apply this file, my score goes down from .531 to .528...</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433202,
          "author_name": "Aleksandr Kiselev",
          "author_url": "",
          "post_date": "2018-12-04T20:39:31.683000",
          "content": "<p><a href=\"/ldm314\">@ldm314</a> do you think <a href=\"/tomomimoriyama\">@tomomimoriyama</a> is trying to miss lead us? :D</p>",
          "votes": -1,
          "replies": []
        },
        {
          "id": 433231,
          "author_name": "Brian",
          "author_url": "",
          "post_date": "2018-12-04T21:47:13.377000",
          "content": "<p>I don't think so. I came up with a similar list independently after <a href=\"/tomomimoriyama\">@tomomimoriyama</a>  demonstrated that the yellow channels were available. I ended up with the same score drop vs the 126 list I came up with from only RGB.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433461,
          "author_name": "Valentina Biryukova",
          "author_url": "",
          "post_date": "2018-12-05T05:18:30.957000",
          "content": "<p>It is very upset that I don't use leak but all my scores were dropped :(</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433563,
      "author_name": "Mark Worrall",
      "author_url": "",
      "post_date": "2018-12-05T07:24:09.127000",
      "content": "<p>I don't think the LB has been updated properly (or something weird is going on) as my score has gone from:</p>\n\n<p>0.482 no leaked data\n0.527 using leaked data\n0.528 post LB update?</p>",
      "votes": 6,
      "replies": [
        {
          "id": 433575,
          "author_name": "Valentina Biryukova",
          "author_url": "",
          "post_date": "2018-12-05T07:39:32.640000",
          "content": "<p>Yes! I suppose the same! My score without leak has dropped!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433576,
          "author_name": "SpruceMoose",
          "author_url": "",
          "post_date": "2018-12-05T07:40:53.367000",
          "content": "<p>Definitely something weird going on</p>\n\n<p>0.525 no leak\n0.561 leak \n0.561 after update</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 433600,
          "author_name": "Aleksandr Kiselev",
          "author_url": "",
          "post_date": "2018-12-05T08:05:46.760000",
          "content": "<p>Drop the samples, they said\nIt would be fair, they said</p>\n\n<p>P.S. no offense :)</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 433603,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-05T08:08:26.280000",
          "content": "<p>Haha - no offence taken. I don't think it's clear what's gone on at the moment but the ids from the csv you provided are still being scored in the public LB somehow.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433610,
          "author_name": "Vladimir Sorokin",
          "author_url": "",
          "post_date": "2018-12-05T08:13:51.300000",
          "content": "<p>Same! Our score without leak has decreased from 0.506 to 0.491 .</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433648,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-05T09:43:38.643000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433653,
          "author_name": "David J. Slate",
          "author_url": "",
          "post_date": "2018-12-05T09:49:56.027000",
          "content": "<p>I did not use the leaked labels, and my highest LB score has dropped from 0.504 to 0.499.  But that doesn't necessarily indicate any scoring error by the organizers.  It could be that even without the leaks, my model did a good job of predicting the affected images, so without them my score dropped a bit.\nThe LB still looks dubious at the high end, however, since those scores seem to have been hardly affected by the \"reset\" at all.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433697,
          "author_name": "wchzh",
          "author_url": "",
          "post_date": "2018-12-05T11:14:43.027000",
          "content": "<p>my score also decreased from 0.501 to 0.495.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433756,
          "author_name": "Guanshuo Xu",
          "author_url": "",
          "post_date": "2018-12-05T12:50:26.927000",
          "content": "<p>0.570 -&gt; 0.553\nWas I using too many leaks? :P</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433963,
          "author_name": "Zhijian Li",
          "author_url": "",
          "post_date": "2018-12-05T17:43:55.623000",
          "content": "<p>haha, I think so.\nmy score dropped from 0.498 to 0.494,\nusing the leaks, it goes to 0.549</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 434133,
      "author_name": "Brian",
      "author_url": "",
      "post_date": "2018-12-06T00:06:11.560000",
      "content": "<p>After the leaderboard changes, in general, models submitted with the 126 samples I have posted went down in score. Models using the list from <a href=\"/tomomimoriyama\">@tomomimoriyama</a> went up in score. What meaning this has, if any, I won't speculate on as some of these posted answers are confirmed to be inaccurate.</p>\n\n<p>Thanks to <a href=\"/philculliton\">@philculliton</a> for cleaning thing up and evening out the field.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 436658,
      "author_name": "yunhai",
      "author_url": "",
      "post_date": "2018-12-10T17:56:22.613000",
      "content": "<p>since this challenge is permited to use the external datasets,  can anyone provide a complete HPA dataset to download including full HPA image and train.csv? don't think we should cost much time exploring external datasets.  </p>",
      "votes": 4,
      "replies": [
        {
          "id": 442604,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-20T07:29:07.140000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 442840,
          "author_name": "",
          "author_url": "",
          "post_date": "2018-12-20T15:39:13.173000",
          "content": "<p>Hi Jarvis, Thank you for these useful information! Are we allowed to use HPA datasets in training a model? Does it not have any implication such as data leakage? Sorry, I didn't check all HPA datasets. Thanks </p>",
          "votes": 9,
          "replies": []
        }
      ]
    },
    {
      "id": 443202,
      "author_name": "Maksim Rodin",
      "author_url": "",
      "post_date": "2018-12-21T07:55:34.103000",
      "content": "<p>Hello guys,\nCould anybody clarify, does it make sense to include the leak in the submission now or not?\nThank you!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 443225,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-21T08:41:46.163000",
          "content": "<p>The advantage provided by the leak will probably vanish in the private leaderboard but it doesnt hurt to include it if only to see where you stand. The hpa data themselves give an extra edge in addition to any leak advantage. I dont see any reason why this extra edge wont remain in the private leaderboard.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433691,
      "author_name": "Renan Rodrigues dos Santos",
      "author_url": "",
      "post_date": "2018-12-05T10:58:30.890000",
      "content": "<p><a href=\"/philculliton\">@philculliton</a> After LB rescoring, my leak only submission (no training) scored 0.085, now it is scoring 0.062 (about two rare classes with F1 close to 1.00), so definitely there are still leakage in public LB. Don't know about private LB though.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 432185,
      "author_name": "DStjhb",
      "author_url": "",
      "post_date": "2018-12-03T14:25:42.633000",
      "content": "<p>I concur.</p>\n\n<p>If they make this data official and give us an easy zipped download of the HPA in a smaller image size (512p), that would be the best approach. If HPA can afford to have people trawling their servers for tens of thousands of 2048p images from individual urls they can afford to give a direct link to a compressed archive. They must have known that the HPA would be the obvious first point of call for anybody looking for external data.</p>\n\n<p><a href=\"/tomomimoriyama\">@tomomimoriyama</a> and <a href=\"/ldm314\">@ldm314</a> (Brian) have kindly provided access to parsing scripts that allow download of the larger images and label them accordingly. Having to download at mostly 2048px and resize is still a pain. <a href=\"/artemtprv\">@artemtprv</a> provided a web-scraping kernel to access 'preview' images from the HPA website that are 800x800, but these appear to be either cropped (or at high magnification), and there are far fewer than in the XML urls.</p>\n\n<p>I am guessing the organisers pre-reduced the training data for the competition to 512px to allow easy download and training for amateur enthusiasts and typical users. Some people don't even have the connection or hdd space to acquire the larger original image set, let alone access to memory and GPUs capable of training models with them properly. Having to download a full size external dataset (even if resized later) to make it onto the top half of the leaderboard is a little extreme.</p>\n\n<p>There's still a good amount of time left to fix this so please address it. Even if you just exclude the leaked images, that leaves the option for kind community members to upload the data somewhere independent at a reduced image size with accompanied csv files.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 432072,
      "author_name": "Allan Wang",
      "author_url": "",
      "post_date": "2018-12-03T11:14:54.117000",
      "content": "<p>Cannot agree more with you. The competition needs to focus on methods and skills instead of \"cheating\" to get a high score.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 432160,
      "author_name": "Khoi Nguyen",
      "author_url": "",
      "post_date": "2018-12-03T13:46:03.287000",
      "content": "<p>Wait, what? There is a leak?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 433965,
      "author_name": "Robert",
      "author_url": "",
      "post_date": "2018-12-05T17:48:11.510000",
      "content": "<p>Hmm... based on the scores that people are posting it looks like most of the leaks are still being scored.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 433974,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-12-05T18:00:19.807000",
          "content": "<p>Hi Robert -</p>\n\n<p>We've changed the public LB exposure for two rare classes, which is having a noticeable effect on public LB scores.  An overwhelming majority of the leaked samples are no longer being scored at all.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435873,
          "author_name": "zymale",
          "author_url": "",
          "post_date": "2018-12-09T00:00:58.267000",
          "content": "<p>Is there any mistake here? Before you adjusted, my score was 0.529. After the adjustment, it was 0.513. I didn't use data leak before, but when I use data leak,now my score is 0.555. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435874,
          "author_name": "pete",
          "author_url": "",
          "post_date": "2018-12-09T00:05:24.707000",
          "content": "<p>It sounds like reasonable behaviour. I would say it's OK.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 435877,
          "author_name": "zymale",
          "author_url": "",
          "post_date": "2018-12-09T00:12:00.187000",
          "content": "<p>Hi,pete.\nI don't think it's a problem either. I'm mainly concerned about whether there will be a second adjustment, because the adjustment still has a great impact on the score. This may cause some trouble for me to adjust the model. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 439847,
          "author_name": "Wang Xinliang",
          "author_url": "",
          "post_date": "2018-12-16T14:09:50.683000",
          "content": "<p>Yes, I'm agree with you. This will cause trouble to adjust the model.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 433469,
      "author_name": "phalanx",
      "author_url": "",
      "post_date": "2018-12-05T05:30:00.543000",
      "content": "<p>Leaderboard have been reseted, but leak is applied to public LB.<br>\n* no leak submission: public LB. 0.496<br>\n* no leak submission + leak: public LB. 0.502<br>\nThe effective of leak is small, but public LB is improved.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 433472,
          "author_name": "Khoi Nguyen",
          "author_url": "",
          "post_date": "2018-12-05T05:39:05.337000",
          "content": "<p>From my perspective they didn't fix anything much.</p>\n\n<ul>\n<li><p>Before reset + leak: 0.556</p></li>\n<li><p>After reset + leak: 0.555</p></li>\n</ul>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 433477,
          "author_name": "phalanx",
          "author_url": "",
          "post_date": "2018-12-05T05:47:22.053000",
          "content": "<p>Thanks Khoi. I think so too.<br>\nIt is strange that the scores of many participants have not changed.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 432876,
      "author_name": "Aleksandr Kiselev",
      "author_url": "",
      "post_date": "2018-12-04T12:28:23.320000",
      "content": "<p>I want to point out that some classes are very rare in train and test set, so removing them from test set will reduce them even more and score will be affected a lot by correct/incorrect predictions for that classes. IMHO it will cause a big shakeup later.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 432892,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-04T12:43:51.427000",
          "content": "<p>Well this is only the case if the HPA data contains the answers to some of the rare classes - that's the only case where we need to remove from test. The flip side of this is that the leak is even more serious as the impact is large.</p>\n\n<p>To be honest, I'm not sure why we are worrying about LB stability when it's just been revealed part of the answers are externally available - this compromises the integrity of the competition which has a metric defined precisely to try to encourage rare class predictions.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432908,
          "author_name": "Aleksandr Kiselev",
          "author_url": "",
          "post_date": "2018-12-04T13:07:35.870000",
          "content": "<p>\"Leaked\" data have been published already here <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534</a> (you can download images like described here <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984</a> ) and available to everyone. So it should be +- ok if everyone use same labels for that samples, as it doesn't change overall rating. So yes, I'm worrying about stability more.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 432915,
          "author_name": "DStjhb",
          "author_url": "",
          "post_date": "2018-12-04T13:12:59.543000",
          "content": "<p>The leaked test examples have a relative majority of class 16 (Cytokinetic bridge), which is a pretty rare class overall, so it is a bit worrying. <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#423736\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/69984#423736</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432919,
          "author_name": "Mark Worrall",
          "author_url": "",
          "post_date": "2018-12-04T13:20:49.757000",
          "content": "<p>Part of the outcome of this competition is meant to be a Nature Methods paper - hardly worth much if the metric is biased as some people found the answers online.</p>\n<p>This is exacerbated as the extent isn't confirmed officially (it's being left to drag out) though I agree if the organisers can confirm the extent and either give a csv with those answers in or exclude them I think that is fine.</p>\n<p><strong>Just to be clear:</strong> in my view asking people to parse and download massive amounts of extra data just to plug a leak is a weak fix. You can worry all you like about stability but if the integrity is compromised it undermines the whole competition.</p>\n<p>Mark</p>\n<p>P. S. Alexander Kiselev: if you truly want to help out why not provide a csv with all the images from the test set and their labels from HPA if you have them? We can all get back on with the problem at hand then.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 432968,
          "author_name": "Aleksandr Kiselev",
          "author_url": "",
          "post_date": "2018-12-04T14:24:38.563000",
          "content": "<p>That nice csv file have been already sincerely provided by Tomomi Moriyama (TestEtraMatchingUnder_259_R14_G12_B10.csv from <a href=\"https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534\">https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534</a> ) all you need is just join it with another csv... Ok let me do it.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433039,
          "author_name": "Phil Culliton",
          "author_url": "",
          "post_date": "2018-12-04T15:42:35.720000",
          "content": "<p>Hi Alexander - thanks for bringing this up!  The host and I are aware of the presence / impact of rare classes here.  I'll be doing what I can to mitigate the impact - there will be some shifts made outside of the affected samples.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 433601,
          "author_name": "FlYM",
          "author_url": "",
          "post_date": "2018-12-05T08:06:14.010000",
          "content": "<p>I guess you have done these shifts because I have not trained yet with the new data leak and there's been a huge jump with my LB.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 432301,
      "author_name": "Robert",
      "author_url": "",
      "post_date": "2018-12-03T17:44:01.497000",
      "content": "<p>This is so annoying, another leak of the solutions to the test set.  So it looks like getting a high score is about finding the leaked test set data rather than building a good model.\nI agree with some of the suggestions for making this extra data public from someone who already has it.  The organizers could then limit external data to this extra data.  It would also need to be done soon so that we have time to train on the new data.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "431652": "Hi all,\n\nI realise this might not make me the most popular person on this forum but it's something I think needs to be addressed and kaggle (e.g @philculliton) should provide an answer.\n\nTucked away in a discussion [here][1] is the revelation that the (difficult to access) HPA data has been used to boost @tomomimoriyama (plus I imagine others) to the top of the LB. In all fairness to the few people discussing this topic they have acknowledged it's uncertain as to whether it's allowed or not but I for one would like this explicit and clarified. \n\n**Use of the HPA data**: I'm only going on what little is mentioned but it appears the HPA data is not being used in the sense I understand external data should to be used (e.g. word embeddings) in order to supplement a problem but instead people have basically found the same images in the HPA data as appear in the train and test sets. i.e. they've found some of the answers. I cannot imagine this is what the competition organisers intend when they challenge us to automate *'biomedical image analysis to accelerate the understanding of human cells and disease'.*\n\nFurther, from what I can gather, the HPA data is a pain to access and as pointed out by @stecasasso (Chase the Trane) this sort of external data has been banned in previous competitions. I can't force the kaggle or the organisers to ban it in this case but I can safely say that if a prerequisite for tackling image competitions with huge data is to have to parse even more hard to reach huge data to simply uncover test images it's definitely not what I had in mind when I decided to start this competition and use machine learning for good. \n\nA final point: if I've gotten any of the above mixed up I am happy to amend my post for inaccuracies or clarify misunderstanding. I'd also like to thank those using the HPA data for being honest about finding similar images in it, particularly @ldm314 (Brian) and @tomomimoriyama.\n\nCheers,\n\nMark\n\nP. S. also mentioned [here][2]\n\n\n  [1]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/70206\n  [2]: https://www.kaggle.com/c/human-protein-atlas-image-classification/discussion/72534",
    "433361": "It's a mistake to use competition participants' findings (the leaks) for test set update. Just ban the external data.",
    "433019": "Hi all - thanks for bringing this to my attention. The samples in question will no longer be included in scoring. I'll be making that change (and rescoring / resetting the leaderboard) today.\n\nMark, @ldm314 and @tomomimoriyama, I'll be in touch with you via DM. I appreciate your efforts on this.",
    "433234": "Here is the 126 I've been using, this time with the image IDS. ",
    "431818": "My proposal to even out the playing field is to take all of the v18 HPA and include it here with the training data. Remove the leaked images from the test set. I found and use ~125 matches and @tomomimoriyama found over twice that. With the test set being 11k images, removing the overlap would leave plenty to test with.",
    "433830": "Hey all -\n\nI updated and reset the LB last night.  Here's what changed:\n\n1) The majority of the impacted samples are no longer included in scoring.\n\n2) A small number of samples with very rare labels were not removed from scoring, but were shifted to the public leaderboard.\n\nBoth of these moves will affect public LB scores, especially for less common labels.  If you're concerned about your score changing after this fix was implemented, the fix did impact everyone's scores - whether they were using the leak or not.  Any shift on the leaderboard, whether ignoring samples or moving them between public and private LBs, will affect scores. I do understand the confusion expressed about scores changing: I sincerely apologize and hope this clears things up!\n\nThanks again to everyone for your feedback and help!  It's much appreciated.",
    "432969": "Test matches found by Tomomi Moriyama with labels from HPA.",
    "433563": "I don't think the LB has been updated properly (or something weird is going on) as my score has gone from:\n\n0.482 no leaked data\n0.527 using leaked data\n0.528 post LB update?",
    "434133": "After the leaderboard changes, in general, models submitted with the 126 samples I have posted went down in score. Models using the list from @tomomimoriyama went up in score. What meaning this has, if any, I won't speculate on as some of these posted answers are confirmed to be inaccurate.\n\nThanks to @philculliton for cleaning thing up and evening out the field.\n\n",
    "436658": "since this challenge is permited to use the external datasets,  can anyone provide a complete HPA dataset to download including full HPA image and train.csv? don't think we should cost much time exploring external datasets.  ",
    "443202": "Hello guys,\nCould anybody clarify, does it make sense to include the leak in the submission now or not?\nThank you!",
    "433691": "@philculliton After LB rescoring, my leak only submission (no training) scored 0.085, now it is scoring 0.062 (about two rare classes with F1 close to 1.00), so definitely there are still leakage in public LB. Don't know about private LB though.",
    "432185": "I concur.\n\nIf they make this data official and give us an easy zipped download of the HPA in a smaller image size (512p), that would be the best approach. If HPA can afford to have people trawling their servers for tens of thousands of 2048p images from individual urls they can afford to give a direct link to a compressed archive. They must have known that the HPA would be the obvious first point of call for anybody looking for external data.\n\n@tomomimoriyama and @ldm314 (Brian) have kindly provided access to parsing scripts that allow download of the larger images and label them accordingly. Having to download at mostly 2048px and resize is still a pain. @artemtprv provided a web-scraping kernel to access 'preview' images from the HPA website that are 800x800, but these appear to be either cropped (or at high magnification), and there are far fewer than in the XML urls.\n\nI am guessing the organisers pre-reduced the training data for the competition to 512px to allow easy download and training for amateur enthusiasts and typical users. Some people don't even have the connection or hdd space to acquire the larger original image set, let alone access to memory and GPUs capable of training models with them properly. Having to download a full size external dataset (even if resized later) to make it onto the top half of the leaderboard is a little extreme.\n\nThere's still a good amount of time left to fix this so please address it. Even if you just exclude the leaked images, that leaves the option for kind community members to upload the data somewhere independent at a reduced image size with accompanied csv files.",
    "432072": "Cannot agree more with you. The competition needs to focus on methods and skills instead of \"cheating\" to get a high score.",
    "432160": "Wait, what? There is a leak?",
    "433965": "Hmm... based on the scores that people are posting it looks like most of the leaks are still being scored.",
    "433469": "Leaderboard have been reseted, but leak is applied to public LB.<br>\n* no leak submission: public LB. 0.496<br>\n* no leak submission + leak: public LB. 0.502<br>\nThe effective of leak is small, but public LB is improved.",
    "432876": "I want to point out that some classes are very rare in train and test set, so removing them from test set will reduce them even more and score will be affected a lot by correct/incorrect predictions for that classes. IMHO it will cause a big shakeup later.",
    "432301": "This is so annoying, another leak of the solutions to the test set.  So it looks like getting a high score is about finding the leaked test set data rather than building a good model.\nI agree with some of the suggestions for making this extra data public from someone who already has it.  The organizers could then limit external data to this extra data.  It would also need to be done soon so that we have time to train on the new data."
  }
}