{
  "id": 336766,
  "title": "Question for the hosts about rationale for the public (not private) LB",
  "url": "/competitions/hubmap-organ-segmentation/discussion/336766",
  "author_name": "",
  "post_date": "2022-07-12T22:23:12.954917900Z",
  "votes": 11,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Dear hosts,</p>\n<p>I was initially nonplussed about the fact that the final dataset uses a different data source/stain/slide thickness/scale from the training dataset. However, on reflection, I actually think it's quite interesting and cool - it gets people to develop more generalisable methods.</p>\n<p>However, what I cannot understand is why you decided to include HPA data in the public leaderboard. It means 'overfitting' (which isn't really a fair term here, but still) to HPA will improve your public LB score, but harm your private one.</p>\n<p>Usually when people are overfitting to the public LB, the GMs come out and shake their heads and say \"tut tut, you should have trusted in your CV\" - but we can't, because the whole point of the competition is we don't have access to any HuBMAP data - we HAVE to use the public LB. </p>\n<p>The best strategy, as far as I can tell, is to predict 0 for every single HPA example during submission, using the meta-data, and then only predict properly on HuBMAP data. Only this way can you get a reasonable idea of how well yours models are doing on HuBMAP alone.</p>\n<p>This therefore means that understanding of LB probing/gaming is paramount for this competition, which seems like a huge misjudgement, I feel.</p>\n<p>Hosts/organisers, have I missed something? If not, might you consider removing the HPA data from the public LB? If you don't, many people will just write hacks around it during submission anyway, I suspect…</p>",
  "messages": [
    {
      "id": "1853443",
      "postDate": "07/12/2022 22:23:12",
      "content": "<p>Dear hosts,</p>\n<p>I was initially nonplussed about the fact that the final dataset uses a different data source/stain/slide thickness/scale from the training dataset. However, on reflection, I actually think it's quite interesting and cool - it gets people to develop more generalisable methods.</p>\n<p>However, what I cannot understand is why you decided to include HPA data in the public leaderboard. It means 'overfitting' (which isn't really a fair term here, but still) to HPA will improve your public LB score, but harm your private one.</p>\n<p>Usually when people are overfitting to the public LB, the GMs come out and shake their heads and say \"tut tut, you should have trusted in your CV\" - but we can't, because the whole point of the competition is we don't have access to any HuBMAP data - we HAVE to use the public LB. </p>\n<p>The best strategy, as far as I can tell, is to predict 0 for every single HPA example during submission, using the meta-data, and then only predict properly on HuBMAP data. Only this way can you get a reasonable idea of how well yours models are doing on HuBMAP alone.</p>\n<p>This therefore means that understanding of LB probing/gaming is paramount for this competition, which seems like a huge misjudgement, I feel.</p>\n<p>Hosts/organisers, have I missed something? If not, might you consider removing the HPA data from the public LB? If you don't, many people will just write hacks around it during submission anyway, I suspect…</p>",
      "rawMarkdown": "Dear hosts,\n\nI was initially nonplussed about the fact that the final dataset uses a different data source/stain/slide thickness/scale from the training dataset. However, on reflection, I actually think it's quite interesting and cool - it gets people to develop more generalisable methods.\n\nHowever, what I cannot understand is why you decided to include HPA data in the public leaderboard. It means 'overfitting' (which isn't really a fair term here, but still) to HPA will improve your public LB score, but harm your private one.\n\nUsually when people are overfitting to the public LB, the GMs come out and shake their heads and say \"tut tut, you should have trusted in your CV\" - but we can't, because the whole point of the competition is we don't have access to any HuBMAP data - we HAVE to use the public LB. \n\nThe best strategy, as far as I can tell, is to predict 0 for every single HPA example during submission, using the meta-data, and then only predict properly on HuBMAP data. Only this way can you get a reasonable idea of how well yours models are doing on HuBMAP alone.\n\nThis therefore means that understanding of LB probing/gaming is paramount for this competition, which seems like a huge misjudgement, I feel.\n\nHosts/organisers, have I missed something? If not, might you consider removing the HPA data from the public LB? If you don't, many people will just write hacks around it during submission anyway, I suspect...",
      "votes": null
    },
    {
      "id": "1853809",
      "postDate": "07/13/2022 06:53:49",
      "content": "<p>As you said, It will lead to leaderboard probing anyway regardless of including hpa in public test set. I don't think hpa should be removed though. Everyone should do their own research and find out why cv/lb discrepancy occurs.</p>\n<p>I think private test set should include both hpa and hubmap because organizers said that they want generalization, right? At this point, our model's doesn't necessarily have to be generalizable. They only need to overfit to hubmap distribution.</p>",
      "rawMarkdown": "As you said, It will lead to leaderboard probing anyway regardless of including hpa in public test set. I don't think hpa should be removed though. Everyone should do their own research and find out why cv/lb discrepancy occurs.\n\nI think private test set should include both hpa and hubmap because organizers said that they want generalization, right? At this point, our model's doesn't necessarily have to be generalizable. They only need to overfit to hubmap distribution.",
      "votes": null
    },
    {
      "id": "1853813",
      "postDate": "07/13/2022 07:01:26",
      "content": "<p>The optimal solution will be the one where CV will follow the same distribution as public and private board. which obviously means that you need external data to even make a good CV strategy. <br>\nI have tried with one publicly available but it didn't seem to work. <br>\nThe data on Hubmap's website is not publicly available and requires permission to use. So someone using that might be disqualified. <br>\nI think best thing we can wish for is competition host provide some external data of hubmap just to let us know some basic idea of how should we design our models.</p>",
      "rawMarkdown": "The optimal solution will be the one where CV will follow the same distribution as public and private board. which obviously means that you need external data to even make a good CV strategy. \nI have tried with one publicly available but it didn't seem to work. \nThe data on Hubmap's website is not publicly available and requires permission to use. So someone using that might be disqualified. \nI think best thing we can wish for is competition host provide some external data of hubmap just to let us know some basic idea of how should we design our models.",
      "votes": null
    },
    {
      "id": "1854153",
      "postDate": "07/13/2022 13:14:10",
      "content": "<p>Yeah, that's the other option - include HPA in the private test.</p>\n<p>What I cannot understand is why you'd make the private and public LBs differ - I cannot think of a reason to justify it.</p>",
      "rawMarkdown": "Yeah, that's the other option - include HPA in the private test.\n\nWhat I cannot understand is why you'd make the private and public LBs differ - I cannot think of a reason to justify it.",
      "votes": null
    },
    {
      "id": "1854170",
      "postDate": "07/13/2022 13:25:19",
      "content": "<p>I guess it gives some kind of a \"twist\" to competition :D</p>",
      "rawMarkdown": "I guess it gives some kind of a \"twist\" to competition :D",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1853809,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "07/13/2022 06:53:49",
      "content": "<p>As you said, It will lead to leaderboard probing anyway regardless of including hpa in public test set. I don't think hpa should be removed though. Everyone should do their own research and find out why cv/lb discrepancy occurs.</p>\n<p>I think private test set should include both hpa and hubmap because organizers said that they want generalization, right? At this point, our model's doesn't necessarily have to be generalizable. They only need to overfit to hubmap distribution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1854153,
          "author_name": "jamesphoward",
          "author_url": "",
          "post_date": "07/13/2022 13:14:10",
          "content": "<p>Yeah, that's the other option - include HPA in the private test.</p>\n<p>What I cannot understand is why you'd make the private and public LBs differ - I cannot think of a reason to justify it.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1854170,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "07/13/2022 13:25:19",
          "content": "<p>I guess it gives some kind of a \"twist\" to competition :D</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1853813,
      "author_name": "muhammad4hmed",
      "author_url": "",
      "post_date": "07/13/2022 07:01:26",
      "content": "<p>The optimal solution will be the one where CV will follow the same distribution as public and private board. which obviously means that you need external data to even make a good CV strategy. <br>\nI have tried with one publicly available but it didn't seem to work. <br>\nThe data on Hubmap's website is not publicly available and requires permission to use. So someone using that might be disqualified. <br>\nI think best thing we can wish for is competition host provide some external data of hubmap just to let us know some basic idea of how should we design our models.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1853443": "Dear hosts,\n\nI was initially nonplussed about the fact that the final dataset uses a different data source/stain/slide thickness/scale from the training dataset. However, on reflection, I actually think it's quite interesting and cool - it gets people to develop more generalisable methods.\n\nHowever, what I cannot understand is why you decided to include HPA data in the public leaderboard. It means 'overfitting' (which isn't really a fair term here, but still) to HPA will improve your public LB score, but harm your private one.\n\nUsually when people are overfitting to the public LB, the GMs come out and shake their heads and say \"tut tut, you should have trusted in your CV\" - but we can't, because the whole point of the competition is we don't have access to any HuBMAP data - we HAVE to use the public LB. \n\nThe best strategy, as far as I can tell, is to predict 0 for every single HPA example during submission, using the meta-data, and then only predict properly on HuBMAP data. Only this way can you get a reasonable idea of how well yours models are doing on HuBMAP alone.\n\nThis therefore means that understanding of LB probing/gaming is paramount for this competition, which seems like a huge misjudgement, I feel.\n\nHosts/organisers, have I missed something? If not, might you consider removing the HPA data from the public LB? If you don't, many people will just write hacks around it during submission anyway, I suspect...",
    "1853809": "As you said, It will lead to leaderboard probing anyway regardless of including hpa in public test set. I don't think hpa should be removed though. Everyone should do their own research and find out why cv/lb discrepancy occurs.\n\nI think private test set should include both hpa and hubmap because organizers said that they want generalization, right? At this point, our model's doesn't necessarily have to be generalizable. They only need to overfit to hubmap distribution.",
    "1853813": "The optimal solution will be the one where CV will follow the same distribution as public and private board. which obviously means that you need external data to even make a good CV strategy. \nI have tried with one publicly available but it didn't seem to work. \nThe data on Hubmap's website is not publicly available and requires permission to use. So someone using that might be disqualified. \nI think best thing we can wish for is competition host provide some external data of hubmap just to let us know some basic idea of how should we design our models.",
    "1854153": "Yeah, that's the other option - include HPA in the private test.\n\nWhat I cannot understand is why you'd make the private and public LBs differ - I cannot think of a reason to justify it.",
    "1854170": "I guess it gives some kind of a \"twist\" to competition :D"
  },
  "source": "meta"
}