{
  "id": 238131,
  "title": "Zero private LB score and the deepflash2 submission notebook",
  "url": "/competitions/hubmap-kidney-segmentation/discussion/238131",
  "author_name": "",
  "post_date": "2021-05-11T10:18:22.279095500Z",
  "votes": 17,
  "comment_count": 6,
  "views": 0,
  "content": "<p>We realized that a big discussion is going on about zero private LB scores and the <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub\" target=\"_blank\">deepflash2 submission</a> notebook - the submission of a single unet model trained with efficient sampling. Here’s what happened:</p>\n<p><strong>0 private LB score</strong><br>\nWe were working with v14 (the 0 private LB score version) until one week ago as we realized that the runtime of the notebook (public data) and submission (public and private data) were very similar.<br>\nWe tried to figure out what was going on but couldn't locate the error (probably caused by the ‘tifffile’ library and large private test images, but no OOM error was thrown;  On the private data set, the sample submission file was saved as <code>submission.csv</code> but empty.)</p>\n<ul>\n<li>We discussed this issue publicly (see the <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments\" target=\"_blank\">comments</a>)</li>\n<li><a href=\"https://www.kaggle.com/theudas\" target=\"_blank\">@theudas</a> contacted the Kaggle staff to get information on this matter <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </li>\n</ul>\n<blockquote>\n  <p>Dear Addison,<br>\n  I hope it is fine, that I message you out of the blue.<br>\n  We are the authors of the deepflash2 package, that is used by many in the Hubmap Challenge.<br>\n  Inference Notebook: <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub\" target=\"_blank\">https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub</a><br>\n  Questions appeared, if inference on the hidden testdata works because our submission notebook is quite fast.<br>\n  Our assumption is, that during private testing, all image ids are present in the sample_submission.csv file.<br>\n  (in the csv file gets swapped during private testing)</p>\n  <p>If that is not the case we need to rewrite our notebook as soon as possible so anyone who has forked it can adjust their copy.</p>\n  <p>Sample Submission of ours:<br>\n  <a href=\"https://www.kaggle.com/matjes/pseudo-sub-hubmap-efficient-sampling-deepflash2?scriptVersionId=61744184\" target=\"_blank\">https://www.kaggle.com/matjes/pseudo-sub-hubmap-efficient-sampling-deepflash2?scriptVersionId=61744184</a></p>\n  <p>I wish you a wonderful day<br>\n  Phil</p>\n  <ul>\n  <li>We won't publish the reply without the consent of <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> , but in summary it did not help to understand the error.</li>\n  </ul>\n</blockquote>\n<p>Finally, we decided to do a complete rewrite of the notebook, based on the kernels of ( <a href=\"https://www.kaggle.com/leighplt\" target=\"_blank\">@leighplt</a> <a href=\"https://www.kaggle.com/leighplt/pytorch-fcn-resnet50\" target=\"_blank\">kernel</a> and <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> (<a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub\" target=\"_blank\">kernel</a>).<br>\nWe made the new kernel publicly available as soon as possible (4 days before the end of the challenge), still not knowing if there was actually a problem with v14 or not.</p>\n<p><strong>The high private LB score of v15</strong><br>\nAt the time we published the new inference version (v15), the notebook scored 0.922 on the public LB, which was by far not the highest-ranking public kernel. We used the model from the public <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-train\" target=\"_blank\">training notebook</a>, trained a month ago. We didn't know that this kernel would finally have a private LB score of 0.945 (bronze medal) and even 0.946  private LB (0.921 public LB ) at the standard threshold of 0.5 (34th private LB).</p>\n<p>The lower public and the higher private score is probably caused by the d488c759a image, which should have been removed from the public LB evaluation!</p>",
  "messages": [
    {
      "id": "1301937",
      "postDate": "05/11/2021 10:18:22",
      "content": "<p>We realized that a big discussion is going on about zero private LB scores and the <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub\" target=\"_blank\">deepflash2 submission</a> notebook - the submission of a single unet model trained with efficient sampling. Here’s what happened:</p>\n<p><strong>0 private LB score</strong><br>\nWe were working with v14 (the 0 private LB score version) until one week ago as we realized that the runtime of the notebook (public data) and submission (public and private data) were very similar.<br>\nWe tried to figure out what was going on but couldn't locate the error (probably caused by the ‘tifffile’ library and large private test images, but no OOM error was thrown;  On the private data set, the sample submission file was saved as <code>submission.csv</code> but empty.)</p>\n<ul>\n<li>We discussed this issue publicly (see the <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments\" target=\"_blank\">comments</a>)</li>\n<li><a href=\"https://www.kaggle.com/theudas\" target=\"_blank\">@theudas</a> contacted the Kaggle staff to get information on this matter <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> </li>\n</ul>\n<blockquote>\n  <p>Dear Addison,<br>\n  I hope it is fine, that I message you out of the blue.<br>\n  We are the authors of the deepflash2 package, that is used by many in the Hubmap Challenge.<br>\n  Inference Notebook: <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub\" target=\"_blank\">https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub</a><br>\n  Questions appeared, if inference on the hidden testdata works because our submission notebook is quite fast.<br>\n  Our assumption is, that during private testing, all image ids are present in the sample_submission.csv file.<br>\n  (in the csv file gets swapped during private testing)</p>\n  <p>If that is not the case we need to rewrite our notebook as soon as possible so anyone who has forked it can adjust their copy.</p>\n  <p>Sample Submission of ours:<br>\n  <a href=\"https://www.kaggle.com/matjes/pseudo-sub-hubmap-efficient-sampling-deepflash2?scriptVersionId=61744184\" target=\"_blank\">https://www.kaggle.com/matjes/pseudo-sub-hubmap-efficient-sampling-deepflash2?scriptVersionId=61744184</a></p>\n  <p>I wish you a wonderful day<br>\n  Phil</p>\n  <ul>\n  <li>We won't publish the reply without the consent of <a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> , but in summary it did not help to understand the error.</li>\n  </ul>\n</blockquote>\n<p>Finally, we decided to do a complete rewrite of the notebook, based on the kernels of ( <a href=\"https://www.kaggle.com/leighplt\" target=\"_blank\">@leighplt</a> <a href=\"https://www.kaggle.com/leighplt/pytorch-fcn-resnet50\" target=\"_blank\">kernel</a> and <a href=\"https://www.kaggle.com/iafoss\" target=\"_blank\">@iafoss</a> (<a href=\"https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub\" target=\"_blank\">kernel</a>).<br>\nWe made the new kernel publicly available as soon as possible (4 days before the end of the challenge), still not knowing if there was actually a problem with v14 or not.</p>\n<p><strong>The high private LB score of v15</strong><br>\nAt the time we published the new inference version (v15), the notebook scored 0.922 on the public LB, which was by far not the highest-ranking public kernel. We used the model from the public <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-train\" target=\"_blank\">training notebook</a>, trained a month ago. We didn't know that this kernel would finally have a private LB score of 0.945 (bronze medal) and even 0.946  private LB (0.921 public LB ) at the standard threshold of 0.5 (34th private LB).</p>\n<p>The lower public and the higher private score is probably caused by the d488c759a image, which should have been removed from the public LB evaluation!</p>",
      "rawMarkdown": "We realized that a big discussion is going on about zero private LB scores and the [deepflash2 submission](https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub) notebook - the submission of a single unet model trained with efficient sampling. Here’s what happened:\n\n**0 private LB score**\nWe were working with v14 (the 0 private LB score version) until one week ago as we realized that the runtime of the notebook (public data) and submission (public and private data) were very similar.\nWe tried to figure out what was going on but couldn't locate the error (probably caused by the ‘tifffile’ library and large private test images, but no OOM error was thrown;  On the private data set, the sample submission file was saved as `submission.csv` but empty.)\n- We discussed this issue publicly (see the [comments](https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments))\n- @theudas contacted the Kaggle staff to get information on this matter @addisonhoward \n>Dear Addison,\n>I hope it is fine, that I message you out of the blue.\n> We are the authors of the deepflash2 package, that is used by many in the Hubmap Challenge.\n> Inference Notebook: https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub\n>Questions appeared, if inference on the hidden testdata works because our submission notebook is quite fast.\n> Our assumption is, that during private testing, all image ids are present in the sample_submission.csv file.\n>(in the csv file gets swapped during private testing)\n\n> If that is not the case we need to rewrite our notebook as soon as possible so anyone who has forked it can adjust their copy.\n\n>Sample Submission of ours:\nhttps://www.kaggle.com/matjes/pseudo-sub-hubmap-efficient-sampling-deepflash2?scriptVersionId=61744184\n\n> I wish you a wonderful day\n> Phil\n- We won't publish the reply without the consent of @addisonhoward , but in summary it did not help to understand the error.\n\nFinally, we decided to do a complete rewrite of the notebook, based on the kernels of ( @leighplt [kernel](https://www.kaggle.com/leighplt/pytorch-fcn-resnet50) and @iafoss ([kernel](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub)).\nWe made the new kernel publicly available as soon as possible (4 days before the end of the challenge), still not knowing if there was actually a problem with v14 or not.\n\n**The high private LB score of v15**\nAt the time we published the new inference version (v15), the notebook scored 0.922 on the public LB, which was by far not the highest-ranking public kernel. We used the model from the public [training notebook](https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-train), trained a month ago. We didn't know that this kernel would finally have a private LB score of 0.945 (bronze medal) and even 0.946  private LB (0.921 public LB ) at the standard threshold of 0.5 (34th private LB).\n\nThe lower public and the higher private score is probably caused by the d488c759a image, which should have been removed from the public LB evaluation!",
      "votes": null
    },
    {
      "id": "1302191",
      "postDate": "05/11/2021 12:36:14",
      "content": "<p>Hi, thanks for posting this. All test image IDs were present in the private / hidden test set's <code>sample_submission.csv</code> file. The CSV does get swapped during private testing.</p>",
      "rawMarkdown": "Hi, thanks for posting this. All test image IDs were present in the private / hidden test set's `sample_submission.csv` file. The CSV does get swapped during private testing.",
      "votes": null
    },
    {
      "id": "1302198",
      "postDate": "05/11/2021 12:40:25",
      "content": "<p>Thanks. Just to clarify. Both sample_submission csvs?</p>",
      "rawMarkdown": "Thanks. Just to clarify. Both sample_submission csvs?",
      "votes": null
    },
    {
      "id": "1302224",
      "postDate": "05/11/2021 12:51:11",
      "content": "<p>That's correct. They both get swapped to the version with the full set of test image IDs.</p>",
      "rawMarkdown": "That's correct. They both get swapped to the version with the full set of test image IDs.",
      "votes": null
    },
    {
      "id": "1302233",
      "postDate": "05/11/2021 12:56:01",
      "content": "<p>Thanks, Phil!<br>\nIs there an understanding of what could have happened so that many of us have a perfectly valid public LB score and zero private score for the full inference code (not a dummy csv submission)?</p>",
      "rawMarkdown": "Thanks, Phil!\nIs there an understanding of what could have happened so that many of us have a perfectly valid public LB score and zero private score for the full inference code (not a dummy csv submission)?",
      "votes": null
    },
    {
      "id": "1302602",
      "postDate": "05/11/2021 16:13:28",
      "content": "<p>You're welcome! We're looking into it currently. We'd like to understand what happened as well.</p>",
      "rawMarkdown": "You're welcome! We're looking into it currently. We'd like to understand what happened as well.",
      "votes": null
    },
    {
      "id": "1302915",
      "postDate": "05/11/2021 19:26:43",
      "content": "<p><a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> I believe the reason is posted <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments#1301770\" target=\"_blank\">here</a>. There is a for loop which infers each test image one at a time. The public test images are first in the rows of the sample submission file, so the 5 public test images get inferred corrected and added to the list variable. Next the first private test image throws an error (probably file image too large using tiff library). This stops the execution of that code cell.</p>\n<p>The notebook continues to finish and converts the list (which only contains predictions for public test) into a submission file.</p>",
      "rawMarkdown": "sakvaua I believe the reason is posted [here][1]. There is a for loop which infers each test image one at a time. The public test images are first in the rows of the sample submission file, so the 5 public test images get inferred corrected and added to the list variable. Next the first private test image throws an error (probably file image too large using tiff library). This stops the execution of that code cell.\n\nThe notebook continues to finish and converts the list (which only contains predictions for public test) into a submission file.\n\n[1]: https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments#1301770",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1302191,
      "author_name": "philculliton",
      "author_url": "",
      "post_date": "05/11/2021 12:36:14",
      "content": "<p>Hi, thanks for posting this. All test image IDs were present in the private / hidden test set's <code>sample_submission.csv</code> file. The CSV does get swapped during private testing.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1302198,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "05/11/2021 12:40:25",
          "content": "<p>Thanks. Just to clarify. Both sample_submission csvs?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1302224,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "05/11/2021 12:51:11",
          "content": "<p>That's correct. They both get swapped to the version with the full set of test image IDs.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1302233,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "05/11/2021 12:56:01",
          "content": "<p>Thanks, Phil!<br>\nIs there an understanding of what could have happened so that many of us have a perfectly valid public LB score and zero private score for the full inference code (not a dummy csv submission)?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1302602,
          "author_name": "philculliton",
          "author_url": "",
          "post_date": "05/11/2021 16:13:28",
          "content": "<p>You're welcome! We're looking into it currently. We'd like to understand what happened as well.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1302915,
          "author_name": "cdeotte",
          "author_url": "",
          "post_date": "05/11/2021 19:26:43",
          "content": "<p><a href=\"https://www.kaggle.com/sakvaua\" target=\"_blank\">@sakvaua</a> I believe the reason is posted <a href=\"https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments#1301770\" target=\"_blank\">here</a>. There is a for loop which infers each test image one at a time. The public test images are first in the rows of the sample submission file, so the 5 public test images get inferred corrected and added to the list variable. Next the first private test image throws an error (probably file image too large using tiff library). This stops the execution of that code cell.</p>\n<p>The notebook continues to finish and converts the list (which only contains predictions for public test) into a submission file.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1301937": "We realized that a big discussion is going on about zero private LB scores and the [deepflash2 submission](https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub) notebook - the submission of a single unet model trained with efficient sampling. Here’s what happened:\n\n**0 private LB score**\nWe were working with v14 (the 0 private LB score version) until one week ago as we realized that the runtime of the notebook (public data) and submission (public and private data) were very similar.\nWe tried to figure out what was going on but couldn't locate the error (probably caused by the ‘tifffile’ library and large private test images, but no OOM error was thrown;  On the private data set, the sample submission file was saved as `submission.csv` but empty.)\n- We discussed this issue publicly (see the [comments](https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments))\n- @theudas contacted the Kaggle staff to get information on this matter @addisonhoward \n>Dear Addison,\n>I hope it is fine, that I message you out of the blue.\n> We are the authors of the deepflash2 package, that is used by many in the Hubmap Challenge.\n> Inference Notebook: https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub\n>Questions appeared, if inference on the hidden testdata works because our submission notebook is quite fast.\n> Our assumption is, that during private testing, all image ids are present in the sample_submission.csv file.\n>(in the csv file gets swapped during private testing)\n\n> If that is not the case we need to rewrite our notebook as soon as possible so anyone who has forked it can adjust their copy.\n\n>Sample Submission of ours:\nhttps://www.kaggle.com/matjes/pseudo-sub-hubmap-efficient-sampling-deepflash2?scriptVersionId=61744184\n\n> I wish you a wonderful day\n> Phil\n- We won't publish the reply without the consent of @addisonhoward , but in summary it did not help to understand the error.\n\nFinally, we decided to do a complete rewrite of the notebook, based on the kernels of ( @leighplt [kernel](https://www.kaggle.com/leighplt/pytorch-fcn-resnet50) and @iafoss ([kernel](https://www.kaggle.com/iafoss/hubmap-pytorch-fast-ai-starter-sub)).\nWe made the new kernel publicly available as soon as possible (4 days before the end of the challenge), still not knowing if there was actually a problem with v14 or not.\n\n**The high private LB score of v15**\nAt the time we published the new inference version (v15), the notebook scored 0.922 on the public LB, which was by far not the highest-ranking public kernel. We used the model from the public [training notebook](https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-train), trained a month ago. We didn't know that this kernel would finally have a private LB score of 0.945 (bronze medal) and even 0.946  private LB (0.921 public LB ) at the standard threshold of 0.5 (34th private LB).\n\nThe lower public and the higher private score is probably caused by the d488c759a image, which should have been removed from the public LB evaluation!",
    "1302191": "Hi, thanks for posting this. All test image IDs were present in the private / hidden test set's `sample_submission.csv` file. The CSV does get swapped during private testing.",
    "1302198": "Thanks. Just to clarify. Both sample_submission csvs?",
    "1302224": "That's correct. They both get swapped to the version with the full set of test image IDs.",
    "1302233": "Thanks, Phil!\nIs there an understanding of what could have happened so that many of us have a perfectly valid public LB score and zero private score for the full inference code (not a dummy csv submission)?",
    "1302602": "You're welcome! We're looking into it currently. We'd like to understand what happened as well.",
    "1302915": "sakvaua I believe the reason is posted [here][1]. There is a for loop which infers each test image one at a time. The public test images are first in the rows of the sample submission file, so the 5 public test images get inferred corrected and added to the list variable. Next the first private test image throws an error (probably file image too large using tiff library). This stops the execution of that code cell.\n\nThe notebook continues to finish and converts the list (which only contains predictions for public test) into a submission file.\n\n[1]: https://www.kaggle.com/matjes/hubmap-efficient-sampling-deepflash2-sub/comments#1301770"
  },
  "source": "meta"
}