{
  "id": 466551,
  "title": "Trouble with Submission Scoring Errors: Testing the Scoring System",
  "url": "/competitions/blood-vessel-segmentation/discussion/466551",
  "author_name": "",
  "post_date": "2024-01-09T05:31:33.740692700Z",
  "votes": 5,
  "comment_count": 14,
  "views": 0,
  "content": "<p>I'm having trouble overcoming the <em>Submission Scoring Errors</em>, and a quick look at the submission forum suggests I'm not alone.</p>\n<p>For this reason, I made a skeleton of our notebook <a href=\"https://www.kaggle.com/cyberian516/testing-the-scoring-system\" target=\"_blank\">which can be found here</a>. This notebook creates nearly empty labels for each image in the dataset. Then it scores itself using the <a href=\"https://www.kaggle.com/code/metric/surface-dice-metric/notebook\" target=\"_blank\">source code of the Surface Dice Metric</a> suggested in the <a href=\"www.kaggle.com/competitions/blood-vessel-segmentation/overview/evaluation\" target=\"_blank\">competition's main page</a> without any issues (except for a low score), but it also has a submission scoring error.</p>\n<p>The following issues seem to have occurred in other notebooks:</p>\n<ul>\n<li>Incorrect RLE implementation [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463224#2572213\" target=\"_blank\">1</a>]</li>\n<li>Forgetting to binarize predicted labels [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464104\" target=\"_blank\">1</a>]</li>\n<li>Out Of Memory errors from too noisy labels [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464287#2580964\" target=\"_blank\">1</a>]</li>\n<li>Incorrect data loading [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/465824#2588855\" target=\"_blank\">1</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454732#2521239\" target=\"_blank\">2</a>]</li>\n</ul>\n<p>It has also been suggested that empty masks can also cause submission scoring errors. [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008\" target=\"_blank\">1</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454799#2523059\" target=\"_blank\">2</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455062\" target=\"_blank\">3</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001\" target=\"_blank\">4</a>] </p>\n<p>However, I don't seem to run into any of these issues when running the scoring script <em>in</em> our notebook.</p>\n<p>Using the suggested scoring notebook, The skeleton notebook  <strong>doesn't have sufficient memory to compute the score for a set of labels the size of <code>kidney_1_dense</code></strong></p>\n<p>As a first-time Kaggler, this issue has left me stumped. If anyone else is running into the same vague submission scoring errors, this discussion page might be another good place to diagnose them together.</p>\n<h2>Personal notes:</h2>\n<p><strong>23/01/09</strong></p>\n<ul>\n<li><p>The 3D surface dice method is used, starting in version 7 of the notebook. Thanks, <a href=\"https://www.kaggle.com/E\" target=\"_blank\">@E</a>/S Pronk!</p></li>\n<li><p>Some experiments were run in the notebook:<br>\n<em>Experiment 1</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code>, <code>kidney_1_voi</code>, and <code>kidney_2</code></li>\n<li>Submission:  a set of empty labels in the shape of <code>kidney_1_dense</code>, <code>kidney_1_voi</code>, and <code>kidney_2</code></li>\n<li><code>resize_fraction</code>: 1.0</li>\n<li>Score: out of memory error</li></ul>\n<p><em>Experiment 2</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  a set of empty labels in the shape of <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 1.0</li>\n<li>Score: out of memory error</li></ul>\n<p><em>Experiment 3</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 1.0</li>\n<li>Score: out of memory error</li></ul>\n<p><em>Experiment 4</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 0.2</li>\n<li>Score: 1.0 (max RAM about 4.1GB)</li></ul>\n<p><em>Experiment 5</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 0.5</li>\n<li>Score: 1.0 (max RAM about 9.6GB)</li></ul>\n<p><em>Experiment 5</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 0.75</li>\n<li>Score: 1.0 (max RAM about 21.7GB)</li></ul></li>\n</ul>\n<p><strong>23/01/08</strong></p>\n<ul>\n<li>Submission scoring error happens between 2-3 minutes after the notebook is submitted. </li>\n<li>When I run the notebook, it is finished processing kidney_1 and loading kidney_1_voi between minutes 2 and 3. </li>\n<li>This may also be about the amount of time it takes to load a large dataset. </li>\n<li><code>submission.csv</code> has not been created yet.</li>\n</ul>",
  "messages": [
    {
      "id": "2593189",
      "postDate": "01/09/2024 05:31:33",
      "content": "<p>I'm having trouble overcoming the <em>Submission Scoring Errors</em>, and a quick look at the submission forum suggests I'm not alone.</p>\n<p>For this reason, I made a skeleton of our notebook <a href=\"https://www.kaggle.com/cyberian516/testing-the-scoring-system\" target=\"_blank\">which can be found here</a>. This notebook creates nearly empty labels for each image in the dataset. Then it scores itself using the <a href=\"https://www.kaggle.com/code/metric/surface-dice-metric/notebook\" target=\"_blank\">source code of the Surface Dice Metric</a> suggested in the <a href=\"www.kaggle.com/competitions/blood-vessel-segmentation/overview/evaluation\" target=\"_blank\">competition's main page</a> without any issues (except for a low score), but it also has a submission scoring error.</p>\n<p>The following issues seem to have occurred in other notebooks:</p>\n<ul>\n<li>Incorrect RLE implementation [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463224#2572213\" target=\"_blank\">1</a>]</li>\n<li>Forgetting to binarize predicted labels [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464104\" target=\"_blank\">1</a>]</li>\n<li>Out Of Memory errors from too noisy labels [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464287#2580964\" target=\"_blank\">1</a>]</li>\n<li>Incorrect data loading [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/465824#2588855\" target=\"_blank\">1</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454732#2521239\" target=\"_blank\">2</a>]</li>\n</ul>\n<p>It has also been suggested that empty masks can also cause submission scoring errors. [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008\" target=\"_blank\">1</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454799#2523059\" target=\"_blank\">2</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455062\" target=\"_blank\">3</a>] [<a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001\" target=\"_blank\">4</a>] </p>\n<p>However, I don't seem to run into any of these issues when running the scoring script <em>in</em> our notebook.</p>\n<p>Using the suggested scoring notebook, The skeleton notebook  <strong>doesn't have sufficient memory to compute the score for a set of labels the size of <code>kidney_1_dense</code></strong></p>\n<p>As a first-time Kaggler, this issue has left me stumped. If anyone else is running into the same vague submission scoring errors, this discussion page might be another good place to diagnose them together.</p>\n<h2>Personal notes:</h2>\n<p><strong>23/01/09</strong></p>\n<ul>\n<li><p>The 3D surface dice method is used, starting in version 7 of the notebook. Thanks, <a href=\"https://www.kaggle.com/E\" target=\"_blank\">@E</a>/S Pronk!</p></li>\n<li><p>Some experiments were run in the notebook:<br>\n<em>Experiment 1</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code>, <code>kidney_1_voi</code>, and <code>kidney_2</code></li>\n<li>Submission:  a set of empty labels in the shape of <code>kidney_1_dense</code>, <code>kidney_1_voi</code>, and <code>kidney_2</code></li>\n<li><code>resize_fraction</code>: 1.0</li>\n<li>Score: out of memory error</li></ul>\n<p><em>Experiment 2</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  a set of empty labels in the shape of <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 1.0</li>\n<li>Score: out of memory error</li></ul>\n<p><em>Experiment 3</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 1.0</li>\n<li>Score: out of memory error</li></ul>\n<p><em>Experiment 4</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 0.2</li>\n<li>Score: 1.0 (max RAM about 4.1GB)</li></ul>\n<p><em>Experiment 5</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 0.5</li>\n<li>Score: 1.0 (max RAM about 9.6GB)</li></ul>\n<p><em>Experiment 5</em></p>\n<ul>\n<li>Solution: <code>submission.csv</code>, truncated to only include <code>kidney_1_dense</code></li>\n<li>Submission:  Same as <code>kidney_1_dense</code></li>\n<li><code>resize_fraction</code>: 0.75</li>\n<li>Score: 1.0 (max RAM about 21.7GB)</li></ul></li>\n</ul>\n<p><strong>23/01/08</strong></p>\n<ul>\n<li>Submission scoring error happens between 2-3 minutes after the notebook is submitted. </li>\n<li>When I run the notebook, it is finished processing kidney_1 and loading kidney_1_voi between minutes 2 and 3. </li>\n<li>This may also be about the amount of time it takes to load a large dataset. </li>\n<li><code>submission.csv</code> has not been created yet.</li>\n</ul>",
      "rawMarkdown": "I'm having trouble overcoming the *Submission Scoring Errors*, and a quick look at the submission forum suggests I'm not alone.\n\nFor this reason, I made a skeleton of our notebook [which can be found here](https://www.kaggle.com/cyberian516/testing-the-scoring-system). This notebook creates nearly empty labels for each image in the dataset. Then it scores itself using the [source code of the Surface Dice Metric](https://www.kaggle.com/code/metric/surface-dice-metric/notebook) suggested in the [competition's main page](www.kaggle.com/competitions/blood-vessel-segmentation/overview/evaluation) without any issues (except for a low score), but it also has a submission scoring error.\n\nThe following issues seem to have occurred in other notebooks:\n* Incorrect RLE implementation [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463224#2572213)]\n* Forgetting to binarize predicted labels [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464104)]\n* Out Of Memory errors from too noisy labels [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464287#2580964)]\n* Incorrect data loading [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/465824#2588855)] [[2](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454732#2521239)]\n\nIt has also been suggested that empty masks can also cause submission scoring errors. [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008)] [[2](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454799#2523059)] [[3](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455062)] [[4](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001)] \n\nHowever, I don't seem to run into any of these issues when running the scoring script *in* our notebook.\n\nUsing the suggested scoring notebook, The skeleton notebook ~~receives a score of 0.0002930597061 on `kidney_1`, `kidney_1_voi`, and `kidney_2`, but \"Submission Scoring Errors\" prevent it from receiving an official score.~~ **doesn't have sufficient memory to compute the score for a set of labels the size of `kidney_1_dense`**\n\nAs a first-time Kaggler, this issue has left me stumped. If anyone else is running into the same vague submission scoring errors, this discussion page might be another good place to diagnose them together.\n\n\n\n\n## Personal notes:\n**23/01/09**\n- The 3D surface dice method is used, starting in version 7 of the notebook. Thanks, @E/S Pronk!\n- Some experiments were run in the notebook:\n *Experiment 1*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`, `kidney_1_voi`, and `kidney_2`\n    - Submission:  a set of empty labels in the shape of `kidney_1_dense`, `kidney_1_voi`, and `kidney_2`\n    - `resize_fraction`: 1.0\n    - Score: out of memory error\n\n *Experiment 2*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  a set of empty labels in the shape of `kidney_1_dense`\n    - `resize_fraction`: 1.0\n    - Score: out of memory error\n\n *Experiment 3*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 1.0\n    - Score: out of memory error\n\n *Experiment 4*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 0.2\n    - Score: 1.0 (max RAM about 4.1GB)\n\n *Experiment 5*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 0.5\n    - Score: 1.0 (max RAM about 9.6GB)\n\n *Experiment 5*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 0.75\n    - Score: 1.0 (max RAM about 21.7GB)\n\n**23/01/08**\n- Submission scoring error happens between 2-3 minutes after the notebook is submitted. \n- When I run the notebook, it is finished processing kidney_1 and loading kidney_1_voi between minutes 2 and 3. \n- This may also be about the amount of time it takes to load a large dataset. \n- `submission.csv` has not been created yet.",
      "votes": null
    },
    {
      "id": "2594414",
      "postDate": "01/09/2024 21:13:27",
      "content": "<p>It seems one difference between your code and the actual surface dice scoring method used for the leaderboard is that the leaderboard uses 3d surface dice, and you seem too be calling 2d surface dice.</p>\n<p>See: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461879\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461879</a></p>",
      "rawMarkdown": "It seems one difference between your code and the actual surface dice scoring method used for the leaderboard is that the leaderboard uses 3d surface dice, and you seem too be calling 2d surface dice.\n\nSee: https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461879",
      "votes": null
    },
    {
      "id": "2595342",
      "postDate": "01/10/2024 11:17:59",
      "content": "<p>Hello, <br>\nI am having problems with the submission as well. However my submission fails after around 35 minutes with a \"Scoring Error\".</p>\n<p>I am trying to run my inference notebook on larger datasets to see if there is a out of memory error, but when I do so the notebook<br>\ntakes only up to 48% of the vRAM…</p>\n<p>This is the link to my notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/fmgsf12/unet-submission</a></p>",
      "rawMarkdown": "Hello, \nI am having problems with the submission as well. However my submission fails after around 35 minutes with a \"Scoring Error\".\n\nI am trying to run my inference notebook on larger datasets to see if there is a out of memory error, but when I do so the notebook\ntakes only up to 48% of the vRAM...\n\nThis is the link to my notebook [https://www.kaggle.com/code/fmgsf12/unet-submission](url)",
      "votes": null
    },
    {
      "id": "2595360",
      "postDate": "01/10/2024 11:27:54",
      "content": "<p>I think you should probably threshold your prediction before rle_encoding it. The rle_encode checks for consecutive values that are the same. If your output is straight from a linear layer or sigmoid activation it will be floating point containing either logits or probabilities.<br>\nDepending on your model output you could threshold simply by: <code>prediction = prediction &gt; threshold</code></p>",
      "rawMarkdown": "I think you should probably threshold your prediction before rle_encoding it. The rle_encode checks for consecutive values that are the same. If your output is straight from a linear layer or sigmoid activation it will be floating point containing either logits or probabilities.\nDepending on your model output you could threshold simply by: `prediction = prediction > threshold`",
      "votes": null
    },
    {
      "id": "2595395",
      "postDate": "01/10/2024 11:53:16",
      "content": "<p>Thanks for your reply.</p>\n<p>This is the transformation I am applying to the predictions</p>\n<pre><code>post_trans = Compose([Activations(=), AsDiscrete(=0.5)]), \n</code></pre>\n<p>Then I make the predictions with</p>\n<pre><code>val_outputs = sliding\nprediction = pad)  i  val_outputs])\nprediction = prediction.detach.cpu.numpy\nrle_predictions = rle\nid_ = f\nsubmission_data = {: id_, : rle_predictions}\ncurrent_submission = pd.\nsubmission.append(current_submission)\n</code></pre>\n<p>So they should already be thresholded.<br>\nI also checked that the minimum and maximum of the predictions were respectively 0 and 1 (float32)<br>\nand this was the case.<br>\ncould the problem be that they are given as floating points? <br>\nBut then this should cause the notebook to fail as well on the example test data, but there it works just fine in writing the <code>submission.csv</code> file</p>",
      "rawMarkdown": "Thanks for your reply.\n\nThis is the transformation I am applying to the predictions\n\n```\npost_trans = Compose([Activations(sigmoid=True), AsDiscrete(threshold=0.5)]), \n```\n\nThen I make the predictions with\n\n```\nval_outputs = sliding_window_inference(image, roi_size, sw_batch_size, best_model)\nprediction = pad_list_data_collate([torch.squeeze(post_trans(i)) for i in val_outputs])\nprediction = prediction.detach().cpu().numpy()\nrle_predictions = rle_encode(prediction)\nid_ = f\"{dataset_name}_{idx:04}\"\nsubmission_data = {\"id\": id_, \"rle\": rle_predictions}\ncurrent_submission = pd.DataFrame(data= submission_data, index=[0])\nsubmission.append(current_submission)\n```\n\n\nSo they should already be thresholded.\nI also checked that the minimum and maximum of the predictions were respectively 0 and 1 (float32)\nand this was the case.\ncould the problem be that they are given as floating points? \nBut then this should cause the notebook to fail as well on the example test data, but there it works just fine in writing the `submission.csv` file",
      "votes": null
    },
    {
      "id": "2595465",
      "postDate": "01/10/2024 13:00:43",
      "content": "<p>ah ok, I saw only the second piece of code didn't know you had a layer that applies thresholding. It shouldn't be a problem that they are floating point. One other explanation I can think of is that the prediction contains too many ones. Like if the final layer before post_trans gave mostly positive numbers, or was normalized for instance, which would cause the subsequent sigmoid to be higher than threshold. Maybe check the output before post_trans or try with a higher threshold? </p>",
      "rawMarkdown": "ah ok, I saw only the second piece of code didn't know you had a layer that applies thresholding. It shouldn't be a problem that they are floating point. One other explanation I can think of is that the prediction contains too many ones. Like if the final layer before post_trans gave mostly positive numbers, or was normalized for instance, which would cause the subsequent sigmoid to be higher than threshold. Maybe check the output before post_trans or try with a higher threshold?",
      "votes": null
    },
    {
      "id": "2595471",
      "postDate": "01/10/2024 13:08:18",
      "content": "<p>Thanks again for your reply. But if the predictions contains too many ones, shouldn't it just yield a lower score, without errors? <br>\nThis is still useful to check, but I visualized some images and it doesn't look like it is \"oversegmenting\" it</p>",
      "rawMarkdown": "Thanks again for your reply. But if the predictions contains too many ones, shouldn't it just yield a lower score, without errors? \nThis is still useful to check, but I visualized some images and it doesn't look like it is \"oversegmenting\" it",
      "votes": null
    },
    {
      "id": "2595483",
      "postDate": "01/10/2024 13:17:28",
      "content": "<p>ideally it would give a lower score, but I think there are some limits in what it can score, because internally it tries to calculate the distance between the prediction and the ground truth, I speculate is that if there are too many ones it becomes too hard or resource intensive to calculate. Anyway, I've seen it before when I accidentally produced a prediction with a lot of noise.</p>",
      "rawMarkdown": "ideally it would give a lower score, but I think there are some limits in what it can score, because internally it tries to calculate the distance between the prediction and the ground truth, I speculate is that if there are too many ones it becomes too hard or resource intensive to calculate. Anyway, I've seen it before when I accidentally produced a prediction with a lot of noise.",
      "votes": null
    },
    {
      "id": "2595638",
      "postDate": "01/10/2024 14:39:31",
      "content": "<p>Unfortunately, this is not the case either, the resulting segmentation are still sparse and the majority of the values before the sigmoid activation are negative… I ran the scoring notebook on my predictions for the kidney_1_dense dataset against the 'train_rle.csv' file and initially it didn't work because it was expeting the \"width\" and \"height\" columns in the merged dataframe with the predictions and the ground truth segmentations. From the example submission it doesn't seem like these columns are required in the file, so I guess they'll alredy be in the ground truth solutions… I am new to notebook competitions procedures and probably there is something that I am still missing, but I still cannot see why my notebook fails the scoring after running for 34 minutes. There is no way to see the internal logs from the scoring notebook, right? Anyway, Im doing a submission with only empty predictions to see if that at least works EDIT: It breaks even if I don't make any predictions and submit only empty inferred masks….</p>",
      "rawMarkdown": "Unfortunately, this is not the case either, the resulting segmentation are still sparse and the majority of the values before the sigmoid activation are negative... I ran the scoring notebook on my predictions for the kidney_1_dense dataset against the 'train_rle.csv' file and initially it didn't work because it was expeting the \"width\" and \"height\" columns in the merged dataframe with the predictions and the ground truth segmentations. From the example submission it doesn't seem like these columns are required in the file, so I guess they'll alredy be in the ground truth solutions... I am new to notebook competitions procedures and probably there is something that I am still missing, but I still cannot see why my notebook fails the scoring after running for 34 minutes. There is no way to see the internal logs from the scoring notebook, right? Anyway, Im doing a submission with only empty predictions to see if that at least works EDIT: It breaks even if I don't make any predictions and submit only empty inferred masks....",
      "votes": null
    },
    {
      "id": "2595740",
      "postDate": "01/10/2024 15:58:38",
      "content": "<p>ok, i see something else. You are using idx from an enumerate</p>\n<pre><code>for idx, batch in (test_data_loader):\n</code></pre>\n<p>but maybe the filenames don't start with 0000.tif in the public test set, it could start with something like 0500.tif just like the kidney_3_dense dataset. So you'd have to store the image names alongside the data in order to generate the correct ids for the csv</p>",
      "rawMarkdown": "ok, i see something else. You are using idx from an enumerate\n```\nfor idx, batch in enumerate(test_data_loader):\n```\n\nbut maybe the filenames don't start with 0000.tif in the public test set, it could start with something like 0500.tif just like the kidney_3_dense dataset. So you'd have to store the image names alongside the data in order to generate the correct ids for the csv",
      "votes": null
    },
    {
      "id": "2596025",
      "postDate": "01/10/2024 19:53:57",
      "content": "<p>This error is exceedingly frustrating. I think mine has something to do with file naming? But I have no idea. I made it such that my code runs through each folder in the '/test' folder and then updates the rle dictionary accordingly. Is this the right way to do it? </p>\n<p>Is there any example notebooks that literally just submit a simple working submission (no model etc.) that I could use to debug?</p>",
      "rawMarkdown": "This error is exceedingly frustrating. I think mine has something to do with file naming? But I have no idea. I made it such that my code runs through each folder in the '/test' folder and then updates the rle dictionary accordingly. Is this the right way to do it? \n\nIs there any example notebooks that literally just submit a simple working submission (no model etc.) that I could use to debug?",
      "votes": null
    },
    {
      "id": "2596038",
      "postDate": "01/10/2024 20:08:57",
      "content": "<p>I don't have any submissions left to test with, but this should work:</p>\n<pre><code>import os\nimport glob\n\ndata_path = \ndatasets = \n\ncsv_lines = \n dataset  datasets:\n     file  (glob(os(dataset, ,))):\n        id = os(dataset)\n        id = id +  + os(file)\n        id = id(, )\n\n        rle = \n        csv_lines(f)\n\nwith (, ) as stream:\n    stream(csv_lines)\n</code></pre>",
      "rawMarkdown": "I don't have any submissions left to test with, but this should work:\n\n```\nimport os\nimport glob\n\ndata_path = \"../..\"\ndatasets = [os.path.join(data_path, ds) for ds in (\"test/kidney_5\", \"test/kidney_6\")]\n\ncsv_lines = [\"id,rle\\n\"]\nfor dataset in datasets:\n    for file in sorted(glob.glob(os.path.join(dataset, \"images\",\"*.tif\"))):\n        id = os.path.split(dataset)[-1]\n        id = id + \"_\" + os.path.split(file)[-1]\n        id = id.replace(\".tif\", \"\")\n        \n        rle = \"1 1\"\n        csv_lines.append(f\"{id},{rle}\\n\")\n\nwith open(\"submission.csv\", \"w\") as stream:\n    stream.writelines(csv_lines)\n```",
      "votes": null
    },
    {
      "id": "2596042",
      "postDate": "01/10/2024 20:09:44",
      "content": "<p>update data path to the relevant path on kaggle or your local system</p>",
      "rawMarkdown": "update data path to the relevant path on kaggle or your local system",
      "votes": null
    },
    {
      "id": "2596201",
      "postDate": "01/10/2024 23:51:14",
      "content": "<p>So if your code runs on the practice submission with kidney_5 and kidney_6 hardcoded in as the file paths, then it would also in theory run on the actual submission? </p>",
      "rawMarkdown": "So if your code runs on the practice submission with kidney_5 and kidney_6 hardcoded in as the file paths, then it would also in theory run on the actual submission?",
      "votes": null
    },
    {
      "id": "2596540",
      "postDate": "01/11/2024 06:39:38",
      "content": "<p>Thanks, good catch! Version 7 and beyond should use the 3D surface dice metric now.</p>",
      "rawMarkdown": "Thanks, good catch! Version 7 and beyond should use the 3D surface dice metric now.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2594414,
      "author_name": "limitz",
      "author_url": "",
      "post_date": "01/09/2024 21:13:27",
      "content": "<p>It seems one difference between your code and the actual surface dice scoring method used for the leaderboard is that the leaderboard uses 3d surface dice, and you seem too be calling 2d surface dice.</p>\n<p>See: <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461879\" target=\"_blank\">https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461879</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2596540,
          "author_name": "cyberian516",
          "author_url": "",
          "post_date": "01/11/2024 06:39:38",
          "content": "<p>Thanks, good catch! Version 7 and beyond should use the 3D surface dice metric now.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2595342,
      "author_name": "fmgsf12",
      "author_url": "",
      "post_date": "01/10/2024 11:17:59",
      "content": "<p>Hello, <br>\nI am having problems with the submission as well. However my submission fails after around 35 minutes with a \"Scoring Error\".</p>\n<p>I am trying to run my inference notebook on larger datasets to see if there is a out of memory error, but when I do so the notebook<br>\ntakes only up to 48% of the vRAM…</p>\n<p>This is the link to my notebook <a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/fmgsf12/unet-submission</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 2595360,
          "author_name": "limitz",
          "author_url": "",
          "post_date": "01/10/2024 11:27:54",
          "content": "<p>I think you should probably threshold your prediction before rle_encoding it. The rle_encode checks for consecutive values that are the same. If your output is straight from a linear layer or sigmoid activation it will be floating point containing either logits or probabilities.<br>\nDepending on your model output you could threshold simply by: <code>prediction = prediction &gt; threshold</code></p>",
          "votes": null,
          "replies": [
            {
              "id": 2595395,
              "author_name": "fmgsf12",
              "author_url": "",
              "post_date": "01/10/2024 11:53:16",
              "content": "<p>Thanks for your reply.</p>\n<p>This is the transformation I am applying to the predictions</p>\n<pre><code>post_trans = Compose([Activations(=), AsDiscrete(=0.5)]), \n</code></pre>\n<p>Then I make the predictions with</p>\n<pre><code>val_outputs = sliding\nprediction = pad)  i  val_outputs])\nprediction = prediction.detach.cpu.numpy\nrle_predictions = rle\nid_ = f\nsubmission_data = {: id_, : rle_predictions}\ncurrent_submission = pd.\nsubmission.append(current_submission)\n</code></pre>\n<p>So they should already be thresholded.<br>\nI also checked that the minimum and maximum of the predictions were respectively 0 and 1 (float32)<br>\nand this was the case.<br>\ncould the problem be that they are given as floating points? <br>\nBut then this should cause the notebook to fail as well on the example test data, but there it works just fine in writing the <code>submission.csv</code> file</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2595465,
                  "author_name": "limitz",
                  "author_url": "",
                  "post_date": "01/10/2024 13:00:43",
                  "content": "<p>ah ok, I saw only the second piece of code didn't know you had a layer that applies thresholding. It shouldn't be a problem that they are floating point. One other explanation I can think of is that the prediction contains too many ones. Like if the final layer before post_trans gave mostly positive numbers, or was normalized for instance, which would cause the subsequent sigmoid to be higher than threshold. Maybe check the output before post_trans or try with a higher threshold? </p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2595471,
                      "author_name": "fmgsf12",
                      "author_url": "",
                      "post_date": "01/10/2024 13:08:18",
                      "content": "<p>Thanks again for your reply. But if the predictions contains too many ones, shouldn't it just yield a lower score, without errors? <br>\nThis is still useful to check, but I visualized some images and it doesn't look like it is \"oversegmenting\" it</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2595483,
                          "author_name": "limitz",
                          "author_url": "",
                          "post_date": "01/10/2024 13:17:28",
                          "content": "<p>ideally it would give a lower score, but I think there are some limits in what it can score, because internally it tries to calculate the distance between the prediction and the ground truth, I speculate is that if there are too many ones it becomes too hard or resource intensive to calculate. Anyway, I've seen it before when I accidentally produced a prediction with a lot of noise.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2595638,
                              "author_name": "fmgsf12",
                              "author_url": "",
                              "post_date": "01/10/2024 14:39:31",
                              "content": "<p>Unfortunately, this is not the case either, the resulting segmentation are still sparse and the majority of the values before the sigmoid activation are negative… I ran the scoring notebook on my predictions for the kidney_1_dense dataset against the 'train_rle.csv' file and initially it didn't work because it was expeting the \"width\" and \"height\" columns in the merged dataframe with the predictions and the ground truth segmentations. From the example submission it doesn't seem like these columns are required in the file, so I guess they'll alredy be in the ground truth solutions… I am new to notebook competitions procedures and probably there is something that I am still missing, but I still cannot see why my notebook fails the scoring after running for 34 minutes. There is no way to see the internal logs from the scoring notebook, right? Anyway, Im doing a submission with only empty predictions to see if that at least works EDIT: It breaks even if I don't make any predictions and submit only empty inferred masks….</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2595740,
                                  "author_name": "limitz",
                                  "author_url": "",
                                  "post_date": "01/10/2024 15:58:38",
                                  "content": "<p>ok, i see something else. You are using idx from an enumerate</p>\n<pre><code>for idx, batch in (test_data_loader):\n</code></pre>\n<p>but maybe the filenames don't start with 0000.tif in the public test set, it could start with something like 0500.tif just like the kidney_3_dense dataset. So you'd have to store the image names alongside the data in order to generate the correct ids for the csv</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2596025,
      "author_name": "jekasm19",
      "author_url": "",
      "post_date": "01/10/2024 19:53:57",
      "content": "<p>This error is exceedingly frustrating. I think mine has something to do with file naming? But I have no idea. I made it such that my code runs through each folder in the '/test' folder and then updates the rle dictionary accordingly. Is this the right way to do it? </p>\n<p>Is there any example notebooks that literally just submit a simple working submission (no model etc.) that I could use to debug?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2596038,
          "author_name": "limitz",
          "author_url": "",
          "post_date": "01/10/2024 20:08:57",
          "content": "<p>I don't have any submissions left to test with, but this should work:</p>\n<pre><code>import os\nimport glob\n\ndata_path = \ndatasets = \n\ncsv_lines = \n dataset  datasets:\n     file  (glob(os(dataset, ,))):\n        id = os(dataset)\n        id = id +  + os(file)\n        id = id(, )\n\n        rle = \n        csv_lines(f)\n\nwith (, ) as stream:\n    stream(csv_lines)\n</code></pre>",
          "votes": null,
          "replies": [
            {
              "id": 2596042,
              "author_name": "limitz",
              "author_url": "",
              "post_date": "01/10/2024 20:09:44",
              "content": "<p>update data path to the relevant path on kaggle or your local system</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2596201,
                  "author_name": "jekasm19",
                  "author_url": "",
                  "post_date": "01/10/2024 23:51:14",
                  "content": "<p>So if your code runs on the practice submission with kidney_5 and kidney_6 hardcoded in as the file paths, then it would also in theory run on the actual submission? </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2593189": "I'm having trouble overcoming the *Submission Scoring Errors*, and a quick look at the submission forum suggests I'm not alone.\n\nFor this reason, I made a skeleton of our notebook [which can be found here](https://www.kaggle.com/cyberian516/testing-the-scoring-system). This notebook creates nearly empty labels for each image in the dataset. Then it scores itself using the [source code of the Surface Dice Metric](https://www.kaggle.com/code/metric/surface-dice-metric/notebook) suggested in the [competition's main page](www.kaggle.com/competitions/blood-vessel-segmentation/overview/evaluation) without any issues (except for a low score), but it also has a submission scoring error.\n\nThe following issues seem to have occurred in other notebooks:\n* Incorrect RLE implementation [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/463224#2572213)]\n* Forgetting to binarize predicted labels [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464104)]\n* Out Of Memory errors from too noisy labels [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/464287#2580964)]\n* Incorrect data loading [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/465824#2588855)] [[2](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454732#2521239)]\n\nIt has also been suggested that empty masks can also cause submission scoring errors. [[1](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455008)] [[2](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/454799#2523059)] [[3](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455062)] [[4](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455001)] \n\nHowever, I don't seem to run into any of these issues when running the scoring script *in* our notebook.\n\nUsing the suggested scoring notebook, The skeleton notebook ~~receives a score of 0.0002930597061 on `kidney_1`, `kidney_1_voi`, and `kidney_2`, but \"Submission Scoring Errors\" prevent it from receiving an official score.~~ **doesn't have sufficient memory to compute the score for a set of labels the size of `kidney_1_dense`**\n\nAs a first-time Kaggler, this issue has left me stumped. If anyone else is running into the same vague submission scoring errors, this discussion page might be another good place to diagnose them together.\n\n\n\n\n## Personal notes:\n**23/01/09**\n- The 3D surface dice method is used, starting in version 7 of the notebook. Thanks, @E/S Pronk!\n- Some experiments were run in the notebook:\n *Experiment 1*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`, `kidney_1_voi`, and `kidney_2`\n    - Submission:  a set of empty labels in the shape of `kidney_1_dense`, `kidney_1_voi`, and `kidney_2`\n    - `resize_fraction`: 1.0\n    - Score: out of memory error\n\n *Experiment 2*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  a set of empty labels in the shape of `kidney_1_dense`\n    - `resize_fraction`: 1.0\n    - Score: out of memory error\n\n *Experiment 3*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 1.0\n    - Score: out of memory error\n\n *Experiment 4*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 0.2\n    - Score: 1.0 (max RAM about 4.1GB)\n\n *Experiment 5*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 0.5\n    - Score: 1.0 (max RAM about 9.6GB)\n\n *Experiment 5*\n    - Solution: `submission.csv`, truncated to only include `kidney_1_dense`\n    - Submission:  Same as `kidney_1_dense`\n    - `resize_fraction`: 0.75\n    - Score: 1.0 (max RAM about 21.7GB)\n\n**23/01/08**\n- Submission scoring error happens between 2-3 minutes after the notebook is submitted. \n- When I run the notebook, it is finished processing kidney_1 and loading kidney_1_voi between minutes 2 and 3. \n- This may also be about the amount of time it takes to load a large dataset. \n- `submission.csv` has not been created yet.",
    "2594414": "It seems one difference between your code and the actual surface dice scoring method used for the leaderboard is that the leaderboard uses 3d surface dice, and you seem too be calling 2d surface dice.\n\nSee: https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461879",
    "2595342": "Hello, \nI am having problems with the submission as well. However my submission fails after around 35 minutes with a \"Scoring Error\".\n\nI am trying to run my inference notebook on larger datasets to see if there is a out of memory error, but when I do so the notebook\ntakes only up to 48% of the vRAM...\n\nThis is the link to my notebook [https://www.kaggle.com/code/fmgsf12/unet-submission](url)",
    "2595360": "I think you should probably threshold your prediction before rle_encoding it. The rle_encode checks for consecutive values that are the same. If your output is straight from a linear layer or sigmoid activation it will be floating point containing either logits or probabilities.\nDepending on your model output you could threshold simply by: `prediction = prediction > threshold`",
    "2595395": "Thanks for your reply.\n\nThis is the transformation I am applying to the predictions\n\n```\npost_trans = Compose([Activations(sigmoid=True), AsDiscrete(threshold=0.5)]), \n```\n\nThen I make the predictions with\n\n```\nval_outputs = sliding_window_inference(image, roi_size, sw_batch_size, best_model)\nprediction = pad_list_data_collate([torch.squeeze(post_trans(i)) for i in val_outputs])\nprediction = prediction.detach().cpu().numpy()\nrle_predictions = rle_encode(prediction)\nid_ = f\"{dataset_name}_{idx:04}\"\nsubmission_data = {\"id\": id_, \"rle\": rle_predictions}\ncurrent_submission = pd.DataFrame(data= submission_data, index=[0])\nsubmission.append(current_submission)\n```\n\n\nSo they should already be thresholded.\nI also checked that the minimum and maximum of the predictions were respectively 0 and 1 (float32)\nand this was the case.\ncould the problem be that they are given as floating points? \nBut then this should cause the notebook to fail as well on the example test data, but there it works just fine in writing the `submission.csv` file",
    "2595465": "ah ok, I saw only the second piece of code didn't know you had a layer that applies thresholding. It shouldn't be a problem that they are floating point. One other explanation I can think of is that the prediction contains too many ones. Like if the final layer before post_trans gave mostly positive numbers, or was normalized for instance, which would cause the subsequent sigmoid to be higher than threshold. Maybe check the output before post_trans or try with a higher threshold?",
    "2595471": "Thanks again for your reply. But if the predictions contains too many ones, shouldn't it just yield a lower score, without errors? \nThis is still useful to check, but I visualized some images and it doesn't look like it is \"oversegmenting\" it",
    "2595483": "ideally it would give a lower score, but I think there are some limits in what it can score, because internally it tries to calculate the distance between the prediction and the ground truth, I speculate is that if there are too many ones it becomes too hard or resource intensive to calculate. Anyway, I've seen it before when I accidentally produced a prediction with a lot of noise.",
    "2595638": "Unfortunately, this is not the case either, the resulting segmentation are still sparse and the majority of the values before the sigmoid activation are negative... I ran the scoring notebook on my predictions for the kidney_1_dense dataset against the 'train_rle.csv' file and initially it didn't work because it was expeting the \"width\" and \"height\" columns in the merged dataframe with the predictions and the ground truth segmentations. From the example submission it doesn't seem like these columns are required in the file, so I guess they'll alredy be in the ground truth solutions... I am new to notebook competitions procedures and probably there is something that I am still missing, but I still cannot see why my notebook fails the scoring after running for 34 minutes. There is no way to see the internal logs from the scoring notebook, right? Anyway, Im doing a submission with only empty predictions to see if that at least works EDIT: It breaks even if I don't make any predictions and submit only empty inferred masks....",
    "2595740": "ok, i see something else. You are using idx from an enumerate\n```\nfor idx, batch in enumerate(test_data_loader):\n```\n\nbut maybe the filenames don't start with 0000.tif in the public test set, it could start with something like 0500.tif just like the kidney_3_dense dataset. So you'd have to store the image names alongside the data in order to generate the correct ids for the csv",
    "2596025": "This error is exceedingly frustrating. I think mine has something to do with file naming? But I have no idea. I made it such that my code runs through each folder in the '/test' folder and then updates the rle dictionary accordingly. Is this the right way to do it? \n\nIs there any example notebooks that literally just submit a simple working submission (no model etc.) that I could use to debug?",
    "2596038": "I don't have any submissions left to test with, but this should work:\n\n```\nimport os\nimport glob\n\ndata_path = \"../..\"\ndatasets = [os.path.join(data_path, ds) for ds in (\"test/kidney_5\", \"test/kidney_6\")]\n\ncsv_lines = [\"id,rle\\n\"]\nfor dataset in datasets:\n    for file in sorted(glob.glob(os.path.join(dataset, \"images\",\"*.tif\"))):\n        id = os.path.split(dataset)[-1]\n        id = id + \"_\" + os.path.split(file)[-1]\n        id = id.replace(\".tif\", \"\")\n        \n        rle = \"1 1\"\n        csv_lines.append(f\"{id},{rle}\\n\")\n\nwith open(\"submission.csv\", \"w\") as stream:\n    stream.writelines(csv_lines)\n```",
    "2596042": "update data path to the relevant path on kaggle or your local system",
    "2596201": "So if your code runs on the practice submission with kidney_5 and kidney_6 hardcoded in as the file paths, then it would also in theory run on the actual submission?",
    "2596540": "Thanks, good catch! Version 7 and beyond should use the 3D surface dice metric now."
  },
  "source": "meta"
}