{
  "id": 456761,
  "title": "Fixing the metric computation code",
  "url": "/competitions/blood-vessel-segmentation/discussion/456761",
  "author_name": "Optimo",
  "post_date": "2023-11-21T15:21:40.969000",
  "votes": 19,
  "comment_count": 44,
  "views": 0,
  "content": "<p>There is currently an OOM error occurring depending on the \"complexity\" of the predictions made.</p>\n<p>See this <a href=\"https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline\" target=\"_blank\">notebook</a> or this <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455220\" target=\"_blank\">discussion</a> to get more info.</p>\n<p>I'm creating this thread so that we can all share concrete improvements that can be done on the <a href=\"https://www.kaggle.com/code/metric/surface-dice-metric/notebook\" target=\"_blank\">official metric code</a> in order to reduce first memory usage, and also speed if possible.</p>\n<p>A correct metric accepting any submission in correct format will ensure a fair competition and the best possible outcome for kagglers and organizers.</p>\n<p>Thanks for your help!</p>",
  "messages": [
    {
      "id": 2533084,
      "postDate": "2023-11-21T15:21:40.970Z",
      "content": "<p>There is currently an OOM error occurring depending on the \"complexity\" of the predictions made.</p>\n<p>See this <a href=\"https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline\" target=\"_blank\">notebook</a> or this <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455220\" target=\"_blank\">discussion</a> to get more info.</p>\n<p>I'm creating this thread so that we can all share concrete improvements that can be done on the <a href=\"https://www.kaggle.com/code/metric/surface-dice-metric/notebook\" target=\"_blank\">official metric code</a> in order to reduce first memory usage, and also speed if possible.</p>\n<p>A correct metric accepting any submission in correct format will ensure a fair competition and the best possible outcome for kagglers and organizers.</p>\n<p>Thanks for your help!</p>",
      "rawMarkdown": "There is currently an OOM error occurring depending on the \"complexity\" of the predictions made.\n\nSee this [notebook](https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline) or this [discussion](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455220) to get more info.\n\nI'm creating this thread so that we can all share concrete improvements that can be done on the [official metric code](https://www.kaggle.com/code/metric/surface-dice-metric/notebook) in order to reduce first memory usage, and also speed if possible.\n\nA correct metric accepting any submission in correct format will ensure a fair competition and the best possible outcome for kagglers and organizers.\n\nThanks for your help!",
      "votes": 19
    },
    {
      "id": 2554559,
      "postDate": "2023-12-09T07:38:50.897Z",
      "content": "<p>As hengck23 said, the code with tolerance 0 can be greatly simplified and reduce memory usage. If I understand correctly, only 2 slices need to be on RAM, not the whole volume, and this is exact, not an approximation.</p>\n<p>My implementation is this: <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a></p>",
      "rawMarkdown": "As hengck23 said, the code with tolerance 0 can be greatly simplified and reduce memory usage. If I understand correctly, only 2 slices need to be on RAM, not the whole volume, and this is exact, not an approximation.\n\nMy implementation is this: https://www.kaggle.com/code/junkoda/fast-surface-dice-computation",
      "votes": 14,
      "replies": [
        {
          "id": 2554827,
          "postDate": "2023-12-09T12:46:06.107Z",
          "content": "<p>thanks for the implementation.<br>\nSurface dice computation is reduced from 15 min to a few sec for me. </p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nyou may want to take a look at this</p>",
          "rawMarkdown": "thanks for the implementation.\nSurface dice computation is reduced from 15 min to a few sec for me. \n\n\n@ryanholbrook \nyou may want to take a look at this",
          "votes": 3
        }
      ]
    },
    {
      "id": 2551160,
      "postDate": "2023-12-06T14:26:03.507Z",
      "content": "<p>I just updated the metric to include the dtype downcasting suggested by <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>. In my tests, this reduces memory consumption by about 20% while preserving scores out to six decimals. Please let me know if you observe any unexpected changes to submissions scores though.</p>",
      "rawMarkdown": "I just updated the metric to include the dtype downcasting suggested by @optimo. In my tests, this reduces memory consumption by about 20% while preserving scores out to six decimals. Please let me know if you observe any unexpected changes to submissions scores though.",
      "votes": 9,
      "replies": [
        {
          "id": 2551420,
          "postDate": "2023-12-06T17:27:03.690Z",
          "content": "<p>worked on my case! 🎉</p>",
          "rawMarkdown": "worked on my case! 🎉",
          "votes": 1,
          "replies": [
            {
              "id": 2559212,
              "postDate": "2023-12-12T16:56:22.370Z",
              "content": "<p>Actually <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> it worked on my first submission case, now my second attempt with different model also gets OOM.</p>\n<p>If we are not planning to change the competition metric (as discussed <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041\" target=\"_blank\">here</a>) I think it is worth using <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s implementation which is both much faster and uses less memory with the same scores on my internal tests.</p>",
              "rawMarkdown": "Actually @ryanholbrook it worked on my first submission case, now my second attempt with different model also gets OOM.\n\nIf we are not planning to change the competition metric (as discussed [here](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041)) I think it is worth using @junkoda's implementation which is both much faster and uses less memory with the same scores on my internal tests.",
              "votes": 3
            },
            {
              "id": 2567370,
              "postDate": "2023-12-19T15:15:21.083Z",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> What kind of tests and code do you need ? So that we can prepare it for you to try and validate when you'll have time after santa's competition? <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">code</a> looks already pretty good but we can add a few more checks if you want.</p>",
              "rawMarkdown": "@ryanholbrook What kind of tests and code do you need ? So that we can prepare it for you to try and validate when you'll have time after santa's competition? @junkoda's [code](https://www.kaggle.com/code/junkoda/fast-surface-dice-computation) looks already pretty good but we can add a few more checks if you want."
            },
            {
              "id": 2570545,
              "postDate": "2023-12-22T10:22:14.503Z",
              "content": "<p>Actually I had a bug in my pipeline, the error was not OOM.</p>",
              "rawMarkdown": "Actually I had a bug in my pipeline, the error was not OOM.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2553075,
          "postDate": "2023-12-08T01:57:12.610Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    },
    {
      "id": 2544522,
      "postDate": "2023-11-30T22:14:37.667Z",
      "content": "<p>Hi everyone,</p>\n<p>I just made a change to metric that I think will allow more submissions to succeed without OOM. The main change is to the <code>_sort_distances_surfels</code> function, which now sorts in place. It seems that some submissions can create a large number of surface areas, which caused this function to allocate a great deal of memory unnecessarily.</p>\n<p>I'm also going to run some tests on the dtype casting suggested by <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>, but wanted to release this first since it didn't require any changes to numerics.</p>\n<p>Thank you everyone for your patience. I know this has been less than ideal. I'm hoping these improvements will make this a less frustrating experience.</p>",
      "rawMarkdown": "Hi everyone,\n\nI just made a change to metric that I think will allow more submissions to succeed without OOM. The main change is to the `_sort_distances_surfels` function, which now sorts in place. It seems that some submissions can create a large number of surface areas, which caused this function to allocate a great deal of memory unnecessarily.\n\nI'm also going to run some tests on the dtype casting suggested by @optimo, but wanted to release this first since it didn't require any changes to numerics.\n\nThank you everyone for your patience. I know this has been less than ideal. I'm hoping these improvements will make this a less frustrating experience.",
      "votes": 3,
      "replies": [
        {
          "id": 2544949,
          "postDate": "2023-12-01T07:39:03.607Z",
          "content": "<p>I really appreciate your ongoing efforts. 😃<br>\nHowever, I continue to encounter a submission scoring error…😭</p>",
          "rawMarkdown": "I really appreciate your ongoing efforts. 😃\nHowever, I continue to encounter a submission scoring error...😭",
          "votes": 1,
          "replies": [
            {
              "id": 2544970,
              "postDate": "2023-12-01T08:06:44.503Z",
              "content": "<p>I see the same.</p>\n<p>I had the same notebook run yesterday and today and I don't see a change in the scoring error. Not sure if my predictions are still poor/complex.</p>",
              "rawMarkdown": "I see the same.\n\nI had the same notebook run yesterday and today and I don't see a change in the scoring error. Not sure if my predictions are still poor/complex.",
              "votes": 1
            }
          ]
        },
        {
          "id": 2545109,
          "postDate": "2023-12-01T09:54:02.610Z",
          "content": "<p>I just tried thresholds 0.5 and 0.85 that still get OOM (threshold 0.95 runs ok).<br>\nI think the dtype casting is safe to use -&gt; also note that I delete the original images after the crop in my code since they are not used anymore after that.</p>",
          "rawMarkdown": "I just tried thresholds 0.5 and 0.85 that still get OOM (threshold 0.95 runs ok).\nI think the dtype casting is safe to use -> also note that I delete the original images after the crop in my code since they are not used anymore after that."
        },
        {
          "id": 2550019,
          "postDate": "2023-12-05T17:38:55.060Z",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, have you been able to dig more on the dtype casting proposal ?</p>",
          "rawMarkdown": "Hello @ryanholbrook, have you been able to dig more on the dtype casting proposal ?",
          "replies": [
            {
              "id": 2550036,
              "postDate": "2023-12-05T17:51:25.063Z",
              "content": "<p>My apologies for piling on this but any update on this issue from the kaggle team would really help. Thank you.</p>",
              "rawMarkdown": "My apologies for piling on this but any update on this issue from the kaggle team would really help. Thank you.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2537642,
      "postDate": "2023-11-25T11:31:29.417Z",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I made a code update proposal here : <a href=\"https://www.kaggle.com/optimo/surface-dice-metric-low-memory/\" target=\"_blank\">https://www.kaggle.com/optimo/surface-dice-metric-low-memory/</a></p>\n<p>It's only minimal changes but it should reduce the bottle neck matrices footprints by two (float64 -&gt; float32). Some explanations are available at the beginning of the notebook.</p>\n<p>Could you have a look and let me know what you think?</p>\n<p>If you are willing to test it on real submissions feel free to have a look at my failed submissions with OOM which should hopefully run fine with the proposed changes.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "@ryanholbrook I made a code update proposal here : https://www.kaggle.com/optimo/surface-dice-metric-low-memory/\n\nIt's only minimal changes but it should reduce the bottle neck matrices footprints by two (float64 -> float32). Some explanations are available at the beginning of the notebook.\n\nCould you have a look and let me know what you think?\n\nIf you are willing to test it on real submissions feel free to have a look at my failed submissions with OOM which should hopefully run fine with the proposed changes.\n\nThank you!",
      "votes": 4,
      "replies": [
        {
          "id": 2537692,
          "postDate": "2023-11-25T12:15:37.530Z",
          "content": "<p>Really good work, In your experience, is it faster or slower while reducing memory, because if calculations are happening in float32, they can be faster, also lowering it to float16 might also be a viable option as the masks are only 0 and 1 but some calculations can take longer with float16</p>",
          "rawMarkdown": "Really good work, In your experience, is it faster or slower while reducing memory, because if calculations are happening in float32, they can be faster, also lowering it to float16 might also be a viable option as the masks are only 0 and 1 but some calculations can take longer with float16",
          "replies": [
            {
              "id": 2537809,
              "postDate": "2023-11-25T13:56:23Z",
              "content": "<p>I did not notice any speed improvement with float32 in my experiments. I tried to switch everything to float16 but ended up with Nan score.</p>",
              "rawMarkdown": "I did not notice any speed improvement with float32 in my experiments. I tried to switch everything to float16 but ended up with Nan score."
            }
          ]
        },
        {
          "id": 2541485,
          "postDate": "2023-11-28T14:36:12.837Z",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> any chance you could have a look at this proposal ? 🙏</p>",
          "rawMarkdown": "@ryanholbrook any chance you could have a look at this proposal ? 🙏",
          "votes": 1
        },
        {
          "id": 2542035,
          "postDate": "2023-11-29T03:01:15.243Z",
          "content": "<p>I ran the code and it significantly reduced the memory usage. Nice work!</p>",
          "rawMarkdown": "I ran the code and it significantly reduced the memory usage. Nice work!",
          "votes": 1,
          "replies": [
            {
              "id": 2542690,
              "postDate": "2023-11-29T12:36:38.590Z",
              "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> maybe ? 🙂 it seems quite straightforward to test the new implementation!</p>",
              "rawMarkdown": "@inversion maybe ? 🙂 it seems quite straightforward to test the new implementation!",
              "votes": 1
            },
            {
              "id": 2543563,
              "postDate": "2023-11-30T07:17:17.120Z",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, it appears that the recent 12 submissions are encountering submission scoring errors, potentially due to OOM issues… Even if <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> 's code doesn't resolve the OOM issue, I would greatly appreciate the prompt presentation of alternative solutions.😭</p>",
              "rawMarkdown": "@ryanholbrook, it appears that the recent 12 submissions are encountering submission scoring errors, potentially due to OOM issues… Even if @optimo 's code doesn't resolve the OOM issue, I would greatly appreciate the prompt presentation of alternative solutions.😭"
            }
          ]
        },
        {
          "id": 2542724,
          "postDate": "2023-11-29T13:09:17.313Z",
          "content": "<p>Hi, Optimo</p>\n<p>Did you try comparing the final score for fp64 and fp32? I wonder what is the scale of differences between the two. </p>",
          "rawMarkdown": "Hi, Optimo\n\nDid you try comparing the final score for fp64 and fp32? I wonder what is the scale of differences between the two. ",
          "replies": [
            {
              "id": 2542739,
              "postDate": "2023-11-29T13:30:22.257Z",
              "content": "<p>I had the exact same score (on the cases I tried), I believe the computation is exactly the same as the final computation is done in float32 anyway, it's just that some matrices are stored with float64 precision because <code>np.array()</code> creates them like that by default.</p>",
              "rawMarkdown": "I had the exact same score (on the cases I tried), I believe the computation is exactly the same as the final computation is done in float32 anyway, it's just that some matrices are stored with float64 precision because `np.array()` creates them like that by default.",
              "votes": 2
            },
            {
              "id": 2542782,
              "postDate": "2023-11-29T14:03:01.093Z",
              "content": "<p>Great, thanks!</p>",
              "rawMarkdown": "Great, thanks!"
            }
          ]
        },
        {
          "id": 2544246,
          "postDate": "2023-11-30T17:24:12.820Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>,</p>\n<p>Thank you for your work, and apologies for the delayed response -- I caught a respiratory infection and I've been unable to work for a little bit.</p>\n<p>I'll run some tests with the metric and update if everything looks good. Thanks again.</p>",
          "rawMarkdown": "Hi @optimo,\n\nThank you for your work, and apologies for the delayed response -- I caught a respiratory infection and I've been unable to work for a little bit.\n\nI'll run some tests with the metric and update if everything looks good. Thanks again.",
          "votes": 3,
          "replies": [
            {
              "id": 2544276,
              "postDate": "2023-11-30T17:59:50.057Z",
              "content": "<p>I thought of the possibility of you being sick, hope you’ll recover quickly! Thanks for looking into it!</p>",
              "rawMarkdown": "I thought of the possibility of you being sick, hope you’ll recover quickly! Thanks for looking into it!",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2544771,
      "postDate": "2023-12-01T04:35:51.730Z",
      "content": "<p>the surface dice metric code is very slow to run to search for parameters in local experiments.</p>\n<p>since the  tolerance is zero, i suggest one can use :</p>\n<ol>\n<li>convert volume to surface voxel </li>\n<li>use normal volume dice on truth and predicted surface voxel  to approximate surface dice  in experiments</li>\n</ol>",
      "rawMarkdown": "the surface dice metric code is very slow to run to search for parameters in local experiments.\n\nsince the  tolerance is zero, i suggest one can use :\n1. convert volume to surface voxel \n2. use normal volume dice on truth and predicted surface voxel  to approximate surface dice  in experiments",
      "votes": 1
    },
    {
      "id": 2543696,
      "postDate": "2023-11-30T09:52:48.450Z",
      "content": "<p>Thank you for your work, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> .</p>\n<p>In my tests, I see a 3GB improvement in memory usage and the validation scores differ in 6th digit. I'm hoping it remains the same for multiple thresholds and complexities of predictions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F140209%2F3f256e4c4f44da27884f964b7100c04b%2FScreenshot%202023-11-30%20at%203.21.12%20PM.png?generation=1701337897018946&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you for your work, @optimo .\n\nIn my tests, I see a 3GB improvement in memory usage and the validation scores differ in 6th digit. I'm hoping it remains the same for multiple thresholds and complexities of predictions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F140209%2F3f256e4c4f44da27884f964b7100c04b%2FScreenshot%202023-11-30%20at%203.21.12%20PM.png?generation=1701337897018946&alt=media)\n\n",
      "votes": 2
    },
    {
      "id": 2609540,
      "postDate": "2024-01-19T14:46:22.567Z",
      "content": "<p>Why does my submission kernel still confront with OOM error???</p>",
      "rawMarkdown": "Why does my submission kernel still confront with OOM error???"
    },
    {
      "id": 2533215,
      "postDate": "2023-11-21T17:07:37.483Z",
      "content": "<p>I was having consistent submission failures (~9) that would occur any time I used my model predictions that went away once I replaced my <code>rle_encode/decode</code> functions with the \"proper\" ones. It turned out that I had copied an incorrectly implemented rle encoder/decoder for my own and they were encoding the masks incorrectly. With the proper functions in place, I've been able to drop my score threshold down to 0.05 and haven't encountered any scoring errors.</p>\n<p>I believe I made this fix around the same time the hosts pushed a metric update, so my \"fix\" of replacing my RLE functions and the apparent resolution may be unrelated.</p>",
      "rawMarkdown": "I was having consistent submission failures (~9) that would occur any time I used my model predictions that went away once I replaced my `rle_encode/decode` functions with the \"proper\" ones. It turned out that I had copied an incorrectly implemented rle encoder/decoder for my own and they were encoding the masks incorrectly. With the proper functions in place, I've been able to drop my score threshold down to 0.05 and haven't encountered any scoring errors.\n\nI believe I made this fix around the same time the hosts pushed a metric update, so my \"fix\" of replacing my RLE functions and the apparent resolution may be unrelated.",
      "replies": [
        {
          "id": 2533236,
          "postDate": "2023-11-21T17:24:46.390Z",
          "content": "<p>I currently have a notebook where simply changing the outputs threshold creates OOM or not, so it's not about RLE function.</p>",
          "rawMarkdown": "I currently have a notebook where simply changing the outputs threshold creates OOM or not, so it's not about RLE function.",
          "votes": 1,
          "replies": [
            {
              "id": 2533320,
              "postDate": "2023-11-21T19:19:37.020Z",
              "content": "<p>You would be correct. I just ran in to the OOM error after modifying my model input slightly, which increased sensitivity. I could get it to work increasing my score threshold, but the cost to leaderboard performance isn't worth it. I'll wait until the host sort out the issues with the metric.</p>",
              "rawMarkdown": "You would be correct. I just ran in to the OOM error after modifying my model input slightly, which increased sensitivity. I could get it to work increasing my score threshold, but the cost to leaderboard performance isn't worth it. I'll wait until the host sort out the issues with the metric."
            }
          ]
        }
      ]
    },
    {
      "id": 2533107,
      "postDate": "2023-11-21T15:33:31.477Z",
      "content": "<p>Ok I'll start.</p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> between the current version (22) of the competition metric and the previous one that I copy pasted a week ago or so I see only two differences:</p>\n<ul>\n<li>new version contains new but unused functions <code>compute_surface_overlap_at_tolerance</code>,  <code>compute_robust_hausdorff</code>, <code>compute_average_surface_distance</code> and <code>compute_dice_coefficient</code>.</li>\n<li>new version adds a call to <code>np.nan_to_num</code> to <code>surface_distances[\"distances_gt_to_pred\"]</code>:  this seems to be the only real change to the code. <strong>May it solves the OOM error to simply add <code>copy=False</code>  to <code>np.nan_to_num</code> ?</strong></li>\n</ul>",
      "rawMarkdown": "Ok I'll start.\n\n@ryanholbrook between the current version (22) of the competition metric and the previous one that I copy pasted a week ago or so I see only two differences:\n- new version contains new but unused functions `compute_surface_overlap_at_tolerance`,  `compute_robust_hausdorff`, `compute_average_surface_distance` and `compute_dice_coefficient`.\n- new version adds a call to `np.nan_to_num` to `surface_distances[\"distances_gt_to_pred\"]`:  this seems to be the only real change to the code. **May it solves the OOM error to simply add `copy=False`  to `np.nan_to_num` ?**",
      "replies": [
        {
          "id": 2533213,
          "postDate": "2023-11-21T17:07:03.910Z",
          "content": "<p>Would <code>np.int32</code> be enough in our case in <code>distance_transform_edt</code> function for <code>pts</code> and <code>roots</code> ?</p>",
          "rawMarkdown": "Would `np.int32` be enough in our case in `distance_transform_edt` function for `pts` and `roots` ?",
          "replies": [
            {
              "id": 2533230,
              "postDate": "2023-11-21T17:19:36.007Z",
              "content": "<p>In <code>create_table_neighbour_code_to_surface_area</code> can we set the dtype to <code>np.float32</code> when defining <code>neighbour_code_to_surface_area</code> which will be duplicated by the prediction surface area ?</p>",
              "rawMarkdown": "In `create_table_neighbour_code_to_surface_area` can we set the dtype to `np.float32` when defining `neighbour_code_to_surface_area` which will be duplicated by the prediction surface area ?"
            }
          ]
        },
        {
          "id": 2533221,
          "postDate": "2023-11-21T17:13:33.117Z",
          "content": "<p>Hi Optimo, thank you for the thread. I think you might have the versions switched. Version 22 (the current version) no longer has the <code>np.nan_to_num</code> functions. This is the same as the original metric when the competition launched. It also lacks the unused functions. I removed these just to clean up the code some.</p>\n<p>When I ran a memory profiler on the <a href=\"https://github.com/google-deepmind/surface-distance\" target=\"_blank\">original DeepMind code</a>, the problem was largely with the <code>distance_transform_edt</code> function as implemented in <code>scipy.ndimage</code>. I swapped it for an implementation from <code>imagepy</code> which yielded some improvement. I also rearranged the computation in <code>compute_surface_distances</code> to delete objects once they were no longer needed. My guess is those two places are where to focus, but I'd need to do more profiling.</p>",
          "rawMarkdown": "Hi Optimo, thank you for the thread. I think you might have the versions switched. Version 22 (the current version) no longer has the `np.nan_to_num` functions. This is the same as the original metric when the competition launched. It also lacks the unused functions. I removed these just to clean up the code some.\n\nWhen I ran a memory profiler on the [original DeepMind code](https://github.com/google-deepmind/surface-distance), the problem was largely with the `distance_transform_edt` function as implemented in `scipy.ndimage`. I swapped it for an implementation from `imagepy` which yielded some improvement. I also rearranged the computation in `compute_surface_distances` to delete objects once they were no longer needed. My guess is those two places are where to focus, but I'd need to do more profiling.",
          "votes": 2,
          "replies": [
            {
              "id": 2533233,
              "postDate": "2023-11-21T17:22:52.197Z",
              "content": "<p>Yes you are right sorry I switched the two versions so we can forget about <code>nan_to_num</code> what about <code>create_table_neighbour_code_to_surface_area</code> to <code>np.float32</code> this looks promising to me.</p>",
              "rawMarkdown": "Yes you are right sorry I switched the two versions so we can forget about `nan_to_num` what about `create_table_neighbour_code_to_surface_area` to `np.float32` this looks promising to me."
            },
            {
              "id": 2533240,
              "postDate": "2023-11-21T17:29:16.233Z",
              "content": "<p>I think that might have made the difference, I went from 0.659 to 0.737 because I tried the submission today and did not need to remove small areas anymore, I did a fully successful submit</p>",
              "rawMarkdown": "I think that might have made the difference, I went from 0.659 to 0.737 because I tried the submission today and did not need to remove small areas anymore, I did a fully successful submit"
            },
            {
              "id": 2533250,
              "postDate": "2023-11-21T17:41:10.757Z",
              "content": "<p>I had OOM error 5 hours ago with threshold 0.85 but not with threshold 0.95.</p>",
              "rawMarkdown": "I had OOM error 5 hours ago with threshold 0.85 but not with threshold 0.95."
            },
            {
              "id": 2533489,
              "postDate": "2023-11-22T01:04:29.500Z",
              "content": "<p>i run the  Version 22 of the metric code on my local validation, there is division by zero at:</p>\n<pre><code>def compute_surface_dice_at_tolerance(surface_distances, tolerance_mm):\n       \n    surface_dice = (overlap_gt + overlap_pred) / ((np.sum(surfel_areas_gt) +\n                                                  np.sum(surfel_areas_pred))) #should  eps\n    return surface_dice\n\n\n    (score(\n        =solution_df,\n        =submission_df,\n        row_id_column_name =,\n        rle_column_name =,\n        =0,\n    ))\n</code></pre>",
              "rawMarkdown": "i run the  Version 22 of the metric code on my local validation, there is division by zero at:\n\n```\n\ndef compute_surface_dice_at_tolerance(surface_distances, tolerance_mm):\n       ....\n\tsurface_dice = (overlap_gt + overlap_pred) / ((np.sum(surfel_areas_gt) +\n\t                                              np.sum(surfel_areas_pred))) #should add eps\n\treturn surface_dice\n\n#calling function\n\tprint(score(\n\t\tsolution=solution_df,\n\t\tsubmission=submission_df,\n\t\trow_id_column_name ='id',\n\t\trle_column_name ='rle',\n\t\ttolerance=0,\n\t))\n```",
              "votes": 1
            },
            {
              "id": 2533548,
              "postDate": "2023-11-22T03:15:55.850Z",
              "content": "<p>lowering the threshold from 0.5 to 0.3 had susubmission scoring error</p>",
              "rawMarkdown": "lowering the threshold from 0.5 to 0.3 had susubmission scoring error"
            },
            {
              "id": 2533568,
              "postDate": "2023-11-22T03:46:34.050Z",
              "content": "<p>You may forget to input the group and slice params to score function, the 3d metric function calling should be</p>\n<pre><code>score(sol.copy(), sub.copy(), ,, , , )\n</code></pre>\n<p>Otherwise it is used for 2d metric. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>",
              "rawMarkdown": "You may forget to input the group and slice params to score function, the 3d metric function calling should be\n```python\nscore(sol.copy(), sub.copy(), 'id','rle', 0, 'group', 'slice')\n```\nOtherwise it is used for 2d metric. @hengck23"
            },
            {
              "id": 2533637,
              "postDate": "2023-11-22T05:17:46.767Z",
              "content": "<p><a href=\"https://www.kaggle.com/snorfyang\" target=\"_blank\">@snorfyang</a> <br>\nthanks it works for me</p>",
              "rawMarkdown": "@snorfyang \nthanks it works for me"
            },
            {
              "id": 2535704,
              "postDate": "2023-11-23T14:45:13.847Z",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> any updates on this topic ?</p>",
              "rawMarkdown": "@ryanholbrook any updates on this topic ?"
            }
          ]
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2554559,
      "author_name": "🐢 Jun Koda",
      "author_url": "",
      "post_date": "2023-12-09T07:38:50.897000",
      "content": "<p>As hengck23 said, the code with tolerance 0 can be greatly simplified and reduce memory usage. If I understand correctly, only 2 slices need to be on RAM, not the whole volume, and this is exact, not an approximation.</p>\n<p>My implementation is this: <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">https://www.kaggle.com/code/junkoda/fast-surface-dice-computation</a></p>",
      "votes": 14,
      "replies": [
        {
          "id": 2554827,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2023-12-09T12:46:06.107000",
          "content": "<p>thanks for the implementation.<br>\nSurface dice computation is reduced from 15 min to a few sec for me. </p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> <br>\nyou may want to take a look at this</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 2551160,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2023-12-06T14:26:03.507000",
      "content": "<p>I just updated the metric to include the dtype downcasting suggested by <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>. In my tests, this reduces memory consumption by about 20% while preserving scores out to six decimals. Please let me know if you observe any unexpected changes to submissions scores though.</p>",
      "votes": 9,
      "replies": [
        {
          "id": 2551420,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-12-06T17:27:03.690000",
          "content": "<p>worked on my case! 🎉</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2559212,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-12-12T16:56:22.370000",
              "content": "<p>Actually <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> it worked on my first submission case, now my second attempt with different model also gets OOM.</p>\n<p>If we are not planning to change the competition metric (as discussed <a href=\"https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/461041\" target=\"_blank\">here</a>) I think it is worth using <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s implementation which is both much faster and uses less memory with the same scores on my internal tests.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2567370,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-12-19T15:15:21.083000",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> What kind of tests and code do you need ? So that we can prepare it for you to try and validate when you'll have time after santa's competition? <a href=\"https://www.kaggle.com/junkoda\" target=\"_blank\">@junkoda</a>'s <a href=\"https://www.kaggle.com/code/junkoda/fast-surface-dice-computation\" target=\"_blank\">code</a> looks already pretty good but we can add a few more checks if you want.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2570545,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-12-22T10:22:14.503000",
              "content": "<p>Actually I had a bug in my pipeline, the error was not OOM.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2553075,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-12-08T01:57:12.610000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2544522,
      "author_name": "Ryan Holbrook",
      "author_url": "",
      "post_date": "2023-11-30T22:14:37.667000",
      "content": "<p>Hi everyone,</p>\n<p>I just made a change to metric that I think will allow more submissions to succeed without OOM. The main change is to the <code>_sort_distances_surfels</code> function, which now sorts in place. It seems that some submissions can create a large number of surface areas, which caused this function to allocate a great deal of memory unnecessarily.</p>\n<p>I'm also going to run some tests on the dtype casting suggested by <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>, but wanted to release this first since it didn't require any changes to numerics.</p>\n<p>Thank you everyone for your patience. I know this has been less than ideal. I'm hoping these improvements will make this a less frustrating experience.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2544949,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-12-01T07:39:03.607000",
          "content": "<p>I really appreciate your ongoing efforts. 😃<br>\nHowever, I continue to encounter a submission scoring error…😭</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2544970,
              "author_name": "binga",
              "author_url": "",
              "post_date": "2023-12-01T08:06:44.503000",
              "content": "<p>I see the same.</p>\n<p>I had the same notebook run yesterday and today and I don't see a change in the scoring error. Not sure if my predictions are still poor/complex.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 2545109,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-12-01T09:54:02.610000",
          "content": "<p>I just tried thresholds 0.5 and 0.85 that still get OOM (threshold 0.95 runs ok).<br>\nI think the dtype casting is safe to use -&gt; also note that I delete the original images after the crop in my code since they are not used anymore after that.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2550019,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-12-05T17:38:55.060000",
          "content": "<p>Hello <a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, have you been able to dig more on the dtype casting proposal ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2550036,
              "author_name": "binga",
              "author_url": "",
              "post_date": "2023-12-05T17:51:25.063000",
              "content": "<p>My apologies for piling on this but any update on this issue from the kaggle team would really help. Thank you.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2537642,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-11-25T11:31:29.417000",
      "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> I made a code update proposal here : <a href=\"https://www.kaggle.com/optimo/surface-dice-metric-low-memory/\" target=\"_blank\">https://www.kaggle.com/optimo/surface-dice-metric-low-memory/</a></p>\n<p>It's only minimal changes but it should reduce the bottle neck matrices footprints by two (float64 -&gt; float32). Some explanations are available at the beginning of the notebook.</p>\n<p>Could you have a look and let me know what you think?</p>\n<p>If you are willing to test it on real submissions feel free to have a look at my failed submissions with OOM which should hopefully run fine with the proposed changes.</p>\n<p>Thank you!</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2537692,
          "author_name": "Harshit Sheoran",
          "author_url": "",
          "post_date": "2023-11-25T12:15:37.530000",
          "content": "<p>Really good work, In your experience, is it faster or slower while reducing memory, because if calculations are happening in float32, they can be faster, also lowering it to float16 might also be a viable option as the masks are only 0 and 1 but some calculations can take longer with float16</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2537809,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-25T13:56:23",
              "content": "<p>I did not notice any speed improvement with float32 in my experiments. I tried to switch everything to float16 but ended up with Nan score.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2541485,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-11-28T14:36:12.837000",
          "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> any chance you could have a look at this proposal ? 🙏</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 2542035,
          "author_name": "siwooyong",
          "author_url": "",
          "post_date": "2023-11-29T03:01:15.243000",
          "content": "<p>I ran the code and it significantly reduced the memory usage. Nice work!</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2542690,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-29T12:36:38.590000",
              "content": "<p><a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> maybe ? 🙂 it seems quite straightforward to test the new implementation!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2543563,
              "author_name": "siwooyong",
              "author_url": "",
              "post_date": "2023-11-30T07:17:17.120000",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a>, it appears that the recent 12 submissions are encountering submission scoring errors, potentially due to OOM issues… Even if <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> 's code doesn't resolve the OOM issue, I would greatly appreciate the prompt presentation of alternative solutions.😭</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2542724,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2023-11-29T13:09:17.313000",
          "content": "<p>Hi, Optimo</p>\n<p>Did you try comparing the final score for fp64 and fp32? I wonder what is the scale of differences between the two. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2542739,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-29T13:30:22.257000",
              "content": "<p>I had the exact same score (on the cases I tried), I believe the computation is exactly the same as the final computation is done in float32 anyway, it's just that some matrices are stored with float64 precision because <code>np.array()</code> creates them like that by default.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2542782,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2023-11-29T14:03:01.093000",
              "content": "<p>Great, thanks!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2544246,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-11-30T17:24:12.820000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a>,</p>\n<p>Thank you for your work, and apologies for the delayed response -- I caught a respiratory infection and I've been unable to work for a little bit.</p>\n<p>I'll run some tests with the metric and update if everything looks good. Thanks again.</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2544276,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-30T17:59:50.057000",
              "content": "<p>I thought of the possibility of you being sick, hope you’ll recover quickly! Thanks for looking into it!</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2544771,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2023-12-01T04:35:51.730000",
      "content": "<p>the surface dice metric code is very slow to run to search for parameters in local experiments.</p>\n<p>since the  tolerance is zero, i suggest one can use :</p>\n<ol>\n<li>convert volume to surface voxel </li>\n<li>use normal volume dice on truth and predicted surface voxel  to approximate surface dice  in experiments</li>\n</ol>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2543696,
      "author_name": "binga",
      "author_url": "",
      "post_date": "2023-11-30T09:52:48.450000",
      "content": "<p>Thank you for your work, <a href=\"https://www.kaggle.com/optimo\" target=\"_blank\">@optimo</a> .</p>\n<p>In my tests, I see a 3GB improvement in memory usage and the validation scores differ in 6th digit. I'm hoping it remains the same for multiple thresholds and complexities of predictions.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F140209%2F3f256e4c4f44da27884f964b7100c04b%2FScreenshot%202023-11-30%20at%203.21.12%20PM.png?generation=1701337897018946&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2609540,
      "author_name": "豆柴金鯱",
      "author_url": "",
      "post_date": "2024-01-19T14:46:22.567000",
      "content": "<p>Why does my submission kernel still confront with OOM error???</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2533215,
      "author_name": "kcetskcaz",
      "author_url": "",
      "post_date": "2023-11-21T17:07:37.483000",
      "content": "<p>I was having consistent submission failures (~9) that would occur any time I used my model predictions that went away once I replaced my <code>rle_encode/decode</code> functions with the \"proper\" ones. It turned out that I had copied an incorrectly implemented rle encoder/decoder for my own and they were encoding the masks incorrectly. With the proper functions in place, I've been able to drop my score threshold down to 0.05 and haven't encountered any scoring errors.</p>\n<p>I believe I made this fix around the same time the hosts pushed a metric update, so my \"fix\" of replacing my RLE functions and the apparent resolution may be unrelated.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2533236,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-11-21T17:24:46.390000",
          "content": "<p>I currently have a notebook where simply changing the outputs threshold creates OOM or not, so it's not about RLE function.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2533320,
              "author_name": "kcetskcaz",
              "author_url": "",
              "post_date": "2023-11-21T19:19:37.020000",
              "content": "<p>You would be correct. I just ran in to the OOM error after modifying my model input slightly, which increased sensitivity. I could get it to work increasing my score threshold, but the cost to leaderboard performance isn't worth it. I'll wait until the host sort out the issues with the metric.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2533107,
      "author_name": "Optimo",
      "author_url": "",
      "post_date": "2023-11-21T15:33:31.477000",
      "content": "<p>Ok I'll start.</p>\n<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> between the current version (22) of the competition metric and the previous one that I copy pasted a week ago or so I see only two differences:</p>\n<ul>\n<li>new version contains new but unused functions <code>compute_surface_overlap_at_tolerance</code>,  <code>compute_robust_hausdorff</code>, <code>compute_average_surface_distance</code> and <code>compute_dice_coefficient</code>.</li>\n<li>new version adds a call to <code>np.nan_to_num</code> to <code>surface_distances[\"distances_gt_to_pred\"]</code>:  this seems to be the only real change to the code. <strong>May it solves the OOM error to simply add <code>copy=False</code>  to <code>np.nan_to_num</code> ?</strong></li>\n</ul>",
      "votes": 0,
      "replies": [
        {
          "id": 2533213,
          "author_name": "Optimo",
          "author_url": "",
          "post_date": "2023-11-21T17:07:03.910000",
          "content": "<p>Would <code>np.int32</code> be enough in our case in <code>distance_transform_edt</code> function for <code>pts</code> and <code>roots</code> ?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2533230,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-21T17:19:36.007000",
              "content": "<p>In <code>create_table_neighbour_code_to_surface_area</code> can we set the dtype to <code>np.float32</code> when defining <code>neighbour_code_to_surface_area</code> which will be duplicated by the prediction surface area ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2533221,
          "author_name": "Ryan Holbrook",
          "author_url": "",
          "post_date": "2023-11-21T17:13:33.117000",
          "content": "<p>Hi Optimo, thank you for the thread. I think you might have the versions switched. Version 22 (the current version) no longer has the <code>np.nan_to_num</code> functions. This is the same as the original metric when the competition launched. It also lacks the unused functions. I removed these just to clean up the code some.</p>\n<p>When I ran a memory profiler on the <a href=\"https://github.com/google-deepmind/surface-distance\" target=\"_blank\">original DeepMind code</a>, the problem was largely with the <code>distance_transform_edt</code> function as implemented in <code>scipy.ndimage</code>. I swapped it for an implementation from <code>imagepy</code> which yielded some improvement. I also rearranged the computation in <code>compute_surface_distances</code> to delete objects once they were no longer needed. My guess is those two places are where to focus, but I'd need to do more profiling.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 2533233,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-21T17:22:52.197000",
              "content": "<p>Yes you are right sorry I switched the two versions so we can forget about <code>nan_to_num</code> what about <code>create_table_neighbour_code_to_surface_area</code> to <code>np.float32</code> this looks promising to me.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2533240,
              "author_name": "Harshit Sheoran",
              "author_url": "",
              "post_date": "2023-11-21T17:29:16.233000",
              "content": "<p>I think that might have made the difference, I went from 0.659 to 0.737 because I tried the submission today and did not need to remove small areas anymore, I did a fully successful submit</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2533250,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-21T17:41:10.757000",
              "content": "<p>I had OOM error 5 hours ago with threshold 0.85 but not with threshold 0.95.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2533489,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-11-22T01:04:29.500000",
              "content": "<p>i run the  Version 22 of the metric code on my local validation, there is division by zero at:</p>\n<pre><code>def compute_surface_dice_at_tolerance(surface_distances, tolerance_mm):\n       \n    surface_dice = (overlap_gt + overlap_pred) / ((np.sum(surfel_areas_gt) +\n                                                  np.sum(surfel_areas_pred))) #should  eps\n    return surface_dice\n\n\n    (score(\n        =solution_df,\n        =submission_df,\n        row_id_column_name =,\n        rle_column_name =,\n        =0,\n    ))\n</code></pre>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2533548,
              "author_name": "Arunodhayan",
              "author_url": "",
              "post_date": "2023-11-22T03:15:55.850000",
              "content": "<p>lowering the threshold from 0.5 to 0.3 had susubmission scoring error</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2533568,
              "author_name": "Snorf",
              "author_url": "",
              "post_date": "2023-11-22T03:46:34.050000",
              "content": "<p>You may forget to input the group and slice params to score function, the 3d metric function calling should be</p>\n<pre><code>score(sol.copy(), sub.copy(), ,, , , )\n</code></pre>\n<p>Otherwise it is used for 2d metric. <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2533637,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2023-11-22T05:17:46.767000",
              "content": "<p><a href=\"https://www.kaggle.com/snorfyang\" target=\"_blank\">@snorfyang</a> <br>\nthanks it works for me</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2535704,
              "author_name": "Optimo",
              "author_url": "",
              "post_date": "2023-11-23T14:45:13.847000",
              "content": "<p><a href=\"https://www.kaggle.com/ryanholbrook\" target=\"_blank\">@ryanholbrook</a> any updates on this topic ?</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2533084": "There is currently an OOM error occurring depending on the \"complexity\" of the predictions made.\n\nSee this [notebook](https://www.kaggle.com/code/kashiwaba/sennet-hoa-inference-unet-simple-baseline) or this [discussion](https://www.kaggle.com/competitions/blood-vessel-segmentation/discussion/455220) to get more info.\n\nI'm creating this thread so that we can all share concrete improvements that can be done on the [official metric code](https://www.kaggle.com/code/metric/surface-dice-metric/notebook) in order to reduce first memory usage, and also speed if possible.\n\nA correct metric accepting any submission in correct format will ensure a fair competition and the best possible outcome for kagglers and organizers.\n\nThanks for your help!",
    "2554559": "As hengck23 said, the code with tolerance 0 can be greatly simplified and reduce memory usage. If I understand correctly, only 2 slices need to be on RAM, not the whole volume, and this is exact, not an approximation.\n\nMy implementation is this: https://www.kaggle.com/code/junkoda/fast-surface-dice-computation",
    "2551160": "I just updated the metric to include the dtype downcasting suggested by @optimo. In my tests, this reduces memory consumption by about 20% while preserving scores out to six decimals. Please let me know if you observe any unexpected changes to submissions scores though.",
    "2544522": "Hi everyone,\n\nI just made a change to metric that I think will allow more submissions to succeed without OOM. The main change is to the `_sort_distances_surfels` function, which now sorts in place. It seems that some submissions can create a large number of surface areas, which caused this function to allocate a great deal of memory unnecessarily.\n\nI'm also going to run some tests on the dtype casting suggested by @optimo, but wanted to release this first since it didn't require any changes to numerics.\n\nThank you everyone for your patience. I know this has been less than ideal. I'm hoping these improvements will make this a less frustrating experience.",
    "2537642": "@ryanholbrook I made a code update proposal here : https://www.kaggle.com/optimo/surface-dice-metric-low-memory/\n\nIt's only minimal changes but it should reduce the bottle neck matrices footprints by two (float64 -> float32). Some explanations are available at the beginning of the notebook.\n\nCould you have a look and let me know what you think?\n\nIf you are willing to test it on real submissions feel free to have a look at my failed submissions with OOM which should hopefully run fine with the proposed changes.\n\nThank you!",
    "2544771": "the surface dice metric code is very slow to run to search for parameters in local experiments.\n\nsince the  tolerance is zero, i suggest one can use :\n1. convert volume to surface voxel \n2. use normal volume dice on truth and predicted surface voxel  to approximate surface dice  in experiments",
    "2543696": "Thank you for your work, @optimo .\n\nIn my tests, I see a 3GB improvement in memory usage and the validation scores differ in 6th digit. I'm hoping it remains the same for multiple thresholds and complexities of predictions.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F140209%2F3f256e4c4f44da27884f964b7100c04b%2FScreenshot%202023-11-30%20at%203.21.12%20PM.png?generation=1701337897018946&alt=media)\n\n",
    "2609540": "Why does my submission kernel still confront with OOM error???",
    "2533215": "I was having consistent submission failures (~9) that would occur any time I used my model predictions that went away once I replaced my `rle_encode/decode` functions with the \"proper\" ones. It turned out that I had copied an incorrectly implemented rle encoder/decoder for my own and they were encoding the masks incorrectly. With the proper functions in place, I've been able to drop my score threshold down to 0.05 and haven't encountered any scoring errors.\n\nI believe I made this fix around the same time the hosts pushed a metric update, so my \"fix\" of replacing my RLE functions and the apparent resolution may be unrelated.",
    "2533107": "Ok I'll start.\n\n@ryanholbrook between the current version (22) of the competition metric and the previous one that I copy pasted a week ago or so I see only two differences:\n- new version contains new but unused functions `compute_surface_overlap_at_tolerance`,  `compute_robust_hausdorff`, `compute_average_surface_distance` and `compute_dice_coefficient`.\n- new version adds a call to `np.nan_to_num` to `surface_distances[\"distances_gt_to_pred\"]`:  this seems to be the only real change to the code. **May it solves the OOM error to simply add `copy=False`  to `np.nan_to_num` ?**"
  }
}