{
  "id": 547044,
  "title": "Very long submission time. 6+ hours and still scoring",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/547044",
  "author_name": "DennisSakva",
  "post_date": "2024-11-19T13:59:02.003000",
  "votes": 7,
  "comment_count": 23,
  "views": 0,
  "content": "<p>How large is the test dataset? My submission script takes 15 minutes on the \"visible\" test dataset with three tomograms but it's already been six hours and my submission is still not finished. I wonder if that is due to some bug or if the hidden test set is so much larger (More than 70 tomograms).</p>",
  "messages": [
    {
      "id": 3049776,
      "postDate": "2024-11-19T13:59:02.003Z",
      "content": "<p>How large is the test dataset? My submission script takes 15 minutes on the \"visible\" test dataset with three tomograms but it's already been six hours and my submission is still not finished. I wonder if that is due to some bug or if the hidden test set is so much larger (More than 70 tomograms).</p>",
      "rawMarkdown": "How large is the test dataset? My submission script takes 15 minutes on the \"visible\" test dataset with three tomograms but it's already been six hours and my submission is still not finished. I wonder if that is due to some bug or if the hidden test set is so much larger (More than 70 tomograms).",
      "votes": 7
    },
    {
      "id": 3052405,
      "postDate": "2024-11-22T12:10:15.913Z",
      "content": "<p>seems that chatgpt offer another solution without CCL. Use pytorch distance transform:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb45e4d2e7577cc28d56107a3be9ea5f%2FSelection_721.png?generation=1732277368555041&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbcdaadebd6e89e7348a175a75e4a999b%2FSelection_723.png?generation=1732277379369178&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "seems that chatgpt offer another solution without CCL. Use pytorch distance transform:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb45e4d2e7577cc28d56107a3be9ea5f%2FSelection_721.png?generation=1732277368555041&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbcdaadebd6e89e7348a175a75e4a999b%2FSelection_723.png?generation=1732277379369178&alt=media)",
      "votes": 1
    },
    {
      "id": 3052394,
      "postDate": "2024-11-22T11:47:48.597Z",
      "content": "<p>one trick to speedup connected component labeling (CCL) is to make your predicted object (threshold binary sphere) small, i.e. smaller radius. there is less +ve pixel.</p>\n<p>alternatively,use scale =0 for small particles, and scale=1  for larger particles (just resize from 0 to 1 for CCL. for segmentation we can still use scale 0 for all)</p>\n<p>now segmentation networks typically take about 10 sec to predict per pixel label.<br>\nspeedup GPU CL processing is about 3 sec.</p>\n<p>to process all 500 hidden test tomography,that will be about 2hr</p>",
      "rawMarkdown": "one trick to speedup connected component labeling (CCL) is to make your predicted object (threshold binary sphere) small, i.e. smaller radius. there is less +ve pixel.\n\nalternatively,use scale =0 for small particles, and scale=1  for larger particles (just resize from 0 to 1 for CCL. for segmentation we can still use scale 0 for all)\n\nnow segmentation networks typically take about 10 sec to predict per pixel label.\nspeedup GPU CL processing is about 3 sec.\n\nto process all 500 hidden test tomography,that will be about 2hr",
      "votes": 1,
      "replies": [
        {
          "id": 3052845,
          "postDate": "2024-11-22T22:00:38.973Z",
          "content": "<p><a href=\"https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\" target=\"_blank\">https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta</a></p>\n<p>4x speedup and reduce 5hr to 1~1.5hr for submission by post-processing at 0.5 scale for connected component labeling.<br>\nIt loses about 0.003 accuracy.</p>\n<p>one can use it for fast development and revert to the original version for final submission.</p>",
          "rawMarkdown": "https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\n\n4x speedup and reduce 5hr to 1~1.5hr for submission by post-processing at 0.5 scale for connected component labeling.\nIt loses about 0.003 accuracy.\n\none can use it for fast development and revert to the original version for final submission.",
          "votes": 2
        }
      ]
    },
    {
      "id": 3052382,
      "postDate": "2024-11-22T11:40:25.773Z",
      "content": "<p>if that was not fast enough, then ….</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03cbf6b7524d93bfa4b2de2eb0605d3a%2FSelection_720.png?generation=1732275583993758&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75501691217b805a99f99dc4312a039c%2FSelection_719.png?generation=1732275596479102&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "if that was not fast enough, then ....\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03cbf6b7524d93bfa4b2de2eb0605d3a%2FSelection_720.png?generation=1732275583993758&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75501691217b805a99f99dc4312a039c%2FSelection_719.png?generation=1732275596479102&alt=media)\n",
      "votes": 1
    },
    {
      "id": 3049892,
      "postDate": "2024-11-19T15:38:06.110Z",
      "content": "<p>It is about 500 hidden test tomographs. mine takes about 4hr to score. you do not need cocomplicated network</p>",
      "rawMarkdown": "It is about 500 hidden test tomographs. mine takes about 4hr to score. you do not need cocomplicated network",
      "votes": 1,
      "replies": [
        {
          "id": 3050407,
          "postDate": "2024-11-20T07:10:59.280Z",
          "content": "<p>Nice! I have a pretty small model. It's the postprocessing part to extract centroids that takes most of the time.</p>",
          "rawMarkdown": "Nice! I have a pretty small model. It's the postprocessing part to extract centroids that takes most of the time.",
          "replies": [
            {
              "id": 3050435,
              "postDate": "2024-11-20T07:56:16.560Z",
              "content": "<p>what post- processing are you using?</p>\n<p>the fasting is using connected component + centroid<br>\n(14 sec for one tomo for one cpu. you can break tomo into 4 parts and run at 4 threads or compile a multi-thread c code)<br>\n<a href=\"https://www.kaggle.com/code/hengck23/speed-up-connected-component-analysis-with-pytorch\" target=\"_blank\">https://www.kaggle.com/code/hengck23/speed-up-connected-component-analysis-with-pytorch</a></p>",
              "rawMarkdown": "what post- processing are you using?\n\nthe fasting is using connected component + centroid\n(14 sec for one tomo for one cpu. you can break tomo into 4 parts and run at 4 threads or compile a multi-thread c code)\nhttps://www.kaggle.com/code/hengck23/speed-up-connected-component-analysis-with-pytorch",
              "votes": 3
            },
            {
              "id": 3050449,
              "postDate": "2024-11-20T08:12:34.950Z",
              "content": "<p><a href=\"https://github.com/DanuserLab/u-segment3D\" target=\"_blank\">https://github.com/DanuserLab/u-segment3D</a></p>\n<p>A general algorithm for consensus 3D cell segmentation from 2D segmented stacks <br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2024.05.03.592249v2.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.05.03.592249v2.full.pdf</a></p>\n<p>i think you can find other parallel implementation. i haven't got time to check in details yet</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7bb3a78b8021255c284fbf9383a767d%2FSelection_999(6864).png?generation=1732090310624912&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2e48c3bf8bff69b576e905f53744ac7f%2FSelection_999(6863).png?generation=1732090319628289&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><a href=\"https://wrfranklin.org/nikola/pages/connect/\" target=\"_blank\">https://wrfranklin.org/nikola/pages/connect/</a> <br>\n<a href=\"https://wrfranklin.org/pmwiki/pmwiki.php/Research/ConnectedComponents\" target=\"_blank\">https://wrfranklin.org/pmwiki/pmwiki.php/Research/ConnectedComponents</a></p>\n<p>1024x1088x1088  <br>\n2GHz IBM T43p laptop   <br>\n51 CPU seconds  </p>",
              "rawMarkdown": "https://github.com/DanuserLab/u-segment3D\n\nA general algorithm for consensus 3D cell segmentation from 2D segmented stacks \nhttps://www.biorxiv.org/content/10.1101/2024.05.03.592249v2.full.pdf\n\n\ni think you can find other parallel implementation. i haven't got time to check in details yet\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7bb3a78b8021255c284fbf9383a767d%2FSelection_999(6864).png?generation=1732090310624912&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2e48c3bf8bff69b576e905f53744ac7f%2FSelection_999(6863).png?generation=1732090319628289&alt=media)\n\n---\nhttps://wrfranklin.org/nikola/pages/connect/ \nhttps://wrfranklin.org/pmwiki/pmwiki.php/Research/ConnectedComponents\n\n1024x1088x1088  \n2GHz IBM T43p laptop   \n51 CPU seconds  ",
              "votes": 2
            },
            {
              "id": 3050461,
              "postDate": "2024-11-20T08:25:17.287Z",
              "content": "<p>The longest part for now is a watershed algorithm to split touching components. Will probably remove it and see how it impacts the final score.</p>",
              "rawMarkdown": "The longest part for now is a watershed algorithm to split touching components. Will probably remove it and see how it impacts the final score."
            },
            {
              "id": 3050469,
              "postDate": "2024-11-20T08:33:21.050Z",
              "content": "<p>i think watershed is not required:</p>\n<ul>\n<li>make groundtruth by taking t*radius sphere (where t is between and 0 and 1, it can be different values for different particles)</li>\n<li>check if sphere are touching each other (or check the distance of nearest neighbour)</li>\n<li>confirm they are not touching each other and feed this ground truth to your post processing algorithm to see if you can recover your locations xyz from these spheres.</li>\n<li>if the answer is yes, then you can set these spheres as segmenation targets</li>\n</ul>\n<p>this explain why the particles don't touch each other in 2d or 3d or in physical setup:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547091\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547091</a></p>",
              "rawMarkdown": "i think watershed is not required:\n- make groundtruth by taking t*radius sphere (where t is between and 0 and 1, it can be different values for different particles)\n- check if sphere are touching each other (or check the distance of nearest neighbour)\n- confirm they are not touching each other and feed this ground truth to your post processing algorithm to see if you can recover your locations xyz from these spheres.\n- if the answer is yes, then you can set these spheres as segmenation targets\n\nthis explain why the particles don't touch each other in 2d or 3d or in physical setup:\nhttps://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547091\n"
            },
            {
              "id": 3051046,
              "postDate": "2024-11-20T19:17:06.180Z",
              "content": "<p>I am curious about how much speedup you were able to get with the gpu-based connected components algorithm. Would it be possible to share some numbers? </p>",
              "rawMarkdown": "I am curious about how much speedup you were able to get with the gpu-based connected components algorithm. Would it be possible to share some numbers? ",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3049825,
      "postDate": "2024-11-19T14:48:30.190Z",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544867\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544867</a> Apparently 500</p>",
      "rawMarkdown": "https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544867 Apparently 500",
      "votes": 1,
      "replies": [
        {
          "id": 3049834,
          "postDate": "2024-11-19T14:55:01.943Z",
          "content": "<p>Holly molly, that leaves us with less than 90 seconds per image. And it means that the test 3 images should be finished in less than 5 minutes.</p>",
          "rawMarkdown": "Holly molly, that leaves us with less than 90 seconds per image. And it means that the test 3 images should be finished in less than 5 minutes.",
          "votes": 1,
          "replies": [
            {
              "id": 3049840,
              "postDate": "2024-11-19T14:58:48.593Z",
              "content": "<p>You're right. 500 images.<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702#3038910\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702#3038910</a><br>\nWill probably have to use multiprocessing and the like to the max to fit in the allowed timeframe.</p>",
              "rawMarkdown": "You're right. 500 images.\nhttps://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702#3038910\nWill probably have to use multiprocessing and the like to the max to fit in the allowed timeframe.",
              "votes": 2
            }
          ]
        },
        {
          "id": 3050089,
          "postDate": "2024-11-19T20:03:31.203Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true,
          "replies": [
            {
              "id": 3050094,
              "postDate": "2024-11-19T20:12:08.427Z",
              "content": "<p>All test is processed and scored at submission. Although only public score is shown.</p>",
              "rawMarkdown": "All test is processed and scored at submission. Although only public score is shown.",
              "votes": 1
            },
            {
              "id": 3050827,
              "postDate": "2024-11-20T15:22:26.830Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 3050929,
              "postDate": "2024-11-20T16:51:48.447Z",
              "content": "<p>Again. No, they've already calculated, simply not shown yet.</p>",
              "rawMarkdown": "Again. No, they've already calculated, simply not shown yet.",
              "votes": 1
            },
            {
              "id": 3050980,
              "postDate": "2024-11-20T17:36:29.470Z",
              "content": "<p>\"Again. No, they've already calculated, simply not shown yet.\"<br>\nThis statement is correct.</p>",
              "rawMarkdown": "\"Again. No, they've already calculated, simply not shown yet.\"\nThis statement is correct.",
              "votes": 4
            }
          ]
        }
      ]
    },
    {
      "id": 3051874,
      "postDate": "2024-11-21T18:11:20.557Z",
      "content": "<p>hey man do you still have this problem and if you dont how can you solve it</p>",
      "rawMarkdown": "hey man do you still have this problem and if you dont how can you solve it"
    },
    {
      "id": 3049814,
      "postDate": "2024-11-19T14:38:08.323Z",
      "content": "<p>My visible test takes also around 15 minuts. About to run my first submission. I'll let you know when it's finished. If it finish.</p>\n<p>EDIT: I'm currently running 5 2D model, 7 fold and 4 rot90 inference because I though it was fast enough.</p>",
      "rawMarkdown": "My visible test takes also around 15 minuts. About to run my first submission. I'll let you know when it's finished. If it finish.\n\nEDIT: I'm currently running 5 2D model, 7 fold and 4 rot90 inference because I though it was fast enough.",
      "replies": [
        {
          "id": 3049820,
          "postDate": "2024-11-19T14:43:53.707Z",
          "content": "<p>That's a lot of models :)</p>",
          "rawMarkdown": "That's a lot of models :)"
        },
        {
          "id": 3049895,
          "postDate": "2024-11-19T15:39:28.773Z",
          "content": "<p>multiple 2d encode and a few 3d/2d encoder</p>",
          "rawMarkdown": "multiple 2d encode and a few 3d/2d encoder"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3052405,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-22T12:10:15.913000",
      "content": "<p>seems that chatgpt offer another solution without CCL. Use pytorch distance transform:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb45e4d2e7577cc28d56107a3be9ea5f%2FSelection_721.png?generation=1732277368555041&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbcdaadebd6e89e7348a175a75e4a999b%2FSelection_723.png?generation=1732277379369178&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3052394,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-22T11:47:48.597000",
      "content": "<p>one trick to speedup connected component labeling (CCL) is to make your predicted object (threshold binary sphere) small, i.e. smaller radius. there is less +ve pixel.</p>\n<p>alternatively,use scale =0 for small particles, and scale=1  for larger particles (just resize from 0 to 1 for CCL. for segmentation we can still use scale 0 for all)</p>\n<p>now segmentation networks typically take about 10 sec to predict per pixel label.<br>\nspeedup GPU CL processing is about 3 sec.</p>\n<p>to process all 500 hidden test tomography,that will be about 2hr</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3052845,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-22T22:00:38.973000",
          "content": "<p><a href=\"https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta\" target=\"_blank\">https://www.kaggle.com/code/hengck23/1-hr-fast-2d-3d-unet-resnet34d-scanner-tta</a></p>\n<p>4x speedup and reduce 5hr to 1~1.5hr for submission by post-processing at 0.5 scale for connected component labeling.<br>\nIt loses about 0.003 accuracy.</p>\n<p>one can use it for fast development and revert to the original version for final submission.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3052382,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-22T11:40:25.773000",
      "content": "<p>if that was not fast enough, then ….</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03cbf6b7524d93bfa4b2de2eb0605d3a%2FSelection_720.png?generation=1732275583993758&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75501691217b805a99f99dc4312a039c%2FSelection_719.png?generation=1732275596479102&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3049892,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-19T15:38:06.110000",
      "content": "<p>It is about 500 hidden test tomographs. mine takes about 4hr to score. you do not need cocomplicated network</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3050407,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-11-20T07:10:59.280000",
          "content": "<p>Nice! I have a pretty small model. It's the postprocessing part to extract centroids that takes most of the time.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3050435,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-20T07:56:16.560000",
              "content": "<p>what post- processing are you using?</p>\n<p>the fasting is using connected component + centroid<br>\n(14 sec for one tomo for one cpu. you can break tomo into 4 parts and run at 4 threads or compile a multi-thread c code)<br>\n<a href=\"https://www.kaggle.com/code/hengck23/speed-up-connected-component-analysis-with-pytorch\" target=\"_blank\">https://www.kaggle.com/code/hengck23/speed-up-connected-component-analysis-with-pytorch</a></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3050449,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-20T08:12:34.950000",
              "content": "<p><a href=\"https://github.com/DanuserLab/u-segment3D\" target=\"_blank\">https://github.com/DanuserLab/u-segment3D</a></p>\n<p>A general algorithm for consensus 3D cell segmentation from 2D segmented stacks <br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2024.05.03.592249v2.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.05.03.592249v2.full.pdf</a></p>\n<p>i think you can find other parallel implementation. i haven't got time to check in details yet</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Ff7bb3a78b8021255c284fbf9383a767d%2FSelection_999(6864).png?generation=1732090310624912&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2e48c3bf8bff69b576e905f53744ac7f%2FSelection_999(6863).png?generation=1732090319628289&amp;alt=media\" alt=\"\"></p>\n<hr>\n<p><a href=\"https://wrfranklin.org/nikola/pages/connect/\" target=\"_blank\">https://wrfranklin.org/nikola/pages/connect/</a> <br>\n<a href=\"https://wrfranklin.org/pmwiki/pmwiki.php/Research/ConnectedComponents\" target=\"_blank\">https://wrfranklin.org/pmwiki/pmwiki.php/Research/ConnectedComponents</a></p>\n<p>1024x1088x1088  <br>\n2GHz IBM T43p laptop   <br>\n51 CPU seconds  </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3050461,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-11-20T08:25:17.287000",
              "content": "<p>The longest part for now is a watershed algorithm to split touching components. Will probably remove it and see how it impacts the final score.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050469,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-20T08:33:21.050000",
              "content": "<p>i think watershed is not required:</p>\n<ul>\n<li>make groundtruth by taking t*radius sphere (where t is between and 0 and 1, it can be different values for different particles)</li>\n<li>check if sphere are touching each other (or check the distance of nearest neighbour)</li>\n<li>confirm they are not touching each other and feed this ground truth to your post processing algorithm to see if you can recover your locations xyz from these spheres.</li>\n<li>if the answer is yes, then you can set these spheres as segmenation targets</li>\n</ul>\n<p>this explain why the particles don't touch each other in 2d or 3d or in physical setup:<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547091\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/547091</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3051046,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-20T19:17:06.180000",
              "content": "<p>I am curious about how much speedup you were able to get with the gpu-based connected components algorithm. Would it be possible to share some numbers? </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3049825,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-11-19T14:48:30.190000",
      "content": "<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544867\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544867</a> Apparently 500</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3049834,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-11-19T14:55:01.943000",
          "content": "<p>Holly molly, that leaves us with less than 90 seconds per image. And it means that the test 3 images should be finished in less than 5 minutes.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3049840,
              "author_name": "DennisSakva",
              "author_url": "",
              "post_date": "2024-11-19T14:58:48.593000",
              "content": "<p>You're right. 500 images.<br>\n<a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702#3038910\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702#3038910</a><br>\nWill probably have to use multiprocessing and the like to the max to fit in the allowed timeframe.</p>",
              "votes": 2,
              "replies": []
            }
          ]
        },
        {
          "id": 3050089,
          "author_name": "",
          "author_url": "",
          "post_date": "2024-11-19T20:03:31.203000",
          "content": "",
          "votes": -1,
          "replies": [
            {
              "id": 3050094,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-19T20:12:08.427000",
              "content": "<p>All test is processed and scored at submission. Although only public score is shown.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050827,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-20T15:22:26.830000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050929,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-20T16:51:48.447000",
              "content": "<p>Again. No, they've already calculated, simply not shown yet.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3050980,
              "author_name": "Reza Paraan",
              "author_url": "",
              "post_date": "2024-11-20T17:36:29.470000",
              "content": "<p>\"Again. No, they've already calculated, simply not shown yet.\"<br>\nThis statement is correct.</p>",
              "votes": 4,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3051874,
      "author_name": "ulasdesouza",
      "author_url": "",
      "post_date": "2024-11-21T18:11:20.557000",
      "content": "<p>hey man do you still have this problem and if you dont how can you solve it</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3049814,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2024-11-19T14:38:08.323000",
      "content": "<p>My visible test takes also around 15 minuts. About to run my first submission. I'll let you know when it's finished. If it finish.</p>\n<p>EDIT: I'm currently running 5 2D model, 7 fold and 4 rot90 inference because I though it was fast enough.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3049820,
          "author_name": "DennisSakva",
          "author_url": "",
          "post_date": "2024-11-19T14:43:53.707000",
          "content": "<p>That's a lot of models :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3049895,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-19T15:39:28.773000",
          "content": "<p>multiple 2d encode and a few 3d/2d encoder</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3049776": "How large is the test dataset? My submission script takes 15 minutes on the \"visible\" test dataset with three tomograms but it's already been six hours and my submission is still not finished. I wonder if that is due to some bug or if the hidden test set is so much larger (More than 70 tomograms).",
    "3052405": "seems that chatgpt offer another solution without CCL. Use pytorch distance transform:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbb45e4d2e7577cc28d56107a3be9ea5f%2FSelection_721.png?generation=1732277368555041&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbcdaadebd6e89e7348a175a75e4a999b%2FSelection_723.png?generation=1732277379369178&alt=media)",
    "3052394": "one trick to speedup connected component labeling (CCL) is to make your predicted object (threshold binary sphere) small, i.e. smaller radius. there is less +ve pixel.\n\nalternatively,use scale =0 for small particles, and scale=1  for larger particles (just resize from 0 to 1 for CCL. for segmentation we can still use scale 0 for all)\n\nnow segmentation networks typically take about 10 sec to predict per pixel label.\nspeedup GPU CL processing is about 3 sec.\n\nto process all 500 hidden test tomography,that will be about 2hr",
    "3052382": "if that was not fast enough, then ....\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03cbf6b7524d93bfa4b2de2eb0605d3a%2FSelection_720.png?generation=1732275583993758&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F75501691217b805a99f99dc4312a039c%2FSelection_719.png?generation=1732275596479102&alt=media)\n",
    "3049892": "It is about 500 hidden test tomographs. mine takes about 4hr to score. you do not need cocomplicated network",
    "3049825": "https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544867 Apparently 500",
    "3051874": "hey man do you still have this problem and if you dont how can you solve it",
    "3049814": "My visible test takes also around 15 minuts. About to run my first submission. I'll let you know when it's finished. If it finish.\n\nEDIT: I'm currently running 5 2D model, 7 fold and 4 rot90 inference because I though it was fast enough."
  }
}