{
  "id": 568681,
  "title": "Pre-processing data for segmentation masks",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/568681",
  "author_name": "",
  "post_date": "2025-03-17T10:51:46.924860500Z",
  "votes": 1,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Wondering what the approach is to pre-processing data for segmentation. If we take using binary segmentation masks as an example for the 648 tomograms, is it sensible to preprocess these maps, store them as numpy arrays, and then read + augment upon training? There are some tricks to storing compressed amounts of the masks (i.e. storing only the indexes of the mask pixels) - but I am finding that this is quite a large amount of data to store on file.</p>\n<p>Is this a reasonable approach? I suspect many modelling approaches will be using YOLO where the data storage requirements are far less, but will appreciate any feedback on what's generally accepted :)</p>",
  "messages": [
    {
      "id": "3152034",
      "postDate": "03/17/2025 10:51:46",
      "content": "<p>Wondering what the approach is to pre-processing data for segmentation. If we take using binary segmentation masks as an example for the 648 tomograms, is it sensible to preprocess these maps, store them as numpy arrays, and then read + augment upon training? There are some tricks to storing compressed amounts of the masks (i.e. storing only the indexes of the mask pixels) - but I am finding that this is quite a large amount of data to store on file.</p>\n<p>Is this a reasonable approach? I suspect many modelling approaches will be using YOLO where the data storage requirements are far less, but will appreciate any feedback on what's generally accepted :)</p>",
      "rawMarkdown": "Wondering what the approach is to pre-processing data for segmentation. If we take using binary segmentation masks as an example for the 648 tomograms, is it sensible to preprocess these maps, store them as numpy arrays, and then read + augment upon training? There are some tricks to storing compressed amounts of the masks (i.e. storing only the indexes of the mask pixels) - but I am finding that this is quite a large amount of data to store on file.\n\nIs this a reasonable approach? I suspect many modelling approaches will be using YOLO where the data storage requirements are far less, but will appreciate any feedback on what's generally accepted :)",
      "votes": null
    },
    {
      "id": "3153406",
      "postDate": "03/18/2025 19:08:18",
      "content": "<p>Yes, this is a feasible solution.  You can then use something like cc3d to convert the segmentation masks into points.  Just keep in mind you will likely need to break the 3d volume into smaller voxels for your model to process with some overlay between the voxels.  </p>\n<p>For inference I recommend building the 3d volume at runtime instead of trying to store them because the read and write will be expensive and likely go over the 12 hour limit.  Does this make sense or did I misunderstand the question?</p>",
      "rawMarkdown": "Yes, this is a feasible solution.  You can then use something like cc3d to convert the segmentation masks into points.  Just keep in mind you will likely need to break the 3d volume into smaller voxels for your model to process with some overlay between the voxels.  \n\nFor inference I recommend building the 3d volume at runtime instead of trying to store them because the read and write will be expensive and likely go over the 12 hour limit.  Does this make sense or did I misunderstand the question?",
      "votes": null
    },
    {
      "id": "3153497",
      "postDate": "03/18/2025 22:42:35",
      "content": "<p>Thanks, yes I participated in the previous CryoET competition, inference will be fine because I won't require the segmentation masks. It is more so for training, this competition has 600 more tomograms than last, and I'm wondering how people are pre-processing the segmentation masks (if doing segmentation) when training. If we have say 100 3d masks the storage for this already gets far too large to fit into VRAM, and then even still storing on disk requires a very large amount of storage, the dataset alone is 236GB. So do we pre-process and then store on file and then read in during training - do people just have very large disk spaces? </p>",
      "rawMarkdown": "Thanks, yes I participated in the previous CryoET competition, inference will be fine because I won't require the segmentation masks. It is more so for training, this competition has 600 more tomograms than last, and I'm wondering how people are pre-processing the segmentation masks (if doing segmentation) when training. If we have say 100 3d masks the storage for this already gets far too large to fit into VRAM, and then even still storing on disk requires a very large amount of storage, the dataset alone is 236GB. So do we pre-process and then store on file and then read in during training - do people just have very large disk spaces?",
      "votes": null
    },
    {
      "id": "3153945",
      "postDate": "03/19/2025 11:03:13",
      "content": "<p>If you are using 3D U-Net for training, generating masks one sample at a time in the dataset seems more cost-effective in terms of disk space compared to precomputing all masks directly.</p>",
      "rawMarkdown": "If you are using 3D U-Net for training, generating masks one sample at a time in the dataset seems more cost-effective in terms of disk space compared to precomputing all masks directly.",
      "votes": null
    },
    {
      "id": "3154433",
      "postDate": "03/19/2025 23:35:51",
      "content": "<p>Yes, though it is not super fast to generate masks, and this will lead to slower training. Maybe that is the best option given constraints on disk space.</p>",
      "rawMarkdown": "Yes, though it is not super fast to generate masks, and this will lead to slower training. Maybe that is the best option given constraints on disk space.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3153406,
      "author_name": "connorjd",
      "author_url": "",
      "post_date": "03/18/2025 19:08:18",
      "content": "<p>Yes, this is a feasible solution.  You can then use something like cc3d to convert the segmentation masks into points.  Just keep in mind you will likely need to break the 3d volume into smaller voxels for your model to process with some overlay between the voxels.  </p>\n<p>For inference I recommend building the 3d volume at runtime instead of trying to store them because the read and write will be expensive and likely go over the 12 hour limit.  Does this make sense or did I misunderstand the question?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3153497,
          "author_name": "homiecal",
          "author_url": "",
          "post_date": "03/18/2025 22:42:35",
          "content": "<p>Thanks, yes I participated in the previous CryoET competition, inference will be fine because I won't require the segmentation masks. It is more so for training, this competition has 600 more tomograms than last, and I'm wondering how people are pre-processing the segmentation masks (if doing segmentation) when training. If we have say 100 3d masks the storage for this already gets far too large to fit into VRAM, and then even still storing on disk requires a very large amount of storage, the dataset alone is 236GB. So do we pre-process and then store on file and then read in during training - do people just have very large disk spaces? </p>",
          "votes": null,
          "replies": [
            {
              "id": 3153945,
              "author_name": "switch9527",
              "author_url": "",
              "post_date": "03/19/2025 11:03:13",
              "content": "<p>If you are using 3D U-Net for training, generating masks one sample at a time in the dataset seems more cost-effective in terms of disk space compared to precomputing all masks directly.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3154433,
                  "author_name": "homiecal",
                  "author_url": "",
                  "post_date": "03/19/2025 23:35:51",
                  "content": "<p>Yes, though it is not super fast to generate masks, and this will lead to slower training. Maybe that is the best option given constraints on disk space.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3152034": "Wondering what the approach is to pre-processing data for segmentation. If we take using binary segmentation masks as an example for the 648 tomograms, is it sensible to preprocess these maps, store them as numpy arrays, and then read + augment upon training? There are some tricks to storing compressed amounts of the masks (i.e. storing only the indexes of the mask pixels) - but I am finding that this is quite a large amount of data to store on file.\n\nIs this a reasonable approach? I suspect many modelling approaches will be using YOLO where the data storage requirements are far less, but will appreciate any feedback on what's generally accepted :)",
    "3153406": "Yes, this is a feasible solution.  You can then use something like cc3d to convert the segmentation masks into points.  Just keep in mind you will likely need to break the 3d volume into smaller voxels for your model to process with some overlay between the voxels.  \n\nFor inference I recommend building the 3d volume at runtime instead of trying to store them because the read and write will be expensive and likely go over the 12 hour limit.  Does this make sense or did I misunderstand the question?",
    "3153497": "Thanks, yes I participated in the previous CryoET competition, inference will be fine because I won't require the segmentation masks. It is more so for training, this competition has 600 more tomograms than last, and I'm wondering how people are pre-processing the segmentation masks (if doing segmentation) when training. If we have say 100 3d masks the storage for this already gets far too large to fit into VRAM, and then even still storing on disk requires a very large amount of storage, the dataset alone is 236GB. So do we pre-process and then store on file and then read in during training - do people just have very large disk spaces?",
    "3153945": "If you are using 3D U-Net for training, generating masks one sample at a time in the dataset seems more cost-effective in terms of disk space compared to precomputing all masks directly.",
    "3154433": "Yes, though it is not super fast to generate masks, and this will lead to slower training. Maybe that is the best option given constraints on disk space."
  },
  "source": "meta"
}