{
  "id": 569996,
  "title": "Best preprocessing approach",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/569996",
  "author_name": "",
  "post_date": "2025-03-25T11:19:56.680320100Z",
  "votes": 13,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I haven’t started working on this competition yet, but I’ve been thinking about it a lot, and I’m completely unsure about the best way to approach data preprocessing, especially for models like U-net. Obviously, the best method will likely be determined through experimentation, but some reasoning can still be applied, and I wanted to share my thoughts with you. I’d be happy to discuss your opinions in the comments.  </p>\n<p>The first obvious approach is not to resize the original data in any way and simply split the train and test objects into overlapping patches and train on them. The clear issues here are too many zero-class labels and an excessively large volume of data, making it difficult and time-consuming to train the model without access to more powerful GPUs than those provided on Kaggle.  </p>\n<p>The other methods involve either resizing the objects, cropping them, or both, followed by splitting into patches and so on.  </p>\n<ol>\n<li><p><strong>Taking a certain number of slices above and below the motor slice for training and rescaling the slices to the desired size</strong>  <br>\nBut then it’s unclear what to do with tomograms that don’t contain a motor to keep the dataset representative. Additionally, for the test set, you can’t just take a chunk—you’d have to process the entire tomogram, which is resource-intensive and might negatively impact model performance.  </p></li>\n<li><p><strong>Uniformly sampling a certain number of slices across the entire tomogram and rescaling all three dimensions.</strong>  <br>\nThe problem here is that the motor might not clearly appear in the selected slice, especially when there are multiple motors. And we are also straightly discard some data, losing information. However, the advantage of this method is its universality compared to method 1 and its efficiency compared to the baseline approach where raw, unprocessed data is used.  </p></li>\n<li><p><strong>Rescaling all data to a higher voxel spacing to reduce the size of the input data.</strong>  <br>\nThe obvious issue is that we don’t know the voxel spacing for the test data, so it’s unclear how to rescale them. One option is to rescale everything to the average voxel spacing of the training set, but that doesn’t sound reliable. The advantage, though, is that we aren’t directly discarding data, as is done in approach 2.  </p></li>\n<li><p><strong>Extending the idea from approach 3, we could ignore voxel spacing altogether and simply interpolate everything to a fixed size.</strong>  <br>\nThe problem here is that the tomograms are initially very different in size, so they might end up overly compressed. But, like approach 3, we aren’t directly discarding data as in approach 2. This method is also as universal as approach 2.  </p></li>\n</ol>\n<p>These are more or less all the methods I’ve seen on the forum or thought about myself. I haven’t yet considered how these methods might affect target transformation. But the most promising approaches seem to be approaches 2 and 4.</p>",
  "messages": [
    {
      "id": "3159189",
      "postDate": "03/25/2025 11:19:56",
      "content": "<p>I haven’t started working on this competition yet, but I’ve been thinking about it a lot, and I’m completely unsure about the best way to approach data preprocessing, especially for models like U-net. Obviously, the best method will likely be determined through experimentation, but some reasoning can still be applied, and I wanted to share my thoughts with you. I’d be happy to discuss your opinions in the comments.  </p>\n<p>The first obvious approach is not to resize the original data in any way and simply split the train and test objects into overlapping patches and train on them. The clear issues here are too many zero-class labels and an excessively large volume of data, making it difficult and time-consuming to train the model without access to more powerful GPUs than those provided on Kaggle.  </p>\n<p>The other methods involve either resizing the objects, cropping them, or both, followed by splitting into patches and so on.  </p>\n<ol>\n<li><p><strong>Taking a certain number of slices above and below the motor slice for training and rescaling the slices to the desired size</strong>  <br>\nBut then it’s unclear what to do with tomograms that don’t contain a motor to keep the dataset representative. Additionally, for the test set, you can’t just take a chunk—you’d have to process the entire tomogram, which is resource-intensive and might negatively impact model performance.  </p></li>\n<li><p><strong>Uniformly sampling a certain number of slices across the entire tomogram and rescaling all three dimensions.</strong>  <br>\nThe problem here is that the motor might not clearly appear in the selected slice, especially when there are multiple motors. And we are also straightly discard some data, losing information. However, the advantage of this method is its universality compared to method 1 and its efficiency compared to the baseline approach where raw, unprocessed data is used.  </p></li>\n<li><p><strong>Rescaling all data to a higher voxel spacing to reduce the size of the input data.</strong>  <br>\nThe obvious issue is that we don’t know the voxel spacing for the test data, so it’s unclear how to rescale them. One option is to rescale everything to the average voxel spacing of the training set, but that doesn’t sound reliable. The advantage, though, is that we aren’t directly discarding data, as is done in approach 2.  </p></li>\n<li><p><strong>Extending the idea from approach 3, we could ignore voxel spacing altogether and simply interpolate everything to a fixed size.</strong>  <br>\nThe problem here is that the tomograms are initially very different in size, so they might end up overly compressed. But, like approach 3, we aren’t directly discarding data as in approach 2. This method is also as universal as approach 2.  </p></li>\n</ol>\n<p>These are more or less all the methods I’ve seen on the forum or thought about myself. I haven’t yet considered how these methods might affect target transformation. But the most promising approaches seem to be approaches 2 and 4.</p>",
      "rawMarkdown": "I haven’t started working on this competition yet, but I’ve been thinking about it a lot, and I’m completely unsure about the best way to approach data preprocessing, especially for models like U-net. Obviously, the best method will likely be determined through experimentation, but some reasoning can still be applied, and I wanted to share my thoughts with you. I’d be happy to discuss your opinions in the comments.  \n\nThe first obvious approach is not to resize the original data in any way and simply split the train and test objects into overlapping patches and train on them. The clear issues here are too many zero-class labels and an excessively large volume of data, making it difficult and time-consuming to train the model without access to more powerful GPUs than those provided on Kaggle.  \n\nThe other methods involve either resizing the objects, cropping them, or both, followed by splitting into patches and so on.  \n\n1. **Taking a certain number of slices above and below the motor slice for training and rescaling the slices to the desired size**  \n   But then it’s unclear what to do with tomograms that don’t contain a motor to keep the dataset representative. Additionally, for the test set, you can’t just take a chunk—you’d have to process the entire tomogram, which is resource-intensive and might negatively impact model performance.  \n\n2. **Uniformly sampling a certain number of slices across the entire tomogram and rescaling all three dimensions.**  \n   The problem here is that the motor might not clearly appear in the selected slice, especially when there are multiple motors. And we are also straightly discard some data, losing information. However, the advantage of this method is its universality compared to method 1 and its efficiency compared to the baseline approach where raw, unprocessed data is used.  \n\n3. **Rescaling all data to a higher voxel spacing to reduce the size of the input data.**  \n   The obvious issue is that we don’t know the voxel spacing for the test data, so it’s unclear how to rescale them. One option is to rescale everything to the average voxel spacing of the training set, but that doesn’t sound reliable. The advantage, though, is that we aren’t directly discarding data, as is done in approach 2.  \n\n4. **Extending the idea from approach 3, we could ignore voxel spacing altogether and simply interpolate everything to a fixed size.**  \n   The problem here is that the tomograms are initially very different in size, so they might end up overly compressed. But, like approach 3, we aren’t directly discarding data as in approach 2. This method is also as universal as approach 2.  \n\nThese are more or less all the methods I’ve seen on the forum or thought about myself. I haven’t yet considered how these methods might affect target transformation. But the most promising approaches seem to be approaches 2 and 4.",
      "votes": null
    },
    {
      "id": "3163186",
      "postDate": "03/30/2025 12:10:59",
      "content": "<p>Have you thought about adding denoising to the pre processing ? I just read this notebook about denoising. <a href=\"https://www.kaggle.com/code/andreipaulavets/byu-denoising-cryo-et-with-noise2void\" target=\"_blank\">BYU Denoising Cryo-ET with Noise2Void</a></p>",
      "rawMarkdown": "Have you thought about adding denoising to the pre processing ? I just read this notebook about denoising. [BYU Denoising Cryo-ET with Noise2Void](https://www.kaggle.com/code/andreipaulavets/byu-denoising-cryo-et-with-noise2void)",
      "votes": null
    },
    {
      "id": "3163304",
      "postDate": "03/30/2025 15:41:52",
      "content": "<p>I have considered denoising and normalization. However, I did not include these topics in the post because I wanted to focus more on data splitting and standardization rather than on normalization, augmentation, and other stuff.</p>",
      "rawMarkdown": "I have considered denoising and normalization. However, I did not include these topics in the post because I wanted to focus more on data splitting and standardization rather than on normalization, augmentation, and other stuff.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3163186,
      "author_name": "mathieuduverne",
      "author_url": "",
      "post_date": "03/30/2025 12:10:59",
      "content": "<p>Have you thought about adding denoising to the pre processing ? I just read this notebook about denoising. <a href=\"https://www.kaggle.com/code/andreipaulavets/byu-denoising-cryo-et-with-noise2void\" target=\"_blank\">BYU Denoising Cryo-ET with Noise2Void</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 3163304,
          "author_name": "eikyou",
          "author_url": "",
          "post_date": "03/30/2025 15:41:52",
          "content": "<p>I have considered denoising and normalization. However, I did not include these topics in the post because I wanted to focus more on data splitting and standardization rather than on normalization, augmentation, and other stuff.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3159189": "I haven’t started working on this competition yet, but I’ve been thinking about it a lot, and I’m completely unsure about the best way to approach data preprocessing, especially for models like U-net. Obviously, the best method will likely be determined through experimentation, but some reasoning can still be applied, and I wanted to share my thoughts with you. I’d be happy to discuss your opinions in the comments.  \n\nThe first obvious approach is not to resize the original data in any way and simply split the train and test objects into overlapping patches and train on them. The clear issues here are too many zero-class labels and an excessively large volume of data, making it difficult and time-consuming to train the model without access to more powerful GPUs than those provided on Kaggle.  \n\nThe other methods involve either resizing the objects, cropping them, or both, followed by splitting into patches and so on.  \n\n1. **Taking a certain number of slices above and below the motor slice for training and rescaling the slices to the desired size**  \n   But then it’s unclear what to do with tomograms that don’t contain a motor to keep the dataset representative. Additionally, for the test set, you can’t just take a chunk—you’d have to process the entire tomogram, which is resource-intensive and might negatively impact model performance.  \n\n2. **Uniformly sampling a certain number of slices across the entire tomogram and rescaling all three dimensions.**  \n   The problem here is that the motor might not clearly appear in the selected slice, especially when there are multiple motors. And we are also straightly discard some data, losing information. However, the advantage of this method is its universality compared to method 1 and its efficiency compared to the baseline approach where raw, unprocessed data is used.  \n\n3. **Rescaling all data to a higher voxel spacing to reduce the size of the input data.**  \n   The obvious issue is that we don’t know the voxel spacing for the test data, so it’s unclear how to rescale them. One option is to rescale everything to the average voxel spacing of the training set, but that doesn’t sound reliable. The advantage, though, is that we aren’t directly discarding data, as is done in approach 2.  \n\n4. **Extending the idea from approach 3, we could ignore voxel spacing altogether and simply interpolate everything to a fixed size.**  \n   The problem here is that the tomograms are initially very different in size, so they might end up overly compressed. But, like approach 3, we aren’t directly discarding data as in approach 2. This method is also as universal as approach 2.  \n\nThese are more or less all the methods I’ve seen on the forum or thought about myself. I haven’t yet considered how these methods might affect target transformation. But the most promising approaches seem to be approaches 2 and 4.",
    "3163186": "Have you thought about adding denoising to the pre processing ? I just read this notebook about denoising. [BYU Denoising Cryo-ET with Noise2Void](https://www.kaggle.com/code/andreipaulavets/byu-denoising-cryo-et-with-noise2void)",
    "3163304": "I have considered denoising and normalization. However, I did not include these topics in the post because I wanted to focus more on data splitting and standardization rather than on normalization, augmentation, and other stuff."
  },
  "source": "meta"
}