{
  "id": 555247,
  "title": "Has Anyone Been Successful With Anything But Denoised Data?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/555247",
  "author_name": "",
  "post_date": "2025-01-06T08:06:18.241199800Z",
  "votes": 15,
  "comment_count": 16,
  "views": 0,
  "content": "<p>I've tried both the synthetic and the IsoNet corrected data.  In both cases, everything else being equal, I ended up with a worse model than with Denoised data alone.  The synthetic data makes sense given how noisy it is, but the IsoNet corrected I thought was supposed to be a processing step beyond Denoised.  I'm wondering if this experience is universal or if there's something specific to my setup that's interfering.  Thanks in advance for any info people are willing to share!</p>",
  "messages": [
    {
      "id": "3089557",
      "postDate": "01/06/2025 08:06:18",
      "content": "<p>I've tried both the synthetic and the IsoNet corrected data.  In both cases, everything else being equal, I ended up with a worse model than with Denoised data alone.  The synthetic data makes sense given how noisy it is, but the IsoNet corrected I thought was supposed to be a processing step beyond Denoised.  I'm wondering if this experience is universal or if there's something specific to my setup that's interfering.  Thanks in advance for any info people are willing to share!</p>",
      "rawMarkdown": "I've tried both the synthetic and the IsoNet corrected data.  In both cases, everything else being equal, I ended up with a worse model than with Denoised data alone.  The synthetic data makes sense given how noisy it is, but the IsoNet corrected I thought was supposed to be a processing step beyond Denoised.  I'm wondering if this experience is universal or if there's something specific to my setup that's interfering.  Thanks in advance for any info people are willing to share!",
      "votes": null
    },
    {
      "id": "3089572",
      "postDate": "01/06/2025 08:20:17",
      "content": "<p>I'm always nervous about extra processing steps, since they can never create new information and may destroy information. I would actually expect the best results on the rawest data (WBP I think?), and this may be particularly important when trying to do transfer learning from the synthetic data. I haven't investigated this though, because as I understand it we only have 'denoised' avialable in the test set (did anyone actually check this?).</p>",
      "rawMarkdown": "I'm always nervous about extra processing steps, since they can never create new information and may destroy information. I would actually expect the best results on the rawest data (WBP I think?), and this may be particularly important when trying to do transfer learning from the synthetic data. I haven't investigated this though, because as I understand it we only have 'denoised' avialable in the test set (did anyone actually check this?).",
      "votes": null
    },
    {
      "id": "3089603",
      "postDate": "01/06/2025 09:27:39",
      "content": "<p>Don't be. I once finished the competition badly because I focused too much on 16bit unprocessed 7-channels images (RAW), while the winning solution used JPGs (Well done in steak terms). NN sees things very differently and unless you work with vanishingly small signals pre-processing can greatly help your network.</p>",
      "rawMarkdown": "Don't be. I once finished the competition badly because I focused too much on 16bit unprocessed 7-channels images (RAW), while the winning solution used JPGs (Well done in steak terms). NN sees things very differently and unless you work with vanishingly small signals pre-processing can greatly help your network.",
      "votes": null
    },
    {
      "id": "3089627",
      "postDate": "01/06/2025 10:33:55",
      "content": "<p>As a suggestion. You can use UNET style NN to predict other types of preprocessing (IsoNet, etc) from the denoised one and then finetune NN on the real task. But that would be useful if you have more data (denoised vs IsoNet, etc, but no segmentations are needed) than just the train set. </p>",
      "rawMarkdown": "As a suggestion. You can use UNET style NN to predict other types of preprocessing (IsoNet, etc) from the denoised one and then finetune NN on the real task. But that would be useful if you have more data (denoised vs IsoNet, etc, but no segmentations are needed) than just the train set.",
      "votes": null
    },
    {
      "id": "3089666",
      "postDate": "01/06/2025 11:49:20",
      "content": "<p>Oh I absolutely agree that preprocessing is needed. But by using the denoised data, we are using a specific type of denoising that may or not be optimal for our models. I'd rather work on the raw data and try out different denoising algorithms.</p>",
      "rawMarkdown": "Oh I absolutely agree that preprocessing is needed. But by using the denoised data, we are using a specific type of denoising that may or not be optimal for our models. I'd rather work on the raw data and try out different denoising algorithms.",
      "votes": null
    },
    {
      "id": "3089881",
      "postDate": "01/06/2025 16:29:09",
      "content": "<p>Because the lack of data I've attempted to pretrain in synthetic. Nothing spectacular, nothing terrible. But the models obtained were all blind to experimental virus, luckyly virus is easy to catch.</p>\n<p>EDIT: Just to clarify. In my attempts, pretraining apparently helped a bit at start of real training but didn't reached better results at the end.</p>",
      "rawMarkdown": "Because the lack of data I've attempted to pretrain in synthetic. Nothing spectacular, nothing terrible. But the models obtained were all blind to experimental virus, luckyly virus is easy to catch.\n\nEDIT: Just to clarify. In my attempts, pretraining apparently helped a bit at start of real training but didn't reached better results at the end.",
      "votes": null
    },
    {
      "id": "3090009",
      "postDate": "01/06/2025 19:13:01",
      "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> <br>\nI didn't spend much time looking through it, but the data portal may have the original raw data.  It's definitely organized differently that what we received for the competition.<br>\n<a href=\"https://cryoetdataportal.czscience.com/datasets/10440\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10440</a></p>",
      "rawMarkdown": "jeroencottaar \nI didn't spend much time looking through it, but the data portal may have the original raw data.  It's definitely organized differently that what we received for the competition.\n[https://cryoetdataportal.czscience.com/datasets/10440](https://cryoetdataportal.czscience.com/datasets/10440)",
      "votes": null
    },
    {
      "id": "3090017",
      "postDate": "01/06/2025 19:27:19",
      "content": "<p>Yes, the Portal has the raw tomograms available for the competition data set. You can learn about the data schema <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#data-organization\" target=\"_blank\">here</a>, and in particular, here is the <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#run-download-options\" target=\"_blank\">folder structure for runs</a>.</p>",
      "rawMarkdown": "Yes, the Portal has the raw tomograms available for the competition data set. You can learn about the data schema [here](https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#data-organization), and in particular, here is the [folder structure for runs](https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#run-download-options).",
      "votes": null
    },
    {
      "id": "3090021",
      "postDate": "01/06/2025 19:35:47",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/danniellemccarthy\" target=\"_blank\">@danniellemccarthy</a>! BTW, is the software that was used to process the tomograms to create the denoised data available anywhere?  Might be useful to try to run that against the synthetic datasets to create denoised versions of those.</p>",
      "rawMarkdown": "Thanks @danniellemccarthy! BTW, is the software that was used to process the tomograms to create the denoised data available anywhere?  Might be useful to try to run that against the synthetic datasets to create denoised versions of those.",
      "votes": null
    },
    {
      "id": "3090124",
      "postDate": "01/06/2025 22:57:12",
      "content": "<p>Is your model focusing on 2D in the xy plane? I've found that IsoNet's real improvement is in the side views. The raw data is a tilt series of images and then the tomogram reconstruction creates a missing wedge which in turn creates \"streak\"-like artifacts in one of the side views (x-z I think). IsoNet's main purpose is missing wedge reconstruction, not necessarily noise reduction (based on what I've read in some Cryo EM papers) so you'll see most of the improvement in the yz and xz planes like in this image. I found this webpage to be very helpful in understanding how these tomograms were made: <a href=\"https://cryoem101.org/cryoet-chapter-5/\" target=\"_blank\">https://cryoem101.org/cryoet-chapter-5/</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F689ac688a7d90db481e5fc6d88fcd330%2Fcryoem.PNG?generation=1736204147224101&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Is your model focusing on 2D in the xy plane? I've found that IsoNet's real improvement is in the side views. The raw data is a tilt series of images and then the tomogram reconstruction creates a missing wedge which in turn creates \"streak\"-like artifacts in one of the side views (x-z I think). IsoNet's main purpose is missing wedge reconstruction, not necessarily noise reduction (based on what I've read in some Cryo EM papers) so you'll see most of the improvement in the yz and xz planes like in this image. I found this webpage to be very helpful in understanding how these tomograms were made: https://cryoem101.org/cryoet-chapter-5/\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F689ac688a7d90db481e5fc6d88fcd330%2Fcryoem.PNG?generation=1736204147224101&alt=media)",
      "votes": null
    },
    {
      "id": "3090132",
      "postDate": "01/06/2025 23:28:52",
      "content": "<p><a href=\"https://www.kaggle.com/williamtsisson\" target=\"_blank\">@williamtsisson</a> Yeah, this is discussed in one of the other threads.  IsoNet is supposed to partially correct the z-axis deformation, but for some reason, we don't seem to be seeing it in the data.</p>\n<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742</a></p>\n<p>You have to hunt for it, but here's an image comparing a composite of all the denoised images vs. isonet corrected.  I've seen images of what's supposed to happen, but in our dataset the effect is much more subtle bordering on non-existent.  It's even harder to see in the individual images.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Feecfd805df90dea8cfc39359054d0e78%2Fisonet.jpg?generation=1736206078333719&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "williamtsisson Yeah, this is discussed in one of the other threads.  IsoNet is supposed to partially correct the z-axis deformation, but for some reason, we don't seem to be seeing it in the data.\n\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742)\n\nYou have to hunt for it, but here's an image comparing a composite of all the denoised images vs. isonet corrected.  I've seen images of what's supposed to happen, but in our dataset the effect is much more subtle bordering on non-existent.  It's even harder to see in the individual images.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Feecfd805df90dea8cfc39359054d0e78%2Fisonet.jpg?generation=1736206078333719&alt=media)",
      "votes": null
    },
    {
      "id": "3090162",
      "postDate": "01/07/2025 00:52:05",
      "content": "<p>Yeah the difference on our dataset is much more obvious if you zoom out.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F10929099862f93f39c25c5bebb11a2cf%2FScreenshot%202025-01-06%20at%206.23.40PM.png?generation=1736209456756106&amp;alt=media\" alt=\"\"></p>\n<p>But yeah the isonet correction seems like a mixed bag. Like in this image of Thyroglobulin (denoised on left, isonet on right), I don't think it would help for the top particle but it might help for the bottom particle to separate it from the blob in the bottom left <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F48b358b19ae6a3f0e24f0769d94ca51f%2FScreenshot%202025-01-06%20at%206.49.38PM.png?generation=1736211079822804&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Yeah the difference on our dataset is much more obvious if you zoom out.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F10929099862f93f39c25c5bebb11a2cf%2FScreenshot%202025-01-06%20at%206.23.40PM.png?generation=1736209456756106&alt=media)\n\nBut yeah the isonet correction seems like a mixed bag. Like in this image of Thyroglobulin (denoised on left, isonet on right), I don't think it would help for the top particle but it might help for the bottom particle to separate it from the blob in the bottom left \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F48b358b19ae6a3f0e24f0769d94ca51f%2FScreenshot%202025-01-06%20at%206.49.38PM.png?generation=1736211079822804&alt=media)",
      "votes": null
    },
    {
      "id": "3090169",
      "postDate": "01/07/2025 01:08:25",
      "content": "<p>Ahh, cool!  Thanks for the zoom out suggestion.  I can certainly see it there (and in your other photo.) It's weird that it seems to affect some things but not others, but maybe that's a Fourier wavelength thing.</p>",
      "rawMarkdown": "Ahh, cool!  Thanks for the zoom out suggestion.  I can certainly see it there (and in your other photo.) It's weird that it seems to affect some things but not others, but maybe that's a Fourier wavelength thing.",
      "votes": null
    },
    {
      "id": "3091019",
      "postDate": "01/07/2025 23:01:00",
      "content": "<p>Hi all. I can say a couple of things about the isonet-corrected tomograms. Isonet is a great tool and we found that in order to achieve the same results as the ones in their paper, we have to do sit down and figure out the right set of parameters. We definitely spent some time doing that, but the results are not the best they can be, as you have seen and reported here. Why did we include them in the challenge? We were not convinced that we shouldn't include them based on our tests. So the bottom line is: please take the isonet-corrected tomograms with a grain of salt.</p>",
      "rawMarkdown": "Hi all. I can say a couple of things about the isonet-corrected tomograms. Isonet is a great tool and we found that in order to achieve the same results as the ones in their paper, we have to do sit down and figure out the right set of parameters. We definitely spent some time doing that, but the results are not the best they can be, as you have seen and reported here. Why did we include them in the challenge? We were not convinced that we shouldn't include them based on our tests. So the bottom line is: please take the isonet-corrected tomograms with a grain of salt.",
      "votes": null
    },
    {
      "id": "3094211",
      "postDate": "01/11/2025 18:29:21",
      "content": "<p>I am also interested in this! The famous DenoisET 😆</p>",
      "rawMarkdown": "I am also interested in this! The famous DenoisET 😆",
      "votes": null
    },
    {
      "id": "3094308",
      "postDate": "01/11/2025 22:30:26",
      "content": "<p>I've tried a lot of different things using the YOLO approach that <a href=\"https://www.kaggle.com/ITK8191\" target=\"_blank\">@ITK8191</a> shared. The best model I trained was with the additional data. In another discussion I think that I mentioned that I denoised the data using only gaussian denoising. The images don't change much from the original when denoised with Gaussian, but I think the particles become more visible at least.</p>\n<p>I've shared the YOLO synthetic dataset here for anyone who finds it useful: <a href=\"https://www.kaggle.com/code/sersasj/czii-making-datasets-for-yolo-synthetic-data\" target=\"_blank\">YOLO Synthetic Dataset</a>.</p>\n<p>Maybe it was just a coincidence that the model ended up in a better local minimum, but the YOLO model using only the original data achieved around ~0.65, and ~0.68 with synthetic data.</p>",
      "rawMarkdown": "I've tried a lot of different things using the YOLO approach that @ITK8191 shared. The best model I trained was with the additional data. In another discussion I think that I mentioned that I denoised the data using only gaussian denoising. The images don't change much from the original when denoised with Gaussian, but I think the particles become more visible at least.\n\nI've shared the YOLO synthetic dataset here for anyone who finds it useful: [YOLO Synthetic Dataset](https://www.kaggle.com/code/sersasj/czii-making-datasets-for-yolo-synthetic-data).\n\nMaybe it was just a coincidence that the model ended up in a better local minimum, but the YOLO model using only the original data achieved around ~0.65, and ~0.68 with synthetic data.",
      "votes": null
    },
    {
      "id": "3094337",
      "postDate": "01/12/2025 00:48:05",
      "content": "<p>That's a pretty big difference!  If it were random, I would expect something smaller unless there was something wrong with the 0.65 model.</p>",
      "rawMarkdown": "That's a pretty big difference!  If it were random, I would expect something smaller unless there was something wrong with the 0.65 model.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3089572,
      "author_name": "jeroencottaar",
      "author_url": "",
      "post_date": "01/06/2025 08:20:17",
      "content": "<p>I'm always nervous about extra processing steps, since they can never create new information and may destroy information. I would actually expect the best results on the rawest data (WBP I think?), and this may be particularly important when trying to do transfer learning from the synthetic data. I haven't investigated this though, because as I understand it we only have 'denoised' avialable in the test set (did anyone actually check this?).</p>",
      "votes": null,
      "replies": [
        {
          "id": 3089603,
          "author_name": "sakvaua",
          "author_url": "",
          "post_date": "01/06/2025 09:27:39",
          "content": "<p>Don't be. I once finished the competition badly because I focused too much on 16bit unprocessed 7-channels images (RAW), while the winning solution used JPGs (Well done in steak terms). NN sees things very differently and unless you work with vanishingly small signals pre-processing can greatly help your network.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3089666,
              "author_name": "jeroencottaar",
              "author_url": "",
              "post_date": "01/06/2025 11:49:20",
              "content": "<p>Oh I absolutely agree that preprocessing is needed. But by using the denoised data, we are using a specific type of denoising that may or not be optimal for our models. I'd rather work on the raw data and try out different denoising algorithms.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3090009,
                  "author_name": "davidlist",
                  "author_url": "",
                  "post_date": "01/06/2025 19:13:01",
                  "content": "<p><a href=\"https://www.kaggle.com/jeroencottaar\" target=\"_blank\">@jeroencottaar</a> <br>\nI didn't spend much time looking through it, but the data portal may have the original raw data.  It's definitely organized differently that what we received for the competition.<br>\n<a href=\"https://cryoetdataportal.czscience.com/datasets/10440\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10440</a></p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3090017,
                      "author_name": "danniellemccarthy",
                      "author_url": "",
                      "post_date": "01/06/2025 19:27:19",
                      "content": "<p>Yes, the Portal has the raw tomograms available for the competition data set. You can learn about the data schema <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#data-organization\" target=\"_blank\">here</a>, and in particular, here is the <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#run-download-options\" target=\"_blank\">folder structure for runs</a>.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3090021,
                          "author_name": "davidlist",
                          "author_url": "",
                          "post_date": "01/06/2025 19:35:47",
                          "content": "<p>Thanks <a href=\"https://www.kaggle.com/danniellemccarthy\" target=\"_blank\">@danniellemccarthy</a>! BTW, is the software that was used to process the tomograms to create the denoised data available anywhere?  Might be useful to try to run that against the synthetic datasets to create denoised versions of those.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3094211,
                              "author_name": "carlosperez97",
                              "author_url": "",
                              "post_date": "01/11/2025 18:29:21",
                              "content": "<p>I am also interested in this! The famous DenoisET 😆</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3089627,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "01/06/2025 10:33:55",
      "content": "<p>As a suggestion. You can use UNET style NN to predict other types of preprocessing (IsoNet, etc) from the denoised one and then finetune NN on the real task. But that would be useful if you have more data (denoised vs IsoNet, etc, but no segmentations are needed) than just the train set. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3089881,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "01/06/2025 16:29:09",
      "content": "<p>Because the lack of data I've attempted to pretrain in synthetic. Nothing spectacular, nothing terrible. But the models obtained were all blind to experimental virus, luckyly virus is easy to catch.</p>\n<p>EDIT: Just to clarify. In my attempts, pretraining apparently helped a bit at start of real training but didn't reached better results at the end.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3090124,
      "author_name": "williamtsisson",
      "author_url": "",
      "post_date": "01/06/2025 22:57:12",
      "content": "<p>Is your model focusing on 2D in the xy plane? I've found that IsoNet's real improvement is in the side views. The raw data is a tilt series of images and then the tomogram reconstruction creates a missing wedge which in turn creates \"streak\"-like artifacts in one of the side views (x-z I think). IsoNet's main purpose is missing wedge reconstruction, not necessarily noise reduction (based on what I've read in some Cryo EM papers) so you'll see most of the improvement in the yz and xz planes like in this image. I found this webpage to be very helpful in understanding how these tomograms were made: <a href=\"https://cryoem101.org/cryoet-chapter-5/\" target=\"_blank\">https://cryoem101.org/cryoet-chapter-5/</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F689ac688a7d90db481e5fc6d88fcd330%2Fcryoem.PNG?generation=1736204147224101&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3090132,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "01/06/2025 23:28:52",
          "content": "<p><a href=\"https://www.kaggle.com/williamtsisson\" target=\"_blank\">@williamtsisson</a> Yeah, this is discussed in one of the other threads.  IsoNet is supposed to partially correct the z-axis deformation, but for some reason, we don't seem to be seeing it in the data.</p>\n<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742\" target=\"_blank\">https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742</a></p>\n<p>You have to hunt for it, but here's an image comparing a composite of all the denoised images vs. isonet corrected.  I've seen images of what's supposed to happen, but in our dataset the effect is much more subtle bordering on non-existent.  It's even harder to see in the individual images.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Feecfd805df90dea8cfc39359054d0e78%2Fisonet.jpg?generation=1736206078333719&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 3090162,
              "author_name": "williamtsisson",
              "author_url": "",
              "post_date": "01/07/2025 00:52:05",
              "content": "<p>Yeah the difference on our dataset is much more obvious if you zoom out.  <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F10929099862f93f39c25c5bebb11a2cf%2FScreenshot%202025-01-06%20at%206.23.40PM.png?generation=1736209456756106&amp;alt=media\" alt=\"\"></p>\n<p>But yeah the isonet correction seems like a mixed bag. Like in this image of Thyroglobulin (denoised on left, isonet on right), I don't think it would help for the top particle but it might help for the bottom particle to separate it from the blob in the bottom left <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F48b358b19ae6a3f0e24f0769d94ca51f%2FScreenshot%202025-01-06%20at%206.49.38PM.png?generation=1736211079822804&amp;alt=media\" alt=\"\"></p>",
              "votes": null,
              "replies": [
                {
                  "id": 3090169,
                  "author_name": "davidlist",
                  "author_url": "",
                  "post_date": "01/07/2025 01:08:25",
                  "content": "<p>Ahh, cool!  Thanks for the zoom out suggestion.  I can certainly see it there (and in your other photo.) It's weird that it seems to affect some things but not others, but maybe that's a Fourier wavelength thing.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3091019,
      "author_name": "rezaparaan",
      "author_url": "",
      "post_date": "01/07/2025 23:01:00",
      "content": "<p>Hi all. I can say a couple of things about the isonet-corrected tomograms. Isonet is a great tool and we found that in order to achieve the same results as the ones in their paper, we have to do sit down and figure out the right set of parameters. We definitely spent some time doing that, but the results are not the best they can be, as you have seen and reported here. Why did we include them in the challenge? We were not convinced that we shouldn't include them based on our tests. So the bottom line is: please take the isonet-corrected tomograms with a grain of salt.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3094308,
      "author_name": "sersasj",
      "author_url": "",
      "post_date": "01/11/2025 22:30:26",
      "content": "<p>I've tried a lot of different things using the YOLO approach that <a href=\"https://www.kaggle.com/ITK8191\" target=\"_blank\">@ITK8191</a> shared. The best model I trained was with the additional data. In another discussion I think that I mentioned that I denoised the data using only gaussian denoising. The images don't change much from the original when denoised with Gaussian, but I think the particles become more visible at least.</p>\n<p>I've shared the YOLO synthetic dataset here for anyone who finds it useful: <a href=\"https://www.kaggle.com/code/sersasj/czii-making-datasets-for-yolo-synthetic-data\" target=\"_blank\">YOLO Synthetic Dataset</a>.</p>\n<p>Maybe it was just a coincidence that the model ended up in a better local minimum, but the YOLO model using only the original data achieved around ~0.65, and ~0.68 with synthetic data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3094337,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "01/12/2025 00:48:05",
          "content": "<p>That's a pretty big difference!  If it were random, I would expect something smaller unless there was something wrong with the 0.65 model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3089557": "I've tried both the synthetic and the IsoNet corrected data.  In both cases, everything else being equal, I ended up with a worse model than with Denoised data alone.  The synthetic data makes sense given how noisy it is, but the IsoNet corrected I thought was supposed to be a processing step beyond Denoised.  I'm wondering if this experience is universal or if there's something specific to my setup that's interfering.  Thanks in advance for any info people are willing to share!",
    "3089572": "I'm always nervous about extra processing steps, since they can never create new information and may destroy information. I would actually expect the best results on the rawest data (WBP I think?), and this may be particularly important when trying to do transfer learning from the synthetic data. I haven't investigated this though, because as I understand it we only have 'denoised' avialable in the test set (did anyone actually check this?).",
    "3089603": "Don't be. I once finished the competition badly because I focused too much on 16bit unprocessed 7-channels images (RAW), while the winning solution used JPGs (Well done in steak terms). NN sees things very differently and unless you work with vanishingly small signals pre-processing can greatly help your network.",
    "3089627": "As a suggestion. You can use UNET style NN to predict other types of preprocessing (IsoNet, etc) from the denoised one and then finetune NN on the real task. But that would be useful if you have more data (denoised vs IsoNet, etc, but no segmentations are needed) than just the train set.",
    "3089666": "Oh I absolutely agree that preprocessing is needed. But by using the denoised data, we are using a specific type of denoising that may or not be optimal for our models. I'd rather work on the raw data and try out different denoising algorithms.",
    "3089881": "Because the lack of data I've attempted to pretrain in synthetic. Nothing spectacular, nothing terrible. But the models obtained were all blind to experimental virus, luckyly virus is easy to catch.\n\nEDIT: Just to clarify. In my attempts, pretraining apparently helped a bit at start of real training but didn't reached better results at the end.",
    "3090009": "jeroencottaar \nI didn't spend much time looking through it, but the data portal may have the original raw data.  It's definitely organized differently that what we received for the competition.\n[https://cryoetdataportal.czscience.com/datasets/10440](https://cryoetdataportal.czscience.com/datasets/10440)",
    "3090017": "Yes, the Portal has the raw tomograms available for the competition data set. You can learn about the data schema [here](https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#data-organization), and in particular, here is the [folder structure for runs](https://chanzuckerberg.github.io/cryoet-data-portal/cryoet_data_portal_docsite_data.html#run-download-options).",
    "3090021": "Thanks @danniellemccarthy! BTW, is the software that was used to process the tomograms to create the denoised data available anywhere?  Might be useful to try to run that against the synthetic datasets to create denoised versions of those.",
    "3090124": "Is your model focusing on 2D in the xy plane? I've found that IsoNet's real improvement is in the side views. The raw data is a tilt series of images and then the tomogram reconstruction creates a missing wedge which in turn creates \"streak\"-like artifacts in one of the side views (x-z I think). IsoNet's main purpose is missing wedge reconstruction, not necessarily noise reduction (based on what I've read in some Cryo EM papers) so you'll see most of the improvement in the yz and xz planes like in this image. I found this webpage to be very helpful in understanding how these tomograms were made: https://cryoem101.org/cryoet-chapter-5/\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F689ac688a7d90db481e5fc6d88fcd330%2Fcryoem.PNG?generation=1736204147224101&alt=media)",
    "3090132": "williamtsisson Yeah, this is discussed in one of the other threads.  IsoNet is supposed to partially correct the z-axis deformation, but for some reason, we don't seem to be seeing it in the data.\n\n[https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545742)\n\nYou have to hunt for it, but here's an image comparing a composite of all the denoised images vs. isonet corrected.  I've seen images of what's supposed to happen, but in our dataset the effect is much more subtle bordering on non-existent.  It's even harder to see in the individual images.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F10704200%2Feecfd805df90dea8cfc39359054d0e78%2Fisonet.jpg?generation=1736206078333719&alt=media)",
    "3090162": "Yeah the difference on our dataset is much more obvious if you zoom out.  \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F10929099862f93f39c25c5bebb11a2cf%2FScreenshot%202025-01-06%20at%206.23.40PM.png?generation=1736209456756106&alt=media)\n\nBut yeah the isonet correction seems like a mixed bag. Like in this image of Thyroglobulin (denoised on left, isonet on right), I don't think it would help for the top particle but it might help for the bottom particle to separate it from the blob in the bottom left \n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F24058428%2F48b358b19ae6a3f0e24f0769d94ca51f%2FScreenshot%202025-01-06%20at%206.49.38PM.png?generation=1736211079822804&alt=media)",
    "3090169": "Ahh, cool!  Thanks for the zoom out suggestion.  I can certainly see it there (and in your other photo.) It's weird that it seems to affect some things but not others, but maybe that's a Fourier wavelength thing.",
    "3091019": "Hi all. I can say a couple of things about the isonet-corrected tomograms. Isonet is a great tool and we found that in order to achieve the same results as the ones in their paper, we have to do sit down and figure out the right set of parameters. We definitely spent some time doing that, but the results are not the best they can be, as you have seen and reported here. Why did we include them in the challenge? We were not convinced that we shouldn't include them based on our tests. So the bottom line is: please take the isonet-corrected tomograms with a grain of salt.",
    "3094211": "I am also interested in this! The famous DenoisET 😆",
    "3094308": "I've tried a lot of different things using the YOLO approach that @ITK8191 shared. The best model I trained was with the additional data. In another discussion I think that I mentioned that I denoised the data using only gaussian denoising. The images don't change much from the original when denoised with Gaussian, but I think the particles become more visible at least.\n\nI've shared the YOLO synthetic dataset here for anyone who finds it useful: [YOLO Synthetic Dataset](https://www.kaggle.com/code/sersasj/czii-making-datasets-for-yolo-synthetic-data).\n\nMaybe it was just a coincidence that the model ended up in a better local minimum, but the YOLO model using only the original data achieved around ~0.65, and ~0.68 with synthetic data.",
    "3094337": "That's a pretty big difference!  If it were random, I would expect something smaller unless there was something wrong with the 0.65 model."
  },
  "source": "meta"
}