{
  "id": 556103,
  "title": "Does anyone have results with the 10441 external dataset works?",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/556103",
  "author_name": "ITK8191",
  "post_date": "2025-01-11T08:56:54.042000",
  "votes": 6,
  "comment_count": 12,
  "views": 0,
  "content": "<p>Hello everyone, we are now in the second half of the competition and the fight for LBs is becoming fierce.<br>\nI now suspect that the low accuracy of the model is due to a lack of dataset size, and am trying to use an external dataset. As has already been mentioned several times by other competitors, the data called <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">DS-10441</a> is also involved with <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>, the host of this competition, and is considered close to DS-10440 (the data in this competition).<br>\nData 10441 can be easily downloaded from <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551222\" target=\"_blank\">this discussion</a> and <a href=\"https://www.kaggle.com/datasets/drxc75/czcdp10441\" target=\"_blank\">his dataset</a>. </p>\n<p>btw, All volume images were saved in wbp format, which differs from the denoised data used in this competition.<br>\nIs anyone using this dataset to improve your CV or LB? <br>\nI'm not doing well till now.</p>\n<p>Here I will share some experiments I have tried that did not work.</p>\n<ol>\n<li>I tried mixing the wbp data with the denoised data and training it. The validation data was only denoised. CV remained flat.</li>\n<li>In addition to 1, I also used ctfdeconvolved and isonetcorrected on the 10440 dataset for training. This also did not work enough.</li>\n<li>Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.</li>\n</ol>\n<p>When converting from wbp to denoised in 3 exp, some external datasets had artifacts and the conversion could not be done successfully. The following image is an example: a slice of TS_20 at z=425. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F2bb0d50b5bfe2ce5868af3ece04ef38d%2FTS_20_0420_denoised.png?generation=1736585054092517&amp;alt=media\" alt=\"TS_20_0420_denoised\"><br>\nThe appearance of this artifact made more than half of the volume images unusable.</p>\n<p>Again, has anyone used an external dataset to improve CV or LB? Is there a way to avoid artifacts in the wbp conversion?<br>\nI'd be happy to hear your thoughts and opinions, even if they're not solutions or tips!</p>",
  "messages": [
    {
      "id": 3093676,
      "postDate": "2025-01-11T08:56:54.043Z",
      "content": "<p>Hello everyone, we are now in the second half of the competition and the fight for LBs is becoming fierce.<br>\nI now suspect that the low accuracy of the model is due to a lack of dataset size, and am trying to use an external dataset. As has already been mentioned several times by other competitors, the data called <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">DS-10441</a> is also involved with <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>, the host of this competition, and is considered close to DS-10440 (the data in this competition).<br>\nData 10441 can be easily downloaded from <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551222\" target=\"_blank\">this discussion</a> and <a href=\"https://www.kaggle.com/datasets/drxc75/czcdp10441\" target=\"_blank\">his dataset</a>. </p>\n<p>btw, All volume images were saved in wbp format, which differs from the denoised data used in this competition.<br>\nIs anyone using this dataset to improve your CV or LB? <br>\nI'm not doing well till now.</p>\n<p>Here I will share some experiments I have tried that did not work.</p>\n<ol>\n<li>I tried mixing the wbp data with the denoised data and training it. The validation data was only denoised. CV remained flat.</li>\n<li>In addition to 1, I also used ctfdeconvolved and isonetcorrected on the 10440 dataset for training. This also did not work enough.</li>\n<li>Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.</li>\n</ol>\n<p>When converting from wbp to denoised in 3 exp, some external datasets had artifacts and the conversion could not be done successfully. The following image is an example: a slice of TS_20 at z=425. <br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F2bb0d50b5bfe2ce5868af3ece04ef38d%2FTS_20_0420_denoised.png?generation=1736585054092517&amp;alt=media\" alt=\"TS_20_0420_denoised\"><br>\nThe appearance of this artifact made more than half of the volume images unusable.</p>\n<p>Again, has anyone used an external dataset to improve CV or LB? Is there a way to avoid artifacts in the wbp conversion?<br>\nI'd be happy to hear your thoughts and opinions, even if they're not solutions or tips!</p>",
      "rawMarkdown": "Hello everyone, we are now in the second half of the competition and the fight for LBs is becoming fierce.\nI now suspect that the low accuracy of the model is due to a lack of dataset size, and am trying to use an external dataset. As has already been mentioned several times by other competitors, the data called [DS-10441](https://cryoetdataportal.czscience.com/datasets/10441) is also involved with @kharrington, the host of this competition, and is considered close to DS-10440 (the data in this competition).\nData 10441 can be easily downloaded from [this discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551222) and [his dataset](https://www.kaggle.com/datasets/drxc75/czcdp10441). \n\nbtw, All volume images were saved in wbp format, which differs from the denoised data used in this competition.\nIs anyone using this dataset to improve your CV or LB? \nI'm not doing well till now.\n\nHere I will share some experiments I have tried that did not work.\n1. I tried mixing the wbp data with the denoised data and training it. The validation data was only denoised. CV remained flat.\n2.  In addition to 1, I also used ctfdeconvolved and isonetcorrected on the 10440 dataset for training. This also did not work enough.\n3. Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.\n\nWhen converting from wbp to denoised in 3 exp, some external datasets had artifacts and the conversion could not be done successfully. The following image is an example: a slice of TS_20 at z=425. \n![TS_20_0420_denoised](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F2bb0d50b5bfe2ce5868af3ece04ef38d%2FTS_20_0420_denoised.png?generation=1736585054092517&alt=media)\nThe appearance of this artifact made more than half of the volume images unusable.\n\nAgain, has anyone used an external dataset to improve CV or LB? Is there a way to avoid artifacts in the wbp conversion?\nI'd be happy to hear your thoughts and opinions, even if they're not solutions or tips!",
      "votes": 6
    },
    {
      "id": 3093831,
      "postDate": "2025-01-11T11:56:33.320Z",
      "content": "<p>I've tried to pretrain in raw, it helped at beggining. But didn't improved the final results. Also mention that the pretrained models were blind to experimental virus. May be because wbp format you say. I'm not sure what do you mean with that, the files I've downloaded were all zarr files. I've uploaded them toguether coordinates <a href=\"https://www.kaggle.com/datasets/sacuscreed/czii-synth-data\" target=\"_blank\">here</a>.</p>",
      "rawMarkdown": "I've tried to pretrain in raw, it helped at beggining. But didn't improved the final results. Also mention that the pretrained models were blind to experimental virus. May be because wbp format you say. I'm not sure what do you mean with that, the files I've downloaded were all zarr files. I've uploaded them toguether coordinates [here](https://www.kaggle.com/datasets/sacuscreed/czii-synth-data).",
      "votes": 1,
      "replies": [
        {
          "id": 3094353,
          "postDate": "2025-01-12T01:37:08.360Z",
          "content": "<p>I also used an external dataset for pre-training but it didn't work, I don't know why.<br>\nBy wbp I meant volumetric images with no image preprocessing. The types of image preprocessing used in this competition are wbp (no preprocessing), denoised, ctfdeconvolved and isonetcorrected. And all volumes have the extension zarr. I looked at your dataset and the images are closest to wbp.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F0aef7c4d68a6d276b6e1301689d660ba%2FTS_5_0920.png?generation=1736645645496849&amp;alt=media\" alt=\"\"><br>\nThis image is at z=925 in TS_5 of your dataset.</p>",
          "rawMarkdown": "I also used an external dataset for pre-training but it didn't work, I don't know why.\nBy wbp I meant volumetric images with no image preprocessing. The types of image preprocessing used in this competition are wbp (no preprocessing), denoised, ctfdeconvolved and isonetcorrected. And all volumes have the extension zarr. I looked at your dataset and the images are closest to wbp.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F0aef7c4d68a6d276b6e1301689d660ba%2FTS_5_0920.png?generation=1736645645496849&alt=media)\nThis image is at z=925 in TS_5 of your dataset.",
          "votes": 1,
          "replies": [
            {
              "id": 3094360,
              "postDate": "2025-01-12T01:49:54.713Z",
              "content": "<p>I think more than processing is the fact that they are simulated. They are a good aproximation but they differ from experimental in more than processing steps since they're based (I think) in idealized readings. In my case synthetic virus were empty shells, while experimental have a blurry core. So the models could extrapolate in experimental other particles reasonably well, but no virus at all.</p>",
              "rawMarkdown": "I think more than processing is the fact that they are simulated. They are a good aproximation but they differ from experimental in more than processing steps since they're based (I think) in idealized readings. In my case synthetic virus were empty shells, while experimental have a blurry core. So the models could extrapolate in experimental other particles reasonably well, but no virus at all.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3094199,
      "postDate": "2025-01-11T18:18:38.223Z",
      "content": "<p><strong>Regarding point 3:</strong> Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.<br>\nDid you manage to obtain decent results with the Unet? </p>\n<p>I also convert synthetic datasets to \"denoise\" versions but when incorporating this new data to the training, the model is just doing well in experimental datasets and is is not being able to understand what is happening in synthetic. I saw different distribution ranges in wbp raw synthetic and wbp raw experimentals :(</p>\n<p>I was also trying to find the model or algo for denoising the data (in paper they mentioned it is DenoisET) but no success so far. </p>\n<p>Here is an example of the \"denoise\" model. (values in the title are just mean and standard deviation of the image)</p>\n<ul>\n<li>Noisy image corresponds to wbp raw</li>\n<li>Output it is the model output</li>\n<li>Clean image is the ground truth denoised image</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4238644%2Fbb183db52403b6379f1cdc2852f6c6df%2Fdownload.png?generation=1736619308388119&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "**Regarding point 3:** Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.\nDid you manage to obtain decent results with the Unet? \n\nI also convert synthetic datasets to \"denoise\" versions but when incorporating this new data to the training, the model is just doing well in experimental datasets and is is not being able to understand what is happening in synthetic. I saw different distribution ranges in wbp raw synthetic and wbp raw experimentals :(\n\nI was also trying to find the model or algo for denoising the data (in paper they mentioned it is DenoisET) but no success so far. \n\nHere is an example of the \"denoise\" model. (values in the title are just mean and standard deviation of the image)\n- Noisy image corresponds to wbp raw\n- Output it is the model output\n- Clean image is the ground truth denoised image\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4238644%2Fbb183db52403b6379f1cdc2852f6c6df%2Fdownload.png?generation=1736619308388119&alt=media)",
      "votes": 2,
      "replies": [
        {
          "id": 3094367,
          "postDate": "2025-01-12T02:08:29.637Z",
          "content": "<p>I have been able to recover denoised images from wbp with UNet at least with sufficient accuracy. I don't have detailed statistics, but I subjectively judge it to be visually sufficient.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F3104ddd5ca15e506102efe488b8c33df%2F50_rec.png?generation=1736647656278605&amp;alt=media\" alt=\"\"><br>\nleft is input wbp image, center is predicted denoised image by Unet and right is Ground Truth image.<br>\nIf you want to achieve results like this, consider using larger networks and/or pretrained weights.<br>\nHowever, as mentioned above, this UNet produces artifacts in more than half of the 10441 dataset.</p>",
          "rawMarkdown": "I have been able to recover denoised images from wbp with UNet at least with sufficient accuracy. I don't have detailed statistics, but I subjectively judge it to be visually sufficient.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F3104ddd5ca15e506102efe488b8c33df%2F50_rec.png?generation=1736647656278605&alt=media)\nleft is input wbp image, center is predicted denoised image by Unet and right is Ground Truth image.\nIf you want to achieve results like this, consider using larger networks and/or pretrained weights.\nHowever, as mentioned above, this UNet produces artifacts in more than half of the 10441 dataset.",
          "votes": 2,
          "replies": [
            {
              "id": 3096028,
              "postDate": "2025-01-14T02:35:53.340Z",
              "content": "<p>Hi again. I've been recently reading about this and Is not <a href=\"https://github.com/tbepler/topaz\" target=\"_blank\">Topaz</a> beign used to denoise wbp?</p>",
              "rawMarkdown": "Hi again. I've been recently reading about this and Is not [Topaz](https://github.com/tbepler/topaz) beign used to denoise wbp?"
            },
            {
              "id": 3097993,
              "postDate": "2025-01-16T00:35:23.293Z",
              "content": "<p>I heard Topaz first time, and again, I think a simple UNet would be able to get a decent denoised image.</p>",
              "rawMarkdown": "I heard Topaz first time, and again, I think a simple UNet would be able to get a decent denoised image."
            }
          ]
        }
      ]
    },
    {
      "id": 3096898,
      "postDate": "2025-01-14T19:42:16.273Z",
      "content": "<p><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.full\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.full</a></p>\n<p>\"DenoisET implements the Noise2Noise algorithm37 for denoising cryoET tomograms. This ML algorithm learns to denoise imaging data from training on paired noisy measurements and is thus suitable for methods like cryoET in which the clean signal is difficult to realistically simulate and cannot be measured. However, because cryoET data are collected as a series of frames of the same underlying field-of-view but with different realizations of the stochastic noise process, training data can be readily generated by reconstructing paired tomograms after splitting the acquired frames into half-sets. Our implementation relies on a similar U-Net architecture as used in Topaz-Denoise38 and leverages the contrast transfer function-deconvolved tomograms produced by AreTomo326 to improve contrast enhancement. The increased SNR provided by denoising proved critical for our manual annotation efforts and benefited our machine learning labeling workflows as well (Fig. 1C).\"</p>\n<p>So I assume DenoisET is not open.</p>",
      "rawMarkdown": "https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.full\n\n\"DenoisET implements the Noise2Noise algorithm37 for denoising cryoET tomograms. This ML algorithm learns to denoise imaging data from training on paired noisy measurements and is thus suitable for methods like cryoET in which the clean signal is difficult to realistically simulate and cannot be measured. However, because cryoET data are collected as a series of frames of the same underlying field-of-view but with different realizations of the stochastic noise process, training data can be readily generated by reconstructing paired tomograms after splitting the acquired frames into half-sets. Our implementation relies on a similar U-Net architecture as used in Topaz-Denoise38 and leverages the contrast transfer function-deconvolved tomograms produced by AreTomo326 to improve contrast enhancement. The increased SNR provided by denoising proved critical for our manual annotation efforts and benefited our machine learning labeling workflows as well (Fig. 1C).\"\n\nSo I assume DenoisET is not open.",
      "replies": [
        {
          "id": 3097994,
          "postDate": "2025-01-16T00:38:49.363Z",
          "content": "<p>This is the first time I've seen this paper. I've certainly thought about using Noise2Noise or Noise2Void to get the Denoised image, but it didn't work well enough, thanks for the info.</p>",
          "rawMarkdown": "This is the first time I've seen this paper. I've certainly thought about using Noise2Noise or Noise2Void to get the Denoised image, but it didn't work well enough, thanks for the info.",
          "replies": [
            {
              "id": 3098033,
              "postDate": "2025-01-16T02:11:20.953Z",
              "content": "<p>I've faced same issue. I could train a \"DenoisET\" from topaz that works well on experimental wbp. But synthetic… Playing with DCGAN right now, but I've never used it before.</p>",
              "rawMarkdown": "I've faced same issue. I could train a \"DenoisET\" from topaz that works well on experimental wbp. But synthetic... Playing with DCGAN right now, but I've never used it before."
            },
            {
              "id": 3098081,
              "postDate": "2025-01-16T03:52:58.167Z",
              "content": "<p>I have also tried some GANs, especially CycleGAN, which converts between Denoised and wbp, but the results were awful. This architecture is supposed to perform the task of denoising while preserving the shape to some extent, but I am encountering the phenomenon that the objects that should be detected are moved or erased.</p>",
              "rawMarkdown": "I have also tried some GANs, especially CycleGAN, which converts between Denoised and wbp, but the results were awful. This architecture is supposed to perform the task of denoising while preserving the shape to some extent, but I am encountering the phenomenon that the objects that should be detected are moved or erased."
            }
          ]
        }
      ]
    },
    {
      "id": 3098085,
      "postDate": "2025-01-16T03:59:54.153Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3093831,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-01-11T11:56:33.320000",
      "content": "<p>I've tried to pretrain in raw, it helped at beggining. But didn't improved the final results. Also mention that the pretrained models were blind to experimental virus. May be because wbp format you say. I'm not sure what do you mean with that, the files I've downloaded were all zarr files. I've uploaded them toguether coordinates <a href=\"https://www.kaggle.com/datasets/sacuscreed/czii-synth-data\" target=\"_blank\">here</a>.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3094353,
          "author_name": "ITK8191",
          "author_url": "",
          "post_date": "2025-01-12T01:37:08.360000",
          "content": "<p>I also used an external dataset for pre-training but it didn't work, I don't know why.<br>\nBy wbp I meant volumetric images with no image preprocessing. The types of image preprocessing used in this competition are wbp (no preprocessing), denoised, ctfdeconvolved and isonetcorrected. And all volumes have the extension zarr. I looked at your dataset and the images are closest to wbp.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F0aef7c4d68a6d276b6e1301689d660ba%2FTS_5_0920.png?generation=1736645645496849&amp;alt=media\" alt=\"\"><br>\nThis image is at z=925 in TS_5 of your dataset.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3094360,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-01-12T01:49:54.713000",
              "content": "<p>I think more than processing is the fact that they are simulated. They are a good aproximation but they differ from experimental in more than processing steps since they're based (I think) in idealized readings. In my case synthetic virus were empty shells, while experimental have a blurry core. So the models could extrapolate in experimental other particles reasonably well, but no virus at all.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3094199,
      "author_name": "Carlos Pérez Ricardo",
      "author_url": "",
      "post_date": "2025-01-11T18:18:38.223000",
      "content": "<p><strong>Regarding point 3:</strong> Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.<br>\nDid you manage to obtain decent results with the Unet? </p>\n<p>I also convert synthetic datasets to \"denoise\" versions but when incorporating this new data to the training, the model is just doing well in experimental datasets and is is not being able to understand what is happening in synthetic. I saw different distribution ranges in wbp raw synthetic and wbp raw experimentals :(</p>\n<p>I was also trying to find the model or algo for denoising the data (in paper they mentioned it is DenoisET) but no success so far. </p>\n<p>Here is an example of the \"denoise\" model. (values in the title are just mean and standard deviation of the image)</p>\n<ul>\n<li>Noisy image corresponds to wbp raw</li>\n<li>Output it is the model output</li>\n<li>Clean image is the ground truth denoised image</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4238644%2Fbb183db52403b6379f1cdc2852f6c6df%2Fdownload.png?generation=1736619308388119&amp;alt=media\" alt=\"\"></p>",
      "votes": 2,
      "replies": [
        {
          "id": 3094367,
          "author_name": "ITK8191",
          "author_url": "",
          "post_date": "2025-01-12T02:08:29.637000",
          "content": "<p>I have been able to recover denoised images from wbp with UNet at least with sufficient accuracy. I don't have detailed statistics, but I subjectively judge it to be visually sufficient.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F3104ddd5ca15e506102efe488b8c33df%2F50_rec.png?generation=1736647656278605&amp;alt=media\" alt=\"\"><br>\nleft is input wbp image, center is predicted denoised image by Unet and right is Ground Truth image.<br>\nIf you want to achieve results like this, consider using larger networks and/or pretrained weights.<br>\nHowever, as mentioned above, this UNet produces artifacts in more than half of the 10441 dataset.</p>",
          "votes": 2,
          "replies": [
            {
              "id": 3096028,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-01-14T02:35:53.340000",
              "content": "<p>Hi again. I've been recently reading about this and Is not <a href=\"https://github.com/tbepler/topaz\" target=\"_blank\">Topaz</a> beign used to denoise wbp?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3097993,
              "author_name": "ITK8191",
              "author_url": "",
              "post_date": "2025-01-16T00:35:23.293000",
              "content": "<p>I heard Topaz first time, and again, I think a simple UNet would be able to get a decent denoised image.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3096898,
      "author_name": "Ángel Jacinto Sánchez Ruiz",
      "author_url": "",
      "post_date": "2025-01-14T19:42:16.273000",
      "content": "<p><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.full\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.full</a></p>\n<p>\"DenoisET implements the Noise2Noise algorithm37 for denoising cryoET tomograms. This ML algorithm learns to denoise imaging data from training on paired noisy measurements and is thus suitable for methods like cryoET in which the clean signal is difficult to realistically simulate and cannot be measured. However, because cryoET data are collected as a series of frames of the same underlying field-of-view but with different realizations of the stochastic noise process, training data can be readily generated by reconstructing paired tomograms after splitting the acquired frames into half-sets. Our implementation relies on a similar U-Net architecture as used in Topaz-Denoise38 and leverages the contrast transfer function-deconvolved tomograms produced by AreTomo326 to improve contrast enhancement. The increased SNR provided by denoising proved critical for our manual annotation efforts and benefited our machine learning labeling workflows as well (Fig. 1C).\"</p>\n<p>So I assume DenoisET is not open.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3097994,
          "author_name": "ITK8191",
          "author_url": "",
          "post_date": "2025-01-16T00:38:49.363000",
          "content": "<p>This is the first time I've seen this paper. I've certainly thought about using Noise2Noise or Noise2Void to get the Denoised image, but it didn't work well enough, thanks for the info.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3098033,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2025-01-16T02:11:20.953000",
              "content": "<p>I've faced same issue. I could train a \"DenoisET\" from topaz that works well on experimental wbp. But synthetic… Playing with DCGAN right now, but I've never used it before.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3098081,
              "author_name": "ITK8191",
              "author_url": "",
              "post_date": "2025-01-16T03:52:58.167000",
              "content": "<p>I have also tried some GANs, especially CycleGAN, which converts between Denoised and wbp, but the results were awful. This architecture is supposed to perform the task of denoising while preserving the shape to some extent, but I am encountering the phenomenon that the objects that should be detected are moved or erased.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3098085,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-01-16T03:59:54.153000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3093676": "Hello everyone, we are now in the second half of the competition and the fight for LBs is becoming fierce.\nI now suspect that the low accuracy of the model is due to a lack of dataset size, and am trying to use an external dataset. As has already been mentioned several times by other competitors, the data called [DS-10441](https://cryoetdataportal.czscience.com/datasets/10441) is also involved with @kharrington, the host of this competition, and is considered close to DS-10440 (the data in this competition).\nData 10441 can be easily downloaded from [this discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/551222) and [his dataset](https://www.kaggle.com/datasets/drxc75/czcdp10441). \n\nbtw, All volume images were saved in wbp format, which differs from the denoised data used in this competition.\nIs anyone using this dataset to improve your CV or LB? \nI'm not doing well till now.\n\nHere I will share some experiments I have tried that did not work.\n1. I tried mixing the wbp data with the denoised data and training it. The validation data was only denoised. CV remained flat.\n2.  In addition to 1, I also used ctfdeconvolved and isonetcorrected on the 10440 dataset for training. This also did not work enough.\n3. Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.\n\nWhen converting from wbp to denoised in 3 exp, some external datasets had artifacts and the conversion could not be done successfully. The following image is an example: a slice of TS_20 at z=425. \n![TS_20_0420_denoised](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2294613%2F2bb0d50b5bfe2ce5868af3ece04ef38d%2FTS_20_0420_denoised.png?generation=1736585054092517&alt=media)\nThe appearance of this artifact made more than half of the volume images unusable.\n\nAgain, has anyone used an external dataset to improve CV or LB? Is there a way to avoid artifacts in the wbp conversion?\nI'd be happy to hear your thoughts and opinions, even if they're not solutions or tips!",
    "3093831": "I've tried to pretrain in raw, it helped at beggining. But didn't improved the final results. Also mention that the pretrained models were blind to experimental virus. May be because wbp format you say. I'm not sure what do you mean with that, the files I've downloaded were all zarr files. I've uploaded them toguether coordinates [here](https://www.kaggle.com/datasets/sacuscreed/czii-synth-data).",
    "3094199": "**Regarding point 3:** Converting wbp to denoised with a simple UNet and using it for training. This also did not improve CV.\nDid you manage to obtain decent results with the Unet? \n\nI also convert synthetic datasets to \"denoise\" versions but when incorporating this new data to the training, the model is just doing well in experimental datasets and is is not being able to understand what is happening in synthetic. I saw different distribution ranges in wbp raw synthetic and wbp raw experimentals :(\n\nI was also trying to find the model or algo for denoising the data (in paper they mentioned it is DenoisET) but no success so far. \n\nHere is an example of the \"denoise\" model. (values in the title are just mean and standard deviation of the image)\n- Noisy image corresponds to wbp raw\n- Output it is the model output\n- Clean image is the ground truth denoised image\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4238644%2Fbb183db52403b6379f1cdc2852f6c6df%2Fdownload.png?generation=1736619308388119&alt=media)",
    "3096898": "https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.full\n\n\"DenoisET implements the Noise2Noise algorithm37 for denoising cryoET tomograms. This ML algorithm learns to denoise imaging data from training on paired noisy measurements and is thus suitable for methods like cryoET in which the clean signal is difficult to realistically simulate and cannot be measured. However, because cryoET data are collected as a series of frames of the same underlying field-of-view but with different realizations of the stochastic noise process, training data can be readily generated by reconstructing paired tomograms after splitting the acquired frames into half-sets. Our implementation relies on a similar U-Net architecture as used in Topaz-Denoise38 and leverages the contrast transfer function-deconvolved tomograms produced by AreTomo326 to improve contrast enhancement. The increased SNR provided by denoising proved critical for our manual annotation efforts and benefited our machine learning labeling workflows as well (Fig. 1C).\"\n\nSo I assume DenoisET is not open.",
    "3098085": ""
  }
}