{
  "id": 550656,
  "title": "Questions on data preprocessing methods",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/550656",
  "author_name": "",
  "post_date": "2024-12-08T19:13:10.339739100Z",
  "votes": 6,
  "comment_count": 4,
  "views": 0,
  "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> : some questions on data preprocessing that I think many of us are facing, but are to my knowledge not answered yet.</p>\n<ul>\n<li><p>It seems the WBP/CTF-deconvolved/denoised/isonet data are steps of a sequential processing pipeline, in that order. This seems to indicate that we could make isonet data on test if we so desired, but not WBP and CTF-deconvolved. Is there a specific reason why only denoised is provided on test (or is that not actually the case)?</p></li>\n<li><p>The synthetic data provided seems to have none of these steps applied (or otherwise looks 'rawer' than the real data). Is that indeed the case? Can we easily apply these steps ourselves? The competition paper mentions packages to use, but without knowing how exactly to use these and what parameters to use reconstructing the flow is not really feasible.</p></li>\n</ul>",
  "messages": [
    {
      "id": "3067039",
      "postDate": "12/08/2024 19:13:10",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> : some questions on data preprocessing that I think many of us are facing, but are to my knowledge not answered yet.</p>\n<ul>\n<li><p>It seems the WBP/CTF-deconvolved/denoised/isonet data are steps of a sequential processing pipeline, in that order. This seems to indicate that we could make isonet data on test if we so desired, but not WBP and CTF-deconvolved. Is there a specific reason why only denoised is provided on test (or is that not actually the case)?</p></li>\n<li><p>The synthetic data provided seems to have none of these steps applied (or otherwise looks 'rawer' than the real data). Is that indeed the case? Can we easily apply these steps ourselves? The competition paper mentions packages to use, but without knowing how exactly to use these and what parameters to use reconstructing the flow is not really feasible.</p></li>\n</ul>",
      "rawMarkdown": "kharrington : some questions on data preprocessing that I think many of us are facing, but are to my knowledge not answered yet.\n\n- It seems the WBP/CTF-deconvolved/denoised/isonet data are steps of a sequential processing pipeline, in that order. This seems to indicate that we could make isonet data on test if we so desired, but not WBP and CTF-deconvolved. Is there a specific reason why only denoised is provided on test (or is that not actually the case)?\n\n- The synthetic data provided seems to have none of these steps applied (or otherwise looks 'rawer' than the real data). Is that indeed the case? Can we easily apply these steps ourselves? The competition paper mentions packages to use, but without knowing how exactly to use these and what parameters to use reconstructing the flow is not really feasible.",
      "votes": null
    },
    {
      "id": "3067092",
      "postDate": "12/08/2024 21:33:33",
      "content": "<p>instead of setting 'denoise' as the target, one can try to use 'WBP' as target and try the following experiments</p>\n<ol>\n<li>train and validate with phantom experimental data only</li>\n<li>train and validate with simulated experimental data only</li>\n<li>validate =  phantom, train = phantom+simulated</li>\n</ol>\n<p>you have to verify the effectiveness of simulated data first</p>",
      "rawMarkdown": "instead of setting 'denoise' as the target, one can try to use 'WBP' as target and try the following experiments\n1. train and validate with phantom experimental data only\n2. train and validate with simulated experimental data only\n3. validate =  phantom, train = phantom+simulated\n\nyou have to verify the effectiveness of simulated data first",
      "votes": null
    },
    {
      "id": "3067101",
      "postDate": "12/08/2024 21:56:49",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Phantom == WBP in your experiments 1,2, and 3, correct?</p>",
      "rawMarkdown": "hengck23 Phantom == WBP in your experiments 1,2, and 3, correct?",
      "votes": null
    },
    {
      "id": "3067104",
      "postDate": "12/08/2024 22:04:55",
      "content": "<p>phantom means kaggle data (synthetic cell, real captured tomograph).<br>\nsimulated means extra data downloaded (digital cell, digital captured tomograph).</p>",
      "rawMarkdown": "phantom means kaggle data (synthetic cell, real captured tomograph).\nsimulated means extra data downloaded (digital cell, digital captured tomograph).",
      "votes": null
    },
    {
      "id": "3067112",
      "postDate": "12/08/2024 22:34:07",
      "content": "<p>Ahh…  I didn't realize the training data was phantom samples…  Or that creating phantom samples as stand-ins for real tissue was something you could do.  Thanks!</p>\n<p><strong>Note:</strong>  For anyone else reading this perplexity.ai query \"how are phantom samples created in cryoET\" is pretty interesting.</p>",
      "rawMarkdown": "Ahh...  I didn't realize the training data was phantom samples...  Or that creating phantom samples as stand-ins for real tissue was something you could do.  Thanks!\n\n**Note:**  For anyone else reading this perplexity.ai query \"how are phantom samples created in cryoET\" is pretty interesting.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3067092,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "12/08/2024 21:33:33",
      "content": "<p>instead of setting 'denoise' as the target, one can try to use 'WBP' as target and try the following experiments</p>\n<ol>\n<li>train and validate with phantom experimental data only</li>\n<li>train and validate with simulated experimental data only</li>\n<li>validate =  phantom, train = phantom+simulated</li>\n</ol>\n<p>you have to verify the effectiveness of simulated data first</p>",
      "votes": null,
      "replies": [
        {
          "id": 3067101,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "12/08/2024 21:56:49",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Phantom == WBP in your experiments 1,2, and 3, correct?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3067104,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "12/08/2024 22:04:55",
              "content": "<p>phantom means kaggle data (synthetic cell, real captured tomograph).<br>\nsimulated means extra data downloaded (digital cell, digital captured tomograph).</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3067112,
                  "author_name": "davidlist",
                  "author_url": "",
                  "post_date": "12/08/2024 22:34:07",
                  "content": "<p>Ahh…  I didn't realize the training data was phantom samples…  Or that creating phantom samples as stand-ins for real tissue was something you could do.  Thanks!</p>\n<p><strong>Note:</strong>  For anyone else reading this perplexity.ai query \"how are phantom samples created in cryoET\" is pretty interesting.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067039": "kharrington : some questions on data preprocessing that I think many of us are facing, but are to my knowledge not answered yet.\n\n- It seems the WBP/CTF-deconvolved/denoised/isonet data are steps of a sequential processing pipeline, in that order. This seems to indicate that we could make isonet data on test if we so desired, but not WBP and CTF-deconvolved. Is there a specific reason why only denoised is provided on test (or is that not actually the case)?\n\n- The synthetic data provided seems to have none of these steps applied (or otherwise looks 'rawer' than the real data). Is that indeed the case? Can we easily apply these steps ourselves? The competition paper mentions packages to use, but without knowing how exactly to use these and what parameters to use reconstructing the flow is not really feasible.",
    "3067092": "instead of setting 'denoise' as the target, one can try to use 'WBP' as target and try the following experiments\n1. train and validate with phantom experimental data only\n2. train and validate with simulated experimental data only\n3. validate =  phantom, train = phantom+simulated\n\nyou have to verify the effectiveness of simulated data first",
    "3067101": "hengck23 Phantom == WBP in your experiments 1,2, and 3, correct?",
    "3067104": "phantom means kaggle data (synthetic cell, real captured tomograph).\nsimulated means extra data downloaded (digital cell, digital captured tomograph).",
    "3067112": "Ahh...  I didn't realize the training data was phantom samples...  Or that creating phantom samples as stand-ins for real tissue was something you could do.  Thanks!\n\n**Note:**  For anyone else reading this perplexity.ai query \"how are phantom samples created in cryoET\" is pretty interesting."
  },
  "source": "meta"
}