{
  "id": 544702,
  "title": "Welcome to the CryoET Object Identification Challenge!",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/544702",
  "author_name": "Kyle Harrington",
  "post_date": "2024-11-06T15:58:55.062000",
  "votes": 55,
  "comment_count": 48,
  "views": 0,
  "content": "<p><strong>Hello and welcome!</strong></p>\n<p>We are excited to invite you to join the CZII - CryoET Object Identification Challenge, where we need new cutting-edge machine learning algorithms to tackle a critical bottleneck in biomedical discovery—the annotation of complex 3D cellular images captured by cryo-electron tomography (cryoET).</p>\n<p><strong>About the Competition</strong><br>\nIn this challenge, your goal is to create machine learning algorithms capable of robustly annotating a diverse set of protein complexes and particles in 3D tomograms. CryoET offers near-atomic resolution of cellular components, providing unparalleled insight into how proteins function in their native environments. However, despite its power, the process of manually annotating these images remains slow and labor-intensive. With over 15,000 publicly available tomograms in the <a href=\"https://cryoetdataportal.czscience.com/\" target=\"_blank\">CryoET Data Portal</a> and only 5% annotated, there's an urgent need to automate this process.</p>\n<p><strong>Your Challenge</strong><br>\nYou'll be tasked with developing ML algorithms that can reliably annotate five classes of particles within hundreds of real-world tomograms. These particles vary greatly in size and shape, from virus-like particles to beta-galactosidase, and are similar to those found in crowded, complex cellular environments. Your algorithm will need to distinguish these particles from non-targets like nucleosomes, filaments, and membrane-bound proteins. To ensure your solution is generalizable, you’ll have access to a limited set of annotated tomograms for training, but you are welcome to incorporate additional synthetic or experimental data to enhance your model’s performance.</p>\n<p><strong>Why Does This Matter?</strong><br>\nProteins are essential to every aspect of cell function, and understanding their interactions within the cellular architecture is key to advancing human health. CryoET has the potential to unlock critical insights, but the lack of efficient annotation tools is a major barrier. By contributing to this challenge, you will help develop tools that could accelerate discoveries in cell biology and open new avenues for disease treatments. Moreover, your algorithms could be directly applied to the vast corpus of data on the CryoET Data Portal, enabling widespread use and impact beyond this competition.</p>\n<p><strong>The Impact</strong><br>\nIn addition to competing for one of ten cash prizes, your work will be part of a larger movement to revolutionize how scientists analyze cryoET data. Your algorithm could provide benchmarks for the field, improve data pipelines, and fuel new collaborations between the cryoET and machine learning communities. Ultimately, we believe that better tools for annotating cryoET data could lead to breakthroughs in cell biology—similar to how AlphaFold transformed the understanding of protein structures.</p>\n<p><strong>Join Us!</strong><br>\nWe are excited to see how you tackle this challenge and what innovative solutions you develop. Along with our team of hosts, we will be active on the Discussion boards, ready to assist and answer any questions you have. This is a unique opportunity to contribute to a rapidly growing field and make a lasting impact on biomedical science.</p>\n<p><strong>Resources</strong><br>\nWe have developed a number of resources to help folks who are interested in participating, and/or want to learn more about cryoET:</p>\n<ul>\n<li>Example notebooks (on <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\" target=\"_blank\">Github</a>, and on <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/code\" target=\"_blank\">Kaggle</a>)<ul>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/3d_unet_monai\" target=\"_blank\">MONAI-based 3D UNet</a></li>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/DeepFindET\" target=\"_blank\">DeepFindET</a></li>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/tomotwin_picking_notebook\" target=\"_blank\">TomoTwin</a></li></ul></li>\n<li><a href=\"https://cryoetdataportal.czscience.com/competition\" target=\"_blank\">CZ CryoET Data Portal resources</a> for this competition</li>\n<li><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1\" target=\"_blank\">Preprint about this competition and the dataset</a></li>\n<li><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621608v1\" target=\"_blank\">Preprint about open source tools to support competitors</a></li>\n</ul>\n<p>Good luck, and we can’t wait to see your solutions!</p>\n<p>Kyle and Reza</p>",
  "messages": [
    {
      "id": 3038120,
      "postDate": "2024-11-06T15:58:55.063Z",
      "content": "<p><strong>Hello and welcome!</strong></p>\n<p>We are excited to invite you to join the CZII - CryoET Object Identification Challenge, where we need new cutting-edge machine learning algorithms to tackle a critical bottleneck in biomedical discovery—the annotation of complex 3D cellular images captured by cryo-electron tomography (cryoET).</p>\n<p><strong>About the Competition</strong><br>\nIn this challenge, your goal is to create machine learning algorithms capable of robustly annotating a diverse set of protein complexes and particles in 3D tomograms. CryoET offers near-atomic resolution of cellular components, providing unparalleled insight into how proteins function in their native environments. However, despite its power, the process of manually annotating these images remains slow and labor-intensive. With over 15,000 publicly available tomograms in the <a href=\"https://cryoetdataportal.czscience.com/\" target=\"_blank\">CryoET Data Portal</a> and only 5% annotated, there's an urgent need to automate this process.</p>\n<p><strong>Your Challenge</strong><br>\nYou'll be tasked with developing ML algorithms that can reliably annotate five classes of particles within hundreds of real-world tomograms. These particles vary greatly in size and shape, from virus-like particles to beta-galactosidase, and are similar to those found in crowded, complex cellular environments. Your algorithm will need to distinguish these particles from non-targets like nucleosomes, filaments, and membrane-bound proteins. To ensure your solution is generalizable, you’ll have access to a limited set of annotated tomograms for training, but you are welcome to incorporate additional synthetic or experimental data to enhance your model’s performance.</p>\n<p><strong>Why Does This Matter?</strong><br>\nProteins are essential to every aspect of cell function, and understanding their interactions within the cellular architecture is key to advancing human health. CryoET has the potential to unlock critical insights, but the lack of efficient annotation tools is a major barrier. By contributing to this challenge, you will help develop tools that could accelerate discoveries in cell biology and open new avenues for disease treatments. Moreover, your algorithms could be directly applied to the vast corpus of data on the CryoET Data Portal, enabling widespread use and impact beyond this competition.</p>\n<p><strong>The Impact</strong><br>\nIn addition to competing for one of ten cash prizes, your work will be part of a larger movement to revolutionize how scientists analyze cryoET data. Your algorithm could provide benchmarks for the field, improve data pipelines, and fuel new collaborations between the cryoET and machine learning communities. Ultimately, we believe that better tools for annotating cryoET data could lead to breakthroughs in cell biology—similar to how AlphaFold transformed the understanding of protein structures.</p>\n<p><strong>Join Us!</strong><br>\nWe are excited to see how you tackle this challenge and what innovative solutions you develop. Along with our team of hosts, we will be active on the Discussion boards, ready to assist and answer any questions you have. This is a unique opportunity to contribute to a rapidly growing field and make a lasting impact on biomedical science.</p>\n<p><strong>Resources</strong><br>\nWe have developed a number of resources to help folks who are interested in participating, and/or want to learn more about cryoET:</p>\n<ul>\n<li>Example notebooks (on <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\" target=\"_blank\">Github</a>, and on <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/code\" target=\"_blank\">Kaggle</a>)<ul>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/3d_unet_monai\" target=\"_blank\">MONAI-based 3D UNet</a></li>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/DeepFindET\" target=\"_blank\">DeepFindET</a></li>\n<li><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/tomotwin_picking_notebook\" target=\"_blank\">TomoTwin</a></li></ul></li>\n<li><a href=\"https://cryoetdataportal.czscience.com/competition\" target=\"_blank\">CZ CryoET Data Portal resources</a> for this competition</li>\n<li><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1\" target=\"_blank\">Preprint about this competition and the dataset</a></li>\n<li><a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621608v1\" target=\"_blank\">Preprint about open source tools to support competitors</a></li>\n</ul>\n<p>Good luck, and we can’t wait to see your solutions!</p>\n<p>Kyle and Reza</p>",
      "rawMarkdown": "**Hello and welcome!**\n\nWe are excited to invite you to join the CZII - CryoET Object Identification Challenge, where we need new cutting-edge machine learning algorithms to tackle a critical bottleneck in biomedical discovery—the annotation of complex 3D cellular images captured by cryo-electron tomography (cryoET).\n\n**About the Competition**\nIn this challenge, your goal is to create machine learning algorithms capable of robustly annotating a diverse set of protein complexes and particles in 3D tomograms. CryoET offers near-atomic resolution of cellular components, providing unparalleled insight into how proteins function in their native environments. However, despite its power, the process of manually annotating these images remains slow and labor-intensive. With over 15,000 publicly available tomograms in the [CryoET Data Portal](https://cryoetdataportal.czscience.com/) and only 5% annotated, there's an urgent need to automate this process.\n\n**Your Challenge**\nYou'll be tasked with developing ML algorithms that can reliably annotate five classes of particles within hundreds of real-world tomograms. These particles vary greatly in size and shape, from virus-like particles to beta-galactosidase, and are similar to those found in crowded, complex cellular environments. Your algorithm will need to distinguish these particles from non-targets like nucleosomes, filaments, and membrane-bound proteins. To ensure your solution is generalizable, you’ll have access to a limited set of annotated tomograms for training, but you are welcome to incorporate additional synthetic or experimental data to enhance your model’s performance.\n\n**Why Does This Matter?**\nProteins are essential to every aspect of cell function, and understanding their interactions within the cellular architecture is key to advancing human health. CryoET has the potential to unlock critical insights, but the lack of efficient annotation tools is a major barrier. By contributing to this challenge, you will help develop tools that could accelerate discoveries in cell biology and open new avenues for disease treatments. Moreover, your algorithms could be directly applied to the vast corpus of data on the CryoET Data Portal, enabling widespread use and impact beyond this competition.\n\n**The Impact**\nIn addition to competing for one of ten cash prizes, your work will be part of a larger movement to revolutionize how scientists analyze cryoET data. Your algorithm could provide benchmarks for the field, improve data pipelines, and fuel new collaborations between the cryoET and machine learning communities. Ultimately, we believe that better tools for annotating cryoET data could lead to breakthroughs in cell biology—similar to how AlphaFold transformed the understanding of protein structures.\n\n**Join Us!**\nWe are excited to see how you tackle this challenge and what innovative solutions you develop. Along with our team of hosts, we will be active on the Discussion boards, ready to assist and answer any questions you have. This is a unique opportunity to contribute to a rapidly growing field and make a lasting impact on biomedical science.\n\n**Resources**\nWe have developed a number of resources to help folks who are interested in participating, and/or want to learn more about cryoET:\n- Example notebooks (on [Github](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks), and on [Kaggle](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/code))\n    - [MONAI-based 3D UNet](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/3d_unet_monai)\n    - [DeepFindET](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/DeepFindET)\n    - [TomoTwin](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/tomotwin_picking_notebook)\n- [CZ CryoET Data Portal resources](https://cryoetdataportal.czscience.com/competition) for this competition\n- [Preprint about this competition and the dataset](https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1)\n- [Preprint about open source tools to support competitors](https://www.biorxiv.org/content/10.1101/2024.11.04.621608v1)\n\nGood luck, and we can’t wait to see your solutions!\n\nKyle and Reza\n",
      "votes": 54
    },
    {
      "id": 3038451,
      "postDate": "2024-11-07T00:24:29.507Z",
      "content": "<p>Hi Kyle,</p>\n<p>I am wondering if you could explain in more detail the reasoning behind only including 7 training files and reserving ~500 for the test set. Are you looking to find out whether synthetic data can reliably used train deep learning models on this problem? </p>\n<p>Looking through CryoET Data Portal I noticed there is already some simulated data for this competition: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a>. It might be helpful to have info on the availability of extra data like this more visible. </p>\n<p>Also if there is additional experimental data that could benefit from more labeling effort elevating its visibility would be very helpful.</p>\n<p>I was about to walk away from this competition due to small training dataset. Others might respond similarly to seeing the training set size. Surfacing any relevant synthetic data, scripts, or other resources to increase amount of training data would be helpful as that seems critical for this competition. </p>\n<p>From the description of the competition it seems to be focused on development of new ML methods for this problem but with the set up of the train and test data it appears to be more about getting more data or using synthetic data for the problem. I am interested to know which is the actual challenge currently?</p>",
      "rawMarkdown": "Hi Kyle,\n\nI am wondering if you could explain in more detail the reasoning behind only including 7 training files and reserving ~500 for the test set. Are you looking to find out whether synthetic data can reliably used train deep learning models on this problem? \n\nLooking through CryoET Data Portal I noticed there is already some simulated data for this competition: https://cryoetdataportal.czscience.com/datasets/10441. It might be helpful to have info on the availability of extra data like this more visible. \n\nAlso if there is additional experimental data that could benefit from more labeling effort elevating its visibility would be very helpful.\n\nI was about to walk away from this competition due to small training dataset. Others might respond similarly to seeing the training set size. Surfacing any relevant synthetic data, scripts, or other resources to increase amount of training data would be helpful as that seems critical for this competition. \n\nFrom the description of the competition it seems to be focused on development of new ML methods for this problem but with the set up of the train and test data it appears to be more about getting more data or using synthetic data for the problem. I am interested to know which is the actual challenge currently?",
      "votes": 5,
      "replies": [
        {
          "id": 3038463,
          "postDate": "2024-11-07T01:08:54.277Z",
          "content": "<p>Hi Devin, Thanks for your comment. Kyle will chime in too. Our team found out with deepfindET (notebook available) that with even a small training dataset we were getting reasonable performance from the model. This was done by incrementally reducing the size of the training set and monitoring the model performance. We reasoned that better models would do better with similar or smaller training sets compared to the minimum size we found. The data is homogenous enough such that these 7 tomograms represent the data very well. We are also trying to recreate a real-world situation where a cryoET researcher can only afford to make a small training set. This is because getting good ground truth for cryoET is very time-consuming. So the aim of the challenge is not data generation, but models better than what's currently available. You have correctly identified the cryoET data portal as a good source of additional training data. Please let us know if you have more questions. </p>",
          "rawMarkdown": "Hi Devin, Thanks for your comment. Kyle will chime in too. Our team found out with deepfindET (notebook available) that with even a small training dataset we were getting reasonable performance from the model. This was done by incrementally reducing the size of the training set and monitoring the model performance. We reasoned that better models would do better with similar or smaller training sets compared to the minimum size we found. The data is homogenous enough such that these 7 tomograms represent the data very well. We are also trying to recreate a real-world situation where a cryoET researcher can only afford to make a small training set. This is because getting good ground truth for cryoET is very time-consuming. So the aim of the challenge is not data generation, but models better than what's currently available. You have correctly identified the cryoET data portal as a good source of additional training data. Please let us know if you have more questions. ",
          "votes": 8,
          "replies": [
            {
              "id": 3038465,
              "postDate": "2024-11-07T01:29:51.387Z",
              "content": "<p>Hi Reza, Thanks for the info! The fact that you got reasonable performance on small datasets alleviates most of my concerns, and I can see the value in simulating real world data availability. It sounds like you have more then enough data to run this competition. </p>",
              "rawMarkdown": "Hi Reza, Thanks for the info! The fact that you got reasonable performance on small datasets alleviates most of my concerns, and I can see the value in simulating real world data availability. It sounds like you have more then enough data to run this competition. "
            },
            {
              "id": 3038925,
              "postDate": "2024-11-07T14:12:05.040Z",
              "content": "<p>Exactly as <a href=\"https://www.kaggle.com/rezaparaan\" target=\"_blank\">@rezaparaan</a> said, and beyond generally using the cryoET data portal for extra training data, we also have provided some synthetic data that matches the competition: </p>\n<p><a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a></p>\n<p>It is also possible to generate more synthetic data.</p>",
              "rawMarkdown": "Exactly as @rezaparaan said, and beyond generally using the cryoET data portal for extra training data, we also have provided some synthetic data that matches the competition: \n\nhttps://cryoetdataportal.czscience.com/datasets/10441\n\nIt is also possible to generate more synthetic data."
            }
          ]
        }
      ]
    },
    {
      "id": 3040316,
      "postDate": "2024-11-09T02:36:39.247Z",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> <br>\nhi, would it be possible for the host to release the code for the synthetic data generator?</p>\n<p>i want to build a foundation model using synthetic data. I would need to find a way to:</p>\n<ul>\n<li>mimic actual noise, artifacts for tomography images</li>\n<li>as a foundation model, I want to detect more particle objects (apart from the 6 targets of this competition) at pertaining satge.<br>\n(Then i will investigate if we can later do few shot transfer learning using just a few real data from the kaggle train dataset with the foundation model)</li>\n</ul>",
      "rawMarkdown": "@kharrington \nhi, would it be possible for the host to release the code for the synthetic data generator?\n\ni want to build a foundation model using synthetic data. I would need to find a way to:\n- mimic actual noise, artifacts for tomography images\n- as a foundation model, I want to detect more particle objects (apart from the 6 targets of this competition) at pertaining satge.\n(Then i will investigate if we can later do few shot transfer learning using just a few real data from the kaggle train dataset with the foundation model)",
      "votes": 4,
      "replies": [
        {
          "id": 3041698,
          "postDate": "2024-11-10T17:57:41.063Z",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> yes, it is available here: <a href=\"https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py\" target=\"_blank\">https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py</a></p>",
          "rawMarkdown": "@hengck23 yes, it is available here: https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py",
          "votes": 6,
          "replies": [
            {
              "id": 3041730,
              "postDate": "2024-11-10T18:49:07.550Z",
              "content": "<p>thanks a lot! i will use it!</p>",
              "rawMarkdown": "thanks a lot! i will use it!"
            }
          ]
        }
      ]
    },
    {
      "id": 3038745,
      "postDate": "2024-11-07T09:55:46.107Z",
      "content": "<p>Could you confirm that the 7 images in the training dataset and the 500 (?) images in the test dataset are real images that have been hand-labeled? Or are they for example synthetically generated?</p>\n<p>If they are hand-labeled, do you have any indication of possible labeling errors? For example, how different do the labels tend to be if two different people or teams do the labeling?</p>",
      "rawMarkdown": "Could you confirm that the 7 images in the training dataset and the 500 (?) images in the test dataset are real images that have been hand-labeled? Or are they for example synthetically generated?\n\nIf they are hand-labeled, do you have any indication of possible labeling errors? For example, how different do the labels tend to be if two different people or teams do the labeling?",
      "votes": 2,
      "replies": [
        {
          "id": 3038910,
          "postDate": "2024-11-07T14:06:55.867Z",
          "content": "<p>Correct, there are 7 training samples here, and 500 test images. Both training and test are real images that have been hand labeled. We also have synthetic tomograms available: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a> and more synthetic tomograms can be generated. I'll add a better explanation about <em>why</em> the dataset is split like this.</p>\n<p>The real tomograms are hand labeled by domain specialists in cryoET. The data is really challenging to annotate because of the noise and particle sizes. We found that manual labeling errors tended to be underannotation (e.g. some particles were not labeled because the annotators couldn't find them visually). </p>",
          "rawMarkdown": "Correct, there are 7 training samples here, and 500 test images. Both training and test are real images that have been hand labeled. We also have synthetic tomograms available: https://cryoetdataportal.czscience.com/datasets/10441 and more synthetic tomograms can be generated. I'll add a better explanation about *why* the dataset is split like this.\n\nThe real tomograms are hand labeled by domain specialists in cryoET. The data is really challenging to annotate because of the noise and particle sizes. We found that manual labeling errors tended to be underannotation (e.g. some particles were not labeled because the annotators couldn't find them visually). ",
          "votes": 5,
          "replies": [
            {
              "id": 3039501,
              "postDate": "2024-11-08T05:41:21.817Z",
              "content": "<p>Hi Kyle, <br>\nSo to confirm, underannotations can also be present in the train data, right?. <br>\nIf yes, is there any way to handle it without manually doing it ?</p>",
              "rawMarkdown": "Hi Kyle, \nSo to confirm, underannotations can also be present in the train data, right?. \nIf yes, is there any way to handle it without manually doing it ?"
            },
            {
              "id": 3039963,
              "postDate": "2024-11-08T14:41:28.080Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/arunimbasak\" target=\"_blank\">@arunimbasak</a> </p>\n<p>To be clear, we have been able to train useful models with this training data as is. There are some methods in the literature that you could try for training with partially annotated data as well. We do not believe that it is necessary to do extra annotation on the training data (but it is possible that some folks could get a little boost from doing that).</p>\n<p>For methods that require perfectly annotated training data, you can use the synthetic data we provide: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310</a></p>\n<p>Cheers!</p>",
              "rawMarkdown": "Hi @arunimbasak \n\nTo be clear, we have been able to train useful models with this training data as is. There are some methods in the literature that you could try for training with partially annotated data as well. We do not believe that it is necessary to do extra annotation on the training data (but it is possible that some folks could get a little boost from doing that).\n\nFor methods that require perfectly annotated training data, you can use the synthetic data we provide: https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310\n\nCheers!",
              "votes": 2
            },
            {
              "id": 3040698,
              "postDate": "2024-11-09T13:17:40.180Z",
              "content": "<p>Yeah, that's also make sense. <br>\nThanks for providing the synthetic data. That would be helpful.</p>",
              "rawMarkdown": "Yeah, that's also make sense. \nThanks for providing the synthetic data. That would be helpful."
            },
            {
              "id": 3041767,
              "postDate": "2024-11-10T20:16:26.783Z",
              "content": "<p>Hi! This might be a totally unnecessary comment but I'll add it anyway 😅</p>\n<p>I think this can be taken care of in two main ways:</p>\n<ol>\n<li><p>The authors of the competition already took a step to minimize the effect of underannotation on evaluating our model performance; by using F4 score instead of F1 or AUC. This means that if you have a clever model that detects particles that were missed by the human annotators, it will not be penalized that much by the competition metric. On the other hand, if you have a model that misses an almost-surely identified particle by the human annotators, it will be heavily penalized.</p></li>\n<li><p>A possible solution that you can investigate is trying a psudolabeling approach, meaning you can try to train a model A to try and label the unannotated particles for you, and take its confident labels and append them to the ground-truth annotations. This sometimes improves model performance, and I think it's worth trying.</p></li>\n</ol>",
              "rawMarkdown": "Hi! This might be a totally unnecessary comment but I'll add it anyway 😅\n\nI think this can be taken care of in two main ways:\n\n1. The authors of the competition already took a step to minimize the effect of underannotation on evaluating our model performance; by using F4 score instead of F1 or AUC. This means that if you have a clever model that detects particles that were missed by the human annotators, it will not be penalized that much by the competition metric. On the other hand, if you have a model that misses an almost-surely identified particle by the human annotators, it will be heavily penalized.\n \n2. A possible solution that you can investigate is trying a psudolabeling approach, meaning you can try to train a model A to try and label the unannotated particles for you, and take its confident labels and append them to the ground-truth annotations. This sometimes improves model performance, and I think it's worth trying.",
              "votes": 4
            },
            {
              "id": 3052576,
              "postDate": "2024-11-22T15:10:23.753Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> , do you have any idea how \"under-labelled\" the dataset is? </p>\n<p>I.e. is there a way to estimate what fraction of particles are labelled vs. unlabelled (at least on average?). </p>\n<p>Would this differ per particle type? </p>",
              "rawMarkdown": "Hi @kharrington , do you have any idea how \"under-labelled\" the dataset is? \n\nI.e. is there a way to estimate what fraction of particles are labelled vs. unlabelled (at least on average?). \n\nWould this differ per particle type? \n\n"
            },
            {
              "id": 3052758,
              "postDate": "2024-11-22T19:53:02.290Z",
              "content": "<p>We do not have an estimate of how \"under-labelled\" the dataset is. </p>\n<p>The number of unlabelled particles is expected to be different per particle type. The larger particles were easier to label. </p>\n<p>We've discussed some of the challenges around this in the paper: <a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.abstract\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.abstract</a> including Table 1 which says the average count of particle per tomogram. </p>\n<p>However, the variability is large, so it isn't possible to say something like: \"Table 1 says 6 VLP particles on average per tomogram. I have only found 1 in this tomogram, so I should look for 5 more then stop.\" It can be meaningful to use the average particle count numbers across all predictions, but I would only use that information loosely.</p>",
              "rawMarkdown": "We do not have an estimate of how \"under-labelled\" the dataset is. \n\nThe number of unlabelled particles is expected to be different per particle type. The larger particles were easier to label. \n\nWe've discussed some of the challenges around this in the paper: https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.abstract including Table 1 which says the average count of particle per tomogram. \n\nHowever, the variability is large, so it isn't possible to say something like: \"Table 1 says 6 VLP particles on average per tomogram. I have only found 1 in this tomogram, so I should look for 5 more then stop.\" It can be meaningful to use the average particle count numbers across all predictions, but I would only use that information loosely.",
              "votes": 3
            },
            {
              "id": 3052762,
              "postDate": "2024-11-22T19:58:22.280Z",
              "content": "<p>Thanks!</p>\n<p>All good, averages is what I was after!</p>",
              "rawMarkdown": "Thanks!\n\nAll good, averages is what I was after!"
            }
          ]
        }
      ]
    },
    {
      "id": 3038246,
      "postDate": "2024-11-06T18:47:07.423Z",
      "content": "<p>Dear Kyle, <br>\nThe links to the GitHub notebooks are not working.</p>",
      "rawMarkdown": "Dear Kyle, \nThe links to the GitHub notebooks are not working.",
      "votes": 2,
      "replies": [
        {
          "id": 3038454,
          "postDate": "2024-11-07T00:40:54.973Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/taherehzarratehsan\" target=\"_blank\">@taherehzarratehsan</a>, sorry about that, we had quite a few things to coordinate to make resources public on the launch day. The notebooks repo on GH should be available now.</p>",
          "rawMarkdown": "Hi @taherehzarratehsan, sorry about that, we had quite a few things to coordinate to make resources public on the launch day. The notebooks repo on GH should be available now.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3121174,
      "postDate": "2025-02-11T10:42:55.323Z",
      "content": "<p>Hi Kyle,</p>\n<p>I wonder if the runtime limitation is also for scoring the test set. I have a submission that runs successfully on the fake test set of the 3 tomograms however it is not scored for the competition and throws runtime error. Does the notebook need to run on the test set also on 12 hours? <br>\nThank you for your support I am just trying to still get my score now so that I can estimate the performance.</p>",
      "rawMarkdown": "Hi Kyle,\n\nI wonder if the runtime limitation is also for scoring the test set. I have a submission that runs successfully on the fake test set of the 3 tomograms however it is not scored for the competition and throws runtime error. Does the notebook need to run on the test set also on 12 hours? \nThank you for your support I am just trying to still get my score now so that I can estimate the performance.\n"
    },
    {
      "id": 3095270,
      "postDate": "2025-01-13T06:25:14.273Z",
      "content": "<p>Hello Kyle or Reza,<br>\nCan you please explain the labels?</p>\n<p>From what I understand, each tomogram is just a space with some objects it in, and we are trying to identify those objects and what type of objects they are. However, the labels (x, y, z) are real-valued (not integers), and have coordinates like (2000, 700, 1000) when the tomograms are only of size (630, 630, 184).<br>\nSo why are the labels out of space?<br>\nWhy are the labels just coordinates and not maybe a segmentation mask or bounding cubes?</p>",
      "rawMarkdown": "Hello Kyle or Reza,\nCan you please explain the labels?\n\nFrom what I understand, each tomogram is just a space with some objects it in, and we are trying to identify those objects and what type of objects they are. However, the labels (x, y, z) are real-valued (not integers), and have coordinates like (2000, 700, 1000) when the tomograms are only of size (630, 630, 184).\nSo why are the labels out of space?\nWhy are the labels just coordinates and not maybe a segmentation mask or bounding cubes?",
      "replies": [
        {
          "id": 3095796,
          "postDate": "2025-01-13T18:04:20.080Z",
          "content": "<p>Hello. The tomograms are made of voxels and their dimensions are the number of voxels. The coordinates are measured in units of distance in Angstroms. The voxels have physical sizes that let you switch from voxel counting to distance counting, depending on which resolution of the tomogram you are working with. If you search through the discussions in the forum you will find more detailed answers. Thanks for your question.</p>",
          "rawMarkdown": "Hello. The tomograms are made of voxels and their dimensions are the number of voxels. The coordinates are measured in units of distance in Angstroms. The voxels have physical sizes that let you switch from voxel counting to distance counting, depending on which resolution of the tomogram you are working with. If you search through the discussions in the forum you will find more detailed answers. Thanks for your question."
        }
      ]
    },
    {
      "id": 3090982,
      "postDate": "2025-01-07T21:32:47.843Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> ,  for submission, does it have to be CPU notebook ( and run within 12 hours)? or either CPU or GPU notebook is fine?</p>\n<p>Thanks!!</p>",
      "rawMarkdown": "Hi @kharrington ,  for submission, does it have to be CPU notebook ( and run within 12 hours)? or either CPU or GPU notebook is fine?\n\nThanks!!"
    },
    {
      "id": 3072798,
      "postDate": "2024-12-15T16:33:51.273Z",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> Please excuse my ignorance, is the missing wedge artifact present in the dataset? I could not figure it out from the competition paper.</p>",
      "rawMarkdown": "@kharrington Please excuse my ignorance, is the missing wedge artifact present in the dataset? I could not figure it out from the competition paper."
    },
    {
      "id": 3071571,
      "postDate": "2024-12-14T00:24:22.613Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> ,</p>\n<p>A question about how CryoET works. So if I have a sample that is  1k nm (length) x 1k nm (width) x 200 nm (thick, z-axis). Do CryoET only able to scan the surface of this Rectangular cuboid? Or it can penetrate into the middle of the sample?<br>\nThe reason I am asking is because I see in the reference (<a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full)\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full)</a>, the thickness of the sample is about 200nm. if CryoET can only image the surface of the sample, we can't get sample that has resolution &lt; 200nm, at least for z-axis. Is that correct?</p>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @kharrington ,\n\nA question about how CryoET works. So if I have a sample that is  1k nm (length) x 1k nm (width) x 200 nm (thick, z-axis). Do CryoET only able to scan the surface of this Rectangular cuboid? Or it can penetrate into the middle of the sample?\nThe reason I am asking is because I see in the reference (https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full), the thickness of the sample is about 200nm. if CryoET can only image the surface of the sample, we can't get sample that has resolution < 200nm, at least for z-axis. Is that correct?\n\nThanks!\n",
      "replies": [
        {
          "id": 3071594,
          "postDate": "2024-12-14T01:37:41.037Z",
          "content": "<p>Hi Hui Liu, Indeed cryoET is a volume imaging method not a surface imaging method.  It is analogous to CT and MRI imaging.  In those methods you take projection images through a solid object (e.g. a brain) and use computational methods to reconstruct a full 3D volume (not a surface) of the entire object.  Some other details are at: <a href=\"https://en.wikipedia.org/wiki/Cryogenic_electron_tomography\" target=\"_blank\">https://en.wikipedia.org/wiki/Cryogenic_electron_tomography</a> and also here: <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/about.html\" target=\"_blank\">https://chanzuckerberg.github.io/cryoet-data-portal/about.html</a>    Thanks for your interest in this challenge, we are very eager to encourage the participation of this community.  Good luck and happy to answer further questions.  Bridget</p>",
          "rawMarkdown": "Hi Hui Liu, Indeed cryoET is a volume imaging method not a surface imaging method.  It is analogous to CT and MRI imaging.  In those methods you take projection images through a solid object (e.g. a brain) and use computational methods to reconstruct a full 3D volume (not a surface) of the entire object.  Some other details are at: https://en.wikipedia.org/wiki/Cryogenic_electron_tomography and also here: https://chanzuckerberg.github.io/cryoet-data-portal/about.html    Thanks for your interest in this challenge, we are very eager to encourage the participation of this community.  Good luck and happy to answer further questions.  Bridget"
        }
      ]
    },
    {
      "id": 3069073,
      "postDate": "2024-12-11T03:09:59.927Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> ,</p>\n<ol>\n<li><p>It seems like the predictions should be the center of particle. May I know how the center is defined? Suppose I know the 3D segmentation of the particle. Is the center coordinate defined by the average of coordinates of all the points segmented to be part of the particle?</p></li>\n<li><p>What is the value range of tomogram in zarr? And does the value represent something like light intensity?  In other words, does higher value means background (water), and lower value means some objects. </p></li>\n</ol>\n<p>Thanks!</p>",
      "rawMarkdown": "Hi @kharrington ,\n\n1. It seems like the predictions should be the center of particle. May I know how the center is defined? Suppose I know the 3D segmentation of the particle. Is the center coordinate defined by the average of coordinates of all the points segmented to be part of the particle?\n\n2. What is the value range of tomogram in zarr? And does the value represent something like light intensity?  In other words, does higher value means background (water), and lower value means some objects. \n\nThanks!",
      "replies": [
        {
          "id": 3069386,
          "postDate": "2024-12-11T12:34:04.650Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/huiliu0\" target=\"_blank\">@huiliu0</a> </p>\n<ol>\n<li><p>The predictions can be thought of as centers of particles, but were curated by many approaches including humans manually placing some coordinates. Note that our detection threshold is 0.5x the radius of the particle, so slight differences in how the center coordinate is calculated should not have a significant effect on performance.</p></li>\n<li><p>The images are taken using electron microscopy instead of light, which most of us are familiar with. While electron microscopy is a complex technique, the basic idea is: an electron beam is directed at the sample, and the denser the material at a particular location, the more it scatters the electrons. In our reconstruction, regions with higher density appear darker. Proteins and other biological structures are denser than the surrounding embedding medium, allowing us to visualize them in the tomogram.</p></li>\n</ol>\n<p>Cheers,<br>\nKyle</p>",
          "rawMarkdown": "Hi @huiliu0 \n\n1. The predictions can be thought of as centers of particles, but were curated by many approaches including humans manually placing some coordinates. Note that our detection threshold is 0.5x the radius of the particle, so slight differences in how the center coordinate is calculated should not have a significant effect on performance.\n\n2. The images are taken using electron microscopy instead of light, which most of us are familiar with. While electron microscopy is a complex technique, the basic idea is: an electron beam is directed at the sample, and the denser the material at a particular location, the more it scatters the electrons. In our reconstruction, regions with higher density appear darker. Proteins and other biological structures are denser than the surrounding embedding medium, allowing us to visualize them in the tomogram.\n\nCheers,\nKyle",
          "votes": 1
        }
      ]
    },
    {
      "id": 3058840,
      "postDate": "2024-11-30T03:58:15.243Z",
      "content": "<p>I think I've found an error in the co-pick utils code. Line 67 in <a href=\"https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/segmentation_from_picks.py\" target=\"_blank\">https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/segmentation_from_picks.py</a></p>\n<p><code>cx, cy, cz = pick.location.z / voxel_spacing, pick.location.y / voxel_spacing, pick.location.x / voxel_spacing</code></p>\n<p>Shouldn't this be:</p>\n<p><code>cx, cy, cz = pick.location.x / voxel_spacing, pick.location.y / voxel_spacing, pick.location.z / voxel_spacing</code></p>",
      "rawMarkdown": "I think I've found an error in the co-pick utils code. Line 67 in https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/segmentation_from_picks.py\n\n`cx, cy, cz = pick.location.z / voxel_spacing, pick.location.y / voxel_spacing, pick.location.x / voxel_spacing`\n\nShouldn't this be:\n\n`cx, cy, cz = pick.location.x / voxel_spacing, pick.location.y / voxel_spacing, pick.location.z / voxel_spacing`\n",
      "replies": [
        {
          "id": 3066672,
          "postDate": "2024-12-08T11:03:17.340Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a> <br>\nI too have faced an error (shape cannot be negative) here, that is due to the <em>voxel_spacing</em> that is being passed to <em>from_picks</em> method. It is taking the value of 10 for all the scales, so for lower scales of (92, 315, 315) and (46, 158, 158) it should be ~20 and ~40 respectively<br>\n<code>'datasets': [{'coordinateTransformations': [{'scale': [10.012444196428572,\n                                                                         10.012444196428572,\n                                                                         10.012444537618887],\n                                                               'type': 'scale'}],</code><br>\nSo by fixing this  line I got it working. But for lower levels the masks of tiny particles are completely missing</p>",
          "rawMarkdown": "Hi @homiecal \nI too have faced an error (shape cannot be negative) here, that is due to the *voxel_spacing* that is being passed to *from_picks* method. It is taking the value of 10 for all the scales, so for lower scales of (92, 315, 315) and (46, 158, 158) it should be ~20 and ~40 respectively\n`'datasets': [{'coordinateTransformations': [{'scale': [10.012444196428572,\n                                                                         10.012444196428572,\n                                                                         10.012444537618887],\n                                                               'type': 'scale'}],`\nSo by fixing this  line I got it working. But for lower levels the masks of tiny particles are completely missing",
          "replies": [
            {
              "id": 3066674,
              "postDate": "2024-12-08T11:04:21.107Z",
              "content": "<p><a href=\"https://github.com/copick/copick-utils/blob/8936b5b9379bf4f58d49ac58a9669076baeeff5b/src/copick_utils/segmentation/segmentation_from_picks.py#L151\" target=\"_blank\">https://github.com/copick/copick-utils/blob/8936b5b9379bf4f58d49ac58a9669076baeeff5b/src/copick_utils/segmentation/segmentation_from_picks.py#L151</a></p>",
              "rawMarkdown": "https://github.com/copick/copick-utils/blob/8936b5b9379bf4f58d49ac58a9669076baeeff5b/src/copick_utils/segmentation/segmentation_from_picks.py#L151\n"
            },
            {
              "id": 3066976,
              "postDate": "2024-12-08T17:08:28.887Z",
              "content": "<p><a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a> and <a href=\"https://www.kaggle.com/bharat0\" target=\"_blank\">@bharat0</a> thank you for pointing this out. Sorry about the delay, I'm still figuring out how to get enough notifications from Kaggle. </p>\n<p>I've started the fix here: <a href=\"https://github.com/copick/copick-utils/pull/9\" target=\"_blank\">https://github.com/copick/copick-utils/pull/9</a></p>",
              "rawMarkdown": "@homiecal and @bharat0 thank you for pointing this out. Sorry about the delay, I'm still figuring out how to get enough notifications from Kaggle. \n\nI've started the fix here: https://github.com/copick/copick-utils/pull/9",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3058350,
      "postDate": "2024-11-29T11:23:10.150Z",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>  <br>\nhi, Does the competition provide both noisy and noise-free tomographic images?<br>\nI would like to incorporate denoising training into the network, but I do not have denoised labels for the training data.</p>",
      "rawMarkdown": "@kharrington  \nhi, Does the competition provide both noisy and noise-free tomographic images?\nI would like to incorporate denoising training into the network, but I do not have denoised labels for the training data.\n",
      "replies": [
        {
          "id": 3058820,
          "postDate": "2024-11-30T02:55:16.027Z",
          "content": "<p><a href=\"https://www.kaggle.com/liuxiaoonline\" target=\"_blank\">@liuxiaoonline</a> </p>\n<p>We actually provide synthetic data which might be what you are looking for: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310</a></p>",
          "rawMarkdown": "@liuxiaoonline \n\nWe actually provide synthetic data which might be what you are looking for: https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310",
          "votes": 1
        },
        {
          "id": 3058841,
          "postDate": "2024-11-30T04:01:45.617Z",
          "content": "<p><a href=\"https://www.kaggle.com/liuxiaoonline\" target=\"_blank\">@liuxiaoonline</a> can you please help me understand what you mean by denoised labels? In the dataset we are provided with picks, and denoised tomograms, the picks are set as our ground truth, are these not denoised labels?</p>",
          "rawMarkdown": "@liuxiaoonline can you please help me understand what you mean by denoised labels? In the dataset we are provided with picks, and denoised tomograms, the picks are set as our ground truth, are these not denoised labels?"
        }
      ]
    },
    {
      "id": 3045115,
      "postDate": "2024-11-14T07:23:20.200Z",
      "content": "<p>Hi! I am new to the field of machine learning. I have a question related to obtaining images through CryoET. How do you collect data by slicing or rotating the specimen?</p>",
      "rawMarkdown": "Hi! I am new to the field of machine learning. I have a question related to obtaining images through CryoET. How do you collect data by slicing or rotating the specimen?",
      "replies": [
        {
          "id": 3058819,
          "postDate": "2024-11-30T02:53:55.357Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rustambazarbayev\" target=\"_blank\">@rustambazarbayev</a> </p>\n<p>Check out this page: <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/intro_cryoet.html#cryoet-intro\" target=\"_blank\">https://chanzuckerberg.github.io/cryoet-data-portal/intro_cryoet.html#cryoet-intro</a> it covers a lot of the fundamentals of CryoET</p>",
          "rawMarkdown": "Hi @rustambazarbayev \n\nCheck out this page: https://chanzuckerberg.github.io/cryoet-data-portal/intro_cryoet.html#cryoet-intro it covers a lot of the fundamentals of CryoET",
          "votes": 2
        }
      ]
    },
    {
      "id": 3042134,
      "postDate": "2024-11-11T08:52:41.113Z",
      "content": "<p>Appreciate the support.</p>",
      "rawMarkdown": "Appreciate the support."
    },
    {
      "id": 3038483,
      "postDate": "2024-11-07T02:57:57.447Z",
      "content": "<p>Hi Kyle,</p>\n<p>What does benchmark.csv on the LB mean? Is it your best result so far that we need to surpass?</p>",
      "rawMarkdown": "Hi Kyle,\n\nWhat does benchmark.csv on the LB mean? Is it your best result so far that we need to surpass?",
      "replies": [
        {
          "id": 3038918,
          "postDate": "2024-11-07T14:09:33.793Z",
          "content": "<p>benchmark.csv are results we obtained with <a href=\"https://github.com/copick/DeepFindET/tree/main\" target=\"_blank\">DeepFindET</a>. We're hoping folks can get past this.</p>",
          "rawMarkdown": "benchmark.csv are results we obtained with [DeepFindET](https://github.com/copick/DeepFindET/tree/main). We're hoping folks can get past this.",
          "votes": 4,
          "replies": [
            {
              "id": 3040070,
              "postDate": "2024-11-08T16:49:55.483Z",
              "content": "<p>Hello Kyle!</p>\n<p>I find that inference using GPU P100 on Kaggle takes way longer than the times referenced in the example notebooks (I get &gt;100s per tomo vs. 25-30s in <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/inference.ipynb\" target=\"_blank\">DeepFindET Inference</a>). Extracting coordinates is also slower. Was the 12 hours run-time optimized for?</p>\n<p>Thanks!</p>",
              "rawMarkdown": "Hello Kyle!\n\nI find that inference using GPU P100 on Kaggle takes way longer than the times referenced in the example notebooks (I get >100s per tomo vs. 25-30s in [DeepFindET Inference](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/inference.ipynb)). Extracting coordinates is also slower. Was the 12 hours run-time optimized for?\n\nThanks!"
            },
            {
              "id": 3040075,
              "postDate": "2024-11-08T16:54:49.310Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>,</p>\n<p>Runtimes are definitely one of the challenges in this competition. I've had better luck using the T4s on Kaggle, but some optimizations are still needed to get a full DeepFindET run working (that is one of the reasons I posted the BlobDetector code). </p>\n<p>A couple suggestions that could help:</p>\n<ul>\n<li>figuring out what regions of the image can easily be skipped</li>\n<li>moving the segmentation mask-&gt;coordinate calculation to GPU</li>\n</ul>\n<p>If we get any new code improvements on our side that can help folks out, then we'll post them in the discussion here!</p>",
              "rawMarkdown": "Hi @andreizamfir,\n\nRuntimes are definitely one of the challenges in this competition. I've had better luck using the T4s on Kaggle, but some optimizations are still needed to get a full DeepFindET run working (that is one of the reasons I posted the BlobDetector code). \n\nA couple suggestions that could help:\n- figuring out what regions of the image can easily be skipped\n- moving the segmentation mask->coordinate calculation to GPU\n\nIf we get any new code improvements on our side that can help folks out, then we'll post them in the discussion here!",
              "votes": 3
            }
          ]
        },
        {
          "id": 3041861,
          "postDate": "2024-11-11T00:32:00.580Z",
          "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> </p>\n<p>Hi Kyle,  I think we're all suspecting there is something off about the evaluation implementation.  <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381\" target=\"_blank\">As per this discussion</a>.  Would it be possible for you to confirm if that benchmark value was produced by running DeepFindET directly in Kaggle (without the time limit), or have you evaluated it separately and just posted the score?   Thanks.</p>",
          "rawMarkdown": "@kharrington \n\nHi Kyle,  I think we're all suspecting there is something off about the evaluation implementation.  [As per this discussion](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381).  Would it be possible for you to confirm if that benchmark value was produced by running DeepFindET directly in Kaggle (without the time limit), or have you evaluated it separately and just posted the score?   Thanks.",
          "replies": [
            {
              "id": 3041923,
              "postDate": "2024-11-11T02:23:30.653Z",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a>, that benchmark was created by running DeepFindET externally. Some optimization needs to be done to get DeepFindET running within the compute constraints on Kaggle. </p>\n<p>The main things I think are worth trying now are getting proper multi-GPU with the T4 nodes on Kaggle, and getting the segmentation mask-&gt;position code to run on GPU instead of CPU.</p>",
              "rawMarkdown": "Hi @ollypowell, that benchmark was created by running DeepFindET externally. Some optimization needs to be done to get DeepFindET running within the compute constraints on Kaggle. \n\nThe main things I think are worth trying now are getting proper multi-GPU with the T4 nodes on Kaggle, and getting the segmentation mask->position code to run on GPU instead of CPU.",
              "votes": 1
            },
            {
              "id": 3044134,
              "postDate": "2024-11-13T04:40:55.140Z",
              "content": "<p>Thanks for confirming.  I think until the bug in the evaluation server is fixed I will hold off trying anything.  But eventually I'd like to use a combination of UNet for segmentation and something faster like your blob detector to save running it over lots of empty space.</p>",
              "rawMarkdown": "Thanks for confirming.  I think until the bug in the evaluation server is fixed I will hold off trying anything.  But eventually I'd like to use a combination of UNet for segmentation and something faster like your blob detector to save running it over lots of empty space."
            },
            {
              "id": 3044142,
              "postDate": "2024-11-13T05:04:40.760Z",
              "content": "<p>I too suspect that there is something wrong with the evaluation system. The best scoring model only gets a score of 0.003 which might as well be zero… </p>",
              "rawMarkdown": "I too suspect that there is something wrong with the evaluation system. The best scoring model only gets a score of 0.003 which might as well be zero... "
            },
            {
              "id": 3044887,
              "postDate": "2024-11-13T23:40:10.750Z",
              "content": "<p>hey <a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a>, </p>\n<p>We have identified the bug in evaluation and will update soon.</p>\n<p>thank you for sticking around!<br>\nKyle</p>",
              "rawMarkdown": "hey @ollypowell, \n\nWe have identified the bug in evaluation and will update soon.\n\nthank you for sticking around!\nKyle",
              "votes": 5
            },
            {
              "id": 3044906,
              "postDate": "2024-11-14T00:33:58.637Z",
              "content": "<p>Awesome news Kyle, thanks,  looking forward to getting stuck into the challenge!</p>",
              "rawMarkdown": "Awesome news Kyle, thanks,  looking forward to getting stuck into the challenge!"
            }
          ]
        }
      ]
    },
    {
      "id": 3073045,
      "postDate": "2024-12-16T01:28:21.733Z",
      "content": "<p>Given the limited annotated training dataset available for cryoET, are there existing standardized data augmentation methods specifically for cryoET data? Particularly for 3D image data, what are the most efficient ways to generate synthetic data or leverage existing datasets for data augmentation to ensure the model's generalization ability?</p>",
      "rawMarkdown": "Given the limited annotated training dataset available for cryoET, are there existing standardized data augmentation methods specifically for cryoET data? Particularly for 3D image data, what are the most efficient ways to generate synthetic data or leverage existing datasets for data augmentation to ensure the model's generalization ability?",
      "votes": 2,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3038451,
      "author_name": "Devin Anzelmo",
      "author_url": "",
      "post_date": "2024-11-07T00:24:29.507000",
      "content": "<p>Hi Kyle,</p>\n<p>I am wondering if you could explain in more detail the reasoning behind only including 7 training files and reserving ~500 for the test set. Are you looking to find out whether synthetic data can reliably used train deep learning models on this problem? </p>\n<p>Looking through CryoET Data Portal I noticed there is already some simulated data for this competition: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a>. It might be helpful to have info on the availability of extra data like this more visible. </p>\n<p>Also if there is additional experimental data that could benefit from more labeling effort elevating its visibility would be very helpful.</p>\n<p>I was about to walk away from this competition due to small training dataset. Others might respond similarly to seeing the training set size. Surfacing any relevant synthetic data, scripts, or other resources to increase amount of training data would be helpful as that seems critical for this competition. </p>\n<p>From the description of the competition it seems to be focused on development of new ML methods for this problem but with the set up of the train and test data it appears to be more about getting more data or using synthetic data for the problem. I am interested to know which is the actual challenge currently?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3038463,
          "author_name": "Reza Paraan",
          "author_url": "",
          "post_date": "2024-11-07T01:08:54.277000",
          "content": "<p>Hi Devin, Thanks for your comment. Kyle will chime in too. Our team found out with deepfindET (notebook available) that with even a small training dataset we were getting reasonable performance from the model. This was done by incrementally reducing the size of the training set and monitoring the model performance. We reasoned that better models would do better with similar or smaller training sets compared to the minimum size we found. The data is homogenous enough such that these 7 tomograms represent the data very well. We are also trying to recreate a real-world situation where a cryoET researcher can only afford to make a small training set. This is because getting good ground truth for cryoET is very time-consuming. So the aim of the challenge is not data generation, but models better than what's currently available. You have correctly identified the cryoET data portal as a good source of additional training data. Please let us know if you have more questions. </p>",
          "votes": 8,
          "replies": [
            {
              "id": 3038465,
              "author_name": "Devin Anzelmo",
              "author_url": "",
              "post_date": "2024-11-07T01:29:51.387000",
              "content": "<p>Hi Reza, Thanks for the info! The fact that you got reasonable performance on small datasets alleviates most of my concerns, and I can see the value in simulating real world data availability. It sounds like you have more then enough data to run this competition. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3038925,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-07T14:12:05.040000",
              "content": "<p>Exactly as <a href=\"https://www.kaggle.com/rezaparaan\" target=\"_blank\">@rezaparaan</a> said, and beyond generally using the cryoET data portal for extra training data, we also have provided some synthetic data that matches the competition: </p>\n<p><a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a></p>\n<p>It is also possible to generate more synthetic data.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3040316,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-09T02:36:39.247000",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> <br>\nhi, would it be possible for the host to release the code for the synthetic data generator?</p>\n<p>i want to build a foundation model using synthetic data. I would need to find a way to:</p>\n<ul>\n<li>mimic actual noise, artifacts for tomography images</li>\n<li>as a foundation model, I want to detect more particle objects (apart from the 6 targets of this competition) at pertaining satge.<br>\n(Then i will investigate if we can later do few shot transfer learning using just a few real data from the kaggle train dataset with the foundation model)</li>\n</ul>",
      "votes": 4,
      "replies": [
        {
          "id": 3041698,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-10T17:57:41.063000",
          "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> yes, it is available here: <a href=\"https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py\" target=\"_blank\">https://github.com/copick/copick-catalog/blob/main/solutions/polnet/generate-copick-project/solution.py</a></p>",
          "votes": 6,
          "replies": [
            {
              "id": 3041730,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-10T18:49:07.550000",
              "content": "<p>thanks a lot! i will use it!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3038745,
      "author_name": "Jeroen Cottaar",
      "author_url": "",
      "post_date": "2024-11-07T09:55:46.107000",
      "content": "<p>Could you confirm that the 7 images in the training dataset and the 500 (?) images in the test dataset are real images that have been hand-labeled? Or are they for example synthetically generated?</p>\n<p>If they are hand-labeled, do you have any indication of possible labeling errors? For example, how different do the labels tend to be if two different people or teams do the labeling?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3038910,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-07T14:06:55.867000",
          "content": "<p>Correct, there are 7 training samples here, and 500 test images. Both training and test are real images that have been hand labeled. We also have synthetic tomograms available: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441</a> and more synthetic tomograms can be generated. I'll add a better explanation about <em>why</em> the dataset is split like this.</p>\n<p>The real tomograms are hand labeled by domain specialists in cryoET. The data is really challenging to annotate because of the noise and particle sizes. We found that manual labeling errors tended to be underannotation (e.g. some particles were not labeled because the annotators couldn't find them visually). </p>",
          "votes": 5,
          "replies": [
            {
              "id": 3039501,
              "author_name": "Arunim Basak",
              "author_url": "",
              "post_date": "2024-11-08T05:41:21.817000",
              "content": "<p>Hi Kyle, <br>\nSo to confirm, underannotations can also be present in the train data, right?. <br>\nIf yes, is there any way to handle it without manually doing it ?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3039963,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-08T14:41:28.080000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/arunimbasak\" target=\"_blank\">@arunimbasak</a> </p>\n<p>To be clear, we have been able to train useful models with this training data as is. There are some methods in the literature that you could try for training with partially annotated data as well. We do not believe that it is necessary to do extra annotation on the training data (but it is possible that some folks could get a little boost from doing that).</p>\n<p>For methods that require perfectly annotated training data, you can use the synthetic data we provide: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310</a></p>\n<p>Cheers!</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3040698,
              "author_name": "Arunim Basak",
              "author_url": "",
              "post_date": "2024-11-09T13:17:40.180000",
              "content": "<p>Yeah, that's also make sense. <br>\nThanks for providing the synthetic data. That would be helpful.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3041767,
              "author_name": "Mohamed AbdulMaksoud",
              "author_url": "",
              "post_date": "2024-11-10T20:16:26.783000",
              "content": "<p>Hi! This might be a totally unnecessary comment but I'll add it anyway 😅</p>\n<p>I think this can be taken care of in two main ways:</p>\n<ol>\n<li><p>The authors of the competition already took a step to minimize the effect of underannotation on evaluating our model performance; by using F4 score instead of F1 or AUC. This means that if you have a clever model that detects particles that were missed by the human annotators, it will not be penalized that much by the competition metric. On the other hand, if you have a model that misses an almost-surely identified particle by the human annotators, it will be heavily penalized.</p></li>\n<li><p>A possible solution that you can investigate is trying a psudolabeling approach, meaning you can try to train a model A to try and label the unannotated particles for you, and take its confident labels and append them to the ground-truth annotations. This sometimes improves model performance, and I think it's worth trying.</p></li>\n</ol>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 3052576,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2024-11-22T15:10:23.753000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> , do you have any idea how \"under-labelled\" the dataset is? </p>\n<p>I.e. is there a way to estimate what fraction of particles are labelled vs. unlabelled (at least on average?). </p>\n<p>Would this differ per particle type? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3052758,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-22T19:53:02.290000",
              "content": "<p>We do not have an estimate of how \"under-labelled\" the dataset is. </p>\n<p>The number of unlabelled particles is expected to be different per particle type. The larger particles were easier to label. </p>\n<p>We've discussed some of the challenges around this in the paper: <a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.abstract\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v2.abstract</a> including Table 1 which says the average count of particle per tomogram. </p>\n<p>However, the variability is large, so it isn't possible to say something like: \"Table 1 says 6 VLP particles on average per tomogram. I have only found 1 in this tomogram, so I should look for 5 more then stop.\" It can be meaningful to use the average particle count numbers across all predictions, but I would only use that information loosely.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3052762,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2024-11-22T19:58:22.280000",
              "content": "<p>Thanks!</p>\n<p>All good, averages is what I was after!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3038246,
      "author_name": "Tahereh Zarrat Ehsan",
      "author_url": "",
      "post_date": "2024-11-06T18:47:07.423000",
      "content": "<p>Dear Kyle, <br>\nThe links to the GitHub notebooks are not working.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3038454,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-07T00:40:54.973000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/taherehzarratehsan\" target=\"_blank\">@taherehzarratehsan</a>, sorry about that, we had quite a few things to coordinate to make resources public on the launch day. The notebooks repo on GH should be available now.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3121174,
      "author_name": "Yousef Metwally",
      "author_url": "",
      "post_date": "2025-02-11T10:42:55.323000",
      "content": "<p>Hi Kyle,</p>\n<p>I wonder if the runtime limitation is also for scoring the test set. I have a submission that runs successfully on the fake test set of the 3 tomograms however it is not scored for the competition and throws runtime error. Does the notebook need to run on the test set also on 12 hours? <br>\nThank you for your support I am just trying to still get my score now so that I can estimate the performance.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3095270,
      "author_name": "FAYOLZEUDOM",
      "author_url": "",
      "post_date": "2025-01-13T06:25:14.273000",
      "content": "<p>Hello Kyle or Reza,<br>\nCan you please explain the labels?</p>\n<p>From what I understand, each tomogram is just a space with some objects it in, and we are trying to identify those objects and what type of objects they are. However, the labels (x, y, z) are real-valued (not integers), and have coordinates like (2000, 700, 1000) when the tomograms are only of size (630, 630, 184).<br>\nSo why are the labels out of space?<br>\nWhy are the labels just coordinates and not maybe a segmentation mask or bounding cubes?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3095796,
          "author_name": "Reza Paraan",
          "author_url": "",
          "post_date": "2025-01-13T18:04:20.080000",
          "content": "<p>Hello. The tomograms are made of voxels and their dimensions are the number of voxels. The coordinates are measured in units of distance in Angstroms. The voxels have physical sizes that let you switch from voxel counting to distance counting, depending on which resolution of the tomogram you are working with. If you search through the discussions in the forum you will find more detailed answers. Thanks for your question.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3090982,
      "author_name": "Hui Liu",
      "author_url": "",
      "post_date": "2025-01-07T21:32:47.843000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> ,  for submission, does it have to be CPU notebook ( and run within 12 hours)? or either CPU or GPU notebook is fine?</p>\n<p>Thanks!!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3072798,
      "author_name": "Andrei Zamfir",
      "author_url": "",
      "post_date": "2024-12-15T16:33:51.273000",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> Please excuse my ignorance, is the missing wedge artifact present in the dataset? I could not figure it out from the competition paper.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3071571,
      "author_name": "Hui Liu",
      "author_url": "",
      "post_date": "2024-12-14T00:24:22.613000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> ,</p>\n<p>A question about how CryoET works. So if I have a sample that is  1k nm (length) x 1k nm (width) x 200 nm (thick, z-axis). Do CryoET only able to scan the surface of this Rectangular cuboid? Or it can penetrate into the middle of the sample?<br>\nThe reason I am asking is because I see in the reference (<a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full)\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full)</a>, the thickness of the sample is about 200nm. if CryoET can only image the surface of the sample, we can't get sample that has resolution &lt; 200nm, at least for z-axis. Is that correct?</p>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3071594,
          "author_name": "Bridget Carragher",
          "author_url": "",
          "post_date": "2024-12-14T01:37:41.037000",
          "content": "<p>Hi Hui Liu, Indeed cryoET is a volume imaging method not a surface imaging method.  It is analogous to CT and MRI imaging.  In those methods you take projection images through a solid object (e.g. a brain) and use computational methods to reconstruct a full 3D volume (not a surface) of the entire object.  Some other details are at: <a href=\"https://en.wikipedia.org/wiki/Cryogenic_electron_tomography\" target=\"_blank\">https://en.wikipedia.org/wiki/Cryogenic_electron_tomography</a> and also here: <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/about.html\" target=\"_blank\">https://chanzuckerberg.github.io/cryoet-data-portal/about.html</a>    Thanks for your interest in this challenge, we are very eager to encourage the participation of this community.  Good luck and happy to answer further questions.  Bridget</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3069073,
      "author_name": "Hui Liu",
      "author_url": "",
      "post_date": "2024-12-11T03:09:59.927000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> ,</p>\n<ol>\n<li><p>It seems like the predictions should be the center of particle. May I know how the center is defined? Suppose I know the 3D segmentation of the particle. Is the center coordinate defined by the average of coordinates of all the points segmented to be part of the particle?</p></li>\n<li><p>What is the value range of tomogram in zarr? And does the value represent something like light intensity?  In other words, does higher value means background (water), and lower value means some objects. </p></li>\n</ol>\n<p>Thanks!</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3069386,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-12-11T12:34:04.650000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/huiliu0\" target=\"_blank\">@huiliu0</a> </p>\n<ol>\n<li><p>The predictions can be thought of as centers of particles, but were curated by many approaches including humans manually placing some coordinates. Note that our detection threshold is 0.5x the radius of the particle, so slight differences in how the center coordinate is calculated should not have a significant effect on performance.</p></li>\n<li><p>The images are taken using electron microscopy instead of light, which most of us are familiar with. While electron microscopy is a complex technique, the basic idea is: an electron beam is directed at the sample, and the denser the material at a particular location, the more it scatters the electrons. In our reconstruction, regions with higher density appear darker. Proteins and other biological structures are denser than the surrounding embedding medium, allowing us to visualize them in the tomogram.</p></li>\n</ol>\n<p>Cheers,<br>\nKyle</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3058840,
      "author_name": "homiecal",
      "author_url": "",
      "post_date": "2024-11-30T03:58:15.243000",
      "content": "<p>I think I've found an error in the co-pick utils code. Line 67 in <a href=\"https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/segmentation_from_picks.py\" target=\"_blank\">https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/segmentation_from_picks.py</a></p>\n<p><code>cx, cy, cz = pick.location.z / voxel_spacing, pick.location.y / voxel_spacing, pick.location.x / voxel_spacing</code></p>\n<p>Shouldn't this be:</p>\n<p><code>cx, cy, cz = pick.location.x / voxel_spacing, pick.location.y / voxel_spacing, pick.location.z / voxel_spacing</code></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3066672,
          "author_name": "Bharat",
          "author_url": "",
          "post_date": "2024-12-08T11:03:17.340000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a> <br>\nI too have faced an error (shape cannot be negative) here, that is due to the <em>voxel_spacing</em> that is being passed to <em>from_picks</em> method. It is taking the value of 10 for all the scales, so for lower scales of (92, 315, 315) and (46, 158, 158) it should be ~20 and ~40 respectively<br>\n<code>'datasets': [{'coordinateTransformations': [{'scale': [10.012444196428572,\n                                                                         10.012444196428572,\n                                                                         10.012444537618887],\n                                                               'type': 'scale'}],</code><br>\nSo by fixing this  line I got it working. But for lower levels the masks of tiny particles are completely missing</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3066674,
              "author_name": "Bharat",
              "author_url": "",
              "post_date": "2024-12-08T11:04:21.107000",
              "content": "<p><a href=\"https://github.com/copick/copick-utils/blob/8936b5b9379bf4f58d49ac58a9669076baeeff5b/src/copick_utils/segmentation/segmentation_from_picks.py#L151\" target=\"_blank\">https://github.com/copick/copick-utils/blob/8936b5b9379bf4f58d49ac58a9669076baeeff5b/src/copick_utils/segmentation/segmentation_from_picks.py#L151</a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3066976,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-12-08T17:08:28.887000",
              "content": "<p><a href=\"https://www.kaggle.com/homiecal\" target=\"_blank\">@homiecal</a> and <a href=\"https://www.kaggle.com/bharat0\" target=\"_blank\">@bharat0</a> thank you for pointing this out. Sorry about the delay, I'm still figuring out how to get enough notifications from Kaggle. </p>\n<p>I've started the fix here: <a href=\"https://github.com/copick/copick-utils/pull/9\" target=\"_blank\">https://github.com/copick/copick-utils/pull/9</a></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3058350,
      "author_name": "liuxiaoonline",
      "author_url": "",
      "post_date": "2024-11-29T11:23:10.150000",
      "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>  <br>\nhi, Does the competition provide both noisy and noise-free tomographic images?<br>\nI would like to incorporate denoising training into the network, but I do not have denoised labels for the training data.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3058820,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-30T02:55:16.027000",
          "content": "<p><a href=\"https://www.kaggle.com/liuxiaoonline\" target=\"_blank\">@liuxiaoonline</a> </p>\n<p>We actually provide synthetic data which might be what you are looking for: <a href=\"https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310\" target=\"_blank\">https://cryoetdataportal.czscience.com/datasets/10441?deposition-id=10310</a></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3058841,
          "author_name": "homiecal",
          "author_url": "",
          "post_date": "2024-11-30T04:01:45.617000",
          "content": "<p><a href=\"https://www.kaggle.com/liuxiaoonline\" target=\"_blank\">@liuxiaoonline</a> can you please help me understand what you mean by denoised labels? In the dataset we are provided with picks, and denoised tomograms, the picks are set as our ground truth, are these not denoised labels?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3045115,
      "author_name": "Rustam Bazarbayev",
      "author_url": "",
      "post_date": "2024-11-14T07:23:20.200000",
      "content": "<p>Hi! I am new to the field of machine learning. I have a question related to obtaining images through CryoET. How do you collect data by slicing or rotating the specimen?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3058819,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-30T02:53:55.357000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/rustambazarbayev\" target=\"_blank\">@rustambazarbayev</a> </p>\n<p>Check out this page: <a href=\"https://chanzuckerberg.github.io/cryoet-data-portal/intro_cryoet.html#cryoet-intro\" target=\"_blank\">https://chanzuckerberg.github.io/cryoet-data-portal/intro_cryoet.html#cryoet-intro</a> it covers a lot of the fundamentals of CryoET</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 3042134,
      "author_name": "vishnu gopal P",
      "author_url": "",
      "post_date": "2024-11-11T08:52:41.113000",
      "content": "<p>Appreciate the support.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3038483,
      "author_name": "Pavel Orlov",
      "author_url": "",
      "post_date": "2024-11-07T02:57:57.447000",
      "content": "<p>Hi Kyle,</p>\n<p>What does benchmark.csv on the LB mean? Is it your best result so far that we need to surpass?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3038918,
          "author_name": "Kyle Harrington",
          "author_url": "",
          "post_date": "2024-11-07T14:09:33.793000",
          "content": "<p>benchmark.csv are results we obtained with <a href=\"https://github.com/copick/DeepFindET/tree/main\" target=\"_blank\">DeepFindET</a>. We're hoping folks can get past this.</p>",
          "votes": 4,
          "replies": [
            {
              "id": 3040070,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-11-08T16:49:55.483000",
              "content": "<p>Hello Kyle!</p>\n<p>I find that inference using GPU P100 on Kaggle takes way longer than the times referenced in the example notebooks (I get &gt;100s per tomo vs. 25-30s in <a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/blob/main/DeepFindET/inference.ipynb\" target=\"_blank\">DeepFindET Inference</a>). Extracting coordinates is also slower. Was the 12 hours run-time optimized for?</p>\n<p>Thanks!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3040075,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-08T16:54:49.310000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/andreizamfir\" target=\"_blank\">@andreizamfir</a>,</p>\n<p>Runtimes are definitely one of the challenges in this competition. I've had better luck using the T4s on Kaggle, but some optimizations are still needed to get a full DeepFindET run working (that is one of the reasons I posted the BlobDetector code). </p>\n<p>A couple suggestions that could help:</p>\n<ul>\n<li>figuring out what regions of the image can easily be skipped</li>\n<li>moving the segmentation mask-&gt;coordinate calculation to GPU</li>\n</ul>\n<p>If we get any new code improvements on our side that can help folks out, then we'll post them in the discussion here!</p>",
              "votes": 3,
              "replies": []
            }
          ]
        },
        {
          "id": 3041861,
          "author_name": "Olly Powell",
          "author_url": "",
          "post_date": "2024-11-11T00:32:00.580000",
          "content": "<p><a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a> </p>\n<p>Hi Kyle,  I think we're all suspecting there is something off about the evaluation implementation.  <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545381\" target=\"_blank\">As per this discussion</a>.  Would it be possible for you to confirm if that benchmark value was produced by running DeepFindET directly in Kaggle (without the time limit), or have you evaluated it separately and just posted the score?   Thanks.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3041923,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-11T02:23:30.653000",
              "content": "<p>Hi <a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a>, that benchmark was created by running DeepFindET externally. Some optimization needs to be done to get DeepFindET running within the compute constraints on Kaggle. </p>\n<p>The main things I think are worth trying now are getting proper multi-GPU with the T4 nodes on Kaggle, and getting the segmentation mask-&gt;position code to run on GPU instead of CPU.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3044134,
              "author_name": "Olly Powell",
              "author_url": "",
              "post_date": "2024-11-13T04:40:55.140000",
              "content": "<p>Thanks for confirming.  I think until the bug in the evaluation server is fixed I will hold off trying anything.  But eventually I'd like to use a combination of UNet for segmentation and something faster like your blob detector to save running it over lots of empty space.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3044142,
              "author_name": "Max Campbell",
              "author_url": "",
              "post_date": "2024-11-13T05:04:40.760000",
              "content": "<p>I too suspect that there is something wrong with the evaluation system. The best scoring model only gets a score of 0.003 which might as well be zero… </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3044887,
              "author_name": "Kyle Harrington",
              "author_url": "",
              "post_date": "2024-11-13T23:40:10.750000",
              "content": "<p>hey <a href=\"https://www.kaggle.com/ollypowell\" target=\"_blank\">@ollypowell</a>, </p>\n<p>We have identified the bug in evaluation and will update soon.</p>\n<p>thank you for sticking around!<br>\nKyle</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3044906,
              "author_name": "Olly Powell",
              "author_url": "",
              "post_date": "2024-11-14T00:33:58.637000",
              "content": "<p>Awesome news Kyle, thanks,  looking forward to getting stuck into the challenge!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3073045,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-12-16T01:28:21.733000",
      "content": "<p>Given the limited annotated training dataset available for cryoET, are there existing standardized data augmentation methods specifically for cryoET data? Particularly for 3D image data, what are the most efficient ways to generate synthetic data or leverage existing datasets for data augmentation to ensure the model's generalization ability?</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3038120": "**Hello and welcome!**\n\nWe are excited to invite you to join the CZII - CryoET Object Identification Challenge, where we need new cutting-edge machine learning algorithms to tackle a critical bottleneck in biomedical discovery—the annotation of complex 3D cellular images captured by cryo-electron tomography (cryoET).\n\n**About the Competition**\nIn this challenge, your goal is to create machine learning algorithms capable of robustly annotating a diverse set of protein complexes and particles in 3D tomograms. CryoET offers near-atomic resolution of cellular components, providing unparalleled insight into how proteins function in their native environments. However, despite its power, the process of manually annotating these images remains slow and labor-intensive. With over 15,000 publicly available tomograms in the [CryoET Data Portal](https://cryoetdataportal.czscience.com/) and only 5% annotated, there's an urgent need to automate this process.\n\n**Your Challenge**\nYou'll be tasked with developing ML algorithms that can reliably annotate five classes of particles within hundreds of real-world tomograms. These particles vary greatly in size and shape, from virus-like particles to beta-galactosidase, and are similar to those found in crowded, complex cellular environments. Your algorithm will need to distinguish these particles from non-targets like nucleosomes, filaments, and membrane-bound proteins. To ensure your solution is generalizable, you’ll have access to a limited set of annotated tomograms for training, but you are welcome to incorporate additional synthetic or experimental data to enhance your model’s performance.\n\n**Why Does This Matter?**\nProteins are essential to every aspect of cell function, and understanding their interactions within the cellular architecture is key to advancing human health. CryoET has the potential to unlock critical insights, but the lack of efficient annotation tools is a major barrier. By contributing to this challenge, you will help develop tools that could accelerate discoveries in cell biology and open new avenues for disease treatments. Moreover, your algorithms could be directly applied to the vast corpus of data on the CryoET Data Portal, enabling widespread use and impact beyond this competition.\n\n**The Impact**\nIn addition to competing for one of ten cash prizes, your work will be part of a larger movement to revolutionize how scientists analyze cryoET data. Your algorithm could provide benchmarks for the field, improve data pipelines, and fuel new collaborations between the cryoET and machine learning communities. Ultimately, we believe that better tools for annotating cryoET data could lead to breakthroughs in cell biology—similar to how AlphaFold transformed the understanding of protein structures.\n\n**Join Us!**\nWe are excited to see how you tackle this challenge and what innovative solutions you develop. Along with our team of hosts, we will be active on the Discussion boards, ready to assist and answer any questions you have. This is a unique opportunity to contribute to a rapidly growing field and make a lasting impact on biomedical science.\n\n**Resources**\nWe have developed a number of resources to help folks who are interested in participating, and/or want to learn more about cryoET:\n- Example notebooks (on [Github](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks), and on [Kaggle](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/code))\n    - [MONAI-based 3D UNet](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/3d_unet_monai)\n    - [DeepFindET](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/DeepFindET)\n    - [TomoTwin](https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks/tree/main/tomotwin_picking_notebook)\n- [CZ CryoET Data Portal resources](https://cryoetdataportal.czscience.com/competition) for this competition\n- [Preprint about this competition and the dataset](https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1)\n- [Preprint about open source tools to support competitors](https://www.biorxiv.org/content/10.1101/2024.11.04.621608v1)\n\nGood luck, and we can’t wait to see your solutions!\n\nKyle and Reza\n",
    "3038451": "Hi Kyle,\n\nI am wondering if you could explain in more detail the reasoning behind only including 7 training files and reserving ~500 for the test set. Are you looking to find out whether synthetic data can reliably used train deep learning models on this problem? \n\nLooking through CryoET Data Portal I noticed there is already some simulated data for this competition: https://cryoetdataportal.czscience.com/datasets/10441. It might be helpful to have info on the availability of extra data like this more visible. \n\nAlso if there is additional experimental data that could benefit from more labeling effort elevating its visibility would be very helpful.\n\nI was about to walk away from this competition due to small training dataset. Others might respond similarly to seeing the training set size. Surfacing any relevant synthetic data, scripts, or other resources to increase amount of training data would be helpful as that seems critical for this competition. \n\nFrom the description of the competition it seems to be focused on development of new ML methods for this problem but with the set up of the train and test data it appears to be more about getting more data or using synthetic data for the problem. I am interested to know which is the actual challenge currently?",
    "3040316": "@kharrington \nhi, would it be possible for the host to release the code for the synthetic data generator?\n\ni want to build a foundation model using synthetic data. I would need to find a way to:\n- mimic actual noise, artifacts for tomography images\n- as a foundation model, I want to detect more particle objects (apart from the 6 targets of this competition) at pertaining satge.\n(Then i will investigate if we can later do few shot transfer learning using just a few real data from the kaggle train dataset with the foundation model)",
    "3038745": "Could you confirm that the 7 images in the training dataset and the 500 (?) images in the test dataset are real images that have been hand-labeled? Or are they for example synthetically generated?\n\nIf they are hand-labeled, do you have any indication of possible labeling errors? For example, how different do the labels tend to be if two different people or teams do the labeling?",
    "3038246": "Dear Kyle, \nThe links to the GitHub notebooks are not working.",
    "3121174": "Hi Kyle,\n\nI wonder if the runtime limitation is also for scoring the test set. I have a submission that runs successfully on the fake test set of the 3 tomograms however it is not scored for the competition and throws runtime error. Does the notebook need to run on the test set also on 12 hours? \nThank you for your support I am just trying to still get my score now so that I can estimate the performance.\n",
    "3095270": "Hello Kyle or Reza,\nCan you please explain the labels?\n\nFrom what I understand, each tomogram is just a space with some objects it in, and we are trying to identify those objects and what type of objects they are. However, the labels (x, y, z) are real-valued (not integers), and have coordinates like (2000, 700, 1000) when the tomograms are only of size (630, 630, 184).\nSo why are the labels out of space?\nWhy are the labels just coordinates and not maybe a segmentation mask or bounding cubes?",
    "3090982": "Hi @kharrington ,  for submission, does it have to be CPU notebook ( and run within 12 hours)? or either CPU or GPU notebook is fine?\n\nThanks!!",
    "3072798": "@kharrington Please excuse my ignorance, is the missing wedge artifact present in the dataset? I could not figure it out from the competition paper.",
    "3071571": "Hi @kharrington ,\n\nA question about how CryoET works. So if I have a sample that is  1k nm (length) x 1k nm (width) x 200 nm (thick, z-axis). Do CryoET only able to scan the surface of this Rectangular cuboid? Or it can penetrate into the middle of the sample?\nThe reason I am asking is because I see in the reference (https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full), the thickness of the sample is about 200nm. if CryoET can only image the surface of the sample, we can't get sample that has resolution < 200nm, at least for z-axis. Is that correct?\n\nThanks!\n",
    "3069073": "Hi @kharrington ,\n\n1. It seems like the predictions should be the center of particle. May I know how the center is defined? Suppose I know the 3D segmentation of the particle. Is the center coordinate defined by the average of coordinates of all the points segmented to be part of the particle?\n\n2. What is the value range of tomogram in zarr? And does the value represent something like light intensity?  In other words, does higher value means background (water), and lower value means some objects. \n\nThanks!",
    "3058840": "I think I've found an error in the co-pick utils code. Line 67 in https://github.com/copick/copick-utils/blob/main/src/copick_utils/segmentation/segmentation_from_picks.py\n\n`cx, cy, cz = pick.location.z / voxel_spacing, pick.location.y / voxel_spacing, pick.location.x / voxel_spacing`\n\nShouldn't this be:\n\n`cx, cy, cz = pick.location.x / voxel_spacing, pick.location.y / voxel_spacing, pick.location.z / voxel_spacing`\n",
    "3058350": "@kharrington  \nhi, Does the competition provide both noisy and noise-free tomographic images?\nI would like to incorporate denoising training into the network, but I do not have denoised labels for the training data.\n",
    "3045115": "Hi! I am new to the field of machine learning. I have a question related to obtaining images through CryoET. How do you collect data by slicing or rotating the specimen?",
    "3042134": "Appreciate the support.",
    "3038483": "Hi Kyle,\n\nWhat does benchmark.csv on the LB mean? Is it your best result so far that we need to surpass?",
    "3073045": "Given the limited annotated training dataset available for cryoET, are there existing standardized data augmentation methods specifically for cryoET data? Particularly for 3D image data, what are the most efficient ways to generate synthetic data or leverage existing datasets for data augmentation to ensure the model's generalization ability?"
  }
}