{
  "id": 545221,
  "title": "[lb0.748] experiment resuts on cyroET foundation model build with synthetic data",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/545221",
  "author_name": "hengck23",
  "post_date": "2024-11-09T02:58:52.493000",
  "votes": 81,
  "comment_count": 67,
  "views": 0,
  "content": "<p>this is in progress and will be updated often.</p>\n<p>inspiration:</p>\n<ul>\n<li>did you know that you can pretrain imagenet model without using real images? This is one famous work using fractal images:</li>\n</ul>\n<p><a href=\"https://hirokatsukataoka16.github.io/Pretraining-without-Natural-Images/\" target=\"_blank\">https://hirokatsukataoka16.github.io/Pretraining-without-Natural-Images/</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbeaecb6153485f554f602842238e6e95%2FSelection_649.png?generation=1731120390177754&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>pertaining with synthetic images is valuable when it is hard to obtain real images (and their labels). Examples include seismic data, tomography data, etc</li>\n</ul>\n<hr>\n<p>my plan:</p>\n<ol>\n<li>synthetic data :  <ul>\n<li>from CryoET Data Portal</li>\n<li>find synthetic data generator to create my own data with more noise, more object orinetation and more particle objects</li></ul></li>\n<li>some efficient self-supervised methods?  </li>\n<li>transformer-based foundation model. Transformers perform very well if there is massive data.</li>\n<li>Transfer learning to n-shot using real kaggle train data and make submission</li>\n</ol>",
  "messages": [
    {
      "id": 3040321,
      "postDate": "2024-11-09T02:58:52.493Z",
      "content": "<p>this is in progress and will be updated often.</p>\n<p>inspiration:</p>\n<ul>\n<li>did you know that you can pretrain imagenet model without using real images? This is one famous work using fractal images:</li>\n</ul>\n<p><a href=\"https://hirokatsukataoka16.github.io/Pretraining-without-Natural-Images/\" target=\"_blank\">https://hirokatsukataoka16.github.io/Pretraining-without-Natural-Images/</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbeaecb6153485f554f602842238e6e95%2FSelection_649.png?generation=1731120390177754&amp;alt=media\" alt=\"\"></p>\n<ul>\n<li>pertaining with synthetic images is valuable when it is hard to obtain real images (and their labels). Examples include seismic data, tomography data, etc</li>\n</ul>\n<hr>\n<p>my plan:</p>\n<ol>\n<li>synthetic data :  <ul>\n<li>from CryoET Data Portal</li>\n<li>find synthetic data generator to create my own data with more noise, more object orinetation and more particle objects</li></ul></li>\n<li>some efficient self-supervised methods?  </li>\n<li>transformer-based foundation model. Transformers perform very well if there is massive data.</li>\n<li>Transfer learning to n-shot using real kaggle train data and make submission</li>\n</ol>",
      "rawMarkdown": "this is in progress and will be updated often.\n\ninspiration:\n- did you know that you can pretrain imagenet model without using real images? This is one famous work using fractal images:\n\nhttps://hirokatsukataoka16.github.io/Pretraining-without-Natural-Images/\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbeaecb6153485f554f602842238e6e95%2FSelection_649.png?generation=1731120390177754&alt=media)\n\n- pertaining with synthetic images is valuable when it is hard to obtain real images (and their labels). Examples include seismic data, tomography data, etc\n\n---\n\nmy plan:\n1. synthetic data :  \n   - from CryoET Data Portal\n   - find synthetic data generator to create my own data with more noise, more object orinetation and more particle objects\n2. some efficient self-supervised methods?  \n3. transformer-based foundation model. Transformers perform very well if there is massive data.\n3. Transfer learning to n-shot using real kaggle train data and make submission\n",
      "votes": 81
    },
    {
      "id": 3057240,
      "postDate": "2024-11-27T22:38:55.387Z",
      "content": "<p>a list of winning models</p>\n<p>SHREC 2021: CLASSIFICATION IN CRYO-ELECTRON TOMOGRAMS<br>\n<a href=\"https://arxiv.org/pdf/2203.10035\" target=\"_blank\">https://arxiv.org/pdf/2203.10035</a></p>\n<p>SHREC’19 Track: Classification in Cryo-Electron Tomograms<br>\n<a href=\"https://webspace.science.uu.nl/~veltk101/publications/art/shrec2019-cryo.pdf\" target=\"_blank\">https://webspace.science.uu.nl/~veltk101/publications/art/shrec2019-cryo.pdf</a></p>",
      "rawMarkdown": "a list of winning models\n\nSHREC 2021: CLASSIFICATION IN CRYO-ELECTRON TOMOGRAMS\nhttps://arxiv.org/pdf/2203.10035\n\n\nSHREC’19 Track: Classification in Cryo-Electron Tomograms\nhttps://webspace.science.uu.nl/~veltk101/publications/art/shrec2019-cryo.pdf\n",
      "votes": 10
    },
    {
      "id": 3054691,
      "postDate": "2024-11-25T02:22:13.190Z",
      "content": "<p>most likely, the winning solution would be two-stage approach.<br>\nstage 1:</p>\n<ul>\n<li>segmentation net to predict coords of candidate particles<br>\nstage 2:</li>\n<li>crop the volume to verify: rank the most likely candidates and refine coordinate</li>\n</ul>\n<p>for stage2:</p>\n<ul>\n<li>size/volume of prediction is important for filtering results</li>\n<li>you can upsize the crop for better accuracy</li>\n<li>how to combine scores of stage.1 and stage.2?</li>\n</ul>\n<hr>\n<p>you probably need to do a lot of probing. we actually need to know the recall and precision of the lb submission.<br>\nthis is how:</p>\n<ul>\n<li>submit for one particle: one equation for recall and precision</li>\n<li>same as above, but submit known number of false positives(e.g. just take corner point or background point or just out of volume point or (0,0,0)) : another equation for recall and precision</li>\n</ul>",
      "rawMarkdown": "most likely, the winning solution would be two-stage approach.\nstage 1:\n- segmentation net to predict coords of candidate particles\nstage 2:\n- crop the volume to verify: rank the most likely candidates and refine coordinate\n\nfor stage2:\n- size/volume of prediction is important for filtering results\n- you can upsize the crop for better accuracy\n- how to combine scores of stage.1 and stage.2?\n\n----\n\nyou probably need to do a lot of probing. we actually need to know the recall and precision of the lb submission.\nthis is how:\n- submit for one particle: one equation for recall and precision\n- same as above, but submit known number of false positives(e.g. just take corner point or background point or just out of volume point or (0,0,0)) : another equation for recall and precision\n\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 3085686,
          "postDate": "2025-01-01T12:48:21.140Z",
          "content": "<h3>Any guidance on approach in stage 2</h3>\n<p>The original way seems to be very complex and computationally expensive.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F0392b522e900fb7bec916c7aa890a353%2FScreenshot%202025-01-01%20181656.png?generation=1735735646288793&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "### Any guidance on approach in stage 2\n\nThe original way seems to be very complex and computationally expensive.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F0392b522e900fb7bec916c7aa890a353%2FScreenshot%202025-01-01%20181656.png?generation=1735735646288793&alt=media)"
        }
      ]
    },
    {
      "id": 3054647,
      "postDate": "2024-11-25T00:02:51.337Z",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F749225bf11d4e2dc3359c05ebb1b2030%2Fezgif-3-a60846ca1e.gif?generation=1732492937995569&amp;alt=media\" alt=\"\"></p>\n<p>how it look in 3d<br>\n(volume rendering using pyvista)</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F749225bf11d4e2dc3359c05ebb1b2030%2Fezgif-3-a60846ca1e.gif?generation=1732492937995569&alt=media)\n\nhow it look in 3d\n(volume rendering using pyvista)",
      "votes": 7
    },
    {
      "id": 3056473,
      "postDate": "2024-11-27T02:18:10.780Z",
      "content": "<p>contribution of lb score:</p>\n<p>baseline(ttax2) 0.702<br>\nbetter model parameters(ttax2): 0.718<br>\ntta, 2xrot90 : add 0.05<br>\ntta, 4xrot90 : add 0.10<br>\nensemble (same fold, different base encoder):  0.718+0.713 = 0.743</p>",
      "rawMarkdown": "contribution of lb score:\n\nbaseline(ttax2) 0.702\nbetter model parameters(ttax2): 0.718\ntta, 2xrot90 : add 0.05\ntta, 4xrot90 : add 0.10\nensemble (same fold, different base encoder):  0.718+0.713 = 0.743\n",
      "votes": 5
    },
    {
      "id": 3049275,
      "postDate": "2024-11-18T22:54:22.680Z",
      "content": "<p>handling border artifacts is always a trick to win kaggle<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65e270024f58c752eaf67d9ad7a183fd%2FSelection_705.png?generation=1731970460522047&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "handling border artifacts is always a trick to win kaggle\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65e270024f58c752eaf67d9ad7a183fd%2FSelection_705.png?generation=1731970460522047&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 3051295,
          "postDate": "2024-11-21T06:09:44.063Z",
          "content": "<p>Hello author, I would like to ask if this artifact is the one produced during the 3D reconstruction process of your synthetic data?</p>",
          "rawMarkdown": "Hello author, I would like to ask if this artifact is the one produced during the 3D reconstruction process of your synthetic data?\n",
          "replies": [
            {
              "id": 3051300,
              "postDate": "2024-11-21T06:13:27.460Z",
              "content": "<p>results of </p>\n<pre><code>            probability[:, z:z + num_slice, : + image_size, : + image_size] += prob\n            count[:, z:z + num_slice, : + image_size, : + image_size] += \n</code></pre>\n<p>correct version</p>\n<pre><code>            probability[:, z:z + num_slice, : + image_size, : + image_size] += weight * prob\n            count[:, z:z + num_slice, : + image_size, : + image_size] += weight  #weight is shape =(:,num_slice,image_size, image_size)\n</code></pre>",
              "rawMarkdown": "results of \n```\n\n\t\t\tprobability[:, z:z + num_slice, y:y + image_size, x:x + image_size] += prob\n\t\t\tcount[:, z:z + num_slice, y:y + image_size, x:x + image_size] += 1\n\n```\n\ncorrect version\n\n```\n\n\t\t\tprobability[:, z:z + num_slice, y:y + image_size, x:x + image_size] += weight * prob\n\t\t\tcount[:, z:z + num_slice, y:y + image_size, x:x + image_size] += weight  #weight is shape =(:,num_slice,image_size, image_size)\n\n```",
              "votes": 2
            },
            {
              "id": 3051306,
              "postDate": "2024-11-21T06:18:53.037Z",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab423841d1da42cd294c127bedf8582f%2FSelection_717.png?generation=1732169930925251&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab423841d1da42cd294c127bedf8582f%2FSelection_717.png?generation=1732169930925251&alt=media)",
              "votes": 3
            },
            {
              "id": 3086673,
              "postDate": "2025-01-02T15:17:57.330Z",
              "content": "<p>Pls share some resources for post prediction of overlapping patches…</p>\n<p>Superly needed (if its right grammatically)<br>\nthanks</p>",
              "rawMarkdown": "Pls share some resources for post prediction of overlapping patches...\n\nSuperly needed (if its right grammatically)\nthanks"
            }
          ]
        }
      ]
    },
    {
      "id": 3047384,
      "postDate": "2024-11-16T16:00:50.687Z",
      "content": "<p>first baseline experiment results are up:<br>\nupdated on 17-oct</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F806ebf376130887c9acf397bfeee9865%2FSelection_691.png?generation=1731784529142531&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "first baseline experiment results are up:\nupdated on 17-oct\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F806ebf376130887c9acf397bfeee9865%2FSelection_691.png?generation=1731784529142531&alt=media)",
      "votes": 5,
      "replies": [
        {
          "id": 3047391,
          "postDate": "2024-11-16T16:14:05Z",
          "content": "<p>Using a 2D encoder and a 3D decoder (I think that's what you're doing?) is pretty interesting! Nice work <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I always enjoy learning from your discussion posts.</p>",
          "rawMarkdown": "Using a 2D encoder and a 3D decoder (I think that's what you're doing?) is pretty interesting! Nice work @hengck23, I always enjoy learning from your discussion posts.",
          "replies": [
            {
              "id": 3047416,
              "postDate": "2024-11-16T16:35:46.897Z",
              "content": "<p><a href=\"https://www.kaggle.com/code/hengck23/2d-to-3d-unet-demo\" target=\"_blank\">https://www.kaggle.com/code/hengck23/2d-to-3d-unet-demo</a><br>\nThis is an old version. </p>\n<p>The cryoET version is slightly different. you can refer to <a href=\"https://www.kaggle.com/datasets/hengck23/hengck-czii-cryo-et-01\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/hengck-czii-cryo-et-01</a></p>",
              "rawMarkdown": "https://www.kaggle.com/code/hengck23/2d-to-3d-unet-demo\nThis is an old version. \n\nThe cryoET version is slightly different. you can refer to https://www.kaggle.com/datasets/hengck23/hengck-czii-cryo-et-01",
              "votes": 1
            },
            {
              "id": 3048958,
              "postDate": "2024-11-18T14:54:30.110Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 3047420,
          "postDate": "2024-11-16T16:42:43.733Z",
          "content": "<p>training log<br>\n(see file attached)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c1fac7d7ad70a4d42b6d15b8bc88dac%2FSelection_686.png?generation=1731775361123507&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "training log\n(see file attached)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c1fac7d7ad70a4d42b6d15b8bc88dac%2FSelection_686.png?generation=1731775361123507&alt=media)",
          "replies": [
            {
              "id": 3047425,
              "postDate": "2024-11-16T16:48:51.253Z",
              "content": "<p>sorry if it’s too obvious, what do these 92 iterations per epoch mean?</p>",
              "rawMarkdown": "sorry if it’s too obvious, what do these 92 iterations per epoch mean?"
            },
            {
              "id": 3050682,
              "postDate": "2024-11-20T13:06:12.140Z",
              "content": "<p>May I ask how did you process the data? How did you get 92 iterations per epoch? Thanks!</p>",
              "rawMarkdown": "May I ask how did you process the data? How did you get 92 iterations per epoch? Thanks!",
              "votes": 1
            },
            {
              "id": 3051141,
              "postDate": "2024-11-20T22:57:50.270Z",
              "content": "<p>that is not important. it just says that for one epoch, 92 batches are randomly created. each batch has 3 samples  (volume crop)</p>",
              "rawMarkdown": "that is not important. it just says that for one epoch, 92 batches are randomly created. each batch has 3 samples  (volume crop)"
            }
          ]
        },
        {
          "id": 3047477,
          "postDate": "2024-11-16T18:10:31.870Z",
          "content": "<p>threshold versus  f-beta score</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F10aa8ad2df95befcc332ca721731bdc8%2FSelection_690.png?generation=1731780624993380&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "threshold versus  f-beta score\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F10aa8ad2df95befcc332ca721731bdc8%2FSelection_690.png?generation=1731780624993380&alt=media)"
        },
        {
          "id": 3057922,
          "postDate": "2024-11-28T20:14:53.280Z",
          "content": "<p>Do you report \"local LB\" (which is another name for validation, I guess) for predictions and GT divided by 10 or in the original space (similar to how you would post-process predictions during submit)? </p>",
          "rawMarkdown": "Do you report \"local LB\" (which is another name for validation, I guess) for predictions and GT divided by 10 or in the original space (similar to how you would post-process predictions during submit)? ",
          "replies": [
            {
              "id": 3057976,
              "postDate": "2024-11-28T21:22:00.267Z",
              "content": "<p>yes, local LB is local CV. using  official mertic  code <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">https://www.kaggle.com/code/metric/czi-cryoet-84969</a><br>\nso it is what you submitted in csv </p>",
              "rawMarkdown": "yes, local LB is local CV. using  official mertic  code https://www.kaggle.com/code/metric/czi-cryoet-84969\nso it is what you submitted in csv \n"
            },
            {
              "id": 3058232,
              "postDate": "2024-11-29T08:46:53.857Z",
              "content": "<p>Correct me if I'm wrong. But I've been recently checking metric, and since positives are predictions inside .5 times particle radius, It doesn't matter the scale (they just need both be the same, GT and predictions).</p>",
              "rawMarkdown": "Correct me if I'm wrong. But I've been recently checking metric, and since positives are predictions inside .5 times particle radius, It doesn't matter the scale (they just need both be the same, GT and predictions)."
            },
            {
              "id": 3058279,
              "postDate": "2024-11-29T09:28:14.227Z",
              "content": "<p>slight numerical error exists though both metric will be very close</p>",
              "rawMarkdown": "slight numerical error exists though both metric will be very close",
              "votes": 1
            },
            {
              "id": 3058691,
              "postDate": "2024-11-29T19:45:06.543Z",
              "content": "<p>Doesn't seem right to me. </p>\n<p>Suppose your predictions are off by 5 pixels in 10A. It means that they are off by 50 pixels in 1A. So if you're using the same evaluation code with the same radius values, metrics will be vastly different. </p>",
              "rawMarkdown": "Doesn't seem right to me. \n\nSuppose your predictions are off by 5 pixels in 10A. It means that they are off by 50 pixels in 1A. So if you're using the same evaluation code with the same radius values, metrics will be vastly different. "
            },
            {
              "id": 3058717,
              "postDate": "2024-11-29T20:41:25.263Z",
              "content": "<p>In or out is with respect .5 times the particle radius. The particle radius scales with the predictions. You just need to esure that particle radius inside metric are in the same scale than your predictions. At kaggle submission is only one way. But at home, you chose.</p>",
              "rawMarkdown": "In or out is with respect .5 times the particle radius. The particle radius scales with the predictions. You just need to esure that particle radius inside metric are in the same scale than your predictions. At kaggle submission is only one way. But at home, you chose."
            },
            {
              "id": 3058719,
              "postDate": "2024-11-29T20:44:06.433Z",
              "content": "<p>Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake. </p>",
              "rawMarkdown": "Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake. ",
              "votes": 2
            },
            {
              "id": 3058726,
              "postDate": "2024-11-29T20:56:00.997Z",
              "content": "<p>And that's exactly what I did haha I've noticed my mistake right after answer. Thanks for point it out.</p>",
              "rawMarkdown": "And that's exactly what I did haha I've noticed my mistake right after answer. Thanks for point it out.",
              "votes": 2
            },
            {
              "id": 3058728,
              "postDate": "2024-11-29T21:03:39.967Z",
              "content": "<p>I did it as well… It was painful to find the mistake </p>",
              "rawMarkdown": "I did it as well... It was painful to find the mistake "
            },
            {
              "id": 3060110,
              "postDate": "2024-12-01T12:05:08.380Z",
              "content": "<p>We've got two score fixes already so will probably everything be fine and it's just me but… Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).</p>\n<p>Public LB and overoptimistic ~.5</p>\n<p>correct scales ~.1</p>",
              "rawMarkdown": "We've got two score fixes already so will probably everything be fine and it's just me but... Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).\n\nPublic LB and overoptimistic ~.5\n\ncorrect scales ~.1"
            }
          ]
        }
      ]
    },
    {
      "id": 3041344,
      "postDate": "2024-11-10T07:47:27.680Z",
      "content": "<p>i will come back to this later. but we do need a fast and robust 3d peak detector<br>\n<a href=\"https://stackoverflow.com/questions/3684484/peak-detection-in-a-2d-array\" target=\"_blank\">https://stackoverflow.com/questions/3684484/peak-detection-in-a-2d-array</a></p>",
      "rawMarkdown": "i will come back to this later. but we do need a fast and robust 3d peak detector\nhttps://stackoverflow.com/questions/3684484/peak-detection-in-a-2d-array",
      "votes": 6
    },
    {
      "id": 3057131,
      "postDate": "2024-11-27T18:35:00.257Z",
      "content": "<p>3d object detection toolkit, with 3d roi align and nms<br>\n<a href=\"https://github.com/TimothyZero/MedVision/tree/main\" target=\"_blank\">https://github.com/TimothyZero/MedVision/tree/main</a><br>\n<a href=\"https://github.com/pytorch/vision/issues/2402\" target=\"_blank\">https://github.com/pytorch/vision/issues/2402</a></p>",
      "rawMarkdown": "3d object detection toolkit, with 3d roi align and nms\nhttps://github.com/TimothyZero/MedVision/tree/main\nhttps://github.com/pytorch/vision/issues/2402",
      "votes": 3,
      "replies": [
        {
          "id": 3091158,
          "postDate": "2025-01-08T06:09:05.413Z",
          "content": "<p><a href=\"https://github.com/MIC-DKFZ/batchgenerators\" target=\"_blank\">https://github.com/MIC-DKFZ/batchgenerators</a></p>",
          "rawMarkdown": "https://github.com/MIC-DKFZ/batchgenerators"
        }
      ]
    },
    {
      "id": 3055872,
      "postDate": "2024-11-26T08:09:23.903Z",
      "content": "<p><a href=\"https://github.com/icthrm/SC-Net\" target=\"_blank\">https://github.com/icthrm/SC-Net</a><br>\nself supervised denoise</p>",
      "rawMarkdown": "https://github.com/icthrm/SC-Net\nself supervised denoise",
      "votes": 3
    },
    {
      "id": 3051142,
      "postDate": "2024-11-20T23:00:23.383Z",
      "content": "<p>my early experiment:<br>\nresnet3d + unet-decoder3d (66mb)<br>\n-local cv 0.694, lb=0.666</p>\n<p>as reference<br>\nresnet18d + unet-decoder3d (88mb)<br>\nresnet34d + unet-decoder3d (123mb)</p>",
      "rawMarkdown": "my early experiment:\nresnet3d + unet-decoder3d (66mb)\n-local cv 0.694, lb=0.666\n\n\nas reference\nresnet18d + unet-decoder3d (88mb)\nresnet34d + unet-decoder3d (123mb)",
      "votes": 3
    },
    {
      "id": 3053670,
      "postDate": "2024-11-23T18:15:29.710Z",
      "content": "<p>learning curve: </p>\n<ul>\n<li>if you submit a non-overfitted model, you will see cv vs lb correlation</li>\n<li>it is easy to get high cv (e.g. in the range of 0.77 or even 0.82). but we are not interested in that.</li>\n<li>rather, we are interested to get a cv that <strong>we know</strong> is close to lb</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F485fa0824f57310ba5e8422e81edc678%2FSelection_731.png?generation=1732385724424357&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "learning curve: \n- if you submit a non-overfitted model, you will see cv vs lb correlation\n- it is easy to get high cv (e.g. in the range of 0.77 or even 0.82). but we are not interested in that.\n- rather, we are interested to get a cv that **we know** is close to lb\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F485fa0824f57310ba5e8422e81edc678%2FSelection_731.png?generation=1732385724424357&alt=media)",
      "votes": 4
    },
    {
      "id": 3041079,
      "postDate": "2024-11-09T22:40:30.117Z",
      "content": "<p>CryoTransformer: A Transformer Model for Picking Protein Particles from Cryo-EM Micrographs<br>\n<a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC10634673/pdf/nihpp-2023.10.19.563155v1.pdf\" target=\"_blank\">https://pmc.ncbi.nlm.nih.gov/articles/PMC10634673/pdf/nihpp-2023.10.19.563155v1.pdf</a><br>\n<a href=\"https://github.com/jianlin-cheng/CryoTransformer\" target=\"_blank\">https://github.com/jianlin-cheng/CryoTransformer</a></p>",
      "rawMarkdown": "CryoTransformer: A Transformer Model for Picking Protein Particles from Cryo-EM Micrographs\nhttps://pmc.ncbi.nlm.nih.gov/articles/PMC10634673/pdf/nihpp-2023.10.19.563155v1.pdf\nhttps://github.com/jianlin-cheng/CryoTransformer",
      "votes": 3
    },
    {
      "id": 3040609,
      "postDate": "2024-11-09T10:44:59.937Z",
      "content": "<p>make your own data?<br>\n<a href=\"https://github.com/cellcanvas/album-catalog/blob/main/solutions/polnet/generate-tomogram/solution.py\" target=\"_blank\">https://github.com/cellcanvas/album-catalog/blob/main/solutions/polnet/generate-tomogram/solution.py</a></p>\n<p><a href=\"https://github.com/anmartinezs/polnet\" target=\"_blank\">https://github.com/anmartinezs/polnet</a><br>\n[1] Martinez-Sanchez A.*, and Lamm L., Jasnin M. and Phelippeau H. (2024) \"Simulating the cellular context in synthetic datasets for cryo-electron tomography\" IEEE Transactions on Medical Imaging </p>\n<hr>\n<p><a href=\"https://github.com/phonchi/Computational-CryoET\" target=\"_blank\">https://github.com/phonchi/Computational-CryoET</a></p>",
      "rawMarkdown": "make your own data?\nhttps://github.com/cellcanvas/album-catalog/blob/main/solutions/polnet/generate-tomogram/solution.py\n\nhttps://github.com/anmartinezs/polnet\n[1] Martinez-Sanchez A.*, and Lamm L., Jasnin M. and Phelippeau H. (2024) \"Simulating the cellular context in synthetic datasets for cryo-electron tomography\" IEEE Transactions on Medical Imaging \n\n---\nhttps://github.com/phonchi/Computational-CryoET",
      "votes": 3
    },
    {
      "id": 3040323,
      "postDate": "2024-11-09T03:00:21.273Z",
      "content": "<p>a simple notebook for benchmarking:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder</a></p>",
      "rawMarkdown": "a simple notebook for benchmarking:\nhttps://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder",
      "votes": 3
    },
    {
      "id": 3041568,
      "postDate": "2024-11-10T14:19:40.523Z",
      "content": "<p><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks</a></p>\n<p>more example network!!!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F50564f4f210d4ed09f9c14ce35ed513d%2FSelection_663.png?generation=1731248372945421&amp;alt=media\" alt=\"\"></p>\n<p>Open-source Tools for CryoET Particle Picking Machine Learning Competitions<br>\n<a href=\"https://www.researchgate.net/publication/385559021_Open-source_Tools_for_CryoET_Particle_Picking_Machine_Learning_Competitions\" target=\"_blank\">https://www.researchgate.net/publication/385559021_Open-source_Tools_for_CryoET_Particle_Picking_Machine_Learning_Competitions</a> </p>",
      "rawMarkdown": "https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\n\nmore example network!!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F50564f4f210d4ed09f9c14ce35ed513d%2FSelection_663.png?generation=1731248372945421&alt=media)\n\n\nOpen-source Tools for CryoET Particle Picking Machine Learning Competitions\nhttps://www.researchgate.net/publication/385559021_Open-source_Tools_for_CryoET_Particle_Picking_Machine_Learning_Competitions ",
      "votes": 4
    },
    {
      "id": 3041550,
      "postDate": "2024-11-10T14:04:22.950Z",
      "content": "<p>important details from dataset paper:</p>\n<p>Annotating CryoET Volumes: A Machine Learning Challenge <br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full.pdf</a></p>\n<p>\"This set was generated by mixing varying percentages (in 10% 378 increments) of picks from a fine-tuned DeepFindET model and ground truth data. By comparing their F-beta scores, we observed that datasets containing 80% or more ground-truth results yielded mock datasets that score better than DeepFindET\"</p>\n<p>is DeepFindET used in annotations? YES! … and other tools …</p>\n<p>\"The tools described above, and other published software packages were stitched together into several workflows to generate “ground truth” labels \"</p>\n<hr>",
      "rawMarkdown": "important details from dataset paper:\n\nAnnotating CryoET Volumes: A Machine Learning Challenge \nhttps://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full.pdf\n\n\"This set was generated by mixing varying percentages (in 10% 378 increments) of picks from a fine-tuned DeepFindET model and ground truth data. By comparing their F-beta scores, we observed that datasets containing 80% or more ground-truth results yielded mock datasets that score better than DeepFindET\"\n\nis DeepFindET used in annotations? YES! ... and other tools ...\n\n\"The tools described above, and other published software packages were stitched together into several workflows to generate “ground truth” labels \"\n\n\n----\n ",
      "votes": 4,
      "replies": [
        {
          "id": 3043440,
          "postDate": "2024-11-12T12:23:42.650Z",
          "content": "<p>please read the paper in details. use chatgpt to help you!</p>\n<p>i was at first pretty confused about phatom datset (is it real or not, is it pyhsical or software simulated ) ….</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F258d203f385aaf78e7780047f73ca4a7%2FSelection_999(6791).png?generation=1731414132696051&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F406d112248a2f25a2fab3c504e1acb01%2FSelection_999(6792).png?generation=1731414144754747&amp;alt=media\" alt=\"\"></p>\n<p>you solution may be quite of reverse enginerring the annotation process</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f5523228d52061170267d223c3d66fe%2FSelection_999(6793).png?generation=1731414155492008&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1c2fb5e7eee99120d17dae885e4529b9%2FSelection_999(6794).png?generation=1731414220848937&amp;alt=media\" alt=\"\"></p>",
          "rawMarkdown": "please read the paper in details. use chatgpt to help you!\n\ni was at first pretty confused about phatom datset (is it real or not, is it pyhsical or software simulated ) ....\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F258d203f385aaf78e7780047f73ca4a7%2FSelection_999(6791).png?generation=1731414132696051&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F406d112248a2f25a2fab3c504e1acb01%2FSelection_999(6792).png?generation=1731414144754747&alt=media)\n\nyou solution may be quite of reverse enginerring the annotation process\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f5523228d52061170267d223c3d66fe%2FSelection_999(6793).png?generation=1731414155492008&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1c2fb5e7eee99120d17dae885e4529b9%2FSelection_999(6794).png?generation=1731414220848937&alt=media)\n",
          "votes": 2
        },
        {
          "id": 3083301,
          "postDate": "2024-12-29T09:16:35.210Z",
          "content": "<p>This conversation may be useful:</p>\n<p><a href=\"https://chatgpt.com/share/6771132f-6638-8009-b1e8-c36b76d36af4\" target=\"_blank\">https://chatgpt.com/share/6771132f-6638-8009-b1e8-c36b76d36af4</a></p>",
          "rawMarkdown": "This conversation may be useful:\n\nhttps://chatgpt.com/share/6771132f-6638-8009-b1e8-c36b76d36af4"
        }
      ]
    },
    {
      "id": 3041062,
      "postDate": "2024-11-09T22:09:00.543Z",
      "content": "<p>introduction video<br>\n<a href=\"https://cryoem101.org/chapter-1-et/\" target=\"_blank\">https://cryoem101.org/chapter-1-et/</a></p>\n<p>Electron Tomography<br>\n<a href=\"https://www.youtube.com/watch?v=M2D3lDIv9jY\" target=\"_blank\">https://www.youtube.com/watch?v=M2D3lDIv9jY</a></p>\n<p>Cryo EM Tomography<br>\n<a href=\"https://www.youtube.com/watch?v=frgjc-ZNhOY\" target=\"_blank\">https://www.youtube.com/watch?v=frgjc-ZNhOY</a></p>\n<p>A 3 minute introduction to CryoEM<br>\n<a href=\"https://www.youtube.com/watch?v=BJKkC0W-6Qk\" target=\"_blank\">https://www.youtube.com/watch?v=BJKkC0W-6Qk</a></p>\n<p>Minute Biophysics - Cryo-Electron Tomography (Cryo-E.T.), Tiffany<br>\n<a href=\"https://www.youtube.com/watch?v=lf-tWHSbr8Q\" target=\"_blank\">https://www.youtube.com/watch?v=lf-tWHSbr8Q</a></p>\n<p>Non-averaged 3D structure of a single tetra-nucleosome arrays by Na+ and H1 by cryo-ET and IPET<br>\n<a href=\"https://www.youtube.com/watch?v=RIFGnKzK0rk\" target=\"_blank\">https://www.youtube.com/watch?v=RIFGnKzK0rk</a></p>\n<p>tilt series<br>\n<a href=\"https://www.youtube.com/watch?v=SbEvuSgskWw\" target=\"_blank\">https://www.youtube.com/watch?v=SbEvuSgskWw</a></p>",
      "rawMarkdown": "introduction video\nhttps://cryoem101.org/chapter-1-et/\n\nElectron Tomography\nhttps://www.youtube.com/watch?v=M2D3lDIv9jY\n\nCryo EM Tomography\nhttps://www.youtube.com/watch?v=frgjc-ZNhOY\n\n\nA 3 minute introduction to CryoEM\nhttps://www.youtube.com/watch?v=BJKkC0W-6Qk\n\nMinute Biophysics - Cryo-Electron Tomography (Cryo-E.T.), Tiffany\nhttps://www.youtube.com/watch?v=lf-tWHSbr8Q\n\nNon-averaged 3D structure of a single tetra-nucleosome arrays by Na+ and H1 by cryo-ET and IPET\nhttps://www.youtube.com/watch?v=RIFGnKzK0rk\n\ntilt series\nhttps://www.youtube.com/watch?v=SbEvuSgskWw",
      "votes": 4
    },
    {
      "id": 3042209,
      "postDate": "2024-11-11T10:16:35.373Z",
      "content": "<p>SHREC’20 Benchmark: Classification in cryo-electron tomograms<br>\n<a href=\"https://www.shrec.net/cryo-et/2020/shrec_cryoet2020_preprint.pdf\" target=\"_blank\">https://www.shrec.net/cryo-et/2020/shrec_cryoet2020_preprint.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7df65c7d2c33bb7617cbf7ddc6d31bd6%2FSelection_667.png?generation=1731320168774378&amp;alt=media\" alt=\"\"></p>\n<p>U-net Multi-task Cascade (UMC) seems to be a power house</p>",
      "rawMarkdown": "SHREC’20 Benchmark: Classification in cryo-electron tomograms\nhttps://www.shrec.net/cryo-et/2020/shrec_cryoet2020_preprint.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7df65c7d2c33bb7617cbf7ddc6d31bd6%2FSelection_667.png?generation=1731320168774378&alt=media)\n\nU-net Multi-task Cascade (UMC) seems to be a power house",
      "votes": 1
    },
    {
      "id": 3071138,
      "postDate": "2024-12-13T11:54:43.670Z",
      "content": "<p>interesting paper:</p>\n<p>A surrogate loss function for optimization of F-beta score in binary classification with imbalanced data<br>\n<a href=\"https://paperswithcode.com/paper/a-surrogate-loss-function-for-optimization-of\" target=\"_blank\">https://paperswithcode.com/paper/a-surrogate-loss-function-for-optimization-of</a></p>",
      "rawMarkdown": "interesting paper:\n\nA surrogate loss function for optimization of F-beta score in binary classification with imbalanced data\nhttps://paperswithcode.com/paper/a-surrogate-loss-function-for-optimization-of\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 3071373,
          "postDate": "2024-12-13T17:25:59.440Z",
          "content": "<p>Interesting function, I played around with Tversky loss adjusting beta to increase focus on recall however I haven't seen significant differences. What do you think about another model to be applied after post processing (ignoring computational time for now) to refine objects on which we cc3d to derive the centroid?</p>",
          "rawMarkdown": "Interesting function, I played around with Tversky loss adjusting beta to increase focus on recall however I haven't seen significant differences. What do you think about another model to be applied after post processing (ignoring computational time for now) to refine objects on which we cc3d to derive the centroid?"
        }
      ]
    },
    {
      "id": 3040630,
      "postDate": "2024-11-09T11:24:28.803Z",
      "content": "<blockquote>\n  <p>some efficient self-supervised methods? </p>\n</blockquote>\n<p>I'm feeling this might be the way to go. <br>\nThere is quite a lot of data <a href=\"https://cryoetdataportal.czscience.com/browse-data/datasets\" target=\"_blank\">here</a> (and possible somewhere else?) that is unlabelled, or labelled for different purposes. </p>\n<p>Can we do some sort of MAE pre-training? </p>",
      "rawMarkdown": "> some efficient self-supervised methods? \n\nI'm feeling this might be the way to go. \nThere is quite a lot of data [here](https://cryoetdataportal.czscience.com/browse-data/datasets) (and possible somewhere else?) that is unlabelled, or labelled for different purposes. \n\nCan we do some sort of MAE pre-training? \n",
      "votes": 1,
      "replies": [
        {
          "id": 3040651,
          "postDate": "2024-11-09T12:12:17.487Z",
          "content": "<p>something simpler since we now have the synthetic data generator</p>",
          "rawMarkdown": "something simpler since we now have the synthetic data generator",
          "votes": 1,
          "replies": [
            {
              "id": 3040700,
              "postDate": "2024-11-09T13:19:31.590Z",
              "content": "<p>If you trust the synthetic data then yes you are right. </p>\n<p>In the paper you linked the results look OK, but they also only trained on synthetic. So train on synthetic and then fine-tune on real seems like a good bet as well. </p>",
              "rawMarkdown": "If you trust the synthetic data then yes you are right. \n\nIn the paper you linked the results look OK, but they also only trained on synthetic. So train on synthetic and then fine-tune on real seems like a good bet as well. "
            },
            {
              "id": 3041345,
              "postDate": "2024-11-10T07:48:20.367Z",
              "content": "<p>this might be of interest to you:<br>\nCryoMAE: Few-Shot Cryo-EM Particle Picking with Masked Autoencoders<br>\n<a href=\"https://arxiv.org/pdf/2404.10178\" target=\"_blank\">https://arxiv.org/pdf/2404.10178</a></p>",
              "rawMarkdown": "this might be of interest to you:\nCryoMAE: Few-Shot Cryo-EM Particle Picking with Masked Autoencoders\nhttps://arxiv.org/pdf/2404.10178",
              "votes": 5
            },
            {
              "id": 3041444,
              "postDate": "2024-11-10T11:10:29.217Z",
              "content": "<p>Nice! Thanks, yeah that's pretty close to what I was thinking!</p>",
              "rawMarkdown": "Nice! Thanks, yeah that's pretty close to what I was thinking!"
            }
          ]
        }
      ]
    },
    {
      "id": 3042706,
      "postDate": "2024-11-11T17:31:56.723Z",
      "content": "<p>it seems that my model is just identifying the object by its size. I check the pdb_id 3d model, it seems that there is no noticeable appearance difference (all just look like a ball of proteins?)</p>",
      "rawMarkdown": "it seems that my model is just identifying the object by its size. I check the pdb_id 3d model, it seems that there is no noticeable appearance difference (all just look like a ball of proteins?)",
      "votes": 2,
      "replies": [
        {
          "id": 3042827,
          "postDate": "2024-11-11T19:39:08.977Z",
          "content": "<p>What I have been wondering is if we can sidestep the segmentation approach. </p>\n<p>I.e. the approach I (and most people so far) are using creates segmentation masks for training, trains a segmentation model, and then tries to find the centroids of the segmentation. </p>\n<p>The one problem with this approach is that the segmentation masks aren't given, we're creating spheres based on the type ourselves, and they don't necessarily align with what is in the tomograms.  </p>\n<p>Could we bypass this step and have the model predict the centroids directly? <br>\nI don't have a great idea for how exactly to do so yet. </p>",
          "rawMarkdown": "What I have been wondering is if we can sidestep the segmentation approach. \n\nI.e. the approach I (and most people so far) are using creates segmentation masks for training, trains a segmentation model, and then tries to find the centroids of the segmentation. \n\nThe one problem with this approach is that the segmentation masks aren't given, we're creating spheres based on the type ourselves, and they don't necessarily align with what is in the tomograms.  \n\nCould we bypass this step and have the model predict the centroids directly? \nI don't have a great idea for how exactly to do so yet. ",
          "replies": [
            {
              "id": 3042840,
              "postDate": "2024-11-11T19:55:33.113Z",
              "content": "<p>i did not do segmentation.  just predict centroid</p>",
              "rawMarkdown": "i did not do segmentation.  just predict centroid",
              "votes": 3
            },
            {
              "id": 3042853,
              "postDate": "2024-11-11T20:13:59.887Z",
              "content": "<p>I'm basing my comment on your public notebook, but you still predict class probabilities (or at least logits) per voxel right? </p>\n<p>I see you don't actually predict the class per voxel but just the centroid right away. </p>\n<p>But how do you format your training data? <br>\nDo you do it like the example code and create segmentation masks? </p>",
              "rawMarkdown": "I'm basing my comment on your public notebook, but you still predict class probabilities (or at least logits) per voxel right? \n\nI see you don't actually predict the class per voxel but just the centroid right away. \n\nBut how do you format your training data? \nDo you do it like the example code and create segmentation masks? "
            },
            {
              "id": 3042956,
              "postDate": "2024-11-12T00:47:47.313Z",
              "content": "<p>each voxel is labelled either if it is the centroid of the 6 classes, hence the CE loss</p>\n<p>closet work: detect cell center for counting<br>\n<a href=\"https://www.nature.com/articles/s41598-021-96067-3\" target=\"_blank\">https://www.nature.com/articles/s41598-021-96067-3</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03d1e832f82f5b95ae5fa3e27661e0a6%2FSelection_669.png?generation=1731372455697570&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "each voxel is labelled either if it is the centroid of the 6 classes, hence the CE loss\n\ncloset work: detect cell center for counting\nhttps://www.nature.com/articles/s41598-021-96067-3\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03d1e832f82f5b95ae5fa3e27661e0a6%2FSelection_669.png?generation=1731372455697570&alt=media)",
              "votes": 6
            },
            {
              "id": 3051233,
              "postDate": "2024-11-21T04:00:29.677Z",
              "content": "<p>Hi! I noticed that the model output in your notebook has 7 channels. I assume 6 of them correspond to the centroids of the 6 classes. Could you clarify the purpose of the extra channel? Is it for the background? If so, how should I label it during training?</p>",
              "rawMarkdown": "Hi! I noticed that the model output in your notebook has 7 channels. I assume 6 of them correspond to the centroids of the 6 classes. Could you clarify the purpose of the extra channel? Is it for the background? If so, how should I label it during training?"
            },
            {
              "id": 3051258,
              "postDate": "2024-11-21T05:11:44.750Z",
              "content": "<p>if u use softmax=7 channel, <br>\nor u can use sigmoid=6 channel</p>",
              "rawMarkdown": "if u use softmax=7 channel, \nor u can use sigmoid=6 channel"
            },
            {
              "id": 3055088,
              "postDate": "2024-11-25T14:09:00.190Z",
              "content": "<p>But those masks look as heatmaps. I mean, they're continuous, how to use CE loss to non discrete labels?</p>\n<p>EDIT: Nvm, I've just need to open my idea of CE Loss</p>\n<p><a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html</a></p>\n<p>\"The target that this criterion expects should contain either:</p>\n<p>Class indices in the range <br>\n…</p>\n<p>Probabilities for each class; useful when labels beyond a single class per minibatch item are required, such as for blended labels, label smoothing, etc. The unreduced (i.e. with reduction set to 'none') loss for this case can be described as…\"</p>",
              "rawMarkdown": "But those masks look as heatmaps. I mean, they're continuous, how to use CE loss to non discrete labels?\n\nEDIT: Nvm, I've just need to open my idea of CE Loss\n\nhttps://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html\n\n\"The target that this criterion expects should contain either:\n\nClass indices in the range \n...\n\nProbabilities for each class; useful when labels beyond a single class per minibatch item are required, such as for blended labels, label smoothing, etc. The unreduced (i.e. with reduction set to 'none') loss for this case can be described as...\""
            },
            {
              "id": 3055096,
              "postDate": "2024-11-25T14:16:51.280Z",
              "content": "<p>if you want to use CE, then it is 0,1 binary mask.</p>\n<p>if you want to use gaussian mask, then use JS divergence loss</p>",
              "rawMarkdown": "if you want to use CE, then it is 0,1 binary mask.\n\nif you want to use gaussian mask, then use JS divergence loss",
              "votes": 1
            },
            {
              "id": 3055171,
              "postDate": "2024-11-25T15:55:03.893Z",
              "content": "<p>Thanks for fast answer. I will be checking the sources but meanwhile something I don't understand then is</p>\n<p>\"i did not do segmentation. just predict centroid\"</p>\n<p>\"each voxel is labelled either if it is the centroid of the 6 classes, hence the CE loss\"</p>\n<p>You've been at  some point training with binary masks with only one positive pixel per particle and CE loss? That's not too unforgiving for closer pixels to centroid?</p>\n<p>EDIT: From the source they used</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc2052b6ba87111ec9a6f73fb821b9a2e%2Floss.png?generation=1732550580013733&amp;alt=media\" alt=\"\"></p>",
              "rawMarkdown": "Thanks for fast answer. I will be checking the sources but meanwhile something I don't understand then is\n\n\"i did not do segmentation. just predict centroid\"\n\n\"each voxel is labelled either if it is the centroid of the 6 classes, hence the CE loss\"\n\n You've been at  some point training with binary masks with only one positive pixel per particle and CE loss? That's not too unforgiving for closer pixels to centroid?\n\nEDIT: From the source they used\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc2052b6ba87111ec9a6f73fb821b9a2e%2Floss.png?generation=1732550580013733&alt=media)",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 3111966,
      "postDate": "2025-01-31T20:27:54.480Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , I tried to train a 3d ResNet + U-Net decoder. The 3d ResNet is a pre-trained video model. But I am getting very poor performance. My assumptions was that using pre-trained models should boost performance. Should I freeze or unfreeze the entire ResNet encoder, what is your advice?</p>",
      "rawMarkdown": "@hengck23 , I tried to train a 3d ResNet + U-Net decoder. The 3d ResNet is a pre-trained video model. But I am getting very poor performance. My assumptions was that using pre-trained models should boost performance. Should I freeze or unfreeze the entire ResNet encoder, what is your advice?"
    },
    {
      "id": 3048760,
      "postDate": "2024-11-18T10:34:49.797Z",
      "content": "<p>I'm curious how far did you manage to go with synthetic data? Are the simulators good enough to allow training on simulated data to generalize on real data? Thanks!</p>\n<p>Another TEM simulator exists here:</p>\n<p><a href=\"https://github.com/timothygrant80/cisTEM\" target=\"_blank\">https://github.com/timothygrant80/cisTEM</a></p>\n<p>And is described in the following publication:</p>\n<p><a href=\"https://journals.iucr.org/m/issues/2021/06/00/rq5007/\" target=\"_blank\">https://journals.iucr.org/m/issues/2021/06/00/rq5007/</a></p>",
      "rawMarkdown": "I'm curious how far did you manage to go with synthetic data? Are the simulators good enough to allow training on simulated data to generalize on real data? Thanks!\n\nAnother TEM simulator exists here:\n\nhttps://github.com/timothygrant80/cisTEM\n\nAnd is described in the following publication:\n\nhttps://journals.iucr.org/m/issues/2021/06/00/rq5007/",
      "replies": [
        {
          "id": 3048773,
          "postDate": "2024-11-18T10:45:17.467Z",
          "content": "<p>More useful resources, especially pertaining synthetic data generation, can be found here:</p>\n<p><a href=\"https://github.com/phonchi/Computational-CryoET\" target=\"_blank\">https://github.com/phonchi/Computational-CryoET</a></p>",
          "rawMarkdown": "More useful resources, especially pertaining synthetic data generation, can be found here:\n\nhttps://github.com/phonchi/Computational-CryoET\n"
        },
        {
          "id": 3048806,
          "postDate": "2024-11-18T11:20:22.383Z",
          "content": "<p>i have started working on the simulated data.<br>\nmy judgement is that the \"proper\" use of sythetic data will be a booster to the score.<br>\ni am estimating someting like:<br>\nwithout sythetic : lb score=0.75<br>\nwith sythetic: add 0.05 to 0.10</p>",
          "rawMarkdown": "i have started working on the simulated data.\nmy judgement is that the \"proper\" use of sythetic data will be a booster to the score.\ni am estimating someting like:\nwithout sythetic : lb score=0.75\nwith sythetic: add 0.05 to 0.10",
          "votes": 5,
          "replies": [
            {
              "id": 3063236,
              "postDate": "2024-12-04T09:42:41.087Z",
              "content": "<p>Thanks for your sharing :) </p>",
              "rawMarkdown": "Thanks for your sharing :) "
            }
          ]
        }
      ]
    },
    {
      "id": 3078338,
      "postDate": "2024-12-22T07:20:53.113Z",
      "content": "<p>Thank you very much for sharing</p>",
      "rawMarkdown": "Thank you very much for sharing"
    }
  ],
  "comments": [
    {
      "id": 3057240,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-27T22:38:55.387000",
      "content": "<p>a list of winning models</p>\n<p>SHREC 2021: CLASSIFICATION IN CRYO-ELECTRON TOMOGRAMS<br>\n<a href=\"https://arxiv.org/pdf/2203.10035\" target=\"_blank\">https://arxiv.org/pdf/2203.10035</a></p>\n<p>SHREC’19 Track: Classification in Cryo-Electron Tomograms<br>\n<a href=\"https://webspace.science.uu.nl/~veltk101/publications/art/shrec2019-cryo.pdf\" target=\"_blank\">https://webspace.science.uu.nl/~veltk101/publications/art/shrec2019-cryo.pdf</a></p>",
      "votes": 10,
      "replies": []
    },
    {
      "id": 3054691,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-25T02:22:13.190000",
      "content": "<p>most likely, the winning solution would be two-stage approach.<br>\nstage 1:</p>\n<ul>\n<li>segmentation net to predict coords of candidate particles<br>\nstage 2:</li>\n<li>crop the volume to verify: rank the most likely candidates and refine coordinate</li>\n</ul>\n<p>for stage2:</p>\n<ul>\n<li>size/volume of prediction is important for filtering results</li>\n<li>you can upsize the crop for better accuracy</li>\n<li>how to combine scores of stage.1 and stage.2?</li>\n</ul>\n<hr>\n<p>you probably need to do a lot of probing. we actually need to know the recall and precision of the lb submission.<br>\nthis is how:</p>\n<ul>\n<li>submit for one particle: one equation for recall and precision</li>\n<li>same as above, but submit known number of false positives(e.g. just take corner point or background point or just out of volume point or (0,0,0)) : another equation for recall and precision</li>\n</ul>",
      "votes": 8,
      "replies": [
        {
          "id": 3085686,
          "author_name": "Keesari Vigneshwar Reddy",
          "author_url": "",
          "post_date": "2025-01-01T12:48:21.140000",
          "content": "<h3>Any guidance on approach in stage 2</h3>\n<p>The original way seems to be very complex and computationally expensive.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F15856017%2F0392b522e900fb7bec916c7aa890a353%2FScreenshot%202025-01-01%20181656.png?generation=1735735646288793&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3054647,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-25T00:02:51.337000",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F749225bf11d4e2dc3359c05ebb1b2030%2Fezgif-3-a60846ca1e.gif?generation=1732492937995569&amp;alt=media\" alt=\"\"></p>\n<p>how it look in 3d<br>\n(volume rendering using pyvista)</p>",
      "votes": 7,
      "replies": []
    },
    {
      "id": 3056473,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-27T02:18:10.780000",
      "content": "<p>contribution of lb score:</p>\n<p>baseline(ttax2) 0.702<br>\nbetter model parameters(ttax2): 0.718<br>\ntta, 2xrot90 : add 0.05<br>\ntta, 4xrot90 : add 0.10<br>\nensemble (same fold, different base encoder):  0.718+0.713 = 0.743</p>",
      "votes": 5,
      "replies": []
    },
    {
      "id": 3049275,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-18T22:54:22.680000",
      "content": "<p>handling border artifacts is always a trick to win kaggle<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65e270024f58c752eaf67d9ad7a183fd%2FSelection_705.png?generation=1731970460522047&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 3051295,
          "author_name": "liuxiaoonline",
          "author_url": "",
          "post_date": "2024-11-21T06:09:44.063000",
          "content": "<p>Hello author, I would like to ask if this artifact is the one produced during the 3D reconstruction process of your synthetic data?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3051300,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-21T06:13:27.460000",
              "content": "<p>results of </p>\n<pre><code>            probability[:, z:z + num_slice, : + image_size, : + image_size] += prob\n            count[:, z:z + num_slice, : + image_size, : + image_size] += \n</code></pre>\n<p>correct version</p>\n<pre><code>            probability[:, z:z + num_slice, : + image_size, : + image_size] += weight * prob\n            count[:, z:z + num_slice, : + image_size, : + image_size] += weight  #weight is shape =(:,num_slice,image_size, image_size)\n</code></pre>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3051306,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-21T06:18:53.037000",
              "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fab423841d1da42cd294c127bedf8582f%2FSelection_717.png?generation=1732169930925251&amp;alt=media\" alt=\"\"></p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3086673,
              "author_name": "Sarvagya Porwal",
              "author_url": "",
              "post_date": "2025-01-02T15:17:57.330000",
              "content": "<p>Pls share some resources for post prediction of overlapping patches…</p>\n<p>Superly needed (if its right grammatically)<br>\nthanks</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3047384,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-16T16:00:50.687000",
      "content": "<p>first baseline experiment results are up:<br>\nupdated on 17-oct</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F806ebf376130887c9acf397bfeee9865%2FSelection_691.png?generation=1731784529142531&amp;alt=media\" alt=\"\"></p>",
      "votes": 5,
      "replies": [
        {
          "id": 3047391,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-11-16T16:14:05",
          "content": "<p>Using a 2D encoder and a 3D decoder (I think that's what you're doing?) is pretty interesting! Nice work <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, I always enjoy learning from your discussion posts.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3047416,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-16T16:35:46.897000",
              "content": "<p><a href=\"https://www.kaggle.com/code/hengck23/2d-to-3d-unet-demo\" target=\"_blank\">https://www.kaggle.com/code/hengck23/2d-to-3d-unet-demo</a><br>\nThis is an old version. </p>\n<p>The cryoET version is slightly different. you can refer to <a href=\"https://www.kaggle.com/datasets/hengck23/hengck-czii-cryo-et-01\" target=\"_blank\">https://www.kaggle.com/datasets/hengck23/hengck-czii-cryo-et-01</a></p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3048958,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-18T14:54:30.110000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3047420,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-16T16:42:43.733000",
          "content": "<p>training log<br>\n(see file attached)</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F2c1fac7d7ad70a4d42b6d15b8bc88dac%2FSelection_686.png?generation=1731775361123507&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": [
            {
              "id": 3047425,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-11-16T16:48:51.253000",
              "content": "<p>sorry if it’s too obvious, what do these 92 iterations per epoch mean?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3050682,
              "author_name": "Steven_Y",
              "author_url": "",
              "post_date": "2024-11-20T13:06:12.140000",
              "content": "<p>May I ask how did you process the data? How did you get 92 iterations per epoch? Thanks!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3051141,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-20T22:57:50.270000",
              "content": "<p>that is not important. it just says that for one epoch, 92 batches are randomly created. each batch has 3 samples  (volume crop)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3047477,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-16T18:10:31.870000",
          "content": "<p>threshold versus  f-beta score</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F10aa8ad2df95befcc332ca721731bdc8%2FSelection_690.png?generation=1731780624993380&amp;alt=media\" alt=\"\"></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3057922,
          "author_name": "Ivan Panshin",
          "author_url": "",
          "post_date": "2024-11-28T20:14:53.280000",
          "content": "<p>Do you report \"local LB\" (which is another name for validation, I guess) for predictions and GT divided by 10 or in the original space (similar to how you would post-process predictions during submit)? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3057976,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-28T21:22:00.267000",
              "content": "<p>yes, local LB is local CV. using  official mertic  code <a href=\"https://www.kaggle.com/code/metric/czi-cryoet-84969\" target=\"_blank\">https://www.kaggle.com/code/metric/czi-cryoet-84969</a><br>\nso it is what you submitted in csv </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3058232,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-29T08:46:53.857000",
              "content": "<p>Correct me if I'm wrong. But I've been recently checking metric, and since positives are predictions inside .5 times particle radius, It doesn't matter the scale (they just need both be the same, GT and predictions).</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3058279,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-29T09:28:14.227000",
              "content": "<p>slight numerical error exists though both metric will be very close</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3058691,
              "author_name": "Ivan Panshin",
              "author_url": "",
              "post_date": "2024-11-29T19:45:06.543000",
              "content": "<p>Doesn't seem right to me. </p>\n<p>Suppose your predictions are off by 5 pixels in 10A. It means that they are off by 50 pixels in 1A. So if you're using the same evaluation code with the same radius values, metrics will be vastly different. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3058717,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-29T20:41:25.263000",
              "content": "<p>In or out is with respect .5 times the particle radius. The particle radius scales with the predictions. You just need to esure that particle radius inside metric are in the same scale than your predictions. At kaggle submission is only one way. But at home, you chose.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3058719,
              "author_name": "Ivan Panshin",
              "author_url": "",
              "post_date": "2024-11-29T20:44:06.433000",
              "content": "<p>Yeah, exactly, my point is that radius values are given for 1A and it's easy to use them for metric calculation at 10A by mistake. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3058726,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-29T20:56:00.997000",
              "content": "<p>And that's exactly what I did haha I've noticed my mistake right after answer. Thanks for point it out.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3058728,
              "author_name": "Ivan Panshin",
              "author_url": "",
              "post_date": "2024-11-29T21:03:39.967000",
              "content": "<p>I did it as well… It was painful to find the mistake </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3060110,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-12-01T12:05:08.380000",
              "content": "<p>We've got two score fixes already so will probably everything be fine and it's just me but… Why my CV scores with overoptimistic predictions (that's with predictions and GT in 10A and radius in 1A, a radius 10 times bigger than it should be) correlate better with public LB than the correct ones (everything in 1A).</p>\n<p>Public LB and overoptimistic ~.5</p>\n<p>correct scales ~.1</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3041344,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-10T07:47:27.680000",
      "content": "<p>i will come back to this later. but we do need a fast and robust 3d peak detector<br>\n<a href=\"https://stackoverflow.com/questions/3684484/peak-detection-in-a-2d-array\" target=\"_blank\">https://stackoverflow.com/questions/3684484/peak-detection-in-a-2d-array</a></p>",
      "votes": 6,
      "replies": []
    },
    {
      "id": 3057131,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-27T18:35:00.257000",
      "content": "<p>3d object detection toolkit, with 3d roi align and nms<br>\n<a href=\"https://github.com/TimothyZero/MedVision/tree/main\" target=\"_blank\">https://github.com/TimothyZero/MedVision/tree/main</a><br>\n<a href=\"https://github.com/pytorch/vision/issues/2402\" target=\"_blank\">https://github.com/pytorch/vision/issues/2402</a></p>",
      "votes": 3,
      "replies": [
        {
          "id": 3091158,
          "author_name": "Keesari Vigneshwar Reddy",
          "author_url": "",
          "post_date": "2025-01-08T06:09:05.413000",
          "content": "<p><a href=\"https://github.com/MIC-DKFZ/batchgenerators\" target=\"_blank\">https://github.com/MIC-DKFZ/batchgenerators</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3055872,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-26T08:09:23.903000",
      "content": "<p><a href=\"https://github.com/icthrm/SC-Net\" target=\"_blank\">https://github.com/icthrm/SC-Net</a><br>\nself supervised denoise</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3051142,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-20T23:00:23.383000",
      "content": "<p>my early experiment:<br>\nresnet3d + unet-decoder3d (66mb)<br>\n-local cv 0.694, lb=0.666</p>\n<p>as reference<br>\nresnet18d + unet-decoder3d (88mb)<br>\nresnet34d + unet-decoder3d (123mb)</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3053670,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-23T18:15:29.710000",
      "content": "<p>learning curve: </p>\n<ul>\n<li>if you submit a non-overfitted model, you will see cv vs lb correlation</li>\n<li>it is easy to get high cv (e.g. in the range of 0.77 or even 0.82). but we are not interested in that.</li>\n<li>rather, we are interested to get a cv that <strong>we know</strong> is close to lb</li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F485fa0824f57310ba5e8422e81edc678%2FSelection_731.png?generation=1732385724424357&amp;alt=media\" alt=\"\"></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3041079,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-09T22:40:30.117000",
      "content": "<p>CryoTransformer: A Transformer Model for Picking Protein Particles from Cryo-EM Micrographs<br>\n<a href=\"https://pmc.ncbi.nlm.nih.gov/articles/PMC10634673/pdf/nihpp-2023.10.19.563155v1.pdf\" target=\"_blank\">https://pmc.ncbi.nlm.nih.gov/articles/PMC10634673/pdf/nihpp-2023.10.19.563155v1.pdf</a><br>\n<a href=\"https://github.com/jianlin-cheng/CryoTransformer\" target=\"_blank\">https://github.com/jianlin-cheng/CryoTransformer</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3040609,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-09T10:44:59.937000",
      "content": "<p>make your own data?<br>\n<a href=\"https://github.com/cellcanvas/album-catalog/blob/main/solutions/polnet/generate-tomogram/solution.py\" target=\"_blank\">https://github.com/cellcanvas/album-catalog/blob/main/solutions/polnet/generate-tomogram/solution.py</a></p>\n<p><a href=\"https://github.com/anmartinezs/polnet\" target=\"_blank\">https://github.com/anmartinezs/polnet</a><br>\n[1] Martinez-Sanchez A.*, and Lamm L., Jasnin M. and Phelippeau H. (2024) \"Simulating the cellular context in synthetic datasets for cryo-electron tomography\" IEEE Transactions on Medical Imaging </p>\n<hr>\n<p><a href=\"https://github.com/phonchi/Computational-CryoET\" target=\"_blank\">https://github.com/phonchi/Computational-CryoET</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3040323,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-09T03:00:21.273000",
      "content": "<p>a simple notebook for benchmarking:<br>\n<a href=\"https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder\" target=\"_blank\">https://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 3041568,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-10T14:19:40.523000",
      "content": "<p><a href=\"https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\" target=\"_blank\">https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks</a></p>\n<p>more example network!!!<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F50564f4f210d4ed09f9c14ce35ed513d%2FSelection_663.png?generation=1731248372945421&amp;alt=media\" alt=\"\"></p>\n<p>Open-source Tools for CryoET Particle Picking Machine Learning Competitions<br>\n<a href=\"https://www.researchgate.net/publication/385559021_Open-source_Tools_for_CryoET_Particle_Picking_Machine_Learning_Competitions\" target=\"_blank\">https://www.researchgate.net/publication/385559021_Open-source_Tools_for_CryoET_Particle_Picking_Machine_Learning_Competitions</a> </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3041550,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-10T14:04:22.950000",
      "content": "<p>important details from dataset paper:</p>\n<p>Annotating CryoET Volumes: A Machine Learning Challenge <br>\n<a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full.pdf\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full.pdf</a></p>\n<p>\"This set was generated by mixing varying percentages (in 10% 378 increments) of picks from a fine-tuned DeepFindET model and ground truth data. By comparing their F-beta scores, we observed that datasets containing 80% or more ground-truth results yielded mock datasets that score better than DeepFindET\"</p>\n<p>is DeepFindET used in annotations? YES! … and other tools …</p>\n<p>\"The tools described above, and other published software packages were stitched together into several workflows to generate “ground truth” labels \"</p>\n<hr>",
      "votes": 4,
      "replies": [
        {
          "id": 3043440,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-12T12:23:42.650000",
          "content": "<p>please read the paper in details. use chatgpt to help you!</p>\n<p>i was at first pretty confused about phatom datset (is it real or not, is it pyhsical or software simulated ) ….</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F258d203f385aaf78e7780047f73ca4a7%2FSelection_999(6791).png?generation=1731414132696051&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F406d112248a2f25a2fab3c504e1acb01%2FSelection_999(6792).png?generation=1731414144754747&amp;alt=media\" alt=\"\"></p>\n<p>you solution may be quite of reverse enginerring the annotation process</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7f5523228d52061170267d223c3d66fe%2FSelection_999(6793).png?generation=1731414155492008&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F1c2fb5e7eee99120d17dae885e4529b9%2FSelection_999(6794).png?generation=1731414220848937&amp;alt=media\" alt=\"\"></p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 3083301,
          "author_name": "Keesari Vigneshwar Reddy",
          "author_url": "",
          "post_date": "2024-12-29T09:16:35.210000",
          "content": "<p>This conversation may be useful:</p>\n<p><a href=\"https://chatgpt.com/share/6771132f-6638-8009-b1e8-c36b76d36af4\" target=\"_blank\">https://chatgpt.com/share/6771132f-6638-8009-b1e8-c36b76d36af4</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3041062,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-09T22:09:00.543000",
      "content": "<p>introduction video<br>\n<a href=\"https://cryoem101.org/chapter-1-et/\" target=\"_blank\">https://cryoem101.org/chapter-1-et/</a></p>\n<p>Electron Tomography<br>\n<a href=\"https://www.youtube.com/watch?v=M2D3lDIv9jY\" target=\"_blank\">https://www.youtube.com/watch?v=M2D3lDIv9jY</a></p>\n<p>Cryo EM Tomography<br>\n<a href=\"https://www.youtube.com/watch?v=frgjc-ZNhOY\" target=\"_blank\">https://www.youtube.com/watch?v=frgjc-ZNhOY</a></p>\n<p>A 3 minute introduction to CryoEM<br>\n<a href=\"https://www.youtube.com/watch?v=BJKkC0W-6Qk\" target=\"_blank\">https://www.youtube.com/watch?v=BJKkC0W-6Qk</a></p>\n<p>Minute Biophysics - Cryo-Electron Tomography (Cryo-E.T.), Tiffany<br>\n<a href=\"https://www.youtube.com/watch?v=lf-tWHSbr8Q\" target=\"_blank\">https://www.youtube.com/watch?v=lf-tWHSbr8Q</a></p>\n<p>Non-averaged 3D structure of a single tetra-nucleosome arrays by Na+ and H1 by cryo-ET and IPET<br>\n<a href=\"https://www.youtube.com/watch?v=RIFGnKzK0rk\" target=\"_blank\">https://www.youtube.com/watch?v=RIFGnKzK0rk</a></p>\n<p>tilt series<br>\n<a href=\"https://www.youtube.com/watch?v=SbEvuSgskWw\" target=\"_blank\">https://www.youtube.com/watch?v=SbEvuSgskWw</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 3042209,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-11T10:16:35.373000",
      "content": "<p>SHREC’20 Benchmark: Classification in cryo-electron tomograms<br>\n<a href=\"https://www.shrec.net/cryo-et/2020/shrec_cryoet2020_preprint.pdf\" target=\"_blank\">https://www.shrec.net/cryo-et/2020/shrec_cryoet2020_preprint.pdf</a></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7df65c7d2c33bb7617cbf7ddc6d31bd6%2FSelection_667.png?generation=1731320168774378&amp;alt=media\" alt=\"\"></p>\n<p>U-net Multi-task Cascade (UMC) seems to be a power house</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3071138,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-12-13T11:54:43.670000",
      "content": "<p>interesting paper:</p>\n<p>A surrogate loss function for optimization of F-beta score in binary classification with imbalanced data<br>\n<a href=\"https://paperswithcode.com/paper/a-surrogate-loss-function-for-optimization-of\" target=\"_blank\">https://paperswithcode.com/paper/a-surrogate-loss-function-for-optimization-of</a></p>",
      "votes": 2,
      "replies": [
        {
          "id": 3071373,
          "author_name": "Andrei Zamfir",
          "author_url": "",
          "post_date": "2024-12-13T17:25:59.440000",
          "content": "<p>Interesting function, I played around with Tversky loss adjusting beta to increase focus on recall however I haven't seen significant differences. What do you think about another model to be applied after post processing (ignoring computational time for now) to refine objects on which we cc3d to derive the centroid?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3040630,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2024-11-09T11:24:28.803000",
      "content": "<blockquote>\n  <p>some efficient self-supervised methods? </p>\n</blockquote>\n<p>I'm feeling this might be the way to go. <br>\nThere is quite a lot of data <a href=\"https://cryoetdataportal.czscience.com/browse-data/datasets\" target=\"_blank\">here</a> (and possible somewhere else?) that is unlabelled, or labelled for different purposes. </p>\n<p>Can we do some sort of MAE pre-training? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3040651,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-09T12:12:17.487000",
          "content": "<p>something simpler since we now have the synthetic data generator</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3040700,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2024-11-09T13:19:31.590000",
              "content": "<p>If you trust the synthetic data then yes you are right. </p>\n<p>In the paper you linked the results look OK, but they also only trained on synthetic. So train on synthetic and then fine-tune on real seems like a good bet as well. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3041345,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-10T07:48:20.367000",
              "content": "<p>this might be of interest to you:<br>\nCryoMAE: Few-Shot Cryo-EM Particle Picking with Masked Autoencoders<br>\n<a href=\"https://arxiv.org/pdf/2404.10178\" target=\"_blank\">https://arxiv.org/pdf/2404.10178</a></p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 3041444,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2024-11-10T11:10:29.217000",
              "content": "<p>Nice! Thanks, yeah that's pretty close to what I was thinking!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3042706,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-11T17:31:56.723000",
      "content": "<p>it seems that my model is just identifying the object by its size. I check the pdb_id 3d model, it seems that there is no noticeable appearance difference (all just look like a ball of proteins?)</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3042827,
          "author_name": "fnands",
          "author_url": "",
          "post_date": "2024-11-11T19:39:08.977000",
          "content": "<p>What I have been wondering is if we can sidestep the segmentation approach. </p>\n<p>I.e. the approach I (and most people so far) are using creates segmentation masks for training, trains a segmentation model, and then tries to find the centroids of the segmentation. </p>\n<p>The one problem with this approach is that the segmentation masks aren't given, we're creating spheres based on the type ourselves, and they don't necessarily align with what is in the tomograms.  </p>\n<p>Could we bypass this step and have the model predict the centroids directly? <br>\nI don't have a great idea for how exactly to do so yet. </p>",
          "votes": 0,
          "replies": [
            {
              "id": 3042840,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-11T19:55:33.113000",
              "content": "<p>i did not do segmentation.  just predict centroid</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3042853,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2024-11-11T20:13:59.887000",
              "content": "<p>I'm basing my comment on your public notebook, but you still predict class probabilities (or at least logits) per voxel right? </p>\n<p>I see you don't actually predict the class per voxel but just the centroid right away. </p>\n<p>But how do you format your training data? <br>\nDo you do it like the example code and create segmentation masks? </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3042956,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-12T00:47:47.313000",
              "content": "<p>each voxel is labelled either if it is the centroid of the 6 classes, hence the CE loss</p>\n<p>closet work: detect cell center for counting<br>\n<a href=\"https://www.nature.com/articles/s41598-021-96067-3\" target=\"_blank\">https://www.nature.com/articles/s41598-021-96067-3</a><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F03d1e832f82f5b95ae5fa3e27661e0a6%2FSelection_669.png?generation=1731372455697570&amp;alt=media\" alt=\"\"></p>",
              "votes": 6,
              "replies": []
            },
            {
              "id": 3051233,
              "author_name": "Bianco Chiu",
              "author_url": "",
              "post_date": "2024-11-21T04:00:29.677000",
              "content": "<p>Hi! I noticed that the model output in your notebook has 7 channels. I assume 6 of them correspond to the centroids of the 6 classes. Could you clarify the purpose of the extra channel? Is it for the background? If so, how should I label it during training?</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3051258,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-21T05:11:44.750000",
              "content": "<p>if u use softmax=7 channel, <br>\nor u can use sigmoid=6 channel</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3055088,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-25T14:09:00.190000",
              "content": "<p>But those masks look as heatmaps. I mean, they're continuous, how to use CE loss to non discrete labels?</p>\n<p>EDIT: Nvm, I've just need to open my idea of CE Loss</p>\n<p><a href=\"https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html\" target=\"_blank\">https://pytorch.org/docs/stable/generated/torch.nn.CrossEntropyLoss.html</a></p>\n<p>\"The target that this criterion expects should contain either:</p>\n<p>Class indices in the range <br>\n…</p>\n<p>Probabilities for each class; useful when labels beyond a single class per minibatch item are required, such as for blended labels, label smoothing, etc. The unreduced (i.e. with reduction set to 'none') loss for this case can be described as…\"</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3055096,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-25T14:16:51.280000",
              "content": "<p>if you want to use CE, then it is 0,1 binary mask.</p>\n<p>if you want to use gaussian mask, then use JS divergence loss</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3055171,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-25T15:55:03.893000",
              "content": "<p>Thanks for fast answer. I will be checking the sources but meanwhile something I don't understand then is</p>\n<p>\"i did not do segmentation. just predict centroid\"</p>\n<p>\"each voxel is labelled either if it is the centroid of the 6 classes, hence the CE loss\"</p>\n<p>You've been at  some point training with binary masks with only one positive pixel per particle and CE loss? That's not too unforgiving for closer pixels to centroid?</p>\n<p>EDIT: From the source they used</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8722753%2Fc2052b6ba87111ec9a6f73fb821b9a2e%2Floss.png?generation=1732550580013733&amp;alt=media\" alt=\"\"></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3111966,
      "author_name": "Success Moses",
      "author_url": "",
      "post_date": "2025-01-31T20:27:54.480000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , I tried to train a 3d ResNet + U-Net decoder. The 3d ResNet is a pre-trained video model. But I am getting very poor performance. My assumptions was that using pre-trained models should boost performance. Should I freeze or unfreeze the entire ResNet encoder, what is your advice?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3048760,
      "author_name": "umarell",
      "author_url": "",
      "post_date": "2024-11-18T10:34:49.797000",
      "content": "<p>I'm curious how far did you manage to go with synthetic data? Are the simulators good enough to allow training on simulated data to generalize on real data? Thanks!</p>\n<p>Another TEM simulator exists here:</p>\n<p><a href=\"https://github.com/timothygrant80/cisTEM\" target=\"_blank\">https://github.com/timothygrant80/cisTEM</a></p>\n<p>And is described in the following publication:</p>\n<p><a href=\"https://journals.iucr.org/m/issues/2021/06/00/rq5007/\" target=\"_blank\">https://journals.iucr.org/m/issues/2021/06/00/rq5007/</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 3048773,
          "author_name": "umarell",
          "author_url": "",
          "post_date": "2024-11-18T10:45:17.467000",
          "content": "<p>More useful resources, especially pertaining synthetic data generation, can be found here:</p>\n<p><a href=\"https://github.com/phonchi/Computational-CryoET\" target=\"_blank\">https://github.com/phonchi/Computational-CryoET</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3048806,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-18T11:20:22.383000",
          "content": "<p>i have started working on the simulated data.<br>\nmy judgement is that the \"proper\" use of sythetic data will be a booster to the score.<br>\ni am estimating someting like:<br>\nwithout sythetic : lb score=0.75<br>\nwith sythetic: add 0.05 to 0.10</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3063236,
              "author_name": "SenTran",
              "author_url": "",
              "post_date": "2024-12-04T09:42:41.087000",
              "content": "<p>Thanks for your sharing :) </p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3078338,
      "author_name": "hukaixin",
      "author_url": "",
      "post_date": "2024-12-22T07:20:53.113000",
      "content": "<p>Thank you very much for sharing</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3040321": "this is in progress and will be updated often.\n\ninspiration:\n- did you know that you can pretrain imagenet model without using real images? This is one famous work using fractal images:\n\nhttps://hirokatsukataoka16.github.io/Pretraining-without-Natural-Images/\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2Fbeaecb6153485f554f602842238e6e95%2FSelection_649.png?generation=1731120390177754&alt=media)\n\n- pertaining with synthetic images is valuable when it is hard to obtain real images (and their labels). Examples include seismic data, tomography data, etc\n\n---\n\nmy plan:\n1. synthetic data :  \n   - from CryoET Data Portal\n   - find synthetic data generator to create my own data with more noise, more object orinetation and more particle objects\n2. some efficient self-supervised methods?  \n3. transformer-based foundation model. Transformers perform very well if there is massive data.\n3. Transfer learning to n-shot using real kaggle train data and make submission\n",
    "3057240": "a list of winning models\n\nSHREC 2021: CLASSIFICATION IN CRYO-ELECTRON TOMOGRAMS\nhttps://arxiv.org/pdf/2203.10035\n\n\nSHREC’19 Track: Classification in Cryo-Electron Tomograms\nhttps://webspace.science.uu.nl/~veltk101/publications/art/shrec2019-cryo.pdf\n",
    "3054691": "most likely, the winning solution would be two-stage approach.\nstage 1:\n- segmentation net to predict coords of candidate particles\nstage 2:\n- crop the volume to verify: rank the most likely candidates and refine coordinate\n\nfor stage2:\n- size/volume of prediction is important for filtering results\n- you can upsize the crop for better accuracy\n- how to combine scores of stage.1 and stage.2?\n\n----\n\nyou probably need to do a lot of probing. we actually need to know the recall and precision of the lb submission.\nthis is how:\n- submit for one particle: one equation for recall and precision\n- same as above, but submit known number of false positives(e.g. just take corner point or background point or just out of volume point or (0,0,0)) : another equation for recall and precision\n\n\n",
    "3054647": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F749225bf11d4e2dc3359c05ebb1b2030%2Fezgif-3-a60846ca1e.gif?generation=1732492937995569&alt=media)\n\nhow it look in 3d\n(volume rendering using pyvista)",
    "3056473": "contribution of lb score:\n\nbaseline(ttax2) 0.702\nbetter model parameters(ttax2): 0.718\ntta, 2xrot90 : add 0.05\ntta, 4xrot90 : add 0.10\nensemble (same fold, different base encoder):  0.718+0.713 = 0.743\n",
    "3049275": "handling border artifacts is always a trick to win kaggle\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F65e270024f58c752eaf67d9ad7a183fd%2FSelection_705.png?generation=1731970460522047&alt=media)",
    "3047384": "first baseline experiment results are up:\nupdated on 17-oct\n\n ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F806ebf376130887c9acf397bfeee9865%2FSelection_691.png?generation=1731784529142531&alt=media)",
    "3041344": "i will come back to this later. but we do need a fast and robust 3d peak detector\nhttps://stackoverflow.com/questions/3684484/peak-detection-in-a-2d-array",
    "3057131": "3d object detection toolkit, with 3d roi align and nms\nhttps://github.com/TimothyZero/MedVision/tree/main\nhttps://github.com/pytorch/vision/issues/2402",
    "3055872": "https://github.com/icthrm/SC-Net\nself supervised denoise",
    "3051142": "my early experiment:\nresnet3d + unet-decoder3d (66mb)\n-local cv 0.694, lb=0.666\n\n\nas reference\nresnet18d + unet-decoder3d (88mb)\nresnet34d + unet-decoder3d (123mb)",
    "3053670": "learning curve: \n- if you submit a non-overfitted model, you will see cv vs lb correlation\n- it is easy to get high cv (e.g. in the range of 0.77 or even 0.82). but we are not interested in that.\n- rather, we are interested to get a cv that **we know** is close to lb\n\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F485fa0824f57310ba5e8422e81edc678%2FSelection_731.png?generation=1732385724424357&alt=media)",
    "3041079": "CryoTransformer: A Transformer Model for Picking Protein Particles from Cryo-EM Micrographs\nhttps://pmc.ncbi.nlm.nih.gov/articles/PMC10634673/pdf/nihpp-2023.10.19.563155v1.pdf\nhttps://github.com/jianlin-cheng/CryoTransformer",
    "3040609": "make your own data?\nhttps://github.com/cellcanvas/album-catalog/blob/main/solutions/polnet/generate-tomogram/solution.py\n\nhttps://github.com/anmartinezs/polnet\n[1] Martinez-Sanchez A.*, and Lamm L., Jasnin M. and Phelippeau H. (2024) \"Simulating the cellular context in synthetic datasets for cryo-electron tomography\" IEEE Transactions on Medical Imaging \n\n---\nhttps://github.com/phonchi/Computational-CryoET",
    "3040323": "a simple notebook for benchmarking:\nhttps://www.kaggle.com/code/hengck23/3d-unet-using-2d-image-encoder",
    "3041568": "https://github.com/czimaginginstitute/2024_czii_mlchallenge_notebooks\n\nmore example network!!!\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F50564f4f210d4ed09f9c14ce35ed513d%2FSelection_663.png?generation=1731248372945421&alt=media)\n\n\nOpen-source Tools for CryoET Particle Picking Machine Learning Competitions\nhttps://www.researchgate.net/publication/385559021_Open-source_Tools_for_CryoET_Particle_Picking_Machine_Learning_Competitions ",
    "3041550": "important details from dataset paper:\n\nAnnotating CryoET Volumes: A Machine Learning Challenge \nhttps://www.biorxiv.org/content/10.1101/2024.11.04.621686v1.full.pdf\n\n\"This set was generated by mixing varying percentages (in 10% 378 increments) of picks from a fine-tuned DeepFindET model and ground truth data. By comparing their F-beta scores, we observed that datasets containing 80% or more ground-truth results yielded mock datasets that score better than DeepFindET\"\n\nis DeepFindET used in annotations? YES! ... and other tools ...\n\n\"The tools described above, and other published software packages were stitched together into several workflows to generate “ground truth” labels \"\n\n\n----\n ",
    "3041062": "introduction video\nhttps://cryoem101.org/chapter-1-et/\n\nElectron Tomography\nhttps://www.youtube.com/watch?v=M2D3lDIv9jY\n\nCryo EM Tomography\nhttps://www.youtube.com/watch?v=frgjc-ZNhOY\n\n\nA 3 minute introduction to CryoEM\nhttps://www.youtube.com/watch?v=BJKkC0W-6Qk\n\nMinute Biophysics - Cryo-Electron Tomography (Cryo-E.T.), Tiffany\nhttps://www.youtube.com/watch?v=lf-tWHSbr8Q\n\nNon-averaged 3D structure of a single tetra-nucleosome arrays by Na+ and H1 by cryo-ET and IPET\nhttps://www.youtube.com/watch?v=RIFGnKzK0rk\n\ntilt series\nhttps://www.youtube.com/watch?v=SbEvuSgskWw",
    "3042209": "SHREC’20 Benchmark: Classification in cryo-electron tomograms\nhttps://www.shrec.net/cryo-et/2020/shrec_cryoet2020_preprint.pdf\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F113660%2F7df65c7d2c33bb7617cbf7ddc6d31bd6%2FSelection_667.png?generation=1731320168774378&alt=media)\n\nU-net Multi-task Cascade (UMC) seems to be a power house",
    "3071138": "interesting paper:\n\nA surrogate loss function for optimization of F-beta score in binary classification with imbalanced data\nhttps://paperswithcode.com/paper/a-surrogate-loss-function-for-optimization-of\n\n",
    "3040630": "> some efficient self-supervised methods? \n\nI'm feeling this might be the way to go. \nThere is quite a lot of data [here](https://cryoetdataportal.czscience.com/browse-data/datasets) (and possible somewhere else?) that is unlabelled, or labelled for different purposes. \n\nCan we do some sort of MAE pre-training? \n",
    "3042706": "it seems that my model is just identifying the object by its size. I check the pdb_id 3d model, it seems that there is no noticeable appearance difference (all just look like a ball of proteins?)",
    "3111966": "@hengck23 , I tried to train a 3d ResNet + U-Net decoder. The 3d ResNet is a pre-trained video model. But I am getting very poor performance. My assumptions was that using pre-trained models should boost performance. Should I freeze or unfreeze the entire ResNet encoder, what is your advice?",
    "3048760": "I'm curious how far did you manage to go with synthetic data? Are the simulators good enough to allow training on simulated data to generalize on real data? Thanks!\n\nAnother TEM simulator exists here:\n\nhttps://github.com/timothygrant80/cisTEM\n\nAnd is described in the following publication:\n\nhttps://journals.iucr.org/m/issues/2021/06/00/rq5007/",
    "3078338": "Thank you very much for sharing"
  }
}