{
  "id": 547350,
  "title": "if you are stuck below the benchmark.csv, read this!",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/547350",
  "author_name": "hengck23",
  "post_date": "2024-11-21T04:12:21.130000",
  "votes": 98,
  "comment_count": 34,
  "views": 0,
  "content": "<p>a comfortable lb score without many bells and whistles is about 0.630. If you are getting less than benchmark.csv score of 0.587, the below may help</p>\n<h2>1. design a workable end-to-end framework. Most important details are:</h2>\n<h3>scanning window:</h3>\n<ul>\n<li>it is not possible to process 184x630x630 directly and you are likely to break it into smaller scanning \"windows\" and feed them into your network, and then fuse the predictions together into one later.</li>\n</ul>\n<h3>coordinates extraction</h3>\n<ul>\n<li>you need to post-process the predicted heatmap from the model and transform them to coordinates for submission.csv </li>\n</ul>\n<p>These two details are very important. To ensure the framework is correct, you should</p>\n<ul>\n<li>verify the end-to-end framework work on the <strong>train set.</strong> this gives the upper bound of your post processing (because when using the train set, you factor out problems due to generalisation error between train/validation set)</li>\n<li>if you set <strong>predicted heatmap  = your target ground truth</strong> (for backpropagation), you MUST get a perfect score of near <strong>1.00 in f-score</strong>.</li>\n</ul>\n<p>let me illustrate a few examples. In the extreme case, the 32x32x32 scanning window is not a good choice. there are border artifacts. For resnet34, it has 4 scales and its context is larger than 32. One should aim for the largest scanning window (in appropriate dimension) whenever possible.</p>\n<hr>\n<h2>2. Select good model:</h2>\n<p>You have choices of 2d, 2.5d, 3d encoder. you also have choices of 2d, 3d/ seqenuce decoder. It is important to understand their context. e.g. if you use 2.5d unet then w,h&gt;&gt;d in the scan window (e.g. 630x630x16). if you use 3d decoder, then the depth d of your window should be large (e.g. 256x256x160). you should at least try various window size on:</p>\n<ol>\n<li>pure 2d unet</li>\n<li>pure 3d unet</li>\n<li>hybrid of 2d/3d</li>\n</ol>\n<p>why try 2d? this is because the dataset document mentions that some annotations are done in 2d. \"Try\" means proper experiment for local validation and compare correlation for public lb score.</p>\n<p>never assume that 3d is always better than 2d (and vice versus). let the experiment results speak for themselves.</p>\n<hr>\n<h2>3. Dataset processing.</h2>\n<p>First, you need to normalise value to (0,1) or(-1,1). i think the experiment data all come from the <strong>same</strong> source and hence their intensity values are very smiliar. you have the choice of normalise by min/max, by std, by percentile. you can normalise by slice, subvolume,full volume  or share normalising parameters for all volumes.<br>\nplease think about this.</p>\n<p>please read the dataset paper pdf. it is very important.</p>\n<hr>\n<p>how many volumes should I use for validation, one or two or three?  how to stop training early?</p>\n<p>try to use all denosied, wbp, isonet, etc … for training.</p>\n<p>try standard augmentation for now but better augmentation does give better results. </p>\n<p>how about extra synthetic dataset? you can consider that later because my lb score of 0.700 is without synthetic yet (although I will use them later and expect some boost of 5%)</p>",
  "messages": [
    {
      "id": 3051236,
      "postDate": "2024-11-21T04:12:21.130Z",
      "content": "<p>a comfortable lb score without many bells and whistles is about 0.630. If you are getting less than benchmark.csv score of 0.587, the below may help</p>\n<h2>1. design a workable end-to-end framework. Most important details are:</h2>\n<h3>scanning window:</h3>\n<ul>\n<li>it is not possible to process 184x630x630 directly and you are likely to break it into smaller scanning \"windows\" and feed them into your network, and then fuse the predictions together into one later.</li>\n</ul>\n<h3>coordinates extraction</h3>\n<ul>\n<li>you need to post-process the predicted heatmap from the model and transform them to coordinates for submission.csv </li>\n</ul>\n<p>These two details are very important. To ensure the framework is correct, you should</p>\n<ul>\n<li>verify the end-to-end framework work on the <strong>train set.</strong> this gives the upper bound of your post processing (because when using the train set, you factor out problems due to generalisation error between train/validation set)</li>\n<li>if you set <strong>predicted heatmap  = your target ground truth</strong> (for backpropagation), you MUST get a perfect score of near <strong>1.00 in f-score</strong>.</li>\n</ul>\n<p>let me illustrate a few examples. In the extreme case, the 32x32x32 scanning window is not a good choice. there are border artifacts. For resnet34, it has 4 scales and its context is larger than 32. One should aim for the largest scanning window (in appropriate dimension) whenever possible.</p>\n<hr>\n<h2>2. Select good model:</h2>\n<p>You have choices of 2d, 2.5d, 3d encoder. you also have choices of 2d, 3d/ seqenuce decoder. It is important to understand their context. e.g. if you use 2.5d unet then w,h&gt;&gt;d in the scan window (e.g. 630x630x16). if you use 3d decoder, then the depth d of your window should be large (e.g. 256x256x160). you should at least try various window size on:</p>\n<ol>\n<li>pure 2d unet</li>\n<li>pure 3d unet</li>\n<li>hybrid of 2d/3d</li>\n</ol>\n<p>why try 2d? this is because the dataset document mentions that some annotations are done in 2d. \"Try\" means proper experiment for local validation and compare correlation for public lb score.</p>\n<p>never assume that 3d is always better than 2d (and vice versus). let the experiment results speak for themselves.</p>\n<hr>\n<h2>3. Dataset processing.</h2>\n<p>First, you need to normalise value to (0,1) or(-1,1). i think the experiment data all come from the <strong>same</strong> source and hence their intensity values are very smiliar. you have the choice of normalise by min/max, by std, by percentile. you can normalise by slice, subvolume,full volume  or share normalising parameters for all volumes.<br>\nplease think about this.</p>\n<p>please read the dataset paper pdf. it is very important.</p>\n<hr>\n<p>how many volumes should I use for validation, one or two or three?  how to stop training early?</p>\n<p>try to use all denosied, wbp, isonet, etc … for training.</p>\n<p>try standard augmentation for now but better augmentation does give better results. </p>\n<p>how about extra synthetic dataset? you can consider that later because my lb score of 0.700 is without synthetic yet (although I will use them later and expect some boost of 5%)</p>",
      "rawMarkdown": "a comfortable lb score without many bells and whistles is about 0.630. If you are getting less than benchmark.csv score of 0.587, the below may help\n\n##1. design a workable end-to-end framework. Most important details are:\n### scanning window:\n  - it is not possible to process 184x630x630 directly and you are likely to break it into smaller scanning \"windows\" and feed them into your network, and then fuse the predictions together into one later.\n### coordinates extraction\n  - you need to post-process the predicted heatmap from the model and transform them to coordinates for submission.csv \n\nThese two details are very important. To ensure the framework is correct, you should\n- verify the end-to-end framework work on the **train set.** this gives the upper bound of your post processing (because when using the train set, you factor out problems due to generalisation error between train/validation set)\n- if you set **predicted heatmap  = your target ground truth** (for backpropagation), you MUST get a perfect score of near **1.00 in f-score**.\n\nlet me illustrate a few examples. In the extreme case, the 32x32x32 scanning window is not a good choice. there are border artifacts. For resnet34, it has 4 scales and its context is larger than 32. One should aim for the largest scanning window (in appropriate dimension) whenever possible.\n\n\n\n----\n\n\n##2. Select good model:\n\nYou have choices of 2d, 2.5d, 3d encoder. you also have choices of 2d, 3d/ seqenuce decoder. It is important to understand their context. e.g. if you use 2.5d unet then w,h>>d in the scan window (e.g. 630x630x16). if you use 3d decoder, then the depth d of your window should be large (e.g. 256x256x160). you should at least try various window size on:\n1. pure 2d unet\n2. pure 3d unet\n3. hybrid of 2d/3d\n\nwhy try 2d? this is because the dataset document mentions that some annotations are done in 2d. \"Try\" means proper experiment for local validation and compare correlation for public lb score.\n\nnever assume that 3d is always better than 2d (and vice versus). let the experiment results speak for themselves.\n\n\n\n\n----\n\n##3. Dataset processing.\nFirst, you need to normalise value to (0,1) or(-1,1). i think the experiment data all come from the **same** source and hence their intensity values are very smiliar. you have the choice of normalise by min/max, by std, by percentile. you can normalise by slice, subvolume,full volume  or share normalising parameters for all volumes.\nplease think about this.\n\nplease read the dataset paper pdf. it is very important.\n\n---\n\nhow many volumes should I use for validation, one or two or three?  how to stop training early?\n\ntry to use all denosied, wbp, isonet, etc ... for training.\n\ntry standard augmentation for now but better augmentation does give better results. \n\nhow about extra synthetic dataset? you can consider that later because my lb score of 0.700 is without synthetic yet (although I will use them later and expect some boost of 5%)",
      "votes": 97
    },
    {
      "id": 3053819,
      "postDate": "2024-11-24T00:34:37.330Z",
      "content": "<p>yes, it is possible to get zero for training (i.e. model not learning anything or just predicting all pixels as negative)</p>\n<h2>Training process</h2>\n<ul>\n<li>before you start training, compute the positive-to-negative ratio of the labeled pixels.</li>\n<li>you already know that it is likely to cause an imbalance problem.</li>\n<li>ideally, we want the segmentation targets to be disjoint (easier for post-processing later). But we want it to be large for a better positive-to-negative ratio to drive the back-propagation</li>\n<li>Please think about it and need to search for the best size</li>\n<li>alternatively, you can have 2 targets, one for post-processing, another for driving the back-propagation </li>\n</ul>\n<hr>\n<ul>\n<li>but there is another more serious problem. the target signal is quite weak.</li>\n<li>if you are good at optimizer, you need to warm them up correctly (adam variants) and reset them to prevent them from getting stuck at the all-zero (and other) local minimum. Or you can use plain SGD which may be easier to control manually.</li>\n<li>in my experiment, there is more than one minimum. <ul>\n<li>if I leave train/validation data and model unchanged and just change the behavior of the optimizer, you can end up minimum (i.e. where training peaks or stabilizes) at cv=0, 0.5, 0.65, 0.75, 0.85</li></ul></li>\n</ul>",
      "rawMarkdown": "yes, it is possible to get zero for training (i.e. model not learning anything or just predicting all pixels as negative)\n\n## Training process\n- before you start training, compute the positive-to-negative ratio of the labeled pixels.\n- you already know that it is likely to cause an imbalance problem.\n- ideally, we want the segmentation targets to be disjoint (easier for post-processing later). But we want it to be large for a better positive-to-negative ratio to drive the back-propagation\n- Please think about it and need to search for the best size\n- alternatively, you can have 2 targets, one for post-processing, another for driving the back-propagation \n\n---\n\n- but there is another more serious problem. the target signal is quite weak.\n- if you are good at optimizer, you need to warm them up correctly (adam variants) and reset them to prevent them from getting stuck at the all-zero (and other) local minimum. Or you can use plain SGD which may be easier to control manually.\n- in my experiment, there is more than one minimum. \n   - if I leave train/validation data and model unchanged and just change the behavior of the optimizer, you can end up minimum (i.e. where training peaks or stabilizes) at cv=0, 0.5, 0.65, 0.75, 0.85\n\n",
      "votes": 8,
      "replies": [
        {
          "id": 3054236,
          "postDate": "2024-11-24T13:07:12.563Z",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , this is very helpful :). Playing with the optimizer is something I don't have much experience with (aside from changing learning rate) so I am looking forward to digging a little deeper here.</p>\n<p>I think I understand what you mean by warmup, but I don't quite follow the \"reset to zero\" part. Could you elaborate a little more?</p>",
          "rawMarkdown": "Hi @hengck23 , this is very helpful :). Playing with the optimizer is something I don't have much experience with (aside from changing learning rate) so I am looking forward to digging a little deeper here.\n\nI think I understand what you mean by warmup, but I don't quite follow the \"reset to zero\" part. Could you elaborate a little more?",
          "replies": [
            {
              "id": 3054925,
              "postDate": "2024-11-25T10:42:04.160Z",
              "content": "<p>try asking chatgpt: it is said resetting optimizer in training improves vlidation loss. can you explain. please show some code. can u search the web and show some papers?</p>\n<p>come back to here if u don't good results.</p>",
              "rawMarkdown": "try asking chatgpt: it is said resetting optimizer in training improves vlidation loss. can you explain. please show some code. can u search the web and show some papers?\n\ncome back to here if u don't good results.",
              "votes": 7
            },
            {
              "id": 3055000,
              "postDate": "2024-11-25T12:07:17.097Z",
              "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> :) I'm trying some experiments using cosine annealing with warm restarts. These experiments are still running. It looks very hyperparameter sensitive so I'm running a pretty thorough search.</p>\n<p>If I understand correctly what you're suggesting here is a bit different than warm restarts though. It kind sounds like you're suggesting we do cold restarts with a completely fresh optimizer! This is new to me, I'll do some experiments to see how this goes :). ChatGPT was helpful here as you suspected.</p>",
              "rawMarkdown": "Thank you @hengck23 :) I'm trying some experiments using cosine annealing with warm restarts. These experiments are still running. It looks very hyperparameter sensitive so I'm running a pretty thorough search.\n\nIf I understand correctly what you're suggesting here is a bit different than warm restarts though. It kind sounds like you're suggesting we do cold restarts with a completely fresh optimizer! This is new to me, I'll do some experiments to see how this goes :). ChatGPT was helpful here as you suspected."
            }
          ]
        },
        {
          "id": 3054321,
          "postDate": "2024-11-24T15:06:17.517Z",
          "content": "<p>Looks like the likelihood of getting stuck in a local minimum is high. Mutations? Hmm</p>",
          "rawMarkdown": "Looks like the likelihood of getting stuck in a local minimum is high. Mutations? Hmm"
        }
      ]
    },
    {
      "id": 3055516,
      "postDate": "2024-11-25T20:44:54.893Z",
      "content": "<blockquote>\n  <p>if you set predicted heatmap = your target ground truth (for backpropagation), you MUST get a perfect score of near 1.00 in f-score.</p>\n</blockquote>\n<p>Do you currently achieve this? </p>\n<p>My current approach is a bit simplistic and can likely be improved, but I am at: </p>\n<pre><code> for TS_73_6: .\n for TS_69_2: .\n for TS_6_4: .\n for TS_6_6: .\n for TS_86_3: .\n for TS_99_9: .\n for TS_5_4: .\n Score: .\n</code></pre>",
      "rawMarkdown": "> if you set predicted heatmap = your target ground truth (for backpropagation), you MUST get a perfect score of near 1.00 in f-score.\n\nDo you currently achieve this? \n\nMy current approach is a bit simplistic and can likely be improved, but I am at: \n```\nScore for TS_73_6: 0.9792993955946692\nScore for TS_69_2: 0.998953427524856\nScore for TS_6_4: 1.0\nScore for TS_6_6: 0.9980008567756676\nScore for TS_86_3: 0.9882783882783883\nScore for TS_99_9: 0.9651746583029144\nScore for TS_5_4: 1.0\nTotal Score: 0.9872025639815435\n```\n\n\n",
      "votes": 1,
      "replies": [
        {
          "id": 3056534,
          "postDate": "2024-11-27T04:50:15.567Z",
          "content": "<p>you need to get 1.00 for simple methods.<br>\nconnected component labeling (CCL) assumes that objects are dis-connected.<br>\nyou need to reduce the size of your ground truth targets so that they are not connected.<br>\nyour objective is to detect coordinates (centroid) and not to segment the whole object.</p>\n<p>if your target is connected, then you should not have used CCL. Other methods like watershed and distance transform may work better.</p>\n<p>if only a small part of your targets are connected. then you need to \"classify\" your segmentation results as being single or multiple particles, and then apply corresponding methods to separate touching particles. This is not difficult as we roughly know the particle size (radius) but troublesome.</p>\n<hr>\n<p>my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0</p>",
          "rawMarkdown": "you need to get 1.00 for simple methods.\nconnected component labeling (CCL) assumes that objects are dis-connected.\nyou need to reduce the size of your ground truth targets so that they are not connected.\nyour objective is to detect coordinates (centroid) and not to segment the whole object.\n\nif your target is connected, then you should not have used CCL. Other methods like watershed and distance transform may work better.\n\nif only a small part of your targets are connected. then you need to \"classify\" your segmentation results as being single or multiple particles, and then apply corresponding methods to separate touching particles. This is not difficult as we roughly know the particle size (radius) but troublesome.\n\n---\n\nmy current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0",
          "votes": 5,
          "replies": [
            {
              "id": 3056648,
              "postDate": "2024-11-27T08:27:08.163Z",
              "content": "<p>Hi!</p>\n<p>Thanks yeah I am not using CCL, more of an erosion and nonmax-supression. </p>\n<blockquote>\n  <p>my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0</p>\n</blockquote>\n<p>Thanks! I'll try and optimize further. I was just wondering whether 1.0 across the board was possible. </p>",
              "rawMarkdown": "Hi!\n\nThanks yeah I am not using CCL, more of an erosion and nonmax-supression. \n\n> my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0\n\nThanks! I'll try and optimize further. I was just wondering whether 1.0 across the board was possible. \n"
            },
            {
              "id": 3057615,
              "postDate": "2024-11-28T11:20:48.117Z",
              "content": "<blockquote>\n  <p>my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0</p>\n</blockquote>\n<p>When you say uses disconnected segmentation targets, how was this achieved? From my limited understanding, the data itself influences whether you can perform disconnected/connected. If cell centres are really close to each other then it is not feasible to perform disconnected on the raw data alone.</p>\n<p>Please let me know if I am misunderstanding anything.</p>",
              "rawMarkdown": ">my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0\n\nWhen you say uses disconnected segmentation targets, how was this achieved? From my limited understanding, the data itself influences whether you can perform disconnected/connected. If cell centres are really close to each other then it is not feasible to perform disconnected on the raw data alone.\n\nPlease let me know if I am misunderstanding anything.",
              "votes": 1
            },
            {
              "id": 3057632,
              "postDate": "2024-11-28T11:44:41.393Z",
              "content": "<p>please read the dataset pdf and plot the ground truth coordinates for visualisation.</p>",
              "rawMarkdown": "please read the dataset pdf and plot the ground truth coordinates for visualisation.",
              "votes": 1
            },
            {
              "id": 3058342,
              "postDate": "2024-11-29T11:07:19.020Z",
              "content": "<p>*.5 or factor small enough.</p>",
              "rawMarkdown": "*.5 or factor small enough.",
              "votes": 1
            },
            {
              "id": 3064271,
              "postDate": "2024-12-05T12:51:54.920Z",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> when you talk about the segmentation targets being disjoint, do you mean for each individual particle (so all apo-ferritins within tomo 0) or overall for the particles within that tomo? (so all particles being disjoint from each other within tomo 0), say we have a multi-class segmentation mask, if we sum along the class dimension, no overlap? thanks!!</p>",
              "rawMarkdown": "@hengck23 when you talk about the segmentation targets being disjoint, do you mean for each individual particle (so all apo-ferritins within tomo 0) or overall for the particles within that tomo? (so all particles being disjoint from each other within tomo 0), say we have a multi-class segmentation mask, if we sum along the class dimension, no overlap? thanks!!"
            },
            {
              "id": 3064278,
              "postDate": "2024-12-05T12:57:41.300Z",
              "content": "<p>for each class, particles do not overlap.  <br>\nThis can be verified plotting out centroid values in 3d xyz</p>",
              "rawMarkdown": "for each class, particles do not overlap.  \nThis can be verified plotting out centroid values in 3d xyz",
              "votes": 2
            }
          ]
        }
      ]
    },
    {
      "id": 3053548,
      "postDate": "2024-11-23T16:04:44Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Having particle being splitted in the cropped 3D images I guess is not ideal. Should I enforce boundary conditions when cropping? I am currently using Monoai \"RandCropByLabelClassesd\" for cropping and my results are not great scoring under the benchmark with aprox 0.4. I have checked the cropped images and they always crop labels on the edges. </p>",
      "rawMarkdown": "@hengck23  Having particle being splitted in the cropped 3D images I guess is not ideal. Should I enforce boundary conditions when cropping? I am currently using Monoai \"RandCropByLabelClassesd\" for cropping and my results are not great scoring under the benchmark with aprox 0.4. I have checked the cropped images and they always crop labels on the edges. \n",
      "votes": 1,
      "replies": [
        {
          "id": 3054936,
          "postDate": "2024-11-25T10:51:46.280Z",
          "content": "<p>let us temporaily forget about the competition</p>\n<p>consider experiment with only 2d images. train 2d unet on whole images and compare unet trained on tiles.<br>\ntrain different models with different settings, eg handling boundary or not, remove partial label at boundary or not, etc</p>\n<p>infer and validate.  what if infer conditions is not the same as training?<br>\nhow is inference behaviour different on train, compared to validation images?</p>\n<p>in datascience, trust only your data and experiment results</p>",
          "rawMarkdown": "let us temporaily forget about the competition\n\nconsider experiment with only 2d images. train 2d unet on whole images and compare unet trained on tiles.\ntrain different models with different settings, eg handling boundary or not, remove partial label at boundary or not, etc\n\ninfer and validate.  what if infer conditions is not the same as training?\nhow is inference behaviour different on train, compared to validation images?\n\nin datascience, trust only your data and experiment results\n\n",
          "votes": 4
        }
      ]
    },
    {
      "id": 3055316,
      "postDate": "2024-11-25T17:25:25.633Z",
      "content": "<p>蛙哥蛙哥我偶像！每次阅读都是享受。</p>",
      "rawMarkdown": "蛙哥蛙哥我偶像！每次阅读都是享受。",
      "votes": -3
    },
    {
      "id": 3082686,
      "postDate": "2024-12-28T13:36:36.327Z",
      "content": "<pre><code>One aim for the largest window (in appropriate whenever possible.\n</code></pre>\n<p>Considering batch_size = 1</p>\n<p>What is the shape of best scanning window that fits in T4 GPU?</p>\n<p><a href=\"https://www.kaggle.com/code/kharrington/deepfindet-train\" target=\"_blank\">https://www.kaggle.com/code/kharrington/deepfindet-train</a> : The demo notebook used (72, 72, 72)<br>\nBut I think we can do better.</p>\n<p>So I tried (64, 630, 630), (23, 630, 630), (23, 600, 600).</p>\n<p>When I tried (23, 600, 600)</p>\n<pre><code>: CUDA out of memory. Tried to allocate . GiB. GPU  has a total capacity of . GiB of which . MiB is free. Process  has . GiB memory in use. Of the allocated memory . GiB is allocated by PyTorch, and . MiB is reserved by PyTorch but unallocated.\n</code></pre>\n<p>What largest scanning window worked for you?</p>",
      "rawMarkdown": "```\nOne should aim for the largest scanning window (in appropriate dimension) whenever possible.\n```\n\nConsidering batch_size = 1\n\nWhat is the shape of best scanning window that fits in T4 GPU?\n\nhttps://www.kaggle.com/code/kharrington/deepfindet-train : The demo notebook used (72, 72, 72)\nBut I think we can do better.\n\nSo I tried (64, 630, 630), (23, 630, 630), (23, 600, 600).\n\nWhen I tried (23, 600, 600)\n```\nOutOfMemoryError: CUDA out of memory. Tried to allocate 1.97 GiB. GPU 0 has a total capacity of 15.89 GiB of which 985.12 MiB is free. Process 2802 has 14.92 GiB memory in use. Of the allocated memory 14.52 GiB is allocated by PyTorch, and 118.38 MiB is reserved by PyTorch but unallocated.\n```\n\nWhat largest scanning window worked for you?",
      "replies": [
        {
          "id": 3082731,
          "postDate": "2024-12-28T14:39:01.093Z",
          "content": "<p>There's more aspects to consider besides simply the size of one batch. You can obviously work with batches bigger than 72x72x72 (I work with 1 batch at a time too, around 40 times larger than in the demo notebook), but your model architecture will play an important role in the memory consumption for the tensors involved in a forward pass of the network. </p>\n<p>You also have to take into account the number of encoding blocks and the dimensions of each batch, as they all need to be divisible by the down sampling factor (23 in your example is prime, for example, so you'd require padding). There's multiple constraints that need to be juggled together to satisfy memory consumption and performance.</p>",
          "rawMarkdown": "There's more aspects to consider besides simply the size of one batch. You can obviously work with batches bigger than 72x72x72 (I work with 1 batch at a time too, around 40 times larger than in the demo notebook), but your model architecture will play an important role in the memory consumption for the tensors involved in a forward pass of the network. \n\nYou also have to take into account the number of encoding blocks and the dimensions of each batch, as they all need to be divisible by the down sampling factor (23 in your example is prime, for example, so you'd require padding). There's multiple constraints that need to be juggled together to satisfy memory consumption and performance.",
          "votes": 1
        }
      ]
    },
    {
      "id": 3069068,
      "postDate": "2024-12-11T02:54:19.227Z",
      "content": "<blockquote>\n  <p>please read the dataset paper pdf. it is very important.</p>\n</blockquote>\n<p>where would i find this</p>",
      "rawMarkdown": ">please read the dataset paper pdf. it is very important.\n\nwhere would i find this",
      "replies": [
        {
          "id": 3069143,
          "postDate": "2024-12-11T05:41:32.017Z",
          "content": "<p>Dataset paper <a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1\" target=\"_blank\">here</a>. Sourced from a post from  the host <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702\" target=\"_blank\">here</a>.</p>",
          "rawMarkdown": "Dataset paper [here](https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1). Sourced from a post from  the host [here](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702)."
        }
      ]
    },
    {
      "id": 3064141,
      "postDate": "2024-12-05T09:18:13.583Z",
      "content": "<p>Hello, I have another question. The competition data includes four types of scan images（denoised.zarr, ctfdeconvolved.zarr, isonetcorrected.zarr, wbp.zarr）, but only the denoised type is included when submitting the code. So, during model training, should we include the other three types in the training dataset as well? I'm not sure if this will improve the model's performance, but I think it might be worth trying.</p>",
      "rawMarkdown": "Hello, I have another question. The competition data includes four types of scan images（denoised.zarr, ctfdeconvolved.zarr, isonetcorrected.zarr, wbp.zarr）, but only the denoised type is included when submitting the code. So, during model training, should we include the other three types in the training dataset as well? I'm not sure if this will improve the model's performance, but I think it might be worth trying."
    },
    {
      "id": 3064137,
      "postDate": "2024-12-05T09:09:59.263Z",
      "content": "<p>Hello, I was wondering if including data with different scales and resolutions ((46, 158, 158), (92, 315, 315), (84, 630, 630)) in the model training process, or enabling the model to support inputs with varying resolutions, could potentially improve its performance?</p>\n<p>I believe that data with different resolutions contain different types of information. For example, low-resolution data (shape: 46, 158, 158) might capture more global information, while high-resolution data focuses on local details. Different resolutions may provide complementary information.</p>\n<p>I'm not sure if this idea is correct 🤣.</p>",
      "rawMarkdown": "Hello, I was wondering if including data with different scales and resolutions ((46, 158, 158), (92, 315, 315), (84, 630, 630)) in the model training process, or enabling the model to support inputs with varying resolutions, could potentially improve its performance?\n\nI believe that data with different resolutions contain different types of information. For example, low-resolution data (shape: 46, 158, 158) might capture more global information, while high-resolution data focuses on local details. Different resolutions may provide complementary information.\n\nI'm not sure if this idea is correct 🤣."
    },
    {
      "id": 3058299,
      "postDate": "2024-11-29T09:57:10.380Z",
      "content": "<p>Hi. First of all, thank for taking your time helping other people on the competition. Very useful info!<br>\nI'm familiar with conventional (2d) unets but never stepped into a problem that required dividing input data into windows, so i have some doubts. My goal asking this questions is not to skip experimentation and self training on the problem but to understand certain intuitions . My approach is to make a dataset that is able to split volumes given a certain 3d window shape and 3d stride ( if stride shape is lower than window shape there's some overlap in the scanning windows ), then my unet implementation its a 3dunet so it abstracts all possible cases (2d, 2.5d and 3d) (for example for 2d i just have to set the window shape to (1, width, height)). Finally my ground-truth is made by setting gaussian distributions (taking in account particle radius) into the windows to make heatmaps with shape (depth, width, height, num_particles). </p>\n<p>My question is how applying different overlaps and padding would affect heatmap estimation accuracy. Thank you in advance.    </p>",
      "rawMarkdown": "Hi. First of all, thank for taking your time helping other people on the competition. Very useful info!\nI'm familiar with conventional (2d) unets but never stepped into a problem that required dividing input data into windows, so i have some doubts. My goal asking this questions is not to skip experimentation and self training on the problem but to understand certain intuitions . My approach is to make a dataset that is able to split volumes given a certain 3d window shape and 3d stride ( if stride shape is lower than window shape there's some overlap in the scanning windows ), then my unet implementation its a 3dunet so it abstracts all possible cases (2d, 2.5d and 3d) (for example for 2d i just have to set the window shape to (1, width, height)). Finally my ground-truth is made by setting gaussian distributions (taking in account particle radius) into the windows to make heatmaps with shape (depth, width, height, num_particles). \n\nMy question is how applying different overlaps and padding would affect heatmap estimation accuracy. Thank you in advance.    ",
      "replies": [
        {
          "id": 3061799,
          "postDate": "2024-12-03T01:58:15.283Z",
          "content": "<p>my suggestion:<br>\nyou just need to do or imagine several experiments. you can start with 2d Unet becuase it is easy to implement.<br>\n1) train = full image (e.g. 630x630), validation= full image (e.g. 630x630)<br>\n2) train = full image (e.g. 630x630), validation= crop image (try crop=512,256,128,64,32)<br>\n3) train = crop image (try crop=512,256,128,64,32), validation= full image (e.g. 630x630)<br>\n4) train = crop image (try crop=512,256,128,64,32), validation= crop image (try crop=512,256,128,64,32)</p>\n<p>in addition, try to use train set as validation set, e.g. train=full image, infer at train at crop</p>\n<hr>\n<p>for each experiment, are interested in:</p>\n<ol>\n<li>train/validation log loss (back propagation)</li>\n<li>train/validation fbeta score metric (actually i would monitor all precision,recall, hit,miss,fp, num_predict)</li>\n<li>visualisation of results (the segmentation map, probability and binarised)</li>\n</ol>\n<p>after doing a few times, you would know what is the issue.<br>\nin summary:</p>\n<ul>\n<li>when a model decide if a pixel(x,y) is +ve or -neg, it looks at a context square around (x,y) for information.</li>\n<li>conisder x,y=(315,315) in full image image (630,630).</li>\n<li>now you make a crop so that the SAME pixel now become (0,0) in crop (128,128). the model still look at context square around (0,0) but fill in the mssing context pixels with zero.</li>\n<li>the input context information is different and obviously, the model prediction would be \"different\" for the \"SAME\" pixel</li>\n<li>so if we input an image SxS into the model, only the interior prediction is most relieable. prediction at the border are more prone to error.</li>\n<li>so if you need to divide a large image into crops, the interior of each crops is the most relieable. you should try to use the interior (instead of all) of crop to assemble back the results into the full image size</li>\n</ul>\n<hr>\n<p>on a side note, CNN has implicit postion encoding. when you present a pixel = x,y,v (x,y are 2d coords and v is pixel intensity value) to the network, he actually \"sees\" x,y in additional to v. Although in textbook, it is mentioned that convloution operation is translationat invariant, CNN network is actuall not. you can search papers on this.</p>\n<p>during training/inference, he not only see context around (x,y), he also see distance of (x,y) from image boundary to make decision</p>",
          "rawMarkdown": "my suggestion:\nyou just need to do or imagine several experiments. you can start with 2d Unet becuase it is easy to implement.\n1) train = full image (e.g. 630x630), validation= full image (e.g. 630x630)\n2) train = full image (e.g. 630x630), validation= crop image (try crop=512,256,128,64,32)\n3) train = crop image (try crop=512,256,128,64,32), validation= full image (e.g. 630x630)\n4) train = crop image (try crop=512,256,128,64,32), validation= crop image (try crop=512,256,128,64,32)\n\nin addition, try to use train set as validation set, e.g. train=full image, infer at train at crop\n\n---\n\nfor each experiment, are interested in:\n1. train/validation log loss (back propagation)\n2. train/validation fbeta score metric (actually i would monitor all precision,recall, hit,miss,fp, num_predict)\n3. visualisation of results (the segmentation map, probability and binarised)\n\nafter doing a few times, you would know what is the issue.\nin summary:\n- when a model decide if a pixel(x,y) is +ve or -neg, it looks at a context square around (x,y) for information.\n- conisder x,y=(315,315) in full image image (630,630).\n- now you make a crop so that the SAME pixel now become (0,0) in crop (128,128). the model still look at context square around (0,0) but fill in the mssing context pixels with zero.\n- the input context information is different and obviously, the model prediction would be \"different\" for the \"SAME\" pixel\n- so if we input an image SxS into the model, only the interior prediction is most relieable. prediction at the border are more prone to error.\n- so if you need to divide a large image into crops, the interior of each crops is the most relieable. you should try to use the interior (instead of all) of crop to assemble back the results into the full image size\n\n---\n\non a side note, CNN has implicit postion encoding. when you present a pixel = x,y,v (x,y are 2d coords and v is pixel intensity value) to the network, he actually \"sees\" x,y in additional to v. Although in textbook, it is mentioned that convloution operation is translationat invariant, CNN network is actuall not. you can search papers on this.\n\nduring training/inference, he not only see context around (x,y), he also see distance of (x,y) from image boundary to make decision",
          "votes": 6
        }
      ]
    },
    {
      "id": 3052527,
      "postDate": "2024-11-22T14:17:47.570Z",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for all these tips. Do you get similar values between the score in validation set vs LB? </p>",
      "rawMarkdown": "@hengck23 Thanks for all these tips. Do you get similar values between the score in validation set vs LB? ",
      "replies": [
        {
          "id": 3052649,
          "postDate": "2024-11-22T16:53:21.673Z",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221\" target=\"_blank\">Hengck's Results/Ideas</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> is one of the really helpful grandmasters around here. They usually start a running thread at the beginning of the competition that shows data, correlations with the LB, and ideas. Check it out!</p>",
          "rawMarkdown": "[Hengck's Results/Ideas](https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221) @hengck23 is one of the really helpful grandmasters around here. They usually start a running thread at the beginning of the competition that shows data, correlations with the LB, and ideas. Check it out!"
        }
      ]
    },
    {
      "id": 3051333,
      "postDate": "2024-11-21T07:37:22.687Z",
      "content": "<p>Thanks for all those tips! Do you manage to do inference with 3D models within the time limits? If not do you think model quantization could be worth trying?</p>",
      "rawMarkdown": "Thanks for all those tips! Do you manage to do inference with 3D models within the time limits? If not do you think model quantization could be worth trying?",
      "replies": [
        {
          "id": 3051420,
          "postDate": "2024-11-21T09:25:44.167Z",
          "content": "<p>a small 3d unet is sufficient. it can run easily within the notebook limit</p>",
          "rawMarkdown": "a small 3d unet is sufficient. it can run easily within the notebook limit",
          "votes": 1
        }
      ]
    },
    {
      "id": 3057289,
      "postDate": "2024-11-28T01:35:19.447Z",
      "rawMarkdown": "",
      "votes": -3,
      "isDeleted": true,
      "replies": [
        {
          "id": 3057771,
          "postDate": "2024-11-28T15:43:49.490Z",
          "content": "<p>Prompt= \"Respond to the above text in a thankful way\" </p>",
          "rawMarkdown": "Prompt= \"Respond to the above text in a thankful way\" ",
          "votes": 9,
          "replies": [
            {
              "id": 3057860,
              "postDate": "2024-11-28T18:10:49.040Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        }
      ]
    },
    {
      "id": 3055420,
      "postDate": "2024-11-25T19:00:09.357Z",
      "content": "<p>Thanks for these tips!!!</p>",
      "rawMarkdown": "Thanks for these tips!!!",
      "votes": 1
    },
    {
      "id": 3078337,
      "postDate": "2024-12-22T07:19:30.857Z",
      "content": "<p>Thank you for sharing your suggestions.</p>",
      "rawMarkdown": "Thank you for sharing your suggestions."
    },
    {
      "id": 3051403,
      "postDate": "2024-11-21T09:11:27.953Z",
      "content": "<p>Thank you for sharing your suggestions.</p>",
      "rawMarkdown": "Thank you for sharing your suggestions."
    }
  ],
  "comments": [
    {
      "id": 3053819,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-11-24T00:34:37.330000",
      "content": "<p>yes, it is possible to get zero for training (i.e. model not learning anything or just predicting all pixels as negative)</p>\n<h2>Training process</h2>\n<ul>\n<li>before you start training, compute the positive-to-negative ratio of the labeled pixels.</li>\n<li>you already know that it is likely to cause an imbalance problem.</li>\n<li>ideally, we want the segmentation targets to be disjoint (easier for post-processing later). But we want it to be large for a better positive-to-negative ratio to drive the back-propagation</li>\n<li>Please think about it and need to search for the best size</li>\n<li>alternatively, you can have 2 targets, one for post-processing, another for driving the back-propagation </li>\n</ul>\n<hr>\n<ul>\n<li>but there is another more serious problem. the target signal is quite weak.</li>\n<li>if you are good at optimizer, you need to warm them up correctly (adam variants) and reset them to prevent them from getting stuck at the all-zero (and other) local minimum. Or you can use plain SGD which may be easier to control manually.</li>\n<li>in my experiment, there is more than one minimum. <ul>\n<li>if I leave train/validation data and model unchanged and just change the behavior of the optimizer, you can end up minimum (i.e. where training peaks or stabilizes) at cv=0, 0.5, 0.65, 0.75, 0.85</li></ul></li>\n</ul>",
      "votes": 8,
      "replies": [
        {
          "id": 3054236,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-11-24T13:07:12.563000",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> , this is very helpful :). Playing with the optimizer is something I don't have much experience with (aside from changing learning rate) so I am looking forward to digging a little deeper here.</p>\n<p>I think I understand what you mean by warmup, but I don't quite follow the \"reset to zero\" part. Could you elaborate a little more?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3054925,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-25T10:42:04.160000",
              "content": "<p>try asking chatgpt: it is said resetting optimizer in training improves vlidation loss. can you explain. please show some code. can u search the web and show some papers?</p>\n<p>come back to here if u don't good results.</p>",
              "votes": 7,
              "replies": []
            },
            {
              "id": 3055000,
              "author_name": "chemdatafarmer",
              "author_url": "",
              "post_date": "2024-11-25T12:07:17.097000",
              "content": "<p>Thank you <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> :) I'm trying some experiments using cosine annealing with warm restarts. These experiments are still running. It looks very hyperparameter sensitive so I'm running a pretty thorough search.</p>\n<p>If I understand correctly what you're suggesting here is a bit different than warm restarts though. It kind sounds like you're suggesting we do cold restarts with a completely fresh optimizer! This is new to me, I'll do some experiments to see how this goes :). ChatGPT was helpful here as you suspected.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 3054321,
          "author_name": "Paula Livingstone",
          "author_url": "",
          "post_date": "2024-11-24T15:06:17.517000",
          "content": "<p>Looks like the likelihood of getting stuck in a local minimum is high. Mutations? Hmm</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3055516,
      "author_name": "fnands",
      "author_url": "",
      "post_date": "2024-11-25T20:44:54.893000",
      "content": "<blockquote>\n  <p>if you set predicted heatmap = your target ground truth (for backpropagation), you MUST get a perfect score of near 1.00 in f-score.</p>\n</blockquote>\n<p>Do you currently achieve this? </p>\n<p>My current approach is a bit simplistic and can likely be improved, but I am at: </p>\n<pre><code> for TS_73_6: .\n for TS_69_2: .\n for TS_6_4: .\n for TS_6_6: .\n for TS_86_3: .\n for TS_99_9: .\n for TS_5_4: .\n Score: .\n</code></pre>",
      "votes": 1,
      "replies": [
        {
          "id": 3056534,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-27T04:50:15.567000",
          "content": "<p>you need to get 1.00 for simple methods.<br>\nconnected component labeling (CCL) assumes that objects are dis-connected.<br>\nyou need to reduce the size of your ground truth targets so that they are not connected.<br>\nyour objective is to detect coordinates (centroid) and not to segment the whole object.</p>\n<p>if your target is connected, then you should not have used CCL. Other methods like watershed and distance transform may work better.</p>\n<p>if only a small part of your targets are connected. then you need to \"classify\" your segmentation results as being single or multiple particles, and then apply corresponding methods to separate touching particles. This is not difficult as we roughly know the particle size (radius) but troublesome.</p>\n<hr>\n<p>my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0</p>",
          "votes": 5,
          "replies": [
            {
              "id": 3056648,
              "author_name": "fnands",
              "author_url": "",
              "post_date": "2024-11-27T08:27:08.163000",
              "content": "<p>Hi!</p>\n<p>Thanks yeah I am not using CCL, more of an erosion and nonmax-supression. </p>\n<blockquote>\n  <p>my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0</p>\n</blockquote>\n<p>Thanks! I'll try and optimize further. I was just wondering whether 1.0 across the board was possible. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3057615,
              "author_name": "homiecal",
              "author_url": "",
              "post_date": "2024-11-28T11:20:48.117000",
              "content": "<blockquote>\n  <p>my current solution uses disconnected segmentation targets. and I can get perfect fbeta score =1.0</p>\n</blockquote>\n<p>When you say uses disconnected segmentation targets, how was this achieved? From my limited understanding, the data itself influences whether you can perform disconnected/connected. If cell centres are really close to each other then it is not feasible to perform disconnected on the raw data alone.</p>\n<p>Please let me know if I am misunderstanding anything.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3057632,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-11-28T11:44:41.393000",
              "content": "<p>please read the dataset pdf and plot the ground truth coordinates for visualisation.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3058342,
              "author_name": "Ángel Jacinto Sánchez Ruiz",
              "author_url": "",
              "post_date": "2024-11-29T11:07:19.020000",
              "content": "<p>*.5 or factor small enough.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3064271,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2024-12-05T12:51:54.920000",
              "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> when you talk about the segmentation targets being disjoint, do you mean for each individual particle (so all apo-ferritins within tomo 0) or overall for the particles within that tomo? (so all particles being disjoint from each other within tomo 0), say we have a multi-class segmentation mask, if we sum along the class dimension, no overlap? thanks!!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3064278,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "2024-12-05T12:57:41.300000",
              "content": "<p>for each class, particles do not overlap.  <br>\nThis can be verified plotting out centroid values in 3d xyz</p>",
              "votes": 2,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3053548,
      "author_name": "Ignasi Alemany",
      "author_url": "",
      "post_date": "2024-11-23T16:04:44",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>  Having particle being splitted in the cropped 3D images I guess is not ideal. Should I enforce boundary conditions when cropping? I am currently using Monoai \"RandCropByLabelClassesd\" for cropping and my results are not great scoring under the benchmark with aprox 0.4. I have checked the cropped images and they always crop labels on the edges. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 3054936,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-25T10:51:46.280000",
          "content": "<p>let us temporaily forget about the competition</p>\n<p>consider experiment with only 2d images. train 2d unet on whole images and compare unet trained on tiles.<br>\ntrain different models with different settings, eg handling boundary or not, remove partial label at boundary or not, etc</p>\n<p>infer and validate.  what if infer conditions is not the same as training?<br>\nhow is inference behaviour different on train, compared to validation images?</p>\n<p>in datascience, trust only your data and experiment results</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3055316,
      "author_name": "FakeOrange",
      "author_url": "",
      "post_date": "2024-11-25T17:25:25.633000",
      "content": "<p>蛙哥蛙哥我偶像！每次阅读都是享受。</p>",
      "votes": -3,
      "replies": []
    },
    {
      "id": 3082686,
      "author_name": "Keesari Vigneshwar Reddy",
      "author_url": "",
      "post_date": "2024-12-28T13:36:36.327000",
      "content": "<pre><code>One aim for the largest window (in appropriate whenever possible.\n</code></pre>\n<p>Considering batch_size = 1</p>\n<p>What is the shape of best scanning window that fits in T4 GPU?</p>\n<p><a href=\"https://www.kaggle.com/code/kharrington/deepfindet-train\" target=\"_blank\">https://www.kaggle.com/code/kharrington/deepfindet-train</a> : The demo notebook used (72, 72, 72)<br>\nBut I think we can do better.</p>\n<p>So I tried (64, 630, 630), (23, 630, 630), (23, 600, 600).</p>\n<p>When I tried (23, 600, 600)</p>\n<pre><code>: CUDA out of memory. Tried to allocate . GiB. GPU  has a total capacity of . GiB of which . MiB is free. Process  has . GiB memory in use. Of the allocated memory . GiB is allocated by PyTorch, and . MiB is reserved by PyTorch but unallocated.\n</code></pre>\n<p>What largest scanning window worked for you?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3082731,
          "author_name": "Andrei Zamfir",
          "author_url": "",
          "post_date": "2024-12-28T14:39:01.093000",
          "content": "<p>There's more aspects to consider besides simply the size of one batch. You can obviously work with batches bigger than 72x72x72 (I work with 1 batch at a time too, around 40 times larger than in the demo notebook), but your model architecture will play an important role in the memory consumption for the tensors involved in a forward pass of the network. </p>\n<p>You also have to take into account the number of encoding blocks and the dimensions of each batch, as they all need to be divisible by the down sampling factor (23 in your example is prime, for example, so you'd require padding). There's multiple constraints that need to be juggled together to satisfy memory consumption and performance.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3069068,
      "author_name": "NathanTuttle",
      "author_url": "",
      "post_date": "2024-12-11T02:54:19.227000",
      "content": "<blockquote>\n  <p>please read the dataset paper pdf. it is very important.</p>\n</blockquote>\n<p>where would i find this</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3069143,
          "author_name": "Bartley",
          "author_url": "",
          "post_date": "2024-12-11T05:41:32.017000",
          "content": "<p>Dataset paper <a href=\"https://www.biorxiv.org/content/10.1101/2024.11.04.621686v1\" target=\"_blank\">here</a>. Sourced from a post from  the host <a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/544702\" target=\"_blank\">here</a>.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3064141,
      "author_name": "Switch9527",
      "author_url": "",
      "post_date": "2024-12-05T09:18:13.583000",
      "content": "<p>Hello, I have another question. The competition data includes four types of scan images（denoised.zarr, ctfdeconvolved.zarr, isonetcorrected.zarr, wbp.zarr）, but only the denoised type is included when submitting the code. So, during model training, should we include the other three types in the training dataset as well? I'm not sure if this will improve the model's performance, but I think it might be worth trying.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3064137,
      "author_name": "Switch9527",
      "author_url": "",
      "post_date": "2024-12-05T09:09:59.263000",
      "content": "<p>Hello, I was wondering if including data with different scales and resolutions ((46, 158, 158), (92, 315, 315), (84, 630, 630)) in the model training process, or enabling the model to support inputs with varying resolutions, could potentially improve its performance?</p>\n<p>I believe that data with different resolutions contain different types of information. For example, low-resolution data (shape: 46, 158, 158) might capture more global information, while high-resolution data focuses on local details. Different resolutions may provide complementary information.</p>\n<p>I'm not sure if this idea is correct 🤣.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3058299,
      "author_name": "alvaro blom-dahl oliver",
      "author_url": "",
      "post_date": "2024-11-29T09:57:10.380000",
      "content": "<p>Hi. First of all, thank for taking your time helping other people on the competition. Very useful info!<br>\nI'm familiar with conventional (2d) unets but never stepped into a problem that required dividing input data into windows, so i have some doubts. My goal asking this questions is not to skip experimentation and self training on the problem but to understand certain intuitions . My approach is to make a dataset that is able to split volumes given a certain 3d window shape and 3d stride ( if stride shape is lower than window shape there's some overlap in the scanning windows ), then my unet implementation its a 3dunet so it abstracts all possible cases (2d, 2.5d and 3d) (for example for 2d i just have to set the window shape to (1, width, height)). Finally my ground-truth is made by setting gaussian distributions (taking in account particle radius) into the windows to make heatmaps with shape (depth, width, height, num_particles). </p>\n<p>My question is how applying different overlaps and padding would affect heatmap estimation accuracy. Thank you in advance.    </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3061799,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-12-03T01:58:15.283000",
          "content": "<p>my suggestion:<br>\nyou just need to do or imagine several experiments. you can start with 2d Unet becuase it is easy to implement.<br>\n1) train = full image (e.g. 630x630), validation= full image (e.g. 630x630)<br>\n2) train = full image (e.g. 630x630), validation= crop image (try crop=512,256,128,64,32)<br>\n3) train = crop image (try crop=512,256,128,64,32), validation= full image (e.g. 630x630)<br>\n4) train = crop image (try crop=512,256,128,64,32), validation= crop image (try crop=512,256,128,64,32)</p>\n<p>in addition, try to use train set as validation set, e.g. train=full image, infer at train at crop</p>\n<hr>\n<p>for each experiment, are interested in:</p>\n<ol>\n<li>train/validation log loss (back propagation)</li>\n<li>train/validation fbeta score metric (actually i would monitor all precision,recall, hit,miss,fp, num_predict)</li>\n<li>visualisation of results (the segmentation map, probability and binarised)</li>\n</ol>\n<p>after doing a few times, you would know what is the issue.<br>\nin summary:</p>\n<ul>\n<li>when a model decide if a pixel(x,y) is +ve or -neg, it looks at a context square around (x,y) for information.</li>\n<li>conisder x,y=(315,315) in full image image (630,630).</li>\n<li>now you make a crop so that the SAME pixel now become (0,0) in crop (128,128). the model still look at context square around (0,0) but fill in the mssing context pixels with zero.</li>\n<li>the input context information is different and obviously, the model prediction would be \"different\" for the \"SAME\" pixel</li>\n<li>so if we input an image SxS into the model, only the interior prediction is most relieable. prediction at the border are more prone to error.</li>\n<li>so if you need to divide a large image into crops, the interior of each crops is the most relieable. you should try to use the interior (instead of all) of crop to assemble back the results into the full image size</li>\n</ul>\n<hr>\n<p>on a side note, CNN has implicit postion encoding. when you present a pixel = x,y,v (x,y are 2d coords and v is pixel intensity value) to the network, he actually \"sees\" x,y in additional to v. Although in textbook, it is mentioned that convloution operation is translationat invariant, CNN network is actuall not. you can search papers on this.</p>\n<p>during training/inference, he not only see context around (x,y), he also see distance of (x,y) from image boundary to make decision</p>",
          "votes": 6,
          "replies": []
        }
      ]
    },
    {
      "id": 3052527,
      "author_name": "Ignasi Alemany",
      "author_url": "",
      "post_date": "2024-11-22T14:17:47.570000",
      "content": "<p><a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> Thanks for all these tips. Do you get similar values between the score in validation set vs LB? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 3052649,
          "author_name": "chemdatafarmer",
          "author_url": "",
          "post_date": "2024-11-22T16:53:21.673000",
          "content": "<p><a href=\"https://www.kaggle.com/competitions/czii-cryo-et-object-identification/discussion/545221\" target=\"_blank\">Hengck's Results/Ideas</a> <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> is one of the really helpful grandmasters around here. They usually start a running thread at the beginning of the competition that shows data, correlations with the LB, and ideas. Check it out!</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3051333,
      "author_name": "Konstantina",
      "author_url": "",
      "post_date": "2024-11-21T07:37:22.687000",
      "content": "<p>Thanks for all those tips! Do you manage to do inference with 3D models within the time limits? If not do you think model quantization could be worth trying?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3051420,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2024-11-21T09:25:44.167000",
          "content": "<p>a small 3d unet is sufficient. it can run easily within the notebook limit</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 3057289,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-11-28T01:35:19.447000",
      "content": "",
      "votes": -3,
      "replies": [
        {
          "id": 3057771,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2024-11-28T15:43:49.490000",
          "content": "<p>Prompt= \"Respond to the above text in a thankful way\" </p>",
          "votes": 9,
          "replies": [
            {
              "id": 3057860,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-11-28T18:10:49.040000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3055420,
      "author_name": "Param2007",
      "author_url": "",
      "post_date": "2024-11-25T19:00:09.357000",
      "content": "<p>Thanks for these tips!!!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3078337,
      "author_name": "hukaixin",
      "author_url": "",
      "post_date": "2024-12-22T07:19:30.857000",
      "content": "<p>Thank you for sharing your suggestions.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3051403,
      "author_name": "Switch9527",
      "author_url": "",
      "post_date": "2024-11-21T09:11:27.953000",
      "content": "<p>Thank you for sharing your suggestions.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3051236": "a comfortable lb score without many bells and whistles is about 0.630. If you are getting less than benchmark.csv score of 0.587, the below may help\n\n##1. design a workable end-to-end framework. Most important details are:\n### scanning window:\n  - it is not possible to process 184x630x630 directly and you are likely to break it into smaller scanning \"windows\" and feed them into your network, and then fuse the predictions together into one later.\n### coordinates extraction\n  - you need to post-process the predicted heatmap from the model and transform them to coordinates for submission.csv \n\nThese two details are very important. To ensure the framework is correct, you should\n- verify the end-to-end framework work on the **train set.** this gives the upper bound of your post processing (because when using the train set, you factor out problems due to generalisation error between train/validation set)\n- if you set **predicted heatmap  = your target ground truth** (for backpropagation), you MUST get a perfect score of near **1.00 in f-score**.\n\nlet me illustrate a few examples. In the extreme case, the 32x32x32 scanning window is not a good choice. there are border artifacts. For resnet34, it has 4 scales and its context is larger than 32. One should aim for the largest scanning window (in appropriate dimension) whenever possible.\n\n\n\n----\n\n\n##2. Select good model:\n\nYou have choices of 2d, 2.5d, 3d encoder. you also have choices of 2d, 3d/ seqenuce decoder. It is important to understand their context. e.g. if you use 2.5d unet then w,h>>d in the scan window (e.g. 630x630x16). if you use 3d decoder, then the depth d of your window should be large (e.g. 256x256x160). you should at least try various window size on:\n1. pure 2d unet\n2. pure 3d unet\n3. hybrid of 2d/3d\n\nwhy try 2d? this is because the dataset document mentions that some annotations are done in 2d. \"Try\" means proper experiment for local validation and compare correlation for public lb score.\n\nnever assume that 3d is always better than 2d (and vice versus). let the experiment results speak for themselves.\n\n\n\n\n----\n\n##3. Dataset processing.\nFirst, you need to normalise value to (0,1) or(-1,1). i think the experiment data all come from the **same** source and hence their intensity values are very smiliar. you have the choice of normalise by min/max, by std, by percentile. you can normalise by slice, subvolume,full volume  or share normalising parameters for all volumes.\nplease think about this.\n\nplease read the dataset paper pdf. it is very important.\n\n---\n\nhow many volumes should I use for validation, one or two or three?  how to stop training early?\n\ntry to use all denosied, wbp, isonet, etc ... for training.\n\ntry standard augmentation for now but better augmentation does give better results. \n\nhow about extra synthetic dataset? you can consider that later because my lb score of 0.700 is without synthetic yet (although I will use them later and expect some boost of 5%)",
    "3053819": "yes, it is possible to get zero for training (i.e. model not learning anything or just predicting all pixels as negative)\n\n## Training process\n- before you start training, compute the positive-to-negative ratio of the labeled pixels.\n- you already know that it is likely to cause an imbalance problem.\n- ideally, we want the segmentation targets to be disjoint (easier for post-processing later). But we want it to be large for a better positive-to-negative ratio to drive the back-propagation\n- Please think about it and need to search for the best size\n- alternatively, you can have 2 targets, one for post-processing, another for driving the back-propagation \n\n---\n\n- but there is another more serious problem. the target signal is quite weak.\n- if you are good at optimizer, you need to warm them up correctly (adam variants) and reset them to prevent them from getting stuck at the all-zero (and other) local minimum. Or you can use plain SGD which may be easier to control manually.\n- in my experiment, there is more than one minimum. \n   - if I leave train/validation data and model unchanged and just change the behavior of the optimizer, you can end up minimum (i.e. where training peaks or stabilizes) at cv=0, 0.5, 0.65, 0.75, 0.85\n\n",
    "3055516": "> if you set predicted heatmap = your target ground truth (for backpropagation), you MUST get a perfect score of near 1.00 in f-score.\n\nDo you currently achieve this? \n\nMy current approach is a bit simplistic and can likely be improved, but I am at: \n```\nScore for TS_73_6: 0.9792993955946692\nScore for TS_69_2: 0.998953427524856\nScore for TS_6_4: 1.0\nScore for TS_6_6: 0.9980008567756676\nScore for TS_86_3: 0.9882783882783883\nScore for TS_99_9: 0.9651746583029144\nScore for TS_5_4: 1.0\nTotal Score: 0.9872025639815435\n```\n\n\n",
    "3053548": "@hengck23  Having particle being splitted in the cropped 3D images I guess is not ideal. Should I enforce boundary conditions when cropping? I am currently using Monoai \"RandCropByLabelClassesd\" for cropping and my results are not great scoring under the benchmark with aprox 0.4. I have checked the cropped images and they always crop labels on the edges. \n",
    "3055316": "蛙哥蛙哥我偶像！每次阅读都是享受。",
    "3082686": "```\nOne should aim for the largest scanning window (in appropriate dimension) whenever possible.\n```\n\nConsidering batch_size = 1\n\nWhat is the shape of best scanning window that fits in T4 GPU?\n\nhttps://www.kaggle.com/code/kharrington/deepfindet-train : The demo notebook used (72, 72, 72)\nBut I think we can do better.\n\nSo I tried (64, 630, 630), (23, 630, 630), (23, 600, 600).\n\nWhen I tried (23, 600, 600)\n```\nOutOfMemoryError: CUDA out of memory. Tried to allocate 1.97 GiB. GPU 0 has a total capacity of 15.89 GiB of which 985.12 MiB is free. Process 2802 has 14.92 GiB memory in use. Of the allocated memory 14.52 GiB is allocated by PyTorch, and 118.38 MiB is reserved by PyTorch but unallocated.\n```\n\nWhat largest scanning window worked for you?",
    "3069068": ">please read the dataset paper pdf. it is very important.\n\nwhere would i find this",
    "3064141": "Hello, I have another question. The competition data includes four types of scan images（denoised.zarr, ctfdeconvolved.zarr, isonetcorrected.zarr, wbp.zarr）, but only the denoised type is included when submitting the code. So, during model training, should we include the other three types in the training dataset as well? I'm not sure if this will improve the model's performance, but I think it might be worth trying.",
    "3064137": "Hello, I was wondering if including data with different scales and resolutions ((46, 158, 158), (92, 315, 315), (84, 630, 630)) in the model training process, or enabling the model to support inputs with varying resolutions, could potentially improve its performance?\n\nI believe that data with different resolutions contain different types of information. For example, low-resolution data (shape: 46, 158, 158) might capture more global information, while high-resolution data focuses on local details. Different resolutions may provide complementary information.\n\nI'm not sure if this idea is correct 🤣.",
    "3058299": "Hi. First of all, thank for taking your time helping other people on the competition. Very useful info!\nI'm familiar with conventional (2d) unets but never stepped into a problem that required dividing input data into windows, so i have some doubts. My goal asking this questions is not to skip experimentation and self training on the problem but to understand certain intuitions . My approach is to make a dataset that is able to split volumes given a certain 3d window shape and 3d stride ( if stride shape is lower than window shape there's some overlap in the scanning windows ), then my unet implementation its a 3dunet so it abstracts all possible cases (2d, 2.5d and 3d) (for example for 2d i just have to set the window shape to (1, width, height)). Finally my ground-truth is made by setting gaussian distributions (taking in account particle radius) into the windows to make heatmaps with shape (depth, width, height, num_particles). \n\nMy question is how applying different overlaps and padding would affect heatmap estimation accuracy. Thank you in advance.    ",
    "3052527": "@hengck23 Thanks for all these tips. Do you get similar values between the score in validation set vs LB? ",
    "3051333": "Thanks for all those tips! Do you manage to do inference with 3D models within the time limits? If not do you think model quantization could be worth trying?",
    "3057289": "",
    "3055420": "Thanks for these tips!!!",
    "3078337": "Thank you for sharing your suggestions.",
    "3051403": "Thank you for sharing your suggestions."
  }
}