{
  "id": 548053,
  "title": "CZII CryoET: 630x630 PNG Dataset 8/16bit",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/548053",
  "author_name": "",
  "post_date": "2024-11-24T21:12:07.499557300Z",
  "votes": 15,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi, I saw most of you are struggling with the type of data, I have converted the competition dataset into 8/16bit PNG format. It includes Tomogram images frame by frame and extracted label (mask) images.</p>\n<p>Dataset Link: <a href=\"https://www.kaggle.com/datasets/gowrishankarp/czii-cryoet-630x630-png-dataset-816bit\" target=\"_blank\">CZII CryoET: 630x630 PNG Dataset 8/16bit</a><br>\nNotebook Link: <a href=\"https://www.kaggle.com/code/gowrishankarp/czii-data-preparation-8-16bit-png\" target=\"_blank\">CZII CryoET - Data Preparation 8/16bit (PNG)</a></p>\n<p>Images are captured at <code>voxel_size=10</code>, No changes were made to the dimensions of the images (630x630).</p>\n<p>Tomograms and segmentations were loaded using <code>get_tomogram</code> &amp; <code>segmentation_from_picks</code> using copick library,</p>\n<pre><code>tomogram = run.get_voxel_spacing(voxel_size).get_tomogram(tomo_type).numpy()\n</code></pre>\n<pre><code>pick = run.get_picks(object_name=pickable_object.name, user_id=)\ntarget = segmentation_from_picks.from_picks(\n    pick[], \n    target, \n    target_objects[pickable_object.name][] * ,\n    target_objects[pickable_object.name][],\n    voxel_spacing=voxel_size\n)\n</code></pre>\n<p>All the tomograms are normalized using,<br>\n<strong>Update</strong>: From the comments, I have changed min norm percentile to <strong>1</strong>, to capture more info. For more precision to capture more detail, you can use the 16bit images.</p>\n<pre><code> ():\n     = np.percentile(data,)\n     = np.percentile(data,)\n    data = np.clip(data,,)\n    data = (data-)/(-)\n     data\n</code></pre>\n<p>Label and corresponding mapping with its <code>pdb_id</code>, <code>particle_type</code> &amp; <code>radius</code>,</p>\n<pre><code> \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n        \n    \n</code></pre>\n<p>Preview of few tomograms with masks,<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F309e3882d5f2ad83b0237683878ae5ff%2FUntitled.png?generation=1732481149169182&amp;alt=media\" alt=\"TS_69_2\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F63b916ea77d5eb7c0b1f7159e50dff49%2Funtitled%201.png?generation=1732482382298180&amp;alt=media\" alt=\"TS_5_4\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F057aa7251cbdcd5cbd68b64137a2d84c%2Funtitled%202.png?generation=1732482575046448&amp;alt=media\" alt=\"TS_6_6\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2Fec5a71cc45301af906c60a5934728398%2Funtitled%203.png?generation=1732482621947948&amp;alt=media\" alt=\"TS_99_9\"><br>\n&gt;<br>\nThanks</p>\n<p>Ref: <br>\n<a href=\"https://www.kaggle.com/code/ahsuna123/3d-u-net-training-only\" target=\"_blank\">https://www.kaggle.com/code/ahsuna123/3d-u-net-training-only</a><br>\n<a href=\"https://www.kaggle.com/code/kharrington/blobdetector\" target=\"_blank\">https://www.kaggle.com/code/kharrington/blobdetector</a></p>",
  "messages": [
    {
      "id": "3054592",
      "postDate": "11/24/2024 21:12:07",
      "content": "<p>Hi, I saw most of you are struggling with the type of data, I have converted the competition dataset into 8/16bit PNG format. It includes Tomogram images frame by frame and extracted label (mask) images.</p>\n<p>Dataset Link: <a href=\"https://www.kaggle.com/datasets/gowrishankarp/czii-cryoet-630x630-png-dataset-816bit\" target=\"_blank\">CZII CryoET: 630x630 PNG Dataset 8/16bit</a><br>\nNotebook Link: <a href=\"https://www.kaggle.com/code/gowrishankarp/czii-data-preparation-8-16bit-png\" target=\"_blank\">CZII CryoET - Data Preparation 8/16bit (PNG)</a></p>\n<p>Images are captured at <code>voxel_size=10</code>, No changes were made to the dimensions of the images (630x630).</p>\n<p>Tomograms and segmentations were loaded using <code>get_tomogram</code> &amp; <code>segmentation_from_picks</code> using copick library,</p>\n<pre><code>tomogram = run.get_voxel_spacing(voxel_size).get_tomogram(tomo_type).numpy()\n</code></pre>\n<pre><code>pick = run.get_picks(object_name=pickable_object.name, user_id=)\ntarget = segmentation_from_picks.from_picks(\n    pick[], \n    target, \n    target_objects[pickable_object.name][] * ,\n    target_objects[pickable_object.name][],\n    voxel_spacing=voxel_size\n)\n</code></pre>\n<p>All the tomograms are normalized using,<br>\n<strong>Update</strong>: From the comments, I have changed min norm percentile to <strong>1</strong>, to capture more info. For more precision to capture more detail, you can use the 16bit images.</p>\n<pre><code> ():\n     = np.percentile(data,)\n     = np.percentile(data,)\n    data = np.clip(data,,)\n    data = (data-)/(-)\n     data\n</code></pre>\n<p>Label and corresponding mapping with its <code>pdb_id</code>, <code>particle_type</code> &amp; <code>radius</code>,</p>\n<pre><code> \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n             \n        \n        \n             \n             \n             \n             \n        \n    \n</code></pre>\n<p>Preview of few tomograms with masks,<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F309e3882d5f2ad83b0237683878ae5ff%2FUntitled.png?generation=1732481149169182&amp;alt=media\" alt=\"TS_69_2\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F63b916ea77d5eb7c0b1f7159e50dff49%2Funtitled%201.png?generation=1732482382298180&amp;alt=media\" alt=\"TS_5_4\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F057aa7251cbdcd5cbd68b64137a2d84c%2Funtitled%202.png?generation=1732482575046448&amp;alt=media\" alt=\"TS_6_6\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2Fec5a71cc45301af906c60a5934728398%2Funtitled%203.png?generation=1732482621947948&amp;alt=media\" alt=\"TS_99_9\"><br>\n&gt;<br>\nThanks</p>\n<p>Ref: <br>\n<a href=\"https://www.kaggle.com/code/ahsuna123/3d-u-net-training-only\" target=\"_blank\">https://www.kaggle.com/code/ahsuna123/3d-u-net-training-only</a><br>\n<a href=\"https://www.kaggle.com/code/kharrington/blobdetector\" target=\"_blank\">https://www.kaggle.com/code/kharrington/blobdetector</a></p>",
      "rawMarkdown": "Hi, I saw most of you are struggling with the type of data, I have converted the competition dataset into 8/16bit PNG format. It includes Tomogram images frame by frame and extracted label (mask) images.\n\nDataset Link: [CZII CryoET: 630x630 PNG Dataset 8/16bit](https://www.kaggle.com/datasets/gowrishankarp/czii-cryoet-630x630-png-dataset-816bit)\nNotebook Link: [CZII CryoET - Data Preparation 8/16bit (PNG)](https://www.kaggle.com/code/gowrishankarp/czii-data-preparation-8-16bit-png)\n\nImages are captured at `voxel_size=10`, No changes were made to the dimensions of the images (630x630).\n\nTomograms and segmentations were loaded using `get_tomogram` & `segmentation_from_picks` using copick library,\n```python\ntomogram = run.get_voxel_spacing(voxel_size).get_tomogram(tomo_type).numpy()\n```\n```python\npick = run.get_picks(object_name=pickable_object.name, user_id=\"curation\")\ntarget = segmentation_from_picks.from_picks(\n    pick[0], \n    target, \n    target_objects[pickable_object.name]['radius'] * 0.8,\n    target_objects[pickable_object.name]['label'],\n    voxel_spacing=voxel_size\n)\n```\n\nAll the tomograms are normalized using,\n**Update**: From the comments, I have changed min norm percentile to **1**, to capture more info. For more precision to capture more detail, you can use the 16bit images.\n\n```python\ndef normalise_by_percentile(data, min=1, max=99):\n    min = np.percentile(data,min)\n    max = np.percentile(data,max)\n    data = np.clip(data,min,max)\n    data = (data-min)/(max-min)\n    return data\n```\n\n\nLabel and corresponding mapping with its `pdb_id`, `particle_type` & `radius`,\n```\n\"pickable_objects\": [\n        {\n            \"name\": \"apo-ferritin\",\n            \"is_particle\": true,\n            \"pdb_id\": \"4V1W\",\n            \"label\": 1,\n            \"radius\": 60,\n        },\n        {\n            \"name\": \"beta-amylase\",\n            \"is_particle\": true,\n            \"pdb_id\": \"1FA2\",\n            \"label\": 2,\n            \"radius\": 65,\n        },\n        {\n            \"name\": \"beta-galactosidase\",\n            \"is_particle\": true,\n            \"pdb_id\": \"6X1Q\",\n            \"label\": 3,\n            \"radius\": 90,\n        },\n        {\n            \"name\": \"ribosome\",\n            \"is_particle\": true,\n            \"pdb_id\": \"6EK0\",\n            \"label\": 4,\n            \"radius\": 150,\n        },\n        {\n            \"name\": \"thyroglobulin\",\n            \"is_particle\": true,\n            \"pdb_id\": \"6SCJ\",\n            \"label\": 5,\n            \"radius\": 130,\n        },\n        {\n            \"name\": \"virus-like-particle\",\n            \"is_particle\": true,\n            \"label\": 6,\n            \"radius\": 135,\n        }\n    ],\n```\n\nPreview of few tomograms with masks,\n![TS_69_2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F309e3882d5f2ad83b0237683878ae5ff%2FUntitled.png?generation=1732481149169182&alt=media)\n![TS_5_4](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F63b916ea77d5eb7c0b1f7159e50dff49%2Funtitled%201.png?generation=1732482382298180&alt=media)\n![TS_6_6](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F057aa7251cbdcd5cbd68b64137a2d84c%2Funtitled%202.png?generation=1732482575046448&alt=media)\n![TS_99_9](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2Fec5a71cc45301af906c60a5934728398%2Funtitled%203.png?generation=1732482621947948&alt=media)\n>\nThanks\n\nRef: \nhttps://www.kaggle.com/code/ahsuna123/3d-u-net-training-only\nhttps://www.kaggle.com/code/kharrington/blobdetector",
      "votes": null
    },
    {
      "id": "3054608",
      "postDate": "11/24/2024 21:45:45",
      "content": "<p>What is a good strategy to find the minim and max percentile to use when normalizing? </p>",
      "rawMarkdown": "What is a good strategy to find the minim and max percentile to use when normalizing?",
      "votes": null
    },
    {
      "id": "3054651",
      "postDate": "11/25/2024 00:13:10",
      "content": "<p>Experimentation.</p>\n<p>EDIT: The reason of using  percentiles as min and max scaling is to keep out outlier values, the tails of intensities distribution. As more tail you cut, more probable is you'll cut outliers, but more probable is too you'll cut valuable information. Actually you don't cut anything. Just leave those values out 0 to 1 range.</p>\n<p>I think around 1-5 and 95-99 are reasonable, but you'll have to try.</p>",
      "rawMarkdown": "Experimentation.\n\nEDIT: The reason of using  percentiles as min and max scaling is to keep out outlier values, the tails of intensities distribution. As more tail you cut, more probable is you'll cut outliers, but more probable is too you'll cut valuable information. Actually you don't cut anything. Just leave those values out 0 to 1 range.\n\nI think around 1-5 and 95-99 are reasonable, but you'll have to try.",
      "votes": null
    },
    {
      "id": "3054653",
      "postDate": "11/25/2024 00:14:04",
      "content": "<p>Hi. Thanks for share. Personally I've been tunning radius visually. Can I ask from where you took these ones?</p>",
      "rawMarkdown": "Hi. Thanks for share. Personally I've been tunning radius visually. Can I ask from where you took these ones?",
      "votes": null
    },
    {
      "id": "3054737",
      "postDate": "11/25/2024 04:19:43",
      "content": "<p>Hi, Thanks for the comment, the values were taken from the pinned notebook <a href=\"https://www.kaggle.com/code/kharrington/blobdetector\" target=\"_blank\">BlobDetector</a> by competition organizers <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>.</p>",
      "rawMarkdown": "Hi, Thanks for the comment, the values were taken from the pinned notebook [BlobDetector](https://www.kaggle.com/code/kharrington/blobdetector) by competition organizers @kharrington.",
      "votes": null
    },
    {
      "id": "3054894",
      "postDate": "11/25/2024 10:10:01",
      "content": "<p>Yes that is what I was thinking. You are actually not cutting anything just squeezing all on the tails. So you need to make sure I guess the pixels being squashed are not labels as you are losing information</p>",
      "rawMarkdown": "Yes that is what I was thinking. You are actually not cutting anything just squeezing all on the tails. So you need to make sure I guess the pixels being squashed are not labels as you are losing information",
      "votes": null
    },
    {
      "id": "3054908",
      "postDate": "11/25/2024 10:25:28",
      "content": "<p>my suggestion is cut to 0,1 for visualiation only. for modeling you can keep the values. if u want to cut the values for modeling, the support should should come from experimental results.</p>\n<p>in datascience, data speaks for itself. that is the best strategy. just need to make sure your experiment results are not bias.</p>",
      "rawMarkdown": "my suggestion is cut to 0,1 for visualiation only. for modeling you can keep the values. if u want to cut the values for modeling, the support should should come from experimental results.\n\nin datascience, data speaks for itself. that is the best strategy. just need to make sure your experiment results are not bias.",
      "votes": null
    },
    {
      "id": "3054910",
      "postDate": "11/25/2024 10:27:19",
      "content": "<p>one note: you can probe hidden test data to see if outliers values occur</p>",
      "rawMarkdown": "one note: you can probe hidden test data to see if outliers values occur",
      "votes": null
    },
    {
      "id": "3055236",
      "postDate": "11/25/2024 16:40:47",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thanks for the comment, instead of directly trimming, maybe we can try something like this, inspired from previous SenNet + HOA Competition.</p>\n<pre><code> ():\n     = np.percentile(data,)\n     = np.percentile(data,)\n    data[data&gt;] = (data[data&gt;]-)* + \n    data[data&lt;] = (data[data&lt;]-)* + \n    data = (data-)/(-+smooth)\n     data\n</code></pre>",
      "rawMarkdown": "Hi @hengck23, thanks for the comment, instead of directly trimming, maybe we can try something like this, inspired from previous SenNet + HOA Competition.\n\n```python\ndef normalise_by_percentile(data, min=5, max=99, smooth=1e-5):\n    min = np.percentile(data,min)\n    max = np.percentile(data,max)\n    data[data>max] = (data[data>max]-max)*1e-3 + max\n    data[data<min] = (data[data<min]-min)*1e-3 + min\n    data = (data-min)/(max-min+smooth)\n    return data\n```",
      "votes": null
    },
    {
      "id": "3061809",
      "postDate": "12/03/2024 02:14:07",
      "content": "<p>HI, I have also added <strong>Bounding_Boxes</strong> for each frame and each label <code>[x_center, y_center, height, width]</code>. Which can help in performing any Object Detection task.</p>",
      "rawMarkdown": "HI, I have also added **Bounding_Boxes** for each frame and each label `[x_center, y_center, height, width]`. Which can help in performing any Object Detection task.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3054608,
      "author_name": "ignasialemany",
      "author_url": "",
      "post_date": "11/24/2024 21:45:45",
      "content": "<p>What is a good strategy to find the minim and max percentile to use when normalizing? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3054651,
          "author_name": "sacuscreed",
          "author_url": "",
          "post_date": "11/25/2024 00:13:10",
          "content": "<p>Experimentation.</p>\n<p>EDIT: The reason of using  percentiles as min and max scaling is to keep out outlier values, the tails of intensities distribution. As more tail you cut, more probable is you'll cut outliers, but more probable is too you'll cut valuable information. Actually you don't cut anything. Just leave those values out 0 to 1 range.</p>\n<p>I think around 1-5 and 95-99 are reasonable, but you'll have to try.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3054894,
              "author_name": "ignasialemany",
              "author_url": "",
              "post_date": "11/25/2024 10:10:01",
              "content": "<p>Yes that is what I was thinking. You are actually not cutting anything just squeezing all on the tails. So you need to make sure I guess the pixels being squashed are not labels as you are losing information</p>",
              "votes": null,
              "replies": []
            },
            {
              "id": 3054908,
              "author_name": "hengck23",
              "author_url": "",
              "post_date": "11/25/2024 10:25:28",
              "content": "<p>my suggestion is cut to 0,1 for visualiation only. for modeling you can keep the values. if u want to cut the values for modeling, the support should should come from experimental results.</p>\n<p>in datascience, data speaks for itself. that is the best strategy. just need to make sure your experiment results are not bias.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3054910,
                  "author_name": "hengck23",
                  "author_url": "",
                  "post_date": "11/25/2024 10:27:19",
                  "content": "<p>one note: you can probe hidden test data to see if outliers values occur</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 3055236,
                  "author_name": "gowrishankarp",
                  "author_url": "",
                  "post_date": "11/25/2024 16:40:47",
                  "content": "<p>Hi <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a>, thanks for the comment, instead of directly trimming, maybe we can try something like this, inspired from previous SenNet + HOA Competition.</p>\n<pre><code> ():\n     = np.percentile(data,)\n     = np.percentile(data,)\n    data[data&gt;] = (data[data&gt;]-)* + \n    data[data&lt;] = (data[data&lt;]-)* + \n    data = (data-)/(-+smooth)\n     data\n</code></pre>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3054653,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "11/25/2024 00:14:04",
      "content": "<p>Hi. Thanks for share. Personally I've been tunning radius visually. Can I ask from where you took these ones?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3054737,
          "author_name": "gowrishankarp",
          "author_url": "",
          "post_date": "11/25/2024 04:19:43",
          "content": "<p>Hi, Thanks for the comment, the values were taken from the pinned notebook <a href=\"https://www.kaggle.com/code/kharrington/blobdetector\" target=\"_blank\">BlobDetector</a> by competition organizers <a href=\"https://www.kaggle.com/kharrington\" target=\"_blank\">@kharrington</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3061809,
      "author_name": "gowrishankarp",
      "author_url": "",
      "post_date": "12/03/2024 02:14:07",
      "content": "<p>HI, I have also added <strong>Bounding_Boxes</strong> for each frame and each label <code>[x_center, y_center, height, width]</code>. Which can help in performing any Object Detection task.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3054592": "Hi, I saw most of you are struggling with the type of data, I have converted the competition dataset into 8/16bit PNG format. It includes Tomogram images frame by frame and extracted label (mask) images.\n\nDataset Link: [CZII CryoET: 630x630 PNG Dataset 8/16bit](https://www.kaggle.com/datasets/gowrishankarp/czii-cryoet-630x630-png-dataset-816bit)\nNotebook Link: [CZII CryoET - Data Preparation 8/16bit (PNG)](https://www.kaggle.com/code/gowrishankarp/czii-data-preparation-8-16bit-png)\n\nImages are captured at `voxel_size=10`, No changes were made to the dimensions of the images (630x630).\n\nTomograms and segmentations were loaded using `get_tomogram` & `segmentation_from_picks` using copick library,\n```python\ntomogram = run.get_voxel_spacing(voxel_size).get_tomogram(tomo_type).numpy()\n```\n```python\npick = run.get_picks(object_name=pickable_object.name, user_id=\"curation\")\ntarget = segmentation_from_picks.from_picks(\n    pick[0], \n    target, \n    target_objects[pickable_object.name]['radius'] * 0.8,\n    target_objects[pickable_object.name]['label'],\n    voxel_spacing=voxel_size\n)\n```\n\nAll the tomograms are normalized using,\n**Update**: From the comments, I have changed min norm percentile to **1**, to capture more info. For more precision to capture more detail, you can use the 16bit images.\n\n```python\ndef normalise_by_percentile(data, min=1, max=99):\n    min = np.percentile(data,min)\n    max = np.percentile(data,max)\n    data = np.clip(data,min,max)\n    data = (data-min)/(max-min)\n    return data\n```\n\n\nLabel and corresponding mapping with its `pdb_id`, `particle_type` & `radius`,\n```\n\"pickable_objects\": [\n        {\n            \"name\": \"apo-ferritin\",\n            \"is_particle\": true,\n            \"pdb_id\": \"4V1W\",\n            \"label\": 1,\n            \"radius\": 60,\n        },\n        {\n            \"name\": \"beta-amylase\",\n            \"is_particle\": true,\n            \"pdb_id\": \"1FA2\",\n            \"label\": 2,\n            \"radius\": 65,\n        },\n        {\n            \"name\": \"beta-galactosidase\",\n            \"is_particle\": true,\n            \"pdb_id\": \"6X1Q\",\n            \"label\": 3,\n            \"radius\": 90,\n        },\n        {\n            \"name\": \"ribosome\",\n            \"is_particle\": true,\n            \"pdb_id\": \"6EK0\",\n            \"label\": 4,\n            \"radius\": 150,\n        },\n        {\n            \"name\": \"thyroglobulin\",\n            \"is_particle\": true,\n            \"pdb_id\": \"6SCJ\",\n            \"label\": 5,\n            \"radius\": 130,\n        },\n        {\n            \"name\": \"virus-like-particle\",\n            \"is_particle\": true,\n            \"label\": 6,\n            \"radius\": 135,\n        }\n    ],\n```\n\nPreview of few tomograms with masks,\n![TS_69_2](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F309e3882d5f2ad83b0237683878ae5ff%2FUntitled.png?generation=1732481149169182&alt=media)\n![TS_5_4](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F63b916ea77d5eb7c0b1f7159e50dff49%2Funtitled%201.png?generation=1732482382298180&alt=media)\n![TS_6_6](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2F057aa7251cbdcd5cbd68b64137a2d84c%2Funtitled%202.png?generation=1732482575046448&alt=media)\n![TS_99_9](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F8096091%2Fec5a71cc45301af906c60a5934728398%2Funtitled%203.png?generation=1732482621947948&alt=media)\n>\nThanks\n\nRef: \nhttps://www.kaggle.com/code/ahsuna123/3d-u-net-training-only\nhttps://www.kaggle.com/code/kharrington/blobdetector",
    "3054608": "What is a good strategy to find the minim and max percentile to use when normalizing?",
    "3054651": "Experimentation.\n\nEDIT: The reason of using  percentiles as min and max scaling is to keep out outlier values, the tails of intensities distribution. As more tail you cut, more probable is you'll cut outliers, but more probable is too you'll cut valuable information. Actually you don't cut anything. Just leave those values out 0 to 1 range.\n\nI think around 1-5 and 95-99 are reasonable, but you'll have to try.",
    "3054653": "Hi. Thanks for share. Personally I've been tunning radius visually. Can I ask from where you took these ones?",
    "3054737": "Hi, Thanks for the comment, the values were taken from the pinned notebook [BlobDetector](https://www.kaggle.com/code/kharrington/blobdetector) by competition organizers @kharrington.",
    "3054894": "Yes that is what I was thinking. You are actually not cutting anything just squeezing all on the tails. So you need to make sure I guess the pixels being squashed are not labels as you are losing information",
    "3054908": "my suggestion is cut to 0,1 for visualiation only. for modeling you can keep the values. if u want to cut the values for modeling, the support should should come from experimental results.\n\nin datascience, data speaks for itself. that is the best strategy. just need to make sure your experiment results are not bias.",
    "3054910": "one note: you can probe hidden test data to see if outliers values occur",
    "3055236": "Hi @hengck23, thanks for the comment, instead of directly trimming, maybe we can try something like this, inspired from previous SenNet + HOA Competition.\n\n```python\ndef normalise_by_percentile(data, min=5, max=99, smooth=1e-5):\n    min = np.percentile(data,min)\n    max = np.percentile(data,max)\n    data[data>max] = (data[data>max]-max)*1e-3 + max\n    data[data<min] = (data[data<min]-min)*1e-3 + min\n    data = (data-min)/(max-min+smooth)\n    return data\n```",
    "3061809": "HI, I have also added **Bounding_Boxes** for each frame and each label `[x_center, y_center, height, width]`. Which can help in performing any Object Detection task."
  },
  "source": "meta"
}