{
  "id": 575028,
  "title": "CryoET Dataset with Pixel Anomalies Corrected",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028",
  "author_name": "",
  "post_date": "2025-04-25T13:06:27.096647700Z",
  "votes": 34,
  "comment_count": 15,
  "views": 0,
  "content": "<h2>Summary</h2>\n<p>I released the corrected dataset of pixel anomalies found on the <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s external dataset.</p>\n<p>link: <a href=\"https://www.kaggle.com/datasets/tatamikenn/byu-cryoet-dataset-with-pixel-anomalies-corrected/data\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/byu-cryoet-dataset-with-pixel-anomalies-corrected/data</a></p>\n<p>Some of the original data (from the CryoET Data Portal) appear to have been incorrectly quantized, resulting in pixel values ranging from -128 to 127 — as if they were stored as <code>int8</code>, even though the expected data range should be from 0 to 255 like <code>uint8</code>.<br>\nWhich causes significant corruption of source image (see the pictures in visualization section).</p>\n<p>The key correction applied is as follows:</p>\n<pre><code>x = x.astype().astype()\n</code></pre>\n<p>For more detail, you can refer to the following notebooks:</p>\n<ul>\n<li>EDA: <a href=\"https://www.kaggle.com/code/tatamikenn/byu-eda-on-the-irregular-data\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/byu-eda-on-the-irregular-data</a></li>\n<li>Process: <a href=\"https://www.kaggle.com/code/tatamikenn/byu-download-irregular-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/byu-download-irregular-data/notebook</a></li>\n</ul>\n<h2>Visualization</h2>\n<p><strong>Before correction</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8d2166da33e933acf89e5acc9c865b74%2FScreenshot%202025-04-25%20at%2022.04.25.png?generation=1745586325626721&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1c29e40e7d196d4c741b0118d2031752%2FScreenshot%202025-04-25%20at%2022.04.47.png?generation=1745586341084858&amp;alt=media\" alt=\"\"></p>\n<p><strong>After correction</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbc8601a58e8092323221848400cbfc03%2FScreenshot%202025-04-25%20at%2022.04.36.png?generation=1745586366776285&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fb7da5f6b1117cc4de4b7510b9e379a29%2FScreenshot%202025-04-25%20at%2022.04.57.png?generation=1745586379171915&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3187031",
      "postDate": "04/25/2025 13:06:27",
      "content": "<h2>Summary</h2>\n<p>I released the corrected dataset of pixel anomalies found on the <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a>'s external dataset.</p>\n<p>link: <a href=\"https://www.kaggle.com/datasets/tatamikenn/byu-cryoet-dataset-with-pixel-anomalies-corrected/data\" target=\"_blank\">https://www.kaggle.com/datasets/tatamikenn/byu-cryoet-dataset-with-pixel-anomalies-corrected/data</a></p>\n<p>Some of the original data (from the CryoET Data Portal) appear to have been incorrectly quantized, resulting in pixel values ranging from -128 to 127 — as if they were stored as <code>int8</code>, even though the expected data range should be from 0 to 255 like <code>uint8</code>.<br>\nWhich causes significant corruption of source image (see the pictures in visualization section).</p>\n<p>The key correction applied is as follows:</p>\n<pre><code>x = x.astype().astype()\n</code></pre>\n<p>For more detail, you can refer to the following notebooks:</p>\n<ul>\n<li>EDA: <a href=\"https://www.kaggle.com/code/tatamikenn/byu-eda-on-the-irregular-data\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/byu-eda-on-the-irregular-data</a></li>\n<li>Process: <a href=\"https://www.kaggle.com/code/tatamikenn/byu-download-irregular-data/notebook\" target=\"_blank\">https://www.kaggle.com/code/tatamikenn/byu-download-irregular-data/notebook</a></li>\n</ul>\n<h2>Visualization</h2>\n<p><strong>Before correction</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8d2166da33e933acf89e5acc9c865b74%2FScreenshot%202025-04-25%20at%2022.04.25.png?generation=1745586325626721&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1c29e40e7d196d4c741b0118d2031752%2FScreenshot%202025-04-25%20at%2022.04.47.png?generation=1745586341084858&amp;alt=media\" alt=\"\"></p>\n<p><strong>After correction</strong>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbc8601a58e8092323221848400cbfc03%2FScreenshot%202025-04-25%20at%2022.04.36.png?generation=1745586366776285&amp;alt=media\" alt=\"\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fb7da5f6b1117cc4de4b7510b9e379a29%2FScreenshot%202025-04-25%20at%2022.04.57.png?generation=1745586379171915&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "## Summary\n\nI released the corrected dataset of pixel anomalies found on the @brendanartley's external dataset.\n\nlink: https://www.kaggle.com/datasets/tatamikenn/byu-cryoet-dataset-with-pixel-anomalies-corrected/data\n\nSome of the original data (from the CryoET Data Portal) appear to have been incorrectly quantized, resulting in pixel values ranging from -128 to 127 — as if they were stored as `int8`, even though the expected data range should be from 0 to 255 like `uint8`.\nWhich causes significant corruption of source image (see the pictures in visualization section).\n\nThe key correction applied is as follows:\n\n```python\nx = x.astype(\"uint8\").astype(\"float32\")\n```\n\nFor more detail, you can refer to the following notebooks:\n- EDA: https://www.kaggle.com/code/tatamikenn/byu-eda-on-the-irregular-data\n- Process: https://www.kaggle.com/code/tatamikenn/byu-download-irregular-data/notebook\n\n## Visualization\n\n**Before correction**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8d2166da33e933acf89e5acc9c865b74%2FScreenshot%202025-04-25%20at%2022.04.25.png?generation=1745586325626721&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1c29e40e7d196d4c741b0118d2031752%2FScreenshot%202025-04-25%20at%2022.04.47.png?generation=1745586341084858&alt=media)\n\n**After correction**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbc8601a58e8092323221848400cbfc03%2FScreenshot%202025-04-25%20at%2022.04.36.png?generation=1745586366776285&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fb7da5f6b1117cc4de4b7510b9e379a29%2FScreenshot%202025-04-25%20at%2022.04.57.png?generation=1745586379171915&alt=media)",
      "votes": null
    },
    {
      "id": "3187038",
      "postDate": "04/25/2025 13:27:41",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>!</p>",
      "rawMarkdown": "Thanks @tatamikenn!",
      "votes": null
    },
    {
      "id": "3187056",
      "postDate": "04/25/2025 13:55:58",
      "content": "<p>Thanks man- there seem to be a few blackened images like this in the original training set as well? Will the same correction work there?</p>",
      "rawMarkdown": "Thanks man- there seem to be a few blackened images like this in the original training set as well? Will the same correction work there?",
      "votes": null
    },
    {
      "id": "3187090",
      "postDate": "04/25/2025 14:50:37",
      "content": "<p>Thank you for the helpful information.</p>",
      "rawMarkdown": "Thank you for the helpful information.",
      "votes": null
    },
    {
      "id": "3187178",
      "postDate": "04/25/2025 16:41:40",
      "content": "<p>thanks man god job </p>",
      "rawMarkdown": "thanks man god job",
      "votes": null
    },
    {
      "id": "3187364",
      "postDate": "04/25/2025 22:06:19",
      "content": "<p>Thank you and great work!<br>\nIn your dataset in incorrect labels, there are tomogram names and z_slice values for those tomograms, are only THOSE slices incorrectly quantized? <br>\nI would generally expect not just one , but nearby slices to be incorrectly quantized too….  </p>",
      "rawMarkdown": "Thank you and great work!\nIn your dataset in incorrect labels, there are tomogram names and z_slice values for those tomograms, are only THOSE slices incorrectly quantized? \nI would generally expect not just one , but nearby slices to be incorrectly quantized too....",
      "votes": null
    },
    {
      "id": "3187385",
      "postDate": "04/25/2025 23:46:21",
      "content": "<p>Don't care much of this column. These are slices which exist annotations. I added it to check the image around motors.</p>",
      "rawMarkdown": "Don't care much of this column. These are slices which exist annotations. I added it to check the image around motors.",
      "votes": null
    },
    {
      "id": "3188959",
      "postDate": "04/28/2025 14:37:14",
      "content": "<p>Thanks for sharings it helps alot</p>",
      "rawMarkdown": "Thanks for sharings it helps alot",
      "votes": null
    },
    {
      "id": "3198930",
      "postDate": "05/10/2025 07:53:09",
      "content": "<p>Hi Thanks for the insights!</p>\n<p>But I have a quesiton. Let's see we will implement a data normalization for preprocessing. <br>\nI.e.  Normalize slice data using the 2nd and 98th percentiles and Normalized image in the range [0, 255]. In such case, I think this irregularity will not be an issue any more. Is it?</p>\n<p>Best regards<br>\nLeo</p>",
      "rawMarkdown": "Hi Thanks for the insights!\n\nBut I have a quesiton. Let's see we will implement a data normalization for preprocessing. \nI.e.  Normalize slice data using the 2nd and 98th percentiles and Normalized image in the range [0, 255]. In such case, I think this irregularity will not be an issue any more. Is it?\n\nBest regards\nLeo",
      "votes": null
    },
    {
      "id": "3217394",
      "postDate": "06/05/2025 01:28:14",
      "content": "<p>I think fixing the anomalies of the test was key to get stable score in private. We didn't do that in our final subs, and unfortunately shaked down significantly. It is sad how this was not really represented in public lb tomos leading to all our work get punished severely :(</p>",
      "rawMarkdown": "I think fixing the anomalies of the test was key to get stable score in private. We didn't do that in our final subs, and unfortunately shaked down significantly. It is sad how this was not really represented in public lb tomos leading to all our work get punished severely :(",
      "votes": null
    },
    {
      "id": "3217405",
      "postDate": "06/05/2025 01:53:26",
      "content": "<p>I'm realized that I forget to add these logics to correct the wrong quantiled tomogram in submission code too, right before clicking \"Submit\" button for the last submission, but decided not gonna fix and still continue to submit caused of 12h deadline limitation. My reasons for this:</p>\n<ul>\n<li>I've submit that logic and observe no change in public LB, like yours</li>\n<li>Tomograms from public/private test seem to be produced/belong privately to a lab, with good and unified data collecting/preprocessing techniques, so such simple errors seem not very likely</li>\n<li>Host had train models, visualize, label, review the tomogram again. If I were the host, I will fix this before release the competition dataset instead of forcing Kaggler to \"probe/guess the dataset\"</li>\n</ul>\n<p>But, now I'm making late submissions to verify these as well.</p>",
      "rawMarkdown": "I'm realized that I forget to add these logics to correct the wrong quantiled tomogram in submission code too, right before clicking \"Submit\" button for the last submission, but decided not gonna fix and still continue to submit caused of 12h deadline limitation. My reasons for this:\n- I've submit that logic and observe no change in public LB, like yours\n- Tomograms from public/private test seem to be produced/belong privately to a lab, with good and unified data collecting/preprocessing techniques, so such simple errors seem not very likely\n- Host had train models, visualize, label, review the tomogram again. If I were the host, I will fix this before release the competition dataset instead of forcing Kaggler to \"probe/guess the dataset\"\n\nBut, now I'm making late submissions to verify these as well.",
      "votes": null
    },
    {
      "id": "3217406",
      "postDate": "06/05/2025 01:56:57",
      "content": "<blockquote>\n  <p>I've submit that logic and observe no change in public LB, like yours</p>\n</blockquote>\n<p>I observed similar result in the public LB. So this correction was not applied in my final sub. However, not sure in private LB.</p>",
      "rawMarkdown": "> I've submit that logic and observe no change in public LB, like yours\n\nI observed similar result in the public LB. So this correction was not applied in my final sub. However, not sure in private LB.",
      "votes": null
    },
    {
      "id": "3217410",
      "postDate": "06/05/2025 02:02:16",
      "content": "<p>I believe correction by quantization is not robust way of fixing this anomaly because it depends on the choice of the quantiles. Note that the absolute error value is at most 127.</p>",
      "rawMarkdown": "I believe correction by quantization is not robust way of fixing this anomaly because it depends on the choice of the quantiles. Note that the absolute error value is at most 127.",
      "votes": null
    },
    {
      "id": "3217413",
      "postDate": "06/05/2025 02:05:12",
      "content": "<p>Our private lb scores honestly looks very random. I don't see clear correlation, or why we shaked down.<br>\nSome random ensembles produced much better results than others without clear reason. But the only sub that has clearly a different pattern was the one with the anomalies fix added. That's why i believe it was the reason.</p>",
      "rawMarkdown": "Our private lb scores honestly looks very random. I don't see clear correlation, or why we shaked down.\nSome random ensembles produced much better results than others without clear reason. But the only sub that has clearly a different pattern was the one with the anomalies fix added. That's why i believe it was the reason.",
      "votes": null
    },
    {
      "id": "3217416",
      "postDate": "06/05/2025 02:07:19",
      "content": "<p>The difference is huge. I am talking about smth like this:<br>\nPublic: 0.841, private 0.785 (w/o)<br>\nPublic: 0.841, private 0.818 (with fix)</p>",
      "rawMarkdown": "The difference is huge. I am talking about smth like this:\nPublic: 0.841, private 0.785 (w/o)\nPublic: 0.841, private 0.818 (with fix)",
      "votes": null
    },
    {
      "id": "3217426",
      "postDate": "06/05/2025 02:12:18",
      "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> thank you for sharing the result. It is noticeable gain.<br>\nMaybe test tomogram also contains this anomaly, but I think we need more submissions, considering high randomness in this competition.</p>",
      "rawMarkdown": "mohammad2012191 thank you for sharing the result. It is noticeable gain.\nMaybe test tomogram also contains this anomaly, but I think we need more submissions, considering high randomness in this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3187038,
      "author_name": "brendanartley",
      "author_url": "",
      "post_date": "04/25/2025 13:27:41",
      "content": "<p>Thanks <a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3187056,
      "author_name": "yahalom",
      "author_url": "",
      "post_date": "04/25/2025 13:55:58",
      "content": "<p>Thanks man- there seem to be a few blackened images like this in the original training set as well? Will the same correction work there?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3187090,
      "author_name": "sohilaelmassry",
      "author_url": "",
      "post_date": "04/25/2025 14:50:37",
      "content": "<p>Thank you for the helpful information.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3187178,
      "author_name": "analyticaobscura",
      "author_url": "",
      "post_date": "04/25/2025 16:41:40",
      "content": "<p>thanks man god job </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3187364,
      "author_name": "iamudit",
      "author_url": "",
      "post_date": "04/25/2025 22:06:19",
      "content": "<p>Thank you and great work!<br>\nIn your dataset in incorrect labels, there are tomogram names and z_slice values for those tomograms, are only THOSE slices incorrectly quantized? <br>\nI would generally expect not just one , but nearby slices to be incorrectly quantized too….  </p>",
      "votes": null,
      "replies": [
        {
          "id": 3187385,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "04/25/2025 23:46:21",
          "content": "<p>Don't care much of this column. These are slices which exist annotations. I added it to check the image around motors.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3188959,
      "author_name": "saraharshadbcs",
      "author_url": "",
      "post_date": "04/28/2025 14:37:14",
      "content": "<p>Thanks for sharings it helps alot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3198930,
      "author_name": "yxyyxy",
      "author_url": "",
      "post_date": "05/10/2025 07:53:09",
      "content": "<p>Hi Thanks for the insights!</p>\n<p>But I have a quesiton. Let's see we will implement a data normalization for preprocessing. <br>\nI.e.  Normalize slice data using the 2nd and 98th percentiles and Normalized image in the range [0, 255]. In such case, I think this irregularity will not be an issue any more. Is it?</p>\n<p>Best regards<br>\nLeo</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217410,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "06/05/2025 02:02:16",
          "content": "<p>I believe correction by quantization is not robust way of fixing this anomaly because it depends on the choice of the quantiles. Note that the absolute error value is at most 127.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3217394,
      "author_name": "mohammad2012191",
      "author_url": "",
      "post_date": "06/05/2025 01:28:14",
      "content": "<p>I think fixing the anomalies of the test was key to get stable score in private. We didn't do that in our final subs, and unfortunately shaked down significantly. It is sad how this was not really represented in public lb tomos leading to all our work get punished severely :(</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217405,
          "author_name": "dangnh0611",
          "author_url": "",
          "post_date": "06/05/2025 01:53:26",
          "content": "<p>I'm realized that I forget to add these logics to correct the wrong quantiled tomogram in submission code too, right before clicking \"Submit\" button for the last submission, but decided not gonna fix and still continue to submit caused of 12h deadline limitation. My reasons for this:</p>\n<ul>\n<li>I've submit that logic and observe no change in public LB, like yours</li>\n<li>Tomograms from public/private test seem to be produced/belong privately to a lab, with good and unified data collecting/preprocessing techniques, so such simple errors seem not very likely</li>\n<li>Host had train models, visualize, label, review the tomogram again. If I were the host, I will fix this before release the competition dataset instead of forcing Kaggler to \"probe/guess the dataset\"</li>\n</ul>\n<p>But, now I'm making late submissions to verify these as well.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3217406,
              "author_name": "tatamikenn",
              "author_url": "",
              "post_date": "06/05/2025 01:56:57",
              "content": "<blockquote>\n  <p>I've submit that logic and observe no change in public LB, like yours</p>\n</blockquote>\n<p>I observed similar result in the public LB. So this correction was not applied in my final sub. However, not sure in private LB.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3217413,
                  "author_name": "mohammad2012191",
                  "author_url": "",
                  "post_date": "06/05/2025 02:05:12",
                  "content": "<p>Our private lb scores honestly looks very random. I don't see clear correlation, or why we shaked down.<br>\nSome random ensembles produced much better results than others without clear reason. But the only sub that has clearly a different pattern was the one with the anomalies fix added. That's why i believe it was the reason.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3217416,
                      "author_name": "mohammad2012191",
                      "author_url": "",
                      "post_date": "06/05/2025 02:07:19",
                      "content": "<p>The difference is huge. I am talking about smth like this:<br>\nPublic: 0.841, private 0.785 (w/o)<br>\nPublic: 0.841, private 0.818 (with fix)</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3217426,
                          "author_name": "tatamikenn",
                          "author_url": "",
                          "post_date": "06/05/2025 02:12:18",
                          "content": "<p><a href=\"https://www.kaggle.com/mohammad2012191\" target=\"_blank\">@mohammad2012191</a> thank you for sharing the result. It is noticeable gain.<br>\nMaybe test tomogram also contains this anomaly, but I think we need more submissions, considering high randomness in this competition.</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3187031": "## Summary\n\nI released the corrected dataset of pixel anomalies found on the @brendanartley's external dataset.\n\nlink: https://www.kaggle.com/datasets/tatamikenn/byu-cryoet-dataset-with-pixel-anomalies-corrected/data\n\nSome of the original data (from the CryoET Data Portal) appear to have been incorrectly quantized, resulting in pixel values ranging from -128 to 127 — as if they were stored as `int8`, even though the expected data range should be from 0 to 255 like `uint8`.\nWhich causes significant corruption of source image (see the pictures in visualization section).\n\nThe key correction applied is as follows:\n\n```python\nx = x.astype(\"uint8\").astype(\"float32\")\n```\n\nFor more detail, you can refer to the following notebooks:\n- EDA: https://www.kaggle.com/code/tatamikenn/byu-eda-on-the-irregular-data\n- Process: https://www.kaggle.com/code/tatamikenn/byu-download-irregular-data/notebook\n\n## Visualization\n\n**Before correction**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F8d2166da33e933acf89e5acc9c865b74%2FScreenshot%202025-04-25%20at%2022.04.25.png?generation=1745586325626721&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1c29e40e7d196d4c741b0118d2031752%2FScreenshot%202025-04-25%20at%2022.04.47.png?generation=1745586341084858&alt=media)\n\n**After correction**:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fbc8601a58e8092323221848400cbfc03%2FScreenshot%202025-04-25%20at%2022.04.36.png?generation=1745586366776285&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fb7da5f6b1117cc4de4b7510b9e379a29%2FScreenshot%202025-04-25%20at%2022.04.57.png?generation=1745586379171915&alt=media)",
    "3187038": "Thanks @tatamikenn!",
    "3187056": "Thanks man- there seem to be a few blackened images like this in the original training set as well? Will the same correction work there?",
    "3187090": "Thank you for the helpful information.",
    "3187178": "thanks man god job",
    "3187364": "Thank you and great work!\nIn your dataset in incorrect labels, there are tomogram names and z_slice values for those tomograms, are only THOSE slices incorrectly quantized? \nI would generally expect not just one , but nearby slices to be incorrectly quantized too....",
    "3187385": "Don't care much of this column. These are slices which exist annotations. I added it to check the image around motors.",
    "3188959": "Thanks for sharings it helps alot",
    "3198930": "Hi Thanks for the insights!\n\nBut I have a quesiton. Let's see we will implement a data normalization for preprocessing. \nI.e.  Normalize slice data using the 2nd and 98th percentiles and Normalized image in the range [0, 255]. In such case, I think this irregularity will not be an issue any more. Is it?\n\nBest regards\nLeo",
    "3217394": "I think fixing the anomalies of the test was key to get stable score in private. We didn't do that in our final subs, and unfortunately shaked down significantly. It is sad how this was not really represented in public lb tomos leading to all our work get punished severely :(",
    "3217405": "I'm realized that I forget to add these logics to correct the wrong quantiled tomogram in submission code too, right before clicking \"Submit\" button for the last submission, but decided not gonna fix and still continue to submit caused of 12h deadline limitation. My reasons for this:\n- I've submit that logic and observe no change in public LB, like yours\n- Tomograms from public/private test seem to be produced/belong privately to a lab, with good and unified data collecting/preprocessing techniques, so such simple errors seem not very likely\n- Host had train models, visualize, label, review the tomogram again. If I were the host, I will fix this before release the competition dataset instead of forcing Kaggler to \"probe/guess the dataset\"\n\nBut, now I'm making late submissions to verify these as well.",
    "3217406": "> I've submit that logic and observe no change in public LB, like yours\n\nI observed similar result in the public LB. So this correction was not applied in my final sub. However, not sure in private LB.",
    "3217410": "I believe correction by quantization is not robust way of fixing this anomaly because it depends on the choice of the quantiles. Note that the absolute error value is at most 127.",
    "3217413": "Our private lb scores honestly looks very random. I don't see clear correlation, or why we shaked down.\nSome random ensembles produced much better results than others without clear reason. But the only sub that has clearly a different pattern was the one with the anomalies fix added. That's why i believe it was the reason.",
    "3217416": "The difference is huge. I am talking about smth like this:\nPublic: 0.841, private 0.785 (w/o)\nPublic: 0.841, private 0.818 (with fix)",
    "3217426": "mohammad2012191 thank you for sharing the result. It is noticeable gain.\nMaybe test tomogram also contains this anomaly, but I think we need more submissions, considering high randomness in this competition."
  },
  "source": "meta"
}