{
  "id": 582948,
  "title": "Post-Competition Benchmark Dataset",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/582948",
  "author_name": "",
  "post_date": "2025-06-03T20:35:46.736383700Z",
  "votes": 8,
  "comment_count": 4,
  "views": 0,
  "content": "<p>With the competition ending in a day, we're currently working on getting everything in order to wrap things up. One of our main tasks is to clean up the noise in the annotations and post all the data publicly on Kaggle. Similar to the leaderboard rescore, we are using the top submissions to thoroughly verify the test dataset and make it as clean as possible. </p>\n<p>Note that this dataset is only for public community use and won't be used to decide final standings on the leaderboard. We believe this to be the best decision for the community as the current noisy (but not too noisy due to the rescore) dataset is the foundation for everyone's work in this competition and we don't want to invalidate the competitive nature of the competition. </p>\n<p>But the reason I'm making this post is that we have no reasonable way to clean the train dataset besides waiting for your code to be published and retraining your models on the test set in order to clean the train. But we don't need to reinvent the wheel necessarily. If anybody has their own cleaned version of the train dataset, we'd love to see your work and use it to make the community dataset. We won't release it until after the competition so don't worry about losing your competitive advantage. Anybody whose work we use in making the dataset will be attributed in our final paper. </p>\n<p>If anybody is interested, please reach out! <a href=\"mailto:andrewjdarley@gmail.com\">andrewjdarley@gmail.com</a></p>",
  "messages": [
    {
      "id": "3216617",
      "postDate": "06/03/2025 20:35:46",
      "content": "<p>With the competition ending in a day, we're currently working on getting everything in order to wrap things up. One of our main tasks is to clean up the noise in the annotations and post all the data publicly on Kaggle. Similar to the leaderboard rescore, we are using the top submissions to thoroughly verify the test dataset and make it as clean as possible. </p>\n<p>Note that this dataset is only for public community use and won't be used to decide final standings on the leaderboard. We believe this to be the best decision for the community as the current noisy (but not too noisy due to the rescore) dataset is the foundation for everyone's work in this competition and we don't want to invalidate the competitive nature of the competition. </p>\n<p>But the reason I'm making this post is that we have no reasonable way to clean the train dataset besides waiting for your code to be published and retraining your models on the test set in order to clean the train. But we don't need to reinvent the wheel necessarily. If anybody has their own cleaned version of the train dataset, we'd love to see your work and use it to make the community dataset. We won't release it until after the competition so don't worry about losing your competitive advantage. Anybody whose work we use in making the dataset will be attributed in our final paper. </p>\n<p>If anybody is interested, please reach out! <a href=\"mailto:andrewjdarley@gmail.com\">andrewjdarley@gmail.com</a></p>",
      "rawMarkdown": "With the competition ending in a day, we're currently working on getting everything in order to wrap things up. One of our main tasks is to clean up the noise in the annotations and post all the data publicly on Kaggle. Similar to the leaderboard rescore, we are using the top submissions to thoroughly verify the test dataset and make it as clean as possible. \n\nNote that this dataset is only for public community use and won't be used to decide final standings on the leaderboard. We believe this to be the best decision for the community as the current noisy (but not too noisy due to the rescore) dataset is the foundation for everyone's work in this competition and we don't want to invalidate the competitive nature of the competition. \n\nBut the reason I'm making this post is that we have no reasonable way to clean the train dataset besides waiting for your code to be published and retraining your models on the test set in order to clean the train. But we don't need to reinvent the wheel necessarily. If anybody has their own cleaned version of the train dataset, we'd love to see your work and use it to make the community dataset. We won't release it until after the competition so don't worry about losing your competitive advantage. Anybody whose work we use in making the dataset will be attributed in our final paper. \n\nIf anybody is interested, please reach out! andrewjdarley@gmail.com",
      "votes": null
    },
    {
      "id": "3217038",
      "postDate": "06/04/2025 13:00:15",
      "content": "<p>Does that mean that the test set has more flawed annotations that you are aware of but that will not be corrected for the final scoring? That is unfortunate as this will benefit teams who performed leaderboard overfitting and will not reward the team that actually produces the most correct prediction. But I of course understand where you are coming from, especially after some people got upset after the rescore :-/<br>\nWe do have some corrections for the training data, bartleys data as well as annotations for a few additional publicly available tomograms that we can share. Would take us some time though as they are not in the original coordinate system and I am on vacation for 2+ weeks</p>",
      "rawMarkdown": "Does that mean that the test set has more flawed annotations that you are aware of but that will not be corrected for the final scoring? That is unfortunate as this will benefit teams who performed leaderboard overfitting and will not reward the team that actually produces the most correct prediction. But I of course understand where you are coming from, especially after some people got upset after the rescore :-/\nWe do have some corrections for the training data, bartleys data as well as annotations for a few additional publicly available tomograms that we can share. Would take us some time though as they are not in the original coordinate system and I am on vacation for 2+ weeks",
      "votes": null
    },
    {
      "id": "3217064",
      "postDate": "06/04/2025 13:42:14",
      "content": "<p>We would appreciate any extra data you have. Thanks! </p>\n<p>I understand your concern regarding our decision to leave the test set as is for the leaderboard. But the public test set needs to remain representative of the test set as a whole in order for there to be any sort of useful metric to build models this comp. This was advised to us by the Kaggle staff and is a common practice in many competitions.</p>",
      "rawMarkdown": "We would appreciate any extra data you have. Thanks! \n\nI understand your concern regarding our decision to leave the test set as is for the leaderboard. But the public test set needs to remain representative of the test set as a whole in order for there to be any sort of useful metric to build models this comp. This was advised to us by the Kaggle staff and is a common practice in many competitions.",
      "votes": null
    },
    {
      "id": "3217403",
      "postDate": "06/05/2025 01:49:17",
      "content": "<p>I found few anomaly samples both CryoET database and this competition's dataset.<br>\nI believe this anomaly is due to wrong process of quantization. I tried to correct wrongly quantized pixel value range, but obviously, it is better to fix before quantization is applied to avoid information loss. So I also recommend checking quantization code.<br>\nSee also the previous discussions:</p>\n<p><a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028</a><br>\n<a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579659\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579659</a></p>",
      "rawMarkdown": "I found few anomaly samples both CryoET database and this competition's dataset.\nI believe this anomaly is due to wrong process of quantization. I tried to correct wrongly quantized pixel value range, but obviously, it is better to fix before quantization is applied to avoid information loss. So I also recommend checking quantization code.\nSee also the previous discussions:\n\nhttps://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028\nhttps://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579659",
      "votes": null
    },
    {
      "id": "3222444",
      "postDate": "06/12/2025 07:45:09",
      "content": "<p>Interested </p>",
      "rawMarkdown": "Interested",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217038,
      "author_name": "fabianisensee",
      "author_url": "",
      "post_date": "06/04/2025 13:00:15",
      "content": "<p>Does that mean that the test set has more flawed annotations that you are aware of but that will not be corrected for the final scoring? That is unfortunate as this will benefit teams who performed leaderboard overfitting and will not reward the team that actually produces the most correct prediction. But I of course understand where you are coming from, especially after some people got upset after the rescore :-/<br>\nWe do have some corrections for the training data, bartleys data as well as annotations for a few additional publicly available tomograms that we can share. Would take us some time though as they are not in the original coordinate system and I am on vacation for 2+ weeks</p>",
      "votes": null,
      "replies": [
        {
          "id": 3217064,
          "author_name": "andrewjdarley",
          "author_url": "",
          "post_date": "06/04/2025 13:42:14",
          "content": "<p>We would appreciate any extra data you have. Thanks! </p>\n<p>I understand your concern regarding our decision to leave the test set as is for the leaderboard. But the public test set needs to remain representative of the test set as a whole in order for there to be any sort of useful metric to build models this comp. This was advised to us by the Kaggle staff and is a common practice in many competitions.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3217403,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "06/05/2025 01:49:17",
      "content": "<p>I found few anomaly samples both CryoET database and this competition's dataset.<br>\nI believe this anomaly is due to wrong process of quantization. I tried to correct wrongly quantized pixel value range, but obviously, it is better to fix before quantization is applied to avoid information loss. So I also recommend checking quantization code.<br>\nSee also the previous discussions:</p>\n<p><a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028</a><br>\n<a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579659\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579659</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3222444,
      "author_name": "msinghbe22thaparedu",
      "author_url": "",
      "post_date": "06/12/2025 07:45:09",
      "content": "<p>Interested </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3216617": "With the competition ending in a day, we're currently working on getting everything in order to wrap things up. One of our main tasks is to clean up the noise in the annotations and post all the data publicly on Kaggle. Similar to the leaderboard rescore, we are using the top submissions to thoroughly verify the test dataset and make it as clean as possible. \n\nNote that this dataset is only for public community use and won't be used to decide final standings on the leaderboard. We believe this to be the best decision for the community as the current noisy (but not too noisy due to the rescore) dataset is the foundation for everyone's work in this competition and we don't want to invalidate the competitive nature of the competition. \n\nBut the reason I'm making this post is that we have no reasonable way to clean the train dataset besides waiting for your code to be published and retraining your models on the test set in order to clean the train. But we don't need to reinvent the wheel necessarily. If anybody has their own cleaned version of the train dataset, we'd love to see your work and use it to make the community dataset. We won't release it until after the competition so don't worry about losing your competitive advantage. Anybody whose work we use in making the dataset will be attributed in our final paper. \n\nIf anybody is interested, please reach out! andrewjdarley@gmail.com",
    "3217038": "Does that mean that the test set has more flawed annotations that you are aware of but that will not be corrected for the final scoring? That is unfortunate as this will benefit teams who performed leaderboard overfitting and will not reward the team that actually produces the most correct prediction. But I of course understand where you are coming from, especially after some people got upset after the rescore :-/\nWe do have some corrections for the training data, bartleys data as well as annotations for a few additional publicly available tomograms that we can share. Would take us some time though as they are not in the original coordinate system and I am on vacation for 2+ weeks",
    "3217064": "We would appreciate any extra data you have. Thanks! \n\nI understand your concern regarding our decision to leave the test set as is for the leaderboard. But the public test set needs to remain representative of the test set as a whole in order for there to be any sort of useful metric to build models this comp. This was advised to us by the Kaggle staff and is a common practice in many competitions.",
    "3217403": "I found few anomaly samples both CryoET database and this competition's dataset.\nI believe this anomaly is due to wrong process of quantization. I tried to correct wrongly quantized pixel value range, but obviously, it is better to fix before quantization is applied to avoid information loss. So I also recommend checking quantization code.\nSee also the previous discussions:\n\nhttps://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575028\nhttps://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/579659",
    "3222444": "Interested"
  },
  "source": "meta"
}