{
  "id": 575769,
  "title": "Preprint for Our Project!",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575769",
  "author_name": "",
  "post_date": "2025-04-30T20:41:18.074074800Z",
  "votes": 18,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Hi everyone,</p>\n<p>We’re excited to share that a preprint describing the background and methods of this competition is now available on bioRxiv! </p>\n<p>You can check it out here: <a href=\"https://www.biorxiv.org/content/10.1101/2025.04.23.650258v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.04.23.650258v1</a></p>\n<p>The manuscript covers the biological motivation, data collection and annotation processes, as well as some preliminary analysis and benchmarks. We hope it offers deeper context for those curious about the origins of the dataset and the goals of the challenge.</p>\n<p>As always, we welcome feedback (especially if you happen to spot any errors). In fact, we've already caught one ourselves: the number of tomograms (as represented in <code>train_labels.csv</code>) was recently updated. There are now 648 tomograms total:</p>\n<ul>\n<li>362 contain motor/motors</li>\n<li>286 do not</li>\n</ul>\n<p>We’ll be updating the preprint to reflect this correction and any others that come up. While this isn’t a call for formal peer review, we do appreciate the collective eyes of this thoughtful community. If something seems off, feel free to let us know!</p>\n<p>Thanks for your continued participation, we’re thrilled to see what you all uncover in this dataset!</p>\n<p>— The BYU Team</p>",
  "messages": [
    {
      "id": "3190563",
      "postDate": "04/30/2025 20:41:18",
      "content": "<p>Hi everyone,</p>\n<p>We’re excited to share that a preprint describing the background and methods of this competition is now available on bioRxiv! </p>\n<p>You can check it out here: <a href=\"https://www.biorxiv.org/content/10.1101/2025.04.23.650258v1\" target=\"_blank\">https://www.biorxiv.org/content/10.1101/2025.04.23.650258v1</a></p>\n<p>The manuscript covers the biological motivation, data collection and annotation processes, as well as some preliminary analysis and benchmarks. We hope it offers deeper context for those curious about the origins of the dataset and the goals of the challenge.</p>\n<p>As always, we welcome feedback (especially if you happen to spot any errors). In fact, we've already caught one ourselves: the number of tomograms (as represented in <code>train_labels.csv</code>) was recently updated. There are now 648 tomograms total:</p>\n<ul>\n<li>362 contain motor/motors</li>\n<li>286 do not</li>\n</ul>\n<p>We’ll be updating the preprint to reflect this correction and any others that come up. While this isn’t a call for formal peer review, we do appreciate the collective eyes of this thoughtful community. If something seems off, feel free to let us know!</p>\n<p>Thanks for your continued participation, we’re thrilled to see what you all uncover in this dataset!</p>\n<p>— The BYU Team</p>",
      "rawMarkdown": "Hi everyone,\n\nWe’re excited to share that a preprint describing the background and methods of this competition is now available on bioRxiv! \n\nYou can check it out here: https://www.biorxiv.org/content/10.1101/2025.04.23.650258v1\n\nThe manuscript covers the biological motivation, data collection and annotation processes, as well as some preliminary analysis and benchmarks. We hope it offers deeper context for those curious about the origins of the dataset and the goals of the challenge.\n\nAs always, we welcome feedback (especially if you happen to spot any errors). In fact, we've already caught one ourselves: the number of tomograms (as represented in `train_labels.csv`) was recently updated. There are now 648 tomograms total:\n- 362 contain motor/motors\n- 286 do not\n\nWe’ll be updating the preprint to reflect this correction and any others that come up. While this isn’t a call for formal peer review, we do appreciate the collective eyes of this thoughtful community. If something seems off, feel free to let us know!\n\nThanks for your continued participation, we’re thrilled to see what you all uncover in this dataset!\n\n &mdash; The BYU Team",
      "votes": null
    },
    {
      "id": "3191418",
      "postDate": "05/01/2025 17:48:33",
      "content": "<p>Thank you for sharing the preprint. I will read it properly next week and might have more comments, but here are some quick observations.</p>\n<p>It seems like you are quoting your own model scores in this paper, which is fine. However, that is likely very misleading as to what is possible given how far ahead the competitors are right now on the LB.</p>\n<p>When someone is among the co-authors ( <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> ), that eliminates the need to mention them in acknowledgements. The authorship is an acknowledgement on its own.</p>\n<p>Speaking of authorship, you may want to consider including <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for providing a large external dataset.</p>\n<p>For this dataset to be a widely used benchmark, it is paramount that there are no obvious errors in it. There are several tomograms that are classified as having zero motors, while it is easy to show that they do. See an example below for <code>tomo_47ac94</code>. This is where the AI truly shines: the best models should be able to identify some motors that were missed by humans.</p>\n<p><img src=\"https://i.ibb.co/HfRSTxgq/tomo-47ac94-slice-0085-false.jpg\" alt=\"false negative\"></p>",
      "rawMarkdown": "Thank you for sharing the preprint. I will read it properly next week and might have more comments, but here are some quick observations.\n\nIt seems like you are quoting your own model scores in this paper, which is fine. However, that is likely very misleading as to what is possible given how far ahead the competitors are right now on the LB.\n\nWhen someone is among the co-authors ( @inversion ), that eliminates the need to mention them in acknowledgements. The authorship is an acknowledgement on its own.\n\nSpeaking of authorship, you may want to consider including @brendanartley for providing a large external dataset.\n\nFor this dataset to be a widely used benchmark, it is paramount that there are no obvious errors in it. There are several tomograms that are classified as having zero motors, while it is easy to show that they do. See an example below for `tomo_47ac94`. This is where the AI truly shines: the best models should be able to identify some motors that were missed by humans.\n\n![false negative](https://i.ibb.co/HfRSTxgq/tomo-47ac94-slice-0085-false.jpg)",
      "votes": null
    },
    {
      "id": "3197300",
      "postDate": "05/08/2025 03:16:34",
      "content": "<p>Thank you for providing the preprint. I plan to write a paper about my solution after the competition ends and submit it to a journal. May I ask if the full hidden test set will be released after the competition closes? Thank you!</p>",
      "rawMarkdown": "Thank you for providing the preprint. I plan to write a paper about my solution after the competition ends and submit it to a journal. May I ask if the full hidden test set will be released after the competition closes? Thank you!",
      "votes": null
    },
    {
      "id": "3198270",
      "postDate": "05/09/2025 08:12:34",
      "content": "<p>Thank you for sharing. This competition holds deep personal meaning for me. Last year, I was infected with H. pylori and suffered internal bleeding. I’ll never forget the moment I fainted from the severe blood loss—it was a truly life-altering experience.</p>\n<p>I didn’t intend to be sentimental, but I wanted to express my gratitude for this challenge. It has taught me so much, and I genuinely hope that, together, we can one day develop methods to control bacteria to such an extent that we see a real reduction in annual deaths.</p>\n<p>After all, every long journey begins with a small step—and I believe this is one of them. </p>",
      "rawMarkdown": "Thank you for sharing. This competition holds deep personal meaning for me. Last year, I was infected with H. pylori and suffered internal bleeding. I’ll never forget the moment I fainted from the severe blood loss—it was a truly life-altering experience.\n\nI didn’t intend to be sentimental, but I wanted to express my gratitude for this challenge. It has taught me so much, and I genuinely hope that, together, we can one day develop methods to control bacteria to such an extent that we see a real reduction in annual deaths.\n\nAfter all, every long journey begins with a small step—and I believe this is one of them.",
      "votes": null
    },
    {
      "id": "3199256",
      "postDate": "05/10/2025 17:34:26",
      "content": "<p>Thanks for sharing this preprint.</p>",
      "rawMarkdown": "Thanks for sharing this preprint.",
      "votes": null
    },
    {
      "id": "3202115",
      "postDate": "05/14/2025 20:59:38",
      "content": "<p>Interesting observations <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, it's also the case with mislabels for tomos which are marked with 1 motor, where there's obviously more present. </p>\n<p>Slice 84 in tomo_c649f8 (not labelled): <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F4c6620b4a28b9045d67a7cbbca787787%2FScreenshot%202025-05-14%20235057.png?generation=1747255874843839&amp;alt=media\" alt=\"\"><br>\nSlice 146 in tomo_c649f8 (labelled):<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F3b5ae15adcc3381e1edd31e40f4380a0%2FScreenshot%202025-05-14%20235240.png?generation=1747255974038921&amp;alt=media\" alt=\"\"><br>\nI was analyzing some very confident FNs (this one was not within the 1000 Angstroms range of the motor at 146, per the labels, so result turns to -1,-1,-1, thus becoming an FN), and our model picked up on both motors but since the requirement is of at most 1 prediction per tomo, the one at slice 84 was picked by our aggregation algorithm. </p>\n<p>How can we be sure that this is not occurring in the test set as well? Makes it hard to make decisions without confidence in the data.</p>",
      "rawMarkdown": "Interesting observations @tilii7, it's also the case with mislabels for tomos which are marked with 1 motor, where there's obviously more present. \n\nSlice 84 in tomo_c649f8 (not labelled): ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F4c6620b4a28b9045d67a7cbbca787787%2FScreenshot%202025-05-14%20235057.png?generation=1747255874843839&alt=media)\nSlice 146 in tomo_c649f8 (labelled):![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F3b5ae15adcc3381e1edd31e40f4380a0%2FScreenshot%202025-05-14%20235240.png?generation=1747255974038921&alt=media)\nI was analyzing some very confident FNs (this one was not within the 1000 Angstroms range of the motor at 146, per the labels, so result turns to -1,-1,-1, thus becoming an FN), and our model picked up on both motors but since the requirement is of at most 1 prediction per tomo, the one at slice 84 was picked by our aggregation algorithm. \n\nHow can we be sure that this is not occurring in the test set as well? Makes it hard to make decisions without confidence in the data.",
      "votes": null
    },
    {
      "id": "3202123",
      "postDate": "05/14/2025 21:22:07",
      "content": "<p>To me this seems like the same motor, but the slice on top is more convincing.</p>",
      "rawMarkdown": "To me this seems like the same motor, but the slice on top is more convincing.",
      "votes": null
    },
    {
      "id": "3202126",
      "postDate": "05/14/2025 21:34:52",
      "content": "<p>Thanks for pointing that out, I've included a side by side so it's easier to grasp the XY coords. I've shifted the second motor to slice 154, where the flagellum associated with the second motor is more visible. Cycling through Napari makes it more obvious.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2Fe65736b2026d33c6e394f4801d666aae%2Fsidebyside.jpg?generation=1747258386905784&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thanks for pointing that out, I've included a side by side so it's easier to grasp the XY coords. I've shifted the second motor to slice 154, where the flagellum associated with the second motor is more visible. Cycling through Napari makes it more obvious.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2Fe65736b2026d33c6e394f4801d666aae%2Fsidebyside.jpg?generation=1747258386905784&alt=media)",
      "votes": null
    },
    {
      "id": "3204515",
      "postDate": "05/18/2025 12:38:57",
      "content": "<p>Thanks, for sharing this preprint.</p>",
      "rawMarkdown": "Thanks, for sharing this preprint.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3191418,
      "author_name": "tilii7",
      "author_url": "",
      "post_date": "05/01/2025 17:48:33",
      "content": "<p>Thank you for sharing the preprint. I will read it properly next week and might have more comments, but here are some quick observations.</p>\n<p>It seems like you are quoting your own model scores in this paper, which is fine. However, that is likely very misleading as to what is possible given how far ahead the competitors are right now on the LB.</p>\n<p>When someone is among the co-authors ( <a href=\"https://www.kaggle.com/inversion\" target=\"_blank\">@inversion</a> ), that eliminates the need to mention them in acknowledgements. The authorship is an acknowledgement on its own.</p>\n<p>Speaking of authorship, you may want to consider including <a href=\"https://www.kaggle.com/brendanartley\" target=\"_blank\">@brendanartley</a> for providing a large external dataset.</p>\n<p>For this dataset to be a widely used benchmark, it is paramount that there are no obvious errors in it. There are several tomograms that are classified as having zero motors, while it is easy to show that they do. See an example below for <code>tomo_47ac94</code>. This is where the AI truly shines: the best models should be able to identify some motors that were missed by humans.</p>\n<p><img src=\"https://i.ibb.co/HfRSTxgq/tomo-47ac94-slice-0085-false.jpg\" alt=\"false negative\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 3202115,
          "author_name": "andreizamfir",
          "author_url": "",
          "post_date": "05/14/2025 20:59:38",
          "content": "<p>Interesting observations <a href=\"https://www.kaggle.com/tilii7\" target=\"_blank\">@tilii7</a>, it's also the case with mislabels for tomos which are marked with 1 motor, where there's obviously more present. </p>\n<p>Slice 84 in tomo_c649f8 (not labelled): <img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F4c6620b4a28b9045d67a7cbbca787787%2FScreenshot%202025-05-14%20235057.png?generation=1747255874843839&amp;alt=media\" alt=\"\"><br>\nSlice 146 in tomo_c649f8 (labelled):<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F3b5ae15adcc3381e1edd31e40f4380a0%2FScreenshot%202025-05-14%20235240.png?generation=1747255974038921&amp;alt=media\" alt=\"\"><br>\nI was analyzing some very confident FNs (this one was not within the 1000 Angstroms range of the motor at 146, per the labels, so result turns to -1,-1,-1, thus becoming an FN), and our model picked up on both motors but since the requirement is of at most 1 prediction per tomo, the one at slice 84 was picked by our aggregation algorithm. </p>\n<p>How can we be sure that this is not occurring in the test set as well? Makes it hard to make decisions without confidence in the data.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3202123,
              "author_name": "tilii7",
              "author_url": "",
              "post_date": "05/14/2025 21:22:07",
              "content": "<p>To me this seems like the same motor, but the slice on top is more convincing.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3202126,
                  "author_name": "andreizamfir",
                  "author_url": "",
                  "post_date": "05/14/2025 21:34:52",
                  "content": "<p>Thanks for pointing that out, I've included a side by side so it's easier to grasp the XY coords. I've shifted the second motor to slice 154, where the flagellum associated with the second motor is more visible. Cycling through Napari makes it more obvious.</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2Fe65736b2026d33c6e394f4801d666aae%2Fsidebyside.jpg?generation=1747258386905784&amp;alt=media\" alt=\"\"></p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3197300,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "05/08/2025 03:16:34",
      "content": "<p>Thank you for providing the preprint. I plan to write a paper about my solution after the competition ends and submit it to a journal. May I ask if the full hidden test set will be released after the competition closes? Thank you!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3198270,
      "author_name": "edonisraci",
      "author_url": "",
      "post_date": "05/09/2025 08:12:34",
      "content": "<p>Thank you for sharing. This competition holds deep personal meaning for me. Last year, I was infected with H. pylori and suffered internal bleeding. I’ll never forget the moment I fainted from the severe blood loss—it was a truly life-altering experience.</p>\n<p>I didn’t intend to be sentimental, but I wanted to express my gratitude for this challenge. It has taught me so much, and I genuinely hope that, together, we can one day develop methods to control bacteria to such an extent that we see a real reduction in annual deaths.</p>\n<p>After all, every long journey begins with a small step—and I believe this is one of them. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3199256,
      "author_name": "khushikyad001",
      "author_url": "",
      "post_date": "05/10/2025 17:34:26",
      "content": "<p>Thanks for sharing this preprint.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3204515,
      "author_name": "elhassenselahy",
      "author_url": "",
      "post_date": "05/18/2025 12:38:57",
      "content": "<p>Thanks, for sharing this preprint.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3190563": "Hi everyone,\n\nWe’re excited to share that a preprint describing the background and methods of this competition is now available on bioRxiv! \n\nYou can check it out here: https://www.biorxiv.org/content/10.1101/2025.04.23.650258v1\n\nThe manuscript covers the biological motivation, data collection and annotation processes, as well as some preliminary analysis and benchmarks. We hope it offers deeper context for those curious about the origins of the dataset and the goals of the challenge.\n\nAs always, we welcome feedback (especially if you happen to spot any errors). In fact, we've already caught one ourselves: the number of tomograms (as represented in `train_labels.csv`) was recently updated. There are now 648 tomograms total:\n- 362 contain motor/motors\n- 286 do not\n\nWe’ll be updating the preprint to reflect this correction and any others that come up. While this isn’t a call for formal peer review, we do appreciate the collective eyes of this thoughtful community. If something seems off, feel free to let us know!\n\nThanks for your continued participation, we’re thrilled to see what you all uncover in this dataset!\n\n &mdash; The BYU Team",
    "3191418": "Thank you for sharing the preprint. I will read it properly next week and might have more comments, but here are some quick observations.\n\nIt seems like you are quoting your own model scores in this paper, which is fine. However, that is likely very misleading as to what is possible given how far ahead the competitors are right now on the LB.\n\nWhen someone is among the co-authors ( @inversion ), that eliminates the need to mention them in acknowledgements. The authorship is an acknowledgement on its own.\n\nSpeaking of authorship, you may want to consider including @brendanartley for providing a large external dataset.\n\nFor this dataset to be a widely used benchmark, it is paramount that there are no obvious errors in it. There are several tomograms that are classified as having zero motors, while it is easy to show that they do. See an example below for `tomo_47ac94`. This is where the AI truly shines: the best models should be able to identify some motors that were missed by humans.\n\n![false negative](https://i.ibb.co/HfRSTxgq/tomo-47ac94-slice-0085-false.jpg)",
    "3197300": "Thank you for providing the preprint. I plan to write a paper about my solution after the competition ends and submit it to a journal. May I ask if the full hidden test set will be released after the competition closes? Thank you!",
    "3198270": "Thank you for sharing. This competition holds deep personal meaning for me. Last year, I was infected with H. pylori and suffered internal bleeding. I’ll never forget the moment I fainted from the severe blood loss—it was a truly life-altering experience.\n\nI didn’t intend to be sentimental, but I wanted to express my gratitude for this challenge. It has taught me so much, and I genuinely hope that, together, we can one day develop methods to control bacteria to such an extent that we see a real reduction in annual deaths.\n\nAfter all, every long journey begins with a small step—and I believe this is one of them.",
    "3199256": "Thanks for sharing this preprint.",
    "3202115": "Interesting observations @tilii7, it's also the case with mislabels for tomos which are marked with 1 motor, where there's obviously more present. \n\nSlice 84 in tomo_c649f8 (not labelled): ![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F4c6620b4a28b9045d67a7cbbca787787%2FScreenshot%202025-05-14%20235057.png?generation=1747255874843839&alt=media)\nSlice 146 in tomo_c649f8 (labelled):![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2F3b5ae15adcc3381e1edd31e40f4380a0%2FScreenshot%202025-05-14%20235240.png?generation=1747255974038921&alt=media)\nI was analyzing some very confident FNs (this one was not within the 1000 Angstroms range of the motor at 146, per the labels, so result turns to -1,-1,-1, thus becoming an FN), and our model picked up on both motors but since the requirement is of at most 1 prediction per tomo, the one at slice 84 was picked by our aggregation algorithm. \n\nHow can we be sure that this is not occurring in the test set as well? Makes it hard to make decisions without confidence in the data.",
    "3202123": "To me this seems like the same motor, but the slice on top is more convincing.",
    "3202126": "Thanks for pointing that out, I've included a side by side so it's easier to grasp the XY coords. I've shifted the second motor to slice 154, where the flagellum associated with the second motor is more visible. Cycling through Napari makes it more obvious.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F20416701%2Fe65736b2026d33c6e394f4801d666aae%2Fsidebyside.jpg?generation=1747258386905784&alt=media)",
    "3204515": "Thanks, for sharing this preprint."
  },
  "source": "meta"
}