{
  "id": 583244,
  "title": "easy solution of LB 0.80745: using only YOLO and the original dataset && summarize the approach to this competition ",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/writeups/0-3-cvpr-easy-solution-of-lb-0-80745-using-only-yo",
  "author_name": "",
  "post_date": "2025-06-11T13:45:46.980Z",
  "votes": 6,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Thanks to this competition and the numerous public discussion information, we learned a lot. This is our first time participating in an image detection competition, so our insights may be limited and we write down the entire exploration process and hope everyone will discuss it together.</p>\n<h3>Conclusion</h3>\n<p>Feature engineering is the first, the model is the one that is fixed ! </p>\n<h3>Data preprocessing</h3>\n<p>The preprocessing process mentioned next is to fix the model and train hyper-parameters, that is, the pre-trained YOLOv8m model and epoch are set to 40.<br>\nFirst of all, everyone should be clear that this is a medical image. Its typical feature is that the data distribution is sparse, most of which are negative samples, and the signal-to-noise ratio is very low. So we cannot directly use the feature engineering of natural images, at least their  hyper-parameters cannot be copied directly.  <br>\nWe have tried many methods in feature engineering. First, we directly use all slices of all datasets to send them into the model (any YOLO or UNet version), and the result is very poor, and the LB may be only 0.4-0.5.<br>\nWe analyzed the data and found that the id ratio of motors is 1 to 1, and we found that about 90% of the positive samples have only 1 motor, and the pixel proportion is very small, which forced us to regard the dataset as a serious imbalance of positive and negative samples. The first thing we thought of was to use focal loss, and the target score is F_2, which pays more attention to Recall, so we abandoned a lot of negative samples.<br>\nHowever, using focal loss for this task is not a good choice, the LB score is only 0.2 , and we didn't find the reason. Abandoning negative samples does bring a lot of gain. At the beginning, we abandoned negative samples and got a score of about 0.65, but there are more and more FPs. Later, we made a lot of sampling attempts around 2D slices, but the effect was stuck around 0.7.<br>\nTherefore, we consider the 2.5D input and hope to make up for the missing continuous information in 2D by stacking the inputs in succession slices. Next, we still refer to the 2D sampling process and made the following attempts:<br>\nPositive sample:</p>\n<ul>\n<li>For each motor-existing position (x,y,z), slices of the z±TRUST range are collected, each with a target box.<br>\nNegative sample sampling (two categories):</li>\n</ul>\n<p>To control category balance, the following two negative samples were introduced:</p>\n<ul>\n<li>Type A: In the same tomo, avoid the motor slices and randomly take NEG_PER_TOMO targetless image from the remaining slices.</li>\n<li>Type B: From tomogram without any motors, NEG_EMPTY_TOMO images are randomly taken.<br>\nAnd we adopted a hierarchical sampling strategy based on the number of motors and ensure that the number of training sets and verification sets is consistent. The above operation has brought our score to 0.74.  Here again, all the above data processing is carried out under the premise of fixing the training model and parameters. We have limited resources, and we use pre-trained  v8m and the epoch is set to 40.  By the way, we saw that the scores of many people in the public list were above 0.8, so here we have reason to believe that we have not overfitted, and we did not take cross-validation, hahaha~ </li>\n</ul>\n<h3>Train and Infer</h3>\n<p>During the public discussion, many great grandmasters scoffed at YOLO. But I want to say that YOLO or other specific model such as 3D UNET is not the whole game, and YOLO has already achieved quite high results in various testing competitions, so why should we reject it?  We can do many feature projects that are suitable for tasks around YOLO.<br>\nWhen data enhancement, in addition to conventional enhancements, we try to add Gaussian noise and Poisson noise, as cryo-electron microscopy may carry these noises.  After determining the preprocessing process, we now try the v8m and v10l versions, with epoch set to 70.  When predicting, we adopted the model integration strategy, but it often timed out.   Now we find that there are many solutions in the public discussion, which also reminds beginners to read more discussions.  We have also adopted too multi-scale prediction strategies, but the effect is not as good as a single model.   Finally we submitted the model with wbf and h-flip + v-flip.</p>\n<h3>Some feelings</h3>\n<ul>\n<li>Through a lot of public discussion, we can obtain a lot of new information, such as data sets newly marked by grandmasters , data correction tools, and some speed up  strategy. </li>\n<li>We can fully explore the upper limit of a model's feature engineering. It may not be cost-effective to worry about the model first.</li>\n</ul>\n<h3>TO DO</h3>\n<p>We are  going to  summarize the approach of the top-ranked teams here ,  that is, we want to try and combine the top-ranked team approaches, and we will show them here  ..   </p>",
  "messages": [
    {
      "id": "3217971",
      "postDate": "06/05/2025 15:53:57",
      "content": "<p>Thanks to this competition and the numerous public discussion information, we learned a lot. This is our first time participating in an image detection competition, so our insights may be limited and we write down the entire exploration process and hope everyone will discuss it together.</p>\n<h3>Conclusion</h3>\n<p>Feature engineering is the first, the model is the one that is fixed ! </p>\n<h3>Data preprocessing</h3>\n<p>The preprocessing process mentioned next is to fix the model and train hyper-parameters, that is, the pre-trained YOLOv8m model and epoch are set to 40.<br>\nFirst of all, everyone should be clear that this is a medical image. Its typical feature is that the data distribution is sparse, most of which are negative samples, and the signal-to-noise ratio is very low. So we cannot directly use the feature engineering of natural images, at least their  hyper-parameters cannot be copied directly.  <br>\nWe have tried many methods in feature engineering. First, we directly use all slices of all datasets to send them into the model (any YOLO or UNet version), and the result is very poor, and the LB may be only 0.4-0.5.<br>\nWe analyzed the data and found that the id ratio of motors is 1 to 1, and we found that about 90% of the positive samples have only 1 motor, and the pixel proportion is very small, which forced us to regard the dataset as a serious imbalance of positive and negative samples. The first thing we thought of was to use focal loss, and the target score is F_2, which pays more attention to Recall, so we abandoned a lot of negative samples.<br>\nHowever, using focal loss for this task is not a good choice, the LB score is only 0.2 , and we didn't find the reason. Abandoning negative samples does bring a lot of gain. At the beginning, we abandoned negative samples and got a score of about 0.65, but there are more and more FPs. Later, we made a lot of sampling attempts around 2D slices, but the effect was stuck around 0.7.<br>\nTherefore, we consider the 2.5D input and hope to make up for the missing continuous information in 2D by stacking the inputs in succession slices. Next, we still refer to the 2D sampling process and made the following attempts:<br>\nPositive sample:</p>\n<ul>\n<li>For each motor-existing position (x,y,z), slices of the z±TRUST range are collected, each with a target box.<br>\nNegative sample sampling (two categories):</li>\n</ul>\n<p>To control category balance, the following two negative samples were introduced:</p>\n<ul>\n<li>Type A: In the same tomo, avoid the motor slices and randomly take NEG_PER_TOMO targetless image from the remaining slices.</li>\n<li>Type B: From tomogram without any motors, NEG_EMPTY_TOMO images are randomly taken.<br>\nAnd we adopted a hierarchical sampling strategy based on the number of motors and ensure that the number of training sets and verification sets is consistent. The above operation has brought our score to 0.74.  Here again, all the above data processing is carried out under the premise of fixing the training model and parameters. We have limited resources, and we use pre-trained  v8m and the epoch is set to 40.  By the way, we saw that the scores of many people in the public list were above 0.8, so here we have reason to believe that we have not overfitted, and we did not take cross-validation, hahaha~ </li>\n</ul>\n<h3>Train and Infer</h3>\n<p>During the public discussion, many great grandmasters scoffed at YOLO. But I want to say that YOLO or other specific model such as 3D UNET is not the whole game, and YOLO has already achieved quite high results in various testing competitions, so why should we reject it?  We can do many feature projects that are suitable for tasks around YOLO.<br>\nWhen data enhancement, in addition to conventional enhancements, we try to add Gaussian noise and Poisson noise, as cryo-electron microscopy may carry these noises.  After determining the preprocessing process, we now try the v8m and v10l versions, with epoch set to 70.  When predicting, we adopted the model integration strategy, but it often timed out.   Now we find that there are many solutions in the public discussion, which also reminds beginners to read more discussions.  We have also adopted too multi-scale prediction strategies, but the effect is not as good as a single model.   Finally we submitted the model with wbf and h-flip + v-flip.</p>\n<h3>Some feelings</h3>\n<ul>\n<li>Through a lot of public discussion, we can obtain a lot of new information, such as data sets newly marked by grandmasters , data correction tools, and some speed up  strategy. </li>\n<li>We can fully explore the upper limit of a model's feature engineering. It may not be cost-effective to worry about the model first.</li>\n</ul>\n<h3>TO DO</h3>\n<p>We are  going to  summarize the approach of the top-ranked teams here ,  that is, we want to try and combine the top-ranked team approaches, and we will show them here  ..   </p>",
      "rawMarkdown": "Thanks to this competition and the numerous public discussion information, we learned a lot. This is our first time participating in an image detection competition, so our insights may be limited and we write down the entire exploration process and hope everyone will discuss it together.\n### Conclusion\nFeature engineering is the first, the model is the one that is fixed ! \n### Data preprocessing\nThe preprocessing process mentioned next is to fix the model and train hyper-parameters, that is, the pre-trained YOLOv8m model and epoch are set to 40.\nFirst of all, everyone should be clear that this is a medical image. Its typical feature is that the data distribution is sparse, most of which are negative samples, and the signal-to-noise ratio is very low. So we cannot directly use the feature engineering of natural images, at least their  hyper-parameters cannot be copied directly.  \nWe have tried many methods in feature engineering. First, we directly use all slices of all datasets to send them into the model (any YOLO or UNet version), and the result is very poor, and the LB may be only 0.4-0.5.\nWe analyzed the data and found that the id ratio of motors is 1 to 1, and we found that about 90% of the positive samples have only 1 motor, and the pixel proportion is very small, which forced us to regard the dataset as a serious imbalance of positive and negative samples. The first thing we thought of was to use focal loss, and the target score is F_2, which pays more attention to Recall, so we abandoned a lot of negative samples.\nHowever, using focal loss for this task is not a good choice, the LB score is only 0.2 , and we didn't find the reason. Abandoning negative samples does bring a lot of gain. At the beginning, we abandoned negative samples and got a score of about 0.65, but there are more and more FPs. Later, we made a lot of sampling attempts around 2D slices, but the effect was stuck around 0.7.\nTherefore, we consider the 2.5D input and hope to make up for the missing continuous information in 2D by stacking the inputs in succession slices. Next, we still refer to the 2D sampling process and made the following attempts:\nPositive sample:\n- For each motor-existing position (x,y,z), slices of the z±TRUST range are collected, each with a target box.\nNegative sample sampling (two categories):\n\nTo control category balance, the following two negative samples were introduced:\n- Type A: In the same tomo, avoid the motor slices and randomly take NEG_PER_TOMO targetless image from the remaining slices.\n- Type B: From tomogram without any motors, NEG_EMPTY_TOMO images are randomly taken.\nAnd we adopted a hierarchical sampling strategy based on the number of motors and ensure that the number of training sets and verification sets is consistent. The above operation has brought our score to 0.74.  Here again, all the above data processing is carried out under the premise of fixing the training model and parameters. We have limited resources, and we use pre-trained  v8m and the epoch is set to 40.  By the way, we saw that the scores of many people in the public list were above 0.8, so here we have reason to believe that we have not overfitted, and we did not take cross-validation, hahaha~ \n \n### Train and Infer\nDuring the public discussion, many great grandmasters scoffed at YOLO. But I want to say that YOLO or other specific model such as 3D UNET is not the whole game, and YOLO has already achieved quite high results in various testing competitions, so why should we reject it?  We can do many feature projects that are suitable for tasks around YOLO.\nWhen data enhancement, in addition to conventional enhancements, we try to add Gaussian noise and Poisson noise, as cryo-electron microscopy may carry these noises.  After determining the preprocessing process, we now try the v8m and v10l versions, with epoch set to 70.  When predicting, we adopted the model integration strategy, but it often timed out.   Now we find that there are many solutions in the public discussion, which also reminds beginners to read more discussions.  We have also adopted too multi-scale prediction strategies, but the effect is not as good as a single model.   Finally we submitted the model with wbf and h-flip + v-flip.\n\n### Some feelings\n- Through a lot of public discussion, we can obtain a lot of new information, such as data sets newly marked by grandmasters , data correction tools, and some speed up  strategy. \n- We can fully explore the upper limit of a model's feature engineering. It may not be cost-effective to worry about the model first.\n\n### TO DO\nWe are  going to  summarize the approach of the top-ranked teams here ,  that is, we want to try and combine the top-ranked team approaches, and we will show them here  ..",
      "votes": null
    },
    {
      "id": "3217976",
      "postDate": "06/05/2025 16:02:44",
      "content": "<p>Hope you all enjoy your Kaggle journey!</p>",
      "rawMarkdown": "Hope you all enjoy your Kaggle journey!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3217976,
      "author_name": "wym2024",
      "author_url": "",
      "post_date": "06/05/2025 16:02:44",
      "content": "<p>Hope you all enjoy your Kaggle journey!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217971": "Thanks to this competition and the numerous public discussion information, we learned a lot. This is our first time participating in an image detection competition, so our insights may be limited and we write down the entire exploration process and hope everyone will discuss it together.\n### Conclusion\nFeature engineering is the first, the model is the one that is fixed ! \n### Data preprocessing\nThe preprocessing process mentioned next is to fix the model and train hyper-parameters, that is, the pre-trained YOLOv8m model and epoch are set to 40.\nFirst of all, everyone should be clear that this is a medical image. Its typical feature is that the data distribution is sparse, most of which are negative samples, and the signal-to-noise ratio is very low. So we cannot directly use the feature engineering of natural images, at least their  hyper-parameters cannot be copied directly.  \nWe have tried many methods in feature engineering. First, we directly use all slices of all datasets to send them into the model (any YOLO or UNet version), and the result is very poor, and the LB may be only 0.4-0.5.\nWe analyzed the data and found that the id ratio of motors is 1 to 1, and we found that about 90% of the positive samples have only 1 motor, and the pixel proportion is very small, which forced us to regard the dataset as a serious imbalance of positive and negative samples. The first thing we thought of was to use focal loss, and the target score is F_2, which pays more attention to Recall, so we abandoned a lot of negative samples.\nHowever, using focal loss for this task is not a good choice, the LB score is only 0.2 , and we didn't find the reason. Abandoning negative samples does bring a lot of gain. At the beginning, we abandoned negative samples and got a score of about 0.65, but there are more and more FPs. Later, we made a lot of sampling attempts around 2D slices, but the effect was stuck around 0.7.\nTherefore, we consider the 2.5D input and hope to make up for the missing continuous information in 2D by stacking the inputs in succession slices. Next, we still refer to the 2D sampling process and made the following attempts:\nPositive sample:\n- For each motor-existing position (x,y,z), slices of the z±TRUST range are collected, each with a target box.\nNegative sample sampling (two categories):\n\nTo control category balance, the following two negative samples were introduced:\n- Type A: In the same tomo, avoid the motor slices and randomly take NEG_PER_TOMO targetless image from the remaining slices.\n- Type B: From tomogram without any motors, NEG_EMPTY_TOMO images are randomly taken.\nAnd we adopted a hierarchical sampling strategy based on the number of motors and ensure that the number of training sets and verification sets is consistent. The above operation has brought our score to 0.74.  Here again, all the above data processing is carried out under the premise of fixing the training model and parameters. We have limited resources, and we use pre-trained  v8m and the epoch is set to 40.  By the way, we saw that the scores of many people in the public list were above 0.8, so here we have reason to believe that we have not overfitted, and we did not take cross-validation, hahaha~ \n \n### Train and Infer\nDuring the public discussion, many great grandmasters scoffed at YOLO. But I want to say that YOLO or other specific model such as 3D UNET is not the whole game, and YOLO has already achieved quite high results in various testing competitions, so why should we reject it?  We can do many feature projects that are suitable for tasks around YOLO.\nWhen data enhancement, in addition to conventional enhancements, we try to add Gaussian noise and Poisson noise, as cryo-electron microscopy may carry these noises.  After determining the preprocessing process, we now try the v8m and v10l versions, with epoch set to 70.  When predicting, we adopted the model integration strategy, but it often timed out.   Now we find that there are many solutions in the public discussion, which also reminds beginners to read more discussions.  We have also adopted too multi-scale prediction strategies, but the effect is not as good as a single model.   Finally we submitted the model with wbf and h-flip + v-flip.\n\n### Some feelings\n- Through a lot of public discussion, we can obtain a lot of new information, such as data sets newly marked by grandmasters , data correction tools, and some speed up  strategy. \n- We can fully explore the upper limit of a model's feature engineering. It may not be cost-effective to worry about the model first.\n\n### TO DO\nWe are  going to  summarize the approach of the top-ranked teams here ,  that is, we want to try and combine the top-ranked team approaches, and we will show them here  ..",
    "3217976": "Hope you all enjoy your Kaggle journey!"
  },
  "source": "meta"
}