{
  "id": 576756,
  "title": "YOLO - Does best.pt checkpoint = best LB? ",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/576756",
  "author_name": "",
  "post_date": "2025-05-06T22:23:24.487633300Z",
  "votes": 31,
  "comment_count": 35,
  "views": 0,
  "content": "<p>I've shared some early results here at <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575799\" target=\"_blank\">@eikyou discussion</a></p>\n<p>But it started getting a bit long, so I’m opening another discussion.</p>\n<p>For the past two weeks, I’ve been training various YOLO models and always submitted the <code>best.pt</code> checkpoint, assuming that Ultralytics' default behavior  <a href=\"https://github.com/ultralytics/ultralytics/issues/14137\" target=\"_blank\">0.1 × mAP@0.5 + 0.9 × mAP@0.5:0.95</a> would consistently result in the best public LB result.</p>\n<p>I was wrong.</p>\n<p>In one of my recent runs, I got the best.pt checkpoints at different epochs (not 100% about the epochs):</p>\n<ul>\n<li>Around ~ 28 → LB: <strong>0.816</strong></li>\n<li>Around ~ 40 → LB: <strong>0.752</strong></li>\n<li>Around ~ 50 → LB: <strong>0.744</strong></li>\n</ul>\n<p>This made me realize that some of my past experiments might’ve had better LB potential if I had just submitted a different checkpoint. So I started investigating what metric I should trust to pick the best epoch… but I haven’t found a clear pattern yet.</p>\n<p>Here’s some of my current results (each line a model/checkpoint) trained on 4/5 and evaluated on 1/5 of only positive images :</p>\n<pre><code>, precision, recall, f1, mAP50, mAP50-\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n</code></pre>\n<p>I also computed the correlation between fbeta and other metrics:</p>\n<pre><code>:   -.  \n:       .  \n:           .  \n:       -.  \n-:    -.  \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fed3a4581733209200f5202ce75cb3158%2FcoefMatrix.jfif?generation=1746570087097235&amp;alt=media\" alt=\"Correlation Matrix\"></p>\n<p>Still experimenting, but I thought this might be helpful if others are blindly trusting <code>best.pt</code> like I was.</p>",
  "messages": [
    {
      "id": "3195326",
      "postDate": "05/06/2025 22:23:24",
      "content": "<p>I've shared some early results here at <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575799\" target=\"_blank\">@eikyou discussion</a></p>\n<p>But it started getting a bit long, so I’m opening another discussion.</p>\n<p>For the past two weeks, I’ve been training various YOLO models and always submitted the <code>best.pt</code> checkpoint, assuming that Ultralytics' default behavior  <a href=\"https://github.com/ultralytics/ultralytics/issues/14137\" target=\"_blank\">0.1 × mAP@0.5 + 0.9 × mAP@0.5:0.95</a> would consistently result in the best public LB result.</p>\n<p>I was wrong.</p>\n<p>In one of my recent runs, I got the best.pt checkpoints at different epochs (not 100% about the epochs):</p>\n<ul>\n<li>Around ~ 28 → LB: <strong>0.816</strong></li>\n<li>Around ~ 40 → LB: <strong>0.752</strong></li>\n<li>Around ~ 50 → LB: <strong>0.744</strong></li>\n</ul>\n<p>This made me realize that some of my past experiments might’ve had better LB potential if I had just submitted a different checkpoint. So I started investigating what metric I should trust to pick the best epoch… but I haven’t found a clear pattern yet.</p>\n<p>Here’s some of my current results (each line a model/checkpoint) trained on 4/5 and evaluated on 1/5 of only positive images :</p>\n<pre><code>, precision, recall, f1, mAP50, mAP50-\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n., ., ., ., ., .\n</code></pre>\n<p>I also computed the correlation between fbeta and other metrics:</p>\n<pre><code>:   -.  \n:       .  \n:           .  \n:       -.  \n-:    -.  \n</code></pre>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fed3a4581733209200f5202ce75cb3158%2FcoefMatrix.jfif?generation=1746570087097235&amp;alt=media\" alt=\"Correlation Matrix\"></p>\n<p>Still experimenting, but I thought this might be helpful if others are blindly trusting <code>best.pt</code> like I was.</p>",
      "rawMarkdown": "I've shared some early results here at [@eikyou discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575799)\n\nBut it started getting a bit long, so I’m opening another discussion.\n\nFor the past two weeks, I’ve been training various YOLO models and always submitted the `best.pt` checkpoint, assuming that Ultralytics' default behavior  [0.1 × mAP@0.5 + 0.9 × mAP@0.5:0.95](https://github.com/ultralytics/ultralytics/issues/14137) would consistently result in the best public LB result.\n\nI was wrong.\n\nIn one of my recent runs, I got the best.pt checkpoints at different epochs (not 100% about the epochs):\n\n* Around ~ 28 → LB: **0.816**\n* Around ~ 40 → LB: **0.752**\n* Around ~ 50 → LB: **0.744**\n\nThis made me realize that some of my past experiments might’ve had better LB potential if I had just submitted a different checkpoint. So I started investigating what metric I should trust to pick the best epoch... but I haven’t found a clear pattern yet.\n\nHere’s some of my current results (each line a model/checkpoint) trained on 4/5 and evaluated on 1/5 of only positive images :\n\n```\nfbeta, precision, recall, f1, mAP50, mAP50-95\n0.816, 0.9651, 0.9594, 0.9623, 0.9812, 0.7316\n0.788, 0.9797, 0.9506, 0.9649, 0.9826, 0.6589\n0.749, 0.9696, 0.9572, 0.9633, 0.9843, 0.7268\n0.706, 0.9758, 0.9629, 0.9693, 0.9806, 0.7020\n0.744, 0.9747, 0.9542, 0.9644, 0.9888, 0.7621\n0.792, 0.9723, 0.9770, 0.9747, 0.9822, 0.7349\n0.714, 0.9782, 0.9523, 0.9651, 0.9818, 0.7161\n0.747, 0.9663, 0.9400, 0.9530, 0.9859, 0.7301\n```\n\nI also computed the correlation between fbeta and other metrics:\n\n```\nprecision:   -0.397  \nrecall:       0.258  \nf1:           0.055  \nmAP50:       -0.116  \nmAP50-95:    -0.038  \n```\n![Correlation Matrix](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fed3a4581733209200f5202ce75cb3158%2FcoefMatrix.jfif?generation=1746570087097235&alt=media)\n\nStill experimenting, but I thought this might be helpful if others are blindly trusting `best.pt` like I was.",
      "votes": null
    },
    {
      "id": "3195355",
      "postDate": "05/07/2025 00:01:38",
      "content": "<p>I have recently tried saving the model for each epoch with save_period=1 and selecting the model for the smallest epoch of dfl-loss in addition to best.pt. Generally, but not always, the model with the smallest dfl-loss tends to have the higher LB.</p>",
      "rawMarkdown": "I have recently tried saving the model for each epoch with save_period=1 and selecting the model for the smallest epoch of dfl-loss in addition to best.pt. Generally, but not always, the model with the smallest dfl-loss tends to have the higher LB.",
      "votes": null
    },
    {
      "id": "3195364",
      "postDate": "05/07/2025 00:25:23",
      "content": "<p>can you trust this result not shake up in the end? That's my biggest concern for tunning yolo like that. </p>",
      "rawMarkdown": "can you trust this result not shake up in the end? That's my biggest concern for tunning yolo like that.",
      "votes": null
    },
    {
      "id": "3195365",
      "postDate": "05/07/2025 00:25:56",
      "content": "<p>How did u split the data into train and val? Do the slices from the same tomo appear both in train and val or do you do split based on tomos?</p>",
      "rawMarkdown": "How did u split the data into train and val? Do the slices from the same tomo appear both in train and val or do you do split based on tomos?",
      "votes": null
    },
    {
      "id": "3195372",
      "postDate": "05/07/2025 00:48:43",
      "content": "<p>I split the data by tomogram. I'm currently only using tomograms that contain one motor, which gave me 250 tomograms for training and 63 for validation</p>",
      "rawMarkdown": "I split the data by tomogram. I'm currently only using tomograms that contain one motor, which gave me 250 tomograms for training and 63 for validation",
      "votes": null
    },
    {
      "id": "3195379",
      "postDate": "05/07/2025 00:55:12",
      "content": "<p>Yeah, that’s definitely something I’m worried about too. I’m planning to add Bartley external dataset as a validation set (just need to handle the preprocessing to match mine) to see if I can find any consistent correlation between my best models on the LB and some metric.</p>\n<p>Really hoping it helps, because things feel too random right now, and it makes testing new ideas pretty hard</p>",
      "rawMarkdown": "Yeah, that’s definitely something I’m worried about too. I’m planning to add Bartley external dataset as a validation set (just need to handle the preprocessing to match mine) to see if I can find any consistent correlation between my best models on the LB and some metric.\n\nReally hoping it helps, because things feel too random right now, and it makes testing new ideas pretty hard",
      "votes": null
    },
    {
      "id": "3195384",
      "postDate": "05/07/2025 00:59:05",
      "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> one thing you can do is to do augmentation on the validation set. The result is quite stable to cv and lb.</p>",
      "rawMarkdown": "sersasj one thing you can do is to do augmentation on the validation set. The result is quite stable to cv and lb.",
      "votes": null
    },
    {
      "id": "3196434",
      "postDate": "05/07/2025 03:34:14",
      "content": "<p>how do you ensemble those two models?</p>",
      "rawMarkdown": "how do you ensemble those two models?",
      "votes": null
    },
    {
      "id": "3196452",
      "postDate": "05/07/2025 04:09:08",
      "content": "<p>No ensemble will be made.<br>\nSubmit best.pt and dfl_loss_best.pt separately to observe LB.</p>",
      "rawMarkdown": "No ensemble will be made.\nSubmit best.pt and dfl_loss_best.pt separately to observe LB.",
      "votes": null
    },
    {
      "id": "3196457",
      "postDate": "05/07/2025 04:15:55",
      "content": "<p>Could I ask about your training data? Do you use samples where the number of motors is greater than 0, or only those where the number of motors equals 1?</p>",
      "rawMarkdown": "Could I ask about your training data? Do you use samples where the number of motors is greater than 0, or only those where the number of motors equals 1?",
      "votes": null
    },
    {
      "id": "3196462",
      "postDate": "05/07/2025 04:21:38",
      "content": "<p>Could I ask which version of Ultralytics you're using? I'm using v8.3.111, and in my experiments, different versions of the codebase produce different validation results on the same dataset.</p>",
      "rawMarkdown": "Could I ask which version of Ultralytics you're using? I'm using v8.3.111, and in my experiments, different versions of the codebase produce different validation results on the same dataset.",
      "votes": null
    },
    {
      "id": "3196467",
      "postDate": "05/07/2025 04:29:13",
      "content": "<p>The number of motors in the sample is greater than or equal to 1. I used more than 1 because the LB score was not good with only a sample with a number of motors equals 1.</p>",
      "rawMarkdown": "The number of motors in the sample is greater than or equal to 1. I used more than 1 because the LB score was not good with only a sample with a number of motors equals 1.",
      "votes": null
    },
    {
      "id": "3196644",
      "postDate": "05/07/2025 08:57:39",
      "content": "<p>For us as well, the same model exhibits significant randomness across different runs, and this issue still persists to this day.</p>",
      "rawMarkdown": "For us as well, the same model exhibits significant randomness across different runs, and this issue still persists to this day.",
      "votes": null
    },
    {
      "id": "3196734",
      "postDate": "05/07/2025 11:28:57",
      "content": "<p><a href=\"https://www.kaggle.com/fangsionfang\" target=\"_blank\">@fangsionfang</a> I'm using version 8.3.121. I actually cloned the Ultralytics repo and have been modifying a few things, changing the activation function, loss, trying out timm backbones with YOLO head, etc. </p>",
      "rawMarkdown": "fangsionfang I'm using version 8.3.121. I actually cloned the Ultralytics repo and have been modifying a few things, changing the activation function, loss, trying out timm backbones with YOLO head, etc.",
      "votes": null
    },
    {
      "id": "3197039",
      "postDate": "05/07/2025 18:50:53",
      "content": "<p>Massive +1 to this — blindly trusting best.pt cost me some decent submissions too.</p>\n<p>Ultralytics’ default metric might be great for general object detection, but when LB scoring diverges (like heavy class imbalance or FBeta-weighted targets), mAP just doesn’t cut it as a reliable proxy.</p>\n<p>I’ve started saving checkpoints every few epochs and scoring them manually on a holdout LB-style fold using the comp metric (e.g., fbeta), then picking the actual best. Bit more work, but worth it.</p>\n<p>Appreciate you sharing the correlation matrix — confirms that intuition beats default metrics sometimes 🔍</p>\n<p>Anyone tried using LB proxy models to predict the best checkpoint?</p>",
      "rawMarkdown": "Massive +1 to this — blindly trusting best.pt cost me some decent submissions too.\n\nUltralytics’ default metric might be great for general object detection, but when LB scoring diverges (like heavy class imbalance or FBeta-weighted targets), mAP just doesn’t cut it as a reliable proxy.\n\nI’ve started saving checkpoints every few epochs and scoring them manually on a holdout LB-style fold using the comp metric (e.g., fbeta), then picking the actual best. Bit more work, but worth it.\n\nAppreciate you sharing the correlation matrix — confirms that intuition beats default metrics sometimes 🔍\n\nAnyone tried using LB proxy models to predict the best checkpoint?",
      "votes": null
    },
    {
      "id": "3197135",
      "postDate": "05/07/2025 20:53:03",
      "content": "<p>Tried submitting dfl_loss_best with different ckpts today but result was much worse.</p>",
      "rawMarkdown": "Tried submitting dfl_loss_best with different ckpts today but result was much worse.",
      "votes": null
    },
    {
      "id": "3197235",
      "postDate": "05/08/2025 00:05:06",
      "content": "<p>My experience is that dfl_best_loss models sometimes score lower than best.pt, but it has never been extremely bad. try optimizing epoch's model. The optimization code can be found at <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/577949\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/577949</a></p>",
      "rawMarkdown": "My experience is that dfl_best_loss models sometimes score lower than best.pt, but it has never been extremely bad. try optimizing epoch's model. The optimization code can be found at https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/577949",
      "votes": null
    },
    {
      "id": "3197553",
      "postDate": "05/08/2025 09:24:47",
      "content": "<p>I dont know if what i do is the most robust method but i have splited the train data into train_val and test. Then i measure everything in my test set which i treat as LB (thats why i have very minimal submissions and a low score yet). That way i hope that i will manage to do the hyper param tunning without overfitting to the public LB. Anybody thoughts about this kind of methodology? </p>",
      "rawMarkdown": "I dont know if what i do is the most robust method but i have splited the train data into train_val and test. Then i measure everything in my test set which i treat as LB (thats why i have very minimal submissions and a low score yet). That way i hope that i will manage to do the hyper param tunning without overfitting to the public LB. Anybody thoughts about this kind of methodology?",
      "votes": null
    },
    {
      "id": "3197570",
      "postDate": "05/08/2025 10:05:49",
      "content": "<p>Sound trustful enough on condition there is enough data. Are u sure metrics won't be too noisy due to small amount of data?</p>",
      "rawMarkdown": "Sound trustful enough on condition there is enough data. Are u sure metrics won't be too noisy due to small amount of data?",
      "votes": null
    },
    {
      "id": "3197585",
      "postDate": "05/08/2025 10:23:45",
      "content": "<p>No i am not sure, i just go with this strategy for now</p>",
      "rawMarkdown": "No i am not sure, i just go with this strategy for now",
      "votes": null
    },
    {
      "id": "3198133",
      "postDate": "05/09/2025 05:10:03",
      "content": "<p>Thank you very much. <br>\nNormally, augmentation of validation data is not done, but I found it useful when the data is small, as in this case. Certainly the dfl_loss of the validation set in training is stable.</p>",
      "rawMarkdown": "Thank you very much. \nNormally, augmentation of validation data is not done, but I found it useful when the data is small, as in this case. Certainly the dfl_loss of the validation set in training is stable.",
      "votes": null
    },
    {
      "id": "3198166",
      "postDate": "05/09/2025 05:57:35",
      "content": "<p>Thank you for sharing your insights. I have some doubts about using the DFL loss as the main metric for model evaluation. Fundamentally, DFL loss functions similarly to bounding box regression loss. However, in this motor detection task, the critical requirement is precise center localization - bounding box accuracy is actually of secondary importance.</p>\n<p>In practice, model achieves satisfactory performance early in training (demonstrating good center-point detection and classification capability). The subsequent optimization might lead to overfitting through excessive refinement of box edges. So continuous optimization of dfl loss afterwards does not necessarily imply an improvement in model performance</p>\n<p>So I take it that when the dfl loss is lowest does it mean that the model is not necessarily the best?</p>",
      "rawMarkdown": "Thank you for sharing your insights. I have some doubts about using the DFL loss as the main metric for model evaluation. Fundamentally, DFL loss functions similarly to bounding box regression loss. However, in this motor detection task, the critical requirement is precise center localization - bounding box accuracy is actually of secondary importance.\n\nIn practice, model achieves satisfactory performance early in training (demonstrating good center-point detection and classification capability). The subsequent optimization might lead to overfitting through excessive refinement of box edges. So continuous optimization of dfl loss afterwards does not necessarily imply an improvement in model performance\n\nSo I take it that when the dfl loss is lowest does it mean that the model is not necessarily the best?",
      "votes": null
    },
    {
      "id": "3198181",
      "postDate": "05/09/2025 06:19:30",
      "content": "<p>Thank you for your input.<br>\nI don't consider the detection of the center position so strictly based on the evaluation index which is positive within 1000 Å. Since most of the Voxel spacing is 10-20 Å, I interpret a discrepancy of up to 50 pixels as OK (please correct me if I'm wrong). <br>\nSince Yolo is excellent, I believe that the motor decision (0 or 1) affects the score more than the exact position.</p>",
      "rawMarkdown": "Thank you for your input.\nI don't consider the detection of the center position so strictly based on the evaluation index which is positive within 1000 Å. Since most of the Voxel spacing is 10-20 Å, I interpret a discrepancy of up to 50 pixels as OK (please correct me if I'm wrong). \nSince Yolo is excellent, I believe that the motor decision (0 or 1) affects the score more than the exact position.",
      "votes": null
    },
    {
      "id": "3198188",
      "postDate": "05/09/2025 06:28:07",
      "content": "<p>Yes, improving the model's ability to distinguish between positive and negative samples is a crucial direction, while localization capability is indeed less critical in this context.</p>",
      "rawMarkdown": "Yes, improving the model's ability to distinguish between positive and negative samples is a crucial direction, while localization capability is indeed less critical in this context.",
      "votes": null
    },
    {
      "id": "3199008",
      "postDate": "05/10/2025 10:52:43",
      "content": "<p>you can add callback function to save dfl_loss_best model</p>",
      "rawMarkdown": "you can add callback function to save dfl_loss_best model",
      "votes": null
    },
    {
      "id": "3199061",
      "postDate": "05/10/2025 11:44:56",
      "content": "<p>hey if you don't mind telling me. can you please tell me how to reduce the false positives. i am using yolo11x with pretrained. its a different competition i am talking about in which the train images has only single object without any other background objects. but the test data has multiple images. i think this is confusing my model and its sometimes detecting the false objects with higher accuracy sometimes higher than actual images. can you give me a nice suggestion.</p>",
      "rawMarkdown": "hey if you don't mind telling me. can you please tell me how to reduce the false positives. i am using yolo11x with pretrained. its a different competition i am talking about in which the train images has only single object without any other background objects. but the test data has multiple images. i think this is confusing my model and its sometimes detecting the false objects with higher accuracy sometimes higher than actual images. can you give me a nice suggestion.",
      "votes": null
    },
    {
      "id": "3199067",
      "postDate": "05/10/2025 11:50:37",
      "content": "<p>Reducing the strength of augmentation.</p>",
      "rawMarkdown": "Reducing the strength of augmentation.",
      "votes": null
    },
    {
      "id": "3199245",
      "postDate": "05/10/2025 17:14:45",
      "content": "<p>Hi Tom, I am new to this field and participating in a competition with over 50 participants. To secure a top 3 position, should I focus on parameter tuning and augmentation, or would it be better to try advanced techniques like changing the YOLO backbone or model ensembling? The top participant achieved 97%, but they didn't share their approach. I would really appreciate your guidance on this.</p>",
      "rawMarkdown": "Hi Tom, I am new to this field and participating in a competition with over 50 participants. To secure a top 3 position, should I focus on parameter tuning and augmentation, or would it be better to try advanced techniques like changing the YOLO backbone or model ensembling? The top participant achieved 97%, but they didn't share their approach. I would really appreciate your guidance on this.",
      "votes": null
    },
    {
      "id": "3199347",
      "postDate": "05/10/2025 20:49:23",
      "content": "<p>good job👍</p>",
      "rawMarkdown": "good job👍",
      "votes": null
    },
    {
      "id": "3199626",
      "postDate": "05/11/2025 08:46:44",
      "content": "<p><a href=\"https://www.kaggle.com/mohanapavanbezawada\" target=\"_blank\">@mohanapavanbezawada</a> for begining you can take a look at the yolo notebook provided by host. I would say ensembling is challenging for object localization cuz you need to consider the perception of each model and merge them. You can refer the 1st solution in crytoET comp. Their ensembling method is great. Maybe I'll public my best yolo ensembling notebook later. You can copy and edit it. Changing backbone also works, but it needs engineering effort.</p>",
      "rawMarkdown": "mohanapavanbezawada for begining you can take a look at the yolo notebook provided by host. I would say ensembling is challenging for object localization cuz you need to consider the perception of each model and merge them. You can refer the 1st solution in crytoET comp. Their ensembling method is great. Maybe I'll public my best yolo ensembling notebook later. You can copy and edit it. Changing backbone also works, but it needs engineering effort.",
      "votes": null
    },
    {
      "id": "3199670",
      "postDate": "05/11/2025 10:10:58",
      "content": "<p>Have you noticed any specific patterns or conditions where the model with the lowest dfl-loss doesn't perform best on the leaderboard?</p>",
      "rawMarkdown": "Have you noticed any specific patterns or conditions where the model with the lowest dfl-loss doesn't perform best on the leaderboard?",
      "votes": null
    },
    {
      "id": "3199958",
      "postDate": "05/11/2025 18:48:29",
      "content": "<p>you can add back function</p>",
      "rawMarkdown": "you can add back function",
      "votes": null
    },
    {
      "id": "3200107",
      "postDate": "05/12/2025 04:41:42",
      "content": "<p>how to do it？</p>",
      "rawMarkdown": "how to do it？",
      "votes": null
    },
    {
      "id": "3200215",
      "postDate": "05/12/2025 09:06:44",
      "content": "<p>An example code:</p>\n<pre><code>best_dfl_loss = ()\n ():\n     best_dfl_loss\n    save_dir = \n\n    val_metrics = (trainer, , )\n\n    current_dfl_loss = \n     trainer.epoch == : \n         ()\n\n    current_dfl_loss = val_metrics[]\n\n     current_dfl_loss   :\n        ()\n         current_dfl_loss &lt; best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, )\n\n            ckpt = {\n                : trainer.epoch,\n                : , \n                : trainer.ema.updates  trainer.ema  ,\n                : trainer.optimizer.state_dict(),\n                : (trainer.args), \n                : best_dfl_loss\n            }\n\n             trainer.ema  (trainer.ema, ):\n                ckpt[] = trainer.ema.ema.()\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model\n            :\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model \n            torch.save(ckpt, ckpt_path)\n\n            ()\n\nmodel = YOLO()\n\nmodel.add_callback(, on_fit_epoch_end)\n\nmodel.train()\n</code></pre>",
      "rawMarkdown": "An example code:\n\n```\nbest_dfl_loss = float('inf')\ndef on_fit_epoch_end(trainer):\n    global best_dfl_loss\n    save_dir = 'run/train/weights'\n    \n    val_metrics = getattr(trainer, 'metrics', None)\n\n    current_dfl_loss = None\n    if trainer.epoch == 0: \n         print(f\"Debug: Available validation metric keys at epoch {trainer.epoch}: {list(val_metrics.keys())}\")\n\n    current_dfl_loss = val_metrics['val/dfl_loss']\n\n    if current_dfl_loss is not None:\n        print(f\"Epoch {trainer.epoch}: Current DFL Loss: {current_dfl_loss:.4f}, Best DFL Loss: {best_dfl_loss:.4f}\")\n        if current_dfl_loss < best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, f'best_dfl_loss_epoch_{trainer.epoch}_loss_{current_dfl_loss:.4f}.pt')\n\n            ckpt = {\n                'epoch': trainer.epoch,\n                'best_fitness': None, \n                'updates': trainer.ema.updates if trainer.ema else 0,\n                'optimizer': trainer.optimizer.state_dict(),\n                'train_args': vars(trainer.args), \n                'dfl_loss_for_best': best_dfl_loss\n            }\n    \n            if trainer.ema and hasattr(trainer.ema, 'ema'):\n                ckpt['ema'] = trainer.ema.ema.float()\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model\n            else:\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model \n            torch.save(ckpt, ckpt_path)\n\n            print(f\"Epoch {trainer.epoch+1}: Saved new best model (dfl_loss): {ckpt_path} with DFL Loss {best_dfl_loss:.4f}\")\n\nmodel = YOLO(\"yolov8n.pt\")\n\nmodel.add_callback(\"on_fit_epoch_end\", on_fit_epoch_end)\n\nmodel.train()\n```",
      "votes": null
    },
    {
      "id": "3202345",
      "postDate": "05/15/2025 08:45:07",
      "content": "<p>Below is a more robust implementation with corrected grammar and syntax:</p>\n<p>Main Modifications:<br>\nAutomatically save weights to the project directory.<br>\n'motor_detector' is used as the name in the model.train() function.</p>\n<p>At the end of training, the 'val/dfl_loss' metric does not exist, which raises a KeyError.<br>\nI handled this issue by adding a conditional check and a print statement to gracefully skip it.</p>\n<pre><code>best_dfl_loss = ()\n ():\n     best_dfl_loss\n     (trainer, )  (trainer.args, ):\n        save_dir = os.path.join(trainer.args.project, , )\n    :\n        save_dir = \n\n    os.makedirs(save_dir, exist_ok=)\n\n    val_metrics = (trainer, , )\n    current_dfl_loss = \n\n     trainer.epoch == :\n        ()\n\n     val_metrics    val_metrics:\n        current_dfl_loss = val_metrics[]\n\n        ()\n         current_dfl_loss &lt; best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, )\n\n            ckpt = {\n                : trainer.epoch,\n                : ,\n                : trainer.ema.updates  trainer.ema  ,\n                : trainer.optimizer.state_dict(),\n                : (trainer.args),\n                : best_dfl_loss\n            }\n\n             trainer.ema  (trainer.ema, ):\n                ckpt[] = trainer.ema.ema.()\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model\n            :\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model\n\n            torch.save(ckpt, ckpt_path)\n\n            ()\n    :\n        ()\n</code></pre>",
      "rawMarkdown": "Below is a more robust implementation with corrected grammar and syntax:\n\nMain Modifications:\nAutomatically save weights to the project directory.\n'motor_detector' is used as the name in the model.train() function.\n\nAt the end of training, the 'val/dfl_loss' metric does not exist, which raises a KeyError.\nI handled this issue by adding a conditional check and a print statement to gracefully skip it.\n\n\n```\nbest_dfl_loss = float('inf')\ndef on_fit_epoch_end(trainer):\n    global best_dfl_loss\n    if hasattr(trainer, 'args') and hasattr(trainer.args, 'project'):\n        save_dir = os.path.join(trainer.args.project, 'motor_detector', 'weights')\n    else:\n        save_dir = 'runs/train/weights'\n\n    os.makedirs(save_dir, exist_ok=True)\n\n    val_metrics = getattr(trainer, 'metrics', None)\n    current_dfl_loss = None\n\n    if trainer.epoch == 0:\n        print(f\"Debug: Available validation metric keys at epoch {trainer.epoch}: {list(val_metrics.keys()) if val_metrics else 'None'}\")\n\n    if val_metrics and 'val/dfl_loss' in val_metrics:\n        current_dfl_loss = val_metrics['val/dfl_loss']\n\n        print(f\"Epoch {trainer.epoch}: Current DFL Loss: {current_dfl_loss:.4f}, Best DFL Loss: {best_dfl_loss:.4f}\")\n        if current_dfl_loss < best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, f'best_dfl_loss_epoch_{trainer.epoch}_loss_{current_dfl_loss:.4f}.pt')\n\n            ckpt = {\n                'epoch': trainer.epoch,\n                'best_fitness': None,\n                'updates': trainer.ema.updates if trainer.ema else 0,\n                'optimizer': trainer.optimizer.state_dict(),\n                'train_args': vars(trainer.args),\n                'dfl_loss_for_best': best_dfl_loss\n            }\n\n            if trainer.ema and hasattr(trainer.ema, 'ema'):\n                ckpt['ema'] = trainer.ema.ema.float()\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model\n            else:\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model\n\n            torch.save(ckpt, ckpt_path)\n\n            print(f\"Epoch {trainer.epoch+1}: Saved new best model (dfl_loss): {ckpt_path} with DFL Loss {best_dfl_loss:.4f}\")\n    else:\n        print(f\"Epoch {trainer.epoch}: val/dfl_loss not found in metrics. Available keys: {list(val_metrics.keys()) if val_metrics else 'None'}\")\n```",
      "votes": null
    },
    {
      "id": "3202390",
      "postDate": "05/15/2025 09:50:13",
      "content": "<p>Thanks for sharing with us!</p>",
      "rawMarkdown": "Thanks for sharing with us!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3195355,
      "author_name": "minfuka",
      "author_url": "",
      "post_date": "05/07/2025 00:01:38",
      "content": "<p>I have recently tried saving the model for each epoch with save_period=1 and selecting the model for the smallest epoch of dfl-loss in addition to best.pt. Generally, but not always, the model with the smallest dfl-loss tends to have the higher LB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3196434,
          "author_name": "konohayui",
          "author_url": "",
          "post_date": "05/07/2025 03:34:14",
          "content": "<p>how do you ensemble those two models?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3196452,
              "author_name": "minfuka",
              "author_url": "",
              "post_date": "05/07/2025 04:09:08",
              "content": "<p>No ensemble will be made.<br>\nSubmit best.pt and dfl_loss_best.pt separately to observe LB.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3196457,
                  "author_name": "fangsionfang",
                  "author_url": "",
                  "post_date": "05/07/2025 04:15:55",
                  "content": "<p>Could I ask about your training data? Do you use samples where the number of motors is greater than 0, or only those where the number of motors equals 1?</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3196467,
                      "author_name": "minfuka",
                      "author_url": "",
                      "post_date": "05/07/2025 04:29:13",
                      "content": "<p>The number of motors in the sample is greater than or equal to 1. I used more than 1 because the LB score was not good with only a sample with a number of motors equals 1.</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3197135,
                          "author_name": "mohammad2012191",
                          "author_url": "",
                          "post_date": "05/07/2025 20:53:03",
                          "content": "<p>Tried submitting dfl_loss_best with different ckpts today but result was much worse.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3197235,
                              "author_name": "minfuka",
                              "author_url": "",
                              "post_date": "05/08/2025 00:05:06",
                              "content": "<p>My experience is that dfl_best_loss models sometimes score lower than best.pt, but it has never been extremely bad. try optimizing epoch's model. The optimization code can be found at <a href=\"https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/577949\" target=\"_blank\">https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/577949</a></p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3199008,
          "author_name": "i2nfinit3y",
          "author_url": "",
          "post_date": "05/10/2025 10:52:43",
          "content": "<p>you can add callback function to save dfl_loss_best model</p>",
          "votes": null,
          "replies": [
            {
              "id": 3200107,
              "author_name": "daifanhao",
              "author_url": "",
              "post_date": "05/12/2025 04:41:42",
              "content": "<p>how to do it？</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3200215,
                  "author_name": "i2nfinit3y",
                  "author_url": "",
                  "post_date": "05/12/2025 09:06:44",
                  "content": "<p>An example code:</p>\n<pre><code>best_dfl_loss = ()\n ():\n     best_dfl_loss\n    save_dir = \n\n    val_metrics = (trainer, , )\n\n    current_dfl_loss = \n     trainer.epoch == : \n         ()\n\n    current_dfl_loss = val_metrics[]\n\n     current_dfl_loss   :\n        ()\n         current_dfl_loss &lt; best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, )\n\n            ckpt = {\n                : trainer.epoch,\n                : , \n                : trainer.ema.updates  trainer.ema  ,\n                : trainer.optimizer.state_dict(),\n                : (trainer.args), \n                : best_dfl_loss\n            }\n\n             trainer.ema  (trainer.ema, ):\n                ckpt[] = trainer.ema.ema.()\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model\n            :\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model \n            torch.save(ckpt, ckpt_path)\n\n            ()\n\nmodel = YOLO()\n\nmodel.add_callback(, on_fit_epoch_end)\n\nmodel.train()\n</code></pre>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3202345,
                      "author_name": "seeingtimes",
                      "author_url": "",
                      "post_date": "05/15/2025 08:45:07",
                      "content": "<p>Below is a more robust implementation with corrected grammar and syntax:</p>\n<p>Main Modifications:<br>\nAutomatically save weights to the project directory.<br>\n'motor_detector' is used as the name in the model.train() function.</p>\n<p>At the end of training, the 'val/dfl_loss' metric does not exist, which raises a KeyError.<br>\nI handled this issue by adding a conditional check and a print statement to gracefully skip it.</p>\n<pre><code>best_dfl_loss = ()\n ():\n     best_dfl_loss\n     (trainer, )  (trainer.args, ):\n        save_dir = os.path.join(trainer.args.project, , )\n    :\n        save_dir = \n\n    os.makedirs(save_dir, exist_ok=)\n\n    val_metrics = (trainer, , )\n    current_dfl_loss = \n\n     trainer.epoch == :\n        ()\n\n     val_metrics    val_metrics:\n        current_dfl_loss = val_metrics[]\n\n        ()\n         current_dfl_loss &lt; best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, )\n\n            ckpt = {\n                : trainer.epoch,\n                : ,\n                : trainer.ema.updates  trainer.ema  ,\n                : trainer.optimizer.state_dict(),\n                : (trainer.args),\n                : best_dfl_loss\n            }\n\n             trainer.ema  (trainer.ema, ):\n                ckpt[] = trainer.ema.ema.()\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model\n            :\n                ckpt[] = deepcopy(trainer.model).half()  (trainer.model, nn.Module)  trainer.model\n\n            torch.save(ckpt, ckpt_path)\n\n            ()\n    :\n        ()\n</code></pre>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3202390,
                          "author_name": "towhid121",
                          "author_url": "",
                          "post_date": "05/15/2025 09:50:13",
                          "content": "<p>Thanks for sharing with us!</p>",
                          "votes": null,
                          "replies": []
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 3199670,
          "author_name": "towhid121",
          "author_url": "",
          "post_date": "05/11/2025 10:10:58",
          "content": "<p>Have you noticed any specific patterns or conditions where the model with the lowest dfl-loss doesn't perform best on the leaderboard?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3195364,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "05/07/2025 00:25:23",
      "content": "<p>can you trust this result not shake up in the end? That's my biggest concern for tunning yolo like that. </p>",
      "votes": null,
      "replies": [
        {
          "id": 3195379,
          "author_name": "sersasj",
          "author_url": "",
          "post_date": "05/07/2025 00:55:12",
          "content": "<p>Yeah, that’s definitely something I’m worried about too. I’m planning to add Bartley external dataset as a validation set (just need to handle the preprocessing to match mine) to see if I can find any consistent correlation between my best models on the LB and some metric.</p>\n<p>Really hoping it helps, because things feel too random right now, and it makes testing new ideas pretty hard</p>",
          "votes": null,
          "replies": [
            {
              "id": 3195384,
              "author_name": "tom99763",
              "author_url": "",
              "post_date": "05/07/2025 00:59:05",
              "content": "<p><a href=\"https://www.kaggle.com/sersasj\" target=\"_blank\">@sersasj</a> one thing you can do is to do augmentation on the validation set. The result is quite stable to cv and lb.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3198133,
                  "author_name": "minfuka",
                  "author_url": "",
                  "post_date": "05/09/2025 05:10:03",
                  "content": "<p>Thank you very much. <br>\nNormally, augmentation of validation data is not done, but I found it useful when the data is small, as in this case. Certainly the dfl_loss of the validation set in training is stable.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 3198166,
                      "author_name": "yyyy0201",
                      "author_url": "",
                      "post_date": "05/09/2025 05:57:35",
                      "content": "<p>Thank you for sharing your insights. I have some doubts about using the DFL loss as the main metric for model evaluation. Fundamentally, DFL loss functions similarly to bounding box regression loss. However, in this motor detection task, the critical requirement is precise center localization - bounding box accuracy is actually of secondary importance.</p>\n<p>In practice, model achieves satisfactory performance early in training (demonstrating good center-point detection and classification capability). The subsequent optimization might lead to overfitting through excessive refinement of box edges. So continuous optimization of dfl loss afterwards does not necessarily imply an improvement in model performance</p>\n<p>So I take it that when the dfl loss is lowest does it mean that the model is not necessarily the best?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 3198181,
                          "author_name": "minfuka",
                          "author_url": "",
                          "post_date": "05/09/2025 06:19:30",
                          "content": "<p>Thank you for your input.<br>\nI don't consider the detection of the center position so strictly based on the evaluation index which is positive within 1000 Å. Since most of the Voxel spacing is 10-20 Å, I interpret a discrepancy of up to 50 pixels as OK (please correct me if I'm wrong). <br>\nSince Yolo is excellent, I believe that the motor decision (0 or 1) affects the score more than the exact position.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 3198188,
                              "author_name": "yyyy0201",
                              "author_url": "",
                              "post_date": "05/09/2025 06:28:07",
                              "content": "<p>Yes, improving the model's ability to distinguish between positive and negative samples is a crucial direction, while localization capability is indeed less critical in this context.</p>",
                              "votes": null,
                              "replies": []
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3195365,
      "author_name": "eikyou",
      "author_url": "",
      "post_date": "05/07/2025 00:25:56",
      "content": "<p>How did u split the data into train and val? Do the slices from the same tomo appear both in train and val or do you do split based on tomos?</p>",
      "votes": null,
      "replies": [
        {
          "id": 3195372,
          "author_name": "sersasj",
          "author_url": "",
          "post_date": "05/07/2025 00:48:43",
          "content": "<p>I split the data by tomogram. I'm currently only using tomograms that contain one motor, which gave me 250 tomograms for training and 63 for validation</p>",
          "votes": null,
          "replies": [
            {
              "id": 3196462,
              "author_name": "fangsionfang",
              "author_url": "",
              "post_date": "05/07/2025 04:21:38",
              "content": "<p>Could I ask which version of Ultralytics you're using? I'm using v8.3.111, and in my experiments, different versions of the codebase produce different validation results on the same dataset.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3196734,
                  "author_name": "sersasj",
                  "author_url": "",
                  "post_date": "05/07/2025 11:28:57",
                  "content": "<p><a href=\"https://www.kaggle.com/fangsionfang\" target=\"_blank\">@fangsionfang</a> I'm using version 8.3.121. I actually cloned the Ultralytics repo and have been modifying a few things, changing the activation function, loss, trying out timm backbones with YOLO head, etc. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3196644,
      "author_name": "yyyy0201",
      "author_url": "",
      "post_date": "05/07/2025 08:57:39",
      "content": "<p>For us as well, the same model exhibits significant randomness across different runs, and this issue still persists to this day.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3197039,
      "author_name": "siddharth776",
      "author_url": "",
      "post_date": "05/07/2025 18:50:53",
      "content": "<p>Massive +1 to this — blindly trusting best.pt cost me some decent submissions too.</p>\n<p>Ultralytics’ default metric might be great for general object detection, but when LB scoring diverges (like heavy class imbalance or FBeta-weighted targets), mAP just doesn’t cut it as a reliable proxy.</p>\n<p>I’ve started saving checkpoints every few epochs and scoring them manually on a holdout LB-style fold using the comp metric (e.g., fbeta), then picking the actual best. Bit more work, but worth it.</p>\n<p>Appreciate you sharing the correlation matrix — confirms that intuition beats default metrics sometimes 🔍</p>\n<p>Anyone tried using LB proxy models to predict the best checkpoint?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3197553,
      "author_name": "vasileioscharatsidis",
      "author_url": "",
      "post_date": "05/08/2025 09:24:47",
      "content": "<p>I dont know if what i do is the most robust method but i have splited the train data into train_val and test. Then i measure everything in my test set which i treat as LB (thats why i have very minimal submissions and a low score yet). That way i hope that i will manage to do the hyper param tunning without overfitting to the public LB. Anybody thoughts about this kind of methodology? </p>",
      "votes": null,
      "replies": [
        {
          "id": 3197570,
          "author_name": "eikyou",
          "author_url": "",
          "post_date": "05/08/2025 10:05:49",
          "content": "<p>Sound trustful enough on condition there is enough data. Are u sure metrics won't be too noisy due to small amount of data?</p>",
          "votes": null,
          "replies": [
            {
              "id": 3197585,
              "author_name": "vasileioscharatsidis",
              "author_url": "",
              "post_date": "05/08/2025 10:23:45",
              "content": "<p>No i am not sure, i just go with this strategy for now</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3199061,
      "author_name": "mohanapavanbezawada",
      "author_url": "",
      "post_date": "05/10/2025 11:44:56",
      "content": "<p>hey if you don't mind telling me. can you please tell me how to reduce the false positives. i am using yolo11x with pretrained. its a different competition i am talking about in which the train images has only single object without any other background objects. but the test data has multiple images. i think this is confusing my model and its sometimes detecting the false objects with higher accuracy sometimes higher than actual images. can you give me a nice suggestion.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3199067,
          "author_name": "tom99763",
          "author_url": "",
          "post_date": "05/10/2025 11:50:37",
          "content": "<p>Reducing the strength of augmentation.</p>",
          "votes": null,
          "replies": [
            {
              "id": 3199245,
              "author_name": "mohanapavanbezawada",
              "author_url": "",
              "post_date": "05/10/2025 17:14:45",
              "content": "<p>Hi Tom, I am new to this field and participating in a competition with over 50 participants. To secure a top 3 position, should I focus on parameter tuning and augmentation, or would it be better to try advanced techniques like changing the YOLO backbone or model ensembling? The top participant achieved 97%, but they didn't share their approach. I would really appreciate your guidance on this.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 3199626,
                  "author_name": "tom99763",
                  "author_url": "",
                  "post_date": "05/11/2025 08:46:44",
                  "content": "<p><a href=\"https://www.kaggle.com/mohanapavanbezawada\" target=\"_blank\">@mohanapavanbezawada</a> for begining you can take a look at the yolo notebook provided by host. I would say ensembling is challenging for object localization cuz you need to consider the perception of each model and merge them. You can refer the 1st solution in crytoET comp. Their ensembling method is great. Maybe I'll public my best yolo ensembling notebook later. You can copy and edit it. Changing backbone also works, but it needs engineering effort.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 3199347,
      "author_name": "abdallahabdallahatef",
      "author_url": "",
      "post_date": "05/10/2025 20:49:23",
      "content": "<p>good job👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3199958,
      "author_name": "khankhantmg",
      "author_url": "",
      "post_date": "05/11/2025 18:48:29",
      "content": "<p>you can add back function</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3195326": "I've shared some early results here at [@eikyou discussion](https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/575799)\n\nBut it started getting a bit long, so I’m opening another discussion.\n\nFor the past two weeks, I’ve been training various YOLO models and always submitted the `best.pt` checkpoint, assuming that Ultralytics' default behavior  [0.1 × mAP@0.5 + 0.9 × mAP@0.5:0.95](https://github.com/ultralytics/ultralytics/issues/14137) would consistently result in the best public LB result.\n\nI was wrong.\n\nIn one of my recent runs, I got the best.pt checkpoints at different epochs (not 100% about the epochs):\n\n* Around ~ 28 → LB: **0.816**\n* Around ~ 40 → LB: **0.752**\n* Around ~ 50 → LB: **0.744**\n\nThis made me realize that some of my past experiments might’ve had better LB potential if I had just submitted a different checkpoint. So I started investigating what metric I should trust to pick the best epoch... but I haven’t found a clear pattern yet.\n\nHere’s some of my current results (each line a model/checkpoint) trained on 4/5 and evaluated on 1/5 of only positive images :\n\n```\nfbeta, precision, recall, f1, mAP50, mAP50-95\n0.816, 0.9651, 0.9594, 0.9623, 0.9812, 0.7316\n0.788, 0.9797, 0.9506, 0.9649, 0.9826, 0.6589\n0.749, 0.9696, 0.9572, 0.9633, 0.9843, 0.7268\n0.706, 0.9758, 0.9629, 0.9693, 0.9806, 0.7020\n0.744, 0.9747, 0.9542, 0.9644, 0.9888, 0.7621\n0.792, 0.9723, 0.9770, 0.9747, 0.9822, 0.7349\n0.714, 0.9782, 0.9523, 0.9651, 0.9818, 0.7161\n0.747, 0.9663, 0.9400, 0.9530, 0.9859, 0.7301\n```\n\nI also computed the correlation between fbeta and other metrics:\n\n```\nprecision:   -0.397  \nrecall:       0.258  \nf1:           0.055  \nmAP50:       -0.116  \nmAP50-95:    -0.038  \n```\n![Correlation Matrix](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2221915%2Fed3a4581733209200f5202ce75cb3158%2FcoefMatrix.jfif?generation=1746570087097235&alt=media)\n\nStill experimenting, but I thought this might be helpful if others are blindly trusting `best.pt` like I was.",
    "3195355": "I have recently tried saving the model for each epoch with save_period=1 and selecting the model for the smallest epoch of dfl-loss in addition to best.pt. Generally, but not always, the model with the smallest dfl-loss tends to have the higher LB.",
    "3195364": "can you trust this result not shake up in the end? That's my biggest concern for tunning yolo like that.",
    "3195365": "How did u split the data into train and val? Do the slices from the same tomo appear both in train and val or do you do split based on tomos?",
    "3195372": "I split the data by tomogram. I'm currently only using tomograms that contain one motor, which gave me 250 tomograms for training and 63 for validation",
    "3195379": "Yeah, that’s definitely something I’m worried about too. I’m planning to add Bartley external dataset as a validation set (just need to handle the preprocessing to match mine) to see if I can find any consistent correlation between my best models on the LB and some metric.\n\nReally hoping it helps, because things feel too random right now, and it makes testing new ideas pretty hard",
    "3195384": "sersasj one thing you can do is to do augmentation on the validation set. The result is quite stable to cv and lb.",
    "3196434": "how do you ensemble those two models?",
    "3196452": "No ensemble will be made.\nSubmit best.pt and dfl_loss_best.pt separately to observe LB.",
    "3196457": "Could I ask about your training data? Do you use samples where the number of motors is greater than 0, or only those where the number of motors equals 1?",
    "3196462": "Could I ask which version of Ultralytics you're using? I'm using v8.3.111, and in my experiments, different versions of the codebase produce different validation results on the same dataset.",
    "3196467": "The number of motors in the sample is greater than or equal to 1. I used more than 1 because the LB score was not good with only a sample with a number of motors equals 1.",
    "3196644": "For us as well, the same model exhibits significant randomness across different runs, and this issue still persists to this day.",
    "3196734": "fangsionfang I'm using version 8.3.121. I actually cloned the Ultralytics repo and have been modifying a few things, changing the activation function, loss, trying out timm backbones with YOLO head, etc.",
    "3197039": "Massive +1 to this — blindly trusting best.pt cost me some decent submissions too.\n\nUltralytics’ default metric might be great for general object detection, but when LB scoring diverges (like heavy class imbalance or FBeta-weighted targets), mAP just doesn’t cut it as a reliable proxy.\n\nI’ve started saving checkpoints every few epochs and scoring them manually on a holdout LB-style fold using the comp metric (e.g., fbeta), then picking the actual best. Bit more work, but worth it.\n\nAppreciate you sharing the correlation matrix — confirms that intuition beats default metrics sometimes 🔍\n\nAnyone tried using LB proxy models to predict the best checkpoint?",
    "3197135": "Tried submitting dfl_loss_best with different ckpts today but result was much worse.",
    "3197235": "My experience is that dfl_best_loss models sometimes score lower than best.pt, but it has never been extremely bad. try optimizing epoch's model. The optimization code can be found at https://www.kaggle.com/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/577949",
    "3197553": "I dont know if what i do is the most robust method but i have splited the train data into train_val and test. Then i measure everything in my test set which i treat as LB (thats why i have very minimal submissions and a low score yet). That way i hope that i will manage to do the hyper param tunning without overfitting to the public LB. Anybody thoughts about this kind of methodology?",
    "3197570": "Sound trustful enough on condition there is enough data. Are u sure metrics won't be too noisy due to small amount of data?",
    "3197585": "No i am not sure, i just go with this strategy for now",
    "3198133": "Thank you very much. \nNormally, augmentation of validation data is not done, but I found it useful when the data is small, as in this case. Certainly the dfl_loss of the validation set in training is stable.",
    "3198166": "Thank you for sharing your insights. I have some doubts about using the DFL loss as the main metric for model evaluation. Fundamentally, DFL loss functions similarly to bounding box regression loss. However, in this motor detection task, the critical requirement is precise center localization - bounding box accuracy is actually of secondary importance.\n\nIn practice, model achieves satisfactory performance early in training (demonstrating good center-point detection and classification capability). The subsequent optimization might lead to overfitting through excessive refinement of box edges. So continuous optimization of dfl loss afterwards does not necessarily imply an improvement in model performance\n\nSo I take it that when the dfl loss is lowest does it mean that the model is not necessarily the best?",
    "3198181": "Thank you for your input.\nI don't consider the detection of the center position so strictly based on the evaluation index which is positive within 1000 Å. Since most of the Voxel spacing is 10-20 Å, I interpret a discrepancy of up to 50 pixels as OK (please correct me if I'm wrong). \nSince Yolo is excellent, I believe that the motor decision (0 or 1) affects the score more than the exact position.",
    "3198188": "Yes, improving the model's ability to distinguish between positive and negative samples is a crucial direction, while localization capability is indeed less critical in this context.",
    "3199008": "you can add callback function to save dfl_loss_best model",
    "3199061": "hey if you don't mind telling me. can you please tell me how to reduce the false positives. i am using yolo11x with pretrained. its a different competition i am talking about in which the train images has only single object without any other background objects. but the test data has multiple images. i think this is confusing my model and its sometimes detecting the false objects with higher accuracy sometimes higher than actual images. can you give me a nice suggestion.",
    "3199067": "Reducing the strength of augmentation.",
    "3199245": "Hi Tom, I am new to this field and participating in a competition with over 50 participants. To secure a top 3 position, should I focus on parameter tuning and augmentation, or would it be better to try advanced techniques like changing the YOLO backbone or model ensembling? The top participant achieved 97%, but they didn't share their approach. I would really appreciate your guidance on this.",
    "3199347": "good job👍",
    "3199626": "mohanapavanbezawada for begining you can take a look at the yolo notebook provided by host. I would say ensembling is challenging for object localization cuz you need to consider the perception of each model and merge them. You can refer the 1st solution in crytoET comp. Their ensembling method is great. Maybe I'll public my best yolo ensembling notebook later. You can copy and edit it. Changing backbone also works, but it needs engineering effort.",
    "3199670": "Have you noticed any specific patterns or conditions where the model with the lowest dfl-loss doesn't perform best on the leaderboard?",
    "3199958": "you can add back function",
    "3200107": "how to do it？",
    "3200215": "An example code:\n\n```\nbest_dfl_loss = float('inf')\ndef on_fit_epoch_end(trainer):\n    global best_dfl_loss\n    save_dir = 'run/train/weights'\n    \n    val_metrics = getattr(trainer, 'metrics', None)\n\n    current_dfl_loss = None\n    if trainer.epoch == 0: \n         print(f\"Debug: Available validation metric keys at epoch {trainer.epoch}: {list(val_metrics.keys())}\")\n\n    current_dfl_loss = val_metrics['val/dfl_loss']\n\n    if current_dfl_loss is not None:\n        print(f\"Epoch {trainer.epoch}: Current DFL Loss: {current_dfl_loss:.4f}, Best DFL Loss: {best_dfl_loss:.4f}\")\n        if current_dfl_loss < best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, f'best_dfl_loss_epoch_{trainer.epoch}_loss_{current_dfl_loss:.4f}.pt')\n\n            ckpt = {\n                'epoch': trainer.epoch,\n                'best_fitness': None, \n                'updates': trainer.ema.updates if trainer.ema else 0,\n                'optimizer': trainer.optimizer.state_dict(),\n                'train_args': vars(trainer.args), \n                'dfl_loss_for_best': best_dfl_loss\n            }\n    \n            if trainer.ema and hasattr(trainer.ema, 'ema'):\n                ckpt['ema'] = trainer.ema.ema.float()\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model\n            else:\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model \n            torch.save(ckpt, ckpt_path)\n\n            print(f\"Epoch {trainer.epoch+1}: Saved new best model (dfl_loss): {ckpt_path} with DFL Loss {best_dfl_loss:.4f}\")\n\nmodel = YOLO(\"yolov8n.pt\")\n\nmodel.add_callback(\"on_fit_epoch_end\", on_fit_epoch_end)\n\nmodel.train()\n```",
    "3202345": "Below is a more robust implementation with corrected grammar and syntax:\n\nMain Modifications:\nAutomatically save weights to the project directory.\n'motor_detector' is used as the name in the model.train() function.\n\nAt the end of training, the 'val/dfl_loss' metric does not exist, which raises a KeyError.\nI handled this issue by adding a conditional check and a print statement to gracefully skip it.\n\n\n```\nbest_dfl_loss = float('inf')\ndef on_fit_epoch_end(trainer):\n    global best_dfl_loss\n    if hasattr(trainer, 'args') and hasattr(trainer.args, 'project'):\n        save_dir = os.path.join(trainer.args.project, 'motor_detector', 'weights')\n    else:\n        save_dir = 'runs/train/weights'\n\n    os.makedirs(save_dir, exist_ok=True)\n\n    val_metrics = getattr(trainer, 'metrics', None)\n    current_dfl_loss = None\n\n    if trainer.epoch == 0:\n        print(f\"Debug: Available validation metric keys at epoch {trainer.epoch}: {list(val_metrics.keys()) if val_metrics else 'None'}\")\n\n    if val_metrics and 'val/dfl_loss' in val_metrics:\n        current_dfl_loss = val_metrics['val/dfl_loss']\n\n        print(f\"Epoch {trainer.epoch}: Current DFL Loss: {current_dfl_loss:.4f}, Best DFL Loss: {best_dfl_loss:.4f}\")\n        if current_dfl_loss < best_dfl_loss:\n            best_dfl_loss = current_dfl_loss\n            ckpt_path = os.path.join(save_dir, f'best_dfl_loss_epoch_{trainer.epoch}_loss_{current_dfl_loss:.4f}.pt')\n\n            ckpt = {\n                'epoch': trainer.epoch,\n                'best_fitness': None,\n                'updates': trainer.ema.updates if trainer.ema else 0,\n                'optimizer': trainer.optimizer.state_dict(),\n                'train_args': vars(trainer.args),\n                'dfl_loss_for_best': best_dfl_loss\n            }\n\n            if trainer.ema and hasattr(trainer.ema, 'ema'):\n                ckpt['ema'] = trainer.ema.ema.float()\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model\n            else:\n                ckpt['model'] = deepcopy(trainer.model).half() if isinstance(trainer.model, nn.Module) else trainer.model\n\n            torch.save(ckpt, ckpt_path)\n\n            print(f\"Epoch {trainer.epoch+1}: Saved new best model (dfl_loss): {ckpt_path} with DFL Loss {best_dfl_loss:.4f}\")\n    else:\n        print(f\"Epoch {trainer.epoch}: val/dfl_loss not found in metrics. Available keys: {list(val_metrics.keys()) if val_metrics else 'None'}\")\n```",
    "3202390": "Thanks for sharing with us!"
  },
  "source": "meta"
}