{
  "id": 581247,
  "title": "How to avoid overfitting.",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/581247",
  "author_name": "Handudu",
  "post_date": "2025-05-29T12:02:50.119000",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Different folds and inference parameters result in a significant gap in the LB scores for each submission. Is there any method that can enhance model stability, such as Test Time Augmentation (TTA)? I believe the leaderboard will experience a big shake-up.</p>",
  "messages": [
    {
      "id": 3212161,
      "postDate": "2025-05-29T12:02:50.120Z",
      "content": "<p>Different folds and inference parameters result in a significant gap in the LB scores for each submission. Is there any method that can enhance model stability, such as Test Time Augmentation (TTA)? I believe the leaderboard will experience a big shake-up.</p>",
      "rawMarkdown": "Different folds and inference parameters result in a significant gap in the LB scores for each submission. Is there any method that can enhance model stability, such as Test Time Augmentation (TTA)? I believe the leaderboard will experience a big shake-up.",
      "votes": 5
    },
    {
      "id": 3212166,
      "postDate": "2025-05-29T12:09:39.427Z",
      "content": "<p><a href=\"https://www.kaggle.com/handudu\" target=\"_blank\">@handudu</a> Try to produce false positives as much as you can.</p>",
      "rawMarkdown": "@handudu Try to produce false positives as much as you can.",
      "votes": 1,
      "replies": [
        {
          "id": 3212207,
          "postDate": "2025-05-29T13:27:09.350Z",
          "content": "<p>hey <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> , could you please explain how producing more FP will lead to better correlation between CV and LB?</p>",
          "rawMarkdown": "hey @tom99763 , could you please explain how producing more FP will lead to better correlation between CV and LB?",
          "replies": [
            {
              "id": 3212232,
              "postDate": "2025-05-29T14:23:52.043Z",
              "content": "<p>To be honest, I do not built a solid local CV for this competition. Because I cannot simulate the 0/1 sample distribution in the test data locally. However, I believe that those who scored high did this work.</p>",
              "rawMarkdown": "To be honest, I do not built a solid local CV for this competition. Because I cannot simulate the 0/1 sample distribution in the test data locally. However, I believe that those who scored high did this work."
            },
            {
              "id": 3212240,
              "postDate": "2025-05-29T14:36:01.343Z",
              "content": "<p>I also expect a shake-up, considering that public and private lb are randomly split. It's also a problem with a low number of positive samples, in the few hundreds, so that's also introducing volatility. The selection of the subs is also hard!</p>",
              "rawMarkdown": "I also expect a shake-up, considering that public and private lb are randomly split. It's also a problem with a low number of positive samples, in the few hundreds, so that's also introducing volatility. The selection of the subs is also hard!",
              "votes": 1
            },
            {
              "id": 3212248,
              "postDate": "2025-05-29T14:47:16.560Z",
              "content": "<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> If you experiment enough with YOLO8 to YOLO11, you'll find that one of them shows a strong CV-LB correlation within a confidence threshold of 0.05 to 0.3. And based on the evaluation metric, accepting more false positives is penalized less than accepting more false negatives. There's no way to risk the latter in the final submission.</p>",
              "rawMarkdown": "@iamparadox If you experiment enough with YOLO8 to YOLO11, you'll find that one of them shows a strong CV-LB correlation within a confidence threshold of 0.05 to 0.3. And based on the evaluation metric, accepting more false positives is penalized less than accepting more false negatives. There's no way to risk the latter in the final submission.",
              "votes": 2
            },
            {
              "id": 3212313,
              "postDate": "2025-05-29T16:01:41.097Z",
              "content": "<p>Initially my aim was to find a robust way to correlate CV and LB. I tried various ways to simulate the domain shift, for example I split the training data based on voxel size meaning I hold out tomograph of certain voxel sizes completely and then validate on that, but it didn't yield any correlation and so does adding noise to the validation set. All of our models are able to adapt to these domain shifts, furthermore our models CV is very high like 0.97-0.98 (as is the case for others), so when our CV increases a little bit I am not sure to attribute this to randomness or the change in model</p>\n<p>But that's why this competition is so much fun and a great learning experience !</p>",
              "rawMarkdown": "Initially my aim was to find a robust way to correlate CV and LB. I tried various ways to simulate the domain shift, for example I split the training data based on voxel size meaning I hold out tomograph of certain voxel sizes completely and then validate on that, but it didn't yield any correlation and so does adding noise to the validation set. All of our models are able to adapt to these domain shifts, furthermore our models CV is very high like 0.97-0.98 (as is the case for others), so when our CV increases a little bit I am not sure to attribute this to randomness or the change in model\n\nBut that's why this competition is so much fun and a great learning experience !",
              "votes": 2
            },
            {
              "id": 3213400,
              "postDate": "2025-05-30T00:38:06.993Z",
              "content": "<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> <br>\nMy cv scores are also located in 0.97-0.98+. It's quite hard to build good cv since we are predicting the labeler. I actually have experience handling this kind of task when I studied my m.s. degree. I have industry cooperation with Micron Technology, handling HBM anomaly detection. This project is much hard than the problem in this comp. They ask me using only 88 scanned images to build a very good detector to replace one of their product process. The labels are annotated by their expert where some of the ground truths are labeled as positives in their AOI machine but accepted as negatives from the expert 💀 So the actual task is modeling the preference of expert to reduce their production cost.<br>\nMy solution in this comp is based on the part of this project. The whole project is more complex and innovative than the approach I use in this comp. I will also introduce it in my solution write-up.</p>",
              "rawMarkdown": "@iamparadox \nMy cv scores are also located in 0.97-0.98+. It's quite hard to build good cv since we are predicting the labeler. I actually have experience handling this kind of task when I studied my m.s. degree. I have industry cooperation with Micron Technology, handling HBM anomaly detection. This project is much hard than the problem in this comp. They ask me using only 88 scanned images to build a very good detector to replace one of their product process. The labels are annotated by their expert where some of the ground truths are labeled as positives in their AOI machine but accepted as negatives from the expert 💀 So the actual task is modeling the preference of expert to reduce their production cost.\nMy solution in this comp is based on the part of this project. The whole project is more complex and innovative than the approach I use in this comp. I will also introduce it in my solution write-up."
            }
          ]
        }
      ]
    },
    {
      "id": 3212216,
      "postDate": "2025-05-29T14:00:38.440Z",
      "rawMarkdown": "",
      "votes": -5,
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3212166,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-05-29T12:09:39.427000",
      "content": "<p><a href=\"https://www.kaggle.com/handudu\" target=\"_blank\">@handudu</a> Try to produce false positives as much as you can.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3212207,
          "author_name": "IAmParadox",
          "author_url": "",
          "post_date": "2025-05-29T13:27:09.350000",
          "content": "<p>hey <a href=\"https://www.kaggle.com/tom99763\" target=\"_blank\">@tom99763</a> , could you please explain how producing more FP will lead to better correlation between CV and LB?</p>",
          "votes": 0,
          "replies": [
            {
              "id": 3212232,
              "author_name": "Handudu",
              "author_url": "",
              "post_date": "2025-05-29T14:23:52.043000",
              "content": "<p>To be honest, I do not built a solid local CV for this competition. Because I cannot simulate the 0/1 sample distribution in the test data locally. However, I believe that those who scored high did this work.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3212240,
              "author_name": "Andrei Zamfir",
              "author_url": "",
              "post_date": "2025-05-29T14:36:01.343000",
              "content": "<p>I also expect a shake-up, considering that public and private lb are randomly split. It's also a problem with a low number of positive samples, in the few hundreds, so that's also introducing volatility. The selection of the subs is also hard!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3212248,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-05-29T14:47:16.560000",
              "content": "<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> If you experiment enough with YOLO8 to YOLO11, you'll find that one of them shows a strong CV-LB correlation within a confidence threshold of 0.05 to 0.3. And based on the evaluation metric, accepting more false positives is penalized less than accepting more false negatives. There's no way to risk the latter in the final submission.</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3212313,
              "author_name": "IAmParadox",
              "author_url": "",
              "post_date": "2025-05-29T16:01:41.097000",
              "content": "<p>Initially my aim was to find a robust way to correlate CV and LB. I tried various ways to simulate the domain shift, for example I split the training data based on voxel size meaning I hold out tomograph of certain voxel sizes completely and then validate on that, but it didn't yield any correlation and so does adding noise to the validation set. All of our models are able to adapt to these domain shifts, furthermore our models CV is very high like 0.97-0.98 (as is the case for others), so when our CV increases a little bit I am not sure to attribute this to randomness or the change in model</p>\n<p>But that's why this competition is so much fun and a great learning experience !</p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 3213400,
              "author_name": "Tom",
              "author_url": "",
              "post_date": "2025-05-30T00:38:06.993000",
              "content": "<p><a href=\"https://www.kaggle.com/iamparadox\" target=\"_blank\">@iamparadox</a> <br>\nMy cv scores are also located in 0.97-0.98+. It's quite hard to build good cv since we are predicting the labeler. I actually have experience handling this kind of task when I studied my m.s. degree. I have industry cooperation with Micron Technology, handling HBM anomaly detection. This project is much hard than the problem in this comp. They ask me using only 88 scanned images to build a very good detector to replace one of their product process. The labels are annotated by their expert where some of the ground truths are labeled as positives in their AOI machine but accepted as negatives from the expert 💀 So the actual task is modeling the preference of expert to reduce their production cost.<br>\nMy solution in this comp is based on the part of this project. The whole project is more complex and innovative than the approach I use in this comp. I will also introduce it in my solution write-up.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3212216,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-29T14:00:38.440000",
      "content": "",
      "votes": -5,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3212161": "Different folds and inference parameters result in a significant gap in the LB scores for each submission. Is there any method that can enhance model stability, such as Test Time Augmentation (TTA)? I believe the leaderboard will experience a big shake-up.",
    "3212166": "@handudu Try to produce false positives as much as you can.",
    "3212216": ""
  }
}