{
  "id": 583163,
  "title": "Alternate Validation Strategy",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/583163",
  "author_name": "",
  "post_date": "2025-06-05T06:02:18.823652700Z",
  "votes": 5,
  "comment_count": 1,
  "views": 0,
  "content": "<p>One of the things that made this competition very difficult was measuring performance with multi-motor data and then trying to transfer that to the leaderboard. What I saw most people doing was just dropping the multi motor and only including single motor tomograms or 1 and 0 motor tomograms. This dropped quite a bit of data in a competition that was already pretty short of it. Most people using this method eventually hit a wall because validating against this subset of the data was too easy and yielded deceptively high scores. </p>\n<p>In the last week or so I figured out a strategy that seems to carry over to the leaderboard a lot better. I modified the competition metric and my postprocessing code so that it could pass through multiple motor predictions and I used all of the data to validate against. With the increased motor and tomogram count and using 5 folds I was able to get numbers that very closely resembled the leaderboard and even yielded thresholds that seemed to line up with my expectations and experience on the leaderboard. I wish I had figured this out sooner so that I didn't have so much noise while climbing. </p>\n<p>One of the things I did along with this was visualize the threshold sweeps. My postprocessing used 2 thresholds, slice confidence to filter out weak detections and then also a cluster sum threshold, adding together confidence from close and overlapping predictions. From this I made some 3d plots that showed the surface. It made it more clear why sometimes things were extremely fragile to thresholds on the leaderboard. Comparing the raw detection outputs and a different second layer classifier. It only gets a slightly higher peak but the bowl is not nearly as sharp. Landing exactly the right threshold on the first is very difficult. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6889ff3b8302b7421ce4e6e76007a6be%2FScreenshot%202025-06-04%20at%2010.57.02PM.png?generation=1749103061419933&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffd53abda52a7b734df12f0970e8db7d4%2FScreenshot%202025-06-04%20at%2010.57.12PM.png?generation=1749103046621062&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "3217547",
      "postDate": "06/05/2025 06:02:18",
      "content": "<p>One of the things that made this competition very difficult was measuring performance with multi-motor data and then trying to transfer that to the leaderboard. What I saw most people doing was just dropping the multi motor and only including single motor tomograms or 1 and 0 motor tomograms. This dropped quite a bit of data in a competition that was already pretty short of it. Most people using this method eventually hit a wall because validating against this subset of the data was too easy and yielded deceptively high scores. </p>\n<p>In the last week or so I figured out a strategy that seems to carry over to the leaderboard a lot better. I modified the competition metric and my postprocessing code so that it could pass through multiple motor predictions and I used all of the data to validate against. With the increased motor and tomogram count and using 5 folds I was able to get numbers that very closely resembled the leaderboard and even yielded thresholds that seemed to line up with my expectations and experience on the leaderboard. I wish I had figured this out sooner so that I didn't have so much noise while climbing. </p>\n<p>One of the things I did along with this was visualize the threshold sweeps. My postprocessing used 2 thresholds, slice confidence to filter out weak detections and then also a cluster sum threshold, adding together confidence from close and overlapping predictions. From this I made some 3d plots that showed the surface. It made it more clear why sometimes things were extremely fragile to thresholds on the leaderboard. Comparing the raw detection outputs and a different second layer classifier. It only gets a slightly higher peak but the bowl is not nearly as sharp. Landing exactly the right threshold on the first is very difficult. </p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6889ff3b8302b7421ce4e6e76007a6be%2FScreenshot%202025-06-04%20at%2010.57.02PM.png?generation=1749103061419933&amp;alt=media\" alt=\"\"></p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffd53abda52a7b734df12f0970e8db7d4%2FScreenshot%202025-06-04%20at%2010.57.12PM.png?generation=1749103046621062&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "One of the things that made this competition very difficult was measuring performance with multi-motor data and then trying to transfer that to the leaderboard. What I saw most people doing was just dropping the multi motor and only including single motor tomograms or 1 and 0 motor tomograms. This dropped quite a bit of data in a competition that was already pretty short of it. Most people using this method eventually hit a wall because validating against this subset of the data was too easy and yielded deceptively high scores. \n\nIn the last week or so I figured out a strategy that seems to carry over to the leaderboard a lot better. I modified the competition metric and my postprocessing code so that it could pass through multiple motor predictions and I used all of the data to validate against. With the increased motor and tomogram count and using 5 folds I was able to get numbers that very closely resembled the leaderboard and even yielded thresholds that seemed to line up with my expectations and experience on the leaderboard. I wish I had figured this out sooner so that I didn't have so much noise while climbing. \n\nOne of the things I did along with this was visualize the threshold sweeps. My postprocessing used 2 thresholds, slice confidence to filter out weak detections and then also a cluster sum threshold, adding together confidence from close and overlapping predictions. From this I made some 3d plots that showed the surface. It made it more clear why sometimes things were extremely fragile to thresholds on the leaderboard. Comparing the raw detection outputs and a different second layer classifier. It only gets a slightly higher peak but the bowl is not nearly as sharp. Landing exactly the right threshold on the first is very difficult. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6889ff3b8302b7421ce4e6e76007a6be%2FScreenshot%202025-06-04%20at%2010.57.02PM.png?generation=1749103061419933&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffd53abda52a7b734df12f0970e8db7d4%2FScreenshot%202025-06-04%20at%2010.57.12PM.png?generation=1749103046621062&alt=media)",
      "votes": null
    },
    {
      "id": "3248712",
      "postDate": "07/15/2025 04:33:06",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Good insights and analysis. </p>",
      "rawMarkdown": "ryches Good insights and analysis.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3248712,
      "author_name": "tom99763",
      "author_url": "",
      "post_date": "07/15/2025 04:33:06",
      "content": "<p><a href=\"https://www.kaggle.com/ryches\" target=\"_blank\">@ryches</a> Good insights and analysis. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3217547": "One of the things that made this competition very difficult was measuring performance with multi-motor data and then trying to transfer that to the leaderboard. What I saw most people doing was just dropping the multi motor and only including single motor tomograms or 1 and 0 motor tomograms. This dropped quite a bit of data in a competition that was already pretty short of it. Most people using this method eventually hit a wall because validating against this subset of the data was too easy and yielded deceptively high scores. \n\nIn the last week or so I figured out a strategy that seems to carry over to the leaderboard a lot better. I modified the competition metric and my postprocessing code so that it could pass through multiple motor predictions and I used all of the data to validate against. With the increased motor and tomogram count and using 5 folds I was able to get numbers that very closely resembled the leaderboard and even yielded thresholds that seemed to line up with my expectations and experience on the leaderboard. I wish I had figured this out sooner so that I didn't have so much noise while climbing. \n\nOne of the things I did along with this was visualize the threshold sweeps. My postprocessing used 2 thresholds, slice confidence to filter out weak detections and then also a cluster sum threshold, adding together confidence from close and overlapping predictions. From this I made some 3d plots that showed the surface. It made it more clear why sometimes things were extremely fragile to thresholds on the leaderboard. Comparing the raw detection outputs and a different second layer classifier. It only gets a slightly higher peak but the bowl is not nearly as sharp. Landing exactly the right threshold on the first is very difficult. \n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2F6889ff3b8302b7421ce4e6e76007a6be%2FScreenshot%202025-06-04%20at%2010.57.02PM.png?generation=1749103061419933&alt=media)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1035002%2Ffd53abda52a7b734df12f0970e8db7d4%2FScreenshot%202025-06-04%20at%2010.57.12PM.png?generation=1749103046621062&alt=media)",
    "3248712": "ryches Good insights and analysis."
  },
  "source": "meta"
}