{
  "id": 572731,
  "title": "a few tricks for test domain shift in kaggle competition",
  "url": "/competitions/byu-locating-bacterial-flagellar-motors-2025/discussion/572731",
  "author_name": "hengck23",
  "post_date": "2025-04-11T03:51:33.476000",
  "votes": 49,
  "comment_count": 9,
  "views": 0,
  "content": "<p>i make some submission and do some simple test, it is obviously that:<br>\n(1) your model give has lower  lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.<br>\n(2) despite domain shift, most model are good at predicting negative label (background image)</p>\n<ul>\n<li>you can prove this by: set very low threshold. make submission for object only for &lt;0.1, then for &lt;0.2 ….. we want to find the point when lb stop giving zero (lb is zero when both precision and recall are zero). compare this with your training.</li>\n</ul>\n<p>(3) if the false accept rate is low or zero, then you can do test probing. e.g. if you flip the image intensity and if  there is some \"lb score\", you can conclude some image have intensity flip (like the ones in the train). you can proceed to test for larger scale, smaller scale, blur,  etc.. in short you can probe test domain shift.</p>\n<p>I recall a kaggle competition for coral reef object detection where is hidden test label has \"bigger aspect box label\" (test and train labels are differently).  while many kagglers are puzzled why their lb is worse than lb, some kagglers find/probe this in the competition</p>\n<p>(3) normally if your model works robust for image +noise, it is resilient to domain shift. try to use noise augmentation.</p>\n<p>(4) what is fatal to model:</p>\n<ul>\n<li>when test scale is different from train</li>\n<li>when test intensity level is different from train (different histogram distribution)</li>\n<li>when test noise level is different from train</li>\n<li>artifacts in test</li>\n<li>different object structure/appearance in test</li>\n</ul>\n<p>good luck!</p>\n<hr>\n<p>if you are developing AI commercial product, you imagine your customer problem and think of how to develop AI tech to help them.</p>\n<p>If you are writing AI paper, you imagine who others will use your work to advance even more and you think of solving fundamental problems</p>\n<p>if you are taking part in AI competition, you imagine how the hidden test data will look like and think of how that can help you to be better than others</p>",
  "messages": [
    {
      "id": 3176187,
      "postDate": "2025-04-11T03:51:33.477Z",
      "content": "<p>i make some submission and do some simple test, it is obviously that:<br>\n(1) your model give has lower  lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.<br>\n(2) despite domain shift, most model are good at predicting negative label (background image)</p>\n<ul>\n<li>you can prove this by: set very low threshold. make submission for object only for &lt;0.1, then for &lt;0.2 ….. we want to find the point when lb stop giving zero (lb is zero when both precision and recall are zero). compare this with your training.</li>\n</ul>\n<p>(3) if the false accept rate is low or zero, then you can do test probing. e.g. if you flip the image intensity and if  there is some \"lb score\", you can conclude some image have intensity flip (like the ones in the train). you can proceed to test for larger scale, smaller scale, blur,  etc.. in short you can probe test domain shift.</p>\n<p>I recall a kaggle competition for coral reef object detection where is hidden test label has \"bigger aspect box label\" (test and train labels are differently).  while many kagglers are puzzled why their lb is worse than lb, some kagglers find/probe this in the competition</p>\n<p>(3) normally if your model works robust for image +noise, it is resilient to domain shift. try to use noise augmentation.</p>\n<p>(4) what is fatal to model:</p>\n<ul>\n<li>when test scale is different from train</li>\n<li>when test intensity level is different from train (different histogram distribution)</li>\n<li>when test noise level is different from train</li>\n<li>artifacts in test</li>\n<li>different object structure/appearance in test</li>\n</ul>\n<p>good luck!</p>\n<hr>\n<p>if you are developing AI commercial product, you imagine your customer problem and think of how to develop AI tech to help them.</p>\n<p>If you are writing AI paper, you imagine who others will use your work to advance even more and you think of solving fundamental problems</p>\n<p>if you are taking part in AI competition, you imagine how the hidden test data will look like and think of how that can help you to be better than others</p>",
      "rawMarkdown": "i make some submission and do some simple test, it is obviously that:\n(1) your model give has lower  lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.\n(2) despite domain shift, most model are good at predicting negative label (background image)\n- you can prove this by: set very low threshold. make submission for object only for <0.1, then for <0.2 ..... we want to find the point when lb stop giving zero (lb is zero when both precision and recall are zero). compare this with your training.\n\n(3) if the false accept rate is low or zero, then you can do test probing. e.g. if you flip the image intensity and if  there is some \"lb score\", you can conclude some image have intensity flip (like the ones in the train). you can proceed to test for larger scale, smaller scale, blur,  etc.. in short you can probe test domain shift.\n\nI recall a kaggle competition for coral reef object detection where is hidden test label has \"bigger aspect box label\" (test and train labels are differently).  while many kagglers are puzzled why their lb is worse than lb, some kagglers find/probe this in the competition\n\n(3) normally if your model works robust for image +noise, it is resilient to domain shift. try to use noise augmentation.\n\n(4) what is fatal to model:\n- when test scale is different from train\n- when test intensity level is different from train (different histogram distribution)\n- when test noise level is different from train\n- artifacts in test\n- different object structure/appearance in test\n\ngood luck!\n\n\n----\n\nif you are developing AI commercial product, you imagine your customer problem and think of how to develop AI tech to help them.\n\nIf you are writing AI paper, you imagine who others will use your work to advance even more and you think of solving fundamental problems\n\nif you are taking part in AI competition, you imagine how the hidden test data will look like and think of how that can help you to be better than others",
      "votes": 49
    },
    {
      "id": 3177086,
      "postDate": "2025-04-12T08:25:41.217Z",
      "content": "<p>In my case, my model development is based on keypoint detection (like orb in opencv, I think it ignores domain shift cause it is not a ml model). After probing many times, I found that my current model detects too much false positives on hidden testset. But its recall is great on my local test, almost cannot miss any motor.</p>",
      "rawMarkdown": "In my case, my model development is based on keypoint detection (like orb in opencv, I think it ignores domain shift cause it is not a ml model). After probing many times, I found that my current model detects too much false positives on hidden testset. But its recall is great on my local test, almost cannot miss any motor.",
      "votes": 3,
      "replies": [
        {
          "id": 3178100,
          "postDate": "2025-04-13T17:58:37.853Z",
          "content": "<p>Yeah from what I have heard about your solutions for this comp so far I am super excited to see your work this time around! </p>",
          "rawMarkdown": "Yeah from what I have heard about your solutions for this comp so far I am super excited to see your work this time around! "
        }
      ]
    },
    {
      "id": 3177501,
      "postDate": "2025-04-12T21:24:15.173Z",
      "content": "<p>Super insightful! Love the practical tricks—especially the probing methods and the mindset tips. Really sharp advice . Thanks for sharing!</p>",
      "rawMarkdown": "Super insightful! Love the practical tricks—especially the probing methods and the mindset tips. Really sharp advice . Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 3176255,
      "postDate": "2025-04-11T06:38:32.080Z",
      "content": "<p>This is more evident in my current Lb. I tried lowering the threshold by 0.05, lb 0.610 -&gt; 0.725, but I know there is still a lot more I need to do.</p>\n<blockquote>\n  <p>(1) your model give has lower lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.</p>\n</blockquote>",
      "rawMarkdown": "This is more evident in my current Lb. I tried lowering the threshold by 0.05, lb 0.610 -> 0.725, but I know there is still a lot more I need to do.\n>(1) your model give has lower lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.",
      "votes": 1
    },
    {
      "id": 3180011,
      "postDate": "2025-04-16T03:17:51.563Z",
      "content": "<p>maybe you know the % of pos samples in the test. then you can do dynamic thresholding</p>",
      "rawMarkdown": "maybe you know the % of pos samples in the test. then you can do dynamic thresholding",
      "votes": 2
    },
    {
      "id": 3180062,
      "postDate": "2025-04-16T04:35:44.103Z",
      "content": "<p>Amazing and informative </p>",
      "rawMarkdown": "Amazing and informative "
    },
    {
      "id": 3176416,
      "postDate": "2025-04-11T10:21:21.783Z",
      "content": "<p>I've been experimenting with the conf thresholds for several day by far. I think I found a way to get the optimal LB threshold based on cv only. Tested that on 4 different pipelines (for each one i submitted several thresholds and the one my cv says is optimal) and it gave best results in all. But choosing which model is better based on cv is still not solved for me. So i think we need to depend on lb a bit :)</p>",
      "rawMarkdown": "I've been experimenting with the conf thresholds for several day by far. I think I found a way to get the optimal LB threshold based on cv only. Tested that on 4 different pipelines (for each one i submitted several thresholds and the one my cv says is optimal) and it gave best results in all. But choosing which model is better based on cv is still not solved for me. So i think we need to depend on lb a bit :)",
      "replies": [
        {
          "id": 3176437,
          "postDate": "2025-04-11T10:50:00.710Z",
          "content": "<p>wokring on threshold will not win competition.<br>\nmaking model insensitive/robust against threshold will</p>",
          "rawMarkdown": "wokring on threshold will not win competition.\nmaking model insensitive/robust against threshold will",
          "votes": 4
        }
      ]
    },
    {
      "id": 3178148,
      "postDate": "2025-04-13T19:29:21.103Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 3177086,
      "author_name": "Tom",
      "author_url": "",
      "post_date": "2025-04-12T08:25:41.217000",
      "content": "<p>In my case, my model development is based on keypoint detection (like orb in opencv, I think it ignores domain shift cause it is not a ml model). After probing many times, I found that my current model detects too much false positives on hidden testset. But its recall is great on my local test, almost cannot miss any motor.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 3178100,
          "author_name": "Cody_Null",
          "author_url": "",
          "post_date": "2025-04-13T17:58:37.853000",
          "content": "<p>Yeah from what I have heard about your solutions for this comp so far I am super excited to see your work this time around! </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3177501,
      "author_name": "MEHUL MRIDUL",
      "author_url": "",
      "post_date": "2025-04-12T21:24:15.173000",
      "content": "<p>Super insightful! Love the practical tricks—especially the probing methods and the mindset tips. Really sharp advice . Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3176255,
      "author_name": "colt",
      "author_url": "",
      "post_date": "2025-04-11T06:38:32.080000",
      "content": "<p>This is more evident in my current Lb. I tried lowering the threshold by 0.05, lb 0.610 -&gt; 0.725, but I know there is still a lot more I need to do.</p>\n<blockquote>\n  <p>(1) your model give has lower lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.</p>\n</blockquote>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 3180011,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2025-04-16T03:17:51.563000",
      "content": "<p>maybe you know the % of pos samples in the test. then you can do dynamic thresholding</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3180062,
      "author_name": "Memoona Qaiser",
      "author_url": "",
      "post_date": "2025-04-16T04:35:44.103000",
      "content": "<p>Amazing and informative </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3176416,
      "author_name": "Mohamed Eltayeb",
      "author_url": "",
      "post_date": "2025-04-11T10:21:21.783000",
      "content": "<p>I've been experimenting with the conf thresholds for several day by far. I think I found a way to get the optimal LB threshold based on cv only. Tested that on 4 different pipelines (for each one i submitted several thresholds and the one my cv says is optimal) and it gave best results in all. But choosing which model is better based on cv is still not solved for me. So i think we need to depend on lb a bit :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 3176437,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2025-04-11T10:50:00.710000",
          "content": "<p>wokring on threshold will not win competition.<br>\nmaking model insensitive/robust against threshold will</p>",
          "votes": 4,
          "replies": []
        }
      ]
    },
    {
      "id": 3178148,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-04-13T19:29:21.103000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3176187": "i make some submission and do some simple test, it is obviously that:\n(1) your model give has lower  lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.\n(2) despite domain shift, most model are good at predicting negative label (background image)\n- you can prove this by: set very low threshold. make submission for object only for <0.1, then for <0.2 ..... we want to find the point when lb stop giving zero (lb is zero when both precision and recall are zero). compare this with your training.\n\n(3) if the false accept rate is low or zero, then you can do test probing. e.g. if you flip the image intensity and if  there is some \"lb score\", you can conclude some image have intensity flip (like the ones in the train). you can proceed to test for larger scale, smaller scale, blur,  etc.. in short you can probe test domain shift.\n\nI recall a kaggle competition for coral reef object detection where is hidden test label has \"bigger aspect box label\" (test and train labels are differently).  while many kagglers are puzzled why their lb is worse than lb, some kagglers find/probe this in the competition\n\n(3) normally if your model works robust for image +noise, it is resilient to domain shift. try to use noise augmentation.\n\n(4) what is fatal to model:\n- when test scale is different from train\n- when test intensity level is different from train (different histogram distribution)\n- when test noise level is different from train\n- artifacts in test\n- different object structure/appearance in test\n\ngood luck!\n\n\n----\n\nif you are developing AI commercial product, you imagine your customer problem and think of how to develop AI tech to help them.\n\nIf you are writing AI paper, you imagine who others will use your work to advance even more and you think of solving fundamental problems\n\nif you are taking part in AI competition, you imagine how the hidden test data will look like and think of how that can help you to be better than others",
    "3177086": "In my case, my model development is based on keypoint detection (like orb in opencv, I think it ignores domain shift cause it is not a ml model). After probing many times, I found that my current model detects too much false positives on hidden testset. But its recall is great on my local test, almost cannot miss any motor.",
    "3177501": "Super insightful! Love the practical tricks—especially the probing methods and the mindset tips. Really sharp advice . Thanks for sharing!",
    "3176255": "This is more evident in my current Lb. I tried lowering the threshold by 0.05, lb 0.610 -> 0.725, but I know there is still a lot more I need to do.\n>(1) your model give has lower lb then cv because of less detection (not more FP). your lb improve if you lower your threshold.",
    "3180011": "maybe you know the % of pos samples in the test. then you can do dynamic thresholding",
    "3180062": "Amazing and informative ",
    "3176416": "I've been experimenting with the conf thresholds for several day by far. I think I found a way to get the optimal LB threshold based on cv only. Tested that on 4 different pipelines (for each one i submitted several thresholds and the one my cv says is optimal) and it gave best results in all. But choosing which model is better based on cv is still not solved for me. So i think we need to depend on lb a bit :)",
    "3178148": ""
  }
}