{
  "id": 307607,
  "title": "My BBox annotation theory that might explain the CV-LB discrepancy and the shakeup",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/307607",
  "author_name": "",
  "post_date": "2022-02-15T00:53:04.587185800Z",
  "votes": 28,
  "comment_count": 4,
  "views": 0,
  "content": "<p>There has been many discussions here about this. The key question is why does the CV and LB not match? Related questions are:</p>\n<ul>\n<li>Why does inference at 2-3x training resolution benefit the public LB (in my case, 1.5x for YOLOX) ?</li>\n<li>Why did this trick not work for private LB? (At least it did not work for me).</li>\n</ul>\n<p>In the last week or so, I discovered something that might explain all this. It is all about how the bounding boxes are annotated. Because, this competition is not strictly about high recall in predicting the COTS. It is actually about predicting the annotations in the private LB. Specifically, it is predicting the locations and size of the ground truth (GT) annotated bboxes!</p>\n<p>I had made some benchmarks that tests this theory. I made a very simple post-inference \"trick\" - scaling all bboxes by a multiplier (but keeping the bbox centers the same). Here are my results:</p>\n<p>epoch10 model:<br>\nInference @ 1.0x training: </p>\n<ul>\n<li>bbox scale 1.0x: public LB (0.636), private LB (0.712)</li>\n<li>0.9x: public LB (0.725), private LB (0.685)</li>\n<li>0.875x: public LB (0.736), private LB (0.671)</li>\n</ul>\n<p>epoch12 model:<br>\nInference @ 1.0x training: </p>\n<ul>\n<li>bbox scale 0.95x: public LB (0.692), private LB (0.715)</li>\n<li>bbox scale 0.875x: public LB (0.737), private LB (0.673)</li>\n</ul>\n<p>I conclude that the GT bboxes in the public LB were smaller (i.e. tighter fit around the COTS) than those predicted in my model, and that simply scaling all the predicted bboxes achieved a better result because the dimensions of the scaled bboxes were closer to the GT bboxes in the public LB. HOWEVER, the GT bboxes in the private LB are NOT the same! In fact, I suspect they are larger than predicted by my epoch10 model. I should have tried 1.1x bbox scaling but i didn't.</p>\n<p>Then, this explains why inference at higher res than training improved public LB. If you examined the videos and the predicted bboxes, you may have noticed that inferencing at higher multiples resulted in tighter-fitting bboxes, especially for the larger COTS. Below my results:</p>\n<p>epoch10 model:<br>\nInference @ 1.0x training: public LB (0.636), private LB (0.712)<br>\nInference @ 1.3x training: public LB (0.750), private LB (0.702)<br>\nInference @ 1.4x training: public LB (0.756), private LB (0.690)<br>\nInference @ 1.5x training:</p>\n<ul>\n<li>bbox scaled 1.0x: public LB (0.765), private LB (0.668)</li>\n<li>0.95x: public LB (0.754), private LB (0.607)</li>\n<li>1.05x: public LB (0.723), private LB (0.668)</li>\n</ul>\n<p>Does this bbox scaling affect larger or smaller bboxes? I benched this also, using size-selective scaling:</p>\n<p>epoch10 model:<br>\nInference @ 1.0x training: </p>\n<ul>\n<li>bbox scale 1.0x: public LB (0.636), private LB (0.712)</li>\n<li>0.9x: public LB (0.725), private LB (0.685)</li>\n<li>0.9x only for predicted bboxes &lt; 40^2 in area: public LB (0.684), private LB (0.545)</li>\n<li>0.9x only for predicted bboxes &gt; 40^2 in area: public LB (0.665), private LB (0.696)</li>\n</ul>\n<p>This suggests that the public LB were tighter fitting for both large and small COTS. HOWEVER, the private LB results suggest that the smaller COTS were annotated more loosely (i.e. bigger bboxes, larger margins around the actual COTS), compared with the larger COTS. This is because 0.9x scaling of small bboxes did more damage (0.712 -&gt; 0.545), compared to 0.9x scaling of large bboxes (0.712 -&gt; 0.696).</p>\n<p>Given these results, to fit the private LB, I would inference @ 1.0x training, and scale the bboxes by 1.1x (only for the small COTS).</p>\n<p>(WIP placeholder - I may have more to say after I analyze all of my submissions)</p>\n<p>What do you guys think?</p>",
  "messages": [
    {
      "id": "1690484",
      "postDate": "02/15/2022 00:53:04",
      "content": "<p>There has been many discussions here about this. The key question is why does the CV and LB not match? Related questions are:</p>\n<ul>\n<li>Why does inference at 2-3x training resolution benefit the public LB (in my case, 1.5x for YOLOX) ?</li>\n<li>Why did this trick not work for private LB? (At least it did not work for me).</li>\n</ul>\n<p>In the last week or so, I discovered something that might explain all this. It is all about how the bounding boxes are annotated. Because, this competition is not strictly about high recall in predicting the COTS. It is actually about predicting the annotations in the private LB. Specifically, it is predicting the locations and size of the ground truth (GT) annotated bboxes!</p>\n<p>I had made some benchmarks that tests this theory. I made a very simple post-inference \"trick\" - scaling all bboxes by a multiplier (but keeping the bbox centers the same). Here are my results:</p>\n<p>epoch10 model:<br>\nInference @ 1.0x training: </p>\n<ul>\n<li>bbox scale 1.0x: public LB (0.636), private LB (0.712)</li>\n<li>0.9x: public LB (0.725), private LB (0.685)</li>\n<li>0.875x: public LB (0.736), private LB (0.671)</li>\n</ul>\n<p>epoch12 model:<br>\nInference @ 1.0x training: </p>\n<ul>\n<li>bbox scale 0.95x: public LB (0.692), private LB (0.715)</li>\n<li>bbox scale 0.875x: public LB (0.737), private LB (0.673)</li>\n</ul>\n<p>I conclude that the GT bboxes in the public LB were smaller (i.e. tighter fit around the COTS) than those predicted in my model, and that simply scaling all the predicted bboxes achieved a better result because the dimensions of the scaled bboxes were closer to the GT bboxes in the public LB. HOWEVER, the GT bboxes in the private LB are NOT the same! In fact, I suspect they are larger than predicted by my epoch10 model. I should have tried 1.1x bbox scaling but i didn't.</p>\n<p>Then, this explains why inference at higher res than training improved public LB. If you examined the videos and the predicted bboxes, you may have noticed that inferencing at higher multiples resulted in tighter-fitting bboxes, especially for the larger COTS. Below my results:</p>\n<p>epoch10 model:<br>\nInference @ 1.0x training: public LB (0.636), private LB (0.712)<br>\nInference @ 1.3x training: public LB (0.750), private LB (0.702)<br>\nInference @ 1.4x training: public LB (0.756), private LB (0.690)<br>\nInference @ 1.5x training:</p>\n<ul>\n<li>bbox scaled 1.0x: public LB (0.765), private LB (0.668)</li>\n<li>0.95x: public LB (0.754), private LB (0.607)</li>\n<li>1.05x: public LB (0.723), private LB (0.668)</li>\n</ul>\n<p>Does this bbox scaling affect larger or smaller bboxes? I benched this also, using size-selective scaling:</p>\n<p>epoch10 model:<br>\nInference @ 1.0x training: </p>\n<ul>\n<li>bbox scale 1.0x: public LB (0.636), private LB (0.712)</li>\n<li>0.9x: public LB (0.725), private LB (0.685)</li>\n<li>0.9x only for predicted bboxes &lt; 40^2 in area: public LB (0.684), private LB (0.545)</li>\n<li>0.9x only for predicted bboxes &gt; 40^2 in area: public LB (0.665), private LB (0.696)</li>\n</ul>\n<p>This suggests that the public LB were tighter fitting for both large and small COTS. HOWEVER, the private LB results suggest that the smaller COTS were annotated more loosely (i.e. bigger bboxes, larger margins around the actual COTS), compared with the larger COTS. This is because 0.9x scaling of small bboxes did more damage (0.712 -&gt; 0.545), compared to 0.9x scaling of large bboxes (0.712 -&gt; 0.696).</p>\n<p>Given these results, to fit the private LB, I would inference @ 1.0x training, and scale the bboxes by 1.1x (only for the small COTS).</p>\n<p>(WIP placeholder - I may have more to say after I analyze all of my submissions)</p>\n<p>What do you guys think?</p>",
      "rawMarkdown": "There has been many discussions here about this. The key question is why does the CV and LB not match? Related questions are:\n- Why does inference at 2-3x training resolution benefit the public LB (in my case, 1.5x for YOLOX) ?\n- Why did this trick not work for private LB? (At least it did not work for me).\n\nIn the last week or so, I discovered something that might explain all this. It is all about how the bounding boxes are annotated. Because, this competition is not strictly about high recall in predicting the COTS. It is actually about predicting the annotations in the private LB. Specifically, it is predicting the locations and size of the ground truth (GT) annotated bboxes!\n\nI had made some benchmarks that tests this theory. I made a very simple post-inference \"trick\" - scaling all bboxes by a multiplier (but keeping the bbox centers the same). Here are my results:\n\nepoch10 model:\nInference @ 1.0x training: \n- bbox scale 1.0x: public LB (0.636), private LB (0.712)\n- 0.9x: public LB (0.725), private LB (0.685)\n- 0.875x: public LB (0.736), private LB (0.671)\n\nepoch12 model:\nInference @ 1.0x training: \n- bbox scale 0.95x: public LB (0.692), private LB (0.715)\n- bbox scale 0.875x: public LB (0.737), private LB (0.673)\n\nI conclude that the GT bboxes in the public LB were smaller (i.e. tighter fit around the COTS) than those predicted in my model, and that simply scaling all the predicted bboxes achieved a better result because the dimensions of the scaled bboxes were closer to the GT bboxes in the public LB. HOWEVER, the GT bboxes in the private LB are NOT the same! In fact, I suspect they are larger than predicted by my epoch10 model. I should have tried 1.1x bbox scaling but i didn't.\n\nThen, this explains why inference at higher res than training improved public LB. If you examined the videos and the predicted bboxes, you may have noticed that inferencing at higher multiples resulted in tighter-fitting bboxes, especially for the larger COTS. Below my results:\n\nepoch10 model:\nInference @ 1.0x training: public LB (0.636), private LB (0.712)\nInference @ 1.3x training: public LB (0.750), private LB (0.702)\nInference @ 1.4x training: public LB (0.756), private LB (0.690)\nInference @ 1.5x training:\n- bbox scaled 1.0x: public LB (0.765), private LB (0.668)\n- 0.95x: public LB (0.754), private LB (0.607)\n- 1.05x: public LB (0.723), private LB (0.668)\n\nDoes this bbox scaling affect larger or smaller bboxes? I benched this also, using size-selective scaling:\n\nepoch10 model:\nInference @ 1.0x training: \n- bbox scale 1.0x: public LB (0.636), private LB (0.712)\n- 0.9x: public LB (0.725), private LB (0.685)\n- 0.9x only for predicted bboxes < 40^2 in area: public LB (0.684), private LB (0.545)\n- 0.9x only for predicted bboxes > 40^2 in area: public LB (0.665), private LB (0.696)\n\nThis suggests that the public LB were tighter fitting for both large and small COTS. HOWEVER, the private LB results suggest that the smaller COTS were annotated more loosely (i.e. bigger bboxes, larger margins around the actual COTS), compared with the larger COTS. This is because 0.9x scaling of small bboxes did more damage (0.712 -> 0.545), compared to 0.9x scaling of large bboxes (0.712 -> 0.696).\n\nGiven these results, to fit the private LB, I would inference @ 1.0x training, and scale the bboxes by 1.1x (only for the small COTS).\n\n(WIP placeholder - I may have more to say after I analyze all of my submissions)\n\nWhat do you guys think?",
      "votes": null
    },
    {
      "id": "1690933",
      "postDate": "02/15/2022 06:57:07",
      "content": "<p>I think focusing on  CV, multi-stage model or ensemble models would achieve the goal.  <br>\ninstead of purely fitting the metric.</p>",
      "rawMarkdown": "I think focusing on  CV, multi-stage model or ensemble models would achieve the goal.  \ninstead of purely fitting the metric.",
      "votes": null
    },
    {
      "id": "1691295",
      "postDate": "02/15/2022 10:36:55",
      "content": "<p>how about resubmitting for only predicted bboxes for certain sizes (only one size)?<br>\nmaybe this probing can solve the mystery</p>",
      "rawMarkdown": "how about resubmitting for only predicted bboxes for certain sizes (only one size)?\nmaybe this probing can solve the mystery",
      "votes": null
    },
    {
      "id": "1691478",
      "postDate": "02/15/2022 12:37:13",
      "content": "<p>Good advice! Use late submission to probe test data. Don't think there's any limit, since I noticed the late submission button on older competitions. </p>",
      "rawMarkdown": "Good advice! Use late submission to probe test data. Don't think there's any limit, since I noticed the late submission button on older competitions.",
      "votes": null
    },
    {
      "id": "1691552",
      "postDate": "02/15/2022 13:22:34",
      "content": "<p>Wow, great detective work. I also tried scaling bboxes. I tried 0.95x but i was also using 2.5x image size and the result lowered my public LB. It's probably a good thing that I didn't get this to work because if someone uses this on their 4 final submissions, it looks like it hurts on private LB.</p>",
      "rawMarkdown": "Wow, great detective work. I also tried scaling bboxes. I tried 0.95x but i was also using 2.5x image size and the result lowered my public LB. It's probably a good thing that I didn't get this to work because if someone uses this on their 4 final submissions, it looks like it hurts on private LB.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1690933,
      "author_name": "dragonzhang",
      "author_url": "",
      "post_date": "02/15/2022 06:57:07",
      "content": "<p>I think focusing on  CV, multi-stage model or ensemble models would achieve the goal.  <br>\ninstead of purely fitting the metric.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1691295,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "02/15/2022 10:36:55",
      "content": "<p>how about resubmitting for only predicted bboxes for certain sizes (only one size)?<br>\nmaybe this probing can solve the mystery</p>",
      "votes": null,
      "replies": [
        {
          "id": 1691478,
          "author_name": "liminchen1",
          "author_url": "",
          "post_date": "02/15/2022 12:37:13",
          "content": "<p>Good advice! Use late submission to probe test data. Don't think there's any limit, since I noticed the late submission button on older competitions. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1691552,
      "author_name": "cdeotte",
      "author_url": "",
      "post_date": "02/15/2022 13:22:34",
      "content": "<p>Wow, great detective work. I also tried scaling bboxes. I tried 0.95x but i was also using 2.5x image size and the result lowered my public LB. It's probably a good thing that I didn't get this to work because if someone uses this on their 4 final submissions, it looks like it hurts on private LB.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1690484": "There has been many discussions here about this. The key question is why does the CV and LB not match? Related questions are:\n- Why does inference at 2-3x training resolution benefit the public LB (in my case, 1.5x for YOLOX) ?\n- Why did this trick not work for private LB? (At least it did not work for me).\n\nIn the last week or so, I discovered something that might explain all this. It is all about how the bounding boxes are annotated. Because, this competition is not strictly about high recall in predicting the COTS. It is actually about predicting the annotations in the private LB. Specifically, it is predicting the locations and size of the ground truth (GT) annotated bboxes!\n\nI had made some benchmarks that tests this theory. I made a very simple post-inference \"trick\" - scaling all bboxes by a multiplier (but keeping the bbox centers the same). Here are my results:\n\nepoch10 model:\nInference @ 1.0x training: \n- bbox scale 1.0x: public LB (0.636), private LB (0.712)\n- 0.9x: public LB (0.725), private LB (0.685)\n- 0.875x: public LB (0.736), private LB (0.671)\n\nepoch12 model:\nInference @ 1.0x training: \n- bbox scale 0.95x: public LB (0.692), private LB (0.715)\n- bbox scale 0.875x: public LB (0.737), private LB (0.673)\n\nI conclude that the GT bboxes in the public LB were smaller (i.e. tighter fit around the COTS) than those predicted in my model, and that simply scaling all the predicted bboxes achieved a better result because the dimensions of the scaled bboxes were closer to the GT bboxes in the public LB. HOWEVER, the GT bboxes in the private LB are NOT the same! In fact, I suspect they are larger than predicted by my epoch10 model. I should have tried 1.1x bbox scaling but i didn't.\n\nThen, this explains why inference at higher res than training improved public LB. If you examined the videos and the predicted bboxes, you may have noticed that inferencing at higher multiples resulted in tighter-fitting bboxes, especially for the larger COTS. Below my results:\n\nepoch10 model:\nInference @ 1.0x training: public LB (0.636), private LB (0.712)\nInference @ 1.3x training: public LB (0.750), private LB (0.702)\nInference @ 1.4x training: public LB (0.756), private LB (0.690)\nInference @ 1.5x training:\n- bbox scaled 1.0x: public LB (0.765), private LB (0.668)\n- 0.95x: public LB (0.754), private LB (0.607)\n- 1.05x: public LB (0.723), private LB (0.668)\n\nDoes this bbox scaling affect larger or smaller bboxes? I benched this also, using size-selective scaling:\n\nepoch10 model:\nInference @ 1.0x training: \n- bbox scale 1.0x: public LB (0.636), private LB (0.712)\n- 0.9x: public LB (0.725), private LB (0.685)\n- 0.9x only for predicted bboxes < 40^2 in area: public LB (0.684), private LB (0.545)\n- 0.9x only for predicted bboxes > 40^2 in area: public LB (0.665), private LB (0.696)\n\nThis suggests that the public LB were tighter fitting for both large and small COTS. HOWEVER, the private LB results suggest that the smaller COTS were annotated more loosely (i.e. bigger bboxes, larger margins around the actual COTS), compared with the larger COTS. This is because 0.9x scaling of small bboxes did more damage (0.712 -> 0.545), compared to 0.9x scaling of large bboxes (0.712 -> 0.696).\n\nGiven these results, to fit the private LB, I would inference @ 1.0x training, and scale the bboxes by 1.1x (only for the small COTS).\n\n(WIP placeholder - I may have more to say after I analyze all of my submissions)\n\nWhat do you guys think?",
    "1690933": "I think focusing on  CV, multi-stage model or ensemble models would achieve the goal.  \ninstead of purely fitting the metric.",
    "1691295": "how about resubmitting for only predicted bboxes for certain sizes (only one size)?\nmaybe this probing can solve the mystery",
    "1691478": "Good advice! Use late submission to probe test data. Don't think there's any limit, since I noticed the late submission button on older competitions.",
    "1691552": "Wow, great detective work. I also tried scaling bboxes. I tried 0.95x but i was also using 2.5x image size and the result lowered my public LB. It's probably a good thing that I didn't get this to work because if someone uses this on their 4 final submissions, it looks like it hurts on private LB."
  },
  "source": "meta"
}