{
  "id": 306143,
  "title": "About the Uncertainty of the Winning Model",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/306143",
  "author_name": "Bilzard",
  "post_date": "2022-02-08T09:42:05.270000",
  "votes": 9,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I have a concern about this competitions evaluation will lead to overfitting to test dataset. In other words, the model which happened to just fit the private test data is likely to win.</p>\n<p>The rationale for this is as follows.</p>\n<ol>\n<li>First, the competition metrics is based on F2 score which depends on FP and FN to the test dataset.</li>\n<li>Suppose number of FPs is proportional to number of background frames, the F2 score depends on the ratio of background frames in the test set.</li>\n<li>No one knows about the background frame ratio of the test set. Some model will fit to the hidden test set, and other ones will not.</li>\n<li>Thus, the model which happened to fit to the background ratio of hidden test data likely to win</li>\n</ol>\n<p>In my opinion, if we are under such uncertain condition, the host might have had these options:</p>\n<ol>\n<li>to design evaluation metrics which is more robust to background frame ratio (I'm not sure, but mAP might have been better.)</li>\n<li>to disclose the FP ratio of the hidden test set (and other statics of hidden data). This will make situation even for everyone, and will reduce the risk that a model that happens to be lucky will win.</li>\n</ol>\n<p>I don't intend to write this post to demanding something during this competition since the competition is almost close to end.<br>\nIf the people organizing this competition read this post, I hope they will find it useful in organizing future competitions.</p>\n<p>Also, I welcome the opposing opinion which point out the flaw of the above discussion.</p>\n<h1>Summary</h1>\n<p><strong>Bad Thing</strong></p>\n<ul>\n<li>The final score of a model depends on the background ratio of hidden test data. That means a model that happened to fit to the background ratio of hidden test data likely to win. We can make this risk as small as possible, but we can't reduce it to zero.</li>\n</ul>\n<p><strong>Good Thing</strong></p>\n<ul>\n<li>No one can probe private data (because most of the private test frame comes after the public frames and predictions per frame is limited up to 100). That means the condition is even for everyone.</li>\n</ul>",
  "messages": [
    {
      "id": 1681178,
      "postDate": "2022-02-08T09:42:05.270Z",
      "content": "<p>I have a concern about this competitions evaluation will lead to overfitting to test dataset. In other words, the model which happened to just fit the private test data is likely to win.</p>\n<p>The rationale for this is as follows.</p>\n<ol>\n<li>First, the competition metrics is based on F2 score which depends on FP and FN to the test dataset.</li>\n<li>Suppose number of FPs is proportional to number of background frames, the F2 score depends on the ratio of background frames in the test set.</li>\n<li>No one knows about the background frame ratio of the test set. Some model will fit to the hidden test set, and other ones will not.</li>\n<li>Thus, the model which happened to fit to the background ratio of hidden test data likely to win</li>\n</ol>\n<p>In my opinion, if we are under such uncertain condition, the host might have had these options:</p>\n<ol>\n<li>to design evaluation metrics which is more robust to background frame ratio (I'm not sure, but mAP might have been better.)</li>\n<li>to disclose the FP ratio of the hidden test set (and other statics of hidden data). This will make situation even for everyone, and will reduce the risk that a model that happens to be lucky will win.</li>\n</ol>\n<p>I don't intend to write this post to demanding something during this competition since the competition is almost close to end.<br>\nIf the people organizing this competition read this post, I hope they will find it useful in organizing future competitions.</p>\n<p>Also, I welcome the opposing opinion which point out the flaw of the above discussion.</p>\n<h1>Summary</h1>\n<p><strong>Bad Thing</strong></p>\n<ul>\n<li>The final score of a model depends on the background ratio of hidden test data. That means a model that happened to fit to the background ratio of hidden test data likely to win. We can make this risk as small as possible, but we can't reduce it to zero.</li>\n</ul>\n<p><strong>Good Thing</strong></p>\n<ul>\n<li>No one can probe private data (because most of the private test frame comes after the public frames and predictions per frame is limited up to 100). That means the condition is even for everyone.</li>\n</ul>",
      "rawMarkdown": "I have a concern about this competitions evaluation will lead to overfitting to test dataset. In other words, the model which happened to just fit the private test data is likely to win.\n\nThe rationale for this is as follows.\n\n1. First, the competition metrics is based on F2 score which depends on FP and FN to the test dataset.\n2. Suppose number of FPs is proportional to number of background frames, the F2 score depends on the ratio of background frames in the test set.\n3. No one knows about the background frame ratio of the test set. Some model will fit to the hidden test set, and other ones will not.\n4. Thus, the model which happened to fit to the background ratio of hidden test data likely to win\n\nIn my opinion, if we are under such uncertain condition, the host might have had these options:\n\n1. to design evaluation metrics which is more robust to background frame ratio (I'm not sure, but mAP might have been better.)\n2. to disclose the FP ratio of the hidden test set (and other statics of hidden data). This will make situation even for everyone, and will reduce the risk that a model that happens to be lucky will win.\n\nI don't intend to write this post to demanding something during this competition since the competition is almost close to end.\nIf the people organizing this competition read this post, I hope they will find it useful in organizing future competitions.\n\nAlso, I welcome the opposing opinion which point out the flaw of the above discussion.\n\n# Summary\n\n**Bad Thing**\n\n* The final score of a model depends on the background ratio of hidden test data. That means a model that happened to fit to the background ratio of hidden test data likely to win. We can make this risk as small as possible, but we can't reduce it to zero.\n\n**Good Thing**\n\n* No one can probe private data (because most of the private test frame comes after the public frames and predictions per frame is limited up to 100). That means the condition is even for everyone.",
      "votes": 9
    },
    {
      "id": 1687028,
      "postDate": "2022-02-12T14:53:34.390Z",
      "content": "<p>I think the shake-up will not be as big as you think in terms of LB rankings<br>\nanyway most of us use large scale inference</p>",
      "rawMarkdown": "I think the shake-up will not be as big as you think in terms of LB rankings\nanyway most of us use large scale inference\n\n",
      "votes": 1
    },
    {
      "id": 1681210,
      "postDate": "2022-02-08T10:18:46.490Z",
      "content": "<p>After saw shake up in Jigsaw competition, as a newbie idk what's going on. But how you can ensure your model have the generalization level enough but not overfitting public lb? Like you trust your CV mostly and lb for just testing public lb?</p>",
      "rawMarkdown": "After saw shake up in Jigsaw competition, as a newbie idk what's going on. But how you can ensure your model have the generalization level enough but not overfitting public lb? Like you trust your CV mostly and lb for just testing public lb?",
      "replies": [
        {
          "id": 1681214,
          "postDate": "2022-02-08T10:25:44.960Z",
          "content": "<p>Thank you for commenting.</p>\n<p>I'm not sure the best way, but what I'm currently planning is to make submission with models which are trained under several background frame ratio.<br>\nFortunately, we have 4 submission in this competition, which allows us to hedge our risks better.</p>",
          "rawMarkdown": "Thank you for commenting.\n\nI'm not sure the best way, but what I'm currently planning is to make submission with models which are trained under several background frame ratio.\nFortunately, we have 4 submission in this competition, which allows us to hedge our risks better.",
          "votes": 1
        },
        {
          "id": 1681217,
          "postDate": "2022-02-08T10:31:40.907Z",
          "content": "<p>I think the best way, if possible, is to create a model that is robust to the background frame ratio without sacrificing the F2 score. However, I think there is a trade-off between F2 and background ratio.</p>",
          "rawMarkdown": "I think the best way, if possible, is to create a model that is robust to the background frame ratio without sacrificing the F2 score. However, I think there is a trade-off between F2 and background ratio.",
          "votes": 1
        },
        {
          "id": 1681230,
          "postDate": "2022-02-08T10:52:38.643Z",
          "content": "<p>Note: here I only discussed about the risk of variating background ratio.<br>\nOf course we have to hedge other risks like variating target size distribution, different background textures, unseen confusing objects etc.</p>",
          "rawMarkdown": "Note: here I only discussed about the risk of variating background ratio.\nOf course we have to hedge other risks like variating target size distribution, different background textures, unseen confusing objects etc.",
          "votes": 1
        },
        {
          "id": 1681537,
          "postDate": "2022-02-08T14:39:55.367Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1686092,
          "postDate": "2022-02-11T18:46:26.690Z",
          "rawMarkdown": "",
          "votes": -2,
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1687028,
      "author_name": "kaggler",
      "author_url": "",
      "post_date": "2022-02-12T14:53:34.390000",
      "content": "<p>I think the shake-up will not be as big as you think in terms of LB rankings<br>\nanyway most of us use large scale inference</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1681210,
      "author_name": "Bao Loc Pham",
      "author_url": "",
      "post_date": "2022-02-08T10:18:46.490000",
      "content": "<p>After saw shake up in Jigsaw competition, as a newbie idk what's going on. But how you can ensure your model have the generalization level enough but not overfitting public lb? Like you trust your CV mostly and lb for just testing public lb?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1681214,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-08T10:25:44.960000",
          "content": "<p>Thank you for commenting.</p>\n<p>I'm not sure the best way, but what I'm currently planning is to make submission with models which are trained under several background frame ratio.<br>\nFortunately, we have 4 submission in this competition, which allows us to hedge our risks better.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1681217,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-08T10:31:40.907000",
          "content": "<p>I think the best way, if possible, is to create a model that is robust to the background frame ratio without sacrificing the F2 score. However, I think there is a trade-off between F2 and background ratio.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1681230,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-02-08T10:52:38.643000",
          "content": "<p>Note: here I only discussed about the risk of variating background ratio.<br>\nOf course we have to hedge other risks like variating target size distribution, different background textures, unseen confusing objects etc.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1681537,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-08T14:39:55.367000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1686092,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-02-11T18:46:26.690000",
          "content": "",
          "votes": -2,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1681178": "I have a concern about this competitions evaluation will lead to overfitting to test dataset. In other words, the model which happened to just fit the private test data is likely to win.\n\nThe rationale for this is as follows.\n\n1. First, the competition metrics is based on F2 score which depends on FP and FN to the test dataset.\n2. Suppose number of FPs is proportional to number of background frames, the F2 score depends on the ratio of background frames in the test set.\n3. No one knows about the background frame ratio of the test set. Some model will fit to the hidden test set, and other ones will not.\n4. Thus, the model which happened to fit to the background ratio of hidden test data likely to win\n\nIn my opinion, if we are under such uncertain condition, the host might have had these options:\n\n1. to design evaluation metrics which is more robust to background frame ratio (I'm not sure, but mAP might have been better.)\n2. to disclose the FP ratio of the hidden test set (and other statics of hidden data). This will make situation even for everyone, and will reduce the risk that a model that happens to be lucky will win.\n\nI don't intend to write this post to demanding something during this competition since the competition is almost close to end.\nIf the people organizing this competition read this post, I hope they will find it useful in organizing future competitions.\n\nAlso, I welcome the opposing opinion which point out the flaw of the above discussion.\n\n# Summary\n\n**Bad Thing**\n\n* The final score of a model depends on the background ratio of hidden test data. That means a model that happened to fit to the background ratio of hidden test data likely to win. We can make this risk as small as possible, but we can't reduce it to zero.\n\n**Good Thing**\n\n* No one can probe private data (because most of the private test frame comes after the public frames and predictions per frame is limited up to 100). That means the condition is even for everyone.",
    "1687028": "I think the shake-up will not be as big as you think in terms of LB rankings\nanyway most of us use large scale inference\n\n",
    "1681210": "After saw shake up in Jigsaw competition, as a newbie idk what's going on. But how you can ensure your model have the generalization level enough but not overfitting public lb? Like you trust your CV mostly and lb for just testing public lb?"
  }
}