{
  "id": 298877,
  "title": "conf = 0.13 too small?",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/298877",
  "author_name": "",
  "post_date": "2022-01-05T04:53:33.318774900Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>The conf threshold I am currently using for the LB maximum (0.547, YOLOv5) is 0.13.<br>\nEven though F2score emphasizes Recall, I'm wondering if the number in this conf is too low. Overfitting to PublicLB?</p>",
  "messages": [
    {
      "id": "1638833",
      "postDate": "01/05/2022 04:53:33",
      "content": "<p>The conf threshold I am currently using for the LB maximum (0.547, YOLOv5) is 0.13.<br>\nEven though F2score emphasizes Recall, I'm wondering if the number in this conf is too low. Overfitting to PublicLB?</p>",
      "rawMarkdown": "The conf threshold I am currently using for the LB maximum (0.547, YOLOv5) is 0.13.\nEven though F2score emphasizes Recall, I'm wondering if the number in this conf is too low. Overfitting to PublicLB?",
      "votes": null
    },
    {
      "id": "1638947",
      "postDate": "01/05/2022 07:31:31",
      "content": "<p>The confidence threshold value itself isn’t a problem; it depends on the model (low conf threshold is no problem if the model can distinguish positive and negative bounding box well).<br>\nThe problem is, the parameter fitting your validation set has generally applicable to public/private test set.</p>\n<p>For me, the discussion below has much insight:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/298427\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/298427</a></p>\n<p>For other strategy, though it’s computationally intense, nested cross validation for fitting hyper parameter might reduce overfitting.</p>",
      "rawMarkdown": "The confidence threshold value itself isn’t a problem; it depends on the model (low conf threshold is no problem if the model can distinguish positive and negative bounding box well).\nThe problem is, the parameter fitting your validation set has generally applicable to public/private test set.\n\nFor me, the discussion below has much insight:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/298427\n\nFor other strategy, though it’s computationally intense, nested cross validation for fitting hyper parameter might reduce overfitting.",
      "votes": null
    },
    {
      "id": "1639110",
      "postDate": "01/05/2022 11:39:42",
      "content": "<p>the f2 formula is:</p>\n<pre><code>def f_beta(tp, fp, fn, beta=2):\n    f2 = (1 + beta ** 2) * tp / ((1 + beta ** 2) * tp + beta ** 2 * fn + fp)\n    # f2 = 5 * tp / (5 * tp + 4 * fn + fp)\n\nreturn f2\n</code></pre>\n<p>for example, if i lower the conf and i get an improvement of one tp (and a corresponding reduction of one fn) but a degradation one more fp, i still gain in f2. </p>\n<p>this is because of the beta multiplier in the formula, one improvement of tp becomes 5x tp.</p>\n<p>but if lower the conf leads to e.g 5x more fp than tp, the f2 score degrades.<br>\n(the actual number may not be 5x, you should use python to make a a graph f2 vs tp,trp,fn)</p>\n<p>for your trained model, you should have a graph of conf vs fp, conf vs tp , conf vs tn.<br>\nthen you should estimate be able to estimate the performance of your model in test (since you know how many test images and test objects)</p>",
      "rawMarkdown": "the f2 formula is:\n\n```\ndef f_beta(tp, fp, fn, beta=2):\n    f2 = (1 + beta ** 2) * tp / ((1 + beta ** 2) * tp + beta ** 2 * fn + fp)\n    # f2 = 5 * tp / (5 * tp + 4 * fn + fp)\n\nreturn f2\n\n```\n\nfor example, if i lower the conf and i get an improvement of one tp (and a corresponding reduction of one fn) but a degradation one more fp, i still gain in f2. \n\nthis is because of the beta multiplier in the formula, one improvement of tp becomes 5x tp.\n\nbut if lower the conf leads to e.g 5x more fp than tp, the f2 score degrades.\n(the actual number may not be 5x, you should use python to make a a graph f2 vs tp,trp,fn)\n\nfor your trained model, you should have a graph of conf vs fp, conf vs tp , conf vs tn.\nthen you should estimate be able to estimate the performance of your model in test (since you know how many test images and test objects)",
      "votes": null
    },
    {
      "id": "1639411",
      "postDate": "01/05/2022 16:02:16",
      "content": "<p>Thanks for the interesting comments.<br>\nI will read the discussion you shared.</p>",
      "rawMarkdown": "Thanks for the interesting comments.\nI will read the discussion you shared.",
      "votes": null
    },
    {
      "id": "1639417",
      "postDate": "01/05/2022 16:07:04",
      "content": "<p>thank you.<br>\nWhen evaluating the validation data, I would like to examine the relationship with conf.</p>",
      "rawMarkdown": "thank you.\nWhen evaluating the validation data, I would like to examine the relationship with conf.",
      "votes": null
    },
    {
      "id": "1639550",
      "postDate": "01/05/2022 18:15:26",
      "content": "<p>in theory, when detection is more important than fp, you can modify your binary cross entropy loss to weigh more on positive class.</p>\n<p>alternatively, you can increase the population of positive objects by cut and paster / mosaic augmentation for one image.</p>\n<p>then the optimal confidence threshold should increase in value.</p>",
      "rawMarkdown": "in theory, when detection is more important than fp, you can modify your binary cross entropy loss to weigh more on positive class.\n\nalternatively, you can increase the population of positive objects by cut and paster / mosaic augmentation for one image.\n\nthen the optimal confidence threshold should increase in value.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1638947,
      "author_name": "tatamikenn",
      "author_url": "",
      "post_date": "01/05/2022 07:31:31",
      "content": "<p>The confidence threshold value itself isn’t a problem; it depends on the model (low conf threshold is no problem if the model can distinguish positive and negative bounding box well).<br>\nThe problem is, the parameter fitting your validation set has generally applicable to public/private test set.</p>\n<p>For me, the discussion below has much insight:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/298427\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/298427</a></p>\n<p>For other strategy, though it’s computationally intense, nested cross validation for fitting hyper parameter might reduce overfitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1639411,
          "author_name": "syurenuko",
          "author_url": "",
          "post_date": "01/05/2022 16:02:16",
          "content": "<p>Thanks for the interesting comments.<br>\nI will read the discussion you shared.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1639110,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "01/05/2022 11:39:42",
      "content": "<p>the f2 formula is:</p>\n<pre><code>def f_beta(tp, fp, fn, beta=2):\n    f2 = (1 + beta ** 2) * tp / ((1 + beta ** 2) * tp + beta ** 2 * fn + fp)\n    # f2 = 5 * tp / (5 * tp + 4 * fn + fp)\n\nreturn f2\n</code></pre>\n<p>for example, if i lower the conf and i get an improvement of one tp (and a corresponding reduction of one fn) but a degradation one more fp, i still gain in f2. </p>\n<p>this is because of the beta multiplier in the formula, one improvement of tp becomes 5x tp.</p>\n<p>but if lower the conf leads to e.g 5x more fp than tp, the f2 score degrades.<br>\n(the actual number may not be 5x, you should use python to make a a graph f2 vs tp,trp,fn)</p>\n<p>for your trained model, you should have a graph of conf vs fp, conf vs tp , conf vs tn.<br>\nthen you should estimate be able to estimate the performance of your model in test (since you know how many test images and test objects)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1639417,
          "author_name": "syurenuko",
          "author_url": "",
          "post_date": "01/05/2022 16:07:04",
          "content": "<p>thank you.<br>\nWhen evaluating the validation data, I would like to examine the relationship with conf.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1639550,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "01/05/2022 18:15:26",
          "content": "<p>in theory, when detection is more important than fp, you can modify your binary cross entropy loss to weigh more on positive class.</p>\n<p>alternatively, you can increase the population of positive objects by cut and paster / mosaic augmentation for one image.</p>\n<p>then the optimal confidence threshold should increase in value.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1638833": "The conf threshold I am currently using for the LB maximum (0.547, YOLOv5) is 0.13.\nEven though F2score emphasizes Recall, I'm wondering if the number in this conf is too low. Overfitting to PublicLB?",
    "1638947": "The confidence threshold value itself isn’t a problem; it depends on the model (low conf threshold is no problem if the model can distinguish positive and negative bounding box well).\nThe problem is, the parameter fitting your validation set has generally applicable to public/private test set.\n\nFor me, the discussion below has much insight:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/298427\n\nFor other strategy, though it’s computationally intense, nested cross validation for fitting hyper parameter might reduce overfitting.",
    "1639110": "the f2 formula is:\n\n```\ndef f_beta(tp, fp, fn, beta=2):\n    f2 = (1 + beta ** 2) * tp / ((1 + beta ** 2) * tp + beta ** 2 * fn + fp)\n    # f2 = 5 * tp / (5 * tp + 4 * fn + fp)\n\nreturn f2\n\n```\n\nfor example, if i lower the conf and i get an improvement of one tp (and a corresponding reduction of one fn) but a degradation one more fp, i still gain in f2. \n\nthis is because of the beta multiplier in the formula, one improvement of tp becomes 5x tp.\n\nbut if lower the conf leads to e.g 5x more fp than tp, the f2 score degrades.\n(the actual number may not be 5x, you should use python to make a a graph f2 vs tp,trp,fn)\n\nfor your trained model, you should have a graph of conf vs fp, conf vs tp , conf vs tn.\nthen you should estimate be able to estimate the performance of your model in test (since you know how many test images and test objects)",
    "1639411": "Thanks for the interesting comments.\nI will read the discussion you shared.",
    "1639417": "thank you.\nWhen evaluating the validation data, I would like to examine the relationship with conf.",
    "1639550": "in theory, when detection is more important than fp, you can modify your binary cross entropy loss to weigh more on positive class.\n\nalternatively, you can increase the population of positive objects by cut and paster / mosaic augmentation for one image.\n\nthen the optimal confidence threshold should increase in value."
  },
  "source": "meta"
}