{
  "id": 302632,
  "title": "[Solved] Something Odd in the LB Evaluation - 10K of Additional FPs Affect Little on the LB Score",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/302632",
  "author_name": "Bilzard",
  "post_date": "2022-01-23T13:10:05.743000",
  "votes": 46,
  "comment_count": 29,
  "views": 0,
  "content": "<h1>In Short</h1>\n<p>I added 10,000 FP bbox to one frame, and submit this model. After calculated LB, the score is dropped 0.006 (0.601 -&gt; 0.595). Given the total frame number is around 3,500 (0.25 * 14,000), and number of GT label is ~2450[1], the score drop is apparently too little.</p>\n<p>On the contrary, if I add one FP bbox on every frame, the score is dropped as expected (0.601 -&gt; 0.426).</p>\n<p></p>\n<h1>Assumed Cause</h1>\n<p></p>\n<p>We have now highly rational explanation on this issue.<br>\nIt seems that the box threshold per image is set to ~129 to avoid system load.</p>\n<p>Details are explained in this thread:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980</a></p>\n<h1>Actual Cause</h1>\n<p>These are clarified by Kaggle staff:</p>\n<blockquote>\n  <p>No, the F2 score is calculated globally.<br>\n  There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1663027\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1663027</a></p>\n<h1>Actions</h1>\n<p></p>\n<p></p>\n<p></p>\n<p></p>\n<h1>Code</h1>\n<p>The test code I used is as below:</p>\n<pre><code>P_KEEP = 1.0\nCONF = 0.50\nFP_COUNT = 10_000\n</code></pre>\n<pre><code>def drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp &lt;= p_keep]\n    confs = confs[pp &lt;= p_keep]\n    return bboxes, confs\n\ndef add_fp(bboxes, confs, fp_count=0, fp_conf=1e-4):\n    '''\n    add false negative to the frame\n    '''\n    if fp_count == 0:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    FPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n    bboxes = np.append(bboxes, FPs, axis=0)\n    assert bboxes.dtype == int\n\n    # we need to set sufficiently small confidence value as for not to affect TP\n    conf_FPs = [fp_conf] * fp_count\n    confs = np.append(confs, conf_FPs)\n    assert confs.dtype == float\n    return bboxes, confs    \n\n\ndef predict(model, img, size=768, augment=False, p_keep=1.0):\n    height, width = img.shape[:2]\n    results = model(img, size=size, augment=augment)  # custom inference size\n    preds   = results.pandas().xyxy[0]\n    bboxes  = preds[['xmin','ymin','xmax','ymax']].values\n    if len(bboxes):\n        bboxes  = voc2coco(bboxes, height, width).astype(int)\n        confs   = preds.confidence.values\n\n        bboxes, confs = drop_pred(bboxes, confs, p_keep=p_keep)\n        return bboxes, confs\n    else:\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n        return bboxes, confs\n\ndef format_prediction(bboxes, confs):\n    annot = ''\n    if len(bboxes) &gt; 0:\n        for idx in range(len(bboxes)):\n            xmin, ymin, w, h = bboxes[idx]\n            conf = confs[idx]\n            annot += f'{conf:.8f} {xmin} {ymin} {w} {h}'\n            annot +=' '\n        annot = annot.strip(' ')\n    return annot\n</code></pre>\n<pre><code>for idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=5200 is possible to be a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5200\n    if idx == fp_add_frame:\n        bboxes, confs = add_fp(bboxes, confs, fp_count=FP_COUNT)\n        print(f\"{len(bboxes)}, {len(confs)}\")\n        print(f\"{bboxes[0]}, {confs[0]}\")\n\n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n    if idx &lt; 3:\n        display(show_img(img, bboxes, bbox_format='coco'))\n</code></pre>\n<h1>Reference</h1>\n<p>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302605\" target=\"_blank\">LB probing result: GT Labels per Frame in the Public LB</a></p>\n<h1>Update Note</h1>\n<ul>\n<li>2022/1/24 fix typo FN-&gt;FP</li>\n<li>2022/1/25 actual cause is clarified by Kaggle staff</li>\n</ul>",
  "messages": [
    {
      "id": 1661417,
      "postDate": "2022-01-23T13:10:05.743Z",
      "content": "<h1>In Short</h1>\n<p>I added 10,000 FP bbox to one frame, and submit this model. After calculated LB, the score is dropped 0.006 (0.601 -&gt; 0.595). Given the total frame number is around 3,500 (0.25 * 14,000), and number of GT label is ~2450[1], the score drop is apparently too little.</p>\n<p>On the contrary, if I add one FP bbox on every frame, the score is dropped as expected (0.601 -&gt; 0.426).</p>\n<p></p>\n<h1>Assumed Cause</h1>\n<p></p>\n<p>We have now highly rational explanation on this issue.<br>\nIt seems that the box threshold per image is set to ~129 to avoid system load.</p>\n<p>Details are explained in this thread:<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980</a></p>\n<h1>Actual Cause</h1>\n<p>These are clarified by Kaggle staff:</p>\n<blockquote>\n  <p>No, the F2 score is calculated globally.<br>\n  There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?</p>\n</blockquote>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1663027\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1663027</a></p>\n<h1>Actions</h1>\n<p></p>\n<p></p>\n<p></p>\n<p></p>\n<h1>Code</h1>\n<p>The test code I used is as below:</p>\n<pre><code>P_KEEP = 1.0\nCONF = 0.50\nFP_COUNT = 10_000\n</code></pre>\n<pre><code>def drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp &lt;= p_keep]\n    confs = confs[pp &lt;= p_keep]\n    return bboxes, confs\n\ndef add_fp(bboxes, confs, fp_count=0, fp_conf=1e-4):\n    '''\n    add false negative to the frame\n    '''\n    if fp_count == 0:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    FPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n    bboxes = np.append(bboxes, FPs, axis=0)\n    assert bboxes.dtype == int\n\n    # we need to set sufficiently small confidence value as for not to affect TP\n    conf_FPs = [fp_conf] * fp_count\n    confs = np.append(confs, conf_FPs)\n    assert confs.dtype == float\n    return bboxes, confs    \n\n\ndef predict(model, img, size=768, augment=False, p_keep=1.0):\n    height, width = img.shape[:2]\n    results = model(img, size=size, augment=augment)  # custom inference size\n    preds   = results.pandas().xyxy[0]\n    bboxes  = preds[['xmin','ymin','xmax','ymax']].values\n    if len(bboxes):\n        bboxes  = voc2coco(bboxes, height, width).astype(int)\n        confs   = preds.confidence.values\n\n        bboxes, confs = drop_pred(bboxes, confs, p_keep=p_keep)\n        return bboxes, confs\n    else:\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n        return bboxes, confs\n\ndef format_prediction(bboxes, confs):\n    annot = ''\n    if len(bboxes) &gt; 0:\n        for idx in range(len(bboxes)):\n            xmin, ymin, w, h = bboxes[idx]\n            conf = confs[idx]\n            annot += f'{conf:.8f} {xmin} {ymin} {w} {h}'\n            annot +=' '\n        annot = annot.strip(' ')\n    return annot\n</code></pre>\n<pre><code>for idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=5200 is possible to be a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5200\n    if idx == fp_add_frame:\n        bboxes, confs = add_fp(bboxes, confs, fp_count=FP_COUNT)\n        print(f\"{len(bboxes)}, {len(confs)}\")\n        print(f\"{bboxes[0]}, {confs[0]}\")\n\n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n    if idx &lt; 3:\n        display(show_img(img, bboxes, bbox_format='coco'))\n</code></pre>\n<h1>Reference</h1>\n<p>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302605\" target=\"_blank\">LB probing result: GT Labels per Frame in the Public LB</a></p>\n<h1>Update Note</h1>\n<ul>\n<li>2022/1/24 fix typo FN-&gt;FP</li>\n<li>2022/1/25 actual cause is clarified by Kaggle staff</li>\n</ul>",
      "rawMarkdown": "# In Short\n\nI added 10,000 FP bbox to one frame, and submit this model. After calculated LB, the score is dropped 0.006 (0.601 -> 0.595). Given the total frame number is around 3,500 (0.25 * 14,000), and number of GT label is ~2450[1], the score drop is apparently too little.\n\nOn the contrary, if I add one FP bbox on every frame, the score is dropped as expected (0.601 -> 0.426).\n\n~~I suspect the algorithm used in the leaderboard evaluation code contains bugs or, at least, it might not be as expected for us.~~\n\n# Assumed Cause\n\n~~One hypothesis I came up with is the F2 score is calculated per image, and then average each of them. However, if it is true, we can't reasonably define F2 score for the background (no-annotation) frames.\nAnother possibility is skipping background image and only calculating for annotated frame. However, if it is true, the model which produces a lot of FP wouldn't be considered properly with this algorithm.~~\n\nWe have now highly rational explanation on this issue.\nIt seems that the box threshold per image is set to ~129 to avoid system load.\n\nDetails are explained in this thread:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\n\n# Actual Cause\n\nThese are clarified by Kaggle staff:\n\n> No, the F2 score is calculated globally.\n> There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?\n\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1663027\n\n# Actions\n\n~~Kagglers, please share ideas -- possible cause, reasonable explanations and the way to test it.\nAnyone who is kindly to reproduce this issue is also welcome.~~\n\n~~However, considering it is close to the competition end, the best way is to ask Kaggle staff to disclose LB calculation code, unless we have to consume a lot of time to test and resolve for the issue.If the kaggle staffs are reading this post, please consider for that.~~\n\n~~I hope this would be just my misunderstanding or my code bug.~~\n\n~~I'm now waiting for Kaggle staff officially replies to this post.~~\n\n# Code\n\nThe test code I used is as below:\n\n```\nP_KEEP = 1.0\nCONF = 0.50\nFP_COUNT = 10_000\n```\n\n```\ndef drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp <= p_keep]\n    confs = confs[pp <= p_keep]\n    return bboxes, confs\n\ndef add_fp(bboxes, confs, fp_count=0, fp_conf=1e-4):\n    '''\n    add false negative to the frame\n    '''\n    if fp_count == 0:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    FPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n    bboxes = np.append(bboxes, FPs, axis=0)\n    assert bboxes.dtype == int\n    \n    # we need to set sufficiently small confidence value as for not to affect TP\n    conf_FPs = [fp_conf] * fp_count\n    confs = np.append(confs, conf_FPs)\n    assert confs.dtype == float\n    return bboxes, confs    \n    \n\ndef predict(model, img, size=768, augment=False, p_keep=1.0):\n    height, width = img.shape[:2]\n    results = model(img, size=size, augment=augment)  # custom inference size\n    preds   = results.pandas().xyxy[0]\n    bboxes  = preds[['xmin','ymin','xmax','ymax']].values\n    if len(bboxes):\n        bboxes  = voc2coco(bboxes, height, width).astype(int)\n        confs   = preds.confidence.values\n        \n        bboxes, confs = drop_pred(bboxes, confs, p_keep=p_keep)\n        return bboxes, confs\n    else:\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n        return bboxes, confs\n    \ndef format_prediction(bboxes, confs):\n    annot = ''\n    if len(bboxes) > 0:\n        for idx in range(len(bboxes)):\n            xmin, ymin, w, h = bboxes[idx]\n            conf = confs[idx]\n            annot += f'{conf:.8f} {xmin} {ymin} {w} {h}'\n            annot +=' '\n        annot = annot.strip(' ')\n    return annot\n```\n\n```\nfor idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=5200 is possible to be a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5200\n    if idx == fp_add_frame:\n        bboxes, confs = add_fp(bboxes, confs, fp_count=FP_COUNT)\n        print(f\"{len(bboxes)}, {len(confs)}\")\n        print(f\"{bboxes[0]}, {confs[0]}\")\n    \n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n    if idx < 3:\n        display(show_img(img, bboxes, bbox_format='coco'))\n```\n\n# Reference\n\n[1] [LB probing result: GT Labels per Frame in the Public LB](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302605)\n\n# Update Note\n\n- 2022/1/24 fix typo FN->FP\n- 2022/1/25 actual cause is clarified by Kaggle staff",
      "votes": 46
    },
    {
      "id": 1661980,
      "postDate": "2022-01-23T22:15:14.407Z",
      "content": "<p>I think they just cut-off too many predicted boxes per image to not overload the system. The drop in your score would be smaller if it would be local metric. The only explanation I can see is that it is cut-off. Or maybe you have duplicate boxes they drop, I did not check your code. All this is just a guess.</p>",
      "rawMarkdown": "I think they just cut-off too many predicted boxes per image to not overload the system. The drop in your score would be smaller if it would be local metric. The only explanation I can see is that it is cut-off. Or maybe you have duplicate boxes they drop, I did not check your code. All this is just a guess.",
      "votes": 6,
      "replies": [
        {
          "id": 1661999,
          "postDate": "2022-01-23T23:19:51.210Z",
          "content": "<p>Thank you for pointing this out.</p>\n<p>I had thought of this once, but never delved into it deeply.</p>\n<p>We can estimate the cutoff threshold by the following equation [1].<br>\nα means the number of FPs added to the prediction.</p>\n<p>$$<br>\n\\frac{TP}{\\alpha} = \\frac{1}{5} \\left( \\frac{1}{F_2^\\alpha} - \\frac{1}{F_2} \\right)^{-1} \\tag{5}<br>\n$$</p>\n<p>Using the same formula as adding one FP to every image, I estimated the TP score to be ~1538.<br>\nSubstituting this into equation (5), the estimated value of α for the cutoff for each image is ~129. This is a very likely value.</p>\n<p>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156</a></p>",
          "rawMarkdown": "Thank you for pointing this out.\n\nI had thought of this once, but never delved into it deeply.\n\nWe can estimate the cutoff threshold by the following equation [1].\nα means the number of FPs added to the prediction.\n\n$$\n\\frac{TP}{\\alpha} = \\frac{1}{5} \\left( \\frac{1}{F_2^\\alpha} - \\frac{1}{F_2} \\right)^{-1} \\tag{5}\n$$\n\nUsing the same formula as adding one FP to every image, I estimated the TP score to be ~1538.\nSubstituting this into equation (5), the estimated value of α for the cutoff for each image is ~129. This is a very likely value.\n\n[1] https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156",
          "votes": 3
        },
        {
          "id": 1662011,
          "postDate": "2022-01-23T23:39:06.630Z",
          "content": "<blockquote>\n  <p>Or maybe you have duplicate boxes they drop</p>\n</blockquote>\n<p>I took duplicated boxes into account. The above code generates unique boxes.</p>",
          "rawMarkdown": "> Or maybe you have duplicate boxes they drop\n\nI took duplicated boxes into account. The above code generates unique boxes.",
          "votes": 1
        },
        {
          "id": 1662347,
          "postDate": "2022-01-24T07:46:54.110Z",
          "content": "<p>so probably maxdets=100 :)</p>",
          "rawMarkdown": "so probably maxdets=100 :)",
          "votes": 1
        },
        {
          "id": 1663022,
          "postDate": "2022-01-24T18:02:59.850Z",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> is correct. Capping the number of predictions per image let us use simple pairwise comparisons in the metric without worrying about excessive scoring runtimes. </p>",
          "rawMarkdown": "@philippsinger is correct. Capping the number of predictions per image let us use simple pairwise comparisons in the metric without worrying about excessive scoring runtimes. ",
          "votes": 3
        },
        {
          "id": 1663299,
          "postDate": "2022-01-25T00:47:34.503Z",
          "content": "<p>What is the criteria for cutting off? If it is by score, what happens if box are of the same score? </p>\n<p>Say of i submit 1000  prediction box of same score1.0, the evalution selects The first 100 or the best 100?  </p>\n<p>It is better to set a rule that at most 100 box predicted per image in this case</p>",
          "rawMarkdown": "What is the criteria for cutting off? If it is by score, what happens if box are of the same score? \n\nSay of i submit 1000  prediction box of same score1.0, the evalution selects The first 100 or the best 100?  \n\n\nIt is better to set a rule that at most 100 box predicted per image in this case",
          "votes": 1
        },
        {
          "id": 1664600,
          "postDate": "2022-01-26T05:01:12.133Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Hi, how about the question <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> asked above? I want to know the answer too.</p>",
          "rawMarkdown": "@sohier Hi, how about the question @hengck23 asked above? I want to know the answer too."
        }
      ]
    },
    {
      "id": 1661440,
      "postDate": "2022-01-23T13:21:05.573Z",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nDear Kaggle staffs, please read through the post and consider to disclose the evaluation code, unless we have to spend a lot of time to solve this issue. Considering the competition end is closing, it would be tough to all of us.</p>",
      "rawMarkdown": "@addisonhoward @sohier \nDear Kaggle staffs, please read through the post and consider to disclose the evaluation code, unless we have to spend a lot of time to solve this issue. Considering the competition end is closing, it would be tough to all of us.",
      "votes": 6,
      "replies": [
        {
          "id": 1661974,
          "postDate": "2022-01-23T22:01:46.460Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 1661993,
          "postDate": "2022-01-23T23:05:11.253Z",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Some quick questions:</p>\n<ul>\n<li>does evaluation conducted per image, i.e. first calculate F2 per image, and then mean them to get the final score?</li>\n<li>If the above question is right, how can you define F2 on the background (non-annotated) frame?</li>\n<li>Do you use cutoff too many bounding box per image? If it is, what is the threshold?</li>\n<li>How cutoff is conducted?  Predictions with lowest confidence values?</li>\n</ul>",
          "rawMarkdown": "@addisonhoward @sohier \n\nSome quick questions:\n\n* does evaluation conducted per image, i.e. first calculate F2 per image, and then mean them to get the final score?\n* If the above question is right, how can you define F2 on the background (non-annotated) frame?\n* Do you use cutoff too many bounding box per image? If it is, what is the threshold?\n* How cutoff is conducted?  Predictions with lowest confidence values?"
        },
        {
          "id": 1663027,
          "postDate": "2022-01-24T18:06:29.547Z",
          "content": "<ul>\n<li>No, the F2 score is calculated globally. </li>\n<li>There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?</li>\n</ul>",
          "rawMarkdown": "- No, the F2 score is calculated globally. \n- There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?",
          "votes": 3
        },
        {
          "id": 1663116,
          "postDate": "2022-01-24T19:50:05.007Z",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thank you for clarification.</p>",
          "rawMarkdown": "@sohier Thank you for clarification."
        }
      ]
    },
    {
      "id": 1662294,
      "postDate": "2022-01-24T07:01:54.533Z",
      "content": "<p>I agree we found it, some weeks ago but couldn't prove it. Lots of better models got worse score. We made FP reduction manuever but still score didn't increased much or sometimes dropped. Second observation everytime we jump inference resolution val F2 always drops but on LB almost always explodes to higher value. </p>",
      "rawMarkdown": "I agree we found it, some weeks ago but couldn't prove it. Lots of better models got worse score. We made FP reduction manuever but still score didn't increased much or sometimes dropped. Second observation everytime we jump inference resolution val F2 always drops but on LB almost always explodes to higher value. ",
      "votes": 1,
      "replies": [
        {
          "id": 1662338,
          "postDate": "2022-01-24T07:40:34.773Z",
          "content": "<blockquote>\n  <p>Lots of better models got worse score</p>\n</blockquote>\n<p>I think this problem is caused by something else. The issue I encountered is now have rational explanation.</p>\n<p>Please read this thread.<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980</a></p>",
          "rawMarkdown": "> Lots of better models got worse score\n\nI think this problem is caused by something else. The issue I encountered is now have rational explanation.\n\nPlease read this thread.\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980"
        }
      ]
    },
    {
      "id": 1662884,
      "postDate": "2022-01-24T15:57:32.110Z",
      "content": "<p>kagglers are inference the evaluation way, too. 👍👍</p>",
      "rawMarkdown": "kagglers are inference the evaluation way, too. 👍👍"
    },
    {
      "id": 1662566,
      "postDate": "2022-01-24T11:54:17.990Z",
      "content": "<p>just a quick note: <br>\nto submit really FP,  try (x,y,1,1)</p>",
      "rawMarkdown": "just a quick note: \nto submit really FP,  try (x,y,1,1)",
      "replies": [
        {
          "id": 1662599,
          "postDate": "2022-01-24T12:15:52.380Z",
          "content": "<p>Yes. I already use this as below:</p>\n<pre><code>FPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n</code></pre>",
          "rawMarkdown": "Yes. I already use this as below:\n\n```\nFPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n```"
        }
      ]
    },
    {
      "id": 1662368,
      "postDate": "2022-01-24T08:12:59.727Z",
      "content": "<p>But still 129 Detections should significantly drop F2, for throwing FP boxes</p>",
      "rawMarkdown": "But still 129 Detections should significantly drop F2, for throwing FP boxes",
      "replies": [
        {
          "id": 1662388,
          "postDate": "2022-01-24T08:34:37.817Z",
          "content": "<p>No. It’s as expected. Please calculate by yourself using eq (5) shared above.</p>",
          "rawMarkdown": "No. It’s as expected. Please calculate by yourself using eq (5) shared above.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1662303,
      "postDate": "2022-01-24T07:08:22.360Z",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> Maybe I am dumb but how you add 10000 FN to one frame ?:) False Negative is not detected starfish, eventually remove TP which we don't know. I missed something obvious ?</p>",
      "rawMarkdown": "@tatamikenn Maybe I am dumb but how you add 10000 FN to one frame ?:) False Negative is not detected starfish, eventually remove TP which we don't know. I missed something obvious ?",
      "replies": [
        {
          "id": 1662334,
          "postDate": "2022-01-24T07:35:38.387Z",
          "content": "<p>Sorry, FN is typo. I've corrected the notation now(FN-&gt;FP)</p>",
          "rawMarkdown": "Sorry, FN is typo. I've corrected the notation now(FN->FP)",
          "votes": 1
        }
      ]
    },
    {
      "id": 1662217,
      "postDate": "2022-01-24T05:59:31.063Z",
      "content": "<blockquote>\n  <p>One hypothesis I came up with is the F2 score is calculated per image, and then average each of them.</p>\n</blockquote>\n<p>I would say this hypothesis is not true.</p>\n<p>I conducted another experiment which put zero prediction on frame idx=5,200 (we denote this experiment as (A)).<br>\nIf the hypothesis is true, then the calculated F2 value of the frame is zero.<br>\nOn the contrary, if we add 10,000 FP to the frame, the calculated F2 value of the frame is equal to or greater than zero.<br>\nThus, the observed LB score in the latter case is equal to or greater than the former case (since F2 of a frame is equal to or greater than the other, whereas the other frame's F2 are the same).<br>\nTherefore, the LB score of experiment (A) should be less than or equal to the case of +10K * FP case.</p>\n<p>However, in the experiment (A), I observed zero drop in LB score (0.601 -&gt; 0.601) which is greater than +10K * FP case(0.595).<br>\nThis contradicts what we discussed above.</p>\n<p>the code is here:</p>\n<pre><code>for idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=3000 is a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5_200\n    if idx == fp_add_frame:\n        # blank prediction\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n\n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n</code></pre>",
      "rawMarkdown": "> One hypothesis I came up with is the F2 score is calculated per image, and then average each of them.\n\nI would say this hypothesis is not true.\n\nI conducted another experiment which put zero prediction on frame idx=5,200 (we denote this experiment as (A)).\nIf the hypothesis is true, then the calculated F2 value of the frame is zero.\nOn the contrary, if we add 10,000 FP to the frame, the calculated F2 value of the frame is equal to or greater than zero.\nThus, the observed LB score in the latter case is equal to or greater than the former case (since F2 of a frame is equal to or greater than the other, whereas the other frame's F2 are the same).\nTherefore, the LB score of experiment (A) should be less than or equal to the case of +10K * FP case.\n\nHowever, in the experiment (A), I observed zero drop in LB score (0.601 -> 0.601) which is greater than +10K * FP case(0.595).\nThis contradicts what we discussed above.\n\nthe code is here:\n\n```\nfor idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=3000 is a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5_200\n    if idx == fp_add_frame:\n        # blank prediction\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n    \n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n```"
    },
    {
      "id": 1661953,
      "postDate": "2022-01-23T21:17:20.170Z",
      "content": "<p>I think, that's right.</p>\n<blockquote>\n  <p>One hypothesis I came up with is the F2 score is calculated per image, and then average each of them. </p>\n</blockquote>\n<p>We can actually test it. Add 10000 FN boxes to every image, or to every second image. If F2 drops dramatically, then your theory is right. If not, then something strange is going on here.</p>\n<p>But personally, I think, averaging F2 over each image, is a good thing to do 😃.</p>",
      "rawMarkdown": "I think, that's right.\n> One hypothesis I came up with is the F2 score is calculated per image, and then average each of them. \n\nWe can actually test it. Add 10000 FN boxes to every image, or to every second image. If F2 drops dramatically, then your theory is right. If not, then something strange is going on here.\n\nBut personally, I think, averaging F2 over each image, is a good thing to do 😃.",
      "replies": [
        {
          "id": 1662007,
          "postDate": "2022-01-23T23:33:48.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/vadbeg\" target=\"_blank\">@vadbeg</a> </p>\n<blockquote>\n  <p>But personally, I think, averaging F2 over each image, is a good thing to do</p>\n</blockquote>\n<p>How can you define F2 on the background (non-annotated) image?</p>",
          "rawMarkdown": "@vadbeg \n\n>But personally, I think, averaging F2 over each image, is a good thing to do\n\nHow can you define F2 on the background (non-annotated) image?"
        },
        {
          "id": 1662020,
          "postDate": "2022-01-23T23:58:19.507Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>Maybe it skips if there are no detections on the background and sets F2 to zero for average if there were detections on a background image. But I have zero clue if I am correct and I am not sure how to test for that. Some clarification is needed.</p>",
          "rawMarkdown": "@tatamikenn \n\nMaybe it skips if there are no detections on the background and sets F2 to zero for average if there were detections on a background image. But I have zero clue if I am correct and I am not sure how to test for that. Some clarification is needed."
        },
        {
          "id": 1662023,
          "postDate": "2022-01-24T00:01:34.427Z",
          "content": "<p><a href=\"https://www.kaggle.com/outwrest\" target=\"_blank\">@outwrest</a> </p>\n<blockquote>\n  <p>Maybe it skips if there are no detections on the background and sets F2 to zero for average if there were detections on a background image.</p>\n</blockquote>\n<p>This is possible, but I think this evaluation method is nonsense. It can't give low score for the model which generates plenty of FPs to the background image.</p>",
          "rawMarkdown": "@outwrest \n\n> Maybe it skips if there are no detections on the background and sets F2 to zero for average if there were detections on a background image.\n\nThis is possible, but I think this evaluation method is nonsense. It can't give low score for the model which generates plenty of FPs to the background image."
        },
        {
          "id": 1662024,
          "postDate": "2022-01-24T00:02:19.060Z",
          "content": "<p>See also this thread. I think this option is more probable.</p>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980</a></p>",
          "rawMarkdown": "See also this thread. I think this option is more probable.\n\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980"
        },
        {
          "id": 1662027,
          "postDate": "2022-01-24T00:06:39.513Z",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> I see the thread and that is a better explanation. </p>\n<blockquote>\n  <p>evaluation method is nonsense</p>\n</blockquote>\n<p>I don't see why F2 was chosen for object detection, mAP would've sufficed and proven to be easier to implement and test. I wasted a lot of time implementing and worrying about it.  </p>",
          "rawMarkdown": "@tatamikenn I see the thread and that is a better explanation. \n\n>  evaluation method is nonsense\n\nI don't see why F2 was chosen for object detection, mAP would've sufficed and proven to be easier to implement and test. I wasted a lot of time implementing and worrying about it.  "
        },
        {
          "id": 1662029,
          "postDate": "2022-01-24T00:12:04.497Z",
          "content": "<blockquote>\n  <p>I don't see why F2 was chosen for object detection, mAP would've sufficed and proven to be easier to implement and test.</p>\n</blockquote>\n<p>I partly agree with that but, as you know, mAP is not appropriate for real application since it doesn't considers confidence threshold.<br>\nYou set confidence threshold as small as possible for local evaluation, right?</p>\n<p>I think they want production model, not academic ones.<br>\nThey considers setting proper confidence threshold is our task to submit better prediction model as a product.</p>",
          "rawMarkdown": "> I don't see why F2 was chosen for object detection, mAP would've sufficed and proven to be easier to implement and test.\n\nI partly agree with that but, as you know, mAP is not appropriate for real application since it doesn't considers confidence threshold.\nYou set confidence threshold as small as possible for local evaluation, right?\n\nI think they want production model, not academic ones.\nThey considers setting proper confidence threshold is our task to submit better prediction model as a product.",
          "votes": 5
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1661980,
      "author_name": "Psi",
      "author_url": "",
      "post_date": "2022-01-23T22:15:14.407000",
      "content": "<p>I think they just cut-off too many predicted boxes per image to not overload the system. The drop in your score would be smaller if it would be local metric. The only explanation I can see is that it is cut-off. Or maybe you have duplicate boxes they drop, I did not check your code. All this is just a guess.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1661999,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-23T23:19:51.210000",
          "content": "<p>Thank you for pointing this out.</p>\n<p>I had thought of this once, but never delved into it deeply.</p>\n<p>We can estimate the cutoff threshold by the following equation [1].<br>\nα means the number of FPs added to the prediction.</p>\n<p>$$<br>\n\\frac{TP}{\\alpha} = \\frac{1}{5} \\left( \\frac{1}{F_2^\\alpha} - \\frac{1}{F_2} \\right)^{-1} \\tag{5}<br>\n$$</p>\n<p>Using the same formula as adding one FP to every image, I estimated the TP score to be ~1538.<br>\nSubstituting this into equation (5), the estimated value of α for the cutoff for each image is ~129. This is a very likely value.</p>\n<p>[1] <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156</a></p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1662011,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-23T23:39:06.630000",
          "content": "<blockquote>\n  <p>Or maybe you have duplicate boxes they drop</p>\n</blockquote>\n<p>I took duplicated boxes into account. The above code generates unique boxes.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1662347,
          "author_name": "Lukasz Borecki",
          "author_url": "",
          "post_date": "2022-01-24T07:46:54.110000",
          "content": "<p>so probably maxdets=100 :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1663022,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2022-01-24T18:02:59.850000",
          "content": "<p><a href=\"https://www.kaggle.com/philippsinger\" target=\"_blank\">@philippsinger</a> is correct. Capping the number of predictions per image let us use simple pairwise comparisons in the metric without worrying about excessive scoring runtimes. </p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1663299,
          "author_name": "hengck23",
          "author_url": "",
          "post_date": "2022-01-25T00:47:34.503000",
          "content": "<p>What is the criteria for cutting off? If it is by score, what happens if box are of the same score? </p>\n<p>Say of i submit 1000  prediction box of same score1.0, the evalution selects The first 100 or the best 100?  </p>\n<p>It is better to set a rule that at most 100 box predicted per image in this case</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1664600,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-26T05:01:12.133000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Hi, how about the question <a href=\"https://www.kaggle.com/hengck23\" target=\"_blank\">@hengck23</a> asked above? I want to know the answer too.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1661440,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-23T13:21:05.573000",
      "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> <br>\nDear Kaggle staffs, please read through the post and consider to disclose the evaluation code, unless we have to spend a lot of time to solve this issue. Considering the competition end is closing, it would be tough to all of us.</p>",
      "votes": 6,
      "replies": [
        {
          "id": 1661974,
          "author_name": "",
          "author_url": "",
          "post_date": "2022-01-23T22:01:46.460000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1661993,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-23T23:05:11.253000",
          "content": "<p><a href=\"https://www.kaggle.com/addisonhoward\" target=\"_blank\">@addisonhoward</a> <a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> </p>\n<p>Some quick questions:</p>\n<ul>\n<li>does evaluation conducted per image, i.e. first calculate F2 per image, and then mean them to get the final score?</li>\n<li>If the above question is right, how can you define F2 on the background (non-annotated) frame?</li>\n<li>Do you use cutoff too many bounding box per image? If it is, what is the threshold?</li>\n<li>How cutoff is conducted?  Predictions with lowest confidence values?</li>\n</ul>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1663027,
          "author_name": "Sohier Dane",
          "author_url": "",
          "post_date": "2022-01-24T18:06:29.547000",
          "content": "<ul>\n<li>No, the F2 score is calculated globally. </li>\n<li>There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?</li>\n</ul>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1663116,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T19:50:05.007000",
          "content": "<p><a href=\"https://www.kaggle.com/sohier\" target=\"_blank\">@sohier</a> Thank you for clarification.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1662294,
      "author_name": "Lukasz Borecki",
      "author_url": "",
      "post_date": "2022-01-24T07:01:54.533000",
      "content": "<p>I agree we found it, some weeks ago but couldn't prove it. Lots of better models got worse score. We made FP reduction manuever but still score didn't increased much or sometimes dropped. Second observation everytime we jump inference resolution val F2 always drops but on LB almost always explodes to higher value. </p>",
      "votes": 1,
      "replies": [
        {
          "id": 1662338,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T07:40:34.773000",
          "content": "<blockquote>\n  <p>Lots of better models got worse score</p>\n</blockquote>\n<p>I think this problem is caused by something else. The issue I encountered is now have rational explanation.</p>\n<p>Please read this thread.<br>\n<a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1662884,
      "author_name": "Eugene J. Ryu",
      "author_url": "",
      "post_date": "2022-01-24T15:57:32.110000",
      "content": "<p>kagglers are inference the evaluation way, too. 👍👍</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1662566,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2022-01-24T11:54:17.990000",
      "content": "<p>just a quick note: <br>\nto submit really FP,  try (x,y,1,1)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1662599,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T12:15:52.380000",
          "content": "<p>Yes. I already use this as below:</p>\n<pre><code>FPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n</code></pre>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1662368,
      "author_name": "Lukasz Borecki",
      "author_url": "",
      "post_date": "2022-01-24T08:12:59.727000",
      "content": "<p>But still 129 Detections should significantly drop F2, for throwing FP boxes</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1662388,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T08:34:37.817000",
          "content": "<p>No. It’s as expected. Please calculate by yourself using eq (5) shared above.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1662303,
      "author_name": "Lukasz Borecki",
      "author_url": "",
      "post_date": "2022-01-24T07:08:22.360000",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> Maybe I am dumb but how you add 10000 FN to one frame ?:) False Negative is not detected starfish, eventually remove TP which we don't know. I missed something obvious ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1662334,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T07:35:38.387000",
          "content": "<p>Sorry, FN is typo. I've corrected the notation now(FN-&gt;FP)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1662217,
      "author_name": "Bilzard",
      "author_url": "",
      "post_date": "2022-01-24T05:59:31.063000",
      "content": "<blockquote>\n  <p>One hypothesis I came up with is the F2 score is calculated per image, and then average each of them.</p>\n</blockquote>\n<p>I would say this hypothesis is not true.</p>\n<p>I conducted another experiment which put zero prediction on frame idx=5,200 (we denote this experiment as (A)).<br>\nIf the hypothesis is true, then the calculated F2 value of the frame is zero.<br>\nOn the contrary, if we add 10,000 FP to the frame, the calculated F2 value of the frame is equal to or greater than zero.<br>\nThus, the observed LB score in the latter case is equal to or greater than the former case (since F2 of a frame is equal to or greater than the other, whereas the other frame's F2 are the same).<br>\nTherefore, the LB score of experiment (A) should be less than or equal to the case of +10K * FP case.</p>\n<p>However, in the experiment (A), I observed zero drop in LB score (0.601 -&gt; 0.601) which is greater than +10K * FP case(0.595).<br>\nThis contradicts what we discussed above.</p>\n<p>the code is here:</p>\n<pre><code>for idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=3000 is a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5_200\n    if idx == fp_add_frame:\n        # blank prediction\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n\n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n</code></pre>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1661953,
      "author_name": "Vadzim Tsitko",
      "author_url": "",
      "post_date": "2022-01-23T21:17:20.170000",
      "content": "<p>I think, that's right.</p>\n<blockquote>\n  <p>One hypothesis I came up with is the F2 score is calculated per image, and then average each of them. </p>\n</blockquote>\n<p>We can actually test it. Add 10000 FN boxes to every image, or to every second image. If F2 drops dramatically, then your theory is right. If not, then something strange is going on here.</p>\n<p>But personally, I think, averaging F2 over each image, is a good thing to do 😃.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1662007,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-23T23:33:48.460000",
          "content": "<p><a href=\"https://www.kaggle.com/vadbeg\" target=\"_blank\">@vadbeg</a> </p>\n<blockquote>\n  <p>But personally, I think, averaging F2 over each image, is a good thing to do</p>\n</blockquote>\n<p>How can you define F2 on the background (non-annotated) image?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1662020,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "2022-01-23T23:58:19.507000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> </p>\n<p>Maybe it skips if there are no detections on the background and sets F2 to zero for average if there were detections on a background image. But I have zero clue if I am correct and I am not sure how to test for that. Some clarification is needed.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1662023,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T00:01:34.427000",
          "content": "<p><a href=\"https://www.kaggle.com/outwrest\" target=\"_blank\">@outwrest</a> </p>\n<blockquote>\n  <p>Maybe it skips if there are no detections on the background and sets F2 to zero for average if there were detections on a background image.</p>\n</blockquote>\n<p>This is possible, but I think this evaluation method is nonsense. It can't give low score for the model which generates plenty of FPs to the background image.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1662024,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T00:02:19.060000",
          "content": "<p>See also this thread. I think this option is more probable.</p>\n<p><a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\" target=\"_blank\">https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980</a></p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1662027,
          "author_name": "outwrest",
          "author_url": "",
          "post_date": "2022-01-24T00:06:39.513000",
          "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a> I see the thread and that is a better explanation. </p>\n<blockquote>\n  <p>evaluation method is nonsense</p>\n</blockquote>\n<p>I don't see why F2 was chosen for object detection, mAP would've sufficed and proven to be easier to implement and test. I wasted a lot of time implementing and worrying about it.  </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1662029,
          "author_name": "Bilzard",
          "author_url": "",
          "post_date": "2022-01-24T00:12:04.497000",
          "content": "<blockquote>\n  <p>I don't see why F2 was chosen for object detection, mAP would've sufficed and proven to be easier to implement and test.</p>\n</blockquote>\n<p>I partly agree with that but, as you know, mAP is not appropriate for real application since it doesn't considers confidence threshold.<br>\nYou set confidence threshold as small as possible for local evaluation, right?</p>\n<p>I think they want production model, not academic ones.<br>\nThey considers setting proper confidence threshold is our task to submit better prediction model as a product.</p>",
          "votes": 5,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1661417": "# In Short\n\nI added 10,000 FP bbox to one frame, and submit this model. After calculated LB, the score is dropped 0.006 (0.601 -> 0.595). Given the total frame number is around 3,500 (0.25 * 14,000), and number of GT label is ~2450[1], the score drop is apparently too little.\n\nOn the contrary, if I add one FP bbox on every frame, the score is dropped as expected (0.601 -> 0.426).\n\n~~I suspect the algorithm used in the leaderboard evaluation code contains bugs or, at least, it might not be as expected for us.~~\n\n# Assumed Cause\n\n~~One hypothesis I came up with is the F2 score is calculated per image, and then average each of them. However, if it is true, we can't reasonably define F2 score for the background (no-annotation) frames.\nAnother possibility is skipping background image and only calculating for annotated frame. However, if it is true, the model which produces a lot of FP wouldn't be considered properly with this algorithm.~~\n\nWe have now highly rational explanation on this issue.\nIt seems that the box threshold per image is set to ~129 to avoid system load.\n\nDetails are explained in this thread:\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1661980\n\n# Actual Cause\n\nThese are clarified by Kaggle staff:\n\n> No, the F2 score is calculated globally.\n> There is a cutoff of 100 predictions per image. Is this low enough that it's relevant for your actual model outputs?\n\nhttps://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302632#1663027\n\n# Actions\n\n~~Kagglers, please share ideas -- possible cause, reasonable explanations and the way to test it.\nAnyone who is kindly to reproduce this issue is also welcome.~~\n\n~~However, considering it is close to the competition end, the best way is to ask Kaggle staff to disclose LB calculation code, unless we have to consume a lot of time to test and resolve for the issue.If the kaggle staffs are reading this post, please consider for that.~~\n\n~~I hope this would be just my misunderstanding or my code bug.~~\n\n~~I'm now waiting for Kaggle staff officially replies to this post.~~\n\n# Code\n\nThe test code I used is as below:\n\n```\nP_KEEP = 1.0\nCONF = 0.50\nFP_COUNT = 10_000\n```\n\n```\ndef drop_pred(bboxes, confs, p_keep):\n    '''\n    randomly drop prediction for the probability (1 - p_keep).\n    '''\n    if p_keep == 1:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    pp = np.random.uniform(size=len(bboxes))\n    bboxes = bboxes[pp <= p_keep]\n    confs = confs[pp <= p_keep]\n    return bboxes, confs\n\ndef add_fp(bboxes, confs, fp_count=0, fp_conf=1e-4):\n    '''\n    add false negative to the frame\n    '''\n    if fp_count == 0:\n        return bboxes, confs\n    bboxes = bboxes.copy()\n    confs = confs.copy()\n    assert len(bboxes) == len(confs)\n    FPs = [[i % 1280, i // 1280, 1, 1] for i in range(fp_count)]\n    bboxes = np.append(bboxes, FPs, axis=0)\n    assert bboxes.dtype == int\n    \n    # we need to set sufficiently small confidence value as for not to affect TP\n    conf_FPs = [fp_conf] * fp_count\n    confs = np.append(confs, conf_FPs)\n    assert confs.dtype == float\n    return bboxes, confs    \n    \n\ndef predict(model, img, size=768, augment=False, p_keep=1.0):\n    height, width = img.shape[:2]\n    results = model(img, size=size, augment=augment)  # custom inference size\n    preds   = results.pandas().xyxy[0]\n    bboxes  = preds[['xmin','ymin','xmax','ymax']].values\n    if len(bboxes):\n        bboxes  = voc2coco(bboxes, height, width).astype(int)\n        confs   = preds.confidence.values\n        \n        bboxes, confs = drop_pred(bboxes, confs, p_keep=p_keep)\n        return bboxes, confs\n    else:\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n        return bboxes, confs\n    \ndef format_prediction(bboxes, confs):\n    annot = ''\n    if len(bboxes) > 0:\n        for idx in range(len(bboxes)):\n            xmin, ymin, w, h = bboxes[idx]\n            conf = confs[idx]\n            annot += f'{conf:.8f} {xmin} {ymin} {w} {h}'\n            annot +=' '\n        annot = annot.strip(' ')\n    return annot\n```\n\n```\nfor idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=5200 is possible to be a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5200\n    if idx == fp_add_frame:\n        bboxes, confs = add_fp(bboxes, confs, fp_count=FP_COUNT)\n        print(f\"{len(bboxes)}, {len(confs)}\")\n        print(f\"{bboxes[0]}, {confs[0]}\")\n    \n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n    if idx < 3:\n        display(show_img(img, bboxes, bbox_format='coco'))\n```\n\n# Reference\n\n[1] [LB probing result: GT Labels per Frame in the Public LB](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302605)\n\n# Update Note\n\n- 2022/1/24 fix typo FN->FP\n- 2022/1/25 actual cause is clarified by Kaggle staff",
    "1661980": "I think they just cut-off too many predicted boxes per image to not overload the system. The drop in your score would be smaller if it would be local metric. The only explanation I can see is that it is cut-off. Or maybe you have duplicate boxes they drop, I did not check your code. All this is just a guess.",
    "1661440": "@addisonhoward @sohier \nDear Kaggle staffs, please read through the post and consider to disclose the evaluation code, unless we have to spend a lot of time to solve this issue. Considering the competition end is closing, it would be tough to all of us.",
    "1662294": "I agree we found it, some weeks ago but couldn't prove it. Lots of better models got worse score. We made FP reduction manuever but still score didn't increased much or sometimes dropped. Second observation everytime we jump inference resolution val F2 always drops but on LB almost always explodes to higher value. ",
    "1662884": "kagglers are inference the evaluation way, too. 👍👍",
    "1662566": "just a quick note: \nto submit really FP,  try (x,y,1,1)",
    "1662368": "But still 129 Detections should significantly drop F2, for throwing FP boxes",
    "1662303": "@tatamikenn Maybe I am dumb but how you add 10000 FN to one frame ?:) False Negative is not detected starfish, eventually remove TP which we don't know. I missed something obvious ?",
    "1662217": "> One hypothesis I came up with is the F2 score is calculated per image, and then average each of them.\n\nI would say this hypothesis is not true.\n\nI conducted another experiment which put zero prediction on frame idx=5,200 (we denote this experiment as (A)).\nIf the hypothesis is true, then the calculated F2 value of the frame is zero.\nOn the contrary, if we add 10,000 FP to the frame, the calculated F2 value of the frame is equal to or greater than zero.\nThus, the observed LB score in the latter case is equal to or greater than the former case (since F2 of a frame is equal to or greater than the other, whereas the other frame's F2 are the same).\nTherefore, the LB score of experiment (A) should be less than or equal to the case of +10K * FP case.\n\nHowever, in the experiment (A), I observed zero drop in LB score (0.601 -> 0.601) which is greater than +10K * FP case(0.595).\nThis contradicts what we discussed above.\n\nthe code is here:\n\n```\nfor idx, (img, pred_df) in enumerate(tqdm(iter_test)):\n    bboxes, confs  = predict(model, img, size=IMG_SIZE, augment=AUGMENT, p_keep=P_KEEP)\n\n    # idx=3000 is a public test frame\n    #   - ref: https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302057\n    fp_add_frame = 5_200\n    if idx == fp_add_frame:\n        # blank prediction\n        bboxes, confs = np.zeros((0, 4), dtype=int), np.zeros((0), dtype=int)\n    \n    annot = format_prediction(bboxes, confs)    \n    pred_df['annotations'] = annot\n    env.predict(pred_df)\n```",
    "1661953": "I think, that's right.\n> One hypothesis I came up with is the F2 score is calculated per image, and then average each of them. \n\nWe can actually test it. Add 10000 FN boxes to every image, or to every second image. If F2 drops dramatically, then your theory is right. If not, then something strange is going on here.\n\nBut personally, I think, averaging F2 over each image, is a good thing to do 😃."
  }
}