{
  "id": 302605,
  "title": "LB probing result: GT Labels per Frame in the Public LB",
  "url": "/competitions/tensorflow-great-barrier-reef/discussion/302605",
  "author_name": "",
  "post_date": "2022-01-23T10:30:59.352262200Z",
  "votes": 14,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I estimated GT labels per frame in the public LB using [1].</p>\n<p>In short, it is estimated to contain about <strong>0.705 COTS per frame</strong> in the public test set, which is close to that of video-1 in the train dataset (which is high-COTS video compare to other two videos).</p>\n<p>It indicates high-recall model tend to gets higher in the public LB.</p>\n<p>If they contains lower COTS in the private test frames, the model tend to shake down if your model produces a lot of FNs.</p>\n<p>We have to care precision as well as recall in order not to shake down.</p>\n<hr>\n<p>Estimated COTS per frame in public LB:</p>\n<pre><code>    GT/I    TP/I    FN/I    FP/I\n0    0.716   0.439   0.276   0.354\n1    0.697   0.439   0.257   0.429\n2    0.703   0.439   0.264   0.403\nmean    0.705   0.439   0.266   0.396\nstd    0.010   0.000   0.010   0.038\n</code></pre>\n<p>COTS per frame in the train videos:</p>\n<pre><code>         sum_cots  duration  mean_cots\nfold_id                               \n0            3065      6708   0.456917\n1            6384      8232   0.775510\n2            2449      8561   0.286065\n</code></pre>\n<h1>Reference</h1>\n<p>[1]: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156\" target=\"_blank\">Magic Two: We Can Estimate Number of GT Labels on the Public Test Data</a></p>",
  "messages": [
    {
      "id": "1661252",
      "postDate": "01/23/2022 10:30:59",
      "content": "<p>I estimated GT labels per frame in the public LB using [1].</p>\n<p>In short, it is estimated to contain about <strong>0.705 COTS per frame</strong> in the public test set, which is close to that of video-1 in the train dataset (which is high-COTS video compare to other two videos).</p>\n<p>It indicates high-recall model tend to gets higher in the public LB.</p>\n<p>If they contains lower COTS in the private test frames, the model tend to shake down if your model produces a lot of FNs.</p>\n<p>We have to care precision as well as recall in order not to shake down.</p>\n<hr>\n<p>Estimated COTS per frame in public LB:</p>\n<pre><code>    GT/I    TP/I    FN/I    FP/I\n0    0.716   0.439   0.276   0.354\n1    0.697   0.439   0.257   0.429\n2    0.703   0.439   0.264   0.403\nmean    0.705   0.439   0.266   0.396\nstd    0.010   0.000   0.010   0.038\n</code></pre>\n<p>COTS per frame in the train videos:</p>\n<pre><code>         sum_cots  duration  mean_cots\nfold_id                               \n0            3065      6708   0.456917\n1            6384      8232   0.775510\n2            2449      8561   0.286065\n</code></pre>\n<h1>Reference</h1>\n<p>[1]: <a href=\"https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156\" target=\"_blank\">Magic Two: We Can Estimate Number of GT Labels on the Public Test Data</a></p>",
      "rawMarkdown": "I estimated GT labels per frame in the public LB using [1].\n\nIn short, it is estimated to contain about **0.705 COTS per frame** in the public test set, which is close to that of video-1 in the train dataset (which is high-COTS video compare to other two videos).\n\nIt indicates high-recall model tend to gets higher in the public LB.\n\nIf they contains lower COTS in the private test frames, the model tend to shake down if your model produces a lot of FNs.\n\nWe have to care precision as well as recall in order not to shake down.\n\n---\n\nEstimated COTS per frame in public LB:\n```\n\tGT/I\tTP/I\tFN/I\tFP/I\n0\t0.716\t0.439\t0.276\t0.354\n1\t0.697\t0.439\t0.257\t0.429\n2\t0.703\t0.439\t0.264\t0.403\nmean\t0.705\t0.439\t0.266\t0.396\nstd\t0.010\t0.000\t0.010\t0.038\n```\n\nCOTS per frame in the train videos:\n```\n         sum_cots  duration  mean_cots\nfold_id                               \n0            3065      6708   0.456917\n1            6384      8232   0.775510\n2            2449      8561   0.286065\n```\n\n# Reference\n\n[1]: [Magic Two: We Can Estimate Number of GT Labels on the Public Test Data](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156)",
      "votes": null
    },
    {
      "id": "1661263",
      "postDate": "01/23/2022 10:39:21",
      "content": "<p>Thanks for sharing.<br>\nMy validation score of video-1 is closest to public LB, which is consistent with your result.</p>",
      "rawMarkdown": "Thanks for sharing.\nMy validation score of video-1 is closest to public LB, which is consistent with your result.",
      "votes": null
    },
    {
      "id": "1661272",
      "postDate": "01/23/2022 10:43:18",
      "content": "<p>I see. Thank you for sharing your observation.</p>",
      "rawMarkdown": "I see. Thank you for sharing your observation.",
      "votes": null
    },
    {
      "id": "1661289",
      "postDate": "01/23/2022 10:58:54",
      "content": "<p>My findings are also consistent with this! Thanks for sharing.</p>",
      "rawMarkdown": "My findings are also consistent with this! Thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1661263,
      "author_name": "drtausamaru",
      "author_url": "",
      "post_date": "01/23/2022 10:39:21",
      "content": "<p>Thanks for sharing.<br>\nMy validation score of video-1 is closest to public LB, which is consistent with your result.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1661272,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "01/23/2022 10:43:18",
          "content": "<p>I see. Thank you for sharing your observation.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1661289,
      "author_name": "dauriel",
      "author_url": "",
      "post_date": "01/23/2022 10:58:54",
      "content": "<p>My findings are also consistent with this! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1661252": "I estimated GT labels per frame in the public LB using [1].\n\nIn short, it is estimated to contain about **0.705 COTS per frame** in the public test set, which is close to that of video-1 in the train dataset (which is high-COTS video compare to other two videos).\n\nIt indicates high-recall model tend to gets higher in the public LB.\n\nIf they contains lower COTS in the private test frames, the model tend to shake down if your model produces a lot of FNs.\n\nWe have to care precision as well as recall in order not to shake down.\n\n---\n\nEstimated COTS per frame in public LB:\n```\n\tGT/I\tTP/I\tFN/I\tFP/I\n0\t0.716\t0.439\t0.276\t0.354\n1\t0.697\t0.439\t0.257\t0.429\n2\t0.703\t0.439\t0.264\t0.403\nmean\t0.705\t0.439\t0.266\t0.396\nstd\t0.010\t0.000\t0.010\t0.038\n```\n\nCOTS per frame in the train videos:\n```\n         sum_cots  duration  mean_cots\nfold_id                               \n0            3065      6708   0.456917\n1            6384      8232   0.775510\n2            2449      8561   0.286065\n```\n\n# Reference\n\n[1]: [Magic Two: We Can Estimate Number of GT Labels on the Public Test Data](https://www.kaggle.com/c/tensorflow-great-barrier-reef/discussion/302156)",
    "1661263": "Thanks for sharing.\nMy validation score of video-1 is closest to public LB, which is consistent with your result.",
    "1661272": "I see. Thank you for sharing your observation.",
    "1661289": "My findings are also consistent with this! Thanks for sharing."
  },
  "source": "meta"
}