{
  "id": 312071,
  "title": "CV vs LB          ",
  "url": "/competitions/birdclef-2022/discussion/312071",
  "author_name": "",
  "post_date": "2022-03-10T08:05:06.532417Z",
  "votes": 8,
  "comment_count": 6,
  "views": 0,
  "content": "<p>tf_efficientnet_b7 (single model)<br>\nn_mels: 256<br>\nCV (F1@0.30): 0.64<br>\nLB : 0.58</p>",
  "messages": [
    {
      "id": "1717774",
      "postDate": "03/10/2022 08:05:06",
      "content": "<p>tf_efficientnet_b7 (single model)<br>\nn_mels: 256<br>\nCV (F1@0.30): 0.64<br>\nLB : 0.58</p>",
      "rawMarkdown": "tf_efficientnet_b7 (single model)\nn_mels: 256\nCV (F1@0.30): 0.64\nLB : 0.58",
      "votes": null
    },
    {
      "id": "1717991",
      "postDate": "03/10/2022 12:12:57",
      "content": "<p>tf_efficientnet_b0_ns (single model, 5 folds but only one trained and scored yet)<br>\nCV: 0.63<br>\nLB: 0.54</p>\n<p>Need to fix my CV strategy……</p>",
      "rawMarkdown": "tf_efficientnet_b0_ns (single model, 5 folds but only one trained and scored yet)\nCV: 0.63\nLB: 0.54\n\nNeed to fix my CV strategy......",
      "votes": null
    },
    {
      "id": "1720183",
      "postDate": "03/12/2022 14:37:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kfk42kfk\" target=\"_blank\">@kfk42kfk</a>  what is F1@0.30 ? If i know then i can share my CV</p>",
      "rawMarkdown": "Hi @kfk42kfk  what is F1@0.30 ? If i know then i can share my CV",
      "votes": null
    },
    {
      "id": "1720856",
      "postDate": "03/13/2022 07:52:01",
      "content": "<p>I calculated mine using this code snippet from this notebook: <a href=\"https://www.kaggle.com/kaerunantoka/birdclef2022-n001-training\" target=\"_blank\">https://www.kaggle.com/kaerunantoka/birdclef2022-n001-training</a></p>\n<pre><code>class MetricMeter(object):\n    def __init__(self):\n        self.reset()\n\n    def reset(self):\n        self.y_true = []\n        self.y_pred = []\n\n    def update(self, y_true, y_pred):\n        self.y_true.extend(y_true.cpu().detach().numpy().tolist())\n        # self.y_pred.extend(torch.sigmoid(y_pred).cpu().detach().numpy().tolist())\n        # self.y_pred.extend(y_pred[\"clipwise_output\"].max(axis=1)[0].cpu().detach().numpy().tolist())\n        self.y_pred.extend(y_pred[\"clipwise_output\"].cpu().detach().numpy().tolist())\n\n    @property\n    def avg(self):\n        self.f1_03 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) &gt; 0.3, average=\"micro\")\n        self.f1_05 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) &gt; 0.5, average=\"micro\")\n\n        return {\n            \"f1_at_03\" : self.f1_03,\n            \"f1_at_05\" : self.f1_05,\n        }\n</code></pre>",
      "rawMarkdown": "I calculated mine using this code snippet from this notebook: https://www.kaggle.com/kaerunantoka/birdclef2022-n001-training\n```\nclass MetricMeter(object):\n    def __init__(self):\n        self.reset()\n    \n    def reset(self):\n        self.y_true = []\n        self.y_pred = []\n    \n    def update(self, y_true, y_pred):\n        self.y_true.extend(y_true.cpu().detach().numpy().tolist())\n        # self.y_pred.extend(torch.sigmoid(y_pred).cpu().detach().numpy().tolist())\n        # self.y_pred.extend(y_pred[\"clipwise_output\"].max(axis=1)[0].cpu().detach().numpy().tolist())\n        self.y_pred.extend(y_pred[\"clipwise_output\"].cpu().detach().numpy().tolist())\n\n    @property\n    def avg(self):\n        self.f1_03 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) > 0.3, average=\"micro\")\n        self.f1_05 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) > 0.5, average=\"micro\")\n        \n        return {\n            \"f1_at_03\" : self.f1_03,\n            \"f1_at_05\" : self.f1_05,\n        }\n```",
      "votes": null
    },
    {
      "id": "1723765",
      "postDate": "03/15/2022 17:37:24",
      "content": "<p>Why would you use \"micro\" ? Isn't the metric \"macro\"?</p>\n<blockquote>\n  <p>Submissions are evaluated on a metric that is most similar to the <strong>macro F1 score</strong>. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.</p>\n</blockquote>",
      "rawMarkdown": "Why would you use \"micro\" ? Isn't the metric \"macro\"?\n\n> Submissions are evaluated on a metric that is most similar to the **macro F1 score**. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.",
      "votes": null
    },
    {
      "id": "1782219",
      "postDate": "05/09/2022 10:57:03",
      "content": "<p>Thanks for sharing you CV and LB scores.<br>\nThere seems to be a big gap between the CV and LB. Do you solve the gap?</p>",
      "rawMarkdown": "Thanks for sharing you CV and LB scores.\nThere seems to be a big gap between the CV and LB. Do you solve the gap?",
      "votes": null
    },
    {
      "id": "1786552",
      "postDate": "05/13/2022 02:22:39",
      "content": "<p>Yes, my CV and LB also have a big gap. In fact, I'm not sure if I should trust CV or LB when submitting the final 2 results.</p>\n<p>As I said in the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/324304#1784539\" target=\"_blank\">discussion</a>, there are some reasons for each view, so I'm not sure how making a decision.</p>",
      "rawMarkdown": "Yes, my CV and LB also have a big gap. In fact, I'm not sure if I should trust CV or LB when submitting the final 2 results.\n\nAs I said in the [discussion](https://www.kaggle.com/competitions/birdclef-2022/discussion/324304#1784539), there are some reasons for each view, so I'm not sure how making a decision.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1717991,
      "author_name": "hinepo",
      "author_url": "",
      "post_date": "03/10/2022 12:12:57",
      "content": "<p>tf_efficientnet_b0_ns (single model, 5 folds but only one trained and scored yet)<br>\nCV: 0.63<br>\nLB: 0.54</p>\n<p>Need to fix my CV strategy……</p>",
      "votes": null,
      "replies": [
        {
          "id": 1782219,
          "author_name": "zhoumichael",
          "author_url": "",
          "post_date": "05/09/2022 10:57:03",
          "content": "<p>Thanks for sharing you CV and LB scores.<br>\nThere seems to be a big gap between the CV and LB. Do you solve the gap?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1720183,
      "author_name": "ulrich07",
      "author_url": "",
      "post_date": "03/12/2022 14:37:10",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/kfk42kfk\" target=\"_blank\">@kfk42kfk</a>  what is F1@0.30 ? If i know then i can share my CV</p>",
      "votes": null,
      "replies": [
        {
          "id": 1720856,
          "author_name": "kfk42kfk",
          "author_url": "",
          "post_date": "03/13/2022 07:52:01",
          "content": "<p>I calculated mine using this code snippet from this notebook: <a href=\"https://www.kaggle.com/kaerunantoka/birdclef2022-n001-training\" target=\"_blank\">https://www.kaggle.com/kaerunantoka/birdclef2022-n001-training</a></p>\n<pre><code>class MetricMeter(object):\n    def __init__(self):\n        self.reset()\n\n    def reset(self):\n        self.y_true = []\n        self.y_pred = []\n\n    def update(self, y_true, y_pred):\n        self.y_true.extend(y_true.cpu().detach().numpy().tolist())\n        # self.y_pred.extend(torch.sigmoid(y_pred).cpu().detach().numpy().tolist())\n        # self.y_pred.extend(y_pred[\"clipwise_output\"].max(axis=1)[0].cpu().detach().numpy().tolist())\n        self.y_pred.extend(y_pred[\"clipwise_output\"].cpu().detach().numpy().tolist())\n\n    @property\n    def avg(self):\n        self.f1_03 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) &gt; 0.3, average=\"micro\")\n        self.f1_05 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) &gt; 0.5, average=\"micro\")\n\n        return {\n            \"f1_at_03\" : self.f1_03,\n            \"f1_at_05\" : self.f1_05,\n        }\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1723765,
          "author_name": "dschettler8845",
          "author_url": "",
          "post_date": "03/15/2022 17:37:24",
          "content": "<p>Why would you use \"micro\" ? Isn't the metric \"macro\"?</p>\n<blockquote>\n  <p>Submissions are evaluated on a metric that is most similar to the <strong>macro F1 score</strong>. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.</p>\n</blockquote>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1786552,
      "author_name": "jiedengsc",
      "author_url": "",
      "post_date": "05/13/2022 02:22:39",
      "content": "<p>Yes, my CV and LB also have a big gap. In fact, I'm not sure if I should trust CV or LB when submitting the final 2 results.</p>\n<p>As I said in the <a href=\"https://www.kaggle.com/competitions/birdclef-2022/discussion/324304#1784539\" target=\"_blank\">discussion</a>, there are some reasons for each view, so I'm not sure how making a decision.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1717774": "tf_efficientnet_b7 (single model)\nn_mels: 256\nCV (F1@0.30): 0.64\nLB : 0.58",
    "1717991": "tf_efficientnet_b0_ns (single model, 5 folds but only one trained and scored yet)\nCV: 0.63\nLB: 0.54\n\nNeed to fix my CV strategy......",
    "1720183": "Hi @kfk42kfk  what is F1@0.30 ? If i know then i can share my CV",
    "1720856": "I calculated mine using this code snippet from this notebook: https://www.kaggle.com/kaerunantoka/birdclef2022-n001-training\n```\nclass MetricMeter(object):\n    def __init__(self):\n        self.reset()\n    \n    def reset(self):\n        self.y_true = []\n        self.y_pred = []\n    \n    def update(self, y_true, y_pred):\n        self.y_true.extend(y_true.cpu().detach().numpy().tolist())\n        # self.y_pred.extend(torch.sigmoid(y_pred).cpu().detach().numpy().tolist())\n        # self.y_pred.extend(y_pred[\"clipwise_output\"].max(axis=1)[0].cpu().detach().numpy().tolist())\n        self.y_pred.extend(y_pred[\"clipwise_output\"].cpu().detach().numpy().tolist())\n\n    @property\n    def avg(self):\n        self.f1_03 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) > 0.3, average=\"micro\")\n        self.f1_05 = metrics.f1_score(np.array(self.y_true), np.array(self.y_pred) > 0.5, average=\"micro\")\n        \n        return {\n            \"f1_at_03\" : self.f1_03,\n            \"f1_at_05\" : self.f1_05,\n        }\n```",
    "1723765": "Why would you use \"micro\" ? Isn't the metric \"macro\"?\n\n> Submissions are evaluated on a metric that is most similar to the **macro F1 score**. Given the amount of audio data used in this competition it wasn't feasible to label every single species found in every soundscape. Instead only a subset of species are actually scored for any given audio file. After dropping all of the un-scored rows we technically run a weighted classification accuracy with the weights set such that all of the species are assigned the same total weight and the true negatives and true positives for each species have the same weight. The extra complexity exists purely to allow us to have a great deal of control over which birds are scored for a given soundscape. For offline cross validation purposes, the macro F1 is the closest analogue to the actual metric.",
    "1782219": "Thanks for sharing you CV and LB scores.\nThere seems to be a big gap between the CV and LB. Do you solve the gap?",
    "1786552": "Yes, my CV and LB also have a big gap. In fact, I'm not sure if I should trust CV or LB when submitting the final 2 results.\n\nAs I said in the [discussion](https://www.kaggle.com/competitions/birdclef-2022/discussion/324304#1784539), there are some reasons for each view, so I'm not sure how making a decision."
  },
  "source": "meta"
}