{
  "id": 319582,
  "title": "The reason SED model's optimal threshold differs between validation and public score",
  "url": "/competitions/birdclef-2022/discussion/319582",
  "author_name": "",
  "post_date": "2022-04-18T04:00:36.708455800Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>In my experience, optimal threshold to calculate F1 score in validation is different from that of the public score, and I was wondering why it happens. Now, I have a hypothesis about it.</p>\n<p>Clipwise output calculated in AttBlockV2 layer is below.<br>\n(reference: <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter</a>)</p>\n<pre><code>    def forward(self, x):\n        # x: (n_samples, n_in, n_time)\n        norm_att = torch.softmax(torch.tanh(self.att(x)), dim=-1)\n        cla = self.nonlinear_transform(self.cla(x))\n        x = torch.sum(norm_att * cla, dim=2)\n        return x, norm_att, cla\n</code></pre>\n<p>As you can see, x (that is equivalent to clipwise output) is represented as a summation of (norm_att × cla) along dimension 2. </p>\n<p>If you set parameter : period larger 5 seconds to deal with weak label problem, the dimension 2 get bigger value, and as a result, the value of x can also be bigger.<br>\n(You can learn about weak label and sound event detection in this paper) <a href=\"https://arxiv.org/pdf/2107.05463.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.05463.pdf</a>)</p>\n<p>That is why the optimal threshold for the validation step is bigger than that used in the inference step when we set larger parameter : period in training step in general.</p>\n<p>(I think other reasons  can exist, but I don't come up with.)</p>",
  "messages": [
    {
      "id": "1758775",
      "postDate": "04/18/2022 04:00:36",
      "content": "<p>In my experience, optimal threshold to calculate F1 score in validation is different from that of the public score, and I was wondering why it happens. Now, I have a hypothesis about it.</p>\n<p>Clipwise output calculated in AttBlockV2 layer is below.<br>\n(reference: <a href=\"https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter\" target=\"_blank\">https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter</a>)</p>\n<pre><code>    def forward(self, x):\n        # x: (n_samples, n_in, n_time)\n        norm_att = torch.softmax(torch.tanh(self.att(x)), dim=-1)\n        cla = self.nonlinear_transform(self.cla(x))\n        x = torch.sum(norm_att * cla, dim=2)\n        return x, norm_att, cla\n</code></pre>\n<p>As you can see, x (that is equivalent to clipwise output) is represented as a summation of (norm_att × cla) along dimension 2. </p>\n<p>If you set parameter : period larger 5 seconds to deal with weak label problem, the dimension 2 get bigger value, and as a result, the value of x can also be bigger.<br>\n(You can learn about weak label and sound event detection in this paper) <a href=\"https://arxiv.org/pdf/2107.05463.pdf\" target=\"_blank\">https://arxiv.org/pdf/2107.05463.pdf</a>)</p>\n<p>That is why the optimal threshold for the validation step is bigger than that used in the inference step when we set larger parameter : period in training step in general.</p>\n<p>(I think other reasons  can exist, but I don't come up with.)</p>",
      "rawMarkdown": "In my experience, optimal threshold to calculate F1 score in validation is different from that of the public score, and I was wondering why it happens. Now, I have a hypothesis about it.\n\nClipwise output calculated in AttBlockV2 layer is below.\n(reference: https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter)\n\n```\n    def forward(self, x):\n        # x: (n_samples, n_in, n_time)\n        norm_att = torch.softmax(torch.tanh(self.att(x)), dim=-1)\n        cla = self.nonlinear_transform(self.cla(x))\n        x = torch.sum(norm_att * cla, dim=2)\n        return x, norm_att, cla\n```\n\nAs you can see, x (that is equivalent to clipwise output) is represented as a summation of (norm_att × cla) along dimension 2. \n\nIf you set parameter : period larger 5 seconds to deal with weak label problem, the dimension 2 get bigger value, and as a result, the value of x can also be bigger.\n(You can learn about weak label and sound event detection in this paper) https://arxiv.org/pdf/2107.05463.pdf)\n\nThat is why the optimal threshold for the validation step is bigger than that used in the inference step when we set larger parameter : period in training step in general.\n\n(I think other reasons  can exist, but I don't come up with.)",
      "votes": null
    },
    {
      "id": "1764024",
      "postDate": "04/22/2022 04:46:00",
      "content": "<p>if my public score is low (like 0.53) while I'm getting a high score when predicting on my validation dataset<br>\n(around 0.89), do you mean I may get a high score on the private score?</p>",
      "rawMarkdown": "if my public score is low (like 0.53) while I'm getting a high score when predicting on my validation dataset\n(around 0.89), do you mean I may get a high score on the private score?",
      "votes": null
    },
    {
      "id": "1764131",
      "postDate": "04/22/2022 07:26:04",
      "content": "<p>There is a big dissociation between the validation score and the public score.<br>\nPossibly, domain shift between train data and test data cause this dissociation<br>\n(the difference of sound collection method can cause domain shift.)</p>",
      "rawMarkdown": "There is a big dissociation between the validation score and the public score.\nPossibly, domain shift between train data and test data cause this dissociation\n(the difference of sound collection method can cause domain shift.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1764024,
      "author_name": "xv7d111",
      "author_url": "",
      "post_date": "04/22/2022 04:46:00",
      "content": "<p>if my public score is low (like 0.53) while I'm getting a high score when predicting on my validation dataset<br>\n(around 0.89), do you mean I may get a high score on the private score?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1764131,
          "author_name": "kotanoda",
          "author_url": "",
          "post_date": "04/22/2022 07:26:04",
          "content": "<p>There is a big dissociation between the validation score and the public score.<br>\nPossibly, domain shift between train data and test data cause this dissociation<br>\n(the difference of sound collection method can cause domain shift.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1758775": "In my experience, optimal threshold to calculate F1 score in validation is different from that of the public score, and I was wondering why it happens. Now, I have a hypothesis about it.\n\nClipwise output calculated in AttBlockV2 layer is below.\n(reference: https://www.kaggle.com/code/hidehisaarai1213/pytorch-training-birdclef2021-starter)\n\n```\n    def forward(self, x):\n        # x: (n_samples, n_in, n_time)\n        norm_att = torch.softmax(torch.tanh(self.att(x)), dim=-1)\n        cla = self.nonlinear_transform(self.cla(x))\n        x = torch.sum(norm_att * cla, dim=2)\n        return x, norm_att, cla\n```\n\nAs you can see, x (that is equivalent to clipwise output) is represented as a summation of (norm_att × cla) along dimension 2. \n\nIf you set parameter : period larger 5 seconds to deal with weak label problem, the dimension 2 get bigger value, and as a result, the value of x can also be bigger.\n(You can learn about weak label and sound event detection in this paper) https://arxiv.org/pdf/2107.05463.pdf)\n\nThat is why the optimal threshold for the validation step is bigger than that used in the inference step when we set larger parameter : period in training step in general.\n\n(I think other reasons  can exist, but I don't come up with.)",
    "1764024": "if my public score is low (like 0.53) while I'm getting a high score when predicting on my validation dataset\n(around 0.89), do you mean I may get a high score on the private score?",
    "1764131": "There is a big dissociation between the validation score and the public score.\nPossibly, domain shift between train data and test data cause this dissociation\n(the difference of sound collection method can cause domain shift.)"
  },
  "source": "meta"
}