{
  "id": 237697,
  "title": "Is anyone else using tfa.metrics.FBetaScore?",
  "url": "/competitions/birdclef-2021/discussion/237697",
  "author_name": "",
  "post_date": "2021-05-09T21:02:15.866260500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Is anyone else using tfa.metrics.FBetaScore in their development to see the performance of their model?</p>\n<p>I'm importing tensorflow_addons as tfa and using the following code snippet:</p>\n<pre><code>model.compile(optimizer=optimiser,\n              loss='sparse_categorical_crossentropy',\n              metrics=['accuracy',tfa.metrics.FBetaScore(num_classes=15, average=\"micro\")])\n</code></pre>\n<p>but I end up with the same FBetaScore (0.1322) for every epoch even though my accuracy has improved from 32% during the first epoch to 79% for the last (I'm using a subset of the data right now with only 15 classes).</p>\n<p>I imagine my FBeta score should improve with each epoch since my accuracy is but that could be wrong. If anyone has any suggestions please let me know.</p>\n<p>Thank you!</p>",
  "messages": [
    {
      "id": "1299604",
      "postDate": "05/09/2021 21:02:15",
      "content": "<p>Is anyone else using tfa.metrics.FBetaScore in their development to see the performance of their model?</p>\n<p>I'm importing tensorflow_addons as tfa and using the following code snippet:</p>\n<pre><code>model.compile(optimizer=optimiser,\n              loss='sparse_categorical_crossentropy',\n              metrics=['accuracy',tfa.metrics.FBetaScore(num_classes=15, average=\"micro\")])\n</code></pre>\n<p>but I end up with the same FBetaScore (0.1322) for every epoch even though my accuracy has improved from 32% during the first epoch to 79% for the last (I'm using a subset of the data right now with only 15 classes).</p>\n<p>I imagine my FBeta score should improve with each epoch since my accuracy is but that could be wrong. If anyone has any suggestions please let me know.</p>\n<p>Thank you!</p>",
      "rawMarkdown": "Is anyone else using tfa.metrics.FBetaScore in their development to see the performance of their model?\n\nI'm importing tensorflow_addons as tfa and using the following code snippet:\n\n    model.compile(optimizer=optimiser,\n                  loss='sparse_categorical_crossentropy',\n                  metrics=['accuracy',tfa.metrics.FBetaScore(num_classes=15, average=\"micro\")])\n\nbut I end up with the same FBetaScore (0.1322) for every epoch even though my accuracy has improved from 32% during the first epoch to 79% for the last (I'm using a subset of the data right now with only 15 classes).\n\nI imagine my FBeta score should improve with each epoch since my accuracy is but that could be wrong. If anyone has any suggestions please let me know.\n\nThank you!",
      "votes": null
    },
    {
      "id": "1299660",
      "postDate": "05/09/2021 23:06:54",
      "content": "<p>This looks weird: <code>num_classes=15</code></p>\n<p>There are 397 bird species here.</p>",
      "rawMarkdown": "This looks weird: `num_classes=15`\n\nThere are 397 bird species here.",
      "votes": null
    },
    {
      "id": "1299709",
      "postDate": "05/10/2021 00:57:28",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I'm actually using a small subset of the classes right now. The accuracies don't represent the performance over the entire data set.</p>\n<p>I'm just trying to find ways to evaluate the performance in terms of mean F-score beta since I believe that's what the leaderboards are ranking us on. Did you do something similar?</p>",
      "rawMarkdown": "cpmpml I'm actually using a small subset of the classes right now. The accuracies don't represent the performance over the entire data set.\n\nI'm just trying to find ways to evaluate the performance in terms of mean F-score beta since I believe that's what the leaderboards are ranking us on. Did you do something similar?",
      "votes": null
    },
    {
      "id": "1299907",
      "postDate": "05/10/2021 06:23:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/megilgallant\" target=\"_blank\">@megilgallant</a> </p>\n<p>I always like to have fun with tfa(tensorflow_addons)</p>\n<p>I suggest paying with  <code>threshold</code>  Args like this :</p>\n<p>pip install tfa-nightly</p>\n<pre><code>metrics=['accuracy',\n              tfa.metrics.FBetaScore(name = 'FBetaScore_Default',  num_classes=15, average=\"micro\"),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.9',  num_classes=15, average=\"micro\" , threshold=0.9),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.5',  num_classes=15, average=\"micro\" , threshold=0.5),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.3',  num_classes=15, average=\"micro\" , threshold=0.3),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.1',  num_classes=15, average=\"micro\" , threshold=0.1),\n              ]) \n</code></pre>",
      "rawMarkdown": "Hi @megilgallant \n\nI always like to have fun with tfa(tensorflow_addons)\n\nI suggest paying with  `threshold`  Args like this :\n\n pip install tfa-nightly\n\n\n```\nmetrics=['accuracy',\n              tfa.metrics.FBetaScore(name = 'FBetaScore_Default',  num_classes=15, average=\"micro\"),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.9',  num_classes=15, average=\"micro\" , threshold=0.9),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.5',  num_classes=15, average=\"micro\" , threshold=0.5),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.3',  num_classes=15, average=\"micro\" , threshold=0.3),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.1',  num_classes=15, average=\"micro\" , threshold=0.1),\n              ]) \n```",
      "votes": null
    },
    {
      "id": "1300158",
      "postDate": "05/10/2021 09:58:11",
      "content": "<blockquote>\n  <p>I'm actually using a small subset of the classes right now.</p>\n</blockquote>\n<p>This may be the cause of the behavior you observe.</p>\n<p>Also, it looks like this metric does not implement the 'samples' average which is used in this competition.</p>",
      "rawMarkdown": "> I'm actually using a small subset of the classes right now.\n\nThis may be the cause of the behavior you observe.\n\nAlso, it looks like this metric does not implement the 'samples' average which is used in this competition.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1299660,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "05/09/2021 23:06:54",
      "content": "<p>This looks weird: <code>num_classes=15</code></p>\n<p>There are 397 bird species here.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1299709,
      "author_name": "megilgallant",
      "author_url": "",
      "post_date": "05/10/2021 00:57:28",
      "content": "<p><a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> I'm actually using a small subset of the classes right now. The accuracies don't represent the performance over the entire data set.</p>\n<p>I'm just trying to find ways to evaluate the performance in terms of mean F-score beta since I believe that's what the leaderboards are ranking us on. Did you do something similar?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1300158,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "05/10/2021 09:58:11",
          "content": "<blockquote>\n  <p>I'm actually using a small subset of the classes right now.</p>\n</blockquote>\n<p>This may be the cause of the behavior you observe.</p>\n<p>Also, it looks like this metric does not implement the 'samples' average which is used in this competition.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1299907,
      "author_name": "faisalalsrheed",
      "author_url": "",
      "post_date": "05/10/2021 06:23:27",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/megilgallant\" target=\"_blank\">@megilgallant</a> </p>\n<p>I always like to have fun with tfa(tensorflow_addons)</p>\n<p>I suggest paying with  <code>threshold</code>  Args like this :</p>\n<p>pip install tfa-nightly</p>\n<pre><code>metrics=['accuracy',\n              tfa.metrics.FBetaScore(name = 'FBetaScore_Default',  num_classes=15, average=\"micro\"),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.9',  num_classes=15, average=\"micro\" , threshold=0.9),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.5',  num_classes=15, average=\"micro\" , threshold=0.5),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.3',  num_classes=15, average=\"micro\" , threshold=0.3),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.1',  num_classes=15, average=\"micro\" , threshold=0.1),\n              ]) \n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1299604": "Is anyone else using tfa.metrics.FBetaScore in their development to see the performance of their model?\n\nI'm importing tensorflow_addons as tfa and using the following code snippet:\n\n    model.compile(optimizer=optimiser,\n                  loss='sparse_categorical_crossentropy',\n                  metrics=['accuracy',tfa.metrics.FBetaScore(num_classes=15, average=\"micro\")])\n\nbut I end up with the same FBetaScore (0.1322) for every epoch even though my accuracy has improved from 32% during the first epoch to 79% for the last (I'm using a subset of the data right now with only 15 classes).\n\nI imagine my FBeta score should improve with each epoch since my accuracy is but that could be wrong. If anyone has any suggestions please let me know.\n\nThank you!",
    "1299660": "This looks weird: `num_classes=15`\n\nThere are 397 bird species here.",
    "1299709": "cpmpml I'm actually using a small subset of the classes right now. The accuracies don't represent the performance over the entire data set.\n\nI'm just trying to find ways to evaluate the performance in terms of mean F-score beta since I believe that's what the leaderboards are ranking us on. Did you do something similar?",
    "1299907": "Hi @megilgallant \n\nI always like to have fun with tfa(tensorflow_addons)\n\nI suggest paying with  `threshold`  Args like this :\n\n pip install tfa-nightly\n\n\n```\nmetrics=['accuracy',\n              tfa.metrics.FBetaScore(name = 'FBetaScore_Default',  num_classes=15, average=\"micro\"),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.9',  num_classes=15, average=\"micro\" , threshold=0.9),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.5',  num_classes=15, average=\"micro\" , threshold=0.5),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.3',  num_classes=15, average=\"micro\" , threshold=0.3),\n              tfa.metrics.FBetaScore(name = 'FBetaScore_threshold=0.1',  num_classes=15, average=\"micro\" , threshold=0.1),\n              ]) \n```",
    "1300158": "> I'm actually using a small subset of the classes right now.\n\nThis may be the cause of the behavior you observe.\n\nAlso, it looks like this metric does not implement the 'samples' average which is used in this competition."
  },
  "source": "meta"
}