{
  "id": 160320,
  "title": "[Competition Metrics]",
  "url": "/competitions/birdsong-recognition/discussion/160320",
  "author_name": "",
  "post_date": "2020-06-20T19:03:45.339272700Z",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi everyone!</p>\n\n<p>I have created kernel with explanation and implementation of competition metrics:</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/competition-metrics\">[Competition Metrics]</a></p>\n\n<p>Hope it helps you!</p>\n\n<h3>If you find any imprecision in calculating metrics, please, let me know!</h3>",
  "messages": [
    {
      "id": "894798",
      "postDate": "06/20/2020 19:03:45",
      "content": "<p>Hi everyone!</p>\n\n<p>I have created kernel with explanation and implementation of competition metrics:</p>\n\n<p><a href=\"https://www.kaggle.com/shonenkov/competition-metrics\">[Competition Metrics]</a></p>\n\n<p>Hope it helps you!</p>\n\n<h3>If you find any imprecision in calculating metrics, please, let me know!</h3>",
      "rawMarkdown": "Hi everyone!\n\nI have created kernel with explanation and implementation of competition metrics:\n\n[[Competition Metrics]](https://www.kaggle.com/shonenkov/competition-metrics)\n\nHope it helps you!\n\n### If you find any imprecision in calculating metrics, please, let me know!",
      "votes": null
    },
    {
      "id": "895144",
      "postDate": "06/21/2020 06:12:32",
      "content": "<p>example:\n| actual | pred | TP | FN | FP | row's micro F1 |\n| --- | --- | --- | ---| ---| ---| \n|  \"a b\"|  \"a b c\" | 2 | 0 | 1 | 4/5 | \n|  \"m n o\"|  \"m p a\" | 1 | 2 | 2 | 1/3 | \n| Total  |    |3| 2| 3| |</p>\n\n<p>(1)  AVG of row wise F1  = 17/30\n(2) F1 using total TP, FN, FP = 6/11</p>\n\n<p>I am thinking comp metric is (1) but I can be wrong.</p>",
      "rawMarkdown": "example:\n| actual | pred | TP | FN | FP | row's micro F1 |\n| --- | --- | --- | ---| ---| ---| \n|  \"a b\"|  \"a b c\" | 2 | 0 | 1 | 4/5 | \n|  \"m n o\"|  \"m p a\" | 1 | 2 | 2 | 1/3 | \n| Total  |    |3| 2| 3| |\n\n(1)  AVG of row wise F1  = 17/30\n(2) F1 using total TP, FN, FP = 6/11\n\nI am thinking comp metric is (1) but I can be wrong.",
      "votes": null
    },
    {
      "id": "895161",
      "postDate": "06/21/2020 06:38:38",
      "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> thank you for example!</p>\n\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">In sklearn documentation about f1 score</a>:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F91ebe60d08b67b1f0953362d338ed46b%2FScreenshot%202020-06-21%20at%2009.30.23.png?generation=1592721081921068&amp;alt=media\" alt=\"\"></p>\n\n<blockquote>\n  <p>Submissions will be evaluated based on their row-wise <strong>micro averaged F1 score</strong>.</p>\n</blockquote>\n\n<p>So I suppose that correct variant is (2) (in your example). Globally counting the total TP, FN and FP. But I can be wrong also. </p>\n\n<p>What are you thinking about competition metric? <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> <a href=\"/tomdenton\">@tomdenton</a></p>",
      "rawMarkdown": "dhananjay3 thank you for example!\n\n[In sklearn documentation about f1 score](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F91ebe60d08b67b1f0953362d338ed46b%2FScreenshot%202020-06-21%20at%2009.30.23.png?generation=1592721081921068&amp;alt=media)\n\n&gt; Submissions will be evaluated based on their row-wise **micro averaged F1 score**.\n\nSo I suppose that correct variant is (2) (in your example). Globally counting the total TP, FN and FP. But I can be wrong also. \n\nWhat are you thinking about competition metric? @stefankahl @sohier @tomdenton",
      "votes": null
    },
    {
      "id": "895177",
      "postDate": "06/21/2020 06:52:18",
      "content": "<p>if you observe closely while we are calculating row wise F1 we are counting total TP, FP, FN in this row. and that's what I believe row wise micro F1 will be. for comparison row wise Macro F1 will be average of all F1 scores for all 264 classes for this row.</p>",
      "rawMarkdown": "if you observe closely while we are calculating row wise F1 we are counting total TP, FP, FN in this row. and that's what I believe row wise micro F1 will be. for comparison row wise Macro F1 will be average of all F1 scores for all 264 classes for this row.",
      "votes": null
    },
    {
      "id": "895184",
      "postDate": "06/21/2020 06:56:31",
      "content": "<p>I did a little experiment to check this exact thing. given an all 'nocall' submission scores 0.54. and if we modify few rows to predict all 264 classes that row's F1 will be nearly zero but number of this type of rows is low then (1) score should not be affected much. but in case (2) total FN and FP will be greatly increased and score (2) should be highly affected.</p>\n\n<p>Edit: looks like I forgot to import random and out of submissions for today. so I will update results tomorrow.</p>",
      "rawMarkdown": "I did a little experiment to check this exact thing. given an all 'nocall' submission scores 0.54. and if we modify few rows to predict all 264 classes that row's F1 will be nearly zero but number of this type of rows is low then (1) score should not be affected much. but in case (2) total FN and FP will be greatly increased and score (2) should be highly affected.\n\nEdit: looks like I forgot to import random and out of submissions for today. so I will update results tomorrow.",
      "votes": null
    },
    {
      "id": "895265",
      "postDate": "06/21/2020 08:03:48",
      "content": "<p><a href=\"/carriesmi\">@carriesmi</a> good idea, thank you! I have tried to make submission with this update:</p>\n\n<p><code>\nfor i in range(min([10, submission.shape[0]])):\n    submission.iloc[i]['birds'] = ' '.join(all_birds)\n</code></p>\n\n<p>It gives 0.54. I suppose (1) is correct finally. I will make some corrections in calculating metrics.</p>\n\n<p>Thank you <a href=\"/carriesmi\">@carriesmi</a> <a href=\"/dhananjay3\">@dhananjay3</a> for getting right answer together! :)</p>\n\n<p>I think you should make corrections in description metrics, current version is not clearly understanding - many competitors (not only novices) will ask these questions again and again, thanks in advance  <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> <a href=\"/tomdenton\">@tomdenton</a></p>",
      "rawMarkdown": "carriesmi good idea, thank you! I have tried to make submission with this update:\n\n```\nfor i in range(min([10, submission.shape[0]])):\n    submission.iloc[i]['birds'] = ' '.join(all_birds)\n```\n\nIt gives 0.54. I suppose (1) is correct finally. I will make some corrections in calculating metrics.\n\nThank you @carriesmi @dhananjay3 for getting right answer together! :)\n\nI think you should make corrections in description metrics, current version is not clearly understanding - many competitors (not only novices) will ask these questions again and again, thanks in advance  @stefankahl @sohier @tomdenton",
      "votes": null
    },
    {
      "id": "895978",
      "postDate": "06/21/2020 18:20:00",
      "content": "<p>Do I understand right that 0.54 for nocall means exactly that 54% of all observations have nocall, i.e. only 46% observations have birds in them?</p>",
      "rawMarkdown": "Do I understand right that 0.54 for nocall means exactly that 54% of all observations have nocall, i.e. only 46% observations have birds in them?",
      "votes": null
    },
    {
      "id": "979157",
      "postDate": "08/20/2020 16:58:47",
      "content": "<p>In the public test set, it seems possible. But I think for private set it would be completely different.</p>",
      "rawMarkdown": "In the public test set, it seems possible. But I think for private set it would be completely different.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 895144,
      "author_name": "dhananjay3",
      "author_url": "",
      "post_date": "06/21/2020 06:12:32",
      "content": "<p>example:\n| actual | pred | TP | FN | FP | row's micro F1 |\n| --- | --- | --- | ---| ---| ---| \n|  \"a b\"|  \"a b c\" | 2 | 0 | 1 | 4/5 | \n|  \"m n o\"|  \"m p a\" | 1 | 2 | 2 | 1/3 | \n| Total  |    |3| 2| 3| |</p>\n\n<p>(1)  AVG of row wise F1  = 17/30\n(2) F1 using total TP, FN, FP = 6/11</p>\n\n<p>I am thinking comp metric is (1) but I can be wrong.</p>",
      "votes": null,
      "replies": [
        {
          "id": 895161,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/21/2020 06:38:38",
          "content": "<p><a href=\"/dhananjay3\">@dhananjay3</a> thank you for example!</p>\n\n<p><a href=\"https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html\">In sklearn documentation about f1 score</a>:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F91ebe60d08b67b1f0953362d338ed46b%2FScreenshot%202020-06-21%20at%2009.30.23.png?generation=1592721081921068&amp;alt=media\" alt=\"\"></p>\n\n<blockquote>\n  <p>Submissions will be evaluated based on their row-wise <strong>micro averaged F1 score</strong>.</p>\n</blockquote>\n\n<p>So I suppose that correct variant is (2) (in your example). Globally counting the total TP, FN and FP. But I can be wrong also. </p>\n\n<p>What are you thinking about competition metric? <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> <a href=\"/tomdenton\">@tomdenton</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 895177,
          "author_name": "dhananjay3",
          "author_url": "",
          "post_date": "06/21/2020 06:52:18",
          "content": "<p>if you observe closely while we are calculating row wise F1 we are counting total TP, FP, FN in this row. and that's what I believe row wise micro F1 will be. for comparison row wise Macro F1 will be average of all F1 scores for all 264 classes for this row.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 895184,
          "author_name": "carriesmi",
          "author_url": "",
          "post_date": "06/21/2020 06:56:31",
          "content": "<p>I did a little experiment to check this exact thing. given an all 'nocall' submission scores 0.54. and if we modify few rows to predict all 264 classes that row's F1 will be nearly zero but number of this type of rows is low then (1) score should not be affected much. but in case (2) total FN and FP will be greatly increased and score (2) should be highly affected.</p>\n\n<p>Edit: looks like I forgot to import random and out of submissions for today. so I will update results tomorrow.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 895265,
          "author_name": "shonenkov",
          "author_url": "",
          "post_date": "06/21/2020 08:03:48",
          "content": "<p><a href=\"/carriesmi\">@carriesmi</a> good idea, thank you! I have tried to make submission with this update:</p>\n\n<p><code>\nfor i in range(min([10, submission.shape[0]])):\n    submission.iloc[i]['birds'] = ' '.join(all_birds)\n</code></p>\n\n<p>It gives 0.54. I suppose (1) is correct finally. I will make some corrections in calculating metrics.</p>\n\n<p>Thank you <a href=\"/carriesmi\">@carriesmi</a> <a href=\"/dhananjay3\">@dhananjay3</a> for getting right answer together! :)</p>\n\n<p>I think you should make corrections in description metrics, current version is not clearly understanding - many competitors (not only novices) will ask these questions again and again, thanks in advance  <a href=\"/stefankahl\">@stefankahl</a> <a href=\"/sohier\">@sohier</a> <a href=\"/tomdenton\">@tomdenton</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 895978,
      "author_name": "snovik1975",
      "author_url": "",
      "post_date": "06/21/2020 18:20:00",
      "content": "<p>Do I understand right that 0.54 for nocall means exactly that 54% of all observations have nocall, i.e. only 46% observations have birds in them?</p>",
      "votes": null,
      "replies": [
        {
          "id": 979157,
          "author_name": "rsinda",
          "author_url": "",
          "post_date": "08/20/2020 16:58:47",
          "content": "<p>In the public test set, it seems possible. But I think for private set it would be completely different.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "894798": "Hi everyone!\n\nI have created kernel with explanation and implementation of competition metrics:\n\n[[Competition Metrics]](https://www.kaggle.com/shonenkov/competition-metrics)\n\nHope it helps you!\n\n### If you find any imprecision in calculating metrics, please, let me know!",
    "895144": "example:\n| actual | pred | TP | FN | FP | row's micro F1 |\n| --- | --- | --- | ---| ---| ---| \n|  \"a b\"|  \"a b c\" | 2 | 0 | 1 | 4/5 | \n|  \"m n o\"|  \"m p a\" | 1 | 2 | 2 | 1/3 | \n| Total  |    |3| 2| 3| |\n\n(1)  AVG of row wise F1  = 17/30\n(2) F1 using total TP, FN, FP = 6/11\n\nI am thinking comp metric is (1) but I can be wrong.",
    "895161": "dhananjay3 thank you for example!\n\n[In sklearn documentation about f1 score](https://scikit-learn.org/stable/modules/generated/sklearn.metrics.f1_score.html):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1920073%2F91ebe60d08b67b1f0953362d338ed46b%2FScreenshot%202020-06-21%20at%2009.30.23.png?generation=1592721081921068&amp;alt=media)\n\n&gt; Submissions will be evaluated based on their row-wise **micro averaged F1 score**.\n\nSo I suppose that correct variant is (2) (in your example). Globally counting the total TP, FN and FP. But I can be wrong also. \n\nWhat are you thinking about competition metric? @stefankahl @sohier @tomdenton",
    "895177": "if you observe closely while we are calculating row wise F1 we are counting total TP, FP, FN in this row. and that's what I believe row wise micro F1 will be. for comparison row wise Macro F1 will be average of all F1 scores for all 264 classes for this row.",
    "895184": "I did a little experiment to check this exact thing. given an all 'nocall' submission scores 0.54. and if we modify few rows to predict all 264 classes that row's F1 will be nearly zero but number of this type of rows is low then (1) score should not be affected much. but in case (2) total FN and FP will be greatly increased and score (2) should be highly affected.\n\nEdit: looks like I forgot to import random and out of submissions for today. so I will update results tomorrow.",
    "895265": "carriesmi good idea, thank you! I have tried to make submission with this update:\n\n```\nfor i in range(min([10, submission.shape[0]])):\n    submission.iloc[i]['birds'] = ' '.join(all_birds)\n```\n\nIt gives 0.54. I suppose (1) is correct finally. I will make some corrections in calculating metrics.\n\nThank you @carriesmi @dhananjay3 for getting right answer together! :)\n\nI think you should make corrections in description metrics, current version is not clearly understanding - many competitors (not only novices) will ask these questions again and again, thanks in advance  @stefankahl @sohier @tomdenton",
    "895978": "Do I understand right that 0.54 for nocall means exactly that 54% of all observations have nocall, i.e. only 46% observations have birds in them?",
    "979157": "In the public test set, it seems possible. But I think for private set it would be completely different."
  },
  "source": "meta"
}