{
  "id": 310335,
  "title": "Expected test data distribution",
  "url": "/competitions/birdclef-2022/discussion/310335",
  "author_name": "",
  "post_date": "2022-02-28T19:27:27.380214800Z",
  "votes": 10,
  "comment_count": 9,
  "views": 0,
  "content": "<p>I found quite an interesting pattern with sending fake/synthetic prediction data:</p>\n<ul>\n<li>all <strong>True</strong>: score <strong>0.51</strong></li>\n<li>all <strong>False</strong>: score <strong>0.48</strong></li>\n<li>random: score <strong>0.50</strong></li>\n</ul>\n<p><a href=\"https://www.kaggle.com/ethanwharris\" target=\"_blank\">@ethanwharris</a> guess generally we can say that the TP rate is the number of positive cases in the test set (since we always predict positive) and that the FN rate is zero so then it’s: <code>positive_cases / positive_cases + 0.5 * (num_cases - positive_cases)</code>. Then it’s: <code>positive_cases / 0.5 * (positive_cases + num_cases)</code>. Then convert to percentage so: <code>positive_fraction / 0.5 * positive_fraction + 0.5</code>.<br>\nThen set that equal to 0.5, we have 1/4 = 3/4 positive_fraction. So the number of positive cases in the predicted data is like 1/3…</p>",
  "messages": [
    {
      "id": "1707787",
      "postDate": "02/28/2022 19:27:27",
      "content": "<p>I found quite an interesting pattern with sending fake/synthetic prediction data:</p>\n<ul>\n<li>all <strong>True</strong>: score <strong>0.51</strong></li>\n<li>all <strong>False</strong>: score <strong>0.48</strong></li>\n<li>random: score <strong>0.50</strong></li>\n</ul>\n<p><a href=\"https://www.kaggle.com/ethanwharris\" target=\"_blank\">@ethanwharris</a> guess generally we can say that the TP rate is the number of positive cases in the test set (since we always predict positive) and that the FN rate is zero so then it’s: <code>positive_cases / positive_cases + 0.5 * (num_cases - positive_cases)</code>. Then it’s: <code>positive_cases / 0.5 * (positive_cases + num_cases)</code>. Then convert to percentage so: <code>positive_fraction / 0.5 * positive_fraction + 0.5</code>.<br>\nThen set that equal to 0.5, we have 1/4 = 3/4 positive_fraction. So the number of positive cases in the predicted data is like 1/3…</p>",
      "rawMarkdown": "I found quite an interesting pattern with sending fake/synthetic prediction data:\n\n- all **True**: score **0.51**\n- all **False**: score **0.48**\n- random: score **0.50**\n\n@ethanwharris guess generally we can say that the TP rate is the number of positive cases in the test set (since we always predict positive) and that the FN rate is zero so then it’s: `positive_cases / positive_cases + 0.5 * (num_cases - positive_cases)`. Then it’s: `positive_cases / 0.5 * (positive_cases + num_cases)`. Then convert to percentage so: `positive_fraction / 0.5 * positive_fraction + 0.5`.\nThen set that equal to 0.5, we have 1/4 = 3/4 positive_fraction. So the number of positive cases in the predicted data is like 1/3…",
      "votes": null
    },
    {
      "id": "1707799",
      "postDate": "02/28/2022 19:53:54",
      "content": "<p>How do you explain that all false yields 0.48 then?</p>",
      "rawMarkdown": "How do you explain that all false yields 0.48 then?",
      "votes": null
    },
    {
      "id": "1707801",
      "postDate": "02/28/2022 20:06:30",
      "content": "<p>simple experiment:</p>\n<pre><code>from sklearn.metrics import f1_score\n\ndef compute_f1(lb=0):\n    y_pred = [lb] * 100\n\n    f1 = []\n    ratio = []\n    for r in range(100):\n        y_true = [0] * r + [1] * (100 - r)\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n\nplt.plot(*compute_f1(0), label=\"all False\")\nplt.plot(*compute_f1(1), label=\"all True\")\nplt.grid(), plt.legend()\nplt.xlabel(\"ratio for 0 label\")\nplt.ylabel(\"F1 score\")\n</code></pre>",
      "rawMarkdown": "simple experiment:\n```\nfrom sklearn.metrics import f1_score\n\ndef compute_f1(lb=0):\n    y_pred = [lb] * 100\n\n    f1 = []\n    ratio = []\n    for r in range(100):\n        y_true = [0] * r + [1] * (100 - r)\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n\nplt.plot(*compute_f1(0), label=\"all False\")\nplt.plot(*compute_f1(1), label=\"all True\")\nplt.grid(), plt.legend()\nplt.xlabel(\"ratio for 0 label\")\nplt.ylabel(\"F1 score\")\n```",
      "votes": null
    },
    {
      "id": "1707803",
      "postDate": "02/28/2022 20:08:05",
      "content": "<p>not sure, but it is the result of this competition evaluation…</p>",
      "rawMarkdown": "not sure, but it is the result of this competition evaluation...",
      "votes": null
    },
    {
      "id": "1707819",
      "postDate": "02/28/2022 20:36:28",
      "content": "<p>I know it is the result.  And it is incompatible with f1 score as you define dit in your post unless mistaken.</p>\n<p>I am not teasing you, I honestly have doubts about the metric definition.  </p>",
      "rawMarkdown": "I know it is the result.  And it is incompatible with f1 score as you define dit in your post unless mistaken.\n\nI am not teasing you, I honestly have doubts about the metric definition.",
      "votes": null
    },
    {
      "id": "1707827",
      "postDate": "02/28/2022 20:54:06",
      "content": "<p>From this experiment there is no ratio that explains the observed results.</p>",
      "rawMarkdown": "From this experiment there is no ratio that explains the observed results.",
      "votes": null
    },
    {
      "id": "1707845",
      "postDate": "02/28/2022 21:41:11",
      "content": "<p>well, I have a few more open questions about this competition, including <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/309001\" target=\"_blank\">Dicrepency in test data</a></p>",
      "rawMarkdown": "well, I have a few more open questions about this competition, including [Dicrepency in test data](https://www.kaggle.com/c/birdclef-2022/discussion/309001)",
      "votes": null
    },
    {
      "id": "1707851",
      "postDate": "02/28/2022 21:48:36",
      "content": "<p>exactly, which is why I start this topic and hope that one of the hosts could bring some lite to the evaluation… 🙏</p>",
      "rawMarkdown": "exactly, which is why I start this topic and hope that one of the hosts could bring some lite to the evaluation... 🙏",
      "votes": null
    },
    {
      "id": "1707894",
      "postDate": "02/28/2022 23:23:39",
      "content": "<p>You assumed a single class, which is not the case here.</p>",
      "rawMarkdown": "You assumed a single class, which is not the case here.",
      "votes": null
    },
    {
      "id": "1708173",
      "postDate": "03/01/2022 07:14:50",
      "content": "<p>fair point but if I adjust it to be a random shuffle for binary encoding, <code>sklearn</code> gives 0 for all False predictions regardless of the True/False dataset distribution…</p>\n<pre><code>def compute_f1(lb=0):\n    y_pred = [[lb] * 100] * 1000\n\n    f1 = []\n    ratio = []\n    for r in tqdm(range(100)):\n        y = [0] * r + [1] * (100 - r)\n        y_true = []\n        for k in range(1000):\n            np.random.shuffle(y)\n            y_true.append(list(y))\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n</code></pre>",
      "rawMarkdown": "fair point but if I adjust it to be a random shuffle for binary encoding, `sklearn` gives 0 for all False predictions regardless of the True/False dataset distribution...\n```\ndef compute_f1(lb=0):\n    y_pred = [[lb] * 100] * 1000\n\n    f1 = []\n    ratio = []\n    for r in tqdm(range(100)):\n        y = [0] * r + [1] * (100 - r)\n        y_true = []\n        for k in range(1000):\n            np.random.shuffle(y)\n            y_true.append(list(y))\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n```",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1707799,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "02/28/2022 19:53:54",
      "content": "<p>How do you explain that all false yields 0.48 then?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1707803,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "02/28/2022 20:08:05",
          "content": "<p>not sure, but it is the result of this competition evaluation…</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1707819,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/28/2022 20:36:28",
          "content": "<p>I know it is the result.  And it is incompatible with f1 score as you define dit in your post unless mistaken.</p>\n<p>I am not teasing you, I honestly have doubts about the metric definition.  </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1707845,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "02/28/2022 21:41:11",
          "content": "<p>well, I have a few more open questions about this competition, including <a href=\"https://www.kaggle.com/c/birdclef-2022/discussion/309001\" target=\"_blank\">Dicrepency in test data</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1707801,
      "author_name": "jirkaborovec",
      "author_url": "",
      "post_date": "02/28/2022 20:06:30",
      "content": "<p>simple experiment:</p>\n<pre><code>from sklearn.metrics import f1_score\n\ndef compute_f1(lb=0):\n    y_pred = [lb] * 100\n\n    f1 = []\n    ratio = []\n    for r in range(100):\n        y_true = [0] * r + [1] * (100 - r)\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n\nplt.plot(*compute_f1(0), label=\"all False\")\nplt.plot(*compute_f1(1), label=\"all True\")\nplt.grid(), plt.legend()\nplt.xlabel(\"ratio for 0 label\")\nplt.ylabel(\"F1 score\")\n</code></pre>",
      "votes": null,
      "replies": [
        {
          "id": 1707827,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/28/2022 20:54:06",
          "content": "<p>From this experiment there is no ratio that explains the observed results.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1707851,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "02/28/2022 21:48:36",
          "content": "<p>exactly, which is why I start this topic and hope that one of the hosts could bring some lite to the evaluation… 🙏</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1707894,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "02/28/2022 23:23:39",
          "content": "<p>You assumed a single class, which is not the case here.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1708173,
          "author_name": "jirkaborovec",
          "author_url": "",
          "post_date": "03/01/2022 07:14:50",
          "content": "<p>fair point but if I adjust it to be a random shuffle for binary encoding, <code>sklearn</code> gives 0 for all False predictions regardless of the True/False dataset distribution…</p>\n<pre><code>def compute_f1(lb=0):\n    y_pred = [[lb] * 100] * 1000\n\n    f1 = []\n    ratio = []\n    for r in tqdm(range(100)):\n        y = [0] * r + [1] * (100 - r)\n        y_true = []\n        for k in range(1000):\n            np.random.shuffle(y)\n            y_true.append(list(y))\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n</code></pre>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1707787": "I found quite an interesting pattern with sending fake/synthetic prediction data:\n\n- all **True**: score **0.51**\n- all **False**: score **0.48**\n- random: score **0.50**\n\n@ethanwharris guess generally we can say that the TP rate is the number of positive cases in the test set (since we always predict positive) and that the FN rate is zero so then it’s: `positive_cases / positive_cases + 0.5 * (num_cases - positive_cases)`. Then it’s: `positive_cases / 0.5 * (positive_cases + num_cases)`. Then convert to percentage so: `positive_fraction / 0.5 * positive_fraction + 0.5`.\nThen set that equal to 0.5, we have 1/4 = 3/4 positive_fraction. So the number of positive cases in the predicted data is like 1/3…",
    "1707799": "How do you explain that all false yields 0.48 then?",
    "1707801": "simple experiment:\n```\nfrom sklearn.metrics import f1_score\n\ndef compute_f1(lb=0):\n    y_pred = [lb] * 100\n\n    f1 = []\n    ratio = []\n    for r in range(100):\n        y_true = [0] * r + [1] * (100 - r)\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n\nplt.plot(*compute_f1(0), label=\"all False\")\nplt.plot(*compute_f1(1), label=\"all True\")\nplt.grid(), plt.legend()\nplt.xlabel(\"ratio for 0 label\")\nplt.ylabel(\"F1 score\")\n```",
    "1707803": "not sure, but it is the result of this competition evaluation...",
    "1707819": "I know it is the result.  And it is incompatible with f1 score as you define dit in your post unless mistaken.\n\nI am not teasing you, I honestly have doubts about the metric definition.",
    "1707827": "From this experiment there is no ratio that explains the observed results.",
    "1707845": "well, I have a few more open questions about this competition, including [Dicrepency in test data](https://www.kaggle.com/c/birdclef-2022/discussion/309001)",
    "1707851": "exactly, which is why I start this topic and hope that one of the hosts could bring some lite to the evaluation... 🙏",
    "1707894": "You assumed a single class, which is not the case here.",
    "1708173": "fair point but if I adjust it to be a random shuffle for binary encoding, `sklearn` gives 0 for all False predictions regardless of the True/False dataset distribution...\n```\ndef compute_f1(lb=0):\n    y_pred = [[lb] * 100] * 1000\n\n    f1 = []\n    ratio = []\n    for r in tqdm(range(100)):\n        y = [0] * r + [1] * (100 - r)\n        y_true = []\n        for k in range(1000):\n            np.random.shuffle(y)\n            y_true.append(list(y))\n        f1.append(f1_score(y_true, y_pred, average='macro'))\n        ratio.append(r / 100.)\n    return ratio, f1\n```"
  },
  "source": "meta"
}