{
  "id": 173276,
  "title": "Clarify competition metric",
  "url": "/competitions/birdsong-recognition/discussion/173276",
  "author_name": "Volodymyr",
  "post_date": "2020-08-08T15:06:11.643000",
  "votes": 3,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Also maybe we have the same discussion earlier. But I did not find it. So maybe someone can clarify is such metric computation is correct?</p>\n\n<p>```\ndef compute_raw_wise_f1micro(\n    y_true: np.ndarray, \n    y_pred: np.ndarray, \n    threshold: int = 0.5\n):\n    \"\"\"\n    Competion metric</p>\n\n<pre><code>Parameters\n----------\ny_true: np.ndarray\n    Array with indexes of true bird classes\n    Dim : [N]\ny_pred: np.ndarray\n    Array with predicted probabilities for each row\n    Dim: [N,n_classes]\nthreshold: int\n    Threshold for binirization of predicted classes\n\"\"\"\nscore = []\nn_labels = y_pred.shape[1]\n\nfor t, p in zip(y_true, y_pred):\n    # Here we have only one bird class\n    # So we wrap it into list for correct f1 computation\n    t = [t]\n    # Get all bird classes that we predict\n    # higher then threshold\n    p = np.where(p &amp;gt; threshold)[0]\n    # If we do not have any predicted classes\n    # then we put class `nocall`\n    if len(p) == 0:\n        p = [n_labels]   \n    # For all extra predicted classes\n    # we padd true values with `nocall` class\n    if len(t) &amp;lt; len(p):\n        t = t + [n_labels]*(len(p)-len(t))\n    # Compute micro F1\n    score.append(f1_score(t,p, average='micro'))\n\n# Avarage among rows\nreturn np.mean(score)\n</code></pre>\n\n<p>```</p>\n\n<p>And If it is correct, then when we submit <code>nocall</code> for all rows and get 0.544 score. Then we state that we have 1 for all nocall rows and 0 for all rows with any bird. So we have 54.4 % of <code>nocall</code> rows in Public set.</p>\n\n<p>Thanks for your answer!</p>",
  "messages": [
    {
      "id": 962946,
      "postDate": "2020-08-08T15:06:11.643Z",
      "content": "<p>Also maybe we have the same discussion earlier. But I did not find it. So maybe someone can clarify is such metric computation is correct?</p>\n\n<p>```\ndef compute_raw_wise_f1micro(\n    y_true: np.ndarray, \n    y_pred: np.ndarray, \n    threshold: int = 0.5\n):\n    \"\"\"\n    Competion metric</p>\n\n<pre><code>Parameters\n----------\ny_true: np.ndarray\n    Array with indexes of true bird classes\n    Dim : [N]\ny_pred: np.ndarray\n    Array with predicted probabilities for each row\n    Dim: [N,n_classes]\nthreshold: int\n    Threshold for binirization of predicted classes\n\"\"\"\nscore = []\nn_labels = y_pred.shape[1]\n\nfor t, p in zip(y_true, y_pred):\n    # Here we have only one bird class\n    # So we wrap it into list for correct f1 computation\n    t = [t]\n    # Get all bird classes that we predict\n    # higher then threshold\n    p = np.where(p &amp;gt; threshold)[0]\n    # If we do not have any predicted classes\n    # then we put class `nocall`\n    if len(p) == 0:\n        p = [n_labels]   \n    # For all extra predicted classes\n    # we padd true values with `nocall` class\n    if len(t) &amp;lt; len(p):\n        t = t + [n_labels]*(len(p)-len(t))\n    # Compute micro F1\n    score.append(f1_score(t,p, average='micro'))\n\n# Avarage among rows\nreturn np.mean(score)\n</code></pre>\n\n<p>```</p>\n\n<p>And If it is correct, then when we submit <code>nocall</code> for all rows and get 0.544 score. Then we state that we have 1 for all nocall rows and 0 for all rows with any bird. So we have 54.4 % of <code>nocall</code> rows in Public set.</p>\n\n<p>Thanks for your answer!</p>",
      "rawMarkdown": "Also maybe we have the same discussion earlier. But I did not find it. So maybe someone can clarify is such metric computation is correct?\n\n```\ndef compute_raw_wise_f1micro(\n    y_true: np.ndarray, \n    y_pred: np.ndarray, \n    threshold: int = 0.5\n):\n    \"\"\"\n    Competion metric\n    \n    Parameters\n    ----------\n    y_true: np.ndarray\n        Array with indexes of true bird classes\n        Dim : [N]\n    y_pred: np.ndarray\n        Array with predicted probabilities for each row\n        Dim: [N,n_classes]\n    threshold: int\n        Threshold for binirization of predicted classes\n    \"\"\"\n    score = []\n    n_labels = y_pred.shape[1]\n    \n    for t, p in zip(y_true, y_pred):\n        # Here we have only one bird class\n        # So we wrap it into list for correct f1 computation\n        t = [t]\n        # Get all bird classes that we predict\n        # higher then threshold\n        p = np.where(p &gt; threshold)[0]\n        # If we do not have any predicted classes\n        # then we put class `nocall`\n        if len(p) == 0:\n            p = [n_labels]   \n        # For all extra predicted classes\n        # we padd true values with `nocall` class\n        if len(t) &lt; len(p):\n            t = t + [n_labels]*(len(p)-len(t))\n        # Compute micro F1\n        score.append(f1_score(t,p, average='micro'))\n        \n    # Avarage among rows\n    return np.mean(score)\n```\n\nAnd If it is correct, then when we submit `nocall` for all rows and get 0.544 score. Then we state that we have 1 for all nocall rows and 0 for all rows with any bird. So we have 54.4 % of `nocall` rows in Public set.\n\nThanks for your answer!",
      "votes": 3
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "962946": "Also maybe we have the same discussion earlier. But I did not find it. So maybe someone can clarify is such metric computation is correct?\n\n```\ndef compute_raw_wise_f1micro(\n    y_true: np.ndarray, \n    y_pred: np.ndarray, \n    threshold: int = 0.5\n):\n    \"\"\"\n    Competion metric\n    \n    Parameters\n    ----------\n    y_true: np.ndarray\n        Array with indexes of true bird classes\n        Dim : [N]\n    y_pred: np.ndarray\n        Array with predicted probabilities for each row\n        Dim: [N,n_classes]\n    threshold: int\n        Threshold for binirization of predicted classes\n    \"\"\"\n    score = []\n    n_labels = y_pred.shape[1]\n    \n    for t, p in zip(y_true, y_pred):\n        # Here we have only one bird class\n        # So we wrap it into list for correct f1 computation\n        t = [t]\n        # Get all bird classes that we predict\n        # higher then threshold\n        p = np.where(p &gt; threshold)[0]\n        # If we do not have any predicted classes\n        # then we put class `nocall`\n        if len(p) == 0:\n            p = [n_labels]   \n        # For all extra predicted classes\n        # we padd true values with `nocall` class\n        if len(t) &lt; len(p):\n            t = t + [n_labels]*(len(p)-len(t))\n        # Compute micro F1\n        score.append(f1_score(t,p, average='micro'))\n        \n    # Avarage among rows\n    return np.mean(score)\n```\n\nAnd If it is correct, then when we submit `nocall` for all rows and get 0.544 score. Then we state that we have 1 for all nocall rows and 0 for all rows with any bird. So we have 54.4 % of `nocall` rows in Public set.\n\nThanks for your answer!"
  }
}