{
  "id": 492070,
  "title": "Competition metric",
  "url": "/competitions/leash-BELKA/discussion/492070",
  "author_name": "",
  "post_date": "2024-04-08T14:10:45.962308700Z",
  "votes": 8,
  "comment_count": 10,
  "views": 0,
  "content": "<blockquote>\n  <p>This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.</p>\n</blockquote>\n<p>I can't get validation and test metric even close to LB score with <code>average_precision_score(average=\"micro\")</code>:</p>\n<pre><code>(\n    df_test.binds,\n    predictor.((df_test)),\n    average = ,\n    sample_weight = df_test.weights\n)\n# \n</code></pre>\n<p><code>average=\"macro\"</code>  looks better:</p>\n<pre><code>(\n    df_test.binds,\n    predictor.((df_test)),\n    average = ,\n    sample_weight = df_test.weights\n)\n# \n</code></pre>\n<p>0.31 test score VS 0.45 LB score can be explained taking into account different class balance in test/train datasets, so I'm looking for clarification about competition metric actually computed at LB. If it is something custom, demo notebook like this one <a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook</a> will be very helpful.</p>\n<p>UPD<br>\nIt seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.</p>",
  "messages": [
    {
      "id": "2741656",
      "postDate": "04/08/2024 14:10:45",
      "content": "<blockquote>\n  <p>This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.</p>\n</blockquote>\n<p>I can't get validation and test metric even close to LB score with <code>average_precision_score(average=\"micro\")</code>:</p>\n<pre><code>(\n    df_test.binds,\n    predictor.((df_test)),\n    average = ,\n    sample_weight = df_test.weights\n)\n# \n</code></pre>\n<p><code>average=\"macro\"</code>  looks better:</p>\n<pre><code>(\n    df_test.binds,\n    predictor.((df_test)),\n    average = ,\n    sample_weight = df_test.weights\n)\n# \n</code></pre>\n<p>0.31 test score VS 0.45 LB score can be explained taking into account different class balance in test/train datasets, so I'm looking for clarification about competition metric actually computed at LB. If it is something custom, demo notebook like this one <a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook</a> will be very helpful.</p>\n<p>UPD<br>\nIt seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.</p>",
      "rawMarkdown": ">This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.\n\nI can't get validation and test metric even close to LB score with `average_precision_score(average=\"micro\")`:\n\n```{python}\nsklearn.metrics.average_precision_score(\n    df_test.binds,\n    predictor.predict_proba(TabularDataset(df_test)),\n    average = \"micro\",\n    sample_weight = df_test.weights\n)\n# 0.98\n```\n\n`average=\"macro\"`  looks better:\n```{python}\nsklearn.metrics.average_precision_score(\n    df_test.binds,\n    predictor.predict_proba(TabularDataset(df_test)),\n    average = \"macro\",\n    sample_weight = df_test.weights\n)\n# 0.31\n```\n0.31 test score VS 0.45 LB score can be explained taking into account different class balance in test/train datasets, so I'm looking for clarification about competition metric actually computed at LB. If it is something custom, demo notebook like this one https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook will be very helpful.\n\nUPD\nIt seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.",
      "votes": null
    },
    {
      "id": "2741674",
      "postDate": "04/08/2024 14:29:45",
      "content": "<p>The metric is simply <code>sklearn.metrics.average_precision_score</code> with <code>average=\"micro\"</code>. </p>",
      "rawMarkdown": "The metric is simply `sklearn.metrics.average_precision_score` with `average=\"micro\"`.",
      "votes": null
    },
    {
      "id": "2741701",
      "postDate": "04/08/2024 14:58:30",
      "content": "<p>Thanks, but it is still not very clear how such scores are possible.<br>\nLet's generate some data with expected class distribution (~0.5% of molecules are binders for each protein):</p>\n<pre><code>y_true = .concatenate(\n    [.([] * 100_000), .([, , ] * )]\n)\n</code></pre>\n<p>And we will predict 0 class (not binder) all the time:</p>\n<pre><code>y_scores = np(\n       (y_true)] \n)\n</code></pre>\n<p><code>sklearn.metrics.average_precision_score</code> with <code>average=\"micro\"</code> gives us</p>\n<pre><code>average\n# \n</code></pre>\n<p>But all-zero submission are scored as low as 0.008.</p>",
      "rawMarkdown": "Thanks, but it is still not very clear how such scores are possible.\nLet's generate some data with expected class distribution (~0.5% of molecules are binders for each protein):\n\n```{python}\ny_true = np.concatenate(\n    [np.array([0] * 100_000), np.array([1, 2, 3] * 500)]\n)\n```\nAnd we will predict 0 class (not binder) all the time:\n```{python}\ny_scores = np.stack(\n   [[1, 0, 0, 0] for i in range(y_true.shape[0])] \n)\n```\n`sklearn.metrics.average_precision_score` with `average=\"micro\"` gives us\n```{python}\naverage_precision_score(y_true, y_scores, average = \"micro\")\n# 0.97\n```\nBut all-zero submission are scored as low as 0.008.",
      "votes": null
    },
    {
      "id": "2741765",
      "postDate": "04/08/2024 15:27:21",
      "content": "<p>I did it like this</p>\n<p><code>y_true = np.concatenate([np.array([0] * 99_500), np.array([1] * 500)])</code></p>\n<p>and</p>\n<p><code>y_scores = np.zeros(y_true.shape[0])</code></p>\n<p>then</p>\n<p><code>average_precision_score(y_true, y_scores, average=\"micro\")</code></p>\n<p>gives <code>0.005</code> which is close to the <code>0.008</code>observed from leaderboard. There are slightly more positive examples in public leaderboard then the train set</p>",
      "rawMarkdown": "I did it like this\n\n`y_true = np.concatenate([np.array([0] * 99_500), np.array([1] * 500)])`\n\nand\n\n`y_scores = np.zeros(y_true.shape[0])`\n\nthen\n\n`average_precision_score(y_true, y_scores, average=\"micro\")`\n\ngives `0.005` which is close to the `0.008`observed from leaderboard. There are slightly more positive examples in public leaderboard then the train set",
      "votes": null
    },
    {
      "id": "2741771",
      "postDate": "04/08/2024 15:37:05",
      "content": "<p>It is true, but why we should do micro (or macro, or None - it is all the same in this case) averaging for 2-class classification? We just get simple average_precision_score for binary target whereas competition target is not binary.</p>",
      "rawMarkdown": "It is true, but why we should do micro (or macro, or None - it is all the same in this case) averaging for 2-class classification? We just get simple average_precision_score for binary target whereas competition target is not binary.",
      "votes": null
    },
    {
      "id": "2741774",
      "postDate": "04/08/2024 15:39:24",
      "content": "<p>Maybe there is some confusion? I believe we are predicting whether the molecule binds which is a binary problem. Maybe you are thinking of the three different proteins? </p>",
      "rawMarkdown": "Maybe there is some confusion? I believe we are predicting whether the molecule binds which is a binary problem. Maybe you are thinking of the three different proteins?",
      "votes": null
    },
    {
      "id": "2741825",
      "postDate": "04/08/2024 15:55:17",
      "content": "<blockquote>\n  <p>Here, you’ll build predictive models to estimate the binding affinity of unknown chemical compounds to specified protein targets.</p>\n</blockquote>\n<p>and</p>\n<blockquote>\n  <p>This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.</p>\n</blockquote>\n<p>Micro averaging only makes sense for &gt;2 classes.<br>\nThis left no room for other interpetation except 4-class classification (molecule binds or not to one of 3 proteins or not binds at all).<br>\nIf it's still implied binary classification, it would be better to fix relevant section of the competition page.</p>",
      "rawMarkdown": ">Here, you’ll build predictive models to estimate the binding affinity of unknown chemical compounds to specified protein targets.\n\nand\n\n>This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.\n\nMicro averaging only makes sense for >2 classes.\nThis left no room for other interpetation except 4-class classification (molecule binds or not to one of 3 proteins or not binds at all).\nIf it's still implied binary classification, it would be better to fix relevant section of the competition page.",
      "votes": null
    },
    {
      "id": "2741859",
      "postDate": "04/08/2024 16:16:34",
      "content": "<p>It seems you are thinking about this much deeper then I am. I just looked at the sample submission file and saw one number to predict and matched it to the column full of 1's and 0's in the training data. </p>\n<p>I don't understand <code>map</code> as well as I should, but I did notice that the <code>macro</code> and <code>micro</code> give the same scores.  </p>\n<p>The sample submission shows one value to predict so i am pretty sure binary is correct. </p>",
      "rawMarkdown": "It seems you are thinking about this much deeper then I am. I just looked at the sample submission file and saw one number to predict and matched it to the column full of 1's and 0's in the training data. \n\nI don't understand `map` as well as I should, but I did notice that the `macro` and `micro` give the same scores.  \n\nThe sample submission shows one value to predict so i am pretty sure binary is correct.",
      "votes": null
    },
    {
      "id": "2741890",
      "postDate": "04/08/2024 16:28:40",
      "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Fc8c450a14bd51c02795b4876f1b315a9%2Ftest.PNG?generation=1712593571904972&amp;alt=media\"><br>\nTest file has 3 rows (sometimes 1 or 2 - perhaps because not all molecules were screened against all proteins) per molecule.</p>",
      "rawMarkdown": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Fc8c450a14bd51c02795b4876f1b315a9%2Ftest.PNG?generation=1712593571904972&alt=media)\nTest file has 3 rows (sometimes 1 or 2 - perhaps because not all molecules were screened against all proteins) per molecule.",
      "votes": null
    },
    {
      "id": "2741927",
      "postDate": "04/08/2024 16:48:59",
      "content": "<p>One possible explanation would be, - these metrics correlate pretty well [for the given task] and orgs decided to choose the simplest version.</p>\n<p>I think it is the case, since we have 300kk training samples and only 1.5kk test samples, - the orgs must think that this test size is enough to test generalizability of the models.</p>",
      "rawMarkdown": "One possible explanation would be, - these metrics correlate pretty well [for the given task] and orgs decided to choose the simplest version.\n\nI think it is the case, since we have 300kk training samples and only 1.5kk test samples, - the orgs must think that this test size is enough to test generalizability of the models.",
      "votes": null
    },
    {
      "id": "2742900",
      "postDate": "04/09/2024 06:15:05",
      "content": "<p>It seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.</p>",
      "rawMarkdown": "It seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2741674,
      "author_name": "inversion",
      "author_url": "",
      "post_date": "04/08/2024 14:29:45",
      "content": "<p>The metric is simply <code>sklearn.metrics.average_precision_score</code> with <code>average=\"micro\"</code>. </p>",
      "votes": null,
      "replies": [
        {
          "id": 2741701,
          "author_name": "ogurtsov",
          "author_url": "",
          "post_date": "04/08/2024 14:58:30",
          "content": "<p>Thanks, but it is still not very clear how such scores are possible.<br>\nLet's generate some data with expected class distribution (~0.5% of molecules are binders for each protein):</p>\n<pre><code>y_true = .concatenate(\n    [.([] * 100_000), .([, , ] * )]\n)\n</code></pre>\n<p>And we will predict 0 class (not binder) all the time:</p>\n<pre><code>y_scores = np(\n       (y_true)] \n)\n</code></pre>\n<p><code>sklearn.metrics.average_precision_score</code> with <code>average=\"micro\"</code> gives us</p>\n<pre><code>average\n# \n</code></pre>\n<p>But all-zero submission are scored as low as 0.008.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2741765,
          "author_name": "devinanzelmo",
          "author_url": "",
          "post_date": "04/08/2024 15:27:21",
          "content": "<p>I did it like this</p>\n<p><code>y_true = np.concatenate([np.array([0] * 99_500), np.array([1] * 500)])</code></p>\n<p>and</p>\n<p><code>y_scores = np.zeros(y_true.shape[0])</code></p>\n<p>then</p>\n<p><code>average_precision_score(y_true, y_scores, average=\"micro\")</code></p>\n<p>gives <code>0.005</code> which is close to the <code>0.008</code>observed from leaderboard. There are slightly more positive examples in public leaderboard then the train set</p>",
          "votes": null,
          "replies": [
            {
              "id": 2741771,
              "author_name": "ogurtsov",
              "author_url": "",
              "post_date": "04/08/2024 15:37:05",
              "content": "<p>It is true, but why we should do micro (or macro, or None - it is all the same in this case) averaging for 2-class classification? We just get simple average_precision_score for binary target whereas competition target is not binary.</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2741927,
                  "author_name": "martynoveduard",
                  "author_url": "",
                  "post_date": "04/08/2024 16:48:59",
                  "content": "<p>One possible explanation would be, - these metrics correlate pretty well [for the given task] and orgs decided to choose the simplest version.</p>\n<p>I think it is the case, since we have 300kk training samples and only 1.5kk test samples, - the orgs must think that this test size is enough to test generalizability of the models.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2741774,
              "author_name": "devinanzelmo",
              "author_url": "",
              "post_date": "04/08/2024 15:39:24",
              "content": "<p>Maybe there is some confusion? I believe we are predicting whether the molecule binds which is a binary problem. Maybe you are thinking of the three different proteins? </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2741825,
                  "author_name": "ogurtsov",
                  "author_url": "",
                  "post_date": "04/08/2024 15:55:17",
                  "content": "<blockquote>\n  <p>Here, you’ll build predictive models to estimate the binding affinity of unknown chemical compounds to specified protein targets.</p>\n</blockquote>\n<p>and</p>\n<blockquote>\n  <p>This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.</p>\n</blockquote>\n<p>Micro averaging only makes sense for &gt;2 classes.<br>\nThis left no room for other interpetation except 4-class classification (molecule binds or not to one of 3 proteins or not binds at all).<br>\nIf it's still implied binary classification, it would be better to fix relevant section of the competition page.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2741859,
              "author_name": "devinanzelmo",
              "author_url": "",
              "post_date": "04/08/2024 16:16:34",
              "content": "<p>It seems you are thinking about this much deeper then I am. I just looked at the sample submission file and saw one number to predict and matched it to the column full of 1's and 0's in the training data. </p>\n<p>I don't understand <code>map</code> as well as I should, but I did notice that the <code>macro</code> and <code>micro</code> give the same scores.  </p>\n<p>The sample submission shows one value to predict so i am pretty sure binary is correct. </p>",
              "votes": null,
              "replies": [
                {
                  "id": 2741890,
                  "author_name": "ogurtsov",
                  "author_url": "",
                  "post_date": "04/08/2024 16:28:40",
                  "content": "<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Fc8c450a14bd51c02795b4876f1b315a9%2Ftest.PNG?generation=1712593571904972&amp;alt=media\"><br>\nTest file has 3 rows (sometimes 1 or 2 - perhaps because not all molecules were screened against all proteins) per molecule.</p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2742900,
      "author_name": "ogurtsov",
      "author_url": "",
      "post_date": "04/09/2024 06:15:05",
      "content": "<p>It seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2741656": ">This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.\n\nI can't get validation and test metric even close to LB score with `average_precision_score(average=\"micro\")`:\n\n```{python}\nsklearn.metrics.average_precision_score(\n    df_test.binds,\n    predictor.predict_proba(TabularDataset(df_test)),\n    average = \"micro\",\n    sample_weight = df_test.weights\n)\n# 0.98\n```\n\n`average=\"macro\"`  looks better:\n```{python}\nsklearn.metrics.average_precision_score(\n    df_test.binds,\n    predictor.predict_proba(TabularDataset(df_test)),\n    average = \"macro\",\n    sample_weight = df_test.weights\n)\n# 0.31\n```\n0.31 test score VS 0.45 LB score can be explained taking into account different class balance in test/train datasets, so I'm looking for clarification about competition metric actually computed at LB. If it is something custom, demo notebook like this one https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook will be very helpful.\n\nUPD\nIt seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging.",
    "2741674": "The metric is simply `sklearn.metrics.average_precision_score` with `average=\"micro\"`.",
    "2741701": "Thanks, but it is still not very clear how such scores are possible.\nLet's generate some data with expected class distribution (~0.5% of molecules are binders for each protein):\n\n```{python}\ny_true = np.concatenate(\n    [np.array([0] * 100_000), np.array([1, 2, 3] * 500)]\n)\n```\nAnd we will predict 0 class (not binder) all the time:\n```{python}\ny_scores = np.stack(\n   [[1, 0, 0, 0] for i in range(y_true.shape[0])] \n)\n```\n`sklearn.metrics.average_precision_score` with `average=\"micro\"` gives us\n```{python}\naverage_precision_score(y_true, y_scores, average = \"micro\")\n# 0.97\n```\nBut all-zero submission are scored as low as 0.008.",
    "2741765": "I did it like this\n\n`y_true = np.concatenate([np.array([0] * 99_500), np.array([1] * 500)])`\n\nand\n\n`y_scores = np.zeros(y_true.shape[0])`\n\nthen\n\n`average_precision_score(y_true, y_scores, average=\"micro\")`\n\ngives `0.005` which is close to the `0.008`observed from leaderboard. There are slightly more positive examples in public leaderboard then the train set",
    "2741771": "It is true, but why we should do micro (or macro, or None - it is all the same in this case) averaging for 2-class classification? We just get simple average_precision_score for binary target whereas competition target is not binary.",
    "2741774": "Maybe there is some confusion? I believe we are predicting whether the molecule binds which is a binary problem. Maybe you are thinking of the three different proteins?",
    "2741825": ">Here, you’ll build predictive models to estimate the binding affinity of unknown chemical compounds to specified protein targets.\n\nand\n\n>This metric for this competition is the Mean Average Precision (micro) between the predicted probability and the observed target.\n\nMicro averaging only makes sense for >2 classes.\nThis left no room for other interpetation except 4-class classification (molecule binds or not to one of 3 proteins or not binds at all).\nIf it's still implied binary classification, it would be better to fix relevant section of the competition page.",
    "2741859": "It seems you are thinking about this much deeper then I am. I just looked at the sample submission file and saw one number to predict and matched it to the column full of 1's and 0's in the training data. \n\nI don't understand `map` as well as I should, but I did notice that the `macro` and `micro` give the same scores.  \n\nThe sample submission shows one value to predict so i am pretty sure binary is correct.",
    "2741890": "![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F682731%2Fc8c450a14bd51c02795b4876f1b315a9%2Ftest.PNG?generation=1712593571904972&alt=media)\nTest file has 3 rows (sometimes 1 or 2 - perhaps because not all molecules were screened against all proteins) per molecule.",
    "2741927": "One possible explanation would be, - these metrics correlate pretty well [for the given task] and orgs decided to choose the simplest version.\n\nI think it is the case, since we have 300kk training samples and only 1.5kk test samples, - the orgs must think that this test size is enough to test generalizability of the models.",
    "2742900": "It seems like LB score function takes into account only 3 class predictions of 4. Outputs for class 0 (probability that molecule doesn't bind to any protein) are omitted. So instead of score predictions like (0.7, 0.1, 0.1, 0.1) against ground truth like (1, 0, 0, 0) we should score (0.1, 0.1, 0.1) against (0, 0, 0). Moreover, predictions and ground truth are collapsed into one binary vector, so it is no difference between micro- and macro-averaging."
  },
  "source": "meta"
}