{
  "id": 198418,
  "title": "Correct Metric in Pytorch: LWLRAP",
  "url": "/competitions/rfcx-species-audio-detection/discussion/198418",
  "author_name": "",
  "post_date": "2020-11-21T06:45:57.195034200Z",
  "votes": 60,
  "comment_count": 19,
  "views": 0,
  "content": "<p>Hi everyone! I'm new to this comp and the metric looks very foreign and interesting. There exists a LRAP implementation in sklearn but it's actually different from the competition metric (specifically when there are differing number of ground-truth species across samples)</p>\n<p>I wrote both LRAP and Label-weighted LRAP (LWLRAP) in pytorch. The LRAP implementation is equal to sklearn LRAP on some basic tests. Hope this helps. Feel free to correct me if my implementation is wrong anywhere :) We're all here to learn!</p>\n<pre><code># LRAP. Instance-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = (scores.sum(-1) / labels.sum(-1)).mean()\n    return score.item()\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LWLRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = scores.sum() / labels.sum()\n    return score.item()\n\n# Sample usage\ny_true = torch.tensor(np.array([[1, 1, 0], [1, 0, 1], [0, 0, 1]]))\ny_score = torch.tensor(np.random.randn(3, 3))\nprint(LRAP(y_score, y_true), LWLRAP(y_score, y_true))\n</code></pre>",
  "messages": [
    {
      "id": "1085731",
      "postDate": "11/21/2020 06:45:57",
      "content": "<p>Hi everyone! I'm new to this comp and the metric looks very foreign and interesting. There exists a LRAP implementation in sklearn but it's actually different from the competition metric (specifically when there are differing number of ground-truth species across samples)</p>\n<p>I wrote both LRAP and Label-weighted LRAP (LWLRAP) in pytorch. The LRAP implementation is equal to sklearn LRAP on some basic tests. Hope this helps. Feel free to correct me if my implementation is wrong anywhere :) We're all here to learn!</p>\n<pre><code># LRAP. Instance-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = (scores.sum(-1) / labels.sum(-1)).mean()\n    return score.item()\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LWLRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = scores.sum() / labels.sum()\n    return score.item()\n\n# Sample usage\ny_true = torch.tensor(np.array([[1, 1, 0], [1, 0, 1], [0, 0, 1]]))\ny_score = torch.tensor(np.random.randn(3, 3))\nprint(LRAP(y_score, y_true), LWLRAP(y_score, y_true))\n</code></pre>",
      "rawMarkdown": "Hi everyone! I'm new to this comp and the metric looks very foreign and interesting. There exists a LRAP implementation in sklearn but it's actually different from the competition metric (specifically when there are differing number of ground-truth species across samples)\n\nI wrote both LRAP and Label-weighted LRAP (LWLRAP) in pytorch. The LRAP implementation is equal to sklearn LRAP on some basic tests. Hope this helps. Feel free to correct me if my implementation is wrong anywhere :) We're all here to learn!\n\n```\n# LRAP. Instance-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = (scores.sum(-1) / labels.sum(-1)).mean()\n    return score.item()\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LWLRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = scores.sum() / labels.sum()\n    return score.item()\n\n# Sample usage\ny_true = torch.tensor(np.array([[1, 1, 0], [1, 0, 1], [0, 0, 1]]))\ny_score = torch.tensor(np.random.randn(3, 3))\nprint(LRAP(y_score, y_true), LWLRAP(y_score, y_true))\n```",
      "votes": null
    },
    {
      "id": "1086063",
      "postDate": "11/21/2020 10:37:40",
      "content": "<p>Here is the code that was used for the FreeSound competition, that also used the lwlrap :<br>\nIt's for numpy arrays though.</p>\n<pre><code>def _one_sample_positive_class_precisions(scores, truth):\n    num_classes = scores.shape[0]\n    pos_class_indices = np.flatnonzero(truth &gt; 0)\n\n    if not len(pos_class_indices):\n        return pos_class_indices, np.zeros(0)\n\n    retrieved_classes = np.argsort(scores)[::-1]\n\n    class_rankings = np.zeros(num_classes, dtype=np.int)\n    class_rankings[retrieved_classes] = range(num_classes)\n\n    retrieved_class_true = np.zeros(num_classes, dtype=np.bool)\n    retrieved_class_true[class_rankings[pos_class_indices]] = True\n\n    retrieved_cumulative_hits = np.cumsum(retrieved_class_true)\n\n    precision_at_hits = (\n            retrieved_cumulative_hits[class_rankings[pos_class_indices]] /\n            (1 + class_rankings[pos_class_indices].astype(np.float)))\n    return pos_class_indices, precision_at_hits\n\ndef lwlrap(truth, scores):\n    assert truth.shape == scores.shape\n    num_samples, num_classes = scores.shape\n    precisions_for_samples_by_classes = np.zeros((num_samples, num_classes))\n    for sample_num in range(num_samples):\n        pos_class_indices, precision_at_hits = _one_sample_positive_class_precisions(scores[sample_num, :], truth[sample_num, :])\n        precisions_for_samples_by_classes[sample_num, pos_class_indices] = precision_at_hits\n\n    labels_per_class = np.sum(truth &gt; 0, axis=0)\n    weight_per_class = labels_per_class / float(np.sum(labels_per_class))\n\n    per_class_lwlrap = (np.sum(precisions_for_samples_by_classes, axis=0) /\n                        np.maximum(1, labels_per_class))\n    return per_class_lwlrap, weight_per_class\n\ny_true = np.array([[1, 0, 0], [0, 0, 1]])\ny_score = np.array([[0.75, 0.5, 1], [1, 0.2, 0.1]])\n\nscore_class, weight = lwlrap(y_true, y_score)\nscore = (score_class * weight).sum()\n</code></pre>\n<p>It returns the same thing as your code up to the 6th decimal so your code has to be correct.</p>",
      "rawMarkdown": "Here is the code that was used for the FreeSound competition, that also used the lwlrap :\nIt's for numpy arrays though.\n\n```\ndef _one_sample_positive_class_precisions(scores, truth):\n    num_classes = scores.shape[0]\n    pos_class_indices = np.flatnonzero(truth > 0)\n\n    if not len(pos_class_indices):\n        return pos_class_indices, np.zeros(0)\n\n    retrieved_classes = np.argsort(scores)[::-1]\n\n    class_rankings = np.zeros(num_classes, dtype=np.int)\n    class_rankings[retrieved_classes] = range(num_classes)\n\n    retrieved_class_true = np.zeros(num_classes, dtype=np.bool)\n    retrieved_class_true[class_rankings[pos_class_indices]] = True\n\n    retrieved_cumulative_hits = np.cumsum(retrieved_class_true)\n\n    precision_at_hits = (\n            retrieved_cumulative_hits[class_rankings[pos_class_indices]] /\n            (1 + class_rankings[pos_class_indices].astype(np.float)))\n    return pos_class_indices, precision_at_hits\n\ndef lwlrap(truth, scores):\n    assert truth.shape == scores.shape\n    num_samples, num_classes = scores.shape\n    precisions_for_samples_by_classes = np.zeros((num_samples, num_classes))\n    for sample_num in range(num_samples):\n        pos_class_indices, precision_at_hits = _one_sample_positive_class_precisions(scores[sample_num, :], truth[sample_num, :])\n        precisions_for_samples_by_classes[sample_num, pos_class_indices] = precision_at_hits\n        \n    labels_per_class = np.sum(truth > 0, axis=0)\n    weight_per_class = labels_per_class / float(np.sum(labels_per_class))\n\n    per_class_lwlrap = (np.sum(precisions_for_samples_by_classes, axis=0) /\n                        np.maximum(1, labels_per_class))\n    return per_class_lwlrap, weight_per_class\n\ny_true = np.array([[1, 0, 0], [0, 0, 1]])\ny_score = np.array([[0.75, 0.5, 1], [1, 0.2, 0.1]])\n\nscore_class, weight = lwlrap(y_true, y_score)\nscore = (score_class * weight).sum()\n```\n\nIt returns the same thing as your code up to the 6th decimal so your code has to be correct.",
      "votes": null
    },
    {
      "id": "1086179",
      "postDate": "11/21/2020 13:01:17",
      "content": "<p>Good to hear. Thanks, Viel!</p>",
      "rawMarkdown": "Good to hear. Thanks, Viel!",
      "votes": null
    },
    {
      "id": "1089452",
      "postDate": "11/24/2020 14:09:39",
      "content": "<p>If you're like me and still learning Pytorch, you can implement this in fastai via the custom metric class</p>\n<pre><code>import sklearn.metrics as skm\ndef _accumulate(self, learn):\n    #pred = learn.pred.argmax(dim=self.dim_argmax) if self.dim_argmax else learn.pred\n    m = nn.Sigmoid()\n    pred = learn.pred\n    pred = torch.round(m(pred))\n    targ = learn.y\n    pred,targ = to_detach(pred),to_detach(targ)\n    self.preds.append(pred)\n    self.targs.append(targ)\n\nAccumMetric.accumulate = _accumulate\n\ndef LRAP():\n    return skm_to_fastai(skm.label_ranking_average_precision_score)\n</code></pre>\n<p>src: <a href=\"url\" target=\"_blank\">https://forums.fast.ai/t/custom-metric-in-fastai2/68572/4</a></p>",
      "rawMarkdown": "If you're like me and still learning Pytorch, you can implement this in fastai via the custom metric class\n\n \n```\nimport sklearn.metrics as skm\ndef _accumulate(self, learn):\n    #pred = learn.pred.argmax(dim=self.dim_argmax) if self.dim_argmax else learn.pred\n    m = nn.Sigmoid()\n    pred = learn.pred\n    pred = torch.round(m(pred))\n    targ = learn.y\n    pred,targ = to_detach(pred),to_detach(targ)\n    self.preds.append(pred)\n    self.targs.append(targ)\n\nAccumMetric.accumulate = _accumulate\n\ndef LRAP():\n    return skm_to_fastai(skm.label_ranking_average_precision_score)\n```\nsrc: [https://forums.fast.ai/t/custom-metric-in-fastai2/68572/4](url)",
      "votes": null
    },
    {
      "id": "1089917",
      "postDate": "11/24/2020 22:47:27",
      "content": "<p><a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> I see many notebooks using AUC as metric, do you know why they use AUC instead of the competition's metric?</p>",
      "rawMarkdown": "roguekk007 I see many notebooks using AUC as metric, do you know why they use AUC instead of the competition's metric?",
      "votes": null
    },
    {
      "id": "1090217",
      "postDate": "11/25/2020 07:03:14",
      "content": "<p>I think this comp uses LWLRAP primarily because it is a natural metric for multi-label tasks. The difference between AUC and LWLRAP is actually not so large since both are rank-based. The intuition is that you have to score positive classes above negative classes. </p>",
      "rawMarkdown": "I think this comp uses LWLRAP primarily because it is a natural metric for multi-label tasks. The difference between AUC and LWLRAP is actually not so large since both are rank-based. The intuition is that you have to score positive classes above negative classes.",
      "votes": null
    },
    {
      "id": "1090859",
      "postDate": "11/25/2020 16:22:05",
      "content": "<p>If you want to use a custom function that averages by label, (like the one from the OP) you can simply do it like this:</p>\n<pre><code>Learner(..., metrics = AccumMetric(LWRAP))\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "If you want to use a custom function that averages by label, (like the one from the OP) you can simply do it like this:\n```\nLearner(..., metrics = AccumMetric(LWRAP))\n````",
      "votes": null
    },
    {
      "id": "1094760",
      "postDate": "11/28/2020 22:21:39",
      "content": "<p>Thanks for sharing</p>",
      "rawMarkdown": "Thanks for sharing",
      "votes": null
    },
    {
      "id": "1100773",
      "postDate": "12/03/2020 10:26:17",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a>, thanks for sharing this. I'm trying to understand this metric. While I've understood the theory, I was trying to understand your code in the LRAP function. When comparing it with the sklearn's inbuilt <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#label-ranking-average-precision\" target=\"_blank\">LRAP function</a>, I found a few discripancies.</p>\n<pre><code>from sklearn.metrics import label_ranking_average_precision_score \n\ny_true = torch.tensor([[1, 1, 0], [1, 0, 1],[0, 0, 1]])\ny_score = torch.tensor([[0.4156, 0.1749, 0.2129],[0.2964, 0.8753, 0.5021], [0.4369, 0.6327, 0.8359]])\nprint(LRAP(y_score, y_true))\nprint(label_ranking_average_precision_score(y_true, y_score))\n</code></pre>\n<p>Firstly, there's an <code>RuntimeError: Can only calculate the mean of floating types. Got Long instead.</code> error in your LRAP implementation most probably because of torch version mismatch (I'm using torch 1.7.0), so I fix it by replacing the score computation line by this:</p>\n<p><code>score = (scores.sum(-1) / labels.sum(-1)).float().mean()</code></p>\n<p>After this fix, your LRAP returns <code>0.33</code> while sklearn's LRAP returns <code>0.80</code>, I've manually computed the LRAP value and it matches with sklearn's output. I'm not sure what's the bug with your implementation though. (Or maybe I might have misunderstood something)</p>",
      "rawMarkdown": "Hey @roguekk007, thanks for sharing this. I'm trying to understand this metric. While I've understood the theory, I was trying to understand your code in the LRAP function. When comparing it with the sklearn's inbuilt [LRAP function](https://scikit-learn.org/stable/modules/model_evaluation.html#label-ranking-average-precision), I found a few discripancies.\n\n```\nfrom sklearn.metrics import label_ranking_average_precision_score \n\ny_true = torch.tensor([[1, 1, 0], [1, 0, 1],[0, 0, 1]])\ny_score = torch.tensor([[0.4156, 0.1749, 0.2129],[0.2964, 0.8753, 0.5021], [0.4369, 0.6327, 0.8359]])\nprint(LRAP(y_score, y_true))\nprint(label_ranking_average_precision_score(y_true, y_score))\n```\n\nFirstly, there's an `RuntimeError: Can only calculate the mean of floating types. Got Long instead.` error in your LRAP implementation most probably because of torch version mismatch (I'm using torch 1.7.0), so I fix it by replacing the score computation line by this:\n\n```score = (scores.sum(-1) / labels.sum(-1)).float().mean()```\n\nAfter this fix, your LRAP returns `0.33` while sklearn's LRAP returns `0.80`, I've manually computed the LRAP value and it matches with sklearn's output. I'm not sure what's the bug with your implementation though. (Or maybe I might have misunderstood something)",
      "votes": null
    },
    {
      "id": "1103488",
      "postDate": "12/06/2020 00:38:18",
      "content": "<p><a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> can you explain how this function can be used as regular PyTorch loss function (i.e. nn.BCELoss)? I don't see how I could call the back propagation method on it</p>",
      "rawMarkdown": "roguekk007 can you explain how this function can be used as regular PyTorch loss function (i.e. nn.BCELoss)? I don't see how I could call the back propagation method on it",
      "votes": null
    },
    {
      "id": "1107045",
      "postDate": "12/09/2020 10:43:22",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/moth\" target=\"_blank\">@moth</a> ! Sorry for the late reply. This is a metric function; is not differentiable (like categorization accuracy is not differentiable) w.r.t. to the input and should be used for validation only. You would have to use a preferably surrogate loss function to optimize this metric (BCE would be a good place to start). Hope this helps :) </p>",
      "rawMarkdown": "Hi @moth ! Sorry for the late reply. This is a metric function; is not differentiable (like categorization accuracy is not differentiable) w.r.t. to the input and should be used for validation only. You would have to use a preferably surrogate loss function to optimize this metric (BCE would be a good place to start). Hope this helps :)",
      "votes": null
    },
    {
      "id": "1107048",
      "postDate": "12/09/2020 10:47:06",
      "content": "<p>Hi Rishabh! Sorry for the late reply. My LWLRAP assumes both preds and labels to be pytorch float type (so if you plug Long type in there it throws error, you did a correct fix). Sklearn's LRAP metric is actually not totally the same with this competition's LWLRAP metric, thus the difference in computed values. If you call my LRAP function it should return same value as sklearn's LRAP metric. <br>\nLWLRAP and LRAP are different in that for example, if sample 1 has 2 positive labels, sample 2 has 1 positive labels, and sample 3 has 3 positive labels, LRAP would compute LRAP on each sample then weigh them by 1:1:1. However, LWLRAP would weigh them by 2:1:3. The difference can be seen in the minor differences between my implementation of LRAP and LWLRAP. Hope this helps :) </p>",
      "rawMarkdown": "Hi Rishabh! Sorry for the late reply. My LWLRAP assumes both preds and labels to be pytorch float type (so if you plug Long type in there it throws error, you did a correct fix). Sklearn's LRAP metric is actually not totally the same with this competition's LWLRAP metric, thus the difference in computed values. If you call my LRAP function it should return same value as sklearn's LRAP metric. \nLWLRAP and LRAP are different in that for example, if sample 1 has 2 positive labels, sample 2 has 1 positive labels, and sample 3 has 3 positive labels, LRAP would compute LRAP on each sample then weigh them by 1:1:1. However, LWLRAP would weigh them by 2:1:3. The difference can be seen in the minor differences between my implementation of LRAP and LWLRAP. Hope this helps :)",
      "votes": null
    },
    {
      "id": "1107666",
      "postDate": "12/09/2020 21:06:14",
      "content": "<p><a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> if sample 1 has 2 positive labels but they are the same class, would it get counted as 1 or 2?</p>",
      "rawMarkdown": "roguekk007 if sample 1 has 2 positive labels but they are the same class, would it get counted as 1 or 2?",
      "votes": null
    },
    {
      "id": "1107817",
      "postDate": "12/10/2020 01:30:42",
      "content": "<p>Thank you very much for your reply!</p>",
      "rawMarkdown": "Thank you very much for your reply!",
      "votes": null
    },
    {
      "id": "1107855",
      "postDate": "12/10/2020 02:33:45",
      "content": "<p>I don't understand there, how do you have two positive labels for the same class for a single sample…?</p>",
      "rawMarkdown": "I don't understand there, how do you have two positive labels for the same class for a single sample...?",
      "votes": null
    },
    {
      "id": "1107879",
      "postDate": "12/10/2020 03:13:03",
      "content": "<p>Sorry it was not clear. I have seen a clip containing two true positives of the same species_id (and they are near each other i think, so they both got cut into the same 10s segment), i was just wondering, in this case, would it only get counted as one label of this class or…? i was a bit confused because it was two events although they came from the same class</p>",
      "rawMarkdown": "Sorry it was not clear. I have seen a clip containing two true positives of the same species_id (and they are near each other i think, so they both got cut into the same 10s segment), i was just wondering, in this case, would it only get counted as one label of this class or...? i was a bit confused because it was two events although they came from the same class",
      "votes": null
    },
    {
      "id": "1158954",
      "postDate": "01/18/2021 22:27:19",
      "content": "<p>I am using the lwlrap function now in a custom LWLRAP Metric class from pytorch lightning. This class computes the average LWLRAP metric. When using it however, I sometimes get NANs in the output. </p>\n<p>Here is what I did: <br>\n`<br>\nclass LWLRAP(Metric):</p>\n<pre><code>def __init__(self, dist_sync_on_step=False):\n    super().__init__(dist_sync_on_step=dist_sync_on_step)\n    self.add_state(\"total\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n    self.add_state(\"count\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\ndef update(self, preds: torch.Tensor, target: torch.Tensor):\n    #preds, target = self._input_format(preds, target)\n    assert preds.shape == target.shape\n    value = self._lwlrap(preds, target)\n    self.total += value\n    self.count += 1\n\ndef compute(self):\n    return self.total.float()/self.count\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef _lwlrap(self, preds, labels):\n    device = preds.device\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = scores.sum() / labels.sum()\n    return score.item()\n</code></pre>\n<p>`</p>\n<p>Note that I needed to use the device because in some cases tensors where not allocated on the same device that I was training on. </p>\n<p>I based the implementation on the example at <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/metrics.html\" target=\"_blank\">the pytorch lighning website</a>.</p>\n<p>I am going to look into the NANs a bit later. Anyone any ideas already why these could be occurring?</p>",
      "rawMarkdown": "I am using the lwlrap function now in a custom LWLRAP Metric class from pytorch lightning. This class computes the average LWLRAP metric. When using it however, I sometimes get NANs in the output. \n\nHere is what I did: \n`\nclass LWLRAP(Metric):\n\n    def __init__(self, dist_sync_on_step=False):\n        super().__init__(dist_sync_on_step=dist_sync_on_step)\n        self.add_state(\"total\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n        self.add_state(\"count\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\n    def update(self, preds: torch.Tensor, target: torch.Tensor):\n        #preds, target = self._input_format(preds, target)\n        assert preds.shape == target.shape\n        value = self._lwlrap(preds, target)\n        self.total += value\n        self.count += 1\n\n    def compute(self):\n        return self.total.float()/self.count\n\n    # label-level average\n    # Assume float preds [BxC], labels [BxC] of 0 or 1\n    def _lwlrap(self, preds, labels):\n        device = preds.device\n        # Ranks of the predictions\n        ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n        # i, j corresponds to rank of prediction in row i\n        class_ranks = torch.zeros_like(ranked_classes)\n        for i in range(ranked_classes.size(0)):\n            for j in range(ranked_classes.size(1)):\n                class_ranks[i, ranked_classes[i][j]] = j + 1\n        # Mask out to only use the ranks of relevant GT labels\n        ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n        # All the GT ranks are in front now\n        sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n        # Number of GT labels per instance\n        num_labels = labels.sum(-1)\n        pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n        score_matrix = pos_matrix / sorted_ground_truth_ranks\n        score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n        scores = score_matrix * score_mask_matrix\n        score = scores.sum() / labels.sum()\n        return score.item()\n`\n\nNote that I needed to use the device because in some cases tensors where not allocated on the same device that I was training on. \n\nI based the implementation on the example at [the pytorch lighning website](https://pytorch-lightning.readthedocs.io/en/latest/metrics.html).\n\nI am going to look into the NANs a bit later. Anyone any ideas already why these could be occurring?",
      "votes": null
    },
    {
      "id": "1167102",
      "postDate": "01/24/2021 03:59:40",
      "content": "<p><a href=\"https://www.kaggle.com/erikbrakkee\" target=\"_blank\">@erikbrakkee</a> , I am not sure this is the cause, but if you input a sample data with no positive label to this function, the result can be <code>nan.</code><br>\nBecause in such a case the <code>label.sum()</code> returns <code>0</code> and the result end up a zero division at the line <code>score = scores.sum() / labels.sum()</code>.<br>\nThis situation may happen if you include fp data since some fp samples have no true positive labels.</p>",
      "rawMarkdown": "erikbrakkee , I am not sure this is the cause, but if you input a sample data with no positive label to this function, the result can be `nan.`\nBecause in such a case the `label.sum()` returns `0` and the result end up a zero division at the line `score = scores.sum() / labels.sum()`.\nThis situation may happen if you include fp data since some fp samples have no true positive labels.",
      "votes": null
    },
    {
      "id": "1205959",
      "postDate": "02/17/2021 04:21:04",
      "content": "<p>Thanks for putting this together.  I had a go myself, but still a bit stuck so happy borrow yours with time running out now.</p>\n<p>Apologies for such a basic question, but to use this just as a validation metric, I realise that I don't need to keep the autograd.  What I'm not clear on is whether I need to take the model outputs and label targets, and bring them back onto CPU.  Or is there some clever way to move this whole function onto GPU.  Something like:</p>\n<p><code>if torch.cuda.is_available():</code><br>\n<code>val_loss_function = LWLRAP('some stuff I can't figure out').cuda()</code></p>\n<p>Then later when it comes time to use it, where output is my model predictions, target is the ground truths:<br>\n<code>loss = val_loss_function (output,target)</code></p>",
      "rawMarkdown": "Thanks for putting this together.  I had a go myself, but still a bit stuck so happy borrow yours with time running out now.\n\nApologies for such a basic question, but to use this just as a validation metric, I realise that I don't need to keep the autograd.  What I'm not clear on is whether I need to take the model outputs and label targets, and bring them back onto CPU.  Or is there some clever way to move this whole function onto GPU.  Something like:\n\n`if torch.cuda.is_available():`\n`   val_loss_function = LWLRAP('some stuff I can't figure out').cuda()`\n\n\nThen later when it comes time to use it, where output is my model predictions, target is the ground truths:\n`loss = val_loss_function (output,target)`",
      "votes": null
    },
    {
      "id": "1238197",
      "postDate": "03/14/2021 18:28:06",
      "content": "<p>I have since modified the class a bit so it accumulates the results over all batches. This eliminates the NaN problem. I am now returning ths scores.sum() and labels.sum() separately and doing the division only when the result is required. So over the entire dataset, the labels.sum() is always nonzero. </p>\n<p>`<br>\nclass LWLRAP(Metric):</p>\n<pre><code>def __init__(self, dist_sync_on_step=False):\n    super().__init__(dist_sync_on_step=dist_sync_on_step)\n    self.add_state(\"scores_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n    self.add_state(\"labels_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\ndef update(self, preds: torch.Tensor, target: torch.Tensor):\n    #preds, target = self._input_format(preds, target)\n    assert preds.shape == target.shape\n    scores, labels = self._batch_compute(preds, target)\n    self.scores_sum += scores\n    self.labels_sum += labels\n\ndef compute(self):\n    res = self.scores_sum.float()/self.labels_sum.float()\n    return res\n\ndef __call__(self, preds, labels):\n    scores, labels = self._batch_compute(preds, labels)\n    self.scores_sum += scores\n    self.labels_sum += labels\n    res = scores.float()/labels.float()\n    return res\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef _batch_compute(self, preds, labels):\n    device = preds.device\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n\n    return scores.sum(), labels.sum()\n</code></pre>\n<p>`</p>",
      "rawMarkdown": "I have since modified the class a bit so it accumulates the results over all batches. This eliminates the NaN problem. I am now returning ths scores.sum() and labels.sum() separately and doing the division only when the result is required. So over the entire dataset, the labels.sum() is always nonzero. \n\n`\nclass LWLRAP(Metric):\n\n    def __init__(self, dist_sync_on_step=False):\n        super().__init__(dist_sync_on_step=dist_sync_on_step)\n        self.add_state(\"scores_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n        self.add_state(\"labels_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\n    def update(self, preds: torch.Tensor, target: torch.Tensor):\n        #preds, target = self._input_format(preds, target)\n        assert preds.shape == target.shape\n        scores, labels = self._batch_compute(preds, target)\n        self.scores_sum += scores\n        self.labels_sum += labels\n\n    def compute(self):\n        res = self.scores_sum.float()/self.labels_sum.float()\n        return res\n\n    def __call__(self, preds, labels):\n        scores, labels = self._batch_compute(preds, labels)\n        self.scores_sum += scores\n        self.labels_sum += labels\n        res = scores.float()/labels.float()\n        return res\n\n    # label-level average\n    # Assume float preds [BxC], labels [BxC] of 0 or 1\n    def _batch_compute(self, preds, labels):\n        device = preds.device\n        # Ranks of the predictions\n        ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n        # i, j corresponds to rank of prediction in row i\n        class_ranks = torch.zeros_like(ranked_classes)\n        for i in range(ranked_classes.size(0)):\n            for j in range(ranked_classes.size(1)):\n                class_ranks[i, ranked_classes[i][j]] = j + 1\n        # Mask out to only use the ranks of relevant GT labels\n        ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n        # All the GT ranks are in front now\n        sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n        # Number of GT labels per instance\n        num_labels = labels.sum(-1)\n        pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n        score_matrix = pos_matrix / sorted_ground_truth_ranks\n        score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n        scores = score_matrix * score_mask_matrix\n\n        return scores.sum(), labels.sum()\n`",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1086063,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/21/2020 10:37:40",
      "content": "<p>Here is the code that was used for the FreeSound competition, that also used the lwlrap :<br>\nIt's for numpy arrays though.</p>\n<pre><code>def _one_sample_positive_class_precisions(scores, truth):\n    num_classes = scores.shape[0]\n    pos_class_indices = np.flatnonzero(truth &gt; 0)\n\n    if not len(pos_class_indices):\n        return pos_class_indices, np.zeros(0)\n\n    retrieved_classes = np.argsort(scores)[::-1]\n\n    class_rankings = np.zeros(num_classes, dtype=np.int)\n    class_rankings[retrieved_classes] = range(num_classes)\n\n    retrieved_class_true = np.zeros(num_classes, dtype=np.bool)\n    retrieved_class_true[class_rankings[pos_class_indices]] = True\n\n    retrieved_cumulative_hits = np.cumsum(retrieved_class_true)\n\n    precision_at_hits = (\n            retrieved_cumulative_hits[class_rankings[pos_class_indices]] /\n            (1 + class_rankings[pos_class_indices].astype(np.float)))\n    return pos_class_indices, precision_at_hits\n\ndef lwlrap(truth, scores):\n    assert truth.shape == scores.shape\n    num_samples, num_classes = scores.shape\n    precisions_for_samples_by_classes = np.zeros((num_samples, num_classes))\n    for sample_num in range(num_samples):\n        pos_class_indices, precision_at_hits = _one_sample_positive_class_precisions(scores[sample_num, :], truth[sample_num, :])\n        precisions_for_samples_by_classes[sample_num, pos_class_indices] = precision_at_hits\n\n    labels_per_class = np.sum(truth &gt; 0, axis=0)\n    weight_per_class = labels_per_class / float(np.sum(labels_per_class))\n\n    per_class_lwlrap = (np.sum(precisions_for_samples_by_classes, axis=0) /\n                        np.maximum(1, labels_per_class))\n    return per_class_lwlrap, weight_per_class\n\ny_true = np.array([[1, 0, 0], [0, 0, 1]])\ny_score = np.array([[0.75, 0.5, 1], [1, 0.2, 0.1]])\n\nscore_class, weight = lwlrap(y_true, y_score)\nscore = (score_class * weight).sum()\n</code></pre>\n<p>It returns the same thing as your code up to the 6th decimal so your code has to be correct.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1086179,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "11/21/2020 13:01:17",
          "content": "<p>Good to hear. Thanks, Viel!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1094760,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/28/2020 22:21:39",
          "content": "<p>Thanks for sharing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1089452,
      "author_name": "danielvach",
      "author_url": "",
      "post_date": "11/24/2020 14:09:39",
      "content": "<p>If you're like me and still learning Pytorch, you can implement this in fastai via the custom metric class</p>\n<pre><code>import sklearn.metrics as skm\ndef _accumulate(self, learn):\n    #pred = learn.pred.argmax(dim=self.dim_argmax) if self.dim_argmax else learn.pred\n    m = nn.Sigmoid()\n    pred = learn.pred\n    pred = torch.round(m(pred))\n    targ = learn.y\n    pred,targ = to_detach(pred),to_detach(targ)\n    self.preds.append(pred)\n    self.targs.append(targ)\n\nAccumMetric.accumulate = _accumulate\n\ndef LRAP():\n    return skm_to_fastai(skm.label_ranking_average_precision_score)\n</code></pre>\n<p>src: <a href=\"url\" target=\"_blank\">https://forums.fast.ai/t/custom-metric-in-fastai2/68572/4</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 1090859,
          "author_name": "slawekbiel",
          "author_url": "",
          "post_date": "11/25/2020 16:22:05",
          "content": "<p>If you want to use a custom function that averages by label, (like the one from the OP) you can simply do it like this:</p>\n<pre><code>Learner(..., metrics = AccumMetric(LWRAP))\n</code></pre>\n<p>`</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1089917,
      "author_name": "daniellga",
      "author_url": "",
      "post_date": "11/24/2020 22:47:27",
      "content": "<p><a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> I see many notebooks using AUC as metric, do you know why they use AUC instead of the competition's metric?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1090217,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "11/25/2020 07:03:14",
          "content": "<p>I think this comp uses LWLRAP primarily because it is a natural metric for multi-label tasks. The difference between AUC and LWLRAP is actually not so large since both are rank-based. The intuition is that you have to score positive classes above negative classes. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1100773,
      "author_name": "rishabhiitbhu",
      "author_url": "",
      "post_date": "12/03/2020 10:26:17",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a>, thanks for sharing this. I'm trying to understand this metric. While I've understood the theory, I was trying to understand your code in the LRAP function. When comparing it with the sklearn's inbuilt <a href=\"https://scikit-learn.org/stable/modules/model_evaluation.html#label-ranking-average-precision\" target=\"_blank\">LRAP function</a>, I found a few discripancies.</p>\n<pre><code>from sklearn.metrics import label_ranking_average_precision_score \n\ny_true = torch.tensor([[1, 1, 0], [1, 0, 1],[0, 0, 1]])\ny_score = torch.tensor([[0.4156, 0.1749, 0.2129],[0.2964, 0.8753, 0.5021], [0.4369, 0.6327, 0.8359]])\nprint(LRAP(y_score, y_true))\nprint(label_ranking_average_precision_score(y_true, y_score))\n</code></pre>\n<p>Firstly, there's an <code>RuntimeError: Can only calculate the mean of floating types. Got Long instead.</code> error in your LRAP implementation most probably because of torch version mismatch (I'm using torch 1.7.0), so I fix it by replacing the score computation line by this:</p>\n<p><code>score = (scores.sum(-1) / labels.sum(-1)).float().mean()</code></p>\n<p>After this fix, your LRAP returns <code>0.33</code> while sklearn's LRAP returns <code>0.80</code>, I've manually computed the LRAP value and it matches with sklearn's output. I'm not sure what's the bug with your implementation though. (Or maybe I might have misunderstood something)</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107048,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "12/09/2020 10:47:06",
          "content": "<p>Hi Rishabh! Sorry for the late reply. My LWLRAP assumes both preds and labels to be pytorch float type (so if you plug Long type in there it throws error, you did a correct fix). Sklearn's LRAP metric is actually not totally the same with this competition's LWLRAP metric, thus the difference in computed values. If you call my LRAP function it should return same value as sklearn's LRAP metric. <br>\nLWLRAP and LRAP are different in that for example, if sample 1 has 2 positive labels, sample 2 has 1 positive labels, and sample 3 has 3 positive labels, LRAP would compute LRAP on each sample then weigh them by 1:1:1. However, LWLRAP would weigh them by 2:1:3. The difference can be seen in the minor differences between my implementation of LRAP and LWLRAP. Hope this helps :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107666,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "12/09/2020 21:06:14",
          "content": "<p><a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> if sample 1 has 2 positive labels but they are the same class, would it get counted as 1 or 2?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107855,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "12/10/2020 02:33:45",
          "content": "<p>I don't understand there, how do you have two positive labels for the same class for a single sample…?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107879,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "12/10/2020 03:13:03",
          "content": "<p>Sorry it was not clear. I have seen a clip containing two true positives of the same species_id (and they are near each other i think, so they both got cut into the same 10s segment), i was just wondering, in this case, would it only get counted as one label of this class or…? i was a bit confused because it was two events although they came from the same class</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1103488,
      "author_name": "alejopaullier",
      "author_url": "",
      "post_date": "12/06/2020 00:38:18",
      "content": "<p><a href=\"https://www.kaggle.com/roguekk007\" target=\"_blank\">@roguekk007</a> can you explain how this function can be used as regular PyTorch loss function (i.e. nn.BCELoss)? I don't see how I could call the back propagation method on it</p>",
      "votes": null,
      "replies": [
        {
          "id": 1107045,
          "author_name": "roguekk007",
          "author_url": "",
          "post_date": "12/09/2020 10:43:22",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/moth\" target=\"_blank\">@moth</a> ! Sorry for the late reply. This is a metric function; is not differentiable (like categorization accuracy is not differentiable) w.r.t. to the input and should be used for validation only. You would have to use a preferably surrogate loss function to optimize this metric (BCE would be a good place to start). Hope this helps :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1107817,
          "author_name": "alejopaullier",
          "author_url": "",
          "post_date": "12/10/2020 01:30:42",
          "content": "<p>Thank you very much for your reply!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1158954,
      "author_name": "erikbrakkee",
      "author_url": "",
      "post_date": "01/18/2021 22:27:19",
      "content": "<p>I am using the lwlrap function now in a custom LWLRAP Metric class from pytorch lightning. This class computes the average LWLRAP metric. When using it however, I sometimes get NANs in the output. </p>\n<p>Here is what I did: <br>\n`<br>\nclass LWLRAP(Metric):</p>\n<pre><code>def __init__(self, dist_sync_on_step=False):\n    super().__init__(dist_sync_on_step=dist_sync_on_step)\n    self.add_state(\"total\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n    self.add_state(\"count\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\ndef update(self, preds: torch.Tensor, target: torch.Tensor):\n    #preds, target = self._input_format(preds, target)\n    assert preds.shape == target.shape\n    value = self._lwlrap(preds, target)\n    self.total += value\n    self.count += 1\n\ndef compute(self):\n    return self.total.float()/self.count\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef _lwlrap(self, preds, labels):\n    device = preds.device\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = scores.sum() / labels.sum()\n    return score.item()\n</code></pre>\n<p>`</p>\n<p>Note that I needed to use the device because in some cases tensors where not allocated on the same device that I was training on. </p>\n<p>I based the implementation on the example at <a href=\"https://pytorch-lightning.readthedocs.io/en/latest/metrics.html\" target=\"_blank\">the pytorch lighning website</a>.</p>\n<p>I am going to look into the NANs a bit later. Anyone any ideas already why these could be occurring?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1167102,
          "author_name": "ishicamo",
          "author_url": "",
          "post_date": "01/24/2021 03:59:40",
          "content": "<p><a href=\"https://www.kaggle.com/erikbrakkee\" target=\"_blank\">@erikbrakkee</a> , I am not sure this is the cause, but if you input a sample data with no positive label to this function, the result can be <code>nan.</code><br>\nBecause in such a case the <code>label.sum()</code> returns <code>0</code> and the result end up a zero division at the line <code>score = scores.sum() / labels.sum()</code>.<br>\nThis situation may happen if you include fp data since some fp samples have no true positive labels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1238197,
          "author_name": "erikbrakkee",
          "author_url": "",
          "post_date": "03/14/2021 18:28:06",
          "content": "<p>I have since modified the class a bit so it accumulates the results over all batches. This eliminates the NaN problem. I am now returning ths scores.sum() and labels.sum() separately and doing the division only when the result is required. So over the entire dataset, the labels.sum() is always nonzero. </p>\n<p>`<br>\nclass LWLRAP(Metric):</p>\n<pre><code>def __init__(self, dist_sync_on_step=False):\n    super().__init__(dist_sync_on_step=dist_sync_on_step)\n    self.add_state(\"scores_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n    self.add_state(\"labels_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\ndef update(self, preds: torch.Tensor, target: torch.Tensor):\n    #preds, target = self._input_format(preds, target)\n    assert preds.shape == target.shape\n    scores, labels = self._batch_compute(preds, target)\n    self.scores_sum += scores\n    self.labels_sum += labels\n\ndef compute(self):\n    res = self.scores_sum.float()/self.labels_sum.float()\n    return res\n\ndef __call__(self, preds, labels):\n    scores, labels = self._batch_compute(preds, labels)\n    self.scores_sum += scores\n    self.labels_sum += labels\n    res = scores.float()/labels.float()\n    return res\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef _batch_compute(self, preds, labels):\n    device = preds.device\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n\n    return scores.sum(), labels.sum()\n</code></pre>\n<p>`</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1205959,
      "author_name": "ollypowell",
      "author_url": "",
      "post_date": "02/17/2021 04:21:04",
      "content": "<p>Thanks for putting this together.  I had a go myself, but still a bit stuck so happy borrow yours with time running out now.</p>\n<p>Apologies for such a basic question, but to use this just as a validation metric, I realise that I don't need to keep the autograd.  What I'm not clear on is whether I need to take the model outputs and label targets, and bring them back onto CPU.  Or is there some clever way to move this whole function onto GPU.  Something like:</p>\n<p><code>if torch.cuda.is_available():</code><br>\n<code>val_loss_function = LWLRAP('some stuff I can't figure out').cuda()</code></p>\n<p>Then later when it comes time to use it, where output is my model predictions, target is the ground truths:<br>\n<code>loss = val_loss_function (output,target)</code></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1085731": "Hi everyone! I'm new to this comp and the metric looks very foreign and interesting. There exists a LRAP implementation in sklearn but it's actually different from the competition metric (specifically when there are differing number of ground-truth species across samples)\n\nI wrote both LRAP and Label-weighted LRAP (LWLRAP) in pytorch. The LRAP implementation is equal to sklearn LRAP on some basic tests. Hope this helps. Feel free to correct me if my implementation is wrong anywhere :) We're all here to learn!\n\n```\n# LRAP. Instance-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = (scores.sum(-1) / labels.sum(-1)).mean()\n    return score.item()\n\n# label-level average\n# Assume float preds [BxC], labels [BxC] of 0 or 1\ndef LWLRAP(preds, labels):\n    # Ranks of the predictions\n    ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n    # i, j corresponds to rank of prediction in row i\n    class_ranks = torch.zeros_like(ranked_classes)\n    for i in range(ranked_classes.size(0)):\n        for j in range(ranked_classes.size(1)):\n            class_ranks[i, ranked_classes[i][j]] = j + 1\n    # Mask out to only use the ranks of relevant GT labels\n    ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n    # All the GT ranks are in front now\n    sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n    # Number of GT labels per instance\n    num_labels = labels.sum(-1)\n    pos_matrix = torch.tensor(np.array([i+1 for i in range(labels.size(-1))])).unsqueeze(0)\n    score_matrix = pos_matrix / sorted_ground_truth_ranks\n    score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n    scores = score_matrix * score_mask_matrix\n    score = scores.sum() / labels.sum()\n    return score.item()\n\n# Sample usage\ny_true = torch.tensor(np.array([[1, 1, 0], [1, 0, 1], [0, 0, 1]]))\ny_score = torch.tensor(np.random.randn(3, 3))\nprint(LRAP(y_score, y_true), LWLRAP(y_score, y_true))\n```",
    "1086063": "Here is the code that was used for the FreeSound competition, that also used the lwlrap :\nIt's for numpy arrays though.\n\n```\ndef _one_sample_positive_class_precisions(scores, truth):\n    num_classes = scores.shape[0]\n    pos_class_indices = np.flatnonzero(truth > 0)\n\n    if not len(pos_class_indices):\n        return pos_class_indices, np.zeros(0)\n\n    retrieved_classes = np.argsort(scores)[::-1]\n\n    class_rankings = np.zeros(num_classes, dtype=np.int)\n    class_rankings[retrieved_classes] = range(num_classes)\n\n    retrieved_class_true = np.zeros(num_classes, dtype=np.bool)\n    retrieved_class_true[class_rankings[pos_class_indices]] = True\n\n    retrieved_cumulative_hits = np.cumsum(retrieved_class_true)\n\n    precision_at_hits = (\n            retrieved_cumulative_hits[class_rankings[pos_class_indices]] /\n            (1 + class_rankings[pos_class_indices].astype(np.float)))\n    return pos_class_indices, precision_at_hits\n\ndef lwlrap(truth, scores):\n    assert truth.shape == scores.shape\n    num_samples, num_classes = scores.shape\n    precisions_for_samples_by_classes = np.zeros((num_samples, num_classes))\n    for sample_num in range(num_samples):\n        pos_class_indices, precision_at_hits = _one_sample_positive_class_precisions(scores[sample_num, :], truth[sample_num, :])\n        precisions_for_samples_by_classes[sample_num, pos_class_indices] = precision_at_hits\n        \n    labels_per_class = np.sum(truth > 0, axis=0)\n    weight_per_class = labels_per_class / float(np.sum(labels_per_class))\n\n    per_class_lwlrap = (np.sum(precisions_for_samples_by_classes, axis=0) /\n                        np.maximum(1, labels_per_class))\n    return per_class_lwlrap, weight_per_class\n\ny_true = np.array([[1, 0, 0], [0, 0, 1]])\ny_score = np.array([[0.75, 0.5, 1], [1, 0.2, 0.1]])\n\nscore_class, weight = lwlrap(y_true, y_score)\nscore = (score_class * weight).sum()\n```\n\nIt returns the same thing as your code up to the 6th decimal so your code has to be correct.",
    "1086179": "Good to hear. Thanks, Viel!",
    "1089452": "If you're like me and still learning Pytorch, you can implement this in fastai via the custom metric class\n\n \n```\nimport sklearn.metrics as skm\ndef _accumulate(self, learn):\n    #pred = learn.pred.argmax(dim=self.dim_argmax) if self.dim_argmax else learn.pred\n    m = nn.Sigmoid()\n    pred = learn.pred\n    pred = torch.round(m(pred))\n    targ = learn.y\n    pred,targ = to_detach(pred),to_detach(targ)\n    self.preds.append(pred)\n    self.targs.append(targ)\n\nAccumMetric.accumulate = _accumulate\n\ndef LRAP():\n    return skm_to_fastai(skm.label_ranking_average_precision_score)\n```\nsrc: [https://forums.fast.ai/t/custom-metric-in-fastai2/68572/4](url)",
    "1089917": "roguekk007 I see many notebooks using AUC as metric, do you know why they use AUC instead of the competition's metric?",
    "1090217": "I think this comp uses LWLRAP primarily because it is a natural metric for multi-label tasks. The difference between AUC and LWLRAP is actually not so large since both are rank-based. The intuition is that you have to score positive classes above negative classes.",
    "1090859": "If you want to use a custom function that averages by label, (like the one from the OP) you can simply do it like this:\n```\nLearner(..., metrics = AccumMetric(LWRAP))\n````",
    "1094760": "Thanks for sharing",
    "1100773": "Hey @roguekk007, thanks for sharing this. I'm trying to understand this metric. While I've understood the theory, I was trying to understand your code in the LRAP function. When comparing it with the sklearn's inbuilt [LRAP function](https://scikit-learn.org/stable/modules/model_evaluation.html#label-ranking-average-precision), I found a few discripancies.\n\n```\nfrom sklearn.metrics import label_ranking_average_precision_score \n\ny_true = torch.tensor([[1, 1, 0], [1, 0, 1],[0, 0, 1]])\ny_score = torch.tensor([[0.4156, 0.1749, 0.2129],[0.2964, 0.8753, 0.5021], [0.4369, 0.6327, 0.8359]])\nprint(LRAP(y_score, y_true))\nprint(label_ranking_average_precision_score(y_true, y_score))\n```\n\nFirstly, there's an `RuntimeError: Can only calculate the mean of floating types. Got Long instead.` error in your LRAP implementation most probably because of torch version mismatch (I'm using torch 1.7.0), so I fix it by replacing the score computation line by this:\n\n```score = (scores.sum(-1) / labels.sum(-1)).float().mean()```\n\nAfter this fix, your LRAP returns `0.33` while sklearn's LRAP returns `0.80`, I've manually computed the LRAP value and it matches with sklearn's output. I'm not sure what's the bug with your implementation though. (Or maybe I might have misunderstood something)",
    "1103488": "roguekk007 can you explain how this function can be used as regular PyTorch loss function (i.e. nn.BCELoss)? I don't see how I could call the back propagation method on it",
    "1107045": "Hi @moth ! Sorry for the late reply. This is a metric function; is not differentiable (like categorization accuracy is not differentiable) w.r.t. to the input and should be used for validation only. You would have to use a preferably surrogate loss function to optimize this metric (BCE would be a good place to start). Hope this helps :)",
    "1107048": "Hi Rishabh! Sorry for the late reply. My LWLRAP assumes both preds and labels to be pytorch float type (so if you plug Long type in there it throws error, you did a correct fix). Sklearn's LRAP metric is actually not totally the same with this competition's LWLRAP metric, thus the difference in computed values. If you call my LRAP function it should return same value as sklearn's LRAP metric. \nLWLRAP and LRAP are different in that for example, if sample 1 has 2 positive labels, sample 2 has 1 positive labels, and sample 3 has 3 positive labels, LRAP would compute LRAP on each sample then weigh them by 1:1:1. However, LWLRAP would weigh them by 2:1:3. The difference can be seen in the minor differences between my implementation of LRAP and LWLRAP. Hope this helps :)",
    "1107666": "roguekk007 if sample 1 has 2 positive labels but they are the same class, would it get counted as 1 or 2?",
    "1107817": "Thank you very much for your reply!",
    "1107855": "I don't understand there, how do you have two positive labels for the same class for a single sample...?",
    "1107879": "Sorry it was not clear. I have seen a clip containing two true positives of the same species_id (and they are near each other i think, so they both got cut into the same 10s segment), i was just wondering, in this case, would it only get counted as one label of this class or...? i was a bit confused because it was two events although they came from the same class",
    "1158954": "I am using the lwlrap function now in a custom LWLRAP Metric class from pytorch lightning. This class computes the average LWLRAP metric. When using it however, I sometimes get NANs in the output. \n\nHere is what I did: \n`\nclass LWLRAP(Metric):\n\n    def __init__(self, dist_sync_on_step=False):\n        super().__init__(dist_sync_on_step=dist_sync_on_step)\n        self.add_state(\"total\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n        self.add_state(\"count\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\n    def update(self, preds: torch.Tensor, target: torch.Tensor):\n        #preds, target = self._input_format(preds, target)\n        assert preds.shape == target.shape\n        value = self._lwlrap(preds, target)\n        self.total += value\n        self.count += 1\n\n    def compute(self):\n        return self.total.float()/self.count\n\n    # label-level average\n    # Assume float preds [BxC], labels [BxC] of 0 or 1\n    def _lwlrap(self, preds, labels):\n        device = preds.device\n        # Ranks of the predictions\n        ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n        # i, j corresponds to rank of prediction in row i\n        class_ranks = torch.zeros_like(ranked_classes)\n        for i in range(ranked_classes.size(0)):\n            for j in range(ranked_classes.size(1)):\n                class_ranks[i, ranked_classes[i][j]] = j + 1\n        # Mask out to only use the ranks of relevant GT labels\n        ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n        # All the GT ranks are in front now\n        sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n        # Number of GT labels per instance\n        num_labels = labels.sum(-1)\n        pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n        score_matrix = pos_matrix / sorted_ground_truth_ranks\n        score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n        scores = score_matrix * score_mask_matrix\n        score = scores.sum() / labels.sum()\n        return score.item()\n`\n\nNote that I needed to use the device because in some cases tensors where not allocated on the same device that I was training on. \n\nI based the implementation on the example at [the pytorch lighning website](https://pytorch-lightning.readthedocs.io/en/latest/metrics.html).\n\nI am going to look into the NANs a bit later. Anyone any ideas already why these could be occurring?",
    "1167102": "erikbrakkee , I am not sure this is the cause, but if you input a sample data with no positive label to this function, the result can be `nan.`\nBecause in such a case the `label.sum()` returns `0` and the result end up a zero division at the line `score = scores.sum() / labels.sum()`.\nThis situation may happen if you include fp data since some fp samples have no true positive labels.",
    "1205959": "Thanks for putting this together.  I had a go myself, but still a bit stuck so happy borrow yours with time running out now.\n\nApologies for such a basic question, but to use this just as a validation metric, I realise that I don't need to keep the autograd.  What I'm not clear on is whether I need to take the model outputs and label targets, and bring them back onto CPU.  Or is there some clever way to move this whole function onto GPU.  Something like:\n\n`if torch.cuda.is_available():`\n`   val_loss_function = LWLRAP('some stuff I can't figure out').cuda()`\n\n\nThen later when it comes time to use it, where output is my model predictions, target is the ground truths:\n`loss = val_loss_function (output,target)`",
    "1238197": "I have since modified the class a bit so it accumulates the results over all batches. This eliminates the NaN problem. I am now returning ths scores.sum() and labels.sum() separately and doing the division only when the result is required. So over the entire dataset, the labels.sum() is always nonzero. \n\n`\nclass LWLRAP(Metric):\n\n    def __init__(self, dist_sync_on_step=False):\n        super().__init__(dist_sync_on_step=dist_sync_on_step)\n        self.add_state(\"scores_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n        self.add_state(\"labels_sum\", default=torch.tensor(0.), dist_reduce_fx=\"sum\")\n\n    def update(self, preds: torch.Tensor, target: torch.Tensor):\n        #preds, target = self._input_format(preds, target)\n        assert preds.shape == target.shape\n        scores, labels = self._batch_compute(preds, target)\n        self.scores_sum += scores\n        self.labels_sum += labels\n\n    def compute(self):\n        res = self.scores_sum.float()/self.labels_sum.float()\n        return res\n\n    def __call__(self, preds, labels):\n        scores, labels = self._batch_compute(preds, labels)\n        self.scores_sum += scores\n        self.labels_sum += labels\n        res = scores.float()/labels.float()\n        return res\n\n    # label-level average\n    # Assume float preds [BxC], labels [BxC] of 0 or 1\n    def _batch_compute(self, preds, labels):\n        device = preds.device\n        # Ranks of the predictions\n        ranked_classes = torch.argsort(preds, dim=-1, descending=True)\n        # i, j corresponds to rank of prediction in row i\n        class_ranks = torch.zeros_like(ranked_classes)\n        for i in range(ranked_classes.size(0)):\n            for j in range(ranked_classes.size(1)):\n                class_ranks[i, ranked_classes[i][j]] = j + 1\n        # Mask out to only use the ranks of relevant GT labels\n        ground_truth_ranks = class_ranks * labels + (1e6) * (1 - labels)\n        # All the GT ranks are in front now\n        sorted_ground_truth_ranks, _ = torch.sort(ground_truth_ranks, dim=-1, descending=False)\n        # Number of GT labels per instance\n        num_labels = labels.sum(-1)\n        pos_matrix = torch.tensor(np.array([i + 1 for i in range(labels.size(-1))]), device = device).unsqueeze(0)\n        score_matrix = pos_matrix / sorted_ground_truth_ranks\n        score_mask_matrix, _ = torch.sort(labels, dim=-1, descending=True)\n        scores = score_matrix * score_mask_matrix\n\n        return scores.sum(), labels.sum()\n`"
  },
  "source": "meta"
}