{
  "id": 402992,
  "title": "Angular Loss Function",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/402992",
  "author_name": "JungleBeastDS",
  "post_date": "2023-04-20T15:28:41.531000",
  "votes": 6,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Thanks to Kaggle and IceCube guys for hosting this fun competition.</p>\n<p>We ran into the pitfall of trying to optimize the graph neural network structure and not try transformers.</p>\n<p>However, our best loss function was directly using the competition's evaluation metric: angular score, which beat other stuff we tried like L1, VMF, cosine similarity, etc…. I'm not sure if I have seen other people use it, but maybe it could improve results even more on transformers.</p>\n<p>To avoid exploding gradients, we clamp cos(angle) to 1-eps. So eps = 1e-3, means we count any angular loss less than .044725 (arccos(.999)) during training as 0.04425. We found 1e-5 to be the best eps.</p>\n<pre><code>class DifferentiableClamp(torch.autograd.Function):\n    \"\"\"\n    In the forward pass this operation behaves like torch.clamp.\n    But in the backward pass its gradient is 1 everywhere, as if instead of clamp one had used the identity function.\n    \"\"\"\n    @staticmethod\n    def forward(ctx, x, min_val, max_val):\n        return x.clamp(min_val, max_val)\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        return grad_output.clone(), None, None #need none, because of the optional arguments min_val and max v_val\n\nclass Angularloss(pl.LightningModule): \n    def __init__(self, eps = .001): #gradients explode, so have to add eps\n        super().__init__()\n        self.high =1-eps\n        self.low = -1+eps\n        self.Clamp = DifferentiableClamp()\n    def forward(self, y_pred, y_true):\n        scalar_prod = torch.sum(y_pred[:,:3]*y_true,dim = 1)\n        scalar_prod = self.Clamp.apply(scalar_prod, self.low, self.high)\n        return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n</code></pre>",
  "messages": [
    {
      "id": 2228491,
      "postDate": "2023-04-20T15:28:41.533Z",
      "content": "<p>Thanks to Kaggle and IceCube guys for hosting this fun competition.</p>\n<p>We ran into the pitfall of trying to optimize the graph neural network structure and not try transformers.</p>\n<p>However, our best loss function was directly using the competition's evaluation metric: angular score, which beat other stuff we tried like L1, VMF, cosine similarity, etc…. I'm not sure if I have seen other people use it, but maybe it could improve results even more on transformers.</p>\n<p>To avoid exploding gradients, we clamp cos(angle) to 1-eps. So eps = 1e-3, means we count any angular loss less than .044725 (arccos(.999)) during training as 0.04425. We found 1e-5 to be the best eps.</p>\n<pre><code>class DifferentiableClamp(torch.autograd.Function):\n    \"\"\"\n    In the forward pass this operation behaves like torch.clamp.\n    But in the backward pass its gradient is 1 everywhere, as if instead of clamp one had used the identity function.\n    \"\"\"\n    @staticmethod\n    def forward(ctx, x, min_val, max_val):\n        return x.clamp(min_val, max_val)\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        return grad_output.clone(), None, None #need none, because of the optional arguments min_val and max v_val\n\nclass Angularloss(pl.LightningModule): \n    def __init__(self, eps = .001): #gradients explode, so have to add eps\n        super().__init__()\n        self.high =1-eps\n        self.low = -1+eps\n        self.Clamp = DifferentiableClamp()\n    def forward(self, y_pred, y_true):\n        scalar_prod = torch.sum(y_pred[:,:3]*y_true,dim = 1)\n        scalar_prod = self.Clamp.apply(scalar_prod, self.low, self.high)\n        return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n</code></pre>",
      "rawMarkdown": "Thanks to Kaggle and IceCube guys for hosting this fun competition.\n\nWe ran into the pitfall of trying to optimize the graph neural network structure and not try transformers.\n\nHowever, our best loss function was directly using the competition's evaluation metric: angular score, which beat other stuff we tried like L1, VMF, cosine similarity, etc.... I'm not sure if I have seen other people use it, but maybe it could improve results even more on transformers.\n\nTo avoid exploding gradients, we clamp cos(angle) to 1-eps. So eps = 1e-3, means we count any angular loss less than .044725 (arccos(.999)) during training as 0.04425. We found 1e-5 to be the best eps.\n\n\n```\nclass DifferentiableClamp(torch.autograd.Function):\n    \"\"\"\n    In the forward pass this operation behaves like torch.clamp.\n    But in the backward pass its gradient is 1 everywhere, as if instead of clamp one had used the identity function.\n    \"\"\"\n    @staticmethod\n    def forward(ctx, x, min_val, max_val):\n        return x.clamp(min_val, max_val)\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        return grad_output.clone(), None, None #need none, because of the optional arguments min_val and max v_val\n    \nclass Angularloss(pl.LightningModule): \n    def __init__(self, eps = .001): #gradients explode, so have to add eps\n        super().__init__()\n        self.high =1-eps\n        self.low = -1+eps\n        self.Clamp = DifferentiableClamp()\n    def forward(self, y_pred, y_true):\n        scalar_prod = torch.sum(y_pred[:,:3]*y_true,dim = 1)\n        scalar_prod = self.Clamp.apply(scalar_prod, self.low, self.high)\n        return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n\n```",
      "votes": 5
    },
    {
      "id": 2229578,
      "postDate": "2023-04-21T13:38:06.160Z",
      "content": "<p>Thanks for sharing. I tried to use angular loss with tensorflow on TPU and I couldn't get it to work. I never tried it with pytorch, because I was running experiments with tf on TPU to preserve the GPU quota. The problem originated in the division by the norm, but I couldn't figure out why. Even with a huge eps (1e-1), gradients would still explode, which puzzled me. I eventually gave up and went back to VMF. It didn't occur to me that the issue might be on the clipping (in the presence of a division). I'm not sure if this would fix the issue on tf/TPU, but I'll keep it in mind for the future.</p>",
      "rawMarkdown": "Thanks for sharing. I tried to use angular loss with tensorflow on TPU and I couldn't get it to work. I never tried it with pytorch, because I was running experiments with tf on TPU to preserve the GPU quota. The problem originated in the division by the norm, but I couldn't figure out why. Even with a huge eps (1e-1), gradients would still explode, which puzzled me. I eventually gave up and went back to VMF. It didn't occur to me that the issue might be on the clipping (in the presence of a division). I'm not sure if this would fix the issue on tf/TPU, but I'll keep it in mind for the future."
    },
    {
      "id": 2229363,
      "postDate": "2023-04-21T10:00:35.027Z",
      "content": "<p>Thank you for sharing! We have also found that this loss outperforms other losses we tried, such as L1 loss, particularly in the final stages of training.</p>\n<p>What kind of issues did you encounter with the original torch.clamp function?</p>\n<p>We may have been somewhat naive in using this implementation:</p>\n<pre><code>def angular_dist_score_unit_vectors(n_true, n_pred, epsilon=1e-4):\n    scalar_prod = torch.sum(n_true * n_pred, axis=1)\n    scalar_prod = torch.clip(scalar_prod, -1+epsilon, 1-epsilon)\n    return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n</code></pre>",
      "rawMarkdown": "Thank you for sharing! We have also found that this loss outperforms other losses we tried, such as L1 loss, particularly in the final stages of training.\n\nWhat kind of issues did you encounter with the original torch.clamp function?\n\nWe may have been somewhat naive in using this implementation:\n```\ndef angular_dist_score_unit_vectors(n_true, n_pred, epsilon=1e-4):\n    scalar_prod = torch.sum(n_true * n_pred, axis=1)\n    scalar_prod = torch.clip(scalar_prod, -1+epsilon, 1-epsilon)\n    return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n```",
      "replies": [
        {
          "id": 2229537,
          "postDate": "2023-04-21T13:03:44.360Z",
          "content": "<p>That loss function is only correct if n_true and n_pred are unit length.  How do you ensure that?  That is where the non-differentiability comes in.</p>",
          "rawMarkdown": "That loss function is only correct if n_true and n_pred are unit length.  How do you ensure that?  That is where the non-differentiability comes in.",
          "replies": [
            {
              "id": 2229748,
              "postDate": "2023-04-21T16:41:40.360Z",
              "content": "<p>We used the direction task reconstruction method from the Graphnet repo, which normalized the vector.</p>",
              "rawMarkdown": "We used the direction task reconstruction method from the Graphnet repo, which normalized the vector."
            },
            {
              "id": 2229937,
              "postDate": "2023-04-21T20:38:16.180Z",
              "content": "<p>Well, this is what we used to convert the truth:</p>\n<pre><code> ():\n     torch.stack([\n        torch.cos(azimuth) * torch.sin(zenith),\n        torch.sin(azimuth) * torch.sin(zenith),\n        torch.cos(zenith)\n    ], dim=)\n</code></pre>\n<p>and here is a short snippet demonstrating the forward path and the gradient. I just checked and it works</p>\n<pre><code>dense_size = \nhidden_size = \n\ny_true = torch.randn(, )\nn_true = angles_to_unit_vector(y_true[:,], y_true[:,])\n\nhidden_concat = Variable(torch.randn(, *hidden_size), requires_grad=)\n\nfc1 = nn.Linear(*hidden_size, dense_size)\nfc2 = nn.Linear(dense_size, )\nrelu = nn.ReLU()\n\n\nx = fc1(hidden_concat)\nx = relu(x)\nx = fc2(x)\nx = F.normalize(x, p=, dim=)\n\n\n\n\n\nloss = angular_dist_score_unit_vectors(n_true, x)\nloss.backward()\n\nhidden_concat.grad\n</code></pre>",
              "rawMarkdown": "Well, this is what we used to convert the truth:\n```python\ndef angles_to_unit_vector(azimuth, zenith):\n    return torch.stack([\n        torch.cos(azimuth) * torch.sin(zenith),\n        torch.sin(azimuth) * torch.sin(zenith),\n        torch.cos(zenith)\n    ], dim=1)\n```\nand here is a short snippet demonstrating the forward path and the gradient. I just checked and it works\n\n```python\ndense_size = 512\nhidden_size = 160\n\ny_true = torch.randn(10, 2)\nn_true = angles_to_unit_vector(y_true[:,0], y_true[:,1])\n\nhidden_concat = Variable(torch.randn(10, 2*hidden_size), requires_grad=True)\n\nfc1 = nn.Linear(2*hidden_size, dense_size)\nfc2 = nn.Linear(dense_size, 3)\nrelu = nn.ReLU()\n\n\nx = fc1(hidden_concat)\nx = relu(x)\nx = fc2(x)\nx = F.normalize(x, p=2, dim=1)\n\n# this also works:\n# norm = torch.sqrt(torch.sum(x**2, dim=1, keepdims=True))\n# x = x / norm\n\nloss = angular_dist_score_unit_vectors(n_true, x)\nloss.backward()\n\nhidden_concat.grad\n```"
            }
          ]
        },
        {
          "id": 2229745,
          "postDate": "2023-04-21T16:38:33.273Z",
          "content": "<p>The problem with the original clamp functions is for values that get clipped, it makes the gradient 0.</p>\n<p>So if you get a really high or low angular score on a sample, the gradient becomes 0 and it will be a non-factor in back propagation. So, you can still train with the function without exploding gradients as long as you set an epsilon, but it will ignore any samples that had to be clamped.</p>\n<p>To address this, we set the gradient of the clamping function to 1 everywhere. In other words we cap the gradients instead of completely ignoring samples with large gradients.</p>",
          "rawMarkdown": "The problem with the original clamp functions is for values that get clipped, it makes the gradient 0.\n\nSo if you get a really high or low angular score on a sample, the gradient becomes 0 and it will be a non-factor in back propagation. So, you can still train with the function without exploding gradients as long as you set an epsilon, but it will ignore any samples that had to be clamped.\n\nTo address this, we set the gradient of the clamping function to 1 everywhere. In other words we cap the gradients instead of completely ignoring samples with large gradients.",
          "replies": [
            {
              "id": 2229941,
              "postDate": "2023-04-21T20:43:33.380Z",
              "content": "<p>Oh, I see now! Thank you very much for clarification and for bringing this topic.<br>\nNow I see that we effectively ignored events where our score was &lt; 0.0141 or &gt; 3.1274 (this corresponds to eps = 1e-4</p>\n<p>I will try to run our model with you version of the loss.</p>",
              "rawMarkdown": "Oh, I see now! Thank you very much for clarification and for bringing this topic.\nNow I see that we effectively ignored events where our score was < 0.0141 or > 3.1274 (this corresponds to eps = 1e-4\n\nI will try to run our model with you version of the loss."
            }
          ]
        }
      ]
    },
    {
      "id": 2228717,
      "postDate": "2023-04-20T18:31:37.960Z",
      "content": "<p>Wow - I had problems with angular error as a loss function and tracked it down to non-differentiability of clamp function.  You cracked it!  Thanks</p>",
      "rawMarkdown": "Wow - I had problems with angular error as a loss function and tracked it down to non-differentiability of clamp function.  You cracked it!  Thanks"
    },
    {
      "id": 2228660,
      "postDate": "2023-04-20T17:41:56.883Z",
      "content": "<p>Hi, thanks for sharing…</p>\n<p>I used that loss function too, on a very simple neural network (64 neurons, 'relu', 1 hidden layer). The network couldn't go below ~1.18 (when using more than 1 batch during training, more neurons, etc.)… but in any case, the 'mean angular difference' loss function always gave the best results, and with only 1 batch (and without using charge), it quickly converged to ~1.2.</p>",
      "rawMarkdown": "Hi, thanks for sharing...\n\nI used that loss function too, on a very simple neural network (64 neurons, 'relu', 1 hidden layer). The network couldn't go below ~1.18 (when using more than 1 batch during training, more neurons, etc.)... but in any case, the 'mean angular difference' loss function always gave the best results, and with only 1 batch (and without using charge), it quickly converged to ~1.2.",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2229578,
      "author_name": "vialactea",
      "author_url": "",
      "post_date": "2023-04-21T13:38:06.160000",
      "content": "<p>Thanks for sharing. I tried to use angular loss with tensorflow on TPU and I couldn't get it to work. I never tried it with pytorch, because I was running experiments with tf on TPU to preserve the GPU quota. The problem originated in the division by the norm, but I couldn't figure out why. Even with a huge eps (1e-1), gradients would still explode, which puzzled me. I eventually gave up and went back to VMF. It didn't occur to me that the issue might be on the clipping (in the presence of a division). I'm not sure if this would fix the issue on tf/TPU, but I'll keep it in mind for the future.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2229363,
      "author_name": "Inar Timiryasov",
      "author_url": "",
      "post_date": "2023-04-21T10:00:35.027000",
      "content": "<p>Thank you for sharing! We have also found that this loss outperforms other losses we tried, such as L1 loss, particularly in the final stages of training.</p>\n<p>What kind of issues did you encounter with the original torch.clamp function?</p>\n<p>We may have been somewhat naive in using this implementation:</p>\n<pre><code>def angular_dist_score_unit_vectors(n_true, n_pred, epsilon=1e-4):\n    scalar_prod = torch.sum(n_true * n_pred, axis=1)\n    scalar_prod = torch.clip(scalar_prod, -1+epsilon, 1-epsilon)\n    return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n</code></pre>",
      "votes": 0,
      "replies": [
        {
          "id": 2229537,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2023-04-21T13:03:44.360000",
          "content": "<p>That loss function is only correct if n_true and n_pred are unit length.  How do you ensure that?  That is where the non-differentiability comes in.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2229748,
              "author_name": "JungleBeastDS",
              "author_url": "",
              "post_date": "2023-04-21T16:41:40.360000",
              "content": "<p>We used the direction task reconstruction method from the Graphnet repo, which normalized the vector.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2229937,
              "author_name": "Inar Timiryasov",
              "author_url": "",
              "post_date": "2023-04-21T20:38:16.180000",
              "content": "<p>Well, this is what we used to convert the truth:</p>\n<pre><code> ():\n     torch.stack([\n        torch.cos(azimuth) * torch.sin(zenith),\n        torch.sin(azimuth) * torch.sin(zenith),\n        torch.cos(zenith)\n    ], dim=)\n</code></pre>\n<p>and here is a short snippet demonstrating the forward path and the gradient. I just checked and it works</p>\n<pre><code>dense_size = \nhidden_size = \n\ny_true = torch.randn(, )\nn_true = angles_to_unit_vector(y_true[:,], y_true[:,])\n\nhidden_concat = Variable(torch.randn(, *hidden_size), requires_grad=)\n\nfc1 = nn.Linear(*hidden_size, dense_size)\nfc2 = nn.Linear(dense_size, )\nrelu = nn.ReLU()\n\n\nx = fc1(hidden_concat)\nx = relu(x)\nx = fc2(x)\nx = F.normalize(x, p=, dim=)\n\n\n\n\n\nloss = angular_dist_score_unit_vectors(n_true, x)\nloss.backward()\n\nhidden_concat.grad\n</code></pre>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2229745,
          "author_name": "JungleBeastDS",
          "author_url": "",
          "post_date": "2023-04-21T16:38:33.273000",
          "content": "<p>The problem with the original clamp functions is for values that get clipped, it makes the gradient 0.</p>\n<p>So if you get a really high or low angular score on a sample, the gradient becomes 0 and it will be a non-factor in back propagation. So, you can still train with the function without exploding gradients as long as you set an epsilon, but it will ignore any samples that had to be clamped.</p>\n<p>To address this, we set the gradient of the clamping function to 1 everywhere. In other words we cap the gradients instead of completely ignoring samples with large gradients.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2229941,
              "author_name": "Inar Timiryasov",
              "author_url": "",
              "post_date": "2023-04-21T20:43:33.380000",
              "content": "<p>Oh, I see now! Thank you very much for clarification and for bringing this topic.<br>\nNow I see that we effectively ignored events where our score was &lt; 0.0141 or &gt; 3.1274 (this corresponds to eps = 1e-4</p>\n<p>I will try to run our model with you version of the loss.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2228717,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2023-04-20T18:31:37.960000",
      "content": "<p>Wow - I had problems with angular error as a loss function and tracked it down to non-differentiability of clamp function.  You cracked it!  Thanks</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2228660,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-04-20T17:41:56.883000",
      "content": "<p>Hi, thanks for sharing…</p>\n<p>I used that loss function too, on a very simple neural network (64 neurons, 'relu', 1 hidden layer). The network couldn't go below ~1.18 (when using more than 1 batch during training, more neurons, etc.)… but in any case, the 'mean angular difference' loss function always gave the best results, and with only 1 batch (and without using charge), it quickly converged to ~1.2.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2228491": "Thanks to Kaggle and IceCube guys for hosting this fun competition.\n\nWe ran into the pitfall of trying to optimize the graph neural network structure and not try transformers.\n\nHowever, our best loss function was directly using the competition's evaluation metric: angular score, which beat other stuff we tried like L1, VMF, cosine similarity, etc.... I'm not sure if I have seen other people use it, but maybe it could improve results even more on transformers.\n\nTo avoid exploding gradients, we clamp cos(angle) to 1-eps. So eps = 1e-3, means we count any angular loss less than .044725 (arccos(.999)) during training as 0.04425. We found 1e-5 to be the best eps.\n\n\n```\nclass DifferentiableClamp(torch.autograd.Function):\n    \"\"\"\n    In the forward pass this operation behaves like torch.clamp.\n    But in the backward pass its gradient is 1 everywhere, as if instead of clamp one had used the identity function.\n    \"\"\"\n    @staticmethod\n    def forward(ctx, x, min_val, max_val):\n        return x.clamp(min_val, max_val)\n\n    @staticmethod\n    def backward(ctx, grad_output):\n        return grad_output.clone(), None, None #need none, because of the optional arguments min_val and max v_val\n    \nclass Angularloss(pl.LightningModule): \n    def __init__(self, eps = .001): #gradients explode, so have to add eps\n        super().__init__()\n        self.high =1-eps\n        self.low = -1+eps\n        self.Clamp = DifferentiableClamp()\n    def forward(self, y_pred, y_true):\n        scalar_prod = torch.sum(y_pred[:,:3]*y_true,dim = 1)\n        scalar_prod = self.Clamp.apply(scalar_prod, self.low, self.high)\n        return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n\n```",
    "2229578": "Thanks for sharing. I tried to use angular loss with tensorflow on TPU and I couldn't get it to work. I never tried it with pytorch, because I was running experiments with tf on TPU to preserve the GPU quota. The problem originated in the division by the norm, but I couldn't figure out why. Even with a huge eps (1e-1), gradients would still explode, which puzzled me. I eventually gave up and went back to VMF. It didn't occur to me that the issue might be on the clipping (in the presence of a division). I'm not sure if this would fix the issue on tf/TPU, but I'll keep it in mind for the future.",
    "2229363": "Thank you for sharing! We have also found that this loss outperforms other losses we tried, such as L1 loss, particularly in the final stages of training.\n\nWhat kind of issues did you encounter with the original torch.clamp function?\n\nWe may have been somewhat naive in using this implementation:\n```\ndef angular_dist_score_unit_vectors(n_true, n_pred, epsilon=1e-4):\n    scalar_prod = torch.sum(n_true * n_pred, axis=1)\n    scalar_prod = torch.clip(scalar_prod, -1+epsilon, 1-epsilon)\n    return torch.mean(torch.abs(torch.arccos(scalar_prod)))\n```",
    "2228717": "Wow - I had problems with angular error as a loss function and tracked it down to non-differentiability of clamp function.  You cracked it!  Thanks",
    "2228660": "Hi, thanks for sharing...\n\nI used that loss function too, on a very simple neural network (64 neurons, 'relu', 1 hidden layer). The network couldn't go below ~1.18 (when using more than 1 batch during training, more neurons, etc.)... but in any case, the 'mean angular difference' loss function always gave the best results, and with only 1 batch (and without using charge), it quickly converged to ~1.2."
  }
}