{
  "id": 392096,
  "title": "∞ Explanation and Improvement von Mises-Fisher Loss",
  "url": "/competitions/icecube-neutrinos-in-deep-ice/discussion/392096",
  "author_name": "Sergey Stepanov",
  "post_date": "2023-03-03T17:30:04.976000",
  "votes": 39,
  "comment_count": 13,
  "views": 0,
  "content": "<p>von Mises-Fisher loss (vMF) is used in the GraphNet architecture.  Our experiments have shown its  usefulness for models with other architectures as well. This message explains what vMF-loss is and how to improve its Python implementation.</p>\n<p>For reasons beyond my comprehension, not everyone enjoys math, so  I will reverse the logical sequence of this post. We'll start with the final relations, their empirical explanation, and the code. And only after that I will give mathematical details justifying the \"empirical formulas\".</p>\n<h2>Introduction</h2>\n<p>When reconstructing the tracks of neutrinos (to be precise, the leptons generated by them) we solve the regression problem with two output variables (azimuth and zenith). As the angles are periodic and thus ambiguous variables, it is reasonable to shift from the angles to the three components of the unit direction vector: \\(\\mathbf{n}^2=1\\). Accordingly, the aim of our ML-model is to produce a vector \\(\\hat{\\mathbf{n}}\\) whose angle with \\(\\mathbf{n}\\) is 0 (zero error). In contrast to the true (target) vector \\(\\mathbf{n}\\), the length of the predicted vector \\(\\hat{\\mathbf{n}}\\) may be different from 1.</p>\n<h2>Practical part</h2>\n<p>The vMF-loss is equal to the scalar product of the true and predicted vectors with a minus sign and adds a \"regularization term\" that depends on the length of the predicted vector:<br>\n$$<br>\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = - \\hat{\\mathbf{n}}\\,\\mathbf{n}<br>\n+\\log (2\\pi)<br>\n+\\kappa - \\log\\kappa +  \\log \\bigr(1-e^{-2\\kappa} \\bigr),\\,\\,\\,\\,\\,\\,\\,\\,\\kappa=|\\hat{\\mathbf{n}}|.<br>\n$$<br>\nThe implementation of this vMF-loss is very simple (I omit the constant \\(2\\pi\\)):</p>\n<pre><code>def vMF_loss(n_pred, n_true, eps = 1e-8):\n    \"\"\"  von Mises-Fisher Loss: n_true is unit vector ! \"\"\"\n    kappa = torch.norm(n_pred, dim=1)        \n    logC  = -kappa + torch.log( ( kappa+eps )/( 1-torch.exp(-2*kappa)+2*eps ) )\n    return -( (n_true*n_pred).sum(dim=1) + logC ).mean() \n</code></pre>\n<p>It is worth comparing it with the original implementation from GraphNet, given, for example, <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388619\">in this post</a>. First of all, they are numerically equivalent.  However, GraphNet uses the general vMF-loss for m-dimensional space.  In three-dimensional space, the modified Bessel functions are reduced to elementary functions. In addition, there is no need to partition the range of variation \\(\\kappa\\) and use an approximate  formula to eliminate the overflow.  This loss function is also well behaved at zero \\(\\kappa=0\\), unlike the one used in GraphNet.</p>\n<h2>The empirical part</h2>\n<p>Before we dive into the math, let's understand the empirical meaning of the above expression. Let's express the scalar product as a function of the error angle between the vectors  \\(\\mathbf{n}\\,\\hat{\\mathbf{n}}=\\kappa\\,\\cos\\Psi\\) and plot the loss as a function of the angle \\(\\Psi\\) and the vector length \\(\\kappa\\).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F057658d9e8ff1e72c0f3c71aa7e8ceec%2FkMF-kaggle.png?generation=1677863643658249&amp;alt=media\" alt=\"\"></p>\n<p>As can be seen, the loss decreases as the angle tends to zero with a simultaneous increase in \\(\\kappa\\). For the gradient to reach the minimum, we don't need \\(\\kappa\\) itself to grow, although in a notebook <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\">dedicated to GraphNet</a>, large values of \\(\\kappa\\) are used to select well-trained examples. However, the main thing is that the decrease in loss occurs in the direction of decreasing the error angle; and that's exactly what we need.</p>\n<h2>Theoretical part</h2>\n<p>Let's now give a mathematical justification <a href=\"https://arxiv.org/pdf/1812.04616.pdf\">von Mises-Fisher Loss</a>. Let the probability of deviation of a 3-dimensional <i>unit</i> vector \\(\\hat{\\mathbf{n}}\\) from a fixed <i>unit</i> vector \\(\\mathbf{n}\\) be described by the von Mises-Fisher (vMF) distribution :<br>\n$$<br>\np(\\hat{\\mathbf{n}}| \\mathbf{n}) = C(\\kappa)\\, e^{\\kappa\\, \\hat{\\mathbf{n}}\\cdot\\mathbf{n}} = <br>\nC(\\kappa)\\, e^{\\kappa\\, \\cos\\Psi}.<br>\n$$<br>\nIf the vectors are the same: \\(\\hat{\\mathbf{n}}\\mathbf{n}=1\\) - the probability is maximal; if they have the opposite direction, \\(\\hat{\\mathbf{n}}\\mathbf{n}=-1\\) is minimal. The probability decay rate is characterized by the \\(\\kappa\\) (concentration parameter). The value of \\(\\kappa=0\\) defines a uniform distribution over the hypersphere and \\(\\kappa = \\infty\\) defines a point distribution at \\(\\mathbf{n}\\). If the angle \\(\\Psi\\) between the vectors \\(\\hat{\\mathbf{n}}\\) and \\(\\mathbf{n}\\) is small, then we simply get a normal distribution: \\(p(\\hat{\\mathbf{ n}}| \\mathbf{n}) \\sim e^{-\\kappa\\, \\Psi^2/2}\\).</p>\n<p>To find the normalization constant \\(C\\) of the distribution, it is necessary to calculate the integral over the entire space \\(\\hat{\\mathbf{n}}\\). Since this vector is a unit vector, in three-dimensional space it is easy to do this in a spherical coordinate system by choosing the \\(z\\) axis along the vector \\(\\mathbf{n}\\):<br>\n$$<br>\nC(\\kappa)\\, \\int^{2\\pi}_0 d\\phi \\int^\\pi_0 \\sin\\Psi \\, d\\Psi\\,\\,e^{\\kappa \\cos\\Psi} =1.<br>\n$$<br>\nIn an m-dimensional space, the result of such an integration would not be possible to express in elementary functions. In three-dimensional space, we get the simple expression:<br>\n$$<br>\nC(\\kappa) = \\frac{1}{2\\pi}\\,\\frac{\\kappa\\,e^{-\\kappa}}{1-e^{-2\\kappa} }.<br>\n$$<br>\nNote that the normalization factor is finite at zero: \\(C(0)=1/4\\pi\\).</p>\n<p>Let the model output vector \\(\\hat{\\mathbf{n}}\\), as well as the target vector \\(\\mathbf{n}\\) be normalized to one. We will maximize the probability \\(p(\\hat{\\mathbf{n}}| \\mathbf{n})\\), or, equivalently, minimize its negative logarithm (log-likelihood):<br>\n$$<br>\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = -\\log p(\\hat{\\mathbf{n}}| \\mathbf{n}) = -\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) - \\log C(\\kappa).<br>\n$$<br>\nAt this point, it is necessary to move away from mathematics and turn to magic. Above, we considered that \\(\\kappa\\) is a constant characterizing the sharpness of the probability distribution. Let us now assume that this is not a constant, but the length of the predicted vector \\(\\kappa=|\\hat{\\mathbf{n}}|\\) before its normalization. As a result, the distribution will be close to uniform (the loss is large) for small lengths \\(\\hat{\\mathbf{n}}\\). For large \\(\\kappa\\), the loss will be smaller. Now we can consider that the vector \\(\\hat{\\mathbf{n}}\\) is not actually normalized, assuming that \\(\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) \\mapsto \\hat{\\mathbf{n}}\\cdot\\mathbf{n}\\) and \\(\\hat{\\mathbf{n}}\\) has an arbitrary length equal to \\(\\kappa\\).</p>\n<p>That's all.<br>\nLet the global, deep minimum be with you :)<br>\nGood luck!</p>",
  "messages": [
    {
      "id": 2167750,
      "postDate": "2023-03-03T17:30:04.977Z",
      "content": "<p>von Mises-Fisher loss (vMF) is used in the GraphNet architecture.  Our experiments have shown its  usefulness for models with other architectures as well. This message explains what vMF-loss is and how to improve its Python implementation.</p>\n<p>For reasons beyond my comprehension, not everyone enjoys math, so  I will reverse the logical sequence of this post. We'll start with the final relations, their empirical explanation, and the code. And only after that I will give mathematical details justifying the \"empirical formulas\".</p>\n<h2>Introduction</h2>\n<p>When reconstructing the tracks of neutrinos (to be precise, the leptons generated by them) we solve the regression problem with two output variables (azimuth and zenith). As the angles are periodic and thus ambiguous variables, it is reasonable to shift from the angles to the three components of the unit direction vector: \\(\\mathbf{n}^2=1\\). Accordingly, the aim of our ML-model is to produce a vector \\(\\hat{\\mathbf{n}}\\) whose angle with \\(\\mathbf{n}\\) is 0 (zero error). In contrast to the true (target) vector \\(\\mathbf{n}\\), the length of the predicted vector \\(\\hat{\\mathbf{n}}\\) may be different from 1.</p>\n<h2>Practical part</h2>\n<p>The vMF-loss is equal to the scalar product of the true and predicted vectors with a minus sign and adds a \"regularization term\" that depends on the length of the predicted vector:<br>\n$$<br>\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = - \\hat{\\mathbf{n}}\\,\\mathbf{n}<br>\n+\\log (2\\pi)<br>\n+\\kappa - \\log\\kappa +  \\log \\bigr(1-e^{-2\\kappa} \\bigr),\\,\\,\\,\\,\\,\\,\\,\\,\\kappa=|\\hat{\\mathbf{n}}|.<br>\n$$<br>\nThe implementation of this vMF-loss is very simple (I omit the constant \\(2\\pi\\)):</p>\n<pre><code>def vMF_loss(n_pred, n_true, eps = 1e-8):\n    \"\"\"  von Mises-Fisher Loss: n_true is unit vector ! \"\"\"\n    kappa = torch.norm(n_pred, dim=1)        \n    logC  = -kappa + torch.log( ( kappa+eps )/( 1-torch.exp(-2*kappa)+2*eps ) )\n    return -( (n_true*n_pred).sum(dim=1) + logC ).mean() \n</code></pre>\n<p>It is worth comparing it with the original implementation from GraphNet, given, for example, <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388619\">in this post</a>. First of all, they are numerically equivalent.  However, GraphNet uses the general vMF-loss for m-dimensional space.  In three-dimensional space, the modified Bessel functions are reduced to elementary functions. In addition, there is no need to partition the range of variation \\(\\kappa\\) and use an approximate  formula to eliminate the overflow.  This loss function is also well behaved at zero \\(\\kappa=0\\), unlike the one used in GraphNet.</p>\n<h2>The empirical part</h2>\n<p>Before we dive into the math, let's understand the empirical meaning of the above expression. Let's express the scalar product as a function of the error angle between the vectors  \\(\\mathbf{n}\\,\\hat{\\mathbf{n}}=\\kappa\\,\\cos\\Psi\\) and plot the loss as a function of the angle \\(\\Psi\\) and the vector length \\(\\kappa\\).</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F057658d9e8ff1e72c0f3c71aa7e8ceec%2FkMF-kaggle.png?generation=1677863643658249&amp;alt=media\" alt=\"\"></p>\n<p>As can be seen, the loss decreases as the angle tends to zero with a simultaneous increase in \\(\\kappa\\). For the gradient to reach the minimum, we don't need \\(\\kappa\\) itself to grow, although in a notebook <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\">dedicated to GraphNet</a>, large values of \\(\\kappa\\) are used to select well-trained examples. However, the main thing is that the decrease in loss occurs in the direction of decreasing the error angle; and that's exactly what we need.</p>\n<h2>Theoretical part</h2>\n<p>Let's now give a mathematical justification <a href=\"https://arxiv.org/pdf/1812.04616.pdf\">von Mises-Fisher Loss</a>. Let the probability of deviation of a 3-dimensional <i>unit</i> vector \\(\\hat{\\mathbf{n}}\\) from a fixed <i>unit</i> vector \\(\\mathbf{n}\\) be described by the von Mises-Fisher (vMF) distribution :<br>\n$$<br>\np(\\hat{\\mathbf{n}}| \\mathbf{n}) = C(\\kappa)\\, e^{\\kappa\\, \\hat{\\mathbf{n}}\\cdot\\mathbf{n}} = <br>\nC(\\kappa)\\, e^{\\kappa\\, \\cos\\Psi}.<br>\n$$<br>\nIf the vectors are the same: \\(\\hat{\\mathbf{n}}\\mathbf{n}=1\\) - the probability is maximal; if they have the opposite direction, \\(\\hat{\\mathbf{n}}\\mathbf{n}=-1\\) is minimal. The probability decay rate is characterized by the \\(\\kappa\\) (concentration parameter). The value of \\(\\kappa=0\\) defines a uniform distribution over the hypersphere and \\(\\kappa = \\infty\\) defines a point distribution at \\(\\mathbf{n}\\). If the angle \\(\\Psi\\) between the vectors \\(\\hat{\\mathbf{n}}\\) and \\(\\mathbf{n}\\) is small, then we simply get a normal distribution: \\(p(\\hat{\\mathbf{ n}}| \\mathbf{n}) \\sim e^{-\\kappa\\, \\Psi^2/2}\\).</p>\n<p>To find the normalization constant \\(C\\) of the distribution, it is necessary to calculate the integral over the entire space \\(\\hat{\\mathbf{n}}\\). Since this vector is a unit vector, in three-dimensional space it is easy to do this in a spherical coordinate system by choosing the \\(z\\) axis along the vector \\(\\mathbf{n}\\):<br>\n$$<br>\nC(\\kappa)\\, \\int^{2\\pi}_0 d\\phi \\int^\\pi_0 \\sin\\Psi \\, d\\Psi\\,\\,e^{\\kappa \\cos\\Psi} =1.<br>\n$$<br>\nIn an m-dimensional space, the result of such an integration would not be possible to express in elementary functions. In three-dimensional space, we get the simple expression:<br>\n$$<br>\nC(\\kappa) = \\frac{1}{2\\pi}\\,\\frac{\\kappa\\,e^{-\\kappa}}{1-e^{-2\\kappa} }.<br>\n$$<br>\nNote that the normalization factor is finite at zero: \\(C(0)=1/4\\pi\\).</p>\n<p>Let the model output vector \\(\\hat{\\mathbf{n}}\\), as well as the target vector \\(\\mathbf{n}\\) be normalized to one. We will maximize the probability \\(p(\\hat{\\mathbf{n}}| \\mathbf{n})\\), or, equivalently, minimize its negative logarithm (log-likelihood):<br>\n$$<br>\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = -\\log p(\\hat{\\mathbf{n}}| \\mathbf{n}) = -\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) - \\log C(\\kappa).<br>\n$$<br>\nAt this point, it is necessary to move away from mathematics and turn to magic. Above, we considered that \\(\\kappa\\) is a constant characterizing the sharpness of the probability distribution. Let us now assume that this is not a constant, but the length of the predicted vector \\(\\kappa=|\\hat{\\mathbf{n}}|\\) before its normalization. As a result, the distribution will be close to uniform (the loss is large) for small lengths \\(\\hat{\\mathbf{n}}\\). For large \\(\\kappa\\), the loss will be smaller. Now we can consider that the vector \\(\\hat{\\mathbf{n}}\\) is not actually normalized, assuming that \\(\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) \\mapsto \\hat{\\mathbf{n}}\\cdot\\mathbf{n}\\) and \\(\\hat{\\mathbf{n}}\\) has an arbitrary length equal to \\(\\kappa\\).</p>\n<p>That's all.<br>\nLet the global, deep minimum be with you :)<br>\nGood luck!</p>",
      "rawMarkdown": "von Mises-Fisher loss (vMF) is used in the GraphNet architecture.  Our experiments have shown its  usefulness for models with other architectures as well. This message explains what vMF-loss is and how to improve its Python implementation.\n\nFor reasons beyond my comprehension, not everyone enjoys math, so  I will reverse the logical sequence of this post. We'll start with the final relations, their empirical explanation, and the code. And only after that I will give mathematical details justifying the \"empirical formulas\".\n\n## Introduction\n\nWhen reconstructing the tracks of neutrinos (to be precise, the leptons generated by them) we solve the regression problem with two output variables (azimuth and zenith). As the angles are periodic and thus ambiguous variables, it is reasonable to shift from the angles to the three components of the unit direction vector: \\\\(\\mathbf{n}^2=1\\\\). Accordingly, the aim of our ML-model is to produce a vector \\\\(\\hat{\\mathbf{n}}\\\\) whose angle with \\\\(\\mathbf{n}\\\\) is 0 (zero error). In contrast to the true (target) vector \\\\(\\mathbf{n}\\\\), the length of the predicted vector \\\\(\\hat{\\mathbf{n}}\\\\) may be different from 1.\n\n## Practical part\n\nThe vMF-loss is equal to the scalar product of the true and predicted vectors with a minus sign and adds a \"regularization term\" that depends on the length of the predicted vector:\n$$\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = - \\hat{\\mathbf{n}}\\,\\mathbf{n}\n+\\log (2\\pi)\n+\\kappa - \\log\\kappa +  \\log \\bigr(1-e^{-2\\kappa} \\bigr),\\,\\,\\,\\,\\,\\,\\,\\,\\kappa=\\|\\hat{\\mathbf{n}}\\|.\n$$\nThe implementation of this vMF-loss is very simple (I omit the constant \\\\(2\\pi\\\\)):\n```\ndef vMF_loss(n_pred, n_true, eps = 1e-8):\n    \"\"\"  von Mises-Fisher Loss: n_true is unit vector ! \"\"\"\n    kappa = torch.norm(n_pred, dim=1)        \n    logC  = -kappa + torch.log( ( kappa+eps )/( 1-torch.exp(-2*kappa)+2*eps ) )\n    return -( (n_true*n_pred).sum(dim=1) + logC ).mean() \n```\nIt is worth comparing it with the original implementation from GraphNet, given, for example, <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388619\">in this post</a>. First of all, they are numerically equivalent.  However, GraphNet uses the general vMF-loss for m-dimensional space.  In three-dimensional space, the modified Bessel functions are reduced to elementary functions. In addition, there is no need to partition the range of variation \\\\(\\kappa\\\\) and use an approximate  formula to eliminate the overflow.  This loss function is also well behaved at zero \\\\(\\kappa=0\\\\), unlike the one used in GraphNet.\n\n## The empirical part\n\nBefore we dive into the math, let's understand the empirical meaning of the above expression. Let's express the scalar product as a function of the error angle between the vectors  \\\\(\\mathbf{n}\\,\\hat{\\mathbf{n}}=\\kappa\\,\\cos\\Psi\\\\) and plot the loss as a function of the angle \\\\(\\Psi\\\\) and the vector length \\\\(\\kappa\\\\).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F057658d9e8ff1e72c0f3c71aa7e8ceec%2FkMF-kaggle.png?generation=1677863643658249&alt=media)\n\nAs can be seen, the loss decreases as the angle tends to zero with a simultaneous increase in \\\\(\\kappa\\\\). For the gradient to reach the minimum, we don't need \\\\(\\kappa\\\\) itself to grow, although in a notebook <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\">dedicated to GraphNet</a>, large values of \\\\(\\kappa\\\\) are used to select well-trained examples. However, the main thing is that the decrease in loss occurs in the direction of decreasing the error angle; and that's exactly what we need.\n\n## Theoretical part\n\nLet's now give a mathematical justification <a href=\"https://arxiv.org/pdf/1812.04616.pdf\">von Mises-Fisher Loss</a>. Let the probability of deviation of a 3-dimensional <i>unit</i> vector \\\\(\\hat{\\mathbf{n}}\\\\) from a fixed <i>unit</i> vector \\\\(\\mathbf{n}\\\\) be described by the von Mises-Fisher (vMF) distribution :\n$$\np(\\hat{\\mathbf{n}}| \\mathbf{n}) = C(\\kappa)\\, e^{\\kappa\\, \\hat{\\mathbf{n}}\\cdot\\mathbf{n}} = \nC(\\kappa)\\, e^{\\kappa\\, \\cos\\Psi}.\n$$\nIf the vectors are the same: \\\\(\\hat{\\mathbf{n}}\\mathbf{n}=1\\\\) - the probability is maximal; if they have the opposite direction, \\\\(\\hat{\\mathbf{n}}\\mathbf{n}=-1\\\\) is minimal. The probability decay rate is characterized by the \\\\(\\kappa\\\\) (concentration parameter). The value of \\\\(\\kappa=0\\\\) defines a uniform distribution over the hypersphere and \\\\(\\kappa = \\infty\\\\) defines a point distribution at \\\\(\\mathbf{n}\\\\). If the angle \\\\(\\Psi\\\\) between the vectors \\\\(\\hat{\\mathbf{n}}\\\\) and \\\\(\\mathbf{n}\\\\) is small, then we simply get a normal distribution: \\\\(p(\\hat{\\mathbf{ n}}| \\mathbf{n}) \\sim e^{-\\kappa\\, \\Psi^2/2}\\\\).\n\nTo find the normalization constant \\\\(C\\\\) of the distribution, it is necessary to calculate the integral over the entire space \\\\(\\hat{\\mathbf{n}}\\\\). Since this vector is a unit vector, in three-dimensional space it is easy to do this in a spherical coordinate system by choosing the \\\\(z\\\\) axis along the vector \\\\(\\mathbf{n}\\\\):\n$$\nC(\\kappa)\\, \\int^{2\\pi}_0 d\\phi \\int^\\pi_0 \\sin\\Psi \\, d\\Psi\\,\\,e^{\\kappa \\cos\\Psi} =1.\n$$\nIn an m-dimensional space, the result of such an integration would not be possible to express in elementary functions. In three-dimensional space, we get the simple expression:\n$$\nC(\\kappa) = \\frac{1}{2\\pi}\\,\\frac{\\kappa\\,e^{-\\kappa}}{1-e^{-2\\kappa} }.\n$$\nNote that the normalization factor is finite at zero: \\\\(C(0)=1/4\\pi\\\\).\n\nLet the model output vector \\\\(\\hat{\\mathbf{n}}\\\\), as well as the target vector \\\\(\\mathbf{n}\\\\) be normalized to one. We will maximize the probability \\\\(p(\\hat{\\mathbf{n}}| \\mathbf{n})\\\\), or, equivalently, minimize its negative logarithm (log-likelihood):\n$$\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = -\\log p(\\hat{\\mathbf{n}}| \\mathbf{n}) = -\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) - \\log C(\\kappa).\n$$\nAt this point, it is necessary to move away from mathematics and turn to magic. Above, we considered that \\\\(\\kappa\\\\) is a constant characterizing the sharpness of the probability distribution. Let us now assume that this is not a constant, but the length of the predicted vector \\\\(\\kappa=\\|\\hat{\\mathbf{n}}\\|\\\\) before its normalization. As a result, the distribution will be close to uniform (the loss is large) for small lengths \\\\(\\hat{\\mathbf{n}}\\\\). For large \\\\(\\kappa\\\\), the loss will be smaller. Now we can consider that the vector \\\\(\\hat{\\mathbf{n}}\\\\) is not actually normalized, assuming that \\\\(\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) \\mapsto \\hat{\\mathbf{n}}\\cdot\\mathbf{n}\\\\) and \\\\(\\hat{\\mathbf{n}}\\\\) has an arbitrary length equal to \\\\(\\kappa\\\\).\n\nThat's all.\nLet the global, deep minimum be with you :)\nGood luck!\n\n",
      "votes": 39
    },
    {
      "id": 2171142,
      "postDate": "2023-03-06T14:41:59.343Z",
      "content": "<p>Very nice explanation.  The graphnet vonMises-Fisher loss function you mention predicts 4 numbers:  [x,y,z,kappa], while your formula computes kappa from the norm of (x,y,z).  Can you explain the difference?</p>",
      "rawMarkdown": "Very nice explanation.  The graphnet vonMises-Fisher loss function you mention predicts 4 numbers:  [x,y,z,kappa], while your formula computes kappa from the norm of (x,y,z).  Can you explain the difference?",
      "replies": [
        {
          "id": 2171178,
          "postDate": "2023-03-06T15:06:41.073Z",
          "content": "<p>Nevermind 😬 I see that DirectionReconstructionWithKappa computes kappa as |v| then passes normalized v and kappa to the loss.</p>",
          "rawMarkdown": "Nevermind 😬 I see that DirectionReconstructionWithKappa computes kappa as |v| then passes normalized v and kappa to the loss.",
          "votes": 2
        }
      ]
    },
    {
      "id": 2170907,
      "postDate": "2023-03-06T11:33:03.977Z",
      "content": "<p><a href=\"https://www.kaggle.com/synset\" target=\"_blank\">@synset</a> does anyone have the explicit form of the 2D vMF loss too plz ? i am interested by any reference too. Thx</p>",
      "rawMarkdown": "@synset does anyone have the explicit form of the 2D vMF loss too plz ? i am interested by any reference too. Thx",
      "replies": [
        {
          "id": 2171052,
          "postDate": "2023-03-06T13:51:10.850Z",
          "content": "<p>Unfortunately, in a two-dimensional space, the normalization constant \\(C(\\kappa)\\) cannot be expressed in terms of elementary functions. It is equal to \\(1/2\\pi I_0(\\kappa)\\), where \\(I_0\\) is the modified Bessel function of order zero. Why do you need two-dimensional space? Our space is apparently three-dimensional. 😏</p>",
          "rawMarkdown": "Unfortunately, in a two-dimensional space, the normalization constant \\\\(C(\\kappa)\\\\) cannot be expressed in terms of elementary functions. It is equal to \\\\(1/2\\pi I_0(\\kappa)\\\\), where \\\\(I_0\\\\) is the modified Bessel function of order zero. Why do you need two-dimensional space? Our space is apparently three-dimensional. 😏",
          "replies": [
            {
              "id": 2171094,
              "postDate": "2023-03-06T14:25:57.390Z",
              "content": "<p>In the Graphnet paper, they go for 2D vmf loss for azimuth and MAE loss for zenith. I want to give it a try, with my personal implementation. So, I don't understand how to compute it, in the 2D case, except using their heavy and complicated vMF loss. </p>",
              "rawMarkdown": "In the Graphnet paper, they go for 2D vmf loss for azimuth and MAE loss for zenith. I want to give it a try, with my personal implementation. So, I don't understand how to compute it, in the 2D case, except using their heavy and complicated vMF loss. "
            },
            {
              "id": 2171099,
              "postDate": "2023-03-06T14:28:30.907Z",
              "content": "<p><a href=\"https://www.kaggle.com/synset\" target=\"_blank\">@synset</a> And I am also considering this new approach, bc my model's perf didn't improve using your vMF loss, do you remember if it drastically improved yours ? (I am stuck around 1.07)</p>",
              "rawMarkdown": "@synset And I am also considering this new approach, bc my model's perf didn't improve using your vMF loss, do you remember if it drastically improved yours ? (I am stuck around 1.07)"
            },
            {
              "id": 2171308,
              "postDate": "2023-03-06T17:12:27.720Z",
              "content": "<p>The loss function is not the Holy Grail. When training a model, you can switch between different loss functions. For example from vMF to cosine loss and vice versa. This sometimes helps to get the model out of local minima.</p>\n<p>I am convinced that the model should not predict angles. If the output model for the azimuth predicts 0.01, and the target is 6.28, is it bad or good?</p>",
              "rawMarkdown": "The loss function is not the Holy Grail. When training a model, you can switch between different loss functions. For example from vMF to cosine loss and vice versa. This sometimes helps to get the model out of local minima.\n\nI am convinced that the model should not predict angles. If the output model for the azimuth predicts 0.01, and the target is 6.28, is it bad or good?",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2170104,
      "postDate": "2023-03-05T17:26:45.013Z",
      "content": "<p>I tried your loss, my model converges faster but performs as good as before (using a MAE loss for both angles).<br>\nIs that normal, or should I seek for a mistake in my code ?</p>",
      "rawMarkdown": "I tried your loss, my model converges faster but performs as good as before (using a MAE loss for both angles).\nIs that normal, or should I seek for a mistake in my code ?",
      "replies": [
        {
          "id": 2171396,
          "postDate": "2023-03-06T18:39:59.430Z",
          "content": "<p>From my experiments, vMF loss converges slower, but perhaps a little more accurately.<br>\nP.S. I used MSE loss for (x, y, z) for comparision.</p>",
          "rawMarkdown": "From my experiments, vMF loss converges slower, but perhaps a little more accurately.\nP.S. I used MSE loss for (x, y, z) for comparision.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2168863,
      "postDate": "2023-03-04T15:39:03.253Z",
      "content": "<p>Great! Not it's clear.</p>",
      "rawMarkdown": "Great! Not it's clear."
    },
    {
      "id": 2167758,
      "postDate": "2023-03-03T17:38:54.470Z",
      "content": "<p>pure gold thread, ty !</p>",
      "rawMarkdown": "pure gold thread, ty !"
    },
    {
      "id": 2169034,
      "postDate": "2023-03-04T18:15:11.507Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 2170051,
          "postDate": "2023-03-05T16:42:42.287Z",
          "rawMarkdown": "",
          "isDeleted": true
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2171142,
      "author_name": "SolverWorld",
      "author_url": "",
      "post_date": "2023-03-06T14:41:59.343000",
      "content": "<p>Very nice explanation.  The graphnet vonMises-Fisher loss function you mention predicts 4 numbers:  [x,y,z,kappa], while your formula computes kappa from the norm of (x,y,z).  Can you explain the difference?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2171178,
          "author_name": "SolverWorld",
          "author_url": "",
          "post_date": "2023-03-06T15:06:41.073000",
          "content": "<p>Nevermind 😬 I see that DirectionReconstructionWithKappa computes kappa as |v| then passes normalized v and kappa to the loss.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 2170907,
      "author_name": "Luigi Stf",
      "author_url": "",
      "post_date": "2023-03-06T11:33:03.977000",
      "content": "<p><a href=\"https://www.kaggle.com/synset\" target=\"_blank\">@synset</a> does anyone have the explicit form of the 2D vMF loss too plz ? i am interested by any reference too. Thx</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2171052,
          "author_name": "Sergey Stepanov",
          "author_url": "",
          "post_date": "2023-03-06T13:51:10.850000",
          "content": "<p>Unfortunately, in a two-dimensional space, the normalization constant \\(C(\\kappa)\\) cannot be expressed in terms of elementary functions. It is equal to \\(1/2\\pi I_0(\\kappa)\\), where \\(I_0\\) is the modified Bessel function of order zero. Why do you need two-dimensional space? Our space is apparently three-dimensional. 😏</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2171094,
              "author_name": "Luigi Stf",
              "author_url": "",
              "post_date": "2023-03-06T14:25:57.390000",
              "content": "<p>In the Graphnet paper, they go for 2D vmf loss for azimuth and MAE loss for zenith. I want to give it a try, with my personal implementation. So, I don't understand how to compute it, in the 2D case, except using their heavy and complicated vMF loss. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2171099,
              "author_name": "Luigi Stf",
              "author_url": "",
              "post_date": "2023-03-06T14:28:30.907000",
              "content": "<p><a href=\"https://www.kaggle.com/synset\" target=\"_blank\">@synset</a> And I am also considering this new approach, bc my model's perf didn't improve using your vMF loss, do you remember if it drastically improved yours ? (I am stuck around 1.07)</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2171308,
              "author_name": "Sergey Stepanov",
              "author_url": "",
              "post_date": "2023-03-06T17:12:27.720000",
              "content": "<p>The loss function is not the Holy Grail. When training a model, you can switch between different loss functions. For example from vMF to cosine loss and vice versa. This sometimes helps to get the model out of local minima.</p>\n<p>I am convinced that the model should not predict angles. If the output model for the azimuth predicts 0.01, and the target is 6.28, is it bad or good?</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2170104,
      "author_name": "Luigi Stf",
      "author_url": "",
      "post_date": "2023-03-05T17:26:45.013000",
      "content": "<p>I tried your loss, my model converges faster but performs as good as before (using a MAE loss for both angles).<br>\nIs that normal, or should I seek for a mistake in my code ?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2171396,
          "author_name": "Sergey Zlobin",
          "author_url": "",
          "post_date": "2023-03-06T18:39:59.430000",
          "content": "<p>From my experiments, vMF loss converges slower, but perhaps a little more accurately.<br>\nP.S. I used MSE loss for (x, y, z) for comparision.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2168863,
      "author_name": "Sergey Zlobin",
      "author_url": "",
      "post_date": "2023-03-04T15:39:03.253000",
      "content": "<p>Great! Not it's clear.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2167758,
      "author_name": "Luigi Stf",
      "author_url": "",
      "post_date": "2023-03-03T17:38:54.470000",
      "content": "<p>pure gold thread, ty !</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2169034,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-03-04T18:15:11.507000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 2170051,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-05T16:42:42.287000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2167750": "von Mises-Fisher loss (vMF) is used in the GraphNet architecture.  Our experiments have shown its  usefulness for models with other architectures as well. This message explains what vMF-loss is and how to improve its Python implementation.\n\nFor reasons beyond my comprehension, not everyone enjoys math, so  I will reverse the logical sequence of this post. We'll start with the final relations, their empirical explanation, and the code. And only after that I will give mathematical details justifying the \"empirical formulas\".\n\n## Introduction\n\nWhen reconstructing the tracks of neutrinos (to be precise, the leptons generated by them) we solve the regression problem with two output variables (azimuth and zenith). As the angles are periodic and thus ambiguous variables, it is reasonable to shift from the angles to the three components of the unit direction vector: \\\\(\\mathbf{n}^2=1\\\\). Accordingly, the aim of our ML-model is to produce a vector \\\\(\\hat{\\mathbf{n}}\\\\) whose angle with \\\\(\\mathbf{n}\\\\) is 0 (zero error). In contrast to the true (target) vector \\\\(\\mathbf{n}\\\\), the length of the predicted vector \\\\(\\hat{\\mathbf{n}}\\\\) may be different from 1.\n\n## Practical part\n\nThe vMF-loss is equal to the scalar product of the true and predicted vectors with a minus sign and adds a \"regularization term\" that depends on the length of the predicted vector:\n$$\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = - \\hat{\\mathbf{n}}\\,\\mathbf{n}\n+\\log (2\\pi)\n+\\kappa - \\log\\kappa +  \\log \\bigr(1-e^{-2\\kappa} \\bigr),\\,\\,\\,\\,\\,\\,\\,\\,\\kappa=\\|\\hat{\\mathbf{n}}\\|.\n$$\nThe implementation of this vMF-loss is very simple (I omit the constant \\\\(2\\pi\\\\)):\n```\ndef vMF_loss(n_pred, n_true, eps = 1e-8):\n    \"\"\"  von Mises-Fisher Loss: n_true is unit vector ! \"\"\"\n    kappa = torch.norm(n_pred, dim=1)        \n    logC  = -kappa + torch.log( ( kappa+eps )/( 1-torch.exp(-2*kappa)+2*eps ) )\n    return -( (n_true*n_pred).sum(dim=1) + logC ).mean() \n```\nIt is worth comparing it with the original implementation from GraphNet, given, for example, <a href=\"https://www.kaggle.com/competitions/icecube-neutrinos-in-deep-ice/discussion/388619\">in this post</a>. First of all, they are numerically equivalent.  However, GraphNet uses the general vMF-loss for m-dimensional space.  In three-dimensional space, the modified Bessel functions are reduced to elementary functions. In addition, there is no need to partition the range of variation \\\\(\\kappa\\\\) and use an approximate  formula to eliminate the overflow.  This loss function is also well behaved at zero \\\\(\\kappa=0\\\\), unlike the one used in GraphNet.\n\n## The empirical part\n\nBefore we dive into the math, let's understand the empirical meaning of the above expression. Let's express the scalar product as a function of the error angle between the vectors  \\\\(\\mathbf{n}\\,\\hat{\\mathbf{n}}=\\kappa\\,\\cos\\Psi\\\\) and plot the loss as a function of the angle \\\\(\\Psi\\\\) and the vector length \\\\(\\kappa\\\\).\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F1394009%2F057658d9e8ff1e72c0f3c71aa7e8ceec%2FkMF-kaggle.png?generation=1677863643658249&alt=media)\n\nAs can be seen, the loss decreases as the angle tends to zero with a simultaneous increase in \\\\(\\kappa\\\\). For the gradient to reach the minimum, we don't need \\\\(\\kappa\\\\) itself to grow, although in a notebook <a href=\"https://www.kaggle.com/code/rasmusrse/graphnet-example\">dedicated to GraphNet</a>, large values of \\\\(\\kappa\\\\) are used to select well-trained examples. However, the main thing is that the decrease in loss occurs in the direction of decreasing the error angle; and that's exactly what we need.\n\n## Theoretical part\n\nLet's now give a mathematical justification <a href=\"https://arxiv.org/pdf/1812.04616.pdf\">von Mises-Fisher Loss</a>. Let the probability of deviation of a 3-dimensional <i>unit</i> vector \\\\(\\hat{\\mathbf{n}}\\\\) from a fixed <i>unit</i> vector \\\\(\\mathbf{n}\\\\) be described by the von Mises-Fisher (vMF) distribution :\n$$\np(\\hat{\\mathbf{n}}| \\mathbf{n}) = C(\\kappa)\\, e^{\\kappa\\, \\hat{\\mathbf{n}}\\cdot\\mathbf{n}} = \nC(\\kappa)\\, e^{\\kappa\\, \\cos\\Psi}.\n$$\nIf the vectors are the same: \\\\(\\hat{\\mathbf{n}}\\mathbf{n}=1\\\\) - the probability is maximal; if they have the opposite direction, \\\\(\\hat{\\mathbf{n}}\\mathbf{n}=-1\\\\) is minimal. The probability decay rate is characterized by the \\\\(\\kappa\\\\) (concentration parameter). The value of \\\\(\\kappa=0\\\\) defines a uniform distribution over the hypersphere and \\\\(\\kappa = \\infty\\\\) defines a point distribution at \\\\(\\mathbf{n}\\\\). If the angle \\\\(\\Psi\\\\) between the vectors \\\\(\\hat{\\mathbf{n}}\\\\) and \\\\(\\mathbf{n}\\\\) is small, then we simply get a normal distribution: \\\\(p(\\hat{\\mathbf{ n}}| \\mathbf{n}) \\sim e^{-\\kappa\\, \\Psi^2/2}\\\\).\n\nTo find the normalization constant \\\\(C\\\\) of the distribution, it is necessary to calculate the integral over the entire space \\\\(\\hat{\\mathbf{n}}\\\\). Since this vector is a unit vector, in three-dimensional space it is easy to do this in a spherical coordinate system by choosing the \\\\(z\\\\) axis along the vector \\\\(\\mathbf{n}\\\\):\n$$\nC(\\kappa)\\, \\int^{2\\pi}_0 d\\phi \\int^\\pi_0 \\sin\\Psi \\, d\\Psi\\,\\,e^{\\kappa \\cos\\Psi} =1.\n$$\nIn an m-dimensional space, the result of such an integration would not be possible to express in elementary functions. In three-dimensional space, we get the simple expression:\n$$\nC(\\kappa) = \\frac{1}{2\\pi}\\,\\frac{\\kappa\\,e^{-\\kappa}}{1-e^{-2\\kappa} }.\n$$\nNote that the normalization factor is finite at zero: \\\\(C(0)=1/4\\pi\\\\).\n\nLet the model output vector \\\\(\\hat{\\mathbf{n}}\\\\), as well as the target vector \\\\(\\mathbf{n}\\\\) be normalized to one. We will maximize the probability \\\\(p(\\hat{\\mathbf{n}}| \\mathbf{n})\\\\), or, equivalently, minimize its negative logarithm (log-likelihood):\n$$\n\\mathcal{L}(\\hat{\\mathbf{n}},\\,\\mathbf{n}) = -\\log p(\\hat{\\mathbf{n}}| \\mathbf{n}) = -\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) - \\log C(\\kappa).\n$$\nAt this point, it is necessary to move away from mathematics and turn to magic. Above, we considered that \\\\(\\kappa\\\\) is a constant characterizing the sharpness of the probability distribution. Let us now assume that this is not a constant, but the length of the predicted vector \\\\(\\kappa=\\|\\hat{\\mathbf{n}}\\|\\\\) before its normalization. As a result, the distribution will be close to uniform (the loss is large) for small lengths \\\\(\\hat{\\mathbf{n}}\\\\). For large \\\\(\\kappa\\\\), the loss will be smaller. Now we can consider that the vector \\\\(\\hat{\\mathbf{n}}\\\\) is not actually normalized, assuming that \\\\(\\kappa\\,(\\hat{\\mathbf{n}}\\cdot\\mathbf{n}) \\mapsto \\hat{\\mathbf{n}}\\cdot\\mathbf{n}\\\\) and \\\\(\\hat{\\mathbf{n}}\\\\) has an arbitrary length equal to \\\\(\\kappa\\\\).\n\nThat's all.\nLet the global, deep minimum be with you :)\nGood luck!\n\n",
    "2171142": "Very nice explanation.  The graphnet vonMises-Fisher loss function you mention predicts 4 numbers:  [x,y,z,kappa], while your formula computes kappa from the norm of (x,y,z).  Can you explain the difference?",
    "2170907": "@synset does anyone have the explicit form of the 2D vMF loss too plz ? i am interested by any reference too. Thx",
    "2170104": "I tried your loss, my model converges faster but performs as good as before (using a MAE loss for both angles).\nIs that normal, or should I seek for a mistake in my code ?",
    "2168863": "Great! Not it's clear.",
    "2167758": "pure gold thread, ty !",
    "2169034": ""
  }
}