{
  "id": 473706,
  "title": "Jensen-Shannon divergence instead of Kullback-Leibler divergence",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/473706",
  "author_name": "",
  "post_date": "2024-02-05T20:09:26.503159500Z",
  "votes": 9,
  "comment_count": 2,
  "views": 0,
  "content": "<p>The formula of the  KL Divergence (KLD) is:</p>\n<h1>KLD = Σ P(i) * log(P(i)/Q(i))</h1>\n<p>where:</p>\n<ul>\n<li>Σ means the sum for all i, i being the classes</li>\n<li>P(i) is the predicted probability of i</li>\n<li>Q(i) is the true probability of i</li>\n</ul>\n<p>KLD is strictly positive and has, from what I have understood, 4 assumptions:</p>\n<ul>\n<li>Σ P(i) = 1</li>\n<li>Σ Q(i) = 1</li>\n<li>P(i) &gt; 0 for all i</li>\n<li>Q(i) &gt; 0 for all i</li>\n</ul>\n<p>The train set does not fill these conditions since we can find some cases where Q(i) = 0 in almost every row.<br>\nThe KL divergence function cited on the main page of the compeition (<a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook</a>) ignores the cases when Q(i) = 0, but this can lead to Σ P(i) &lt; 1 which is also a problem. When trying the proposed KLD function, I have had many cases where KLD was negative, results which should not happen since KLD is supposed to be strictly positive. </p>\n<p>The Jensen-Shannon divergence (JSD) is another metric aiming to calculate the similarity between two distributions. </p>\n<h1>JSD = 1/2 * Σ [ P(i) * log( 2 * P(i) / M(i) ) ] + 1/2 * Σ [ Q(i) * log( 2 * Q(i) / M(i) ) ]</h1>\n<p>where:</p>\n<ul>\n<li>Σ means the sum for all i, i being the classes</li>\n<li>P(i) is the predicted probability of i</li>\n<li>Q(i) is the true probability of i</li>\n<li>M(i) = P(i) + Q(i)</li>\n</ul>\n<p>The JSD is more robust to null values than the KL divergence because the denominator of the term in the log function is not simply Q(i) but M(i) = (P(i) + Q(i)). By simply clipping all P(i) = 0 to P(i) = 1e-15, as done in many competitions, we ensure that M(i) is never 0. We can apply the same process on Q(i) to ensure that log(0) never occurs, or use a if condition to prevent it and set the result to 0. </p>\n<p>To any one interested in understanding JSD, I highly recommend to read this great article by Aparna Dhinakaran:<br>\n<a href=\"https://towardsdatascience.com/how-to-understand-and-use-jensen-shannon-divergence-b10e11b03fd6\" target=\"_blank\">https://towardsdatascience.com/how-to-understand-and-use-jensen-shannon-divergence-b10e11b03fd6</a></p>",
  "messages": [
    {
      "id": "2637675",
      "postDate": "02/05/2024 20:09:26",
      "content": "<p>The formula of the  KL Divergence (KLD) is:</p>\n<h1>KLD = Σ P(i) * log(P(i)/Q(i))</h1>\n<p>where:</p>\n<ul>\n<li>Σ means the sum for all i, i being the classes</li>\n<li>P(i) is the predicted probability of i</li>\n<li>Q(i) is the true probability of i</li>\n</ul>\n<p>KLD is strictly positive and has, from what I have understood, 4 assumptions:</p>\n<ul>\n<li>Σ P(i) = 1</li>\n<li>Σ Q(i) = 1</li>\n<li>P(i) &gt; 0 for all i</li>\n<li>Q(i) &gt; 0 for all i</li>\n</ul>\n<p>The train set does not fill these conditions since we can find some cases where Q(i) = 0 in almost every row.<br>\nThe KL divergence function cited on the main page of the compeition (<a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook</a>) ignores the cases when Q(i) = 0, but this can lead to Σ P(i) &lt; 1 which is also a problem. When trying the proposed KLD function, I have had many cases where KLD was negative, results which should not happen since KLD is supposed to be strictly positive. </p>\n<p>The Jensen-Shannon divergence (JSD) is another metric aiming to calculate the similarity between two distributions. </p>\n<h1>JSD = 1/2 * Σ [ P(i) * log( 2 * P(i) / M(i) ) ] + 1/2 * Σ [ Q(i) * log( 2 * Q(i) / M(i) ) ]</h1>\n<p>where:</p>\n<ul>\n<li>Σ means the sum for all i, i being the classes</li>\n<li>P(i) is the predicted probability of i</li>\n<li>Q(i) is the true probability of i</li>\n<li>M(i) = P(i) + Q(i)</li>\n</ul>\n<p>The JSD is more robust to null values than the KL divergence because the denominator of the term in the log function is not simply Q(i) but M(i) = (P(i) + Q(i)). By simply clipping all P(i) = 0 to P(i) = 1e-15, as done in many competitions, we ensure that M(i) is never 0. We can apply the same process on Q(i) to ensure that log(0) never occurs, or use a if condition to prevent it and set the result to 0. </p>\n<p>To any one interested in understanding JSD, I highly recommend to read this great article by Aparna Dhinakaran:<br>\n<a href=\"https://towardsdatascience.com/how-to-understand-and-use-jensen-shannon-divergence-b10e11b03fd6\" target=\"_blank\">https://towardsdatascience.com/how-to-understand-and-use-jensen-shannon-divergence-b10e11b03fd6</a></p>",
      "rawMarkdown": "The formula of the  KL Divergence (KLD) is:\n\n# KLD = Σ P(i) * log(P(i)/Q(i))\n\nwhere:\n- Σ means the sum for all i, i being the classes\n- P(i) is the predicted probability of i\n- Q(i) is the true probability of i\n\nKLD is strictly positive and has, from what I have understood, 4 assumptions:\n- Σ P(i) = 1\n- Σ Q(i) = 1\n- P(i) > 0 for all i\n- Q(i) > 0 for all i\n\nThe train set does not fill these conditions since we can find some cases where Q(i) = 0 in almost every row.\nThe KL divergence function cited on the main page of the compeition (https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook) ignores the cases when Q(i) = 0, but this can lead to Σ P(i) < 1 which is also a problem. When trying the proposed KLD function, I have had many cases where KLD was negative, results which should not happen since KLD is supposed to be strictly positive. \n\n\nThe Jensen-Shannon divergence (JSD) is another metric aiming to calculate the similarity between two distributions. \n\n# JSD = 1/2 * Σ [ P(i) * log( 2 * P(i) / M(i) ) ] + 1/2 * Σ [ Q(i) * log( 2 * Q(i) / M(i) ) ]\n\nwhere:\n- Σ means the sum for all i, i being the classes\n- P(i) is the predicted probability of i\n- Q(i) is the true probability of i\n- M(i) = P(i) + Q(i)\n\nThe JSD is more robust to null values than the KL divergence because the denominator of the term in the log function is not simply Q(i) but M(i) = (P(i) + Q(i)). By simply clipping all P(i) = 0 to P(i) = 1e-15, as done in many competitions, we ensure that M(i) is never 0. We can apply the same process on Q(i) to ensure that log(0) never occurs, or use a if condition to prevent it and set the result to 0. \n\n\nTo any one interested in understanding JSD, I highly recommend to read this great article by Aparna Dhinakaran:\nhttps://towardsdatascience.com/how-to-understand-and-use-jensen-shannon-divergence-b10e11b03fd6",
      "votes": null
    },
    {
      "id": "2643843",
      "postDate": "02/09/2024 06:24:46",
      "content": "<p>Just for fun I tried this in the training of the efficientnetb0 in the ensemble <a href=\"https://www.kaggle.com/code/cody11null/quick-ensemble\" target=\"_blank\">Quick Ensemble</a>, but it didn't improve the score.</p>",
      "rawMarkdown": "Just for fun I tried this in the training of the efficientnetb0 in the ensemble [Quick Ensemble](https://www.kaggle.com/code/cody11null/quick-ensemble), but it didn't improve the score.",
      "votes": null
    },
    {
      "id": "2644358",
      "postDate": "02/09/2024 12:16:17",
      "content": "<p>I haven't tried efficientnet models yet. Did it make the results worse than KL divergence, or were they about the same? </p>\n<p>Unfortunately, the submissions on the LB are evaluated with KL divergence even if this metric is not ideal, reason why JS divergence may not improve the score. However, a metric does not tell us for sure which model is better, it only tells us which model is better according to this metric (this is a kind of metric subjectivity, but objectively, we cannot know which model is better). In other circumstances than in this competition, if I encounter cases with many Q(i) and P(i) = 0, I would trust more JS divergence.</p>",
      "rawMarkdown": "I haven't tried efficientnet models yet. Did it make the results worse than KL divergence, or were they about the same? \n\nUnfortunately, the submissions on the LB are evaluated with KL divergence even if this metric is not ideal, reason why JS divergence may not improve the score. However, a metric does not tell us for sure which model is better, it only tells us which model is better according to this metric (this is a kind of metric subjectivity, but objectively, we cannot know which model is better). In other circumstances than in this competition, if I encounter cases with many Q(i) and P(i) = 0, I would trust more JS divergence.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2643843,
      "author_name": "tc23p0",
      "author_url": "",
      "post_date": "02/09/2024 06:24:46",
      "content": "<p>Just for fun I tried this in the training of the efficientnetb0 in the ensemble <a href=\"https://www.kaggle.com/code/cody11null/quick-ensemble\" target=\"_blank\">Quick Ensemble</a>, but it didn't improve the score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2644358,
          "author_name": "fabienpv",
          "author_url": "",
          "post_date": "02/09/2024 12:16:17",
          "content": "<p>I haven't tried efficientnet models yet. Did it make the results worse than KL divergence, or were they about the same? </p>\n<p>Unfortunately, the submissions on the LB are evaluated with KL divergence even if this metric is not ideal, reason why JS divergence may not improve the score. However, a metric does not tell us for sure which model is better, it only tells us which model is better according to this metric (this is a kind of metric subjectivity, but objectively, we cannot know which model is better). In other circumstances than in this competition, if I encounter cases with many Q(i) and P(i) = 0, I would trust more JS divergence.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2637675": "The formula of the  KL Divergence (KLD) is:\n\n# KLD = Σ P(i) * log(P(i)/Q(i))\n\nwhere:\n- Σ means the sum for all i, i being the classes\n- P(i) is the predicted probability of i\n- Q(i) is the true probability of i\n\nKLD is strictly positive and has, from what I have understood, 4 assumptions:\n- Σ P(i) = 1\n- Σ Q(i) = 1\n- P(i) > 0 for all i\n- Q(i) > 0 for all i\n\nThe train set does not fill these conditions since we can find some cases where Q(i) = 0 in almost every row.\nThe KL divergence function cited on the main page of the compeition (https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook) ignores the cases when Q(i) = 0, but this can lead to Σ P(i) < 1 which is also a problem. When trying the proposed KLD function, I have had many cases where KLD was negative, results which should not happen since KLD is supposed to be strictly positive. \n\n\nThe Jensen-Shannon divergence (JSD) is another metric aiming to calculate the similarity between two distributions. \n\n# JSD = 1/2 * Σ [ P(i) * log( 2 * P(i) / M(i) ) ] + 1/2 * Σ [ Q(i) * log( 2 * Q(i) / M(i) ) ]\n\nwhere:\n- Σ means the sum for all i, i being the classes\n- P(i) is the predicted probability of i\n- Q(i) is the true probability of i\n- M(i) = P(i) + Q(i)\n\nThe JSD is more robust to null values than the KL divergence because the denominator of the term in the log function is not simply Q(i) but M(i) = (P(i) + Q(i)). By simply clipping all P(i) = 0 to P(i) = 1e-15, as done in many competitions, we ensure that M(i) is never 0. We can apply the same process on Q(i) to ensure that log(0) never occurs, or use a if condition to prevent it and set the result to 0. \n\n\nTo any one interested in understanding JSD, I highly recommend to read this great article by Aparna Dhinakaran:\nhttps://towardsdatascience.com/how-to-understand-and-use-jensen-shannon-divergence-b10e11b03fd6",
    "2643843": "Just for fun I tried this in the training of the efficientnetb0 in the ensemble [Quick Ensemble](https://www.kaggle.com/code/cody11null/quick-ensemble), but it didn't improve the score.",
    "2644358": "I haven't tried efficientnet models yet. Did it make the results worse than KL divergence, or were they about the same? \n\nUnfortunately, the submissions on the LB are evaluated with KL divergence even if this metric is not ideal, reason why JS divergence may not improve the score. However, a metric does not tell us for sure which model is better, it only tells us which model is better according to this metric (this is a kind of metric subjectivity, but objectively, we cannot know which model is better). In other circumstances than in this competition, if I encounter cases with many Q(i) and P(i) = 0, I would trust more JS divergence."
  },
  "source": "meta"
}