{
  "id": 470642,
  "title": "Kullback-Leibler (KL) Divergence Properties and its Relationship to Statistical Inference",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/470642",
  "author_name": "",
  "post_date": "2024-01-24T21:42:42.022954500Z",
  "votes": 7,
  "comment_count": 1,
  "views": 0,
  "content": "<p>Suppose we have two discrete probability distributions, \\(\\mathbf{P}\\) and \\(\\mathbf{Q}\\), with a common sample space \\(E\\). In the case of this competition, let \\(\\mathbf{P}\\) represent the solution (ground truth) and let \\(\\mathbf{Q}\\) represent the submission (model output). The formula that computes the KL divergence between these distributions is as follows:</p>\n<p>$$\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = \\sum _{x \\in E} p(x) \\ln \\left( \\frac{p(x)}{q(x)} \\right),$$</p>\n<p>Where the sum is only over the support of \\(\\mathbf{P}\\) (since anything outside the support of \\(\\mathbf{P}\\) but where \\(q(x) \\neq 0\\) results in a summand of 0). \\(E\\) is the set of all possible draws from the distributions (e.g. seizure_vote, lpd_vote, etc). There are six possible outcomes in \\(E\\), then, which could be encoded as numbers ranging from 1 to 6. Thus, in the equation above, there are 6 different summands. This formula corresponds to the following code snippet from the Kullback Leibler Divergence utility:</p>\n<p><code>solution.loc[y_nonzero_indices, col] = solution.loc[y_nonzero_indices, col] * np.log(solution.loc[y_nonzero_indices, col] / submission.loc[y_nonzero_indices, col])</code></p>\n<p>Where the computation is happening per column between the solution and submission dataframes. Note that, in the utility, maxs and mins of \\(\\mathbf{Q}\\) are clipped to prevent undesirable behavior. The utility returns either the average or weighted average:</p>\n<p>if micro_average:</p>\n<pre><code>`return np.average(solution.sum(=1), =sample_weights)`\n</code></pre>\n<p>else:</p>\n<pre><code>` .average(solution.())`\n</code></pre>\n<p>Properties of the KL divergence are as follows:</p>\n<p>\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\neq \\text {KL}(\\mathbf{Q}, \\mathbf{P})\\) in general<br>\n\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\geq 0\\)<br>\nIf \\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = 0\\), then \\(\\mathbf{P} = \\mathbf{Q})\\) (iff the model is identifiable; definite)<br>\n\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\nleq \\text {KL}(\\mathbf{P}, \\mathbf{R}) + \\text {KL}(\\mathbf{R}, \\mathbf{Q})\\) (triangle inequality; in general)</p>\n<p>Due to the first and last properties, the KL divergence cannot properly be called a distance, but instead a <strong>divergence</strong>. From a statistical inference perspective, the asymmetry of the KL divergence is key to the ability to estimate it (via expectations and the law of large numbers). In fact, finding the parameters of \\(\\mathbf{Q}\\) that minimizes estimate for \\(\\text {KL}(\\mathbf{P}, \\mathbf{Q})\\) is exactly the same as finding the parameters of \\(\\mathbf{Q}\\) that maximizes the likelihood of the observed data (i.e. the <strong>maximum likelihood estimate</strong>; model must be identifiable).</p>\n<p>Sources:<br>\n<em>MITx 18.6501x Fundamentals of Statistics - Unit 3 Material</em><br>\n<em>Kullback Leibler Divergence competition metric notebook</em> (<a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook</a>)</p>",
  "messages": [
    {
      "id": "2618646",
      "postDate": "01/24/2024 21:42:42",
      "content": "<p>Suppose we have two discrete probability distributions, \\(\\mathbf{P}\\) and \\(\\mathbf{Q}\\), with a common sample space \\(E\\). In the case of this competition, let \\(\\mathbf{P}\\) represent the solution (ground truth) and let \\(\\mathbf{Q}\\) represent the submission (model output). The formula that computes the KL divergence between these distributions is as follows:</p>\n<p>$$\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = \\sum _{x \\in E} p(x) \\ln \\left( \\frac{p(x)}{q(x)} \\right),$$</p>\n<p>Where the sum is only over the support of \\(\\mathbf{P}\\) (since anything outside the support of \\(\\mathbf{P}\\) but where \\(q(x) \\neq 0\\) results in a summand of 0). \\(E\\) is the set of all possible draws from the distributions (e.g. seizure_vote, lpd_vote, etc). There are six possible outcomes in \\(E\\), then, which could be encoded as numbers ranging from 1 to 6. Thus, in the equation above, there are 6 different summands. This formula corresponds to the following code snippet from the Kullback Leibler Divergence utility:</p>\n<p><code>solution.loc[y_nonzero_indices, col] = solution.loc[y_nonzero_indices, col] * np.log(solution.loc[y_nonzero_indices, col] / submission.loc[y_nonzero_indices, col])</code></p>\n<p>Where the computation is happening per column between the solution and submission dataframes. Note that, in the utility, maxs and mins of \\(\\mathbf{Q}\\) are clipped to prevent undesirable behavior. The utility returns either the average or weighted average:</p>\n<p>if micro_average:</p>\n<pre><code>`return np.average(solution.sum(=1), =sample_weights)`\n</code></pre>\n<p>else:</p>\n<pre><code>` .average(solution.())`\n</code></pre>\n<p>Properties of the KL divergence are as follows:</p>\n<p>\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\neq \\text {KL}(\\mathbf{Q}, \\mathbf{P})\\) in general<br>\n\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\geq 0\\)<br>\nIf \\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = 0\\), then \\(\\mathbf{P} = \\mathbf{Q})\\) (iff the model is identifiable; definite)<br>\n\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\nleq \\text {KL}(\\mathbf{P}, \\mathbf{R}) + \\text {KL}(\\mathbf{R}, \\mathbf{Q})\\) (triangle inequality; in general)</p>\n<p>Due to the first and last properties, the KL divergence cannot properly be called a distance, but instead a <strong>divergence</strong>. From a statistical inference perspective, the asymmetry of the KL divergence is key to the ability to estimate it (via expectations and the law of large numbers). In fact, finding the parameters of \\(\\mathbf{Q}\\) that minimizes estimate for \\(\\text {KL}(\\mathbf{P}, \\mathbf{Q})\\) is exactly the same as finding the parameters of \\(\\mathbf{Q}\\) that maximizes the likelihood of the observed data (i.e. the <strong>maximum likelihood estimate</strong>; model must be identifiable).</p>\n<p>Sources:<br>\n<em>MITx 18.6501x Fundamentals of Statistics - Unit 3 Material</em><br>\n<em>Kullback Leibler Divergence competition metric notebook</em> (<a href=\"https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook\" target=\"_blank\">https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook</a>)</p>",
      "rawMarkdown": "Suppose we have two discrete probability distributions, \\\\(\\mathbf{P}\\\\) and \\\\(\\mathbf{Q}\\\\), with a common sample space \\\\(E\\\\). In the case of this competition, let \\\\(\\mathbf{P}\\\\) represent the solution (ground truth) and let \\\\(\\mathbf{Q}\\\\) represent the submission (model output). The formula that computes the KL divergence between these distributions is as follows:\n\n$$\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = \\sum _{x \\in E} p(x) \\ln \\left( \\frac{p(x)}{q(x)} \\right),$$\n\nWhere the sum is only over the support of \\\\(\\mathbf{P}\\\\) (since anything outside the support of \\\\(\\mathbf{P}\\\\) but where \\\\(q(x) \\neq 0\\\\) results in a summand of 0). \\\\(E\\\\) is the set of all possible draws from the distributions (e.g. seizure_vote, lpd_vote, etc). There are six possible outcomes in \\\\(E\\\\), then, which could be encoded as numbers ranging from 1 to 6. Thus, in the equation above, there are 6 different summands. This formula corresponds to the following code snippet from the Kullback Leibler Divergence utility:\n\n`solution.loc[y_nonzero_indices, col] = solution.loc[y_nonzero_indices, col] * np.log(solution.loc[y_nonzero_indices, col] / submission.loc[y_nonzero_indices, col])`\n\nWhere the computation is happening per column between the solution and submission dataframes. Note that, in the utility, maxs and mins of \\\\(\\mathbf{Q}\\\\) are clipped to prevent undesirable behavior. The utility returns either the average or weighted average:\n\nif micro_average:\n\n    `return np.average(solution.sum(axis=1), weights=sample_weights)`\n\nelse:\n\n    `return np.average(solution.mean())`\n\n Properties of the KL divergence are as follows:\n\n\\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\neq \\text {KL}(\\mathbf{Q}, \\mathbf{P})\\\\) in general\n\\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\geq 0\\\\)\nIf \\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = 0\\\\), then \\\\(\\mathbf{P} = \\mathbf{Q})\\\\) (iff the model is identifiable; definite)\n\\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\nleq \\text {KL}(\\mathbf{P}, \\mathbf{R}) + \\text {KL}(\\mathbf{R}, \\mathbf{Q})\\\\) (triangle inequality; in general)\n\nDue to the first and last properties, the KL divergence cannot properly be called a distance, but instead a **divergence**. From a statistical inference perspective, the asymmetry of the KL divergence is key to the ability to estimate it (via expectations and the law of large numbers). In fact, finding the parameters of \\\\(\\mathbf{Q}\\\\) that minimizes estimate for \\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q})\\\\) is exactly the same as finding the parameters of \\\\(\\mathbf{Q}\\\\) that maximizes the likelihood of the observed data (i.e. the **maximum likelihood estimate**; model must be identifiable).\n\nSources:\n*MITx 18.6501x Fundamentals of Statistics - Unit 3 Material*\n*Kullback Leibler Divergence competition metric notebook* (https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook)",
      "votes": null
    },
    {
      "id": "2671297",
      "postDate": "02/27/2024 12:31:48",
      "content": "<p><a href=\"https://www.kaggle.com/lmhongkhnh\" target=\"_blank\">@lmhongkhnh</a> </p>",
      "rawMarkdown": "lmhongkhnh",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2671297,
      "author_name": "levinguyen02",
      "author_url": "",
      "post_date": "02/27/2024 12:31:48",
      "content": "<p><a href=\"https://www.kaggle.com/lmhongkhnh\" target=\"_blank\">@lmhongkhnh</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2618646": "Suppose we have two discrete probability distributions, \\\\(\\mathbf{P}\\\\) and \\\\(\\mathbf{Q}\\\\), with a common sample space \\\\(E\\\\). In the case of this competition, let \\\\(\\mathbf{P}\\\\) represent the solution (ground truth) and let \\\\(\\mathbf{Q}\\\\) represent the submission (model output). The formula that computes the KL divergence between these distributions is as follows:\n\n$$\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = \\sum _{x \\in E} p(x) \\ln \\left( \\frac{p(x)}{q(x)} \\right),$$\n\nWhere the sum is only over the support of \\\\(\\mathbf{P}\\\\) (since anything outside the support of \\\\(\\mathbf{P}\\\\) but where \\\\(q(x) \\neq 0\\\\) results in a summand of 0). \\\\(E\\\\) is the set of all possible draws from the distributions (e.g. seizure_vote, lpd_vote, etc). There are six possible outcomes in \\\\(E\\\\), then, which could be encoded as numbers ranging from 1 to 6. Thus, in the equation above, there are 6 different summands. This formula corresponds to the following code snippet from the Kullback Leibler Divergence utility:\n\n`solution.loc[y_nonzero_indices, col] = solution.loc[y_nonzero_indices, col] * np.log(solution.loc[y_nonzero_indices, col] / submission.loc[y_nonzero_indices, col])`\n\nWhere the computation is happening per column between the solution and submission dataframes. Note that, in the utility, maxs and mins of \\\\(\\mathbf{Q}\\\\) are clipped to prevent undesirable behavior. The utility returns either the average or weighted average:\n\nif micro_average:\n\n    `return np.average(solution.sum(axis=1), weights=sample_weights)`\n\nelse:\n\n    `return np.average(solution.mean())`\n\n Properties of the KL divergence are as follows:\n\n\\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\neq \\text {KL}(\\mathbf{Q}, \\mathbf{P})\\\\) in general\n\\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\geq 0\\\\)\nIf \\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) = 0\\\\), then \\\\(\\mathbf{P} = \\mathbf{Q})\\\\) (iff the model is identifiable; definite)\n\\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q}) \\nleq \\text {KL}(\\mathbf{P}, \\mathbf{R}) + \\text {KL}(\\mathbf{R}, \\mathbf{Q})\\\\) (triangle inequality; in general)\n\nDue to the first and last properties, the KL divergence cannot properly be called a distance, but instead a **divergence**. From a statistical inference perspective, the asymmetry of the KL divergence is key to the ability to estimate it (via expectations and the law of large numbers). In fact, finding the parameters of \\\\(\\mathbf{Q}\\\\) that minimizes estimate for \\\\(\\text {KL}(\\mathbf{P}, \\mathbf{Q})\\\\) is exactly the same as finding the parameters of \\\\(\\mathbf{Q}\\\\) that maximizes the likelihood of the observed data (i.e. the **maximum likelihood estimate**; model must be identifiable).\n\nSources:\n*MITx 18.6501x Fundamentals of Statistics - Unit 3 Material*\n*Kullback Leibler Divergence competition metric notebook* (https://www.kaggle.com/code/metric/kullback-leibler-divergence/notebook)",
    "2671297": "lmhongkhnh"
  },
  "source": "meta"
}