{
  "id": 478650,
  "title": "Error Analysis for Probability Distribution: Conditional Probability",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/478650",
  "author_name": "",
  "post_date": "2024-02-21T17:16:05.537901200Z",
  "votes": 9,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I came up with the way analyzing typical label errors, and share it here.</p>\n<p>Unlike usual classification tasks, this competition's labels are provided by probability distribution.<br>\nSo we can't naively apply error analysis with confusion matrix.</p>\n<p>The possible alternatives is <em>conditional probability</em>.</p>\n<h2>Confusion Matrix by Conditional Probability</h2>\n<p>The process of obtaining conditional probability is like this:</p>\n<ol>\n<li>calculate joint probability for all possible label pairs (C, C) and take average for all EEG IDs.</li>\n<li>normalize probability with marginal probability to obtain conditional probability</li>\n</ol>\n<p>There are two type of conditional probability matrix whether normalization is taken for predictions / GTs:</p>\n<p>$$<br>\nP(y_\\text{pred} = j | y_\\text{GT} = i) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{GT}}(i)} \\\\<br>\nP_{y_\\text{GT}}(i) := \\sum_j P(y_\\text{GT} = i, y_\\text{pred} = j)<br>\n$$</p>\n<p>$$<br>\nP(y_\\text{GT} = i | y_\\text{pred} = j) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{pred}}(j)} \\\\<br>\nP_{y_\\text{pred}}(j) := \\sum_i P(y_\\text{GT} = i, y_\\text{pred} = j)<br>\n$$</p>\n<p>The notation of marginal probability is followed as [1].</p>\n<p>Although I believe this formulation is rational for this competition, feel free to point out if you find any flaws, or come up with alternative method to qualify typical label errors.</p>\n<h2>Example Plot</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F26035fefea603679dc6a40cc940b2bd2%2Ftotal_prob.jpg?generation=1708575995899880&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1d7ebbde30b42bbad525ec85d6e3b749%2Fconditional_by_gt.jpg?generation=1708576009908935&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe37f3f665bff3232739f834fd66fc3c3%2Fconditional_by_pred.jpg?generation=1708576020298774&amp;alt=media\"></p>\n<h2>Sample Code</h2>\n<pre><code> ():\n    \n    gts = gts[:, :, np.newaxis]\n    preds = preds[:, np.newaxis, :]\n     weights   :\n        weights = weights[:, np.newaxis, np.newaxis]\n        mat = (preds * gts * weights).(axis=) / weights.(axis=)\n    :\n        mat = (preds * gts).mean(axis=)\n\n     normalize:\n        mat = mat / (mat.(axis=norm_axis, keepdims=) + eps)\n     mat\n</code></pre>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://en.wikipedia.org/wiki/Marginal_distribution\" target=\"_blank\">https://en.wikipedia.org/wiki/Marginal_distribution</a></li>\n</ul>",
  "messages": [
    {
      "id": "2662050",
      "postDate": "02/21/2024 17:16:05",
      "content": "<p>I came up with the way analyzing typical label errors, and share it here.</p>\n<p>Unlike usual classification tasks, this competition's labels are provided by probability distribution.<br>\nSo we can't naively apply error analysis with confusion matrix.</p>\n<p>The possible alternatives is <em>conditional probability</em>.</p>\n<h2>Confusion Matrix by Conditional Probability</h2>\n<p>The process of obtaining conditional probability is like this:</p>\n<ol>\n<li>calculate joint probability for all possible label pairs (C, C) and take average for all EEG IDs.</li>\n<li>normalize probability with marginal probability to obtain conditional probability</li>\n</ol>\n<p>There are two type of conditional probability matrix whether normalization is taken for predictions / GTs:</p>\n<p>$$<br>\nP(y_\\text{pred} = j | y_\\text{GT} = i) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{GT}}(i)} \\\\<br>\nP_{y_\\text{GT}}(i) := \\sum_j P(y_\\text{GT} = i, y_\\text{pred} = j)<br>\n$$</p>\n<p>$$<br>\nP(y_\\text{GT} = i | y_\\text{pred} = j) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{pred}}(j)} \\\\<br>\nP_{y_\\text{pred}}(j) := \\sum_i P(y_\\text{GT} = i, y_\\text{pred} = j)<br>\n$$</p>\n<p>The notation of marginal probability is followed as [1].</p>\n<p>Although I believe this formulation is rational for this competition, feel free to point out if you find any flaws, or come up with alternative method to qualify typical label errors.</p>\n<h2>Example Plot</h2>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F26035fefea603679dc6a40cc940b2bd2%2Ftotal_prob.jpg?generation=1708575995899880&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1d7ebbde30b42bbad525ec85d6e3b749%2Fconditional_by_gt.jpg?generation=1708576009908935&amp;alt=media\"><br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe37f3f665bff3232739f834fd66fc3c3%2Fconditional_by_pred.jpg?generation=1708576020298774&amp;alt=media\"></p>\n<h2>Sample Code</h2>\n<pre><code> ():\n    \n    gts = gts[:, :, np.newaxis]\n    preds = preds[:, np.newaxis, :]\n     weights   :\n        weights = weights[:, np.newaxis, np.newaxis]\n        mat = (preds * gts * weights).(axis=) / weights.(axis=)\n    :\n        mat = (preds * gts).mean(axis=)\n\n     normalize:\n        mat = mat / (mat.(axis=norm_axis, keepdims=) + eps)\n     mat\n</code></pre>\n<h2>Reference</h2>\n<ul>\n<li>[1] <a href=\"https://en.wikipedia.org/wiki/Marginal_distribution\" target=\"_blank\">https://en.wikipedia.org/wiki/Marginal_distribution</a></li>\n</ul>",
      "rawMarkdown": "I came up with the way analyzing typical label errors, and share it here.\n\nUnlike usual classification tasks, this competition's labels are provided by probability distribution.\nSo we can't naively apply error analysis with confusion matrix.\n\nThe possible alternatives is *conditional probability*.\n\n## Confusion Matrix by Conditional Probability\n\nThe process of obtaining conditional probability is like this:\n\n1. calculate joint probability for all possible label pairs (C, C) and take average for all EEG IDs.\n2. normalize probability with marginal probability to obtain conditional probability\n\nThere are two type of conditional probability matrix whether normalization is taken for predictions / GTs:\n\n$$\nP(y_\\text{pred} = j | y_\\text{GT} = i) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{GT}}(i)} \\\\\\\\\nP_{y_\\text{GT}}(i) := \\sum_j P(y_\\text{GT} = i, y_\\text{pred} = j)\n$$\n\n$$\nP(y_\\text{GT} = i | y_\\text{pred} = j) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{pred}}(j)} \\\\\\\\\nP_{y_\\text{pred}}(j) := \\sum_i P(y_\\text{GT} = i, y_\\text{pred} = j)\n$$\n\nThe notation of marginal probability is followed as [1].\n\nAlthough I believe this formulation is rational for this competition, feel free to point out if you find any flaws, or come up with alternative method to qualify typical label errors.\n\n## Example Plot\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F26035fefea603679dc6a40cc940b2bd2%2Ftotal_prob.jpg?generation=1708575995899880&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1d7ebbde30b42bbad525ec85d6e3b749%2Fconditional_by_gt.jpg?generation=1708576009908935&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe37f3f665bff3232739f834fd66fc3c3%2Fconditional_by_pred.jpg?generation=1708576020298774&alt=media)\n\n## Sample Code\n\n```python\ndef calc_conditional_prob(\n    preds: np.ndarray,\n    gts: np.ndarray,\n    weights: np.ndarray | None = None,\n    normalize: bool = True,\n    norm_axis: int = 0,\n    eps=1e-4,\n):\n    \"\"\"\n    Parameters\n    ----------\n    preds: (N, C) array of predicted probabilities\n    gts: (N, C) array of ground truth probabilities\n    weights: (N, ) array of weights\n\n    Returns\n    -------\n    conditional_matrix: (C, C) array of conditional probabilities\n    \"\"\"\n    gts = gts[:, :, np.newaxis]\n    preds = preds[:, np.newaxis, :]\n    if weights is not None:\n        weights = weights[:, np.newaxis, np.newaxis]\n        mat = (preds * gts * weights).sum(axis=0) / weights.sum(axis=0)\n    else:\n        mat = (preds * gts).mean(axis=0)\n\n    if normalize:\n        mat = mat / (mat.sum(axis=norm_axis, keepdims=True) + eps)\n    return mat\n```\n\n## Reference\n\n- [1] https://en.wikipedia.org/wiki/Marginal_distribution",
      "votes": null
    },
    {
      "id": "2662143",
      "postDate": "02/21/2024 18:16:46",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, what if you just calculate simple MAE, RMSE, MAPE or SMAPE,  then plot the confusion matrix, by doing this:</p>\n<table>\n<thead>\n<tr>\n<th>seizure_vote</th>\n<th>lpd_vote</th>\n<th>gpd_vote</th>\n<th>lrda_vote</th>\n<th>grda_vote</th>\n<th>other_vote</th>\n<th>id</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.008830102</td>\n<td>0.043293837</td>\n<td>0.043633398</td>\n<td>0.07877874</td>\n<td>0.54204774</td>\n<td>0.28341624</td>\n<td>70922502_true</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0.1111111111</td>\n<td>0.05555555556</td>\n<td>0</td>\n<td>0.2777777778</td>\n<td>0.5555555556</td>\n<td>70922502_pred</td>\n</tr>\n<tr>\n<td>0.008830102</td>\n<td>0.06781727411</td>\n<td>0.01192215756</td>\n<td>0.07877874</td>\n<td>0.2642699622</td>\n<td>0.2721393156</td>\n<td>abs_diff</td>\n</tr>\n</tbody>\n</table>",
      "rawMarkdown": "tatamikenn, what if you just calculate simple MAE, RMSE, MAPE or SMAPE,  then plot the confusion matrix, by doing this:\n| seizure_vote | lpd_vote     | gpd_vote     | lrda_vote    | grda_vote    | other_vote   | id               |\n|--------------|--------------|--------------|--------------|--------------|--------------|------------------|\n| 0.008830102  | 0.043293837  | 0.043633398  | 0.07877874   | 0.54204774   | 0.28341624   | 70922502_true   |\n| 0            | 0.1111111111 | 0.05555555556| 0            | 0.2777777778 | 0.5555555556 | 70922502_pred   |\n| 0.008830102  | 0.06781727411| 0.01192215756| 0.07877874   | 0.2642699622 | 0.2721393156 | abs_diff         |",
      "votes": null
    },
    {
      "id": "2662486",
      "postDate": "02/21/2024 23:12:47",
      "content": "<p>Seems like a slightly weird formulation to me. P(GT=i,pred=j|pred=j) can just be simplified to P(GT=i|pred=j)</p>",
      "rawMarkdown": "Seems like a slightly weird formulation to me. P(GT=i,pred=j|pred=j) can just be simplified to P(GT=i|pred=j)",
      "votes": null
    },
    {
      "id": "2662546",
      "postDate": "02/22/2024 01:17:12",
      "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <br>\nWell, I also logged label-wise mean of (pointwise) KL-divergence.<br>\nHowever, what I wanted to see here is <em>pairwise</em> information over label errors.<br>\nLike this:</p>\n<ul>\n<li><strong>lateral/general confusion</strong>: e.g. if <code>GPD</code> and <code>LPD</code> is common error, we should feed more information on L/R difference to our model.</li>\n<li><strong>signal shape confusion</strong>: e.g. if <code>GPD</code> and <code>GRDA</code> is common error, we should feed more information on signal shape to our model.</li>\n<li><strong>hidden connection over two labels</strong>: e.g. as we can see in the example plot, <code>GRDA</code> and <code>other</code> seems to have close connection. Does this mean these two labels have similar features in training examples? etc.</li>\n</ul>\n<p>So I come up with conditional probability.</p>",
      "rawMarkdown": "sergiosaharovskiy \nWell, I also logged label-wise mean of (pointwise) KL-divergence.\nHowever, what I wanted to see here is *pairwise* information over label errors.\nLike this:\n\n- **lateral/general confusion**: e.g. if `GPD` and `LPD` is common error, we should feed more information on L/R difference to our model.\n- **signal shape confusion**: e.g. if `GPD` and `GRDA` is common error, we should feed more information on signal shape to our model.\n- **hidden connection over two labels**: e.g. as we can see in the example plot, `GRDA` and `other` seems to have close connection. Does this mean these two labels have similar features in training examples? etc.\n\nSo I come up with conditional probability.",
      "votes": null
    },
    {
      "id": "2662550",
      "postDate": "02/22/2024 01:20:22",
      "content": "<p><a href=\"https://www.kaggle.com/caelhasse\" target=\"_blank\">@caelhasse</a> You are right. I found other mistake on math formulation, . fixed. Thank you for pointing out.</p>",
      "rawMarkdown": "caelhasse You are right. I found other mistake on math formulation, ~~so I fix it later~~. fixed. Thank you for pointing out.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2662143,
      "author_name": "sergiosaharovskiy",
      "author_url": "",
      "post_date": "02/21/2024 18:16:46",
      "content": "<p><a href=\"https://www.kaggle.com/tatamikenn\" target=\"_blank\">@tatamikenn</a>, what if you just calculate simple MAE, RMSE, MAPE or SMAPE,  then plot the confusion matrix, by doing this:</p>\n<table>\n<thead>\n<tr>\n<th>seizure_vote</th>\n<th>lpd_vote</th>\n<th>gpd_vote</th>\n<th>lrda_vote</th>\n<th>grda_vote</th>\n<th>other_vote</th>\n<th>id</th>\n</tr>\n</thead>\n<tbody>\n<tr>\n<td>0.008830102</td>\n<td>0.043293837</td>\n<td>0.043633398</td>\n<td>0.07877874</td>\n<td>0.54204774</td>\n<td>0.28341624</td>\n<td>70922502_true</td>\n</tr>\n<tr>\n<td>0</td>\n<td>0.1111111111</td>\n<td>0.05555555556</td>\n<td>0</td>\n<td>0.2777777778</td>\n<td>0.5555555556</td>\n<td>70922502_pred</td>\n</tr>\n<tr>\n<td>0.008830102</td>\n<td>0.06781727411</td>\n<td>0.01192215756</td>\n<td>0.07877874</td>\n<td>0.2642699622</td>\n<td>0.2721393156</td>\n<td>abs_diff</td>\n</tr>\n</tbody>\n</table>",
      "votes": null,
      "replies": [
        {
          "id": 2662546,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "02/22/2024 01:17:12",
          "content": "<p><a href=\"https://www.kaggle.com/sergiosaharovskiy\" target=\"_blank\">@sergiosaharovskiy</a> <br>\nWell, I also logged label-wise mean of (pointwise) KL-divergence.<br>\nHowever, what I wanted to see here is <em>pairwise</em> information over label errors.<br>\nLike this:</p>\n<ul>\n<li><strong>lateral/general confusion</strong>: e.g. if <code>GPD</code> and <code>LPD</code> is common error, we should feed more information on L/R difference to our model.</li>\n<li><strong>signal shape confusion</strong>: e.g. if <code>GPD</code> and <code>GRDA</code> is common error, we should feed more information on signal shape to our model.</li>\n<li><strong>hidden connection over two labels</strong>: e.g. as we can see in the example plot, <code>GRDA</code> and <code>other</code> seems to have close connection. Does this mean these two labels have similar features in training examples? etc.</li>\n</ul>\n<p>So I come up with conditional probability.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2662486,
      "author_name": "caelhasse",
      "author_url": "",
      "post_date": "02/21/2024 23:12:47",
      "content": "<p>Seems like a slightly weird formulation to me. P(GT=i,pred=j|pred=j) can just be simplified to P(GT=i|pred=j)</p>",
      "votes": null,
      "replies": [
        {
          "id": 2662550,
          "author_name": "tatamikenn",
          "author_url": "",
          "post_date": "02/22/2024 01:20:22",
          "content": "<p><a href=\"https://www.kaggle.com/caelhasse\" target=\"_blank\">@caelhasse</a> You are right. I found other mistake on math formulation, . fixed. Thank you for pointing out.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2662050": "I came up with the way analyzing typical label errors, and share it here.\n\nUnlike usual classification tasks, this competition's labels are provided by probability distribution.\nSo we can't naively apply error analysis with confusion matrix.\n\nThe possible alternatives is *conditional probability*.\n\n## Confusion Matrix by Conditional Probability\n\nThe process of obtaining conditional probability is like this:\n\n1. calculate joint probability for all possible label pairs (C, C) and take average for all EEG IDs.\n2. normalize probability with marginal probability to obtain conditional probability\n\nThere are two type of conditional probability matrix whether normalization is taken for predictions / GTs:\n\n$$\nP(y_\\text{pred} = j | y_\\text{GT} = i) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{GT}}(i)} \\\\\\\\\nP_{y_\\text{GT}}(i) := \\sum_j P(y_\\text{GT} = i, y_\\text{pred} = j)\n$$\n\n$$\nP(y_\\text{GT} = i | y_\\text{pred} = j) := \\frac{P(y_\\text{GT} = i, y_\\text{pred} = j)}{P_{y_\\text{pred}}(j)} \\\\\\\\\nP_{y_\\text{pred}}(j) := \\sum_i P(y_\\text{GT} = i, y_\\text{pred} = j)\n$$\n\nThe notation of marginal probability is followed as [1].\n\nAlthough I believe this formulation is rational for this competition, feel free to point out if you find any flaws, or come up with alternative method to qualify typical label errors.\n\n## Example Plot\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F26035fefea603679dc6a40cc940b2bd2%2Ftotal_prob.jpg?generation=1708575995899880&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2F1d7ebbde30b42bbad525ec85d6e3b749%2Fconditional_by_gt.jpg?generation=1708576009908935&alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4910466%2Fe37f3f665bff3232739f834fd66fc3c3%2Fconditional_by_pred.jpg?generation=1708576020298774&alt=media)\n\n## Sample Code\n\n```python\ndef calc_conditional_prob(\n    preds: np.ndarray,\n    gts: np.ndarray,\n    weights: np.ndarray | None = None,\n    normalize: bool = True,\n    norm_axis: int = 0,\n    eps=1e-4,\n):\n    \"\"\"\n    Parameters\n    ----------\n    preds: (N, C) array of predicted probabilities\n    gts: (N, C) array of ground truth probabilities\n    weights: (N, ) array of weights\n\n    Returns\n    -------\n    conditional_matrix: (C, C) array of conditional probabilities\n    \"\"\"\n    gts = gts[:, :, np.newaxis]\n    preds = preds[:, np.newaxis, :]\n    if weights is not None:\n        weights = weights[:, np.newaxis, np.newaxis]\n        mat = (preds * gts * weights).sum(axis=0) / weights.sum(axis=0)\n    else:\n        mat = (preds * gts).mean(axis=0)\n\n    if normalize:\n        mat = mat / (mat.sum(axis=norm_axis, keepdims=True) + eps)\n    return mat\n```\n\n## Reference\n\n- [1] https://en.wikipedia.org/wiki/Marginal_distribution",
    "2662143": "tatamikenn, what if you just calculate simple MAE, RMSE, MAPE or SMAPE,  then plot the confusion matrix, by doing this:\n| seizure_vote | lpd_vote     | gpd_vote     | lrda_vote    | grda_vote    | other_vote   | id               |\n|--------------|--------------|--------------|--------------|--------------|--------------|------------------|\n| 0.008830102  | 0.043293837  | 0.043633398  | 0.07877874   | 0.54204774   | 0.28341624   | 70922502_true   |\n| 0            | 0.1111111111 | 0.05555555556| 0            | 0.2777777778 | 0.5555555556 | 70922502_pred   |\n| 0.008830102  | 0.06781727411| 0.01192215756| 0.07877874   | 0.2642699622 | 0.2721393156 | abs_diff         |",
    "2662486": "Seems like a slightly weird formulation to me. P(GT=i,pred=j|pred=j) can just be simplified to P(GT=i|pred=j)",
    "2662546": "sergiosaharovskiy \nWell, I also logged label-wise mean of (pointwise) KL-divergence.\nHowever, what I wanted to see here is *pairwise* information over label errors.\nLike this:\n\n- **lateral/general confusion**: e.g. if `GPD` and `LPD` is common error, we should feed more information on L/R difference to our model.\n- **signal shape confusion**: e.g. if `GPD` and `GRDA` is common error, we should feed more information on signal shape to our model.\n- **hidden connection over two labels**: e.g. as we can see in the example plot, `GRDA` and `other` seems to have close connection. Does this mean these two labels have similar features in training examples? etc.\n\nSo I come up with conditional probability.",
    "2662550": "caelhasse You are right. I found other mistake on math formulation, ~~so I fix it later~~. fixed. Thank you for pointing out."
  },
  "source": "meta"
}