{
  "id": 472395,
  "title": "KL Divergence Observations",
  "url": "/competitions/hms-harmful-brain-activity-classification/discussion/472395",
  "author_name": "",
  "post_date": "2024-02-01T00:33:20.196991200Z",
  "votes": 27,
  "comment_count": 1,
  "views": 0,
  "content": "<p>It seems it would be helpful for some people to discuss KL Divergence, as this is the metric being used in this competition.  First, let's look at the governing equation found on <a href=\"https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence\" target=\"_blank\">wiki</a>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930191%2F2905636c912d294f4f4d24da83fd7cb5%2FKL%20divergence.PNG?generation=1706681283987742&amp;alt=media\"></p>\n<p>In this equation, P and Q are both probability distributions.  The KL Divergence is calculated by taking the logarithm of the ratio between the two distributions and adding up each position.  This is a little bit different than metrics I usually deal with for performance and this is likely true for others as well. In the KLD paradigm, being confidently wrong is severely penalized.  I made a notebook found <a href=\"https://www.kaggle.com/robbob62287/kl-divergence\" target=\"_blank\">here</a> with some cherry picked examples that shows simply having something that makes very confident predictions (I refer to this as a strong learner) vs. something that makes less confident decisions (I refer to this as a weak learner), it is seen the weaker learner scores higher in KL Divergence even though its likelihood of getting the correct classification is lower. So what does this mean for this competition? I have experienced and seen throughout the discussions that others have also had success with the following ideas:</p>\n<ul>\n<li><p><strong>Early Stopping</strong> - The <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-70\" target=\"_blank\">Wavenet</a> and <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">EfficientNet</a> starter notebooks provided by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> discuss only training for a few Epochs to avoid overfitting.  This helps the networks not get to a point where it is making very confident, but wrong predictions.  I have personally had good success of doing this a little differently by having more epochs but only saving the model with the best KL Divergence.  This is already implemented in <a href=\"https://keras.io/api/callbacks/model_checkpoint/\" target=\"_blank\">TensorFlow</a>.</p></li>\n<li><p><strong>Loss Functions</strong> - Another way to encourage weak learners is to use different loss functions.  The simplest way to do this is to just use KL Divergence as the actual loss function, which is already built into <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/KLDivergence\" target=\"_blank\">TensorFlow</a>.  Another method I have found to work well is to use Label Smoothing with Categorical Crossentropy.  Using label smoothing I have gotten public LB-scores of 0.45 on the signals and 0.44 on the spectrograms as described in the starter notebooks.  I have been reading about Temperature Scaling, which sounds like an interesting idea but I have not gotten it to improve my CV scores to this point. </p></li>\n<li><p><strong>Ensembles of Weak Learners</strong> - It was noted in <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> 's <a href=\"https://www.kaggle.com/code/cody11null/quick-ensemble\" target=\"_blank\">notebook</a> that making an ensemble of two high-performing public models based on ResNet and EfficientNet generated a superior public score.  This is not hard to believe considering that having more than one model is going to cause less confidence in answers when the models do not agree, giving an advantage in the KLD loss paradigm.  I have used this on with several models, specifically to combine the 1D and Spectrogram models and have found I get superior LB scores than any of the independent models, typically better than 0.4.</p></li>\n<li><p><strong>Smaller Models</strong> - My first instinct when seeing the WaveNet and EffNet examples from Chris was to try beefing up the models by making them larger and augmenting the data.  To my surprise, this did not help things and the CV scores I saw in training larger models were worse! This shows Chris going from EffNet B2 down to B0 was not a sort of gamesmanship for this competition, but a very helpful hint, as overfitting is severely penalized because of the evaluation metric.</p></li>\n<li><p><strong>Robust Features</strong> - I first tried using just the raw signals provided, mostly because I know very little about EEGs.  This proved pretty fruitless.  The discussion on EEGs found <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a>, is extremely enlightening.  It was interesting to me that meaningful information is found &lt; 30Hz, as this is where our brains operate.  This is very different than the things I research professionally and understanding the features is invaluable for improving performance.  I have found when researching zero-shot classification methods, that this tends to be the most important factor.  It is also usually the most difficult factor for me when doing Kaggles as I generally know nothing about the subjects when starting (like this competition) and do not have a ton of time to find them.  </p></li>\n</ul>\n<p>Hope this helps!  Would love to have discussions about anything else people have found helpful or have seen in literature that could be helpful in avoiding wrong, but confident answers.</p>",
  "messages": [
    {
      "id": "2629675",
      "postDate": "02/01/2024 00:33:20",
      "content": "<p>It seems it would be helpful for some people to discuss KL Divergence, as this is the metric being used in this competition.  First, let's look at the governing equation found on <a href=\"https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence\" target=\"_blank\">wiki</a>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930191%2F2905636c912d294f4f4d24da83fd7cb5%2FKL%20divergence.PNG?generation=1706681283987742&amp;alt=media\"></p>\n<p>In this equation, P and Q are both probability distributions.  The KL Divergence is calculated by taking the logarithm of the ratio between the two distributions and adding up each position.  This is a little bit different than metrics I usually deal with for performance and this is likely true for others as well. In the KLD paradigm, being confidently wrong is severely penalized.  I made a notebook found <a href=\"https://www.kaggle.com/robbob62287/kl-divergence\" target=\"_blank\">here</a> with some cherry picked examples that shows simply having something that makes very confident predictions (I refer to this as a strong learner) vs. something that makes less confident decisions (I refer to this as a weak learner), it is seen the weaker learner scores higher in KL Divergence even though its likelihood of getting the correct classification is lower. So what does this mean for this competition? I have experienced and seen throughout the discussions that others have also had success with the following ideas:</p>\n<ul>\n<li><p><strong>Early Stopping</strong> - The <a href=\"https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-70\" target=\"_blank\">Wavenet</a> and <a href=\"https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57\" target=\"_blank\">EfficientNet</a> starter notebooks provided by <a href=\"https://www.kaggle.com/cdeotte\" target=\"_blank\">@cdeotte</a> discuss only training for a few Epochs to avoid overfitting.  This helps the networks not get to a point where it is making very confident, but wrong predictions.  I have personally had good success of doing this a little differently by having more epochs but only saving the model with the best KL Divergence.  This is already implemented in <a href=\"https://keras.io/api/callbacks/model_checkpoint/\" target=\"_blank\">TensorFlow</a>.</p></li>\n<li><p><strong>Loss Functions</strong> - Another way to encourage weak learners is to use different loss functions.  The simplest way to do this is to just use KL Divergence as the actual loss function, which is already built into <a href=\"https://www.tensorflow.org/api_docs/python/tf/keras/losses/KLDivergence\" target=\"_blank\">TensorFlow</a>.  Another method I have found to work well is to use Label Smoothing with Categorical Crossentropy.  Using label smoothing I have gotten public LB-scores of 0.45 on the signals and 0.44 on the spectrograms as described in the starter notebooks.  I have been reading about Temperature Scaling, which sounds like an interesting idea but I have not gotten it to improve my CV scores to this point. </p></li>\n<li><p><strong>Ensembles of Weak Learners</strong> - It was noted in <a href=\"https://www.kaggle.com/cody11null\" target=\"_blank\">@cody11null</a> 's <a href=\"https://www.kaggle.com/code/cody11null/quick-ensemble\" target=\"_blank\">notebook</a> that making an ensemble of two high-performing public models based on ResNet and EfficientNet generated a superior public score.  This is not hard to believe considering that having more than one model is going to cause less confidence in answers when the models do not agree, giving an advantage in the KLD loss paradigm.  I have used this on with several models, specifically to combine the 1D and Spectrogram models and have found I get superior LB scores than any of the independent models, typically better than 0.4.</p></li>\n<li><p><strong>Smaller Models</strong> - My first instinct when seeing the WaveNet and EffNet examples from Chris was to try beefing up the models by making them larger and augmenting the data.  To my surprise, this did not help things and the CV scores I saw in training larger models were worse! This shows Chris going from EffNet B2 down to B0 was not a sort of gamesmanship for this competition, but a very helpful hint, as overfitting is severely penalized because of the evaluation metric.</p></li>\n<li><p><strong>Robust Features</strong> - I first tried using just the raw signals provided, mostly because I know very little about EEGs.  This proved pretty fruitless.  The discussion on EEGs found <a href=\"https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760\" target=\"_blank\">here</a>, is extremely enlightening.  It was interesting to me that meaningful information is found &lt; 30Hz, as this is where our brains operate.  This is very different than the things I research professionally and understanding the features is invaluable for improving performance.  I have found when researching zero-shot classification methods, that this tends to be the most important factor.  It is also usually the most difficult factor for me when doing Kaggles as I generally know nothing about the subjects when starting (like this competition) and do not have a ton of time to find them.  </p></li>\n</ul>\n<p>Hope this helps!  Would love to have discussions about anything else people have found helpful or have seen in literature that could be helpful in avoiding wrong, but confident answers.</p>",
      "rawMarkdown": "It seems it would be helpful for some people to discuss KL Divergence, as this is the metric being used in this competition.  First, let's look at the governing equation found on [wiki](https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930191%2F2905636c912d294f4f4d24da83fd7cb5%2FKL%20divergence.PNG?generation=1706681283987742&alt=media)\n\nIn this equation, P and Q are both probability distributions.  The KL Divergence is calculated by taking the logarithm of the ratio between the two distributions and adding up each position.  This is a little bit different than metrics I usually deal with for performance and this is likely true for others as well. In the KLD paradigm, being confidently wrong is severely penalized.  I made a notebook found [here](https://www.kaggle.com/robbob62287/kl-divergence) with some cherry picked examples that shows simply having something that makes very confident predictions (I refer to this as a strong learner) vs. something that makes less confident decisions (I refer to this as a weak learner), it is seen the weaker learner scores higher in KL Divergence even though its likelihood of getting the correct classification is lower. So what does this mean for this competition? I have experienced and seen throughout the discussions that others have also had success with the following ideas:\n\n-  **Early Stopping** - The [Wavenet](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-70) and [EfficientNet](https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57) starter notebooks provided by @cdeotte discuss only training for a few Epochs to avoid overfitting.  This helps the networks not get to a point where it is making very confident, but wrong predictions.  I have personally had good success of doing this a little differently by having more epochs but only saving the model with the best KL Divergence.  This is already implemented in [TensorFlow](https://keras.io/api/callbacks/model_checkpoint/).\n\n-  **Loss Functions** - Another way to encourage weak learners is to use different loss functions.  The simplest way to do this is to just use KL Divergence as the actual loss function, which is already built into [TensorFlow](https://www.tensorflow.org/api_docs/python/tf/keras/losses/KLDivergence).  Another method I have found to work well is to use Label Smoothing with Categorical Crossentropy.  Using label smoothing I have gotten public LB-scores of 0.45 on the signals and 0.44 on the spectrograms as described in the starter notebooks.  I have been reading about Temperature Scaling, which sounds like an interesting idea but I have not gotten it to improve my CV scores to this point. \n\n-  **Ensembles of Weak Learners** - It was noted in @cody11null 's [notebook](https://www.kaggle.com/code/cody11null/quick-ensemble) that making an ensemble of two high-performing public models based on ResNet and EfficientNet generated a superior public score.  This is not hard to believe considering that having more than one model is going to cause less confidence in answers when the models do not agree, giving an advantage in the KLD loss paradigm.  I have used this on with several models, specifically to combine the 1D and Spectrogram models and have found I get superior LB scores than any of the independent models, typically better than 0.4.\n\n-  **Smaller Models** - My first instinct when seeing the WaveNet and EffNet examples from Chris was to try beefing up the models by making them larger and augmenting the data.  To my surprise, this did not help things and the CV scores I saw in training larger models were worse! This shows Chris going from EffNet B2 down to B0 was not a sort of gamesmanship for this competition, but a very helpful hint, as overfitting is severely penalized because of the evaluation metric.\n\n-  **Robust Features** - I first tried using just the raw signals provided, mostly because I know very little about EEGs.  This proved pretty fruitless.  The discussion on EEGs found [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760), is extremely enlightening.  It was interesting to me that meaningful information is found < 30Hz, as this is where our brains operate.  This is very different than the things I research professionally and understanding the features is invaluable for improving performance.  I have found when researching zero-shot classification methods, that this tends to be the most important factor.  It is also usually the most difficult factor for me when doing Kaggles as I generally know nothing about the subjects when starting (like this competition) and do not have a ton of time to find them.  \n\nHope this helps!  Would love to have discussions about anything else people have found helpful or have seen in literature that could be helpful in avoiding wrong, but confident answers.",
      "votes": null
    },
    {
      "id": "2635656",
      "postDate": "02/04/2024 14:41:36",
      "content": "<p>Thank you for the useful information!</p>",
      "rawMarkdown": "Thank you for the useful information!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2635656,
      "author_name": "justinkmoon1",
      "author_url": "",
      "post_date": "02/04/2024 14:41:36",
      "content": "<p>Thank you for the useful information!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2629675": "It seems it would be helpful for some people to discuss KL Divergence, as this is the metric being used in this competition.  First, let's look at the governing equation found on [wiki](https://en.wikipedia.org/wiki/Kullback%E2%80%93Leibler_divergence):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F2930191%2F2905636c912d294f4f4d24da83fd7cb5%2FKL%20divergence.PNG?generation=1706681283987742&alt=media)\n\nIn this equation, P and Q are both probability distributions.  The KL Divergence is calculated by taking the logarithm of the ratio between the two distributions and adding up each position.  This is a little bit different than metrics I usually deal with for performance and this is likely true for others as well. In the KLD paradigm, being confidently wrong is severely penalized.  I made a notebook found [here](https://www.kaggle.com/robbob62287/kl-divergence) with some cherry picked examples that shows simply having something that makes very confident predictions (I refer to this as a strong learner) vs. something that makes less confident decisions (I refer to this as a weak learner), it is seen the weaker learner scores higher in KL Divergence even though its likelihood of getting the correct classification is lower. So what does this mean for this competition? I have experienced and seen throughout the discussions that others have also had success with the following ideas:\n\n-  **Early Stopping** - The [Wavenet](https://www.kaggle.com/code/cdeotte/wavenet-starter-lb-0-70) and [EfficientNet](https://www.kaggle.com/code/cdeotte/efficientnetb2-starter-lb-0-57) starter notebooks provided by @cdeotte discuss only training for a few Epochs to avoid overfitting.  This helps the networks not get to a point where it is making very confident, but wrong predictions.  I have personally had good success of doing this a little differently by having more epochs but only saving the model with the best KL Divergence.  This is already implemented in [TensorFlow](https://keras.io/api/callbacks/model_checkpoint/).\n\n-  **Loss Functions** - Another way to encourage weak learners is to use different loss functions.  The simplest way to do this is to just use KL Divergence as the actual loss function, which is already built into [TensorFlow](https://www.tensorflow.org/api_docs/python/tf/keras/losses/KLDivergence).  Another method I have found to work well is to use Label Smoothing with Categorical Crossentropy.  Using label smoothing I have gotten public LB-scores of 0.45 on the signals and 0.44 on the spectrograms as described in the starter notebooks.  I have been reading about Temperature Scaling, which sounds like an interesting idea but I have not gotten it to improve my CV scores to this point. \n\n-  **Ensembles of Weak Learners** - It was noted in @cody11null 's [notebook](https://www.kaggle.com/code/cody11null/quick-ensemble) that making an ensemble of two high-performing public models based on ResNet and EfficientNet generated a superior public score.  This is not hard to believe considering that having more than one model is going to cause less confidence in answers when the models do not agree, giving an advantage in the KLD loss paradigm.  I have used this on with several models, specifically to combine the 1D and Spectrogram models and have found I get superior LB scores than any of the independent models, typically better than 0.4.\n\n-  **Smaller Models** - My first instinct when seeing the WaveNet and EffNet examples from Chris was to try beefing up the models by making them larger and augmenting the data.  To my surprise, this did not help things and the CV scores I saw in training larger models were worse! This shows Chris going from EffNet B2 down to B0 was not a sort of gamesmanship for this competition, but a very helpful hint, as overfitting is severely penalized because of the evaluation metric.\n\n-  **Robust Features** - I first tried using just the raw signals provided, mostly because I know very little about EEGs.  This proved pretty fruitless.  The discussion on EEGs found [here](https://www.kaggle.com/competitions/hms-harmful-brain-activity-classification/discussion/469760), is extremely enlightening.  It was interesting to me that meaningful information is found < 30Hz, as this is where our brains operate.  This is very different than the things I research professionally and understanding the features is invaluable for improving performance.  I have found when researching zero-shot classification methods, that this tends to be the most important factor.  It is also usually the most difficult factor for me when doing Kaggles as I generally know nothing about the subjects when starting (like this competition) and do not have a ton of time to find them.  \n\nHope this helps!  Would love to have discussions about anything else people have found helpful or have seen in literature that could be helpful in avoiding wrong, but confident answers.",
    "2635656": "Thank you for the useful information!"
  },
  "source": "meta"
}