{
  "id": 248972,
  "title": "Does anyone balance the dataset?",
  "url": "/competitions/seti-breakthrough-listen/discussion/248972",
  "author_name": "Karol Pajak",
  "post_date": "2021-06-25T23:13:29.020000",
  "votes": 5,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Positive to negative examples ratio is 1:9, this is a significant class imbalance.<br>\nMost example datasets don't do anything about it?<br>\nCan someone explain why they get such high scores without balancing the data? (i.e. bringing the ratio to about 50:50)<br>\nThanks</p>",
  "messages": [
    {
      "id": 1365599,
      "postDate": "2021-06-25T23:13:29.020Z",
      "content": "<p>Positive to negative examples ratio is 1:9, this is a significant class imbalance.<br>\nMost example datasets don't do anything about it?<br>\nCan someone explain why they get such high scores without balancing the data? (i.e. bringing the ratio to about 50:50)<br>\nThanks</p>",
      "rawMarkdown": "Positive to negative examples ratio is 1:9, this is a significant class imbalance.\nMost example datasets don't do anything about it?\nCan someone explain why they get such high scores without balancing the data? (i.e. bringing the ratio to about 50:50)\nThanks",
      "votes": 5
    },
    {
      "id": 1366577,
      "postDate": "2021-06-27T00:08:52.557Z",
      "content": "<p>Class Imbalance is not a major issue if model is evaluated on roc-auc metric, but balancing samples might help model to converge faster a bit. ROC-AUC Score only depends on the ordering of predictions and not on their values. </p>\n<p>You can follow the following discussion by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for details:<br>\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892</a></p>",
      "rawMarkdown": "Class Imbalance is not a major issue if model is evaluated on roc-auc metric, but balancing samples might help model to converge faster a bit. ROC-AUC Score only depends on the ordering of predictions and not on their values. \n\nYou can follow the following discussion by @cpmpml for details:\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892",
      "votes": 3
    },
    {
      "id": 1368302,
      "postDate": "2021-06-28T13:35:54.263Z",
      "content": "<p>I posted this in the SIIM discussion by mistake, but I am doing simple oversampling (3x) to improve mixup learning.  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#1368157\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#1368157</a></p>",
      "rawMarkdown": "I posted this in the SIIM discussion by mistake, but I am doing simple oversampling (3x) to improve mixup learning.  https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#1368157",
      "votes": 4
    },
    {
      "id": 1365723,
      "postDate": "2021-06-26T04:22:18.527Z",
      "content": "<p>I use class Weights Argument in model.fit(). My class weights are <code>{0: 1, 1: 5.39553643}</code>. I calculated this using <code>weights = sklearn.utils.class_weight.compute_class_weight(\"balanced\", np.unique(y_train), y_train)</code>. Basically if the model classifies 1 wrong then the model is penalized 5 times more than classifying 0 wrong </p>",
      "rawMarkdown": "I use class Weights Argument in model.fit(). My class weights are `{0: 1, 1: 5.39553643}`. I calculated this using `weights = sklearn.utils.class_weight.compute_class_weight(\"balanced\", np.unique(y_train), y_train)`. Basically if the model classifies 1 wrong then the model is penalized 5 times more than classifying 0 wrong ",
      "votes": 4
    },
    {
      "id": 1370008,
      "postDate": "2021-06-29T20:04:08.940Z",
      "content": "<p>Reading these great replies, especially on the linked thread, reinforce my belief that effective ML is truly a bag of tricks. Step 1 is to acquire knowledge about a bunch of these tricks, Step 2 is to develop the wisdom and intuition through experience (yours or others) on how best and when to use them.</p>",
      "rawMarkdown": "Reading these great replies, especially on the linked thread, reinforce my belief that effective ML is truly a bag of tricks. Step 1 is to acquire knowledge about a bunch of these tricks, Step 2 is to develop the wisdom and intuition through experience (yours or others) on how best and when to use them.",
      "votes": 1
    },
    {
      "id": 1365724,
      "postDate": "2021-06-26T04:25:32.350Z",
      "content": "<p>In machine learning, when building a classification model with data having far more instances of one class than another, the initial default classifier is often unsatisfactory because it classifies almost every case as the majority class. Many articles show you how you could use oversampling (e.g. SMOTE) or sometimes undersampling or simply class-based sample weighting to retrain the model on “rebalanced” data, but this isn’t always necessary. Here we aim instead to show how much you can do without balancing the data or retraining the model.</p>\n<p>We do this by simply adjusting the the threshold for which we say “Class 1” when the model’s predicted probability of Class 1 is above it in two-class classification, rather than naïvely using the default classification rule which chooses which ever class is predicted to be most probable (probability threshold of 0.5). We will see how this gives you the flexibility to make any desired trade-off between false positive and false negative classifications while avoiding problems created by rebalancing the data.</p>",
      "rawMarkdown": "In machine learning, when building a classification model with data having far more instances of one class than another, the initial default classifier is often unsatisfactory because it classifies almost every case as the majority class. Many articles show you how you could use oversampling (e.g. SMOTE) or sometimes undersampling or simply class-based sample weighting to retrain the model on “rebalanced” data, but this isn’t always necessary. Here we aim instead to show how much you can do without balancing the data or retraining the model.\n\nWe do this by simply adjusting the the threshold for which we say “Class 1” when the model’s predicted probability of Class 1 is above it in two-class classification, rather than naïvely using the default classification rule which chooses which ever class is predicted to be most probable (probability threshold of 0.5). We will see how this gives you the flexibility to make any desired trade-off between false positive and false negative classifications while avoiding problems created by rebalancing the data.",
      "votes": 1
    },
    {
      "id": 1377503,
      "postDate": "2021-07-05T22:58:23.457Z",
      "content": "<p>Unless your dataset is too large to work with, I don't see what problem is being solved with balancing the dataset.  If you're using a metric like accuracy, you shouldn't be, it's not a good metric in general.  Good metrics for classification problems, such as AUC or logloss, work well regardless of the ratio of positive to negative examples.</p>",
      "rawMarkdown": "Unless your dataset is too large to work with, I don't see what problem is being solved with balancing the dataset.  If you're using a metric like accuracy, you shouldn't be, it's not a good metric in general.  Good metrics for classification problems, such as AUC or logloss, work well regardless of the ratio of positive to negative examples."
    },
    {
      "id": 1365966,
      "postDate": "2021-06-26T10:08:06.320Z",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/karolpajak\" target=\"_blank\">@karolpajak</a> ,<br>\nI tried to balance the dataset using <code>WeightedRandomSampler</code> while training a efficientnet_b1 model, the gains for the competition metric and loss function were very marginal. So I parked the idea as of now.<br>\nMay come back to it towards the end of the comp if I need marginal increments to the score to jump some places.</p>\n<p>Coming to the reason why it is not necessarily a neat solution is, we can approach it from the other side as well. If the loss function is defined such that it penalizes a wrong classification of the minority class more than the majority class (given a certain ratio) then eventually, given enough passes, the NN will learn to classify both of the classes properly.</p>\n<p>Why I prefer the 2nd method?</p>\n<ul>\n<li>First of all it is easier to use.</li>\n<li>I am leveraging the entire training data without undersampling and missing out on examples.</li>\n</ul>\n<p>Hope it helps! 😊</p>",
      "rawMarkdown": "Hi @karolpajak ,\nI tried to balance the dataset using `WeightedRandomSampler` while training a efficientnet_b1 model, the gains for the competition metric and loss function were very marginal. So I parked the idea as of now.\nMay come back to it towards the end of the comp if I need marginal increments to the score to jump some places.\n\nComing to the reason why it is not necessarily a neat solution is, we can approach it from the other side as well. If the loss function is defined such that it penalizes a wrong classification of the minority class more than the majority class (given a certain ratio) then eventually, given enough passes, the NN will learn to classify both of the classes properly.\n\nWhy I prefer the 2nd method?\n* First of all it is easier to use.\n* I am leveraging the entire training data without undersampling and missing out on examples.\n\nHope it helps! 😊"
    }
  ],
  "comments": [
    {
      "id": 1366577,
      "author_name": "sajwankit",
      "author_url": "",
      "post_date": "2021-06-27T00:08:52.557000",
      "content": "<p>Class Imbalance is not a major issue if model is evaluated on roc-auc metric, but balancing samples might help model to converge faster a bit. ROC-AUC Score only depends on the ordering of predictions and not on their values. </p>\n<p>You can follow the following discussion by <a href=\"https://www.kaggle.com/cpmpml\" target=\"_blank\">@cpmpml</a> for details:<br>\n<a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892</a></p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 1368302,
      "author_name": "nyanp",
      "author_url": "",
      "post_date": "2021-06-28T13:35:54.263000",
      "content": "<p>I posted this in the SIIM discussion by mistake, but I am doing simple oversampling (3x) to improve mixup learning.  <a href=\"https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#1368157\" target=\"_blank\">https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#1368157</a></p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1365723,
      "author_name": "Mithil Salunkhe",
      "author_url": "",
      "post_date": "2021-06-26T04:22:18.527000",
      "content": "<p>I use class Weights Argument in model.fit(). My class weights are <code>{0: 1, 1: 5.39553643}</code>. I calculated this using <code>weights = sklearn.utils.class_weight.compute_class_weight(\"balanced\", np.unique(y_train), y_train)</code>. Basically if the model classifies 1 wrong then the model is penalized 5 times more than classifying 0 wrong </p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 1370008,
      "author_name": "عثمان",
      "author_url": "",
      "post_date": "2021-06-29T20:04:08.940000",
      "content": "<p>Reading these great replies, especially on the linked thread, reinforce my belief that effective ML is truly a bag of tricks. Step 1 is to acquire knowledge about a bunch of these tricks, Step 2 is to develop the wisdom and intuition through experience (yours or others) on how best and when to use them.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1365724,
      "author_name": "rajattomar132",
      "author_url": "",
      "post_date": "2021-06-26T04:25:32.350000",
      "content": "<p>In machine learning, when building a classification model with data having far more instances of one class than another, the initial default classifier is often unsatisfactory because it classifies almost every case as the majority class. Many articles show you how you could use oversampling (e.g. SMOTE) or sometimes undersampling or simply class-based sample weighting to retrain the model on “rebalanced” data, but this isn’t always necessary. Here we aim instead to show how much you can do without balancing the data or retraining the model.</p>\n<p>We do this by simply adjusting the the threshold for which we say “Class 1” when the model’s predicted probability of Class 1 is above it in two-class classification, rather than naïvely using the default classification rule which chooses which ever class is predicted to be most probable (probability threshold of 0.5). We will see how this gives you the flexibility to make any desired trade-off between false positive and false negative classifications while avoiding problems created by rebalancing the data.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1377503,
      "author_name": "Dmitriy Guller",
      "author_url": "",
      "post_date": "2021-07-05T22:58:23.457000",
      "content": "<p>Unless your dataset is too large to work with, I don't see what problem is being solved with balancing the dataset.  If you're using a metric like accuracy, you shouldn't be, it's not a good metric in general.  Good metrics for classification problems, such as AUC or logloss, work well regardless of the ratio of positive to negative examples.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1365966,
      "author_name": "Manav",
      "author_url": "",
      "post_date": "2021-06-26T10:08:06.320000",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/karolpajak\" target=\"_blank\">@karolpajak</a> ,<br>\nI tried to balance the dataset using <code>WeightedRandomSampler</code> while training a efficientnet_b1 model, the gains for the competition metric and loss function were very marginal. So I parked the idea as of now.<br>\nMay come back to it towards the end of the comp if I need marginal increments to the score to jump some places.</p>\n<p>Coming to the reason why it is not necessarily a neat solution is, we can approach it from the other side as well. If the loss function is defined such that it penalizes a wrong classification of the minority class more than the majority class (given a certain ratio) then eventually, given enough passes, the NN will learn to classify both of the classes properly.</p>\n<p>Why I prefer the 2nd method?</p>\n<ul>\n<li>First of all it is easier to use.</li>\n<li>I am leveraging the entire training data without undersampling and missing out on examples.</li>\n</ul>\n<p>Hope it helps! 😊</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1365599": "Positive to negative examples ratio is 1:9, this is a significant class imbalance.\nMost example datasets don't do anything about it?\nCan someone explain why they get such high scores without balancing the data? (i.e. bringing the ratio to about 50:50)\nThanks",
    "1366577": "Class Imbalance is not a major issue if model is evaluated on roc-auc metric, but balancing samples might help model to converge faster a bit. ROC-AUC Score only depends on the ordering of predictions and not on their values. \n\nYou can follow the following discussion by @cpmpml for details:\nhttps://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892",
    "1368302": "I posted this in the SIIM discussion by mistake, but I am doing simple oversampling (3x) to improve mixup learning.  https://www.kaggle.com/c/siim-isic-melanoma-classification/discussion/172892#1368157",
    "1365723": "I use class Weights Argument in model.fit(). My class weights are `{0: 1, 1: 5.39553643}`. I calculated this using `weights = sklearn.utils.class_weight.compute_class_weight(\"balanced\", np.unique(y_train), y_train)`. Basically if the model classifies 1 wrong then the model is penalized 5 times more than classifying 0 wrong ",
    "1370008": "Reading these great replies, especially on the linked thread, reinforce my belief that effective ML is truly a bag of tricks. Step 1 is to acquire knowledge about a bunch of these tricks, Step 2 is to develop the wisdom and intuition through experience (yours or others) on how best and when to use them.",
    "1365724": "In machine learning, when building a classification model with data having far more instances of one class than another, the initial default classifier is often unsatisfactory because it classifies almost every case as the majority class. Many articles show you how you could use oversampling (e.g. SMOTE) or sometimes undersampling or simply class-based sample weighting to retrain the model on “rebalanced” data, but this isn’t always necessary. Here we aim instead to show how much you can do without balancing the data or retraining the model.\n\nWe do this by simply adjusting the the threshold for which we say “Class 1” when the model’s predicted probability of Class 1 is above it in two-class classification, rather than naïvely using the default classification rule which chooses which ever class is predicted to be most probable (probability threshold of 0.5). We will see how this gives you the flexibility to make any desired trade-off between false positive and false negative classifications while avoiding problems created by rebalancing the data.",
    "1377503": "Unless your dataset is too large to work with, I don't see what problem is being solved with balancing the dataset.  If you're using a metric like accuracy, you shouldn't be, it's not a good metric in general.  Good metrics for classification problems, such as AUC or logloss, work well regardless of the ratio of positive to negative examples.",
    "1365966": "Hi @karolpajak ,\nI tried to balance the dataset using `WeightedRandomSampler` while training a efficientnet_b1 model, the gains for the competition metric and loss function were very marginal. So I parked the idea as of now.\nMay come back to it towards the end of the comp if I need marginal increments to the score to jump some places.\n\nComing to the reason why it is not necessarily a neat solution is, we can approach it from the other side as well. If the loss function is defined such that it penalizes a wrong classification of the minority class more than the majority class (given a certain ratio) then eventually, given enough passes, the NN will learn to classify both of the classes properly.\n\nWhy I prefer the 2nd method?\n* First of all it is easier to use.\n* I am leveraging the entire training data without undersampling and missing out on examples.\n\nHope it helps! 😊"
  }
}