{
  "id": 498402,
  "title": "Data Imbalance - How to tackle? ",
  "url": "/competitions/leash-BELKA/discussion/498402",
  "author_name": "AC",
  "post_date": "2024-04-28T05:53:13.112000",
  "votes": 8,
  "comment_count": 8,
  "views": 0,
  "content": "<p>After looking into the data, the binds column is highly imbalanced:<br>\n<strong>0:</strong> 293656924 samples<br>\n<strong>1:</strong> 1589906 samples<br>\nconstituting to ~ 99.46% and ~ 0.53850062% of the total samples respectively. <br>\nI am opening this thread to learn how everybody is tackling data imbalance at this level. All the suggestions are welcome! </p>",
  "messages": [
    {
      "id": 2780281,
      "postDate": "2024-04-28T05:53:13.113Z",
      "content": "<p>After looking into the data, the binds column is highly imbalanced:<br>\n<strong>0:</strong> 293656924 samples<br>\n<strong>1:</strong> 1589906 samples<br>\nconstituting to ~ 99.46% and ~ 0.53850062% of the total samples respectively. <br>\nI am opening this thread to learn how everybody is tackling data imbalance at this level. All the suggestions are welcome! </p>",
      "rawMarkdown": "After looking into the data, the binds column is highly imbalanced:\n**0:** 293656924 samples\n**1:** 1589906 samples\nconstituting to ~ 99.46% and ~ 0.53850062% of the total samples respectively. \nI am opening this thread to learn how everybody is tackling data imbalance at this level. All the suggestions are welcome! \n\n\n\n",
      "votes": 7
    },
    {
      "id": 2780394,
      "postDate": "2024-04-28T07:08:37.057Z",
      "content": "<p>you should start to do experiments.<br>\ne.g. just use the simplest xgboot + fingerprint or chemberta (3-layer trasnformer),<br>\ntry with different pos to neg ratio in sampling your batch.</p>\n<p>maybe there is no difference? who's know?</p>\n<hr>\n<p>in datascience, decision is based on data.<br>\nno experiments results, no plan</p>\n<p>example of experiments:</p>\n<ol>\n<li>Tuning gradient boosting for imbalanced bioassay modelling with custom loss functions<br>\n<a href=\"https://github.com/dahvida/gradient_boosting_CLF\" target=\"_blank\">https://github.com/dahvida/gradient_boosting_CLF</a></li>\n</ol>\n<p>2.<a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\" target=\"_blank\">https://www.tensorflow.org/tutorials/structured_data/imbalanced_data</a></p>",
      "rawMarkdown": "you should start to do experiments.\ne.g. just use the simplest xgboot + fingerprint or chemberta (3-layer trasnformer),\ntry with different pos to neg ratio in sampling your batch.\n\nmaybe there is no difference? who's know?\n\n---\nin datascience, decision is based on data.\nno experiments results, no plan\n\nexample of experiments:\n1. Tuning gradient boosting for imbalanced bioassay modelling with custom loss functions\nhttps://github.com/dahvida/gradient_boosting_CLF\n\n2.https://www.tensorflow.org/tutorials/structured_data/imbalanced_data",
      "votes": 6,
      "replies": [
        {
          "id": 2791117,
          "postDate": "2024-05-03T13:38:48.903Z",
          "content": "<p>Oh! Captain！My Captain~😙</p>",
          "rawMarkdown": "Oh! Captain！My Captain~😙"
        }
      ]
    },
    {
      "id": 2863562,
      "postDate": "2024-06-09T14:00:18.153Z",
      "content": "<p>Trying to deal with the label imbalance by using SMOTE. I'm reading in the training data as a pyspark dataframe, but I can't find a way to apply SMOTE directly on it. It's also too big to convert to a pandas dataframe (my searches are telling me to convert it to a pandas dataframe to apply smote), is there a way to either:<br>\n1) apply SMOTE on the full dataset?<br>\n2) a good way to break up the dataset into batches and perform oversampling on each batch?</p>",
      "rawMarkdown": "Trying to deal with the label imbalance by using SMOTE. I'm reading in the training data as a pyspark dataframe, but I can't find a way to apply SMOTE directly on it. It's also too big to convert to a pandas dataframe (my searches are telling me to convert it to a pandas dataframe to apply smote), is there a way to either:\n1) apply SMOTE on the full dataset?\n2) a good way to break up the dataset into batches and perform oversampling on each batch?",
      "votes": 1,
      "replies": [
        {
          "id": 2864400,
          "postDate": "2024-06-10T06:21:11.047Z",
          "content": "<p>For me, I tried using SMOTE on a chunk of dataset and it's giving me OOM error. Still looking for the way to deal with it. </p>",
          "rawMarkdown": "For me, I tried using SMOTE on a chunk of dataset and it's giving me OOM error. Still looking for the way to deal with it. ",
          "votes": 1,
          "replies": [
            {
              "id": 2893689,
              "postDate": "2024-06-28T02:31:35.137Z",
              "content": "<p>In regards to this, I split the dataset into binds == 0 and binds == 1, and then I took binds == 1 set and randomly sampled it to a size that was manageable, and then I used SMOTE on that set and saved it as a data shard. I repeated this by continuing to take different samples from binds == 1 and used SMOTE and saved it as different shards.</p>",
              "rawMarkdown": "In regards to this, I split the dataset into binds == 0 and binds == 1, and then I took binds == 1 set and randomly sampled it to a size that was manageable, and then I used SMOTE on that set and saved it as a data shard. I repeated this by continuing to take different samples from binds == 1 and used SMOTE and saved it as different shards."
            }
          ]
        }
      ]
    },
    {
      "id": 2786172,
      "postDate": "2024-05-01T07:42:34.390Z",
      "content": "<p>Take a look at the <a href=\"https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\" target=\"_blank\">tutorial notebook</a>. They sample a 1:1 ratio (with an upper limit of 30k samples for each to reduce training times). But it will, most likely, be a well-kept secret on how to sample it best. At least at the start of the competition.</p>",
      "rawMarkdown": "Take a look at the [tutorial notebook](https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest). They sample a 1:1 ratio (with an upper limit of 30k samples for each to reduce training times). But it will, most likely, be a well-kept secret on how to sample it best. At least at the start of the competition.",
      "votes": 2
    },
    {
      "id": 2780567,
      "postDate": "2024-04-28T09:56:17.293Z",
      "content": "<p>Data imbalance needs to be addressed to ensure that the model has even data distribution to learn from and use. SMOTE is one of the techniques to address data imbalance</p>",
      "rawMarkdown": "Data imbalance needs to be addressed to ensure that the model has even data distribution to learn from and use. SMOTE is one of the techniques to address data imbalance"
    },
    {
      "id": 2780390,
      "postDate": "2024-04-28T07:01:20.500Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2780394,
      "author_name": "hengck23",
      "author_url": "",
      "post_date": "2024-04-28T07:08:37.057000",
      "content": "<p>you should start to do experiments.<br>\ne.g. just use the simplest xgboot + fingerprint or chemberta (3-layer trasnformer),<br>\ntry with different pos to neg ratio in sampling your batch.</p>\n<p>maybe there is no difference? who's know?</p>\n<hr>\n<p>in datascience, decision is based on data.<br>\nno experiments results, no plan</p>\n<p>example of experiments:</p>\n<ol>\n<li>Tuning gradient boosting for imbalanced bioassay modelling with custom loss functions<br>\n<a href=\"https://github.com/dahvida/gradient_boosting_CLF\" target=\"_blank\">https://github.com/dahvida/gradient_boosting_CLF</a></li>\n</ol>\n<p>2.<a href=\"https://www.tensorflow.org/tutorials/structured_data/imbalanced_data\" target=\"_blank\">https://www.tensorflow.org/tutorials/structured_data/imbalanced_data</a></p>",
      "votes": 6,
      "replies": [
        {
          "id": 2791117,
          "author_name": "cheese",
          "author_url": "",
          "post_date": "2024-05-03T13:38:48.903000",
          "content": "<p>Oh! Captain！My Captain~😙</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 2863562,
      "author_name": "MARVINMENG",
      "author_url": "",
      "post_date": "2024-06-09T14:00:18.153000",
      "content": "<p>Trying to deal with the label imbalance by using SMOTE. I'm reading in the training data as a pyspark dataframe, but I can't find a way to apply SMOTE directly on it. It's also too big to convert to a pandas dataframe (my searches are telling me to convert it to a pandas dataframe to apply smote), is there a way to either:<br>\n1) apply SMOTE on the full dataset?<br>\n2) a good way to break up the dataset into batches and perform oversampling on each batch?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2864400,
          "author_name": "AC",
          "author_url": "",
          "post_date": "2024-06-10T06:21:11.047000",
          "content": "<p>For me, I tried using SMOTE on a chunk of dataset and it's giving me OOM error. Still looking for the way to deal with it. </p>",
          "votes": 1,
          "replies": [
            {
              "id": 2893689,
              "author_name": "MARVINMENG",
              "author_url": "",
              "post_date": "2024-06-28T02:31:35.137000",
              "content": "<p>In regards to this, I split the dataset into binds == 0 and binds == 1, and then I took binds == 1 set and randomly sampled it to a size that was manageable, and then I used SMOTE on that set and saved it as a data shard. I repeated this by continuing to take different samples from binds == 1 and used SMOTE and saved it as different shards.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2786172,
      "author_name": "Ariel Ebersberger",
      "author_url": "",
      "post_date": "2024-05-01T07:42:34.390000",
      "content": "<p>Take a look at the <a href=\"https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest\" target=\"_blank\">tutorial notebook</a>. They sample a 1:1 ratio (with an upper limit of 30k samples for each to reduce training times). But it will, most likely, be a well-kept secret on how to sample it best. At least at the start of the competition.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2780567,
      "author_name": "Soumitri Kadambi",
      "author_url": "",
      "post_date": "2024-04-28T09:56:17.293000",
      "content": "<p>Data imbalance needs to be addressed to ensure that the model has even data distribution to learn from and use. SMOTE is one of the techniques to address data imbalance</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2780390,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-04-28T07:01:20.500000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2780281": "After looking into the data, the binds column is highly imbalanced:\n**0:** 293656924 samples\n**1:** 1589906 samples\nconstituting to ~ 99.46% and ~ 0.53850062% of the total samples respectively. \nI am opening this thread to learn how everybody is tackling data imbalance at this level. All the suggestions are welcome! \n\n\n\n",
    "2780394": "you should start to do experiments.\ne.g. just use the simplest xgboot + fingerprint or chemberta (3-layer trasnformer),\ntry with different pos to neg ratio in sampling your batch.\n\nmaybe there is no difference? who's know?\n\n---\nin datascience, decision is based on data.\nno experiments results, no plan\n\nexample of experiments:\n1. Tuning gradient boosting for imbalanced bioassay modelling with custom loss functions\nhttps://github.com/dahvida/gradient_boosting_CLF\n\n2.https://www.tensorflow.org/tutorials/structured_data/imbalanced_data",
    "2863562": "Trying to deal with the label imbalance by using SMOTE. I'm reading in the training data as a pyspark dataframe, but I can't find a way to apply SMOTE directly on it. It's also too big to convert to a pandas dataframe (my searches are telling me to convert it to a pandas dataframe to apply smote), is there a way to either:\n1) apply SMOTE on the full dataset?\n2) a good way to break up the dataset into batches and perform oversampling on each batch?",
    "2786172": "Take a look at the [tutorial notebook](https://www.kaggle.com/code/andrewdblevins/leash-tutorial-ecfps-and-random-forest). They sample a 1:1 ratio (with an upper limit of 30k samples for each to reduce training times). But it will, most likely, be a well-kept secret on how to sample it best. At least at the start of the competition.",
    "2780567": "Data imbalance needs to be addressed to ensure that the model has even data distribution to learn from and use. SMOTE is one of the techniques to address data imbalance",
    "2780390": ""
  }
}