{
  "id": 511007,
  "title": "Oversampling for imbalanced data question",
  "url": "/competitions/leash-BELKA/discussion/511007",
  "author_name": "MARVINMENG",
  "post_date": "2024-06-08T16:43:58.181000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Trying to deal with the label imbalance (about .5% of the data has binds == 1) by using SMOTE. I'm reading in the training data as a pyspark dataframe and encoding the SMILE representations as vectors using word2vec, but I can't find a way to apply SMOTE directly on the pyspark dataframe. It's also too big to convert to a pandas dataframe, I'm curious how people are dealing with the imbalanced dataset and if they found a way to make oversampling work.</p>",
  "messages": [
    {
      "id": 2862252,
      "postDate": "2024-06-08T16:43:58.183Z",
      "content": "<p>Trying to deal with the label imbalance (about .5% of the data has binds == 1) by using SMOTE. I'm reading in the training data as a pyspark dataframe and encoding the SMILE representations as vectors using word2vec, but I can't find a way to apply SMOTE directly on the pyspark dataframe. It's also too big to convert to a pandas dataframe, I'm curious how people are dealing with the imbalanced dataset and if they found a way to make oversampling work.</p>",
      "rawMarkdown": "Trying to deal with the label imbalance (about .5% of the data has binds == 1) by using SMOTE. I'm reading in the training data as a pyspark dataframe and encoding the SMILE representations as vectors using word2vec, but I can't find a way to apply SMOTE directly on the pyspark dataframe. It's also too big to convert to a pandas dataframe, I'm curious how people are dealing with the imbalanced dataset and if they found a way to make oversampling work."
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2862252": "Trying to deal with the label imbalance (about .5% of the data has binds == 1) by using SMOTE. I'm reading in the training data as a pyspark dataframe and encoding the SMILE representations as vectors using word2vec, but I can't find a way to apply SMOTE directly on the pyspark dataframe. It's also too big to convert to a pandas dataframe, I'm curious how people are dealing with the imbalanced dataset and if they found a way to make oversampling work."
  }
}