{
  "id": 71486,
  "title": "Should balance the data?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71486",
  "author_name": "",
  "post_date": "2018-11-14T03:43:43.541484300Z",
  "votes": null,
  "comment_count": 2,
  "views": 0,
  "content": "<p>I'm confused about whether use the balancing strategy. a balanced model have a better performance when the data distribution changed, but there is a strong possibility that Public Test data have a same distribution as train.  Will the Priavte Test data have the same distribution?</p>",
  "messages": [
    {
      "id": "420735",
      "postDate": "11/14/2018 03:43:43",
      "content": "<p>I'm confused about whether use the balancing strategy. a balanced model have a better performance when the data distribution changed, but there is a strong possibility that Public Test data have a same distribution as train.  Will the Priavte Test data have the same distribution?</p>",
      "rawMarkdown": "I'm confused about whether use the balancing strategy. a balanced model have a better performance when the data distribution changed, but there is a strong possibility that Public Test data have a same distribution as train.  Will the Priavte Test data have the same distribution?",
      "votes": null
    },
    {
      "id": "422595",
      "postDate": "11/16/2018 13:09:46",
      "content": "<p>I believe that it is a good idea to oversample a bit. However this is not an easy task in NLP. \nI tried SMOTE (<a href=\"https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\">https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote</a>), does not work very well.</p>\n\n<p>Undersampling is generally a bad idea in Kaggles as you want to train on as much data as possible.</p>",
      "rawMarkdown": "I believe that it is a good idea to oversample a bit. However this is not an easy task in NLP. \nI tried SMOTE (https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote), does not work very well.\n\nUndersampling is generally a bad idea in Kaggles as you want to train on as much data as possible.",
      "votes": null
    },
    {
      "id": "425131",
      "postDate": "11/21/2018 07:29:07",
      "content": "<p>Thanks for your replay. Oversample is a good method to overcome the imbalance situation, but I prefer to weight the gradient. I think they are the same thing. By now, my tries on balancing data all failed with lower plb score. I consider the balancing strategy is less help in this competition</p>",
      "rawMarkdown": "Thanks for your replay. Oversample is a good method to overcome the imbalance situation, but I prefer to weight the gradient. I think they are the same thing. By now, my tries on balancing data all failed with lower plb score. I consider the balancing strategy is less help in this competition",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 422595,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "11/16/2018 13:09:46",
      "content": "<p>I believe that it is a good idea to oversample a bit. However this is not an easy task in NLP. \nI tried SMOTE (<a href=\"https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\">https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote</a>), does not work very well.</p>\n\n<p>Undersampling is generally a bad idea in Kaggles as you want to train on as much data as possible.</p>",
      "votes": null,
      "replies": [
        {
          "id": 425131,
          "author_name": "ziliwang",
          "author_url": "",
          "post_date": "11/21/2018 07:29:07",
          "content": "<p>Thanks for your replay. Oversample is a good method to overcome the imbalance situation, but I prefer to weight the gradient. I think they are the same thing. By now, my tries on balancing data all failed with lower plb score. I consider the balancing strategy is less help in this competition</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "420735": "I'm confused about whether use the balancing strategy. a balanced model have a better performance when the data distribution changed, but there is a strong possibility that Public Test data have a same distribution as train.  Will the Priavte Test data have the same distribution?",
    "422595": "I believe that it is a good idea to oversample a bit. However this is not an easy task in NLP. \nI tried SMOTE (https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote), does not work very well.\n\nUndersampling is generally a bad idea in Kaggles as you want to train on as much data as possible.",
    "425131": "Thanks for your replay. Oversample is a good method to overcome the imbalance situation, but I prefer to weight the gradient. I think they are the same thing. By now, my tries on balancing data all failed with lower plb score. I consider the balancing strategy is less help in this competition"
  },
  "source": "meta"
}