{
  "id": 72205,
  "title": "Imbalanced data problem needs to be solved",
  "url": "/competitions/quora-insincere-questions-classification/discussion/72205",
  "author_name": "",
  "post_date": "2018-11-21T08:54:37.588451700Z",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>This competition is based on whether the question is insincere or not. You must find that this dataset is really imbalanced, positive and negative ratio is over 10. How you guys struggle with this problem?\nNow that we all use Deep learning model, but deep model inherent is not good at distinguishing imbalanced datasets. \nThere are some solutions: Downsampling(random get positive data); Oversampling(some guys use SOMTE, my thought is to use GAN).\nHow you solved this problem? Any good idea?</p>",
  "messages": [
    {
      "id": "425183",
      "postDate": "11/21/2018 08:54:37",
      "content": "<p>This competition is based on whether the question is insincere or not. You must find that this dataset is really imbalanced, positive and negative ratio is over 10. How you guys struggle with this problem?\nNow that we all use Deep learning model, but deep model inherent is not good at distinguishing imbalanced datasets. \nThere are some solutions: Downsampling(random get positive data); Oversampling(some guys use SOMTE, my thought is to use GAN).\nHow you solved this problem? Any good idea?</p>",
      "rawMarkdown": "This competition is based on whether the question is insincere or not. You must find that this dataset is really imbalanced, positive and negative ratio is over 10. How you guys struggle with this problem?\nNow that we all use Deep learning model, but deep model inherent is not good at distinguishing imbalanced datasets. \nThere are some solutions: Downsampling(random get positive data); Oversampling(some guys use SOMTE, my thought is to use GAN).\nHow you solved this problem? Any good idea?",
      "votes": null
    },
    {
      "id": "425218",
      "postDate": "11/21/2018 09:41:06",
      "content": "<p>I also come up with a idea for using class_weight during training to give different weights for imbalanced class(I will try, If get better result, I will tell here)</p>",
      "rawMarkdown": "I also come up with a idea for using class_weight during training to give different weights for imbalanced class(I will try, If get better result, I will tell here)",
      "votes": null
    },
    {
      "id": "425233",
      "postDate": "11/21/2018 10:10:47",
      "content": "<p>Divide weights has no effect there</p>",
      "rawMarkdown": "Divide weights has no effect there",
      "votes": null
    },
    {
      "id": "425452",
      "postDate": "11/21/2018 16:14:56",
      "content": "<p>My idea is to check some winning solutions from Toxic since that competition is similar. There are many things we can try but some of them tend to be more useful. </p>\n\n<ul>\n<li>1st : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557</a></li>\n<li>2nd : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612</a></li>\n<li>3rd (single model): <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644</a></li>\n<li>3rd : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52762\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52762</a></li>\n<li>5th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630</a></li>\n<li>12th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52702\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52702</a></li>\n<li>15th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52563\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52563</a></li>\n<li>25th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52647\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52647</a></li>\n<li>27th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52719\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52719</a></li>\n<li>33rd : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52666\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52666</a></li>\n<li>34th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52645\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52645</a></li>\n</ul>",
      "rawMarkdown": "My idea is to check some winning solutions from Toxic since that competition is similar. There are many things we can try but some of them tend to be more useful. \n\n - 1st : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\n - 2nd : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\n - 3rd (single model): https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644\n - 3rd : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52762\n - 5th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\n - 12th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52702\n - 15th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52563\n - 25th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52647\n - 27th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52719\n - 33rd : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52666\n - 34th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52645",
      "votes": null
    },
    {
      "id": "425469",
      "postDate": "11/21/2018 16:46:09",
      "content": "<p>Hi Lu,\nI give each question an \"insincerity score\".\nDuring classification phase I just set a cutoff at the level where around 6.2% of the results will be classified as 1.\nNot yet anything during learning phase.</p>",
      "rawMarkdown": "Hi Lu,\nI give each question an \"insincerity score\".\nDuring classification phase I just set a cutoff at the level where around 6.2% of the results will be classified as 1.\nNot yet anything during learning phase.",
      "votes": null
    },
    {
      "id": "425690",
      "postDate": "11/22/2018 01:25:39",
      "content": "<p>Amazing, I will check all of this kernels you provided to find out what a good way to solve this problem.\nThank you so much for this materials! Thanks!</p>",
      "rawMarkdown": "Amazing, I will check all of this kernels you provided to find out what a good way to solve this problem.\nThank you so much for this materials! Thanks!",
      "votes": null
    },
    {
      "id": "425693",
      "postDate": "11/22/2018 01:30:15",
      "content": "<p>So how you decide to give each question a 'insincerity score' which means that is a weight for each sample, right? Did you give a Gaussian weights or Random for that?</p>",
      "rawMarkdown": "So how you decide to give each question a 'insincerity score' which means that is a weight for each sample, right? Did you give a Gaussian weights or Random for that?",
      "votes": null
    },
    {
      "id": "425696",
      "postDate": "11/22/2018 01:31:27",
      "content": "<p>Yes, it is. I have tried with sklearn 'balanced' class weights, didn't give any improvement.</p>",
      "rawMarkdown": "Yes, it is. I have tried with sklearn 'balanced' class weights, didn't give any improvement.",
      "votes": null
    },
    {
      "id": "425801",
      "postDate": "11/22/2018 06:19:33",
      "content": "<p>I rewrote the cross entropy function: r * y * log(pred_y) + (1 - r) * (1 - y) * log(1 - pred_y). It doesn't work.</p>",
      "rawMarkdown": "I rewrote the cross entropy function: r * y * log(pred_y) + (1 - r) * (1 - y) * log(1 - pred_y). It doesn't work.",
      "votes": null
    },
    {
      "id": "425807",
      "postDate": "11/22/2018 06:32:17",
      "content": "<p>You have tried reweight class another approach for changing objective function, I have also tested with class_weight, either hasn't give any improvement. So for this problem, imbalanced problem doesn't harm our prediction? No need to worry about this imbalanced problem? Or any other good way to solve it?</p>",
      "rawMarkdown": "You have tried reweight class another approach for changing objective function, I have also tested with class_weight, either hasn't give any improvement. So for this problem, imbalanced problem doesn't harm our prediction? No need to worry about this imbalanced problem? Or any other good way to solve it?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 425218,
      "author_name": "manrunning",
      "author_url": "",
      "post_date": "11/21/2018 09:41:06",
      "content": "<p>I also come up with a idea for using class_weight during training to give different weights for imbalanced class(I will try, If get better result, I will tell here)</p>",
      "votes": null,
      "replies": [
        {
          "id": 425233,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "11/21/2018 10:10:47",
          "content": "<p>Divide weights has no effect there</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 425696,
          "author_name": "manrunning",
          "author_url": "",
          "post_date": "11/22/2018 01:31:27",
          "content": "<p>Yes, it is. I have tried with sklearn 'balanced' class weights, didn't give any improvement.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 425801,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "11/22/2018 06:19:33",
          "content": "<p>I rewrote the cross entropy function: r * y * log(pred_y) + (1 - r) * (1 - y) * log(1 - pred_y). It doesn't work.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 425807,
          "author_name": "manrunning",
          "author_url": "",
          "post_date": "11/22/2018 06:32:17",
          "content": "<p>You have tried reweight class another approach for changing objective function, I have also tested with class_weight, either hasn't give any improvement. So for this problem, imbalanced problem doesn't harm our prediction? No need to worry about this imbalanced problem? Or any other good way to solve it?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 425452,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/21/2018 16:14:56",
      "content": "<p>My idea is to check some winning solutions from Toxic since that competition is similar. There are many things we can try but some of them tend to be more useful. </p>\n\n<ul>\n<li>1st : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557</a></li>\n<li>2nd : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612</a></li>\n<li>3rd (single model): <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644</a></li>\n<li>3rd : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52762\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52762</a></li>\n<li>5th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630</a></li>\n<li>12th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52702\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52702</a></li>\n<li>15th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52563\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52563</a></li>\n<li>25th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52647\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52647</a></li>\n<li>27th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52719\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52719</a></li>\n<li>33rd : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52666\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52666</a></li>\n<li>34th : <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52645\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52645</a></li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 425690,
          "author_name": "manrunning",
          "author_url": "",
          "post_date": "11/22/2018 01:25:39",
          "content": "<p>Amazing, I will check all of this kernels you provided to find out what a good way to solve this problem.\nThank you so much for this materials! Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 425469,
      "author_name": "akuropatwinski",
      "author_url": "",
      "post_date": "11/21/2018 16:46:09",
      "content": "<p>Hi Lu,\nI give each question an \"insincerity score\".\nDuring classification phase I just set a cutoff at the level where around 6.2% of the results will be classified as 1.\nNot yet anything during learning phase.</p>",
      "votes": null,
      "replies": [
        {
          "id": 425693,
          "author_name": "manrunning",
          "author_url": "",
          "post_date": "11/22/2018 01:30:15",
          "content": "<p>So how you decide to give each question a 'insincerity score' which means that is a weight for each sample, right? Did you give a Gaussian weights or Random for that?</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "425183": "This competition is based on whether the question is insincere or not. You must find that this dataset is really imbalanced, positive and negative ratio is over 10. How you guys struggle with this problem?\nNow that we all use Deep learning model, but deep model inherent is not good at distinguishing imbalanced datasets. \nThere are some solutions: Downsampling(random get positive data); Oversampling(some guys use SOMTE, my thought is to use GAN).\nHow you solved this problem? Any good idea?",
    "425218": "I also come up with a idea for using class_weight during training to give different weights for imbalanced class(I will try, If get better result, I will tell here)",
    "425233": "Divide weights has no effect there",
    "425452": "My idea is to check some winning solutions from Toxic since that competition is similar. There are many things we can try but some of them tend to be more useful. \n\n - 1st : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52557\n - 2nd : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52612\n - 3rd (single model): https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52644\n - 3rd : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52762\n - 5th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52630\n - 12th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52702\n - 15th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52563\n - 25th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52647\n - 27th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52719\n - 33rd : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52666\n - 34th : https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/52645",
    "425469": "Hi Lu,\nI give each question an \"insincerity score\".\nDuring classification phase I just set a cutoff at the level where around 6.2% of the results will be classified as 1.\nNot yet anything during learning phase.",
    "425690": "Amazing, I will check all of this kernels you provided to find out what a good way to solve this problem.\nThank you so much for this materials! Thanks!",
    "425693": "So how you decide to give each question a 'insincerity score' which means that is a weight for each sample, right? Did you give a Gaussian weights or Random for that?",
    "425696": "Yes, it is. I have tried with sklearn 'balanced' class weights, didn't give any improvement.",
    "425801": "I rewrote the cross entropy function: r * y * log(pred_y) + (1 - r) * (1 - y) * log(1 - pred_y). It doesn't work.",
    "425807": "You have tried reweight class another approach for changing objective function, I have also tested with class_weight, either hasn't give any improvement. So for this problem, imbalanced problem doesn't harm our prediction? No need to worry about this imbalanced problem? Or any other good way to solve it?"
  },
  "source": "meta"
}