{
  "id": 77528,
  "title": "Has anyone tried balancing the data?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/77528",
  "author_name": "",
  "post_date": "2019-01-13T22:41:17.448133100Z",
  "votes": 4,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I'm just starting here,\nI think that the questions that are insincere are just a few compared to the ones that aren't. which is pretty normal.\nMy question is. if I have to balance the dataset before training?</p>",
  "messages": [
    {
      "id": "455440",
      "postDate": "01/13/2019 22:41:17",
      "content": "<p>I'm just starting here,\nI think that the questions that are insincere are just a few compared to the ones that aren't. which is pretty normal.\nMy question is. if I have to balance the dataset before training?</p>",
      "rawMarkdown": "I'm just starting here,\nI think that the questions that are insincere are just a few compared to the ones that aren't. which is pretty normal.\nMy question is. if I have to balance the dataset before training?",
      "votes": null
    },
    {
      "id": "455448",
      "postDate": "01/13/2019 23:28:08",
      "content": "<p>Yes, I heard that it's good to balence the data, so your models don't do overfitting.</p>",
      "rawMarkdown": "Yes, I heard that it's good to balence the data, so your models don't do overfitting.",
      "votes": null
    },
    {
      "id": "456030",
      "postDate": "01/15/2019 02:02:24",
      "content": "<p>I find that using an embedding -&gt; RNN model saturates very quickly. I got much better results just weighting my loss function and using all the data.</p>",
      "rawMarkdown": "I find that using an embedding -&gt; RNN model saturates very quickly. I got much better results just weighting my loss function and using all the data.",
      "votes": null
    },
    {
      "id": "456226",
      "postDate": "01/15/2019 11:32:06",
      "content": "<p>The big problem with balancing this dataset is you end up throwing away lots of data - if you can find a way to keep using that data I'd love to hear it</p>\n\n<p>Another idea I tried early on was to use keras' <code>class_weight='auto'</code> but it didn't yield great results, could be worth a retry though</p>\n\n<p>Another thing you could do is try a custom loss function (maybe based around F1 score given that's the metric we're trying to optimise?)</p>",
      "rawMarkdown": "The big problem with balancing this dataset is you end up throwing away lots of data - if you can find a way to keep using that data I'd love to hear it\n\nAnother idea I tried early on was to use keras' `class_weight='auto'` but it didn't yield great results, could be worth a retry though\n\nAnother thing you could do is try a custom loss function (maybe based around F1 score given that's the metric we're trying to optimise?)",
      "votes": null
    },
    {
      "id": "456416",
      "postDate": "01/15/2019 19:15:09",
      "content": "<p>Thank you so much, I will weighting my loss</p>",
      "rawMarkdown": "Thank you so much, I will weighting my loss",
      "votes": null
    },
    {
      "id": "456500",
      "postDate": "01/15/2019 23:43:24",
      "content": "<p>There are ways to use both of your ideas, but you have to get around the sparse gradient if you're using a RNN model from the kernels.</p>",
      "rawMarkdown": "There are ways to use both of your ideas, but you have to get around the sparse gradient if you're using a RNN model from the kernels.",
      "votes": null
    },
    {
      "id": "456972",
      "postDate": "01/16/2019 20:16:20",
      "content": "<p>I actually had some improvements using the <code>class_weight</code> function of Keras. I set the weights using the <code>class_weight</code> function from <code>sklearn.utils</code>, but I think the <code>auto</code> setting in Keras does kind of the same. Anyway, might help.</p>",
      "rawMarkdown": "I actually had some improvements using the `class_weight` function of Keras. I set the weights using the `class_weight` function from `sklearn.utils`, but I think the `auto` setting in Keras does kind of the same. Anyway, might help.",
      "votes": null
    },
    {
      "id": "457322",
      "postDate": "01/17/2019 08:18:47",
      "content": "<p>yes. equal numbers of sincere and insincere questions, local cv is around 0.8 but lb is around 0.66. However, my model takes the same amount of time to train for each epoch.</p>\n\n<p>Should I trust my cv? I don't know yet.</p>",
      "rawMarkdown": "yes. equal numbers of sincere and insincere questions, local cv is around 0.8 but lb is around 0.66. However, my model takes the same amount of time to train for each epoch.\n\nShould I trust my cv? I don't know yet.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 455448,
      "author_name": "jmexpert",
      "author_url": "",
      "post_date": "01/13/2019 23:28:08",
      "content": "<p>Yes, I heard that it's good to balence the data, so your models don't do overfitting.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 456030,
      "author_name": "stevenknguyen",
      "author_url": "",
      "post_date": "01/15/2019 02:02:24",
      "content": "<p>I find that using an embedding -&gt; RNN model saturates very quickly. I got much better results just weighting my loss function and using all the data.</p>",
      "votes": null,
      "replies": [
        {
          "id": 456416,
          "author_name": "jmourad100",
          "author_url": "",
          "post_date": "01/15/2019 19:15:09",
          "content": "<p>Thank you so much, I will weighting my loss</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 456226,
      "author_name": "hamishdickson",
      "author_url": "",
      "post_date": "01/15/2019 11:32:06",
      "content": "<p>The big problem with balancing this dataset is you end up throwing away lots of data - if you can find a way to keep using that data I'd love to hear it</p>\n\n<p>Another idea I tried early on was to use keras' <code>class_weight='auto'</code> but it didn't yield great results, could be worth a retry though</p>\n\n<p>Another thing you could do is try a custom loss function (maybe based around F1 score given that's the metric we're trying to optimise?)</p>",
      "votes": null,
      "replies": [
        {
          "id": 456500,
          "author_name": "stevenknguyen",
          "author_url": "",
          "post_date": "01/15/2019 23:43:24",
          "content": "<p>There are ways to use both of your ideas, but you have to get around the sparse gradient if you're using a RNN model from the kernels.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 456972,
          "author_name": "timothylucas",
          "author_url": "",
          "post_date": "01/16/2019 20:16:20",
          "content": "<p>I actually had some improvements using the <code>class_weight</code> function of Keras. I set the weights using the <code>class_weight</code> function from <code>sklearn.utils</code>, but I think the <code>auto</code> setting in Keras does kind of the same. Anyway, might help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 457322,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "01/17/2019 08:18:47",
      "content": "<p>yes. equal numbers of sincere and insincere questions, local cv is around 0.8 but lb is around 0.66. However, my model takes the same amount of time to train for each epoch.</p>\n\n<p>Should I trust my cv? I don't know yet.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "455440": "I'm just starting here,\nI think that the questions that are insincere are just a few compared to the ones that aren't. which is pretty normal.\nMy question is. if I have to balance the dataset before training?",
    "455448": "Yes, I heard that it's good to balence the data, so your models don't do overfitting.",
    "456030": "I find that using an embedding -&gt; RNN model saturates very quickly. I got much better results just weighting my loss function and using all the data.",
    "456226": "The big problem with balancing this dataset is you end up throwing away lots of data - if you can find a way to keep using that data I'd love to hear it\n\nAnother idea I tried early on was to use keras' `class_weight='auto'` but it didn't yield great results, could be worth a retry though\n\nAnother thing you could do is try a custom loss function (maybe based around F1 score given that's the metric we're trying to optimise?)",
    "456416": "Thank you so much, I will weighting my loss",
    "456500": "There are ways to use both of your ideas, but you have to get around the sparse gradient if you're using a RNN model from the kernels.",
    "456972": "I actually had some improvements using the `class_weight` function of Keras. I set the weights using the `class_weight` function from `sklearn.utils`, but I think the `auto` setting in Keras does kind of the same. Anyway, might help.",
    "457322": "yes. equal numbers of sincere and insincere questions, local cv is around 0.8 but lb is around 0.66. However, my model takes the same amount of time to train for each epoch.\n\nShould I trust my cv? I don't know yet."
  },
  "source": "meta"
}