{
  "id": 71788,
  "title": "Augmenting Text Data",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71788",
  "author_name": "",
  "post_date": "2018-11-16T14:17:31.687472200Z",
  "votes": 9,
  "comment_count": 10,
  "views": 0,
  "content": "<p>So I tried some data augmentation tricks:</p>\n\n<ul>\n<li>Oversampling with  SMOTE : <a href=\"https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\">https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote</a>  (definitely not the best idea)</li>\n<li>Replacing words with synonyms : <a href=\"https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\">https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation</a></li>\n</ul>\n\n<p>Has anybody tried something else (working/not working) ?  </p>",
  "messages": [
    {
      "id": "422625",
      "postDate": "11/16/2018 14:17:31",
      "content": "<p>So I tried some data augmentation tricks:</p>\n\n<ul>\n<li>Oversampling with  SMOTE : <a href=\"https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote\">https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote</a>  (definitely not the best idea)</li>\n<li>Replacing words with synonyms : <a href=\"https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\">https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation</a></li>\n</ul>\n\n<p>Has anybody tried something else (working/not working) ?  </p>",
      "rawMarkdown": "So I tried some data augmentation tricks:\n\n - Oversampling with  SMOTE : https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote  (definitely not the best idea)\n - Replacing words with synonyms : https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\n\nHas anybody tried something else (working/not working) ?",
      "votes": null
    },
    {
      "id": "423208",
      "postDate": "11/17/2018 17:56:01",
      "content": "<p>Embedding for Augmenting...</p>",
      "rawMarkdown": "Embedding for Augmenting...",
      "votes": null
    },
    {
      "id": "423255",
      "postDate": "11/17/2018 19:35:47",
      "content": "<p>i planned to do the \"replacing words with synonyms\" just for trying after i 've finished with other stuffs :-)</p>",
      "rawMarkdown": "i planned to do the \"replacing words with synonyms\" just for trying after i 've finished with other stuffs :-)",
      "votes": null
    },
    {
      "id": "425175",
      "postDate": "11/21/2018 08:44:51",
      "content": "<p>Maybe use GAN to generate more negative data, but this will also add some random noise to original datasets. But there is two key point: one is how well the GAN data generation; two is how many generated data should be added.\nIf I have time, I will also try this method.</p>",
      "rawMarkdown": "Maybe use GAN to generate more negative data, but this will also add some random noise to original datasets. But there is two key point: one is how well the GAN data generation; two is how many generated data should be added.\nIf I have time, I will also try this method.",
      "votes": null
    },
    {
      "id": "425199",
      "postDate": "11/21/2018 09:18:19",
      "content": "<p>Could work, but as the maximum runtime for the kernel is 2 hours, you won't have anymore time to train your actual classifier.</p>",
      "rawMarkdown": "Could work, but as the maximum runtime for the kernel is 2 hours, you won't have anymore time to train your actual classifier.",
      "votes": null
    },
    {
      "id": "425216",
      "postDate": "11/21/2018 09:39:14",
      "content": "<p>That's right! To train a GAN also need much time to converge. We have to do our best to do much more based on this time limited! Hard problem!</p>",
      "rawMarkdown": "That's right! To train a GAN also need much time to converge. We have to do our best to do much more based on this time limited! Hard problem!",
      "votes": null
    },
    {
      "id": "427642",
      "postDate": "11/25/2018 22:44:46",
      "content": "<p>Randomly replace part of the words with their nearest neighbor in embedding space. Takes 1/2 day to run: <a href=\"https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation\">https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation</a></p>",
      "rawMarkdown": "Randomly replace part of the words with their nearest neighbor in embedding space. Takes 1/2 day to run: https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation",
      "votes": null
    },
    {
      "id": "427646",
      "postDate": "11/25/2018 22:53:16",
      "content": "<p>Interesting, I'll check that out. Although the 2 hours limitation time + no external data makes data augmentation really hard.</p>",
      "rawMarkdown": "Interesting, I'll check that out. Although the 2 hours limitation time + no external data makes data augmentation really hard.",
      "votes": null
    },
    {
      "id": "427648",
      "postDate": "11/25/2018 22:55:57",
      "content": "<p>I think your second approach is more doable. </p>",
      "rawMarkdown": "I think your second approach is more doable.",
      "votes": null
    },
    {
      "id": "427880",
      "postDate": "11/26/2018 10:22:34",
      "content": "<p>Yep, but I did not get satisfying results so far... I'll work on that</p>",
      "rawMarkdown": "Yep, but I did not get satisfying results so far... I'll work on that",
      "votes": null
    },
    {
      "id": "572698",
      "postDate": "07/11/2019 09:26:02",
      "content": "<p>You can use back translation for augmenting Text. Generates meaningful sentences and works well.</p>",
      "rawMarkdown": "You can use back translation for augmenting Text. Generates meaningful sentences and works well.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 423208,
      "author_name": "arunkumarramanan",
      "author_url": "",
      "post_date": "11/17/2018 17:56:01",
      "content": "<p>Embedding for Augmenting...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 423255,
      "author_name": "dkmerona",
      "author_url": "",
      "post_date": "11/17/2018 19:35:47",
      "content": "<p>i planned to do the \"replacing words with synonyms\" just for trying after i 've finished with other stuffs :-)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 425175,
      "author_name": "manrunning",
      "author_url": "",
      "post_date": "11/21/2018 08:44:51",
      "content": "<p>Maybe use GAN to generate more negative data, but this will also add some random noise to original datasets. But there is two key point: one is how well the GAN data generation; two is how many generated data should be added.\nIf I have time, I will also try this method.</p>",
      "votes": null,
      "replies": [
        {
          "id": 425199,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "11/21/2018 09:18:19",
          "content": "<p>Could work, but as the maximum runtime for the kernel is 2 hours, you won't have anymore time to train your actual classifier.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 425216,
          "author_name": "manrunning",
          "author_url": "",
          "post_date": "11/21/2018 09:39:14",
          "content": "<p>That's right! To train a GAN also need much time to converge. We have to do our best to do much more based on this time limited! Hard problem!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 427642,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "11/25/2018 22:44:46",
      "content": "<p>Randomly replace part of the words with their nearest neighbor in embedding space. Takes 1/2 day to run: <a href=\"https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation\">https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation</a></p>",
      "votes": null,
      "replies": [
        {
          "id": 427646,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "11/25/2018 22:53:16",
          "content": "<p>Interesting, I'll check that out. Although the 2 hours limitation time + no external data makes data augmentation really hard.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 427648,
          "author_name": "shujian",
          "author_url": "",
          "post_date": "11/25/2018 22:55:57",
          "content": "<p>I think your second approach is more doable. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 427880,
          "author_name": "theoviel",
          "author_url": "",
          "post_date": "11/26/2018 10:22:34",
          "content": "<p>Yep, but I did not get satisfying results so far... I'll work on that</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 572698,
      "author_name": "chikubee",
      "author_url": "",
      "post_date": "07/11/2019 09:26:02",
      "content": "<p>You can use back translation for augmenting Text. Generates meaningful sentences and works well.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "422625": "So I tried some data augmentation tricks:\n\n - Oversampling with  SMOTE : https://www.kaggle.com/theoviel/dealing-with-class-imbalance-with-smote  (definitely not the best idea)\n - Replacing words with synonyms : https://www.kaggle.com/theoviel/using-word-embeddings-for-data-augmentation\n\nHas anybody tried something else (working/not working) ?",
    "423208": "Embedding for Augmenting...",
    "423255": "i planned to do the \"replacing words with synonyms\" just for trying after i 've finished with other stuffs :-)",
    "425175": "Maybe use GAN to generate more negative data, but this will also add some random noise to original datasets. But there is two key point: one is how well the GAN data generation; two is how many generated data should be added.\nIf I have time, I will also try this method.",
    "425199": "Could work, but as the maximum runtime for the kernel is 2 hours, you won't have anymore time to train your actual classifier.",
    "425216": "That's right! To train a GAN also need much time to converge. We have to do our best to do much more based on this time limited! Hard problem!",
    "427642": "Randomly replace part of the words with their nearest neighbor in embedding space. Takes 1/2 day to run: https://www.kaggle.com/shujian/fake-some-positive-data-data-augmentation",
    "427646": "Interesting, I'll check that out. Although the 2 hours limitation time + no external data makes data augmentation really hard.",
    "427648": "I think your second approach is more doable.",
    "427880": "Yep, but I did not get satisfying results so far... I'll work on that",
    "572698": "You can use back translation for augmenting Text. Generates meaningful sentences and works well."
  },
  "source": "meta"
}