{
  "id": 78841,
  "title": "Ideas for word embeddings augmentation",
  "url": "/competitions/quora-insincere-questions-classification/discussion/78841",
  "author_name": "Dmytro Danevskyi",
  "post_date": "2019-01-28T10:45:54.104000",
  "votes": 56,
  "comment_count": 25,
  "views": 0,
  "content": "<p>Data augmentation is a very popular strategy for preventing overfitting in convolutional networks. It really makes a lot of sense to use flips/rotations/color shifts for such tasks as classification, object detection, pose estimation, etc., since we know that such transformations (unless used with very aggressive parameters) should not affect correct labels much. This simple strategy populates the data space, increases the model's ability to generalize, and also makes the model to be more robust to unusual lighting conditions, partial occlusion, etc.</p>\n\n<p>As for text data, which is in contrast to image data is essentially <em>discrete</em>, there is no obvious way to do the augmentation.</p>\n\n<p>However, there were some successful applications of data augmentation in raw data space. Some of them were already listed <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71083\">here</a> by @Dieter. One way is to replace some words with synonyms. Another (which was successfully applied by the winners of <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\">Toxic Comment Classification Challenge</a>) is to do cross-translation: using some pretrained machine translation model, translate a source sentence to another language, then translate back.</p>\n\n<p>What I want to point out is that there are some ways to do the augmentation in embedding space as well! Since word embeddings consist of real numbers rather than discrete tokens, it opens a lot of interesting ways to augment the data.  </p>\n\n<p>For instance, <a href=\"https://arxiv.org/abs/1804.08166\">this</a> paper studies the effects of gaussian/bernoulli dropout and adversarial examples generation and reports some improvement over the baseline.</p>\n\n<p>Other options include:</p>\n\n<ul>\n<li>Additive gaussian/other noise. Theoretically, if the perturbation doesn't change the list of nearest neighbors of a particular embedding much, a classifier can still be able to learn without any problems. </li>\n<li>If you are using more than one set of embeddings in your pipeline (say, glove+paragram) you may try to create a model that learns the relationship between the same words represented by different embeddings. Then you can randomly replace some embeddings with their versions converted from another embedding set.</li>\n<li>Learn a distribution of embeddings for a word instead of learning fixed vectors. Then, on each train iteration, you can use the learned distribution to sample new but valid versions of the word embeddings. This may be tricky to implement and also not really applicable to the current competition due to time constraints, but I definitely suggest you check the <a href=\"https://arxiv.org/abs/1711.11027\">paper</a> if you are interested.</li>\n</ul>\n\n<p>Let's discuss and happy kaggling!</p>",
  "messages": [
    {
      "id": 462470,
      "postDate": "2019-01-28T10:45:54.103Z",
      "content": "<p>Data augmentation is a very popular strategy for preventing overfitting in convolutional networks. It really makes a lot of sense to use flips/rotations/color shifts for such tasks as classification, object detection, pose estimation, etc., since we know that such transformations (unless used with very aggressive parameters) should not affect correct labels much. This simple strategy populates the data space, increases the model's ability to generalize, and also makes the model to be more robust to unusual lighting conditions, partial occlusion, etc.</p>\n\n<p>As for text data, which is in contrast to image data is essentially <em>discrete</em>, there is no obvious way to do the augmentation.</p>\n\n<p>However, there were some successful applications of data augmentation in raw data space. Some of them were already listed <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71083\">here</a> by @Dieter. One way is to replace some words with synonyms. Another (which was successfully applied by the winners of <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\">Toxic Comment Classification Challenge</a>) is to do cross-translation: using some pretrained machine translation model, translate a source sentence to another language, then translate back.</p>\n\n<p>What I want to point out is that there are some ways to do the augmentation in embedding space as well! Since word embeddings consist of real numbers rather than discrete tokens, it opens a lot of interesting ways to augment the data.  </p>\n\n<p>For instance, <a href=\"https://arxiv.org/abs/1804.08166\">this</a> paper studies the effects of gaussian/bernoulli dropout and adversarial examples generation and reports some improvement over the baseline.</p>\n\n<p>Other options include:</p>\n\n<ul>\n<li>Additive gaussian/other noise. Theoretically, if the perturbation doesn't change the list of nearest neighbors of a particular embedding much, a classifier can still be able to learn without any problems. </li>\n<li>If you are using more than one set of embeddings in your pipeline (say, glove+paragram) you may try to create a model that learns the relationship between the same words represented by different embeddings. Then you can randomly replace some embeddings with their versions converted from another embedding set.</li>\n<li>Learn a distribution of embeddings for a word instead of learning fixed vectors. Then, on each train iteration, you can use the learned distribution to sample new but valid versions of the word embeddings. This may be tricky to implement and also not really applicable to the current competition due to time constraints, but I definitely suggest you check the <a href=\"https://arxiv.org/abs/1711.11027\">paper</a> if you are interested.</li>\n</ul>\n\n<p>Let's discuss and happy kaggling!</p>",
      "rawMarkdown": "Data augmentation is a very popular strategy for preventing overfitting in convolutional networks. It really makes a lot of sense to use flips/rotations/color shifts for such tasks as classification, object detection, pose estimation, etc., since we know that such transformations (unless used with very aggressive parameters) should not affect correct labels much. This simple strategy populates the data space, increases the model's ability to generalize, and also makes the model to be more robust to unusual lighting conditions, partial occlusion, etc.\n\nAs for text data, which is in contrast to image data is essentially *discrete*, there is no obvious way to do the augmentation.\n\nHowever, there were some successful applications of data augmentation in raw data space. Some of them were already listed [here][1] by @Dieter. One way is to replace some words with synonyms. Another (which was successfully applied by the winners of [Toxic Comment Classification Challenge][2]) is to do cross-translation: using some pretrained machine translation model, translate a source sentence to another language, then translate back.\n\nWhat I want to point out is that there are some ways to do the augmentation in embedding space as well! Since word embeddings consist of real numbers rather than discrete tokens, it opens a lot of interesting ways to augment the data.  \n\nFor instance, [this][3] paper studies the effects of gaussian/bernoulli dropout and adversarial examples generation and reports some improvement over the baseline.\n\nOther options include:\n\n - Additive gaussian/other noise. Theoretically, if the perturbation doesn't change the list of nearest neighbors of a particular embedding much, a classifier can still be able to learn without any problems. \n - If you are using more than one set of embeddings in your pipeline (say, glove+paragram) you may try to create a model that learns the relationship between the same words represented by different embeddings. Then you can randomly replace some embeddings with their versions converted from another embedding set.\n - Learn a distribution of embeddings for a word instead of learning fixed vectors. Then, on each train iteration, you can use the learned distribution to sample new but valid versions of the word embeddings. This may be tricky to implement and also not really applicable to the current competition due to time constraints, but I definitely suggest you check the [paper][4] if you are interested.\n \nLet's discuss and happy kaggling!\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71083\n  [2]: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\n  [3]: https://arxiv.org/abs/1804.08166\n  [4]: https://arxiv.org/abs/1711.11027",
      "votes": 55
    },
    {
      "id": 462487,
      "postDate": "2019-01-28T11:08:33.150Z",
      "content": "<p>I added GaussianNoise with 0.1 from keras and my score improved up to 0.002-0.004</p>",
      "rawMarkdown": "I added GaussianNoise with 0.1 from keras and my score improved up to 0.002-0.004",
      "votes": 28,
      "replies": [
        {
          "id": 462797,
          "postDate": "2019-01-28T23:08:55.530Z",
          "content": "<p>May I ask where you added the noise? Right after embedding?</p>",
          "rawMarkdown": "May I ask where you added the noise? Right after embedding?"
        },
        {
          "id": 462893,
          "postDate": "2019-01-29T04:33:07.237Z",
          "content": "<p>You may add it if you engineered a layer (for example you concatenated both 1D Convolutional and Bi-Directional LSTM then you may either apply batch normalization or not, then you can apply GaussianNoise with your prefer value like 0.1</p>\n\n<p>example:\nconc = concatenate([conv1d_result, bi_lstm_result])(previous_layer)\nconc = BatchNormalization()(conc)\nconc = GaussianNoise(0.1)(conc)</p>\n\n<p>output = Dense(1, activation='sigmoid')(conc)</p>\n\n<p>You may experiment it with other setup.. Godd luck! :-)</p>",
          "rawMarkdown": "You may add it if you engineered a layer (for example you concatenated both 1D Convolutional and Bi-Directional LSTM then you may either apply batch normalization or not, then you can apply GaussianNoise with your prefer value like 0.1\n\nexample:\nconc = concatenate([conv1d_result, bi_lstm_result])(previous_layer)\nconc = BatchNormalization()(conc)\nconc = GaussianNoise(0.1)(conc)\n\noutput = Dense(1, activation='sigmoid')(conc)\n\nYou may experiment it with other setup.. Godd luck! :-)",
          "votes": 18
        },
        {
          "id": 463081,
          "postDate": "2019-01-29T11:22:24.950Z",
          "content": "<p>Adding small amount noise after embedding layer is indeed an interesting option. \nIsn't it like creating \"synonyms\" in the embedding space?</p>\n\n<p>However we should be cautious:\nkeras.layers.GaussianNoise(stddev) has no seed argument :-(  thus every train run different random realization of noise will be used and we probably get more variance in f1 scores across runs :-(</p>",
          "rawMarkdown": "Adding small amount noise after embedding layer is indeed an interesting option. \nIsn't it like creating \"synonyms\" in the embedding space?\n\nHowever we should be cautious:\nkeras.layers.GaussianNoise(stddev) has no seed argument :-(  thus every train run different random realization of noise will be used and we probably get more variance in f1 scores across runs :-(",
          "votes": 10,
          "isDeleted": true
        },
        {
          "id": 463326,
          "postDate": "2019-01-29T20:09:09.167Z",
          "content": "<p>I tried adding customized Gaussian noise layer after Embeddings. It hurts the performance, but I used it with embedding dropout. Maybe you have to choose only one of them, but I haven't given it a try yet.</p>",
          "rawMarkdown": "I tried adding customized Gaussian noise layer after Embeddings. It hurts the performance, but I used it with embedding dropout. Maybe you have to choose only one of them, but I haven't given it a try yet.",
          "votes": 3
        },
        {
          "id": 465572,
          "postDate": "2019-02-03T14:07:34.793Z",
          "content": "<p>Yes, in Dongxu and Zhichao's <a href=\"https://arxiv.org/abs/1804.08166\">paper</a>, the intuition is to add perturbation (to use Goodfellow's more accurate definition than \"noise\") to the embeddings. It is a data augmentation technique. So, pretty much as Annabelle is saying, after the embedding layer. I tried it and it didn't help.</p>",
          "rawMarkdown": "Yes, in Dongxu and Zhichao's [paper](https://arxiv.org/abs/1804.08166), the intuition is to add perturbation (to use Goodfellow's more accurate definition than \"noise\") to the embeddings. It is a data augmentation technique. So, pretty much as Annabelle is saying, after the embedding layer. I tried it and it didn't help.",
          "votes": 2
        }
      ]
    },
    {
      "id": 464304,
      "postDate": "2019-01-31T15:19:02.460Z",
      "content": "<p>I added this version of gaussian noise with a stddev of 0.1 after my embeddings and increased my local CV by 0.0006</p>\n\n<pre><code>class GaussianNoise(nn.Module):\n    def __init__(self, stddev):\n        super(GaussianNoise, self).__init__()\n\n        self.stddev = stddev\n\n    def forward(self, x):\n        noise = torch.empty_like(x)\n        noise.normal_(0, self.stddev)\n\n        return x + noise\n\nif self.training:\n    x = self.gaussian_noise(x)\n</code></pre>",
      "rawMarkdown": "I added this version of gaussian noise with a stddev of 0.1 after my embeddings and increased my local CV by 0.0006\n\n    class GaussianNoise(nn.Module):\n        def __init__(self, stddev):\n            super(GaussianNoise, self).__init__()\n        \n            self.stddev = stddev\n        \n        def forward(self, x):\n            noise = torch.empty_like(x)\n            noise.normal_(0, self.stddev)\n        \n            return x + noise\n\n    if self.training:\n        x = self.gaussian_noise(x)",
      "votes": 10,
      "replies": [
        {
          "id": 464612,
          "postDate": "2019-02-01T06:49:07.173Z",
          "content": "<p><a href=\"/bkkaggle\">@bkkaggle</a>, I got similar improvement in CV by adding noise to the embedding layer of my keras model but my LB score went down 3 points. Did you see significant increase in computation time with the addition of noise? Mine resulted in a  negligible increase of about 75 seconds to the total kernel time.</p>\n\n<p>I am new to pytorch, how would you incorporate your GaussianNoise class to a model similar to the public pytorch kernel? I want to give it a try.</p>",
          "rawMarkdown": "@bkkaggle, I got similar improvement in CV by adding noise to the embedding layer of my keras model but my LB score went down 3 points. Did you see significant increase in computation time with the addition of noise? Mine resulted in a  negligible increase of about 75 seconds to the total kernel time.\n\nI am new to pytorch, how would you incorporate your GaussianNoise class to a model similar to the public pytorch kernel? I want to give it a try.\n",
          "votes": 2
        }
      ]
    },
    {
      "id": 466869,
      "postDate": "2019-02-06T05:24:12.173Z",
      "content": "<p>I used synonyms and my CV went from 0.695 to 0.745. There is a catch though. The split may have the conjugate sentences in training and val. I only added 24k sentences.</p>",
      "rawMarkdown": "I used synonyms and my CV went from 0.695 to 0.745. There is a catch though. The split may have the conjugate sentences in training and val. I only added 24k sentences.",
      "votes": 1
    },
    {
      "id": 465657,
      "postDate": "2019-02-03T17:55:20.663Z",
      "content": "<p>Thank you for sharing. I will use some of them in the coming competition. Very interesting.</p>",
      "rawMarkdown": "Thank you for sharing. I will use some of them in the coming competition. Very interesting.",
      "votes": 1
    },
    {
      "id": 463339,
      "postDate": "2019-01-29T20:35:21.517Z",
      "content": "<p>What about using VAE for this task?</p>",
      "rawMarkdown": "What about using VAE for this task?",
      "votes": 1,
      "replies": [
        {
          "id": 463354,
          "postDate": "2019-01-29T21:06:20.800Z",
          "content": "<p>I believe it should be definitely possible. Basically, VAE is just a variant of how one could learn the distribution over embeddings (which I mentioned in the post). BTW, <a href=\"https://arxiv.org/abs/1711.11027\">this paper</a> uses a variant of a VAE in fact.</p>",
          "rawMarkdown": "I believe it should be definitely possible. Basically, VAE is just a variant of how one could learn the distribution over embeddings (which I mentioned in the post). BTW, [this paper][1] uses a variant of a VAE in fact.\n\n\n  [1]: https://arxiv.org/abs/1711.11027"
        },
        {
          "id": 463455,
          "postDate": "2019-01-30T02:38:07.620Z",
          "content": "<p>time limits.</p>",
          "rawMarkdown": "time limits.",
          "votes": 3
        },
        {
          "id": 463679,
          "postDate": "2019-01-30T12:15:21.007Z",
          "content": "<p>But VAE is not always promising. I have wrote about VAE quite a long time back. <a href=\"https://s4sarath.github.io/2016/11/23/variational_autoenocder_for_Natural_Language_Processing\">https://s4sarath.github.io/2016/11/23/variational_autoenocder_for_Natural_Language_Processing</a> </p>",
          "rawMarkdown": "But VAE is not always promising. I have wrote about VAE quite a long time back. https://s4sarath.github.io/2016/11/23/variational_autoenocder_for_Natural_Language_Processing ",
          "votes": 2
        }
      ]
    },
    {
      "id": 465716,
      "postDate": "2019-02-03T20:49:06.737Z",
      "content": "<p>Textblob library can do machine translation between languages, but it does it by using API to Google translator. As web connection is not allowed, that route is off. </p>",
      "rawMarkdown": "Textblob library can do machine translation between languages, but it does it by using API to Google translator. As web connection is not allowed, that route is off. ",
      "votes": 2
    },
    {
      "id": 463652,
      "postDate": "2019-01-30T11:24:34.787Z",
      "content": "<p>My CV has gone up by 0.005. I have inserted two gaussian layers just after embedding and also \n just after concatenation.</p>",
      "rawMarkdown": "My CV has gone up by 0.005. I have inserted two gaussian layers just after embedding and also \n just after concatenation.",
      "votes": 2,
      "replies": [
        {
          "id": 464160,
          "postDate": "2019-01-31T09:22:50.110Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 464237,
          "postDate": "2019-01-31T12:39:49.073Z",
          "content": "<p>Run time goes up from 6705 sec to 6935 sec. I set n_splits to 5  and n_epochs to 5.  I'm also using Pytorch.  </p>\n\n<p>I have inserted the gaussian layer as follows.</p>\n\n<p>///\nif self.training == True: \n   h_embedding = h_embedding + h_embedding.clone().normal_(0, 0.1) \n///</p>",
          "rawMarkdown": "Run time goes up from 6705 sec to 6935 sec. I set n_splits to 5  and n_epochs to 5.  I'm also using Pytorch.  \n\nI have inserted the gaussian layer as follows.\n\n///\nif self.training == True: \n   h_embedding = h_embedding + h_embedding.clone().normal_(0, 0.1) \n///",
          "votes": 4
        },
        {
          "id": 465714,
          "postDate": "2019-02-03T20:45:52.750Z",
          "content": "<p>You mean:</p>\n\n<p>hembedding = hembedding + hembedding.clone().normal_(0, 0.1)</p>\n\n<p>don't you? :-)</p>",
          "rawMarkdown": "You mean:\n\nhembedding = hembedding + hembedding.clone().normal_(0, 0.1)\n\ndon't you? :-)",
          "votes": 2
        },
        {
          "id": 466441,
          "postDate": "2019-02-05T11:49:18.113Z",
          "content": "<p>Yes.\nAnd I think it is better to disable the noise while prediction as the dropout.\nTherefore I've applied the noise during training only.</p>",
          "rawMarkdown": "Yes.\nAnd I think it is better to disable the noise while prediction as the dropout.\nTherefore I've applied the noise during training only.",
          "votes": 1
        }
      ]
    },
    {
      "id": 463600,
      "postDate": "2019-01-30T09:08:53.353Z",
      "content": "<p>Add GaussianNoise my LB went down to 0.686-0.688</p>",
      "rawMarkdown": "Add GaussianNoise my LB went down to 0.686-0.688",
      "votes": 2,
      "replies": [
        {
          "id": 463619,
          "postDate": "2019-01-30T09:51:46.333Z",
          "content": "<p>What about CV? ;)</p>",
          "rawMarkdown": "What about CV? ;)"
        },
        {
          "id": 463774,
          "postDate": "2019-01-30T15:45:49.130Z",
          "content": "<p>CV increased 0.003</p>",
          "rawMarkdown": "CV increased 0.003",
          "votes": 2
        },
        {
          "id": 463780,
          "postDate": "2019-01-30T15:53:34.253Z",
          "content": "<p>wow, so CV increased which would suggest that model generalizes better, but in LB score went down, so it looks like previous model was overfit to public LB set. If it happened in a small set I would suspect the seed used, but I don't think this is the case here?</p>",
          "rawMarkdown": "wow, so CV increased which would suggest that model generalizes better, but in LB score went down, so it looks like previous model was overfit to public LB set. If it happened in a small set I would suspect the seed used, but I don't think this is the case here?"
        }
      ]
    },
    {
      "id": 987979,
      "postDate": "2020-08-27T17:06:53.360Z",
      "content": "<p>Check out EDA<br>\n<a href=\"https://www.kaggle.com/swarajshinde/eda-data-augmentation-techniques-for-text-nlp\" target=\"_blank\">https://www.kaggle.com/swarajshinde/eda-data-augmentation-techniques-for-text-nlp</a></p>",
      "rawMarkdown": "Check out EDA\nhttps://www.kaggle.com/swarajshinde/eda-data-augmentation-techniques-for-text-nlp"
    }
  ],
  "comments": [
    {
      "id": 462487,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-01-28T11:08:33.150000",
      "content": "<p>I added GaussianNoise with 0.1 from keras and my score improved up to 0.002-0.004</p>",
      "votes": 28,
      "replies": [
        {
          "id": 462797,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-01-28T23:08:55.530000",
          "content": "<p>May I ask where you added the noise? Right after embedding?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 462893,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-29T04:33:07.237000",
          "content": "<p>You may add it if you engineered a layer (for example you concatenated both 1D Convolutional and Bi-Directional LSTM then you may either apply batch normalization or not, then you can apply GaussianNoise with your prefer value like 0.1</p>\n\n<p>example:\nconc = concatenate([conv1d_result, bi_lstm_result])(previous_layer)\nconc = BatchNormalization()(conc)\nconc = GaussianNoise(0.1)(conc)</p>\n\n<p>output = Dense(1, activation='sigmoid')(conc)</p>\n\n<p>You may experiment it with other setup.. Godd luck! :-)</p>",
          "votes": 18,
          "replies": []
        },
        {
          "id": 463081,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-29T11:22:24.950000",
          "content": "<p>Adding small amount noise after embedding layer is indeed an interesting option. \nIsn't it like creating \"synonyms\" in the embedding space?</p>\n\n<p>However we should be cautious:\nkeras.layers.GaussianNoise(stddev) has no seed argument :-(  thus every train run different random realization of noise will be used and we probably get more variance in f1 scores across runs :-(</p>",
          "votes": 10,
          "replies": []
        },
        {
          "id": 463326,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-01-29T20:09:09.167000",
          "content": "<p>I tried adding customized Gaussian noise layer after Embeddings. It hurts the performance, but I used it with embedding dropout. Maybe you have to choose only one of them, but I haven't given it a try yet.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 465572,
          "author_name": "Bjenk Ellefsen",
          "author_url": "",
          "post_date": "2019-02-03T14:07:34.793000",
          "content": "<p>Yes, in Dongxu and Zhichao's <a href=\"https://arxiv.org/abs/1804.08166\">paper</a>, the intuition is to add perturbation (to use Goodfellow's more accurate definition than \"noise\") to the embeddings. It is a data augmentation technique. So, pretty much as Annabelle is saying, after the embedding layer. I tried it and it didn't help.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 464304,
      "author_name": "bilal2vec",
      "author_url": "",
      "post_date": "2019-01-31T15:19:02.460000",
      "content": "<p>I added this version of gaussian noise with a stddev of 0.1 after my embeddings and increased my local CV by 0.0006</p>\n\n<pre><code>class GaussianNoise(nn.Module):\n    def __init__(self, stddev):\n        super(GaussianNoise, self).__init__()\n\n        self.stddev = stddev\n\n    def forward(self, x):\n        noise = torch.empty_like(x)\n        noise.normal_(0, self.stddev)\n\n        return x + noise\n\nif self.training:\n    x = self.gaussian_noise(x)\n</code></pre>",
      "votes": 10,
      "replies": [
        {
          "id": 464612,
          "author_name": "YaGana Sheriff-Hussaini",
          "author_url": "",
          "post_date": "2019-02-01T06:49:07.173000",
          "content": "<p><a href=\"/bkkaggle\">@bkkaggle</a>, I got similar improvement in CV by adding noise to the embedding layer of my keras model but my LB score went down 3 points. Did you see significant increase in computation time with the addition of noise? Mine resulted in a  negligible increase of about 75 seconds to the total kernel time.</p>\n\n<p>I am new to pytorch, how would you incorporate your GaussianNoise class to a model similar to the public pytorch kernel? I want to give it a try.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 466869,
      "author_name": "Kepler456b",
      "author_url": "",
      "post_date": "2019-02-06T05:24:12.173000",
      "content": "<p>I used synonyms and my CV went from 0.695 to 0.745. There is a catch though. The split may have the conjugate sentences in training and val. I only added 24k sentences.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 465657,
      "author_name": "Abhishekmamidi",
      "author_url": "",
      "post_date": "2019-02-03T17:55:20.663000",
      "content": "<p>Thank you for sharing. I will use some of them in the coming competition. Very interesting.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 463339,
      "author_name": "Andrey Nikishaev",
      "author_url": "",
      "post_date": "2019-01-29T20:35:21.517000",
      "content": "<p>What about using VAE for this task?</p>",
      "votes": 1,
      "replies": [
        {
          "id": 463354,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-01-29T21:06:20.800000",
          "content": "<p>I believe it should be definitely possible. Basically, VAE is just a variant of how one could learn the distribution over embeddings (which I mentioned in the post). BTW, <a href=\"https://arxiv.org/abs/1711.11027\">this paper</a> uses a variant of a VAE in fact.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 463455,
          "author_name": "Bai",
          "author_url": "",
          "post_date": "2019-01-30T02:38:07.620000",
          "content": "<p>time limits.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 463679,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-01-30T12:15:21.007000",
          "content": "<p>But VAE is not always promising. I have wrote about VAE quite a long time back. <a href=\"https://s4sarath.github.io/2016/11/23/variational_autoenocder_for_Natural_Language_Processing\">https://s4sarath.github.io/2016/11/23/variational_autoenocder_for_Natural_Language_Processing</a> </p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 465716,
      "author_name": "Jannen",
      "author_url": "",
      "post_date": "2019-02-03T20:49:06.737000",
      "content": "<p>Textblob library can do machine translation between languages, but it does it by using API to Google translator. As web connection is not allowed, that route is off. </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 463652,
      "author_name": "Yu Suzuki",
      "author_url": "",
      "post_date": "2019-01-30T11:24:34.787000",
      "content": "<p>My CV has gone up by 0.005. I have inserted two gaussian layers just after embedding and also \n just after concatenation.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 464160,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-01-31T09:22:50.110000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 464237,
          "author_name": "Yu Suzuki",
          "author_url": "",
          "post_date": "2019-01-31T12:39:49.073000",
          "content": "<p>Run time goes up from 6705 sec to 6935 sec. I set n_splits to 5  and n_epochs to 5.  I'm also using Pytorch.  </p>\n\n<p>I have inserted the gaussian layer as follows.</p>\n\n<p>///\nif self.training == True: \n   h_embedding = h_embedding + h_embedding.clone().normal_(0, 0.1) \n///</p>",
          "votes": 4,
          "replies": []
        },
        {
          "id": 465714,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-02-03T20:45:52.750000",
          "content": "<p>You mean:</p>\n\n<p>hembedding = hembedding + hembedding.clone().normal_(0, 0.1)</p>\n\n<p>don't you? :-)</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 466441,
          "author_name": "Yu Suzuki",
          "author_url": "",
          "post_date": "2019-02-05T11:49:18.113000",
          "content": "<p>Yes.\nAnd I think it is better to disable the noise while prediction as the dropout.\nTherefore I've applied the noise during training only.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 463600,
      "author_name": "Jie Wu",
      "author_url": "",
      "post_date": "2019-01-30T09:08:53.353000",
      "content": "<p>Add GaussianNoise my LB went down to 0.686-0.688</p>",
      "votes": 2,
      "replies": [
        {
          "id": 463619,
          "author_name": "Dmytro Danevskyi",
          "author_url": "",
          "post_date": "2019-01-30T09:51:46.333000",
          "content": "<p>What about CV? ;)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 463774,
          "author_name": "Jie Wu",
          "author_url": "",
          "post_date": "2019-01-30T15:45:49.130000",
          "content": "<p>CV increased 0.003</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 463780,
          "author_name": "Andrzej Kuro",
          "author_url": "",
          "post_date": "2019-01-30T15:53:34.253000",
          "content": "<p>wow, so CV increased which would suggest that model generalizes better, but in LB score went down, so it looks like previous model was overfit to public LB set. If it happened in a small set I would suspect the seed used, but I don't think this is the case here?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 987979,
      "author_name": "D1nall",
      "author_url": "",
      "post_date": "2020-08-27T17:06:53.360000",
      "content": "<p>Check out EDA<br>\n<a href=\"https://www.kaggle.com/swarajshinde/eda-data-augmentation-techniques-for-text-nlp\" target=\"_blank\">https://www.kaggle.com/swarajshinde/eda-data-augmentation-techniques-for-text-nlp</a></p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "462470": "Data augmentation is a very popular strategy for preventing overfitting in convolutional networks. It really makes a lot of sense to use flips/rotations/color shifts for such tasks as classification, object detection, pose estimation, etc., since we know that such transformations (unless used with very aggressive parameters) should not affect correct labels much. This simple strategy populates the data space, increases the model's ability to generalize, and also makes the model to be more robust to unusual lighting conditions, partial occlusion, etc.\n\nAs for text data, which is in contrast to image data is essentially *discrete*, there is no obvious way to do the augmentation.\n\nHowever, there were some successful applications of data augmentation in raw data space. Some of them were already listed [here][1] by @Dieter. One way is to replace some words with synonyms. Another (which was successfully applied by the winners of [Toxic Comment Classification Challenge][2]) is to do cross-translation: using some pretrained machine translation model, translate a source sentence to another language, then translate back.\n\nWhat I want to point out is that there are some ways to do the augmentation in embedding space as well! Since word embeddings consist of real numbers rather than discrete tokens, it opens a lot of interesting ways to augment the data.  \n\nFor instance, [this][3] paper studies the effects of gaussian/bernoulli dropout and adversarial examples generation and reports some improvement over the baseline.\n\nOther options include:\n\n - Additive gaussian/other noise. Theoretically, if the perturbation doesn't change the list of nearest neighbors of a particular embedding much, a classifier can still be able to learn without any problems. \n - If you are using more than one set of embeddings in your pipeline (say, glove+paragram) you may try to create a model that learns the relationship between the same words represented by different embeddings. Then you can randomly replace some embeddings with their versions converted from another embedding set.\n - Learn a distribution of embeddings for a word instead of learning fixed vectors. Then, on each train iteration, you can use the learned distribution to sample new but valid versions of the word embeddings. This may be tricky to implement and also not really applicable to the current competition due to time constraints, but I definitely suggest you check the [paper][4] if you are interested.\n \nLet's discuss and happy kaggling!\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71083\n  [2]: https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge\n  [3]: https://arxiv.org/abs/1804.08166\n  [4]: https://arxiv.org/abs/1711.11027",
    "462487": "I added GaussianNoise with 0.1 from keras and my score improved up to 0.002-0.004",
    "464304": "I added this version of gaussian noise with a stddev of 0.1 after my embeddings and increased my local CV by 0.0006\n\n    class GaussianNoise(nn.Module):\n        def __init__(self, stddev):\n            super(GaussianNoise, self).__init__()\n        \n            self.stddev = stddev\n        \n        def forward(self, x):\n            noise = torch.empty_like(x)\n            noise.normal_(0, self.stddev)\n        \n            return x + noise\n\n    if self.training:\n        x = self.gaussian_noise(x)",
    "466869": "I used synonyms and my CV went from 0.695 to 0.745. There is a catch though. The split may have the conjugate sentences in training and val. I only added 24k sentences.",
    "465657": "Thank you for sharing. I will use some of them in the coming competition. Very interesting.",
    "463339": "What about using VAE for this task?",
    "465716": "Textblob library can do machine translation between languages, but it does it by using API to Google translator. As web connection is not allowed, that route is off. ",
    "463652": "My CV has gone up by 0.005. I have inserted two gaussian layers just after embedding and also \n just after concatenation.",
    "463600": "Add GaussianNoise my LB went down to 0.686-0.688",
    "987979": "Check out EDA\nhttps://www.kaggle.com/swarajshinde/eda-data-augmentation-techniques-for-text-nlp"
  }
}