{
  "id": 71083,
  "title": "Augmentation for text",
  "url": "/competitions/quora-insincere-questions-classification/discussion/71083",
  "author_name": "Dieter",
  "post_date": "2018-11-09T23:56:36.220000",
  "votes": 123,
  "comment_count": 30,
  "views": 0,
  "content": "<p>Of course there are a lot of augmentation techniques for images, but what about text? Let's discuss some techniques:</p>\n\n<ol>\n<li>Exchanging words with synonyms (see e.g. <a href=\"https://arxiv.org/pdf/1502.01710.pdf\">https://arxiv.org/pdf/1502.01710.pdf</a> )</li>\n<li>noising in RNN (<a href=\"https://arxiv.org/pdf/1703.02573.pdf\">https://arxiv.org/pdf/1703.02573.pdf</a>)</li>\n<li>Translation to other language and back ( <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038</a> )</li>\n</ol>",
  "messages": [
    {
      "id": 418465,
      "postDate": "2018-11-09T23:56:36.220Z",
      "content": "<p>Of course there are a lot of augmentation techniques for images, but what about text? Let's discuss some techniques:</p>\n\n<ol>\n<li>Exchanging words with synonyms (see e.g. <a href=\"https://arxiv.org/pdf/1502.01710.pdf\">https://arxiv.org/pdf/1502.01710.pdf</a> )</li>\n<li>noising in RNN (<a href=\"https://arxiv.org/pdf/1703.02573.pdf\">https://arxiv.org/pdf/1703.02573.pdf</a>)</li>\n<li>Translation to other language and back ( <a href=\"https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038\">https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038</a> )</li>\n</ol>",
      "rawMarkdown": "Of course there are a lot of augmentation techniques for images, but what about text? Let's discuss some techniques:\n\n 1. Exchanging words with synonyms (see e.g. https://arxiv.org/pdf/1502.01710.pdf )\n 2. noising in RNN (https://arxiv.org/pdf/1703.02573.pdf)\n 3. Translation to other language and back ( https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038 )\n\n  ",
      "votes": 123
    },
    {
      "id": 419860,
      "postDate": "2018-11-12T16:45:19.820Z",
      "content": "<p>Thanks for the post. I have started experimenting with adding noise into CNN and it seems to improve the score.</p>\n\n<p>Noising the train data (assuming using keras' text tokenising and sequence):</p>\n\n<pre><code>def noise_measurement(train, noise_level):\n    noised_train = train.copy()\n    to_transform = np.random.random(train.shape) &amp;lt; noise_level\n    transform_y = np.random.randint(0, max_features, size=train.shape)\n    noised_train[to_transform] = transform_y[to_transform]\n    return noised_train\n\nbatch_size = 256\nepochs = 4\n\nmodel = get_model()\n\nX_tra, X_val, y_tra, y_val = train_test_split(x_train, y_train, train_size=0.85,\n                                              random_state=233)\nF1_Score = F1Evaluation(validation_data=(X_val, y_val), interval=1)\n\nscore = 0\nbest_i = 0\nfor i in range(epochs):\n    hist = model.fit(noise_measurement(X_tra, 0.15), y_tra, batch_size=batch_size, epochs=1,\n                     validation_data=(X_val, y_val),\n                     callbacks=[F1_Score], verbose=2)\n</code></pre>",
      "rawMarkdown": "Thanks for the post. I have started experimenting with adding noise into CNN and it seems to improve the score.\n\nNoising the train data (assuming using keras' text tokenising and sequence):\n\n    def noise_measurement(train, noise_level):\n        noised_train = train.copy()\n        to_transform = np.random.random(train.shape) &lt; noise_level\n        transform_y = np.random.randint(0, max_features, size=train.shape)\n        noised_train[to_transform] = transform_y[to_transform]\n        return noised_train\n    \n    batch_size = 256\n    epochs = 4\n    \n    model = get_model()\n    \n    X_tra, X_val, y_tra, y_val = train_test_split(x_train, y_train, train_size=0.85,\n                                                  random_state=233)\n    F1_Score = F1Evaluation(validation_data=(X_val, y_val), interval=1)\n    \n    score = 0\n    best_i = 0\n    for i in range(epochs):\n        hist = model.fit(noise_measurement(X_tra, 0.15), y_tra, batch_size=batch_size, epochs=1,\n                         validation_data=(X_val, y_val),\n                         callbacks=[F1_Score], verbose=2)",
      "votes": 10,
      "replies": [
        {
          "id": 420396,
          "postDate": "2018-11-13T15:07:05.417Z",
          "content": "<p>Thank you, will check it out</p>",
          "rawMarkdown": "Thank you, will check it out",
          "votes": 1
        },
        {
          "id": 422907,
          "postDate": "2018-11-17T02:51:56.353Z",
          "content": "<p>I think the function of <code>noise_measurement</code> look like SpatialDropout1D.</p>",
          "rawMarkdown": "I think the function of `noise_measurement` look like SpatialDropout1D."
        },
        {
          "id": 430245,
          "postDate": "2018-11-30T04:21:35.507Z",
          "content": "<p>Great script, very helpful. Can you explain <code>noise_level</code>?</p>",
          "rawMarkdown": "Great script, very helpful. Can you explain `noise_level`?"
        },
        {
          "id": 468513,
          "postDate": "2019-02-09T03:02:38.003Z",
          "content": "<p>Thanks. How much did the trick improve your model?</p>",
          "rawMarkdown": "Thanks. How much did the trick improve your model?"
        }
      ]
    },
    {
      "id": 418613,
      "postDate": "2018-11-10T09:12:20.150Z",
      "content": "<p>If you use a CNN structure you can also flip the text. <a href=\"https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#\">https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#</a></p>",
      "rawMarkdown": "If you use a CNN structure you can also flip the text. https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#",
      "votes": 5,
      "replies": [
        {
          "id": 418709,
          "postDate": "2018-11-10T13:28:00.373Z",
          "content": "<p>Are there any theory or research that have seen performance increase with flipping? I personally don't find any reason to do this. Haven't try it myself though.</p>",
          "rawMarkdown": "Are there any theory or research that have seen performance increase with flipping? I personally don't find any reason to do this. Haven't try it myself though.",
          "votes": 3
        },
        {
          "id": 419011,
          "postDate": "2018-11-11T04:44:51.960Z",
          "content": "<p>Actually that is exactly what Bidirectional does for RNN</p>",
          "rawMarkdown": "Actually that is exactly what Bidirectional does for RNN",
          "votes": 7
        }
      ]
    },
    {
      "id": 854691,
      "postDate": "2020-05-20T08:31:27.037Z",
      "content": "<p>Hi,</p>\n\n<p>I recently read the current literature on augmentation and have summarized my findings here: <a href=\"https://amitness.com/2020/05/data-augmentation-for-nlp\">https://amitness.com/2020/05/data-augmentation-for-nlp</a></p>",
      "rawMarkdown": "Hi,\n\nI recently read the current literature on augmentation and have summarized my findings here: https://amitness.com/2020/05/data-augmentation-for-nlp",
      "votes": 1
    },
    {
      "id": 644156,
      "postDate": "2019-10-08T12:27:36.703Z",
      "content": "<p>sentencepiece have one more: using a different tokenization of known words</p>",
      "rawMarkdown": "sentencepiece have one more: using a different tokenization of known words",
      "votes": 1
    },
    {
      "id": 418859,
      "postDate": "2018-11-10T18:40:41.767Z",
      "content": "<p>@Dieter Your InceptionCNN Model was really helpful to me, it helped me jump .668 to .674, with a few tweaks and very less epochs. I'll try to make some more changes.</p>",
      "rawMarkdown": "@Dieter Your InceptionCNN Model was really helpful to me, it helped me jump .668 to .674, with a few tweaks and very less epochs. I'll try to make some more changes.\n",
      "votes": 3
    },
    {
      "id": 443083,
      "postDate": "2018-12-21T01:54:23.757Z",
      "content": "<p>Maybe synthesis some of insincere cases can be helpful</p>",
      "rawMarkdown": "Maybe synthesis some of insincere cases can be helpful",
      "votes": 1
    },
    {
      "id": 425290,
      "postDate": "2018-11-21T11:55:05.257Z",
      "content": "<p>what about using nltk.wordnet and Embeddings(adding top 10 similar words) to expand vocabulary to the training data ??   </p>",
      "rawMarkdown": "what about using nltk.wordnet and Embeddings(adding top 10 similar words) to expand vocabulary to the training data ??   ",
      "votes": 1,
      "replies": [
        {
          "id": 432795,
          "postDate": "2018-12-04T10:14:33.597Z",
          "content": "<p>I tried this in several nlp task, it didn't work for me...</p>",
          "rawMarkdown": "I tried this in several nlp task, it didn't work for me...",
          "votes": 1
        }
      ]
    },
    {
      "id": 420788,
      "postDate": "2018-11-14T05:41:47.050Z",
      "content": "<p>This paper seems to give a great methodology of data augmentation (grammartically controlled paraphrase and adversarial as well) : \n<a href=\"https://arxiv.org/pdf/1804.06059.pdf\">https://arxiv.org/pdf/1804.06059.pdf</a>\ncoming with pretrained model : \n<a href=\"https://github.com/miyyer/scpn/blob/master/README.md\">https://github.com/miyyer/scpn/blob/master/README.md</a></p>\n\n<p>Though, it seems we cannot use it in our problem here.</p>",
      "rawMarkdown": "This paper seems to give a great methodology of data augmentation (grammartically controlled paraphrase and adversarial as well) : \nhttps://arxiv.org/pdf/1804.06059.pdf\ncoming with pretrained model : \nhttps://github.com/miyyer/scpn/blob/master/README.md\n\nThough, it seems we cannot use it in our problem here.",
      "votes": 2
    },
    {
      "id": 1385994,
      "postDate": "2021-07-13T06:54:52.287Z",
      "content": "<p>I am trying to learn NLP, and this is very helpful. Thank you! </p>",
      "rawMarkdown": "I am trying to learn NLP, and this is very helpful. Thank you! "
    },
    {
      "id": 428315,
      "postDate": "2018-11-27T04:17:02.643Z",
      "content": "<p>I assume that thesauruses and translators would count as external data for this competition, no? </p>",
      "rawMarkdown": "I assume that thesauruses and translators would count as external data for this competition, no? ",
      "replies": [
        {
          "id": 456201,
          "postDate": "2019-01-15T10:43:25.677Z",
          "content": "<p>what about long literals or  nltk included wordnet for synonyms?</p>",
          "rawMarkdown": "what about long literals or  nltk included wordnet for synonyms?"
        }
      ]
    },
    {
      "id": 418808,
      "postDate": "2018-11-10T16:41:16.507Z",
      "content": "<p>Can you translate the text by using TextBlob library (3rd alternative)? Internet is required and I'm not sure if it is permmited.</p>",
      "rawMarkdown": "Can you translate the text by using TextBlob library (3rd alternative)? Internet is required and I'm not sure if it is permmited.",
      "replies": [
        {
          "id": 419012,
          "postDate": "2018-11-11T04:46:14.250Z",
          "content": "<p>I am afraid you are not allowed, however I wanted to list it anyway</p>",
          "rawMarkdown": "I am afraid you are not allowed, however I wanted to list it anyway"
        }
      ]
    },
    {
      "id": 418731,
      "postDate": "2018-11-10T14:03:58.850Z",
      "content": "<p>I will try your flipping tips. Embedded sentences can be seen as sort of image (seqlen, embedsize) so flipping seems good. In one of my personal work i used to display such \"image\" ; it looks like colored bar code :-)</p>",
      "rawMarkdown": "I will try your flipping tips. Embedded sentences can be seen as sort of image (seqlen, embedsize) so flipping seems good. In one of my personal work i used to display such \"image\" ; it looks like colored bar code :-)"
    },
    {
      "id": 418474,
      "postDate": "2018-11-10T00:38:07.457Z",
      "content": "<p>interpolating between two text embeddings as shown in \"Generative Adversarial Text to Image Synthesis\"\n<a href=\"https://arxiv.org/abs/1605.05396\">https://arxiv.org/abs/1605.05396</a></p>",
      "rawMarkdown": "interpolating between two text embeddings as shown in \"Generative Adversarial Text to Image Synthesis\"\nhttps://arxiv.org/abs/1605.05396\n"
    },
    {
      "id": 418473,
      "postDate": "2018-11-10T00:36:42.307Z",
      "content": "<p>Semantic Parsing with Semi-Supervised Sequential Autoencoders\n<a href=\"https://arxiv.org/abs/1609.09315\">https://arxiv.org/abs/1609.09315</a></p>",
      "rawMarkdown": "Semantic Parsing with Semi-Supervised Sequential Autoencoders\nhttps://arxiv.org/abs/1609.09315\n\n"
    },
    {
      "id": 3195390,
      "postDate": "2025-05-07T01:20:50.397Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 430266,
      "postDate": "2018-11-30T04:54:44.217Z",
      "rawMarkdown": "",
      "votes": 2,
      "isDeleted": true
    },
    {
      "id": 419378,
      "postDate": "2018-11-11T20:49:30.103Z",
      "rawMarkdown": "",
      "votes": 4,
      "isDeleted": true,
      "replies": [
        {
          "id": 419521,
          "postDate": "2018-11-12T05:25:15.820Z",
          "content": "<p>There is also a deep-learning based method where you add random noise to wikipedia and let the NN correct it <a href=\"https://github.com/MajorTal/DeepSpell\">https://github.com/MajorTal/DeepSpell</a></p>",
          "rawMarkdown": "There is also a deep-learning based method where you add random noise to wikipedia and let the NN correct it https://github.com/MajorTal/DeepSpell",
          "votes": 6
        },
        {
          "id": 430251,
          "postDate": "2018-11-30T04:30:54.300Z",
          "content": "<p>Norvig's solution seems to be pretty slow relative to this dataset.</p>",
          "rawMarkdown": "Norvig's solution seems to be pretty slow relative to this dataset."
        }
      ]
    },
    {
      "id": 1495838,
      "postDate": "2021-08-29T20:36:59.247Z",
      "content": "<p>Thanks for sharing!! very helpful</p>",
      "rawMarkdown": "Thanks for sharing!! very helpful\n\n",
      "votes": 1
    },
    {
      "id": 418531,
      "postDate": "2018-11-10T03:50:24.107Z",
      "content": "<p>thanks</p>",
      "rawMarkdown": "thanks"
    }
  ],
  "comments": [
    {
      "id": 419860,
      "author_name": "Yair Beer",
      "author_url": "",
      "post_date": "2018-11-12T16:45:19.820000",
      "content": "<p>Thanks for the post. I have started experimenting with adding noise into CNN and it seems to improve the score.</p>\n\n<p>Noising the train data (assuming using keras' text tokenising and sequence):</p>\n\n<pre><code>def noise_measurement(train, noise_level):\n    noised_train = train.copy()\n    to_transform = np.random.random(train.shape) &amp;lt; noise_level\n    transform_y = np.random.randint(0, max_features, size=train.shape)\n    noised_train[to_transform] = transform_y[to_transform]\n    return noised_train\n\nbatch_size = 256\nepochs = 4\n\nmodel = get_model()\n\nX_tra, X_val, y_tra, y_val = train_test_split(x_train, y_train, train_size=0.85,\n                                              random_state=233)\nF1_Score = F1Evaluation(validation_data=(X_val, y_val), interval=1)\n\nscore = 0\nbest_i = 0\nfor i in range(epochs):\n    hist = model.fit(noise_measurement(X_tra, 0.15), y_tra, batch_size=batch_size, epochs=1,\n                     validation_data=(X_val, y_val),\n                     callbacks=[F1_Score], verbose=2)\n</code></pre>",
      "votes": 10,
      "replies": [
        {
          "id": 420396,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-13T15:07:05.417000",
          "content": "<p>Thank you, will check it out</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 422907,
          "author_name": "Salon_sai",
          "author_url": "",
          "post_date": "2018-11-17T02:51:56.353000",
          "content": "<p>I think the function of <code>noise_measurement</code> look like SpatialDropout1D.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 430245,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-11-30T04:21:35.507000",
          "content": "<p>Great script, very helpful. Can you explain <code>noise_level</code>?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468513,
          "author_name": "LTDR",
          "author_url": "",
          "post_date": "2019-02-09T03:02:38.003000",
          "content": "<p>Thanks. How much did the trick improve your model?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 418613,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-11-10T09:12:20.150000",
      "content": "<p>If you use a CNN structure you can also flip the text. <a href=\"https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#\">https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#</a></p>",
      "votes": 5,
      "replies": [
        {
          "id": 418709,
          "author_name": "Mekaveli.",
          "author_url": "",
          "post_date": "2018-11-10T13:28:00.373000",
          "content": "<p>Are there any theory or research that have seen performance increase with flipping? I personally don't find any reason to do this. Haven't try it myself though.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 419011,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-11T04:44:51.960000",
          "content": "<p>Actually that is exactly what Bidirectional does for RNN</p>",
          "votes": 7,
          "replies": []
        }
      ]
    },
    {
      "id": 854691,
      "author_name": "Amit Chaudhary",
      "author_url": "",
      "post_date": "2020-05-20T08:31:27.037000",
      "content": "<p>Hi,</p>\n\n<p>I recently read the current literature on augmentation and have summarized my findings here: <a href=\"https://amitness.com/2020/05/data-augmentation-for-nlp\">https://amitness.com/2020/05/data-augmentation-for-nlp</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 644156,
      "author_name": "Cédric Lacrambe",
      "author_url": "",
      "post_date": "2019-10-08T12:27:36.703000",
      "content": "<p>sentencepiece have one more: using a different tokenization of known words</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 418859,
      "author_name": "AshishSinha",
      "author_url": "",
      "post_date": "2018-11-10T18:40:41.767000",
      "content": "<p>@Dieter Your InceptionCNN Model was really helpful to me, it helped me jump .668 to .674, with a few tweaks and very less epochs. I'll try to make some more changes.</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 443083,
      "author_name": "mr007rin",
      "author_url": "",
      "post_date": "2018-12-21T01:54:23.757000",
      "content": "<p>Maybe synthesis some of insincere cases can be helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 425290,
      "author_name": "BalaramB",
      "author_url": "",
      "post_date": "2018-11-21T11:55:05.257000",
      "content": "<p>what about using nltk.wordnet and Embeddings(adding top 10 similar words) to expand vocabulary to the training data ??   </p>",
      "votes": 1,
      "replies": [
        {
          "id": 432795,
          "author_name": "spongebob",
          "author_url": "",
          "post_date": "2018-12-04T10:14:33.597000",
          "content": "<p>I tried this in several nlp task, it didn't work for me...</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 420788,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2018-11-14T05:41:47.050000",
      "content": "<p>This paper seems to give a great methodology of data augmentation (grammartically controlled paraphrase and adversarial as well) : \n<a href=\"https://arxiv.org/pdf/1804.06059.pdf\">https://arxiv.org/pdf/1804.06059.pdf</a>\ncoming with pretrained model : \n<a href=\"https://github.com/miyyer/scpn/blob/master/README.md\">https://github.com/miyyer/scpn/blob/master/README.md</a></p>\n\n<p>Though, it seems we cannot use it in our problem here.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1385994,
      "author_name": "Rizvi Hasan",
      "author_url": "",
      "post_date": "2021-07-13T06:54:52.287000",
      "content": "<p>I am trying to learn NLP, and this is very helpful. Thank you! </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 428315,
      "author_name": "HappyPrancer",
      "author_url": "",
      "post_date": "2018-11-27T04:17:02.643000",
      "content": "<p>I assume that thesauruses and translators would count as external data for this competition, no? </p>",
      "votes": 0,
      "replies": [
        {
          "id": 456201,
          "author_name": "Cédric Lacrambe",
          "author_url": "",
          "post_date": "2019-01-15T10:43:25.677000",
          "content": "<p>what about long literals or  nltk included wordnet for synonyms?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 418808,
      "author_name": "cayala",
      "author_url": "",
      "post_date": "2018-11-10T16:41:16.507000",
      "content": "<p>Can you translate the text by using TextBlob library (3rd alternative)? Internet is required and I'm not sure if it is permmited.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 419012,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-11T04:46:14.250000",
          "content": "<p>I am afraid you are not allowed, however I wanted to list it anyway</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 418731,
      "author_name": "Jean-Marc",
      "author_url": "",
      "post_date": "2018-11-10T14:03:58.850000",
      "content": "<p>I will try your flipping tips. Embedded sentences can be seen as sort of image (seqlen, embedsize) so flipping seems good. In one of my personal work i used to display such \"image\" ; it looks like colored bar code :-)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418474,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-11-10T00:38:07.457000",
      "content": "<p>interpolating between two text embeddings as shown in \"Generative Adversarial Text to Image Synthesis\"\n<a href=\"https://arxiv.org/abs/1605.05396\">https://arxiv.org/abs/1605.05396</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 418473,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2018-11-10T00:36:42.307000",
      "content": "<p>Semantic Parsing with Semi-Supervised Sequential Autoencoders\n<a href=\"https://arxiv.org/abs/1609.09315\">https://arxiv.org/abs/1609.09315</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3195390,
      "author_name": "",
      "author_url": "",
      "post_date": "2025-05-07T01:20:50.397000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 430266,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-30T04:54:44.217000",
      "content": "",
      "votes": 2,
      "replies": []
    },
    {
      "id": 419378,
      "author_name": "",
      "author_url": "",
      "post_date": "2018-11-11T20:49:30.103000",
      "content": "",
      "votes": 4,
      "replies": [
        {
          "id": 419521,
          "author_name": "Dieter",
          "author_url": "",
          "post_date": "2018-11-12T05:25:15.820000",
          "content": "<p>There is also a deep-learning based method where you add random noise to wikipedia and let the NN correct it <a href=\"https://github.com/MajorTal/DeepSpell\">https://github.com/MajorTal/DeepSpell</a></p>",
          "votes": 6,
          "replies": []
        },
        {
          "id": 430251,
          "author_name": "Matthew Anderson",
          "author_url": "",
          "post_date": "2018-11-30T04:30:54.300000",
          "content": "<p>Norvig's solution seems to be pretty slow relative to this dataset.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1495838,
      "author_name": "Fuco",
      "author_url": "",
      "post_date": "2021-08-29T20:36:59.247000",
      "content": "<p>Thanks for sharing!! very helpful</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 418531,
      "author_name": "Nipi",
      "author_url": "",
      "post_date": "2018-11-10T03:50:24.107000",
      "content": "<p>thanks</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "418465": "Of course there are a lot of augmentation techniques for images, but what about text? Let's discuss some techniques:\n\n 1. Exchanging words with synonyms (see e.g. https://arxiv.org/pdf/1502.01710.pdf )\n 2. noising in RNN (https://arxiv.org/pdf/1703.02573.pdf)\n 3. Translation to other language and back ( https://www.kaggle.com/c/jigsaw-toxic-comment-classification-challenge/discussion/48038 )\n\n  ",
    "419860": "Thanks for the post. I have started experimenting with adding noise into CNN and it seems to improve the score.\n\nNoising the train data (assuming using keras' text tokenising and sequence):\n\n    def noise_measurement(train, noise_level):\n        noised_train = train.copy()\n        to_transform = np.random.random(train.shape) &lt; noise_level\n        transform_y = np.random.randint(0, max_features, size=train.shape)\n        noised_train[to_transform] = transform_y[to_transform]\n        return noised_train\n    \n    batch_size = 256\n    epochs = 4\n    \n    model = get_model()\n    \n    X_tra, X_val, y_tra, y_val = train_test_split(x_train, y_train, train_size=0.85,\n                                                  random_state=233)\n    F1_Score = F1Evaluation(validation_data=(X_val, y_val), interval=1)\n    \n    score = 0\n    best_i = 0\n    for i in range(epochs):\n        hist = model.fit(noise_measurement(X_tra, 0.15), y_tra, batch_size=batch_size, epochs=1,\n                         validation_data=(X_val, y_val),\n                         callbacks=[F1_Score], verbose=2)",
    "418613": "If you use a CNN structure you can also flip the text. https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#",
    "854691": "Hi,\n\nI recently read the current literature on augmentation and have summarized my findings here: https://amitness.com/2020/05/data-augmentation-for-nlp",
    "644156": "sentencepiece have one more: using a different tokenization of known words",
    "418859": "@Dieter Your InceptionCNN Model was really helpful to me, it helped me jump .668 to .674, with a few tweaks and very less epochs. I'll try to make some more changes.\n",
    "443083": "Maybe synthesis some of insincere cases can be helpful",
    "425290": "what about using nltk.wordnet and Embeddings(adding top 10 similar words) to expand vocabulary to the training data ??   ",
    "420788": "This paper seems to give a great methodology of data augmentation (grammartically controlled paraphrase and adversarial as well) : \nhttps://arxiv.org/pdf/1804.06059.pdf\ncoming with pretrained model : \nhttps://github.com/miyyer/scpn/blob/master/README.md\n\nThough, it seems we cannot use it in our problem here.",
    "1385994": "I am trying to learn NLP, and this is very helpful. Thank you! ",
    "428315": "I assume that thesauruses and translators would count as external data for this competition, no? ",
    "418808": "Can you translate the text by using TextBlob library (3rd alternative)? Internet is required and I'm not sure if it is permmited.",
    "418731": "I will try your flipping tips. Embedded sentences can be seen as sort of image (seqlen, embedsize) so flipping seems good. In one of my personal work i used to display such \"image\" ; it looks like colored bar code :-)",
    "418474": "interpolating between two text embeddings as shown in \"Generative Adversarial Text to Image Synthesis\"\nhttps://arxiv.org/abs/1605.05396\n",
    "418473": "Semantic Parsing with Semi-Supervised Sequential Autoencoders\nhttps://arxiv.org/abs/1609.09315\n\n",
    "3195390": "",
    "430266": "",
    "419378": "",
    "1495838": "Thanks for sharing!! very helpful\n\n",
    "418531": "thanks"
  }
}