{
  "id": 80577,
  "title": "33rd place solution- FastText embedding",
  "url": "/competitions/quora-insincere-questions-classification/writeups/plan-null-33rd-place-solution-fasttext-embedding",
  "author_name": "",
  "post_date": "2019-02-20T10:46:27.253Z",
  "votes": 12,
  "comment_count": 2,
  "views": 0,
  "content": "<p>Let me start by thanking everyone here who participated and contributed. We learned a lot of tips and tricks from community shared kernels and discussions. The competition was challenging in terms of finding correlated local validation, running the solution in 2 hours, and producing reproducible results, to name a few. We found a nearly correlated validation set before 1 week of the competition end. We used simple averaging where each trained NN model (Pytorch) was different either in terms of learning rate, pre-processing, embedding or architecture to maintain model diversity. One important thing we realised was: given small learning rate and large number of epochs, single FastText embedding based models were beating our glove+paragram models. So we focused on tuning FastText based models which could provide considerable score within 5-6 epochs. We also added Gaussian Noise to some models after embeddings to reduce the overdependence of RNN on specific keywords.</p>\n\n<p><strong>Solution summary:</strong></p>\n\n<ul>\n<li><strong>Runtime</strong>: 6352.2 secs</li>\n<li><strong>Preprocessing</strong>: Cleaning special characters, number pre-processing, misspell cleaning (For some models changed preprocessing sequence to add diversity)</li>\n<li><strong>Embedding</strong>: GLoVe, FastText, Paragram embeddings </li>\n<li><strong>Neural Network architecture</strong> (trained for 5-6 epochs with no fold): \n<ul><li>Stacked LSTM-GRU-128 hidden units, with GloVe+Paragram embedding</li>\n<li>Stacked LSTM-GRU 60 hidden units with attention and capsule, and GloVe+Paragram embedding</li>\n<li>Stacked LSTM-GRU-60 hidden units with GloVe+Paragram embedding</li>\n<li>Stacked LSTM-GRU-60 hidden units with FastText embeddings</li>\n<li>Stacked LSTM-GRU-80 hidden units with FastText embeddings and different preprocessing sequence </li></ul></li>\n<li><strong>Blending</strong>: Averaging prediction of each model with linear regression coefficients </li>\n</ul>\n\n<p><strong>Things that did not work</strong></p>\n\n<ul>\n<li>We tried pseudo labelling in different ways, but it didn't provide major boost considering its running time for it, so we dropped it in the end.</li>\n<li>We tried variety of preprocessing techniques to no avail. All of them tend to decrease the LB score with slight improvement in cv. Fearing overfitting pre-processing to training data with kept it to minimum.</li>\n<li>One trick that we tried was weight saving and retraining. For example, we trained the model and saved its weights before the model reached optimum. Then for the next model we loaded the weights for LSTM and GRU and did not pass gradients through them. This forced the new parts of model like an extra CNN layers or linear units to cover up for this. This saved time as the new model reached optimum within 2 epochs. But this did not add considerable benefit to the ensemble. In my opinion, majority of the information pertaining text was captured by RNN units, leaving little information required to be captured by newly added layers. Do let me know your thoughts on this experiments and its results.</li>\n</ul>\n\n<p>Special thanks to <a href=\"http://www.kaggle.com/shujian\">Shujian</a>, <a href=\"http://www.kaggle.com/bminixhofer\">Benjamin Minixhofer</a>, <a href=\"http://www.kaggle.com/christofhenkel\">Dieter</a>, <a href=\"http://www.kaggle.com/ryches\">Ryches</a>, <a href=\"http://www.kaggle.com/tunguz\">Bojan</a>, to name a few, for great kernels and discussions!</p>\n\n<p>This wouldn’t have been possible without awesome teammates, <a href=\"http://www.kaggle.com/ashish2123\">Ashish</a> and <a href=\"http://www.kaggle.com/rsrade\">Rahul</a>, who put in lot of effort and made this competition a great learning experience.</p>\n\n<p>Thanks for reading, I am planning to release the code in few days after cleaning it.    </p>\n\n<p>Edit: Github repository: <a href=\"https://github.com/soham97/Quora-Insincere-Questions-Classification-Challenge-NLP\">https://github.com/soham97/Quora-Insincere-Questions-Classification-Challenge-NLP</a></p>",
  "messages": [
    {
      "id": "471526",
      "postDate": "02/14/2019 15:21:10",
      "content": "<p>Let me start by thanking everyone here who participated and contributed. We learned a lot of tips and tricks from community shared kernels and discussions. The competition was challenging in terms of finding correlated local validation, running the solution in 2 hours, and producing reproducible results, to name a few. We found a nearly correlated validation set before 1 week of the competition end. We used simple averaging where each trained NN model (Pytorch) was different either in terms of learning rate, pre-processing, embedding or architecture to maintain model diversity. One important thing we realised was: given small learning rate and large number of epochs, single FastText embedding based models were beating our glove+paragram models. So we focused on tuning FastText based models which could provide considerable score within 5-6 epochs. We also added Gaussian Noise to some models after embeddings to reduce the overdependence of RNN on specific keywords.</p>\n\n<p><strong>Solution summary:</strong></p>\n\n<ul>\n<li><strong>Runtime</strong>: 6352.2 secs</li>\n<li><strong>Preprocessing</strong>: Cleaning special characters, number pre-processing, misspell cleaning (For some models changed preprocessing sequence to add diversity)</li>\n<li><strong>Embedding</strong>: GLoVe, FastText, Paragram embeddings </li>\n<li><strong>Neural Network architecture</strong> (trained for 5-6 epochs with no fold): \n<ul><li>Stacked LSTM-GRU-128 hidden units, with GloVe+Paragram embedding</li>\n<li>Stacked LSTM-GRU 60 hidden units with attention and capsule, and GloVe+Paragram embedding</li>\n<li>Stacked LSTM-GRU-60 hidden units with GloVe+Paragram embedding</li>\n<li>Stacked LSTM-GRU-60 hidden units with FastText embeddings</li>\n<li>Stacked LSTM-GRU-80 hidden units with FastText embeddings and different preprocessing sequence </li></ul></li>\n<li><strong>Blending</strong>: Averaging prediction of each model with linear regression coefficients </li>\n</ul>\n\n<p><strong>Things that did not work</strong></p>\n\n<ul>\n<li>We tried pseudo labelling in different ways, but it didn't provide major boost considering its running time for it, so we dropped it in the end.</li>\n<li>We tried variety of preprocessing techniques to no avail. All of them tend to decrease the LB score with slight improvement in cv. Fearing overfitting pre-processing to training data with kept it to minimum.</li>\n<li>One trick that we tried was weight saving and retraining. For example, we trained the model and saved its weights before the model reached optimum. Then for the next model we loaded the weights for LSTM and GRU and did not pass gradients through them. This forced the new parts of model like an extra CNN layers or linear units to cover up for this. This saved time as the new model reached optimum within 2 epochs. But this did not add considerable benefit to the ensemble. In my opinion, majority of the information pertaining text was captured by RNN units, leaving little information required to be captured by newly added layers. Do let me know your thoughts on this experiments and its results.</li>\n</ul>\n\n<p>Special thanks to <a href=\"http://www.kaggle.com/shujian\">Shujian</a>, <a href=\"http://www.kaggle.com/bminixhofer\">Benjamin Minixhofer</a>, <a href=\"http://www.kaggle.com/christofhenkel\">Dieter</a>, <a href=\"http://www.kaggle.com/ryches\">Ryches</a>, <a href=\"http://www.kaggle.com/tunguz\">Bojan</a>, to name a few, for great kernels and discussions!</p>\n\n<p>This wouldn’t have been possible without awesome teammates, <a href=\"http://www.kaggle.com/ashish2123\">Ashish</a> and <a href=\"http://www.kaggle.com/rsrade\">Rahul</a>, who put in lot of effort and made this competition a great learning experience.</p>\n\n<p>Thanks for reading, I am planning to release the code in few days after cleaning it.    </p>\n\n<p>Edit: Github repository: <a href=\"https://github.com/soham97/Quora-Insincere-Questions-Classification-Challenge-NLP\">https://github.com/soham97/Quora-Insincere-Questions-Classification-Challenge-NLP</a></p>",
      "rawMarkdown": "Let me start by thanking everyone here who participated and contributed. We learned a lot of tips and tricks from community shared kernels and discussions. The competition was challenging in terms of finding correlated local validation, running the solution in 2 hours, and producing reproducible results, to name a few. We found a nearly correlated validation set before 1 week of the competition end. We used simple averaging where each trained NN model (Pytorch) was different either in terms of learning rate, pre-processing, embedding or architecture to maintain model diversity. One important thing we realised was: given small learning rate and large number of epochs, single FastText embedding based models were beating our glove+paragram models. So we focused on tuning FastText based models which could provide considerable score within 5-6 epochs. We also added Gaussian Noise to some models after embeddings to reduce the overdependence of RNN on specific keywords.\n\n**Solution summary:**\n\n - **Runtime**: 6352.2 secs\n - **Preprocessing**: Cleaning special characters, number pre-processing, misspell cleaning (For some models changed preprocessing sequence to add diversity)\n - **Embedding**: GLoVe, FastText, Paragram embeddings \n - **Neural Network architecture** (trained for 5-6 epochs with no fold): \n  - Stacked LSTM-GRU-128 hidden units, with GloVe+Paragram embedding\n  - Stacked LSTM-GRU 60 hidden units with attention and capsule, and GloVe+Paragram embedding\n  - Stacked LSTM-GRU-60 hidden units with GloVe+Paragram embedding\n  - Stacked LSTM-GRU-60 hidden units with FastText embeddings\n  - Stacked LSTM-GRU-80 hidden units with FastText embeddings and different preprocessing sequence \n - **Blending**: Averaging prediction of each model with linear regression coefficients \n\n**Things that did not work**\n\n - We tried pseudo labelling in different ways, but it didn't provide major boost considering its running time for it, so we dropped it in the end.\n - We tried variety of preprocessing techniques to no avail. All of them tend to decrease the LB score with slight improvement in cv. Fearing overfitting pre-processing to training data with kept it to minimum.\n - One trick that we tried was weight saving and retraining. For example, we trained the model and saved its weights before the model reached optimum. Then for the next model we loaded the weights for LSTM and GRU and did not pass gradients through them. This forced the new parts of model like an extra CNN layers or linear units to cover up for this. This saved time as the new model reached optimum within 2 epochs. But this did not add considerable benefit to the ensemble. In my opinion, majority of the information pertaining text was captured by RNN units, leaving little information required to be captured by newly added layers. Do let me know your thoughts on this experiments and its results.\n\nSpecial thanks to [Shujian][1], [Benjamin Minixhofer][2], [Dieter][3], [Ryches][4], [Bojan][5], to name a few, for great kernels and discussions!\n\nThis wouldn’t have been possible without awesome teammates, [Ashish][6] and [Rahul][7], who put in lot of effort and made this competition a great learning experience.\n\nThanks for reading, I am planning to release the code in few days after cleaning it. \t\n\nEdit: Github repository: https://github.com/soham97/Quora-Insincere-Questions-Classification-Challenge-NLP\n\n\n  [1]: http://www.kaggle.com/shujian\n  [2]: http://www.kaggle.com/bminixhofer\n  [3]: http://www.kaggle.com/christofhenkel\n  [4]: http://www.kaggle.com/ryches\n  [5]: http://www.kaggle.com/tunguz\n  [6]: http://www.kaggle.com/ashish2123\n  [7]: http://www.kaggle.com/rsrade",
      "votes": null
    },
    {
      "id": "471645",
      "postDate": "02/14/2019 17:58:58",
      "content": "<p>Congratulations @soham. Thanks for sharing the solution.</p>",
      "rawMarkdown": "Congratulations @soham. Thanks for sharing the solution.",
      "votes": null
    },
    {
      "id": "471866",
      "postDate": "02/15/2019 02:19:53",
      "content": "<p>Congratulations.</p>",
      "rawMarkdown": "Congratulations.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471645,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/14/2019 17:58:58",
      "content": "<p>Congratulations @soham. Thanks for sharing the solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471866,
      "author_name": "mamonaka",
      "author_url": "",
      "post_date": "02/15/2019 02:19:53",
      "content": "<p>Congratulations.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "471526": "Let me start by thanking everyone here who participated and contributed. We learned a lot of tips and tricks from community shared kernels and discussions. The competition was challenging in terms of finding correlated local validation, running the solution in 2 hours, and producing reproducible results, to name a few. We found a nearly correlated validation set before 1 week of the competition end. We used simple averaging where each trained NN model (Pytorch) was different either in terms of learning rate, pre-processing, embedding or architecture to maintain model diversity. One important thing we realised was: given small learning rate and large number of epochs, single FastText embedding based models were beating our glove+paragram models. So we focused on tuning FastText based models which could provide considerable score within 5-6 epochs. We also added Gaussian Noise to some models after embeddings to reduce the overdependence of RNN on specific keywords.\n\n**Solution summary:**\n\n - **Runtime**: 6352.2 secs\n - **Preprocessing**: Cleaning special characters, number pre-processing, misspell cleaning (For some models changed preprocessing sequence to add diversity)\n - **Embedding**: GLoVe, FastText, Paragram embeddings \n - **Neural Network architecture** (trained for 5-6 epochs with no fold): \n  - Stacked LSTM-GRU-128 hidden units, with GloVe+Paragram embedding\n  - Stacked LSTM-GRU 60 hidden units with attention and capsule, and GloVe+Paragram embedding\n  - Stacked LSTM-GRU-60 hidden units with GloVe+Paragram embedding\n  - Stacked LSTM-GRU-60 hidden units with FastText embeddings\n  - Stacked LSTM-GRU-80 hidden units with FastText embeddings and different preprocessing sequence \n - **Blending**: Averaging prediction of each model with linear regression coefficients \n\n**Things that did not work**\n\n - We tried pseudo labelling in different ways, but it didn't provide major boost considering its running time for it, so we dropped it in the end.\n - We tried variety of preprocessing techniques to no avail. All of them tend to decrease the LB score with slight improvement in cv. Fearing overfitting pre-processing to training data with kept it to minimum.\n - One trick that we tried was weight saving and retraining. For example, we trained the model and saved its weights before the model reached optimum. Then for the next model we loaded the weights for LSTM and GRU and did not pass gradients through them. This forced the new parts of model like an extra CNN layers or linear units to cover up for this. This saved time as the new model reached optimum within 2 epochs. But this did not add considerable benefit to the ensemble. In my opinion, majority of the information pertaining text was captured by RNN units, leaving little information required to be captured by newly added layers. Do let me know your thoughts on this experiments and its results.\n\nSpecial thanks to [Shujian][1], [Benjamin Minixhofer][2], [Dieter][3], [Ryches][4], [Bojan][5], to name a few, for great kernels and discussions!\n\nThis wouldn’t have been possible without awesome teammates, [Ashish][6] and [Rahul][7], who put in lot of effort and made this competition a great learning experience.\n\nThanks for reading, I am planning to release the code in few days after cleaning it. \t\n\nEdit: Github repository: https://github.com/soham97/Quora-Insincere-Questions-Classification-Challenge-NLP\n\n\n  [1]: http://www.kaggle.com/shujian\n  [2]: http://www.kaggle.com/bminixhofer\n  [3]: http://www.kaggle.com/christofhenkel\n  [4]: http://www.kaggle.com/ryches\n  [5]: http://www.kaggle.com/tunguz\n  [6]: http://www.kaggle.com/ashish2123\n  [7]: http://www.kaggle.com/rsrade",
    "471645": "Congratulations @soham. Thanks for sharing the solution.",
    "471866": "Congratulations."
  },
  "source": "meta"
}