{
  "id": 80542,
  "title": "25th place solution - unfreeze and tune embeddings!",
  "url": "/competitions/quora-insincere-questions-classification/writeups/danzell-25th-place-solution-unfreeze-and-tune-embe",
  "author_name": "",
  "post_date": "2021-01-04T09:42:53.603Z",
  "votes": 26,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi all - I tackled this competition in R &amp; keras. Right after stage 1 docker-images got updated and I had a really bad feeling. I am relieved now that all worked out!</p>\n<h2>Preprocessing:</h2>\n<p>I did some basic preprocessing (replacing common typos and separating special characters) – nothing special here (<a href=\"https://www.kaggle.com/springmanndaniel/preprocessing-in-r\">link to preproc kernel</a>)</p>\n<h2>Embeddings</h2>\n<p>I combined R and Python (Reticulate) to load and merge (GloVe + Para) pretrained embeddings. This way I could save some time. (<a href=\"https://www.kaggle.com/springmanndaniel/combine-r-and-python-to-load-embeddings\">link to embedding kernel</a>)<br>\nVocabulary was built on training-data only (196090 words).<br>\nAll words that did not appear in GloVe/ Para were replaced by zeroes. Towards the end of each models (last epoch) training phase I turned the embedding layer to trainable (see <strong>Boosting</strong>). This way each model overfitted a little bit + the model created some representation for missing words. </p>\n<h2>Keras Model</h2>\n<p>I used a single model architecture and trained it on 6 folds. The final ensemble was a simple average of the six runs. <br>\nThe model was a mix of LSTM, Convolutions and fully connected Layers.  (<a href=\"https://www.kaggle.com/springmanndaniel/25th-place-solution?scriptVersionId=10255060\">link to model kernel</a>)</p>\n<ul>\n<li>Epochs: <strong>4</strong></li>\n<li>Learning-rate: <strong>0.003, 0.003, 0.003, 0.001</strong></li>\n<li>Batch-size: <strong>512x2</strong> on epoch 1-3 and <strong>512x1.5</strong> on epoch 4</li>\n<li>Input sequence length: <strong>60</strong></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F11198a9e0b5f678a58ec8246948a3902%2Fmodel_arch.png?generation=1609753363950890&amp;alt=media\" alt=\"\"></p>\n<p><strong>CV SCORE: ~0.6977</strong>  <br>\n<strong>Fold: 1</strong> Val F1 Score: <strong>0.695</strong> Val Loss: 0.0954 best thresh: 0.4 (unknown words = 51432)</p>\n<p><strong>Fold: 2</strong> Val F1 Score: <strong>0.698</strong> Val Loss: 0.0942 best thresh: 0.38 (unknown words = 7871)</p>\n<p><strong>Fold: 3</strong> Val F1 Score: <strong>0.695</strong> Val Loss: 0.0952 best thresh: 0.38 (unknown words = 20)</p>\n<p><strong>Fold: 4</strong> Val F1 Score: <strong>0.695</strong> Val Loss: 0.093 best thresh: 0.36 (unknown words = 20)</p>\n<p><strong>Fold: 5</strong> Val F1 Score: <strong>0.701</strong> Val Loss: 0.0917 best thresh: 0.41 (unknown words = 20)</p>\n<p><strong>Fold: 6</strong> Val F1 Score: <strong>0.700</strong> Val Loss: 0.0943 best thresh: 0.36 (unknown words = 20)</p>\n<h2>Threshold</h2>\n<p>I calculated the threshold based on validation data. </p>\n<h2>Boost</h2>\n<p>What boosted my model most was unfreezing embeddings towards the end of each run and updating unknown words by their newly learned representations so that subsequent models  could utilize more words for training.<br>\nThis helped because each of the 6 models overfitted slightly - which added more diversity to the final ensemble &amp; helped the model to deal with unknown words.</p>\n<p>It looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F3dc4bde4aadf4c928bd9662cba72826f%2Funfreeze.png?generation=1609751408734952&amp;alt=media\" alt=\"\"></p>\n<p>Cheers<br>\ndan</p>\n<hr>\n<p>edit:</p>\n<ul>\n<li>fixed typos</li>\n<li>added keras network</li>\n</ul>",
  "messages": [
    {
      "id": "471316",
      "postDate": "02/14/2019 09:49:21",
      "content": "<p>Hi all - I tackled this competition in R &amp; keras. Right after stage 1 docker-images got updated and I had a really bad feeling. I am relieved now that all worked out!</p>\n<h2>Preprocessing:</h2>\n<p>I did some basic preprocessing (replacing common typos and separating special characters) – nothing special here (<a href=\"https://www.kaggle.com/springmanndaniel/preprocessing-in-r\">link to preproc kernel</a>)</p>\n<h2>Embeddings</h2>\n<p>I combined R and Python (Reticulate) to load and merge (GloVe + Para) pretrained embeddings. This way I could save some time. (<a href=\"https://www.kaggle.com/springmanndaniel/combine-r-and-python-to-load-embeddings\">link to embedding kernel</a>)<br>\nVocabulary was built on training-data only (196090 words).<br>\nAll words that did not appear in GloVe/ Para were replaced by zeroes. Towards the end of each models (last epoch) training phase I turned the embedding layer to trainable (see <strong>Boosting</strong>). This way each model overfitted a little bit + the model created some representation for missing words. </p>\n<h2>Keras Model</h2>\n<p>I used a single model architecture and trained it on 6 folds. The final ensemble was a simple average of the six runs. <br>\nThe model was a mix of LSTM, Convolutions and fully connected Layers.  (<a href=\"https://www.kaggle.com/springmanndaniel/25th-place-solution?scriptVersionId=10255060\">link to model kernel</a>)</p>\n<ul>\n<li>Epochs: <strong>4</strong></li>\n<li>Learning-rate: <strong>0.003, 0.003, 0.003, 0.001</strong></li>\n<li>Batch-size: <strong>512x2</strong> on epoch 1-3 and <strong>512x1.5</strong> on epoch 4</li>\n<li>Input sequence length: <strong>60</strong></li>\n</ul>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F11198a9e0b5f678a58ec8246948a3902%2Fmodel_arch.png?generation=1609753363950890&amp;alt=media\" alt=\"\"></p>\n<p><strong>CV SCORE: ~0.6977</strong>  <br>\n<strong>Fold: 1</strong> Val F1 Score: <strong>0.695</strong> Val Loss: 0.0954 best thresh: 0.4 (unknown words = 51432)</p>\n<p><strong>Fold: 2</strong> Val F1 Score: <strong>0.698</strong> Val Loss: 0.0942 best thresh: 0.38 (unknown words = 7871)</p>\n<p><strong>Fold: 3</strong> Val F1 Score: <strong>0.695</strong> Val Loss: 0.0952 best thresh: 0.38 (unknown words = 20)</p>\n<p><strong>Fold: 4</strong> Val F1 Score: <strong>0.695</strong> Val Loss: 0.093 best thresh: 0.36 (unknown words = 20)</p>\n<p><strong>Fold: 5</strong> Val F1 Score: <strong>0.701</strong> Val Loss: 0.0917 best thresh: 0.41 (unknown words = 20)</p>\n<p><strong>Fold: 6</strong> Val F1 Score: <strong>0.700</strong> Val Loss: 0.0943 best thresh: 0.36 (unknown words = 20)</p>\n<h2>Threshold</h2>\n<p>I calculated the threshold based on validation data. </p>\n<h2>Boost</h2>\n<p>What boosted my model most was unfreezing embeddings towards the end of each run and updating unknown words by their newly learned representations so that subsequent models  could utilize more words for training.<br>\nThis helped because each of the 6 models overfitted slightly - which added more diversity to the final ensemble &amp; helped the model to deal with unknown words.</p>\n<p>It looks like this:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F3dc4bde4aadf4c928bd9662cba72826f%2Funfreeze.png?generation=1609751408734952&amp;alt=media\" alt=\"\"></p>\n<p>Cheers<br>\ndan</p>\n<hr>\n<p>edit:</p>\n<ul>\n<li>fixed typos</li>\n<li>added keras network</li>\n</ul>",
      "rawMarkdown": "Hi all - I tackled this competition in R &amp; keras. Right after stage 1 docker-images got updated and I had a really bad feeling. I am relieved now that all worked out!\n\n## Preprocessing:\nI did some basic preprocessing (replacing common typos and separating special characters) – nothing special here (<a href=\"https://www.kaggle.com/springmanndaniel/preprocessing-in-r\">link to preproc kernel</a>)\n\n## Embeddings\nI combined R and Python (Reticulate) to load and merge (GloVe + Para) pretrained embeddings. This way I could save some time. (<a href=\"https://www.kaggle.com/springmanndaniel/combine-r-and-python-to-load-embeddings\">link to embedding kernel</a>)\nVocabulary was built on training-data only (196090 words).\nAll words that did not appear in GloVe/ Para were replaced by zeroes. Towards the end of each models (last epoch) training phase I turned the embedding layer to trainable (see **Boosting**). This way each model overfitted a little bit + the model created some representation for missing words. \n\n## Keras Model\nI used a single model architecture and trained it on 6 folds. The final ensemble was a simple average of the six runs. \nThe model was a mix of LSTM, Convolutions and fully connected Layers.  (<a href=\"https://www.kaggle.com/springmanndaniel/25th-place-solution?scriptVersionId=10255060\">link to model kernel</a>)\n- Epochs: **4**\n- Learning-rate: **0.003, 0.003, 0.003, 0.001**\n- Batch-size: **512x2** on epoch 1-3 and **512x1.5** on epoch 4\n- Input sequence length: **60**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F11198a9e0b5f678a58ec8246948a3902%2Fmodel_arch.png?generation=1609753363950890&alt=media)\n\n**CV SCORE: ~0.6977**  \n**Fold: 1** Val F1 Score: **0.695** Val Loss: 0.0954 best thresh: 0.4 (unknown words = 51432)\n\n**Fold: 2** Val F1 Score: **0.698** Val Loss: 0.0942 best thresh: 0.38 (unknown words = 7871)\n\n**Fold: 3** Val F1 Score: **0.695** Val Loss: 0.0952 best thresh: 0.38 (unknown words = 20)\n\n**Fold: 4** Val F1 Score: **0.695** Val Loss: 0.093 best thresh: 0.36 (unknown words = 20)\n\n**Fold: 5** Val F1 Score: **0.701** Val Loss: 0.0917 best thresh: 0.41 (unknown words = 20)\n\n**Fold: 6** Val F1 Score: **0.700** Val Loss: 0.0943 best thresh: 0.36 (unknown words = 20)\n\n\n## Threshold\nI calculated the threshold based on validation data. \n\n## Boost\nWhat boosted my model most was unfreezing embeddings towards the end of each run and updating unknown words by their newly learned representations so that subsequent models  could utilize more words for training.\nThis helped because each of the 6 models overfitted slightly - which added more diversity to the final ensemble &amp; helped the model to deal with unknown words.\n\nIt looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F3dc4bde4aadf4c928bd9662cba72826f%2Funfreeze.png?generation=1609751408734952&alt=media)\n\nCheers\ndan\n\n_____\nedit:\n- fixed typos\n- added keras network",
      "votes": null
    },
    {
      "id": "471320",
      "postDate": "02/14/2019 09:53:25",
      "content": "<p>Congratulations @danzell</p>",
      "rawMarkdown": "Congratulations @danzell",
      "votes": null
    },
    {
      "id": "471332",
      "postDate": "02/14/2019 10:03:11",
      "content": "<p>Happy your stuff ran within time! Congratz!</p>",
      "rawMarkdown": "Happy your stuff ran within time! Congratz!",
      "votes": null
    },
    {
      "id": "471334",
      "postDate": "02/14/2019 10:06:47",
      "content": "<p>Thanks - and congratulations on winning this competition!</p>",
      "rawMarkdown": "Thanks - and congratulations on winning this competition!",
      "votes": null
    },
    {
      "id": "471335",
      "postDate": "02/14/2019 10:07:16",
      "content": "<p>thank you! </p>",
      "rawMarkdown": "thank you!",
      "votes": null
    },
    {
      "id": "542200",
      "postDate": "06/03/2019 15:04:53",
      "content": "<p>Great work! The Keras model link is not working, is it possible to fix it?</p>",
      "rawMarkdown": "Great work! The Keras model link is not working, is it possible to fix it?",
      "votes": null
    },
    {
      "id": "542835",
      "postDate": "06/04/2019 06:15:31",
      "content": "<p>Fixed it. Sorry for the broken link. </p>\n\n<p>_edit:\ncreating and padding sequences within a dplyr-pipe might be the problem! The docker image version changed and runtime seems really bad - will dig into this later.</p>",
      "rawMarkdown": "Fixed it. Sorry for the broken link. \n\n_edit:\ncreating and padding sequences within a dplyr-pipe might be the problem! The docker image version changed and runtime seems really bad - will dig into this later.",
      "votes": null
    },
    {
      "id": "544772",
      "postDate": "06/05/2019 22:33:10",
      "content": "<p>No problem. Happy to see R solutions in NLP, so rare these days. Could you explain me why you had to do a python code for the load_glove function? Couldn't it have been done in R?</p>",
      "rawMarkdown": "No problem. Happy to see R solutions in NLP, so rare these days. Could you explain me why you had to do a python code for the load_glove function? Couldn't it have been done in R?",
      "votes": null
    },
    {
      "id": "544791",
      "postDate": "06/05/2019 23:15:45",
      "content": "<p>Time and memory was limited. Using reticulate + python reduced memory and was faster. Let me know if there is a \"better\" way :)</p>",
      "rawMarkdown": "Time and memory was limited. Using reticulate + python reduced memory and was faster. Let me know if there is a \"better\" way :)",
      "votes": null
    },
    {
      "id": "733658",
      "postDate": "01/31/2020 11:46:05",
      "content": "<p>nice!</p>",
      "rawMarkdown": "nice!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471320,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/14/2019 09:53:25",
      "content": "<p>Congratulations @danzell</p>",
      "votes": null,
      "replies": [
        {
          "id": 471335,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "02/14/2019 10:07:16",
          "content": "<p>thank you! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471332,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "02/14/2019 10:03:11",
      "content": "<p>Happy your stuff ran within time! Congratz!</p>",
      "votes": null,
      "replies": [
        {
          "id": 471334,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "02/14/2019 10:06:47",
          "content": "<p>Thanks - and congratulations on winning this competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 542200,
      "author_name": "daniellga",
      "author_url": "",
      "post_date": "06/03/2019 15:04:53",
      "content": "<p>Great work! The Keras model link is not working, is it possible to fix it?</p>",
      "votes": null,
      "replies": [
        {
          "id": 542835,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "06/04/2019 06:15:31",
          "content": "<p>Fixed it. Sorry for the broken link. </p>\n\n<p>_edit:\ncreating and padding sequences within a dplyr-pipe might be the problem! The docker image version changed and runtime seems really bad - will dig into this later.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544772,
          "author_name": "daniellga",
          "author_url": "",
          "post_date": "06/05/2019 22:33:10",
          "content": "<p>No problem. Happy to see R solutions in NLP, so rare these days. Could you explain me why you had to do a python code for the load_glove function? Couldn't it have been done in R?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 544791,
          "author_name": "springmanndaniel",
          "author_url": "",
          "post_date": "06/05/2019 23:15:45",
          "content": "<p>Time and memory was limited. Using reticulate + python reduced memory and was faster. Let me know if there is a \"better\" way :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 733658,
      "author_name": "flitzifast",
      "author_url": "",
      "post_date": "01/31/2020 11:46:05",
      "content": "<p>nice!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "471316": "Hi all - I tackled this competition in R &amp; keras. Right after stage 1 docker-images got updated and I had a really bad feeling. I am relieved now that all worked out!\n\n## Preprocessing:\nI did some basic preprocessing (replacing common typos and separating special characters) – nothing special here (<a href=\"https://www.kaggle.com/springmanndaniel/preprocessing-in-r\">link to preproc kernel</a>)\n\n## Embeddings\nI combined R and Python (Reticulate) to load and merge (GloVe + Para) pretrained embeddings. This way I could save some time. (<a href=\"https://www.kaggle.com/springmanndaniel/combine-r-and-python-to-load-embeddings\">link to embedding kernel</a>)\nVocabulary was built on training-data only (196090 words).\nAll words that did not appear in GloVe/ Para were replaced by zeroes. Towards the end of each models (last epoch) training phase I turned the embedding layer to trainable (see **Boosting**). This way each model overfitted a little bit + the model created some representation for missing words. \n\n## Keras Model\nI used a single model architecture and trained it on 6 folds. The final ensemble was a simple average of the six runs. \nThe model was a mix of LSTM, Convolutions and fully connected Layers.  (<a href=\"https://www.kaggle.com/springmanndaniel/25th-place-solution?scriptVersionId=10255060\">link to model kernel</a>)\n- Epochs: **4**\n- Learning-rate: **0.003, 0.003, 0.003, 0.001**\n- Batch-size: **512x2** on epoch 1-3 and **512x1.5** on epoch 4\n- Input sequence length: **60**\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F11198a9e0b5f678a58ec8246948a3902%2Fmodel_arch.png?generation=1609753363950890&alt=media)\n\n**CV SCORE: ~0.6977**  \n**Fold: 1** Val F1 Score: **0.695** Val Loss: 0.0954 best thresh: 0.4 (unknown words = 51432)\n\n**Fold: 2** Val F1 Score: **0.698** Val Loss: 0.0942 best thresh: 0.38 (unknown words = 7871)\n\n**Fold: 3** Val F1 Score: **0.695** Val Loss: 0.0952 best thresh: 0.38 (unknown words = 20)\n\n**Fold: 4** Val F1 Score: **0.695** Val Loss: 0.093 best thresh: 0.36 (unknown words = 20)\n\n**Fold: 5** Val F1 Score: **0.701** Val Loss: 0.0917 best thresh: 0.41 (unknown words = 20)\n\n**Fold: 6** Val F1 Score: **0.700** Val Loss: 0.0943 best thresh: 0.36 (unknown words = 20)\n\n\n## Threshold\nI calculated the threshold based on validation data. \n\n## Boost\nWhat boosted my model most was unfreezing embeddings towards the end of each run and updating unknown words by their newly learned representations so that subsequent models  could utilize more words for training.\nThis helped because each of the 6 models overfitted slightly - which added more diversity to the final ensemble &amp; helped the model to deal with unknown words.\n\nIt looks like this:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F606532%2F3dc4bde4aadf4c928bd9662cba72826f%2Funfreeze.png?generation=1609751408734952&alt=media)\n\nCheers\ndan\n\n_____\nedit:\n- fixed typos\n- added keras network",
    "471320": "Congratulations @danzell",
    "471332": "Happy your stuff ran within time! Congratz!",
    "471334": "Thanks - and congratulations on winning this competition!",
    "471335": "thank you!",
    "542200": "Great work! The Keras model link is not working, is it possible to fix it?",
    "542835": "Fixed it. Sorry for the broken link. \n\n_edit:\ncreating and padding sequences within a dplyr-pipe might be the problem! The docker image version changed and runtime seems really bad - will dig into this later.",
    "544772": "No problem. Happy to see R solutions in NLP, so rare these days. Could you explain me why you had to do a python code for the load_glove function? Couldn't it have been done in R?",
    "544791": "Time and memory was limited. Using reticulate + python reduced memory and was faster. Let me know if there is a \"better\" way :)",
    "733658": "nice!"
  },
  "source": "meta"
}