{
  "id": 80527,
  "title": "20th solution - 2 models, various embeds, mixed loss",
  "url": "/competitions/quora-insincere-questions-classification/writeups/i-have-no-idea-20th-solution-2-models-various-embe",
  "author_name": "",
  "post_date": "2019-02-15T08:22:09.907Z",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I was shocked when I saw the final standing. We never passed the 0.7 baseline on public LB and it was really frustrating. I basically gave up and just prayed this competition was about CV instead of public LB. It turned out to be true. I teamed up with <a href=\"https://www.kaggle.com/kukicap\">YangHe</a> in the last week and we decided to make two submissions: one focusing on CV and one focusing on public LB. I was responsible for the CV one. My final submission reached CV 0.698 and public LB 0.699.</p>\n\n<p>Anyway, this is my <a href=\"https://www.kaggle.com/jihangz/20th-solution-4-folds-2-models-mixed-loss\">kernel</a>.</p>\n\n<p><strong>Pre-processing:</strong></p>\n\n<p>Basically the public kernel, with some bug fixed (order of punc clean/contraction clean) and more contraction cleaning. I also used multiprocessing to speed things up. I met a bug when using Keras Tokenizer with PyTorch model: I couldn't set num words=None in the Tokenizer. It would run into some CUDA error during the training phase. So I fitted the Tokenizer locally and set num words = len(tokenizer.word_index) in the kernel.</p>\n\n<p><strong>Model:</strong></p>\n\n<p>I built two models:</p>\n\n<ol>\n<li><p>concat(GloVe, FastText) embedding + LSTM + TextCNN with kernel size [1, 2, 3, 4] + 2 dense layers, with some batch normalizations and dropout layers</p></li>\n<li><p>mean(GloVe, Para) embedding + LSTM + GRU + concat(GlobalAvgPool, GlobalMaxPool) + 2 dense layers, with some dropout layers</p></li>\n</ol>\n\n<p>We noticed the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79911\">bug</a> in the embedding dropout after the submission deadline.</p>\n\n<p><strong>Training:</strong></p>\n\n<p>I split the training data into 4 folds.</p>\n\n<p>Loss: BCE + soft F1 loss. I changed from BCE loss to this mixed loss on the last day and it gave me 0.003 boost on public LB and 0.002 boost on CV. It gave a stabler threshold v. F1 curve at the optimal point and I believe this granted us the 20th position. I also tried BCE + Lovasz, BCE pretrain and Lovasz fine-tune, BCE pretrain and soft f1 fine-tune, etc. Some of them didn't improve the model, others didn't converge at all. The model didn't converge when I was using pure soft F1 loss. This might be due to the imbalance of the label. Oversample might be needed when using soft F1 loss, but I didn't have the time to try. </p>\n\n<p>I used consine schedule with max LR = 0.003, and trained each model 4 epochs. I think consine schedule is better than step schedule and it is my favorite scheduler of all time. Notice that overfitting the training set a little would give a stabler threshold v. F1 curve. That's why all 0.7 public kernels overfit, \nI also tried AdamW with weight_decay = 0.0001, and it indeed gave better result. I didn't use it since it took more time to run.</p>\n\n<p><strong>Post-processing:</strong></p>\n\n<p>Average of all 8 classifiers and set the threshold based on oof prediction. I have made an all positive submission to figure out there were 3376 insincere questions in the public test data. I noticed that a lot of solutions to the past competition set the threshold so that the ratio of predicted label in test set is the same as training set. However, I didn't do that because I felt that it would be dangerous to use the same strategy in a binary classification problem.</p>\n\n<p><strong>Some Takeaways:</strong></p>\n\n<ol>\n<li><p>Stability is the key. You want threshold as insensitive as possible. </p></li>\n<li><p>Model is not the most important thing. The major variation is in the embedding layer.</p></li>\n<li><p>Read discussion, read public kernels, read solutions to similar past competitions, read solutions to different past competitions.</p></li>\n<li><p>When you fork someone's code, read it! It might not be bug-free!</p></li>\n<li><p>Don't give up! The shakeup is REAL!</p></li>\n</ol>",
  "messages": [
    {
      "id": "471251",
      "postDate": "02/14/2019 07:41:21",
      "content": "<p>I was shocked when I saw the final standing. We never passed the 0.7 baseline on public LB and it was really frustrating. I basically gave up and just prayed this competition was about CV instead of public LB. It turned out to be true. I teamed up with <a href=\"https://www.kaggle.com/kukicap\">YangHe</a> in the last week and we decided to make two submissions: one focusing on CV and one focusing on public LB. I was responsible for the CV one. My final submission reached CV 0.698 and public LB 0.699.</p>\n\n<p>Anyway, this is my <a href=\"https://www.kaggle.com/jihangz/20th-solution-4-folds-2-models-mixed-loss\">kernel</a>.</p>\n\n<p><strong>Pre-processing:</strong></p>\n\n<p>Basically the public kernel, with some bug fixed (order of punc clean/contraction clean) and more contraction cleaning. I also used multiprocessing to speed things up. I met a bug when using Keras Tokenizer with PyTorch model: I couldn't set num words=None in the Tokenizer. It would run into some CUDA error during the training phase. So I fitted the Tokenizer locally and set num words = len(tokenizer.word_index) in the kernel.</p>\n\n<p><strong>Model:</strong></p>\n\n<p>I built two models:</p>\n\n<ol>\n<li><p>concat(GloVe, FastText) embedding + LSTM + TextCNN with kernel size [1, 2, 3, 4] + 2 dense layers, with some batch normalizations and dropout layers</p></li>\n<li><p>mean(GloVe, Para) embedding + LSTM + GRU + concat(GlobalAvgPool, GlobalMaxPool) + 2 dense layers, with some dropout layers</p></li>\n</ol>\n\n<p>We noticed the <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79911\">bug</a> in the embedding dropout after the submission deadline.</p>\n\n<p><strong>Training:</strong></p>\n\n<p>I split the training data into 4 folds.</p>\n\n<p>Loss: BCE + soft F1 loss. I changed from BCE loss to this mixed loss on the last day and it gave me 0.003 boost on public LB and 0.002 boost on CV. It gave a stabler threshold v. F1 curve at the optimal point and I believe this granted us the 20th position. I also tried BCE + Lovasz, BCE pretrain and Lovasz fine-tune, BCE pretrain and soft f1 fine-tune, etc. Some of them didn't improve the model, others didn't converge at all. The model didn't converge when I was using pure soft F1 loss. This might be due to the imbalance of the label. Oversample might be needed when using soft F1 loss, but I didn't have the time to try. </p>\n\n<p>I used consine schedule with max LR = 0.003, and trained each model 4 epochs. I think consine schedule is better than step schedule and it is my favorite scheduler of all time. Notice that overfitting the training set a little would give a stabler threshold v. F1 curve. That's why all 0.7 public kernels overfit, \nI also tried AdamW with weight_decay = 0.0001, and it indeed gave better result. I didn't use it since it took more time to run.</p>\n\n<p><strong>Post-processing:</strong></p>\n\n<p>Average of all 8 classifiers and set the threshold based on oof prediction. I have made an all positive submission to figure out there were 3376 insincere questions in the public test data. I noticed that a lot of solutions to the past competition set the threshold so that the ratio of predicted label in test set is the same as training set. However, I didn't do that because I felt that it would be dangerous to use the same strategy in a binary classification problem.</p>\n\n<p><strong>Some Takeaways:</strong></p>\n\n<ol>\n<li><p>Stability is the key. You want threshold as insensitive as possible. </p></li>\n<li><p>Model is not the most important thing. The major variation is in the embedding layer.</p></li>\n<li><p>Read discussion, read public kernels, read solutions to similar past competitions, read solutions to different past competitions.</p></li>\n<li><p>When you fork someone's code, read it! It might not be bug-free!</p></li>\n<li><p>Don't give up! The shakeup is REAL!</p></li>\n</ol>",
      "rawMarkdown": "I was shocked when I saw the final standing. We never passed the 0.7 baseline on public LB and it was really frustrating. I basically gave up and just prayed this competition was about CV instead of public LB. It turned out to be true. I teamed up with [YangHe][1] in the last week and we decided to make two submissions: one focusing on CV and one focusing on public LB. I was responsible for the CV one. My final submission reached CV 0.698 and public LB 0.699.\n\nAnyway, this is my [kernel][2].\n\n\n**Pre-processing:**\n\nBasically the public kernel, with some bug fixed (order of punc clean/contraction clean) and more contraction cleaning. I also used multiprocessing to speed things up. I met a bug when using Keras Tokenizer with PyTorch model: I couldn't set num words=None in the Tokenizer. It would run into some CUDA error during the training phase. So I fitted the Tokenizer locally and set num words = len(tokenizer.word_index) in the kernel.\n\n\n**Model:**\n\nI built two models:\n\n1. concat(GloVe, FastText) embedding + LSTM + TextCNN with kernel size [1, 2, 3, 4] + 2 dense layers, with some batch normalizations and dropout layers\n\n2. mean(GloVe, Para) embedding + LSTM + GRU + concat(GlobalAvgPool, GlobalMaxPool) + 2 dense layers, with some dropout layers\n\nWe noticed the [bug][3] in the embedding dropout after the submission deadline.\n\n\n**Training:**\n\nI split the training data into 4 folds.\n\nLoss: BCE + soft F1 loss. I changed from BCE loss to this mixed loss on the last day and it gave me 0.003 boost on public LB and 0.002 boost on CV. It gave a stabler threshold v. F1 curve at the optimal point and I believe this granted us the 20th position. I also tried BCE + Lovasz, BCE pretrain and Lovasz fine-tune, BCE pretrain and soft f1 fine-tune, etc. Some of them didn't improve the model, others didn't converge at all. The model didn't converge when I was using pure soft F1 loss. This might be due to the imbalance of the label. Oversample might be needed when using soft F1 loss, but I didn't have the time to try. \n\nI used consine schedule with max LR = 0.003, and trained each model 4 epochs. I think consine schedule is better than step schedule and it is my favorite scheduler of all time. Notice that overfitting the training set a little would give a stabler threshold v. F1 curve. That's why all 0.7 public kernels overfit, \nI also tried AdamW with weight_decay = 0.0001, and it indeed gave better result. I didn't use it since it took more time to run.\n\n\n**Post-processing:**\n\nAverage of all 8 classifiers and set the threshold based on oof prediction. I have made an all positive submission to figure out there were 3376 insincere questions in the public test data. I noticed that a lot of solutions to the past competition set the threshold so that the ratio of predicted label in test set is the same as training set. However, I didn't do that because I felt that it would be dangerous to use the same strategy in a binary classification problem.\n\n\n**Some Takeaways:**\n\n1. Stability is the key. You want threshold as insensitive as possible. \n\n2. Model is not the most important thing. The major variation is in the embedding layer.\n\n3. Read discussion, read public kernels, read solutions to similar past competitions, read solutions to different past competitions.\n\n4. When you fork someone's code, read it! It might not be bug-free!\n\n5. Don't give up! The shakeup is REAL!\n\n  [1]: https://www.kaggle.com/kukicap\n  [2]: https://www.kaggle.com/jihangz/20th-solution-4-folds-2-models-mixed-loss\n  [3]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79911",
      "votes": null
    },
    {
      "id": "471290",
      "postDate": "02/14/2019 08:53:44",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!",
      "votes": null
    },
    {
      "id": "471326",
      "postDate": "02/14/2019 09:57:18",
      "content": "<p>Congrats Jihang Zhang, Thanks for sharing your takeaways.</p>",
      "rawMarkdown": "Congrats Jihang Zhang, Thanks for sharing your takeaways.",
      "votes": null
    },
    {
      "id": "471390",
      "postDate": "02/14/2019 11:39:20",
      "content": "<p><a href=\"/jihangz\">@jihangz</a>, congrats and thanks for sharing your solution.</p>",
      "rawMarkdown": "jihangz, congrats and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "471705",
      "postDate": "02/14/2019 19:37:21",
      "content": "<p>Congratulations on the high finish! and thanks for sharing.</p>",
      "rawMarkdown": "Congratulations on the high finish! and thanks for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471290,
      "author_name": "laevatein",
      "author_url": "",
      "post_date": "02/14/2019 08:53:44",
      "content": "<p>Congrats!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471326,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/14/2019 09:57:18",
      "content": "<p>Congrats Jihang Zhang, Thanks for sharing your takeaways.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471390,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "02/14/2019 11:39:20",
      "content": "<p><a href=\"/jihangz\">@jihangz</a>, congrats and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471705,
      "author_name": "mtodisco10",
      "author_url": "",
      "post_date": "02/14/2019 19:37:21",
      "content": "<p>Congratulations on the high finish! and thanks for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "471251": "I was shocked when I saw the final standing. We never passed the 0.7 baseline on public LB and it was really frustrating. I basically gave up and just prayed this competition was about CV instead of public LB. It turned out to be true. I teamed up with [YangHe][1] in the last week and we decided to make two submissions: one focusing on CV and one focusing on public LB. I was responsible for the CV one. My final submission reached CV 0.698 and public LB 0.699.\n\nAnyway, this is my [kernel][2].\n\n\n**Pre-processing:**\n\nBasically the public kernel, with some bug fixed (order of punc clean/contraction clean) and more contraction cleaning. I also used multiprocessing to speed things up. I met a bug when using Keras Tokenizer with PyTorch model: I couldn't set num words=None in the Tokenizer. It would run into some CUDA error during the training phase. So I fitted the Tokenizer locally and set num words = len(tokenizer.word_index) in the kernel.\n\n\n**Model:**\n\nI built two models:\n\n1. concat(GloVe, FastText) embedding + LSTM + TextCNN with kernel size [1, 2, 3, 4] + 2 dense layers, with some batch normalizations and dropout layers\n\n2. mean(GloVe, Para) embedding + LSTM + GRU + concat(GlobalAvgPool, GlobalMaxPool) + 2 dense layers, with some dropout layers\n\nWe noticed the [bug][3] in the embedding dropout after the submission deadline.\n\n\n**Training:**\n\nI split the training data into 4 folds.\n\nLoss: BCE + soft F1 loss. I changed from BCE loss to this mixed loss on the last day and it gave me 0.003 boost on public LB and 0.002 boost on CV. It gave a stabler threshold v. F1 curve at the optimal point and I believe this granted us the 20th position. I also tried BCE + Lovasz, BCE pretrain and Lovasz fine-tune, BCE pretrain and soft f1 fine-tune, etc. Some of them didn't improve the model, others didn't converge at all. The model didn't converge when I was using pure soft F1 loss. This might be due to the imbalance of the label. Oversample might be needed when using soft F1 loss, but I didn't have the time to try. \n\nI used consine schedule with max LR = 0.003, and trained each model 4 epochs. I think consine schedule is better than step schedule and it is my favorite scheduler of all time. Notice that overfitting the training set a little would give a stabler threshold v. F1 curve. That's why all 0.7 public kernels overfit, \nI also tried AdamW with weight_decay = 0.0001, and it indeed gave better result. I didn't use it since it took more time to run.\n\n\n**Post-processing:**\n\nAverage of all 8 classifiers and set the threshold based on oof prediction. I have made an all positive submission to figure out there were 3376 insincere questions in the public test data. I noticed that a lot of solutions to the past competition set the threshold so that the ratio of predicted label in test set is the same as training set. However, I didn't do that because I felt that it would be dangerous to use the same strategy in a binary classification problem.\n\n\n**Some Takeaways:**\n\n1. Stability is the key. You want threshold as insensitive as possible. \n\n2. Model is not the most important thing. The major variation is in the embedding layer.\n\n3. Read discussion, read public kernels, read solutions to similar past competitions, read solutions to different past competitions.\n\n4. When you fork someone's code, read it! It might not be bug-free!\n\n5. Don't give up! The shakeup is REAL!\n\n  [1]: https://www.kaggle.com/kukicap\n  [2]: https://www.kaggle.com/jihangz/20th-solution-4-folds-2-models-mixed-loss\n  [3]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79911",
    "471290": "Congrats!",
    "471326": "Congrats Jihang Zhang, Thanks for sharing your takeaways.",
    "471390": "jihangz, congrats and thanks for sharing your solution.",
    "471705": "Congratulations on the high finish! and thanks for sharing."
  },
  "source": "meta"
}