{
  "id": 80514,
  "title": "22nd Solution - 6 Models and POS Tagging",
  "url": "/competitions/quora-insincere-questions-classification/writeups/superteam-bit-ly-2qbpvnv-22nd-solution-6-models-an",
  "author_name": "",
  "post_date": "2019-02-14T06:10:31.678754200Z",
  "votes": 44,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Thanks to everyone who participated in the awesome kernels and discussions that happened during this competition as well as my brilliant teammates.  Always great to have people to bounce ideas off of. </p>\n\n<p>Here is the link to our final solution: \n<a href=\"https://www.kaggle.com/ryches/22nd-place-solution-6-models-pos-tagging\">https://www.kaggle.com/ryches/22nd-place-solution-6-models-pos-tagging</a></p>\n\n<p>The guts of our solution was largely driven architected the same as the kernels we made public. </p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings</a></li>\n<li><a href=\"https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\">https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis</a></li>\n<li><a href=\"https://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also\">https://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also</a></li>\n<li><a href=\"https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip\">https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip</a></li>\n<li><a href=\"https://www.kaggle.com/christofhenkel/keras-starter\">https://www.kaggle.com/christofhenkel/keras-starter</a></li>\n<li><a href=\"https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\">https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis</a></li>\n</ol>\n\n<p>I have written a relatively comprehensive description of our entire solution in the link above, but to give a summary:</p>\n\n<p>In this competition we were able to train a total of 6 models for a total of 74 epochs. How did we fit so many epochs into our 2 hour limit? We filtered out the easy examples. <a href=\"/christofhenkel\">@christofhenkel</a> figured out by looking at the histogram of our predictions that within a few epochs our models had already confidently classified over 70 percent of our training samples. We trained a model really quickly in order to filter these easy questions. Once we threw those samples away we were able to train models just as accurately only using the 30 percent that remained. </p>\n\n<p>Now that we had this additional time we trained 5 models paired with different embeddings based on how they performed in our offline ensembling. Our hillclimbing found that the best combination with 5 models was:</p>\n\n<ul>\n<li>DPCNN with reversing and glove embeddings</li>\n<li>A bidirectional gru into an lstm with the glove embeddings. (this was very similar to what we used for the toxic comment challenge and was our strongest individual model here as well)</li>\n<li>a parrallel lstm and gru model w/glove embeddings</li>\n<li>parts of speech bidirectional lstm and gru model w/paragram embeddings</li>\n<li>parts of speech parallel lstm and gru model w/news embeddings</li>\n</ul>\n\n<p>These choices actually seemed to make some sense given that we have a CNN model, our strongest LSTM/GRU models, use our strongest embedding 3 times and use POS tagging as an augmentor/differentiator to our weaker embeddings. </p>\n\n<p>The POS models ended up doing worse individually but when ensembled significantly boosted our score. Our second submission used 8 models in total and still got a worse score than our 6 models with two of them being POS. If we did not do the filtering trick then we would not have enough time to do the POS tagging as it is relatively slow. I have a more detailed write-up of the POS models in the parts of speech disambiguation kernel I shared. </p>",
  "messages": [
    {
      "id": "471187",
      "postDate": "02/14/2019 06:10:31",
      "content": "<p>Thanks to everyone who participated in the awesome kernels and discussions that happened during this competition as well as my brilliant teammates.  Always great to have people to bounce ideas off of. </p>\n\n<p>Here is the link to our final solution: \n<a href=\"https://www.kaggle.com/ryches/22nd-place-solution-6-models-pos-tagging\">https://www.kaggle.com/ryches/22nd-place-solution-6-models-pos-tagging</a></p>\n\n<p>The guts of our solution was largely driven architected the same as the kernels we made public. </p>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings\">https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings</a></li>\n<li><a href=\"https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\">https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis</a></li>\n<li><a href=\"https://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also\">https://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also</a></li>\n<li><a href=\"https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip\">https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip</a></li>\n<li><a href=\"https://www.kaggle.com/christofhenkel/keras-starter\">https://www.kaggle.com/christofhenkel/keras-starter</a></li>\n<li><a href=\"https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\">https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis</a></li>\n</ol>\n\n<p>I have written a relatively comprehensive description of our entire solution in the link above, but to give a summary:</p>\n\n<p>In this competition we were able to train a total of 6 models for a total of 74 epochs. How did we fit so many epochs into our 2 hour limit? We filtered out the easy examples. <a href=\"/christofhenkel\">@christofhenkel</a> figured out by looking at the histogram of our predictions that within a few epochs our models had already confidently classified over 70 percent of our training samples. We trained a model really quickly in order to filter these easy questions. Once we threw those samples away we were able to train models just as accurately only using the 30 percent that remained. </p>\n\n<p>Now that we had this additional time we trained 5 models paired with different embeddings based on how they performed in our offline ensembling. Our hillclimbing found that the best combination with 5 models was:</p>\n\n<ul>\n<li>DPCNN with reversing and glove embeddings</li>\n<li>A bidirectional gru into an lstm with the glove embeddings. (this was very similar to what we used for the toxic comment challenge and was our strongest individual model here as well)</li>\n<li>a parrallel lstm and gru model w/glove embeddings</li>\n<li>parts of speech bidirectional lstm and gru model w/paragram embeddings</li>\n<li>parts of speech parallel lstm and gru model w/news embeddings</li>\n</ul>\n\n<p>These choices actually seemed to make some sense given that we have a CNN model, our strongest LSTM/GRU models, use our strongest embedding 3 times and use POS tagging as an augmentor/differentiator to our weaker embeddings. </p>\n\n<p>The POS models ended up doing worse individually but when ensembled significantly boosted our score. Our second submission used 8 models in total and still got a worse score than our 6 models with two of them being POS. If we did not do the filtering trick then we would not have enough time to do the POS tagging as it is relatively slow. I have a more detailed write-up of the POS models in the parts of speech disambiguation kernel I shared. </p>",
      "rawMarkdown": "Thanks to everyone who participated in the awesome kernels and discussions that happened during this competition as well as my brilliant teammates.  Always great to have people to bounce ideas off of. \n\nHere is the link to our final solution: \nhttps://www.kaggle.com/ryches/22nd-place-solution-6-models-pos-tagging\n\nThe guts of our solution was largely driven architected the same as the kernels we made public. \n\n 1. https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings\n 2. https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\n 3. https://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also\n 4. https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip\n 5. https://www.kaggle.com/christofhenkel/keras-starter\n 6. https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\n\nI have written a relatively comprehensive description of our entire solution in the link above, but to give a summary:\n\nIn this competition we were able to train a total of 6 models for a total of 74 epochs. How did we fit so many epochs into our 2 hour limit? We filtered out the easy examples. @christofhenkel figured out by looking at the histogram of our predictions that within a few epochs our models had already confidently classified over 70 percent of our training samples. We trained a model really quickly in order to filter these easy questions. Once we threw those samples away we were able to train models just as accurately only using the 30 percent that remained. \n\nNow that we had this additional time we trained 5 models paired with different embeddings based on how they performed in our offline ensembling. Our hillclimbing found that the best combination with 5 models was:\n\n* DPCNN with reversing and glove embeddings\n* A bidirectional gru into an lstm with the glove embeddings. (this was very similar to what we used for the toxic comment challenge and was our strongest individual model here as well)\n* a parrallel lstm and gru model w/glove embeddings\n* parts of speech bidirectional lstm and gru model w/paragram embeddings\n* parts of speech parallel lstm and gru model w/news embeddings\n\nThese choices actually seemed to make some sense given that we have a CNN model, our strongest LSTM/GRU models, use our strongest embedding 3 times and use POS tagging as an augmentor/differentiator to our weaker embeddings. \n\nThe POS models ended up doing worse individually but when ensembled significantly boosted our score. Our second submission used 8 models in total and still got a worse score than our 6 models with two of them being POS. If we did not do the filtering trick then we would not have enough time to do the POS tagging as it is relatively slow. I have a more detailed write-up of the POS models in the parts of speech disambiguation kernel I shared.",
      "votes": null
    },
    {
      "id": "471199",
      "postDate": "02/14/2019 06:19:12",
      "content": "<p>Congratulations ryches.</p>",
      "rawMarkdown": "Congratulations ryches.",
      "votes": null
    },
    {
      "id": "471203",
      "postDate": "02/14/2019 06:24:32",
      "content": "<p>I want to add that our Public LB is at #322. I knew that due to the very small size the Public LB is unreliable, so we sorely trusted the local cv, which was 0.7+ for the model presented here. Seems that this strategy paid off.</p>",
      "rawMarkdown": "I want to add that our Public LB is at #322. I knew that due to the very small size the Public LB is unreliable, so we sorely trusted the local cv, which was 0.7+ for the model presented here. Seems that this strategy paid off.",
      "votes": null
    },
    {
      "id": "471206",
      "postDate": "02/14/2019 06:29:38",
      "content": "<p>Yeah. I detailed this in the final kernel writeup along with the comparison between our 8 model and 6 model with POS. Looks like it isnt finished running though. We had to take a leap of faith that everyone was overfitting the leaderboard. </p>",
      "rawMarkdown": "Yeah. I detailed this in the final kernel writeup along with the comparison between our 8 model and 6 model with POS. Looks like it isnt finished running though. We had to take a leap of faith that everyone was overfitting the leaderboard.",
      "votes": null
    },
    {
      "id": "471208",
      "postDate": "02/14/2019 06:30:12",
      "content": "<p>Thank you. </p>",
      "rawMarkdown": "Thank you.",
      "votes": null
    },
    {
      "id": "471229",
      "postDate": "02/14/2019 07:00:14",
      "content": "<p>thanks for sharing, learn a lot from your team.</p>",
      "rawMarkdown": "thanks for sharing, learn a lot from your team.",
      "votes": null
    },
    {
      "id": "471269",
      "postDate": "02/14/2019 08:28:14",
      "content": "<p>Amazing</p>",
      "rawMarkdown": "Amazing",
      "votes": null
    },
    {
      "id": "471286",
      "postDate": "02/14/2019 08:47:23",
      "content": "<p>Congratulations and thank you for sharing your approach</p>",
      "rawMarkdown": "Congratulations and thank you for sharing your approach",
      "votes": null
    },
    {
      "id": "471387",
      "postDate": "02/14/2019 11:35:57",
      "content": "<p><a href=\"/ryches\">@ryches</a>, congrats and thanks for sharing your solution.</p>",
      "rawMarkdown": "ryches, congrats and thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "471768",
      "postDate": "02/14/2019 21:31:59",
      "content": "<p>Very Impressive, thank you for sharing.</p>",
      "rawMarkdown": "Very Impressive, thank you for sharing.",
      "votes": null
    },
    {
      "id": "477200",
      "postDate": "02/24/2019 05:23:58",
      "content": "<p>You use POS tags to separate different POS of the word, But your embedding vector is same with different POS. So I don't   know what I don't know what you mean by doing POS tags.</p>",
      "rawMarkdown": "You use POS tags to separate different POS of the word, But your embedding vector is same with different POS. So I don't   know what I don't know what you mean by doing POS tags.",
      "votes": null
    },
    {
      "id": "477201",
      "postDate": "02/24/2019 05:35:23",
      "content": "<p>They are initialized as the same and then are trained. I go into more detail in my pos disambiguation kernel</p>",
      "rawMarkdown": "They are initialized as the same and then are trained. I go into more detail in my pos disambiguation kernel",
      "votes": null
    },
    {
      "id": "477258",
      "postDate": "02/24/2019 08:28:12",
      "content": "<p>I see your method named train_model that model was not trained embedding matrix in epoch[1], but in epoch[2] it was be just trained embedding matrix. So the embedding matrix with pos tags will different without pos tags. Is my idea correct? </p>",
      "rawMarkdown": "I see your method named train_model that model was not trained embedding matrix in epoch[1], but in epoch[2] it was be just trained embedding matrix. So the embedding matrix with pos tags will different without pos tags. Is my idea correct?",
      "votes": null
    },
    {
      "id": "477259",
      "postDate": "02/24/2019 08:30:06",
      "content": "<p>Yes. You are correct</p>",
      "rawMarkdown": "Yes. You are correct",
      "votes": null
    },
    {
      "id": "477281",
      "postDate": "02/24/2019 09:26:33",
      "content": "<p>Thx your relpy. Your POS tags is great idea.</p>",
      "rawMarkdown": "Thx your relpy. Your POS tags is great idea.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471199,
      "author_name": "karthik7395",
      "author_url": "",
      "post_date": "02/14/2019 06:19:12",
      "content": "<p>Congratulations ryches.</p>",
      "votes": null,
      "replies": [
        {
          "id": 471208,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/14/2019 06:30:12",
          "content": "<p>Thank you. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471203,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "02/14/2019 06:24:32",
      "content": "<p>I want to add that our Public LB is at #322. I knew that due to the very small size the Public LB is unreliable, so we sorely trusted the local cv, which was 0.7+ for the model presented here. Seems that this strategy paid off.</p>",
      "votes": null,
      "replies": [
        {
          "id": 471206,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/14/2019 06:29:38",
          "content": "<p>Yeah. I detailed this in the final kernel writeup along with the comparison between our 8 model and 6 model with POS. Looks like it isnt finished running though. We had to take a leap of faith that everyone was overfitting the leaderboard. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 471269,
          "author_name": "gsdeepakkumar",
          "author_url": "",
          "post_date": "02/14/2019 08:28:14",
          "content": "<p>Amazing</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471229,
      "author_name": "konohayui",
      "author_url": "",
      "post_date": "02/14/2019 07:00:14",
      "content": "<p>thanks for sharing, learn a lot from your team.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471286,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "02/14/2019 08:47:23",
      "content": "<p>Congratulations and thank you for sharing your approach</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471387,
      "author_name": "sheriytm",
      "author_url": "",
      "post_date": "02/14/2019 11:35:57",
      "content": "<p><a href=\"/ryches\">@ryches</a>, congrats and thanks for sharing your solution.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 471768,
      "author_name": "danielv7",
      "author_url": "",
      "post_date": "02/14/2019 21:31:59",
      "content": "<p>Very Impressive, thank you for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 477200,
      "author_name": "salonsai",
      "author_url": "",
      "post_date": "02/24/2019 05:23:58",
      "content": "<p>You use POS tags to separate different POS of the word, But your embedding vector is same with different POS. So I don't   know what I don't know what you mean by doing POS tags.</p>",
      "votes": null,
      "replies": [
        {
          "id": 477201,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/24/2019 05:35:23",
          "content": "<p>They are initialized as the same and then are trained. I go into more detail in my pos disambiguation kernel</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477258,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "02/24/2019 08:28:12",
          "content": "<p>I see your method named train_model that model was not trained embedding matrix in epoch[1], but in epoch[2] it was be just trained embedding matrix. So the embedding matrix with pos tags will different without pos tags. Is my idea correct? </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477259,
          "author_name": "ryches",
          "author_url": "",
          "post_date": "02/24/2019 08:30:06",
          "content": "<p>Yes. You are correct</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 477281,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "02/24/2019 09:26:33",
          "content": "<p>Thx your relpy. Your POS tags is great idea.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "471187": "Thanks to everyone who participated in the awesome kernels and discussions that happened during this competition as well as my brilliant teammates.  Always great to have people to bounce ideas off of. \n\nHere is the link to our final solution: \nhttps://www.kaggle.com/ryches/22nd-place-solution-6-models-pos-tagging\n\nThe guts of our solution was largely driven architected the same as the kernels we made public. \n\n 1. https://www.kaggle.com/christofhenkel/how-to-preprocessing-when-using-embeddings\n 2. https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\n 3. https://www.kaggle.com/mihaskalic/lstm-is-all-you-need-well-maybe-embeddings-also\n 4. https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip\n 5. https://www.kaggle.com/christofhenkel/keras-starter\n 6. https://www.kaggle.com/ryches/parts-of-speech-disambiguation-error-analysis\n\nI have written a relatively comprehensive description of our entire solution in the link above, but to give a summary:\n\nIn this competition we were able to train a total of 6 models for a total of 74 epochs. How did we fit so many epochs into our 2 hour limit? We filtered out the easy examples. @christofhenkel figured out by looking at the histogram of our predictions that within a few epochs our models had already confidently classified over 70 percent of our training samples. We trained a model really quickly in order to filter these easy questions. Once we threw those samples away we were able to train models just as accurately only using the 30 percent that remained. \n\nNow that we had this additional time we trained 5 models paired with different embeddings based on how they performed in our offline ensembling. Our hillclimbing found that the best combination with 5 models was:\n\n* DPCNN with reversing and glove embeddings\n* A bidirectional gru into an lstm with the glove embeddings. (this was very similar to what we used for the toxic comment challenge and was our strongest individual model here as well)\n* a parrallel lstm and gru model w/glove embeddings\n* parts of speech bidirectional lstm and gru model w/paragram embeddings\n* parts of speech parallel lstm and gru model w/news embeddings\n\nThese choices actually seemed to make some sense given that we have a CNN model, our strongest LSTM/GRU models, use our strongest embedding 3 times and use POS tagging as an augmentor/differentiator to our weaker embeddings. \n\nThe POS models ended up doing worse individually but when ensembled significantly boosted our score. Our second submission used 8 models in total and still got a worse score than our 6 models with two of them being POS. If we did not do the filtering trick then we would not have enough time to do the POS tagging as it is relatively slow. I have a more detailed write-up of the POS models in the parts of speech disambiguation kernel I shared.",
    "471199": "Congratulations ryches.",
    "471203": "I want to add that our Public LB is at #322. I knew that due to the very small size the Public LB is unreliable, so we sorely trusted the local cv, which was 0.7+ for the model presented here. Seems that this strategy paid off.",
    "471206": "Yeah. I detailed this in the final kernel writeup along with the comparison between our 8 model and 6 model with POS. Looks like it isnt finished running though. We had to take a leap of faith that everyone was overfitting the leaderboard.",
    "471208": "Thank you.",
    "471229": "thanks for sharing, learn a lot from your team.",
    "471269": "Amazing",
    "471286": "Congratulations and thank you for sharing your approach",
    "471387": "ryches, congrats and thanks for sharing your solution.",
    "471768": "Very Impressive, thank you for sharing.",
    "477200": "You use POS tags to separate different POS of the word, But your embedding vector is same with different POS. So I don't   know what I don't know what you mean by doing POS tags.",
    "477201": "They are initialized as the same and then are trained. I go into more detail in my pos disambiguation kernel",
    "477258": "I see your method named train_model that model was not trained embedding matrix in epoch[1], but in epoch[2] it was be just trained embedding matrix. So the embedding matrix with pos tags will different without pos tags. Is my idea correct?",
    "477259": "Yes. You are correct",
    "477281": "Thx your relpy. Your POS tags is great idea."
  },
  "source": "meta"
}