{
  "id": 74214,
  "title": "Improvements reported and what worked for me",
  "url": "/competitions/quora-insincere-questions-classification/discussion/74214",
  "author_name": "",
  "post_date": "2018-12-10T05:33:25.203460Z",
  "votes": 129,
  "comment_count": 64,
  "views": 0,
  "content": "<p>This competition has been challenging with fine tuning models in limited runtime limits and chasing the techniques that give absolute and generalized improvements. This thread is dedicated to improvements on what works and what doesn't based on the community's discussion over discussion threads and kernels. </p>\n\n<h1>Improvements reported</h1>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424011\">K-Fold with 5/4 splits and 3-5 epochs works for many</a> </li>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778#423590\">Concatenation (embeddings) has slightly better results but averaging is more efficient</a> <a href=\"/shujian\">@shujian</a></li>\n<li><a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\">It is the preprocessing I use for my current LB score, and it has helped improving it by a bit</a> <a href=\"/theoviel\">@theoviel</a></li>\n<li><a href=\"https://www.kaggle.com/gmhost/gru-capsule#435536\">Capsule - Compare to gru+cnn, improve a little bit within the same time</a> <a href=\"/gmhost\">@gmhost</a></li>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74150#436197\">I ran an experiment by following this paper, BiLSTM-CNNs-CRF, without any fine tuning, the f1 score is around at 0.684</a>  <a href=\"/konohayui\">@konohayui</a></li>\n<li><a href=\"https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#418819\">Conv1D works slightly better (maybe they work the same, and my results were just random noise). Anyways it doesn't make sense to use Conv2D since the embeddings have no ordering</a>  <a href=\"/christofhenkel\">@christofhenkel</a></li>\n</ol>\n\n<h1>My reportings</h1>\n\n<ol>\n<li>Single model LSTM/GRU scores 0.700 now</li>\n<li>Bit of Text cleaning</li>\n<li>Using custom layers from public kernels improves score</li>\n<li>Updating embeddings with lower case words <strong>doesn't show</strong> improvements (Made CV instable to LB)</li>\n<li>Common miss spelled words replacement also shows <strong>no</strong> improvements</li>\n<li>Capsule improves the score a bit. </li>\n<li>Including Meta-Features <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74152#436001\">Related Thread</a>,  <a href=\"https://www.kaggle.com/shaz13/feature-engineering-for-nlp-classification/notebook\">Kernel</a>  (No major improvement) </li>\n</ol>\n\n<p>Peace</p>",
  "messages": [
    {
      "id": "436309",
      "postDate": "12/10/2018 05:33:25",
      "content": "<p>This competition has been challenging with fine tuning models in limited runtime limits and chasing the techniques that give absolute and generalized improvements. This thread is dedicated to improvements on what works and what doesn't based on the community's discussion over discussion threads and kernels. </p>\n\n<h1>Improvements reported</h1>\n\n<ol>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424011\">K-Fold with 5/4 splits and 3-5 epochs works for many</a> </li>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778#423590\">Concatenation (embeddings) has slightly better results but averaging is more efficient</a> <a href=\"/shujian\">@shujian</a></li>\n<li><a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\">It is the preprocessing I use for my current LB score, and it has helped improving it by a bit</a> <a href=\"/theoviel\">@theoviel</a></li>\n<li><a href=\"https://www.kaggle.com/gmhost/gru-capsule#435536\">Capsule - Compare to gru+cnn, improve a little bit within the same time</a> <a href=\"/gmhost\">@gmhost</a></li>\n<li><a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74150#436197\">I ran an experiment by following this paper, BiLSTM-CNNs-CRF, without any fine tuning, the f1 score is around at 0.684</a>  <a href=\"/konohayui\">@konohayui</a></li>\n<li><a href=\"https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#418819\">Conv1D works slightly better (maybe they work the same, and my results were just random noise). Anyways it doesn't make sense to use Conv2D since the embeddings have no ordering</a>  <a href=\"/christofhenkel\">@christofhenkel</a></li>\n</ol>\n\n<h1>My reportings</h1>\n\n<ol>\n<li>Single model LSTM/GRU scores 0.700 now</li>\n<li>Bit of Text cleaning</li>\n<li>Using custom layers from public kernels improves score</li>\n<li>Updating embeddings with lower case words <strong>doesn't show</strong> improvements (Made CV instable to LB)</li>\n<li>Common miss spelled words replacement also shows <strong>no</strong> improvements</li>\n<li>Capsule improves the score a bit. </li>\n<li>Including Meta-Features <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74152#436001\">Related Thread</a>,  <a href=\"https://www.kaggle.com/shaz13/feature-engineering-for-nlp-classification/notebook\">Kernel</a>  (No major improvement) </li>\n</ol>\n\n<p>Peace</p>",
      "rawMarkdown": "This competition has been challenging with fine tuning models in limited runtime limits and chasing the techniques that give absolute and generalized improvements. This thread is dedicated to improvements on what works and what doesn't based on the community's discussion over discussion threads and kernels. \n\n\n\n# Improvements reported\n\n1. [K-Fold with 5/4 splits and 3-5 epochs works for many][1] \n2. [Concatenation (embeddings) has slightly better results but averaging is more efficient][2] @shujian\n3. [It is the preprocessing I use for my current LB score, and it has helped improving it by a bit][3] @theoviel\n4. [Capsule - Compare to gru+cnn, improve a little bit within the same time][4] @gmhost\n5. [I ran an experiment by following this paper, BiLSTM-CNNs-CRF, without any fine tuning, the f1 score is around at 0.684][5]  @konohayui\n6.  [Conv1D works slightly better (maybe they work the same, and my results were just random noise). Anyways it doesn't make sense to use Conv2D since the embeddings have no ordering][6]  @christofhenkel\n\n\n# My reportings\n1. Single model LSTM/GRU scores 0.700 now\n2. Bit of Text cleaning\n3. Using custom layers from public kernels improves score\n4. Updating embeddings with lower case words **doesn't show** improvements (Made CV instable to LB)\n5. Common miss spelled words replacement also shows **no** improvements\n6. Capsule improves the score a bit. \n7. Including Meta-Features [Related Thread][7],  [Kernel][8]  (No major improvement) \n\n\n\n\nPeace\n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424011\n  [2]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778#423590\n  [3]: https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\n  [4]: https://www.kaggle.com/gmhost/gru-capsule#435536\n  [5]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74150#436197\n  [6]: https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#418819\n  [7]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74152#436001\n  [8]: https://www.kaggle.com/shaz13/feature-engineering-for-nlp-classification/notebook",
      "votes": null
    },
    {
      "id": "436315",
      "postDate": "12/10/2018 05:46:36",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "436320",
      "postDate": "12/10/2018 05:57:49",
      "content": "<p>Nothing compared to your contributions. Thanks :) </p>",
      "rawMarkdown": "Nothing compared to your contributions. Thanks :)",
      "votes": null
    },
    {
      "id": "436325",
      "postDate": "12/10/2018 06:02:15",
      "content": "<p>Reason lower case doesn't work is because most of the embeddings don't use lower case words, except(I guess fasttext), I achieved 0.697 with text cleaning and shuijan's 4folds +clr kernel. Can't replicate the result or improve them as of now even with the same  kernel.</p>",
      "rawMarkdown": "Reason lower case doesn't work is because most of the embeddings don't use lower case words, except(I guess fasttext), I achieved 0.697 with text cleaning and shuijan's 4folds +clr kernel. Can't replicate the result or improve them as of now even with the same  kernel.",
      "votes": null
    },
    {
      "id": "436326",
      "postDate": "12/10/2018 06:07:03",
      "content": "<p>Yeah. I figured that out too. One strange thing I have observed is that any processing makes my KFold CV and LB instable. </p>\n\n<blockquote>\n  <p>Can't replicate the result or improve them as of now even with the same kernel.</p>\n</blockquote>\n\n<p>Keras is unstable and random. I am moving to Pytorch </p>",
      "rawMarkdown": "Yeah. I figured that out too. One strange thing I have observed is that any processing makes my KFold CV and LB instable. \n\n&gt; Can't replicate the result or improve them as of now even with the same kernel.\n\nKeras is unstable and random. I am moving to Pytorch",
      "votes": null
    },
    {
      "id": "436330",
      "postDate": "12/10/2018 06:20:15",
      "content": "<p>Thanks for sharing this info Shahebaz! What is the local validation for some of these findings?</p>",
      "rawMarkdown": "Thanks for sharing this info Shahebaz! What is the local validation for some of these findings?",
      "votes": null
    },
    {
      "id": "436331",
      "postDate": "12/10/2018 06:22:45",
      "content": "<p>Around 0.03 ish on average. Although, I cannot establish benchmark due to non-determinism and randomness with keras. I am porting to Pytorch (still new with it) and later I can make some conclusions :) </p>",
      "rawMarkdown": "Around 0.03 ish on average. Although, I cannot establish benchmark due to non-determinism and randomness with keras. I am porting to Pytorch (still new with it) and later I can make some conclusions :)",
      "votes": null
    },
    {
      "id": "436343",
      "postDate": "12/10/2018 06:37:33",
      "content": "<p>Thanks for the report :) is the 0.03 a typo? 0.93?</p>",
      "rawMarkdown": "Thanks for the report :) is the 0.03 a typo? 0.93?",
      "votes": null
    },
    {
      "id": "436347",
      "postDate": "12/10/2018 06:40:20",
      "content": "<p>I meant -- 0.696 (+- 0.03) ish.. </p>",
      "rawMarkdown": "I meant -- 0.696 (+- 0.03) ish..",
      "votes": null
    },
    {
      "id": "436370",
      "postDate": "12/10/2018 07:32:39",
      "content": "<p>I am also moving to PyTorch. Let see how it turns out.</p>",
      "rawMarkdown": "I am also moving to PyTorch. Let see how it turns out.",
      "votes": null
    },
    {
      "id": "436394",
      "postDate": "12/10/2018 08:20:57",
      "content": "<p>Pytorch with CUDNN is as unstable as Keras.</p>",
      "rawMarkdown": "Pytorch with CUDNN is as unstable as Keras.",
      "votes": null
    },
    {
      "id": "436395",
      "postDate": "12/10/2018 08:22:26",
      "content": "<p>It appears to me that single model with CV is better than blending of different models.</p>",
      "rawMarkdown": "It appears to me that single model with CV is better than blending of different models.",
      "votes": null
    },
    {
      "id": "436402",
      "postDate": "12/10/2018 08:34:15",
      "content": "<blockquote>\n  <p>Including Meta-Features Related Thread, Kernel (Still testing..)</p>\n</blockquote>\n\n<p>I tried to embed some counting features, but noticed no improvement. I guess it's because these features are really easy to learn by a RNN. </p>\n\n<p>May be some very special ( or magic ^^ ) features may help, given the model will not likely be able to learn them quickly in 2 hrs. </p>",
      "rawMarkdown": "&gt; Including Meta-Features Related Thread, Kernel (Still testing..)\n\nI tried to embed some counting features, but noticed no improvement. I guess it's because these features are really easy to learn by a RNN. \n\nMay be some very special ( or magic ^^ ) features may help, given the model will not likely be able to learn them quickly in 2 hrs.",
      "votes": null
    },
    {
      "id": "436441",
      "postDate": "12/10/2018 10:12:16",
      "content": "<p>Great thread, keep us updated !</p>",
      "rawMarkdown": "Great thread, keep us updated !",
      "votes": null
    },
    {
      "id": "436648",
      "postDate": "12/10/2018 17:39:43",
      "content": "<p>IMO, a single robust model with good CV generalises better than - simple train-split ensembles. Anyways, the key in this competition is to tune a model to finest in most optimized time. If one achieves that then we  can get to work on diverse models for ensembling. </p>",
      "rawMarkdown": "IMO, a single robust model with good CV generalises better than - simple train-split ensembles. Anyways, the key in this competition is to tune a model to finest in most optimized time. If one achieves that then we  can get to work on diverse models for ensembling.",
      "votes": null
    },
    {
      "id": "436671",
      "postDate": "12/10/2018 18:14:23",
      "content": "<p>Well put :)</p>",
      "rawMarkdown": "Well put :)",
      "votes": null
    },
    {
      "id": "437001",
      "postDate": "12/11/2018 08:25:52",
      "content": "<p>great work!</p>",
      "rawMarkdown": "great work!",
      "votes": null
    },
    {
      "id": "437002",
      "postDate": "12/11/2018 08:27:32",
      "content": "<p>Thanks. Do you mind sharing few tips that got you in 0.7 game? Perhaps any hints :) </p>",
      "rawMarkdown": "Thanks. Do you mind sharing few tips that got you in 0.7 game? Perhaps any hints :)",
      "votes": null
    },
    {
      "id": "437015",
      "postDate": "12/11/2018 08:53:08",
      "content": "<p>I got the score at 0.7 with blend .  I run the same kernel ,but it only got 0.698.\n  Some details:1) concat emb not mean; 2)max_feature=None when Tokenizer;3)some statistic feature</p>",
      "rawMarkdown": "I got the score at 0.7 with blend .  I run the same kernel ,but it only got 0.698.\n  Some details:1) concat emb not mean; 2)max_feature=None when Tokenizer;3)some statistic feature",
      "votes": null
    },
    {
      "id": "437027",
      "postDate": "12/11/2018 09:15:56",
      "content": "<p>Thank you for summarizing things so well. :)</p>",
      "rawMarkdown": "Thank you for summarizing things so well. :)",
      "votes": null
    },
    {
      "id": "437043",
      "postDate": "12/11/2018 09:41:13",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing.",
      "votes": null
    },
    {
      "id": "437049",
      "postDate": "12/11/2018 09:52:17",
      "content": "<p>I tried to do LDA for some extra feature engineering (eg using the doc-topic-distributions), but turns out this is a waste of time. Impossible to train a topic model with high coherence with the limited time, and with the topic models I tested (bear in mind still had to leave time for Kfold) did not improve results. Magic numbers preprocessing improved results the most.</p>",
      "rawMarkdown": "I tried to do LDA for some extra feature engineering (eg using the doc-topic-distributions), but turns out this is a waste of time. Impossible to train a topic model with high coherence with the limited time, and with the topic models I tested (bear in mind still had to leave time for Kfold) did not improve results. Magic numbers preprocessing improved results the most.",
      "votes": null
    },
    {
      "id": "437076",
      "postDate": "12/11/2018 10:31:04",
      "content": "<blockquote>\n  <p>I run the same kernel ,but it only got 0.698.</p>\n</blockquote>\n\n<p>That's randomness and non-determism we are facing partially with kernel runtimes and also keras. Thanks for sharing! </p>",
      "rawMarkdown": "&gt;  I run the same kernel ,but it only got 0.698.\n\nThat's randomness and non-determism we are facing partially with kernel runtimes and also keras. Thanks for sharing!",
      "votes": null
    },
    {
      "id": "437081",
      "postDate": "12/11/2018 10:50:13",
      "content": "<p>attention cannot improve f1-score in your model?</p>",
      "rawMarkdown": "attention cannot improve f1-score in your model?",
      "votes": null
    },
    {
      "id": "437087",
      "postDate": "12/11/2018 11:08:47",
      "content": "<p>I improved my score to LB 0.701 with some text processing (I noticed an improvement on CV too ) </p>",
      "rawMarkdown": "I improved my score to LB 0.701 with some text processing (I noticed an improvement on CV too )",
      "votes": null
    },
    {
      "id": "437091",
      "postDate": "12/11/2018 11:12:40",
      "content": "<p>That's great. Btw, I have sent you a DM. Let me know if there is an open place up  there </p>",
      "rawMarkdown": "That's great. Btw, I have sent you a DM. Let me know if there is an open place up  there",
      "votes": null
    },
    {
      "id": "437493",
      "postDate": "12/12/2018 02:53:04",
      "content": "<p>Hi Serigne, can you tell us what's your CV score when you got LB 0.701? Thanks :)</p>",
      "rawMarkdown": "Hi Serigne, can you tell us what's your CV score when you got LB 0.701? Thanks :)",
      "votes": null
    },
    {
      "id": "437498",
      "postDate": "12/12/2018 03:12:42",
      "content": "<p>Maybe set the seed is useful.</p>",
      "rawMarkdown": "Maybe set the seed is useful.",
      "votes": null
    },
    {
      "id": "437569",
      "postDate": "12/12/2018 06:21:24",
      "content": "<p>Thank for sharing</p>",
      "rawMarkdown": "Thank for sharing",
      "votes": null
    },
    {
      "id": "437583",
      "postDate": "12/12/2018 06:38:42",
      "content": "<p>I improved my score to LB 0.701 with pseudo-tags, 3 models</p>",
      "rawMarkdown": "I improved my score to LB 0.701 with pseudo-tags, 3 models",
      "votes": null
    },
    {
      "id": "437714",
      "postDate": "12/12/2018 10:57:24",
      "content": "<p>I tried it yesterday, but could not get an ideal result. Maybe I do it wrong...</p>",
      "rawMarkdown": "I tried it yesterday, but could not get an ideal result. Maybe I do it wrong...",
      "votes": null
    },
    {
      "id": "437768",
      "postDate": "12/12/2018 13:14:15",
      "content": "<p>great summary!</p>",
      "rawMarkdown": "great summary!",
      "votes": null
    },
    {
      "id": "437808",
      "postDate": "12/12/2018 14:29:35",
      "content": "<p>Working on some babel  and meta features. Hopefully they will provide a lift.</p>",
      "rawMarkdown": "Working on some babel  and meta features. Hopefully they will provide a lift.",
      "votes": null
    },
    {
      "id": "438642",
      "postDate": "12/14/2018 01:18:30",
      "content": "<p>what does pseudo-tags mean?</p>",
      "rawMarkdown": "what does pseudo-tags mean?",
      "votes": null
    },
    {
      "id": "438836",
      "postDate": "12/14/2018 09:00:50",
      "content": "<p>This is great, thanks for sharing!</p>",
      "rawMarkdown": "This is great, thanks for sharing!",
      "votes": null
    },
    {
      "id": "439259",
      "postDate": "12/15/2018 02:57:22",
      "content": "<p>Hello，chi zhu, my hometown is also chongqing, i am a student at uestc , does your team need a teammate?</p>",
      "rawMarkdown": "Hello，chi zhu, my hometown is also chongqing, i am a student at uestc , does your team need a teammate?",
      "votes": null
    },
    {
      "id": "439537",
      "postDate": "12/15/2018 18:37:41",
      "content": "<p>welcome</p>",
      "rawMarkdown": "welcome",
      "votes": null
    },
    {
      "id": "439546",
      "postDate": "12/15/2018 19:01:48",
      "content": "<p>Yeah. It does. But, attention is so common these days and popular among kernel so I thought everyone might already are using it. </p>",
      "rawMarkdown": "Yeah. It does. But, attention is so common these days and popular among kernel so I thought everyone might already are using it.",
      "votes": null
    },
    {
      "id": "439547",
      "postDate": "12/15/2018 19:02:02",
      "content": "<p>Thanks! </p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "439548",
      "postDate": "12/15/2018 19:03:22",
      "content": "<p>Yeah. I confirm this. I mode to pytorch. This is unstable too. Pinged few kaggle team members. Lets hear them  out if they can run kernels multiple times to take average. </p>",
      "rawMarkdown": "Yeah. I confirm this. I mode to pytorch. This is unstable too. Pinged few kaggle team members. Lets hear them  out if they can run kernels multiple times to take average.",
      "votes": null
    },
    {
      "id": "439566",
      "postDate": "12/15/2018 20:05:41",
      "content": "<p>Would this be fair for someone who has figured out how to run a more stable solution within the time frame though? Instead of average I would be open to the idea of taking the lowest score of multiple runs though.</p>",
      "rawMarkdown": "Would this be fair for someone who has figured out how to run a more stable solution within the time frame though? Instead of average I would be open to the idea of taking the lowest score of multiple runs though.",
      "votes": null
    },
    {
      "id": "440228",
      "postDate": "12/17/2018 08:56:18",
      "content": "<p>what is pseudo-tags?</p>",
      "rawMarkdown": "what is pseudo-tags?",
      "votes": null
    },
    {
      "id": "440400",
      "postDate": "12/17/2018 14:00:03",
      "content": "<p>The basic idea is to train a good model on the training set, then let it predict the labels of the test set. Now add the test set to your training set, using the predicted labels as \"ground truth\". Train again. Repeat.</p>\n\n<p>It may seem counter-intuitive, but works well apparently. Here are some links:</p>\n\n<ul>\n<li><a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a> (original paper)</li>\n<li><a href=\"https://towardsdatascience.com/simple-explanation-of-semi-supervised-learning-and-pseudo-labeling-c2218e8c769b\">https://towardsdatascience.com/simple-explanation-of-semi-supervised-learning-and-pseudo-labeling-c2218e8c769b</a></li>\n</ul>",
      "rawMarkdown": "The basic idea is to train a good model on the training set, then let it predict the labels of the test set. Now add the test set to your training set, using the predicted labels as \"ground truth\". Train again. Repeat.\n\nIt may seem counter-intuitive, but works well apparently. Here are some links:\n\n- http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf (original paper)\n- https://towardsdatascience.com/simple-explanation-of-semi-supervised-learning-and-pseudo-labeling-c2218e8c769b",
      "votes": null
    },
    {
      "id": "440417",
      "postDate": "12/17/2018 14:25:07",
      "content": "<p>Great thanks, Max</p>",
      "rawMarkdown": "Great thanks, Max",
      "votes": null
    },
    {
      "id": "440917",
      "postDate": "12/18/2018 05:00:08",
      "content": "<p>Thanks Max</p>",
      "rawMarkdown": "Thanks Max",
      "votes": null
    },
    {
      "id": "440961",
      "postDate": "12/18/2018 05:57:04",
      "content": "<p>Pseudo-tagging is a risky choice , I found it's hard to finsh in 2 hrs.  Maybe single model with CV is a better choice , I'm on it .</p>",
      "rawMarkdown": "Pseudo-tagging is a risky choice , I found it's hard to finsh in 2 hrs.  Maybe single model with CV is a better choice , I'm on it .",
      "votes": null
    },
    {
      "id": "441098",
      "postDate": "12/18/2018 09:01:38",
      "content": "<p>When go to stage 2 Pseudo-tagging cost more time </p>",
      "rawMarkdown": "When go to stage 2 Pseudo-tagging cost more time",
      "votes": null
    },
    {
      "id": "442825",
      "postDate": "12/20/2018 15:21:40",
      "content": "<p>I find that some model cannot improve LB score with pseudo-tags, beacuse train set has enough samples to train model and stage 1 data is small. </p>",
      "rawMarkdown": "I find that some model cannot improve LB score with pseudo-tags, beacuse train set has enough samples to train model and stage 1 data is small.",
      "votes": null
    },
    {
      "id": "444032",
      "postDate": "12/23/2018 01:24:02",
      "content": "<p>check the discussion above</p>",
      "rawMarkdown": "check the discussion above",
      "votes": null
    },
    {
      "id": "444568",
      "postDate": "12/24/2018 09:40:02",
      "content": "<p>Really nice work!!</p>",
      "rawMarkdown": "Really nice work!!",
      "votes": null
    },
    {
      "id": "445767",
      "postDate": "12/27/2018 03:20:13",
      "content": "<p>I know LB means leaderboard, but what does CV mean? Thanks.</p>",
      "rawMarkdown": "I know LB means leaderboard, but what does CV mean? Thanks.",
      "votes": null
    },
    {
      "id": "445909",
      "postDate": "12/27/2018 07:40:32",
      "content": "<p>what is IMO?</p>",
      "rawMarkdown": "what is IMO?",
      "votes": null
    },
    {
      "id": "445919",
      "postDate": "12/27/2018 08:07:02",
      "content": "<p>Cross Validation, local CV means local cross validation score over n-folds.</p>",
      "rawMarkdown": "Cross Validation, local CV means local cross validation score over n-folds.",
      "votes": null
    },
    {
      "id": "446056",
      "postDate": "12/27/2018 12:32:57",
      "content": "<p>In My Opinion</p>",
      "rawMarkdown": "In My Opinion",
      "votes": null
    },
    {
      "id": "446458",
      "postDate": "12/28/2018 05:22:39",
      "content": "<p>(⊙o⊙)thx</p>",
      "rawMarkdown": "(⊙o⊙)thx",
      "votes": null
    },
    {
      "id": "446717",
      "postDate": "12/28/2018 15:02:26",
      "content": "<p>Great summary! I was wondering how can one include pseudo-tags as their train data? I've downloaded my predictions and added it as a new dataset but the kernel counts it as \"external data\". Any guide on this?</p>",
      "rawMarkdown": "Great summary! I was wondering how can one include pseudo-tags as their train data? I've downloaded my predictions and added it as a new dataset but the kernel counts it as \"external data\". Any guide on this?",
      "votes": null
    },
    {
      "id": "446729",
      "postDate": "12/28/2018 15:22:28",
      "content": "<p>You can't load any additional data that is not given by the kernel. So anything you need has to be created in the kernel, in the two hours time frame.</p>\n\n<p>So the only way to use pseudo-tags is to train one model, let it predict the pseudo-tags, then train another model on those.</p>",
      "rawMarkdown": "You can't load any additional data that is not given by the kernel. So anything you need has to be created in the kernel, in the two hours time frame.\n\nSo the only way to use pseudo-tags is to train one model, let it predict the pseudo-tags, then train another model on those.",
      "votes": null
    },
    {
      "id": "447013",
      "postDate": "12/29/2018 01:52:15",
      "content": "<p>I see, thanks for your response.</p>",
      "rawMarkdown": "I see, thanks for your response.",
      "votes": null
    },
    {
      "id": "447228",
      "postDate": "12/29/2018 12:02:20",
      "content": "<p>But that means that the idea could work pretty well in stage 2. Then the test data would have a larger number of 1's and hence score might improve by adding them.</p>",
      "rawMarkdown": "But that means that the idea could work pretty well in stage 2. Then the test data would have a larger number of 1's and hence score might improve by adding them.",
      "votes": null
    },
    {
      "id": "447619",
      "postDate": "12/30/2018 07:13:51",
      "content": "<p>I tried to get the middle layers of NN model and train with lightgbm rather than fc layers. However it is too slow to finish in 2 hours haha : ).</p>",
      "rawMarkdown": "I tried to get the middle layers of NN model and train with lightgbm rather than fc layers. However it is too slow to finish in 2 hours haha : ).",
      "votes": null
    },
    {
      "id": "448074",
      "postDate": "12/31/2018 07:16:00",
      "content": "<p>ok, thank you.</p>",
      "rawMarkdown": "ok, thank you.",
      "votes": null
    },
    {
      "id": "448119",
      "postDate": "12/31/2018 09:26:21",
      "content": "<p>Did you try it?\nDid it complete within 2 hours?</p>",
      "rawMarkdown": "Did you try it?\nDid it complete within 2 hours?",
      "votes": null
    },
    {
      "id": "448226",
      "postDate": "12/31/2018 14:24:24",
      "content": "<p>Of course, it doesn't take long to predict on the small test set. However, I personally have not seen any improvements, though some here have reported they did.</p>",
      "rawMarkdown": "Of course, it doesn't take long to predict on the small test set. However, I personally have not seen any improvements, though some here have reported they did.",
      "votes": null
    },
    {
      "id": "452254",
      "postDate": "01/08/2019 13:01:55",
      "content": "<p><code>\nUpdating embeddings with lower case words doesn't show improvements (Made CV instable to LB)\n</code>\nI use Keras to build deep models, I was wondering how to update embeddings with lower case words? Could you please show me the demo or some materials? Many thanks!</p>",
      "rawMarkdown": "```\nUpdating embeddings with lower case words doesn't show improvements (Made CV instable to LB)\n```\nI use Keras to build deep models, I was wondering how to update embeddings with lower case words? Could you please show me the demo or some materials? Many thanks!",
      "votes": null
    },
    {
      "id": "452287",
      "postDate": "01/08/2019 13:57:12",
      "content": "<p>Check out the add_lower function in the kernel <a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\">https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing</a></p>",
      "rawMarkdown": "Check out the add_lower function in the kernel https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 436315,
      "author_name": "shujian",
      "author_url": "",
      "post_date": "12/10/2018 05:46:36",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": [
        {
          "id": 436320,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/10/2018 05:57:49",
          "content": "<p>Nothing compared to your contributions. Thanks :) </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436325,
      "author_name": "satian",
      "author_url": "",
      "post_date": "12/10/2018 06:02:15",
      "content": "<p>Reason lower case doesn't work is because most of the embeddings don't use lower case words, except(I guess fasttext), I achieved 0.697 with text cleaning and shuijan's 4folds +clr kernel. Can't replicate the result or improve them as of now even with the same  kernel.</p>",
      "votes": null,
      "replies": [
        {
          "id": 436326,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/10/2018 06:07:03",
          "content": "<p>Yeah. I figured that out too. One strange thing I have observed is that any processing makes my KFold CV and LB instable. </p>\n\n<blockquote>\n  <p>Can't replicate the result or improve them as of now even with the same kernel.</p>\n</blockquote>\n\n<p>Keras is unstable and random. I am moving to Pytorch </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436370,
          "author_name": "satian",
          "author_url": "",
          "post_date": "12/10/2018 07:32:39",
          "content": "<p>I am also moving to PyTorch. Let see how it turns out.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436394,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/10/2018 08:20:57",
          "content": "<p>Pytorch with CUDNN is as unstable as Keras.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437498,
          "author_name": "xiaobai1123q",
          "author_url": "",
          "post_date": "12/12/2018 03:12:42",
          "content": "<p>Maybe set the seed is useful.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439548,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/15/2018 19:03:22",
          "content": "<p>Yeah. I confirm this. I mode to pytorch. This is unstable too. Pinged few kaggle team members. Lets hear them  out if they can run kernels multiple times to take average. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439566,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/15/2018 20:05:41",
          "content": "<p>Would this be fair for someone who has figured out how to run a more stable solution within the time frame though? Instead of average I would be open to the idea of taking the lowest score of multiple runs though.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436330,
      "author_name": "rdizzl3",
      "author_url": "",
      "post_date": "12/10/2018 06:20:15",
      "content": "<p>Thanks for sharing this info Shahebaz! What is the local validation for some of these findings?</p>",
      "votes": null,
      "replies": [
        {
          "id": 436331,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/10/2018 06:22:45",
          "content": "<p>Around 0.03 ish on average. Although, I cannot establish benchmark due to non-determinism and randomness with keras. I am porting to Pytorch (still new with it) and later I can make some conclusions :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436343,
          "author_name": "rdizzl3",
          "author_url": "",
          "post_date": "12/10/2018 06:37:33",
          "content": "<p>Thanks for the report :) is the 0.03 a typo? 0.93?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436347,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/10/2018 06:40:20",
          "content": "<p>I meant -- 0.696 (+- 0.03) ish.. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436395,
      "author_name": "philippsinger",
      "author_url": "",
      "post_date": "12/10/2018 08:22:26",
      "content": "<p>It appears to me that single model with CV is better than blending of different models.</p>",
      "votes": null,
      "replies": [
        {
          "id": 436648,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/10/2018 17:39:43",
          "content": "<p>IMO, a single robust model with good CV generalises better than - simple train-split ensembles. Anyways, the key in this competition is to tune a model to finest in most optimized time. If one achieves that then we  can get to work on diverse models for ensembling. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 436671,
          "author_name": "philippsinger",
          "author_url": "",
          "post_date": "12/10/2018 18:14:23",
          "content": "<p>Well put :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 445909,
          "author_name": "chenshengabc",
          "author_url": "",
          "post_date": "12/27/2018 07:40:32",
          "content": "<p>what is IMO?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446056,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "12/27/2018 12:32:57",
          "content": "<p>In My Opinion</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 446458,
          "author_name": "chenshengabc",
          "author_url": "",
          "post_date": "12/28/2018 05:22:39",
          "content": "<p>(⊙o⊙)thx</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 436402,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/10/2018 08:34:15",
      "content": "<blockquote>\n  <p>Including Meta-Features Related Thread, Kernel (Still testing..)</p>\n</blockquote>\n\n<p>I tried to embed some counting features, but noticed no improvement. I guess it's because these features are really easy to learn by a RNN. </p>\n\n<p>May be some very special ( or magic ^^ ) features may help, given the model will not likely be able to learn them quickly in 2 hrs. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 436441,
      "author_name": "theoviel",
      "author_url": "",
      "post_date": "12/10/2018 10:12:16",
      "content": "<p>Great thread, keep us updated !</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 437001,
      "author_name": "chizhu2018",
      "author_url": "",
      "post_date": "12/11/2018 08:25:52",
      "content": "<p>great work!</p>",
      "votes": null,
      "replies": [
        {
          "id": 437002,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/11/2018 08:27:32",
          "content": "<p>Thanks. Do you mind sharing few tips that got you in 0.7 game? Perhaps any hints :) </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437015,
          "author_name": "chizhu2018",
          "author_url": "",
          "post_date": "12/11/2018 08:53:08",
          "content": "<p>I got the score at 0.7 with blend .  I run the same kernel ,but it only got 0.698.\n  Some details:1) concat emb not mean; 2)max_feature=None when Tokenizer;3)some statistic feature</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437043,
          "author_name": "lmdeliangmi",
          "author_url": "",
          "post_date": "12/11/2018 09:41:13",
          "content": "<p>Thanks for sharing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437076,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/11/2018 10:31:04",
          "content": "<blockquote>\n  <p>I run the same kernel ,but it only got 0.698.</p>\n</blockquote>\n\n<p>That's randomness and non-determism we are facing partially with kernel runtimes and also keras. Thanks for sharing! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 439259,
          "author_name": "chenhao2334",
          "author_url": "",
          "post_date": "12/15/2018 02:57:22",
          "content": "<p>Hello，chi zhu, my hometown is also chongqing, i am a student at uestc , does your team need a teammate?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437027,
      "author_name": "shubhamjain28",
      "author_url": "",
      "post_date": "12/11/2018 09:15:56",
      "content": "<p>Thank you for summarizing things so well. :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 437049,
      "author_name": "eeeedev",
      "author_url": "",
      "post_date": "12/11/2018 09:52:17",
      "content": "<p>I tried to do LDA for some extra feature engineering (eg using the doc-topic-distributions), but turns out this is a waste of time. Impossible to train a topic model with high coherence with the limited time, and with the topic models I tested (bear in mind still had to leave time for Kfold) did not improve results. Magic numbers preprocessing improved results the most.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 437081,
      "author_name": "salonsai",
      "author_url": "",
      "post_date": "12/11/2018 10:50:13",
      "content": "<p>attention cannot improve f1-score in your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 439546,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/15/2018 19:01:48",
          "content": "<p>Yeah. It does. But, attention is so common these days and popular among kernel so I thought everyone might already are using it. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437087,
      "author_name": "serigne",
      "author_url": "",
      "post_date": "12/11/2018 11:08:47",
      "content": "<p>I improved my score to LB 0.701 with some text processing (I noticed an improvement on CV too ) </p>",
      "votes": null,
      "replies": [
        {
          "id": 437091,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/11/2018 11:12:40",
          "content": "<p>That's great. Btw, I have sent you a DM. Let me know if there is an open place up  there </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 437493,
          "author_name": "tw1994",
          "author_url": "",
          "post_date": "12/12/2018 02:53:04",
          "content": "<p>Hi Serigne, can you tell us what's your CV score when you got LB 0.701? Thanks :)</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437569,
      "author_name": "phucbb",
      "author_url": "",
      "post_date": "12/12/2018 06:21:24",
      "content": "<p>Thank for sharing</p>",
      "votes": null,
      "replies": [
        {
          "id": 439537,
          "author_name": "",
          "author_url": "",
          "post_date": "12/15/2018 18:37:41",
          "content": "<p>welcome</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437583,
      "author_name": "mr007rin",
      "author_url": "",
      "post_date": "12/12/2018 06:38:42",
      "content": "<p>I improved my score to LB 0.701 with pseudo-tags, 3 models</p>",
      "votes": null,
      "replies": [
        {
          "id": 437714,
          "author_name": "lmdeliangmi",
          "author_url": "",
          "post_date": "12/12/2018 10:57:24",
          "content": "<p>I tried it yesterday, but could not get an ideal result. Maybe I do it wrong...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 438642,
          "author_name": "genyuan",
          "author_url": "",
          "post_date": "12/14/2018 01:18:30",
          "content": "<p>what does pseudo-tags mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440400,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/17/2018 14:00:03",
          "content": "<p>The basic idea is to train a good model on the training set, then let it predict the labels of the test set. Now add the test set to your training set, using the predicted labels as \"ground truth\". Train again. Repeat.</p>\n\n<p>It may seem counter-intuitive, but works well apparently. Here are some links:</p>\n\n<ul>\n<li><a href=\"http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf\">http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf</a> (original paper)</li>\n<li><a href=\"https://towardsdatascience.com/simple-explanation-of-semi-supervised-learning-and-pseudo-labeling-c2218e8c769b\">https://towardsdatascience.com/simple-explanation-of-semi-supervised-learning-and-pseudo-labeling-c2218e8c769b</a></li>\n</ul>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440417,
          "author_name": "genyuan",
          "author_url": "",
          "post_date": "12/17/2018 14:25:07",
          "content": "<p>Great thanks, Max</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440917,
          "author_name": "msafi04",
          "author_url": "",
          "post_date": "12/18/2018 05:00:08",
          "content": "<p>Thanks Max</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 440961,
          "author_name": "mr007rin",
          "author_url": "",
          "post_date": "12/18/2018 05:57:04",
          "content": "<p>Pseudo-tagging is a risky choice , I found it's hard to finsh in 2 hrs.  Maybe single model with CV is a better choice , I'm on it .</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 441098,
          "author_name": "tw1994",
          "author_url": "",
          "post_date": "12/18/2018 09:01:38",
          "content": "<p>When go to stage 2 Pseudo-tagging cost more time </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 442825,
          "author_name": "salonsai",
          "author_url": "",
          "post_date": "12/20/2018 15:21:40",
          "content": "<p>I find that some model cannot improve LB score with pseudo-tags, beacuse train set has enough samples to train model and stage 1 data is small. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447228,
          "author_name": "mlwhiz",
          "author_url": "",
          "post_date": "12/29/2018 12:02:20",
          "content": "<p>But that means that the idea could work pretty well in stage 2. Then the test data would have a larger number of 1's and hence score might improve by adding them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437768,
      "author_name": "iamgroot42",
      "author_url": "",
      "post_date": "12/12/2018 13:14:15",
      "content": "<p>great summary!</p>",
      "votes": null,
      "replies": [
        {
          "id": 439547,
          "author_name": "shaz13",
          "author_url": "",
          "post_date": "12/15/2018 19:02:02",
          "content": "<p>Thanks! </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 437808,
      "author_name": "venkateshradhakrish",
      "author_url": "",
      "post_date": "12/12/2018 14:29:35",
      "content": "<p>Working on some babel  and meta features. Hopefully they will provide a lift.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 438836,
      "author_name": "justinbt",
      "author_url": "",
      "post_date": "12/14/2018 09:00:50",
      "content": "<p>This is great, thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 440228,
      "author_name": "msafi04",
      "author_url": "",
      "post_date": "12/17/2018 08:56:18",
      "content": "<p>what is pseudo-tags?</p>",
      "votes": null,
      "replies": [
        {
          "id": 444032,
          "author_name": "yabutaka",
          "author_url": "",
          "post_date": "12/23/2018 01:24:02",
          "content": "<p>check the discussion above</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 444568,
      "author_name": "gsreyan",
      "author_url": "",
      "post_date": "12/24/2018 09:40:02",
      "content": "<p>Really nice work!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 445767,
      "author_name": "theredarrow",
      "author_url": "",
      "post_date": "12/27/2018 03:20:13",
      "content": "<p>I know LB means leaderboard, but what does CV mean? Thanks.</p>",
      "votes": null,
      "replies": [
        {
          "id": 445919,
          "author_name": "satian",
          "author_url": "",
          "post_date": "12/27/2018 08:07:02",
          "content": "<p>Cross Validation, local CV means local cross validation score over n-folds.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448074,
          "author_name": "theredarrow",
          "author_url": "",
          "post_date": "12/31/2018 07:16:00",
          "content": "<p>ok, thank you.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 446717,
      "author_name": "joonl04",
      "author_url": "",
      "post_date": "12/28/2018 15:02:26",
      "content": "<p>Great summary! I was wondering how can one include pseudo-tags as their train data? I've downloaded my predictions and added it as a new dataset but the kernel counts it as \"external data\". Any guide on this?</p>",
      "votes": null,
      "replies": [
        {
          "id": 446729,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/28/2018 15:22:28",
          "content": "<p>You can't load any additional data that is not given by the kernel. So anything you need has to be created in the kernel, in the two hours time frame.</p>\n\n<p>So the only way to use pseudo-tags is to train one model, let it predict the pseudo-tags, then train another model on those.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 447013,
          "author_name": "joonl04",
          "author_url": "",
          "post_date": "12/29/2018 01:52:15",
          "content": "<p>I see, thanks for your response.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448119,
          "author_name": "gsreyan",
          "author_url": "",
          "post_date": "12/31/2018 09:26:21",
          "content": "<p>Did you try it?\nDid it complete within 2 hours?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 448226,
          "author_name": "mschumacher",
          "author_url": "",
          "post_date": "12/31/2018 14:24:24",
          "content": "<p>Of course, it doesn't take long to predict on the small test set. However, I personally have not seen any improvements, though some here have reported they did.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 447619,
      "author_name": "kaggleczs",
      "author_url": "",
      "post_date": "12/30/2018 07:13:51",
      "content": "<p>I tried to get the middle layers of NN model and train with lightgbm rather than fc layers. However it is too slow to finish in 2 hours haha : ).</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 452254,
      "author_name": "sunnymarkliu",
      "author_url": "",
      "post_date": "01/08/2019 13:01:55",
      "content": "<p><code>\nUpdating embeddings with lower case words doesn't show improvements (Made CV instable to LB)\n</code>\nI use Keras to build deep models, I was wondering how to update embeddings with lower case words? Could you please show me the demo or some materials? Many thanks!</p>",
      "votes": null,
      "replies": [
        {
          "id": 452287,
          "author_name": "gsreyan",
          "author_url": "",
          "post_date": "01/08/2019 13:57:12",
          "content": "<p>Check out the add_lower function in the kernel <a href=\"https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\">https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "436309": "This competition has been challenging with fine tuning models in limited runtime limits and chasing the techniques that give absolute and generalized improvements. This thread is dedicated to improvements on what works and what doesn't based on the community's discussion over discussion threads and kernels. \n\n\n\n# Improvements reported\n\n1. [K-Fold with 5/4 splits and 3-5 epochs works for many][1] \n2. [Concatenation (embeddings) has slightly better results but averaging is more efficient][2] @shujian\n3. [It is the preprocessing I use for my current LB score, and it has helped improving it by a bit][3] @theoviel\n4. [Capsule - Compare to gru+cnn, improve a little bit within the same time][4] @gmhost\n5. [I ran an experiment by following this paper, BiLSTM-CNNs-CRF, without any fine tuning, the f1 score is around at 0.684][5]  @konohayui\n6.  [Conv1D works slightly better (maybe they work the same, and my results were just random noise). Anyways it doesn't make sense to use Conv2D since the embeddings have no ordering][6]  @christofhenkel\n\n\n# My reportings\n1. Single model LSTM/GRU scores 0.700 now\n2. Bit of Text cleaning\n3. Using custom layers from public kernels improves score\n4. Updating embeddings with lower case words **doesn't show** improvements (Made CV instable to LB)\n5. Common miss spelled words replacement also shows **no** improvements\n6. Capsule improves the score a bit. \n7. Including Meta-Features [Related Thread][7],  [Kernel][8]  (No major improvement) \n\n\n\n\nPeace\n\n\n  [1]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71946#424011\n  [2]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/71778#423590\n  [3]: https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing\n  [4]: https://www.kaggle.com/gmhost/gru-capsule#435536\n  [5]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74150#436197\n  [6]: https://www.kaggle.com/christofhenkel/inceptioncnn-with-flip#418819\n  [7]: https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/74152#436001\n  [8]: https://www.kaggle.com/shaz13/feature-engineering-for-nlp-classification/notebook",
    "436315": "Thanks for sharing!",
    "436320": "Nothing compared to your contributions. Thanks :)",
    "436325": "Reason lower case doesn't work is because most of the embeddings don't use lower case words, except(I guess fasttext), I achieved 0.697 with text cleaning and shuijan's 4folds +clr kernel. Can't replicate the result or improve them as of now even with the same  kernel.",
    "436326": "Yeah. I figured that out too. One strange thing I have observed is that any processing makes my KFold CV and LB instable. \n\n&gt; Can't replicate the result or improve them as of now even with the same kernel.\n\nKeras is unstable and random. I am moving to Pytorch",
    "436330": "Thanks for sharing this info Shahebaz! What is the local validation for some of these findings?",
    "436331": "Around 0.03 ish on average. Although, I cannot establish benchmark due to non-determinism and randomness with keras. I am porting to Pytorch (still new with it) and later I can make some conclusions :)",
    "436343": "Thanks for the report :) is the 0.03 a typo? 0.93?",
    "436347": "I meant -- 0.696 (+- 0.03) ish..",
    "436370": "I am also moving to PyTorch. Let see how it turns out.",
    "436394": "Pytorch with CUDNN is as unstable as Keras.",
    "436395": "It appears to me that single model with CV is better than blending of different models.",
    "436402": "&gt; Including Meta-Features Related Thread, Kernel (Still testing..)\n\nI tried to embed some counting features, but noticed no improvement. I guess it's because these features are really easy to learn by a RNN. \n\nMay be some very special ( or magic ^^ ) features may help, given the model will not likely be able to learn them quickly in 2 hrs.",
    "436441": "Great thread, keep us updated !",
    "436648": "IMO, a single robust model with good CV generalises better than - simple train-split ensembles. Anyways, the key in this competition is to tune a model to finest in most optimized time. If one achieves that then we  can get to work on diverse models for ensembling.",
    "436671": "Well put :)",
    "437001": "great work!",
    "437002": "Thanks. Do you mind sharing few tips that got you in 0.7 game? Perhaps any hints :)",
    "437015": "I got the score at 0.7 with blend .  I run the same kernel ,but it only got 0.698.\n  Some details:1) concat emb not mean; 2)max_feature=None when Tokenizer;3)some statistic feature",
    "437027": "Thank you for summarizing things so well. :)",
    "437043": "Thanks for sharing.",
    "437049": "I tried to do LDA for some extra feature engineering (eg using the doc-topic-distributions), but turns out this is a waste of time. Impossible to train a topic model with high coherence with the limited time, and with the topic models I tested (bear in mind still had to leave time for Kfold) did not improve results. Magic numbers preprocessing improved results the most.",
    "437076": "&gt;  I run the same kernel ,but it only got 0.698.\n\nThat's randomness and non-determism we are facing partially with kernel runtimes and also keras. Thanks for sharing!",
    "437081": "attention cannot improve f1-score in your model?",
    "437087": "I improved my score to LB 0.701 with some text processing (I noticed an improvement on CV too )",
    "437091": "That's great. Btw, I have sent you a DM. Let me know if there is an open place up  there",
    "437493": "Hi Serigne, can you tell us what's your CV score when you got LB 0.701? Thanks :)",
    "437498": "Maybe set the seed is useful.",
    "437569": "Thank for sharing",
    "437583": "I improved my score to LB 0.701 with pseudo-tags, 3 models",
    "437714": "I tried it yesterday, but could not get an ideal result. Maybe I do it wrong...",
    "437768": "great summary!",
    "437808": "Working on some babel  and meta features. Hopefully they will provide a lift.",
    "438642": "what does pseudo-tags mean?",
    "438836": "This is great, thanks for sharing!",
    "439259": "Hello，chi zhu, my hometown is also chongqing, i am a student at uestc , does your team need a teammate?",
    "439537": "welcome",
    "439546": "Yeah. It does. But, attention is so common these days and popular among kernel so I thought everyone might already are using it.",
    "439547": "Thanks!",
    "439548": "Yeah. I confirm this. I mode to pytorch. This is unstable too. Pinged few kaggle team members. Lets hear them  out if they can run kernels multiple times to take average.",
    "439566": "Would this be fair for someone who has figured out how to run a more stable solution within the time frame though? Instead of average I would be open to the idea of taking the lowest score of multiple runs though.",
    "440228": "what is pseudo-tags?",
    "440400": "The basic idea is to train a good model on the training set, then let it predict the labels of the test set. Now add the test set to your training set, using the predicted labels as \"ground truth\". Train again. Repeat.\n\nIt may seem counter-intuitive, but works well apparently. Here are some links:\n\n- http://deeplearning.net/wp-content/uploads/2013/03/pseudo_label_final.pdf (original paper)\n- https://towardsdatascience.com/simple-explanation-of-semi-supervised-learning-and-pseudo-labeling-c2218e8c769b",
    "440417": "Great thanks, Max",
    "440917": "Thanks Max",
    "440961": "Pseudo-tagging is a risky choice , I found it's hard to finsh in 2 hrs.  Maybe single model with CV is a better choice , I'm on it .",
    "441098": "When go to stage 2 Pseudo-tagging cost more time",
    "442825": "I find that some model cannot improve LB score with pseudo-tags, beacuse train set has enough samples to train model and stage 1 data is small.",
    "444032": "check the discussion above",
    "444568": "Really nice work!!",
    "445767": "I know LB means leaderboard, but what does CV mean? Thanks.",
    "445909": "what is IMO?",
    "445919": "Cross Validation, local CV means local cross validation score over n-folds.",
    "446056": "In My Opinion",
    "446458": "(⊙o⊙)thx",
    "446717": "Great summary! I was wondering how can one include pseudo-tags as their train data? I've downloaded my predictions and added it as a new dataset but the kernel counts it as \"external data\". Any guide on this?",
    "446729": "You can't load any additional data that is not given by the kernel. So anything you need has to be created in the kernel, in the two hours time frame.\n\nSo the only way to use pseudo-tags is to train one model, let it predict the pseudo-tags, then train another model on those.",
    "447013": "I see, thanks for your response.",
    "447228": "But that means that the idea could work pretty well in stage 2. Then the test data would have a larger number of 1's and hence score might improve by adding them.",
    "447619": "I tried to get the middle layers of NN model and train with lightgbm rather than fc layers. However it is too slow to finish in 2 hours haha : ).",
    "448074": "ok, thank you.",
    "448119": "Did you try it?\nDid it complete within 2 hours?",
    "448226": "Of course, it doesn't take long to predict on the small test set. However, I personally have not seen any improvements, though some here have reported they did.",
    "452254": "```\nUpdating embeddings with lower case words doesn't show improvements (Made CV instable to LB)\n```\nI use Keras to build deep models, I was wondering how to update embeddings with lower case words? Could you please show me the demo or some materials? Many thanks!",
    "452287": "Check out the add_lower function in the kernel https://www.kaggle.com/theoviel/improve-your-score-with-some-text-preprocessing"
  },
  "source": "meta"
}