{
  "id": 79720,
  "title": "What tricks did you come up with?",
  "url": "/competitions/quora-insincere-questions-classification/discussion/79720",
  "author_name": "Max Schumacher",
  "post_date": "2019-02-06T22:00:03.398000",
  "votes": 30,
  "comment_count": 34,
  "views": 0,
  "content": "<p>Now that the competitive phase is over, let's make the best of what each of us has learned!</p>\n\n<p>Which tricks did you guys come up with that weren't shared in any kernels?</p>\n\n<p>For my part, I realized at some point that I wasn't pushing single model CV score any further. So instead, I put all my effort into diversifying my ensemble without sacrificing too much accuracy. To measure, I used a part of the training data as a \"dev set\". On that, I measured the F1 score of an ensemble of my models.</p>\n\n<p>Here are some of the tricks I came up with:</p>\n\n<ul>\n<li>For each model, average the three embedding matrixes with different weights. I found good and diverse ones through random search and luck</li>\n<li>Randomized sample weights for each model. That way, each model focuses on different examples and should reach different solutions</li>\n<li>Re-initialize the random embedding matrix between runs. Since a lot of words don't have embeddings and thus use random vectors, this also helps diversify</li>\n<li>Add some random features to the embedding, different for each model. The models can overfit very slightly to those, also leading to increased diversification.</li>\n<li>Replace some random embedding features with a random vector for each model</li>\n<li>Overall, train longer than I would do otherwise. A slight overfit always seemed to help my ensemble F1.</li>\n<li>Each model was trained on a different subset of the data and with different layer sizes, dropout values, loss functions, etc.</li>\n</ul>\n\n<p>What about you?</p>\n\n<p><em>EDIT:</em>\nNow a major motion picture: <a href=\"https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness\">https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness</a></p>",
  "messages": [
    {
      "id": 467328,
      "postDate": "2019-02-06T22:00:03.400Z",
      "content": "<p>Now that the competitive phase is over, let's make the best of what each of us has learned!</p>\n\n<p>Which tricks did you guys come up with that weren't shared in any kernels?</p>\n\n<p>For my part, I realized at some point that I wasn't pushing single model CV score any further. So instead, I put all my effort into diversifying my ensemble without sacrificing too much accuracy. To measure, I used a part of the training data as a \"dev set\". On that, I measured the F1 score of an ensemble of my models.</p>\n\n<p>Here are some of the tricks I came up with:</p>\n\n<ul>\n<li>For each model, average the three embedding matrixes with different weights. I found good and diverse ones through random search and luck</li>\n<li>Randomized sample weights for each model. That way, each model focuses on different examples and should reach different solutions</li>\n<li>Re-initialize the random embedding matrix between runs. Since a lot of words don't have embeddings and thus use random vectors, this also helps diversify</li>\n<li>Add some random features to the embedding, different for each model. The models can overfit very slightly to those, also leading to increased diversification.</li>\n<li>Replace some random embedding features with a random vector for each model</li>\n<li>Overall, train longer than I would do otherwise. A slight overfit always seemed to help my ensemble F1.</li>\n<li>Each model was trained on a different subset of the data and with different layer sizes, dropout values, loss functions, etc.</li>\n</ul>\n\n<p>What about you?</p>\n\n<p><em>EDIT:</em>\nNow a major motion picture: <a href=\"https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness\">https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness</a></p>",
      "rawMarkdown": "Now that the competitive phase is over, let's make the best of what each of us has learned!\n\nWhich tricks did you guys come up with that weren't shared in any kernels?\n\nFor my part, I realized at some point that I wasn't pushing single model CV score any further. So instead, I put all my effort into diversifying my ensemble without sacrificing too much accuracy. To measure, I used a part of the training data as a \"dev set\". On that, I measured the F1 score of an ensemble of my models.\n\nHere are some of the tricks I came up with:\n\n- For each model, average the three embedding matrixes with different weights. I found good and diverse ones through random search and luck\n- Randomized sample weights for each model. That way, each model focuses on different examples and should reach different solutions\n- Re-initialize the random embedding matrix between runs. Since a lot of words don't have embeddings and thus use random vectors, this also helps diversify\n- Add some random features to the embedding, different for each model. The models can overfit very slightly to those, also leading to increased diversification.\n- Replace some random embedding features with a random vector for each model\n- Overall, train longer than I would do otherwise. A slight overfit always seemed to help my ensemble F1.\n- Each model was trained on a different subset of the data and with different layer sizes, dropout values, loss functions, etc.\n\nWhat about you?\n\n*EDIT:*\nNow a major motion picture: https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness",
      "votes": 29
    },
    {
      "id": 467633,
      "postDate": "2019-02-07T12:59:00.583Z",
      "content": "<p>One thing i found was that I wanted to submit a single model. But Ensemble has its benefits so didnt need to lose on that. I created predictions of test data after each epoch and did a weighted ensembling of those predictions to arrive at a final submission where i gave lower weights to start epochs. Gave me a boost from 0.68 to 0.69 Local CV.</p>",
      "rawMarkdown": "One thing i found was that I wanted to submit a single model. But Ensemble has its benefits so didnt need to lose on that. I created predictions of test data after each epoch and did a weighted ensembling of those predictions to arrive at a final submission where i gave lower weights to start epochs. Gave me a boost from 0.68 to 0.69 Local CV.",
      "votes": 8,
      "replies": [
        {
          "id": 467709,
          "postDate": "2019-02-07T15:49:59.223Z",
          "content": "<p>Ohh, it's a checkpoint ensemble! I read that paper a while ago but I totally forgot about it :( Great idea!</p>",
          "rawMarkdown": "Ohh, it's a checkpoint ensemble! I read that paper a while ago but I totally forgot about it :( Great idea!",
          "votes": 1
        },
        {
          "id": 467717,
          "postDate": "2019-02-07T16:02:12.793Z",
          "content": "<p>One thing I also noticed that adding random noise as an extra input feature was giving me better LB scores. But didn't experiment a lot with that idea. I see you have used a lot of randomizations in your approach. Any idea, why does it work?</p>",
          "rawMarkdown": "One thing I also noticed that adding random noise as an extra input feature was giving me better LB scores. But didn't experiment a lot with that idea. I see you have used a lot of randomizations in your approach. Any idea, why does it work?"
        },
        {
          "id": 467801,
          "postDate": "2019-02-07T18:33:07.573Z",
          "content": "<p>Cool, a Checkpoint ensemble, sounds like a very smart strategy! You get good models almost for free. Looking forward to experiment with this.</p>",
          "rawMarkdown": "Cool, a Checkpoint ensemble, sounds like a very smart strategy! You get good models almost for free. Looking forward to experiment with this.",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 467844,
          "postDate": "2019-02-07T20:43:01.880Z",
          "content": "<p>Yeah, it's a lot of randomness... The idea is to increase variance in your ensemble, which naturally leads to higher ensembled accuracy. Or intuitively, you try to make each model use different strategies to predict the target. That way, each has different strengths and weaknesses that combine into a superior model when ensembled.</p>\n\n<p>In retrospect, I might have overdone it though :)</p>",
          "rawMarkdown": "Yeah, it's a lot of randomness... The idea is to increase variance in your ensemble, which naturally leads to higher ensembled accuracy. Or intuitively, you try to make each model use different strategies to predict the target. That way, each has different strengths and weaknesses that combine into a superior model when ensembled.\n\nIn retrospect, I might have overdone it though :)",
          "votes": 1
        },
        {
          "id": 467939,
          "postDate": "2019-02-08T02:11:18.223Z",
          "content": "<p>Unfortunately (for me), I tried a lot of this 'snapshot' ensemble from the first month until the last days, but it didn't work for me in all settings :_(</p>\n\n<p>On a small note, I realized on the last day that the prediction is not quite free, since the prediction on the private test set (376k) will take some amount of time, so we decided to drop this idea finally.</p>",
          "rawMarkdown": "Unfortunately (for me), I tried a lot of this 'snapshot' ensemble from the first month until the last days, but it didn't work for me in all settings :_(\n\nOn a small note, I realized on the last day that the prediction is not quite free, since the prediction on the private test set (376k) will take some amount of time, so we decided to drop this idea finally.",
          "votes": 1
        },
        {
          "id": 467993,
          "postDate": "2019-02-08T05:20:36.797Z",
          "content": "<p>How did you try to average predictions out. I myself did a moving average of predictions through epochs and not a regression based approach since regression didn't work... </p>",
          "rawMarkdown": "How did you try to average predictions out. I myself did a moving average of predictions through epochs and not a regression based approach since regression didn't work... ",
          "votes": 1
        },
        {
          "id": 468059,
          "postDate": "2019-02-08T08:00:37.827Z",
          "content": "<p>Hi Rahul, thanks for asking. I also used weighted average where I have weights as hyper-parameters. </p>",
          "rawMarkdown": "Hi Rahul, thanks for asking. I also used weighted average where I have weights as hyper-parameters. "
        }
      ]
    },
    {
      "id": 467802,
      "postDate": "2019-02-07T18:34:23.347Z",
      "content": "<ol>\n<li>Using BCE+soft f1 loss gave me a little boost (0.003) on LB and (0.001) on CV on the last day. It gave stabler f1 score.</li>\n<li>Regarding overfit, training with more epochs gives less sensitive threshold v. f1 score plot at the optimal point. It helps when training and test sets are different. That's why all the 0.7 kernels overfit the training set.</li>\n</ol>",
      "rawMarkdown": "1. Using BCE+soft f1 loss gave me a little boost (0.003) on LB and (0.001) on CV on the last day. It gave stabler f1 score.\n2. Regarding overfit, training with more epochs gives less sensitive threshold v. f1 score plot at the optimal point. It helps when training and test sets are different. That's why all the 0.7 kernels overfit the training set.",
      "votes": 3,
      "replies": [
        {
          "id": 467875,
          "postDate": "2019-02-07T21:45:41.673Z",
          "rawMarkdown": "",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 468478,
          "postDate": "2019-02-09T00:20:04.107Z",
          "content": "<p>&gt; <strong>bestfitting wrote</strong>\n&gt; \n&gt; &gt; I hope it will be helpful to you.<br>\n&gt; \n&gt; \n&gt;     def f2_loss(logits, labels):\n&gt;     __small_value=1e-6\n&gt;     beta = 2\n&gt;     batch_size = logits.size()[0]\n&gt;     p = F.sigmoid(logits)\n&gt;     l = labels\n&gt;     num_pos = torch.sum(p, 1) + __small_value\n&gt;     num_pos_hat = torch.sum(l, 1) + __small_value\n&gt;     tp = torch.sum(l * p, 1)\n&gt;     precise = tp / num_pos\n&gt;     recall = tp / num_pos_hat\n&gt;     fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + __small_value)\n&gt;     loss = fs.sum() / batch_size\n&gt;     return (1 - loss)</p>\n\n<p>As bestfitting shared 2 years ago in here <a href=\"https://www.kaggle.com/c/planet-understanding-the-amazon-from-space/discussion/36809\">https://www.kaggle.com/c/planet-understanding-the-amazon-from-space/discussion/36809</a> . Just change beta from 2 to 1.</p>",
          "rawMarkdown": "\n&gt; **bestfitting wrote**\n&gt; \n&gt; &gt; I hope it will be helpful to you.<br>\n&gt; \n&gt; \n&gt;     def f2_loss(logits, labels):\n&gt;     __small_value=1e-6\n&gt;     beta = 2\n&gt;     batch_size = logits.size()[0]\n&gt;     p = F.sigmoid(logits)\n&gt;     l = labels\n&gt;     num_pos = torch.sum(p, 1) + __small_value\n&gt;     num_pos_hat = torch.sum(l, 1) + __small_value\n&gt;     tp = torch.sum(l * p, 1)\n&gt;     precise = tp / num_pos\n&gt;     recall = tp / num_pos_hat\n&gt;     fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + __small_value)\n&gt;     loss = fs.sum() / batch_size\n&gt;     return (1 - loss)\n\nAs bestfitting shared 2 years ago in here https://www.kaggle.com/c/planet-understanding-the-amazon-from-space/discussion/36809 . Just change beta from 2 to 1.",
          "votes": 3
        }
      ]
    },
    {
      "id": 467358,
      "postDate": "2019-02-06T23:48:12.633Z",
      "content": "<p>I mainly focused on improving Kfold local CV score over LB score. Couple of simple steps that helped were: </p>\n\n<ul>\n<li>Training on trainable embeddings for one epoch after slightly overfitting to fixed embeddings. Reinitialized embeddings after each fold.</li>\n<li>Gaussian noise after embedding layer in network during training, as mentioned in this <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/78841\">discussion</a>.</li>\n<li>Divide vectors in embedding matrix by their frequency in the provided embeddings (Used all embeddings except word2vec since it wasn't improving the score for me). That way if a word was present in only one of the embeddings, I wasn't diminishing it's weight by dividing it by 3.</li>\n</ul>\n\n<p>```</p>\n\n<pre><code>global_embedding = np.random.normal(global_mean, global_std, (nb_words, embed_size))\nembedding_count = np.zeros((nb_words,1))\nfor EMBEDDING_FILE in embedding_list:\n    for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore'):\n        word, vec = o.split(' ', 1)\n        if word not in word_index:\n            continue\n        i = word_index[word]\n        if i &amp;gt;= nb_words:\n            continue\n        embedding_vector = np.asarray(vec.split(' '), dtype='float32')[:embed_size]\n        if len(embedding_vector) == embed_size:\n            global_embedding[i] = (embedding_count[i]*global_embedding[i] + embedding_vector)/(embedding_count[i] + 1)\n            embedding_count[i] += 1\n    del embedding_vector\n    gc.collect()\n</code></pre>\n\n<p>```</p>",
      "rawMarkdown": "I mainly focused on improving Kfold local CV score over LB score. Couple of simple steps that helped were: \n\n- Training on trainable embeddings for one epoch after slightly overfitting to fixed embeddings. Reinitialized embeddings after each fold.\n- Gaussian noise after embedding layer in network during training, as mentioned in this [discussion](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/78841).\n- Divide vectors in embedding matrix by their frequency in the provided embeddings (Used all embeddings except word2vec since it wasn't improving the score for me). That way if a word was present in only one of the embeddings, I wasn't diminishing it's weight by dividing it by 3.\n\n```\n\n    global_embedding = np.random.normal(global_mean, global_std, (nb_words, embed_size))\n    embedding_count = np.zeros((nb_words,1))\n    for EMBEDDING_FILE in embedding_list:\n        for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore'):\n            word, vec = o.split(' ', 1)\n            if word not in word_index:\n                continue\n            i = word_index[word]\n            if i &gt;= nb_words:\n                continue\n            embedding_vector = np.asarray(vec.split(' '), dtype='float32')[:embed_size]\n            if len(embedding_vector) == embed_size:\n                global_embedding[i] = (embedding_count[i]*global_embedding[i] + embedding_vector)/(embedding_count[i] + 1)\n                embedding_count[i] += 1\n        del embedding_vector\n        gc.collect()\n\n```",
      "votes": 4
    },
    {
      "id": 471479,
      "postDate": "2019-02-14T14:17:12.390Z",
      "content": "<p>Hi Max, Annabelle\n<a href=\"/mschumacher\">@mschumacher</a>, <a href=\"/sezaugg\">@sezaugg</a></p>\n\n<p>Allow me to congratulate both of you here for the final performance!\nThanks for all your sharing during the competition, and see you again soon!</p>",
      "rawMarkdown": "Hi Max, Annabelle\n@mschumacher, @sezaugg\n\nAllow me to congratulate both of you here for the final performance!\nThanks for all your sharing during the competition, and see you again soon!",
      "votes": 1,
      "replies": [
        {
          "id": 471563,
          "postDate": "2019-02-14T15:54:01.440Z",
          "content": "<p>Also congrats to you! Thanks for being such a great guy and spreading positive vibes in the Kaggle community :)</p>",
          "rawMarkdown": "Also congrats to you! Thanks for being such a great guy and spreading positive vibes in the Kaggle community :)",
          "votes": 1
        },
        {
          "id": 471586,
          "postDate": "2019-02-14T16:12:46.480Z",
          "content": "<p>Hi Neuron Engineer. Thanks, same to you! I am looking forward for next round of active and constructive discussions with you in another competition.</p>",
          "rawMarkdown": "Hi Neuron Engineer. Thanks, same to you! I am looking forward for next round of active and constructive discussions with you in another competition.",
          "votes": 1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 471566,
      "postDate": "2019-02-14T15:55:52.060Z",
      "content": "<p>For those who were wondering, seems like my ideas weren't completely off the mark! I got in 44th. I published the kernel here: <a href=\"https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness\">https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness</a></p>",
      "rawMarkdown": "For those who were wondering, seems like my ideas weren't completely off the mark! I got in 44th. I published the kernel here: https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness"
    },
    {
      "id": 468044,
      "postDate": "2019-02-08T07:30:43.163Z",
      "content": "<p>Hello! Can we use data augmentation for the insincere questions to train the model?</p>",
      "rawMarkdown": "Hello! Can we use data augmentation for the insincere questions to train the model?",
      "replies": [
        {
          "id": 468102,
          "postDate": "2019-02-08T09:37:00.077Z",
          "content": "<p>Hi Saichand! I hope you've noticed already that the competition is over? :) But to answer your question, you can do anything you want, as long as you don't use additional data sources and stay within the two hour limit.</p>",
          "rawMarkdown": "Hi Saichand! I hope you've noticed already that the competition is over? :) But to answer your question, you can do anything you want, as long as you don't use additional data sources and stay within the two hour limit."
        },
        {
          "id": 468305,
          "postDate": "2019-02-08T16:47:06.370Z",
          "content": "<p>I did use dataug in my last sub. I replaced all country names in insincere questions with “Israel” to create extra sentences. There is technical reasoning behind this. If any sentence is insincere already with another country name, it is bound to be insincere with “Israel” in it. That can be inferred from checking several sentences with different country names. Same could be applied to other names or groups. I am trying to formulate this technique. Anyway, my CV went from 0.695 to 0.745 after this.</p>",
          "rawMarkdown": "I did use dataug in my last sub. I replaced all country names in insincere questions with “Israel” to create extra sentences. There is technical reasoning behind this. If any sentence is insincere already with another country name, it is bound to be insincere with “Israel” in it. That can be inferred from checking several sentences with different country names. Same could be applied to other names or groups. I am trying to formulate this technique. Anyway, my CV went from 0.695 to 0.745 after this.",
          "votes": 1
        },
        {
          "id": 468319,
          "postDate": "2019-02-08T17:03:41.133Z",
          "content": "<p>Wow! That gain seems unreal... How did it translate to LB score?</p>",
          "rawMarkdown": "Wow! That gain seems unreal... How did it translate to LB score?"
        },
        {
          "id": 468407,
          "postDate": "2019-02-08T20:25:22.937Z",
          "content": "<p>Actually did not change the LB :) Maybe the gain in cv was from similar sentences left in train and test. Since it did not hurt LB, I think it is safe to use. It might help in shake up.</p>",
          "rawMarkdown": "Actually did not change the LB :) Maybe the gain in cv was from similar sentences left in train and test. Since it did not hurt LB, I think it is safe to use. It might help in shake up.",
          "votes": 1
        }
      ]
    },
    {
      "id": 467545,
      "postDate": "2019-02-07T09:52:18.033Z",
      "content": "<p>Thanks for the post, it is probably a good time to share insights as in 3 weeks we will all be busy with other competitions.</p>\n\n<ul>\n<li>I also found that 3 epochs generally gave best results for a single NN, but 4-5 epochs better when ensembling several NNs via mean. As you mentioned: slight overfit helps ensemble performance</li>\n<li>CV-trained models have more diversity but none of the models learns from all data. Thus, I switched to train each NN from the ensemble on all training data for submissions. This had a quite significant boost on performance. For local experiments, I used a single split: 90% of data for training and 10% for test. </li>\n<li>I tried many different NN architectures (because it is fun, let's admit it) but found that an adequately sized single Bi-LSTM (or GRU) was the main driver of performance. I finally used a single Bi-LSTM layer. GRU tended to be slightly worse in my particular setting.</li>\n<li>Early,  from public kernels I realized that we must separate words from special characters and numbers. This had a quite significant boost on performance.</li>\n</ul>\n\n<p>EDIT: \n- Too late, I started a \"sort of spell-check\", i.e. using pre-processing pipeline that transformed OOV words to closest word available in embedding to increase coverage of the embedding. It seemed to increase performance consistently, but I did not have time to finalize this.</p>",
      "rawMarkdown": "Thanks for the post, it is probably a good time to share insights as in 3 weeks we will all be busy with other competitions.\n\n- I also found that 3 epochs generally gave best results for a single NN, but 4-5 epochs better when ensembling several NNs via mean. As you mentioned: slight overfit helps ensemble performance\n- CV-trained models have more diversity but none of the models learns from all data. Thus, I switched to train each NN from the ensemble on all training data for submissions. This had a quite significant boost on performance. For local experiments, I used a single split: 90% of data for training and 10% for test. \n- I tried many different NN architectures (because it is fun, let's admit it) but found that an adequately sized single Bi-LSTM (or GRU) was the main driver of performance. I finally used a single Bi-LSTM layer. GRU tended to be slightly worse in my particular setting.\n- Early,  from public kernels I realized that we must separate words from special characters and numbers. This had a quite significant boost on performance.\n\nEDIT: \n- Too late, I started a \"sort of spell-check\", i.e. using pre-processing pipeline that transformed OOV words to closest word available in embedding to increase coverage of the embedding. It seemed to increase performance consistently, but I did not have time to finalize this.\n\n ",
      "votes": 16,
      "isDeleted": true,
      "replies": [
        {
          "id": 467710,
          "postDate": "2019-02-07T15:51:11.573Z",
          "content": "<p>Interesting! For me, I tried many different spell-check preprocessing steps, but they consistently hurt my CV against all common sense...</p>",
          "rawMarkdown": "Interesting! For me, I tried many different spell-check preprocessing steps, but they consistently hurt my CV against all common sense..."
        },
        {
          "id": 467791,
          "postDate": "2019-02-07T18:21:43.573Z",
          "content": "<p>My wording was actually imprecise: I did not classic spellcheck but I transformed OOV words to closest word available in embedding. </p>\n\n<p>I also tried some of the spell-check dicts available in public kernels but could not find any measurable improvement.</p>",
          "rawMarkdown": "My wording was actually imprecise: I did not classic spellcheck but I transformed OOV words to closest word available in embedding. \n\nI also tried some of the spell-check dicts available in public kernels but could not find any measurable improvement.",
          "isDeleted": true
        },
        {
          "id": 467842,
          "postDate": "2019-02-07T20:37:23.623Z",
          "content": "<p>Ah, I see! I was thinking that would be too expensive, so didn't try. What I did though is check for capitalized and upper case versions of words. E.g. if the word \"usa\" didn't have an embedding, I checked if there's \"Usa\" or \"USA\" and used those instead.</p>",
          "rawMarkdown": "Ah, I see! I was thinking that would be too expensive, so didn't try. What I did though is check for capitalized and upper case versions of words. E.g. if the word \"usa\" didn't have an embedding, I checked if there's \"Usa\" or \"USA\" and used those instead."
        },
        {
          "id": 467876,
          "postDate": "2019-02-07T21:50:21.253Z",
          "content": "<p>What helped for me:\n* Add additional features, per question, and concat them with the hidden_size units which the Bi-LSTM/GRU outputs. The only feature which seemed to give consistent, reliable, gains of around +0.004 on local holdout set was:\n    - Positive top500 bi-gram score: Go over all bi-grams in the question and check whether they appear in the top500 bi-grams, which appear in questions with target=0. Sum the number of appearances. Dont consider bi-grams which appear in both top500 bi-grams for target=0 and target=1 questions\n    - Negative top500 bi-gram score: Same as above, but for target=1\n* Interestingly, I did NOT get those improvements, when I used those features in a 2nd level model (LightGBM), together with predictions from the first model</p>\n\n<p>What did not help for me:\n* A lot of features, which seemed indicative of a sincere/incincere question (e.g. number of profanity words; average word frequencies -&gt; there was a nice kernel “magic feature” 1 day before the competition’s end). It seems that the Bi-LSTM network picked up those signals already\n* My way of super intensive spelling correction: I considered how often a candidate for correction appeared in the dataset, and the embeddings, compared to its suggested correction; I included keyboard-distance into the DamerauLevenshtein distance (i.e. the word \"mafic\" can be corrected to \"magic\", but \"mapic\" not) - none of it gave any consistent gains above the natural fluctuations</p>",
          "rawMarkdown": "What helped for me:\n* Add additional features, per question, and concat them with the hidden_size units which the Bi-LSTM/GRU outputs. The only feature which seemed to give consistent, reliable, gains of around +0.004 on local holdout set was:\n    - Positive top500 bi-gram score: Go over all bi-grams in the question and check whether they appear in the top500 bi-grams, which appear in questions with target=0. Sum the number of appearances. Dont consider bi-grams which appear in both top500 bi-grams for target=0 and target=1 questions\n    - Negative top500 bi-gram score: Same as above, but for target=1\n* Interestingly, I did NOT get those improvements, when I used those features in a 2nd level model (LightGBM), together with predictions from the first model\n\nWhat did not help for me:\n* A lot of features, which seemed indicative of a sincere/incincere question (e.g. number of profanity words; average word frequencies -&gt; there was a nice kernel “magic feature” 1 day before the competition’s end). It seems that the Bi-LSTM network picked up those signals already\n* My way of super intensive spelling correction: I considered how often a candidate for correction appeared in the dataset, and the embeddings, compared to its suggested correction; I included keyboard-distance into the DamerauLevenshtein distance (i.e. the word \"mafic\" can be corrected to \"magic\", but \"mapic\" not) - none of it gave any consistent gains above the natural fluctuations\n",
          "votes": 2
        },
        {
          "id": 468064,
          "postDate": "2019-02-08T08:10:26.960Z",
          "content": "<p><a href=\"/annabelle\">@annabelle</a> - Could you please explain this. Sorry I did not get what you mean by this. <code>Thus, I switched to train each NN from the ensemble on all training data for submissions</code></p>",
          "rawMarkdown": "@annabelle - Could you please explain this. Sorry I did not get what you mean by this. ```Thus, I switched to train each NN from the ensemble on all training data for submissions```",
          "votes": 1
        },
        {
          "id": 468100,
          "postDate": "2019-02-08T09:34:32.250Z",
          "content": "<p>I think she means that she used validation sets during experimentation to find good configurations. Then, when submitting, she removed the split into train / val sets, training each model on 100% of the data.</p>",
          "rawMarkdown": "I think she means that she used validation sets during experimentation to find good configurations. Then, when submitting, she removed the split into train / val sets, training each model on 100% of the data."
        },
        {
          "id": 468101,
          "postDate": "2019-02-08T09:34:35.580Z",
          "content": "<p>Using the whole dataset for submissions seems so obvious.... but didn't occurred to me :-)</p>\n\n<p>But then how did you generate diversification between the ensemble models? Just by random noise or random initial conditions maybe?</p>",
          "rawMarkdown": "Using the whole dataset for submissions seems so obvious.... but didn't occurred to me :-)\n\nBut then how did you generate diversification between the ensemble models? Just by random noise or random initial conditions maybe?"
        },
        {
          "id": 468105,
          "postDate": "2019-02-08T09:44:11.323Z",
          "content": "<p>I guess it's always a balance between maximizing single model performance versus minimizing correlation. In this case, the gain in performance for each model may have been high enough to outbalance the loss in diversity. I wish I had used that too for at least one of my final submissions...</p>",
          "rawMarkdown": "I guess it's always a balance between maximizing single model performance versus minimizing correlation. In this case, the gain in performance for each model may have been high enough to outbalance the loss in diversity. I wish I had used that too for at least one of my final submissions..."
        },
        {
          "id": 468106,
          "postDate": "2019-02-08T09:48:58.143Z",
          "content": "<p>Of course:</p>\n\n<p>Say we use an ensemble of 3 Neural Nets.\nNN1 is trained with all train data (n=1306122)\nNN2 is trained with all train data (n=1306122)\nNN3 is trained with all train data (n=1306122)</p>\n\n<p>Then, predictions from 3 Neural Nets are combined (e.g. with mean) to obtain the final predictions.</p>\n\n<p>Of course the above works only for submissions to LB.\nTo do compare your models, you need to split train data via CV or hold-out.</p>\n\n<p>Hope this helps</p>",
          "rawMarkdown": "Of course:\n\nSay we use an ensemble of 3 Neural Nets.\nNN1 is trained with all train data (n=1306122)\nNN2 is trained with all train data (n=1306122)\nNN3 is trained with all train data (n=1306122)\n\nThen, predictions from 3 Neural Nets are combined (e.g. with mean) to obtain the final predictions.\n\nOf course the above works only for submissions to LB.\nTo do compare your models, you need to split train data via CV or hold-out.\n\nHope this helps",
          "votes": 2,
          "isDeleted": true
        },
        {
          "id": 468161,
          "postDate": "2019-02-08T11:46:54.603Z",
          "content": "<p>Thanks. I sort of got it later. I did the same thing, with full training data. First I created 12 models ( most are lstm/gru with different embeddings ) and did the same averaging. Instead of N models, I choose all possible combination of models ( N = 3, then (m1 , m2 , m3 ) , (m1 , m2) , (m2, m3) , (m3,m1)) and then choose the models based on local CV. But, not a single ensembling helps me to cross my single model of 0.702. Best N average LB I got is 0.7.</p>",
          "rawMarkdown": "Thanks. I sort of got it later. I did the same thing, with full training data. First I created 12 models ( most are lstm/gru with different embeddings ) and did the same averaging. Instead of N models, I choose all possible combination of models ( N = 3, then (m1 , m2 , m3 ) , (m1 , m2) , (m2, m3) , (m3,m1)) and then choose the models based on local CV. But, not a single ensembling helps me to cross my single model of 0.702. Best N average LB I got is 0.7.",
          "votes": 1
        },
        {
          "id": 468177,
          "postDate": "2019-02-08T12:41:40.360Z",
          "content": "<p>Hey Sarath, we are on the same boat. Straightforward ensemble is indeed not easy to apply.</p>",
          "rawMarkdown": "Hey Sarath, we are on the same boat. Straightforward ensemble is indeed not easy to apply.",
          "votes": 1
        },
        {
          "id": 468226,
          "postDate": "2019-02-08T14:27:54.720Z",
          "content": "<p>Waiting to see the top place solutions, in 2-3 weeks.</p>",
          "rawMarkdown": "Waiting to see the top place solutions, in 2-3 weeks.",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 467633,
      "author_name": "Rahul Agarwal",
      "author_url": "",
      "post_date": "2019-02-07T12:59:00.583000",
      "content": "<p>One thing i found was that I wanted to submit a single model. But Ensemble has its benefits so didnt need to lose on that. I created predictions of test data after each epoch and did a weighted ensembling of those predictions to arrive at a final submission where i gave lower weights to start epochs. Gave me a boost from 0.68 to 0.69 Local CV.</p>",
      "votes": 8,
      "replies": [
        {
          "id": 467709,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-07T15:49:59.223000",
          "content": "<p>Ohh, it's a checkpoint ensemble! I read that paper a while ago but I totally forgot about it :( Great idea!</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 467717,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2019-02-07T16:02:12.793000",
          "content": "<p>One thing I also noticed that adding random noise as an extra input feature was giving me better LB scores. But didn't experiment a lot with that idea. I see you have used a lot of randomizations in your approach. Any idea, why does it work?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467801,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-07T18:33:07.573000",
          "content": "<p>Cool, a Checkpoint ensemble, sounds like a very smart strategy! You get good models almost for free. Looking forward to experiment with this.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 467844,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-07T20:43:01.880000",
          "content": "<p>Yeah, it's a lot of randomness... The idea is to increase variance in your ensemble, which naturally leads to higher ensembled accuracy. Or intuitively, you try to make each model use different strategies to predict the target. That way, each has different strengths and weaknesses that combine into a superior model when ensembled.</p>\n\n<p>In retrospect, I might have overdone it though :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 467939,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-02-08T02:11:18.223000",
          "content": "<p>Unfortunately (for me), I tried a lot of this 'snapshot' ensemble from the first month until the last days, but it didn't work for me in all settings :_(</p>\n\n<p>On a small note, I realized on the last day that the prediction is not quite free, since the prediction on the private test set (376k) will take some amount of time, so we decided to drop this idea finally.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 467993,
          "author_name": "Rahul Agarwal",
          "author_url": "",
          "post_date": "2019-02-08T05:20:36.797000",
          "content": "<p>How did you try to average predictions out. I myself did a moving average of predictions through epochs and not a regression based approach since regression didn't work... </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 468059,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-02-08T08:00:37.827000",
          "content": "<p>Hi Rahul, thanks for asking. I also used weighted average where I have weights as hyper-parameters. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 467802,
      "author_name": "Darkate",
      "author_url": "",
      "post_date": "2019-02-07T18:34:23.347000",
      "content": "<ol>\n<li>Using BCE+soft f1 loss gave me a little boost (0.003) on LB and (0.001) on CV on the last day. It gave stabler f1 score.</li>\n<li>Regarding overfit, training with more epochs gives less sensitive threshold v. f1 score plot at the optimal point. It helps when training and test sets are different. That's why all the 0.7 kernels overfit the training set.</li>\n</ol>",
      "votes": 3,
      "replies": [
        {
          "id": 467875,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-07T21:45:41.673000",
          "content": "",
          "votes": 2,
          "replies": []
        },
        {
          "id": 468478,
          "author_name": "Darkate",
          "author_url": "",
          "post_date": "2019-02-09T00:20:04.107000",
          "content": "<p>&gt; <strong>bestfitting wrote</strong>\n&gt; \n&gt; &gt; I hope it will be helpful to you.<br>\n&gt; \n&gt; \n&gt;     def f2_loss(logits, labels):\n&gt;     __small_value=1e-6\n&gt;     beta = 2\n&gt;     batch_size = logits.size()[0]\n&gt;     p = F.sigmoid(logits)\n&gt;     l = labels\n&gt;     num_pos = torch.sum(p, 1) + __small_value\n&gt;     num_pos_hat = torch.sum(l, 1) + __small_value\n&gt;     tp = torch.sum(l * p, 1)\n&gt;     precise = tp / num_pos\n&gt;     recall = tp / num_pos_hat\n&gt;     fs = (1 + beta * beta) * precise * recall / (beta * beta * precise + recall + __small_value)\n&gt;     loss = fs.sum() / batch_size\n&gt;     return (1 - loss)</p>\n\n<p>As bestfitting shared 2 years ago in here <a href=\"https://www.kaggle.com/c/planet-understanding-the-amazon-from-space/discussion/36809\">https://www.kaggle.com/c/planet-understanding-the-amazon-from-space/discussion/36809</a> . Just change beta from 2 to 1.</p>",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 467358,
      "author_name": "r41",
      "author_url": "",
      "post_date": "2019-02-06T23:48:12.633000",
      "content": "<p>I mainly focused on improving Kfold local CV score over LB score. Couple of simple steps that helped were: </p>\n\n<ul>\n<li>Training on trainable embeddings for one epoch after slightly overfitting to fixed embeddings. Reinitialized embeddings after each fold.</li>\n<li>Gaussian noise after embedding layer in network during training, as mentioned in this <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/78841\">discussion</a>.</li>\n<li>Divide vectors in embedding matrix by their frequency in the provided embeddings (Used all embeddings except word2vec since it wasn't improving the score for me). That way if a word was present in only one of the embeddings, I wasn't diminishing it's weight by dividing it by 3.</li>\n</ul>\n\n<p>```</p>\n\n<pre><code>global_embedding = np.random.normal(global_mean, global_std, (nb_words, embed_size))\nembedding_count = np.zeros((nb_words,1))\nfor EMBEDDING_FILE in embedding_list:\n    for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore'):\n        word, vec = o.split(' ', 1)\n        if word not in word_index:\n            continue\n        i = word_index[word]\n        if i &amp;gt;= nb_words:\n            continue\n        embedding_vector = np.asarray(vec.split(' '), dtype='float32')[:embed_size]\n        if len(embedding_vector) == embed_size:\n            global_embedding[i] = (embedding_count[i]*global_embedding[i] + embedding_vector)/(embedding_count[i] + 1)\n            embedding_count[i] += 1\n    del embedding_vector\n    gc.collect()\n</code></pre>\n\n<p>```</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 471479,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-02-14T14:17:12.390000",
      "content": "<p>Hi Max, Annabelle\n<a href=\"/mschumacher\">@mschumacher</a>, <a href=\"/sezaugg\">@sezaugg</a></p>\n\n<p>Allow me to congratulate both of you here for the final performance!\nThanks for all your sharing during the competition, and see you again soon!</p>",
      "votes": 1,
      "replies": [
        {
          "id": 471563,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-14T15:54:01.440000",
          "content": "<p>Also congrats to you! Thanks for being such a great guy and spreading positive vibes in the Kaggle community :)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 471586,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-14T16:12:46.480000",
          "content": "<p>Hi Neuron Engineer. Thanks, same to you! I am looking forward for next round of active and constructive discussions with you in another competition.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 471566,
      "author_name": "Max Schumacher",
      "author_url": "",
      "post_date": "2019-02-14T15:55:52.060000",
      "content": "<p>For those who were wondering, seems like my ideas weren't completely off the mark! I got in 44th. I published the kernel here: <a href=\"https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness\">https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness</a></p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 468044,
      "author_name": "Saichand",
      "author_url": "",
      "post_date": "2019-02-08T07:30:43.163000",
      "content": "<p>Hello! Can we use data augmentation for the insincere questions to train the model?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 468102,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-08T09:37:00.077000",
          "content": "<p>Hi Saichand! I hope you've noticed already that the competition is over? :) But to answer your question, you can do anything you want, as long as you don't use additional data sources and stay within the two hour limit.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468305,
          "author_name": "Kepler456b",
          "author_url": "",
          "post_date": "2019-02-08T16:47:06.370000",
          "content": "<p>I did use dataug in my last sub. I replaced all country names in insincere questions with “Israel” to create extra sentences. There is technical reasoning behind this. If any sentence is insincere already with another country name, it is bound to be insincere with “Israel” in it. That can be inferred from checking several sentences with different country names. Same could be applied to other names or groups. I am trying to formulate this technique. Anyway, my CV went from 0.695 to 0.745 after this.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 468319,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-08T17:03:41.133000",
          "content": "<p>Wow! That gain seems unreal... How did it translate to LB score?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468407,
          "author_name": "Kepler456b",
          "author_url": "",
          "post_date": "2019-02-08T20:25:22.937000",
          "content": "<p>Actually did not change the LB :) Maybe the gain in cv was from similar sentences left in train and test. Since it did not hurt LB, I think it is safe to use. It might help in shake up.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 467545,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-07T09:52:18.033000",
      "content": "<p>Thanks for the post, it is probably a good time to share insights as in 3 weeks we will all be busy with other competitions.</p>\n\n<ul>\n<li>I also found that 3 epochs generally gave best results for a single NN, but 4-5 epochs better when ensembling several NNs via mean. As you mentioned: slight overfit helps ensemble performance</li>\n<li>CV-trained models have more diversity but none of the models learns from all data. Thus, I switched to train each NN from the ensemble on all training data for submissions. This had a quite significant boost on performance. For local experiments, I used a single split: 90% of data for training and 10% for test. </li>\n<li>I tried many different NN architectures (because it is fun, let's admit it) but found that an adequately sized single Bi-LSTM (or GRU) was the main driver of performance. I finally used a single Bi-LSTM layer. GRU tended to be slightly worse in my particular setting.</li>\n<li>Early,  from public kernels I realized that we must separate words from special characters and numbers. This had a quite significant boost on performance.</li>\n</ul>\n\n<p>EDIT: \n- Too late, I started a \"sort of spell-check\", i.e. using pre-processing pipeline that transformed OOV words to closest word available in embedding to increase coverage of the embedding. It seemed to increase performance consistently, but I did not have time to finalize this.</p>",
      "votes": 16,
      "replies": [
        {
          "id": 467710,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-07T15:51:11.573000",
          "content": "<p>Interesting! For me, I tried many different spell-check preprocessing steps, but they consistently hurt my CV against all common sense...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467791,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-07T18:21:43.573000",
          "content": "<p>My wording was actually imprecise: I did not classic spellcheck but I transformed OOV words to closest word available in embedding. </p>\n\n<p>I also tried some of the spell-check dicts available in public kernels but could not find any measurable improvement.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467842,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-07T20:37:23.623000",
          "content": "<p>Ah, I see! I was thinking that would be too expensive, so didn't try. What I did though is check for capitalized and upper case versions of words. E.g. if the word \"usa\" didn't have an embedding, I checked if there's \"Usa\" or \"USA\" and used those instead.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 467876,
          "author_name": "Trian",
          "author_url": "",
          "post_date": "2019-02-07T21:50:21.253000",
          "content": "<p>What helped for me:\n* Add additional features, per question, and concat them with the hidden_size units which the Bi-LSTM/GRU outputs. The only feature which seemed to give consistent, reliable, gains of around +0.004 on local holdout set was:\n    - Positive top500 bi-gram score: Go over all bi-grams in the question and check whether they appear in the top500 bi-grams, which appear in questions with target=0. Sum the number of appearances. Dont consider bi-grams which appear in both top500 bi-grams for target=0 and target=1 questions\n    - Negative top500 bi-gram score: Same as above, but for target=1\n* Interestingly, I did NOT get those improvements, when I used those features in a 2nd level model (LightGBM), together with predictions from the first model</p>\n\n<p>What did not help for me:\n* A lot of features, which seemed indicative of a sincere/incincere question (e.g. number of profanity words; average word frequencies -&gt; there was a nice kernel “magic feature” 1 day before the competition’s end). It seems that the Bi-LSTM network picked up those signals already\n* My way of super intensive spelling correction: I considered how often a candidate for correction appeared in the dataset, and the embeddings, compared to its suggested correction; I included keyboard-distance into the DamerauLevenshtein distance (i.e. the word \"mafic\" can be corrected to \"magic\", but \"mapic\" not) - none of it gave any consistent gains above the natural fluctuations</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 468064,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-02-08T08:10:26.960000",
          "content": "<p><a href=\"/annabelle\">@annabelle</a> - Could you please explain this. Sorry I did not get what you mean by this. <code>Thus, I switched to train each NN from the ensemble on all training data for submissions</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 468100,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-08T09:34:32.250000",
          "content": "<p>I think she means that she used validation sets during experimentation to find good configurations. Then, when submitting, she removed the split into train / val sets, training each model on 100% of the data.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468101,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-02-08T09:34:35.580000",
          "content": "<p>Using the whole dataset for submissions seems so obvious.... but didn't occurred to me :-)</p>\n\n<p>But then how did you generate diversification between the ensemble models? Just by random noise or random initial conditions maybe?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468105,
          "author_name": "Max Schumacher",
          "author_url": "",
          "post_date": "2019-02-08T09:44:11.323000",
          "content": "<p>I guess it's always a balance between maximizing single model performance versus minimizing correlation. In this case, the gain in performance for each model may have been high enough to outbalance the loss in diversity. I wish I had used that too for at least one of my final submissions...</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 468106,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-08T09:48:58.143000",
          "content": "<p>Of course:</p>\n\n<p>Say we use an ensemble of 3 Neural Nets.\nNN1 is trained with all train data (n=1306122)\nNN2 is trained with all train data (n=1306122)\nNN3 is trained with all train data (n=1306122)</p>\n\n<p>Then, predictions from 3 Neural Nets are combined (e.g. with mean) to obtain the final predictions.</p>\n\n<p>Of course the above works only for submissions to LB.\nTo do compare your models, you need to split train data via CV or hold-out.</p>\n\n<p>Hope this helps</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 468161,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-02-08T11:46:54.603000",
          "content": "<p>Thanks. I sort of got it later. I did the same thing, with full training data. First I created 12 models ( most are lstm/gru with different embeddings ) and did the same averaging. Instead of N models, I choose all possible combination of models ( N = 3, then (m1 , m2 , m3 ) , (m1 , m2) , (m2, m3) , (m3,m1)) and then choose the models based on local CV. But, not a single ensembling helps me to cross my single model of 0.702. Best N average LB I got is 0.7.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 468177,
          "author_name": "Neuron Engineer",
          "author_url": "",
          "post_date": "2019-02-08T12:41:40.360000",
          "content": "<p>Hey Sarath, we are on the same boat. Straightforward ensemble is indeed not easy to apply.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 468226,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-02-08T14:27:54.720000",
          "content": "<p>Waiting to see the top place solutions, in 2-3 weeks.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "467328": "Now that the competitive phase is over, let's make the best of what each of us has learned!\n\nWhich tricks did you guys come up with that weren't shared in any kernels?\n\nFor my part, I realized at some point that I wasn't pushing single model CV score any further. So instead, I put all my effort into diversifying my ensemble without sacrificing too much accuracy. To measure, I used a part of the training data as a \"dev set\". On that, I measured the F1 score of an ensemble of my models.\n\nHere are some of the tricks I came up with:\n\n- For each model, average the three embedding matrixes with different weights. I found good and diverse ones through random search and luck\n- Randomized sample weights for each model. That way, each model focuses on different examples and should reach different solutions\n- Re-initialize the random embedding matrix between runs. Since a lot of words don't have embeddings and thus use random vectors, this also helps diversify\n- Add some random features to the embedding, different for each model. The models can overfit very slightly to those, also leading to increased diversification.\n- Replace some random embedding features with a random vector for each model\n- Overall, train longer than I would do otherwise. A slight overfit always seemed to help my ensemble F1.\n- Each model was trained on a different subset of the data and with different layer sizes, dropout values, loss functions, etc.\n\nWhat about you?\n\n*EDIT:*\nNow a major motion picture: https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness",
    "467633": "One thing i found was that I wanted to submit a single model. But Ensemble has its benefits so didnt need to lose on that. I created predictions of test data after each epoch and did a weighted ensembling of those predictions to arrive at a final submission where i gave lower weights to start epochs. Gave me a boost from 0.68 to 0.69 Local CV.",
    "467802": "1. Using BCE+soft f1 loss gave me a little boost (0.003) on LB and (0.001) on CV on the last day. It gave stabler f1 score.\n2. Regarding overfit, training with more epochs gives less sensitive threshold v. f1 score plot at the optimal point. It helps when training and test sets are different. That's why all the 0.7 kernels overfit the training set.",
    "467358": "I mainly focused on improving Kfold local CV score over LB score. Couple of simple steps that helped were: \n\n- Training on trainable embeddings for one epoch after slightly overfitting to fixed embeddings. Reinitialized embeddings after each fold.\n- Gaussian noise after embedding layer in network during training, as mentioned in this [discussion](https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/78841).\n- Divide vectors in embedding matrix by their frequency in the provided embeddings (Used all embeddings except word2vec since it wasn't improving the score for me). That way if a word was present in only one of the embeddings, I wasn't diminishing it's weight by dividing it by 3.\n\n```\n\n    global_embedding = np.random.normal(global_mean, global_std, (nb_words, embed_size))\n    embedding_count = np.zeros((nb_words,1))\n    for EMBEDDING_FILE in embedding_list:\n        for o in open(EMBEDDING_FILE, encoding=\"utf8\", errors='ignore'):\n            word, vec = o.split(' ', 1)\n            if word not in word_index:\n                continue\n            i = word_index[word]\n            if i &gt;= nb_words:\n                continue\n            embedding_vector = np.asarray(vec.split(' '), dtype='float32')[:embed_size]\n            if len(embedding_vector) == embed_size:\n                global_embedding[i] = (embedding_count[i]*global_embedding[i] + embedding_vector)/(embedding_count[i] + 1)\n                embedding_count[i] += 1\n        del embedding_vector\n        gc.collect()\n\n```",
    "471479": "Hi Max, Annabelle\n@mschumacher, @sezaugg\n\nAllow me to congratulate both of you here for the final performance!\nThanks for all your sharing during the competition, and see you again soon!",
    "471566": "For those who were wondering, seems like my ideas weren't completely off the mark! I got in 44th. I published the kernel here: https://www.kaggle.com/mschumacher/44th-place-add-all-the-randomness",
    "468044": "Hello! Can we use data augmentation for the insincere questions to train the model?",
    "467545": "Thanks for the post, it is probably a good time to share insights as in 3 weeks we will all be busy with other competitions.\n\n- I also found that 3 epochs generally gave best results for a single NN, but 4-5 epochs better when ensembling several NNs via mean. As you mentioned: slight overfit helps ensemble performance\n- CV-trained models have more diversity but none of the models learns from all data. Thus, I switched to train each NN from the ensemble on all training data for submissions. This had a quite significant boost on performance. For local experiments, I used a single split: 90% of data for training and 10% for test. \n- I tried many different NN architectures (because it is fun, let's admit it) but found that an adequately sized single Bi-LSTM (or GRU) was the main driver of performance. I finally used a single Bi-LSTM layer. GRU tended to be slightly worse in my particular setting.\n- Early,  from public kernels I realized that we must separate words from special characters and numbers. This had a quite significant boost on performance.\n\nEDIT: \n- Too late, I started a \"sort of spell-check\", i.e. using pre-processing pipeline that transformed OOV words to closest word available in embedding to increase coverage of the embedding. It seemed to increase performance consistently, but I did not have time to finalize this.\n\n "
  }
}