{
  "id": 80568,
  "title": "1st place solution",
  "url": "/competitions/quora-insincere-questions-classification/discussion/80568",
  "author_name": "Psi",
  "post_date": "2019-02-14T14:14:24.360000",
  "votes": 398,
  "comment_count": 127,
  "views": 0,
  "content": "<p>First of all, we want to thank Kaggle for hosting the competition and Quora for providing such a large dataset. Last 3 months were quite exhausting for us with a steep learning curve and tons of the ideas we wanted to try out. In the following we try to summarize some of the main points of our solution.</p>\n\n<p><strong>Model Structure</strong>\nWe played around with a variety of different model structures, but in the end resorted to a quite simple one that is very similar to those posted here <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79824\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79824</a>. It’s basically a Single Bi-LSTM 128 followed by a Conv1D with kernel size 1 only and GlobalMaxPooling afterwards plus additional dropout layers with minimal dropout. We additionally use a few statistical features.</p>\n\n<p><img src=\"https://i.imgur.com/zUY9tVN.png\" alt=\"enter image description here\"></p>\n\n<p><strong>Embeddings</strong>\nFirst, we use all tokens from both train and test data for our vocabulary. We do the simple pre-cleaning that was posted in a kernel at the start of the competition and split by space afterwards (spacy and nltk resulted in similar performance). We do not lowercase, but keep uppercase, and do not limit the vocab at all. For embeddings we use glove and para where we weight glove a bit higher. The most important thing now is to find as many embeddings as possible for our vocabulary. We had a few steps to achieve this, like checking singular and plural of the word, checking lowercase embeddings, removing special tokens, etc. For public test data we had around 50k of vocab tokens we did not find in the embeddings afterwards. Even though we tried a few different strategies for handling the OOV tokens, we resorted to a single OOV token with a single random embedding vector. </p>\n\n<p><strong>Threshold</strong>\nWe spent a lot of time trying to figure out good strategies for choosing a good threshold for classification. Over time, we saw that estimating the threshold on validation data and then applying it on test data does not really work. There is a large variation on optimal thresholds. So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data. The visualization below shows that in action (not necessarily our final eval). On the x-axis we plot the different fixed thresholds and on the y-axis we see the deviation from the optimal F1 score across folds using this fixed threshold (see CV chapter below). The blue line is the mean, green is median, purple is minimum, red is maximum, and bars are stds. So for example here, if we choose a threshold in the range of 0.927 we expect the F1 score to be not much worse (around 0.001) compared to choosing the optimal threshold (which we can't do for test data). In practise, this might of course deviate further and we could also see larger deviations on PLB. For further elaboration, please check the comments.</p>\n\n<p><img src=\"https://i.imgur.com/2NPwBIR.png\" alt=\"enter image description here\"></p>\n\n<p><strong>Runtime tricks</strong>\nWe aimed at combining as many models as possible. To do this, we needed to improve runtime and the most important thing to achieve this was the following. We do not pad sequences to the same length based on the whole data, but just on a batch level. That means we conduct padding and truncation on the data generator level for each batch separately, so that length of the sentences in a batch can vary in size. Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. This improved runtime heavily and kept accuracy quite robust on single model level, and improved it by being able to average more models.</p>\n\n<p><strong>Fitting</strong>\nWe use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512. We could achieve similar results by even taking 10 or 20 times higher batch sizes, which goes hand in hand with recent research on fast convergence. With these larger batch sizes we could even fit close to 20 models, but results stabilized close to 10 models which is why we chose to go with the smaller batch size in the end. However, there might still be some room left here if one properly tunes this.</p>\n\n<p><strong>Multiple models</strong>\nIn the end, we managed to fit more than 10 models on the complete training dataset with help of the runtime tricks mentioned before. Our best final private score even had only a runtime of 6000 seconds (I think they used a bit better hardware for running), so there would be space for 1-2 more models. With larger batch sizes even much more might be feasible. As mentioned, we then average the rank predictions of each model and use our specified threshold for prediction.</p>\n\n<p><strong>Embrace the randomness</strong>\nAs it was necessary to utilize CUDNN Layers in this competition, there was some randomness involved that could be quite frustrating from time to time. I saw many people trying to fix seeds etc. and some claiming they could completely remove the randomness by using Pytorch (I still don’t believe this BTW as CUDNN has atomic operations). However, as mentioned before, a well working strategy in this competition was to combine multiple models and to end up with a good ensemble, those models should be a bit different to each other. So having different random initializations etc. can be helpful. Seeing people setting the seed as a hyperparameter is weird.</p>\n\n<p><strong>CV Evaluation</strong>\nWhat I saw many people doing wrongly in this competition, and we also only figured this out after a while, is to trust their single out-of-fold evaluation. However, in this competition, it is crucial to combine (average) multiple models (in our case the same model). That means that our CV evaluation looks like the following. We do a k-fold split (mostly 10-fold) and fit the same model up to v-times on the same training split and then successively evaluate it on the single out of fold. So for the first split, we first fit one model and evaluate it, then a second one and evaluate the average and so forth. We repeat this for all 10 folds, landing us with e.g., 100 model fits overall, and then we can take a look at the median or mean over all folds for v-model-ensembles. The reason for doing this is that f1 scores are very different on the split you have. For one 10% split you might end up with a maximum of 0.72 and for the other you might end up at 0.705 or similar. So repeating the split 10 times, fitting the same model v-times for each split, and then looking at the grand picture gave us the best overall evaluation. This routine helped us to compare individual solutions with each other. BTW our final scores are exactly what we would expect from our CV evaluation, but again this might be lucky :)</p>\n\n<p><strong>Robustness and over/underfitting</strong>\nAround 2 weeks before final submission, our results became so stable that changing things did not alter results much. Things like finding more OOV embedding vectors resulted in same results, using slightly different layers ended up being similar, and other things. This was a bit frustrating, but in the end things worked out. In the end, it was important to find a good balance between over and underfitting (as always). Underfitting too much led to good single model performances, but was worse for combining models, and the other way around. For example, if your model overfits, there can be many different solutions to tackle this, e.g., add dropouts, or reduce the vocab size, or reduce model complexity, etc. So if someone says on kaggle that one things works for him/her, that does not necessarily mean that it will work for you as you might already be doing something similar that has similar effects (a good example is the Gaussian noise discussion).</p>\n\n<p><strong>What did not work for us</strong>\nMostly you only read what worked, but here is an incomplete shortlist of what did not work for us. This does not mean that it doesn’t work at all, but rather that it was worse for our specific solution.</p>\n\n<ul>\n<li>Different optimizers (focal loss was similar though)</li>\n<li>Label smoothing</li>\n<li>Auxiliary learning / multitask learning</li>\n<li>Snapshot learning</li>\n<li>Pseudo labeling</li>\n<li>Fitting own embeddings with gensim</li>\n<li>Spelling correction</li>\n<li>Taking median/percentile of predicitons instead of average</li>\n<li>More complex layers and architectures (Attention, QRNN, Capsule, larger/multiple LSTM layers, larger CNN kernel sizes, LGBM or bag of words)</li>\n<li>Word collocations - Several words put together can bear a completely new meaning, which is not captured by embeddings. Glove turned out to have quite a lot of such collocations with words put together using \"-\" sign. So we replaced examples like \"ethnical cleansing\" with \"ethnical-cleansing\", which is then captured by a more appropriate glove embedding. It showed no improvement on CV.</li>\n<li>Extra statistical features - Presence of statistical features added a little bit to the accuracy based on CV, but we saw no improvement with other extra features, like sentiment or bag-of-words based variables.</li>\n<li>Replacement of words with synonyms - An idea of replacing all nationalities (or e.g. political party) with the same word did not work at all.</li>\n<li>Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.</li>\n</ul>",
  "messages": [
    {
      "id": 471476,
      "postDate": "2019-02-14T14:14:24.360Z",
      "content": "<p>First of all, we want to thank Kaggle for hosting the competition and Quora for providing such a large dataset. Last 3 months were quite exhausting for us with a steep learning curve and tons of the ideas we wanted to try out. In the following we try to summarize some of the main points of our solution.</p>\n\n<p><strong>Model Structure</strong>\nWe played around with a variety of different model structures, but in the end resorted to a quite simple one that is very similar to those posted here <a href=\"https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79824\">https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79824</a>. It’s basically a Single Bi-LSTM 128 followed by a Conv1D with kernel size 1 only and GlobalMaxPooling afterwards plus additional dropout layers with minimal dropout. We additionally use a few statistical features.</p>\n\n<p><img src=\"https://i.imgur.com/zUY9tVN.png\" alt=\"enter image description here\"></p>\n\n<p><strong>Embeddings</strong>\nFirst, we use all tokens from both train and test data for our vocabulary. We do the simple pre-cleaning that was posted in a kernel at the start of the competition and split by space afterwards (spacy and nltk resulted in similar performance). We do not lowercase, but keep uppercase, and do not limit the vocab at all. For embeddings we use glove and para where we weight glove a bit higher. The most important thing now is to find as many embeddings as possible for our vocabulary. We had a few steps to achieve this, like checking singular and plural of the word, checking lowercase embeddings, removing special tokens, etc. For public test data we had around 50k of vocab tokens we did not find in the embeddings afterwards. Even though we tried a few different strategies for handling the OOV tokens, we resorted to a single OOV token with a single random embedding vector. </p>\n\n<p><strong>Threshold</strong>\nWe spent a lot of time trying to figure out good strategies for choosing a good threshold for classification. Over time, we saw that estimating the threshold on validation data and then applying it on test data does not really work. There is a large variation on optimal thresholds. So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data. The visualization below shows that in action (not necessarily our final eval). On the x-axis we plot the different fixed thresholds and on the y-axis we see the deviation from the optimal F1 score across folds using this fixed threshold (see CV chapter below). The blue line is the mean, green is median, purple is minimum, red is maximum, and bars are stds. So for example here, if we choose a threshold in the range of 0.927 we expect the F1 score to be not much worse (around 0.001) compared to choosing the optimal threshold (which we can't do for test data). In practise, this might of course deviate further and we could also see larger deviations on PLB. For further elaboration, please check the comments.</p>\n\n<p><img src=\"https://i.imgur.com/2NPwBIR.png\" alt=\"enter image description here\"></p>\n\n<p><strong>Runtime tricks</strong>\nWe aimed at combining as many models as possible. To do this, we needed to improve runtime and the most important thing to achieve this was the following. We do not pad sequences to the same length based on the whole data, but just on a batch level. That means we conduct padding and truncation on the data generator level for each batch separately, so that length of the sentences in a batch can vary in size. Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. This improved runtime heavily and kept accuracy quite robust on single model level, and improved it by being able to average more models.</p>\n\n<p><strong>Fitting</strong>\nWe use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512. We could achieve similar results by even taking 10 or 20 times higher batch sizes, which goes hand in hand with recent research on fast convergence. With these larger batch sizes we could even fit close to 20 models, but results stabilized close to 10 models which is why we chose to go with the smaller batch size in the end. However, there might still be some room left here if one properly tunes this.</p>\n\n<p><strong>Multiple models</strong>\nIn the end, we managed to fit more than 10 models on the complete training dataset with help of the runtime tricks mentioned before. Our best final private score even had only a runtime of 6000 seconds (I think they used a bit better hardware for running), so there would be space for 1-2 more models. With larger batch sizes even much more might be feasible. As mentioned, we then average the rank predictions of each model and use our specified threshold for prediction.</p>\n\n<p><strong>Embrace the randomness</strong>\nAs it was necessary to utilize CUDNN Layers in this competition, there was some randomness involved that could be quite frustrating from time to time. I saw many people trying to fix seeds etc. and some claiming they could completely remove the randomness by using Pytorch (I still don’t believe this BTW as CUDNN has atomic operations). However, as mentioned before, a well working strategy in this competition was to combine multiple models and to end up with a good ensemble, those models should be a bit different to each other. So having different random initializations etc. can be helpful. Seeing people setting the seed as a hyperparameter is weird.</p>\n\n<p><strong>CV Evaluation</strong>\nWhat I saw many people doing wrongly in this competition, and we also only figured this out after a while, is to trust their single out-of-fold evaluation. However, in this competition, it is crucial to combine (average) multiple models (in our case the same model). That means that our CV evaluation looks like the following. We do a k-fold split (mostly 10-fold) and fit the same model up to v-times on the same training split and then successively evaluate it on the single out of fold. So for the first split, we first fit one model and evaluate it, then a second one and evaluate the average and so forth. We repeat this for all 10 folds, landing us with e.g., 100 model fits overall, and then we can take a look at the median or mean over all folds for v-model-ensembles. The reason for doing this is that f1 scores are very different on the split you have. For one 10% split you might end up with a maximum of 0.72 and for the other you might end up at 0.705 or similar. So repeating the split 10 times, fitting the same model v-times for each split, and then looking at the grand picture gave us the best overall evaluation. This routine helped us to compare individual solutions with each other. BTW our final scores are exactly what we would expect from our CV evaluation, but again this might be lucky :)</p>\n\n<p><strong>Robustness and over/underfitting</strong>\nAround 2 weeks before final submission, our results became so stable that changing things did not alter results much. Things like finding more OOV embedding vectors resulted in same results, using slightly different layers ended up being similar, and other things. This was a bit frustrating, but in the end things worked out. In the end, it was important to find a good balance between over and underfitting (as always). Underfitting too much led to good single model performances, but was worse for combining models, and the other way around. For example, if your model overfits, there can be many different solutions to tackle this, e.g., add dropouts, or reduce the vocab size, or reduce model complexity, etc. So if someone says on kaggle that one things works for him/her, that does not necessarily mean that it will work for you as you might already be doing something similar that has similar effects (a good example is the Gaussian noise discussion).</p>\n\n<p><strong>What did not work for us</strong>\nMostly you only read what worked, but here is an incomplete shortlist of what did not work for us. This does not mean that it doesn’t work at all, but rather that it was worse for our specific solution.</p>\n\n<ul>\n<li>Different optimizers (focal loss was similar though)</li>\n<li>Label smoothing</li>\n<li>Auxiliary learning / multitask learning</li>\n<li>Snapshot learning</li>\n<li>Pseudo labeling</li>\n<li>Fitting own embeddings with gensim</li>\n<li>Spelling correction</li>\n<li>Taking median/percentile of predicitons instead of average</li>\n<li>More complex layers and architectures (Attention, QRNN, Capsule, larger/multiple LSTM layers, larger CNN kernel sizes, LGBM or bag of words)</li>\n<li>Word collocations - Several words put together can bear a completely new meaning, which is not captured by embeddings. Glove turned out to have quite a lot of such collocations with words put together using \"-\" sign. So we replaced examples like \"ethnical cleansing\" with \"ethnical-cleansing\", which is then captured by a more appropriate glove embedding. It showed no improvement on CV.</li>\n<li>Extra statistical features - Presence of statistical features added a little bit to the accuracy based on CV, but we saw no improvement with other extra features, like sentiment or bag-of-words based variables.</li>\n<li>Replacement of words with synonyms - An idea of replacing all nationalities (or e.g. political party) with the same word did not work at all.</li>\n<li>Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.</li>\n</ul>",
      "rawMarkdown": "First of all, we want to thank Kaggle for hosting the competition and Quora for providing such a large dataset. Last 3 months were quite exhausting for us with a steep learning curve and tons of the ideas we wanted to try out. In the following we try to summarize some of the main points of our solution.\n\n**Model Structure**\nWe played around with a variety of different model structures, but in the end resorted to a quite simple one that is very similar to those posted here https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79824. It’s basically a Single Bi-LSTM 128 followed by a Conv1D with kernel size 1 only and GlobalMaxPooling afterwards plus additional dropout layers with minimal dropout. We additionally use a few statistical features.\n\n![enter image description here][1]\n\n**Embeddings**\nFirst, we use all tokens from both train and test data for our vocabulary. We do the simple pre-cleaning that was posted in a kernel at the start of the competition and split by space afterwards (spacy and nltk resulted in similar performance). We do not lowercase, but keep uppercase, and do not limit the vocab at all. For embeddings we use glove and para where we weight glove a bit higher. The most important thing now is to find as many embeddings as possible for our vocabulary. We had a few steps to achieve this, like checking singular and plural of the word, checking lowercase embeddings, removing special tokens, etc. For public test data we had around 50k of vocab tokens we did not find in the embeddings afterwards. Even though we tried a few different strategies for handling the OOV tokens, we resorted to a single OOV token with a single random embedding vector. \n\n**Threshold**\nWe spent a lot of time trying to figure out good strategies for choosing a good threshold for classification. Over time, we saw that estimating the threshold on validation data and then applying it on test data does not really work. There is a large variation on optimal thresholds. So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data. The visualization below shows that in action (not necessarily our final eval). On the x-axis we plot the different fixed thresholds and on the y-axis we see the deviation from the optimal F1 score across folds using this fixed threshold (see CV chapter below). The blue line is the mean, green is median, purple is minimum, red is maximum, and bars are stds. So for example here, if we choose a threshold in the range of 0.927 we expect the F1 score to be not much worse (around 0.001) compared to choosing the optimal threshold (which we can't do for test data). In practise, this might of course deviate further and we could also see larger deviations on PLB. For further elaboration, please check the comments.\n\n![enter image description here][2]\n\n**Runtime tricks**\nWe aimed at combining as many models as possible. To do this, we needed to improve runtime and the most important thing to achieve this was the following. We do not pad sequences to the same length based on the whole data, but just on a batch level. That means we conduct padding and truncation on the data generator level for each batch separately, so that length of the sentences in a batch can vary in size. Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. This improved runtime heavily and kept accuracy quite robust on single model level, and improved it by being able to average more models.\n\n**Fitting**\nWe use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512. We could achieve similar results by even taking 10 or 20 times higher batch sizes, which goes hand in hand with recent research on fast convergence. With these larger batch sizes we could even fit close to 20 models, but results stabilized close to 10 models which is why we chose to go with the smaller batch size in the end. However, there might still be some room left here if one properly tunes this.\n\n**Multiple models**\nIn the end, we managed to fit more than 10 models on the complete training dataset with help of the runtime tricks mentioned before. Our best final private score even had only a runtime of 6000 seconds (I think they used a bit better hardware for running), so there would be space for 1-2 more models. With larger batch sizes even much more might be feasible. As mentioned, we then average the rank predictions of each model and use our specified threshold for prediction.\n\n**Embrace the randomness**\nAs it was necessary to utilize CUDNN Layers in this competition, there was some randomness involved that could be quite frustrating from time to time. I saw many people trying to fix seeds etc. and some claiming they could completely remove the randomness by using Pytorch (I still don’t believe this BTW as CUDNN has atomic operations). However, as mentioned before, a well working strategy in this competition was to combine multiple models and to end up with a good ensemble, those models should be a bit different to each other. So having different random initializations etc. can be helpful. Seeing people setting the seed as a hyperparameter is weird.\n\n**CV Evaluation**\nWhat I saw many people doing wrongly in this competition, and we also only figured this out after a while, is to trust their single out-of-fold evaluation. However, in this competition, it is crucial to combine (average) multiple models (in our case the same model). That means that our CV evaluation looks like the following. We do a k-fold split (mostly 10-fold) and fit the same model up to v-times on the same training split and then successively evaluate it on the single out of fold. So for the first split, we first fit one model and evaluate it, then a second one and evaluate the average and so forth. We repeat this for all 10 folds, landing us with e.g., 100 model fits overall, and then we can take a look at the median or mean over all folds for v-model-ensembles. The reason for doing this is that f1 scores are very different on the split you have. For one 10% split you might end up with a maximum of 0.72 and for the other you might end up at 0.705 or similar. So repeating the split 10 times, fitting the same model v-times for each split, and then looking at the grand picture gave us the best overall evaluation. This routine helped us to compare individual solutions with each other. BTW our final scores are exactly what we would expect from our CV evaluation, but again this might be lucky :)\n\n**Robustness and over/underfitting**\nAround 2 weeks before final submission, our results became so stable that changing things did not alter results much. Things like finding more OOV embedding vectors resulted in same results, using slightly different layers ended up being similar, and other things. This was a bit frustrating, but in the end things worked out. In the end, it was important to find a good balance between over and underfitting (as always). Underfitting too much led to good single model performances, but was worse for combining models, and the other way around. For example, if your model overfits, there can be many different solutions to tackle this, e.g., add dropouts, or reduce the vocab size, or reduce model complexity, etc. So if someone says on kaggle that one things works for him/her, that does not necessarily mean that it will work for you as you might already be doing something similar that has similar effects (a good example is the Gaussian noise discussion).\n\n**What did not work for us**\nMostly you only read what worked, but here is an incomplete shortlist of what did not work for us. This does not mean that it doesn’t work at all, but rather that it was worse for our specific solution.\n\n- Different optimizers (focal loss was similar though)\n-  Label smoothing\n-  Auxiliary learning / multitask learning\n-  Snapshot learning\n-  Pseudo labeling\n-  Fitting own embeddings with gensim\n-  Spelling correction\n-  Taking median/percentile of predicitons instead of average\n-  More complex layers and architectures (Attention, QRNN, Capsule, larger/multiple LSTM layers, larger CNN kernel sizes, LGBM or bag of words)\n-  Word collocations - Several words put together can bear a completely new meaning, which is not captured by embeddings. Glove turned out to have quite a lot of such collocations with words put together using \"-\" sign. So we replaced examples like \"ethnical cleansing\" with \"ethnical-cleansing\", which is then captured by a more appropriate glove embedding. It showed no improvement on CV.\n-  Extra statistical features - Presence of statistical features added a little bit to the accuracy based on CV, but we saw no improvement with other extra features, like sentiment or bag-of-words based variables.\n-  Replacement of words with synonyms - An idea of replacing all nationalities (or e.g. political party) with the same word did not work at all.\n-  Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.\n\n\n  [1]: https://i.imgur.com/zUY9tVN.png\n  [2]: https://i.imgur.com/2NPwBIR.png",
      "votes": 397
    },
    {
      "id": 471814,
      "postDate": "2019-02-14T23:13:19.793Z",
      "content": "<p>congrats, can you open source your kernel so we know exactly what you did? look forward to it!</p>",
      "rawMarkdown": "congrats, can you open source your kernel so we know exactly what you did? look forward to it!",
      "votes": 34
    },
    {
      "id": 1255628,
      "postDate": "2021-03-29T02:10:10.067Z",
      "content": "<p>Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.</p>\n<p>Regarding this point <a href=\"https://www.kaggle.com/Psi\" target=\"_blank\">@Psi</a>, what can be the reason for this? I have also tried this in some other sequence model, the accuracy dropped dramatically.</p>",
      "rawMarkdown": "Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.\n\nRegarding this point @Psi, what can be the reason for this? I have also tried this in some other sequence model, the accuracy dropped dramatically.",
      "votes": 1,
      "replies": [
        {
          "id": 1261809,
          "postDate": "2021-04-03T13:26:15.617Z",
          "content": "<p>The model focuses only on certain sequence lengths in each batch and has no diversity.</p>",
          "rawMarkdown": "The model focuses only on certain sequence lengths in each batch and has no diversity.",
          "votes": 1
        }
      ]
    },
    {
      "id": 472122,
      "postDate": "2019-02-15T11:48:01.430Z",
      "content": "<p>Could you elaborate about\n 1. Which statistical features did you use?\n 2. On the model architecture the 1D conv follows a BiLSTM, what was done there exactly? Did you use all the states (from all the steps)? Or did you concat the last outputs of each LSTM?</p>",
      "rawMarkdown": "Could you elaborate about\n 1. Which statistical features did you use?\n 2. On the model architecture the 1D conv follows a BiLSTM, what was done there exactly? Did you use all the states (from all the steps)? Or did you concat the last outputs of each LSTM?",
      "votes": 5,
      "replies": [
        {
          "id": 472178,
          "postDate": "2019-02-15T13:01:34.987Z",
          "content": "<ol>\n<li>The statistical features were: length of the text, number of capital letters, number of exclamation/question/punctuation marks, number of special symbols, number of smileys, number of words, number of unique words and few derivatives.</li>\n<li>We have a Bidirectional LSTM returning a sequence, then a convolution, followed by max pooling.</li>\n</ol>",
          "rawMarkdown": "1. The statistical features were: length of the text, number of capital letters, number of exclamation/question/punctuation marks, number of special symbols, number of smileys, number of words, number of unique words and few derivatives.\n2. We have a Bidirectional LSTM returning a sequence, then a convolution, followed by max pooling.",
          "votes": 20
        },
        {
          "id": 1975464,
          "postDate": "2022-10-06T19:33:59.370Z",
          "content": "<p>Can you kindly elaborate on how did you come up with the idea of applying convolution layer here (Why did you use it)?</p>",
          "rawMarkdown": "Can you kindly elaborate on how did you come up with the idea of applying convolution layer here (Why did you use it)?"
        }
      ]
    },
    {
      "id": 471482,
      "postDate": "2019-02-14T14:19:13.843Z",
      "content": "<p>Congratulation, and thank you for sharing the solution which is truly educative!!\nYou have done much work during these 3 months and totally deserve it.</p>",
      "rawMarkdown": "Congratulation, and thank you for sharing the solution which is truly educative!!\nYou have done much work during these 3 months and totally deserve it.\n",
      "votes": 4
    },
    {
      "id": 474345,
      "postDate": "2019-02-19T08:46:05.543Z",
      "content": "<p>i like your EDA</p>",
      "rawMarkdown": "i like your EDA",
      "votes": 3
    },
    {
      "id": 471629,
      "postDate": "2019-02-14T17:28:00.223Z",
      "content": "<p>It's very interesting that you guys did the padding on the fly. We tried the same thing, but our batchsizes were much larger so we ended up with almost exactly the same runtime as if we just prepadded. Did not consider reducing batch_size and only doing the top 95 percent. Very clever trick. I wonder with both of our trick combined how many models could be fit. </p>",
      "rawMarkdown": "It's very interesting that you guys did the padding on the fly. We tried the same thing, but our batchsizes were much larger so we ended up with almost exactly the same runtime as if we just prepadded. Did not consider reducing batch_size and only doing the top 95 percent. Very clever trick. I wonder with both of our trick combined how many models could be fit. ",
      "votes": 4,
      "replies": [
        {
          "id": 471633,
          "postDate": "2019-02-14T17:33:42.830Z",
          "content": "<p>As written, we actually tried much larger batch sizes up to 10240 allowing us to fit close to 20 models. Results looked good as well, but we saw convergence after combining around 10 models or so and would have needed to further fine-tune the higher batch size models which is why we submitted the 512 batch size models. But as you say, the higher the batch size to choose, the larger the max sequence length is, ending up in exactly what you observed. </p>",
          "rawMarkdown": "As written, we actually tried much larger batch sizes up to 10240 allowing us to fit close to 20 models. Results looked good as well, but we saw convergence after combining around 10 models or so and would have needed to further fine-tune the higher batch size models which is why we submitted the 512 batch size models. But as you say, the higher the batch size to choose, the larger the max sequence length is, ending up in exactly what you observed. ",
          "votes": 5
        },
        {
          "id": 471852,
          "postDate": "2019-02-15T01:40:05.200Z",
          "content": "<p>Padding on the fly is standard practice. Those who are familiar with Tensorflow, padding happens only dynamically (not only, mostly). This is the problem when people start using keras and some default options. I have shared my kernel here <a href=\"https://www.kaggle.com/s4sarath/cudnngru-best-final\">https://www.kaggle.com/s4sarath/cudnngru-best-final</a> . With a max length of 150 words per sentences, I was able to train a model in 3600 seconds ( 5 fold with overall 23 epochs), As @psi told, I have also created 10-12 models roughly, but was no helping in improving PB. Thanks, @psi, and team for detailed information.</p>",
          "rawMarkdown": "Padding on the fly is standard practice. Those who are familiar with Tensorflow, padding happens only dynamically (not only, mostly). This is the problem when people start using keras and some default options. I have shared my kernel here https://www.kaggle.com/s4sarath/cudnngru-best-final . With a max length of 150 words per sentences, I was able to train a model in 3600 seconds ( 5 fold with overall 23 epochs), As @psi told, I have also created 10-12 models roughly, but was no helping in improving PB. Thanks, @psi, and team for detailed information.",
          "votes": 5
        },
        {
          "id": 530061,
          "postDate": "2019-05-11T15:54:35.523Z",
          "content": "<p>@psi, <a href=\"/ryches\">@ryches</a> What do we mean by \"padding on fly\"? Right now, I am padding sequences as data preparation step, like \nX_train = sequence.pad_sequences(X_train,maxlen=MAX_LEN)\nX_valid = sequence.pad_sequences(X_valid,maxlen=MAX_LEN)</p>",
          "rawMarkdown": "@psi, @ryches What do we mean by \"padding on fly\"? Right now, I am padding sequences as data preparation step, like \nX_train = sequence.pad_sequences(X_train,maxlen=MAX_LEN)\nX_valid = sequence.pad_sequences(X_valid,maxlen=MAX_LEN)\n",
          "votes": 1
        }
      ]
    },
    {
      "id": 476532,
      "postDate": "2019-02-22T09:59:06.270Z",
      "content": "<p>awesome</p>",
      "rawMarkdown": "awesome",
      "votes": 1
    },
    {
      "id": 474510,
      "postDate": "2019-02-19T13:28:46.623Z",
      "content": "<p>Impressive </p>",
      "rawMarkdown": "Impressive ",
      "votes": 1
    },
    {
      "id": 473273,
      "postDate": "2019-02-17T17:41:27.870Z",
      "content": "<p>What an impressive way. Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "What an impressive way. Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472947,
      "postDate": "2019-02-17T00:35:32.077Z",
      "content": "<p>Congrats on winning and fantastic explanation throughout !! </p>",
      "rawMarkdown": "Congrats on winning and fantastic explanation throughout !! ",
      "votes": 1
    },
    {
      "id": 472515,
      "postDate": "2019-02-16T04:33:21.767Z",
      "content": "<p>Congrats! and thanks for sharing your insights.</p>",
      "rawMarkdown": "Congrats! and thanks for sharing your insights.",
      "votes": 1
    },
    {
      "id": 472436,
      "postDate": "2019-02-15T22:59:43.150Z",
      "content": "<p>Many interesting points here, practical advise and great writeup. Thanks a lot for sharing!</p>",
      "rawMarkdown": "Many interesting points here, practical advise and great writeup. Thanks a lot for sharing!",
      "votes": 1
    },
    {
      "id": 472390,
      "postDate": "2019-02-15T20:11:10.663Z",
      "content": "<p>Congrats and thanks for sharing your solution. Your solution is really excellent!</p>",
      "rawMarkdown": "Congrats and thanks for sharing your solution. Your solution is really excellent!",
      "votes": 1
    },
    {
      "id": 472377,
      "postDate": "2019-02-15T19:39:21.477Z",
      "content": "<p>Very interesting model structure.Thanks for sharing!</p>",
      "rawMarkdown": "Very interesting model structure.Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 471917,
      "postDate": "2019-02-15T04:41:40.363Z",
      "content": "<p>Congrast and thanks for sharing your work.</p>",
      "rawMarkdown": "Congrast and thanks for sharing your work.",
      "votes": 1
    },
    {
      "id": 471883,
      "postDate": "2019-02-15T03:10:16.867Z",
      "content": "<p>Thank you for the write ups, and congratulations on the strong finish! </p>",
      "rawMarkdown": "Thank you for the write ups, and congratulations on the strong finish! ",
      "votes": 1
    },
    {
      "id": 471875,
      "postDate": "2019-02-15T02:50:27.690Z",
      "content": "<p>Congrats! Really helped, and thanks for sharing.</p>",
      "rawMarkdown": "Congrats! Really helped, and thanks for sharing.",
      "votes": 1
    },
    {
      "id": 471834,
      "postDate": "2019-02-15T00:53:56.420Z",
      "content": "<p>Congratulations <a href=\"/philippsinger\">@philippsinger</a> and team for wining this amazing competition. Thank you for sharing your solution.</p>",
      "rawMarkdown": "Congratulations @philippsinger and team for wining this amazing competition. Thank you for sharing your solution.",
      "votes": 1
    },
    {
      "id": 471643,
      "postDate": "2019-02-14T17:57:48.483Z",
      "content": "<p>Congratulations. @psi. Thanks for sharing your work and I am really impressed with the idea of bucketing.</p>",
      "rawMarkdown": "Congratulations. @psi. Thanks for sharing your work and I am really impressed with the idea of bucketing.",
      "votes": 1
    },
    {
      "id": 471619,
      "postDate": "2019-02-14T16:57:18.133Z",
      "content": "<p>Congrats! Really nice work!</p>",
      "rawMarkdown": "Congrats! Really nice work!",
      "votes": 1
    },
    {
      "id": 471617,
      "postDate": "2019-02-14T16:50:42.257Z",
      "content": "<p>Congratulations and well done! Thank you for the explanation and what a great idea with the padding!</p>",
      "rawMarkdown": "Congratulations and well done! Thank you for the explanation and what a great idea with the padding!",
      "votes": 1
    },
    {
      "id": 471549,
      "postDate": "2019-02-14T15:42:50.430Z",
      "content": "<p>Congrats, you really deserve it. Thanks for detailed explanation.</p>",
      "rawMarkdown": "Congrats, you really deserve it. Thanks for detailed explanation.",
      "votes": 1
    },
    {
      "id": 474782,
      "postDate": "2019-02-19T19:39:08.670Z",
      "content": "<p>Certainly pointers can be taken from involved data</p>",
      "rawMarkdown": "Certainly pointers can be taken from involved data",
      "votes": 2
    },
    {
      "id": 472794,
      "postDate": "2019-02-16T17:21:04.663Z",
      "content": "<p>Congrats...</p>",
      "rawMarkdown": "Congrats...",
      "votes": 2
    },
    {
      "id": 471796,
      "postDate": "2019-02-14T22:21:58.670Z",
      "content": "<p>Very nice. Congratulations!</p>\n\n<p>What do you mean by:\n\"So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data.\"</p>\n\n<p>Could you explain it with an example maybe?</p>",
      "rawMarkdown": "Very nice. Congratulations!\n\nWhat do you mean by:\n\"So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data.\"\n\nCould you explain it with an example maybe?",
      "votes": 2,
      "replies": [
        {
          "id": 471998,
          "postDate": "2019-02-15T07:40:59.497Z",
          "content": "<p>Regarding rank prediction: Assume you have predicted probabilities for a single model, you then transform them into ranks (e.g., rankdata in numpy). Then you average the ranks when combining individual models and divide by the length so your final predictions again end up between zero and one. For then finding a fixed threshold a simple strategy is to just take the mean best threshold on multiple CV runs or similar simulations. However, there are still outliers in both directions depending on the split so if you are really unlucky your fixed threshold is \"far\" away from the optimal. So what we tried is testing various fixed thresholds and evaluate how far the resulting F1 score is compared to if you would take the optimal threshold for this fold. We then finally chose that threshold that had the least average deviation from the optimal. Of course, you can still get unlucky, but to a lesser extend. Does that make it a bit clearer? I will try to add an image example to the top post.</p>",
          "rawMarkdown": "Regarding rank prediction: Assume you have predicted probabilities for a single model, you then transform them into ranks (e.g., rankdata in numpy). Then you average the ranks when combining individual models and divide by the length so your final predictions again end up between zero and one. For then finding a fixed threshold a simple strategy is to just take the mean best threshold on multiple CV runs or similar simulations. However, there are still outliers in both directions depending on the split so if you are really unlucky your fixed threshold is \"far\" away from the optimal. So what we tried is testing various fixed thresholds and evaluate how far the resulting F1 score is compared to if you would take the optimal threshold for this fold. We then finally chose that threshold that had the least average deviation from the optimal. Of course, you can still get unlucky, but to a lesser extend. Does that make it a bit clearer? I will try to add an image example to the top post.",
          "votes": 14
        },
        {
          "id": 474494,
          "postDate": "2019-02-19T12:46:47.230Z",
          "content": "<p>That clarified it a lot, thanks! Interesting approach.</p>",
          "rawMarkdown": "That clarified it a lot, thanks! Interesting approach."
        }
      ]
    },
    {
      "id": 471618,
      "postDate": "2019-02-14T16:54:44.343Z",
      "content": "<p>Excellent work!\nCongrats! </p>",
      "rawMarkdown": "Excellent work!\nCongrats! ",
      "votes": 2
    },
    {
      "id": 471594,
      "postDate": "2019-02-14T16:19:16.807Z",
      "content": "<p>How much of an improvement  did you get by weigthing glove and para rather than just using glove?</p>",
      "rawMarkdown": "How much of an improvement  did you get by weigthing glove and para rather than just using glove?",
      "votes": 2,
      "replies": [
        {
          "id": 471624,
          "postDate": "2019-02-14T17:12:04.463Z",
          "content": "<p>I can't say for sure how much as we decided on using both quite early in the competition, but according to CV it was definitely worth it. We tried a few other things like concatenating etc. which led to similar results, but worse runtime. In the end, this is another aspect of the overfit/underfit discussion above and utilizing other things might lead to the need for doing the combination of embeddings a bit differently.</p>",
          "rawMarkdown": "I can't say for sure how much as we decided on using both quite early in the competition, but according to CV it was definitely worth it. We tried a few other things like concatenating etc. which led to similar results, but worse runtime. In the end, this is another aspect of the overfit/underfit discussion above and utilizing other things might lead to the need for doing the combination of embeddings a bit differently.",
          "votes": 2
        },
        {
          "id": 471657,
          "postDate": "2019-02-14T18:16:55.013Z",
          "content": "<p>As far as I remember averaged embeddings gave relatively small boost, but quite a stable one. I would say around 0.002 or a bit more.</p>",
          "rawMarkdown": "As far as I remember averaged embeddings gave relatively small boost, but quite a stable one. I would say around 0.002 or a bit more.",
          "votes": 2
        }
      ]
    },
    {
      "id": 471587,
      "postDate": "2019-02-14T16:13:08.157Z",
      "content": "<p>Awesome! Those are some great ideas that I can't wait to try out. Thanks a lot for sharing and congrats!!</p>",
      "rawMarkdown": "Awesome! Those are some great ideas that I can't wait to try out. Thanks a lot for sharing and congrats!!",
      "votes": 2
    },
    {
      "id": 3211178,
      "postDate": "2025-05-28T06:51:11.123Z",
      "content": "<p>cheers man<br>\ngreat job</p>",
      "rawMarkdown": "cheers man\ngreat job"
    },
    {
      "id": 2248146,
      "postDate": "2023-05-06T14:26:04.250Z",
      "content": "<p>Hi, I have a doubt. If padding is done batch-wise, the sequence length will differ for different batches. In that case, input neurons will be different for different batches. How is it possible practically? </p>",
      "rawMarkdown": "Hi, I have a doubt. If padding is done batch-wise, the sequence length will differ for different batches. In that case, input neurons will be different for different batches. How is it possible practically? "
    },
    {
      "id": 698766,
      "postDate": "2019-12-19T17:13:25.843Z",
      "content": "<p>great efforts</p>",
      "rawMarkdown": "great efforts"
    },
    {
      "id": 681462,
      "postDate": "2019-11-26T06:25:40.753Z",
      "content": "<p>Wow. This is very informative. Thanks for sharing this.</p>",
      "rawMarkdown": "Wow. This is very informative. Thanks for sharing this."
    },
    {
      "id": 530025,
      "postDate": "2019-05-11T14:16:20.423Z",
      "content": "<p>@psi Even few minutes faster counts a lot when we have to finish everything within two hours. According to what I have read online Pytorch is faster than Keras. So just wondering if you used Pytorch and if this helped you to get things done faster?</p>",
      "rawMarkdown": "@psi Even few minutes faster counts a lot when we have to finish everything within two hours. According to what I have read online Pytorch is faster than Keras. So just wondering if you used Pytorch and if this helped you to get things done faster?",
      "replies": [
        {
          "id": 534025,
          "postDate": "2019-05-20T13:22:57.090Z",
          "content": "<p>We used Keras in this competition. I shortly tested Pytorch and did not see any speed gains.</p>",
          "rawMarkdown": "We used Keras in this competition. I shortly tested Pytorch and did not see any speed gains.",
          "votes": 1
        }
      ]
    },
    {
      "id": 529935,
      "postDate": "2019-05-11T07:39:55.567Z",
      "content": "<p>Congratulations! This was a great learning experience for me.</p>",
      "rawMarkdown": "Congratulations! This was a great learning experience for me."
    },
    {
      "id": 524655,
      "postDate": "2019-04-29T09:23:36.523Z",
      "content": "<blockquote>\n  <p>We use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512.</p>\n</blockquote>\n\n<p>Hi <a href=\"/philippsinger\">@philippsinger</a>. Could you please help me understand this. I'm not quite sure that total refers to here. Is it based upon the batch size \"total\" or the total considering all the iterations needed to go through all epochs?</p>",
      "rawMarkdown": "&gt; We use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512.\n\nHi @philippsinger. Could you please help me understand this. I'm not quite sure that total refers to here. Is it based upon the batch size \"total\" or the total considering all the iterations needed to go through all epochs?\n",
      "replies": [
        {
          "id": 524660,
          "postDate": "2019-04-29T09:31:32.363Z",
          "content": "<p><a href=\"/learnmower\">@learnmower</a> For half of the epochs you increase the LR, and for half of the epochs you decrease it. So basically step_size is half of the epochs and you only do one full cycle.</p>",
          "rawMarkdown": "@learnmower For half of the epochs you increase the LR, and for half of the epochs you decrease it. So basically step_size is half of the epochs and you only do one full cycle.",
          "votes": 3
        },
        {
          "id": 524861,
          "postDate": "2019-04-29T16:57:49.620Z",
          "content": "<p>Ah. Thanks, Psi, for that explanation. +1 up'd!</p>",
          "rawMarkdown": "Ah. Thanks, Psi, for that explanation. +1 up'd!"
        }
      ]
    },
    {
      "id": 516428,
      "postDate": "2019-04-14T06:54:34.517Z",
      "content": "<p>Great work!  Thanks for your detail explanation.</p>\n\n<p>A question about padding sentence:\nYou said in <code>Runtime tricks</code>: Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. \nand <code>What did not work for us</code>: Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.</p>\n\n<p>Is that means: \n1. Don't sort the sentences by length in whole dataset, even more, we can shuffle it?\n2. On a batch , pad and truncate the sentence by 95% length.</p>\n\n<p>Congratulation!</p>",
      "rawMarkdown": "Great work!  Thanks for your detail explanation.\n\nA question about padding sentence:\nYou said in `Runtime tricks`: Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. \nand `What did not work for us`: Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.\n\nIs that means: \n1. Don't sort the sentences by length in whole dataset, even more, we can shuffle it?\n2. On a batch , pad and truncate the sentence by 95% length.\n\nCongratulation!",
      "replies": [
        {
          "id": 516532,
          "postDate": "2019-04-14T11:21:14.907Z",
          "content": "<p>Exactly and thanks :)</p>",
          "rawMarkdown": "Exactly and thanks :)"
        }
      ]
    },
    {
      "id": 510114,
      "postDate": "2019-04-08T17:01:10.163Z",
      "content": "<p>Congratulations !  And this is an incredible discussion.  Thanks for sharing.</p>",
      "rawMarkdown": "Congratulations !  And this is an incredible discussion.  Thanks for sharing."
    },
    {
      "id": 479682,
      "postDate": "2019-02-27T10:38:05.340Z",
      "content": "<p>Congratulations, tough competition! </p>",
      "rawMarkdown": "Congratulations, tough competition! "
    },
    {
      "id": 477449,
      "postDate": "2019-02-24T15:19:47.620Z",
      "content": "<p>good idea!</p>",
      "rawMarkdown": "good idea!"
    },
    {
      "id": 477282,
      "postDate": "2019-02-24T09:26:34.160Z",
      "content": "<p>Congrats!</p>",
      "rawMarkdown": "Congrats!"
    },
    {
      "id": 477099,
      "postDate": "2019-02-23T21:54:32.517Z",
      "content": "<p>Your EDA is good i have learned a lot.</p>",
      "rawMarkdown": "Your EDA is good i have learned a lot."
    },
    {
      "id": 477047,
      "postDate": "2019-02-23T18:52:38.773Z",
      "content": "<p>Congrats !!</p>",
      "rawMarkdown": "Congrats !!"
    },
    {
      "id": 476894,
      "postDate": "2019-02-23T12:11:47.800Z",
      "content": "<p>Looking forward to the kernel</p>",
      "rawMarkdown": "Looking forward to the kernel"
    },
    {
      "id": 476871,
      "postDate": "2019-02-23T11:06:37Z",
      "content": "<p>Perfect!</p>",
      "rawMarkdown": "Perfect!"
    },
    {
      "id": 476853,
      "postDate": "2019-02-23T10:17:13.590Z",
      "rawMarkdown": ""
    },
    {
      "id": 476030,
      "postDate": "2019-02-21T14:16:53.100Z",
      "content": "<p>Congrats and thk u shake. My understanding of your CV Evaluation is like this. You have 10 models and splits training set 10-folds. So each fold of training set was used to train 10 models. You will get the predictions with shape of (n_model, n_sample), and then get average of models result or the average of ranks on the predictions as your ensemble results. Is my understanding correct? BTW, how do u consider to use ranks on the predictions to find threshold to instead of probability?</p>",
      "rawMarkdown": "Congrats and thk u shake. My understanding of your CV Evaluation is like this. You have 10 models and splits training set 10-folds. So each fold of training set was used to train 10 models. You will get the predictions with shape of (n_model, n_sample), and then get average of models result or the average of ranks on the predictions as your ensemble results. Is my understanding correct? BTW, how do u consider to use ranks on the predictions to find threshold to instead of probability?"
    },
    {
      "id": 476025,
      "postDate": "2019-02-21T14:08:09.757Z",
      "content": "<p>Awesome</p>",
      "rawMarkdown": "Awesome"
    },
    {
      "id": 475983,
      "postDate": "2019-02-21T13:10:42.347Z",
      "content": "<p>Cool</p>",
      "rawMarkdown": "Cool"
    },
    {
      "id": 475716,
      "postDate": "2019-02-21T05:43:49.370Z",
      "content": "<p>Brilliant work !</p>",
      "rawMarkdown": "Brilliant work !"
    },
    {
      "id": 475407,
      "postDate": "2019-02-20T17:59:11.830Z",
      "content": "<p>Fantastic!</p>",
      "rawMarkdown": "Fantastic!"
    },
    {
      "id": 475375,
      "postDate": "2019-02-20T17:28:07.597Z",
      "content": "<p>good idea</p>",
      "rawMarkdown": "good idea"
    },
    {
      "id": 475225,
      "postDate": "2019-02-20T13:25:03.893Z",
      "content": "<p>Congrats on your win m8! An incredibly useful follow-up.</p>",
      "rawMarkdown": "Congrats on your win m8! An incredibly useful follow-up."
    },
    {
      "id": 475182,
      "postDate": "2019-02-20T11:58:52.370Z",
      "content": "<p>Congrats and thanks for your kindly share. I can learn a lot from it.</p>",
      "rawMarkdown": "Congrats and thanks for your kindly share. I can learn a lot from it."
    },
    {
      "id": 475038,
      "postDate": "2019-02-20T06:42:34.607Z",
      "content": "<p>Congrats and thank you for sharing this!</p>",
      "rawMarkdown": "Congrats and thank you for sharing this!"
    },
    {
      "id": 474952,
      "postDate": "2019-02-20T03:20:37.583Z",
      "content": "<p>Thanks for sharing, this is great for a beginner like me to learn from. </p>",
      "rawMarkdown": "Thanks for sharing, this is great for a beginner like me to learn from. "
    },
    {
      "id": 474850,
      "postDate": "2019-02-19T22:13:46.957Z",
      "content": "<p>Congratulations! Great work!!</p>",
      "rawMarkdown": "Congratulations! Great work!!"
    },
    {
      "id": 474792,
      "postDate": "2019-02-19T20:03:11.667Z",
      "content": "<p>Congratulations and thanks for sharing !</p>",
      "rawMarkdown": "Congratulations and thanks for sharing !"
    },
    {
      "id": 474471,
      "postDate": "2019-02-19T11:49:43.977Z",
      "content": "<p>Great Approach. Informative.</p>",
      "rawMarkdown": "Great Approach. Informative."
    },
    {
      "id": 473765,
      "postDate": "2019-02-18T13:32:01.103Z",
      "content": "<p>Good job!</p>",
      "rawMarkdown": "Good job!"
    },
    {
      "id": 473694,
      "postDate": "2019-02-18T11:44:53.697Z",
      "content": "<p>Congrats! That's fantastic.</p>",
      "rawMarkdown": "Congrats! That's fantastic."
    },
    {
      "id": 473648,
      "postDate": "2019-02-18T10:10:29.627Z",
      "content": "<p>congratulations </p>",
      "rawMarkdown": "congratulations "
    },
    {
      "id": 473495,
      "postDate": "2019-02-18T05:07:43.420Z",
      "content": "<p>ขอบคุณค่ะ:pn\n<a href=\"https://slot.im/\">https://slot.im/</a>สล็อต</p>",
      "rawMarkdown": "ขอบคุณค่ะ:pn\nhttps://slot.im/สล็อต"
    },
    {
      "id": 472381,
      "postDate": "2019-02-15T19:53:45.193Z",
      "content": "<p>Congrats, impressive to achieve this result with a quite simple model (no  attention or capsule layers)\nyou said that different models used for ensembling, is it a similar structure of the diagram or totally​ ​different ​?</p>",
      "rawMarkdown": "Congrats, impressive to achieve this result with a quite simple model (no  attention or capsule layers)\nyou said that different models used for ensembling, is it a similar structure of the diagram or totally​ ​different ​?",
      "replies": [
        {
          "id": 472385,
          "postDate": "2019-02-15T20:04:41.633Z",
          "content": "<p>Same model fit several times.</p>",
          "rawMarkdown": "Same model fit several times.",
          "votes": 1
        }
      ]
    },
    {
      "id": 472369,
      "postDate": "2019-02-15T19:22:24.817Z",
      "content": "<p>Congrats </p>",
      "rawMarkdown": "Congrats "
    },
    {
      "id": 472248,
      "postDate": "2019-02-15T15:12:45.160Z",
      "content": "<p>Congrats! Great job, thank you for sharing.</p>",
      "rawMarkdown": "Congrats! Great job, thank you for sharing."
    },
    {
      "id": 471962,
      "postDate": "2019-02-15T06:36:35.817Z",
      "content": "<p>Congrats! @psi\nHow did you decide the weights for embeddings?</p>",
      "rawMarkdown": "Congrats! @psi\nHow did you decide the weights for embeddings?",
      "replies": [
        {
          "id": 471988,
          "postDate": "2019-02-15T07:24:34.313Z",
          "content": "<p>As my guess, They test different weights to select best(This is my way). :)</p>",
          "rawMarkdown": "As my guess, They test different weights to select best(This is my way). :)",
          "votes": 1
        },
        {
          "id": 472000,
          "postDate": "2019-02-15T07:46:40.520Z",
          "content": "<p>Thanks! Hyperparameter tuning based on CV.</p>",
          "rawMarkdown": "Thanks! Hyperparameter tuning based on CV.",
          "votes": 1
        }
      ]
    },
    {
      "id": 471634,
      "postDate": "2019-02-14T17:34:05.200Z",
      "content": "<p>Congrats! You mention that you have 10 models  on complete training data. Are these models all similar to the one you describe in Model Structure Section?</p>",
      "rawMarkdown": "Congrats! You mention that you have 10 models  on complete training data. Are these models all similar to the one you describe in Model Structure Section?",
      "replies": [
        {
          "id": 471635,
          "postDate": "2019-02-14T17:37:41.190Z",
          "content": "<p>Yeah, 10 times the same model.</p>",
          "rawMarkdown": "Yeah, 10 times the same model.",
          "votes": 3
        }
      ]
    },
    {
      "id": 1297775,
      "postDate": "2021-05-08T09:52:43.873Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 486712,
      "postDate": "2019-03-09T09:18:51.330Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 476336,
      "postDate": "2019-02-22T01:48:27.290Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 476238,
      "postDate": "2019-02-21T21:05:55.960Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 474410,
      "postDate": "2019-02-19T10:03:49.997Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 473184,
      "postDate": "2019-02-17T13:57:46.653Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 472514,
      "postDate": "2019-02-16T04:31:19.573Z",
      "rawMarkdown": "",
      "votes": 1,
      "isDeleted": true
    },
    {
      "id": 472242,
      "postDate": "2019-02-15T15:05:49.193Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 471753,
      "postDate": "2019-02-14T20:47:05.130Z",
      "content": "<p>Congrats! and thanks for sharing your insights. \nI am impressed that you managed to train 10 models; I was stuck at 5 and that was a serious limitation.</p>",
      "rawMarkdown": "Congrats! and thanks for sharing your insights. \nI am impressed that you managed to train 10 models; I was stuck at 5 and that was a serious limitation.",
      "votes": 2,
      "isDeleted": true,
      "replies": [
        {
          "id": 471784,
          "postDate": "2019-02-14T21:55:56.400Z",
          "content": "<p>Same for me! I really have to try that batch padding trick.</p>",
          "rawMarkdown": "Same for me! I really have to try that batch padding trick.",
          "votes": 1
        }
      ]
    },
    {
      "id": 473203,
      "postDate": "2019-02-17T14:39:28.843Z",
      "content": "<p>Congrats and Thanks !!!</p>",
      "rawMarkdown": "Congrats and Thanks !!!",
      "votes": 1
    },
    {
      "id": 473199,
      "postDate": "2019-02-17T14:28:39.737Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472963,
      "postDate": "2019-02-17T02:03:48.380Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472669,
      "postDate": "2019-02-16T12:50:09.620Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472665,
      "postDate": "2019-02-16T12:43:14.360Z",
      "content": "<p>Congrats and Thanks for sharing!</p>",
      "rawMarkdown": "Congrats and Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472616,
      "postDate": "2019-02-16T10:01:06.670Z",
      "content": "<p>Congrats and Thanks for sharing!</p>",
      "rawMarkdown": "Congrats and Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472539,
      "postDate": "2019-02-16T05:58:32.957Z",
      "content": "<p>Congrats and Thanks for sharing!</p>",
      "rawMarkdown": "Congrats and Thanks for sharing!",
      "votes": 1
    },
    {
      "id": 472495,
      "postDate": "2019-02-16T03:23:22.573Z",
      "content": "<p>Thanks for sharing and congratulations!</p>",
      "rawMarkdown": "Thanks for sharing and congratulations!",
      "votes": 1
    },
    {
      "id": 471810,
      "postDate": "2019-02-14T22:52:07.163Z",
      "content": "<p>Congrats! and thanks for sharing. </p>",
      "rawMarkdown": "Congrats! and thanks for sharing. ",
      "votes": 1
    },
    {
      "id": 1201938,
      "postDate": "2021-02-15T19:01:54.350Z",
      "content": "<p>Very useful. Thanks for sharing.</p>",
      "rawMarkdown": "Very useful. Thanks for sharing."
    },
    {
      "id": 570793,
      "postDate": "2019-07-08T19:06:16.077Z",
      "content": "<p>Congrats, and Thanks for sharing :)</p>",
      "rawMarkdown": "Congrats, and Thanks for sharing :)"
    },
    {
      "id": 564470,
      "postDate": "2019-06-29T13:32:13.210Z",
      "content": "<p>GOOD!Thanks for your work,i love it.</p>",
      "rawMarkdown": "GOOD!Thanks for your work,i love it."
    },
    {
      "id": 552845,
      "postDate": "2019-06-14T16:27:35.167Z",
      "content": "<p>Thanks for the post</p>",
      "rawMarkdown": "Thanks for the post"
    },
    {
      "id": 487092,
      "postDate": "2019-03-10T05:42:03.127Z",
      "content": "<p>Thank you for sharing this! Congrats!</p>",
      "rawMarkdown": "Thank you for sharing this! Congrats!"
    },
    {
      "id": 478489,
      "postDate": "2019-02-26T08:06:49.257Z",
      "content": "<p>thanks for sharing</p>",
      "rawMarkdown": "thanks for sharing"
    },
    {
      "id": 477831,
      "postDate": "2019-02-25T10:32:09.057Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 477540,
      "postDate": "2019-02-24T20:06:06.937Z",
      "content": "<p>learn a lot. thank you!</p>",
      "rawMarkdown": "learn a lot. thank you!"
    },
    {
      "id": 477064,
      "postDate": "2019-02-23T19:28:27.337Z",
      "content": "<p>Thanks for sharing.</p>",
      "rawMarkdown": "Thanks for sharing."
    },
    {
      "id": 476883,
      "postDate": "2019-02-23T11:42:17.853Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!"
    },
    {
      "id": 476826,
      "postDate": "2019-02-23T09:15:18.917Z",
      "content": "<p>Thanks for sharing and congratulations!</p>",
      "rawMarkdown": "Thanks for sharing and congratulations!"
    },
    {
      "id": 476325,
      "postDate": "2019-02-22T01:30:01.993Z",
      "content": "<p>Congrats and Thanks for sharing!!!</p>",
      "rawMarkdown": "Congrats and Thanks for sharing!!!"
    },
    {
      "id": 476310,
      "postDate": "2019-02-22T00:28:46.983Z",
      "content": "<p>Thanks</p>",
      "rawMarkdown": "Thanks"
    },
    {
      "id": 476078,
      "postDate": "2019-02-21T15:29:00.587Z",
      "content": "<p>congrats and thanks for sharing</p>",
      "rawMarkdown": "congrats and thanks for sharing"
    },
    {
      "id": 475750,
      "postDate": "2019-02-21T06:32:17.937Z",
      "content": "<p>Congrats and thanks for sharing!</p>",
      "rawMarkdown": "Congrats and thanks for sharing!"
    },
    {
      "id": 475215,
      "postDate": "2019-02-20T13:03:24.250Z",
      "content": "<p>Thanks for sharing this!</p>",
      "rawMarkdown": "Thanks for sharing this!"
    },
    {
      "id": 475177,
      "postDate": "2019-02-20T11:50:46.710Z",
      "content": "<p>Congrats and thanks!!!</p>",
      "rawMarkdown": "Congrats and thanks!!!"
    },
    {
      "id": 475065,
      "postDate": "2019-02-20T07:36:58.293Z",
      "content": "<p>Congrats and thanks!</p>",
      "rawMarkdown": "Congrats and thanks!"
    },
    {
      "id": 474385,
      "postDate": "2019-02-19T09:33:26.977Z",
      "content": "<p>Congrats and thanks for the insights</p>",
      "rawMarkdown": "Congrats and thanks for the insights"
    },
    {
      "id": 473970,
      "postDate": "2019-02-18T19:12:37.290Z",
      "content": "<p>This taught me a lot! Thank you</p>",
      "rawMarkdown": "This taught me a lot! Thank you"
    },
    {
      "id": 473954,
      "postDate": "2019-02-18T18:39:09.623Z",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!"
    },
    {
      "id": 473337,
      "postDate": "2019-02-17T19:48:52.430Z",
      "content": "<p>Congrats! Thanks for the write-up</p>",
      "rawMarkdown": "Congrats! Thanks for the write-up"
    },
    {
      "id": 472192,
      "postDate": "2019-02-15T13:30:49.227Z",
      "content": "<p>congrats and thanks for sharing !</p>",
      "rawMarkdown": "congrats and thanks for sharing !"
    }
  ],
  "comments": [
    {
      "id": 471814,
      "author_name": "Shubin",
      "author_url": "",
      "post_date": "2019-02-14T23:13:19.793000",
      "content": "<p>congrats, can you open source your kernel so we know exactly what you did? look forward to it!</p>",
      "votes": 34,
      "replies": []
    },
    {
      "id": 1255628,
      "author_name": "Abhishek Verma",
      "author_url": "",
      "post_date": "2021-03-29T02:10:10.067000",
      "content": "<p>Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.</p>\n<p>Regarding this point <a href=\"https://www.kaggle.com/Psi\" target=\"_blank\">@Psi</a>, what can be the reason for this? I have also tried this in some other sequence model, the accuracy dropped dramatically.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 1261809,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2021-04-03T13:26:15.617000",
          "content": "<p>The model focuses only on certain sequence lengths in each batch and has no diversity.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 472122,
      "author_name": "Elior Cohen",
      "author_url": "",
      "post_date": "2019-02-15T11:48:01.430000",
      "content": "<p>Could you elaborate about\n 1. Which statistical features did you use?\n 2. On the model architecture the 1D conv follows a BiLSTM, what was done there exactly? Did you use all the states (from all the steps)? Or did you concat the last outputs of each LSTM?</p>",
      "votes": 5,
      "replies": [
        {
          "id": 472178,
          "author_name": "dott",
          "author_url": "",
          "post_date": "2019-02-15T13:01:34.987000",
          "content": "<ol>\n<li>The statistical features were: length of the text, number of capital letters, number of exclamation/question/punctuation marks, number of special symbols, number of smileys, number of words, number of unique words and few derivatives.</li>\n<li>We have a Bidirectional LSTM returning a sequence, then a convolution, followed by max pooling.</li>\n</ol>",
          "votes": 20,
          "replies": []
        },
        {
          "id": 1975464,
          "author_name": "Akshat Singhal",
          "author_url": "",
          "post_date": "2022-10-06T19:33:59.370000",
          "content": "<p>Can you kindly elaborate on how did you come up with the idea of applying convolution layer here (Why did you use it)?</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 471482,
      "author_name": "Neuron Engineer",
      "author_url": "",
      "post_date": "2019-02-14T14:19:13.843000",
      "content": "<p>Congratulation, and thank you for sharing the solution which is truly educative!!\nYou have done much work during these 3 months and totally deserve it.</p>",
      "votes": 4,
      "replies": []
    },
    {
      "id": 474345,
      "author_name": "Rohith Mohite",
      "author_url": "",
      "post_date": "2019-02-19T08:46:05.543000",
      "content": "<p>i like your EDA</p>",
      "votes": 3,
      "replies": []
    },
    {
      "id": 471629,
      "author_name": "ryches",
      "author_url": "",
      "post_date": "2019-02-14T17:28:00.223000",
      "content": "<p>It's very interesting that you guys did the padding on the fly. We tried the same thing, but our batchsizes were much larger so we ended up with almost exactly the same runtime as if we just prepadded. Did not consider reducing batch_size and only doing the top 95 percent. Very clever trick. I wonder with both of our trick combined how many models could be fit. </p>",
      "votes": 4,
      "replies": [
        {
          "id": 471633,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-02-14T17:33:42.830000",
          "content": "<p>As written, we actually tried much larger batch sizes up to 10240 allowing us to fit close to 20 models. Results looked good as well, but we saw convergence after combining around 10 models or so and would have needed to further fine-tune the higher batch size models which is why we submitted the 512 batch size models. But as you say, the higher the batch size to choose, the larger the max sequence length is, ending up in exactly what you observed. </p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 471852,
          "author_name": "aintnosunshine",
          "author_url": "",
          "post_date": "2019-02-15T01:40:05.200000",
          "content": "<p>Padding on the fly is standard practice. Those who are familiar with Tensorflow, padding happens only dynamically (not only, mostly). This is the problem when people start using keras and some default options. I have shared my kernel here <a href=\"https://www.kaggle.com/s4sarath/cudnngru-best-final\">https://www.kaggle.com/s4sarath/cudnngru-best-final</a> . With a max length of 150 words per sentences, I was able to train a model in 3600 seconds ( 5 fold with overall 23 epochs), As @psi told, I have also created 10-12 models roughly, but was no helping in improving PB. Thanks, @psi, and team for detailed information.</p>",
          "votes": 5,
          "replies": []
        },
        {
          "id": 530061,
          "author_name": "egdelwonKsIataD",
          "author_url": "",
          "post_date": "2019-05-11T15:54:35.523000",
          "content": "<p>@psi, <a href=\"/ryches\">@ryches</a> What do we mean by \"padding on fly\"? Right now, I am padding sequences as data preparation step, like \nX_train = sequence.pad_sequences(X_train,maxlen=MAX_LEN)\nX_valid = sequence.pad_sequences(X_valid,maxlen=MAX_LEN)</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 476532,
      "author_name": "Mohamed Sheik Ibrahim",
      "author_url": "",
      "post_date": "2019-02-22T09:59:06.270000",
      "content": "<p>awesome</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 474510,
      "author_name": "anil777",
      "author_url": "",
      "post_date": "2019-02-19T13:28:46.623000",
      "content": "<p>Impressive </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 473273,
      "author_name": "Gunhyuk Park",
      "author_url": "",
      "post_date": "2019-02-17T17:41:27.870000",
      "content": "<p>What an impressive way. Congratulations and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472947,
      "author_name": "Vishy",
      "author_url": "",
      "post_date": "2019-02-17T00:35:32.077000",
      "content": "<p>Congrats on winning and fantastic explanation throughout !! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472515,
      "author_name": "Shabbir",
      "author_url": "",
      "post_date": "2019-02-16T04:33:21.767000",
      "content": "<p>Congrats! and thanks for sharing your insights.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472436,
      "author_name": "averagemn",
      "author_url": "",
      "post_date": "2019-02-15T22:59:43.150000",
      "content": "<p>Many interesting points here, practical advise and great writeup. Thanks a lot for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472390,
      "author_name": "Nancy Guo",
      "author_url": "",
      "post_date": "2019-02-15T20:11:10.663000",
      "content": "<p>Congrats and thanks for sharing your solution. Your solution is really excellent!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472377,
      "author_name": "Ilya Bakalets",
      "author_url": "",
      "post_date": "2019-02-15T19:39:21.477000",
      "content": "<p>Very interesting model structure.Thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471917,
      "author_name": "pbcquoc",
      "author_url": "",
      "post_date": "2019-02-15T04:41:40.363000",
      "content": "<p>Congrast and thanks for sharing your work.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471883,
      "author_name": "yukiya",
      "author_url": "",
      "post_date": "2019-02-15T03:10:16.867000",
      "content": "<p>Thank you for the write ups, and congratulations on the strong finish! </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471875,
      "author_name": "Yi Chen Huang",
      "author_url": "",
      "post_date": "2019-02-15T02:50:27.690000",
      "content": "<p>Congrats! Really helped, and thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471834,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2019-02-15T00:53:56.420000",
      "content": "<p>Congratulations <a href=\"/philippsinger\">@philippsinger</a> and team for wining this amazing competition. Thank you for sharing your solution.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471643,
      "author_name": "Karthik Chowdary Tsaliki",
      "author_url": "",
      "post_date": "2019-02-14T17:57:48.483000",
      "content": "<p>Congratulations. @psi. Thanks for sharing your work and I am really impressed with the idea of bucketing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471619,
      "author_name": "XY",
      "author_url": "",
      "post_date": "2019-02-14T16:57:18.133000",
      "content": "<p>Congrats! Really nice work!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471617,
      "author_name": "Hamish",
      "author_url": "",
      "post_date": "2019-02-14T16:50:42.257000",
      "content": "<p>Congratulations and well done! Thank you for the explanation and what a great idea with the padding!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471549,
      "author_name": "Murat Korkmaz",
      "author_url": "",
      "post_date": "2019-02-14T15:42:50.430000",
      "content": "<p>Congrats, you really deserve it. Thanks for detailed explanation.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 474782,
      "author_name": "Anuj Garg",
      "author_url": "",
      "post_date": "2019-02-19T19:39:08.670000",
      "content": "<p>Certainly pointers can be taken from involved data</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 472794,
      "author_name": "S T MOHAMMED",
      "author_url": "",
      "post_date": "2019-02-16T17:21:04.663000",
      "content": "<p>Congrats...</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 471796,
      "author_name": "ManuelSH",
      "author_url": "",
      "post_date": "2019-02-14T22:21:58.670000",
      "content": "<p>Very nice. Congratulations!</p>\n\n<p>What do you mean by:\n\"So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data.\"</p>\n\n<p>Could you explain it with an example maybe?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 471998,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-02-15T07:40:59.497000",
          "content": "<p>Regarding rank prediction: Assume you have predicted probabilities for a single model, you then transform them into ranks (e.g., rankdata in numpy). Then you average the ranks when combining individual models and divide by the length so your final predictions again end up between zero and one. For then finding a fixed threshold a simple strategy is to just take the mean best threshold on multiple CV runs or similar simulations. However, there are still outliers in both directions depending on the split so if you are really unlucky your fixed threshold is \"far\" away from the optimal. So what we tried is testing various fixed thresholds and evaluate how far the resulting F1 score is compared to if you would take the optimal threshold for this fold. We then finally chose that threshold that had the least average deviation from the optimal. Of course, you can still get unlucky, but to a lesser extend. Does that make it a bit clearer? I will try to add an image example to the top post.</p>",
          "votes": 14,
          "replies": []
        },
        {
          "id": 474494,
          "author_name": "ManuelSH",
          "author_url": "",
          "post_date": "2019-02-19T12:46:47.230000",
          "content": "<p>That clarified it a lot, thanks! Interesting approach.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 471618,
      "author_name": "Guanshuo Xu",
      "author_url": "",
      "post_date": "2019-02-14T16:54:44.343000",
      "content": "<p>Excellent work!\nCongrats! </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 471594,
      "author_name": "João Duro",
      "author_url": "",
      "post_date": "2019-02-14T16:19:16.807000",
      "content": "<p>How much of an improvement  did you get by weigthing glove and para rather than just using glove?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 471624,
          "author_name": "Psi",
          "author_url": "",
          "post_date": "2019-02-14T17:12:04.463000",
          "content": "<p>I can't say for sure how much as we decided on using both quite early in the competition, but according to CV it was definitely worth it. We tried a few other things like concatenating etc. which led to similar results, but worse runtime. In the end, this is another aspect of the overfit/underfit discussion above and utilizing other things might lead to the need for doing the combination of embeddings a bit differently.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 471657,
          "author_name": "dott",
          "author_url": "",
          "post_date": "2019-02-14T18:16:55.013000",
          "content": "<p>As far as I remember averaged embeddings gave relatively small boost, but quite a stable one. I would say around 0.002 or a bit more.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 471587,
      "author_name": "Max Schumacher",
      "author_url": "",
      "post_date": "2019-02-14T16:13:08.157000",
      "content": "<p>Awesome! Those are some great ideas that I can't wait to try out. Thanks a lot for sharing and congrats!!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 3211178,
      "author_name": "Hoai Vo",
      "author_url": "",
      "post_date": "2025-05-28T06:51:11.123000",
      "content": "<p>cheers man<br>\ngreat job</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2248146,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-06T14:26:04.250000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 698766,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-12-19T17:13:25.843000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 681462,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-26T06:25:40.753000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 530025,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-11T14:16:20.423000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 534025,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-05-20T13:22:57.090000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 529935,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-11T07:39:55.567000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 524655,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-29T09:23:36.523000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 524660,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-29T09:31:32.363000",
          "content": "",
          "votes": 3,
          "replies": []
        },
        {
          "id": 524861,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-29T16:57:49.620000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 516428,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-14T06:54:34.517000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 516532,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-04-14T11:21:14.907000",
          "content": "",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 510114,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-04-08T17:01:10.163000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 479682,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-27T10:38:05.340000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477449,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-24T15:19:47.620000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477282,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-24T09:26:34.160000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477099,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T21:54:32.517000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477047,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T18:52:38.773000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476894,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T12:11:47.800000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476871,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T11:06:37",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476853,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T10:17:13.590000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476030,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T14:16:53.100000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476025,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T14:08:09.757000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475983,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T13:10:42.347000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475716,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T05:43:49.370000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475407,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T17:59:11.830000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475375,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T17:28:07.597000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475225,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T13:25:03.893000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475182,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T11:58:52.370000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475038,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T06:42:34.607000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474952,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T03:20:37.583000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474850,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-19T22:13:46.957000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474792,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-19T20:03:11.667000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474471,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-19T11:49:43.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473765,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T13:32:01.103000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473694,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T11:44:53.697000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473648,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T10:10:29.627000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473495,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T05:07:43.420000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 472381,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T19:53:45.193000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 472385,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-15T20:04:41.633000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 472369,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T19:22:24.817000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 472248,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T15:12:45.160000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 471962,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T06:36:35.817000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 471988,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-15T07:24:34.313000",
          "content": "",
          "votes": 1,
          "replies": []
        },
        {
          "id": 472000,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-15T07:46:40.520000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 471634,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-14T17:34:05.200000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 471635,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-14T17:37:41.190000",
          "content": "",
          "votes": 3,
          "replies": []
        }
      ]
    },
    {
      "id": 1297775,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-05-08T09:52:43.873000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 486712,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-09T09:18:51.330000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476336,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-22T01:48:27.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476238,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T21:05:55.960000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474410,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-19T10:03:49.997000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473184,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-17T13:57:46.653000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 472514,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T04:31:19.573000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472242,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T15:05:49.193000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 471753,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-14T20:47:05.130000",
      "content": "",
      "votes": 2,
      "replies": [
        {
          "id": 471784,
          "author_name": "",
          "author_url": "",
          "post_date": "2019-02-14T21:55:56.400000",
          "content": "",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 473203,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-17T14:39:28.843000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 473199,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-17T14:28:39.737000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472963,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-17T02:03:48.380000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472669,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T12:50:09.620000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472665,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T12:43:14.360000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472616,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T10:01:06.670000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472539,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T05:58:32.957000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 472495,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-16T03:23:22.573000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 471810,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-14T22:52:07.163000",
      "content": "",
      "votes": 1,
      "replies": []
    },
    {
      "id": 1201938,
      "author_name": "",
      "author_url": "",
      "post_date": "2021-02-15T19:01:54.350000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 570793,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-07-08T19:06:16.077000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 564470,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-29T13:32:13.210000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 552845,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-06-14T16:27:35.167000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 487092,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-03-10T05:42:03.127000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 478489,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-26T08:06:49.257000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477831,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-25T10:32:09.057000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477540,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-24T20:06:06.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 477064,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T19:28:27.337000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476883,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T11:42:17.853000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476826,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-23T09:15:18.917000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476325,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-22T01:30:01.993000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476310,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-22T00:28:46.983000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 476078,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T15:29:00.587000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475750,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-21T06:32:17.937000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475215,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T13:03:24.250000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475177,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T11:50:46.710000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 475065,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-20T07:36:58.293000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 474385,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-19T09:33:26.977000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473970,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T19:12:37.290000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473954,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-18T18:39:09.623000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 473337,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-17T19:48:52.430000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 472192,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-02-15T13:30:49.227000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "471476": "First of all, we want to thank Kaggle for hosting the competition and Quora for providing such a large dataset. Last 3 months were quite exhausting for us with a steep learning curve and tons of the ideas we wanted to try out. In the following we try to summarize some of the main points of our solution.\n\n**Model Structure**\nWe played around with a variety of different model structures, but in the end resorted to a quite simple one that is very similar to those posted here https://www.kaggle.com/c/quora-insincere-questions-classification/discussion/79824. It’s basically a Single Bi-LSTM 128 followed by a Conv1D with kernel size 1 only and GlobalMaxPooling afterwards plus additional dropout layers with minimal dropout. We additionally use a few statistical features.\n\n![enter image description here][1]\n\n**Embeddings**\nFirst, we use all tokens from both train and test data for our vocabulary. We do the simple pre-cleaning that was posted in a kernel at the start of the competition and split by space afterwards (spacy and nltk resulted in similar performance). We do not lowercase, but keep uppercase, and do not limit the vocab at all. For embeddings we use glove and para where we weight glove a bit higher. The most important thing now is to find as many embeddings as possible for our vocabulary. We had a few steps to achieve this, like checking singular and plural of the word, checking lowercase embeddings, removing special tokens, etc. For public test data we had around 50k of vocab tokens we did not find in the embeddings afterwards. Even though we tried a few different strategies for handling the OOV tokens, we resorted to a single OOV token with a single random embedding vector. \n\n**Threshold**\nWe spent a lot of time trying to figure out good strategies for choosing a good threshold for classification. Over time, we saw that estimating the threshold on validation data and then applying it on test data does not really work. There is a large variation on optimal thresholds. So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data. The visualization below shows that in action (not necessarily our final eval). On the x-axis we plot the different fixed thresholds and on the y-axis we see the deviation from the optimal F1 score across folds using this fixed threshold (see CV chapter below). The blue line is the mean, green is median, purple is minimum, red is maximum, and bars are stds. So for example here, if we choose a threshold in the range of 0.927 we expect the F1 score to be not much worse (around 0.001) compared to choosing the optimal threshold (which we can't do for test data). In practise, this might of course deviate further and we could also see larger deviations on PLB. For further elaboration, please check the comments.\n\n![enter image description here][2]\n\n**Runtime tricks**\nWe aimed at combining as many models as possible. To do this, we needed to improve runtime and the most important thing to achieve this was the following. We do not pad sequences to the same length based on the whole data, but just on a batch level. That means we conduct padding and truncation on the data generator level for each batch separately, so that length of the sentences in a batch can vary in size. Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. This improved runtime heavily and kept accuracy quite robust on single model level, and improved it by being able to average more models.\n\n**Fitting**\nWe use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512. We could achieve similar results by even taking 10 or 20 times higher batch sizes, which goes hand in hand with recent research on fast convergence. With these larger batch sizes we could even fit close to 20 models, but results stabilized close to 10 models which is why we chose to go with the smaller batch size in the end. However, there might still be some room left here if one properly tunes this.\n\n**Multiple models**\nIn the end, we managed to fit more than 10 models on the complete training dataset with help of the runtime tricks mentioned before. Our best final private score even had only a runtime of 6000 seconds (I think they used a bit better hardware for running), so there would be space for 1-2 more models. With larger batch sizes even much more might be feasible. As mentioned, we then average the rank predictions of each model and use our specified threshold for prediction.\n\n**Embrace the randomness**\nAs it was necessary to utilize CUDNN Layers in this competition, there was some randomness involved that could be quite frustrating from time to time. I saw many people trying to fix seeds etc. and some claiming they could completely remove the randomness by using Pytorch (I still don’t believe this BTW as CUDNN has atomic operations). However, as mentioned before, a well working strategy in this competition was to combine multiple models and to end up with a good ensemble, those models should be a bit different to each other. So having different random initializations etc. can be helpful. Seeing people setting the seed as a hyperparameter is weird.\n\n**CV Evaluation**\nWhat I saw many people doing wrongly in this competition, and we also only figured this out after a while, is to trust their single out-of-fold evaluation. However, in this competition, it is crucial to combine (average) multiple models (in our case the same model). That means that our CV evaluation looks like the following. We do a k-fold split (mostly 10-fold) and fit the same model up to v-times on the same training split and then successively evaluate it on the single out of fold. So for the first split, we first fit one model and evaluate it, then a second one and evaluate the average and so forth. We repeat this for all 10 folds, landing us with e.g., 100 model fits overall, and then we can take a look at the median or mean over all folds for v-model-ensembles. The reason for doing this is that f1 scores are very different on the split you have. For one 10% split you might end up with a maximum of 0.72 and for the other you might end up at 0.705 or similar. So repeating the split 10 times, fitting the same model v-times for each split, and then looking at the grand picture gave us the best overall evaluation. This routine helped us to compare individual solutions with each other. BTW our final scores are exactly what we would expect from our CV evaluation, but again this might be lucky :)\n\n**Robustness and over/underfitting**\nAround 2 weeks before final submission, our results became so stable that changing things did not alter results much. Things like finding more OOV embedding vectors resulted in same results, using slightly different layers ended up being similar, and other things. This was a bit frustrating, but in the end things worked out. In the end, it was important to find a good balance between over and underfitting (as always). Underfitting too much led to good single model performances, but was worse for combining models, and the other way around. For example, if your model overfits, there can be many different solutions to tackle this, e.g., add dropouts, or reduce the vocab size, or reduce model complexity, etc. So if someone says on kaggle that one things works for him/her, that does not necessarily mean that it will work for you as you might already be doing something similar that has similar effects (a good example is the Gaussian noise discussion).\n\n**What did not work for us**\nMostly you only read what worked, but here is an incomplete shortlist of what did not work for us. This does not mean that it doesn’t work at all, but rather that it was worse for our specific solution.\n\n- Different optimizers (focal loss was similar though)\n-  Label smoothing\n-  Auxiliary learning / multitask learning\n-  Snapshot learning\n-  Pseudo labeling\n-  Fitting own embeddings with gensim\n-  Spelling correction\n-  Taking median/percentile of predicitons instead of average\n-  More complex layers and architectures (Attention, QRNN, Capsule, larger/multiple LSTM layers, larger CNN kernel sizes, LGBM or bag of words)\n-  Word collocations - Several words put together can bear a completely new meaning, which is not captured by embeddings. Glove turned out to have quite a lot of such collocations with words put together using \"-\" sign. So we replaced examples like \"ethnical cleansing\" with \"ethnical-cleansing\", which is then captured by a more appropriate glove embedding. It showed no improvement on CV.\n-  Extra statistical features - Presence of statistical features added a little bit to the accuracy based on CV, but we saw no improvement with other extra features, like sentiment or bag-of-words based variables.\n-  Replacement of words with synonyms - An idea of replacing all nationalities (or e.g. political party) with the same word did not work at all.\n-  Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.\n\n\n  [1]: https://i.imgur.com/zUY9tVN.png\n  [2]: https://i.imgur.com/2NPwBIR.png",
    "471814": "congrats, can you open source your kernel so we know exactly what you did? look forward to it!",
    "1255628": "Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.\n\nRegarding this point @Psi, what can be the reason for this? I have also tried this in some other sequence model, the accuracy dropped dramatically.",
    "472122": "Could you elaborate about\n 1. Which statistical features did you use?\n 2. On the model architecture the 1D conv follows a BiLSTM, what was done there exactly? Did you use all the states (from all the steps)? Or did you concat the last outputs of each LSTM?",
    "471482": "Congratulation, and thank you for sharing the solution which is truly educative!!\nYou have done much work during these 3 months and totally deserve it.\n",
    "474345": "i like your EDA",
    "471629": "It's very interesting that you guys did the padding on the fly. We tried the same thing, but our batchsizes were much larger so we ended up with almost exactly the same runtime as if we just prepadded. Did not consider reducing batch_size and only doing the top 95 percent. Very clever trick. I wonder with both of our trick combined how many models could be fit. ",
    "476532": "awesome",
    "474510": "Impressive ",
    "473273": "What an impressive way. Congratulations and thanks for sharing!",
    "472947": "Congrats on winning and fantastic explanation throughout !! ",
    "472515": "Congrats! and thanks for sharing your insights.",
    "472436": "Many interesting points here, practical advise and great writeup. Thanks a lot for sharing!",
    "472390": "Congrats and thanks for sharing your solution. Your solution is really excellent!",
    "472377": "Very interesting model structure.Thanks for sharing!",
    "471917": "Congrast and thanks for sharing your work.",
    "471883": "Thank you for the write ups, and congratulations on the strong finish! ",
    "471875": "Congrats! Really helped, and thanks for sharing.",
    "471834": "Congratulations @philippsinger and team for wining this amazing competition. Thank you for sharing your solution.",
    "471643": "Congratulations. @psi. Thanks for sharing your work and I am really impressed with the idea of bucketing.",
    "471619": "Congrats! Really nice work!",
    "471617": "Congratulations and well done! Thank you for the explanation and what a great idea with the padding!",
    "471549": "Congrats, you really deserve it. Thanks for detailed explanation.",
    "474782": "Certainly pointers can be taken from involved data",
    "472794": "Congrats...",
    "471796": "Very nice. Congratulations!\n\nWhat do you mean by:\n\"So what we did instead is to try to find a fixed threshold on CV that produces the least deviation for the f1 score from the optimal threshold. We saw that we can get more stable results when we produce ranks on the predicted probability and average the ranks instead of averaging probabilities. For final submission we then chose the best CV threshold. This also allowed us to fit the model on the complete data without the need to rely on a random split and less training data.\"\n\nCould you explain it with an example maybe?",
    "471618": "Excellent work!\nCongrats! ",
    "471594": "How much of an improvement  did you get by weigthing glove and para rather than just using glove?",
    "471587": "Awesome! Those are some great ideas that I can't wait to try out. Thanks a lot for sharing and congrats!!",
    "3211178": "cheers man\ngreat job",
    "2248146": "Hi, I have a doubt. If padding is done batch-wise, the sequence length will differ for different batches. In that case, input neurons will be different for different batches. How is it possible practically? ",
    "698766": "great efforts",
    "681462": "Wow. This is very informative. Thanks for sharing this.",
    "530025": "@psi Even few minutes faster counts a lot when we have to finish everything within two hours. According to what I have read online Pytorch is faster than Keras. So just wondering if you used Pytorch and if this helped you to get things done faster?",
    "529935": "Congratulations! This was a great learning experience for me.",
    "524655": "&gt; We use a one cycle policy with Nadam optimizer (you can do this with the typical CyclicLearningRate implementations by just changing the step size to half your total iterations). We chose a batch size of 512.\n\nHi @philippsinger. Could you please help me understand this. I'm not quite sure that total refers to here. Is it based upon the batch size \"total\" or the total considering all the iterations needed to go through all epochs?\n",
    "516428": "Great work!  Thanks for your detail explanation.\n\nA question about padding sentence:\nYou said in `Runtime tricks`: Additionally, we further improved this by not truncating based on the length of the longest sequence in the batch, but based on the 95% percentile of lengths within the sequence. \nand `What did not work for us`: Order the train data by the length of the sentences - This approach gave a dramatic improvement in the fitting time because each batch contained only sentences with similar sizes, but it hurt the accuracy of the model too much.\n\nIs that means: \n1. Don't sort the sentences by length in whole dataset, even more, we can shuffle it?\n2. On a batch , pad and truncate the sentence by 95% length.\n\nCongratulation!",
    "510114": "Congratulations !  And this is an incredible discussion.  Thanks for sharing.",
    "479682": "Congratulations, tough competition! ",
    "477449": "good idea!",
    "477282": "Congrats!",
    "477099": "Your EDA is good i have learned a lot.",
    "477047": "Congrats !!",
    "476894": "Looking forward to the kernel",
    "476871": "Perfect!",
    "476853": "",
    "476030": "Congrats and thk u shake. My understanding of your CV Evaluation is like this. You have 10 models and splits training set 10-folds. So each fold of training set was used to train 10 models. You will get the predictions with shape of (n_model, n_sample), and then get average of models result or the average of ranks on the predictions as your ensemble results. Is my understanding correct? BTW, how do u consider to use ranks on the predictions to find threshold to instead of probability?",
    "476025": "Awesome",
    "475983": "Cool",
    "475716": "Brilliant work !",
    "475407": "Fantastic!",
    "475375": "good idea",
    "475225": "Congrats on your win m8! An incredibly useful follow-up.",
    "475182": "Congrats and thanks for your kindly share. I can learn a lot from it.",
    "475038": "Congrats and thank you for sharing this!",
    "474952": "Thanks for sharing, this is great for a beginner like me to learn from. ",
    "474850": "Congratulations! Great work!!",
    "474792": "Congratulations and thanks for sharing !",
    "474471": "Great Approach. Informative.",
    "473765": "Good job!",
    "473694": "Congrats! That's fantastic.",
    "473648": "congratulations ",
    "473495": "ขอบคุณค่ะ:pn\nhttps://slot.im/สล็อต",
    "472381": "Congrats, impressive to achieve this result with a quite simple model (no  attention or capsule layers)\nyou said that different models used for ensembling, is it a similar structure of the diagram or totally​ ​different ​?",
    "472369": "Congrats ",
    "472248": "Congrats! Great job, thank you for sharing.",
    "471962": "Congrats! @psi\nHow did you decide the weights for embeddings?",
    "471634": "Congrats! You mention that you have 10 models  on complete training data. Are these models all similar to the one you describe in Model Structure Section?",
    "1297775": "",
    "486712": "",
    "476336": "",
    "476238": "",
    "474410": "",
    "473184": "",
    "472514": "",
    "472242": "",
    "471753": "Congrats! and thanks for sharing your insights. \nI am impressed that you managed to train 10 models; I was stuck at 5 and that was a serious limitation.",
    "473203": "Congrats and Thanks !!!",
    "473199": "Congrats and thanks for sharing!",
    "472963": "Congrats and thanks for sharing!",
    "472669": "Congrats and thanks for sharing!",
    "472665": "Congrats and Thanks for sharing!",
    "472616": "Congrats and Thanks for sharing!",
    "472539": "Congrats and Thanks for sharing!",
    "472495": "Thanks for sharing and congratulations!",
    "471810": "Congrats! and thanks for sharing. ",
    "1201938": "Very useful. Thanks for sharing.",
    "570793": "Congrats, and Thanks for sharing :)",
    "564470": "GOOD!Thanks for your work,i love it.",
    "552845": "Thanks for the post",
    "487092": "Thank you for sharing this! Congrats!",
    "478489": "thanks for sharing",
    "477831": "Thanks for sharing!",
    "477540": "learn a lot. thank you!",
    "477064": "Thanks for sharing.",
    "476883": "Congrats and thanks for sharing!",
    "476826": "Thanks for sharing and congratulations!",
    "476325": "Congrats and Thanks for sharing!!!",
    "476310": "Thanks",
    "476078": "congrats and thanks for sharing",
    "475750": "Congrats and thanks for sharing!",
    "475215": "Thanks for sharing this!",
    "475177": "Congrats and thanks!!!",
    "475065": "Congrats and thanks!",
    "474385": "Congrats and thanks for the insights",
    "473970": "This taught me a lot! Thank you",
    "473954": "Thanks for sharing!",
    "473337": "Congrats! Thanks for the write-up",
    "472192": "congrats and thanks for sharing !"
  }
}