{
  "id": 81137,
  "title": "2nd place solution",
  "url": "/competitions/quora-insincere-questions-classification/writeups/takapt-2nd-place-solution",
  "author_name": "",
  "post_date": "2019-02-19T11:45:42.600030100Z",
  "votes": 80,
  "comment_count": 11,
  "views": 0,
  "content": "<h2>Summary</h2>\n\n<p>I used a single NN with 1-layer Bi-GRU and dense layers for statistical features. I did seed averaging for ensemble. I used PyTorch to write NN.</p>\n\n<p>Key factors of my solution are</p>\n\n<ul>\n<li>Tune hyperparameters based on solid CV </li>\n<li>Train word embeddings on the competition dataset. I guess most participants didn't?</li>\n<li>Faster training techniques to train more models. It's adaptive lengths of sequences to input RNN for each batch</li>\n</ul>\n\n<h2>Preprocessing</h2>\n\n<p>I inserted spaces around characters except alphabets and numbers.\nThen, I used keras tokenizer, which splits by only space.</p>\n\n<p>After tokenization, I applied spell correction to OOV words. The rough idea of the spell correction algorithm is to find words with 0 or 1 levenshtein distance while ignoring cases. Precisely, there are a few heuristics.</p>\n\n<p>In my case, devising preprocessing including the above spell correction did not change CV score so much.</p>\n\n<h2>Model architecture</h2>\n\n<p>The main part is 1-layer Bi-GRU with hidden size 128 followed by the concatenation of max pooling, average pooling and first/last positional outputs. Another part is dense layers for statistical features. The outputs of 2 network parts are concatenated, then fed to dense layers.</p>\n\n<pre>QuoraModel(\n  (embedding): Embedding(222910, 668, padding_idx=0)\n  (text): RNNBlock(\n    (rnn): GRU(668, 128, batch_first=True, bidirectional=True)\n  )\n  (features_dense): Sequential(\n    (0): Linear(in_features=92, out_features=32, bias=True)\n    (1): ReLU(inplace)\n    (2): Linear(in_features=32, out_features=16, bias=True)\n    (3): ReLU(inplace)\n  )\n  (dense): Sequential(\n    (0): BatchNorm1d(1040, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (1): ReLU(inplace)\n    (2): Dropout(p=0.25)\n    (3): Linear(in_features=1040, out_features=64, bias=True)\n    (4): BatchNorm1d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (5): ReLU(inplace)\n    (6): Dropout(p=0.1)\n    (7): Linear(in_features=64, out_features=1, bias=True)\n  )\n)\n</pre>\n\n<h2>Embedding</h2>\n\n<p>I used glove and wiki-news pretrained word embeddings. I also trained 64 dimensional word embeddings on the competition dataset (train+test) with fastText.\nIn addition to them, I added 4 binary features to word embeddings (all upper chars?, first char upper?, only first char upper?, OOV word?).\nFinally, I concatenated all of them. The parameters of word embeddings are freezed during training.</p>\n\n<h2>Statistical features</h2>\n\n<ul>\n<li>the number of words</li>\n<li>the number of unique words</li>\n<li>the number of characters</li>\n<li>the number of upper characters</li>\n<li>Bag of characters: Implemented by <code>CountVectorizer(ngram_range=(1, 1), min_df=1e-4, token_pattern=r'\\w+',analyzer='char')</code></li>\n</ul>\n\n<h2>Length of sequences to input RNN</h2>\n\n<p>For the faster training, I adjusted the lengths of sequences for each batch.\nWhen training, I used the maximum length of sequences in the batch or 55 length by applying pre-truncation if the maximum length over 55. When predicting for test, the truncation is applied if the length is over 70 instead of 55.</p>\n\n<p>Thanks to this trick, I was able to train 6 models on kernel compared 5 models without this trick.</p>\n\n<h2>Training</h2>\n\n<p>I used Adam with learning rate 0.001. The learning rate is multiplied by 0.8 after each epoch.</p>\n\n<p>I got the best CV score with batch size 256. But, batch size has the trade-off between score and training time. As I increase batch size, CV score gets worse and training gets faster. I chose batch size 320 by checking CV score and training time on kernel.</p>\n\n<h2>Ensemble</h2>\n\n<p>I did seed averaging of 6 models. I trained 6 models with different seeds. 5 epochs are spent for each model. Each model is trained on the full train dataset, in other words, I didn't use k-fold split to train different models.</p>\n\n<p>I averaged the predictions of 6 models. Then, I made the final binary predictions with threshold 0.36.</p>\n\n<h2>Local validation</h2>\n\n<p>I did 5 fold CV for the local validation. For each fold, I used predictions after ensemble rather than predictions by 1 model for more stable CV, closer CV score to LB score and more optimal hyperparameter search when ensemble.</p>\n\n<h2>CV score</h2>\n\n<p>I show CV scores of my model used for private LB and several models without some feature.\n0.70974 is the CV score for the model used for private LB.</p>\n\n<pre>|Removed feature                   |score  |\n|----------------------------------|-------|\n|no removal                        |0.70974|\n|4 binary embedding feature        |0.70957|\n|spell correction                  |0.70953|\n|statistical features              |0.70877|\n|word embeddings trained on dataset|0.70794|\n</pre>",
  "messages": [
    {
      "id": "474466",
      "postDate": "02/19/2019 11:45:42",
      "content": "<h2>Summary</h2>\n\n<p>I used a single NN with 1-layer Bi-GRU and dense layers for statistical features. I did seed averaging for ensemble. I used PyTorch to write NN.</p>\n\n<p>Key factors of my solution are</p>\n\n<ul>\n<li>Tune hyperparameters based on solid CV </li>\n<li>Train word embeddings on the competition dataset. I guess most participants didn't?</li>\n<li>Faster training techniques to train more models. It's adaptive lengths of sequences to input RNN for each batch</li>\n</ul>\n\n<h2>Preprocessing</h2>\n\n<p>I inserted spaces around characters except alphabets and numbers.\nThen, I used keras tokenizer, which splits by only space.</p>\n\n<p>After tokenization, I applied spell correction to OOV words. The rough idea of the spell correction algorithm is to find words with 0 or 1 levenshtein distance while ignoring cases. Precisely, there are a few heuristics.</p>\n\n<p>In my case, devising preprocessing including the above spell correction did not change CV score so much.</p>\n\n<h2>Model architecture</h2>\n\n<p>The main part is 1-layer Bi-GRU with hidden size 128 followed by the concatenation of max pooling, average pooling and first/last positional outputs. Another part is dense layers for statistical features. The outputs of 2 network parts are concatenated, then fed to dense layers.</p>\n\n<pre>QuoraModel(\n  (embedding): Embedding(222910, 668, padding_idx=0)\n  (text): RNNBlock(\n    (rnn): GRU(668, 128, batch_first=True, bidirectional=True)\n  )\n  (features_dense): Sequential(\n    (0): Linear(in_features=92, out_features=32, bias=True)\n    (1): ReLU(inplace)\n    (2): Linear(in_features=32, out_features=16, bias=True)\n    (3): ReLU(inplace)\n  )\n  (dense): Sequential(\n    (0): BatchNorm1d(1040, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (1): ReLU(inplace)\n    (2): Dropout(p=0.25)\n    (3): Linear(in_features=1040, out_features=64, bias=True)\n    (4): BatchNorm1d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (5): ReLU(inplace)\n    (6): Dropout(p=0.1)\n    (7): Linear(in_features=64, out_features=1, bias=True)\n  )\n)\n</pre>\n\n<h2>Embedding</h2>\n\n<p>I used glove and wiki-news pretrained word embeddings. I also trained 64 dimensional word embeddings on the competition dataset (train+test) with fastText.\nIn addition to them, I added 4 binary features to word embeddings (all upper chars?, first char upper?, only first char upper?, OOV word?).\nFinally, I concatenated all of them. The parameters of word embeddings are freezed during training.</p>\n\n<h2>Statistical features</h2>\n\n<ul>\n<li>the number of words</li>\n<li>the number of unique words</li>\n<li>the number of characters</li>\n<li>the number of upper characters</li>\n<li>Bag of characters: Implemented by <code>CountVectorizer(ngram_range=(1, 1), min_df=1e-4, token_pattern=r'\\w+',analyzer='char')</code></li>\n</ul>\n\n<h2>Length of sequences to input RNN</h2>\n\n<p>For the faster training, I adjusted the lengths of sequences for each batch.\nWhen training, I used the maximum length of sequences in the batch or 55 length by applying pre-truncation if the maximum length over 55. When predicting for test, the truncation is applied if the length is over 70 instead of 55.</p>\n\n<p>Thanks to this trick, I was able to train 6 models on kernel compared 5 models without this trick.</p>\n\n<h2>Training</h2>\n\n<p>I used Adam with learning rate 0.001. The learning rate is multiplied by 0.8 after each epoch.</p>\n\n<p>I got the best CV score with batch size 256. But, batch size has the trade-off between score and training time. As I increase batch size, CV score gets worse and training gets faster. I chose batch size 320 by checking CV score and training time on kernel.</p>\n\n<h2>Ensemble</h2>\n\n<p>I did seed averaging of 6 models. I trained 6 models with different seeds. 5 epochs are spent for each model. Each model is trained on the full train dataset, in other words, I didn't use k-fold split to train different models.</p>\n\n<p>I averaged the predictions of 6 models. Then, I made the final binary predictions with threshold 0.36.</p>\n\n<h2>Local validation</h2>\n\n<p>I did 5 fold CV for the local validation. For each fold, I used predictions after ensemble rather than predictions by 1 model for more stable CV, closer CV score to LB score and more optimal hyperparameter search when ensemble.</p>\n\n<h2>CV score</h2>\n\n<p>I show CV scores of my model used for private LB and several models without some feature.\n0.70974 is the CV score for the model used for private LB.</p>\n\n<pre>|Removed feature                   |score  |\n|----------------------------------|-------|\n|no removal                        |0.70974|\n|4 binary embedding feature        |0.70957|\n|spell correction                  |0.70953|\n|statistical features              |0.70877|\n|word embeddings trained on dataset|0.70794|\n</pre>",
      "rawMarkdown": "## Summary\nI used a single NN with 1-layer Bi-GRU and dense layers for statistical features. I did seed averaging for ensemble. I used PyTorch to write NN.\n\nKey factors of my solution are\n\n* Tune hyperparameters based on solid CV \n* Train word embeddings on the competition dataset. I guess most participants didn't?\n* Faster training techniques to train more models. It's adaptive lengths of sequences to input RNN for each batch\n\n\n## Preprocessing\nI inserted spaces around characters except alphabets and numbers.\nThen, I used keras tokenizer, which splits by only space.\n\nAfter tokenization, I applied spell correction to OOV words. The rough idea of the spell correction algorithm is to find words with 0 or 1 levenshtein distance while ignoring cases. Precisely, there are a few heuristics.\n\nIn my case, devising preprocessing including the above spell correction did not change CV score so much.\n\n## Model architecture\nThe main part is 1-layer Bi-GRU with hidden size 128 followed by the concatenation of max pooling, average pooling and first/last positional outputs. Another part is dense layers for statistical features. The outputs of 2 network parts are concatenated, then fed to dense layers.\n\n<pre>QuoraModel(\n  (embedding): Embedding(222910, 668, padding_idx=0)\n  (text): RNNBlock(\n    (rnn): GRU(668, 128, batch_first=True, bidirectional=True)\n  )\n  (features_dense): Sequential(\n    (0): Linear(in_features=92, out_features=32, bias=True)\n    (1): ReLU(inplace)\n    (2): Linear(in_features=32, out_features=16, bias=True)\n    (3): ReLU(inplace)\n  )\n  (dense): Sequential(\n    (0): BatchNorm1d(1040, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (1): ReLU(inplace)\n    (2): Dropout(p=0.25)\n    (3): Linear(in_features=1040, out_features=64, bias=True)\n    (4): BatchNorm1d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (5): ReLU(inplace)\n    (6): Dropout(p=0.1)\n    (7): Linear(in_features=64, out_features=1, bias=True)\n  )\n)\n</pre>\n\n## Embedding\nI used glove and wiki-news pretrained word embeddings. I also trained 64 dimensional word embeddings on the competition dataset (train+test) with fastText.\nIn addition to them, I added 4 binary features to word embeddings (all upper chars?, first char upper?, only first char upper?, OOV word?).\nFinally, I concatenated all of them. The parameters of word embeddings are freezed during training.\n\n## Statistical features\n* the number of words\n* the number of unique words\n* the number of characters\n* the number of upper characters\n* Bag of characters: Implemented by `CountVectorizer(ngram_range=(1, 1), min_df=1e-4, token_pattern=r'\\w+',analyzer='char')`\n\n## Length of sequences to input RNN\nFor the faster training, I adjusted the lengths of sequences for each batch.\nWhen training, I used the maximum length of sequences in the batch or 55 length by applying pre-truncation if the maximum length over 55. When predicting for test, the truncation is applied if the length is over 70 instead of 55.\n\nThanks to this trick, I was able to train 6 models on kernel compared 5 models without this trick.\n\n\n## Training\nI used Adam with learning rate 0.001. The learning rate is multiplied by 0.8 after each epoch.\n\nI got the best CV score with batch size 256. But, batch size has the trade-off between score and training time. As I increase batch size, CV score gets worse and training gets faster. I chose batch size 320 by checking CV score and training time on kernel.\n\n## Ensemble\nI did seed averaging of 6 models. I trained 6 models with different seeds. 5 epochs are spent for each model. Each model is trained on the full train dataset, in other words, I didn't use k-fold split to train different models.\n\nI averaged the predictions of 6 models. Then, I made the final binary predictions with threshold 0.36.\n\n## Local validation\nI did 5 fold CV for the local validation. For each fold, I used predictions after ensemble rather than predictions by 1 model for more stable CV, closer CV score to LB score and more optimal hyperparameter search when ensemble.\n\n## CV score\nI show CV scores of my model used for private LB and several models without some feature.\n0.70974 is the CV score for the model used for private LB.\n\n<pre>|Removed feature                   |score  |\n|----------------------------------|-------|\n|no removal                        |0.70974|\n|4 binary embedding feature        |0.70957|\n|spell correction                  |0.70953|\n|statistical features              |0.70877|\n|word embeddings trained on dataset|0.70794|\n</pre>",
      "votes": null
    },
    {
      "id": "474881",
      "postDate": "02/20/2019 00:03:35",
      "content": "<p>Congrats! Both top 2 teams used adaptive sequence length to speed things up. </p>",
      "rawMarkdown": "Congrats! Both top 2 teams used adaptive sequence length to speed things up.",
      "votes": null
    },
    {
      "id": "474914",
      "postDate": "02/20/2019 01:49:12",
      "content": "<p>What a great result, and congratulations on your first prize :) <br>\nCould you elaborate on how you chose the threshold=0.36?</p>",
      "rawMarkdown": "What a great result, and congratulations on your first prize :)  \nCould you elaborate on how you chose the threshold=0.36?",
      "votes": null
    },
    {
      "id": "475005",
      "postDate": "02/20/2019 05:45:09",
      "content": "<p>Thank you for sharing!</p>",
      "rawMarkdown": "Thank you for sharing!",
      "votes": null
    },
    {
      "id": "475234",
      "postDate": "02/20/2019 13:31:04",
      "content": "<p>Thanks!</p>\n\n<p>About threshold selection, there is no fancy thing. It was chosen based on local CV and hard-coded. After obtaining ensembled predictions, I checked CV scores with different thresholds and chose roughly a good threshold.</p>\n\n<p>The CV scores with different thresholds of the model used for private LB are</p>\n\n<pre>|threshold|CV score|\n|---------|--------|\n|0.350    |0.70961 |\n|0.355    |0.70969 |\n|0.360    |0.70974 |\n|0.365    |0.70976 |\n|0.370    |0.70969 |\n|0.375    |0.70959 |\n|0.380    |0.70946 |\n</pre>\n\n<p>The best threshold is 0.365 for the predictions in this case. But, I didn't care about the difference of CV scores with threshold 0.36 and 0.365 because it's small and the best threshold varies around 0.36~0.365 when training models with the completely same hyperparameter setting.</p>",
      "rawMarkdown": "Thanks!\n\nAbout threshold selection, there is no fancy thing. It was chosen based on local CV and hard-coded. After obtaining ensembled predictions, I checked CV scores with different thresholds and chose roughly a good threshold.\n\nThe CV scores with different thresholds of the model used for private LB are\n<pre>|threshold|CV score|\n|---------|--------|\n|0.350    |0.70961 |\n|0.355    |0.70969 |\n|0.360    |0.70974 |\n|0.365    |0.70976 |\n|0.370    |0.70969 |\n|0.375    |0.70959 |\n|0.380    |0.70946 |\n</pre>\n\nThe best threshold is 0.365 for the predictions in this case. But, I didn't care about the difference of CV scores with threshold 0.36 and 0.365 because it's small and the best threshold varies around 0.36~0.365 when training models with the completely same hyperparameter setting.",
      "votes": null
    },
    {
      "id": "475628",
      "postDate": "02/21/2019 02:02:19",
      "content": "<p>Thanks for sharing,i wonder how to use pretrained embedding when you already train the embedding on train-test dataset</p>",
      "rawMarkdown": "Thanks for sharing,i wonder how to use pretrained embedding when you already train the embedding on train-test dataset",
      "votes": null
    },
    {
      "id": "476121",
      "postDate": "02/21/2019 16:21:52",
      "content": "<p>Congrats! You mean just train the 64 dimensional word embeddings with your model architecture?</p>",
      "rawMarkdown": "Congrats! You mean just train the 64 dimensional word embeddings with your model architecture?",
      "votes": null
    },
    {
      "id": "477215",
      "postDate": "02/24/2019 06:13:48",
      "content": "<p>Congratulations!!</p>",
      "rawMarkdown": "Congratulations!!",
      "votes": null
    },
    {
      "id": "487170",
      "postDate": "03/10/2019 09:15:59",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "524351",
      "postDate": "04/28/2019 15:03:36",
      "content": "<p>Can you explain about <code>1040</code> feature input?</p>\n\n<p>Thanks so much!</p>",
      "rawMarkdown": "Can you explain about `1040` feature input?\n\nThanks so much!",
      "votes": null
    },
    {
      "id": "533033",
      "postDate": "05/18/2019 10:00:22",
      "content": "<p>Congratulations. Thank you for sharing</p>",
      "rawMarkdown": "Congratulations. Thank you for sharing",
      "votes": null
    },
    {
      "id": "552221",
      "postDate": "06/13/2019 16:44:28",
      "content": "<p><a href=\"/takapt\">@takapt</a> how did you train the embeddings, do you remember? Especially what loss value did you have after the training? (or how many epochs did you train)</p>",
      "rawMarkdown": "takapt how did you train the embeddings, do you remember? Especially what loss value did you have after the training? (or how many epochs did you train)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 474881,
      "author_name": "jihangz",
      "author_url": "",
      "post_date": "02/20/2019 00:03:35",
      "content": "<p>Congrats! Both top 2 teams used adaptive sequence length to speed things up. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 474914,
      "author_name": "pocketsuteado",
      "author_url": "",
      "post_date": "02/20/2019 01:49:12",
      "content": "<p>What a great result, and congratulations on your first prize :) <br>\nCould you elaborate on how you chose the threshold=0.36?</p>",
      "votes": null,
      "replies": [
        {
          "id": 475234,
          "author_name": "takapt",
          "author_url": "",
          "post_date": "02/20/2019 13:31:04",
          "content": "<p>Thanks!</p>\n\n<p>About threshold selection, there is no fancy thing. It was chosen based on local CV and hard-coded. After obtaining ensembled predictions, I checked CV scores with different thresholds and chose roughly a good threshold.</p>\n\n<p>The CV scores with different thresholds of the model used for private LB are</p>\n\n<pre>|threshold|CV score|\n|---------|--------|\n|0.350    |0.70961 |\n|0.355    |0.70969 |\n|0.360    |0.70974 |\n|0.365    |0.70976 |\n|0.370    |0.70969 |\n|0.375    |0.70959 |\n|0.380    |0.70946 |\n</pre>\n\n<p>The best threshold is 0.365 for the predictions in this case. But, I didn't care about the difference of CV scores with threshold 0.36 and 0.365 because it's small and the best threshold varies around 0.36~0.365 when training models with the completely same hyperparameter setting.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 475005,
      "author_name": "atsunorifujita",
      "author_url": "",
      "post_date": "02/20/2019 05:45:09",
      "content": "<p>Thank you for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 475628,
      "author_name": "liviaxu",
      "author_url": "",
      "post_date": "02/21/2019 02:02:19",
      "content": "<p>Thanks for sharing,i wonder how to use pretrained embedding when you already train the embedding on train-test dataset</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 476121,
      "author_name": "salonsai",
      "author_url": "",
      "post_date": "02/21/2019 16:21:52",
      "content": "<p>Congrats! You mean just train the 64 dimensional word embeddings with your model architecture?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 477215,
      "author_name": "rajshree07",
      "author_url": "",
      "post_date": "02/24/2019 06:13:48",
      "content": "<p>Congratulations!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 487170,
      "author_name": "bucchiiii",
      "author_url": "",
      "post_date": "03/10/2019 09:15:59",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 524351,
      "author_name": "evilpsycho42",
      "author_url": "",
      "post_date": "04/28/2019 15:03:36",
      "content": "<p>Can you explain about <code>1040</code> feature input?</p>\n\n<p>Thanks so much!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 533033,
      "author_name": "schiffer98",
      "author_url": "",
      "post_date": "05/18/2019 10:00:22",
      "content": "<p>Congratulations. Thank you for sharing</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 552221,
      "author_name": "davids1992",
      "author_url": "",
      "post_date": "06/13/2019 16:44:28",
      "content": "<p><a href=\"/takapt\">@takapt</a> how did you train the embeddings, do you remember? Especially what loss value did you have after the training? (or how many epochs did you train)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "474466": "## Summary\nI used a single NN with 1-layer Bi-GRU and dense layers for statistical features. I did seed averaging for ensemble. I used PyTorch to write NN.\n\nKey factors of my solution are\n\n* Tune hyperparameters based on solid CV \n* Train word embeddings on the competition dataset. I guess most participants didn't?\n* Faster training techniques to train more models. It's adaptive lengths of sequences to input RNN for each batch\n\n\n## Preprocessing\nI inserted spaces around characters except alphabets and numbers.\nThen, I used keras tokenizer, which splits by only space.\n\nAfter tokenization, I applied spell correction to OOV words. The rough idea of the spell correction algorithm is to find words with 0 or 1 levenshtein distance while ignoring cases. Precisely, there are a few heuristics.\n\nIn my case, devising preprocessing including the above spell correction did not change CV score so much.\n\n## Model architecture\nThe main part is 1-layer Bi-GRU with hidden size 128 followed by the concatenation of max pooling, average pooling and first/last positional outputs. Another part is dense layers for statistical features. The outputs of 2 network parts are concatenated, then fed to dense layers.\n\n<pre>QuoraModel(\n  (embedding): Embedding(222910, 668, padding_idx=0)\n  (text): RNNBlock(\n    (rnn): GRU(668, 128, batch_first=True, bidirectional=True)\n  )\n  (features_dense): Sequential(\n    (0): Linear(in_features=92, out_features=32, bias=True)\n    (1): ReLU(inplace)\n    (2): Linear(in_features=32, out_features=16, bias=True)\n    (3): ReLU(inplace)\n  )\n  (dense): Sequential(\n    (0): BatchNorm1d(1040, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (1): ReLU(inplace)\n    (2): Dropout(p=0.25)\n    (3): Linear(in_features=1040, out_features=64, bias=True)\n    (4): BatchNorm1d(64, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    (5): ReLU(inplace)\n    (6): Dropout(p=0.1)\n    (7): Linear(in_features=64, out_features=1, bias=True)\n  )\n)\n</pre>\n\n## Embedding\nI used glove and wiki-news pretrained word embeddings. I also trained 64 dimensional word embeddings on the competition dataset (train+test) with fastText.\nIn addition to them, I added 4 binary features to word embeddings (all upper chars?, first char upper?, only first char upper?, OOV word?).\nFinally, I concatenated all of them. The parameters of word embeddings are freezed during training.\n\n## Statistical features\n* the number of words\n* the number of unique words\n* the number of characters\n* the number of upper characters\n* Bag of characters: Implemented by `CountVectorizer(ngram_range=(1, 1), min_df=1e-4, token_pattern=r'\\w+',analyzer='char')`\n\n## Length of sequences to input RNN\nFor the faster training, I adjusted the lengths of sequences for each batch.\nWhen training, I used the maximum length of sequences in the batch or 55 length by applying pre-truncation if the maximum length over 55. When predicting for test, the truncation is applied if the length is over 70 instead of 55.\n\nThanks to this trick, I was able to train 6 models on kernel compared 5 models without this trick.\n\n\n## Training\nI used Adam with learning rate 0.001. The learning rate is multiplied by 0.8 after each epoch.\n\nI got the best CV score with batch size 256. But, batch size has the trade-off between score and training time. As I increase batch size, CV score gets worse and training gets faster. I chose batch size 320 by checking CV score and training time on kernel.\n\n## Ensemble\nI did seed averaging of 6 models. I trained 6 models with different seeds. 5 epochs are spent for each model. Each model is trained on the full train dataset, in other words, I didn't use k-fold split to train different models.\n\nI averaged the predictions of 6 models. Then, I made the final binary predictions with threshold 0.36.\n\n## Local validation\nI did 5 fold CV for the local validation. For each fold, I used predictions after ensemble rather than predictions by 1 model for more stable CV, closer CV score to LB score and more optimal hyperparameter search when ensemble.\n\n## CV score\nI show CV scores of my model used for private LB and several models without some feature.\n0.70974 is the CV score for the model used for private LB.\n\n<pre>|Removed feature                   |score  |\n|----------------------------------|-------|\n|no removal                        |0.70974|\n|4 binary embedding feature        |0.70957|\n|spell correction                  |0.70953|\n|statistical features              |0.70877|\n|word embeddings trained on dataset|0.70794|\n</pre>",
    "474881": "Congrats! Both top 2 teams used adaptive sequence length to speed things up.",
    "474914": "What a great result, and congratulations on your first prize :)  \nCould you elaborate on how you chose the threshold=0.36?",
    "475005": "Thank you for sharing!",
    "475234": "Thanks!\n\nAbout threshold selection, there is no fancy thing. It was chosen based on local CV and hard-coded. After obtaining ensembled predictions, I checked CV scores with different thresholds and chose roughly a good threshold.\n\nThe CV scores with different thresholds of the model used for private LB are\n<pre>|threshold|CV score|\n|---------|--------|\n|0.350    |0.70961 |\n|0.355    |0.70969 |\n|0.360    |0.70974 |\n|0.365    |0.70976 |\n|0.370    |0.70969 |\n|0.375    |0.70959 |\n|0.380    |0.70946 |\n</pre>\n\nThe best threshold is 0.365 for the predictions in this case. But, I didn't care about the difference of CV scores with threshold 0.36 and 0.365 because it's small and the best threshold varies around 0.36~0.365 when training models with the completely same hyperparameter setting.",
    "475628": "Thanks for sharing,i wonder how to use pretrained embedding when you already train the embedding on train-test dataset",
    "476121": "Congrats! You mean just train the 64 dimensional word embeddings with your model architecture?",
    "477215": "Congratulations!!",
    "487170": "Thanks!",
    "524351": "Can you explain about `1040` feature input?\n\nThanks so much!",
    "533033": "Congratulations. Thank you for sharing",
    "552221": "takapt how did you train the embeddings, do you remember? Especially what loss value did you have after the training? (or how many epochs did you train)"
  },
  "source": "meta"
}