{
  "id": 81632,
  "title": "4th place solution (with github)",
  "url": "/competitions/quora-insincere-questions-classification/discussion/81632",
  "author_name": "KF",
  "post_date": "2019-02-23T14:11:34.398000",
  "votes": 65,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi guys, <br>\nIt is a little bit late, but I published my solution as below: <br>\n<a href=\"https://github.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification\">https://github.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification</a> <br>\n<a href=\"https://www.kaggle.com/kfujikawa/4th-place\">https://www.kaggle.com/kfujikawa/4th-place</a> <br>\nHere I will try to summarize some of the main points of my solution.</p>\n\n<h1>Summary</h1>\n\n<p>The key factors of my solution are:</p>\n\n<ul>\n<li>Word2Vec fine-tuning</li>\n<li>400dim random sampling from 600dim word embedding per CV</li>\n<li>Simple 2layer BiLSTM model with maxpooling</li>\n<li>5-fold CV and averaging model outputs</li>\n</ul>\n\n<p><img src=\"https://raw.githubusercontent.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification/master/overview.png\" alt=\"overview\"></p>\n\n<h1>Details</h1>\n\n<h2>Preprocessing</h2>\n\n<p>I refered to the public kernel (<a href=\"https://www.kaggle.com/hengzheng/pytorch-starter\">https://www.kaggle.com/hengzheng/pytorch-starter</a>\n) for the most part, and I made slight modifications as below:</p>\n\n<ul>\n<li>Exclude filter of punctuations that <a href=\"https://github.com/keras-team/keras-preprocessing/blob/master/keras_preprocessing/text.py#L169\">Keras Tokenizer has by default</a></li>\n<li>Apply misspell corrections before punctuation spacing</li>\n<li>Insert spaces around characters except alphabets and numbers</li>\n</ul>\n\n<h2>Embedding</h2>\n\n<p>In order to improve the word embeddings which are frequent in Quora dataset but not included in pretrained vectors (Glove and Paragram), I fine-tuned the word embeddings on the competition dataset (train+test) with Word2Vec (CBOW).\nI show the results of preliminary experiments to confirm whether these word embeddings are improved or not. <br>\n<a href=\"https://www.kaggle.com/kfujikawa/word2vec-fine-tuning\">https://www.kaggle.com/kfujikawa/word2vec-fine-tuning</a></p>\n\n<p>I attempted to use word vectors obtained by concatenating before and after fine-tuning, but it was difficult due to the problem of calculation cost.\nTherefore, I decided to obtain word embeddings from 600 to 400 dimensions randomly for each CV.\nThis approach was effective not only to reduce computational cost but also to increase model diversity among CVs, so contributed to improve the score of the Public LB, although the score of the local CV has decreased.</p>\n\n<h2>Model architecture</h2>\n\n<p>I adopted simple 2layer BiLSTM model with maxpooling.\nModel details are shown as below:</p>\n\n<pre><code>BinaryClassifier(\n  (embedding): Embedding(\n    (module): Embedding(212418, 402)\n    (dropout1d): Dropout(p=0.2)\n  )\n  (encoder): Encoder(\n    (module): LSTMEncoder(\n      (rnns): ModuleList(\n        (0): LSTM(402, 128, batch_first=True, bidirectional=True)\n        (1): LSTM(256, 128, batch_first=True, bidirectional=True)\n      )\n    )\n  )\n  (aggregator): Aggregator(\n    (module): MaxPoolingAggregator()\n  )\n  (mlp): MLP(\n    (layers): Sequential(\n      (0): Linear(in_features=262, out_features=128, bias=True)\n      (1): ReLU(inplace)\n      (2): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (3): Linear(in_features=128, out_features=128, bias=True)\n      (4): ReLU(inplace)\n      (5): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    )\n  )\n  (out): Linear(in_features=128, out_features=1, bias=True)\n  (lossfunc): BCEWithLogitsLoss()\n)\n</code></pre>\n\n<h2>Statistical features for words</h2>\n\n<ul>\n<li>Whether or not the word is included in pretrained embedding</li>\n<li>IDF score</li>\n</ul>\n\n<h2>Statistical features for sentences</h2>\n\n<ul>\n<li>the number of characters</li>\n<li>the number of upper characters</li>\n<li>the rate of upper characters</li>\n<li>the number of words</li>\n<li>the number of unique words</li>\n<li>the rate of unique words</li>\n</ul>",
  "messages": [
    {
      "id": 476954,
      "postDate": "2019-02-23T14:11:34.397Z",
      "content": "<p>Hi guys, <br>\nIt is a little bit late, but I published my solution as below: <br>\n<a href=\"https://github.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification\">https://github.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification</a> <br>\n<a href=\"https://www.kaggle.com/kfujikawa/4th-place\">https://www.kaggle.com/kfujikawa/4th-place</a> <br>\nHere I will try to summarize some of the main points of my solution.</p>\n\n<h1>Summary</h1>\n\n<p>The key factors of my solution are:</p>\n\n<ul>\n<li>Word2Vec fine-tuning</li>\n<li>400dim random sampling from 600dim word embedding per CV</li>\n<li>Simple 2layer BiLSTM model with maxpooling</li>\n<li>5-fold CV and averaging model outputs</li>\n</ul>\n\n<p><img src=\"https://raw.githubusercontent.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification/master/overview.png\" alt=\"overview\"></p>\n\n<h1>Details</h1>\n\n<h2>Preprocessing</h2>\n\n<p>I refered to the public kernel (<a href=\"https://www.kaggle.com/hengzheng/pytorch-starter\">https://www.kaggle.com/hengzheng/pytorch-starter</a>\n) for the most part, and I made slight modifications as below:</p>\n\n<ul>\n<li>Exclude filter of punctuations that <a href=\"https://github.com/keras-team/keras-preprocessing/blob/master/keras_preprocessing/text.py#L169\">Keras Tokenizer has by default</a></li>\n<li>Apply misspell corrections before punctuation spacing</li>\n<li>Insert spaces around characters except alphabets and numbers</li>\n</ul>\n\n<h2>Embedding</h2>\n\n<p>In order to improve the word embeddings which are frequent in Quora dataset but not included in pretrained vectors (Glove and Paragram), I fine-tuned the word embeddings on the competition dataset (train+test) with Word2Vec (CBOW).\nI show the results of preliminary experiments to confirm whether these word embeddings are improved or not. <br>\n<a href=\"https://www.kaggle.com/kfujikawa/word2vec-fine-tuning\">https://www.kaggle.com/kfujikawa/word2vec-fine-tuning</a></p>\n\n<p>I attempted to use word vectors obtained by concatenating before and after fine-tuning, but it was difficult due to the problem of calculation cost.\nTherefore, I decided to obtain word embeddings from 600 to 400 dimensions randomly for each CV.\nThis approach was effective not only to reduce computational cost but also to increase model diversity among CVs, so contributed to improve the score of the Public LB, although the score of the local CV has decreased.</p>\n\n<h2>Model architecture</h2>\n\n<p>I adopted simple 2layer BiLSTM model with maxpooling.\nModel details are shown as below:</p>\n\n<pre><code>BinaryClassifier(\n  (embedding): Embedding(\n    (module): Embedding(212418, 402)\n    (dropout1d): Dropout(p=0.2)\n  )\n  (encoder): Encoder(\n    (module): LSTMEncoder(\n      (rnns): ModuleList(\n        (0): LSTM(402, 128, batch_first=True, bidirectional=True)\n        (1): LSTM(256, 128, batch_first=True, bidirectional=True)\n      )\n    )\n  )\n  (aggregator): Aggregator(\n    (module): MaxPoolingAggregator()\n  )\n  (mlp): MLP(\n    (layers): Sequential(\n      (0): Linear(in_features=262, out_features=128, bias=True)\n      (1): ReLU(inplace)\n      (2): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n      (3): Linear(in_features=128, out_features=128, bias=True)\n      (4): ReLU(inplace)\n      (5): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n    )\n  )\n  (out): Linear(in_features=128, out_features=1, bias=True)\n  (lossfunc): BCEWithLogitsLoss()\n)\n</code></pre>\n\n<h2>Statistical features for words</h2>\n\n<ul>\n<li>Whether or not the word is included in pretrained embedding</li>\n<li>IDF score</li>\n</ul>\n\n<h2>Statistical features for sentences</h2>\n\n<ul>\n<li>the number of characters</li>\n<li>the number of upper characters</li>\n<li>the rate of upper characters</li>\n<li>the number of words</li>\n<li>the number of unique words</li>\n<li>the rate of unique words</li>\n</ul>",
      "rawMarkdown": "Hi guys,  \nIt is a little bit late, but I published my solution as below:  \nhttps://github.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification  \nhttps://www.kaggle.com/kfujikawa/4th-place  \nHere I will try to summarize some of the main points of my solution.\n\n# Summary\n\nThe key factors of my solution are:\n\n- Word2Vec fine-tuning\n- 400dim random sampling from 600dim word embedding per CV\n- Simple 2layer BiLSTM model with maxpooling\n- 5-fold CV and averaging model outputs\n\n![overview](https://raw.githubusercontent.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification/master/overview.png)\n\n# Details\n\n## Preprocessing\n\nI refered to the public kernel (https://www.kaggle.com/hengzheng/pytorch-starter\n) for the most part, and I made slight modifications as below:\n\n- Exclude filter of punctuations that [Keras Tokenizer has by default](https://github.com/keras-team/keras-preprocessing/blob/master/keras_preprocessing/text.py#L169)\n- Apply misspell corrections before punctuation spacing\n- Insert spaces around characters except alphabets and numbers\n\n## Embedding\n\nIn order to improve the word embeddings which are frequent in Quora dataset but not included in pretrained vectors (Glove and Paragram), I fine-tuned the word embeddings on the competition dataset (train+test) with Word2Vec (CBOW).\nI show the results of preliminary experiments to confirm whether these word embeddings are improved or not.  \nhttps://www.kaggle.com/kfujikawa/word2vec-fine-tuning\n\nI attempted to use word vectors obtained by concatenating before and after fine-tuning, but it was difficult due to the problem of calculation cost.\nTherefore, I decided to obtain word embeddings from 600 to 400 dimensions randomly for each CV.\nThis approach was effective not only to reduce computational cost but also to increase model diversity among CVs, so contributed to improve the score of the Public LB, although the score of the local CV has decreased.\n\n## Model architecture\n\nI adopted simple 2layer BiLSTM model with maxpooling.\nModel details are shown as below:\n\n    BinaryClassifier(\n      (embedding): Embedding(\n        (module): Embedding(212418, 402)\n        (dropout1d): Dropout(p=0.2)\n      )\n      (encoder): Encoder(\n        (module): LSTMEncoder(\n          (rnns): ModuleList(\n            (0): LSTM(402, 128, batch_first=True, bidirectional=True)\n            (1): LSTM(256, 128, batch_first=True, bidirectional=True)\n          )\n        )\n      )\n      (aggregator): Aggregator(\n        (module): MaxPoolingAggregator()\n      )\n      (mlp): MLP(\n        (layers): Sequential(\n          (0): Linear(in_features=262, out_features=128, bias=True)\n          (1): ReLU(inplace)\n          (2): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n          (3): Linear(in_features=128, out_features=128, bias=True)\n          (4): ReLU(inplace)\n          (5): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (out): Linear(in_features=128, out_features=1, bias=True)\n      (lossfunc): BCEWithLogitsLoss()\n    )\n\n## Statistical features for words\n\n- Whether or not the word is included in pretrained embedding\n- IDF score\n\n## Statistical features for sentences\n\n- the number of characters\n- the number of upper characters\n- the rate of upper characters\n- the number of words\n- the number of unique words\n- the rate of unique words\n\n\n\n",
      "votes": 65
    },
    {
      "id": 531989,
      "postDate": "2019-05-16T01:18:41.797Z",
      "content": "<p>What do you think about adding Attention to your model ?</p>",
      "rawMarkdown": "What do you think about adding Attention to your model ?",
      "votes": 2,
      "replies": [
        {
          "id": 533225,
          "postDate": "2019-05-18T18:03:22.167Z",
          "content": "<p>In my case, it didn't improve my score, but I think it might work depending on the task or implementation of the attention module.</p>",
          "rawMarkdown": "In my case, it didn't improve my score, but I think it might work depending on the task or implementation of the attention module."
        },
        {
          "id": 533550,
          "postDate": "2019-05-19T13:29:39.587Z",
          "content": "<p>right! I also tried but did not improve the score. Not Free Lunch Theorem ^ ^</p>",
          "rawMarkdown": "right! I also tried but did not improve the score. Not Free Lunch Theorem ^ ^",
          "votes": 2
        }
      ]
    },
    {
      "id": 478500,
      "postDate": "2019-02-26T08:28:01.147Z",
      "content": "<p>Good job!</p>",
      "rawMarkdown": "Good job!"
    },
    {
      "id": 531299,
      "postDate": "2019-05-14T16:04:56.440Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 533030,
      "postDate": "2019-05-18T09:57:13.070Z",
      "content": "<p>NIce. Thank KF</p>",
      "rawMarkdown": "NIce. Thank KF"
    },
    {
      "id": 481198,
      "postDate": "2019-03-01T06:24:34.377Z",
      "content": "<p>Great explanation. Thanks KF.</p>",
      "rawMarkdown": "Great explanation. Thanks KF."
    }
  ],
  "comments": [
    {
      "id": 531989,
      "author_name": "KhanhVD",
      "author_url": "",
      "post_date": "2019-05-16T01:18:41.797000",
      "content": "<p>What do you think about adding Attention to your model ?</p>",
      "votes": 2,
      "replies": [
        {
          "id": 533225,
          "author_name": "KF",
          "author_url": "",
          "post_date": "2019-05-18T18:03:22.167000",
          "content": "<p>In my case, it didn't improve my score, but I think it might work depending on the task or implementation of the attention module.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 533550,
          "author_name": "KhanhVD",
          "author_url": "",
          "post_date": "2019-05-19T13:29:39.587000",
          "content": "<p>right! I also tried but did not improve the score. Not Free Lunch Theorem ^ ^</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 478500,
      "author_name": "zhusleep",
      "author_url": "",
      "post_date": "2019-02-26T08:28:01.147000",
      "content": "<p>Good job!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 531299,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-05-14T16:04:56.440000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 533030,
      "author_name": "Fikhri",
      "author_url": "",
      "post_date": "2019-05-18T09:57:13.070000",
      "content": "<p>NIce. Thank KF</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 481198,
      "author_name": "RaviT",
      "author_url": "",
      "post_date": "2019-03-01T06:24:34.377000",
      "content": "<p>Great explanation. Thanks KF.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "476954": "Hi guys,  \nIt is a little bit late, but I published my solution as below:  \nhttps://github.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification  \nhttps://www.kaggle.com/kfujikawa/4th-place  \nHere I will try to summarize some of the main points of my solution.\n\n# Summary\n\nThe key factors of my solution are:\n\n- Word2Vec fine-tuning\n- 400dim random sampling from 600dim word embedding per CV\n- Simple 2layer BiLSTM model with maxpooling\n- 5-fold CV and averaging model outputs\n\n![overview](https://raw.githubusercontent.com/k-fujikawa/Kaggle-Quora-Insincere-Questions-Classification/master/overview.png)\n\n# Details\n\n## Preprocessing\n\nI refered to the public kernel (https://www.kaggle.com/hengzheng/pytorch-starter\n) for the most part, and I made slight modifications as below:\n\n- Exclude filter of punctuations that [Keras Tokenizer has by default](https://github.com/keras-team/keras-preprocessing/blob/master/keras_preprocessing/text.py#L169)\n- Apply misspell corrections before punctuation spacing\n- Insert spaces around characters except alphabets and numbers\n\n## Embedding\n\nIn order to improve the word embeddings which are frequent in Quora dataset but not included in pretrained vectors (Glove and Paragram), I fine-tuned the word embeddings on the competition dataset (train+test) with Word2Vec (CBOW).\nI show the results of preliminary experiments to confirm whether these word embeddings are improved or not.  \nhttps://www.kaggle.com/kfujikawa/word2vec-fine-tuning\n\nI attempted to use word vectors obtained by concatenating before and after fine-tuning, but it was difficult due to the problem of calculation cost.\nTherefore, I decided to obtain word embeddings from 600 to 400 dimensions randomly for each CV.\nThis approach was effective not only to reduce computational cost but also to increase model diversity among CVs, so contributed to improve the score of the Public LB, although the score of the local CV has decreased.\n\n## Model architecture\n\nI adopted simple 2layer BiLSTM model with maxpooling.\nModel details are shown as below:\n\n    BinaryClassifier(\n      (embedding): Embedding(\n        (module): Embedding(212418, 402)\n        (dropout1d): Dropout(p=0.2)\n      )\n      (encoder): Encoder(\n        (module): LSTMEncoder(\n          (rnns): ModuleList(\n            (0): LSTM(402, 128, batch_first=True, bidirectional=True)\n            (1): LSTM(256, 128, batch_first=True, bidirectional=True)\n          )\n        )\n      )\n      (aggregator): Aggregator(\n        (module): MaxPoolingAggregator()\n      )\n      (mlp): MLP(\n        (layers): Sequential(\n          (0): Linear(in_features=262, out_features=128, bias=True)\n          (1): ReLU(inplace)\n          (2): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n          (3): Linear(in_features=128, out_features=128, bias=True)\n          (4): ReLU(inplace)\n          (5): BatchNorm1d(128, eps=1e-05, momentum=0.1, affine=True, track_running_stats=True)\n        )\n      )\n      (out): Linear(in_features=128, out_features=1, bias=True)\n      (lossfunc): BCEWithLogitsLoss()\n    )\n\n## Statistical features for words\n\n- Whether or not the word is included in pretrained embedding\n- IDF score\n\n## Statistical features for sentences\n\n- the number of characters\n- the number of upper characters\n- the rate of upper characters\n- the number of words\n- the number of unique words\n- the rate of unique words\n\n\n\n",
    "531989": "What do you think about adding Attention to your model ?",
    "478500": "Good job!",
    "531299": "",
    "533030": "NIce. Thank KF",
    "481198": "Great explanation. Thanks KF."
  }
}