{
  "id": 80561,
  "title": "7th place solution - bucketing",
  "url": "/competitions/quora-insincere-questions-classification/writeups/yufuin-7th-place-solution-bucketing",
  "author_name": "",
  "post_date": "2019-02-14T19:21:16.563Z",
  "votes": 45,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hi,</p>\n\n<p>Here I explain my solution.</p>\n\n<p>I think bucketing and checkpoint-ensembling are the key factors of my solution, since my preprocessing and my model are quite basic.</p>\n\n<h1>Preprocess</h1>\n\n<p>The core part here is using NLTK TweetTokenizer.</p>\n\n<ol>\n<li>Split each question_text by \" \" (space).</li>\n<li>Replace words that has \"*\" with \"FWORD\", since NLTK TweetTokenizer will split by \"*\", but I want to use words like \"f**k\" as the single token rather than [\"f\", \"*\", \"*\", \"k\"].</li>\n<li>Join by \" \", then apply NLTK TweetTokenizer.</li>\n<li>Split each word by ' (single quote) and -. e.g., [\"it's\", \"nice\"] -&gt; [\"it\", \"'s\", \"nice\"]</li>\n<li>Load pretrained embeddings. For \"FWORD\", using the average of the embeddings of [\"fuck\", \"shit\", \"*\"]. For OOV, using the average of embeddings.</li>\n</ol>\n\n<p>I think that the preprocessing other than applying TweetTokenizer doesn't make big difference, since whether applying such \"*\"-replacement or not doesn't change the local CV score. The only reason why I subimitted this version is just I couldn't ignore the time I spent for preprocessing. XP</p>\n\n<h1>Model</h1>\n\n<ol>\n<li>Embedding layer. Simple average of Glove and Paragram embeddings (thus dim=300). Keep fixed.</li>\n<li>Dropout (keep_prob=0.6)</li>\n<li>Bi-LSTM (each cell_size=128)</li>\n<li>Bi-LSTM (each cell_size=128)</li>\n<li>Concatenation of the average-pooling of the first Bi-LSTM, the max-pooling of the second Bi-LSTM and attention of the second Bi-LSTM. (thus dim=3*256)</li>\n<li>Dense with tanh (dim=32)</li>\n<li>Output with sigmoid</li>\n</ol>\n\n<h1>Training</h1>\n\n<h2>Use bucketing.</h2>\n\n<p>Bucketing is to make a minibatch from instances that have simillar lengths to alleviate the cost of padding. This makes the training speed more than 3x faster and thus I can run 9 epochs for each split of 5-fold.</p>\n\n<p>I must have seen the TensorFlow tutorial page that describes bucketing (it shoud be the tutorial of \"sequence-to-sequence model\"), however, somehow I couldn't find that page now.</p>\n\n<p>For other training details,</p>\n\n<ul>\n<li>Objective function: vanilla sigmoid_cross_entropy</li>\n<li>Optimizer: Adam with default parameters</li>\n<li>Batch size: 512</li>\n<li>Maximum sequence length of the input: 400</li>\n</ul>\n\n<h1>Postprocess</h1>\n\n<p>For each 5-fold model, apply checkpoint-ensembling to maximize each validation score.\nWithout checkpoint-ensembling, the average validation score is about 0.694. After checkpoint-ensembling, it is about 0.700.</p>\n\n<p>After checkpoint-ensembling, ensemble 5 models by averaging output probabilities and thresholds, then submit.</p>\n\n<p>Finally, Thanks everyone working for this competition! I really enjoyed this competition with such a big data!</p>",
  "messages": [
    {
      "id": "471426",
      "postDate": "02/14/2019 12:50:27",
      "content": "<p>Hi,</p>\n\n<p>Here I explain my solution.</p>\n\n<p>I think bucketing and checkpoint-ensembling are the key factors of my solution, since my preprocessing and my model are quite basic.</p>\n\n<h1>Preprocess</h1>\n\n<p>The core part here is using NLTK TweetTokenizer.</p>\n\n<ol>\n<li>Split each question_text by \" \" (space).</li>\n<li>Replace words that has \"*\" with \"FWORD\", since NLTK TweetTokenizer will split by \"*\", but I want to use words like \"f**k\" as the single token rather than [\"f\", \"*\", \"*\", \"k\"].</li>\n<li>Join by \" \", then apply NLTK TweetTokenizer.</li>\n<li>Split each word by ' (single quote) and -. e.g., [\"it's\", \"nice\"] -&gt; [\"it\", \"'s\", \"nice\"]</li>\n<li>Load pretrained embeddings. For \"FWORD\", using the average of the embeddings of [\"fuck\", \"shit\", \"*\"]. For OOV, using the average of embeddings.</li>\n</ol>\n\n<p>I think that the preprocessing other than applying TweetTokenizer doesn't make big difference, since whether applying such \"*\"-replacement or not doesn't change the local CV score. The only reason why I subimitted this version is just I couldn't ignore the time I spent for preprocessing. XP</p>\n\n<h1>Model</h1>\n\n<ol>\n<li>Embedding layer. Simple average of Glove and Paragram embeddings (thus dim=300). Keep fixed.</li>\n<li>Dropout (keep_prob=0.6)</li>\n<li>Bi-LSTM (each cell_size=128)</li>\n<li>Bi-LSTM (each cell_size=128)</li>\n<li>Concatenation of the average-pooling of the first Bi-LSTM, the max-pooling of the second Bi-LSTM and attention of the second Bi-LSTM. (thus dim=3*256)</li>\n<li>Dense with tanh (dim=32)</li>\n<li>Output with sigmoid</li>\n</ol>\n\n<h1>Training</h1>\n\n<h2>Use bucketing.</h2>\n\n<p>Bucketing is to make a minibatch from instances that have simillar lengths to alleviate the cost of padding. This makes the training speed more than 3x faster and thus I can run 9 epochs for each split of 5-fold.</p>\n\n<p>I must have seen the TensorFlow tutorial page that describes bucketing (it shoud be the tutorial of \"sequence-to-sequence model\"), however, somehow I couldn't find that page now.</p>\n\n<p>For other training details,</p>\n\n<ul>\n<li>Objective function: vanilla sigmoid_cross_entropy</li>\n<li>Optimizer: Adam with default parameters</li>\n<li>Batch size: 512</li>\n<li>Maximum sequence length of the input: 400</li>\n</ul>\n\n<h1>Postprocess</h1>\n\n<p>For each 5-fold model, apply checkpoint-ensembling to maximize each validation score.\nWithout checkpoint-ensembling, the average validation score is about 0.694. After checkpoint-ensembling, it is about 0.700.</p>\n\n<p>After checkpoint-ensembling, ensemble 5 models by averaging output probabilities and thresholds, then submit.</p>\n\n<p>Finally, Thanks everyone working for this competition! I really enjoyed this competition with such a big data!</p>",
      "rawMarkdown": "Hi,\n\nHere I explain my solution.\n\nI think bucketing and checkpoint-ensembling are the key factors of my solution, since my preprocessing and my model are quite basic.\n\n# Preprocess\nThe core part here is using NLTK TweetTokenizer.\n\n1. Split each question_text by \" \" (space).\n1. Replace words that has \"\\*\" with \"FWORD\", since NLTK TweetTokenizer will split by \"*\", but I want to use words like \"f\\*\\*k\" as the single token rather than [\"f\", \"\\*\", \"\\*\", \"k\"].\n1. Join by \" \", then apply NLTK TweetTokenizer.\n1. Split each word by ' (single quote) and -. e.g., [\"it's\", \"nice\"] -&gt; [\"it\", \"'s\", \"nice\"]\n1. Load pretrained embeddings. For \"FWORD\", using the average of the embeddings of [\"fuck\", \"shit\", \"\\*\"]. For OOV, using the average of embeddings.\n\nI think that the preprocessing other than applying TweetTokenizer doesn't make big difference, since whether applying such \"*\"-replacement or not doesn't change the local CV score. The only reason why I subimitted this version is just I couldn't ignore the time I spent for preprocessing. XP\n\n# Model\n\n1.  Embedding layer. Simple average of Glove and Paragram embeddings (thus dim=300). Keep fixed.\n1. Dropout (keep_prob=0.6)\n1. Bi-LSTM (each cell_size=128)\n1. Bi-LSTM (each cell_size=128)\n1. Concatenation of the average-pooling of the first Bi-LSTM, the max-pooling of the second Bi-LSTM and attention of the second Bi-LSTM. (thus dim=3\\*256)\n1. Dense with tanh (dim=32)\n1. Output with sigmoid\n\n# Training\n## Use bucketing.\nBucketing is to make a minibatch from instances that have simillar lengths to alleviate the cost of padding. This makes the training speed more than 3x faster and thus I can run 9 epochs for each split of 5-fold.\n\nI must have seen the TensorFlow tutorial page that describes bucketing (it shoud be the tutorial of \"sequence-to-sequence model\"), however, somehow I couldn't find that page now.\n\nFor other training details,\n\n* Objective function: vanilla sigmoid\\_cross\\_entropy\n* Optimizer: Adam with default parameters\n* Batch size: 512\n* Maximum sequence length of the input: 400\n\n# Postprocess\nFor each 5-fold model, apply checkpoint-ensembling to maximize each validation score.\nWithout checkpoint-ensembling, the average validation score is about 0.694. After checkpoint-ensembling, it is about 0.700.\n\nAfter checkpoint-ensembling, ensemble 5 models by averaging output probabilities and thresholds, then submit.\n\nFinally, Thanks everyone working for this competition! I really enjoyed this competition with such a big data!",
      "votes": null
    },
    {
      "id": "471430",
      "postDate": "02/14/2019 12:54:39",
      "content": "<p>congratulation~</p>",
      "rawMarkdown": "congratulation~",
      "votes": null
    },
    {
      "id": "471484",
      "postDate": "02/14/2019 14:22:14",
      "content": "<p>Thank you!</p>",
      "rawMarkdown": "Thank you!",
      "votes": null
    },
    {
      "id": "471601",
      "postDate": "02/14/2019 16:23:58",
      "content": "<p>Congratulations on the high finish!  Thanks for sharing your solution.</p>",
      "rawMarkdown": "Congratulations on the high finish!  Thanks for sharing your solution.",
      "votes": null
    },
    {
      "id": "472278",
      "postDate": "02/15/2019 16:19:08",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "472494",
      "postDate": "02/16/2019 03:04:58",
      "content": "<p>the idea of using bucketing for training mor epochs is really smart</p>",
      "rawMarkdown": "the idea of using bucketing for training mor epochs is really smart",
      "votes": null
    },
    {
      "id": "472503",
      "postDate": "02/16/2019 03:59:04",
      "content": "<p>Great idea on bucketing. It seems the 1st place solution also has bucketing to enable more training. Also find the bucketing in a stanford course, here is the google slides (slide 25-28): <a href=\"https://docs.google.com/presentation/d/1bJhLv_sV81uR3tW90VAqIWvZHCQokHspGHnFaB5Qv8g/edit?usp=sharing\">https://docs.google.com/presentation/d/1bJhLv_sV81uR3tW90VAqIWvZHCQokHspGHnFaB5Qv8g/edit?usp=sharing</a></p>",
      "rawMarkdown": "Great idea on bucketing. It seems the 1st place solution also has bucketing to enable more training. Also find the bucketing in a stanford course, here is the google slides (slide 25-28): https://docs.google.com/presentation/d/1bJhLv_sV81uR3tW90VAqIWvZHCQokHspGHnFaB5Qv8g/edit?usp=sharing",
      "votes": null
    },
    {
      "id": "472683",
      "postDate": "02/16/2019 13:18:42",
      "content": "<p>Could you elaborate on hat is checkpoint-ensembing and how did you do it? Another thing that was unclear to ne was the application of the bucketing. Does it mean that you do padding dynamically, i.e. on batch level, similar to the 1-st place solution?</p>",
      "rawMarkdown": "Could you elaborate on hat is checkpoint-ensembing and how did you do it? Another thing that was unclear to ne was the application of the bucketing. Does it mean that you do padding dynamically, i.e. on batch level, similar to the 1-st place solution?",
      "votes": null
    },
    {
      "id": "472952",
      "postDate": "02/17/2019 00:47:12",
      "content": "<p>Your bucketing approach is quite unique from the others, Thanks for sharing !!</p>",
      "rawMarkdown": "Your bucketing approach is quite unique from the others, Thanks for sharing !!",
      "votes": null
    },
    {
      "id": "582439",
      "postDate": "07/23/2019 07:28:08",
      "content": "<p>can you share your code, let me see see,We will be very grateful。</p>",
      "rawMarkdown": "can you share your code, let me see see,We will be very grateful。",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 471430,
      "author_name": "",
      "author_url": "",
      "post_date": "02/14/2019 12:54:39",
      "content": "<p>congratulation~</p>",
      "votes": null,
      "replies": [
        {
          "id": 471484,
          "author_name": "yufuin",
          "author_url": "",
          "post_date": "02/14/2019 14:22:14",
          "content": "<p>Thank you!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 471601,
      "author_name": "mtodisco10",
      "author_url": "",
      "post_date": "02/14/2019 16:23:58",
      "content": "<p>Congratulations on the high finish!  Thanks for sharing your solution.</p>",
      "votes": null,
      "replies": [
        {
          "id": 472278,
          "author_name": "yufuin",
          "author_url": "",
          "post_date": "02/15/2019 16:19:08",
          "content": "<p>Thanks!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 472494,
      "author_name": "zjucor",
      "author_url": "",
      "post_date": "02/16/2019 03:04:58",
      "content": "<p>the idea of using bucketing for training mor epochs is really smart</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472503,
      "author_name": "nzholmes",
      "author_url": "",
      "post_date": "02/16/2019 03:59:04",
      "content": "<p>Great idea on bucketing. It seems the 1st place solution also has bucketing to enable more training. Also find the bucketing in a stanford course, here is the google slides (slide 25-28): <a href=\"https://docs.google.com/presentation/d/1bJhLv_sV81uR3tW90VAqIWvZHCQokHspGHnFaB5Qv8g/edit?usp=sharing\">https://docs.google.com/presentation/d/1bJhLv_sV81uR3tW90VAqIWvZHCQokHspGHnFaB5Qv8g/edit?usp=sharing</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472683,
      "author_name": "mlisovyi",
      "author_url": "",
      "post_date": "02/16/2019 13:18:42",
      "content": "<p>Could you elaborate on hat is checkpoint-ensembing and how did you do it? Another thing that was unclear to ne was the application of the bucketing. Does it mean that you do padding dynamically, i.e. on batch level, similar to the 1-st place solution?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 472952,
      "author_name": "viswanathravindran",
      "author_url": "",
      "post_date": "02/17/2019 00:47:12",
      "content": "<p>Your bucketing approach is quite unique from the others, Thanks for sharing !!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 582439,
      "author_name": "kuweideshuye",
      "author_url": "",
      "post_date": "07/23/2019 07:28:08",
      "content": "<p>can you share your code, let me see see,We will be very grateful。</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "471426": "Hi,\n\nHere I explain my solution.\n\nI think bucketing and checkpoint-ensembling are the key factors of my solution, since my preprocessing and my model are quite basic.\n\n# Preprocess\nThe core part here is using NLTK TweetTokenizer.\n\n1. Split each question_text by \" \" (space).\n1. Replace words that has \"\\*\" with \"FWORD\", since NLTK TweetTokenizer will split by \"*\", but I want to use words like \"f\\*\\*k\" as the single token rather than [\"f\", \"\\*\", \"\\*\", \"k\"].\n1. Join by \" \", then apply NLTK TweetTokenizer.\n1. Split each word by ' (single quote) and -. e.g., [\"it's\", \"nice\"] -&gt; [\"it\", \"'s\", \"nice\"]\n1. Load pretrained embeddings. For \"FWORD\", using the average of the embeddings of [\"fuck\", \"shit\", \"\\*\"]. For OOV, using the average of embeddings.\n\nI think that the preprocessing other than applying TweetTokenizer doesn't make big difference, since whether applying such \"*\"-replacement or not doesn't change the local CV score. The only reason why I subimitted this version is just I couldn't ignore the time I spent for preprocessing. XP\n\n# Model\n\n1.  Embedding layer. Simple average of Glove and Paragram embeddings (thus dim=300). Keep fixed.\n1. Dropout (keep_prob=0.6)\n1. Bi-LSTM (each cell_size=128)\n1. Bi-LSTM (each cell_size=128)\n1. Concatenation of the average-pooling of the first Bi-LSTM, the max-pooling of the second Bi-LSTM and attention of the second Bi-LSTM. (thus dim=3\\*256)\n1. Dense with tanh (dim=32)\n1. Output with sigmoid\n\n# Training\n## Use bucketing.\nBucketing is to make a minibatch from instances that have simillar lengths to alleviate the cost of padding. This makes the training speed more than 3x faster and thus I can run 9 epochs for each split of 5-fold.\n\nI must have seen the TensorFlow tutorial page that describes bucketing (it shoud be the tutorial of \"sequence-to-sequence model\"), however, somehow I couldn't find that page now.\n\nFor other training details,\n\n* Objective function: vanilla sigmoid\\_cross\\_entropy\n* Optimizer: Adam with default parameters\n* Batch size: 512\n* Maximum sequence length of the input: 400\n\n# Postprocess\nFor each 5-fold model, apply checkpoint-ensembling to maximize each validation score.\nWithout checkpoint-ensembling, the average validation score is about 0.694. After checkpoint-ensembling, it is about 0.700.\n\nAfter checkpoint-ensembling, ensemble 5 models by averaging output probabilities and thresholds, then submit.\n\nFinally, Thanks everyone working for this competition! I really enjoyed this competition with such a big data!",
    "471430": "congratulation~",
    "471484": "Thank you!",
    "471601": "Congratulations on the high finish!  Thanks for sharing your solution.",
    "472278": "Thanks!",
    "472494": "the idea of using bucketing for training mor epochs is really smart",
    "472503": "Great idea on bucketing. It seems the 1st place solution also has bucketing to enable more training. Also find the bucketing in a stanford course, here is the google slides (slide 25-28): https://docs.google.com/presentation/d/1bJhLv_sV81uR3tW90VAqIWvZHCQokHspGHnFaB5Qv8g/edit?usp=sharing",
    "472683": "Could you elaborate on hat is checkpoint-ensembing and how did you do it? Another thing that was unclear to ne was the application of the bucketing. Does it mean that you do padding dynamically, i.e. on batch level, similar to the 1-st place solution?",
    "472952": "Your bucketing approach is quite unique from the others, Thanks for sharing !!",
    "582439": "can you share your code, let me see see,We will be very grateful。"
  },
  "source": "meta"
}