{
  "id": 127259,
  "title": "7th place solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/127259",
  "author_name": "Bo",
  "post_date": "2020-01-23T04:56:36.910000",
  "votes": 52,
  "comment_count": 16,
  "views": 0,
  "content": "<h1>Framework and hardware</h1>\n\n<p>Initially we set out to try both TF 1.15 and TF 2.0. Unfortunately TF 2.0 pipeline's scores are significantly lower (we probably didn't figure out how to correctly use the TF 2.0 API), so all our top submissions and what is described below were done in TF 1.15. </p>\n\n<p>All the experiments were done in Google cloud TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast and Google is generous enough to give us 5 TPUs, the experiment cycle is dramatically reduced.</p>\n\n<h1>Validation scheme and experiment setup</h1>\n\n<p>We noticed that the evaluation metric is not very stable on smaller validation set, for example better models on dev00 (1600 examples) may not be better on dev01. So we rely solely on the whole dev set (7830 examples) for validation (i.e. selecting checkpoints, selecting models, tuning thresholds, tuning ensemble weights). Same thing goes for the public LB. It only has 346 examples, even less stable than dev00.</p>\n\n<p>Most of our models are based on the official implementation of <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">joint-bert</a>, which is surprisingly hard to beat. For each variation to joint-bert, we ran 5 training sessions on TPU simultaneously, with different batch size and learning rate. We save checkpoints every 500 or 1000 steps. Usually bs=64 and lr=4e-5 for 1 epoch gives the best scores.</p>\n\n<h1>Variations to official joint-bert</h1>\n\n<p>We tried the following variations to joint-bert. The ideas for many of them are drawn from this IBM paper: <a href=\"https://arxiv.org/pdf/1909.05286.pdf\">https://arxiv.org/pdf/1909.05286.pdf</a>\n1) and 2) are the most important. Adding them always help. They contribute to the best single model. 3) to 6) sometimes help, sometimes don't. They depend on each other. Nevertheless, training with these variations did produce many diverse models, which are good for ensembling. </p>\n\n<h2>1) pre-trained weights</h2>\n\n<p>Official joint-bert was trained from \"BERT-Large, Uncased\", but training from \"BERT-Large, Uncased (Whole Word Masking)\" will see a big boost.</p>\n\n<p>We also tried fine-tuning joint-bert on Squad 2.0 before fine-tuning it on NQ.</p>\n\n<h2>2) negative sampling</h2>\n\n<p>Official joint-bert samples 2% negative examples in both answerable questions (i.e. the sliding windows that don't contain an answer) and unanswerable questions. As explained in the IBM paper, joint-bert tends to be overconfident for unanswerable questions, so 1% for answerable and 4% unanswerable seem to be better. We saw about 1 point increase in F1 doing negative sampling.</p>\n\n<h2>3) <code>max_seq_length</code>, <code>doc_stride</code></h2>\n\n<p>Default are 512 and 128 respectively. This means for answers not in the beginning of article, they appear about 4 times after pre-processing (in training the example is seen 4 times per epoch, in inference it's predicted 4 times, then the one with max logits is selected), which seems like an overkill.</p>\n\n<p>So we changed doc_stride to 256 during inference, which doesn't affect score much but reduced inference time in half. </p>\n\n<p>For training, we used default 128 as well 192 and 256. IBM paper claims 192 gives best results, but we didn't see much difference.</p>\n\n<h2>4) max_contexts</h2>\n\n<p>Default is 48. There are some very very long wikepedia articles, so joint-bert only take first 48 paragraphs/tables/lists of each article. We tried different values like 100 and 200. Using a bigger value is tradeoff between more answer coverage v.s. more \"empty\" windows.</p>\n\n<h2>5) sentence order shuffling</h2>\n\n<p>Also proposed in the IBM paper: shuffling all the sentences in the paragraph containing short answers. This is an augmentation method. </p>\n\n<h2>6) cased</h2>\n\n<p><code>do_lower_case=False</code>\nFor this, we generated a new vocab file by adding all the NQ special tokens into the cased BERT vocab.</p>\n\n<h2>7) Attention-over-attention</h2>\n\n<p>Mentioned in the IBM paper as the most important change, but it didn't work for us.</p>\n\n<h1>Ensemble</h1>\n\n<p>In 3 hours, we can do inference for 3 models with doc_stride=256. Luckily, n=3 happen to be the number of our best ensemble: adding a fourth model does not help anymore. The ensemble strategy is simply averaging the probability of each candidate span.</p>\n\n<p>Our 2 submissions consist of the following 5 single models:\na. wwm, stride=256, dev 62.4\nb. wwm, neg sampling, pre-tuned on squad, <strong>dev 64.7</strong> (long 69.5, short 57.8) - best single model\nc. wwm, neg sampling, max_contexts=200, dev 64.5\nd. wwm, neg sampling, stride=192, dev 63.8\ne. wwm, neg sampling, cased, dev 63.3</p>\n\n<p>sub1: ensemble of a,b,c, dev 66.8 (long 71.6, short 59.8), private LB 0.69\nsub2: ensemble of c,d,e, <strong>dev 67.0</strong> (long 71.6, short 59.9), private LB 0.69\n(note: these scores are after post-processing)</p>\n\n<h1>Post-process</h1>\n\n<h3>yes/no thresholds</h3>\n\n<p>These are tuned on dev set as well. If the yes/no logits in the <code>answer_type_logit</code> are over the thresholds, predict \"YES\"/\"NO\" regardless of the short span predictions. This gives 0.5 boost to dev F1.</p>\n\n<h3>max_contexts</h3>\n\n<p>Increase max_contexts from the default 48 to 100 or 80 can squeeze out another 0.3 F1 points, taking advantage of the leftover inference time within the 3 hour limit. For sub1 we did 100; for sub2 we only did 80 because generating features for the cased model <code>e</code> took a little more time.</p>\n\n<h1>Code</h1>\n\n<p>repo: <a href=\"https://github.com/boliu61/tf2qa\">https://github.com/boliu61/tf2qa</a>\ninference notebook and model weights: <a href=\"https://www.kaggle.com/boliu0/7th-place-submission\">https://www.kaggle.com/boliu0/7th-place-submission</a></p>",
  "messages": [
    {
      "id": 726570,
      "postDate": "2020-01-23T04:56:36.910Z",
      "content": "<h1>Framework and hardware</h1>\n\n<p>Initially we set out to try both TF 1.15 and TF 2.0. Unfortunately TF 2.0 pipeline's scores are significantly lower (we probably didn't figure out how to correctly use the TF 2.0 API), so all our top submissions and what is described below were done in TF 1.15. </p>\n\n<p>All the experiments were done in Google cloud TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast and Google is generous enough to give us 5 TPUs, the experiment cycle is dramatically reduced.</p>\n\n<h1>Validation scheme and experiment setup</h1>\n\n<p>We noticed that the evaluation metric is not very stable on smaller validation set, for example better models on dev00 (1600 examples) may not be better on dev01. So we rely solely on the whole dev set (7830 examples) for validation (i.e. selecting checkpoints, selecting models, tuning thresholds, tuning ensemble weights). Same thing goes for the public LB. It only has 346 examples, even less stable than dev00.</p>\n\n<p>Most of our models are based on the official implementation of <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">joint-bert</a>, which is surprisingly hard to beat. For each variation to joint-bert, we ran 5 training sessions on TPU simultaneously, with different batch size and learning rate. We save checkpoints every 500 or 1000 steps. Usually bs=64 and lr=4e-5 for 1 epoch gives the best scores.</p>\n\n<h1>Variations to official joint-bert</h1>\n\n<p>We tried the following variations to joint-bert. The ideas for many of them are drawn from this IBM paper: <a href=\"https://arxiv.org/pdf/1909.05286.pdf\">https://arxiv.org/pdf/1909.05286.pdf</a>\n1) and 2) are the most important. Adding them always help. They contribute to the best single model. 3) to 6) sometimes help, sometimes don't. They depend on each other. Nevertheless, training with these variations did produce many diverse models, which are good for ensembling. </p>\n\n<h2>1) pre-trained weights</h2>\n\n<p>Official joint-bert was trained from \"BERT-Large, Uncased\", but training from \"BERT-Large, Uncased (Whole Word Masking)\" will see a big boost.</p>\n\n<p>We also tried fine-tuning joint-bert on Squad 2.0 before fine-tuning it on NQ.</p>\n\n<h2>2) negative sampling</h2>\n\n<p>Official joint-bert samples 2% negative examples in both answerable questions (i.e. the sliding windows that don't contain an answer) and unanswerable questions. As explained in the IBM paper, joint-bert tends to be overconfident for unanswerable questions, so 1% for answerable and 4% unanswerable seem to be better. We saw about 1 point increase in F1 doing negative sampling.</p>\n\n<h2>3) <code>max_seq_length</code>, <code>doc_stride</code></h2>\n\n<p>Default are 512 and 128 respectively. This means for answers not in the beginning of article, they appear about 4 times after pre-processing (in training the example is seen 4 times per epoch, in inference it's predicted 4 times, then the one with max logits is selected), which seems like an overkill.</p>\n\n<p>So we changed doc_stride to 256 during inference, which doesn't affect score much but reduced inference time in half. </p>\n\n<p>For training, we used default 128 as well 192 and 256. IBM paper claims 192 gives best results, but we didn't see much difference.</p>\n\n<h2>4) max_contexts</h2>\n\n<p>Default is 48. There are some very very long wikepedia articles, so joint-bert only take first 48 paragraphs/tables/lists of each article. We tried different values like 100 and 200. Using a bigger value is tradeoff between more answer coverage v.s. more \"empty\" windows.</p>\n\n<h2>5) sentence order shuffling</h2>\n\n<p>Also proposed in the IBM paper: shuffling all the sentences in the paragraph containing short answers. This is an augmentation method. </p>\n\n<h2>6) cased</h2>\n\n<p><code>do_lower_case=False</code>\nFor this, we generated a new vocab file by adding all the NQ special tokens into the cased BERT vocab.</p>\n\n<h2>7) Attention-over-attention</h2>\n\n<p>Mentioned in the IBM paper as the most important change, but it didn't work for us.</p>\n\n<h1>Ensemble</h1>\n\n<p>In 3 hours, we can do inference for 3 models with doc_stride=256. Luckily, n=3 happen to be the number of our best ensemble: adding a fourth model does not help anymore. The ensemble strategy is simply averaging the probability of each candidate span.</p>\n\n<p>Our 2 submissions consist of the following 5 single models:\na. wwm, stride=256, dev 62.4\nb. wwm, neg sampling, pre-tuned on squad, <strong>dev 64.7</strong> (long 69.5, short 57.8) - best single model\nc. wwm, neg sampling, max_contexts=200, dev 64.5\nd. wwm, neg sampling, stride=192, dev 63.8\ne. wwm, neg sampling, cased, dev 63.3</p>\n\n<p>sub1: ensemble of a,b,c, dev 66.8 (long 71.6, short 59.8), private LB 0.69\nsub2: ensemble of c,d,e, <strong>dev 67.0</strong> (long 71.6, short 59.9), private LB 0.69\n(note: these scores are after post-processing)</p>\n\n<h1>Post-process</h1>\n\n<h3>yes/no thresholds</h3>\n\n<p>These are tuned on dev set as well. If the yes/no logits in the <code>answer_type_logit</code> are over the thresholds, predict \"YES\"/\"NO\" regardless of the short span predictions. This gives 0.5 boost to dev F1.</p>\n\n<h3>max_contexts</h3>\n\n<p>Increase max_contexts from the default 48 to 100 or 80 can squeeze out another 0.3 F1 points, taking advantage of the leftover inference time within the 3 hour limit. For sub1 we did 100; for sub2 we only did 80 because generating features for the cased model <code>e</code> took a little more time.</p>\n\n<h1>Code</h1>\n\n<p>repo: <a href=\"https://github.com/boliu61/tf2qa\">https://github.com/boliu61/tf2qa</a>\ninference notebook and model weights: <a href=\"https://www.kaggle.com/boliu0/7th-place-submission\">https://www.kaggle.com/boliu0/7th-place-submission</a></p>",
      "rawMarkdown": "# Framework and hardware\nInitially we set out to try both TF 1.15 and TF 2.0. Unfortunately TF 2.0 pipeline's scores are significantly lower (we probably didn't figure out how to correctly use the TF 2.0 API), so all our top submissions and what is described below were done in TF 1.15. \n\nAll the experiments were done in Google cloud TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast and Google is generous enough to give us 5 TPUs, the experiment cycle is dramatically reduced.\n\n# Validation scheme and experiment setup\nWe noticed that the evaluation metric is not very stable on smaller validation set, for example better models on dev00 (1600 examples) may not be better on dev01. So we rely solely on the whole dev set (7830 examples) for validation (i.e. selecting checkpoints, selecting models, tuning thresholds, tuning ensemble weights). Same thing goes for the public LB. It only has 346 examples, even less stable than dev00.\n\nMost of our models are based on the official implementation of [joint-bert](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint), which is surprisingly hard to beat. For each variation to joint-bert, we ran 5 training sessions on TPU simultaneously, with different batch size and learning rate. We save checkpoints every 500 or 1000 steps. Usually bs=64 and lr=4e-5 for 1 epoch gives the best scores.\n\n# Variations to official joint-bert\nWe tried the following variations to joint-bert. The ideas for many of them are drawn from this IBM paper: https://arxiv.org/pdf/1909.05286.pdf\n1) and 2) are the most important. Adding them always help. They contribute to the best single model. 3) to 6) sometimes help, sometimes don't. They depend on each other. Nevertheless, training with these variations did produce many diverse models, which are good for ensembling. \n## 1) pre-trained weights\nOfficial joint-bert was trained from \"BERT-Large, Uncased\", but training from \"BERT-Large, Uncased (Whole Word Masking)\" will see a big boost.\n\nWe also tried fine-tuning joint-bert on Squad 2.0 before fine-tuning it on NQ.\n\n## 2) negative sampling\nOfficial joint-bert samples 2% negative examples in both answerable questions (i.e. the sliding windows that don't contain an answer) and unanswerable questions. As explained in the IBM paper, joint-bert tends to be overconfident for unanswerable questions, so 1% for answerable and 4% unanswerable seem to be better. We saw about 1 point increase in F1 doing negative sampling.\n\n## 3) `max_seq_length`, `doc_stride`\nDefault are 512 and 128 respectively. This means for answers not in the beginning of article, they appear about 4 times after pre-processing (in training the example is seen 4 times per epoch, in inference it's predicted 4 times, then the one with max logits is selected), which seems like an overkill.\n\nSo we changed doc_stride to 256 during inference, which doesn't affect score much but reduced inference time in half. \n\nFor training, we used default 128 as well 192 and 256. IBM paper claims 192 gives best results, but we didn't see much difference.\n\n## 4) max_contexts\nDefault is 48. There are some very very long wikepedia articles, so joint-bert only take first 48 paragraphs/tables/lists of each article. We tried different values like 100 and 200. Using a bigger value is tradeoff between more answer coverage v.s. more \"empty\" windows.\n\n## 5) sentence order shuffling\nAlso proposed in the IBM paper: shuffling all the sentences in the paragraph containing short answers. This is an augmentation method. \n\n## 6) cased\n`do_lower_case=False`\nFor this, we generated a new vocab file by adding all the NQ special tokens into the cased BERT vocab.\n\n## 7) Attention-over-attention\nMentioned in the IBM paper as the most important change, but it didn't work for us.\n\n# Ensemble\nIn 3 hours, we can do inference for 3 models with doc_stride=256. Luckily, n=3 happen to be the number of our best ensemble: adding a fourth model does not help anymore. The ensemble strategy is simply averaging the probability of each candidate span.\n\nOur 2 submissions consist of the following 5 single models:\na. wwm, stride=256, dev 62.4\nb. wwm, neg sampling, pre-tuned on squad, **dev 64.7** (long 69.5, short 57.8) - best single model\nc. wwm, neg sampling, max_contexts=200, dev 64.5\nd. wwm, neg sampling, stride=192, dev 63.8\ne. wwm, neg sampling, cased, dev 63.3\n\nsub1: ensemble of a,b,c, dev 66.8 (long 71.6, short 59.8), private LB 0.69\nsub2: ensemble of c,d,e, **dev 67.0** (long 71.6, short 59.9), private LB 0.69\n(note: these scores are after post-processing)\n\n# Post-process\n### yes/no thresholds\nThese are tuned on dev set as well. If the yes/no logits in the `answer_type_logit` are over the thresholds, predict \"YES\"/\"NO\" regardless of the short span predictions. This gives 0.5 boost to dev F1.\n\n### max_contexts\nIncrease max_contexts from the default 48 to 100 or 80 can squeeze out another 0.3 F1 points, taking advantage of the leftover inference time within the 3 hour limit. For sub1 we did 100; for sub2 we only did 80 because generating features for the cased model `e` took a little more time.\n\n# Code\nrepo: https://github.com/boliu61/tf2qa\ninference notebook and model weights: https://www.kaggle.com/boliu0/7th-place-submission",
      "votes": 52
    },
    {
      "id": 728686,
      "postDate": "2020-01-25T05:00:22.847Z",
      "content": "<p><strong>Update</strong>: code is available at <a href=\"https://github.com/boliu61/tf2qa\">https://github.com/boliu61/tf2qa</a></p>\n\n<p>inference notebook and model weights: <a href=\"https://www.kaggle.com/boliu0/7th-place-submission\">https://www.kaggle.com/boliu0/7th-place-submission</a></p>",
      "rawMarkdown": "**Update**: code is available at https://github.com/boliu61/tf2qa\n\ninference notebook and model weights: https://www.kaggle.com/boliu0/7th-place-submission",
      "votes": 1
    },
    {
      "id": 728501,
      "postDate": "2020-01-24T20:02:08.713Z",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 ",
      "votes": 1
    },
    {
      "id": 728364,
      "postDate": "2020-01-24T16:40:08.753Z",
      "content": "<p>Congratulations <a href=\"/boliu0\">@boliu0</a>, I really like the experiments you've conducted. Great write up and I will be waiting for the code. 😉 </p>",
      "rawMarkdown": "Congratulations @boliu0, I really like the experiments you've conducted. Great write up and I will be waiting for the code. 😉 ",
      "votes": 1
    },
    {
      "id": 727053,
      "postDate": "2020-01-23T12:17:02.920Z",
      "content": "<p>Thanks for the details explanation.\n\"7. Attention-over-attention didn't work\" mean that you had implemented it but did not give any improvement, or you are unable to implement it? </p>",
      "rawMarkdown": "Thanks for the details explanation.\n\"7. Attention-over-attention didn't work\" mean that you had implemented it but did not give any improvement, or you are unable to implement it? ",
      "votes": 1,
      "replies": [
        {
          "id": 727230,
          "postDate": "2020-01-23T14:47:40.083Z",
          "content": "<p>We implemented it, but it didn't improve the score.</p>",
          "rawMarkdown": "We implemented it, but it didn't improve the score.",
          "votes": 2
        },
        {
          "id": 727597,
          "postDate": "2020-01-23T21:18:03.077Z",
          "content": "<p>Would you mind to share to me the code with the Attention-over-attention implementation? I am interested to research further on this particular algorithm. </p>",
          "rawMarkdown": "Would you mind to share to me the code with the Attention-over-attention implementation? I am interested to research further on this particular algorithm. "
        },
        {
          "id": 728688,
          "postDate": "2020-01-25T05:02:13.997Z",
          "content": "<p>Sure. Our attempted implementation is at <a href=\"https://github.com/boliu61/tf2qa/blob/master/jb_train_tpu.py#L1054\">https://github.com/boliu61/tf2qa/blob/master/jb_train_tpu.py#L1054</a></p>",
          "rawMarkdown": "Sure. Our attempted implementation is at https://github.com/boliu61/tf2qa/blob/master/jb_train_tpu.py#L1054"
        }
      ]
    },
    {
      "id": 726689,
      "postDate": "2020-01-23T07:04:17.793Z",
      "content": "<p>Thanks for sharing, I learned a lot from your post.  </p>",
      "rawMarkdown": "Thanks for sharing, I learned a lot from your post.  ",
      "votes": 1
    },
    {
      "id": 726617,
      "postDate": "2020-01-23T06:03:54.097Z",
      "content": "<p>Congratulations Bo and thanks for sharing :)</p>",
      "rawMarkdown": "Congratulations Bo and thanks for sharing :)",
      "votes": 1
    },
    {
      "id": 726644,
      "postDate": "2020-01-23T06:27:59.360Z",
      "content": "<p>Congratulations\nThanks for Sharing your Approach &amp; Insights!! <a href=\"/boliu0\">@boliu0</a> </p>",
      "rawMarkdown": "Congratulations\nThanks for Sharing your Approach &amp; Insights!! @boliu0 ",
      "votes": -2
    },
    {
      "id": 739426,
      "postDate": "2020-02-07T20:37:36.873Z",
      "content": "<p><a href=\"/boliu0\">@boliu0</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>) </p>",
      "rawMarkdown": "@boliu0, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques) "
    },
    {
      "id": 726591,
      "postDate": "2020-01-23T05:38:41.050Z",
      "content": "<p>Congrats and thanks for sharing! We discovered this paper too late to really benefit from it.</p>",
      "rawMarkdown": "Congrats and thanks for sharing! We discovered this paper too late to really benefit from it.",
      "replies": [
        {
          "id": 726613,
          "postDate": "2020-01-23T06:02:54.640Z",
          "content": "<p>Thanks and congrats to you too. Looking forward to your solution. </p>\n\n<p>We may have discovered that paper (which claimed SOTA) too early and spent too much time on it 😂</p>",
          "rawMarkdown": "Thanks and congrats to you too. Looking forward to your solution. \n\nWe may have discovered that paper (which claimed SOTA) too early and spent too much time on it 😂",
          "votes": 5
        }
      ]
    },
    {
      "id": 794593,
      "postDate": "2020-04-01T23:29:40.910Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 726589,
      "postDate": "2020-01-23T05:34:33.500Z",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "rawMarkdown": "Congratulations and thanks for sharing!",
      "votes": 1
    },
    {
      "id": 733307,
      "postDate": "2020-01-31T00:18:55.643Z",
      "content": "<p>Congrats and thank you for sharing!</p>",
      "rawMarkdown": "Congrats and thank you for sharing!"
    }
  ],
  "comments": [
    {
      "id": 728686,
      "author_name": "Bo",
      "author_url": "",
      "post_date": "2020-01-25T05:00:22.847000",
      "content": "<p><strong>Update</strong>: code is available at <a href=\"https://github.com/boliu61/tf2qa\">https://github.com/boliu61/tf2qa</a></p>\n\n<p>inference notebook and model weights: <a href=\"https://www.kaggle.com/boliu0/7th-place-submission\">https://www.kaggle.com/boliu0/7th-place-submission</a></p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 728501,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-01-24T20:02:08.713000",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 728364,
      "author_name": "Mourad",
      "author_url": "",
      "post_date": "2020-01-24T16:40:08.753000",
      "content": "<p>Congratulations <a href=\"/boliu0\">@boliu0</a>, I really like the experiments you've conducted. Great write up and I will be waiting for the code. 😉 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 727053,
      "author_name": "Chew Kok Wah",
      "author_url": "",
      "post_date": "2020-01-23T12:17:02.920000",
      "content": "<p>Thanks for the details explanation.\n\"7. Attention-over-attention didn't work\" mean that you had implemented it but did not give any improvement, or you are unable to implement it? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 727230,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-01-23T14:47:40.083000",
          "content": "<p>We implemented it, but it didn't improve the score.</p>",
          "votes": 2,
          "replies": []
        },
        {
          "id": 727597,
          "author_name": "Chew Kok Wah",
          "author_url": "",
          "post_date": "2020-01-23T21:18:03.077000",
          "content": "<p>Would you mind to share to me the code with the Attention-over-attention implementation? I am interested to research further on this particular algorithm. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728688,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-01-25T05:02:13.997000",
          "content": "<p>Sure. Our attempted implementation is at <a href=\"https://github.com/boliu61/tf2qa/blob/master/jb_train_tpu.py#L1054\">https://github.com/boliu61/tf2qa/blob/master/jb_train_tpu.py#L1054</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 726689,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2020-01-23T07:04:17.793000",
      "content": "<p>Thanks for sharing, I learned a lot from your post.  </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 726617,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2020-01-23T06:03:54.097000",
      "content": "<p>Congratulations Bo and thanks for sharing :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 726644,
      "author_name": "Ailurophile",
      "author_url": "",
      "post_date": "2020-01-23T06:27:59.360000",
      "content": "<p>Congratulations\nThanks for Sharing your Approach &amp; Insights!! <a href=\"/boliu0\">@boliu0</a> </p>",
      "votes": -2,
      "replies": []
    },
    {
      "id": 739426,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-07T20:37:36.873000",
      "content": "<p><a href=\"/boliu0\">@boliu0</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>) </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 726591,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-01-23T05:38:41.050000",
      "content": "<p>Congrats and thanks for sharing! We discovered this paper too late to really benefit from it.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 726613,
          "author_name": "Bo",
          "author_url": "",
          "post_date": "2020-01-23T06:02:54.640000",
          "content": "<p>Thanks and congrats to you too. Looking forward to your solution. </p>\n\n<p>We may have discovered that paper (which claimed SOTA) too early and spent too much time on it 😂</p>",
          "votes": 5,
          "replies": []
        }
      ]
    },
    {
      "id": 794593,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-04-01T23:29:40.910000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 726589,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "2020-01-23T05:34:33.500000",
      "content": "<p>Congratulations and thanks for sharing!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 733307,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-31T00:18:55.643000",
      "content": "<p>Congrats and thank you for sharing!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "726570": "# Framework and hardware\nInitially we set out to try both TF 1.15 and TF 2.0. Unfortunately TF 2.0 pipeline's scores are significantly lower (we probably didn't figure out how to correctly use the TF 2.0 API), so all our top submissions and what is described below were done in TF 1.15. \n\nAll the experiments were done in Google cloud TPUs. This is my first time seriously using TPUs and I have to say, it feels so good. Because they are so fast and Google is generous enough to give us 5 TPUs, the experiment cycle is dramatically reduced.\n\n# Validation scheme and experiment setup\nWe noticed that the evaluation metric is not very stable on smaller validation set, for example better models on dev00 (1600 examples) may not be better on dev01. So we rely solely on the whole dev set (7830 examples) for validation (i.e. selecting checkpoints, selecting models, tuning thresholds, tuning ensemble weights). Same thing goes for the public LB. It only has 346 examples, even less stable than dev00.\n\nMost of our models are based on the official implementation of [joint-bert](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint), which is surprisingly hard to beat. For each variation to joint-bert, we ran 5 training sessions on TPU simultaneously, with different batch size and learning rate. We save checkpoints every 500 or 1000 steps. Usually bs=64 and lr=4e-5 for 1 epoch gives the best scores.\n\n# Variations to official joint-bert\nWe tried the following variations to joint-bert. The ideas for many of them are drawn from this IBM paper: https://arxiv.org/pdf/1909.05286.pdf\n1) and 2) are the most important. Adding them always help. They contribute to the best single model. 3) to 6) sometimes help, sometimes don't. They depend on each other. Nevertheless, training with these variations did produce many diverse models, which are good for ensembling. \n## 1) pre-trained weights\nOfficial joint-bert was trained from \"BERT-Large, Uncased\", but training from \"BERT-Large, Uncased (Whole Word Masking)\" will see a big boost.\n\nWe also tried fine-tuning joint-bert on Squad 2.0 before fine-tuning it on NQ.\n\n## 2) negative sampling\nOfficial joint-bert samples 2% negative examples in both answerable questions (i.e. the sliding windows that don't contain an answer) and unanswerable questions. As explained in the IBM paper, joint-bert tends to be overconfident for unanswerable questions, so 1% for answerable and 4% unanswerable seem to be better. We saw about 1 point increase in F1 doing negative sampling.\n\n## 3) `max_seq_length`, `doc_stride`\nDefault are 512 and 128 respectively. This means for answers not in the beginning of article, they appear about 4 times after pre-processing (in training the example is seen 4 times per epoch, in inference it's predicted 4 times, then the one with max logits is selected), which seems like an overkill.\n\nSo we changed doc_stride to 256 during inference, which doesn't affect score much but reduced inference time in half. \n\nFor training, we used default 128 as well 192 and 256. IBM paper claims 192 gives best results, but we didn't see much difference.\n\n## 4) max_contexts\nDefault is 48. There are some very very long wikepedia articles, so joint-bert only take first 48 paragraphs/tables/lists of each article. We tried different values like 100 and 200. Using a bigger value is tradeoff between more answer coverage v.s. more \"empty\" windows.\n\n## 5) sentence order shuffling\nAlso proposed in the IBM paper: shuffling all the sentences in the paragraph containing short answers. This is an augmentation method. \n\n## 6) cased\n`do_lower_case=False`\nFor this, we generated a new vocab file by adding all the NQ special tokens into the cased BERT vocab.\n\n## 7) Attention-over-attention\nMentioned in the IBM paper as the most important change, but it didn't work for us.\n\n# Ensemble\nIn 3 hours, we can do inference for 3 models with doc_stride=256. Luckily, n=3 happen to be the number of our best ensemble: adding a fourth model does not help anymore. The ensemble strategy is simply averaging the probability of each candidate span.\n\nOur 2 submissions consist of the following 5 single models:\na. wwm, stride=256, dev 62.4\nb. wwm, neg sampling, pre-tuned on squad, **dev 64.7** (long 69.5, short 57.8) - best single model\nc. wwm, neg sampling, max_contexts=200, dev 64.5\nd. wwm, neg sampling, stride=192, dev 63.8\ne. wwm, neg sampling, cased, dev 63.3\n\nsub1: ensemble of a,b,c, dev 66.8 (long 71.6, short 59.8), private LB 0.69\nsub2: ensemble of c,d,e, **dev 67.0** (long 71.6, short 59.9), private LB 0.69\n(note: these scores are after post-processing)\n\n# Post-process\n### yes/no thresholds\nThese are tuned on dev set as well. If the yes/no logits in the `answer_type_logit` are over the thresholds, predict \"YES\"/\"NO\" regardless of the short span predictions. This gives 0.5 boost to dev F1.\n\n### max_contexts\nIncrease max_contexts from the default 48 to 100 or 80 can squeeze out another 0.3 F1 points, taking advantage of the leftover inference time within the 3 hour limit. For sub1 we did 100; for sub2 we only did 80 because generating features for the cased model `e` took a little more time.\n\n# Code\nrepo: https://github.com/boliu61/tf2qa\ninference notebook and model weights: https://www.kaggle.com/boliu0/7th-place-submission",
    "728686": "**Update**: code is available at https://github.com/boliu61/tf2qa\n\ninference notebook and model weights: https://www.kaggle.com/boliu0/7th-place-submission",
    "728501": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 ",
    "728364": "Congratulations @boliu0, I really like the experiments you've conducted. Great write up and I will be waiting for the code. 😉 ",
    "727053": "Thanks for the details explanation.\n\"7. Attention-over-attention didn't work\" mean that you had implemented it but did not give any improvement, or you are unable to implement it? ",
    "726689": "Thanks for sharing, I learned a lot from your post.  ",
    "726617": "Congratulations Bo and thanks for sharing :)",
    "726644": "Congratulations\nThanks for Sharing your Approach &amp; Insights!! @boliu0 ",
    "739426": "@boliu0, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques) ",
    "726591": "Congrats and thanks for sharing! We discovered this paper too late to really benefit from it.",
    "794593": "",
    "726589": "Congratulations and thanks for sharing!",
    "733307": "Congrats and thank you for sharing!"
  }
}