{
  "id": 128140,
  "title": "21th place solution, puzzlingly shaking from public LB 3th",
  "url": "/competitions/tensorflow2-question-answering/writeups/21th-place-solution-puzzlingly-shaking-from-public",
  "author_name": "",
  "post_date": "2020-01-29T07:25:51.614021200Z",
  "votes": 5,
  "comment_count": 6,
  "views": 0,
  "content": "<p><a href=\"https://github.com/mikelkl/TF2-QA\">Code for Solution</a></p>\n\n<p>Thanks kaggle for holding this wonderful competition. Big thanks to my awesome teammate <a href=\"https://www.kaggle.com/ewrfcas\">@ewrfcas</a> and <a href=\"https://www.kaggle.com/leolemon214\">@leolemon214</a>. It's little pity for dropping from public LB 3th place,  expect fetching gold medal at future. </p>\n\n<p><strong>We carefully read other top solutions, and still puzzling for the 4% drop of private LB, can someone help to figure out the reason for this shaking?</strong></p>\n\n<p>Below are valid part of our solution, all the following experiments are mainly performed on offline <strong>dev containing 1600 examples</strong>, and some results have been verified in public LB.</p>\n\n<h2>1. Preprocessing</h2>\n\n<p>| No   | Technique                         | Pros                                                         | Cons                              | Effect                                |\n| ---- | --------------------------------- | ------------------------------------------------------------ | --------------------------------- | ------------------------------------- |\n| 1    | TF-IDF paragraph selection       | Shorten doc resulting faster inference speed and better accuracy | May loss some context information | - dev f1 +1.8%,<br>- public LB f1 -1% |\n| 2    | Sample negative features till 1:1 | Balance pos and neg                                          | Cause longer training time        | dev f1 +2.248%                        |\n| 3    | Multi-process preprocessing       | Accelerate preprocessing, especially on training data        | Require multi-core CPU            | xN faster (with N processes)          |</p>\n\n<h2>2. Modeling</h2>\n\n<p>| No   | Model Architecture                                 | Idea                                                         | Performance          |\n| ---- | -------------------------------------------------- | ------------------------------------------------------------ | -------------------- |\n| 1    | Roberta-Large joint with long/short span extractor | 1. Jointly model:<br>- answer type<br>- long span<br>- short span<br>2. Output topk start/end logits/index | dev f1 63.986%       |\n| 2    | Albert-xxlarge joint with short span extractor     | Jointly model:<br>- answer type<br>- short span            | def short-f1 69.364% |</p>\n\n<p>All of above model architectures were pretrained on SQuAD dataset by ourselves.</p>\n\n<h2>3. Trick</h2>\n\n<p>| No   | Trick                                                        | Effect                             |\n| ---- | ------------------------------------------------------------ | ---------------------------------- |\n| 1    | If answer_type is yes/no, output yes/no rather than short span | public LB f1 +6%                   |\n| 2    | 1. If answer_type is short, output long span and short span<br>2. If answer_type is long, output long span only<br>3. If answer_type is none, output neither long span nor short span | public LB f1 +8%                   |\n| 3    | Choose the best long/short answer pair from topk * topk kind of long/short answer combinations | dev f1 +0.435%                     |\n| 4    | <code>long_score = summary.long_span_score - summary.long_cls_score - summary.answer_type_logits[0]</code><br><code>short_score = summary.short_span_score - summary.short_cls_score - summary.answer_type_logits[0]</code> | - dev f1 +2.12%<br>- public LB +2% |\n| 5    | Increase long  [CLS] logits multiplier threshold to increase null long answer | dev long-f1 +3.491%                |\n| 6    | Decrease short answer_type logits divisor threshold to increase null short answer | dev short-f1 ?                     |</p>\n\n<h2>4. Ensemble</h2>\n\n<p>| No   | Idea                                                         | Effect                                                       |\n| ---- | ------------------------------------------------------------ | ------------------------------------------------------------ |\n| 1    | For long answer, We vote long answers of 2 <code>Roberta-Large joint with long/short span extractor</code> models | dev long-f1 +3.341%                                          |\n| 2    | For short answer, use step 1 result to locate predicted long answer candidate as input,  We vote short answers of 2 <code>Roberta-Large joint with long/short span extractor</code> models and 4 <code>Albert-xxlarge joint with short span extractor</code> models | - dev short-f1 +2.842% <br>- dev f1 67.569%,  +2.635% <br>- public LB 71%, +5%<br>- private LB 67% |</p>\n\n<p><a href=\"https://github.com/mikelkl/TF2-QA\">Code for Solution</a></p>",
  "messages": [
    {
      "id": "731886",
      "postDate": "01/29/2020 07:25:51",
      "content": "<p><a href=\"https://github.com/mikelkl/TF2-QA\">Code for Solution</a></p>\n\n<p>Thanks kaggle for holding this wonderful competition. Big thanks to my awesome teammate <a href=\"https://www.kaggle.com/ewrfcas\">@ewrfcas</a> and <a href=\"https://www.kaggle.com/leolemon214\">@leolemon214</a>. It's little pity for dropping from public LB 3th place,  expect fetching gold medal at future. </p>\n\n<p><strong>We carefully read other top solutions, and still puzzling for the 4% drop of private LB, can someone help to figure out the reason for this shaking?</strong></p>\n\n<p>Below are valid part of our solution, all the following experiments are mainly performed on offline <strong>dev containing 1600 examples</strong>, and some results have been verified in public LB.</p>\n\n<h2>1. Preprocessing</h2>\n\n<p>| No   | Technique                         | Pros                                                         | Cons                              | Effect                                |\n| ---- | --------------------------------- | ------------------------------------------------------------ | --------------------------------- | ------------------------------------- |\n| 1    | TF-IDF paragraph selection       | Shorten doc resulting faster inference speed and better accuracy | May loss some context information | - dev f1 +1.8%,<br>- public LB f1 -1% |\n| 2    | Sample negative features till 1:1 | Balance pos and neg                                          | Cause longer training time        | dev f1 +2.248%                        |\n| 3    | Multi-process preprocessing       | Accelerate preprocessing, especially on training data        | Require multi-core CPU            | xN faster (with N processes)          |</p>\n\n<h2>2. Modeling</h2>\n\n<p>| No   | Model Architecture                                 | Idea                                                         | Performance          |\n| ---- | -------------------------------------------------- | ------------------------------------------------------------ | -------------------- |\n| 1    | Roberta-Large joint with long/short span extractor | 1. Jointly model:<br>- answer type<br>- long span<br>- short span<br>2. Output topk start/end logits/index | dev f1 63.986%       |\n| 2    | Albert-xxlarge joint with short span extractor     | Jointly model:<br>- answer type<br>- short span            | def short-f1 69.364% |</p>\n\n<p>All of above model architectures were pretrained on SQuAD dataset by ourselves.</p>\n\n<h2>3. Trick</h2>\n\n<p>| No   | Trick                                                        | Effect                             |\n| ---- | ------------------------------------------------------------ | ---------------------------------- |\n| 1    | If answer_type is yes/no, output yes/no rather than short span | public LB f1 +6%                   |\n| 2    | 1. If answer_type is short, output long span and short span<br>2. If answer_type is long, output long span only<br>3. If answer_type is none, output neither long span nor short span | public LB f1 +8%                   |\n| 3    | Choose the best long/short answer pair from topk * topk kind of long/short answer combinations | dev f1 +0.435%                     |\n| 4    | <code>long_score = summary.long_span_score - summary.long_cls_score - summary.answer_type_logits[0]</code><br><code>short_score = summary.short_span_score - summary.short_cls_score - summary.answer_type_logits[0]</code> | - dev f1 +2.12%<br>- public LB +2% |\n| 5    | Increase long  [CLS] logits multiplier threshold to increase null long answer | dev long-f1 +3.491%                |\n| 6    | Decrease short answer_type logits divisor threshold to increase null short answer | dev short-f1 ?                     |</p>\n\n<h2>4. Ensemble</h2>\n\n<p>| No   | Idea                                                         | Effect                                                       |\n| ---- | ------------------------------------------------------------ | ------------------------------------------------------------ |\n| 1    | For long answer, We vote long answers of 2 <code>Roberta-Large joint with long/short span extractor</code> models | dev long-f1 +3.341%                                          |\n| 2    | For short answer, use step 1 result to locate predicted long answer candidate as input,  We vote short answers of 2 <code>Roberta-Large joint with long/short span extractor</code> models and 4 <code>Albert-xxlarge joint with short span extractor</code> models | - dev short-f1 +2.842% <br>- dev f1 67.569%,  +2.635% <br>- public LB 71%, +5%<br>- private LB 67% |</p>\n\n<p><a href=\"https://github.com/mikelkl/TF2-QA\">Code for Solution</a></p>",
      "rawMarkdown": "[Code for Solution](https://github.com/mikelkl/TF2-QA)\n\nThanks kaggle for holding this wonderful competition. Big thanks to my awesome teammate [@ewrfcas](https://www.kaggle.com/ewrfcas) and [@leolemon214](https://www.kaggle.com/leolemon214). It's little pity for dropping from public LB 3th place,  expect fetching gold medal at future. \n\n**We carefully read other top solutions, and still puzzling for the 4% drop of private LB, can someone help to figure out the reason for this shaking?**\n\nBelow are valid part of our solution, all the following experiments are mainly performed on offline **dev containing 1600 examples**, and some results have been verified in public LB.\n\n## 1. Preprocessing\n\n| No   | Technique                         | Pros                                                         | Cons                              | Effect                                |\n| ---- | --------------------------------- | ------------------------------------------------------------ | --------------------------------- | ------------------------------------- |\n| 1    | TF\\-IDF paragraph selection       | Shorten doc resulting faster inference speed and better accuracy | May loss some context information | - dev f1 +1.8%,<br>- public LB f1 -1% |\n| 2    | Sample negative features till 1:1 | Balance pos and neg                                          | Cause longer training time        | dev f1 +2.248%                        |\n| 3    | Multi-process preprocessing       | Accelerate preprocessing, especially on training data        | Require multi-core CPU            | xN faster (with N processes)          |\n\n## 2. Modeling\n\n| No   | Model Architecture                                 | Idea                                                         | Performance          |\n| ---- | -------------------------------------------------- | ------------------------------------------------------------ | -------------------- |\n| 1    | Roberta-Large joint with long/short span extractor | 1. Jointly model:<br>- answer type<br>- long span<br>- short span<br>2. Output topk start/end logits/index | dev f1 63.986%       |\n| 2    | Albert-xxlarge joint with short span extractor     | Jointly model:<br>- answer type<br>- short span            | def short-f1 69.364% |\n\nAll of above model architectures were pretrained on SQuAD dataset by ourselves.\n\n## 3. Trick\n\n| No   | Trick                                                        | Effect                             |\n| ---- | ------------------------------------------------------------ | ---------------------------------- |\n| 1    | If answer_type is yes/no, output yes/no rather than short span | public LB f1 +6%                   |\n| 2    | 1. If answer_type is short, output long span and short span<br>2. If answer_type is long, output long span only<br>3. If answer_type is none, output neither long span nor short span | public LB f1 +8%                   |\n| 3    | Choose the best long/short answer pair from topk * topk kind of long/short answer combinations | dev f1 +0.435%                     |\n| 4    | `long_score = summary.long_span_score - summary.long_cls_score - summary.answer_type_logits[0]`<br>`short_score = summary.short_span_score - summary.short_cls_score - summary.answer_type_logits[0]` | - dev f1 +2.12%<br>- public LB +2% |\n| 5    | Increase long  [CLS] logits multiplier threshold to increase null long answer | dev long-f1 +3.491%                |\n| 6    | Decrease short answer_type logits divisor threshold to increase null short answer | dev short-f1 ?                     |\n\n## 4. Ensemble\n\n| No   | Idea                                                         | Effect                                                       |\n| ---- | ------------------------------------------------------------ | ------------------------------------------------------------ |\n| 1    | For long answer, We vote long answers of 2 `Roberta-Large joint with long/short span extractor` models | dev long-f1 +3.341%                                          |\n| 2    | For short answer, use step 1 result to locate predicted long answer candidate as input,  We vote short answers of 2 `Roberta-Large joint with long/short span extractor` models and 4 `Albert-xxlarge joint with short span extractor` models | - dev short-f1 +2.842% <br>- dev f1 67.569%,  +2.635% <br>- public LB 71%, +5%<br>- private LB 67% |\n\n\n[Code for Solution](https://github.com/mikelkl/TF2-QA)",
      "votes": null
    },
    {
      "id": "731978",
      "postDate": "01/29/2020 10:25:54",
      "content": "<p>Wow! You guys really put the effort on it! Congratulations.</p>",
      "rawMarkdown": "Wow! You guys really put the effort on it! Congratulations.",
      "votes": null
    },
    {
      "id": "731983",
      "postDate": "01/29/2020 10:32:29",
      "content": "<p>Congratulations!!\nThanks for sharing your Approach &amp; Code <a href=\"/mikelkl\">@mikelkl</a> </p>",
      "rawMarkdown": "Congratulations!!\nThanks for sharing your Approach &amp; Code @mikelkl",
      "votes": null
    },
    {
      "id": "732221",
      "postDate": "01/29/2020 15:29:57",
      "content": "<p>Thanks for sharing!!😄 👍 </p>",
      "rawMarkdown": "Thanks for sharing!!😄 👍",
      "votes": null
    },
    {
      "id": "732361",
      "postDate": "01/29/2020 18:11:13",
      "content": "<p><a href=\"/mikelkl\">@mikelkl</a> My guess is that you over-relied on the answer type logits in postprocessing.. refer to my comments in my solution\n<a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127241\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127241</a></p>\n\n<blockquote>\n  <p>I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n  1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n  2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\n  Based on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing (since it's basically overfitting to Dev or public LB).</p>\n</blockquote>",
      "rawMarkdown": "mikelkl My guess is that you over-relied on the answer type logits in postprocessing.. refer to my comments in my solution\nhttps://www.kaggle.com/c/tensorflow2-question-answering/discussion/127241\n\n&gt; I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\nBased on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing (since it's basically overfitting to Dev or public LB).",
      "votes": null
    },
    {
      "id": "732901",
      "postDate": "01/30/2020 12:47:48",
      "content": "<p>Thks for ur reply, very insightful idea. We did touch answer type logits at least 3 times on postprocessing stage, will check it. If so, do u think train an extra yes/no model with balanced resampling data is a good idea?</p>",
      "rawMarkdown": "Thks for ur reply, very insightful idea. We did touch answer type logits at least 3 times on postprocessing stage, will check it. If so, do u think train an extra yes/no model with balanced resampling data is a good idea?",
      "votes": null
    },
    {
      "id": "733224",
      "postDate": "01/30/2020 21:01:30",
      "content": "<p><a href=\"/mikelkl\">@mikelkl</a> That was exactly what I tried - an extra YES/NO verifier trained with ALBERT-xxlarge with only the YES/NO questions. Slight improvement on Dev (~1 pt) but not on private test. I'd say it's not the most effective way due to the lack of data (1k to 10k samples). Might be better to make use of some external knowledge base rather than NQ alone.</p>",
      "rawMarkdown": "mikelkl That was exactly what I tried - an extra YES/NO verifier trained with ALBERT-xxlarge with only the YES/NO questions. Slight improvement on Dev (~1 pt) but not on private test. I'd say it's not the most effective way due to the lack of data (1k to 10k samples). Might be better to make use of some external knowledge base rather than NQ alone.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 731978,
      "author_name": "brunhs",
      "author_url": "",
      "post_date": "01/29/2020 10:25:54",
      "content": "<p>Wow! You guys really put the effort on it! Congratulations.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 731983,
      "author_name": "veeralakrishna",
      "author_url": "",
      "post_date": "01/29/2020 10:32:29",
      "content": "<p>Congratulations!!\nThanks for sharing your Approach &amp; Code <a href=\"/mikelkl\">@mikelkl</a> </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 732221,
      "author_name": "mashlyn",
      "author_url": "",
      "post_date": "01/29/2020 15:29:57",
      "content": "<p>Thanks for sharing!!😄 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 732361,
      "author_name": "siriuself",
      "author_url": "",
      "post_date": "01/29/2020 18:11:13",
      "content": "<p><a href=\"/mikelkl\">@mikelkl</a> My guess is that you over-relied on the answer type logits in postprocessing.. refer to my comments in my solution\n<a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127241\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/127241</a></p>\n\n<blockquote>\n  <p>I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n  1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n  2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\n  Based on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing (since it's basically overfitting to Dev or public LB).</p>\n</blockquote>",
      "votes": null,
      "replies": [
        {
          "id": 732901,
          "author_name": "mikelkl",
          "author_url": "",
          "post_date": "01/30/2020 12:47:48",
          "content": "<p>Thks for ur reply, very insightful idea. We did touch answer type logits at least 3 times on postprocessing stage, will check it. If so, do u think train an extra yes/no model with balanced resampling data is a good idea?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 733224,
          "author_name": "siriuself",
          "author_url": "",
          "post_date": "01/30/2020 21:01:30",
          "content": "<p><a href=\"/mikelkl\">@mikelkl</a> That was exactly what I tried - an extra YES/NO verifier trained with ALBERT-xxlarge with only the YES/NO questions. Slight improvement on Dev (~1 pt) but not on private test. I'd say it's not the most effective way due to the lack of data (1k to 10k samples). Might be better to make use of some external knowledge base rather than NQ alone.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "731886": "[Code for Solution](https://github.com/mikelkl/TF2-QA)\n\nThanks kaggle for holding this wonderful competition. Big thanks to my awesome teammate [@ewrfcas](https://www.kaggle.com/ewrfcas) and [@leolemon214](https://www.kaggle.com/leolemon214). It's little pity for dropping from public LB 3th place,  expect fetching gold medal at future. \n\n**We carefully read other top solutions, and still puzzling for the 4% drop of private LB, can someone help to figure out the reason for this shaking?**\n\nBelow are valid part of our solution, all the following experiments are mainly performed on offline **dev containing 1600 examples**, and some results have been verified in public LB.\n\n## 1. Preprocessing\n\n| No   | Technique                         | Pros                                                         | Cons                              | Effect                                |\n| ---- | --------------------------------- | ------------------------------------------------------------ | --------------------------------- | ------------------------------------- |\n| 1    | TF\\-IDF paragraph selection       | Shorten doc resulting faster inference speed and better accuracy | May loss some context information | - dev f1 +1.8%,<br>- public LB f1 -1% |\n| 2    | Sample negative features till 1:1 | Balance pos and neg                                          | Cause longer training time        | dev f1 +2.248%                        |\n| 3    | Multi-process preprocessing       | Accelerate preprocessing, especially on training data        | Require multi-core CPU            | xN faster (with N processes)          |\n\n## 2. Modeling\n\n| No   | Model Architecture                                 | Idea                                                         | Performance          |\n| ---- | -------------------------------------------------- | ------------------------------------------------------------ | -------------------- |\n| 1    | Roberta-Large joint with long/short span extractor | 1. Jointly model:<br>- answer type<br>- long span<br>- short span<br>2. Output topk start/end logits/index | dev f1 63.986%       |\n| 2    | Albert-xxlarge joint with short span extractor     | Jointly model:<br>- answer type<br>- short span            | def short-f1 69.364% |\n\nAll of above model architectures were pretrained on SQuAD dataset by ourselves.\n\n## 3. Trick\n\n| No   | Trick                                                        | Effect                             |\n| ---- | ------------------------------------------------------------ | ---------------------------------- |\n| 1    | If answer_type is yes/no, output yes/no rather than short span | public LB f1 +6%                   |\n| 2    | 1. If answer_type is short, output long span and short span<br>2. If answer_type is long, output long span only<br>3. If answer_type is none, output neither long span nor short span | public LB f1 +8%                   |\n| 3    | Choose the best long/short answer pair from topk * topk kind of long/short answer combinations | dev f1 +0.435%                     |\n| 4    | `long_score = summary.long_span_score - summary.long_cls_score - summary.answer_type_logits[0]`<br>`short_score = summary.short_span_score - summary.short_cls_score - summary.answer_type_logits[0]` | - dev f1 +2.12%<br>- public LB +2% |\n| 5    | Increase long  [CLS] logits multiplier threshold to increase null long answer | dev long-f1 +3.491%                |\n| 6    | Decrease short answer_type logits divisor threshold to increase null short answer | dev short-f1 ?                     |\n\n## 4. Ensemble\n\n| No   | Idea                                                         | Effect                                                       |\n| ---- | ------------------------------------------------------------ | ------------------------------------------------------------ |\n| 1    | For long answer, We vote long answers of 2 `Roberta-Large joint with long/short span extractor` models | dev long-f1 +3.341%                                          |\n| 2    | For short answer, use step 1 result to locate predicted long answer candidate as input,  We vote short answers of 2 `Roberta-Large joint with long/short span extractor` models and 4 `Albert-xxlarge joint with short span extractor` models | - dev short-f1 +2.842% <br>- dev f1 67.569%,  +2.635% <br>- public LB 71%, +5%<br>- private LB 67% |\n\n\n[Code for Solution](https://github.com/mikelkl/TF2-QA)",
    "731978": "Wow! You guys really put the effort on it! Congratulations.",
    "731983": "Congratulations!!\nThanks for sharing your Approach &amp; Code @mikelkl",
    "732221": "Thanks for sharing!!😄 👍",
    "732361": "mikelkl My guess is that you over-relied on the answer type logits in postprocessing.. refer to my comments in my solution\nhttps://www.kaggle.com/c/tensorflow2-question-answering/discussion/127241\n\n&gt; I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\nBased on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing (since it's basically overfitting to Dev or public LB).",
    "732901": "Thks for ur reply, very insightful idea. We did touch answer type logits at least 3 times on postprocessing stage, will check it. If so, do u think train an extra yes/no model with balanced resampling data is a good idea?",
    "733224": "mikelkl That was exactly what I tried - an extra YES/NO verifier trained with ALBERT-xxlarge with only the YES/NO questions. Slight improvement on Dev (~1 pt) but not on private test. I'd say it's not the most effective way due to the lack of data (1k to 10k samples). Might be better to make use of some external knowledge base rather than NQ alone."
  },
  "source": "meta"
}