{
  "id": 127241,
  "title": "17th Place solution [bert-disjoint] kernel + all utility scripts",
  "url": "/competitions/tensorflow2-question-answering/writeups/siriuself-17th-place-solution-bert-disjoint-kernel",
  "author_name": "",
  "post_date": "2020-01-23T12:59:15.330Z",
  "votes": 15,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Congrats to all winners!</p>\n\n<p>I've made available my solution (public 0.65 private 0.67).\n<a href=\"https://www.kaggle.com/siriuself/tf-qa-wwm-verifier-forked\">https://www.kaggle.com/siriuself/tf-qa-wwm-verifier-forked</a></p>\n\n<p><a href=\"https://www.kaggle.com/siriuself/bert-disjoint-fn-builder\">https://www.kaggle.com/siriuself/bert-disjoint-fn-builder</a>\n<a href=\"https://www.kaggle.com/siriuself/bert-disjoint-modeling\">https://www.kaggle.com/siriuself/bert-disjoint-modeling</a>\n<a href=\"https://www.kaggle.com/siriuself/bert-disjoint-utils\">https://www.kaggle.com/siriuself/bert-disjoint-utils</a>\n<a href=\"https://www.kaggle.com/siriuself/albert-yes-no-fn-builder\">https://www.kaggle.com/siriuself/albert-yes-no-fn-builder</a>\n<a href=\"https://www.kaggle.com/siriuself/albert-yes-no-modeling\">https://www.kaggle.com/siriuself/albert-yes-no-modeling</a>\n<a href=\"https://www.kaggle.com/siriuself/albert-yes-no-utils\">https://www.kaggle.com/siriuself/albert-yes-no-utils</a>\n<a href=\"https://www.kaggle.com/siriuself/tokenization\">https://www.kaggle.com/siriuself/tokenization</a>\n<a href=\"https://www.kaggle.com/siriuself/create-submission\">https://www.kaggle.com/siriuself/create-submission</a></p>\n\n<p>My model is simple, BERT large whole-word-masking uncased, retrained using the start/end logit loss only without the answer type loss. Using provided nq train tf record with the following setting:\n<strong>batch_size</strong>: 32\n<strong>epoch</strong>: 2\n<strong>alpha</strong>: 2e-5, but use ckpt-15000 (so stop at around 1 epoch)\nNote that because of learning rate warmup/decay, this is different from training with 2e-5 for 1 epoch</p>\n\n<p>Then totally disregard answer type classification (since I don't have it), and rely on threshold setting for long and short questions, tuned on the dev set with nq_eval. And yes I let go all the YES/NO questions.</p>\n\n<p>I did try adding an ALBERT xxlarge yes/no verifier after the BERT stage, which seemed to improve for like 1 pt on dev, but apparently not on the LBs somehow. My kernel includes the ALBERT part too. There isn't much insight in the utility scripts, except the modifications I made in order to restore the checkpoint into some contrib layers forced to be re-written in keras (e.g. LayerNormalization). It was a nightmare..</p>\n\n<p><em><strong></strong></em><strong>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>*</strong><em>some reflections/insights</em><strong><em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**</strong>\nI started the competition way too late and was on the wrong ALBERT-taking-too-long direction for quite a while, ending up with little time for tuning. I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\nBased on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing.</p>\n\n<p>Let me know if you have similar/opposite findings. It's just my feeling anyway, with a bit of confirmation of my score, obtained by doing nothing other than removing the answer type loss in training. This is almost higher than my ALBERT-xxlarge-joint model too.. I almost felt like doing more experiments and writing a paper on the classification power reduction, but guess it's too trivial.</p>\n\n<p>Lastly, we are hiring intermediate/senior NLP engineer/researcher/scientist, with possibility of sponsorship to come to our Canadian headquarter in Waterloo, Ontario for strong candidates. Bilingualism in English and Mandarin is a plus.</p>",
  "messages": [
    {
      "id": "726450",
      "postDate": "01/23/2020 02:34:02",
      "content": "<p>Congrats to all winners!</p>\n\n<p>I've made available my solution (public 0.65 private 0.67).\n<a href=\"https://www.kaggle.com/siriuself/tf-qa-wwm-verifier-forked\">https://www.kaggle.com/siriuself/tf-qa-wwm-verifier-forked</a></p>\n\n<p><a href=\"https://www.kaggle.com/siriuself/bert-disjoint-fn-builder\">https://www.kaggle.com/siriuself/bert-disjoint-fn-builder</a>\n<a href=\"https://www.kaggle.com/siriuself/bert-disjoint-modeling\">https://www.kaggle.com/siriuself/bert-disjoint-modeling</a>\n<a href=\"https://www.kaggle.com/siriuself/bert-disjoint-utils\">https://www.kaggle.com/siriuself/bert-disjoint-utils</a>\n<a href=\"https://www.kaggle.com/siriuself/albert-yes-no-fn-builder\">https://www.kaggle.com/siriuself/albert-yes-no-fn-builder</a>\n<a href=\"https://www.kaggle.com/siriuself/albert-yes-no-modeling\">https://www.kaggle.com/siriuself/albert-yes-no-modeling</a>\n<a href=\"https://www.kaggle.com/siriuself/albert-yes-no-utils\">https://www.kaggle.com/siriuself/albert-yes-no-utils</a>\n<a href=\"https://www.kaggle.com/siriuself/tokenization\">https://www.kaggle.com/siriuself/tokenization</a>\n<a href=\"https://www.kaggle.com/siriuself/create-submission\">https://www.kaggle.com/siriuself/create-submission</a></p>\n\n<p>My model is simple, BERT large whole-word-masking uncased, retrained using the start/end logit loss only without the answer type loss. Using provided nq train tf record with the following setting:\n<strong>batch_size</strong>: 32\n<strong>epoch</strong>: 2\n<strong>alpha</strong>: 2e-5, but use ckpt-15000 (so stop at around 1 epoch)\nNote that because of learning rate warmup/decay, this is different from training with 2e-5 for 1 epoch</p>\n\n<p>Then totally disregard answer type classification (since I don't have it), and rely on threshold setting for long and short questions, tuned on the dev set with nq_eval. And yes I let go all the YES/NO questions.</p>\n\n<p>I did try adding an ALBERT xxlarge yes/no verifier after the BERT stage, which seemed to improve for like 1 pt on dev, but apparently not on the LBs somehow. My kernel includes the ALBERT part too. There isn't much insight in the utility scripts, except the modifications I made in order to restore the checkpoint into some contrib layers forced to be re-written in keras (e.g. LayerNormalization). It was a nightmare..</p>\n\n<p><em><strong></strong></em><strong>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>*</strong><em>some reflections/insights</em><strong><em>*</em>**<em>*</em>**<em>*</em>**<em>*</em>**</strong>\nI started the competition way too late and was on the wrong ALBERT-taking-too-long direction for quite a while, ending up with little time for tuning. I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\nBased on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing.</p>\n\n<p>Let me know if you have similar/opposite findings. It's just my feeling anyway, with a bit of confirmation of my score, obtained by doing nothing other than removing the answer type loss in training. This is almost higher than my ALBERT-xxlarge-joint model too.. I almost felt like doing more experiments and writing a paper on the classification power reduction, but guess it's too trivial.</p>\n\n<p>Lastly, we are hiring intermediate/senior NLP engineer/researcher/scientist, with possibility of sponsorship to come to our Canadian headquarter in Waterloo, Ontario for strong candidates. Bilingualism in English and Mandarin is a plus.</p>",
      "rawMarkdown": "Congrats to all winners!\n\nI've made available my solution (public 0.65 private 0.67).\nhttps://www.kaggle.com/siriuself/tf-qa-wwm-verifier-forked\n\nhttps://www.kaggle.com/siriuself/bert-disjoint-fn-builder\nhttps://www.kaggle.com/siriuself/bert-disjoint-modeling\nhttps://www.kaggle.com/siriuself/bert-disjoint-utils\nhttps://www.kaggle.com/siriuself/albert-yes-no-fn-builder\nhttps://www.kaggle.com/siriuself/albert-yes-no-modeling\nhttps://www.kaggle.com/siriuself/albert-yes-no-utils\nhttps://www.kaggle.com/siriuself/tokenization\nhttps://www.kaggle.com/siriuself/create-submission\n\nMy model is simple, BERT large whole-word-masking uncased, retrained using the start/end logit loss only without the answer type loss. Using provided nq train tf record with the following setting:\n**batch_size**: 32\n**epoch**: 2\n**alpha**: 2e-5, but use ckpt-15000 (so stop at around 1 epoch)\nNote that because of learning rate warmup/decay, this is different from training with 2e-5 for 1 epoch\n\nThen totally disregard answer type classification (since I don't have it), and rely on threshold setting for long and short questions, tuned on the dev set with nq_eval. And yes I let go all the YES/NO questions.\n\nI did try adding an ALBERT xxlarge yes/no verifier after the BERT stage, which seemed to improve for like 1 pt on dev, but apparently not on the LBs somehow. My kernel includes the ALBERT part too. There isn't much insight in the utility scripts, except the modifications I made in order to restore the checkpoint into some contrib layers forced to be re-written in keras (e.g. LayerNormalization). It was a nightmare..\n\n*********************************some reflections/insights*************************\nI started the competition way too late and was on the wrong ALBERT-taking-too-long direction for quite a while, ending up with little time for tuning. I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\nBased on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing.\n\nLet me know if you have similar/opposite findings. It's just my feeling anyway, with a bit of confirmation of my score, obtained by doing nothing other than removing the answer type loss in training. This is almost higher than my ALBERT-xxlarge-joint model too.. I almost felt like doing more experiments and writing a paper on the classification power reduction, but guess it's too trivial.\n\nLastly, we are hiring intermediate/senior NLP engineer/researcher/scientist, with possibility of sponsorship to come to our Canadian headquarter in Waterloo, Ontario for strong candidates. Bilingualism in English and Mandarin is a plus.",
      "votes": null
    },
    {
      "id": "726459",
      "postDate": "01/23/2020 02:44:28",
      "content": "<p>Great Job!</p>",
      "rawMarkdown": "Great Job!",
      "votes": null
    },
    {
      "id": "728480",
      "postDate": "01/24/2020 19:29:57",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍",
      "votes": null
    },
    {
      "id": "733310",
      "postDate": "01/31/2020 00:22:14",
      "content": "<p>Thanks for sharing!</p>",
      "rawMarkdown": "Thanks for sharing!",
      "votes": null
    },
    {
      "id": "734165",
      "postDate": "02/01/2020 03:26:55",
      "content": "<p>Nice! Thanks for the open-source!</p>",
      "rawMarkdown": "Nice! Thanks for the open-source!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 726459,
      "author_name": "renxingkai",
      "author_url": "",
      "post_date": "01/23/2020 02:44:28",
      "content": "<p>Great Job!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 728480,
      "author_name": "mashlyn",
      "author_url": "",
      "post_date": "01/24/2020 19:29:57",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 733310,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/31/2020 00:22:14",
      "content": "<p>Thanks for sharing!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 734165,
      "author_name": "adityaecdrid",
      "author_url": "",
      "post_date": "02/01/2020 03:26:55",
      "content": "<p>Nice! Thanks for the open-source!</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "726450": "Congrats to all winners!\n\nI've made available my solution (public 0.65 private 0.67).\nhttps://www.kaggle.com/siriuself/tf-qa-wwm-verifier-forked\n\nhttps://www.kaggle.com/siriuself/bert-disjoint-fn-builder\nhttps://www.kaggle.com/siriuself/bert-disjoint-modeling\nhttps://www.kaggle.com/siriuself/bert-disjoint-utils\nhttps://www.kaggle.com/siriuself/albert-yes-no-fn-builder\nhttps://www.kaggle.com/siriuself/albert-yes-no-modeling\nhttps://www.kaggle.com/siriuself/albert-yes-no-utils\nhttps://www.kaggle.com/siriuself/tokenization\nhttps://www.kaggle.com/siriuself/create-submission\n\nMy model is simple, BERT large whole-word-masking uncased, retrained using the start/end logit loss only without the answer type loss. Using provided nq train tf record with the following setting:\n**batch_size**: 32\n**epoch**: 2\n**alpha**: 2e-5, but use ckpt-15000 (so stop at around 1 epoch)\nNote that because of learning rate warmup/decay, this is different from training with 2e-5 for 1 epoch\n\nThen totally disregard answer type classification (since I don't have it), and rely on threshold setting for long and short questions, tuned on the dev set with nq_eval. And yes I let go all the YES/NO questions.\n\nI did try adding an ALBERT xxlarge yes/no verifier after the BERT stage, which seemed to improve for like 1 pt on dev, but apparently not on the LBs somehow. My kernel includes the ALBERT part too. There isn't much insight in the utility scripts, except the modifications I made in order to restore the checkpoint into some contrib layers forced to be re-written in keras (e.g. LayerNormalization). It was a nightmare..\n\n*********************************some reflections/insights*************************\nI started the competition way too late and was on the wrong ALBERT-taking-too-long direction for quite a while, ending up with little time for tuning. I still have a very strong feeling that the \"joint\" part of bert-joint might be of little use, since we've already known that:\n1) BERT-like structure is poor at passage ranking, and to make it better we need passages at least as many as in MS-MARCO\n2) We only have like 1-3% YES/NO in our training data. Very unbalanced.\nBased on my inspection and verification experiment this might well be the case, which means the answer type classifier might just be reduced to a question type classifier (a much easier task for BERT to see), which would be a bad indicator of what type of answer the passage contains. It might not be a good idea to include it in training in the first place, and it'd be a disaster if you over-rely on the answer type logits for post-processing.\n\nLet me know if you have similar/opposite findings. It's just my feeling anyway, with a bit of confirmation of my score, obtained by doing nothing other than removing the answer type loss in training. This is almost higher than my ALBERT-xxlarge-joint model too.. I almost felt like doing more experiments and writing a paper on the classification power reduction, but guess it's too trivial.\n\nLastly, we are hiring intermediate/senior NLP engineer/researcher/scientist, with possibility of sponsorship to come to our Canadian headquarter in Waterloo, Ontario for strong candidates. Bilingualism in English and Mandarin is a plus.",
    "726459": "Great Job!",
    "728480": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍",
    "733310": "Thanks for sharing!",
    "734165": "Nice! Thanks for the open-source!"
  },
  "source": "meta"
}