{
  "id": 127371,
  "title": "4th place Solution",
  "url": "/competitions/tensorflow2-question-answering/writeups/toxu-4th-place-solution",
  "author_name": "",
  "post_date": "2020-01-23T14:27:48.362204300Z",
  "votes": 29,
  "comment_count": 11,
  "views": 0,
  "content": "<p>Thanks to Kaggle and the hosts for this competition.  It was my first time to participate a question answering competition. I'm happy I learned a lot by doing researches in this area.</p>\n\n<p>Here is my solution:</p>\n\n<h1>preprocessing</h1>\n\n<ul>\n<li>No preprocessing for Text.</li>\n<li>Different Negative sampling rate. Tried 0.02, 0.04 and 0.06.</li>\n</ul>\n\n<h1>Data Aug</h1>\n\n<ul>\n<li>TTA. not work</li>\n<li>Change the answer by replacing it with similar questions' answer.  not work</li>\n<li>Transform from other question answer datasets like squad and hotpotQA. not work</li>\n</ul>\n\n<h1>Models</h1>\n\n<ul>\n<li>Tried XLNet, Bert Large Uncased/Cased, SpanBert Cased, Bert Large WWM.</li>\n<li>Same loss function and prediction as Bert-joint script.</li>\n</ul>\n\n<p>All cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. WWM BERT Large Uncased performed best in my experiments.</p>\n\n<h1>Knowledge Distillation</h1>\n\n<p>I believe knowledge distillation is the key part in my solution.\n* Trained a combined Bert-large model, by adding bert-large weight and wwm-bert-large weight like 0.8 * wwm-bert-large + 0.2 * bert-large. 1 step, 3e-5 lr.\n* Freeze bert layer only finetune classifier weights. 2 step, 1e-5 lr.\n* Treat the first model as a teacher model and do knowledge distillation to get a student model.\n* Finetune student model with only classifer weights. 3 step 1e-5 lr.</p>\n\n<h1>Validation</h1>\n\n<p>The student model achieves 0.7117955439056357 on dev set. and 0.7 on both public lb and private lb.</p>\n\n<h1>Other works</h1>\n\n<p>I failed to implement Adversarial Training using Estimator framework in TF1.0. But it worth trying if you are using pytorch.</p>",
  "messages": [
    {
      "id": "727201",
      "postDate": "01/23/2020 14:27:48",
      "content": "<p>Thanks to Kaggle and the hosts for this competition.  It was my first time to participate a question answering competition. I'm happy I learned a lot by doing researches in this area.</p>\n\n<p>Here is my solution:</p>\n\n<h1>preprocessing</h1>\n\n<ul>\n<li>No preprocessing for Text.</li>\n<li>Different Negative sampling rate. Tried 0.02, 0.04 and 0.06.</li>\n</ul>\n\n<h1>Data Aug</h1>\n\n<ul>\n<li>TTA. not work</li>\n<li>Change the answer by replacing it with similar questions' answer.  not work</li>\n<li>Transform from other question answer datasets like squad and hotpotQA. not work</li>\n</ul>\n\n<h1>Models</h1>\n\n<ul>\n<li>Tried XLNet, Bert Large Uncased/Cased, SpanBert Cased, Bert Large WWM.</li>\n<li>Same loss function and prediction as Bert-joint script.</li>\n</ul>\n\n<p>All cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. WWM BERT Large Uncased performed best in my experiments.</p>\n\n<h1>Knowledge Distillation</h1>\n\n<p>I believe knowledge distillation is the key part in my solution.\n* Trained a combined Bert-large model, by adding bert-large weight and wwm-bert-large weight like 0.8 * wwm-bert-large + 0.2 * bert-large. 1 step, 3e-5 lr.\n* Freeze bert layer only finetune classifier weights. 2 step, 1e-5 lr.\n* Treat the first model as a teacher model and do knowledge distillation to get a student model.\n* Finetune student model with only classifer weights. 3 step 1e-5 lr.</p>\n\n<h1>Validation</h1>\n\n<p>The student model achieves 0.7117955439056357 on dev set. and 0.7 on both public lb and private lb.</p>\n\n<h1>Other works</h1>\n\n<p>I failed to implement Adversarial Training using Estimator framework in TF1.0. But it worth trying if you are using pytorch.</p>",
      "rawMarkdown": "Thanks to Kaggle and the hosts for this competition.  It was my first time to participate a question answering competition. I'm happy I learned a lot by doing researches in this area.\n\nHere is my solution:\n\n# preprocessing\n\n* No preprocessing for Text.\n* Different Negative sampling rate. Tried 0.02, 0.04 and 0.06.\n\n# Data Aug\n\n* TTA. not work\n* Change the answer by replacing it with similar questions' answer.  not work\n* Transform from other question answer datasets like squad and hotpotQA. not work\n\n# Models\n\n* Tried XLNet, Bert Large Uncased/Cased, SpanBert Cased, Bert Large WWM.\n* Same loss function and prediction as Bert-joint script.\n\nAll cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. WWM BERT Large Uncased performed best in my experiments.\n\n# Knowledge Distillation\n\nI believe knowledge distillation is the key part in my solution.\n* Trained a combined Bert-large model, by adding bert-large weight and wwm-bert-large weight like 0.8 * wwm-bert-large + 0.2 * bert-large. 1 step, 3e-5 lr.\n* Freeze bert layer only finetune classifier weights. 2 step, 1e-5 lr.\n* Treat the first model as a teacher model and do knowledge distillation to get a student model.\n* Finetune student model with only classifer weights. 3 step 1e-5 lr.\n\n# Validation\n\nThe student model achieves 0.7117955439056357 on dev set. and 0.7 on both public lb and private lb.\n\n# Other works\n\nI failed to implement Adversarial Training using Estimator framework in TF1.0. But it worth trying if you are using pytorch.",
      "votes": null
    },
    {
      "id": "727268",
      "postDate": "01/23/2020 15:35:41",
      "content": "<p>Awesome and distinct solution. Thanks for sharing. Can you elaborate more how you trained a combined bert? As far as I understand you averaged the weights (i.e. all model parameters) 0.8 / 0.2? Or did you average the bert model output?</p>",
      "rawMarkdown": "Awesome and distinct solution. Thanks for sharing. Can you elaborate more how you trained a combined bert? As far as I understand you averaged the weights (i.e. all model parameters) 0.8 / 0.2? Or did you average the bert model output?",
      "votes": null
    },
    {
      "id": "727273",
      "postDate": "01/23/2020 15:42:21",
      "content": "<p>Yes, I tried different approaches like averaged bert output, concat, multiply. And weighted average gave me the highest val score.</p>",
      "rawMarkdown": "Yes, I tried different approaches like averaged bert output, concat, multiply. And weighted average gave me the highest val score.",
      "votes": null
    },
    {
      "id": "727310",
      "postDate": "01/23/2020 16:13:04",
      "content": "<p>Thank you for you posting. Want to know more about your \"teacher teaching student\" strategy</p>",
      "rawMarkdown": "Thank you for you posting. Want to know more about your \"teacher teaching student\" strategy",
      "votes": null
    },
    {
      "id": "727388",
      "postDate": "01/23/2020 17:20:23",
      "content": "<p>Thanks for sharing. </p>\n\n<p>What are your teacher model's score on dev and LB? Is it much higher than student model, or similar?</p>\n\n<blockquote>\n  <p>All cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. </p>\n</blockquote>\n\n<p>We observed the same results (although we used a cased model in the ensemble). I think this is expected: joint-bert NQ, Bert Squad, Albert Squad all used uncased version. Explanation would be, extractive QA tasks are not that case sensitive. And by using uncased vocab, you get a much bigger actual vocab size.</p>",
      "rawMarkdown": "Thanks for sharing. \n\nWhat are your teacher model's score on dev and LB? Is it much higher than student model, or similar?\n\n&gt; All cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. \n\nWe observed the same results (although we used a cased model in the ensemble). I think this is expected: joint-bert NQ, Bert Squad, Albert Squad all used uncased version. Explanation would be, extractive QA tasks are not that case sensitive. And by using uncased vocab, you get a much bigger actual vocab size.",
      "votes": null
    },
    {
      "id": "727741",
      "postDate": "01/24/2020 02:12:38",
      "content": "<p>Thank you, you may want to read this paper to understand the idea. <a href=\"https://arxiv.org/abs/1503.02531\">Distilling the Knowledge in a Neural Network</a>.</p>",
      "rawMarkdown": "Thank you, you may want to read this paper to understand the idea. [Distilling the Knowledge in a Neural Network](https://arxiv.org/abs/1503.02531).",
      "votes": null
    },
    {
      "id": "727748",
      "postDate": "01/24/2020 02:19:41",
      "content": "<p>Thank you for your explanation and I do agree with you. \nFor knowledge distillation, teacher's score is lower than student model. \nFor teacher model the local val score is 0.7021821120689655.\nAfter KD, the student model's val score is 0.7137155780231474, and it got 0.69 in private lb.\nAfter finetuning the classifier weights the val score decreased to 0.7117955439056357, but private lb increased to 0.70.</p>",
      "rawMarkdown": "Thank you for your explanation and I do agree with you. \nFor knowledge distillation, teacher's score is lower than student model. \nFor teacher model the local val score is 0.7021821120689655.\nAfter KD, the student model's val score is 0.7137155780231474, and it got 0.69 in private lb.\nAfter finetuning the classifier weights the val score decreased to 0.7117955439056357, but private lb increased to 0.70.",
      "votes": null
    },
    {
      "id": "729016",
      "postDate": "01/25/2020 15:29:59",
      "content": "<p>Thanks <a href=\"/tonyxu\">@tonyxu</a> for the write-up and congrats on the result! Just to clarify, what architecture did you use for the student model?</p>",
      "rawMarkdown": "Thanks @tonyxu for the write-up and congrats on the result! Just to clarify, what architecture did you use for the student model?",
      "votes": null
    },
    {
      "id": "729310",
      "postDate": "01/26/2020 02:05:25",
      "content": "<p>It's a bert large model.</p>",
      "rawMarkdown": "It's a bert large model.",
      "votes": null
    },
    {
      "id": "733309",
      "postDate": "01/31/2020 00:21:30",
      "content": "<p>Congrats! Thanks for sharing.</p>",
      "rawMarkdown": "Congrats! Thanks for sharing.",
      "votes": null
    },
    {
      "id": "739424",
      "postDate": "02/07/2020 20:28:45",
      "content": "<p><a href=\"/tonyxu\">@tonyxu</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "rawMarkdown": "tonyxu, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
      "votes": null
    },
    {
      "id": "1271070",
      "postDate": "04/12/2021 09:14:00",
      "content": "<p>Thanks for sharing!<br>\nMay I ask which negative sampling rate did you find best fit this dataset?</p>",
      "rawMarkdown": "Thanks for sharing!\nMay I ask which negative sampling rate did you find best fit this dataset?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1271070,
      "author_name": "kunruizhu",
      "author_url": "",
      "post_date": "04/12/2021 09:14:00",
      "content": "<p>Thanks for sharing!<br>\nMay I ask which negative sampling rate did you find best fit this dataset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 727268,
      "author_name": "christofhenkel",
      "author_url": "",
      "post_date": "01/23/2020 15:35:41",
      "content": "<p>Awesome and distinct solution. Thanks for sharing. Can you elaborate more how you trained a combined bert? As far as I understand you averaged the weights (i.e. all model parameters) 0.8 / 0.2? Or did you average the bert model output?</p>",
      "votes": null,
      "replies": [
        {
          "id": 727273,
          "author_name": "tonyxu",
          "author_url": "",
          "post_date": "01/23/2020 15:42:21",
          "content": "<p>Yes, I tried different approaches like averaged bert output, concat, multiply. And weighted average gave me the highest val score.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 727310,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "01/23/2020 16:13:04",
      "content": "<p>Thank you for you posting. Want to know more about your \"teacher teaching student\" strategy</p>",
      "votes": null,
      "replies": [
        {
          "id": 727741,
          "author_name": "tonyxu",
          "author_url": "",
          "post_date": "01/24/2020 02:12:38",
          "content": "<p>Thank you, you may want to read this paper to understand the idea. <a href=\"https://arxiv.org/abs/1503.02531\">Distilling the Knowledge in a Neural Network</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 727388,
      "author_name": "boliu0",
      "author_url": "",
      "post_date": "01/23/2020 17:20:23",
      "content": "<p>Thanks for sharing. </p>\n\n<p>What are your teacher model's score on dev and LB? Is it much higher than student model, or similar?</p>\n\n<blockquote>\n  <p>All cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. </p>\n</blockquote>\n\n<p>We observed the same results (although we used a cased model in the ensemble). I think this is expected: joint-bert NQ, Bert Squad, Albert Squad all used uncased version. Explanation would be, extractive QA tasks are not that case sensitive. And by using uncased vocab, you get a much bigger actual vocab size.</p>",
      "votes": null,
      "replies": [
        {
          "id": 727748,
          "author_name": "tonyxu",
          "author_url": "",
          "post_date": "01/24/2020 02:19:41",
          "content": "<p>Thank you for your explanation and I do agree with you. \nFor knowledge distillation, teacher's score is lower than student model. \nFor teacher model the local val score is 0.7021821120689655.\nAfter KD, the student model's val score is 0.7137155780231474, and it got 0.69 in private lb.\nAfter finetuning the classifier weights the val score decreased to 0.7117955439056357, but private lb increased to 0.70.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 729016,
      "author_name": "morganmcg",
      "author_url": "",
      "post_date": "01/25/2020 15:29:59",
      "content": "<p>Thanks <a href=\"/tonyxu\">@tonyxu</a> for the write-up and congrats on the result! Just to clarify, what architecture did you use for the student model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 729310,
          "author_name": "tonyxu",
          "author_url": "",
          "post_date": "01/26/2020 02:05:25",
          "content": "<p>It's a bert large model.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 733309,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "01/31/2020 00:21:30",
      "content": "<p>Congrats! Thanks for sharing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 739424,
      "author_name": "vbmokin",
      "author_url": "",
      "post_date": "02/07/2020 20:28:45",
      "content": "<p><a href=\"/tonyxu\">@tonyxu</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "727201": "Thanks to Kaggle and the hosts for this competition.  It was my first time to participate a question answering competition. I'm happy I learned a lot by doing researches in this area.\n\nHere is my solution:\n\n# preprocessing\n\n* No preprocessing for Text.\n* Different Negative sampling rate. Tried 0.02, 0.04 and 0.06.\n\n# Data Aug\n\n* TTA. not work\n* Change the answer by replacing it with similar questions' answer.  not work\n* Transform from other question answer datasets like squad and hotpotQA. not work\n\n# Models\n\n* Tried XLNet, Bert Large Uncased/Cased, SpanBert Cased, Bert Large WWM.\n* Same loss function and prediction as Bert-joint script.\n\nAll cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. WWM BERT Large Uncased performed best in my experiments.\n\n# Knowledge Distillation\n\nI believe knowledge distillation is the key part in my solution.\n* Trained a combined Bert-large model, by adding bert-large weight and wwm-bert-large weight like 0.8 * wwm-bert-large + 0.2 * bert-large. 1 step, 3e-5 lr.\n* Freeze bert layer only finetune classifier weights. 2 step, 1e-5 lr.\n* Treat the first model as a teacher model and do knowledge distillation to get a student model.\n* Finetune student model with only classifer weights. 3 step 1e-5 lr.\n\n# Validation\n\nThe student model achieves 0.7117955439056357 on dev set. and 0.7 on both public lb and private lb.\n\n# Other works\n\nI failed to implement Adversarial Training using Estimator framework in TF1.0. But it worth trying if you are using pytorch.",
    "727268": "Awesome and distinct solution. Thanks for sharing. Can you elaborate more how you trained a combined bert? As far as I understand you averaged the weights (i.e. all model parameters) 0.8 / 0.2? Or did you average the bert model output?",
    "727273": "Yes, I tried different approaches like averaged bert output, concat, multiply. And weighted average gave me the highest val score.",
    "727310": "Thank you for you posting. Want to know more about your \"teacher teaching student\" strategy",
    "727388": "Thanks for sharing. \n\nWhat are your teacher model's score on dev and LB? Is it much higher than student model, or similar?\n\n&gt; All cased models Perform worse than its Uncased Version. Maybe there are something wrong in my data preparation script. \n\nWe observed the same results (although we used a cased model in the ensemble). I think this is expected: joint-bert NQ, Bert Squad, Albert Squad all used uncased version. Explanation would be, extractive QA tasks are not that case sensitive. And by using uncased vocab, you get a much bigger actual vocab size.",
    "727741": "Thank you, you may want to read this paper to understand the idea. [Distilling the Knowledge in a Neural Network](https://arxiv.org/abs/1503.02531).",
    "727748": "Thank you for your explanation and I do agree with you. \nFor knowledge distillation, teacher's score is lower than student model. \nFor teacher model the local val score is 0.7021821120689655.\nAfter KD, the student model's val score is 0.7137155780231474, and it got 0.69 in private lb.\nAfter finetuning the classifier weights the val score decreased to 0.7117955439056357, but private lb increased to 0.70.",
    "729016": "Thanks @tonyxu for the write-up and congrats on the result! Just to clarify, what architecture did you use for the student model?",
    "729310": "It's a bert large model.",
    "733309": "Congrats! Thanks for sharing.",
    "739424": "tonyxu, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
    "1271070": "Thanks for sharing!\nMay I ask which negative sampling rate did you find best fit this dataset?"
  },
  "source": "meta"
}