{
  "id": 128278,
  "title": "9th place solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/128278",
  "author_name": "Anastasia Karpovich",
  "post_date": "2020-01-30T02:30:02.813000",
  "votes": 38,
  "comment_count": 9,
  "views": 0,
  "content": "<p>First I would like to thank Kaggle Team and  TensorFlow for wonderful competition, TFRC program for TPU credits, Google Cloud for 300$ credits and <a href=\"/prokaj\">@prokaj</a>  for sharing his solution. \nIt was great experience to work on big real-world high quality dataset, use TensorFlow 2.0 and TPUs first time, run inference in couple of seconds, train with batch_size=128 and finally win Gold Medal.</p>\n\n<p>My solution is single model in TF 2.1 trained on TPU. It is Bert Joint with some tweaks and postprocessing. Here are main differences from Bert Joint:\n1)  Pretrained model: Whole-Word-Masking Bert Large\n2)  Tfrecords generated with include_unknowns=0.2 (10 time more examples without answer than in original paper).\n3)  Trained 1 epoch with batch size 128, lr=5e-5 (4-5 hours on TPU).\n4)  Use answer type logits: \n-   If answer_type=1 =&gt; yes_no_answer=’NO’\n-   If answer_type=2 =&gt; yes_no_answer=’YES’\n-   If answer_type=4 =&gt; no short answers</p>\n\n<p>5)  Get some answers with top_level=False</p>\n\n<p>I did EDA and noticed that if 2 long answer candidates contain short answer and one candidate is top_level and another candidate is not top_level and it starts with \"Li\" HTML  token  =&gt; about 70% chance that correct candidate is non top_level one. \nSo I implemented this idea as postprocessing.</p>\n\n<p>6)  Linear regression over 9 logits as answer verifier.\n9 logits included 5 answer type logits, cls_start_logit, cls_end_logit, start_span_logit, end_span_logit.</p>\n\n<p>P.S.\nIn my local metric I had long_<em>non</em>_null__threshold = 1, short_non_null_threshold = 1 but for some reason it didn’t have big influence on leaderboard score (comparing to long_non_null_threshold = 2, short_non_null_threshold = 2).</p>\n\n<p>Inference kernel: <a href=\"https://www.kaggle.com/user189546/tfqa-bert-train-tf2\">https://www.kaggle.com/user189546/tfqa-bert-train-tf2</a>\nModel weights: <a href=\"https://www.kaggle.com/user189546/unk0201128w\">https://www.kaggle.com/user189546/unk0201128w</a>\nTrain code: <a href=\"https://www.kaggle.com/user189546/tfqa-train-code\">https://www.kaggle.com/user189546/tfqa-train-code</a>\nTfrecords: <a href=\"https://www.kaggle.com/user189546/train-tfrecords\">https://www.kaggle.com/user189546/train-tfrecords</a></p>\n\n<p>P.S. I reused code from these sources:\n1. <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">https://www.kaggle.com/prokaj/bert-joint-baseline-notebook</a>\n2. <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a>\n3. <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_gputputpupod_tf2.1.ipynb\">https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_gputputpupod_tf2.1.ipynb</a></p>",
  "messages": [
    {
      "id": 732622,
      "postDate": "2020-01-30T02:30:02.813Z",
      "content": "<p>First I would like to thank Kaggle Team and  TensorFlow for wonderful competition, TFRC program for TPU credits, Google Cloud for 300$ credits and <a href=\"/prokaj\">@prokaj</a>  for sharing his solution. \nIt was great experience to work on big real-world high quality dataset, use TensorFlow 2.0 and TPUs first time, run inference in couple of seconds, train with batch_size=128 and finally win Gold Medal.</p>\n\n<p>My solution is single model in TF 2.1 trained on TPU. It is Bert Joint with some tweaks and postprocessing. Here are main differences from Bert Joint:\n1)  Pretrained model: Whole-Word-Masking Bert Large\n2)  Tfrecords generated with include_unknowns=0.2 (10 time more examples without answer than in original paper).\n3)  Trained 1 epoch with batch size 128, lr=5e-5 (4-5 hours on TPU).\n4)  Use answer type logits: \n-   If answer_type=1 =&gt; yes_no_answer=’NO’\n-   If answer_type=2 =&gt; yes_no_answer=’YES’\n-   If answer_type=4 =&gt; no short answers</p>\n\n<p>5)  Get some answers with top_level=False</p>\n\n<p>I did EDA and noticed that if 2 long answer candidates contain short answer and one candidate is top_level and another candidate is not top_level and it starts with \"Li\" HTML  token  =&gt; about 70% chance that correct candidate is non top_level one. \nSo I implemented this idea as postprocessing.</p>\n\n<p>6)  Linear regression over 9 logits as answer verifier.\n9 logits included 5 answer type logits, cls_start_logit, cls_end_logit, start_span_logit, end_span_logit.</p>\n\n<p>P.S.\nIn my local metric I had long_<em>non</em>_null__threshold = 1, short_non_null_threshold = 1 but for some reason it didn’t have big influence on leaderboard score (comparing to long_non_null_threshold = 2, short_non_null_threshold = 2).</p>\n\n<p>Inference kernel: <a href=\"https://www.kaggle.com/user189546/tfqa-bert-train-tf2\">https://www.kaggle.com/user189546/tfqa-bert-train-tf2</a>\nModel weights: <a href=\"https://www.kaggle.com/user189546/unk0201128w\">https://www.kaggle.com/user189546/unk0201128w</a>\nTrain code: <a href=\"https://www.kaggle.com/user189546/tfqa-train-code\">https://www.kaggle.com/user189546/tfqa-train-code</a>\nTfrecords: <a href=\"https://www.kaggle.com/user189546/train-tfrecords\">https://www.kaggle.com/user189546/train-tfrecords</a></p>\n\n<p>P.S. I reused code from these sources:\n1. <a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\">https://www.kaggle.com/prokaj/bert-joint-baseline-notebook</a>\n2. <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a>\n3. <a href=\"https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_gputputpupod_tf2.1.ipynb\">https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_gputputpupod_tf2.1.ipynb</a></p>",
      "rawMarkdown": "First I would like to thank Kaggle Team and  TensorFlow for wonderful competition, TFRC program for TPU credits, Google Cloud for 300$ credits and @prokaj  for sharing his solution. \nIt was great experience to work on big real-world high quality dataset, use TensorFlow 2.0 and TPUs first time, run inference in couple of seconds, train with batch_size=128 and finally win Gold Medal.\n\nMy solution is single model in TF 2.1 trained on TPU. It is Bert Joint with some tweaks and postprocessing. Here are main differences from Bert Joint:\n1)\tPretrained model: Whole-Word-Masking Bert Large\n2)\tTfrecords generated with include_unknowns=0.2 (10 time more examples without answer than in original paper).\n3)\tTrained 1 epoch with batch size 128, lr=5e-5 (4-5 hours on TPU).\n4)\tUse answer type logits: \n-\tIf answer_type=1 =&gt; yes_no_answer=’NO’\n-\tIf answer_type=2 =&gt; yes_no_answer=’YES’\n-\tIf answer_type=4 =&gt; no short answers\n\n\n5)\tGet some answers with top_level=False\n\nI did EDA and noticed that if 2 long answer candidates contain short answer and one candidate is top_level and another candidate is not top_level and it starts with \"Li\" HTML  token  =&gt; about 70% chance that correct candidate is non top_level one. \nSo I implemented this idea as postprocessing.\n\n\n6)\tLinear regression over 9 logits as answer verifier.\n9 logits included 5 answer type logits, cls_start_logit, cls_end_logit, start_span_logit, end_span_logit.\n\nP.S.\nIn my local metric I had long__non__null__threshold = 1, short_non_null_threshold = 1 but for some reason it didn’t have big influence on leaderboard score (comparing to long_non_null_threshold = 2, short_non_null_threshold = 2).\n\n\nInference kernel: https://www.kaggle.com/user189546/tfqa-bert-train-tf2\nModel weights: https://www.kaggle.com/user189546/unk0201128w\nTrain code: https://www.kaggle.com/user189546/tfqa-train-code\nTfrecords: https://www.kaggle.com/user189546/train-tfrecords\n\n\nP.S. I reused code from these sources:\n1. https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\n2. https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\n3. https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_gputputpupod_tf2.1.ipynb",
      "votes": 38
    },
    {
      "id": 739427,
      "postDate": "2020-02-07T20:38:40.767Z",
      "content": "<p><a href=\"/user189546\">@user189546</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "rawMarkdown": "@user189546, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
      "votes": 1
    },
    {
      "id": 737780,
      "postDate": "2020-02-05T19:07:58.310Z",
      "content": "<p>I expect so much from this community. Upvote Anastasia 9TH TOP Solution. PLEASE DON'T DISAPPOINT ME KAGGLERS!</p>",
      "rawMarkdown": "I expect so much from this community. Upvote Anastasia 9TH TOP Solution. PLEASE DON'T DISAPPOINT ME KAGGLERS!",
      "votes": 2
    },
    {
      "id": 733294,
      "postDate": "2020-01-30T23:46:56.767Z",
      "content": "<p>It's a pleasure to read the solution from a well-succeded Woman in top of Competitions. Congratulations Anastasia.</p>",
      "rawMarkdown": "It's a pleasure to read the solution from a well-succeded Woman in top of Competitions. Congratulations Anastasia."
    },
    {
      "id": 736741,
      "postDate": "2020-02-04T14:13:14.547Z",
      "content": "<p>There should be some stats to compare: Male receiving Votes sharing solutions with Female receiving Votes sharing solutions. Any one interested, just take a look to Discussion topics and their respective number of votes.</p>",
      "rawMarkdown": "There should be some stats to compare: Male receiving Votes sharing solutions with Female receiving Votes sharing solutions. Any one interested, just take a look to Discussion topics and their respective number of votes."
    },
    {
      "id": 746686,
      "postDate": "2020-02-15T11:54:08.703Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 736332,
      "postDate": "2020-02-04T04:29:47.057Z",
      "content": "<p>Thanks everybody!</p>",
      "rawMarkdown": "Thanks everybody!",
      "votes": 1
    },
    {
      "id": 733302,
      "postDate": "2020-01-31T00:05:32.740Z",
      "content": "<p>Congrats and thank you for sharing 🎉 </p>",
      "rawMarkdown": "Congrats and thank you for sharing 🎉 "
    },
    {
      "id": 732710,
      "postDate": "2020-01-30T06:29:31.707Z",
      "content": "<p>Thanks for sharing <a href=\"/user189546\">@user189546</a> </p>",
      "rawMarkdown": "Thanks for sharing @user189546 "
    },
    {
      "id": 732638,
      "postDate": "2020-01-30T03:16:16.490Z",
      "content": "<p>Thanks for sharing😄</p>",
      "rawMarkdown": "Thanks for sharing😄"
    }
  ],
  "comments": [
    {
      "id": 739427,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-07T20:38:40.767000",
      "content": "<p><a href=\"/user189546\">@user189546</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 737780,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2020-02-05T19:07:58.310000",
      "content": "<p>I expect so much from this community. Upvote Anastasia 9TH TOP Solution. PLEASE DON'T DISAPPOINT ME KAGGLERS!</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 733294,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2020-01-30T23:46:56.767000",
      "content": "<p>It's a pleasure to read the solution from a well-succeded Woman in top of Competitions. Congratulations Anastasia.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 736741,
      "author_name": "Marília Prata",
      "author_url": "",
      "post_date": "2020-02-04T14:13:14.547000",
      "content": "<p>There should be some stats to compare: Male receiving Votes sharing solutions with Female receiving Votes sharing solutions. Any one interested, just take a look to Discussion topics and their respective number of votes.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 746686,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-15T11:54:08.703000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 736332,
      "author_name": "Anastasia Karpovich",
      "author_url": "",
      "post_date": "2020-02-04T04:29:47.057000",
      "content": "<p>Thanks everybody!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 733302,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-31T00:05:32.740000",
      "content": "<p>Congrats and thank you for sharing 🎉 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732710,
      "author_name": "Erkan Hatipoğlu",
      "author_url": "",
      "post_date": "2020-01-30T06:29:31.707000",
      "content": "<p>Thanks for sharing <a href=\"/user189546\">@user189546</a> </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732638,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-01-30T03:16:16.490000",
      "content": "<p>Thanks for sharing😄</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "732622": "First I would like to thank Kaggle Team and  TensorFlow for wonderful competition, TFRC program for TPU credits, Google Cloud for 300$ credits and @prokaj  for sharing his solution. \nIt was great experience to work on big real-world high quality dataset, use TensorFlow 2.0 and TPUs first time, run inference in couple of seconds, train with batch_size=128 and finally win Gold Medal.\n\nMy solution is single model in TF 2.1 trained on TPU. It is Bert Joint with some tweaks and postprocessing. Here are main differences from Bert Joint:\n1)\tPretrained model: Whole-Word-Masking Bert Large\n2)\tTfrecords generated with include_unknowns=0.2 (10 time more examples without answer than in original paper).\n3)\tTrained 1 epoch with batch size 128, lr=5e-5 (4-5 hours on TPU).\n4)\tUse answer type logits: \n-\tIf answer_type=1 =&gt; yes_no_answer=’NO’\n-\tIf answer_type=2 =&gt; yes_no_answer=’YES’\n-\tIf answer_type=4 =&gt; no short answers\n\n\n5)\tGet some answers with top_level=False\n\nI did EDA and noticed that if 2 long answer candidates contain short answer and one candidate is top_level and another candidate is not top_level and it starts with \"Li\" HTML  token  =&gt; about 70% chance that correct candidate is non top_level one. \nSo I implemented this idea as postprocessing.\n\n\n6)\tLinear regression over 9 logits as answer verifier.\n9 logits included 5 answer type logits, cls_start_logit, cls_end_logit, start_span_logit, end_span_logit.\n\nP.S.\nIn my local metric I had long__non__null__threshold = 1, short_non_null_threshold = 1 but for some reason it didn’t have big influence on leaderboard score (comparing to long_non_null_threshold = 2, short_non_null_threshold = 2).\n\n\nInference kernel: https://www.kaggle.com/user189546/tfqa-bert-train-tf2\nModel weights: https://www.kaggle.com/user189546/unk0201128w\nTrain code: https://www.kaggle.com/user189546/tfqa-train-code\nTfrecords: https://www.kaggle.com/user189546/train-tfrecords\n\n\nP.S. I reused code from these sources:\n1. https://www.kaggle.com/prokaj/bert-joint-baseline-notebook\n2. https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\n3. https://github.com/GoogleCloudPlatform/training-data-analyst/blob/master/courses/fast-and-lean-data-science/keras_flowers_gputputpupod_tf2.1.ipynb",
    "739427": "@user189546, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
    "737780": "I expect so much from this community. Upvote Anastasia 9TH TOP Solution. PLEASE DON'T DISAPPOINT ME KAGGLERS!",
    "733294": "It's a pleasure to read the solution from a well-succeded Woman in top of Competitions. Congratulations Anastasia.",
    "736741": "There should be some stats to compare: Male receiving Votes sharing solutions with Female receiving Votes sharing solutions. Any one interested, just take a look to Discussion topics and their respective number of votes.",
    "746686": "",
    "736332": "Thanks everybody!",
    "733302": "Congrats and thank you for sharing 🎉 ",
    "732710": "Thanks for sharing @user189546 ",
    "732638": "Thanks for sharing😄"
  }
}