{
  "id": 127521,
  "title": "6th place solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/127521",
  "author_name": "prvi",
  "post_date": "2020-01-24T13:15:39.582000",
  "votes": 26,
  "comment_count": 8,
  "views": 0,
  "content": "<p>The final submission was a single BERT based model. It gave .71 on public data  and 0.69 on private leaderboard. Looking at other solutions it was  a little bit overcomplicated.  </p>\n\n<h2>Preprocessing</h2>\n\n<p>I left out the special tokens introduced in the a baseline script (<code>[ContextId=..][Paragraph=0]</code> etc). Instead I kept the simplified html tags (table tags eg.  contained <code>colspan</code> info which I removed). I also added <code>&amp;lt;*&amp;gt;</code>, <code>&amp;lt;/*&amp;gt;</code>   at the beginning and the end of each segment.  I kept 4 % of the negative examples, and also kept the very long answers that were not contained within one segment. I also processed the entire document text, so the <code>max_contexts</code>  argument of the original script was ignored. </p>\n\n<h2>Model output</h2>\n\n<p>Similarly to the baseline I used the classification head,  and one head for span start and end logits.  With masking this used to get both the long answer, short answer logits.\nI also added ''cross'' head, which is a bilinear function of the pairs of the sequence output of the BERT model. Short span logits then obtained  as the sum of the start and end logits and the corresponding output of the cross head. \nImpossible spans were masked out and <code>softmax</code> gave the span probabilities. For the long span cross entropy criterion was used both for start and end logits. For the short spans the error was negative log of the total probability of positive short spans.  These error terms were  computed only for examples having long, short answers. So the aim here is to learn the position given that there is an answer, the probability of having an answer came from the <code>answer_type</code> output.</p>\n\n<h2>Postprocessing</h2>\n\n<p>For each segment the  long and short spans with maximal probability was computed. From the answer type head the probabilities of having a short or long answer in the segment were computed and these probabilities were assigned to the most likely spans within the segment. These votes were maximized over all segments containing the given span.  Then the spans with highest overall scores was considered for the answer. Thresholds were computed using the development data of the NQ dataset. </p>\n\n<h2>Training</h2>\n\n<p>I trained on tpu for 2 epochs using learning rate 2.5e-5 and batch size 64. Before training on nq data, I fine tuned the BERT model on squad 2.0 dataset with the same setting and preprocessing.</p>\n\n<h2>Code</h2>\n\n<p>The final submission was produced with <br>\n<a href=\"https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5\">https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5</a></p>\n\n<p>Pre and post processing code <br>\n<a href=\"https://www.kaggle.com/prokaj/bert-baseline-pre-and-post-process\">https://www.kaggle.com/prokaj/bert-baseline-pre-and-post-process</a></p>\n\n<p>final model in saved model format <br>\n<a href=\"https://www.kaggle.com/prokaj/tpu-2020-01-22\">https://www.kaggle.com/prokaj/tpu-2020-01-22</a></p>\n\n<p>model code (used on tpu) <br>\n<a href=\"https://www.kaggle.com/prokaj/tpu-code\">https://www.kaggle.com/prokaj/tpu-code</a></p>\n\n<p>BERT implementation from official tensorflow models (preinstalled on TPU)\n<a href=\"https://github.com/tensorflow/models/tree/master/official\">https://github.com/tensorflow/models/tree/master/official</a></p>",
  "messages": [
    {
      "id": 728156,
      "postDate": "2020-01-24T13:15:39.583Z",
      "content": "<p>The final submission was a single BERT based model. It gave .71 on public data  and 0.69 on private leaderboard. Looking at other solutions it was  a little bit overcomplicated.  </p>\n\n<h2>Preprocessing</h2>\n\n<p>I left out the special tokens introduced in the a baseline script (<code>[ContextId=..][Paragraph=0]</code> etc). Instead I kept the simplified html tags (table tags eg.  contained <code>colspan</code> info which I removed). I also added <code>&amp;lt;*&amp;gt;</code>, <code>&amp;lt;/*&amp;gt;</code>   at the beginning and the end of each segment.  I kept 4 % of the negative examples, and also kept the very long answers that were not contained within one segment. I also processed the entire document text, so the <code>max_contexts</code>  argument of the original script was ignored. </p>\n\n<h2>Model output</h2>\n\n<p>Similarly to the baseline I used the classification head,  and one head for span start and end logits.  With masking this used to get both the long answer, short answer logits.\nI also added ''cross'' head, which is a bilinear function of the pairs of the sequence output of the BERT model. Short span logits then obtained  as the sum of the start and end logits and the corresponding output of the cross head. \nImpossible spans were masked out and <code>softmax</code> gave the span probabilities. For the long span cross entropy criterion was used both for start and end logits. For the short spans the error was negative log of the total probability of positive short spans.  These error terms were  computed only for examples having long, short answers. So the aim here is to learn the position given that there is an answer, the probability of having an answer came from the <code>answer_type</code> output.</p>\n\n<h2>Postprocessing</h2>\n\n<p>For each segment the  long and short spans with maximal probability was computed. From the answer type head the probabilities of having a short or long answer in the segment were computed and these probabilities were assigned to the most likely spans within the segment. These votes were maximized over all segments containing the given span.  Then the spans with highest overall scores was considered for the answer. Thresholds were computed using the development data of the NQ dataset. </p>\n\n<h2>Training</h2>\n\n<p>I trained on tpu for 2 epochs using learning rate 2.5e-5 and batch size 64. Before training on nq data, I fine tuned the BERT model on squad 2.0 dataset with the same setting and preprocessing.</p>\n\n<h2>Code</h2>\n\n<p>The final submission was produced with <br>\n<a href=\"https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5\">https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5</a></p>\n\n<p>Pre and post processing code <br>\n<a href=\"https://www.kaggle.com/prokaj/bert-baseline-pre-and-post-process\">https://www.kaggle.com/prokaj/bert-baseline-pre-and-post-process</a></p>\n\n<p>final model in saved model format <br>\n<a href=\"https://www.kaggle.com/prokaj/tpu-2020-01-22\">https://www.kaggle.com/prokaj/tpu-2020-01-22</a></p>\n\n<p>model code (used on tpu) <br>\n<a href=\"https://www.kaggle.com/prokaj/tpu-code\">https://www.kaggle.com/prokaj/tpu-code</a></p>\n\n<p>BERT implementation from official tensorflow models (preinstalled on TPU)\n<a href=\"https://github.com/tensorflow/models/tree/master/official\">https://github.com/tensorflow/models/tree/master/official</a></p>",
      "rawMarkdown": "The final submission was a single BERT based model. It gave .71 on public data  and 0.69 on private leaderboard. Looking at other solutions it was  a little bit overcomplicated.  \n\n## Preprocessing\nI left out the special tokens introduced in the a baseline script (`[ContextId=..][Paragraph=0]` etc). Instead I kept the simplified html tags (table tags eg.  contained `colspan` info which I removed). I also added `&lt;*&gt;`, `&lt;/*&gt;`   at the beginning and the end of each segment.  I kept 4 % of the negative examples, and also kept the very long answers that were not contained within one segment. I also processed the entire document text, so the `max_contexts`  argument of the original script was ignored. \n\n##Model output\nSimilarly to the baseline I used the classification head,  and one head for span start and end logits.  With masking this used to get both the long answer, short answer logits.\nI also added ''cross'' head, which is a bilinear function of the pairs of the sequence output of the BERT model. Short span logits then obtained  as the sum of the start and end logits and the corresponding output of the cross head. \nImpossible spans were masked out and `softmax` gave the span probabilities. For the long span cross entropy criterion was used both for start and end logits. For the short spans the error was negative log of the total probability of positive short spans.  These error terms were  computed only for examples having long, short answers. So the aim here is to learn the position given that there is an answer, the probability of having an answer came from the `answer_type` output.\n\n##Postprocessing\nFor each segment the  long and short spans with maximal probability was computed. From the answer type head the probabilities of having a short or long answer in the segment were computed and these probabilities were assigned to the most likely spans within the segment. These votes were maximized over all segments containing the given span.  Then the spans with highest overall scores was considered for the answer. Thresholds were computed using the development data of the NQ dataset. \n\n##Training\nI trained on tpu for 2 epochs using learning rate 2.5e-5 and batch size 64. Before training on nq data, I fine tuned the BERT model on squad 2.0 dataset with the same setting and preprocessing.\n\n##Code\nThe final submission was produced with  \nhttps://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5\n\nPre and post processing code  \nhttps://www.kaggle.com/prokaj/bert-baseline-pre-and-post-process\n\nfinal model in saved model format  \nhttps://www.kaggle.com/prokaj/tpu-2020-01-22\n\nmodel code (used on tpu)  \nhttps://www.kaggle.com/prokaj/tpu-code\n\nBERT implementation from official tensorflow models (preinstalled on TPU)\nhttps://github.com/tensorflow/models/tree/master/official\n\n",
      "votes": 26
    },
    {
      "id": 769909,
      "postDate": "2020-03-12T11:43:20.897Z",
      "content": "<p>Very elegant solution. Congratulations!</p>",
      "rawMarkdown": "Very elegant solution. Congratulations!"
    },
    {
      "id": 733308,
      "postDate": "2020-01-31T00:19:43.007Z",
      "content": "<p>Congrats and thank you for sharing discussion &amp; code!</p>",
      "rawMarkdown": "Congrats and thank you for sharing discussion &amp; code!"
    },
    {
      "id": 728482,
      "postDate": "2020-01-24T19:31:10.933Z",
      "content": "<p>Thanks for the writeup. This link seems to be 404. Can you take a look?\n<a href=\"https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5\">https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5</a></p>",
      "rawMarkdown": "Thanks for the writeup. This link seems to be 404. Can you take a look?\nhttps://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5",
      "replies": [
        {
          "id": 729785,
          "postDate": "2020-01-26T16:31:02Z",
          "content": "<p>Thanks for spotting this. I forgot to share the notebook.</p>",
          "rawMarkdown": "Thanks for spotting this. I forgot to share the notebook.",
          "votes": 1
        }
      ]
    },
    {
      "id": 728280,
      "postDate": "2020-01-24T14:58:03.673Z",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 "
    },
    {
      "id": 739425,
      "postDate": "2020-02-07T20:33:02.777Z",
      "rawMarkdown": "",
      "votes": -1,
      "isDeleted": true
    },
    {
      "id": 728349,
      "postDate": "2020-01-24T16:30:03.793Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 728367,
      "postDate": "2020-01-24T16:41:04.783Z",
      "content": "<p>Congrats!! Thanks for sharing!!</p>",
      "rawMarkdown": "Congrats!! Thanks for sharing!!"
    }
  ],
  "comments": [
    {
      "id": 769909,
      "author_name": "Rishabh Jha",
      "author_url": "",
      "post_date": "2020-03-12T11:43:20.897000",
      "content": "<p>Very elegant solution. Congratulations!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 733308,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-31T00:19:43.007000",
      "content": "<p>Congrats and thank you for sharing discussion &amp; code!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728482,
      "author_name": "Alon Bochman",
      "author_url": "",
      "post_date": "2020-01-24T19:31:10.933000",
      "content": "<p>Thanks for the writeup. This link seems to be 404. Can you take a look?\n<a href=\"https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5\">https://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 729785,
          "author_name": "prvi",
          "author_url": "",
          "post_date": "2020-01-26T16:31:02",
          "content": "<p>Thanks for spotting this. I forgot to share the notebook.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 728280,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-01-24T14:58:03.673000",
      "content": "<p>Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 739425,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-02-07T20:33:02.777000",
      "content": "",
      "votes": -1,
      "replies": []
    },
    {
      "id": 728349,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-24T16:30:03.793000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728367,
      "author_name": "Pietro Marinelli",
      "author_url": "",
      "post_date": "2020-01-24T16:41:04.783000",
      "content": "<p>Congrats!! Thanks for sharing!!</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "728156": "The final submission was a single BERT based model. It gave .71 on public data  and 0.69 on private leaderboard. Looking at other solutions it was  a little bit overcomplicated.  \n\n## Preprocessing\nI left out the special tokens introduced in the a baseline script (`[ContextId=..][Paragraph=0]` etc). Instead I kept the simplified html tags (table tags eg.  contained `colspan` info which I removed). I also added `&lt;*&gt;`, `&lt;/*&gt;`   at the beginning and the end of each segment.  I kept 4 % of the negative examples, and also kept the very long answers that were not contained within one segment. I also processed the entire document text, so the `max_contexts`  argument of the original script was ignored. \n\n##Model output\nSimilarly to the baseline I used the classification head,  and one head for span start and end logits.  With masking this used to get both the long answer, short answer logits.\nI also added ''cross'' head, which is a bilinear function of the pairs of the sequence output of the BERT model. Short span logits then obtained  as the sum of the start and end logits and the corresponding output of the cross head. \nImpossible spans were masked out and `softmax` gave the span probabilities. For the long span cross entropy criterion was used both for start and end logits. For the short spans the error was negative log of the total probability of positive short spans.  These error terms were  computed only for examples having long, short answers. So the aim here is to learn the position given that there is an answer, the probability of having an answer came from the `answer_type` output.\n\n##Postprocessing\nFor each segment the  long and short spans with maximal probability was computed. From the answer type head the probabilities of having a short or long answer in the segment were computed and these probabilities were assigned to the most likely spans within the segment. These votes were maximized over all segments containing the given span.  Then the spans with highest overall scores was considered for the answer. Thresholds were computed using the development data of the NQ dataset. \n\n##Training\nI trained on tpu for 2 epochs using learning rate 2.5e-5 and batch size 64. Before training on nq data, I fine tuned the BERT model on squad 2.0 dataset with the same setting and preprocessing.\n\n##Code\nThe final submission was produced with  \nhttps://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5\n\nPre and post processing code  \nhttps://www.kaggle.com/prokaj/bert-baseline-pre-and-post-process\n\nfinal model in saved model format  \nhttps://www.kaggle.com/prokaj/tpu-2020-01-22\n\nmodel code (used on tpu)  \nhttps://www.kaggle.com/prokaj/tpu-code\n\nBERT implementation from official tensorflow models (preinstalled on TPU)\nhttps://github.com/tensorflow/models/tree/master/official\n\n",
    "769909": "Very elegant solution. Congratulations!",
    "733308": "Congrats and thank you for sharing discussion &amp; code!",
    "728482": "Thanks for the writeup. This link seems to be 404. Can you take a look?\nhttps://www.kaggle.com/prokaj/fork-of-baseline-html-tokens-v5",
    "728280": "Congrats &amp; Thanks for sharing your solutions🎉 😄 👍 ",
    "739425": "",
    "728349": "",
    "728367": "Congrats!! Thanks for sharing!!"
  }
}