{
  "id": 127333,
  "title": "2nd place solution",
  "url": "/competitions/tensorflow2-question-answering/discussion/127333",
  "author_name": "See--",
  "post_date": "2020-01-23T10:42:30.749000",
  "votes": 92,
  "comment_count": 42,
  "views": 0,
  "content": "<p>Good evening,</p>\n\n<p>first of all, I'd like to thank Kaggle and the hosts for this awesome challenge! It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.</p>\n\n<p>Secondly, congrats to all the winners. I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.</p>\n\n<p>My solution is just a single TF2.0 model. It uses custom heads and a BERT transformers backbone (large version). For modeling and training I am using the great <a href=\"https://github.com/huggingface/transformers\">transformers</a> library. I think that the <a href=\"https://github.com/see--/natural-question-answering/blob/master/models.py#L8-L28\">following snippet</a> is useful to understand the modeling:\n```python\nclass TFBertForNaturalQuestionAnswering(TFBertPreTrainedModel):\n    def <strong>init</strong>(self, config, *inputs, **kwargs):\n        super().<strong>init</strong>(config, *inputs, **kwargs)\n        self.num_labels = config.num_labels</p>\n\n<pre><code>    self.bert = TFBertMainLayer(config, name='bert')\n    self.initializer = get_initializer(config.initializer_range)\n    self.qa_outputs = L.Dense(config.num_labels,\n        kernel_initializer=self.initializer, name='qa_outputs')\n    self.long_outputs = L.Dense(1, kernel_initializer=self.initializer,\n        name='long_outputs')\n\ndef call(self, inputs, **kwargs):\n    outputs = self.bert(inputs, **kwargs)\n    sequence_output = outputs[0]\n    logits = self.qa_outputs(sequence_output)\n    start_logits, end_logits = tf.split(logits, 2, axis=-1)\n    start_logits = tf.squeeze(start_logits, -1)\n    end_logits = tf.squeeze(end_logits, -1)\n    long_logits = tf.squeeze(self.long_outputs(sequence_output), -1)\n    return start_logits, end_logits, long_logits\n</code></pre>\n\n<p>```</p>\n\n<p>As you can see, the natural question answering task can be treated like SQUAD-2.0 with an additional head for long answers. Note that we just need a single output: The opening tag of the HTML bounding box. I guess most competitors used a similar modeling so I think what made the difference to most other solutions is the sampling.</p>\n\n<p>I changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer. I started with a really low empty answer ratio which I got from the <a href=\"https://arxiv.org/abs/1901.08634\"><code>bert-joint</code> paper</a>, but I couldn't reach a good score. I tuned a few hyper parameters but overall I got good results with a wide range of parameters. Adding HTML tags as custom tokens helped a bit. I also tried different start weights and found that:</p>\n\n<p><code>bert-large-uncased</code> (~0.70 LB) &lt; <code>bert-large-uncased-whole-word-masking</code> (~0.72 LB) &lt; <code>bert-large-uncased-whole-word-masking-finetuned-squad</code> (~0.73 LB).</p>\n\n<p>That's about it. Thanks to <a href=\"/boliu0\">@boliu0</a>, <a href=\"/christofhenkel\">@christofhenkel</a> and <a href=\"/kentaronakanishi\">@kentaronakanishi</a> for fixing and providing the metric!</p>\n\n<p>Please refer to my repository for implementation details and instructions to reproduce:\n* <a href=\"https://github.com/see--/natural-question-answering\">https://github.com/see--/natural-question-answering</a></p>\n\n<p>You can find the 2nd place kernel and pretrained weights on Kaggle:\n* <a href=\"https://www.kaggle.com/seesee/submit-full\">https://www.kaggle.com/seesee/submit-full</a>\n* <a href=\"https://www.kaggle.com/seesee/nq-bert-uncased-68\">https://www.kaggle.com/seesee/nq-bert-uncased-68</a></p>\n\n<p>Feel free to ask questions and / or create GitHub issues.</p>",
  "messages": [
    {
      "id": 726961,
      "postDate": "2020-01-23T10:42:30.750Z",
      "content": "<p>Good evening,</p>\n\n<p>first of all, I'd like to thank Kaggle and the hosts for this awesome challenge! It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.</p>\n\n<p>Secondly, congrats to all the winners. I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.</p>\n\n<p>My solution is just a single TF2.0 model. It uses custom heads and a BERT transformers backbone (large version). For modeling and training I am using the great <a href=\"https://github.com/huggingface/transformers\">transformers</a> library. I think that the <a href=\"https://github.com/see--/natural-question-answering/blob/master/models.py#L8-L28\">following snippet</a> is useful to understand the modeling:\n```python\nclass TFBertForNaturalQuestionAnswering(TFBertPreTrainedModel):\n    def <strong>init</strong>(self, config, *inputs, **kwargs):\n        super().<strong>init</strong>(config, *inputs, **kwargs)\n        self.num_labels = config.num_labels</p>\n\n<pre><code>    self.bert = TFBertMainLayer(config, name='bert')\n    self.initializer = get_initializer(config.initializer_range)\n    self.qa_outputs = L.Dense(config.num_labels,\n        kernel_initializer=self.initializer, name='qa_outputs')\n    self.long_outputs = L.Dense(1, kernel_initializer=self.initializer,\n        name='long_outputs')\n\ndef call(self, inputs, **kwargs):\n    outputs = self.bert(inputs, **kwargs)\n    sequence_output = outputs[0]\n    logits = self.qa_outputs(sequence_output)\n    start_logits, end_logits = tf.split(logits, 2, axis=-1)\n    start_logits = tf.squeeze(start_logits, -1)\n    end_logits = tf.squeeze(end_logits, -1)\n    long_logits = tf.squeeze(self.long_outputs(sequence_output), -1)\n    return start_logits, end_logits, long_logits\n</code></pre>\n\n<p>```</p>\n\n<p>As you can see, the natural question answering task can be treated like SQUAD-2.0 with an additional head for long answers. Note that we just need a single output: The opening tag of the HTML bounding box. I guess most competitors used a similar modeling so I think what made the difference to most other solutions is the sampling.</p>\n\n<p>I changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer. I started with a really low empty answer ratio which I got from the <a href=\"https://arxiv.org/abs/1901.08634\"><code>bert-joint</code> paper</a>, but I couldn't reach a good score. I tuned a few hyper parameters but overall I got good results with a wide range of parameters. Adding HTML tags as custom tokens helped a bit. I also tried different start weights and found that:</p>\n\n<p><code>bert-large-uncased</code> (~0.70 LB) &lt; <code>bert-large-uncased-whole-word-masking</code> (~0.72 LB) &lt; <code>bert-large-uncased-whole-word-masking-finetuned-squad</code> (~0.73 LB).</p>\n\n<p>That's about it. Thanks to <a href=\"/boliu0\">@boliu0</a>, <a href=\"/christofhenkel\">@christofhenkel</a> and <a href=\"/kentaronakanishi\">@kentaronakanishi</a> for fixing and providing the metric!</p>\n\n<p>Please refer to my repository for implementation details and instructions to reproduce:\n* <a href=\"https://github.com/see--/natural-question-answering\">https://github.com/see--/natural-question-answering</a></p>\n\n<p>You can find the 2nd place kernel and pretrained weights on Kaggle:\n* <a href=\"https://www.kaggle.com/seesee/submit-full\">https://www.kaggle.com/seesee/submit-full</a>\n* <a href=\"https://www.kaggle.com/seesee/nq-bert-uncased-68\">https://www.kaggle.com/seesee/nq-bert-uncased-68</a></p>\n\n<p>Feel free to ask questions and / or create GitHub issues.</p>",
      "rawMarkdown": "Good evening,\n\nfirst of all, I'd like to thank Kaggle and the hosts for this awesome challenge! It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.\n\nSecondly, congrats to all the winners. I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.\n\nMy solution is just a single TF2.0 model. It uses custom heads and a BERT transformers backbone (large version). For modeling and training I am using the great [transformers](https://github.com/huggingface/transformers) library. I think that the [following snippet](https://github.com/see--/natural-question-answering/blob/master/models.py#L8-L28) is useful to understand the modeling:\n```python\nclass TFBertForNaturalQuestionAnswering(TFBertPreTrainedModel):\n    def __init__(self, config, *inputs, **kwargs):\n        super().__init__(config, *inputs, **kwargs)\n        self.num_labels = config.num_labels\n\n        self.bert = TFBertMainLayer(config, name='bert')\n        self.initializer = get_initializer(config.initializer_range)\n        self.qa_outputs = L.Dense(config.num_labels,\n            kernel_initializer=self.initializer, name='qa_outputs')\n        self.long_outputs = L.Dense(1, kernel_initializer=self.initializer,\n            name='long_outputs')\n\n    def call(self, inputs, **kwargs):\n        outputs = self.bert(inputs, **kwargs)\n        sequence_output = outputs[0]\n        logits = self.qa_outputs(sequence_output)\n        start_logits, end_logits = tf.split(logits, 2, axis=-1)\n        start_logits = tf.squeeze(start_logits, -1)\n        end_logits = tf.squeeze(end_logits, -1)\n        long_logits = tf.squeeze(self.long_outputs(sequence_output), -1)\n        return start_logits, end_logits, long_logits\n```\n\nAs you can see, the natural question answering task can be treated like SQUAD-2.0 with an additional head for long answers. Note that we just need a single output: The opening tag of the HTML bounding box. I guess most competitors used a similar modeling so I think what made the difference to most other solutions is the sampling.\n\nI changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer. I started with a really low empty answer ratio which I got from the [`bert-joint` paper](https://arxiv.org/abs/1901.08634), but I couldn't reach a good score. I tuned a few hyper parameters but overall I got good results with a wide range of parameters. Adding HTML tags as custom tokens helped a bit. I also tried different start weights and found that:\n\n`bert-large-uncased` (~0.70 LB) &lt; `bert-large-uncased-whole-word-masking` (~0.72 LB) &lt; `bert-large-uncased-whole-word-masking-finetuned-squad` (~0.73 LB).\n\nThat's about it. Thanks to @boliu0, @christofhenkel and @kentaronakanishi for fixing and providing the metric!\n\nPlease refer to my repository for implementation details and instructions to reproduce:\n* https://github.com/see--/natural-question-answering\n\nYou can find the 2nd place kernel and pretrained weights on Kaggle:\n* https://www.kaggle.com/seesee/submit-full\n* https://www.kaggle.com/seesee/nq-bert-uncased-68\n\nFeel free to ask questions and / or create GitHub issues.",
      "votes": 92
    },
    {
      "id": 726994,
      "postDate": "2020-01-23T11:04:16.027Z",
      "content": "<p>we went with </p>\n\n<blockquote>\n  <p>taken days on my local 2 x 1080 TI machine</p>\n</blockquote>",
      "rawMarkdown": "we went with \n&gt; taken days on my local 2 x 1080 TI machine",
      "votes": 3,
      "replies": [
        {
          "id": 727000,
          "postDate": "2020-01-23T11:11:28.860Z",
          "content": "<p>haha, I ran out of hyperparameters and ideas 😜</p>",
          "rawMarkdown": "haha, I ran out of hyperparameters and ideas 😜",
          "votes": 1
        }
      ]
    },
    {
      "id": 734457,
      "postDate": "2020-02-01T14:08:53.620Z",
      "content": "<p>Congratulations See!!! You derserve it =]</p>",
      "rawMarkdown": "Congratulations See!!! You derserve it =]",
      "votes": 1
    },
    {
      "id": 731403,
      "postDate": "2020-01-28T15:43:54.543Z",
      "content": "<p>Congratulations Mate.</p>",
      "rawMarkdown": "Congratulations Mate.",
      "votes": 1
    },
    {
      "id": 727510,
      "postDate": "2020-01-23T18:59:17.143Z",
      "content": "<p>Your approach is clean and innovative. I am really enjoying reading it. Thank you so much for sharing your code. You're a good example of what kaggle is all about and you deserve to win the TF2.0 prize as well.</p>",
      "rawMarkdown": "Your approach is clean and innovative. I am really enjoying reading it. Thank you so much for sharing your code. You're a good example of what kaggle is all about and you deserve to win the TF2.0 prize as well.",
      "votes": 1,
      "replies": [
        {
          "id": 728000,
          "postDate": "2020-01-24T10:17:22.790Z",
          "content": "<p>Thanks for the kind words 😄 </p>",
          "rawMarkdown": "Thanks for the kind words 😄 "
        }
      ]
    },
    {
      "id": 727327,
      "postDate": "2020-01-23T16:25:33.677Z",
      "content": "<p>I have two questions. :)</p>\n\n<ol>\n<li><p>Could you please share your long answer F1 and short answer F1 if you evaluated them separately in local validation? I'm wondering how much the <code>long_outputs</code> helps. </p></li>\n<li><p>For the <code>long_logits</code>, did you treat it as a binary classification obj? Whether the token is part of the long answer or not?</p></li>\n</ol>",
      "rawMarkdown": "I have two questions. :)\n\n1. Could you please share your long answer F1 and short answer F1 if you evaluated them separately in local validation? I'm wondering how much the `long_outputs` helps. \n\n2. For the `long_logits`, did you treat it as a binary classification obj? Whether the token is part of the long answer or not?",
      "votes": 1,
      "replies": [
        {
          "id": 727353,
          "postDate": "2020-01-23T16:40:47.217Z",
          "content": "<p>1 Yes I do. The scores are (concat F1, short F1, long F1) and local score is:\n* With ~3500 empty and ~3000 non-empty validation samples: (0.563, 0.481, 0.615)\n* Only ~3000 non-empty: (0.722, 0.576, 0.823)</p>\n\n<p>Maybe check this script if you are using the same metric: <a href=\"https://github.com/see--/natural-question-answering/blob/master/eval_server.py\">https://github.com/see--/natural-question-answering/blob/master/eval_server.py</a></p>\n\n<p>The validation set is here:\n<a href=\"https://github.com/see--/natural-question-answering/blob/master/val_ids.csv\">https://github.com/see--/natural-question-answering/blob/master/val_ids.csv</a></p>\n\n<p>2 No, I treat it as softmax classification. Whether the token is the opening HTML tag and predict [CLS] token if there is no long answer.</p>",
          "rawMarkdown": "1 Yes I do. The scores are (concat F1, short F1, long F1) and local score is:\n* With ~3500 empty and ~3000 non-empty validation samples: (0.563, 0.481, 0.615)\n* Only ~3000 non-empty: (0.722, 0.576, 0.823)\n\nMaybe check this script if you are using the same metric: https://github.com/see--/natural-question-answering/blob/master/eval_server.py\n\nThe validation set is here:\nhttps://github.com/see--/natural-question-answering/blob/master/val_ids.csv\n\n2 No, I treat it as softmax classification. Whether the token is the opening HTML tag and predict [CLS] token if there is no long answer.",
          "votes": 1
        },
        {
          "id": 727516,
          "postDate": "2020-01-23T19:04:15.243Z",
          "content": "<blockquote>\n  <p>2 No, I treat it as softmax classification. Whether the token is the opening HTML tag and predict [CLS] token if there is no long answer.</p>\n</blockquote>\n\n<p>Clever. This and your sampling ratios are big learning points for me so far.</p>",
          "rawMarkdown": "&gt;  2 No, I treat it as softmax classification. Whether the token is the opening HTML tag and predict [CLS] token if there is no long answer.\n\nClever. This and your sampling ratios are big learning points for me so far.",
          "votes": 1
        }
      ]
    },
    {
      "id": 727308,
      "postDate": "2020-01-23T16:11:00.210Z",
      "content": "<p>Big congrats! Thank you so much for the amazing sharing of the source code.</p>",
      "rawMarkdown": "Big congrats! Thank you so much for the amazing sharing of the source code.",
      "votes": 1
    },
    {
      "id": 727204,
      "postDate": "2020-01-23T14:29:53.283Z",
      "content": "<p>Great Solution! Thank you for sharing the code!! :)</p>",
      "rawMarkdown": "Great Solution! Thank you for sharing the code!! :)",
      "votes": 1
    },
    {
      "id": 727176,
      "postDate": "2020-01-23T14:10:41.140Z",
      "content": "<p>Congrats, nice write-up and thanks for sharing the code!</p>",
      "rawMarkdown": "Congrats, nice write-up and thanks for sharing the code!",
      "votes": 1
    },
    {
      "id": 726987,
      "postDate": "2020-01-23T10:54:34.120Z",
      "content": "<p>&gt; I changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer.</p>\n\n<p>Great catch is all I can say. Congratulations <a href=\"/seesee\">@seesee</a> 😃 </p>",
      "rawMarkdown": "&gt; I changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer.\n\nGreat catch is all I can say. Congratulations @seesee 😃 ",
      "votes": 1
    },
    {
      "id": 739421,
      "postDate": "2020-02-07T20:24:44.083Z",
      "content": "<p><a href=\"/seesee\">@seesee</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "rawMarkdown": "@seesee, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
      "votes": 2
    },
    {
      "id": 730804,
      "postDate": "2020-01-28T00:11:15.623Z",
      "content": "<p>Thanks for sharing the code <a href=\"/seesee\">@seesee</a> ! I've been having some fun reproducing your results. Was hoping to ask a few questions about it:</p>\n\n<ol>\n<li><p>Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes. Is that better than selecting the relevant long span among the candidates?</p></li>\n<li><p>Your loss function (below) has a different weight for end_loss than the other components. Did you arrive at this through experimentation? Did it outperform equal-weight?</p></li>\n</ol>\n\n<p><code>\nloss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss) / 2.0) +\n                tf.reduce_mean(long_loss)) / 2.0\n</code></p>\n\n<ol>\n<li>Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?</li>\n</ol>\n\n<p>Thanks,\nAlon</p>",
      "rawMarkdown": "Thanks for sharing the code @seesee ! I've been having some fun reproducing your results. Was hoping to ask a few questions about it:\n\n1. Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes. Is that better than selecting the relevant long span among the candidates?\n\n2. Your loss function (below) has a different weight for end_loss than the other components. Did you arrive at this through experimentation? Did it outperform equal-weight?\n \n```\nloss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss) / 2.0) +\n                tf.reduce_mean(long_loss)) / 2.0\n```\n\n3. Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?\n\nThanks,\nAlon\n\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 731016,
          "postDate": "2020-01-28T08:54:20.183Z",
          "content": "<p>1)\n&gt; Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes.</p>\n\n<p>Yes, the crop length is set so that we get the desired ratio between empty and non-empty training samples. By using crops we remove text that would result in empty samples.</p>\n\n<p>&gt; Is that better than selecting the relevant long span among the candidates?</p>\n\n<p>This was not tested. I don't know.</p>\n\n<p>2) Thank you for writing this post! It's actually a bug. At least, it's not what I thought I was using as loss. My intention was:\n```</p>\n\n<h1>fixed loss</h1>\n\n<p>loss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss)) / 2.0 +\n                tf.reduce_mean(long_loss)) / 2.0\n<code>``\nThis is a better approximation of the metric. Long and short answers should be weighted equally! I think it's better than taking the</code>mean` of all 3 loss terms (untested). After retraining with this fixed loss, public and private LB improve (entries are sorted by private LB but truncated, thus both show 0.71):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fd82197c21fe72cbe2171f7c3a864080f%2Ffix_loss.png?generation=1580200920211234&amp;alt=media\" alt=\"\"></p>\n\n<p>&gt; Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?</p>\n\n<p>[1, 2, 3, 4] were tried and 2 gave the best local LB. Personally, I would not do it. It can help, but there are more important hyper parameters (e.g. sample ratios).</p>",
          "rawMarkdown": "1)\n&gt; Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes.\n\nYes, the crop length is set so that we get the desired ratio between empty and non-empty training samples. By using crops we remove text that would result in empty samples.\n\n&gt; Is that better than selecting the relevant long span among the candidates?\n\nThis was not tested. I don't know.\n\n2) Thank you for writing this post! It's actually a bug. At least, it's not what I thought I was using as loss. My intention was:\n```\n# fixed loss\nloss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss)) / 2.0 +\n                tf.reduce_mean(long_loss)) / 2.0\n```\nThis is a better approximation of the metric. Long and short answers should be weighted equally! I think it's better than taking the `mean` of all 3 loss terms (untested). After retraining with this fixed loss, public and private LB improve (entries are sorted by private LB but truncated, thus both show 0.71):\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fd82197c21fe72cbe2171f7c3a864080f%2Ffix_loss.png?generation=1580200920211234&amp;alt=media)\n\n&gt; Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?\n\n[1, 2, 3, 4] were tried and 2 gave the best local LB. Personally, I would not do it. It can help, but there are more important hyper parameters (e.g. sample ratios).",
          "votes": 1
        },
        {
          "id": 731487,
          "postDate": "2020-01-28T17:31:40.280Z",
          "content": "<blockquote>\n  <p>After retraining with this fixed loss, public and private LB improve</p>\n</blockquote>\n\n<p>Haha... If only we'd partnered up, you might have made an extra few $... </p>\n\n<p>Thanks again for the code and answers.</p>\n\n<p>-Alon</p>",
          "rawMarkdown": "&gt; After retraining with this fixed loss, public and private LB improve\n\nHaha... If only we'd partnered up, you might have made an extra few $... \n\nThanks again for the code and answers.\n\n-Alon"
        }
      ]
    },
    {
      "id": 727375,
      "postDate": "2020-01-23T17:07:43.187Z",
      "content": "<p>Congrats to the double win with regular prize + TF2.0 prize! Such a strong performance with a single model. Thanks for sharing the code so fast.</p>\n\n<p>Also thanks for helping us debug the kernel prediction error in last week.</p>\n\n<blockquote>\n  <p>It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.</p>\n</blockquote>\n\n<p>↑ Cannot agree anymore.</p>\n\n<blockquote>\n  <p>I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.</p>\n</blockquote>\n\n<p>↑ We also locked submission a fews days before deadline. I guess those TPUs can really dry up one's ideas quickly huh? We then watched us dropping every day to 24th place on public LB during last week. Not the most fun experience.</p>",
      "rawMarkdown": "Congrats to the double win with regular prize + TF2.0 prize! Such a strong performance with a single model. Thanks for sharing the code so fast.\n\nAlso thanks for helping us debug the kernel prediction error in last week.\n\n\n&gt; It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.\n\n↑ Cannot agree anymore.\n\n\n&gt; I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.\n\n↑ We also locked submission a fews days before deadline. I guess those TPUs can really dry up one's ideas quickly huh? We then watched us dropping every day to 24th place on public LB during last week. Not the most fun experience.",
      "votes": 2
    },
    {
      "id": 727282,
      "postDate": "2020-01-23T15:50:25.103Z",
      "content": "<p>Thanks and congrats for the result!  It is funny, I was discussing with Dieter about what we could have done better and my first thing was to revisit sampling of empty windows ;)  </p>",
      "rawMarkdown": "Thanks and congrats for the result!  It is funny, I was discussing with Dieter about what we could have done better and my first thing was to revisit sampling of empty windows ;)  ",
      "votes": 2
    },
    {
      "id": 733296,
      "postDate": "2020-01-30T23:53:07.030Z",
      "content": "<p>Congrats 🎉  Thank you for writing up &amp; sharing code!</p>",
      "rawMarkdown": "Congrats 🎉  Thank you for writing up &amp; sharing code!"
    },
    {
      "id": 732664,
      "postDate": "2020-01-30T04:11:37.833Z",
      "content": "<p>This is a very good explanation</p>",
      "rawMarkdown": "This is a very good explanation"
    },
    {
      "id": 732388,
      "postDate": "2020-01-29T19:03:59.947Z",
      "content": "<p>This is perhaps a naive question about adding tokens to a Bert vocabulary:\nIn the <a href=\"https://huggingface.co/transformers/main_classes/tokenizer.html#transformers.PreTrainedTokenizer.add_special_tokens\">Huggingface documentation</a>, you add special tokens as follows:</p>\n\n<p><code>\nnum_added_toks = tokenizer.add_special_tokens(special_tokens_dict)\nprint('We have added', num_added_toks, 'tokens')\nmodel.resize_token_embeddings(len(tokenizer))  # Notice: resize_token_embeddings expect to receive the full size of the new vocabulary, i.e. the length of the tokenizer.\n</code>\nIn this winning solution, I didn't find a call to model.resize_token_embeddings (which is not yet implemented by Huggingface for TF models). Don't we need it? Is the Huggingface documentation wrong?</p>",
      "rawMarkdown": "This is perhaps a naive question about adding tokens to a Bert vocabulary:\nIn the [Huggingface documentation](https://huggingface.co/transformers/main_classes/tokenizer.html#transformers.PreTrainedTokenizer.add_special_tokens), you add special tokens as follows:\n\n```\nnum_added_toks = tokenizer.add_special_tokens(special_tokens_dict)\nprint('We have added', num_added_toks, 'tokens')\nmodel.resize_token_embeddings(len(tokenizer))  # Notice: resize_token_embeddings expect to receive the full size of the new vocabulary, i.e. the length of the tokenizer.\n```\nIn this winning solution, I didn't find a call to model.resize\\_token\\_embeddings (which is not yet implemented by Huggingface for TF models). Don't we need it? Is the Huggingface documentation wrong?"
    },
    {
      "id": 731853,
      "postDate": "2020-01-29T06:18:38.160Z",
      "content": "<p>Great stuff!</p>",
      "rawMarkdown": "Great stuff!\n"
    },
    {
      "id": 731084,
      "postDate": "2020-01-28T10:37:13.637Z",
      "content": "<p>Since I locked my submission some days ago, I am only able to checkin today. Congrats <a href=\"/seesee\">@seesee</a> and well done.</p>\n\n<p>Thanks so much for sharing your solution and code. I will try and reproduce you result in a few days. If I have any question then, I will be back to ask here.</p>",
      "rawMarkdown": "Since I locked my submission some days ago, I am only able to checkin today. Congrats @seesee and well done.\n\nThanks so much for sharing your solution and code. I will try and reproduce you result in a few days. If I have any question then, I will be back to ask here."
    },
    {
      "id": 729176,
      "postDate": "2020-01-25T21:14:13.683Z",
      "content": "<p>Congratulations and thanks for providing a repository with codes, this really helps!</p>",
      "rawMarkdown": "Congratulations and thanks for providing a repository with codes, this really helps!"
    },
    {
      "id": 728513,
      "postDate": "2020-01-24T20:32:31.927Z",
      "content": "<p>Congrats &amp; Thanks for sharing your great solusions🎉 😄 👍 </p>",
      "rawMarkdown": "Congrats &amp; Thanks for sharing your great solusions🎉 😄 👍 "
    },
    {
      "id": 727495,
      "postDate": "2020-01-23T18:46:13.937Z",
      "content": "<p>Congrats <a href=\"/seesee\">@seesee</a> and thanks for sharing! Did you totally disregard the answer type classification?</p>\n\n<p>&gt; Adding HTML tags as custom tokens helped a bit</p>\n\n<p>↑ I totally missed this by removing the HTML tags and setting the long answer target position onto the [SEP] token, which performs very poorly, even when there is only the long answer head. Intuitively using separate models for short/long could get even higher?</p>",
      "rawMarkdown": "Congrats @seesee and thanks for sharing! Did you totally disregard the answer type classification?\n\n&gt; Adding HTML tags as custom tokens helped a bit\n\n↑ I totally missed this by removing the HTML tags and setting the long answer target position onto the [SEP] token, which performs very poorly, even when there is only the long answer head. Intuitively using separate models for short/long could get even higher?",
      "replies": [
        {
          "id": 728008,
          "postDate": "2020-01-24T10:29:18.333Z",
          "content": "<blockquote>\n  <p>Intuitively using separate models for short/long could get even higher?</p>\n</blockquote>\n\n<p>Maybe, I prefer a single model and multiple targets. It's easier to tune, requires less code and less training time.</p>",
          "rawMarkdown": "&gt; Intuitively using separate models for short/long could get even higher?\n\nMaybe, I prefer a single model and multiple targets. It's easier to tune, requires less code and less training time."
        }
      ]
    },
    {
      "id": 727107,
      "postDate": "2020-01-23T13:16:29.217Z",
      "content": "<p>Fantastic , kudos :-)</p>",
      "rawMarkdown": "Fantastic , kudos :-)"
    },
    {
      "id": 727097,
      "postDate": "2020-01-23T13:06:31.043Z",
      "content": "<p>Congrats mate <a href=\"/seesee\">@seesee</a>! Just had a look at your code, a very clean code!\nSeems like you have played around with RoBERTa, how was the performance? I had issues recalculating the BPE tokenization using longest-common-substring algorithm. Wondering how was your experience.</p>",
      "rawMarkdown": "Congrats mate @seesee! Just had a look at your code, a very clean code!\nSeems like you have played around with RoBERTa, how was the performance? I had issues recalculating the BPE tokenization using longest-common-substring algorithm. Wondering how was your experience.",
      "replies": [
        {
          "id": 727119,
          "postDate": "2020-01-23T13:24:29.830Z",
          "content": "<p>Thanks. That's a good question. I should have added a <em>What did not work</em> section. You should be able to use the <code>--model_name_or_path roberta-large</code> flag for the training script. However, the performance was just like <code>bert-large-uncased</code> (i.e. ~0.70 LB). As I saw no gains from using this model I switched back to BERT. I still have no real explanation why Roberta didn't perform a lot better.</p>",
          "rawMarkdown": "Thanks. That's a good question. I should have added a *What did not work* section. You should be able to use the `--model_name_or_path roberta-large` flag for the training script. However, the performance was just like `bert-large-uncased` (i.e. ~0.70 LB). As I saw no gains from using this model I switched back to BERT. I still have no real explanation why Roberta didn't perform a lot better.",
          "votes": 2
        }
      ]
    },
    {
      "id": 727050,
      "postDate": "2020-01-23T12:09:51.793Z",
      "content": "<p>Congratulations <a href=\"/seesee\">@seesee</a>. This approach is quite different from which we were using. Thanks for sharing :)</p>",
      "rawMarkdown": "Congratulations @seesee. This approach is quite different from which we were using. Thanks for sharing :)"
    },
    {
      "id": 727045,
      "postDate": "2020-01-23T12:04:51.887Z",
      "content": "<p>Congratulations! It is out of my expect that your approach is so simple but works well.</p>",
      "rawMarkdown": "Congratulations! It is out of my expect that your approach is so simple but works well.",
      "replies": [
        {
          "id": 727089,
          "postDate": "2020-01-23T12:59:37.677Z",
          "content": "<p>Thanks </p>\n\n<p>&gt; Build a model is very simple, but build a simple model is the hardest thing there is -- Guanshuo Xu</p>\n\n<p>🤓 </p>",
          "rawMarkdown": "Thanks \n\n&gt; Build a model is very simple, but build a simple model is the hardest thing there is -- Guanshuo Xu\n\n🤓 ",
          "votes": 1
        },
        {
          "id": 727222,
          "postDate": "2020-01-23T14:38:27.720Z",
          "content": "<p>Hi I have one question: if you don't have threshold, then you will have only 2 situations: blank for both short and long OR slice for both short and long. Right?</p>",
          "rawMarkdown": "Hi I have one question: if you don't have threshold, then you will have only 2 situations: blank for both short and long OR slice for both short and long. Right?"
        },
        {
          "id": 727335,
          "postDate": "2020-01-23T16:31:20.973Z",
          "content": "<p>No. If we ignore <code>YES/NO</code> answers there are 4 (short, long) possibilities: (empty, empty), (empty, span), (span, empty), (span, span). They can all be predicted by the model.</p>",
          "rawMarkdown": "No. If we ignore `YES/NO` answers there are 4 (short, long) possibilities: (empty, empty), (empty, span), (span, empty), (span, span). They can all be predicted by the model."
        },
        {
          "id": 727382,
          "postDate": "2020-01-23T17:16:36.443Z",
          "content": "<p>Thanks.\n<code>\n    min_dist = 1_000_000\n    if long_token != -1:\n      num_searched_long += 1\n      for candidate in candidates[example_id]:\n        cstart, cend = candidate.start_token, candidate.end_token\n        dist = abs(cstart - long_token)\n        if dist &amp;lt; min_dist:\n          min_dist = dist\n        if long_token == cstart:\n          long_answer = f'{cstart}:{cend}'\n          found_long = True\n          break\n</code>\nhave you tried using condition \n<code>cstart&lt;=long_token&lt;=cend</code>\nbesides \n<code>ong_token == cstart</code>? Is score worse?</p>",
          "rawMarkdown": "Thanks.\n```\n    min_dist = 1_000_000\n    if long_token != -1:\n      num_searched_long += 1\n      for candidate in candidates[example_id]:\n        cstart, cend = candidate.start_token, candidate.end_token\n        dist = abs(cstart - long_token)\n        if dist &lt; min_dist:\n          min_dist = dist\n        if long_token == cstart:\n          long_answer = f'{cstart}:{cend}'\n          found_long = True\n          break\n```\nhave you tried using condition \n```cstart&lt;=long_token&lt;=cend ```\nbesides \n```ong_token == cstart```? Is score worse?"
        },
        {
          "id": 728003,
          "postDate": "2020-01-24T10:20:19.043Z",
          "content": "<p>I didn't test it, but I am quite sure that the score won't change. I'd be surprised if it would get better. I checked that the model always predicts one of the opening tags or the CLS token. <code>long_token == cstart</code> should be fine.</p>",
          "rawMarkdown": "I didn't test it, but I am quite sure that the score won't change. I'd be surprised if it would get better. I checked that the model always predicts one of the opening tags or the CLS token. `long_token == cstart` should be fine."
        }
      ]
    },
    {
      "id": 727004,
      "postDate": "2020-01-23T11:16:16.970Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 727920,
      "postDate": "2020-01-24T07:44:20.570Z",
      "content": "<p>Congrats! Thanks for sharing.</p>",
      "rawMarkdown": "Congrats! Thanks for sharing.",
      "votes": 1
    },
    {
      "id": 728691,
      "postDate": "2020-01-25T05:07:50.837Z",
      "content": "<p>Thank you for your sharing!</p>",
      "rawMarkdown": "Thank you for your sharing!"
    },
    {
      "id": 727025,
      "postDate": "2020-01-23T11:46:01.343Z",
      "content": "<p>Congrats! Thanks for sharing.</p>",
      "rawMarkdown": "Congrats! Thanks for sharing."
    }
  ],
  "comments": [
    {
      "id": 726994,
      "author_name": "Dieter",
      "author_url": "",
      "post_date": "2020-01-23T11:04:16.027000",
      "content": "<p>we went with </p>\n\n<blockquote>\n  <p>taken days on my local 2 x 1080 TI machine</p>\n</blockquote>",
      "votes": 3,
      "replies": [
        {
          "id": 727000,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-23T11:11:28.860000",
          "content": "<p>haha, I ran out of hyperparameters and ideas 😜</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 734457,
      "author_name": "Eduardo Rocha de Andrade",
      "author_url": "",
      "post_date": "2020-02-01T14:08:53.620000",
      "content": "<p>Congratulations See!!! You derserve it =]</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 731403,
      "author_name": "Sumit Mishra",
      "author_url": "",
      "post_date": "2020-01-28T15:43:54.543000",
      "content": "<p>Congratulations Mate.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 727510,
      "author_name": "Ken Krige",
      "author_url": "",
      "post_date": "2020-01-23T18:59:17.143000",
      "content": "<p>Your approach is clean and innovative. I am really enjoying reading it. Thank you so much for sharing your code. You're a good example of what kaggle is all about and you deserve to win the TF2.0 prize as well.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 728000,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-24T10:17:22.790000",
          "content": "<p>Thanks for the kind words 😄 </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727327,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2020-01-23T16:25:33.677000",
      "content": "<p>I have two questions. :)</p>\n\n<ol>\n<li><p>Could you please share your long answer F1 and short answer F1 if you evaluated them separately in local validation? I'm wondering how much the <code>long_outputs</code> helps. </p></li>\n<li><p>For the <code>long_logits</code>, did you treat it as a binary classification obj? Whether the token is part of the long answer or not?</p></li>\n</ol>",
      "votes": 1,
      "replies": [
        {
          "id": 727353,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-23T16:40:47.217000",
          "content": "<p>1 Yes I do. The scores are (concat F1, short F1, long F1) and local score is:\n* With ~3500 empty and ~3000 non-empty validation samples: (0.563, 0.481, 0.615)\n* Only ~3000 non-empty: (0.722, 0.576, 0.823)</p>\n\n<p>Maybe check this script if you are using the same metric: <a href=\"https://github.com/see--/natural-question-answering/blob/master/eval_server.py\">https://github.com/see--/natural-question-answering/blob/master/eval_server.py</a></p>\n\n<p>The validation set is here:\n<a href=\"https://github.com/see--/natural-question-answering/blob/master/val_ids.csv\">https://github.com/see--/natural-question-answering/blob/master/val_ids.csv</a></p>\n\n<p>2 No, I treat it as softmax classification. Whether the token is the opening HTML tag and predict [CLS] token if there is no long answer.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 727516,
          "author_name": "Ken Krige",
          "author_url": "",
          "post_date": "2020-01-23T19:04:15.243000",
          "content": "<blockquote>\n  <p>2 No, I treat it as softmax classification. Whether the token is the opening HTML tag and predict [CLS] token if there is no long answer.</p>\n</blockquote>\n\n<p>Clever. This and your sampling ratios are big learning points for me so far.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 727308,
      "author_name": "Jiwei Liu",
      "author_url": "",
      "post_date": "2020-01-23T16:11:00.210000",
      "content": "<p>Big congrats! Thank you so much for the amazing sharing of the source code.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 727204,
      "author_name": "Abhishek Thakur",
      "author_url": "",
      "post_date": "2020-01-23T14:29:53.283000",
      "content": "<p>Great Solution! Thank you for sharing the code!! :)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 727176,
      "author_name": "Giba",
      "author_url": "",
      "post_date": "2020-01-23T14:10:41.140000",
      "content": "<p>Congrats, nice write-up and thanks for sharing the code!</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 726987,
      "author_name": "Mourad",
      "author_url": "",
      "post_date": "2020-01-23T10:54:34.120000",
      "content": "<p>&gt; I changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer.</p>\n\n<p>Great catch is all I can say. Congratulations <a href=\"/seesee\">@seesee</a> 😃 </p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 739421,
      "author_name": "Vitalii Mokin",
      "author_url": "",
      "post_date": "2020-02-07T20:24:44.083000",
      "content": "<p><a href=\"/seesee\">@seesee</a>, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (<a href=\"https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques\">https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques</a>)</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 730804,
      "author_name": "Alon Bochman",
      "author_url": "",
      "post_date": "2020-01-28T00:11:15.623000",
      "content": "<p>Thanks for sharing the code <a href=\"/seesee\">@seesee</a> ! I've been having some fun reproducing your results. Was hoping to ask a few questions about it:</p>\n\n<ol>\n<li><p>Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes. Is that better than selecting the relevant long span among the candidates?</p></li>\n<li><p>Your loss function (below) has a different weight for end_loss than the other components. Did you arrive at this through experimentation? Did it outperform equal-weight?</p></li>\n</ol>\n\n<p><code>\nloss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss) / 2.0) +\n                tf.reduce_mean(long_loss)) / 2.0\n</code></p>\n\n<ol>\n<li>Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?</li>\n</ol>\n\n<p>Thanks,\nAlon</p>",
      "votes": 2,
      "replies": [
        {
          "id": 731016,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-28T08:54:20.183000",
          "content": "<p>1)\n&gt; Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes.</p>\n\n<p>Yes, the crop length is set so that we get the desired ratio between empty and non-empty training samples. By using crops we remove text that would result in empty samples.</p>\n\n<p>&gt; Is that better than selecting the relevant long span among the candidates?</p>\n\n<p>This was not tested. I don't know.</p>\n\n<p>2) Thank you for writing this post! It's actually a bug. At least, it's not what I thought I was using as loss. My intention was:\n```</p>\n\n<h1>fixed loss</h1>\n\n<p>loss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss)) / 2.0 +\n                tf.reduce_mean(long_loss)) / 2.0\n<code>``\nThis is a better approximation of the metric. Long and short answers should be weighted equally! I think it's better than taking the</code>mean` of all 3 loss terms (untested). After retraining with this fixed loss, public and private LB improve (entries are sorted by private LB but truncated, thus both show 0.71):</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F712087%2Fd82197c21fe72cbe2171f7c3a864080f%2Ffix_loss.png?generation=1580200920211234&amp;alt=media\" alt=\"\"></p>\n\n<p>&gt; Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?</p>\n\n<p>[1, 2, 3, 4] were tried and 2 gave the best local LB. Personally, I would not do it. It can help, but there are more important hyper parameters (e.g. sample ratios).</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 731487,
          "author_name": "Alon Bochman",
          "author_url": "",
          "post_date": "2020-01-28T17:31:40.280000",
          "content": "<blockquote>\n  <p>After retraining with this fixed loss, public and private LB improve</p>\n</blockquote>\n\n<p>Haha... If only we'd partnered up, you might have made an extra few $... </p>\n\n<p>Thanks again for the code and answers.</p>\n\n<p>-Alon</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727375,
      "author_name": "Bo",
      "author_url": "",
      "post_date": "2020-01-23T17:07:43.187000",
      "content": "<p>Congrats to the double win with regular prize + TF2.0 prize! Such a strong performance with a single model. Thanks for sharing the code so fast.</p>\n\n<p>Also thanks for helping us debug the kernel prediction error in last week.</p>\n\n<blockquote>\n  <p>It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.</p>\n</blockquote>\n\n<p>↑ Cannot agree anymore.</p>\n\n<blockquote>\n  <p>I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.</p>\n</blockquote>\n\n<p>↑ We also locked submission a fews days before deadline. I guess those TPUs can really dry up one's ideas quickly huh? We then watched us dropping every day to 24th place on public LB during last week. Not the most fun experience.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 727282,
      "author_name": "CPMP",
      "author_url": "",
      "post_date": "2020-01-23T15:50:25.103000",
      "content": "<p>Thanks and congrats for the result!  It is funny, I was discussing with Dieter about what we could have done better and my first thing was to revisit sampling of empty windows ;)  </p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 733296,
      "author_name": "corochann",
      "author_url": "",
      "post_date": "2020-01-30T23:53:07.030000",
      "content": "<p>Congrats 🎉  Thank you for writing up &amp; sharing code!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732664,
      "author_name": "Sapharn",
      "author_url": "",
      "post_date": "2020-01-30T04:11:37.833000",
      "content": "<p>This is a very good explanation</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 732388,
      "author_name": "Alon Bochman",
      "author_url": "",
      "post_date": "2020-01-29T19:03:59.947000",
      "content": "<p>This is perhaps a naive question about adding tokens to a Bert vocabulary:\nIn the <a href=\"https://huggingface.co/transformers/main_classes/tokenizer.html#transformers.PreTrainedTokenizer.add_special_tokens\">Huggingface documentation</a>, you add special tokens as follows:</p>\n\n<p><code>\nnum_added_toks = tokenizer.add_special_tokens(special_tokens_dict)\nprint('We have added', num_added_toks, 'tokens')\nmodel.resize_token_embeddings(len(tokenizer))  # Notice: resize_token_embeddings expect to receive the full size of the new vocabulary, i.e. the length of the tokenizer.\n</code>\nIn this winning solution, I didn't find a call to model.resize_token_embeddings (which is not yet implemented by Huggingface for TF models). Don't we need it? Is the Huggingface documentation wrong?</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 731853,
      "author_name": "Sumit Srivastava",
      "author_url": "",
      "post_date": "2020-01-29T06:18:38.160000",
      "content": "<p>Great stuff!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 731084,
      "author_name": "YaGana Sheriff-Hussaini",
      "author_url": "",
      "post_date": "2020-01-28T10:37:13.637000",
      "content": "<p>Since I locked my submission some days ago, I am only able to checkin today. Congrats <a href=\"/seesee\">@seesee</a> and well done.</p>\n\n<p>Thanks so much for sharing your solution and code. I will try and reproduce you result in a few days. If I have any question then, I will be back to ask here.</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 729176,
      "author_name": "DimitreOliveira",
      "author_url": "",
      "post_date": "2020-01-25T21:14:13.683000",
      "content": "<p>Congratulations and thanks for providing a repository with codes, this really helps!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 728513,
      "author_name": "Miyabon",
      "author_url": "",
      "post_date": "2020-01-24T20:32:31.927000",
      "content": "<p>Congrats &amp; Thanks for sharing your great solusions🎉 😄 👍 </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727495,
      "author_name": "Zhengkai Tu",
      "author_url": "",
      "post_date": "2020-01-23T18:46:13.937000",
      "content": "<p>Congrats <a href=\"/seesee\">@seesee</a> and thanks for sharing! Did you totally disregard the answer type classification?</p>\n\n<p>&gt; Adding HTML tags as custom tokens helped a bit</p>\n\n<p>↑ I totally missed this by removing the HTML tags and setting the long answer target position onto the [SEP] token, which performs very poorly, even when there is only the long answer head. Intuitively using separate models for short/long could get even higher?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 728008,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-24T10:29:18.333000",
          "content": "<blockquote>\n  <p>Intuitively using separate models for short/long could get even higher?</p>\n</blockquote>\n\n<p>Maybe, I prefer a single model and multiple targets. It's easier to tune, requires less code and less training time.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727107,
      "author_name": "aintnosunshine",
      "author_url": "",
      "post_date": "2020-01-23T13:16:29.217000",
      "content": "<p>Fantastic , kudos :-)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727097,
      "author_name": "Kornesh Kanan",
      "author_url": "",
      "post_date": "2020-01-23T13:06:31.043000",
      "content": "<p>Congrats mate <a href=\"/seesee\">@seesee</a>! Just had a look at your code, a very clean code!\nSeems like you have played around with RoBERTa, how was the performance? I had issues recalculating the BPE tokenization using longest-common-substring algorithm. Wondering how was your experience.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 727119,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-23T13:24:29.830000",
          "content": "<p>Thanks. That's a good question. I should have added a <em>What did not work</em> section. You should be able to use the <code>--model_name_or_path roberta-large</code> flag for the training script. However, the performance was just like <code>bert-large-uncased</code> (i.e. ~0.70 LB). As I saw no gains from using this model I switched back to BERT. I still have no real explanation why Roberta didn't perform a lot better.</p>",
          "votes": 2,
          "replies": []
        }
      ]
    },
    {
      "id": 727050,
      "author_name": "Ram Ramrakhya",
      "author_url": "",
      "post_date": "2020-01-23T12:09:51.793000",
      "content": "<p>Congratulations <a href=\"/seesee\">@seesee</a>. This approach is quite different from which we were using. Thanks for sharing :)</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727045,
      "author_name": "SchenbergZ",
      "author_url": "",
      "post_date": "2020-01-23T12:04:51.887000",
      "content": "<p>Congratulations! It is out of my expect that your approach is so simple but works well.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 727089,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-23T12:59:37.677000",
          "content": "<p>Thanks </p>\n\n<p>&gt; Build a model is very simple, but build a simple model is the hardest thing there is -- Guanshuo Xu</p>\n\n<p>🤓 </p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 727222,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2020-01-23T14:38:27.720000",
          "content": "<p>Hi I have one question: if you don't have threshold, then you will have only 2 situations: blank for both short and long OR slice for both short and long. Right?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 727335,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-23T16:31:20.973000",
          "content": "<p>No. If we ignore <code>YES/NO</code> answers there are 4 (short, long) possibilities: (empty, empty), (empty, span), (span, empty), (span, span). They can all be predicted by the model.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 727382,
          "author_name": "SchenbergZ",
          "author_url": "",
          "post_date": "2020-01-23T17:16:36.443000",
          "content": "<p>Thanks.\n<code>\n    min_dist = 1_000_000\n    if long_token != -1:\n      num_searched_long += 1\n      for candidate in candidates[example_id]:\n        cstart, cend = candidate.start_token, candidate.end_token\n        dist = abs(cstart - long_token)\n        if dist &amp;lt; min_dist:\n          min_dist = dist\n        if long_token == cstart:\n          long_answer = f'{cstart}:{cend}'\n          found_long = True\n          break\n</code>\nhave you tried using condition \n<code>cstart&lt;=long_token&lt;=cend</code>\nbesides \n<code>ong_token == cstart</code>? Is score worse?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 728003,
          "author_name": "See--",
          "author_url": "",
          "post_date": "2020-01-24T10:20:19.043000",
          "content": "<p>I didn't test it, but I am quite sure that the score won't change. I'd be surprised if it would get better. I checked that the model always predicts one of the opening tags or the CLS token. <code>long_token == cstart</code> should be fine.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 727004,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-01-23T11:16:16.970000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727920,
      "author_name": "Marek Nurzynski",
      "author_url": "",
      "post_date": "2020-01-24T07:44:20.570000",
      "content": "<p>Congrats! Thanks for sharing.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 728691,
      "author_name": "Takuya Makiyama",
      "author_url": "",
      "post_date": "2020-01-25T05:07:50.837000",
      "content": "<p>Thank you for your sharing!</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 727025,
      "author_name": "Shayekh Islam",
      "author_url": "",
      "post_date": "2020-01-23T11:46:01.343000",
      "content": "<p>Congrats! Thanks for sharing.</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "726961": "Good evening,\n\nfirst of all, I'd like to thank Kaggle and the hosts for this awesome challenge! It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.\n\nSecondly, congrats to all the winners. I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.\n\nMy solution is just a single TF2.0 model. It uses custom heads and a BERT transformers backbone (large version). For modeling and training I am using the great [transformers](https://github.com/huggingface/transformers) library. I think that the [following snippet](https://github.com/see--/natural-question-answering/blob/master/models.py#L8-L28) is useful to understand the modeling:\n```python\nclass TFBertForNaturalQuestionAnswering(TFBertPreTrainedModel):\n    def __init__(self, config, *inputs, **kwargs):\n        super().__init__(config, *inputs, **kwargs)\n        self.num_labels = config.num_labels\n\n        self.bert = TFBertMainLayer(config, name='bert')\n        self.initializer = get_initializer(config.initializer_range)\n        self.qa_outputs = L.Dense(config.num_labels,\n            kernel_initializer=self.initializer, name='qa_outputs')\n        self.long_outputs = L.Dense(1, kernel_initializer=self.initializer,\n            name='long_outputs')\n\n    def call(self, inputs, **kwargs):\n        outputs = self.bert(inputs, **kwargs)\n        sequence_output = outputs[0]\n        logits = self.qa_outputs(sequence_output)\n        start_logits, end_logits = tf.split(logits, 2, axis=-1)\n        start_logits = tf.squeeze(start_logits, -1)\n        end_logits = tf.squeeze(end_logits, -1)\n        long_logits = tf.squeeze(self.long_outputs(sequence_output), -1)\n        return start_logits, end_logits, long_logits\n```\n\nAs you can see, the natural question answering task can be treated like SQUAD-2.0 with an additional head for long answers. Note that we just need a single output: The opening tag of the HTML bounding box. I guess most competitors used a similar modeling so I think what made the difference to most other solutions is the sampling.\n\nI changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer. I started with a really low empty answer ratio which I got from the [`bert-joint` paper](https://arxiv.org/abs/1901.08634), but I couldn't reach a good score. I tuned a few hyper parameters but overall I got good results with a wide range of parameters. Adding HTML tags as custom tokens helped a bit. I also tried different start weights and found that:\n\n`bert-large-uncased` (~0.70 LB) &lt; `bert-large-uncased-whole-word-masking` (~0.72 LB) &lt; `bert-large-uncased-whole-word-masking-finetuned-squad` (~0.73 LB).\n\nThat's about it. Thanks to @boliu0, @christofhenkel and @kentaronakanishi for fixing and providing the metric!\n\nPlease refer to my repository for implementation details and instructions to reproduce:\n* https://github.com/see--/natural-question-answering\n\nYou can find the 2nd place kernel and pretrained weights on Kaggle:\n* https://www.kaggle.com/seesee/submit-full\n* https://www.kaggle.com/seesee/nq-bert-uncased-68\n\nFeel free to ask questions and / or create GitHub issues.",
    "726994": "we went with \n&gt; taken days on my local 2 x 1080 TI machine",
    "734457": "Congratulations See!!! You derserve it =]",
    "731403": "Congratulations Mate.",
    "727510": "Your approach is clean and innovative. I am really enjoying reading it. Thank you so much for sharing your code. You're a good example of what kaggle is all about and you deserve to win the TF2.0 prize as well.",
    "727327": "I have two questions. :)\n\n1. Could you please share your long answer F1 and short answer F1 if you evaluated them separately in local validation? I'm wondering how much the `long_outputs` helps. \n\n2. For the `long_logits`, did you treat it as a binary classification obj? Whether the token is part of the long answer or not?",
    "727308": "Big congrats! Thank you so much for the amazing sharing of the source code.",
    "727204": "Great Solution! Thank you for sharing the code!! :)",
    "727176": "Congrats, nice write-up and thanks for sharing the code!",
    "726987": "&gt; I changed the empty answer ratio so that it is similar to the full dataset. I.e. roughly as many empty answers as answers with a long answer.\n\nGreat catch is all I can say. Congratulations @seesee 😃 ",
    "739421": "@seesee, congratulations! Thanks for sharing your approaches and magic! I added your post to my collection of the best Kaggle kernels and posts of winners of NLP Prize Competitions: [Data Science with DL &amp; NLP: Advanced Techniques] (https://www.kaggle.com/vbmokin/data-science-with-dl-nlp-advanced-techniques)",
    "730804": "Thanks for sharing the code @seesee ! I've been having some fun reproducing your results. Was hoping to ask a few questions about it:\n\n1. Can you explain what the code is doing with crops? Seems like you are creating a text window around the true short answer for training purposes. Is that better than selecting the relevant long span among the candidates?\n\n2. Your loss function (below) has a different weight for end_loss than the other components. Did you arrive at this through experimentation? Did it outperform equal-weight?\n \n```\nloss = ((tf.reduce_mean(start_loss) + tf.reduce_mean(end_loss) / 2.0) +\n                tf.reduce_mean(long_loss)) / 2.0\n```\n\n3. Why did you choose to train for exactly 2 epochs? Would you recommend using validation every X steps as a stopping condition?\n\nThanks,\nAlon\n\n\n",
    "727375": "Congrats to the double win with regular prize + TF2.0 prize! Such a strong performance with a single model. Thanks for sharing the code so fast.\n\nAlso thanks for helping us debug the kernel prediction error in last week.\n\n\n&gt; It was really fun to work with TF2.0 and TPUs. What would have taken days on my local 2 x 1080 TI machine just took a couple of hours. For example, the actual training time for my final model (excluding tokenization and post-processing) is just a little more than 2 hours.\n\n↑ Cannot agree anymore.\n\n\n&gt; I locked my submissions a week ago with +0.04 on public but everyone kept improving. Maybe I should have had continued working on this challenged as well.\n\n↑ We also locked submission a fews days before deadline. I guess those TPUs can really dry up one's ideas quickly huh? We then watched us dropping every day to 24th place on public LB during last week. Not the most fun experience.",
    "727282": "Thanks and congrats for the result!  It is funny, I was discussing with Dieter about what we could have done better and my first thing was to revisit sampling of empty windows ;)  ",
    "733296": "Congrats 🎉  Thank you for writing up &amp; sharing code!",
    "732664": "This is a very good explanation",
    "732388": "This is perhaps a naive question about adding tokens to a Bert vocabulary:\nIn the [Huggingface documentation](https://huggingface.co/transformers/main_classes/tokenizer.html#transformers.PreTrainedTokenizer.add_special_tokens), you add special tokens as follows:\n\n```\nnum_added_toks = tokenizer.add_special_tokens(special_tokens_dict)\nprint('We have added', num_added_toks, 'tokens')\nmodel.resize_token_embeddings(len(tokenizer))  # Notice: resize_token_embeddings expect to receive the full size of the new vocabulary, i.e. the length of the tokenizer.\n```\nIn this winning solution, I didn't find a call to model.resize\\_token\\_embeddings (which is not yet implemented by Huggingface for TF models). Don't we need it? Is the Huggingface documentation wrong?",
    "731853": "Great stuff!\n",
    "731084": "Since I locked my submission some days ago, I am only able to checkin today. Congrats @seesee and well done.\n\nThanks so much for sharing your solution and code. I will try and reproduce you result in a few days. If I have any question then, I will be back to ask here.",
    "729176": "Congratulations and thanks for providing a repository with codes, this really helps!",
    "728513": "Congrats &amp; Thanks for sharing your great solusions🎉 😄 👍 ",
    "727495": "Congrats @seesee and thanks for sharing! Did you totally disregard the answer type classification?\n\n&gt; Adding HTML tags as custom tokens helped a bit\n\n↑ I totally missed this by removing the HTML tags and setting the long answer target position onto the [SEP] token, which performs very poorly, even when there is only the long answer head. Intuitively using separate models for short/long could get even higher?",
    "727107": "Fantastic , kudos :-)",
    "727097": "Congrats mate @seesee! Just had a look at your code, a very clean code!\nSeems like you have played around with RoBERTa, how was the performance? I had issues recalculating the BPE tokenization using longest-common-substring algorithm. Wondering how was your experience.",
    "727050": "Congratulations @seesee. This approach is quite different from which we were using. Thanks for sharing :)",
    "727045": "Congratulations! It is out of my expect that your approach is so simple but works well.",
    "727004": "",
    "727920": "Congrats! Thanks for sharing.",
    "728691": "Thank you for your sharing!",
    "727025": "Congrats! Thanks for sharing."
  }
}