{
  "id": 119957,
  "title": "Reproducing joint BERT baseline",
  "url": "/competitions/tensorflow2-question-answering/discussion/119957",
  "author_name": "",
  "post_date": "2019-12-02T19:38:03.001025Z",
  "votes": 21,
  "comment_count": 33,
  "views": 0,
  "content": "<p>As you might be aware the current baseline is based on this paper <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">https://arxiv.org/pdf/1901.08634.pdf</a> </p>\n\n<p>Nevertheless I have a hard time reproducing the result. For me some of the parameters seem contradictory. For example at one point they write \" found that training for 1 epoch with an initial learning rate of 0.005 was the best setting\" and at the related repo (<a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a> ) is written lr should be 3e-5 </p>",
  "messages": [
    {
      "id": "686108",
      "postDate": "12/02/2019 19:38:03",
      "content": "<p>As you might be aware the current baseline is based on this paper <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">https://arxiv.org/pdf/1901.08634.pdf</a> </p>\n\n<p>Nevertheless I have a hard time reproducing the result. For me some of the parameters seem contradictory. For example at one point they write \" found that training for 1 epoch with an initial learning rate of 0.005 was the best setting\" and at the related repo (<a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a> ) is written lr should be 3e-5 </p>",
      "rawMarkdown": "As you might be aware the current baseline is based on this paper https://arxiv.org/pdf/1901.08634.pdf \n\nNevertheless I have a hard time reproducing the result. For me some of the parameters seem contradictory. For example at one point they write \" found that training for 1 epoch with an initial learning rate of 0.005 was the best setting\" and at the related repo (https://github.com/google-research/language/tree/master/language/question_answering/bert_joint ) is written lr should be 3e-5",
      "votes": null
    },
    {
      "id": "686112",
      "postDate": "12/02/2019 19:45:30",
      "content": "<p>Seems like more fine tuning. The github page discusses more epochs too. What worked for you?</p>",
      "rawMarkdown": "Seems like more fine tuning. The github page discusses more epochs too. What worked for you?",
      "votes": null
    },
    {
      "id": "686116",
      "postDate": "12/02/2019 19:48:53",
      "content": "<p>I am still at 1st epoch. From my experience more epochs have little effect. Want to get 1st epoch to a 0.65+ range first </p>",
      "rawMarkdown": "I am still at 1st epoch. From my experience more epochs have little effect. Want to get 1st epoch to a 0.65+ range first",
      "votes": null
    },
    {
      "id": "686118",
      "postDate": "12/02/2019 19:51:46",
      "content": "<p>Since it seems you are working in pytorch, I’m assuming that you might have implemented the paper.\nCould you explain to me what does the author mean by the following paragraph: </p>\n\n<p><code>For each training instance we compute start and end token indices to represent the target answer span. If all annotated short spans are contained in the instance, we set the start and end target in- dices to point to the smallest span containing all the annotated short answer spans. If there are no annotated short spans but there is an annotated long answer span completely contained in the in- stance, we set the start and end target indices to point to the entire long answer span. If no short or\nlong span can be found in the current instance, we set the target start and end indices to point to the “[CLS]” token.</code></p>\n\n<p>I’m a bit confused to what the author is trying to convey here.</p>\n\n<p>Let’s say start index of document: 256, end index of document: 576. (Removing question tokens). \nStart index of long answer: 310, end index of long answer: 625\nStart index of short answer: 450, end index of short answer: 500.</p>\n\n<p>How would you label these here? Ignore the indices that are outside? Or keep only those answers that fall completely in the range we chose: 256-576 ?</p>",
      "rawMarkdown": "Since it seems you are working in pytorch, I’m assuming that you might have implemented the paper.\nCould you explain to me what does the author mean by the following paragraph: \n\n`For each training instance we compute start and end token indices to represent the target answer span. If all annotated short spans are contained in the instance, we set the start and end target in- dices to point to the smallest span containing all the annotated short answer spans. If there are no annotated short spans but there is an annotated long answer span completely contained in the in- stance, we set the start and end target indices to point to the entire long answer span. If no short or\nlong span can be found in the current instance, we set the target start and end indices to point to the “[CLS]” token. `\n\n\nI’m a bit confused to what the author is trying to convey here.\n\n\nLet’s say start index of document: 256, end index of document: 576. (Removing question tokens). \nStart index of long answer: 310, end index of long answer: 625\nStart index of short answer: 450, end index of short answer: 500.\n\nHow would you label these here? Ignore the indices that are outside? Or keep only those answers that fall completely in the range we chose: 256-576 ?",
      "votes": null
    },
    {
      "id": "686138",
      "postDate": "12/02/2019 20:24:31",
      "content": "<p>as far as I understand label would be 450:500, answer_type = short. Long answer is discarded as its end is out of the span. </p>",
      "rawMarkdown": "as far as I understand label would be 450:500, answer_type = short. Long answer is discarded as its end is out of the span.",
      "votes": null
    },
    {
      "id": "686141",
      "postDate": "12/02/2019 20:26:08",
      "content": "<p>Interesting are those windows which have no answers. According to the default params in repo, 2% of those no_answer windows are fed into training and 98% are ignored</p>",
      "rawMarkdown": "Interesting are those windows which have no answers. According to the default params in repo, 2% of those no_answer windows are fed into training and 98% are ignored",
      "votes": null
    },
    {
      "id": "686148",
      "postDate": "12/02/2019 20:30:53",
      "content": "<p>By that method, aren’t we also excluding (many?) short and long answers?</p>",
      "rawMarkdown": "By that method, aren’t we also excluding (many?) short and long answers?",
      "votes": null
    },
    {
      "id": "686152",
      "postDate": "12/02/2019 20:36:04",
      "content": "<p>No, since the 512 windows overlap (by default with 128) which is way longer than normal long/ short answers, the answers are contained in at least one. </p>",
      "rawMarkdown": "No, since the 512 windows overlap (by default with 128) which is way longer than normal long/ short answers, the answers are contained in at least one.",
      "votes": null
    },
    {
      "id": "686180",
      "postDate": "12/02/2019 21:27:15",
      "content": "<p>Which bert configuration and batch size are you using in your training?  On this page: </p>\n\n<p><a href=\"https://github.com/google-research/bert\">https://github.com/google-research/bert</a></p>\n\n<p>in the \"out of memory issues\" section, there is a statement that \"Unfortunately, these max batch sizes for BERT-Large are so small that they will actually harm the model accuracy, regardless of the learning rate used.\"</p>",
      "rawMarkdown": "Which bert configuration and batch size are you using in your training?  On this page: \n\nhttps://github.com/google-research/bert\n\nin the \"out of memory issues\" section, there is a statement that \"Unfortunately, these max batch sizes for BERT-Large are so small that they will actually harm the model accuracy, regardless of the learning rate used.\"",
      "votes": null
    },
    {
      "id": "686348",
      "postDate": "12/03/2019 03:48:46",
      "content": "<p>could you share your hardware details and how long it takes to run one epoch <a href=\"/christofhenkel\">@christofhenkel</a> </p>",
      "rawMarkdown": "could you share your hardware details and how long it takes to run one epoch @christofhenkel",
      "votes": null
    },
    {
      "id": "686407",
      "postDate": "12/03/2019 05:46:31",
      "content": "<p>15h GTX 1080 Ti</p>",
      "rawMarkdown": "15h GTX 1080 Ti",
      "votes": null
    },
    {
      "id": "686505",
      "postDate": "12/03/2019 08:26:30",
      "content": "<blockquote>\n  <p>for 1 epoch with an initial learning rate of 0.005 was the best setting</p>\n</blockquote>\n\n<p>I think it's just a mistake, maybe they confused 5e-3 and 3e-5. 0.005 is a too high learning rate for transformers.  </p>",
      "rawMarkdown": "&gt; for 1 epoch with an initial learning rate of 0.005 was the best setting\n\nI think it's just a mistake, maybe they confused 5e-3 and 3e-5. 0.005 is a too high learning rate for transformers.",
      "votes": null
    },
    {
      "id": "686513",
      "postDate": "12/03/2019 08:35:56",
      "content": "<p>thats my guess too. But there are more inconsistencies and intransparencies. For example related to optimizer used and if a warmup is applied</p>",
      "rawMarkdown": "thats my guess too. But there are more inconsistencies and intransparencies. For example related to optimizer used and if a warmup is applied",
      "votes": null
    },
    {
      "id": "686517",
      "postDate": "12/03/2019 08:40:32",
      "content": "<p>gradient accumulation helps</p>",
      "rawMarkdown": "gradient accumulation helps",
      "votes": null
    },
    {
      "id": "686520",
      "postDate": "12/03/2019 08:44:36",
      "content": "<p>As usual in academic papers :) It's also about version control - it's hard to align all changes in code with paragraphs in the paper describing the approach. \nps. didn't try to implement myself, sticking to Tf 1 for now. </p>",
      "rawMarkdown": "As usual in academic papers :) It's also about version control - it's hard to align all changes in code with paragraphs in the paper describing the approach. \nps. didn't try to implement myself, sticking to Tf 1 for now.",
      "votes": null
    },
    {
      "id": "686533",
      "postDate": "12/03/2019 09:06:00",
      "content": "<p>&gt; sticking to Tf 1 for now</p>\n\n<p>what does that mean?</p>",
      "rawMarkdown": "&gt; sticking to Tf 1 for now\n\nwhat does that mean?",
      "votes": null
    },
    {
      "id": "686544",
      "postDate": "12/03/2019 09:21:18",
      "content": "<p>I'm also trying to reproduce.</p>\n\n<p>The paper says:</p>\n\n<blockquote>\n  <p>We initialized our model from a BERT model already finetuned on SQuAD 1.1 </p>\n</blockquote>\n\n<p>Do you try finetuning BERT on SQuAD firstly, or only on NQ dataset?</p>",
      "rawMarkdown": "I'm also trying to reproduce.\n\nThe paper says:\n&gt; We initialized our model from a BERT model already finetuned on SQuAD 1.1 \n\nDo you try finetuning BERT on SQuAD firstly, or only on NQ dataset?",
      "votes": null
    },
    {
      "id": "686560",
      "postDate": "12/03/2019 09:44:19",
      "content": "<p>The original implementation is in Tensorflow 1.x - I'm using TF 1.13 for now. </p>",
      "rawMarkdown": "The original implementation is in Tensorflow 1.x - I'm using TF 1.13 for now.",
      "votes": null
    },
    {
      "id": "686568",
      "postDate": "12/03/2019 09:54:49",
      "content": "<p>I still have some problems understanding it fully:</p>\n\n<p><code>\nFormally, we define a training set instance as a four-tuple\n    (c, s, e, t)\n</code></p>\n\n<p>So, the inputs that we have are \"c\", which is a string (BERT tokenized), then we have indicies for the answer in the context string.</p>\n\n<p>```\n[CLS] How old is Dieter? [SEP] blah blah who knows maybe he is 85 years old blah blah [SEP]\n----------------------------  0       1       2       3           4       5   6  7     8       9     10    11    ------------</p>\n\n<p>```</p>\n\n<p>So, we have the indicies for answer as (s, e) = (5, 9) [both inclusive].\nAnd we have answer type.</p>\n\n<p>Am I correct?</p>",
      "rawMarkdown": "I still have some problems understanding it fully:\n\n```\nFormally, we define a training set instance as a four-tuple\n    (c, s, e, t)\n```\n\nSo, the inputs that we have are \"c\", which is a string (BERT tokenized), then we have indicies for the answer in the context string.\n\n\n```\n[CLS] How old is Dieter? [SEP] blah blah who knows maybe he is 85 years old blah blah [SEP]\n----------------------------  0       1       2       3           4       5   6  7     8       9     10    11    ------------\n\n```\n\nSo, we have the indicies for answer as (s, e) = (5, 9) [both inclusive].\nAnd we have answer type.\n\nAm I correct?",
      "votes": null
    },
    {
      "id": "686572",
      "postDate": "12/03/2019 09:58:40",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "686574",
      "postDate": "12/03/2019 10:01:04",
      "content": "<p>Yes. And <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-682249\">this topic</a> would help.</p>",
      "rawMarkdown": "Yes. And [this topic](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-682249) would help.",
      "votes": null
    },
    {
      "id": "686582",
      "postDate": "12/03/2019 10:11:54",
      "content": "<p>Thank! So at inference time, we are actually predicting the indicies? (regression problem)</p>",
      "rawMarkdown": "Thank! So at inference time, we are actually predicting the indicies? (regression problem)",
      "votes": null
    },
    {
      "id": "686588",
      "postDate": "12/03/2019 10:14:53",
      "content": "<p>No, the baseline predicts softmax of one-hot encoded indices as a classification problem. </p>",
      "rawMarkdown": "No, the baseline predicts softmax of one-hot encoded indices as a classification problem.",
      "votes": null
    },
    {
      "id": "686589",
      "postDate": "12/03/2019 10:15:02",
      "content": "<p>No, it's a classification task. In your toy example the gold label is <code>5:9</code>, but whether you predict <code>5:10</code> or <code>1:1000</code> - it's equally bad. It's described now better than before on the Evaluation page.</p>",
      "rawMarkdown": "No, it's a classification task. In your toy example the gold label is `5:9`, but whether you predict `5:10` or `1:1000` - it's equally bad. It's described now better than before on the Evaluation page.",
      "votes": null
    },
    {
      "id": "686590",
      "postDate": "12/03/2019 10:15:44",
      "content": "<p>I'm too slow again :)</p>",
      "rawMarkdown": "I'm too slow again :)",
      "votes": null
    },
    {
      "id": "686591",
      "postDate": "12/03/2019 10:16:30",
      "content": "<p>or I am too fast. <a href=\"/abhishek\">@abhishek</a> see the post Yury made here <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119992#686524\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119992#686524</a></p>",
      "rawMarkdown": "or I am too fast. @abhishek see the post Yury made here https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119992#686524",
      "votes": null
    },
    {
      "id": "686604",
      "postDate": "12/03/2019 10:28:18",
      "content": "<p>ohh. so for every token a probability of start and end is predicted. So, there are 512 tokens, we predict:</p>\n\n<p>512 vector for start\n512 vector for end</p>\n\n<p>got it right now?</p>",
      "rawMarkdown": "ohh. so for every token a probability of start and end is predicted. So, there are 512 tokens, we predict:\n\n512 vector for start\n512 vector for end\n\ngot it right now?",
      "votes": null
    },
    {
      "id": "686607",
      "postDate": "12/03/2019 10:30:49",
      "content": "<p>yes</p>",
      "rawMarkdown": "yes",
      "votes": null
    },
    {
      "id": "686608",
      "postDate": "12/03/2019 10:35:46",
      "content": "<p>Thanks! I’m gonna make a pytorch version of this today.</p>",
      "rawMarkdown": "Thanks! I’m gonna make a pytorch version of this today.",
      "votes": null
    },
    {
      "id": "687018",
      "postDate": "12/03/2019 19:41:28",
      "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> , did you ever try your pytorch training with batch size 2 on Kaggle? For my TF2 training kernel, with <code>batch_size=2</code> and <code>accumulation_size=50</code> (so <code>effective batch size=100</code>) will be about 37 hours to finish one epoch ....I just wonder if pytorch run much  faster. </p>",
      "rawMarkdown": "christofhenkel , did you ever try your pytorch training with batch size 2 on Kaggle? For my TF2 training kernel, with `batch_size=2` and `accumulation_size=50` (so `effective batch size=100`) will be about 37 hours to finish one epoch ....I just wonder if pytorch run much  faster.",
      "votes": null
    },
    {
      "id": "687235",
      "postDate": "12/04/2019 05:24:55",
      "content": "<p>It seems that you can finish training with gtx 1080ti in 15h? Do you train your model with tftrain.tfrecords in bert-joint-baseline dataset? Because in paper True Negative instances are downsampled, which make the dataset only 500k instances. Or you train your model with full data (no downsampling, leads to 8 milliion instances)?</p>",
      "rawMarkdown": "It seems that you can finish training with gtx 1080ti in 15h? Do you train your model with tftrain.tfrecords in bert-joint-baseline dataset? Because in paper True Negative instances are downsampled, which make the dataset only 500k instances. Or you train your model with full data (no downsampling, leads to 8 milliion instances)?",
      "votes": null
    },
    {
      "id": "687399",
      "postDate": "12/04/2019 10:16:35",
      "content": "<blockquote>\n  <p>In your toy example the gold label is 5:9, but whether you predict 5:10 or 1:1000 - it's equally bad. It's described now better than before on the Evaluation page.</p>\n</blockquote>\n\n<p>This bothers me a lot.  I would have expected a dice metric, if you include more than necessary, and 0 if you include less than necessary.  </p>",
      "rawMarkdown": "&gt; In your toy example the gold label is 5:9, but whether you predict 5:10 or 1:1000 - it's equally bad. It's described now better than before on the Evaluation page.\n\nThis bothers me a lot.  I would have expected a dice metric, if you include more than necessary, and 0 if you include less than necessary.",
      "votes": null
    },
    {
      "id": "687445",
      "postDate": "12/04/2019 11:17:07",
      "content": "<p>Very Nice. helped me a lot</p>",
      "rawMarkdown": "Very Nice. helped me a lot",
      "votes": null
    },
    {
      "id": "700440",
      "postDate": "12/22/2019 02:38:24",
      "content": "<p>I'm just catching up discussion here.\nFYI this is the parameters I was able to get from the paper and baseline kernel. Please let me know if I'm missing something.</p>\n\n<p>|  | baseline kernel | The paper |\n| --- | --- | --- |\n|max_seq_length| <strong>384</strong> | 512|\n|doc_stride | 128|  128|\n|max_query_length| 64| N/A|\n|train_bach_size| 32| N/A|\n|predict_batch_size| 8| N/A|\n|learning_rate| <strong>5e-5</strong>|3 · 10−5 |\n|num_train_epochs| 3.0| N/A|\n|warmup_proportion| 0.1| N/A\n|n_best_size| 20| N/A|\n|max_answer_length|30|N/A|</p>\n\n<p>And one more from <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a></p>\n\n<p><em>Assuming you have access to a TPU, you should be able to train a model comparable to ours with the following command in a few hours, although you might need to further tune the learning rate between <strong>1e-4 and 1e-5</strong> and the number of train epochs between <strong>1 and 3</strong>. In our paper we initialize our training from a BERT model trained on SQuAD2.0 and then finetune on NQ for only 1 epoch with a learning rate of <strong>3e-5</strong>.</em></p>",
      "rawMarkdown": "I'm just catching up discussion here.\nFYI this is the parameters I was able to get from the paper and baseline kernel. Please let me know if I'm missing something.\n\n|  | baseline kernel | The paper |\n| --- | --- | --- |\n|max_seq_length| __384__ | 512|\n|doc_stride | 128|  128|\n|max_query_length| 64| N/A|\n|train_bach_size| 32| N/A|\n|predict_batch_size| 8| N/A|\n|learning_rate| __5e-5__|3 · 10−5 |\n|num_train_epochs| 3.0| N/A|\n|warmup_proportion| 0.1| N/A\n|n_best_size| 20| N/A|\n|max_answer_length|30|N/A|\n\nAnd one more from [https://github.com/google-research/language/tree/master/language/question_answering/bert_joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint)\n\n*Assuming you have access to a TPU, you should be able to train a model comparable to ours with the following command in a few hours, although you might need to further tune the learning rate between __1e-4 and 1e-5__ and the number of train epochs between __1 and 3__. In our paper we initialize our training from a BERT model trained on SQuAD2.0 and then finetune on NQ for only 1 epoch with a learning rate of __3e-5__.*",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 686112,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "12/02/2019 19:45:30",
      "content": "<p>Seems like more fine tuning. The github page discusses more epochs too. What worked for you?</p>",
      "votes": null,
      "replies": [
        {
          "id": 686116,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/02/2019 19:48:53",
          "content": "<p>I am still at 1st epoch. From my experience more epochs have little effect. Want to get 1st epoch to a 0.65+ range first </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686348,
          "author_name": "seshurajup",
          "author_url": "",
          "post_date": "12/03/2019 03:48:46",
          "content": "<p>could you share your hardware details and how long it takes to run one epoch <a href=\"/christofhenkel\">@christofhenkel</a> </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686407,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 05:46:31",
          "content": "<p>15h GTX 1080 Ti</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 687018,
          "author_name": "yihdarshieh",
          "author_url": "",
          "post_date": "12/03/2019 19:41:28",
          "content": "<p><a href=\"/christofhenkel\">@christofhenkel</a> , did you ever try your pytorch training with batch size 2 on Kaggle? For my TF2 training kernel, with <code>batch_size=2</code> and <code>accumulation_size=50</code> (so <code>effective batch size=100</code>) will be about 37 hours to finish one epoch ....I just wonder if pytorch run much  faster. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 687235,
          "author_name": "httpwwwfszyc",
          "author_url": "",
          "post_date": "12/04/2019 05:24:55",
          "content": "<p>It seems that you can finish training with gtx 1080ti in 15h? Do you train your model with tftrain.tfrecords in bert-joint-baseline dataset? Because in paper True Negative instances are downsampled, which make the dataset only 500k instances. Or you train your model with full data (no downsampling, leads to 8 milliion instances)?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 686118,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "12/02/2019 19:51:46",
      "content": "<p>Since it seems you are working in pytorch, I’m assuming that you might have implemented the paper.\nCould you explain to me what does the author mean by the following paragraph: </p>\n\n<p><code>For each training instance we compute start and end token indices to represent the target answer span. If all annotated short spans are contained in the instance, we set the start and end target in- dices to point to the smallest span containing all the annotated short answer spans. If there are no annotated short spans but there is an annotated long answer span completely contained in the in- stance, we set the start and end target indices to point to the entire long answer span. If no short or\nlong span can be found in the current instance, we set the target start and end indices to point to the “[CLS]” token.</code></p>\n\n<p>I’m a bit confused to what the author is trying to convey here.</p>\n\n<p>Let’s say start index of document: 256, end index of document: 576. (Removing question tokens). \nStart index of long answer: 310, end index of long answer: 625\nStart index of short answer: 450, end index of short answer: 500.</p>\n\n<p>How would you label these here? Ignore the indices that are outside? Or keep only those answers that fall completely in the range we chose: 256-576 ?</p>",
      "votes": null,
      "replies": [
        {
          "id": 686138,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/02/2019 20:24:31",
          "content": "<p>as far as I understand label would be 450:500, answer_type = short. Long answer is discarded as its end is out of the span. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686141,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/02/2019 20:26:08",
          "content": "<p>Interesting are those windows which have no answers. According to the default params in repo, 2% of those no_answer windows are fed into training and 98% are ignored</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686148,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "12/02/2019 20:30:53",
          "content": "<p>By that method, aren’t we also excluding (many?) short and long answers?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686152,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/02/2019 20:36:04",
          "content": "<p>No, since the 512 windows overlap (by default with 128) which is way longer than normal long/ short answers, the answers are contained in at least one. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 686180,
      "author_name": "particlebbq",
      "author_url": "",
      "post_date": "12/02/2019 21:27:15",
      "content": "<p>Which bert configuration and batch size are you using in your training?  On this page: </p>\n\n<p><a href=\"https://github.com/google-research/bert\">https://github.com/google-research/bert</a></p>\n\n<p>in the \"out of memory issues\" section, there is a statement that \"Unfortunately, these max batch sizes for BERT-Large are so small that they will actually harm the model accuracy, regardless of the learning rate used.\"</p>",
      "votes": null,
      "replies": [
        {
          "id": 686517,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 08:40:32",
          "content": "<p>gradient accumulation helps</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 686505,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "12/03/2019 08:26:30",
      "content": "<blockquote>\n  <p>for 1 epoch with an initial learning rate of 0.005 was the best setting</p>\n</blockquote>\n\n<p>I think it's just a mistake, maybe they confused 5e-3 and 3e-5. 0.005 is a too high learning rate for transformers.  </p>",
      "votes": null,
      "replies": [
        {
          "id": 686513,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 08:35:56",
          "content": "<p>thats my guess too. But there are more inconsistencies and intransparencies. For example related to optimizer used and if a warmup is applied</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686520,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 08:44:36",
          "content": "<p>As usual in academic papers :) It's also about version control - it's hard to align all changes in code with paragraphs in the paper describing the approach. \nps. didn't try to implement myself, sticking to Tf 1 for now. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686533,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 09:06:00",
          "content": "<p>&gt; sticking to Tf 1 for now</p>\n\n<p>what does that mean?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686560,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 09:44:19",
          "content": "<p>The original implementation is in Tensorflow 1.x - I'm using TF 1.13 for now. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 686544,
      "author_name": "kentaronakanishi",
      "author_url": "",
      "post_date": "12/03/2019 09:21:18",
      "content": "<p>I'm also trying to reproduce.</p>\n\n<p>The paper says:</p>\n\n<blockquote>\n  <p>We initialized our model from a BERT model already finetuned on SQuAD 1.1 </p>\n</blockquote>\n\n<p>Do you try finetuning BERT on SQuAD firstly, or only on NQ dataset?</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 686568,
      "author_name": "abhishek",
      "author_url": "",
      "post_date": "12/03/2019 09:54:49",
      "content": "<p>I still have some problems understanding it fully:</p>\n\n<p><code>\nFormally, we define a training set instance as a four-tuple\n    (c, s, e, t)\n</code></p>\n\n<p>So, the inputs that we have are \"c\", which is a string (BERT tokenized), then we have indicies for the answer in the context string.</p>\n\n<p>```\n[CLS] How old is Dieter? [SEP] blah blah who knows maybe he is 85 years old blah blah [SEP]\n----------------------------  0       1       2       3           4       5   6  7     8       9     10    11    ------------</p>\n\n<p>```</p>\n\n<p>So, we have the indicies for answer as (s, e) = (5, 9) [both inclusive].\nAnd we have answer type.</p>\n\n<p>Am I correct?</p>",
      "votes": null,
      "replies": [
        {
          "id": 686572,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 09:58:40",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686574,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 10:01:04",
          "content": "<p>Yes. And <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-682249\">this topic</a> would help.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686582,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "12/03/2019 10:11:54",
          "content": "<p>Thank! So at inference time, we are actually predicting the indicies? (regression problem)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686588,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 10:14:53",
          "content": "<p>No, the baseline predicts softmax of one-hot encoded indices as a classification problem. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686589,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 10:15:02",
          "content": "<p>No, it's a classification task. In your toy example the gold label is <code>5:9</code>, but whether you predict <code>5:10</code> or <code>1:1000</code> - it's equally bad. It's described now better than before on the Evaluation page.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686590,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/03/2019 10:15:44",
          "content": "<p>I'm too slow again :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686591,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 10:16:30",
          "content": "<p>or I am too fast. <a href=\"/abhishek\">@abhishek</a> see the post Yury made here <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119992#686524\">https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119992#686524</a></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686604,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "12/03/2019 10:28:18",
          "content": "<p>ohh. so for every token a probability of start and end is predicted. So, there are 512 tokens, we predict:</p>\n\n<p>512 vector for start\n512 vector for end</p>\n\n<p>got it right now?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686607,
          "author_name": "christofhenkel",
          "author_url": "",
          "post_date": "12/03/2019 10:30:49",
          "content": "<p>yes</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 686608,
          "author_name": "abhishek",
          "author_url": "",
          "post_date": "12/03/2019 10:35:46",
          "content": "<p>Thanks! I’m gonna make a pytorch version of this today.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 687399,
          "author_name": "cpmpml",
          "author_url": "",
          "post_date": "12/04/2019 10:16:35",
          "content": "<blockquote>\n  <p>In your toy example the gold label is 5:9, but whether you predict 5:10 or 1:1000 - it's equally bad. It's described now better than before on the Evaluation page.</p>\n</blockquote>\n\n<p>This bothers me a lot.  I would have expected a dice metric, if you include more than necessary, and 0 if you include less than necessary.  </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 687445,
      "author_name": "mobinmirarab",
      "author_url": "",
      "post_date": "12/04/2019 11:17:07",
      "content": "<p>Very Nice. helped me a lot</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 700440,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "12/22/2019 02:38:24",
      "content": "<p>I'm just catching up discussion here.\nFYI this is the parameters I was able to get from the paper and baseline kernel. Please let me know if I'm missing something.</p>\n\n<p>|  | baseline kernel | The paper |\n| --- | --- | --- |\n|max_seq_length| <strong>384</strong> | 512|\n|doc_stride | 128|  128|\n|max_query_length| 64| N/A|\n|train_bach_size| 32| N/A|\n|predict_batch_size| 8| N/A|\n|learning_rate| <strong>5e-5</strong>|3 · 10−5 |\n|num_train_epochs| 3.0| N/A|\n|warmup_proportion| 0.1| N/A\n|n_best_size| 20| N/A|\n|max_answer_length|30|N/A|</p>\n\n<p>And one more from <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a></p>\n\n<p><em>Assuming you have access to a TPU, you should be able to train a model comparable to ours with the following command in a few hours, although you might need to further tune the learning rate between <strong>1e-4 and 1e-5</strong> and the number of train epochs between <strong>1 and 3</strong>. In our paper we initialize our training from a BERT model trained on SQuAD2.0 and then finetune on NQ for only 1 epoch with a learning rate of <strong>3e-5</strong>.</em></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "686108": "As you might be aware the current baseline is based on this paper https://arxiv.org/pdf/1901.08634.pdf \n\nNevertheless I have a hard time reproducing the result. For me some of the parameters seem contradictory. For example at one point they write \" found that training for 1 epoch with an initial learning rate of 0.005 was the best setting\" and at the related repo (https://github.com/google-research/language/tree/master/language/question_answering/bert_joint ) is written lr should be 3e-5",
    "686112": "Seems like more fine tuning. The github page discusses more epochs too. What worked for you?",
    "686116": "I am still at 1st epoch. From my experience more epochs have little effect. Want to get 1st epoch to a 0.65+ range first",
    "686118": "Since it seems you are working in pytorch, I’m assuming that you might have implemented the paper.\nCould you explain to me what does the author mean by the following paragraph: \n\n`For each training instance we compute start and end token indices to represent the target answer span. If all annotated short spans are contained in the instance, we set the start and end target in- dices to point to the smallest span containing all the annotated short answer spans. If there are no annotated short spans but there is an annotated long answer span completely contained in the in- stance, we set the start and end target indices to point to the entire long answer span. If no short or\nlong span can be found in the current instance, we set the target start and end indices to point to the “[CLS]” token. `\n\n\nI’m a bit confused to what the author is trying to convey here.\n\n\nLet’s say start index of document: 256, end index of document: 576. (Removing question tokens). \nStart index of long answer: 310, end index of long answer: 625\nStart index of short answer: 450, end index of short answer: 500.\n\nHow would you label these here? Ignore the indices that are outside? Or keep only those answers that fall completely in the range we chose: 256-576 ?",
    "686138": "as far as I understand label would be 450:500, answer_type = short. Long answer is discarded as its end is out of the span.",
    "686141": "Interesting are those windows which have no answers. According to the default params in repo, 2% of those no_answer windows are fed into training and 98% are ignored",
    "686148": "By that method, aren’t we also excluding (many?) short and long answers?",
    "686152": "No, since the 512 windows overlap (by default with 128) which is way longer than normal long/ short answers, the answers are contained in at least one.",
    "686180": "Which bert configuration and batch size are you using in your training?  On this page: \n\nhttps://github.com/google-research/bert\n\nin the \"out of memory issues\" section, there is a statement that \"Unfortunately, these max batch sizes for BERT-Large are so small that they will actually harm the model accuracy, regardless of the learning rate used.\"",
    "686348": "could you share your hardware details and how long it takes to run one epoch @christofhenkel",
    "686407": "15h GTX 1080 Ti",
    "686505": "&gt; for 1 epoch with an initial learning rate of 0.005 was the best setting\n\nI think it's just a mistake, maybe they confused 5e-3 and 3e-5. 0.005 is a too high learning rate for transformers.",
    "686513": "thats my guess too. But there are more inconsistencies and intransparencies. For example related to optimizer used and if a warmup is applied",
    "686517": "gradient accumulation helps",
    "686520": "As usual in academic papers :) It's also about version control - it's hard to align all changes in code with paragraphs in the paper describing the approach. \nps. didn't try to implement myself, sticking to Tf 1 for now.",
    "686533": "&gt; sticking to Tf 1 for now\n\nwhat does that mean?",
    "686544": "I'm also trying to reproduce.\n\nThe paper says:\n&gt; We initialized our model from a BERT model already finetuned on SQuAD 1.1 \n\nDo you try finetuning BERT on SQuAD firstly, or only on NQ dataset?",
    "686560": "The original implementation is in Tensorflow 1.x - I'm using TF 1.13 for now.",
    "686568": "I still have some problems understanding it fully:\n\n```\nFormally, we define a training set instance as a four-tuple\n    (c, s, e, t)\n```\n\nSo, the inputs that we have are \"c\", which is a string (BERT tokenized), then we have indicies for the answer in the context string.\n\n\n```\n[CLS] How old is Dieter? [SEP] blah blah who knows maybe he is 85 years old blah blah [SEP]\n----------------------------  0       1       2       3           4       5   6  7     8       9     10    11    ------------\n\n```\n\nSo, we have the indicies for answer as (s, e) = (5, 9) [both inclusive].\nAnd we have answer type.\n\nAm I correct?",
    "686572": "yes",
    "686574": "Yes. And [this topic](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-682249) would help.",
    "686582": "Thank! So at inference time, we are actually predicting the indicies? (regression problem)",
    "686588": "No, the baseline predicts softmax of one-hot encoded indices as a classification problem.",
    "686589": "No, it's a classification task. In your toy example the gold label is `5:9`, but whether you predict `5:10` or `1:1000` - it's equally bad. It's described now better than before on the Evaluation page.",
    "686590": "I'm too slow again :)",
    "686591": "or I am too fast. @abhishek see the post Yury made here https://www.kaggle.com/c/tensorflow2-question-answering/discussion/119992#686524",
    "686604": "ohh. so for every token a probability of start and end is predicted. So, there are 512 tokens, we predict:\n\n512 vector for start\n512 vector for end\n\ngot it right now?",
    "686607": "yes",
    "686608": "Thanks! I’m gonna make a pytorch version of this today.",
    "687018": "christofhenkel , did you ever try your pytorch training with batch size 2 on Kaggle? For my TF2 training kernel, with `batch_size=2` and `accumulation_size=50` (so `effective batch size=100`) will be about 37 hours to finish one epoch ....I just wonder if pytorch run much  faster.",
    "687235": "It seems that you can finish training with gtx 1080ti in 15h? Do you train your model with tftrain.tfrecords in bert-joint-baseline dataset? Because in paper True Negative instances are downsampled, which make the dataset only 500k instances. Or you train your model with full data (no downsampling, leads to 8 milliion instances)?",
    "687399": "&gt; In your toy example the gold label is 5:9, but whether you predict 5:10 or 1:1000 - it's equally bad. It's described now better than before on the Evaluation page.\n\nThis bothers me a lot.  I would have expected a dice metric, if you include more than necessary, and 0 if you include less than necessary.",
    "687445": "Very Nice. helped me a lot",
    "700440": "I'm just catching up discussion here.\nFYI this is the parameters I was able to get from the paper and baseline kernel. Please let me know if I'm missing something.\n\n|  | baseline kernel | The paper |\n| --- | --- | --- |\n|max_seq_length| __384__ | 512|\n|doc_stride | 128|  128|\n|max_query_length| 64| N/A|\n|train_bach_size| 32| N/A|\n|predict_batch_size| 8| N/A|\n|learning_rate| __5e-5__|3 · 10−5 |\n|num_train_epochs| 3.0| N/A|\n|warmup_proportion| 0.1| N/A\n|n_best_size| 20| N/A|\n|max_answer_length|30|N/A|\n\nAnd one more from [https://github.com/google-research/language/tree/master/language/question_answering/bert_joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint)\n\n*Assuming you have access to a TPU, you should be able to train a model comparable to ours with the following command in a few hours, although you might need to further tune the learning rate between __1e-4 and 1e-5__ and the number of train epochs between __1 and 3__. In our paper we initialize our training from a BERT model trained on SQuAD2.0 and then finetune on NQ for only 1 epoch with a learning rate of __3e-5__.*"
  },
  "source": "meta"
}