{
  "id": 116306,
  "title": "about the way BERT-joint (baseline) processes long answers",
  "url": "/competitions/tensorflow2-question-answering/discussion/116306",
  "author_name": "Xinyi Zhao",
  "post_date": "2019-11-08T09:02:45.785000",
  "votes": 5,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I was reading the <a href=\"https://arxiv.org/abs/1901.08634\">source paper</a> and the <a href=\"https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/run_nq.py\">original script</a> of the baseline BERT-joint model. The related codes are from line 696.</p>\n\n<p>While training, one entry can be divided into several instances using a sliding window, and each instance has a max length which was 512 when google trained their model. My question is that apparently some of the long answers are (much) longer than 512 tokens, but according to the paper, if one instance doesn't contain a <strong>complete</strong> long answer, it is going to be regarded as \"null instance\" and thrown out of the training dataset.</p>\n\n<p>Is my understanding correct?\nSince longer max sequence length in training would be more time- and memory-consuming, does it mean \"long long answers\" are out of consideration when using BERT-joint model?</p>\n\n<p>Thx in advance!</p>",
  "messages": [
    {
      "id": 668313,
      "postDate": "2019-11-08T09:02:45.787Z",
      "content": "<p>I was reading the <a href=\"https://arxiv.org/abs/1901.08634\">source paper</a> and the <a href=\"https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/run_nq.py\">original script</a> of the baseline BERT-joint model. The related codes are from line 696.</p>\n\n<p>While training, one entry can be divided into several instances using a sliding window, and each instance has a max length which was 512 when google trained their model. My question is that apparently some of the long answers are (much) longer than 512 tokens, but according to the paper, if one instance doesn't contain a <strong>complete</strong> long answer, it is going to be regarded as \"null instance\" and thrown out of the training dataset.</p>\n\n<p>Is my understanding correct?\nSince longer max sequence length in training would be more time- and memory-consuming, does it mean \"long long answers\" are out of consideration when using BERT-joint model?</p>\n\n<p>Thx in advance!</p>",
      "rawMarkdown": "I was reading the [source paper](https://arxiv.org/abs/1901.08634) and the [original script](https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/run_nq.py) of the baseline BERT-joint model. The related codes are from line 696.\n\nWhile training, one entry can be divided into several instances using a sliding window, and each instance has a max length which was 512 when google trained their model. My question is that apparently some of the long answers are (much) longer than 512 tokens, but according to the paper, if one instance doesn't contain a **complete** long answer, it is going to be regarded as \"null instance\" and thrown out of the training dataset.\n\nIs my understanding correct?\nSince longer max sequence length in training would be more time- and memory-consuming, does it mean \"long long answers\" are out of consideration when using BERT-joint model?\n\nThx in advance!",
      "votes": 5
    },
    {
      "id": 698358,
      "postDate": "2019-12-19T05:37:50.077Z",
      "content": "<p><a href=\"/rohitagarwal\">@rohitagarwal</a> nah not really.\ni was trying to understand how the model was trained back then.\nbut so far i haven't done anything about this problem.</p>",
      "rawMarkdown": "@rohitagarwal nah not really.\ni was trying to understand how the model was trained back then.\nbut so far i haven't done anything about this problem.",
      "votes": 1
    },
    {
      "id": 698242,
      "postDate": "2019-12-19T01:32:16.707Z",
      "content": "<p><a href=\"/xinyicc\">@xinyicc</a> You are raising a very valid point. \nHave you been able to come to terms with this?</p>",
      "rawMarkdown": "@xinyicc You are raising a very valid point. \nHave you been able to come to terms with this?",
      "votes": 1
    },
    {
      "id": 699089,
      "postDate": "2019-12-20T03:28:55.030Z",
      "content": "<p>Here is my understanding. Please correct me if I am wrong.  The model always first find the short answer spans, which usually can be less than 512. Then the long answer  is predicted to be the DOM tree top level node containing the predicted short answer span (at the bottom of left on page 3 of the paper). In this way, the length of the long answer can be over 512.  </p>",
      "rawMarkdown": "Here is my understanding. Please correct me if I am wrong.  The model always first find the short answer spans, which usually can be less than 512. Then the long answer  is predicted to be the DOM tree top level node containing the predicted short answer span (at the bottom of left on page 3 of the paper). In this way, the length of the long answer can be over 512.  "
    },
    {
      "id": 669468,
      "postDate": "2019-11-10T05:37:09.383Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 670102,
          "postDate": "2019-11-11T02:31:41.113Z",
          "content": "<p>yeah, answers longer than 512 are definitely minority, but i don't like the idea that answers can only be shorter than 512 and any answer longer than that can never be predicted correctly.</p>\n\n<p>but since they are minority, maybe this is not very important in practical use...</p>",
          "rawMarkdown": "yeah, answers longer than 512 are definitely minority, but i don't like the idea that answers can only be shorter than 512 and any answer longer than that can never be predicted correctly.\n\nbut since they are minority, maybe this is not very important in practical use..."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 698358,
      "author_name": "Xinyi Zhao",
      "author_url": "",
      "post_date": "2019-12-19T05:37:50.077000",
      "content": "<p><a href=\"/rohitagarwal\">@rohitagarwal</a> nah not really.\ni was trying to understand how the model was trained back then.\nbut so far i haven't done anything about this problem.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 698242,
      "author_name": "Rohit Agarwal",
      "author_url": "",
      "post_date": "2019-12-19T01:32:16.707000",
      "content": "<p><a href=\"/xinyicc\">@xinyicc</a> You are raising a very valid point. \nHave you been able to come to terms with this?</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 699089,
      "author_name": "Xiaofei Zheng",
      "author_url": "",
      "post_date": "2019-12-20T03:28:55.030000",
      "content": "<p>Here is my understanding. Please correct me if I am wrong.  The model always first find the short answer spans, which usually can be less than 512. Then the long answer  is predicted to be the DOM tree top level node containing the predicted short answer span (at the bottom of left on page 3 of the paper). In this way, the length of the long answer can be over 512.  </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 669468,
      "author_name": "",
      "author_url": "",
      "post_date": "2019-11-10T05:37:09.383000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 670102,
          "author_name": "Xinyi Zhao",
          "author_url": "",
          "post_date": "2019-11-11T02:31:41.113000",
          "content": "<p>yeah, answers longer than 512 are definitely minority, but i don't like the idea that answers can only be shorter than 512 and any answer longer than that can never be predicted correctly.</p>\n\n<p>but since they are minority, maybe this is not very important in practical use...</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "668313": "I was reading the [source paper](https://arxiv.org/abs/1901.08634) and the [original script](https://github.com/google-research/language/blob/master/language/question_answering/bert_joint/run_nq.py) of the baseline BERT-joint model. The related codes are from line 696.\n\nWhile training, one entry can be divided into several instances using a sliding window, and each instance has a max length which was 512 when google trained their model. My question is that apparently some of the long answers are (much) longer than 512 tokens, but according to the paper, if one instance doesn't contain a **complete** long answer, it is going to be regarded as \"null instance\" and thrown out of the training dataset.\n\nIs my understanding correct?\nSince longer max sequence length in training would be more time- and memory-consuming, does it mean \"long long answers\" are out of consideration when using BERT-joint model?\n\nThx in advance!",
    "698358": "@rohitagarwal nah not really.\ni was trying to understand how the model was trained back then.\nbut so far i haven't done anything about this problem.",
    "698242": "@xinyicc You are raising a very valid point. \nHave you been able to come to terms with this?",
    "699089": "Here is my understanding. Please correct me if I am wrong.  The model always first find the short answer spans, which usually can be less than 512. Then the long answer  is predicted to be the DOM tree top level node containing the predicted short answer span (at the bottom of left on page 3 of the paper). In this way, the length of the long answer can be over 512.  ",
    "669468": ""
  }
}