{
  "id": 123434,
  "title": "HuggingFace question / BertForQuestionAnswering",
  "url": "/competitions/tensorflow2-question-answering/discussion/123434",
  "author_name": "Alon Bochman",
  "post_date": "2019-12-27T16:01:13.523000",
  "votes": 5,
  "comment_count": 8,
  "views": 0,
  "content": "<p>Calling all PyTorch / HuggingFace experts:</p>\n\n<p>The HuggingFace BertForQuestionAnswering sample code creates duplicate [CLS] tokens. Wondering why:\nThis code is from their docs:  <a href=\"https://huggingface.co/transformers/model_doc/bert.html#bertforquestionanswering\">https://huggingface.co/transformers/model_doc/bert.html#bertforquestionanswering</a>\n```\ntokenizer = BertTokenizer.from_pretrained('bert-base-uncased')\nmodel = BertForQuestionAnswering.from_pretrained('bert-large-uncased-whole-word-masking-finetuned-squad')\nquestion, text = \"Who was Jim Henson?\", \"Jim Henson was a nice puppet\"\ninput_text = \"[CLS] \" + question + \" [SEP] \" + text + \" [SEP]\"\ninput_ids = tokenizer.encode(input_text)\ntoken_type_ids = [0 if i &lt;= input_ids.index(102) else 1 for i in range(len(input_ids))]\nstart_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))\nall_tokens = tokenizer.convert_ids_to_tokens(input_ids)\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))</p>\n\n<h1>a nice puppet</h1>\n\n<p>tokenizer.decode(input_ids)</p>\n\n<h1>'[CLS] [CLS] who was jim henson? [SEP] jim henson was a nice puppet [SEP] [SEP]'</h1>\n\n<p>```</p>\n\n<p>If I remove the extra [CLS], the extraction doesn't work. It's exactly two tokens off:\n```\ninput_ids = tokenizer.encode(input_text, add_special_tokens=False)\n...rerun same code as above...\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))</p>\n\n<h1>was a</h1>\n\n<p>```\nWhat am I doing wrong? How can I get the extraction working without duplicate [CLS] tokens? (and duplicate final [SEP] tokens BTW).</p>",
  "messages": [
    {
      "id": 704554,
      "postDate": "2019-12-27T16:01:13.523Z",
      "content": "<p>Calling all PyTorch / HuggingFace experts:</p>\n\n<p>The HuggingFace BertForQuestionAnswering sample code creates duplicate [CLS] tokens. Wondering why:\nThis code is from their docs:  <a href=\"https://huggingface.co/transformers/model_doc/bert.html#bertforquestionanswering\">https://huggingface.co/transformers/model_doc/bert.html#bertforquestionanswering</a>\n```\ntokenizer = BertTokenizer.from_pretrained('bert-base-uncased')\nmodel = BertForQuestionAnswering.from_pretrained('bert-large-uncased-whole-word-masking-finetuned-squad')\nquestion, text = \"Who was Jim Henson?\", \"Jim Henson was a nice puppet\"\ninput_text = \"[CLS] \" + question + \" [SEP] \" + text + \" [SEP]\"\ninput_ids = tokenizer.encode(input_text)\ntoken_type_ids = [0 if i &lt;= input_ids.index(102) else 1 for i in range(len(input_ids))]\nstart_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))\nall_tokens = tokenizer.convert_ids_to_tokens(input_ids)\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))</p>\n\n<h1>a nice puppet</h1>\n\n<p>tokenizer.decode(input_ids)</p>\n\n<h1>'[CLS] [CLS] who was jim henson? [SEP] jim henson was a nice puppet [SEP] [SEP]'</h1>\n\n<p>```</p>\n\n<p>If I remove the extra [CLS], the extraction doesn't work. It's exactly two tokens off:\n```\ninput_ids = tokenizer.encode(input_text, add_special_tokens=False)\n...rerun same code as above...\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))</p>\n\n<h1>was a</h1>\n\n<p>```\nWhat am I doing wrong? How can I get the extraction working without duplicate [CLS] tokens? (and duplicate final [SEP] tokens BTW).</p>",
      "rawMarkdown": "Calling all PyTorch / HuggingFace experts:\n\n\nThe HuggingFace BertForQuestionAnswering sample code creates duplicate [CLS] tokens. Wondering why:\nThis code is from their docs:  https://huggingface.co/transformers/model_doc/bert.html#bertforquestionanswering\n```\ntokenizer = BertTokenizer.from_pretrained('bert-base-uncased')\nmodel = BertForQuestionAnswering.from_pretrained('bert-large-uncased-whole-word-masking-finetuned-squad')\nquestion, text = \"Who was Jim Henson?\", \"Jim Henson was a nice puppet\"\ninput_text = \"[CLS] \" + question + \" [SEP] \" + text + \" [SEP]\"\ninput_ids = tokenizer.encode(input_text)\ntoken_type_ids = [0 if i &lt;= input_ids.index(102) else 1 for i in range(len(input_ids))]\nstart_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))\nall_tokens = tokenizer.convert_ids_to_tokens(input_ids)\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))\n# a nice puppet\ntokenizer.decode(input_ids)\n#'[CLS] [CLS] who was jim henson? [SEP] jim henson was a nice puppet [SEP] [SEP]'  \n```\n\nIf I remove the extra [CLS], the extraction doesn't work. It's exactly two tokens off:\n```\ninput_ids = tokenizer.encode(input_text, add_special_tokens=False)\n...rerun same code as above...\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))\n# was a\n```\nWhat am I doing wrong? How can I get the extraction working without duplicate [CLS] tokens? (and duplicate final [SEP] tokens BTW).\n\n",
      "votes": 5
    },
    {
      "id": 704717,
      "postDate": "2019-12-27T22:35:09.603Z",
      "content": "<p>Here is the result I got</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F69705497e399e974cd6fe84bb392124a%2FCapture.PNG?generation=1577487945372739&amp;alt=media\" alt=\"\"></p>\n\n<p>I didn't do <code>start_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))</code>, not sure if this difference counts.</p>\n\n<p>Probably check your <code>transformers</code> and <code>torch</code> version, and upgrade if necessary.</p>",
      "rawMarkdown": "Here is the result I got\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F69705497e399e974cd6fe84bb392124a%2FCapture.PNG?generation=1577487945372739&amp;alt=media)\n\n\nI didn't do `start_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))`, not sure if this difference counts.\n\nProbably check your `transformers` and `torch` version, and upgrade if necessary.",
      "votes": 4,
      "replies": [
        {
          "id": 852826,
          "postDate": "2020-05-18T17:50:05.900Z",
          "content": "<p>How should I extend the code for making a chatbot.\nI have preprocess the the question and answers.\nEstablished the input pipeline.\nHow should I proceed with modelling.\nUsing TFBertForQuestionAnswering</p>",
          "rawMarkdown": "How should I extend the code for making a chatbot.\nI have preprocess the the question and answers.\nEstablished the input pipeline.\nHow should I proceed with modelling.\nUsing TFBertForQuestionAnswering"
        }
      ]
    },
    {
      "id": 705375,
      "postDate": "2019-12-28T20:53:12.547Z",
      "content": "<p>Thanks guys. I managed to get it working.</p>",
      "rawMarkdown": "Thanks guys. I managed to get it working.",
      "votes": 1
    },
    {
      "id": 704891,
      "postDate": "2019-12-28T06:05:18.750Z",
      "content": "<p>Hope this helps. I think I got correct output here.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fd1ec7864bf568fc2484857ac83744218%2F2019-12-28%2015.04.27.png?generation=1577513095516398&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hope this helps. I think I got correct output here.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fd1ec7864bf568fc2484857ac83744218%2F2019-12-28%2015.04.27.png?generation=1577513095516398&amp;alt=media)\n",
      "votes": 1,
      "replies": [
        {
          "id": 818889,
          "postDate": "2020-04-24T07:30:16.303Z",
          "content": "<p>Can we use this BertQuestionAnswering of huggingface for a large context. For example: I have a large paragraph (approx. 6000 in length) . If no, then how can we use this model to find the correct answers of a really large abstract and also can we put a threshold on the length of answers and also on start_scores and end_scores?</p>",
          "rawMarkdown": "Can we use this BertQuestionAnswering of huggingface for a large context. For example: I have a large paragraph (approx. 6000 in length) . If no, then how can we use this model to find the correct answers of a really large abstract and also can we put a threshold on the length of answers and also on start_scores and end_scores?"
        },
        {
          "id": 894668,
          "postDate": "2020-06-20T16:35:24.940Z",
          "content": "<p>How they mention in the paper is that u need to break large sentences out from a paragraph.</p>\n\n<p>Further, they experimented by : \n1. taking the first half of the sentence (First only)\n2. taking the last half of the paragraph (Last only)\n3. Most appropriate approach by them is taking first 150 + the last words (First + last)</p>\n\n<p>I hope this might help. </p>",
          "rawMarkdown": "How they mention in the paper is that u need to break large sentences out from a paragraph.\n\nFurther, they experimented by : \n1. taking the first half of the sentence (First only)\n2. taking the last half of the paragraph (Last only)\n3. Most appropriate approach by them is taking first 150 + the last words (First + last)\n\nI hope this might help. "
        }
      ]
    },
    {
      "id": 852833,
      "postDate": "2020-05-18T17:56:43.990Z",
      "rawMarkdown": "",
      "isDeleted": true
    },
    {
      "id": 1463731,
      "postDate": "2021-08-10T09:56:45.120Z",
      "content": "<p>Thank You. This was helpful</p>",
      "rawMarkdown": "Thank You. This was helpful"
    }
  ],
  "comments": [
    {
      "id": 704717,
      "author_name": "Yih-Dar SHIEH",
      "author_url": "",
      "post_date": "2019-12-27T22:35:09.603000",
      "content": "<p>Here is the result I got</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F69705497e399e974cd6fe84bb392124a%2FCapture.PNG?generation=1577487945372739&amp;alt=media\" alt=\"\"></p>\n\n<p>I didn't do <code>start_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))</code>, not sure if this difference counts.</p>\n\n<p>Probably check your <code>transformers</code> and <code>torch</code> version, and upgrade if necessary.</p>",
      "votes": 4,
      "replies": [
        {
          "id": 852826,
          "author_name": "Kayvan Shah",
          "author_url": "",
          "post_date": "2020-05-18T17:50:05.900000",
          "content": "<p>How should I extend the code for making a chatbot.\nI have preprocess the the question and answers.\nEstablished the input pipeline.\nHow should I proceed with modelling.\nUsing TFBertForQuestionAnswering</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 705375,
      "author_name": "Alon Bochman",
      "author_url": "",
      "post_date": "2019-12-28T20:53:12.547000",
      "content": "<p>Thanks guys. I managed to get it working.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 704891,
      "author_name": "higepon",
      "author_url": "",
      "post_date": "2019-12-28T06:05:18.750000",
      "content": "<p>Hope this helps. I think I got correct output here.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fd1ec7864bf568fc2484857ac83744218%2F2019-12-28%2015.04.27.png?generation=1577513095516398&amp;alt=media\" alt=\"\"></p>",
      "votes": 1,
      "replies": [
        {
          "id": 818889,
          "author_name": "Aishwarya Verma",
          "author_url": "",
          "post_date": "2020-04-24T07:30:16.303000",
          "content": "<p>Can we use this BertQuestionAnswering of huggingface for a large context. For example: I have a large paragraph (approx. 6000 in length) . If no, then how can we use this model to find the correct answers of a really large abstract and also can we put a threshold on the length of answers and also on start_scores and end_scores?</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 894668,
          "author_name": "Usman Farooq",
          "author_url": "",
          "post_date": "2020-06-20T16:35:24.940000",
          "content": "<p>How they mention in the paper is that u need to break large sentences out from a paragraph.</p>\n\n<p>Further, they experimented by : \n1. taking the first half of the sentence (First only)\n2. taking the last half of the paragraph (Last only)\n3. Most appropriate approach by them is taking first 150 + the last words (First + last)</p>\n\n<p>I hope this might help. </p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 852833,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-05-18T17:56:43.990000",
      "content": "",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1463731,
      "author_name": "Kayvan Shah",
      "author_url": "",
      "post_date": "2021-08-10T09:56:45.120000",
      "content": "<p>Thank You. This was helpful</p>",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "704554": "Calling all PyTorch / HuggingFace experts:\n\n\nThe HuggingFace BertForQuestionAnswering sample code creates duplicate [CLS] tokens. Wondering why:\nThis code is from their docs:  https://huggingface.co/transformers/model_doc/bert.html#bertforquestionanswering\n```\ntokenizer = BertTokenizer.from_pretrained('bert-base-uncased')\nmodel = BertForQuestionAnswering.from_pretrained('bert-large-uncased-whole-word-masking-finetuned-squad')\nquestion, text = \"Who was Jim Henson?\", \"Jim Henson was a nice puppet\"\ninput_text = \"[CLS] \" + question + \" [SEP] \" + text + \" [SEP]\"\ninput_ids = tokenizer.encode(input_text)\ntoken_type_ids = [0 if i &lt;= input_ids.index(102) else 1 for i in range(len(input_ids))]\nstart_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))\nall_tokens = tokenizer.convert_ids_to_tokens(input_ids)\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))\n# a nice puppet\ntokenizer.decode(input_ids)\n#'[CLS] [CLS] who was jim henson? [SEP] jim henson was a nice puppet [SEP] [SEP]'  \n```\n\nIf I remove the extra [CLS], the extraction doesn't work. It's exactly two tokens off:\n```\ninput_ids = tokenizer.encode(input_text, add_special_tokens=False)\n...rerun same code as above...\nprint(' '.join(all_tokens[torch.argmax(start_scores) : torch.argmax(end_scores)+1]))\n# was a\n```\nWhat am I doing wrong? How can I get the extraction working without duplicate [CLS] tokens? (and duplicate final [SEP] tokens BTW).\n\n",
    "704717": "Here is the result I got\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1533864%2F69705497e399e974cd6fe84bb392124a%2FCapture.PNG?generation=1577487945372739&amp;alt=media)\n\n\nI didn't do `start_scores, end_scores = model(torch.tensor([input_ids]), token_type_ids=torch.tensor([token_type_ids]))`, not sure if this difference counts.\n\nProbably check your `transformers` and `torch` version, and upgrade if necessary.",
    "705375": "Thanks guys. I managed to get it working.",
    "704891": "Hope this helps. I think I got correct output here.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2Fd1ec7864bf568fc2484857ac83744218%2F2019-12-28%2015.04.27.png?generation=1577513095516398&amp;alt=media)\n",
    "852833": "",
    "1463731": "Thank You. This was helpful"
  }
}