{
  "id": 117370,
  "title": "Understand how does training input look like",
  "url": "/competitions/tensorflow2-question-answering/discussion/117370",
  "author_name": "",
  "post_date": "2019-11-15T01:16:40.688934900Z",
  "votes": 47,
  "comment_count": 4,
  "views": 0,
  "content": "<h2>Background</h2>\n\n<p>So do you all understand how the training data look like which is fed into BERT for fine-tuning? I actually had no idea. Let me explain my understanding here. Hope this helps and I’d appreciate your feedback.</p>\n\n<h2>References</h2>\n\n<ul>\n<li>Their paper <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">PDF: A BERT Baseline for the Natural Questions</a></li>\n<li>Their <a href=\"https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq\">Tensorflow implementation</a></li>\n</ul>\n\n<h2>Training data</h2>\n\n<p>As you may have already seen we have training jsonl file which has\n- document_text (String)\n- question_text (String)\n- short_answer (Range)\n- long_answer (Range)</p>\n\n<p>According to the paper, a training set instance is four tuple as follows.\n```\n(c, s, e, t)</p>\n\n<ul>\n<li>c: context of 512 wordpiece ids (question, document tokens and markup)</li>\n<li>s: Index pointing to the start of target answer span.</li>\n<li>e: Index pointing to the end of target answer span.</li>\n<li>t: answer type (0:short, 1:long, 2:yes, 3:no or 4:no-answer)\n```</li>\n</ul>\n\n<p>During training, we give c to BERT model and BERT model predicts (s, e, t). We compute loss based on predicted (s, e, t) and ground truth (s’, e’, t’).</p>\n\n<p>But we still don’t know what’s the relationship between the fields from the jsonl file and (c, s, e, t). Specifically how we convert document_text, question_text, short_answer and long_answer to (c, s, e, t)?</p>\n\n<p>In the paper they describe the process. </p>\n\n<p><em>For each document we generate all possible instances, by listing the document content starting at multiples of 128 tokens, effectively slid- ing a 512 token size window over the entire length of the document with a stride of 128 tokens.</em></p>\n\n<p>They are using sliding window approach for each indices 0th, 128th, 256th, ..., they pick 512 tokens from the index and make one training instance.</p>\n\n<p>A training instance may or may not include target answer.\n<code>\nIf all annotated short spans are contained in the instance\n   set the start and end target indices to point to the smallest span containing all the annotated short answer spans\nIf there are no annotated short spans but there is an annotated long answer span completely contained in the instance\n  set the start and end target indices to point to the entire long answer span.\nIf no short or long span can be found in the current instance\n  set the target start and end indices to point to the “[CLS]” token. \n</code></p>\n\n<h3>Markup tokens</h3>\n\n<p>*We introduce special markup tokens in the document to give the model a notion of which part of the document it is reading.</p>\n\n<p>The special tokens we introduced are of the form “[Paragraph=N]”, “[Table=N]”, and “[List=N]” at the beginning of the N-th paragraph,*</p>\n\n<ul>\n<li>[CLS] Mark meaning “tokens for question” follows.</li>\n<li>[SEP] Mark meaning “tokens from document_text” follows.</li>\n<li>Final [SEP] limiting the total size of each instance to 512 tokens.</li>\n</ul>\n\n<h3>Look real example.</h3>\n\n<p>IIUC a training instance look like this.\n<code>\n[CLS], question tokens, [SEP], slide windowed document_text, start index, end_index.\n</code>\nHow can I confirm my understanding? Can I see the actual training input to Bert? Yes you can have print in input_fn and observe the dataset.</p>\n\n<p>input_ids\n<code>\n[  101   104  2040  2003  1996  2148  3060  2152  5849  1999  2414   102\n   259   107   260   159  2152  3222  1997  2148  3088  1999  2414  3295\n 19817 10354  2389  6843  2675  1010  2414  4769 19817 10354  2389  6843\n  2675  1010  2414  1010 15868  2475  2078  1019 18927 12093  4868  1080\n  2382  1531  2382  1005  1005  1050  1014  1080  5718  1531  4261  1005\n  1005  1059  1013  4868  1012  2753  2620  2475  1080  1050  1014  1012\n 14010  2683  1080  1059  1013  4868  1012  2753  2620  2475  1025  1011\n  1014  1012 14010  2683 12093  1024  4868  1080  2382  1531  2382  1005\n  1005  1050  1014  1080  5718  1531  4261  1005  1005  1059  1013  4868\n  1012  2753  2620  2475  1080  1050  1014  1012 14010  2683  1080  1059\n  1013  4868  1012  2753  2620  2475  1025  1011  1014  1012 14010  2683\n  2152  5849 10030   266   109  1996  2152  3222  1997  2148  3088  1999\n  2414  2003  1996  8041  3260  2013  2148  3088  2000  1996  2142  2983\n  1012  2009  2003  2284  2012  2148  3088  2160  1010  1037  2311  2006\n 19817 10354  2389  6843  2675  1010  2414  1012  2004  2092  2004  4820\n  1996  4822  1997  1996  2152  5849  1010  1996  2311  2036  6184  1996\n  2148  3060 19972  1012  2009  2038  2042  1037  3694  2462  1008  3205\n  2311  2144  3196  1012   267   110  2148  3088  2160  2001  2328  2011\n  7935  1010  7658 10224  1004 21987 12474  2015  1999  1996  5687  2006\n  1996  2609  1997  2054  2018  2042 20653  1005  1055  3309  2127  2009\n  2001  7002  1999  4266  1012  1996  2311  2001  2881  2011  2909  7253\n  6243  1010  2007  6549  6743  2011 24873  5339 26261  6038  4059  1998\n  2909  2798 12819  1010  1998  2441  1999  4537  1012  1996  2311  2001\n  3734  2011  1996  2231  1997  2148  3088  2004  2049  2364  8041  3739\n  1999  1996  2866  1012  2076  2088  2162  2462  1010  3539  2704  5553\n 15488 16446  2973  2045  2096  9283  2148  3088  1005  1055  2162  3488\n  1012   268   111  1999  3777  1010  2148  3088  2150  1037  3072  1010\n  1998  6780  2013  1996  5663  2349  2000  2049  3343  1997  5762 18771\n  1012 11914  1010  1996  2311  2150  2019  8408  1010  2738  2084  1037\n  2152  3222  1012  2076  1996  3865  1010  1996  2311  1010  2029  2001\n  2028  1997  1996  2069  2148  3060  8041  6416  1999  1037  2270  2181\n  1010  2001  9416  2011 13337  2013  2105  1996  2088  1012  2076   102]\n</code>\nOkay this is list of ids, let’s make it human readable using the vocab file in <a href=\"https://www.kaggle.com/philculliton/bertjointbaseline\">https://www.kaggle.com/philculliton/bertjointbaseline</a></p>\n\n<p><code>\n['[CLS]', '[Q]', 'who', 'is', 'the', 'south', 'african', 'high', 'commissioner', 'in', 'london', '[SEP]', '[ContextId=-1]', '[NoLongAnswer]', '[ContextId=0]', '[Table=1]', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'location', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', 'address', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', ',', 'wc', '##2', '##n', '5', '##dp', 'coordinates', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'coordinates', ':', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'high', 'commissioner', 'vacant', '[ContextId=6]', '[Paragraph=1]', 'the', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'is', 'the', 'diplomatic', 'mission', 'from', 'south', 'africa', 'to', 'the', 'united', 'kingdom', '.', 'it', 'is', 'located', 'at', 'south', 'africa', 'house', ',', 'a', 'building', 'on', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', '.', 'as', 'well', 'as', 'containing', 'the', 'offices', 'of', 'the', 'high', 'commissioner', ',', 'the', 'building', 'also', 'hosts', 'the', 'south', 'african', 'consulate', '.', 'it', 'has', 'been', 'a', 'grade', 'ii', '*', 'listed', 'building', 'since', '1982', '.', '[ContextId=7]', '[Paragraph=2]', 'south', 'africa', 'house', 'was', 'built', 'by', 'holland', ',', 'han', '##nen', '&amp;', 'cub', '##itt', '##s', 'in', 'the', '1930s', 'on', 'the', 'site', 'of', 'what', 'had', 'been', 'morley', \"'\", 's', 'hotel', 'until', 'it', 'was', 'demolished', 'in', '1936', '.', 'the', 'building', 'was', 'designed', 'by', 'sir', 'herbert', 'baker', ',', 'with', 'architectural', 'sculpture', 'by', 'coe', '##rt', 'ste', '##yn', '##berg', 'and', 'sir', 'charles', 'wheeler', ',', 'and', 'opened', 'in', '1933', '.', 'the', 'building', 'was', 'acquired', 'by', 'the', 'government', 'of', 'south', 'africa', 'as', 'its', 'main', 'diplomatic', 'presence', 'in', 'the', 'uk', '.', 'during', 'world', 'war', 'ii', ',', 'prime', 'minister', 'jan', 'sm', '##uts', 'lived', 'there', 'while', 'conducting', 'south', 'africa', \"'\", 's', 'war', 'plans', '.', '[ContextId=8]', '[Paragraph=3]', 'in', '1961', ',', 'south', 'africa', 'became', 'a', 'republic', ',', 'and', 'withdrew', 'from', 'the', 'commonwealth', 'due', 'to', 'its', 'policy', 'of', 'racial', 'segregation', '.', 'accordingly', ',', 'the', 'building', 'became', 'an', 'embassy', ',', 'rather', 'than', 'a', 'high', 'commission', '.', 'during', 'the', '1980s', ',', 'the', 'building', ',', 'which', 'was', 'one', 'of', 'the', 'only', 'south', 'african', 'diplomatic', 'missions', 'in', 'a', 'public', 'area', ',', 'was', 'targeted', 'by', 'protesters', 'from', 'around', 'the', 'world', '.', 'during', '[SEP]']\n</code>\nOne more example.\n<code>\n['[CLS]', '[Q]', 'the', 'office', 'episode', 'when', 'they', 'sing', 'to', 'michael', '[SEP]', 'dean', '##gel', '##o', 'vickers', '(', 'fe', '##rrell', ')', 'on', 'how', 'to', 'properly', 'host', 'the', 'branch', \"'\", 's', 'annual', 'dun', '##die', 'awards', '.', 'michael', 'soon', 'learns', 'that', 'dean', '##gel', '##o', 'has', 'a', 'terrible', 'problem', 'with', 'speaking', 'in', 'front', 'of', 'others', '.', 'meanwhile', ',', 'erin', 'han', '##non', '(', 'ellie', 'kemp', '##er', ')', 'grows', 'to', 'dislike', 'her', 'boyfriend', ',', 'gabe', 'lewis', '(', 'zach', 'woods', ')', '.', '[ContextId=21]', '[Paragraph=3]', 'the', 'episode', '-', '-', 'which', 'was', 'originally', 'going', 'to', 'be', 'called', '`', '`', 'goodbye', ',', 'michael', 'part', '1', \"'\", \"'\", '-', '-', 'was', 'the', 'first', 'installment', 'in', 'the', 'series', 'to', 'be', 'both', 'written', 'and', 'directed', 'by', 'kali', '##ng', ',', 'who', 'also', 'portrays', 'kelly', 'kapoor', 'on', 'the', 'series', '.', 'the', 'episode', 'also', 'marks', 'the', 'second', 'appearance', 'of', 'fe', '##rrell', 'as', 'dean', '##gel', '##o', 'vickers', ';', 'fe', '##rrell', 'had', 'signed', 'onto', 'the', 'series', 'to', 'make', 'care', '##ll', \"'\", 's', 'exit', 'transition', 'easier', '.', 'the', 'episode', 'received', 'mostly', 'positive', 'reviews', 'from', 'television', 'critics', '.', '`', '`', 'michael', \"'\", 's', 'last', 'dun', '##dies', \"'\", \"'\", 'was', 'viewed', 'by', '6', '.', '84', '##9', 'million', 'viewers', 'and', 'received', 'a', '3', '.', '3', 'rating', 'among', 'adults', 'between', 'the', 'age', 'of', '18', 'and', '49', '.', 'the', 'episode', 'was', 'the', 'highest', '-', 'rated', 'nbc', 'series', 'of', 'the', 'week', 'that', 'it', 'aired', ',', 'as', 'well', 'as', 'the', 'sixth', '-', 'most', 'watched', 'episode', 'in', 'the', '18', '-', '-', '49', 'demographic', 'for', 'the', 'week', 'it', 'aired', '.', '[ContextId=22]', '[Paragraph=4]', 'at', 'the', 'office', ',', 'michael', 'scott', '(', 'steve', 'care', '##ll', ')', 'announces', 'to', 'the', 'employees', 'that', 'dean', '##gel', '##o', 'vickers', '(', 'will', 'fe', '##rrell', ')', 'will', 'be', 'his', 'co', '-', 'host', 'at', 'the', 'dun', '##dies', '.', 'the', 'dun', '##dies', 'are', 'an', 'annual', 'award', 'program', 'created', 'by', 'michael', 'to', 'mo', '##tiv', '##ate', 'his', 'employees', '.', 'the', 'idea', 'of', 'performance', 'is', 'wo', '##rri', '##some', 'to', 'dean', '##gel', '##o', ',', 'but', 'michael', 'insists', 'he', 'take', 'the', 'job', '.', 'michael', 'brings', 'some', 'of', 'the', 'staff', 'together', 'in', 'the', 'conference', 'room', 'to', 'help', 'dean', '##gel', '##o', 'get', 'prepared', 'for', 'the', 'show', ',', 'but', 'he', 'struggles', 'to', 'be', 'humorous', '.', 'andy', 'bernard', '(', 'ed', 'helm', '##s', ')', 'tries', 'to', 'help', 'him', ',', 'saying', 'he', 'should', 'just', 'think', 'of', 'performing', 'like', 'conducting', 'a', 'meeting', ',', 'but', 'michael', 'objects', ',', 'wanting', 'dean', '##gel', '##o', 'to', 'mimic', 'his', 'style', '[SEP]']\n</code></p>\n\n<h2>Some observations</h2>\n\n<ul>\n<li>It looks like they have [Q] token at the beginning of a question.</li>\n<li>There’s ContextId=N, which we should probably figure out in the paper.</li>\n<li>I thought start/end indicies of an answer is at the end of input_ids. but there are not.\n<ul><li>They might be in segment_ids in training data. I’m not entirely sure yet.</li></ul></li>\n</ul>",
  "messages": [
    {
      "id": "673441",
      "postDate": "11/15/2019 01:16:40",
      "content": "<h2>Background</h2>\n\n<p>So do you all understand how the training data look like which is fed into BERT for fine-tuning? I actually had no idea. Let me explain my understanding here. Hope this helps and I’d appreciate your feedback.</p>\n\n<h2>References</h2>\n\n<ul>\n<li>Their paper <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">PDF: A BERT Baseline for the Natural Questions</a></li>\n<li>Their <a href=\"https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq\">Tensorflow implementation</a></li>\n</ul>\n\n<h2>Training data</h2>\n\n<p>As you may have already seen we have training jsonl file which has\n- document_text (String)\n- question_text (String)\n- short_answer (Range)\n- long_answer (Range)</p>\n\n<p>According to the paper, a training set instance is four tuple as follows.\n```\n(c, s, e, t)</p>\n\n<ul>\n<li>c: context of 512 wordpiece ids (question, document tokens and markup)</li>\n<li>s: Index pointing to the start of target answer span.</li>\n<li>e: Index pointing to the end of target answer span.</li>\n<li>t: answer type (0:short, 1:long, 2:yes, 3:no or 4:no-answer)\n```</li>\n</ul>\n\n<p>During training, we give c to BERT model and BERT model predicts (s, e, t). We compute loss based on predicted (s, e, t) and ground truth (s’, e’, t’).</p>\n\n<p>But we still don’t know what’s the relationship between the fields from the jsonl file and (c, s, e, t). Specifically how we convert document_text, question_text, short_answer and long_answer to (c, s, e, t)?</p>\n\n<p>In the paper they describe the process. </p>\n\n<p><em>For each document we generate all possible instances, by listing the document content starting at multiples of 128 tokens, effectively slid- ing a 512 token size window over the entire length of the document with a stride of 128 tokens.</em></p>\n\n<p>They are using sliding window approach for each indices 0th, 128th, 256th, ..., they pick 512 tokens from the index and make one training instance.</p>\n\n<p>A training instance may or may not include target answer.\n<code>\nIf all annotated short spans are contained in the instance\n   set the start and end target indices to point to the smallest span containing all the annotated short answer spans\nIf there are no annotated short spans but there is an annotated long answer span completely contained in the instance\n  set the start and end target indices to point to the entire long answer span.\nIf no short or long span can be found in the current instance\n  set the target start and end indices to point to the “[CLS]” token. \n</code></p>\n\n<h3>Markup tokens</h3>\n\n<p>*We introduce special markup tokens in the document to give the model a notion of which part of the document it is reading.</p>\n\n<p>The special tokens we introduced are of the form “[Paragraph=N]”, “[Table=N]”, and “[List=N]” at the beginning of the N-th paragraph,*</p>\n\n<ul>\n<li>[CLS] Mark meaning “tokens for question” follows.</li>\n<li>[SEP] Mark meaning “tokens from document_text” follows.</li>\n<li>Final [SEP] limiting the total size of each instance to 512 tokens.</li>\n</ul>\n\n<h3>Look real example.</h3>\n\n<p>IIUC a training instance look like this.\n<code>\n[CLS], question tokens, [SEP], slide windowed document_text, start index, end_index.\n</code>\nHow can I confirm my understanding? Can I see the actual training input to Bert? Yes you can have print in input_fn and observe the dataset.</p>\n\n<p>input_ids\n<code>\n[  101   104  2040  2003  1996  2148  3060  2152  5849  1999  2414   102\n   259   107   260   159  2152  3222  1997  2148  3088  1999  2414  3295\n 19817 10354  2389  6843  2675  1010  2414  4769 19817 10354  2389  6843\n  2675  1010  2414  1010 15868  2475  2078  1019 18927 12093  4868  1080\n  2382  1531  2382  1005  1005  1050  1014  1080  5718  1531  4261  1005\n  1005  1059  1013  4868  1012  2753  2620  2475  1080  1050  1014  1012\n 14010  2683  1080  1059  1013  4868  1012  2753  2620  2475  1025  1011\n  1014  1012 14010  2683 12093  1024  4868  1080  2382  1531  2382  1005\n  1005  1050  1014  1080  5718  1531  4261  1005  1005  1059  1013  4868\n  1012  2753  2620  2475  1080  1050  1014  1012 14010  2683  1080  1059\n  1013  4868  1012  2753  2620  2475  1025  1011  1014  1012 14010  2683\n  2152  5849 10030   266   109  1996  2152  3222  1997  2148  3088  1999\n  2414  2003  1996  8041  3260  2013  2148  3088  2000  1996  2142  2983\n  1012  2009  2003  2284  2012  2148  3088  2160  1010  1037  2311  2006\n 19817 10354  2389  6843  2675  1010  2414  1012  2004  2092  2004  4820\n  1996  4822  1997  1996  2152  5849  1010  1996  2311  2036  6184  1996\n  2148  3060 19972  1012  2009  2038  2042  1037  3694  2462  1008  3205\n  2311  2144  3196  1012   267   110  2148  3088  2160  2001  2328  2011\n  7935  1010  7658 10224  1004 21987 12474  2015  1999  1996  5687  2006\n  1996  2609  1997  2054  2018  2042 20653  1005  1055  3309  2127  2009\n  2001  7002  1999  4266  1012  1996  2311  2001  2881  2011  2909  7253\n  6243  1010  2007  6549  6743  2011 24873  5339 26261  6038  4059  1998\n  2909  2798 12819  1010  1998  2441  1999  4537  1012  1996  2311  2001\n  3734  2011  1996  2231  1997  2148  3088  2004  2049  2364  8041  3739\n  1999  1996  2866  1012  2076  2088  2162  2462  1010  3539  2704  5553\n 15488 16446  2973  2045  2096  9283  2148  3088  1005  1055  2162  3488\n  1012   268   111  1999  3777  1010  2148  3088  2150  1037  3072  1010\n  1998  6780  2013  1996  5663  2349  2000  2049  3343  1997  5762 18771\n  1012 11914  1010  1996  2311  2150  2019  8408  1010  2738  2084  1037\n  2152  3222  1012  2076  1996  3865  1010  1996  2311  1010  2029  2001\n  2028  1997  1996  2069  2148  3060  8041  6416  1999  1037  2270  2181\n  1010  2001  9416  2011 13337  2013  2105  1996  2088  1012  2076   102]\n</code>\nOkay this is list of ids, let’s make it human readable using the vocab file in <a href=\"https://www.kaggle.com/philculliton/bertjointbaseline\">https://www.kaggle.com/philculliton/bertjointbaseline</a></p>\n\n<p><code>\n['[CLS]', '[Q]', 'who', 'is', 'the', 'south', 'african', 'high', 'commissioner', 'in', 'london', '[SEP]', '[ContextId=-1]', '[NoLongAnswer]', '[ContextId=0]', '[Table=1]', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'location', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', 'address', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', ',', 'wc', '##2', '##n', '5', '##dp', 'coordinates', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'coordinates', ':', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'high', 'commissioner', 'vacant', '[ContextId=6]', '[Paragraph=1]', 'the', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'is', 'the', 'diplomatic', 'mission', 'from', 'south', 'africa', 'to', 'the', 'united', 'kingdom', '.', 'it', 'is', 'located', 'at', 'south', 'africa', 'house', ',', 'a', 'building', 'on', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', '.', 'as', 'well', 'as', 'containing', 'the', 'offices', 'of', 'the', 'high', 'commissioner', ',', 'the', 'building', 'also', 'hosts', 'the', 'south', 'african', 'consulate', '.', 'it', 'has', 'been', 'a', 'grade', 'ii', '*', 'listed', 'building', 'since', '1982', '.', '[ContextId=7]', '[Paragraph=2]', 'south', 'africa', 'house', 'was', 'built', 'by', 'holland', ',', 'han', '##nen', '&amp;', 'cub', '##itt', '##s', 'in', 'the', '1930s', 'on', 'the', 'site', 'of', 'what', 'had', 'been', 'morley', \"'\", 's', 'hotel', 'until', 'it', 'was', 'demolished', 'in', '1936', '.', 'the', 'building', 'was', 'designed', 'by', 'sir', 'herbert', 'baker', ',', 'with', 'architectural', 'sculpture', 'by', 'coe', '##rt', 'ste', '##yn', '##berg', 'and', 'sir', 'charles', 'wheeler', ',', 'and', 'opened', 'in', '1933', '.', 'the', 'building', 'was', 'acquired', 'by', 'the', 'government', 'of', 'south', 'africa', 'as', 'its', 'main', 'diplomatic', 'presence', 'in', 'the', 'uk', '.', 'during', 'world', 'war', 'ii', ',', 'prime', 'minister', 'jan', 'sm', '##uts', 'lived', 'there', 'while', 'conducting', 'south', 'africa', \"'\", 's', 'war', 'plans', '.', '[ContextId=8]', '[Paragraph=3]', 'in', '1961', ',', 'south', 'africa', 'became', 'a', 'republic', ',', 'and', 'withdrew', 'from', 'the', 'commonwealth', 'due', 'to', 'its', 'policy', 'of', 'racial', 'segregation', '.', 'accordingly', ',', 'the', 'building', 'became', 'an', 'embassy', ',', 'rather', 'than', 'a', 'high', 'commission', '.', 'during', 'the', '1980s', ',', 'the', 'building', ',', 'which', 'was', 'one', 'of', 'the', 'only', 'south', 'african', 'diplomatic', 'missions', 'in', 'a', 'public', 'area', ',', 'was', 'targeted', 'by', 'protesters', 'from', 'around', 'the', 'world', '.', 'during', '[SEP]']\n</code>\nOne more example.\n<code>\n['[CLS]', '[Q]', 'the', 'office', 'episode', 'when', 'they', 'sing', 'to', 'michael', '[SEP]', 'dean', '##gel', '##o', 'vickers', '(', 'fe', '##rrell', ')', 'on', 'how', 'to', 'properly', 'host', 'the', 'branch', \"'\", 's', 'annual', 'dun', '##die', 'awards', '.', 'michael', 'soon', 'learns', 'that', 'dean', '##gel', '##o', 'has', 'a', 'terrible', 'problem', 'with', 'speaking', 'in', 'front', 'of', 'others', '.', 'meanwhile', ',', 'erin', 'han', '##non', '(', 'ellie', 'kemp', '##er', ')', 'grows', 'to', 'dislike', 'her', 'boyfriend', ',', 'gabe', 'lewis', '(', 'zach', 'woods', ')', '.', '[ContextId=21]', '[Paragraph=3]', 'the', 'episode', '-', '-', 'which', 'was', 'originally', 'going', 'to', 'be', 'called', '`', '`', 'goodbye', ',', 'michael', 'part', '1', \"'\", \"'\", '-', '-', 'was', 'the', 'first', 'installment', 'in', 'the', 'series', 'to', 'be', 'both', 'written', 'and', 'directed', 'by', 'kali', '##ng', ',', 'who', 'also', 'portrays', 'kelly', 'kapoor', 'on', 'the', 'series', '.', 'the', 'episode', 'also', 'marks', 'the', 'second', 'appearance', 'of', 'fe', '##rrell', 'as', 'dean', '##gel', '##o', 'vickers', ';', 'fe', '##rrell', 'had', 'signed', 'onto', 'the', 'series', 'to', 'make', 'care', '##ll', \"'\", 's', 'exit', 'transition', 'easier', '.', 'the', 'episode', 'received', 'mostly', 'positive', 'reviews', 'from', 'television', 'critics', '.', '`', '`', 'michael', \"'\", 's', 'last', 'dun', '##dies', \"'\", \"'\", 'was', 'viewed', 'by', '6', '.', '84', '##9', 'million', 'viewers', 'and', 'received', 'a', '3', '.', '3', 'rating', 'among', 'adults', 'between', 'the', 'age', 'of', '18', 'and', '49', '.', 'the', 'episode', 'was', 'the', 'highest', '-', 'rated', 'nbc', 'series', 'of', 'the', 'week', 'that', 'it', 'aired', ',', 'as', 'well', 'as', 'the', 'sixth', '-', 'most', 'watched', 'episode', 'in', 'the', '18', '-', '-', '49', 'demographic', 'for', 'the', 'week', 'it', 'aired', '.', '[ContextId=22]', '[Paragraph=4]', 'at', 'the', 'office', ',', 'michael', 'scott', '(', 'steve', 'care', '##ll', ')', 'announces', 'to', 'the', 'employees', 'that', 'dean', '##gel', '##o', 'vickers', '(', 'will', 'fe', '##rrell', ')', 'will', 'be', 'his', 'co', '-', 'host', 'at', 'the', 'dun', '##dies', '.', 'the', 'dun', '##dies', 'are', 'an', 'annual', 'award', 'program', 'created', 'by', 'michael', 'to', 'mo', '##tiv', '##ate', 'his', 'employees', '.', 'the', 'idea', 'of', 'performance', 'is', 'wo', '##rri', '##some', 'to', 'dean', '##gel', '##o', ',', 'but', 'michael', 'insists', 'he', 'take', 'the', 'job', '.', 'michael', 'brings', 'some', 'of', 'the', 'staff', 'together', 'in', 'the', 'conference', 'room', 'to', 'help', 'dean', '##gel', '##o', 'get', 'prepared', 'for', 'the', 'show', ',', 'but', 'he', 'struggles', 'to', 'be', 'humorous', '.', 'andy', 'bernard', '(', 'ed', 'helm', '##s', ')', 'tries', 'to', 'help', 'him', ',', 'saying', 'he', 'should', 'just', 'think', 'of', 'performing', 'like', 'conducting', 'a', 'meeting', ',', 'but', 'michael', 'objects', ',', 'wanting', 'dean', '##gel', '##o', 'to', 'mimic', 'his', 'style', '[SEP]']\n</code></p>\n\n<h2>Some observations</h2>\n\n<ul>\n<li>It looks like they have [Q] token at the beginning of a question.</li>\n<li>There’s ContextId=N, which we should probably figure out in the paper.</li>\n<li>I thought start/end indicies of an answer is at the end of input_ids. but there are not.\n<ul><li>They might be in segment_ids in training data. I’m not entirely sure yet.</li></ul></li>\n</ul>",
      "rawMarkdown": "## Background\nSo do you all understand how the training data look like which is fed into BERT for fine-tuning? I actually had no idea. Let me explain my understanding here. Hope this helps and I’d appreciate your feedback.\n\n## References\n- Their paper [PDF: A BERT Baseline for the Natural Questions](https://arxiv.org/pdf/1901.08634.pdf)\n- Their [Tensorflow implementation](https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq)\n\n## Training data\nAs you may have already seen we have training jsonl file which has\n- document_text (String)\n- question_text (String)\n- short_answer (Range)\n- long_answer (Range)\n\nAccording to the paper, a training set instance is four tuple as follows.\n```\n(c, s, e, t)\n\n- c: context of 512 wordpiece ids (question, document tokens and markup)\n- s: Index pointing to the start of target answer span.\n- e: Index pointing to the end of target answer span.\n- t: answer type (0:short, 1:long, 2:yes, 3:no or 4:no-answer)\n```\n\nDuring training, we give c to BERT model and BERT model predicts (s, e, t). We compute loss based on predicted (s, e, t) and ground truth (s’, e’, t’).\n\nBut we still don’t know what’s the relationship between the fields from the jsonl file and (c, s, e, t). Specifically how we convert document_text, question_text, short_answer and long_answer to (c, s, e, t)?\n\nIn the paper they describe the process. \n\n*For each document we generate all possible instances, by listing the document content starting at multiples of 128 tokens, effectively slid- ing a 512 token size window over the entire length of the document with a stride of 128 tokens.*\n\nThey are using sliding window approach for each indices 0th, 128th, 256th, ..., they pick 512 tokens from the index and make one training instance.\n\nA training instance may or may not include target answer.\n```\nIf all annotated short spans are contained in the instance\n   set the start and end target indices to point to the smallest span containing all the annotated short answer spans\nIf there are no annotated short spans but there is an annotated long answer span completely contained in the instance\n  set the start and end target indices to point to the entire long answer span.\nIf no short or long span can be found in the current instance\n  set the target start and end indices to point to the “[CLS]” token. \n```\n### Markup tokens\n*We introduce special markup tokens in the document to give the model a notion of which part of the document it is reading.\n\nThe special tokens we introduced are of the form “[Paragraph=N]”, “[Table=N]”, and “[List=N]” at the beginning of the N-th paragraph,*\n\n- [CLS] Mark meaning “tokens for question” follows.\n- [SEP] Mark meaning “tokens from document_text” follows.\n- Final [SEP] limiting the total size of each instance to 512 tokens.\n\n### Look real example.\nIIUC a training instance look like this.\n```\n[CLS], question tokens, [SEP], slide windowed document_text, start index, end_index.\n```\nHow can I confirm my understanding? Can I see the actual training input to Bert? Yes you can have print in input_fn and observe the dataset.\n\ninput_ids\n```\n[  101   104  2040  2003  1996  2148  3060  2152  5849  1999  2414   102\n   259   107   260   159  2152  3222  1997  2148  3088  1999  2414  3295\n 19817 10354  2389  6843  2675  1010  2414  4769 19817 10354  2389  6843\n  2675  1010  2414  1010 15868  2475  2078  1019 18927 12093  4868  1080\n  2382  1531  2382  1005  1005  1050  1014  1080  5718  1531  4261  1005\n  1005  1059  1013  4868  1012  2753  2620  2475  1080  1050  1014  1012\n 14010  2683  1080  1059  1013  4868  1012  2753  2620  2475  1025  1011\n  1014  1012 14010  2683 12093  1024  4868  1080  2382  1531  2382  1005\n  1005  1050  1014  1080  5718  1531  4261  1005  1005  1059  1013  4868\n  1012  2753  2620  2475  1080  1050  1014  1012 14010  2683  1080  1059\n  1013  4868  1012  2753  2620  2475  1025  1011  1014  1012 14010  2683\n  2152  5849 10030   266   109  1996  2152  3222  1997  2148  3088  1999\n  2414  2003  1996  8041  3260  2013  2148  3088  2000  1996  2142  2983\n  1012  2009  2003  2284  2012  2148  3088  2160  1010  1037  2311  2006\n 19817 10354  2389  6843  2675  1010  2414  1012  2004  2092  2004  4820\n  1996  4822  1997  1996  2152  5849  1010  1996  2311  2036  6184  1996\n  2148  3060 19972  1012  2009  2038  2042  1037  3694  2462  1008  3205\n  2311  2144  3196  1012   267   110  2148  3088  2160  2001  2328  2011\n  7935  1010  7658 10224  1004 21987 12474  2015  1999  1996  5687  2006\n  1996  2609  1997  2054  2018  2042 20653  1005  1055  3309  2127  2009\n  2001  7002  1999  4266  1012  1996  2311  2001  2881  2011  2909  7253\n  6243  1010  2007  6549  6743  2011 24873  5339 26261  6038  4059  1998\n  2909  2798 12819  1010  1998  2441  1999  4537  1012  1996  2311  2001\n  3734  2011  1996  2231  1997  2148  3088  2004  2049  2364  8041  3739\n  1999  1996  2866  1012  2076  2088  2162  2462  1010  3539  2704  5553\n 15488 16446  2973  2045  2096  9283  2148  3088  1005  1055  2162  3488\n  1012   268   111  1999  3777  1010  2148  3088  2150  1037  3072  1010\n  1998  6780  2013  1996  5663  2349  2000  2049  3343  1997  5762 18771\n  1012 11914  1010  1996  2311  2150  2019  8408  1010  2738  2084  1037\n  2152  3222  1012  2076  1996  3865  1010  1996  2311  1010  2029  2001\n  2028  1997  1996  2069  2148  3060  8041  6416  1999  1037  2270  2181\n  1010  2001  9416  2011 13337  2013  2105  1996  2088  1012  2076   102]\n```\nOkay this is list of ids, let’s make it human readable using the vocab file in https://www.kaggle.com/philculliton/bertjointbaseline\n\n```\n['[CLS]', '[Q]', 'who', 'is', 'the', 'south', 'african', 'high', 'commissioner', 'in', 'london', '[SEP]', '[ContextId=-1]', '[NoLongAnswer]', '[ContextId=0]', '[Table=1]', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'location', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', 'address', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', ',', 'wc', '##2', '##n', '5', '##dp', 'coordinates', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'coordinates', ':', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'high', 'commissioner', 'vacant', '[ContextId=6]', '[Paragraph=1]', 'the', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'is', 'the', 'diplomatic', 'mission', 'from', 'south', 'africa', 'to', 'the', 'united', 'kingdom', '.', 'it', 'is', 'located', 'at', 'south', 'africa', 'house', ',', 'a', 'building', 'on', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', '.', 'as', 'well', 'as', 'containing', 'the', 'offices', 'of', 'the', 'high', 'commissioner', ',', 'the', 'building', 'also', 'hosts', 'the', 'south', 'african', 'consulate', '.', 'it', 'has', 'been', 'a', 'grade', 'ii', '*', 'listed', 'building', 'since', '1982', '.', '[ContextId=7]', '[Paragraph=2]', 'south', 'africa', 'house', 'was', 'built', 'by', 'holland', ',', 'han', '##nen', '&amp;', 'cub', '##itt', '##s', 'in', 'the', '1930s', 'on', 'the', 'site', 'of', 'what', 'had', 'been', 'morley', \"'\", 's', 'hotel', 'until', 'it', 'was', 'demolished', 'in', '1936', '.', 'the', 'building', 'was', 'designed', 'by', 'sir', 'herbert', 'baker', ',', 'with', 'architectural', 'sculpture', 'by', 'coe', '##rt', 'ste', '##yn', '##berg', 'and', 'sir', 'charles', 'wheeler', ',', 'and', 'opened', 'in', '1933', '.', 'the', 'building', 'was', 'acquired', 'by', 'the', 'government', 'of', 'south', 'africa', 'as', 'its', 'main', 'diplomatic', 'presence', 'in', 'the', 'uk', '.', 'during', 'world', 'war', 'ii', ',', 'prime', 'minister', 'jan', 'sm', '##uts', 'lived', 'there', 'while', 'conducting', 'south', 'africa', \"'\", 's', 'war', 'plans', '.', '[ContextId=8]', '[Paragraph=3]', 'in', '1961', ',', 'south', 'africa', 'became', 'a', 'republic', ',', 'and', 'withdrew', 'from', 'the', 'commonwealth', 'due', 'to', 'its', 'policy', 'of', 'racial', 'segregation', '.', 'accordingly', ',', 'the', 'building', 'became', 'an', 'embassy', ',', 'rather', 'than', 'a', 'high', 'commission', '.', 'during', 'the', '1980s', ',', 'the', 'building', ',', 'which', 'was', 'one', 'of', 'the', 'only', 'south', 'african', 'diplomatic', 'missions', 'in', 'a', 'public', 'area', ',', 'was', 'targeted', 'by', 'protesters', 'from', 'around', 'the', 'world', '.', 'during', '[SEP]']\n```\nOne more example.\n```\n['[CLS]', '[Q]', 'the', 'office', 'episode', 'when', 'they', 'sing', 'to', 'michael', '[SEP]', 'dean', '##gel', '##o', 'vickers', '(', 'fe', '##rrell', ')', 'on', 'how', 'to', 'properly', 'host', 'the', 'branch', \"'\", 's', 'annual', 'dun', '##die', 'awards', '.', 'michael', 'soon', 'learns', 'that', 'dean', '##gel', '##o', 'has', 'a', 'terrible', 'problem', 'with', 'speaking', 'in', 'front', 'of', 'others', '.', 'meanwhile', ',', 'erin', 'han', '##non', '(', 'ellie', 'kemp', '##er', ')', 'grows', 'to', 'dislike', 'her', 'boyfriend', ',', 'gabe', 'lewis', '(', 'zach', 'woods', ')', '.', '[ContextId=21]', '[Paragraph=3]', 'the', 'episode', '-', '-', 'which', 'was', 'originally', 'going', 'to', 'be', 'called', '`', '`', 'goodbye', ',', 'michael', 'part', '1', \"'\", \"'\", '-', '-', 'was', 'the', 'first', 'installment', 'in', 'the', 'series', 'to', 'be', 'both', 'written', 'and', 'directed', 'by', 'kali', '##ng', ',', 'who', 'also', 'portrays', 'kelly', 'kapoor', 'on', 'the', 'series', '.', 'the', 'episode', 'also', 'marks', 'the', 'second', 'appearance', 'of', 'fe', '##rrell', 'as', 'dean', '##gel', '##o', 'vickers', ';', 'fe', '##rrell', 'had', 'signed', 'onto', 'the', 'series', 'to', 'make', 'care', '##ll', \"'\", 's', 'exit', 'transition', 'easier', '.', 'the', 'episode', 'received', 'mostly', 'positive', 'reviews', 'from', 'television', 'critics', '.', '`', '`', 'michael', \"'\", 's', 'last', 'dun', '##dies', \"'\", \"'\", 'was', 'viewed', 'by', '6', '.', '84', '##9', 'million', 'viewers', 'and', 'received', 'a', '3', '.', '3', 'rating', 'among', 'adults', 'between', 'the', 'age', 'of', '18', 'and', '49', '.', 'the', 'episode', 'was', 'the', 'highest', '-', 'rated', 'nbc', 'series', 'of', 'the', 'week', 'that', 'it', 'aired', ',', 'as', 'well', 'as', 'the', 'sixth', '-', 'most', 'watched', 'episode', 'in', 'the', '18', '-', '-', '49', 'demographic', 'for', 'the', 'week', 'it', 'aired', '.', '[ContextId=22]', '[Paragraph=4]', 'at', 'the', 'office', ',', 'michael', 'scott', '(', 'steve', 'care', '##ll', ')', 'announces', 'to', 'the', 'employees', 'that', 'dean', '##gel', '##o', 'vickers', '(', 'will', 'fe', '##rrell', ')', 'will', 'be', 'his', 'co', '-', 'host', 'at', 'the', 'dun', '##dies', '.', 'the', 'dun', '##dies', 'are', 'an', 'annual', 'award', 'program', 'created', 'by', 'michael', 'to', 'mo', '##tiv', '##ate', 'his', 'employees', '.', 'the', 'idea', 'of', 'performance', 'is', 'wo', '##rri', '##some', 'to', 'dean', '##gel', '##o', ',', 'but', 'michael', 'insists', 'he', 'take', 'the', 'job', '.', 'michael', 'brings', 'some', 'of', 'the', 'staff', 'together', 'in', 'the', 'conference', 'room', 'to', 'help', 'dean', '##gel', '##o', 'get', 'prepared', 'for', 'the', 'show', ',', 'but', 'he', 'struggles', 'to', 'be', 'humorous', '.', 'andy', 'bernard', '(', 'ed', 'helm', '##s', ')', 'tries', 'to', 'help', 'him', ',', 'saying', 'he', 'should', 'just', 'think', 'of', 'performing', 'like', 'conducting', 'a', 'meeting', ',', 'but', 'michael', 'objects', ',', 'wanting', 'dean', '##gel', '##o', 'to', 'mimic', 'his', 'style', '[SEP]']\n```\n## Some observations\n- It looks like they have [Q] token at the beginning of a question.\n- There’s ContextId=N, which we should probably figure out in the paper.\n- I thought start/end indicies of an answer is at the end of input_ids. but there are not.\n   - They might be in segment_ids in training data. I’m not entirely sure yet.",
      "votes": null
    },
    {
      "id": "682249",
      "postDate": "11/27/2019 05:58:49",
      "content": "<p>Thanks for this amazing explanation of what's being inputted into the BERT model! This needs more attention!</p>",
      "rawMarkdown": "Thanks for this amazing explanation of what's being inputted into the BERT model! This needs more attention!",
      "votes": null
    },
    {
      "id": "686695",
      "postDate": "12/03/2019 12:20:51",
      "content": "<p>Thank you. I'm glad I could help.</p>",
      "rawMarkdown": "Thank you. I'm glad I could help.",
      "votes": null
    },
    {
      "id": "690321",
      "postDate": "12/08/2019 11:40:53",
      "content": "<p>good thread .. thanks.</p>",
      "rawMarkdown": "good thread .. thanks.",
      "votes": null
    },
    {
      "id": "692204",
      "postDate": "12/11/2019 02:46:17",
      "content": "<p>Really useful, thanks. <a href=\"/higepon\">@higepon</a> </p>",
      "rawMarkdown": "Really useful, thanks. @higepon",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 682249,
      "author_name": "tanlikesmath",
      "author_url": "",
      "post_date": "11/27/2019 05:58:49",
      "content": "<p>Thanks for this amazing explanation of what's being inputted into the BERT model! This needs more attention!</p>",
      "votes": null,
      "replies": [
        {
          "id": 686695,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/03/2019 12:20:51",
          "content": "<p>Thank you. I'm glad I could help.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 690321,
      "author_name": "johnwill225",
      "author_url": "",
      "post_date": "12/08/2019 11:40:53",
      "content": "<p>good thread .. thanks.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 692204,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/11/2019 02:46:17",
      "content": "<p>Really useful, thanks. <a href=\"/higepon\">@higepon</a> </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "673441": "## Background\nSo do you all understand how the training data look like which is fed into BERT for fine-tuning? I actually had no idea. Let me explain my understanding here. Hope this helps and I’d appreciate your feedback.\n\n## References\n- Their paper [PDF: A BERT Baseline for the Natural Questions](https://arxiv.org/pdf/1901.08634.pdf)\n- Their [Tensorflow implementation](https://www.kaggle.com/philculliton/using-tensorflow-2-0-w-bert-on-nq)\n\n## Training data\nAs you may have already seen we have training jsonl file which has\n- document_text (String)\n- question_text (String)\n- short_answer (Range)\n- long_answer (Range)\n\nAccording to the paper, a training set instance is four tuple as follows.\n```\n(c, s, e, t)\n\n- c: context of 512 wordpiece ids (question, document tokens and markup)\n- s: Index pointing to the start of target answer span.\n- e: Index pointing to the end of target answer span.\n- t: answer type (0:short, 1:long, 2:yes, 3:no or 4:no-answer)\n```\n\nDuring training, we give c to BERT model and BERT model predicts (s, e, t). We compute loss based on predicted (s, e, t) and ground truth (s’, e’, t’).\n\nBut we still don’t know what’s the relationship between the fields from the jsonl file and (c, s, e, t). Specifically how we convert document_text, question_text, short_answer and long_answer to (c, s, e, t)?\n\nIn the paper they describe the process. \n\n*For each document we generate all possible instances, by listing the document content starting at multiples of 128 tokens, effectively slid- ing a 512 token size window over the entire length of the document with a stride of 128 tokens.*\n\nThey are using sliding window approach for each indices 0th, 128th, 256th, ..., they pick 512 tokens from the index and make one training instance.\n\nA training instance may or may not include target answer.\n```\nIf all annotated short spans are contained in the instance\n   set the start and end target indices to point to the smallest span containing all the annotated short answer spans\nIf there are no annotated short spans but there is an annotated long answer span completely contained in the instance\n  set the start and end target indices to point to the entire long answer span.\nIf no short or long span can be found in the current instance\n  set the target start and end indices to point to the “[CLS]” token. \n```\n### Markup tokens\n*We introduce special markup tokens in the document to give the model a notion of which part of the document it is reading.\n\nThe special tokens we introduced are of the form “[Paragraph=N]”, “[Table=N]”, and “[List=N]” at the beginning of the N-th paragraph,*\n\n- [CLS] Mark meaning “tokens for question” follows.\n- [SEP] Mark meaning “tokens from document_text” follows.\n- Final [SEP] limiting the total size of each instance to 512 tokens.\n\n### Look real example.\nIIUC a training instance look like this.\n```\n[CLS], question tokens, [SEP], slide windowed document_text, start index, end_index.\n```\nHow can I confirm my understanding? Can I see the actual training input to Bert? Yes you can have print in input_fn and observe the dataset.\n\ninput_ids\n```\n[  101   104  2040  2003  1996  2148  3060  2152  5849  1999  2414   102\n   259   107   260   159  2152  3222  1997  2148  3088  1999  2414  3295\n 19817 10354  2389  6843  2675  1010  2414  4769 19817 10354  2389  6843\n  2675  1010  2414  1010 15868  2475  2078  1019 18927 12093  4868  1080\n  2382  1531  2382  1005  1005  1050  1014  1080  5718  1531  4261  1005\n  1005  1059  1013  4868  1012  2753  2620  2475  1080  1050  1014  1012\n 14010  2683  1080  1059  1013  4868  1012  2753  2620  2475  1025  1011\n  1014  1012 14010  2683 12093  1024  4868  1080  2382  1531  2382  1005\n  1005  1050  1014  1080  5718  1531  4261  1005  1005  1059  1013  4868\n  1012  2753  2620  2475  1080  1050  1014  1012 14010  2683  1080  1059\n  1013  4868  1012  2753  2620  2475  1025  1011  1014  1012 14010  2683\n  2152  5849 10030   266   109  1996  2152  3222  1997  2148  3088  1999\n  2414  2003  1996  8041  3260  2013  2148  3088  2000  1996  2142  2983\n  1012  2009  2003  2284  2012  2148  3088  2160  1010  1037  2311  2006\n 19817 10354  2389  6843  2675  1010  2414  1012  2004  2092  2004  4820\n  1996  4822  1997  1996  2152  5849  1010  1996  2311  2036  6184  1996\n  2148  3060 19972  1012  2009  2038  2042  1037  3694  2462  1008  3205\n  2311  2144  3196  1012   267   110  2148  3088  2160  2001  2328  2011\n  7935  1010  7658 10224  1004 21987 12474  2015  1999  1996  5687  2006\n  1996  2609  1997  2054  2018  2042 20653  1005  1055  3309  2127  2009\n  2001  7002  1999  4266  1012  1996  2311  2001  2881  2011  2909  7253\n  6243  1010  2007  6549  6743  2011 24873  5339 26261  6038  4059  1998\n  2909  2798 12819  1010  1998  2441  1999  4537  1012  1996  2311  2001\n  3734  2011  1996  2231  1997  2148  3088  2004  2049  2364  8041  3739\n  1999  1996  2866  1012  2076  2088  2162  2462  1010  3539  2704  5553\n 15488 16446  2973  2045  2096  9283  2148  3088  1005  1055  2162  3488\n  1012   268   111  1999  3777  1010  2148  3088  2150  1037  3072  1010\n  1998  6780  2013  1996  5663  2349  2000  2049  3343  1997  5762 18771\n  1012 11914  1010  1996  2311  2150  2019  8408  1010  2738  2084  1037\n  2152  3222  1012  2076  1996  3865  1010  1996  2311  1010  2029  2001\n  2028  1997  1996  2069  2148  3060  8041  6416  1999  1037  2270  2181\n  1010  2001  9416  2011 13337  2013  2105  1996  2088  1012  2076   102]\n```\nOkay this is list of ids, let’s make it human readable using the vocab file in https://www.kaggle.com/philculliton/bertjointbaseline\n\n```\n['[CLS]', '[Q]', 'who', 'is', 'the', 'south', 'african', 'high', 'commissioner', 'in', 'london', '[SEP]', '[ContextId=-1]', '[NoLongAnswer]', '[ContextId=0]', '[Table=1]', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'location', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', 'address', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', ',', 'wc', '##2', '##n', '5', '##dp', 'coordinates', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'coordinates', ':', '51', '°', '30', '′', '30', \"'\", \"'\", 'n', '0', '°', '07', '′', '37', \"'\", \"'\", 'w', '/', '51', '.', '50', '##8', '##2', '°', 'n', '0', '.', '126', '##9', '°', 'w', '/', '51', '.', '50', '##8', '##2', ';', '-', '0', '.', '126', '##9', 'high', 'commissioner', 'vacant', '[ContextId=6]', '[Paragraph=1]', 'the', 'high', 'commission', 'of', 'south', 'africa', 'in', 'london', 'is', 'the', 'diplomatic', 'mission', 'from', 'south', 'africa', 'to', 'the', 'united', 'kingdom', '.', 'it', 'is', 'located', 'at', 'south', 'africa', 'house', ',', 'a', 'building', 'on', 'tr', '##af', '##al', '##gar', 'square', ',', 'london', '.', 'as', 'well', 'as', 'containing', 'the', 'offices', 'of', 'the', 'high', 'commissioner', ',', 'the', 'building', 'also', 'hosts', 'the', 'south', 'african', 'consulate', '.', 'it', 'has', 'been', 'a', 'grade', 'ii', '*', 'listed', 'building', 'since', '1982', '.', '[ContextId=7]', '[Paragraph=2]', 'south', 'africa', 'house', 'was', 'built', 'by', 'holland', ',', 'han', '##nen', '&amp;', 'cub', '##itt', '##s', 'in', 'the', '1930s', 'on', 'the', 'site', 'of', 'what', 'had', 'been', 'morley', \"'\", 's', 'hotel', 'until', 'it', 'was', 'demolished', 'in', '1936', '.', 'the', 'building', 'was', 'designed', 'by', 'sir', 'herbert', 'baker', ',', 'with', 'architectural', 'sculpture', 'by', 'coe', '##rt', 'ste', '##yn', '##berg', 'and', 'sir', 'charles', 'wheeler', ',', 'and', 'opened', 'in', '1933', '.', 'the', 'building', 'was', 'acquired', 'by', 'the', 'government', 'of', 'south', 'africa', 'as', 'its', 'main', 'diplomatic', 'presence', 'in', 'the', 'uk', '.', 'during', 'world', 'war', 'ii', ',', 'prime', 'minister', 'jan', 'sm', '##uts', 'lived', 'there', 'while', 'conducting', 'south', 'africa', \"'\", 's', 'war', 'plans', '.', '[ContextId=8]', '[Paragraph=3]', 'in', '1961', ',', 'south', 'africa', 'became', 'a', 'republic', ',', 'and', 'withdrew', 'from', 'the', 'commonwealth', 'due', 'to', 'its', 'policy', 'of', 'racial', 'segregation', '.', 'accordingly', ',', 'the', 'building', 'became', 'an', 'embassy', ',', 'rather', 'than', 'a', 'high', 'commission', '.', 'during', 'the', '1980s', ',', 'the', 'building', ',', 'which', 'was', 'one', 'of', 'the', 'only', 'south', 'african', 'diplomatic', 'missions', 'in', 'a', 'public', 'area', ',', 'was', 'targeted', 'by', 'protesters', 'from', 'around', 'the', 'world', '.', 'during', '[SEP]']\n```\nOne more example.\n```\n['[CLS]', '[Q]', 'the', 'office', 'episode', 'when', 'they', 'sing', 'to', 'michael', '[SEP]', 'dean', '##gel', '##o', 'vickers', '(', 'fe', '##rrell', ')', 'on', 'how', 'to', 'properly', 'host', 'the', 'branch', \"'\", 's', 'annual', 'dun', '##die', 'awards', '.', 'michael', 'soon', 'learns', 'that', 'dean', '##gel', '##o', 'has', 'a', 'terrible', 'problem', 'with', 'speaking', 'in', 'front', 'of', 'others', '.', 'meanwhile', ',', 'erin', 'han', '##non', '(', 'ellie', 'kemp', '##er', ')', 'grows', 'to', 'dislike', 'her', 'boyfriend', ',', 'gabe', 'lewis', '(', 'zach', 'woods', ')', '.', '[ContextId=21]', '[Paragraph=3]', 'the', 'episode', '-', '-', 'which', 'was', 'originally', 'going', 'to', 'be', 'called', '`', '`', 'goodbye', ',', 'michael', 'part', '1', \"'\", \"'\", '-', '-', 'was', 'the', 'first', 'installment', 'in', 'the', 'series', 'to', 'be', 'both', 'written', 'and', 'directed', 'by', 'kali', '##ng', ',', 'who', 'also', 'portrays', 'kelly', 'kapoor', 'on', 'the', 'series', '.', 'the', 'episode', 'also', 'marks', 'the', 'second', 'appearance', 'of', 'fe', '##rrell', 'as', 'dean', '##gel', '##o', 'vickers', ';', 'fe', '##rrell', 'had', 'signed', 'onto', 'the', 'series', 'to', 'make', 'care', '##ll', \"'\", 's', 'exit', 'transition', 'easier', '.', 'the', 'episode', 'received', 'mostly', 'positive', 'reviews', 'from', 'television', 'critics', '.', '`', '`', 'michael', \"'\", 's', 'last', 'dun', '##dies', \"'\", \"'\", 'was', 'viewed', 'by', '6', '.', '84', '##9', 'million', 'viewers', 'and', 'received', 'a', '3', '.', '3', 'rating', 'among', 'adults', 'between', 'the', 'age', 'of', '18', 'and', '49', '.', 'the', 'episode', 'was', 'the', 'highest', '-', 'rated', 'nbc', 'series', 'of', 'the', 'week', 'that', 'it', 'aired', ',', 'as', 'well', 'as', 'the', 'sixth', '-', 'most', 'watched', 'episode', 'in', 'the', '18', '-', '-', '49', 'demographic', 'for', 'the', 'week', 'it', 'aired', '.', '[ContextId=22]', '[Paragraph=4]', 'at', 'the', 'office', ',', 'michael', 'scott', '(', 'steve', 'care', '##ll', ')', 'announces', 'to', 'the', 'employees', 'that', 'dean', '##gel', '##o', 'vickers', '(', 'will', 'fe', '##rrell', ')', 'will', 'be', 'his', 'co', '-', 'host', 'at', 'the', 'dun', '##dies', '.', 'the', 'dun', '##dies', 'are', 'an', 'annual', 'award', 'program', 'created', 'by', 'michael', 'to', 'mo', '##tiv', '##ate', 'his', 'employees', '.', 'the', 'idea', 'of', 'performance', 'is', 'wo', '##rri', '##some', 'to', 'dean', '##gel', '##o', ',', 'but', 'michael', 'insists', 'he', 'take', 'the', 'job', '.', 'michael', 'brings', 'some', 'of', 'the', 'staff', 'together', 'in', 'the', 'conference', 'room', 'to', 'help', 'dean', '##gel', '##o', 'get', 'prepared', 'for', 'the', 'show', ',', 'but', 'he', 'struggles', 'to', 'be', 'humorous', '.', 'andy', 'bernard', '(', 'ed', 'helm', '##s', ')', 'tries', 'to', 'help', 'him', ',', 'saying', 'he', 'should', 'just', 'think', 'of', 'performing', 'like', 'conducting', 'a', 'meeting', ',', 'but', 'michael', 'objects', ',', 'wanting', 'dean', '##gel', '##o', 'to', 'mimic', 'his', 'style', '[SEP]']\n```\n## Some observations\n- It looks like they have [Q] token at the beginning of a question.\n- There’s ContextId=N, which we should probably figure out in the paper.\n- I thought start/end indicies of an answer is at the end of input_ids. but there are not.\n   - They might be in segment_ids in training data. I’m not entirely sure yet.",
    "682249": "Thanks for this amazing explanation of what's being inputted into the BERT model! This needs more attention!",
    "686695": "Thank you. I'm glad I could help.",
    "690321": "good thread .. thanks.",
    "692204": "Really useful, thanks. @higepon"
  },
  "source": "meta"
}