{
  "id": 125079,
  "title": "What is token_map used for?",
  "url": "/competitions/tensorflow2-question-answering/discussion/125079",
  "author_name": "",
  "post_date": "2020-01-08T12:48:29.212090300Z",
  "votes": 2,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I suppose this is some kind of filter, but what is it specifically used for? And why do numbers go in sequential order?</p>",
  "messages": [
    {
      "id": "713567",
      "postDate": "01/08/2020 12:48:29",
      "content": "<p>I suppose this is some kind of filter, but what is it specifically used for? And why do numbers go in sequential order?</p>",
      "rawMarkdown": "I suppose this is some kind of filter, but what is it specifically used for? And why do numbers go in sequential order?",
      "votes": null
    },
    {
      "id": "713689",
      "postDate": "01/08/2020 14:53:59",
      "content": "<p>The token map is the mapping from <code>index of tokens of model input</code> to <code>index of tokens of original document text</code>.\nWhen we use BERT or other similar models, we predict answer spans (start and end index) from model inputs, although we have to submit answer spans of original text index. </p>\n\n<p>Example:</p>\n\n<p>&gt; Question: <code>When did Dickinsonia live ?</code>\n&gt; Document text: <code>Dickinsonia is an extinct genus of fossils of the Ediacaran biota .</code> .\n&gt; Answer: <code>Ediacaran biota</code></p>\n\n<p>In the original document text, you can see the answer span is <code>9:11</code>.\nWhen we use BERT, we concatenate question and document to make input like below.\n<code>[CLS] When is Dickinsonia live ? [SEP] Dickinsonia is an extinct genus of fossils of the Ediacaran biota . [SEP]</code>\nThen we predict start and end index from transformed inputs. Is this case, the predicted answer span in the inputs is <code>16:18</code> (if the prediction is correct).\nThe result token map is <code>token_map = [-1, -1, -1, -1, -1, -1, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, -1]</code>. We can trace original index using this as <code>token_map[16] -&gt; 9, token_map[18] -&gt; 11</code>.</p>\n\n<p>In addition to that, we use wordpieces to tokenize texts in default BERT settings, so the inputs of model is more complex (e.g., <code>Dickinsonia</code> turns to 2 pieces <code>'dickinson', '##ia'</code>).</p>\n\n<p>In order to get answer span of original document text, we need to prepare token map when preprocessing.</p>",
      "rawMarkdown": "The token map is the mapping from `index of tokens of model input` to `index of tokens of original document text`.\nWhen we use BERT or other similar models, we predict answer spans (start and end index) from model inputs, although we have to submit answer spans of original text index. \n\nExample:\n\n&gt; Question: `When did Dickinsonia live ?`\n&gt; Document text: `Dickinsonia is an extinct genus of fossils of the Ediacaran biota .` .\n&gt; Answer: `Ediacaran biota`\n\nIn the original document text, you can see the answer span is `9:11`.\nWhen we use BERT, we concatenate question and document to make input like below.\n`[CLS] When is Dickinsonia live ? [SEP] Dickinsonia is an extinct genus of fossils of the Ediacaran biota . [SEP]`\nThen we predict start and end index from transformed inputs. Is this case, the predicted answer span in the inputs is `16:18` (if the prediction is correct).\nThe result token map is `token_map = [-1, -1, -1, -1, -1, -1, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, -1]`. We can trace original index using this as `token_map[16] -&gt; 9, token_map[18] -&gt; 11`.\n\nIn addition to that, we use wordpieces to tokenize texts in default BERT settings, so the inputs of model is more complex (e.g., `Dickinsonia` turns to 2 pieces `'dickinson', '##ia'`).\n\nIn order to get answer span of original document text, we need to prepare token map when preprocessing.",
      "votes": null
    },
    {
      "id": "713691",
      "postDate": "01/08/2020 14:54:35",
      "content": "<p>Token maps in general are used to map the answer indices to actual indices.</p>\n\n<p>For example, say you have the sentence \"I Mourad\".</p>\n\n<p>You cleaned it of tokens that are useless to the model and it became \"I Mourad\".</p>\n\n<p>The model answered with \"Mourad\".</p>\n\n<p>In the answer \"Mourad\" has a token index of 1. In the original setnence, a token index of 2.</p>\n\n<p>A token_map maps 1 to 2. Hope this helps.</p>",
      "rawMarkdown": "Token maps in general are used to map the answer indices to actual indices.\n\nFor example, say you have the sentence \"I Mourad\".\n\nYou cleaned it of tokens that are useless to the model and it became \"I Mourad\".\n\nThe model answered with \"Mourad\".\n\nIn the answer \"Mourad\" has a token index of 1. In the original setnence, a token index of 2.\n\nA token_map maps 1 to 2. Hope this helps.",
      "votes": null
    },
    {
      "id": "713903",
      "postDate": "01/08/2020 19:28:06",
      "content": "<p>An input token (a word) cna be mapped to several tokens by the model tokenizer.  As a result the indices no longer match.  For instance, this question:</p>\n\n<pre><code>['which',\n 'is',\n 'the',\n 'most',\n 'common',\n 'use',\n 'of',\n 'opt-in',\n 'e-mail',\n 'marketing']\n</code></pre>\n\n<p>is mapped in these tokens by BERT tokenizer (to be very precise, I show the vocabulary entries corresponding to the numerical tokens produced by the tokenizer):</p>\n\n<pre><code>['which',\n 'is',\n 'the',\n 'most',\n 'common',\n 'use',\n 'of',\n 'opt', \n '-', \n 'in',\n 'e', \n '-', \n 'mail',\n 'marketing']\n</code></pre>\n\n<p>We see that the word 'marketing', with index 9 in the input, is mapped to the token 'marketing' with index 13.</p>",
      "rawMarkdown": "An input token (a word) cna be mapped to several tokens by the model tokenizer.  As a result the indices no longer match.  For instance, this question:\n\n    ['which',\n     'is',\n     'the',\n     'most',\n     'common',\n     'use',\n     'of',\n     'opt-in',\n     'e-mail',\n     'marketing']\n\nis mapped in these tokens by BERT tokenizer (to be very precise, I show the vocabulary entries corresponding to the numerical tokens produced by the tokenizer):\n\n    ['which',\n     'is',\n     'the',\n     'most',\n     'common',\n     'use',\n     'of',\n     'opt', \n     '-', \n     'in',\n     'e', \n     '-', \n     'mail',\n     'marketing']\n\nWe see that the word 'marketing', with index 9 in the input, is mapped to the token 'marketing' with index 13.",
      "votes": null
    },
    {
      "id": "713974",
      "postDate": "01/08/2020 21:40:49",
      "content": "<p>Got it, thanks</p>",
      "rawMarkdown": "Got it, thanks",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 713689,
      "author_name": "kentaronakanishi",
      "author_url": "",
      "post_date": "01/08/2020 14:53:59",
      "content": "<p>The token map is the mapping from <code>index of tokens of model input</code> to <code>index of tokens of original document text</code>.\nWhen we use BERT or other similar models, we predict answer spans (start and end index) from model inputs, although we have to submit answer spans of original text index. </p>\n\n<p>Example:</p>\n\n<p>&gt; Question: <code>When did Dickinsonia live ?</code>\n&gt; Document text: <code>Dickinsonia is an extinct genus of fossils of the Ediacaran biota .</code> .\n&gt; Answer: <code>Ediacaran biota</code></p>\n\n<p>In the original document text, you can see the answer span is <code>9:11</code>.\nWhen we use BERT, we concatenate question and document to make input like below.\n<code>[CLS] When is Dickinsonia live ? [SEP] Dickinsonia is an extinct genus of fossils of the Ediacaran biota . [SEP]</code>\nThen we predict start and end index from transformed inputs. Is this case, the predicted answer span in the inputs is <code>16:18</code> (if the prediction is correct).\nThe result token map is <code>token_map = [-1, -1, -1, -1, -1, -1, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, -1]</code>. We can trace original index using this as <code>token_map[16] -&gt; 9, token_map[18] -&gt; 11</code>.</p>\n\n<p>In addition to that, we use wordpieces to tokenize texts in default BERT settings, so the inputs of model is more complex (e.g., <code>Dickinsonia</code> turns to 2 pieces <code>'dickinson', '##ia'</code>).</p>\n\n<p>In order to get answer span of original document text, we need to prepare token map when preprocessing.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 713691,
      "author_name": "msheriey",
      "author_url": "",
      "post_date": "01/08/2020 14:54:35",
      "content": "<p>Token maps in general are used to map the answer indices to actual indices.</p>\n\n<p>For example, say you have the sentence \"I Mourad\".</p>\n\n<p>You cleaned it of tokens that are useless to the model and it became \"I Mourad\".</p>\n\n<p>The model answered with \"Mourad\".</p>\n\n<p>In the answer \"Mourad\" has a token index of 1. In the original setnence, a token index of 2.</p>\n\n<p>A token_map maps 1 to 2. Hope this helps.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 713903,
      "author_name": "cpmpml",
      "author_url": "",
      "post_date": "01/08/2020 19:28:06",
      "content": "<p>An input token (a word) cna be mapped to several tokens by the model tokenizer.  As a result the indices no longer match.  For instance, this question:</p>\n\n<pre><code>['which',\n 'is',\n 'the',\n 'most',\n 'common',\n 'use',\n 'of',\n 'opt-in',\n 'e-mail',\n 'marketing']\n</code></pre>\n\n<p>is mapped in these tokens by BERT tokenizer (to be very precise, I show the vocabulary entries corresponding to the numerical tokens produced by the tokenizer):</p>\n\n<pre><code>['which',\n 'is',\n 'the',\n 'most',\n 'common',\n 'use',\n 'of',\n 'opt', \n '-', \n 'in',\n 'e', \n '-', \n 'mail',\n 'marketing']\n</code></pre>\n\n<p>We see that the word 'marketing', with index 9 in the input, is mapped to the token 'marketing' with index 13.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 713974,
      "author_name": "mxmka87",
      "author_url": "",
      "post_date": "01/08/2020 21:40:49",
      "content": "<p>Got it, thanks</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "713567": "I suppose this is some kind of filter, but what is it specifically used for? And why do numbers go in sequential order?",
    "713689": "The token map is the mapping from `index of tokens of model input` to `index of tokens of original document text`.\nWhen we use BERT or other similar models, we predict answer spans (start and end index) from model inputs, although we have to submit answer spans of original text index. \n\nExample:\n\n&gt; Question: `When did Dickinsonia live ?`\n&gt; Document text: `Dickinsonia is an extinct genus of fossils of the Ediacaran biota .` .\n&gt; Answer: `Ediacaran biota`\n\nIn the original document text, you can see the answer span is `9:11`.\nWhen we use BERT, we concatenate question and document to make input like below.\n`[CLS] When is Dickinsonia live ? [SEP] Dickinsonia is an extinct genus of fossils of the Ediacaran biota . [SEP]`\nThen we predict start and end index from transformed inputs. Is this case, the predicted answer span in the inputs is `16:18` (if the prediction is correct).\nThe result token map is `token_map = [-1, -1, -1, -1, -1, -1, -1, 0, 1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, -1]`. We can trace original index using this as `token_map[16] -&gt; 9, token_map[18] -&gt; 11`.\n\nIn addition to that, we use wordpieces to tokenize texts in default BERT settings, so the inputs of model is more complex (e.g., `Dickinsonia` turns to 2 pieces `'dickinson', '##ia'`).\n\nIn order to get answer span of original document text, we need to prepare token map when preprocessing.",
    "713691": "Token maps in general are used to map the answer indices to actual indices.\n\nFor example, say you have the sentence \"I Mourad\".\n\nYou cleaned it of tokens that are useless to the model and it became \"I Mourad\".\n\nThe model answered with \"Mourad\".\n\nIn the answer \"Mourad\" has a token index of 1. In the original setnence, a token index of 2.\n\nA token_map maps 1 to 2. Hope this helps.",
    "713903": "An input token (a word) cna be mapped to several tokens by the model tokenizer.  As a result the indices no longer match.  For instance, this question:\n\n    ['which',\n     'is',\n     'the',\n     'most',\n     'common',\n     'use',\n     'of',\n     'opt-in',\n     'e-mail',\n     'marketing']\n\nis mapped in these tokens by BERT tokenizer (to be very precise, I show the vocabulary entries corresponding to the numerical tokens produced by the tokenizer):\n\n    ['which',\n     'is',\n     'the',\n     'most',\n     'common',\n     'use',\n     'of',\n     'opt', \n     '-', \n     'in',\n     'e', \n     '-', \n     'mail',\n     'marketing']\n\nWe see that the word 'marketing', with index 9 in the input, is mapped to the token 'marketing' with index 13.",
    "713974": "Got it, thanks"
  },
  "source": "meta"
}