{
  "id": 120803,
  "title": "Question: How to make a prediction from N strides",
  "url": "/competitions/tensorflow2-question-answering/discussion/120803",
  "author_name": "",
  "post_date": "2019-12-08T23:52:52.129325500Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi there,\nI wrote <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-690321\">Understand how does training input look like</a> and now I'm trying to add inference code to @sakami 's <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">great PyTorch kernel</a>.</p>\n\n<p><strong>Question</strong>\nHow do you make a single prediction for multiple N strides?</p>\n\n<p><strong>What I understand</strong>\nWhen training we create multiple training instances using stride of 128 tokens from one doc text and question pair and use then for training. There is a discussion about downsampling null instances but let's forget about it for now. (See <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">PDF: A BERT Baseline for the Natural Questions</a>).</p>\n\n<p>When predicting we also create multiple instances using the stride technique. But how do you choose/make one final prediction from N predictions from the instances? I read <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">PDF: A BERT Baseline for the Natural Questions</a>, but IIUC they don't mention how they make prediction.</p>\n\n<p>I also read <a href=\"https://github.com/google-research/language/blob/0f0841c8dadfa7b3f79dd342237f40bdd30351fb/language/question_answering/bert_joint/run_nq.py#L1201\">run_nq.py</a> but I'm not very sure if I understand it right. They seem choosing the first one?</p>\n\n<p><code>\n  if predictions:\n    score, summary, start_span, end_span = sorted(predictions, reverse=True)[0]\n    short_span = Span(start_span, end_span)\n    for c in example.candidates:\n      start = short_span.start_token_idx\n      end = short_span.end_token_idx\n      if c[\"top_level\"] and c[\"start_token\"] &lt;= start and c[\"end_token\"] &gt;= end:\n        long_span = Span(c[\"start_token\"], c[\"end_token\"])\n        break\n</code></p>\n\n<p>I appreciate your help!\nbtw I'll be participating <strong>Kaggle Days Tokyo</strong> this week (Week of 12/9). Looking forward to seeing some you :)</p>",
  "messages": [
    {
      "id": "690677",
      "postDate": "12/08/2019 23:52:52",
      "content": "<p>Hi there,\nI wrote <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-690321\">Understand how does training input look like</a> and now I'm trying to add inference code to @sakami 's <a href=\"https://www.kaggle.com/sakami/tfqa-pytorch-baseline\">great PyTorch kernel</a>.</p>\n\n<p><strong>Question</strong>\nHow do you make a single prediction for multiple N strides?</p>\n\n<p><strong>What I understand</strong>\nWhen training we create multiple training instances using stride of 128 tokens from one doc text and question pair and use then for training. There is a discussion about downsampling null instances but let's forget about it for now. (See <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">PDF: A BERT Baseline for the Natural Questions</a>).</p>\n\n<p>When predicting we also create multiple instances using the stride technique. But how do you choose/make one final prediction from N predictions from the instances? I read <a href=\"https://arxiv.org/pdf/1901.08634.pdf\">PDF: A BERT Baseline for the Natural Questions</a>, but IIUC they don't mention how they make prediction.</p>\n\n<p>I also read <a href=\"https://github.com/google-research/language/blob/0f0841c8dadfa7b3f79dd342237f40bdd30351fb/language/question_answering/bert_joint/run_nq.py#L1201\">run_nq.py</a> but I'm not very sure if I understand it right. They seem choosing the first one?</p>\n\n<p><code>\n  if predictions:\n    score, summary, start_span, end_span = sorted(predictions, reverse=True)[0]\n    short_span = Span(start_span, end_span)\n    for c in example.candidates:\n      start = short_span.start_token_idx\n      end = short_span.end_token_idx\n      if c[\"top_level\"] and c[\"start_token\"] &lt;= start and c[\"end_token\"] &gt;= end:\n        long_span = Span(c[\"start_token\"], c[\"end_token\"])\n        break\n</code></p>\n\n<p>I appreciate your help!\nbtw I'll be participating <strong>Kaggle Days Tokyo</strong> this week (Week of 12/9). Looking forward to seeing some you :)</p>",
      "rawMarkdown": "Hi there,\nI wrote [Understand how does training input look like](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-690321) and now I'm trying to add inference code to @sakami 's [great PyTorch kernel](https://www.kaggle.com/sakami/tfqa-pytorch-baseline).\n\n**Question**\nHow do you make a single prediction for multiple N strides?\n\n**What I understand**\nWhen training we create multiple training instances using stride of 128 tokens from one doc text and question pair and use then for training. There is a discussion about downsampling null instances but let's forget about it for now. (See [PDF: A BERT Baseline for the Natural Questions](https://arxiv.org/pdf/1901.08634.pdf)).\n\nWhen predicting we also create multiple instances using the stride technique. But how do you choose/make one final prediction from N predictions from the instances? I read [PDF: A BERT Baseline for the Natural Questions](https://arxiv.org/pdf/1901.08634.pdf), but IIUC they don't mention how they make prediction.\n\nI also read [run_nq.py](https://github.com/google-research/language/blob/0f0841c8dadfa7b3f79dd342237f40bdd30351fb/language/question_answering/bert_joint/run_nq.py#L1201) but I'm not very sure if I understand it right. They seem choosing the first one?\n\n```\n  if predictions:\n    score, summary, start_span, end_span = sorted(predictions, reverse=True)[0]\n    short_span = Span(start_span, end_span)\n    for c in example.candidates:\n      start = short_span.start_token_idx\n      end = short_span.end_token_idx\n      if c[\"top_level\"] and c[\"start_token\"] &lt;= start and c[\"end_token\"] &gt;= end:\n        long_span = Span(c[\"start_token\"], c[\"end_token\"])\n        break\n```\n\nI appreciate your help!\nbtw I'll be participating __Kaggle Days Tokyo__ this week (Week of 12/9). Looking forward to seeing some you :)",
      "votes": null
    },
    {
      "id": "690798",
      "postDate": "12/09/2019 06:54:09",
      "content": "<blockquote>\n  <p>choosing the first one?</p>\n</blockquote>\n\n<p>It's choosing the highest scoring valid short answer span out of the <code>FLAGS.n_best_size</code> top scoring spans collected for each document chunk. So if there are 3 chunks, it will collect by default the 20 spans for each chunk and then select the best of those 60 spans as the final prediction. The score for a span is the sum of scores for predicted start and end index minus the sum of CLS token scores for start/end index predictions (where CLS is used to indicate a no answer prediction). Note that it's sorting the predictions by score and then by an increasing id so the last prediction made will win a tie on score.</p>",
      "rawMarkdown": "&gt; choosing the first one?\n\nIt's choosing the highest scoring valid short answer span out of the `FLAGS.n_best_size` top scoring spans collected for each document chunk. So if there are 3 chunks, it will collect by default the 20 spans for each chunk and then select the best of those 60 spans as the final prediction. The score for a span is the sum of scores for predicted start and end index minus the sum of CLS token scores for start/end index predictions (where CLS is used to indicate a no answer prediction). Note that it's sorting the predictions by score and then by an increasing id so the last prediction made will win a tie on score.",
      "votes": null
    },
    {
      "id": "690804",
      "postDate": "12/09/2019 07:18:20",
      "content": "<p>Thank you Thomas for your help!\nI'll think through what you explained to me. I need to understand how scoring related CLS works.</p>\n\n<p>I thought for this part. They are discarding predictions except for first and best prediction?\n<code>\nsorted(predictions, reverse=True)[0]\n</code></p>",
      "rawMarkdown": "Thank you Thomas for your help!\nI'll think through what you explained to me. I need to understand how scoring related CLS works.\n\nI thought for this part. They are discarding predictions except for first and best prediction?\n```\nsorted(predictions, reverse=True)[0]\n```",
      "votes": null
    },
    {
      "id": "690833",
      "postDate": "12/09/2019 08:14:58",
      "content": "<p>Yes, they are selecting the single best prediction. Is that not what you'd expect? Not quite clear what you're unsure about to try and explain better.\nOn the CLS, there's a comment that \"Span logits minus the cls logits seems to be close to the best\". So I guess the choice here is the result of experiments.</p>",
      "rawMarkdown": "Yes, they are selecting the single best prediction. Is that not what you'd expect? Not quite clear what you're unsure about to try and explain better.\nOn the CLS, there's a comment that \"Span logits minus the cls logits seems to be close to the best\". So I guess the choice here is the result of experiments.",
      "votes": null
    },
    {
      "id": "690851",
      "postDate": "12/09/2019 08:42:12",
      "content": "<p>Thank you now I got the logic and I found out it's written in the paper as follows.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4b895f9bf9722f9ee90125f37a407d05%2FScreen%20Shot%202019-12-09%20at%205.40.47%20PM.png?generation=1575880893811815&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Thank you now I got the logic and I found out it's written in the paper as follows.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4b895f9bf9722f9ee90125f37a407d05%2FScreen%20Shot%202019-12-09%20at%205.40.47%20PM.png?generation=1575880893811815&amp;alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 690798,
      "author_name": "thomasbrandon",
      "author_url": "",
      "post_date": "12/09/2019 06:54:09",
      "content": "<blockquote>\n  <p>choosing the first one?</p>\n</blockquote>\n\n<p>It's choosing the highest scoring valid short answer span out of the <code>FLAGS.n_best_size</code> top scoring spans collected for each document chunk. So if there are 3 chunks, it will collect by default the 20 spans for each chunk and then select the best of those 60 spans as the final prediction. The score for a span is the sum of scores for predicted start and end index minus the sum of CLS token scores for start/end index predictions (where CLS is used to indicate a no answer prediction). Note that it's sorting the predictions by score and then by an increasing id so the last prediction made will win a tie on score.</p>",
      "votes": null,
      "replies": [
        {
          "id": 690804,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/09/2019 07:18:20",
          "content": "<p>Thank you Thomas for your help!\nI'll think through what you explained to me. I need to understand how scoring related CLS works.</p>\n\n<p>I thought for this part. They are discarding predictions except for first and best prediction?\n<code>\nsorted(predictions, reverse=True)[0]\n</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 690833,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "12/09/2019 08:14:58",
          "content": "<p>Yes, they are selecting the single best prediction. Is that not what you'd expect? Not quite clear what you're unsure about to try and explain better.\nOn the CLS, there's a comment that \"Span logits minus the cls logits seems to be close to the best\". So I guess the choice here is the result of experiments.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 690851,
          "author_name": "higepon",
          "author_url": "",
          "post_date": "12/09/2019 08:42:12",
          "content": "<p>Thank you now I got the logic and I found out it's written in the paper as follows.</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4b895f9bf9722f9ee90125f37a407d05%2FScreen%20Shot%202019-12-09%20at%205.40.47%20PM.png?generation=1575880893811815&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "690677": "Hi there,\nI wrote [Understand how does training input look like](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/117370#latest-690321) and now I'm trying to add inference code to @sakami 's [great PyTorch kernel](https://www.kaggle.com/sakami/tfqa-pytorch-baseline).\n\n**Question**\nHow do you make a single prediction for multiple N strides?\n\n**What I understand**\nWhen training we create multiple training instances using stride of 128 tokens from one doc text and question pair and use then for training. There is a discussion about downsampling null instances but let's forget about it for now. (See [PDF: A BERT Baseline for the Natural Questions](https://arxiv.org/pdf/1901.08634.pdf)).\n\nWhen predicting we also create multiple instances using the stride technique. But how do you choose/make one final prediction from N predictions from the instances? I read [PDF: A BERT Baseline for the Natural Questions](https://arxiv.org/pdf/1901.08634.pdf), but IIUC they don't mention how they make prediction.\n\nI also read [run_nq.py](https://github.com/google-research/language/blob/0f0841c8dadfa7b3f79dd342237f40bdd30351fb/language/question_answering/bert_joint/run_nq.py#L1201) but I'm not very sure if I understand it right. They seem choosing the first one?\n\n```\n  if predictions:\n    score, summary, start_span, end_span = sorted(predictions, reverse=True)[0]\n    short_span = Span(start_span, end_span)\n    for c in example.candidates:\n      start = short_span.start_token_idx\n      end = short_span.end_token_idx\n      if c[\"top_level\"] and c[\"start_token\"] &lt;= start and c[\"end_token\"] &gt;= end:\n        long_span = Span(c[\"start_token\"], c[\"end_token\"])\n        break\n```\n\nI appreciate your help!\nbtw I'll be participating __Kaggle Days Tokyo__ this week (Week of 12/9). Looking forward to seeing some you :)",
    "690798": "&gt; choosing the first one?\n\nIt's choosing the highest scoring valid short answer span out of the `FLAGS.n_best_size` top scoring spans collected for each document chunk. So if there are 3 chunks, it will collect by default the 20 spans for each chunk and then select the best of those 60 spans as the final prediction. The score for a span is the sum of scores for predicted start and end index minus the sum of CLS token scores for start/end index predictions (where CLS is used to indicate a no answer prediction). Note that it's sorting the predictions by score and then by an increasing id so the last prediction made will win a tie on score.",
    "690804": "Thank you Thomas for your help!\nI'll think through what you explained to me. I need to understand how scoring related CLS works.\n\nI thought for this part. They are discarding predictions except for first and best prediction?\n```\nsorted(predictions, reverse=True)[0]\n```",
    "690833": "Yes, they are selecting the single best prediction. Is that not what you'd expect? Not quite clear what you're unsure about to try and explain better.\nOn the CLS, there's a comment that \"Span logits minus the cls logits seems to be close to the best\". So I guess the choice here is the result of experiments.",
    "690851": "Thank you now I got the logic and I found out it's written in the paper as follows.\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2199749%2F4b895f9bf9722f9ee90125f37a407d05%2FScreen%20Shot%202019-12-09%20at%205.40.47%20PM.png?generation=1575880893811815&amp;alt=media)"
  },
  "source": "meta"
}