{
  "id": 124710,
  "title": "Natural Questions Data Preparation",
  "url": "/competitions/tensorflow2-question-answering/discussion/124710",
  "author_name": "",
  "post_date": "2020-01-06T03:28:31.980961300Z",
  "votes": 4,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Please refer to <a href=\"https://github.com/zzj0402/natural-questions-environment\">https://github.com/zzj0402/natural-questions-environment</a></p>",
  "messages": [
    {
      "id": "711404",
      "postDate": "01/06/2020 03:28:31",
      "content": "<p>Please refer to <a href=\"https://github.com/zzj0402/natural-questions-environment\">https://github.com/zzj0402/natural-questions-environment</a></p>",
      "rawMarkdown": "Please refer to https://github.com/zzj0402/natural-questions-environment",
      "votes": null
    },
    {
      "id": "711418",
      "postDate": "01/06/2020 04:15:37",
      "content": "<p>I am also struggling to convert the simplified jsonl to tfrecords using the <code>prepare_nq_data.py</code> script provided by the <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">bert_joint</a> repository</p>\n\n<p>After debugging the error mentioned <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/121972\">here </a>, I'm now running into a <code>KeyError: document_tokens</code> in the <code>run_nq.create_example_from_jsonl()</code> function call. </p>\n\n<p>Would be very helpful if someone can point to a solution or share a script that works for the competition dataset</p>\n\n<p>UPDATE: A hack that <em>seems</em> to work is to use  <a href=\"/philculliton\">@philculliton</a>'s [script] (<a href=\"https://www.kaggle.com/philculliton/tf2-0-baseline-w-bert\">https://www.kaggle.com/philculliton/tf2-0-baseline-w-bert</a>) implementations of the <code>create_example_from_jsonl()</code> and <code>CreateTFExampleFn()</code> instead of the <code>run_nq.py</code> ones</p>",
      "rawMarkdown": "I am also struggling to convert the simplified jsonl to tfrecords using the `prepare_nq_data.py` script provided by the [bert_joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint) repository\n\nAfter debugging the error mentioned [here ](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/121972), I'm now running into a `KeyError: document_tokens` in the `run_nq.create_example_from_jsonl()` function call. \n\nWould be very helpful if someone can point to a solution or share a script that works for the competition dataset\n\nUPDATE: A hack that *seems* to work is to use  @philculliton's [script] (https://www.kaggle.com/philculliton/tf2-0-baseline-w-bert) implementations of the `create_example_from_jsonl()` and `CreateTFExampleFn()` instead of the `run_nq.py` ones",
      "votes": null
    },
    {
      "id": "711535",
      "postDate": "01/06/2020 07:54:31",
      "content": "<p>Nice. I just figured out what was the problem on this env as well. Turns out my data path missed the /train/.</p>",
      "rawMarkdown": "Nice. I just figured out what was the problem on this env as well. Turns out my data path missed the /train/.",
      "votes": null
    },
    {
      "id": "713911",
      "postDate": "01/08/2020 19:47:28",
      "content": "<p>Does this approach make any attempt to remove HTML tags from the input text?</p>",
      "rawMarkdown": "Does this approach make any attempt to remove HTML tags from the input text?",
      "votes": null
    },
    {
      "id": "714340",
      "postDate": "01/09/2020 10:20:25",
      "content": "<p>Nope.</p>",
      "rawMarkdown": "Nope.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 711418,
      "author_name": "rahulsd91",
      "author_url": "",
      "post_date": "01/06/2020 04:15:37",
      "content": "<p>I am also struggling to convert the simplified jsonl to tfrecords using the <code>prepare_nq_data.py</code> script provided by the <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">bert_joint</a> repository</p>\n\n<p>After debugging the error mentioned <a href=\"https://www.kaggle.com/c/tensorflow2-question-answering/discussion/121972\">here </a>, I'm now running into a <code>KeyError: document_tokens</code> in the <code>run_nq.create_example_from_jsonl()</code> function call. </p>\n\n<p>Would be very helpful if someone can point to a solution or share a script that works for the competition dataset</p>\n\n<p>UPDATE: A hack that <em>seems</em> to work is to use  <a href=\"/philculliton\">@philculliton</a>'s [script] (<a href=\"https://www.kaggle.com/philculliton/tf2-0-baseline-w-bert\">https://www.kaggle.com/philculliton/tf2-0-baseline-w-bert</a>) implementations of the <code>create_example_from_jsonl()</code> and <code>CreateTFExampleFn()</code> instead of the <code>run_nq.py</code> ones</p>",
      "votes": null,
      "replies": [
        {
          "id": 711535,
          "author_name": "zzj040203",
          "author_url": "",
          "post_date": "01/06/2020 07:54:31",
          "content": "<p>Nice. I just figured out what was the problem on this env as well. Turns out my data path missed the /train/.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 713911,
      "author_name": "karlschelhammer",
      "author_url": "",
      "post_date": "01/08/2020 19:47:28",
      "content": "<p>Does this approach make any attempt to remove HTML tags from the input text?</p>",
      "votes": null,
      "replies": [
        {
          "id": 714340,
          "author_name": "zzj040203",
          "author_url": "",
          "post_date": "01/09/2020 10:20:25",
          "content": "<p>Nope.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "711404": "Please refer to https://github.com/zzj0402/natural-questions-environment",
    "711418": "I am also struggling to convert the simplified jsonl to tfrecords using the `prepare_nq_data.py` script provided by the [bert_joint](https://github.com/google-research/language/tree/master/language/question_answering/bert_joint) repository\n\nAfter debugging the error mentioned [here ](https://www.kaggle.com/c/tensorflow2-question-answering/discussion/121972), I'm now running into a `KeyError: document_tokens` in the `run_nq.create_example_from_jsonl()` function call. \n\nWould be very helpful if someone can point to a solution or share a script that works for the competition dataset\n\nUPDATE: A hack that *seems* to work is to use  @philculliton's [script] (https://www.kaggle.com/philculliton/tf2-0-baseline-w-bert) implementations of the `create_example_from_jsonl()` and `CreateTFExampleFn()` instead of the `run_nq.py` ones",
    "711535": "Nice. I just figured out what was the problem on this env as well. Turns out my data path missed the /train/.",
    "713911": "Does this approach make any attempt to remove HTML tags from the input text?",
    "714340": "Nope."
  },
  "source": "meta"
}