{
  "id": 121998,
  "title": "nq-test.tfrecords",
  "url": "/competitions/tensorflow2-question-answering/discussion/121998",
  "author_name": "",
  "post_date": "2019-12-17T01:54:42.909269Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>A lot of succcessful kernels have used <code>nq-test.tfrecords</code> as one of the input files. Where do they source it from and what does it contain?</p>\n\n<p><a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline#nq-test.tfrecords\">https://www.kaggle.com/prokaj/bert-joint-baseline#nq-test.tfrecords</a>\n<a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/data\">https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/data</a>\n<a href=\"https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers/data\">https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers/data</a>\n<a href=\"https://www.kaggle.com/andrewgao/tensorflow-2-0-bert-yes-no-answers/data\">https://www.kaggle.com/andrewgao/tensorflow-2-0-bert-yes-no-answers/data</a>\n<a href=\"https://www.kaggle.com/prokaj/bert-preprocess-notebook/data\">https://www.kaggle.com/prokaj/bert-preprocess-notebook/data</a>\n<a href=\"https://www.kaggle.com/ymcdull/tensorflow-2-0-edited/data\">https://www.kaggle.com/ymcdull/tensorflow-2-0-edited/data</a>\n<a href=\"https://www.kaggle.com/jonathandickson/my-submission-code/data\">https://www.kaggle.com/jonathandickson/my-submission-code/data</a>\n<a href=\"https://www.kaggle.com/kerneler/starter-bert-joint-baseline-9e284fc6-f/data\">https://www.kaggle.com/kerneler/starter-bert-joint-baseline-9e284fc6-f/data</a></p>\n\n<p><a href=\"/prokaj\">@prokaj</a>  <a href=\"/mmmarchetti\">@mmmarchetti</a>  <a href=\"/andrewgao\">@andrewgao</a> <a href=\"/ymcdull\">@ymcdull</a> <a href=\"/jonathandickson\">@jonathandickson</a> <a href=\"/kerneler\">@kerneler</a> </p>\n\n<p>Others - ??</p>",
  "messages": [
    {
      "id": "696739",
      "postDate": "12/17/2019 01:54:42",
      "content": "<p>A lot of succcessful kernels have used <code>nq-test.tfrecords</code> as one of the input files. Where do they source it from and what does it contain?</p>\n\n<p><a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline#nq-test.tfrecords\">https://www.kaggle.com/prokaj/bert-joint-baseline#nq-test.tfrecords</a>\n<a href=\"https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/data\">https://www.kaggle.com/prokaj/bert-joint-baseline-notebook/data</a>\n<a href=\"https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers/data\">https://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers/data</a>\n<a href=\"https://www.kaggle.com/andrewgao/tensorflow-2-0-bert-yes-no-answers/data\">https://www.kaggle.com/andrewgao/tensorflow-2-0-bert-yes-no-answers/data</a>\n<a href=\"https://www.kaggle.com/prokaj/bert-preprocess-notebook/data\">https://www.kaggle.com/prokaj/bert-preprocess-notebook/data</a>\n<a href=\"https://www.kaggle.com/ymcdull/tensorflow-2-0-edited/data\">https://www.kaggle.com/ymcdull/tensorflow-2-0-edited/data</a>\n<a href=\"https://www.kaggle.com/jonathandickson/my-submission-code/data\">https://www.kaggle.com/jonathandickson/my-submission-code/data</a>\n<a href=\"https://www.kaggle.com/kerneler/starter-bert-joint-baseline-9e284fc6-f/data\">https://www.kaggle.com/kerneler/starter-bert-joint-baseline-9e284fc6-f/data</a></p>\n\n<p><a href=\"/prokaj\">@prokaj</a>  <a href=\"/mmmarchetti\">@mmmarchetti</a>  <a href=\"/andrewgao\">@andrewgao</a> <a href=\"/ymcdull\">@ymcdull</a> <a href=\"/jonathandickson\">@jonathandickson</a> <a href=\"/kerneler\">@kerneler</a> </p>\n\n<p>Others - ??</p>",
      "rawMarkdown": "A lot of succcessful kernels have used `nq-test.tfrecords` as one of the input files. Where do they source it from and what does it contain?\n\nhttps://www.kaggle.com/prokaj/bert-joint-baseline#nq-test.tfrecords\nhttps://www.kaggle.com/prokaj/bert-joint-baseline-notebook/data\nhttps://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers/data\nhttps://www.kaggle.com/andrewgao/tensorflow-2-0-bert-yes-no-answers/data\nhttps://www.kaggle.com/prokaj/bert-preprocess-notebook/data\nhttps://www.kaggle.com/ymcdull/tensorflow-2-0-edited/data\nhttps://www.kaggle.com/jonathandickson/my-submission-code/data\nhttps://www.kaggle.com/kerneler/starter-bert-joint-baseline-9e284fc6-f/data\n\n@prokaj  @mmmarchetti  @andrewgao @ymcdull @jonathandickson @kerneler \n\nOthers - ??",
      "votes": null
    },
    {
      "id": "696800",
      "postDate": "12/17/2019 04:42:28",
      "content": "<p>It's from the Google Natural Questions Bert Joint baseline  available at <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a> (those file are from the Google cloud bucket linked in the README).\nit contains pre-computed inputs for the test set, the result of tokenizing and chunking the test documents. As the test set for this competition differs from the NQ dataset I think it would either not be used or would have been re-generated for the Kaggle test set using the provided scripts. Though I think in order to run on the private test set most kernels would not be using it.</p>",
      "rawMarkdown": "It's from the Google Natural Questions Bert Joint baseline  available at https://github.com/google-research/language/tree/master/language/question_answering/bert_joint (those file are from the Google cloud bucket linked in the README).\nit contains pre-computed inputs for the test set, the result of tokenizing and chunking the test documents. As the test set for this competition differs from the NQ dataset I think it would either not be used or would have been re-generated for the Kaggle test set using the provided scripts. Though I think in order to run on the private test set most kernels would not be using it.",
      "votes": null
    },
    {
      "id": "697041",
      "postDate": "12/17/2019 12:09:14",
      "content": "<p>Thanks a lot for your response <a href=\"/thomasbrandon\">@thomasbrandon</a> . </p>\n\n<p>However, I couldn't find any Google cloud bucket linked in the README.\nAlso, why do you think it doesn't make sense to regenerate this for the private test set.</p>\n\n<p>Apologies if my questions are very basic since I'm new to this.</p>",
      "rawMarkdown": "Thanks a lot for your response @thomasbrandon . \n\nHowever, I couldn't find any Google cloud bucket linked in the README.\nAlso, why do you think it doesn't make sense to regenerate this for the private test set.\n\nApologies if my questions are very basic since I'm new to this.",
      "votes": null
    },
    {
      "id": "697312",
      "postDate": "12/17/2019 17:44:10",
      "content": "<p>The cloud bucket is downloaded with:\n<code>gsutil cp -R gs://bert-nq/bert-joint-baseline .</code></p>\n\n<p>Though actually it only contains the training set (<code>nq-train.tfrecords</code>) and model weights, The test records would be the result of running the <code>language.question_answering.bert_joint.prepare_nq_data</code> script on the Kaggle test JSON.</p>\n\n<blockquote>\n  <p>why do you think it doesn't make sense to regenerate this for the private test set.</p>\n</blockquote>\n\n<p>Sorry, I wasn't clear there. I meant that you could either create the nq-test inputs for the public test set and store them in a dataset to access from your kernel. Or, as your kernel has to run on the private test set which you can't store in a dataset (as you don't have access), you could just do your prediction straight off the JSON for both the public and private test set. </p>",
      "rawMarkdown": "The cloud bucket is downloaded with:\n```gsutil cp -R gs://bert-nq/bert-joint-baseline .```\n\nThough actually it only contains the training set (`nq-train.tfrecords`) and model weights, The test records would be the result of running the `language.question_answering.bert_joint.prepare_nq_data` script on the Kaggle test JSON.\n\n&gt; why do you think it doesn't make sense to regenerate this for the private test set.\n\nSorry, I wasn't clear there. I meant that you could either create the nq-test inputs for the public test set and store them in a dataset to access from your kernel. Or, as your kernel has to run on the private test set which you can't store in a dataset (as you don't have access), you could just do your prediction straight off the JSON for both the public and private test set.",
      "votes": null
    },
    {
      "id": "704412",
      "postDate": "12/27/2019 12:56:46",
      "content": "<p>below script in the kernel generate the file</p>\n\n<p>eval_records='nq-test.tfrecords'\nif not os.path.exists(eval_records):\n    # tf2baseline.FLAGS.max_seq_length = 512\n    eval_writer = bert_utils.FeatureWriter(\n        filename=os.path.join(eval_records),\n        is_training=False)</p>\n\n<pre><code>tokenizer = tokenization.FullTokenizer(vocab_file=path+'bert-joint-baseline/vocab-nq.txt', \n                                       do_lower_case=True)\n\nfeatures = []\nconvert = bert_utils.ConvertExamples2Features(tokenizer=tokenizer,\n                                               is_training=False,\n                                               output_fn=eval_writer.process_feature,\n                                               collect_stat=False)\n\nn_examples = 0\ntqdm_notebook= tqdm.tqdm_notebook if not on_kaggle_server else None\nfor examples in bert_utils.nq_examples_iter(input_file=nq_test_file, \n                                       is_training=False,\n                                       tqdm=tqdm_notebook):\n    for example in examples:\n        n_examples += convert(example)\n\neval_writer.close()\nprint('number of test examples: %d, written to file: %d' % (n_examples,eval_writer.num_features))\n</code></pre>",
      "rawMarkdown": "below script in the kernel generate the file\n\n\neval_records='nq-test.tfrecords'\nif not os.path.exists(eval_records):\n    # tf2baseline.FLAGS.max_seq_length = 512\n    eval_writer = bert_utils.FeatureWriter(\n        filename=os.path.join(eval_records),\n        is_training=False)\n\n    tokenizer = tokenization.FullTokenizer(vocab_file=path+'bert-joint-baseline/vocab-nq.txt', \n                                           do_lower_case=True)\n\n    features = []\n    convert = bert_utils.ConvertExamples2Features(tokenizer=tokenizer,\n                                                   is_training=False,\n                                                   output_fn=eval_writer.process_feature,\n                                                   collect_stat=False)\n\n    n_examples = 0\n    tqdm_notebook= tqdm.tqdm_notebook if not on_kaggle_server else None\n    for examples in bert_utils.nq_examples_iter(input_file=nq_test_file, \n                                           is_training=False,\n                                           tqdm=tqdm_notebook):\n        for example in examples:\n            n_examples += convert(example)\n\n    eval_writer.close()\n    print('number of test examples: %d, written to file: %d' % (n_examples,eval_writer.num_features))",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 696800,
      "author_name": "thomasbrandon",
      "author_url": "",
      "post_date": "12/17/2019 04:42:28",
      "content": "<p>It's from the Google Natural Questions Bert Joint baseline  available at <a href=\"https://github.com/google-research/language/tree/master/language/question_answering/bert_joint\">https://github.com/google-research/language/tree/master/language/question_answering/bert_joint</a> (those file are from the Google cloud bucket linked in the README).\nit contains pre-computed inputs for the test set, the result of tokenizing and chunking the test documents. As the test set for this competition differs from the NQ dataset I think it would either not be used or would have been re-generated for the Kaggle test set using the provided scripts. Though I think in order to run on the private test set most kernels would not be using it.</p>",
      "votes": null,
      "replies": [
        {
          "id": 697041,
          "author_name": "rohitagarwal",
          "author_url": "",
          "post_date": "12/17/2019 12:09:14",
          "content": "<p>Thanks a lot for your response <a href=\"/thomasbrandon\">@thomasbrandon</a> . </p>\n\n<p>However, I couldn't find any Google cloud bucket linked in the README.\nAlso, why do you think it doesn't make sense to regenerate this for the private test set.</p>\n\n<p>Apologies if my questions are very basic since I'm new to this.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 697312,
          "author_name": "thomasbrandon",
          "author_url": "",
          "post_date": "12/17/2019 17:44:10",
          "content": "<p>The cloud bucket is downloaded with:\n<code>gsutil cp -R gs://bert-nq/bert-joint-baseline .</code></p>\n\n<p>Though actually it only contains the training set (<code>nq-train.tfrecords</code>) and model weights, The test records would be the result of running the <code>language.question_answering.bert_joint.prepare_nq_data</code> script on the Kaggle test JSON.</p>\n\n<blockquote>\n  <p>why do you think it doesn't make sense to regenerate this for the private test set.</p>\n</blockquote>\n\n<p>Sorry, I wasn't clear there. I meant that you could either create the nq-test inputs for the public test set and store them in a dataset to access from your kernel. Or, as your kernel has to run on the private test set which you can't store in a dataset (as you don't have access), you could just do your prediction straight off the JSON for both the public and private test set. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 704412,
      "author_name": "wentixiaogege",
      "author_url": "",
      "post_date": "12/27/2019 12:56:46",
      "content": "<p>below script in the kernel generate the file</p>\n\n<p>eval_records='nq-test.tfrecords'\nif not os.path.exists(eval_records):\n    # tf2baseline.FLAGS.max_seq_length = 512\n    eval_writer = bert_utils.FeatureWriter(\n        filename=os.path.join(eval_records),\n        is_training=False)</p>\n\n<pre><code>tokenizer = tokenization.FullTokenizer(vocab_file=path+'bert-joint-baseline/vocab-nq.txt', \n                                       do_lower_case=True)\n\nfeatures = []\nconvert = bert_utils.ConvertExamples2Features(tokenizer=tokenizer,\n                                               is_training=False,\n                                               output_fn=eval_writer.process_feature,\n                                               collect_stat=False)\n\nn_examples = 0\ntqdm_notebook= tqdm.tqdm_notebook if not on_kaggle_server else None\nfor examples in bert_utils.nq_examples_iter(input_file=nq_test_file, \n                                       is_training=False,\n                                       tqdm=tqdm_notebook):\n    for example in examples:\n        n_examples += convert(example)\n\neval_writer.close()\nprint('number of test examples: %d, written to file: %d' % (n_examples,eval_writer.num_features))\n</code></pre>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "696739": "A lot of succcessful kernels have used `nq-test.tfrecords` as one of the input files. Where do they source it from and what does it contain?\n\nhttps://www.kaggle.com/prokaj/bert-joint-baseline#nq-test.tfrecords\nhttps://www.kaggle.com/prokaj/bert-joint-baseline-notebook/data\nhttps://www.kaggle.com/mmmarchetti/tensorflow-2-0-bert-yes-no-answers/data\nhttps://www.kaggle.com/andrewgao/tensorflow-2-0-bert-yes-no-answers/data\nhttps://www.kaggle.com/prokaj/bert-preprocess-notebook/data\nhttps://www.kaggle.com/ymcdull/tensorflow-2-0-edited/data\nhttps://www.kaggle.com/jonathandickson/my-submission-code/data\nhttps://www.kaggle.com/kerneler/starter-bert-joint-baseline-9e284fc6-f/data\n\n@prokaj  @mmmarchetti  @andrewgao @ymcdull @jonathandickson @kerneler \n\nOthers - ??",
    "696800": "It's from the Google Natural Questions Bert Joint baseline  available at https://github.com/google-research/language/tree/master/language/question_answering/bert_joint (those file are from the Google cloud bucket linked in the README).\nit contains pre-computed inputs for the test set, the result of tokenizing and chunking the test documents. As the test set for this competition differs from the NQ dataset I think it would either not be used or would have been re-generated for the Kaggle test set using the provided scripts. Though I think in order to run on the private test set most kernels would not be using it.",
    "697041": "Thanks a lot for your response @thomasbrandon . \n\nHowever, I couldn't find any Google cloud bucket linked in the README.\nAlso, why do you think it doesn't make sense to regenerate this for the private test set.\n\nApologies if my questions are very basic since I'm new to this.",
    "697312": "The cloud bucket is downloaded with:\n```gsutil cp -R gs://bert-nq/bert-joint-baseline .```\n\nThough actually it only contains the training set (`nq-train.tfrecords`) and model weights, The test records would be the result of running the `language.question_answering.bert_joint.prepare_nq_data` script on the Kaggle test JSON.\n\n&gt; why do you think it doesn't make sense to regenerate this for the private test set.\n\nSorry, I wasn't clear there. I meant that you could either create the nq-test inputs for the public test set and store them in a dataset to access from your kernel. Or, as your kernel has to run on the private test set which you can't store in a dataset (as you don't have access), you could just do your prediction straight off the JSON for both the public and private test set.",
    "704412": "below script in the kernel generate the file\n\n\neval_records='nq-test.tfrecords'\nif not os.path.exists(eval_records):\n    # tf2baseline.FLAGS.max_seq_length = 512\n    eval_writer = bert_utils.FeatureWriter(\n        filename=os.path.join(eval_records),\n        is_training=False)\n\n    tokenizer = tokenization.FullTokenizer(vocab_file=path+'bert-joint-baseline/vocab-nq.txt', \n                                           do_lower_case=True)\n\n    features = []\n    convert = bert_utils.ConvertExamples2Features(tokenizer=tokenizer,\n                                                   is_training=False,\n                                                   output_fn=eval_writer.process_feature,\n                                                   collect_stat=False)\n\n    n_examples = 0\n    tqdm_notebook= tqdm.tqdm_notebook if not on_kaggle_server else None\n    for examples in bert_utils.nq_examples_iter(input_file=nq_test_file, \n                                           is_training=False,\n                                           tqdm=tqdm_notebook):\n        for example in examples:\n            n_examples += convert(example)\n\n    eval_writer.close()\n    print('number of test examples: %d, written to file: %d' % (n_examples,eval_writer.num_features))"
  },
  "source": "meta"
}