{
  "id": 120461,
  "title": "Visualizing data with Google's nq_browser",
  "url": "/competitions/tensorflow2-question-answering/discussion/120461",
  "author_name": "",
  "post_date": "2019-12-06T08:41:59.073263300Z",
  "votes": 49,
  "comment_count": 25,
  "views": 0,
  "content": "<p>A common mistake in DL tasks is to start and go on only with architectures, training hacks etc but not looking at data at all. In this competition, as of December 6th, the most important thing is to clarify the metric (it'll either be fixed or we'll go on with a non-perfect implementation, thanks to <a href=\"/christofhenkel\">@christofhenkel</a> and <a href=\"/boliu0\">@boliu0</a> for raising this) but then at some point looking at data will also be important, at least for visualizing model errors. </p>\n\n<p>Actually, Google research team already shared a nice tool to visualize NQ data - it's <code>nq_browser.py</code> in <a href=\"https://github.com/google-research-datasets/natural-questions\">natural-questions</a>. I made a <a href=\"https://github.com/Yorko/natural-questions\">fork</a> adapting the script to Python 3.7. </p>\n\n<p>To try it out you can:\n- download the sample 200-long train dataset from <a href=\"https://ai.google.com/research/NaturalQuestions/download\">here</a>\n- install prerequisites listed in <code>nq_browser.py</code> docstring in my fork (the only subtlty there is a specific tornado version)\n- run <code>python nq_browser.py --nq\\_jsonl ../input/v1.0_sample_nq-train-sample.jsonl --nogzipped --max_examples=10</code>\n- go to localhost:8080</p>\n\n<p>You'll see smth like that </p>\n\n<p><img src=\"https://habrastorage.org/webt/q-/hv/v_/q-hvv_se4wbzzz-gwz2m0idcbbg.png\"></p>\n\n<p>You can explore specific examples as well (by clicking Link on the right-hand side)</p>\n\n<p><img src=\"https://habrastorage.org/webt/cs/1o/j7/cs1oj7iwxlzw-hbveanlfv5njm8.png\"></p>\n\n<p>The correct long answer is highlighted in green \n<img src=\"https://habrastorage.org/webt/9b/sj/i5/9bsji5i0d7wagrhkgsdzod6r7zq.png\"></p>\n\n<p>Good luck with the basic step in any ML task - data exploration!</p>\n\n<p>PS. I got a girl born yesterday, will probably step back with this competition till January (anyway the metric needs to be fixed), wanted to share smth useful once Dieter shared his metric implementation :) </p>",
  "messages": [
    {
      "id": "688909",
      "postDate": "12/06/2019 08:41:59",
      "content": "<p>A common mistake in DL tasks is to start and go on only with architectures, training hacks etc but not looking at data at all. In this competition, as of December 6th, the most important thing is to clarify the metric (it'll either be fixed or we'll go on with a non-perfect implementation, thanks to <a href=\"/christofhenkel\">@christofhenkel</a> and <a href=\"/boliu0\">@boliu0</a> for raising this) but then at some point looking at data will also be important, at least for visualizing model errors. </p>\n\n<p>Actually, Google research team already shared a nice tool to visualize NQ data - it's <code>nq_browser.py</code> in <a href=\"https://github.com/google-research-datasets/natural-questions\">natural-questions</a>. I made a <a href=\"https://github.com/Yorko/natural-questions\">fork</a> adapting the script to Python 3.7. </p>\n\n<p>To try it out you can:\n- download the sample 200-long train dataset from <a href=\"https://ai.google.com/research/NaturalQuestions/download\">here</a>\n- install prerequisites listed in <code>nq_browser.py</code> docstring in my fork (the only subtlty there is a specific tornado version)\n- run <code>python nq_browser.py --nq\\_jsonl ../input/v1.0_sample_nq-train-sample.jsonl --nogzipped --max_examples=10</code>\n- go to localhost:8080</p>\n\n<p>You'll see smth like that </p>\n\n<p><img src=\"https://habrastorage.org/webt/q-/hv/v_/q-hvv_se4wbzzz-gwz2m0idcbbg.png\"></p>\n\n<p>You can explore specific examples as well (by clicking Link on the right-hand side)</p>\n\n<p><img src=\"https://habrastorage.org/webt/cs/1o/j7/cs1oj7iwxlzw-hbveanlfv5njm8.png\"></p>\n\n<p>The correct long answer is highlighted in green \n<img src=\"https://habrastorage.org/webt/9b/sj/i5/9bsji5i0d7wagrhkgsdzod6r7zq.png\"></p>\n\n<p>Good luck with the basic step in any ML task - data exploration!</p>\n\n<p>PS. I got a girl born yesterday, will probably step back with this competition till January (anyway the metric needs to be fixed), wanted to share smth useful once Dieter shared his metric implementation :) </p>",
      "rawMarkdown": "A common mistake in DL tasks is to start and go on only with architectures, training hacks etc but not looking at data at all. In this competition, as of December 6th, the most important thing is to clarify the metric (it'll either be fixed or we'll go on with a non-perfect implementation, thanks to @christofhenkel and @boliu0 for raising this) but then at some point looking at data will also be important, at least for visualizing model errors. \n\nActually, Google research team already shared a nice tool to visualize NQ data - it's `nq_browser.py` in [natural-questions](https://github.com/google-research-datasets/natural-questions). I made a [fork](https://github.com/Yorko/natural-questions) adapting the script to Python 3.7. \n\nTo try it out you can:\n- download the sample 200-long train dataset from [here](https://ai.google.com/research/NaturalQuestions/download)\n- install prerequisites listed in `nq_browser.py ` docstring in my fork (the only subtlty there is a specific tornado version)\n- run `python nq_browser.py --nq\\_jsonl ../input/v1.0_sample_nq-train-sample.jsonl --nogzipped --max_examples=10`\n- go to [localhost:8080](localhost:8080)\n\nYou'll see smth like that \n\n<img src=\"https://habrastorage.org/webt/q-/hv/v_/q-hvv_se4wbzzz-gwz2m0idcbbg.png\">\n\nYou can explore specific examples as well (by clicking Link on the right-hand side)\n\n<img src=\"https://habrastorage.org/webt/cs/1o/j7/cs1oj7iwxlzw-hbveanlfv5njm8.png\">\n\nThe correct long answer is highlighted in green \n<img src=\"https://habrastorage.org/webt/9b/sj/i5/9bsji5i0d7wagrhkgsdzod6r7zq.png\">\n\nGood luck with the basic step in any ML task - data exploration!\n\nPS. I got a girl born yesterday, will probably step back with this competition till January (anyway the metric needs to be fixed), wanted to share smth useful once Dieter shared his metric implementation :)",
      "votes": null
    },
    {
      "id": "688983",
      "postDate": "12/06/2019 10:04:06",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Yury thanks for sharing, and really happy for your new-born daughter!! Having daughter is really life-changing experience. Mine is already 3-years old :)</p>",
      "rawMarkdown": "kashnitsky Yury thanks for sharing, and really happy for your new-born daughter!! Having daughter is really life-changing experience. Mine is already 3-years old :)",
      "votes": null
    },
    {
      "id": "689131",
      "postDate": "12/06/2019 14:10:19",
      "content": "<p>Congratulations!</p>",
      "rawMarkdown": "Congratulations!",
      "votes": null
    },
    {
      "id": "689136",
      "postDate": "12/06/2019 14:21:26",
      "content": "<p>This will help for sure! thanks for sharing and congratulations on your daughter!</p>",
      "rawMarkdown": "This will help for sure! thanks for sharing and congratulations on your daughter!",
      "votes": null
    },
    {
      "id": "689164",
      "postDate": "12/06/2019 15:04:45",
      "content": "<p>Congratulations to you and your family Yury!  We will miss your discussions, but family is more important!</p>",
      "rawMarkdown": "Congratulations to you and your family Yury!  We will miss your discussions, but family is more important!",
      "votes": null
    },
    {
      "id": "689176",
      "postDate": "12/06/2019 15:25:58",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Thank you and congratulations! </p>",
      "rawMarkdown": "kashnitsky Thank you and congratulations!",
      "votes": null
    },
    {
      "id": "689297",
      "postDate": "12/06/2019 18:59:50",
      "content": "<p>Big congratulations! Hope you enjoy your life!</p>",
      "rawMarkdown": "Big congratulations! Hope you enjoy your life!",
      "votes": null
    },
    {
      "id": "689302",
      "postDate": "12/06/2019 19:01:59",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Thanks for sharing ..upvote from my side..!!</p>",
      "rawMarkdown": "kashnitsky Thanks for sharing ..upvote from my side..!!",
      "votes": null
    },
    {
      "id": "690052",
      "postDate": "12/07/2019 21:48:17",
      "content": "<p>Great work/life balance 👍</p>",
      "rawMarkdown": "Great work/life balance 👍",
      "votes": null
    },
    {
      "id": "690136",
      "postDate": "12/08/2019 02:57:08",
      "content": "<p>Conglatulations &amp; Thanks for sharing!💕 ✨ 🎉 😄 👎 </p>",
      "rawMarkdown": "Conglatulations &amp; Thanks for sharing!💕 ✨ 🎉 😄 👎",
      "votes": null
    },
    {
      "id": "690679",
      "postDate": "12/08/2019 23:56:51",
      "content": "<p>happy for your new born daughter</p>",
      "rawMarkdown": "happy for your new born daughter",
      "votes": null
    },
    {
      "id": "691733",
      "postDate": "12/10/2019 12:49:35",
      "content": "<p>wonderful job👍 </p>",
      "rawMarkdown": "wonderful job👍",
      "votes": null
    },
    {
      "id": "691911",
      "postDate": "12/10/2019 16:28:07",
      "content": "<p>Thanks!</p>",
      "rawMarkdown": "Thanks!",
      "votes": null
    },
    {
      "id": "692640",
      "postDate": "12/11/2019 15:06:28",
      "content": "<p>Thanks for sharing this and congratulations for your new born girl :)</p>",
      "rawMarkdown": "Thanks for sharing this and congratulations for your new born girl :)",
      "votes": null
    },
    {
      "id": "693381",
      "postDate": "12/12/2019 10:17:47",
      "content": "<p>Congratulation for the new born daughter!</p>",
      "rawMarkdown": "Congratulation for the new born daughter!",
      "votes": null
    },
    {
      "id": "694939",
      "postDate": "12/14/2019 11:38:43",
      "content": "<p>Thanks for all congrats, both already done and those upcoming :) Even though a bit sleep-deprived, I’d still be kaggling in a background daemon mode. </p>",
      "rawMarkdown": "Thanks for all congrats, both already done and those upcoming :) Even though a bit sleep-deprived, I’d still be kaggling in a background daemon mode.",
      "votes": null
    },
    {
      "id": "695153",
      "postDate": "12/14/2019 17:07:09",
      "content": "<p>Good job mann! 💯 </p>",
      "rawMarkdown": "Good job mann! 💯",
      "votes": null
    },
    {
      "id": "699209",
      "postDate": "12/20/2019 07:02:13",
      "content": "<p>Thank you for the sharing, and hope you're doing fine with the baby. Congratulations :) </p>",
      "rawMarkdown": "Thank you for the sharing, and hope you're doing fine with the baby. Congratulations :)",
      "votes": null
    },
    {
      "id": "705623",
      "postDate": "12/29/2019 07:07:20",
      "content": "<p>This is awesome and congrats on your daughter! I also have a young daughter and so limited time for kaggle, so maybe someone can advise - what's the most efficient way of comparing the predictions.json from the baseline model vs. the dev set visualized with this tool?</p>",
      "rawMarkdown": "This is awesome and congrats on your daughter! I also have a young daughter and so limited time for kaggle, so maybe someone can advise - what's the most efficient way of comparing the predictions.json from the baseline model vs. the dev set visualized with this tool?",
      "votes": null
    },
    {
      "id": "706484",
      "postDate": "12/30/2019 13:03:21",
      "content": "<p>Do you want to visualize predictions as well as correct answers? Then there’s no simpler way rather than customizing this <code>nq\\_browser</code> script. </p>",
      "rawMarkdown": "Do you want to visualize predictions as well as correct answers? Then there’s no simpler way rather than customizing this `nq\\_browser` script.",
      "votes": null
    },
    {
      "id": "706488",
      "postDate": "12/30/2019 13:08:35",
      "content": "<p>Yes, almost done with the customization. I can share if interested, but it's very much a quick &amp; dirty solution...</p>",
      "rawMarkdown": "Yes, almost done with the customization. I can share if interested, but it's very much a quick &amp; dirty solution...",
      "votes": null
    },
    {
      "id": "706493",
      "postDate": "12/30/2019 13:15:07",
      "content": "<p>The whole Kaggle is about quick &amp; dirty solutions :)</p>\n\n<p>Yes, I thinks it’s good to share, nice feature. </p>",
      "rawMarkdown": "The whole Kaggle is about quick &amp; dirty solutions :)\n\nYes, I thinks it’s good to share, nice feature.",
      "votes": null
    },
    {
      "id": "706690",
      "postDate": "12/30/2019 17:36:39",
      "content": "<p>Here's a really dirty version of the browser with predictions that works for me. You can start the script like this: \npython nq_browser_pred.py --nq_jsonl=tiny-dev/nq-dev-sample.jsonl --pred_json=predictions.json --nogzipped --dataset=dev --port=8082\nThe predictions are then visible at: <a href=\"http://localhost:8082/expreds\">http://localhost:8082/expreds</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2Fe57e43f00ba00576a43fb5ce00199edd%2FScreen%20Shot%202019-12-30%20at%2018.28.18.png?generation=1577727283316838&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F5f598fc5c040b52957304d1d752402ba%2FScreen%20Shot%202019-12-30%20at%2018.27.51.png?generation=1577727283380397&amp;alt=media\" alt=\"\"></p>\n\n<p>(I removed the json files from tiny-dev to make the attachment smaller)</p>",
      "rawMarkdown": "Here's a really dirty version of the browser with predictions that works for me. You can start the script like this: \npython nq_browser_pred.py --nq_jsonl=tiny-dev/nq-dev-sample.jsonl --pred_json=predictions.json --nogzipped --dataset=dev --port=8082\nThe predictions are then visible at: http://localhost:8082/expreds\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2Fe57e43f00ba00576a43fb5ce00199edd%2FScreen%20Shot%202019-12-30%20at%2018.28.18.png?generation=1577727283316838&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F5f598fc5c040b52957304d1d752402ba%2FScreen%20Shot%202019-12-30%20at%2018.27.51.png?generation=1577727283380397&amp;alt=media)\n\n(I removed the json files from tiny-dev to make the attachment smaller)",
      "votes": null
    },
    {
      "id": "706742",
      "postDate": "12/30/2019 19:01:45",
      "content": "<p>Thanks, looks nice! </p>",
      "rawMarkdown": "Thanks, looks nice!",
      "votes": null
    },
    {
      "id": "716515",
      "postDate": "01/11/2020 19:45:13",
      "content": "<p>How are you getting the links using the test set since document_url isn't provided for test? I'm struggling to view my predictions in context to the actual test set I used.</p>\n\n<p>Could you also please share your nq-dev-sample.jsonl? The files I'm using are inconsistent with what I'm seeing.</p>",
      "rawMarkdown": "How are you getting the links using the test set since document_url isn't provided for test? I'm struggling to view my predictions in context to the actual test set I used.\n\nCould you also please share your nq-dev-sample.jsonl? The files I'm using are inconsistent with what I'm seeing.",
      "votes": null
    },
    {
      "id": "716654",
      "postDate": "01/12/2020 04:09:35",
      "content": "<p><a href=\"/benyap\">@benyap</a> the files I'm using are from the natural questions website rather than kaggle dataset, so indeed there might be some differences in formatting. I'm only visualizing dev, not test. Here's the link I used to download the nq-dev-sample.jsonl: <a href=\"https://storage.cloud.google.com/natural_questions/v1.0/sample/nq-dev-sample.jsonl.gz\">https://storage.cloud.google.com/natural_questions/v1.0/sample/nq-dev-sample.jsonl.gz</a></p>",
      "rawMarkdown": "benyap the files I'm using are from the natural questions website rather than kaggle dataset, so indeed there might be some differences in formatting. I'm only visualizing dev, not test. Here's the link I used to download the nq-dev-sample.jsonl: https://storage.cloud.google.com/natural_questions/v1.0/sample/nq-dev-sample.jsonl.gz",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 688983,
      "author_name": "ratthachat",
      "author_url": "",
      "post_date": "12/06/2019 10:04:06",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Yury thanks for sharing, and really happy for your new-born daughter!! Having daughter is really life-changing experience. Mine is already 3-years old :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689131,
      "author_name": "zhaomeng1126",
      "author_url": "",
      "post_date": "12/06/2019 14:10:19",
      "content": "<p>Congratulations!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689136,
      "author_name": "dimitreoliveira",
      "author_url": "",
      "post_date": "12/06/2019 14:21:26",
      "content": "<p>This will help for sure! thanks for sharing and congratulations on your daughter!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689164,
      "author_name": "boliu0",
      "author_url": "",
      "post_date": "12/06/2019 15:04:45",
      "content": "<p>Congratulations to you and your family Yury!  We will miss your discussions, but family is more important!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689176,
      "author_name": "dskswu",
      "author_url": "",
      "post_date": "12/06/2019 15:25:58",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Thank you and congratulations! </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689297,
      "author_name": "httpwwwfszyc",
      "author_url": "",
      "post_date": "12/06/2019 18:59:50",
      "content": "<p>Big congratulations! Hope you enjoy your life!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 689302,
      "author_name": "saurav9786",
      "author_url": "",
      "post_date": "12/06/2019 19:01:59",
      "content": "<p><a href=\"/kashnitsky\">@kashnitsky</a> Thanks for sharing ..upvote from my side..!!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 690052,
      "author_name": "purplejester",
      "author_url": "",
      "post_date": "12/07/2019 21:48:17",
      "content": "<p>Great work/life balance 👍</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 690136,
      "author_name": "mashlyn",
      "author_url": "",
      "post_date": "12/08/2019 02:57:08",
      "content": "<p>Conglatulations &amp; Thanks for sharing!💕 ✨ 🎉 😄 👎 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 690679,
      "author_name": "",
      "author_url": "",
      "post_date": "12/08/2019 23:56:51",
      "content": "<p>happy for your new born daughter</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 691733,
      "author_name": "diegojohnson",
      "author_url": "",
      "post_date": "12/10/2019 12:49:35",
      "content": "<p>wonderful job👍 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 691911,
      "author_name": "",
      "author_url": "",
      "post_date": "12/10/2019 16:28:07",
      "content": "<p>Thanks!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 692640,
      "author_name": "pedrojlazevedo",
      "author_url": "",
      "post_date": "12/11/2019 15:06:28",
      "content": "<p>Thanks for sharing this and congratulations for your new born girl :)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 693381,
      "author_name": "rohitagarwal",
      "author_url": "",
      "post_date": "12/12/2019 10:17:47",
      "content": "<p>Congratulation for the new born daughter!</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 694939,
      "author_name": "kashnitsky",
      "author_url": "",
      "post_date": "12/14/2019 11:38:43",
      "content": "<p>Thanks for all congrats, both already done and those upcoming :) Even though a bit sleep-deprived, I’d still be kaggling in a background daemon mode. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 695153,
      "author_name": "dasmehdixtr",
      "author_url": "",
      "post_date": "12/14/2019 17:07:09",
      "content": "<p>Good job mann! 💯 </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 699209,
      "author_name": "nyleve",
      "author_url": "",
      "post_date": "12/20/2019 07:02:13",
      "content": "<p>Thank you for the sharing, and hope you're doing fine with the baby. Congratulations :) </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 705623,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "12/29/2019 07:07:20",
      "content": "<p>This is awesome and congrats on your daughter! I also have a young daughter and so limited time for kaggle, so maybe someone can advise - what's the most efficient way of comparing the predictions.json from the baseline model vs. the dev set visualized with this tool?</p>",
      "votes": null,
      "replies": [
        {
          "id": 706484,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/30/2019 13:03:21",
          "content": "<p>Do you want to visualize predictions as well as correct answers? Then there’s no simpler way rather than customizing this <code>nq\\_browser</code> script. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706488,
          "author_name": "thedrcat",
          "author_url": "",
          "post_date": "12/30/2019 13:08:35",
          "content": "<p>Yes, almost done with the customization. I can share if interested, but it's very much a quick &amp; dirty solution...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 706493,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/30/2019 13:15:07",
          "content": "<p>The whole Kaggle is about quick &amp; dirty solutions :)</p>\n\n<p>Yes, I thinks it’s good to share, nice feature. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 706690,
      "author_name": "thedrcat",
      "author_url": "",
      "post_date": "12/30/2019 17:36:39",
      "content": "<p>Here's a really dirty version of the browser with predictions that works for me. You can start the script like this: \npython nq_browser_pred.py --nq_jsonl=tiny-dev/nq-dev-sample.jsonl --pred_json=predictions.json --nogzipped --dataset=dev --port=8082\nThe predictions are then visible at: <a href=\"http://localhost:8082/expreds\">http://localhost:8082/expreds</a>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2Fe57e43f00ba00576a43fb5ce00199edd%2FScreen%20Shot%202019-12-30%20at%2018.28.18.png?generation=1577727283316838&amp;alt=media\" alt=\"\">\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F5f598fc5c040b52957304d1d752402ba%2FScreen%20Shot%202019-12-30%20at%2018.27.51.png?generation=1577727283380397&amp;alt=media\" alt=\"\"></p>\n\n<p>(I removed the json files from tiny-dev to make the attachment smaller)</p>",
      "votes": null,
      "replies": [
        {
          "id": 706742,
          "author_name": "kashnitsky",
          "author_url": "",
          "post_date": "12/30/2019 19:01:45",
          "content": "<p>Thanks, looks nice! </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716515,
          "author_name": "benyap",
          "author_url": "",
          "post_date": "01/11/2020 19:45:13",
          "content": "<p>How are you getting the links using the test set since document_url isn't provided for test? I'm struggling to view my predictions in context to the actual test set I used.</p>\n\n<p>Could you also please share your nq-dev-sample.jsonl? The files I'm using are inconsistent with what I'm seeing.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 716654,
          "author_name": "thedrcat",
          "author_url": "",
          "post_date": "01/12/2020 04:09:35",
          "content": "<p><a href=\"/benyap\">@benyap</a> the files I'm using are from the natural questions website rather than kaggle dataset, so indeed there might be some differences in formatting. I'm only visualizing dev, not test. Here's the link I used to download the nq-dev-sample.jsonl: <a href=\"https://storage.cloud.google.com/natural_questions/v1.0/sample/nq-dev-sample.jsonl.gz\">https://storage.cloud.google.com/natural_questions/v1.0/sample/nq-dev-sample.jsonl.gz</a></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "688909": "A common mistake in DL tasks is to start and go on only with architectures, training hacks etc but not looking at data at all. In this competition, as of December 6th, the most important thing is to clarify the metric (it'll either be fixed or we'll go on with a non-perfect implementation, thanks to @christofhenkel and @boliu0 for raising this) but then at some point looking at data will also be important, at least for visualizing model errors. \n\nActually, Google research team already shared a nice tool to visualize NQ data - it's `nq_browser.py` in [natural-questions](https://github.com/google-research-datasets/natural-questions). I made a [fork](https://github.com/Yorko/natural-questions) adapting the script to Python 3.7. \n\nTo try it out you can:\n- download the sample 200-long train dataset from [here](https://ai.google.com/research/NaturalQuestions/download)\n- install prerequisites listed in `nq_browser.py ` docstring in my fork (the only subtlty there is a specific tornado version)\n- run `python nq_browser.py --nq\\_jsonl ../input/v1.0_sample_nq-train-sample.jsonl --nogzipped --max_examples=10`\n- go to [localhost:8080](localhost:8080)\n\nYou'll see smth like that \n\n<img src=\"https://habrastorage.org/webt/q-/hv/v_/q-hvv_se4wbzzz-gwz2m0idcbbg.png\">\n\nYou can explore specific examples as well (by clicking Link on the right-hand side)\n\n<img src=\"https://habrastorage.org/webt/cs/1o/j7/cs1oj7iwxlzw-hbveanlfv5njm8.png\">\n\nThe correct long answer is highlighted in green \n<img src=\"https://habrastorage.org/webt/9b/sj/i5/9bsji5i0d7wagrhkgsdzod6r7zq.png\">\n\nGood luck with the basic step in any ML task - data exploration!\n\nPS. I got a girl born yesterday, will probably step back with this competition till January (anyway the metric needs to be fixed), wanted to share smth useful once Dieter shared his metric implementation :)",
    "688983": "kashnitsky Yury thanks for sharing, and really happy for your new-born daughter!! Having daughter is really life-changing experience. Mine is already 3-years old :)",
    "689131": "Congratulations!",
    "689136": "This will help for sure! thanks for sharing and congratulations on your daughter!",
    "689164": "Congratulations to you and your family Yury!  We will miss your discussions, but family is more important!",
    "689176": "kashnitsky Thank you and congratulations!",
    "689297": "Big congratulations! Hope you enjoy your life!",
    "689302": "kashnitsky Thanks for sharing ..upvote from my side..!!",
    "690052": "Great work/life balance 👍",
    "690136": "Conglatulations &amp; Thanks for sharing!💕 ✨ 🎉 😄 👎",
    "690679": "happy for your new born daughter",
    "691733": "wonderful job👍",
    "691911": "Thanks!",
    "692640": "Thanks for sharing this and congratulations for your new born girl :)",
    "693381": "Congratulation for the new born daughter!",
    "694939": "Thanks for all congrats, both already done and those upcoming :) Even though a bit sleep-deprived, I’d still be kaggling in a background daemon mode.",
    "695153": "Good job mann! 💯",
    "699209": "Thank you for the sharing, and hope you're doing fine with the baby. Congratulations :)",
    "705623": "This is awesome and congrats on your daughter! I also have a young daughter and so limited time for kaggle, so maybe someone can advise - what's the most efficient way of comparing the predictions.json from the baseline model vs. the dev set visualized with this tool?",
    "706484": "Do you want to visualize predictions as well as correct answers? Then there’s no simpler way rather than customizing this `nq\\_browser` script.",
    "706488": "Yes, almost done with the customization. I can share if interested, but it's very much a quick &amp; dirty solution...",
    "706493": "The whole Kaggle is about quick &amp; dirty solutions :)\n\nYes, I thinks it’s good to share, nice feature.",
    "706690": "Here's a really dirty version of the browser with predictions that works for me. You can start the script like this: \npython nq_browser_pred.py --nq_jsonl=tiny-dev/nq-dev-sample.jsonl --pred_json=predictions.json --nogzipped --dataset=dev --port=8082\nThe predictions are then visible at: http://localhost:8082/expreds\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2Fe57e43f00ba00576a43fb5ce00199edd%2FScreen%20Shot%202019-12-30%20at%2018.28.18.png?generation=1577727283316838&amp;alt=media)\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F1328185%2F5f598fc5c040b52957304d1d752402ba%2FScreen%20Shot%202019-12-30%20at%2018.27.51.png?generation=1577727283380397&amp;alt=media)\n\n(I removed the json files from tiny-dev to make the attachment smaller)",
    "706742": "Thanks, looks nice!",
    "716515": "How are you getting the links using the test set since document_url isn't provided for test? I'm struggling to view my predictions in context to the actual test set I used.\n\nCould you also please share your nq-dev-sample.jsonl? The files I'm using are inconsistent with what I'm seeing.",
    "716654": "benyap the files I'm using are from the natural questions website rather than kaggle dataset, so indeed there might be some differences in formatting. I'm only visualizing dev, not test. Here's the link I used to download the nq-dev-sample.jsonl: https://storage.cloud.google.com/natural_questions/v1.0/sample/nq-dev-sample.jsonl.gz"
  },
  "source": "meta"
}