{
  "id": 123466,
  "title": "Scoring error again",
  "url": "/competitions/tensorflow2-question-answering/discussion/123466",
  "author_name": "",
  "post_date": "2019-12-28T01:49:16.324045500Z",
  "votes": 3,
  "comment_count": 4,
  "views": 0,
  "content": "<p>I am facing the scoring error. I have looked at posts on this topic already and I have tried to fix the problem myself but I still can't figure out what is wrong. I also posted in one of the topics. </p>\n\n<p>Some questions I have:</p>\n\n<pre><code>Where is the private dataset?\nHow is the private dataset loaded into python memory? If my code doesn't load the private dataset, then what code does? Can I get a snippet of the code that loads the data?\nShould we write a function that takes in one row at a time or one dataframe at a time or a path to a file ?\n</code></pre>\n\n<p>Honestly, this is incredibly frustrating! I have spent my 3 out of 5 daily submission limit in 2 hours. I still don't know what I need to do to avoid this problem. The lack of error details means that I am flying blind. Since I can't see the error, I don't know if I'll be able to fix this in the next 2 tries or 2 million tries.</p>\n\n<p>Here is my notebook: <a href=\"https://www.kaggle.com/ankurgupta1985/randomly-selected\">https://www.kaggle.com/ankurgupta1985/randomly-selected</a></p>\n\n<p>I am already chunking the data and not loading all of it in memory at one time. But, I assume that the final submission.csv is small enough that the corresponding dataframe would fit in memory. I have already asked friends who're experienced at Kaggle to help me and they're dumbfounded too. If anyone can help, I promise to consider naming my first born after you. Thank you very much!</p>",
  "messages": [
    {
      "id": "704770",
      "postDate": "12/28/2019 01:49:16",
      "content": "<p>I am facing the scoring error. I have looked at posts on this topic already and I have tried to fix the problem myself but I still can't figure out what is wrong. I also posted in one of the topics. </p>\n\n<p>Some questions I have:</p>\n\n<pre><code>Where is the private dataset?\nHow is the private dataset loaded into python memory? If my code doesn't load the private dataset, then what code does? Can I get a snippet of the code that loads the data?\nShould we write a function that takes in one row at a time or one dataframe at a time or a path to a file ?\n</code></pre>\n\n<p>Honestly, this is incredibly frustrating! I have spent my 3 out of 5 daily submission limit in 2 hours. I still don't know what I need to do to avoid this problem. The lack of error details means that I am flying blind. Since I can't see the error, I don't know if I'll be able to fix this in the next 2 tries or 2 million tries.</p>\n\n<p>Here is my notebook: <a href=\"https://www.kaggle.com/ankurgupta1985/randomly-selected\">https://www.kaggle.com/ankurgupta1985/randomly-selected</a></p>\n\n<p>I am already chunking the data and not loading all of it in memory at one time. But, I assume that the final submission.csv is small enough that the corresponding dataframe would fit in memory. I have already asked friends who're experienced at Kaggle to help me and they're dumbfounded too. If anyone can help, I promise to consider naming my first born after you. Thank you very much!</p>",
      "rawMarkdown": "I am facing the scoring error. I have looked at posts on this topic already and I have tried to fix the problem myself but I still can't figure out what is wrong. I also posted in one of the topics. \n\nSome questions I have:\n\n    Where is the private dataset?\n    How is the private dataset loaded into python memory? If my code doesn't load the private dataset, then what code does? Can I get a snippet of the code that loads the data?\n    Should we write a function that takes in one row at a time or one dataframe at a time or a path to a file ?\n\nHonestly, this is incredibly frustrating! I have spent my 3 out of 5 daily submission limit in 2 hours. I still don't know what I need to do to avoid this problem. The lack of error details means that I am flying blind. Since I can't see the error, I don't know if I'll be able to fix this in the next 2 tries or 2 million tries.\n\nHere is my notebook: https://www.kaggle.com/ankurgupta1985/randomly-selected\n\nI am already chunking the data and not loading all of it in memory at one time. But, I assume that the final submission.csv is small enough that the corresponding dataframe would fit in memory. I have already asked friends who're experienced at Kaggle to help me and they're dumbfounded too. If anyone can help, I promise to consider naming my first born after you. Thank you very much!",
      "votes": null
    },
    {
      "id": "704850",
      "postDate": "12/28/2019 04:45:24",
      "content": "<p>After wasting a ton of time, I finally fixed it. The following is what I think may be happening. Please note that since we don't see the private data or even the schema of private data, or the errors behind the scoring errors, there is no way to know anything with surety.</p>\n\n<p><strong>Contest organizer changes the contents of the same test data file</strong>\nWhen you predict on the test dataset for yourself (i.e., not for submission), the file <code>/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl</code> seems to have 346 examples. This is the public test dataset. When you submit, your notebook is re-run as-is (without any modification to your code) using the same filename <code>/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl</code> but the file is replaced to have many more examples (some people say 10x more). This is the private test dataset. </p>\n\n<p><strong>The schema of private test dataset may not be the same or have the same properties as the public test dataset</strong>\nIt's possible to load the public test dataset using pandas with chunking so that we only load 10 lines/examples at a time and predict on those 10 examples before loading more data. But, the same exact code fails on the private test dataset. This is not a memory issue because we're not loading more than 10 lines at a time. This indicates that something else fails when reading the private test dataset even if reading the public test dataset the exact same way passed. My guess is that pandas chokes on on the private test dataset is due to a schema error. One example could be that different lines/examples in the private test dataset have different number of columns. </p>\n\n<p><strong>What fails on private test dataset but works for public test dataset</strong>\nFrom <a href=\"https://www.kaggle.com/ankurgupta1985/randomly-selected\">https://www.kaggle.com/ankurgupta1985/randomly-selected</a>:\n<code>\ntest_df_reader = pd.read_json('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', lines=True, chunksize=100)\nsubmission_chunks = [generate_submission(chunk_test_df) for chunk_test_df in test_df_reader]\n</code></p>\n\n<p><strong>What works on both public and private test datasets</strong>\nFrom <a href=\"https://www.kaggle.com/ankurgupta1985/alternative-io-read\">https://www.kaggle.com/ankurgupta1985/alternative-io-read</a>:\n<code>\nwith open('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', 'r') as f:\n    for i, line in enumerate(f):\n        parsed_line = json.loads(line)\n        chunk_test_df = pd.DataFrame.from_records([parsed_line], index=[0])\n        submission_chunks = submission_chunks + [generate_submission(chunk_test_df)]  # redefine over append\n</code></p>\n\n<p><strong>Conclusions</strong>\n1. This is not your fault. It's the fault of the contest organizers. \n2. Contest organizers should have provided information on how exactly the private test dataset was being used. This is essential and no amount of security concerns can ameliorate the lack of this information. How do we know that contestants are not making silent errors because of lack of this information? Is the contest really fair if only some people have this information or if only some people wasted their time fixing this artificial issue?\n3. The size of private datasets should be provided upfront. Google engineers know better than most that data size affects choice of tools. How are the users supposed to decide on which tools to use if users don't know what size of data they're going to handle? Users don't get any feedback on what caused a scoring error. Users are expected to divine if it was indeed a data size issue or some other issue.\n4. Given all this secrecy, the least they could do is provide the schema for the private test dataset and ensure that it matches exactly as the public test dataset. </p>\n\n<p>I hope this post helps other people save one day of their lives. </p>",
      "rawMarkdown": "After wasting a ton of time, I finally fixed it. The following is what I think may be happening. Please note that since we don't see the private data or even the schema of private data, or the errors behind the scoring errors, there is no way to know anything with surety.\n\n**Contest organizer changes the contents of the same test data file**\nWhen you predict on the test dataset for yourself (i.e., not for submission), the file `/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl` seems to have 346 examples. This is the public test dataset. When you submit, your notebook is re-run as-is (without any modification to your code) using the same filename `/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl` but the file is replaced to have many more examples (some people say 10x more). This is the private test dataset. \n\n**The schema of private test dataset may not be the same or have the same properties as the public test dataset**\nIt's possible to load the public test dataset using pandas with chunking so that we only load 10 lines/examples at a time and predict on those 10 examples before loading more data. But, the same exact code fails on the private test dataset. This is not a memory issue because we're not loading more than 10 lines at a time. This indicates that something else fails when reading the private test dataset even if reading the public test dataset the exact same way passed. My guess is that pandas chokes on on the private test dataset is due to a schema error. One example could be that different lines/examples in the private test dataset have different number of columns. \n\n**What fails on private test dataset but works for public test dataset**\nFrom https://www.kaggle.com/ankurgupta1985/randomly-selected:\n```\ntest_df_reader = pd.read_json('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', lines=True, chunksize=100)\nsubmission_chunks = [generate_submission(chunk_test_df) for chunk_test_df in test_df_reader]\n```\n\n**What works on both public and private test datasets**\nFrom https://www.kaggle.com/ankurgupta1985/alternative-io-read:\n```\nwith open('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', 'r') as f:\n    for i, line in enumerate(f):\n        parsed_line = json.loads(line)\n        chunk_test_df = pd.DataFrame.from_records([parsed_line], index=[0])\n        submission_chunks = submission_chunks + [generate_submission(chunk_test_df)]  # redefine over append\n```\n\n**Conclusions**\n1. This is not your fault. It's the fault of the contest organizers. \n2. Contest organizers should have provided information on how exactly the private test dataset was being used. This is essential and no amount of security concerns can ameliorate the lack of this information. How do we know that contestants are not making silent errors because of lack of this information? Is the contest really fair if only some people have this information or if only some people wasted their time fixing this artificial issue?\n3. The size of private datasets should be provided upfront. Google engineers know better than most that data size affects choice of tools. How are the users supposed to decide on which tools to use if users don't know what size of data they're going to handle? Users don't get any feedback on what caused a scoring error. Users are expected to divine if it was indeed a data size issue or some other issue.\n4. Given all this secrecy, the least they could do is provide the schema for the private test dataset and ensure that it matches exactly as the public test dataset. \n\nI hope this post helps other people save one day of their lives.",
      "votes": null
    },
    {
      "id": "705422",
      "postDate": "12/28/2019 23:10:05",
      "content": "<p>Thanks for sharing...</p>",
      "rawMarkdown": "Thanks for sharing...",
      "votes": null
    },
    {
      "id": "706752",
      "postDate": "12/30/2019 19:18:48",
      "content": "<p><a href=\"/ankurgupta1985\">@ankurgupta1985</a> Sorry that you're facing difficulties with your submission. We're continuing to refine our submission flow and trying to strike the right balance between a positive user experience and providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering). I will admit that we can do better. Your feedback is noted, and thank you for sharing your pitfalls for others to learn from.</p>",
      "rawMarkdown": "ankurgupta1985 Sorry that you're facing difficulties with your submission. We're continuing to refine our submission flow and trying to strike the right balance between a positive user experience and providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering). I will admit that we can do better. Your feedback is noted, and thank you for sharing your pitfalls for others to learn from.",
      "votes": null
    },
    {
      "id": "706957",
      "postDate": "12/31/2019 04:32:47",
      "content": "<blockquote>\n  <p>providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering)</p>\n</blockquote>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2510328%2Fb98684c58e3dd03ca6793b7e22a6b38b%2FOoh.jpg?generation=1577766687489555&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "&gt; providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2510328%2Fb98684c58e3dd03ca6793b7e22a6b38b%2FOoh.jpg?generation=1577766687489555&amp;alt=media)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 704850,
      "author_name": "ankurgupta1985",
      "author_url": "",
      "post_date": "12/28/2019 04:45:24",
      "content": "<p>After wasting a ton of time, I finally fixed it. The following is what I think may be happening. Please note that since we don't see the private data or even the schema of private data, or the errors behind the scoring errors, there is no way to know anything with surety.</p>\n\n<p><strong>Contest organizer changes the contents of the same test data file</strong>\nWhen you predict on the test dataset for yourself (i.e., not for submission), the file <code>/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl</code> seems to have 346 examples. This is the public test dataset. When you submit, your notebook is re-run as-is (without any modification to your code) using the same filename <code>/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl</code> but the file is replaced to have many more examples (some people say 10x more). This is the private test dataset. </p>\n\n<p><strong>The schema of private test dataset may not be the same or have the same properties as the public test dataset</strong>\nIt's possible to load the public test dataset using pandas with chunking so that we only load 10 lines/examples at a time and predict on those 10 examples before loading more data. But, the same exact code fails on the private test dataset. This is not a memory issue because we're not loading more than 10 lines at a time. This indicates that something else fails when reading the private test dataset even if reading the public test dataset the exact same way passed. My guess is that pandas chokes on on the private test dataset is due to a schema error. One example could be that different lines/examples in the private test dataset have different number of columns. </p>\n\n<p><strong>What fails on private test dataset but works for public test dataset</strong>\nFrom <a href=\"https://www.kaggle.com/ankurgupta1985/randomly-selected\">https://www.kaggle.com/ankurgupta1985/randomly-selected</a>:\n<code>\ntest_df_reader = pd.read_json('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', lines=True, chunksize=100)\nsubmission_chunks = [generate_submission(chunk_test_df) for chunk_test_df in test_df_reader]\n</code></p>\n\n<p><strong>What works on both public and private test datasets</strong>\nFrom <a href=\"https://www.kaggle.com/ankurgupta1985/alternative-io-read\">https://www.kaggle.com/ankurgupta1985/alternative-io-read</a>:\n<code>\nwith open('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', 'r') as f:\n    for i, line in enumerate(f):\n        parsed_line = json.loads(line)\n        chunk_test_df = pd.DataFrame.from_records([parsed_line], index=[0])\n        submission_chunks = submission_chunks + [generate_submission(chunk_test_df)]  # redefine over append\n</code></p>\n\n<p><strong>Conclusions</strong>\n1. This is not your fault. It's the fault of the contest organizers. \n2. Contest organizers should have provided information on how exactly the private test dataset was being used. This is essential and no amount of security concerns can ameliorate the lack of this information. How do we know that contestants are not making silent errors because of lack of this information? Is the contest really fair if only some people have this information or if only some people wasted their time fixing this artificial issue?\n3. The size of private datasets should be provided upfront. Google engineers know better than most that data size affects choice of tools. How are the users supposed to decide on which tools to use if users don't know what size of data they're going to handle? Users don't get any feedback on what caused a scoring error. Users are expected to divine if it was indeed a data size issue or some other issue.\n4. Given all this secrecy, the least they could do is provide the schema for the private test dataset and ensure that it matches exactly as the public test dataset. </p>\n\n<p>I hope this post helps other people save one day of their lives. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 705422,
      "author_name": "abualabed",
      "author_url": "",
      "post_date": "12/28/2019 23:10:05",
      "content": "<p>Thanks for sharing...</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 706752,
      "author_name": "juliaelliott",
      "author_url": "",
      "post_date": "12/30/2019 19:18:48",
      "content": "<p><a href=\"/ankurgupta1985\">@ankurgupta1985</a> Sorry that you're facing difficulties with your submission. We're continuing to refine our submission flow and trying to strike the right balance between a positive user experience and providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering). I will admit that we can do better. Your feedback is noted, and thank you for sharing your pitfalls for others to learn from.</p>",
      "votes": null,
      "replies": [
        {
          "id": 706957,
          "author_name": "msheriey",
          "author_url": "",
          "post_date": "12/31/2019 04:32:47",
          "content": "<blockquote>\n  <p>providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering)</p>\n</blockquote>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2510328%2Fb98684c58e3dd03ca6793b7e22a6b38b%2FOoh.jpg?generation=1577766687489555&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "704770": "I am facing the scoring error. I have looked at posts on this topic already and I have tried to fix the problem myself but I still can't figure out what is wrong. I also posted in one of the topics. \n\nSome questions I have:\n\n    Where is the private dataset?\n    How is the private dataset loaded into python memory? If my code doesn't load the private dataset, then what code does? Can I get a snippet of the code that loads the data?\n    Should we write a function that takes in one row at a time or one dataframe at a time or a path to a file ?\n\nHonestly, this is incredibly frustrating! I have spent my 3 out of 5 daily submission limit in 2 hours. I still don't know what I need to do to avoid this problem. The lack of error details means that I am flying blind. Since I can't see the error, I don't know if I'll be able to fix this in the next 2 tries or 2 million tries.\n\nHere is my notebook: https://www.kaggle.com/ankurgupta1985/randomly-selected\n\nI am already chunking the data and not loading all of it in memory at one time. But, I assume that the final submission.csv is small enough that the corresponding dataframe would fit in memory. I have already asked friends who're experienced at Kaggle to help me and they're dumbfounded too. If anyone can help, I promise to consider naming my first born after you. Thank you very much!",
    "704850": "After wasting a ton of time, I finally fixed it. The following is what I think may be happening. Please note that since we don't see the private data or even the schema of private data, or the errors behind the scoring errors, there is no way to know anything with surety.\n\n**Contest organizer changes the contents of the same test data file**\nWhen you predict on the test dataset for yourself (i.e., not for submission), the file `/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl` seems to have 346 examples. This is the public test dataset. When you submit, your notebook is re-run as-is (without any modification to your code) using the same filename `/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl` but the file is replaced to have many more examples (some people say 10x more). This is the private test dataset. \n\n**The schema of private test dataset may not be the same or have the same properties as the public test dataset**\nIt's possible to load the public test dataset using pandas with chunking so that we only load 10 lines/examples at a time and predict on those 10 examples before loading more data. But, the same exact code fails on the private test dataset. This is not a memory issue because we're not loading more than 10 lines at a time. This indicates that something else fails when reading the private test dataset even if reading the public test dataset the exact same way passed. My guess is that pandas chokes on on the private test dataset is due to a schema error. One example could be that different lines/examples in the private test dataset have different number of columns. \n\n**What fails on private test dataset but works for public test dataset**\nFrom https://www.kaggle.com/ankurgupta1985/randomly-selected:\n```\ntest_df_reader = pd.read_json('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', lines=True, chunksize=100)\nsubmission_chunks = [generate_submission(chunk_test_df) for chunk_test_df in test_df_reader]\n```\n\n**What works on both public and private test datasets**\nFrom https://www.kaggle.com/ankurgupta1985/alternative-io-read:\n```\nwith open('/kaggle/input/tensorflow2-question-answering/simplified-nq-test.jsonl', 'r') as f:\n    for i, line in enumerate(f):\n        parsed_line = json.loads(line)\n        chunk_test_df = pd.DataFrame.from_records([parsed_line], index=[0])\n        submission_chunks = submission_chunks + [generate_submission(chunk_test_df)]  # redefine over append\n```\n\n**Conclusions**\n1. This is not your fault. It's the fault of the contest organizers. \n2. Contest organizers should have provided information on how exactly the private test dataset was being used. This is essential and no amount of security concerns can ameliorate the lack of this information. How do we know that contestants are not making silent errors because of lack of this information? Is the contest really fair if only some people have this information or if only some people wasted their time fixing this artificial issue?\n3. The size of private datasets should be provided upfront. Google engineers know better than most that data size affects choice of tools. How are the users supposed to decide on which tools to use if users don't know what size of data they're going to handle? Users don't get any feedback on what caused a scoring error. Users are expected to divine if it was indeed a data size issue or some other issue.\n4. Given all this secrecy, the least they could do is provide the schema for the private test dataset and ensure that it matches exactly as the public test dataset. \n\nI hope this post helps other people save one day of their lives.",
    "705422": "Thanks for sharing...",
    "706752": "ankurgupta1985 Sorry that you're facing difficulties with your submission. We're continuing to refine our submission flow and trying to strike the right balance between a positive user experience and providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering). I will admit that we can do better. Your feedback is noted, and thank you for sharing your pitfalls for others to learn from.",
    "706957": "&gt; providing too much feedback that will lead to leaks/probing (which the community has proven itself extremely tenacious in uncovering)\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2510328%2Fb98684c58e3dd03ca6793b7e22a6b38b%2FOoh.jpg?generation=1577766687489555&amp;alt=media)"
  },
  "source": "meta"
}