{
  "id": 518377,
  "title": "Submission Scoring Error",
  "url": "/competitions/uspto-explainable-ai/discussion/518377",
  "author_name": "opamusora (Ivan Viakhirev)",
  "post_date": "2024-07-06T10:33:38.918000",
  "votes": 4,
  "comment_count": 14,
  "views": 0,
  "content": "<p>Getting this error constantly, did anyone stumble upon it and how to fix that?? I'm assuming that it could be whoosh queries freezes or something like that cause it first occured when I tried passing a little longer queries (can't reproduce anything alike locally though), but how to fix it without limiting query length too much?</p>",
  "messages": [
    {
      "id": 2916699,
      "postDate": "2024-07-11T07:33:18.050Z",
      "content": "<p>Some possibilities I think for Submission Scoring Error:</p>\n<ul>\n<li>generated submission.csv has the wrong length or is missing rows</li>\n<li>there are other files created in output directory (I had one submission where I accidentally included a generated log file and it failed)</li>\n<li>a particular query causes exception in Whoosh</li>\n</ul>\n<p>An out of memory issue should trigger Notebook Out of Memory and a data access error (for example fetching patent data that doesn't exist) should trigger Notebook Threw Exception, but I think its possible also in rare cases that both types of error cause a Submission Scoring Error, depending on your computation code.</p>\n<p>I think the best way to detect and fix is to generate and run your queries against a validation index of 2000 to 3000 rows from nearest_neighbors.csv with whoosh_utils.py. Each query should be checked by token count and the parser before search. Benchmark the search time, the queries should complete for the index within 9 hours, you can set a threshold for each individual query also. Check the final length of the submission.csv and your default query fallback.</p>\n<p>Once you have your setup in place it should be smooth sailing from there</p>",
      "rawMarkdown": "Some possibilities I think for Submission Scoring Error:\n- generated submission.csv has the wrong length or is missing rows\n- there are other files created in output directory (I had one submission where I accidentally included a generated log file and it failed)\n- a particular query causes exception in Whoosh\n\nAn out of memory issue should trigger Notebook Out of Memory and a data access error (for example fetching patent data that doesn't exist) should trigger Notebook Threw Exception, but I think its possible also in rare cases that both types of error cause a Submission Scoring Error, depending on your computation code.\n\nI think the best way to detect and fix is to generate and run your queries against a validation index of 2000 to 3000 rows from nearest_neighbors.csv with whoosh_utils.py. Each query should be checked by token count and the parser before search. Benchmark the search time, the queries should complete for the index within 9 hours, you can set a threshold for each individual query also. Check the final length of the submission.csv and your default query fallback.\n\nOnce you have your setup in place it should be smooth sailing from there",
      "votes": 4,
      "replies": [
        {
          "id": 2917686,
          "postDate": "2024-07-11T18:16:34.830Z",
          "content": "<p>thanks for the advices! Indeed, i have my temp file in the output directory. However, i still face the same issue after removing all the files from the output dir except submission.cvs. </p>\n<p>I use test.csv to test the generated quires. 10 rows generates 10 valid quires that have been validated by the whoosh validator and token length check. the average response time for the 10 queries is 7.06 seconds. is this a reasonable time? </p>\n<p>I believe the scoring program uses different publications than the 10 in the test.csv. should i extend my test dataset to, saying 1000 from nearest_neighbors.csv? </p>",
          "rawMarkdown": "thanks for the advices! Indeed, i have my temp file in the output directory. However, i still face the same issue after removing all the files from the output dir except submission.cvs. \n\nI use test.csv to test the generated quires. 10 rows generates 10 valid quires that have been validated by the whoosh validator and token length check. the average response time for the 10 queries is 7.06 seconds. is this a reasonable time? \n\nI believe the scoring program uses different publications than the 10 in the test.csv. should i extend my test dataset to, saying 1000 from nearest_neighbors.csv? ",
          "replies": [
            {
              "id": 2918071,
              "postDate": "2024-07-12T02:51:14.933Z",
              "content": "<p>7 seconds for 10 queries is absolutely fine but the sample size is too small, imo 2000 is the minimum to debug both stable performance and scores. If you have scoring error within 1-2 hours, you might want to check your query generation code or final submission.csv output more thoroughly, if its scoring error after that up to 9 hours, it might be a query generated for a specific target patent and neighbors that is tripping up the search</p>",
              "rawMarkdown": "7 seconds for 10 queries is absolutely fine but the sample size is too small, imo 2000 is the minimum to debug both stable performance and scores. If you have scoring error within 1-2 hours, you might want to check your query generation code or final submission.csv output more thoroughly, if its scoring error after that up to 9 hours, it might be a query generated for a specific target patent and neighbors that is tripping up the search"
            },
            {
              "id": 2918376,
              "postDate": "2024-07-12T07:39:01.240Z",
              "content": "<p>apologies that i didn't make it clear. the average processing time for each query is 7.06 for the 10 test queries. Also, i notice that i cannot even load all data from nearest_neighbors.csv at once because it used up all 30gb memory when i do <code>test_df = pd.read_csv('/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv')</code>. i think i can load and test in smaller batches. however, i want to confirm that you also encounter the same issue before i spend too much time if that is caused by something that i did wrong. </p>",
              "rawMarkdown": "apologies that i didn't make it clear. the average processing time for each query is 7.06 for the 10 test queries. Also, i notice that i cannot even load all data from nearest_neighbors.csv at once because it used up all 30gb memory when i do `test_df = pd.read_csv('/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv')`. i think i can load and test in smaller batches. however, i want to confirm that you also encounter the same issue before i spend too much time if that is caused by something that i did wrong. "
            },
            {
              "id": 2919643,
              "postDate": "2024-07-13T04:37:37.903Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2919644,
              "postDate": "2024-07-13T04:38:03.997Z",
              "content": "<p>by using <code>nrows</code>, you can read only a specific number of rows, which can help avoid memory errors.<br>\n<code>test_df = pd.read_csv('/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv', nrows=2000)</code></p>",
              "rawMarkdown": "by using `nrows`, you can read only a specific number of rows, which can help avoid memory errors.\n`test_df = pd.read_csv('/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv', nrows=2000)`",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2908424,
      "postDate": "2024-07-06T10:33:38.917Z",
      "content": "<p>Getting this error constantly, did anyone stumble upon it and how to fix that?? I'm assuming that it could be whoosh queries freezes or something like that cause it first occured when I tried passing a little longer queries (can't reproduce anything alike locally though), but how to fix it without limiting query length too much?</p>",
      "rawMarkdown": "Getting this error constantly, did anyone stumble upon it and how to fix that?? I'm assuming that it could be whoosh queries freezes or something like that cause it first occured when I tried passing a little longer queries (can't reproduce anything alike locally though), but how to fix it without limiting query length too much?",
      "votes": 4
    },
    {
      "id": 2921307,
      "postDate": "2024-07-14T09:58:28.580Z",
      "content": "<p>try to check if any empty query is generated. and if empty (\"\") than replace it with a dum query like (ti:dummy)</p>",
      "rawMarkdown": "try to check if any empty query is generated. and if empty (\"\") than replace it with a dum query like (ti:dummy)",
      "votes": 1
    },
    {
      "id": 2919660,
      "postDate": "2024-07-13T05:16:26.287Z",
      "content": "<p>As mentioned in the Overview, the number of tokens in the query might be the cause.<br>\n<code>Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.</code></p>\n<p><a href=\"http://www.kaggle.com/competitions/uspto-explainable-ai/overview/evaluation\" target=\"_blank\">www.kaggle.com/competitions/uspto-explainable-ai/overview/evaluation</a></p>",
      "rawMarkdown": "As mentioned in the Overview, the number of tokens in the query might be the cause.\n`Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.`\n\nwww.kaggle.com/competitions/uspto-explainable-ai/overview/evaluation",
      "replies": [
        {
          "id": 2919734,
          "postDate": "2024-07-13T07:02:54.537Z",
          "content": "<p>all of my generated queries are strictly less than 50 tokens.</p>",
          "rawMarkdown": "all of my generated queries are strictly less than 50 tokens.",
          "replies": [
            {
              "id": 2920021,
              "postDate": "2024-07-13T11:37:20.977Z",
              "content": "<p>Despite reducing it to under 50 tokens, I was still encountering errors. However, after making adjustments with the following constraints in mind, I was able to make submissions.</p>\n<p><code>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index.</code></p>",
              "rawMarkdown": "Despite reducing it to under 50 tokens, I was still encountering errors. However, after making adjustments with the following constraints in mind, I was able to make submissions.\n\n`The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. `"
            },
            {
              "id": 2920977,
              "postDate": "2024-07-14T01:04:33.233Z",
              "content": "<p>do you know how approximately many queries generated during the submission? I have no idea whether 7s/query is too slow or just ok</p>",
              "rawMarkdown": "do you know how approximately many queries generated during the submission? I have no idea whether 7s/query is too slow or just ok"
            },
            {
              "id": 2921135,
              "postDate": "2024-07-14T05:54:43.070Z",
              "content": "<p>Based on the dataset description, I understand it to be 2500 queries.<br>\n<code>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</code></p>\n<p><a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/data\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/data</a></p>",
              "rawMarkdown": "Based on the dataset description, I understand it to be 2500 queries.\n`test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.`\n\nhttps://www.kaggle.com/competitions/uspto-explainable-ai/data"
            }
          ]
        }
      ]
    },
    {
      "id": 2916666,
      "postDate": "2024-07-11T06:59:41.277Z",
      "content": "<p>facing exactly the same issue here. it doesn't give specific reason or log for the failure. i don't know how to debug.</p>",
      "rawMarkdown": "facing exactly the same issue here. it doesn't give specific reason or log for the failure. i don't know how to debug.",
      "replies": [
        {
          "id": 2932542,
          "postDate": "2024-07-23T05:14:20.010Z",
          "content": "<p>I am encountering the same problem. Maybe, does it result from the requirement below?<br>\n**<br>\nThe metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.**</p>",
          "rawMarkdown": "I am encountering the same problem. Maybe, does it result from the requirement below?\n**\nThe metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.**"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2916699,
      "author_name": "dt",
      "author_url": "",
      "post_date": "2024-07-11T07:33:18.050000",
      "content": "<p>Some possibilities I think for Submission Scoring Error:</p>\n<ul>\n<li>generated submission.csv has the wrong length or is missing rows</li>\n<li>there are other files created in output directory (I had one submission where I accidentally included a generated log file and it failed)</li>\n<li>a particular query causes exception in Whoosh</li>\n</ul>\n<p>An out of memory issue should trigger Notebook Out of Memory and a data access error (for example fetching patent data that doesn't exist) should trigger Notebook Threw Exception, but I think its possible also in rare cases that both types of error cause a Submission Scoring Error, depending on your computation code.</p>\n<p>I think the best way to detect and fix is to generate and run your queries against a validation index of 2000 to 3000 rows from nearest_neighbors.csv with whoosh_utils.py. Each query should be checked by token count and the parser before search. Benchmark the search time, the queries should complete for the index within 9 hours, you can set a threshold for each individual query also. Check the final length of the submission.csv and your default query fallback.</p>\n<p>Once you have your setup in place it should be smooth sailing from there</p>",
      "votes": 4,
      "replies": [
        {
          "id": 2917686,
          "author_name": "Frank6526",
          "author_url": "",
          "post_date": "2024-07-11T18:16:34.830000",
          "content": "<p>thanks for the advices! Indeed, i have my temp file in the output directory. However, i still face the same issue after removing all the files from the output dir except submission.cvs. </p>\n<p>I use test.csv to test the generated quires. 10 rows generates 10 valid quires that have been validated by the whoosh validator and token length check. the average response time for the 10 queries is 7.06 seconds. is this a reasonable time? </p>\n<p>I believe the scoring program uses different publications than the 10 in the test.csv. should i extend my test dataset to, saying 1000 from nearest_neighbors.csv? </p>",
          "votes": 0,
          "replies": [
            {
              "id": 2918071,
              "author_name": "dt",
              "author_url": "",
              "post_date": "2024-07-12T02:51:14.933000",
              "content": "<p>7 seconds for 10 queries is absolutely fine but the sample size is too small, imo 2000 is the minimum to debug both stable performance and scores. If you have scoring error within 1-2 hours, you might want to check your query generation code or final submission.csv output more thoroughly, if its scoring error after that up to 9 hours, it might be a query generated for a specific target patent and neighbors that is tripping up the search</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2918376,
              "author_name": "Frank6526",
              "author_url": "",
              "post_date": "2024-07-12T07:39:01.240000",
              "content": "<p>apologies that i didn't make it clear. the average processing time for each query is 7.06 for the 10 test queries. Also, i notice that i cannot even load all data from nearest_neighbors.csv at once because it used up all 30gb memory when i do <code>test_df = pd.read_csv('/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv')</code>. i think i can load and test in smaller batches. however, i want to confirm that you also encounter the same issue before i spend too much time if that is caused by something that i did wrong. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2919643,
              "author_name": "",
              "author_url": "",
              "post_date": "2024-07-13T04:37:37.903000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2919644,
              "author_name": "uwakuchi",
              "author_url": "",
              "post_date": "2024-07-13T04:38:03.997000",
              "content": "<p>by using <code>nrows</code>, you can read only a specific number of rows, which can help avoid memory errors.<br>\n<code>test_df = pd.read_csv('/kaggle/input/uspto-explainable-ai/nearest_neighbors.csv', nrows=2000)</code></p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2921307,
      "author_name": "Shreyas Bhatt",
      "author_url": "",
      "post_date": "2024-07-14T09:58:28.580000",
      "content": "<p>try to check if any empty query is generated. and if empty (\"\") than replace it with a dum query like (ti:dummy)</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2919660,
      "author_name": "uwakuchi",
      "author_url": "",
      "post_date": "2024-07-13T05:16:26.287000",
      "content": "<p>As mentioned in the Overview, the number of tokens in the query might be the cause.<br>\n<code>Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.</code></p>\n<p><a href=\"http://www.kaggle.com/competitions/uspto-explainable-ai/overview/evaluation\" target=\"_blank\">www.kaggle.com/competitions/uspto-explainable-ai/overview/evaluation</a></p>",
      "votes": 0,
      "replies": [
        {
          "id": 2919734,
          "author_name": "Frank6526",
          "author_url": "",
          "post_date": "2024-07-13T07:02:54.537000",
          "content": "<p>all of my generated queries are strictly less than 50 tokens.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2920021,
              "author_name": "uwakuchi",
              "author_url": "",
              "post_date": "2024-07-13T11:37:20.977000",
              "content": "<p>Despite reducing it to under 50 tokens, I was still encountering errors. However, after making adjustments with the following constraints in mind, I was able to make submissions.</p>\n<p><code>The metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index.</code></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2920977,
              "author_name": "Frank6526",
              "author_url": "",
              "post_date": "2024-07-14T01:04:33.233000",
              "content": "<p>do you know how approximately many queries generated during the submission? I have no idea whether 7s/query is too slow or just ok</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2921135,
              "author_name": "uwakuchi",
              "author_url": "",
              "post_date": "2024-07-14T05:54:43.070000",
              "content": "<p>Based on the dataset description, I understand it to be 2500 queries.<br>\n<code>test.csv A subset of nearest_neighbors.csv that will cover 2,500 patents in the hidden dataset.</code></p>\n<p><a href=\"https://www.kaggle.com/competitions/uspto-explainable-ai/data\" target=\"_blank\">https://www.kaggle.com/competitions/uspto-explainable-ai/data</a></p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2916666,
      "author_name": "Frank6526",
      "author_url": "",
      "post_date": "2024-07-11T06:59:41.277000",
      "content": "<p>facing exactly the same issue here. it doesn't give specific reason or log for the failure. i don't know how to debug.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2932542,
          "author_name": "yich",
          "author_url": "",
          "post_date": "2024-07-23T05:14:20.010000",
          "content": "<p>I am encountering the same problem. Maybe, does it result from the requirement below?<br>\n**<br>\nThe metric notebook must finish running your queries in 60 minutes, not including the time required for loading the whoosh index. Note that the metric uses four Whoosh searchers in parallel. The time required to start the metric notebook and download the data it uses does not count towards the 60 minutes.**</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2916699": "Some possibilities I think for Submission Scoring Error:\n- generated submission.csv has the wrong length or is missing rows\n- there are other files created in output directory (I had one submission where I accidentally included a generated log file and it failed)\n- a particular query causes exception in Whoosh\n\nAn out of memory issue should trigger Notebook Out of Memory and a data access error (for example fetching patent data that doesn't exist) should trigger Notebook Threw Exception, but I think its possible also in rare cases that both types of error cause a Submission Scoring Error, depending on your computation code.\n\nI think the best way to detect and fix is to generate and run your queries against a validation index of 2000 to 3000 rows from nearest_neighbors.csv with whoosh_utils.py. Each query should be checked by token count and the parser before search. Benchmark the search time, the queries should complete for the index within 9 hours, you can set a threshold for each individual query also. Check the final length of the submission.csv and your default query fallback.\n\nOnce you have your setup in place it should be smooth sailing from there",
    "2908424": "Getting this error constantly, did anyone stumble upon it and how to fix that?? I'm assuming that it could be whoosh queries freezes or something like that cause it first occured when I tried passing a little longer queries (can't reproduce anything alike locally though), but how to fix it without limiting query length too much?",
    "2921307": "try to check if any empty query is generated. and if empty (\"\") than replace it with a dum query like (ti:dummy)",
    "2919660": "As mentioned in the Overview, the number of tokens in the query might be the cause.\n`Queries cannot include more than 50 tokens, as measured with whoosh_utils.count_query_tokens.`\n\nwww.kaggle.com/competitions/uspto-explainable-ai/overview/evaluation",
    "2916666": "facing exactly the same issue here. it doesn't give specific reason or log for the failure. i don't know how to debug."
  }
}