{
  "id": 516597,
  "title": "Scoring Timeout",
  "url": "/competitions/uspto-explainable-ai/discussion/516597",
  "author_name": "",
  "post_date": "2024-07-03T02:33:55.432973800Z",
  "votes": 2,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all!</p>\n<p>I am quite new to this Kaggle competition so pardon me if this is a stupid question I'm asking here…</p>\n<p>I have been making a few attempts to submit my notebook and trying to see if I can get a score from my result but I always get a Notebook Timeout error when it comes to scoring. The entire notebook does not take more than 10 minutes to finish running.</p>\n<p>My entire notebook runs successfully and I get the output as you see in the image below which I submitted:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671377%2F1e338b8fd458f4bb3c8b086a05e9cd87%2Fmsg572333411-1148957.000328.jpg?generation=1719973801458474&amp;alt=media\"></p>\n<p>These are all the data I used: </p>\n<pre><code>test_df = pd.read_csv()\n\nfile_path = \nmetadata_df = pd.read_parquet(file_path)\n\n\nparquet_files_path = \n...\n\n\n\nsub = pd.read_csv()\nsub.to_csv(,index=)\n</code></pre>\n<p><strong>Are there any checks I am missing out?</strong><br>\nI understand that: Only a few example rows equivalent to the real test set are available for download. When your submission is scored the test folders will be replaced with versions containing the complete test set.</p>\n<p><strong>Could this be the issue? If so, I might narrow down the heavy processing component to getting all relevant Parquet files. Do y'all have to use Dask/Multiproc to process them?</strong></p>\n<p>I cannot get a single score from any of my submissions. </p>\n<p>Any guidance will be appreciated!! Cheers!</p>",
  "messages": [
    {
      "id": "2901883",
      "postDate": "07/03/2024 02:33:55",
      "content": "<p>Hi all!</p>\n<p>I am quite new to this Kaggle competition so pardon me if this is a stupid question I'm asking here…</p>\n<p>I have been making a few attempts to submit my notebook and trying to see if I can get a score from my result but I always get a Notebook Timeout error when it comes to scoring. The entire notebook does not take more than 10 minutes to finish running.</p>\n<p>My entire notebook runs successfully and I get the output as you see in the image below which I submitted:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671377%2F1e338b8fd458f4bb3c8b086a05e9cd87%2Fmsg572333411-1148957.000328.jpg?generation=1719973801458474&amp;alt=media\"></p>\n<p>These are all the data I used: </p>\n<pre><code>test_df = pd.read_csv()\n\nfile_path = \nmetadata_df = pd.read_parquet(file_path)\n\n\nparquet_files_path = \n...\n\n\n\nsub = pd.read_csv()\nsub.to_csv(,index=)\n</code></pre>\n<p><strong>Are there any checks I am missing out?</strong><br>\nI understand that: Only a few example rows equivalent to the real test set are available for download. When your submission is scored the test folders will be replaced with versions containing the complete test set.</p>\n<p><strong>Could this be the issue? If so, I might narrow down the heavy processing component to getting all relevant Parquet files. Do y'all have to use Dask/Multiproc to process them?</strong></p>\n<p>I cannot get a single score from any of my submissions. </p>\n<p>Any guidance will be appreciated!! Cheers!</p>",
      "rawMarkdown": "Hi all!\n\nI am quite new to this Kaggle competition so pardon me if this is a stupid question I'm asking here...\n\nI have been making a few attempts to submit my notebook and trying to see if I can get a score from my result but I always get a Notebook Timeout error when it comes to scoring. The entire notebook does not take more than 10 minutes to finish running.\n\nMy entire notebook runs successfully and I get the output as you see in the image below which I submitted:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671377%2F1e338b8fd458f4bb3c8b086a05e9cd87%2Fmsg572333411-1148957.000328.jpg?generation=1719973801458474&alt=media)\n\nThese are all the data I used: \n```python\ntest_df = pd.read_csv(\"/kaggle/input/uspto-explainable-ai/test.csv\")\n\nfile_path = \"/kaggle/input/uspto-explainable-ai/patent_metadata.parquet\"\nmetadata_df = pd.read_parquet(file_path)\n\n# Getting all relevant Parquet files\nparquet_files_path = \"/kaggle/input/uspto-explainable-ai/patent_data/\"\n...\n\n# Submission:\n# Assign queries to submission\nsub = pd.read_csv(\"/kaggle/input/uspto-explainable-ai/sample_submission.csv\")\nsub.to_csv(\"submission.csv\",index=False)\n```\n\n**Are there any checks I am missing out?**\nI understand that: Only a few example rows equivalent to the real test set are available for download. When your submission is scored the test folders will be replaced with versions containing the complete test set.\n\n**Could this be the issue? If so, I might narrow down the heavy processing component to getting all relevant Parquet files. Do y'all have to use Dask/Multiproc to process them?**\n\nI cannot get a single score from any of my submissions. \n\nAny guidance will be appreciated!! Cheers!",
      "votes": null
    },
    {
      "id": "2903718",
      "postDate": "07/04/2024 01:20:18",
      "content": "<p>I found that in your output, these cpc_code seems to be missing the number part, you can check your data processing part, whether this part is deleted, the following is the result that can be scored. I'm more than happy if it helps you.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19211697%2Ffa8d469f0858089da0900becf2411c1b%2Fcpc_codes.png?generation=1720056016216060&amp;alt=media\"></p>",
      "rawMarkdown": "I found that in your output, these cpc_code seems to be missing the number part, you can check your data processing part, whether this part is deleted, the following is the result that can be scored. I'm more than happy if it helps you.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19211697%2Ffa8d469f0858089da0900becf2411c1b%2Fcpc_codes.png?generation=1720056016216060&alt=media)",
      "votes": null
    },
    {
      "id": "2905661",
      "postDate": "07/05/2024 05:17:32",
      "content": "<p>Thanks so much for your response and suggestions!</p>\n<p>I realised that my notebook timeout is most likely due to me trying to get all relevant Parquet files for my LLM solution. </p>\n<p>I saw that many public code uses this: <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\" target=\"_blank\">https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data</a></p>\n<p>But it does not have any Claim or Description data in it. </p>\n<p>Are there any workaround it? </p>",
      "rawMarkdown": "Thanks so much for your response and suggestions!\n\nI realised that my notebook timeout is most likely due to me trying to get all relevant Parquet files for my LLM solution. \n\nI saw that many public code uses this: https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\n\nBut it does not have any Claim or Description data in it. \n\nAre there any workaround it?",
      "votes": null
    },
    {
      "id": "2911605",
      "postDate": "07/08/2024 12:28:35",
      "content": "<p>I also have the same problem! Did you solve it <a href=\"https://www.kaggle.com/hopefulleee\" target=\"_blank\">@hopefulleee</a> ?</p>",
      "rawMarkdown": "I also have the same problem! Did you solve it @hopefulleee ?",
      "votes": null
    },
    {
      "id": "2912715",
      "postDate": "07/09/2024 03:42:38",
      "content": "<p>Unfortunately, I did not. I am still testing but my progress is slow. I'll be testing my code using this dataset soon: <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\" target=\"_blank\">https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data</a> </p>",
      "rawMarkdown": "Unfortunately, I did not. I am still testing but my progress is slow. I'll be testing my code using this dataset soon: https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data",
      "votes": null
    },
    {
      "id": "2912824",
      "postDate": "07/09/2024 05:06:39",
      "content": "<p>I have a bit of confusion, if you read all the files into memory, won't it exceed the memory limit?🤔🤔🤔 And if you want to use LLM, you can try this public notebook: <a href=\"https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng.I\" target=\"_blank\">https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng.I</a> hope this public notebook can be of help to you.🥰</p>",
      "rawMarkdown": "I have a bit of confusion, if you read all the files into memory, won't it exceed the memory limit?🤔🤔🤔 And if you want to use LLM, you can try this public notebook: https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng.I hope this public notebook can be of help to you.🥰",
      "votes": null
    },
    {
      "id": "2913060",
      "postDate": "07/09/2024 08:00:30",
      "content": "<p>I have my code optimised for reading as little files as possible but I guess it is still too much. </p>\n<p>Thank for your recommendation! As mentioned, I saw that many public code uses this: <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\" target=\"_blank\">https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data</a> </p>\n<p>Your recommendation uses that too, which I believe is able to avoid the timeout issue, at the expense of Claim and Description information… which might have to be a compromise I need to make. </p>",
      "rawMarkdown": "I have my code optimised for reading as little files as possible but I guess it is still too much. \n\nThank for your recommendation! As mentioned, I saw that many public code uses this: https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data \n\nYour recommendation uses that too, which I believe is able to avoid the timeout issue, at the expense of Claim and Description information... which might have to be a compromise I need to make.",
      "votes": null
    },
    {
      "id": "2916665",
      "postDate": "07/11/2024 06:56:29",
      "content": "<p>If you run your solution against a Whoosh validation index of 3000 rows from nearest neighbors csv and it completes within 9 hrs, it should pass for the submission. You can benchmark the search time for your validation index to check, one possibility is that for some target publication numbers, the cpc codes that neighbours have are very large, leading to long computation time depending on your solution (some patents have very many cpc codes while others have none).</p>\n<p>Btw Whoosh also has a feature to limit search time, it can be useful for testing (<a href=\"https://whoosh.readthedocs.io/en/latest/api/collectors.html#whoosh.collectors.TimeLimitCollector\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/api/collectors.html#whoosh.collectors.TimeLimitCollector</a>)</p>",
      "rawMarkdown": "If you run your solution against a Whoosh validation index of 3000 rows from nearest neighbors csv and it completes within 9 hrs, it should pass for the submission. You can benchmark the search time for your validation index to check, one possibility is that for some target publication numbers, the cpc codes that neighbours have are very large, leading to long computation time depending on your solution (some patents have very many cpc codes while others have none).\n\nBtw Whoosh also has a feature to limit search time, it can be useful for testing (https://whoosh.readthedocs.io/en/latest/api/collectors.html#whoosh.collectors.TimeLimitCollector)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2903718,
      "author_name": "alannikos",
      "author_url": "",
      "post_date": "07/04/2024 01:20:18",
      "content": "<p>I found that in your output, these cpc_code seems to be missing the number part, you can check your data processing part, whether this part is deleted, the following is the result that can be scored. I'm more than happy if it helps you.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19211697%2Ffa8d469f0858089da0900becf2411c1b%2Fcpc_codes.png?generation=1720056016216060&amp;alt=media\"></p>",
      "votes": null,
      "replies": [
        {
          "id": 2905661,
          "author_name": "hopefulleee",
          "author_url": "",
          "post_date": "07/05/2024 05:17:32",
          "content": "<p>Thanks so much for your response and suggestions!</p>\n<p>I realised that my notebook timeout is most likely due to me trying to get all relevant Parquet files for my LLM solution. </p>\n<p>I saw that many public code uses this: <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\" target=\"_blank\">https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data</a></p>\n<p>But it does not have any Claim or Description data in it. </p>\n<p>Are there any workaround it? </p>",
          "votes": null,
          "replies": [
            {
              "id": 2911605,
              "author_name": "octaviograu",
              "author_url": "",
              "post_date": "07/08/2024 12:28:35",
              "content": "<p>I also have the same problem! Did you solve it <a href=\"https://www.kaggle.com/hopefulleee\" target=\"_blank\">@hopefulleee</a> ?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2912715,
                  "author_name": "hopefulleee",
                  "author_url": "",
                  "post_date": "07/09/2024 03:42:38",
                  "content": "<p>Unfortunately, I did not. I am still testing but my progress is slow. I'll be testing my code using this dataset soon: <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\" target=\"_blank\">https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data</a> </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            },
            {
              "id": 2912824,
              "author_name": "alannikos",
              "author_url": "",
              "post_date": "07/09/2024 05:06:39",
              "content": "<p>I have a bit of confusion, if you read all the files into memory, won't it exceed the memory limit?🤔🤔🤔 And if you want to use LLM, you can try this public notebook: <a href=\"https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng.I\" target=\"_blank\">https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng.I</a> hope this public notebook can be of help to you.🥰</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2913060,
                  "author_name": "hopefulleee",
                  "author_url": "",
                  "post_date": "07/09/2024 08:00:30",
                  "content": "<p>I have my code optimised for reading as little files as possible but I guess it is still too much. </p>\n<p>Thank for your recommendation! As mentioned, I saw that many public code uses this: <a href=\"https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\" target=\"_blank\">https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data</a> </p>\n<p>Your recommendation uses that too, which I believe is able to avoid the timeout issue, at the expense of Claim and Description information… which might have to be a compromise I need to make. </p>",
                  "votes": null,
                  "replies": []
                }
              ]
            }
          ]
        }
      ]
    },
    {
      "id": 2916665,
      "author_name": "dilliontan",
      "author_url": "",
      "post_date": "07/11/2024 06:56:29",
      "content": "<p>If you run your solution against a Whoosh validation index of 3000 rows from nearest neighbors csv and it completes within 9 hrs, it should pass for the submission. You can benchmark the search time for your validation index to check, one possibility is that for some target publication numbers, the cpc codes that neighbours have are very large, leading to long computation time depending on your solution (some patents have very many cpc codes while others have none).</p>\n<p>Btw Whoosh also has a feature to limit search time, it can be useful for testing (<a href=\"https://whoosh.readthedocs.io/en/latest/api/collectors.html#whoosh.collectors.TimeLimitCollector\" target=\"_blank\">https://whoosh.readthedocs.io/en/latest/api/collectors.html#whoosh.collectors.TimeLimitCollector</a>)</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2901883": "Hi all!\n\nI am quite new to this Kaggle competition so pardon me if this is a stupid question I'm asking here...\n\nI have been making a few attempts to submit my notebook and trying to see if I can get a score from my result but I always get a Notebook Timeout error when it comes to scoring. The entire notebook does not take more than 10 minutes to finish running.\n\nMy entire notebook runs successfully and I get the output as you see in the image below which I submitted:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F4671377%2F1e338b8fd458f4bb3c8b086a05e9cd87%2Fmsg572333411-1148957.000328.jpg?generation=1719973801458474&alt=media)\n\nThese are all the data I used: \n```python\ntest_df = pd.read_csv(\"/kaggle/input/uspto-explainable-ai/test.csv\")\n\nfile_path = \"/kaggle/input/uspto-explainable-ai/patent_metadata.parquet\"\nmetadata_df = pd.read_parquet(file_path)\n\n# Getting all relevant Parquet files\nparquet_files_path = \"/kaggle/input/uspto-explainable-ai/patent_data/\"\n...\n\n# Submission:\n# Assign queries to submission\nsub = pd.read_csv(\"/kaggle/input/uspto-explainable-ai/sample_submission.csv\")\nsub.to_csv(\"submission.csv\",index=False)\n```\n\n**Are there any checks I am missing out?**\nI understand that: Only a few example rows equivalent to the real test set are available for download. When your submission is scored the test folders will be replaced with versions containing the complete test set.\n\n**Could this be the issue? If so, I might narrow down the heavy processing component to getting all relevant Parquet files. Do y'all have to use Dask/Multiproc to process them?**\n\nI cannot get a single score from any of my submissions. \n\nAny guidance will be appreciated!! Cheers!",
    "2903718": "I found that in your output, these cpc_code seems to be missing the number part, you can check your data processing part, whether this part is deleted, the following is the result that can be scored. I'm more than happy if it helps you.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F19211697%2Ffa8d469f0858089da0900becf2411c1b%2Fcpc_codes.png?generation=1720056016216060&alt=media)",
    "2905661": "Thanks so much for your response and suggestions!\n\nI realised that my notebook timeout is most likely due to me trying to get all relevant Parquet files for my LLM solution. \n\nI saw that many public code uses this: https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data\n\nBut it does not have any Claim or Description data in it. \n\nAre there any workaround it?",
    "2911605": "I also have the same problem! Did you solve it @hopefulleee ?",
    "2912715": "Unfortunately, I did not. I am still testing but my progress is slow. I'll be testing my code using this dataset soon: https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data",
    "2912824": "I have a bit of confusion, if you read all the files into memory, won't it exceed the memory limit?🤔🤔🤔 And if you want to use LLM, you can try this public notebook: https://www.kaggle.com/code/aerdem4/uspto-llm-solution-prompt-eng.I hope this public notebook can be of help to you.🥰",
    "2913060": "I have my code optimised for reading as little files as possible but I guess it is still too much. \n\nThank for your recommendation! As mentioned, I saw that many public code uses this: https://www.kaggle.com/datasets/aerdem4/uspto-all-patents-after-1975/data \n\nYour recommendation uses that too, which I believe is able to avoid the timeout issue, at the expense of Claim and Description information... which might have to be a compromise I need to make.",
    "2916665": "If you run your solution against a Whoosh validation index of 3000 rows from nearest neighbors csv and it completes within 9 hrs, it should pass for the submission. You can benchmark the search time for your validation index to check, one possibility is that for some target publication numbers, the cpc codes that neighbours have are very large, leading to long computation time depending on your solution (some patents have very many cpc codes while others have none).\n\nBtw Whoosh also has a feature to limit search time, it can be useful for testing (https://whoosh.readthedocs.io/en/latest/api/collectors.html#whoosh.collectors.TimeLimitCollector)"
  },
  "source": "meta"
}