{
  "id": 551587,
  "title": "Notebook Inference Error",
  "url": "/competitions/jane-street-real-time-market-data-forecasting/discussion/551587",
  "author_name": "",
  "post_date": "2024-12-14T05:33:45.756920900Z",
  "votes": 2,
  "comment_count": 3,
  "views": 0,
  "content": "<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/yufanshiu/submission-parquet</a></p>\n<p>It's keep scoring for several hours but fails eventually, can someone help me to solve it, thank you.</p>\n<p>If the link don't works, please use copy and paste, I don't know why</p>",
  "messages": [
    {
      "id": "3071678",
      "postDate": "12/14/2024 05:33:45",
      "content": "<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/code/yufanshiu/submission-parquet</a></p>\n<p>It's keep scoring for several hours but fails eventually, can someone help me to solve it, thank you.</p>\n<p>If the link don't works, please use copy and paste, I don't know why</p>",
      "rawMarkdown": "[https://www.kaggle.com/code/yufanshiu/submission-parquet](url)\n\nIt's keep scoring for several hours but fails eventually, can someone help me to solve it, thank you.\n\nIf the link don't works, please use copy and paste, I don't know why",
      "votes": null
    },
    {
      "id": "3072031",
      "postDate": "12/14/2024 15:49:56",
      "content": "<p>I recommend you to verify the Inference Pipeline.<br>\n    - The issue might lie in how the model is applied during scoring. For example:<br>\n         - Predicting on a massive dataset without batching.<br>\n         - The model itself being too large for Kaggle’s GPU/CPU limits. So analyze your runtime and memory constraints. I <br>\n           will show you example code.<br>\n                     import psutil<br>\n                     print(f\"Memory Usage: {psutil.virtual_memory().percent}%\")</p>\n<ul>\n<li>Please check submission.parquet<br>\n      - Ensure the submission.parquet file is valid and correctly formatted.<br>\n      - Load it in a separate notebook to verify:<br>\n               import pandas as pd<br>\n               try:<br>\n                  df = pd.read_parquet(\"submission.parquet\")<br>\n                  print(df.info())<br>\n               except Exception as e:<br>\n                  print(f\"Error loading parquet file: {e}\")</li>\n</ul>",
      "rawMarkdown": "I recommend you to verify the Inference Pipeline.\n    - The issue might lie in how the model is applied during scoring. For example:\n         - Predicting on a massive dataset without batching.\n         - The model itself being too large for Kaggle’s GPU/CPU limits. So analyze your runtime and memory constraints. I \n           will show you example code.\n                     import psutil\n                     print(f\"Memory Usage: {psutil.virtual_memory().percent}%\")\n   - Please check submission.parquet\n          - Ensure the submission.parquet file is valid and correctly formatted.\n          - Load it in a separate notebook to verify:\n                   import pandas as pd\n                   try:\n                      df = pd.read_parquet(\"submission.parquet\")\n                      print(df.info())\n                   except Exception as e:\n                      print(f\"Error loading parquet file: {e}\")",
      "votes": null
    },
    {
      "id": "3072279",
      "postDate": "12/15/2024 01:33:29",
      "content": "<p>I explored the steps suggested below, but I still get this stubborn error after a few hours or so from submission: \"Notebook Inference Server ErrorYour submission notebook's inference server was disconnected unexpectedly, or a request timed out. See more debugging tips\". Is that the problem due to a 60-sec time timeout, or could that be memory issues, etc.?</p>",
      "rawMarkdown": "I explored the steps suggested below, but I still get this stubborn error after a few hours or so from submission: \"Notebook Inference Server ErrorYour submission notebook's inference server was disconnected unexpectedly, or a request timed out. See more debugging tips\". Is that the problem due to a 60-sec time timeout, or could that be memory issues, etc.?",
      "votes": null
    },
    {
      "id": "3072418",
      "postDate": "12/15/2024 06:50:33",
      "content": "<p>Is the real inference is a dataset instead of several batch?</p>",
      "rawMarkdown": "Is the real inference is a dataset instead of several batch?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3072031,
      "author_name": "amunsentom",
      "author_url": "",
      "post_date": "12/14/2024 15:49:56",
      "content": "<p>I recommend you to verify the Inference Pipeline.<br>\n    - The issue might lie in how the model is applied during scoring. For example:<br>\n         - Predicting on a massive dataset without batching.<br>\n         - The model itself being too large for Kaggle’s GPU/CPU limits. So analyze your runtime and memory constraints. I <br>\n           will show you example code.<br>\n                     import psutil<br>\n                     print(f\"Memory Usage: {psutil.virtual_memory().percent}%\")</p>\n<ul>\n<li>Please check submission.parquet<br>\n      - Ensure the submission.parquet file is valid and correctly formatted.<br>\n      - Load it in a separate notebook to verify:<br>\n               import pandas as pd<br>\n               try:<br>\n                  df = pd.read_parquet(\"submission.parquet\")<br>\n                  print(df.info())<br>\n               except Exception as e:<br>\n                  print(f\"Error loading parquet file: {e}\")</li>\n</ul>",
      "votes": null,
      "replies": [
        {
          "id": 3072418,
          "author_name": "yufanshiu",
          "author_url": "",
          "post_date": "12/15/2024 06:50:33",
          "content": "<p>Is the real inference is a dataset instead of several batch?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 3072279,
      "author_name": "izaznov",
      "author_url": "",
      "post_date": "12/15/2024 01:33:29",
      "content": "<p>I explored the steps suggested below, but I still get this stubborn error after a few hours or so from submission: \"Notebook Inference Server ErrorYour submission notebook's inference server was disconnected unexpectedly, or a request timed out. See more debugging tips\". Is that the problem due to a 60-sec time timeout, or could that be memory issues, etc.?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "3071678": "[https://www.kaggle.com/code/yufanshiu/submission-parquet](url)\n\nIt's keep scoring for several hours but fails eventually, can someone help me to solve it, thank you.\n\nIf the link don't works, please use copy and paste, I don't know why",
    "3072031": "I recommend you to verify the Inference Pipeline.\n    - The issue might lie in how the model is applied during scoring. For example:\n         - Predicting on a massive dataset without batching.\n         - The model itself being too large for Kaggle’s GPU/CPU limits. So analyze your runtime and memory constraints. I \n           will show you example code.\n                     import psutil\n                     print(f\"Memory Usage: {psutil.virtual_memory().percent}%\")\n   - Please check submission.parquet\n          - Ensure the submission.parquet file is valid and correctly formatted.\n          - Load it in a separate notebook to verify:\n                   import pandas as pd\n                   try:\n                      df = pd.read_parquet(\"submission.parquet\")\n                      print(df.info())\n                   except Exception as e:\n                      print(f\"Error loading parquet file: {e}\")",
    "3072279": "I explored the steps suggested below, but I still get this stubborn error after a few hours or so from submission: \"Notebook Inference Server ErrorYour submission notebook's inference server was disconnected unexpectedly, or a request timed out. See more debugging tips\". Is that the problem due to a 60-sec time timeout, or could that be memory issues, etc.?",
    "3072418": "Is the real inference is a dataset instead of several batch?"
  },
  "source": "meta"
}