{
  "id": 554378,
  "title": "\"Submission CSV Not Found\" error experienced for all 500 samples, but scoring passes for 10 samples",
  "url": "/competitions/czii-cryo-et-object-identification/discussion/554378",
  "author_name": "",
  "post_date": "2025-01-01T07:20:04.021455200Z",
  "votes": 1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I ran submission for all ~500 test samples - I get the \"Submission CSV Not Found\" error after ~4 hours of scoring time. </p>\n<p>If I run for only 10 test samples, then the notebook passes and I get a score. The only difference between the Submissions are the number of test samples scored.</p>\n<p>I don't understand why I would get the \"Submission CSV Not Found\" error for the version with ~500 test samples - could this error be obfuscating some other error?</p>",
  "messages": [
    {
      "id": "3085451",
      "postDate": "01/01/2025 07:20:04",
      "content": "<p>I ran submission for all ~500 test samples - I get the \"Submission CSV Not Found\" error after ~4 hours of scoring time. </p>\n<p>If I run for only 10 test samples, then the notebook passes and I get a score. The only difference between the Submissions are the number of test samples scored.</p>\n<p>I don't understand why I would get the \"Submission CSV Not Found\" error for the version with ~500 test samples - could this error be obfuscating some other error?</p>",
      "rawMarkdown": "I ran submission for all ~500 test samples - I get the \"Submission CSV Not Found\" error after ~4 hours of scoring time. \n\nIf I run for only 10 test samples, then the notebook passes and I get a score. The only difference between the Submissions are the number of test samples scored.\n\nI don't understand why I would get the \"Submission CSV Not Found\" error for the version with ~500 test samples - could this error be obfuscating some other error?",
      "votes": null
    },
    {
      "id": "3085475",
      "postDate": "01/01/2025 07:55:03",
      "content": "<p>10 of my first 15 submissions were failures due to chasing very difficult to debug notebook issues in notebooks that worked just fine until I submitted them.  There are plenty of ways for things to go wrong that make little intuitive sense. The easiest solution is to start with one of the notebooks that are known to work under the \"Code\" tab and then modify it to incorporate your own code.</p>",
      "rawMarkdown": "10 of my first 15 submissions were failures due to chasing very difficult to debug notebook issues in notebooks that worked just fine until I submitted them.  There are plenty of ways for things to go wrong that make little intuitive sense. The easiest solution is to start with one of the notebooks that are known to work under the \"Code\" tab and then modify it to incorporate your own code.",
      "votes": null
    },
    {
      "id": "3085740",
      "postDate": "01/01/2025 14:23:18",
      "content": "<p>You can try running on the training dataset first and checking if it produces the expected output in the correct format because it has more samples. </p>",
      "rawMarkdown": "You can try running on the training dataset first and checking if it produces the expected output in the correct format because it has more samples.",
      "votes": null
    },
    {
      "id": "3085855",
      "postDate": "01/01/2025 17:06:02",
      "content": "<p>Most of unexplicable submission not found were for not cleaning working folder before writing it.</p>",
      "rawMarkdown": "Most of unexplicable submission not found were for not cleaning working folder before writing it.",
      "votes": null
    },
    {
      "id": "3086300",
      "postDate": "01/02/2025 07:30:46",
      "content": "<p>It seems as though the issue is coming from one of the samples from root.runs[50:100], I now get \"Submission Scoring Error\". </p>\n<p>Runs performed - <strong>all notebooks run in less than 1 hour:</strong><br>\n✅ root.runs[0:25]<br>\n❌ root.runs[0:100] \"Submission Scoring Error\"<br>\n❌ root.runs[50:100] \"Submission Scoring Error\"<br>\n✅ root.runs[0:100] but take the first 10,000 particle picks, for example <code>submission[0:10000].to_csv(\"submission.csv\",index=False)</code> </p>\n<p>root.runs is equal to the list of tomograms, so there will be ~500 root.runs to score</p>\n<p>I have performed all checks on data, columns, etc.</p>\n<pre><code>df_predict = df_predict.drop_duplicates(subset=[, , , ])\ndf_predict = df_predict[df_predict[].isin((classes.keys()))]\ndf_predict = df_predict[[, , , , ]]\ndf_predict.insert(, , ((df_predict)))\n</code></pre>\n<p>So it has to do with scoring the actual competition metric, no direct errors are coming from my pipeline.</p>\n<p>I will continue to debug with some other tests, but I'm almost out of ideas, especially given the error occurs during the Submission Scoring - I've looked at competition metric and I can't see where there would be any errors.</p>\n<p>The next thing I will try is to filter out any NAs that might be in the submission set - I can't see anywhere as to where this would occur in my code, but its worth a shot if this could cause issues with competition metric.</p>",
      "rawMarkdown": "It seems as though the issue is coming from one of the samples from root.runs[50:100], I now get \"Submission Scoring Error\". \n\nRuns performed - **all notebooks run in less than 1 hour:**\n✅ root.runs[0:25]\n❌ root.runs[0:100] \"Submission Scoring Error\"\n❌ root.runs[50:100] \"Submission Scoring Error\"\n✅ root.runs[0:100] but take the first 10,000 particle picks, for example `submission[0:10000].to_csv(\"submission.csv\",index=False)` \n\nroot.runs is equal to the list of tomograms, so there will be ~500 root.runs to score\n\nI have performed all checks on data, columns, etc.\n```python\ndf_predict = df_predict.drop_duplicates(subset=['experiment', 'x', 'y', 'z'])\ndf_predict = df_predict[df_predict['particle_type'].isin(list(classes.keys()))]\ndf_predict = df_predict[['experiment', 'particle_type', 'x', 'y', 'z']]\ndf_predict.insert(0, 'id', range(len(df_predict)))\n```\n\nSo it has to do with scoring the actual competition metric, no direct errors are coming from my pipeline.\n\nI will continue to debug with some other tests, but I'm almost out of ideas, especially given the error occurs during the Submission Scoring - I've looked at competition metric and I can't see where there would be any errors.\n\nThe next thing I will try is to filter out any NAs that might be in the submission set - I can't see anywhere as to where this would occur in my code, but its worth a shot if this could cause issues with competition metric.",
      "votes": null
    },
    {
      "id": "3086322",
      "postDate": "01/02/2025 07:50:30",
      "content": "<p>I'm almost positive I was getting a Submission Scoring Error at one point because my code was throwing an exception before python finished flushing it's output to the disk (i.e. so there was probably a partial line in the output or something else weird like that.)  You can always put a try/catch around your code and then write everything all at once at the end.  (All things I tried.)</p>",
      "rawMarkdown": "I'm almost positive I was getting a Submission Scoring Error at one point because my code was throwing an exception before python finished flushing it's output to the disk (i.e. so there was probably a partial line in the output or something else weird like that.)  You can always put a try/catch around your code and then write everything all at once at the end.  (All things I tried.)",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3085475,
      "author_name": "davidlist",
      "author_url": "",
      "post_date": "01/01/2025 07:55:03",
      "content": "<p>10 of my first 15 submissions were failures due to chasing very difficult to debug notebook issues in notebooks that worked just fine until I submitted them.  There are plenty of ways for things to go wrong that make little intuitive sense. The easiest solution is to start with one of the notebooks that are known to work under the \"Code\" tab and then modify it to incorporate your own code.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3085740,
      "author_name": "snnclsr",
      "author_url": "",
      "post_date": "01/01/2025 14:23:18",
      "content": "<p>You can try running on the training dataset first and checking if it produces the expected output in the correct format because it has more samples. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3085855,
      "author_name": "sacuscreed",
      "author_url": "",
      "post_date": "01/01/2025 17:06:02",
      "content": "<p>Most of unexplicable submission not found were for not cleaning working folder before writing it.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 3086300,
      "author_name": "homiecal",
      "author_url": "",
      "post_date": "01/02/2025 07:30:46",
      "content": "<p>It seems as though the issue is coming from one of the samples from root.runs[50:100], I now get \"Submission Scoring Error\". </p>\n<p>Runs performed - <strong>all notebooks run in less than 1 hour:</strong><br>\n✅ root.runs[0:25]<br>\n❌ root.runs[0:100] \"Submission Scoring Error\"<br>\n❌ root.runs[50:100] \"Submission Scoring Error\"<br>\n✅ root.runs[0:100] but take the first 10,000 particle picks, for example <code>submission[0:10000].to_csv(\"submission.csv\",index=False)</code> </p>\n<p>root.runs is equal to the list of tomograms, so there will be ~500 root.runs to score</p>\n<p>I have performed all checks on data, columns, etc.</p>\n<pre><code>df_predict = df_predict.drop_duplicates(subset=[, , , ])\ndf_predict = df_predict[df_predict[].isin((classes.keys()))]\ndf_predict = df_predict[[, , , , ]]\ndf_predict.insert(, , ((df_predict)))\n</code></pre>\n<p>So it has to do with scoring the actual competition metric, no direct errors are coming from my pipeline.</p>\n<p>I will continue to debug with some other tests, but I'm almost out of ideas, especially given the error occurs during the Submission Scoring - I've looked at competition metric and I can't see where there would be any errors.</p>\n<p>The next thing I will try is to filter out any NAs that might be in the submission set - I can't see anywhere as to where this would occur in my code, but its worth a shot if this could cause issues with competition metric.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3086322,
          "author_name": "davidlist",
          "author_url": "",
          "post_date": "01/02/2025 07:50:30",
          "content": "<p>I'm almost positive I was getting a Submission Scoring Error at one point because my code was throwing an exception before python finished flushing it's output to the disk (i.e. so there was probably a partial line in the output or something else weird like that.)  You can always put a try/catch around your code and then write everything all at once at the end.  (All things I tried.)</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3085451": "I ran submission for all ~500 test samples - I get the \"Submission CSV Not Found\" error after ~4 hours of scoring time. \n\nIf I run for only 10 test samples, then the notebook passes and I get a score. The only difference between the Submissions are the number of test samples scored.\n\nI don't understand why I would get the \"Submission CSV Not Found\" error for the version with ~500 test samples - could this error be obfuscating some other error?",
    "3085475": "10 of my first 15 submissions were failures due to chasing very difficult to debug notebook issues in notebooks that worked just fine until I submitted them.  There are plenty of ways for things to go wrong that make little intuitive sense. The easiest solution is to start with one of the notebooks that are known to work under the \"Code\" tab and then modify it to incorporate your own code.",
    "3085740": "You can try running on the training dataset first and checking if it produces the expected output in the correct format because it has more samples.",
    "3085855": "Most of unexplicable submission not found were for not cleaning working folder before writing it.",
    "3086300": "It seems as though the issue is coming from one of the samples from root.runs[50:100], I now get \"Submission Scoring Error\". \n\nRuns performed - **all notebooks run in less than 1 hour:**\n✅ root.runs[0:25]\n❌ root.runs[0:100] \"Submission Scoring Error\"\n❌ root.runs[50:100] \"Submission Scoring Error\"\n✅ root.runs[0:100] but take the first 10,000 particle picks, for example `submission[0:10000].to_csv(\"submission.csv\",index=False)` \n\nroot.runs is equal to the list of tomograms, so there will be ~500 root.runs to score\n\nI have performed all checks on data, columns, etc.\n```python\ndf_predict = df_predict.drop_duplicates(subset=['experiment', 'x', 'y', 'z'])\ndf_predict = df_predict[df_predict['particle_type'].isin(list(classes.keys()))]\ndf_predict = df_predict[['experiment', 'particle_type', 'x', 'y', 'z']]\ndf_predict.insert(0, 'id', range(len(df_predict)))\n```\n\nSo it has to do with scoring the actual competition metric, no direct errors are coming from my pipeline.\n\nI will continue to debug with some other tests, but I'm almost out of ideas, especially given the error occurs during the Submission Scoring - I've looked at competition metric and I can't see where there would be any errors.\n\nThe next thing I will try is to filter out any NAs that might be in the submission set - I can't see anywhere as to where this would occur in my code, but its worth a shot if this could cause issues with competition metric.",
    "3086322": "I'm almost positive I was getting a Submission Scoring Error at one point because my code was throwing an exception before python finished flushing it's output to the disk (i.e. so there was probably a partial line in the output or something else weird like that.)  You can always put a try/catch around your code and then write everything all at once at the end.  (All things I tried.)"
  },
  "source": "meta"
}