{
  "id": 169370,
  "title": "Cannot submit my notebook",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/169370",
  "author_name": "",
  "post_date": "2020-07-23T16:57:56.693042300Z",
  "votes": 2,
  "comment_count": 27,
  "views": 0,
  "content": "<p>Hi! I'm trying to submit my solution but I'm getting an error. I followed the instructions in these threads (<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166533\">first</a>, <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168276\">second</a>), namely:\n- make sure you iterate over all patients of the test set\n- make sure you cover all weeks from -12 to 133 inclusively\n- turn the internet off in the notebook\n- make sure the name of the submission is \"submission.csv\"  </p>\n\n<p>I also made sure all my column names and formats were OK. Yet, I still can't submit...</p>\n\n<p>Here's the execution info:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa155866847b3d2b75fa6600577a2d717%2FScreenshot%20from%202020-07-23%2012-46-13.png?generation=1595523261224339&amp;alt=media\" alt=\"\"></p>",
  "messages": [
    {
      "id": "942245",
      "postDate": "07/23/2020 16:57:56",
      "content": "<p>Hi! I'm trying to submit my solution but I'm getting an error. I followed the instructions in these threads (<a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166533\">first</a>, <a href=\"https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168276\">second</a>), namely:\n- make sure you iterate over all patients of the test set\n- make sure you cover all weeks from -12 to 133 inclusively\n- turn the internet off in the notebook\n- make sure the name of the submission is \"submission.csv\"  </p>\n\n<p>I also made sure all my column names and formats were OK. Yet, I still can't submit...</p>\n\n<p>Here's the execution info:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa155866847b3d2b75fa6600577a2d717%2FScreenshot%20from%202020-07-23%2012-46-13.png?generation=1595523261224339&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hi! I'm trying to submit my solution but I'm getting an error. I followed the instructions in these threads ([first](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166533), [second](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168276)), namely:\n- make sure you iterate over all patients of the test set\n- make sure you cover all weeks from -12 to 133 inclusively\n- turn the internet off in the notebook\n- make sure the name of the submission is \"submission.csv\"  \n\nI also made sure all my column names and formats were OK. Yet, I still can't submit...\n\nHere's the execution info:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa155866847b3d2b75fa6600577a2d717%2FScreenshot%20from%202020-07-23%2012-46-13.png?generation=1595523261224339&amp;alt=media)",
      "votes": null
    },
    {
      "id": "942433",
      "postDate": "07/23/2020 18:34:14",
      "content": "<p>The Status columns in my Submissions tab says \"Submission CSV Not Found\"... But when I look at the actual notebook, the Output section does contain the CSV...</p>",
      "rawMarkdown": "The Status columns in my Submissions tab says \"Submission CSV Not Found\"... But when I look at the actual notebook, the Output section does contain the CSV...",
      "votes": null
    },
    {
      "id": "942533",
      "postDate": "07/23/2020 19:52:43",
      "content": "<p>This is an unfortunate issue. I would suggest perusing some of the suggested solutions to the same issue encountered by another user in <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/152425\">this thread</a>. Hope this helps -- good luck!</p>",
      "rawMarkdown": "This is an unfortunate issue. I would suggest perusing some of the suggested solutions to the same issue encountered by another user in [this thread](https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/152425). Hope this helps -- good luck!",
      "votes": null
    },
    {
      "id": "942544",
      "postDate": "07/23/2020 19:59:47",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> You might need to verify that the submission file you are producing is named <code>submission.csv</code></p>",
      "rawMarkdown": "archimedus You might need to verify that the submission file you are producing is named `submission.csv`",
      "votes": null
    },
    {
      "id": "942556",
      "postDate": "07/23/2020 20:13:17",
      "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I rechecked and that's exactly how my file is named. Do we need to specify a certain directory when we output? I was using <code>output.to_csv(\"submission.csv\", index=False)</code> but using Kaggle's \"Copy file path\" function next to <em>output &gt; kaggle/working</em> in the right panel's \"Data\" tab seems to suggest I add <code>../input/output/</code> as a prefix to the file name. Does that seem right?</p>",
      "rawMarkdown": "juliaelliott I rechecked and that's exactly how my file is named. Do we need to specify a certain directory when we output? I was using `output.to_csv(\"submission.csv\", index=False)` but using Kaggle's \"Copy file path\" function next to *output &gt; kaggle/working* in the right panel's \"Data\" tab seems to suggest I add `../input/output/` as a prefix to the file name. Does that seem right?",
      "votes": null
    },
    {
      "id": "942558",
      "postDate": "07/23/2020 20:14:43",
      "content": "<p>Thanks <a href=\"/roshanr11\">@roshanr11</a>, I will look into these suggestions!</p>",
      "rawMarkdown": "Thanks @roshanr11, I will look into these suggestions!",
      "votes": null
    },
    {
      "id": "942584",
      "postDate": "07/23/2020 20:57:24",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> At the top of your notebook in the viewer mode (not editing), does it say \"Version X of Y\"?</p>\n\n<p>It's very possible that the latest version has actually failed for a different reason and you're seeing a successfully run version in the viewer (by default it shows you the latest successful one).</p>",
      "rawMarkdown": "archimedus At the top of your notebook in the viewer mode (not editing), does it say \"Version X of Y\"?\n\nIt's very possible that the latest version has actually failed for a different reason and you're seeing a successfully run version in the viewer (by default it shows you the latest successful one).",
      "votes": null
    },
    {
      "id": "942615",
      "postDate": "07/23/2020 21:45:29",
      "content": "<p><a href=\"/herbison\">@herbison</a> actually, there is a discrepancy between the number of versions that seem to exist for the same notebook when looking at the \"My Submissions\" tab (currently 27) and when looking at the viewer/edit mode (currently 32). Seems a bit odd, maybe the issue has something to do with that?</p>",
      "rawMarkdown": "herbison actually, there is a discrepancy between the number of versions that seem to exist for the same notebook when looking at the \"My Submissions\" tab (currently 27) and when looking at the viewer/edit mode (currently 32). Seems a bit odd, maybe the issue has something to do with that?",
      "votes": null
    },
    {
      "id": "943085",
      "postDate": "07/24/2020 06:13:00",
      "content": "<p>Hi <a href=\"/herbison\">@herbison</a>,  I aslo meet the unknown error \"Notebook Exceeded Allowed Compute\". I'm sure that the kernel (the committed version works fine) running time is about 900s without GPU and internet is off as well. </p>",
      "rawMarkdown": "Hi @herbison,  I aslo meet the unknown error \"Notebook Exceeded Allowed Compute\". I'm sure that the kernel (the committed version works fine) running time is about 900s without GPU and internet is off as well.",
      "votes": null
    },
    {
      "id": "943805",
      "postDate": "07/24/2020 15:39:21",
      "content": "<p><a href=\"/dxchen\">@dxchen</a> That means the rerun of your notebook on private data probably ran out of memory, I wouldn't be surprised if there is more private data or something.</p>\n\n<p><a href=\"/archimedus\">@archimedus</a> It does look like your commit version was also missing submission.csv, please make sure you're looking at the correct version (goto the viewer tab, click on Versions, click on the latest one).</p>",
      "rawMarkdown": "dxchen That means the rerun of your notebook on private data probably ran out of memory, I wouldn't be surprised if there is more private data or something.\n\n@archimedus It does look like your commit version was also missing submission.csv, please make sure you're looking at the correct version (goto the viewer tab, click on Versions, click on the latest one).",
      "votes": null
    },
    {
      "id": "943926",
      "postDate": "07/24/2020 17:18:03",
      "content": "<p><a href=\"/herbison\">@herbison</a> My Notebook produces the submission.csv file, but it looks like the summary somehow doesn't realize it:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F366c89889034fed690f1031d2cffb748%2FScreenshot%20from%202020-07-24%2013-04-28.png?generation=1595610616477857&amp;alt=media\" alt=\"\"></p>\n\n<p>Scrolling down...\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa4111e1fb5a7d289e43aa821a4e2680d%2FScreenshot%20from%202020-07-24%2013-13-32.png?generation=1595610925014576&amp;alt=media\" alt=\"\"></p>\n\n<hr>\n\n<p>There also seems to be a discrepancy between what version numbers are available. Maybe this offers a clue to the source of the problem:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F69fccb544beeda38ead1a18ecc9e98f1%2FScreenshot%20from%202020-07-24%2013-05-33.png?generation=1595610717867342&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "herbison My Notebook produces the submission.csv file, but it looks like the summary somehow doesn't realize it:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F366c89889034fed690f1031d2cffb748%2FScreenshot%20from%202020-07-24%2013-04-28.png?generation=1595610616477857&amp;alt=media)\n\nScrolling down...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa4111e1fb5a7d289e43aa821a4e2680d%2FScreenshot%20from%202020-07-24%2013-13-32.png?generation=1595610925014576&amp;alt=media)\n\n\n---\nThere also seems to be a discrepancy between what version numbers are available. Maybe this offers a clue to the source of the problem:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F69fccb544beeda38ead1a18ecc9e98f1%2FScreenshot%20from%202020-07-24%2013-05-33.png?generation=1595610717867342&amp;alt=media)",
      "votes": null
    },
    {
      "id": "943929",
      "postDate": "07/24/2020 17:29:04",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> this has happened to me in previous competitions where i get the error <strong>Submission CSV Not Found</strong>. </p>\n\n<p>The problem is your submission kernel is crashing while predicting on the hidden test data used to calculate the public LB score and the kernel stops before the csv file is created to calculate the score. It is usually when GPU memory allocation exceeds the limit or due to a small syntax error you may have missed.</p>\n\n<p>The  <strong>Notebook Exceeded Allowed Compute</strong> error will be shown if your submission kernel exceeds available RAM, but GPU memory exceeding does not usually trigger this error.</p>",
      "rawMarkdown": "archimedus this has happened to me in previous competitions where i get the error **Submission CSV Not Found**. \n\nThe problem is your submission kernel is crashing while predicting on the hidden test data used to calculate the public LB score and the kernel stops before the csv file is created to calculate the score. It is usually when GPU memory allocation exceeds the limit or due to a small syntax error you may have missed.\n\nThe  **Notebook Exceeded Allowed Compute** error will be shown if your submission kernel exceeds available RAM, but GPU memory exceeding does not usually trigger this error.",
      "votes": null
    },
    {
      "id": "943930",
      "postDate": "07/24/2020 17:30:51",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> Correction, I checked and your recent commit does have the submission.csv, but the private rerun hits an exception in your code just before writing it out, please review your code for possible assumptions about the private data files.</p>",
      "rawMarkdown": "archimedus Correction, I checked and your recent commit does have the submission.csv, but the private rerun hits an exception in your code just before writing it out, please review your code for possible assumptions about the private data files.",
      "votes": null
    },
    {
      "id": "943986",
      "postDate": "07/24/2020 18:17:00",
      "content": "<p><a href=\"/herbison\">@herbison</a> I'm not sure what sort of assumptions might be relevant (I'm a relative novice at all this). To the best of my knowledge I'm not doing anything really special with the files <em>per se</em>.</p>\n\n<p>Are you talking about file paths? For all inputs I have the base path <code>\"../input/osic-pulmonary-fibrosis-progression/\"</code>; my output has no prefix, I just use <code>output.to_csv(\"submission.csv\", index=False)</code> directly. Apart from that:\n- I'm not using any input files other than the default package provided for the competition.\n- I use pydicom.dcmread() (in a <em>with</em> block) to read the image files\n- I use Pandas' read_csv() function to read the input CSV's</p>\n\n<p>Is there some line of code I could insert in my notebook to throw up a message about the sort of exceptions you're thinking of?</p>\n\n<p>Unless perhaps this is what you're referring to?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F0e03793e87b97d14edc2b9f8f0e62d5f%2FScreenshot%20from%202020-07-24%2014-15-30.png?generation=1595614573212920&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "herbison I'm not sure what sort of assumptions might be relevant (I'm a relative novice at all this). To the best of my knowledge I'm not doing anything really special with the files *per se*.\n\nAre you talking about file paths? For all inputs I have the base path `\"../input/osic-pulmonary-fibrosis-progression/\"`; my output has no prefix, I just use `output.to_csv(\"submission.csv\", index=False)` directly. Apart from that:\n- I'm not using any input files other than the default package provided for the competition.\n- I use pydicom.dcmread() (in a *with* block) to read the image files\n- I use Pandas' read_csv() function to read the input CSV's\n\nIs there some line of code I could insert in my notebook to throw up a message about the sort of exceptions you're thinking of?\n\nUnless perhaps this is what you're referring to?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F0e03793e87b97d14edc2b9f8f0e62d5f%2FScreenshot%20from%202020-07-24%2014-15-30.png?generation=1595614573212920&amp;alt=media)",
      "votes": null
    },
    {
      "id": "943995",
      "postDate": "07/24/2020 18:21:56",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> Make sure that every file you try to read actually exists.</p>",
      "rawMarkdown": "archimedus Make sure that every file you try to read actually exists.",
      "votes": null
    },
    {
      "id": "944184",
      "postDate": "07/24/2020 22:42:50",
      "content": "<p><a href=\"/herbison\">@herbison</a> All the files my code tries to access should exist.\n- CSV's don't seem to cause any trouble.\n- For image files my code loops through lists of files obtained using the os library. \n- The file paths are generated partly with a list of patient ID's, but I ran a version of my notebook that printed a line every time a new path was tried to say whether or not it's valid, and all results were \"True\".</p>\n\n<p>I'm currently committing a version in which I commented out a <code>del</code> statement (which follows immediately some file-reading <code>with</code> blocks - I was having trouble with memory usage so I figured I'd try doing some garbage-collection explicitly). Who knows, maybe that'll work.</p>\n\n<p>Another potential source of the problem: when you're talking about private data files, do you mean the testing files that are used to score my model but which I won't see on my end? If so, my code relies on the DCM files' <code>.pixel_array</code>, <code>.Rows</code>, <code>.Columns</code>, <code>.SliceLocation</code> attributes. If too many of a patient's image files lack these attributes, my code just skips those examples. It doesn't cause too much trouble on the training end (apart from wasting a few training examples) and none on the visible test set, but if the same issues surface in the private data files maybe my code ends up not making predictions for one or more test patient, and rather than the submission checker saying that my CSV file is incomplete (because of missing predictions) it just says that my CSV file isn't there.</p>",
      "rawMarkdown": "herbison All the files my code tries to access should exist.\n- CSV's don't seem to cause any trouble.\n- For image files my code loops through lists of files obtained using the os library. \n- The file paths are generated partly with a list of patient ID's, but I ran a version of my notebook that printed a line every time a new path was tried to say whether or not it's valid, and all results were \"True\".\n\nI'm currently committing a version in which I commented out a `del` statement (which follows immediately some file-reading `with` blocks - I was having trouble with memory usage so I figured I'd try doing some garbage-collection explicitly). Who knows, maybe that'll work.\n\nAnother potential source of the problem: when you're talking about private data files, do you mean the testing files that are used to score my model but which I won't see on my end? If so, my code relies on the DCM files' `.pixel_array`, `.Rows`, `.Columns`, `.SliceLocation` attributes. If too many of a patient's image files lack these attributes, my code just skips those examples. It doesn't cause too much trouble on the training end (apart from wasting a few training examples) and none on the visible test set, but if the same issues surface in the private data files maybe my code ends up not making predictions for one or more test patient, and rather than the submission checker saying that my CSV file is incomplete (because of missing predictions) it just says that my CSV file isn't there.",
      "votes": null
    },
    {
      "id": "944189",
      "postDate": "07/24/2020 22:54:05",
      "content": "<p>have you posted a link to your notebook so we can try and help you find your error? Are you looking for dicom files with hard-coded names (like dcm.1, which might not always exist in the test set?)</p>",
      "rawMarkdown": "have you posted a link to your notebook so we can try and help you find your error? Are you looking for dicom files with hard-coded names (like dcm.1, which might not always exist in the test set?)",
      "votes": null
    },
    {
      "id": "944260",
      "postDate": "07/25/2020 01:08:56",
      "content": "<p>getting similar error :(</p>\n\n<p>I'm not sure the log you get when you run your notebook is relevant to the error when you submit your csv file.  I have notebooks that show the same error log after running the notebook but when I submit I get a leaderboard score just fine.  </p>\n\n<p>Unfortunately we don't get a log when we submit for leaderboard scoring, so the only relevant information is \"Submission CSV Not Found\" which makes no sense when the notebook ran without error.</p>\n\n<p>I might just move on, come back to it later and try to run it again (i.e. unplug and re-plug the router so to speak lol).  Might try the copy and paste the code into a new notebook idea as well, maybe clean up the code as best I can.</p>",
      "rawMarkdown": "getting similar error :(\n\nI'm not sure the log you get when you run your notebook is relevant to the error when you submit your csv file.  I have notebooks that show the same error log after running the notebook but when I submit I get a leaderboard score just fine.  \n\nUnfortunately we don't get a log when we submit for leaderboard scoring, so the only relevant information is \"Submission CSV Not Found\" which makes no sense when the notebook ran without error.\n\nI might just move on, come back to it later and try to run it again (i.e. unplug and re-plug the router so to speak lol).  Might try the copy and paste the code into a new notebook idea as well, maybe clean up the code as best I can.",
      "votes": null
    },
    {
      "id": "946817",
      "postDate": "07/26/2020 21:30:27",
      "content": "<p>So I found a solution to my problem, but it doesn't make much sense.  Probably something very particular to this competition, but this notebook\\ person (<a href=\"https://www.kaggle.com/titericz/tabular-simple-eda-linear-model\">https://www.kaggle.com/titericz/tabular-simple-eda-linear-model</a>) merged the test data with the training data and then used that to merge with the submission file, and that worked for me.  I kept all of the same features, but that data pipeline worked.  No idea why, but I'll take it I guess.</p>",
      "rawMarkdown": "So I found a solution to my problem, but it doesn't make much sense.  Probably something very particular to this competition, but this notebook\\ person (https://www.kaggle.com/titericz/tabular-simple-eda-linear-model) merged the test data with the training data and then used that to merge with the submission file, and that worked for me.  I kept all of the same features, but that data pipeline worked.  No idea why, but I'll take it I guess.",
      "votes": null
    },
    {
      "id": "946875",
      "postDate": "07/26/2020 23:12:53",
      "content": "<p>FYI, I tried replacing my use of the <code>.SliceLocation</code> attributes with <code>.SliceThickness</code> and <code>.InstanceNumber</code> (after verifying that every single image file in the train set that doesn't require the GDCM library to open has these), but it didn't change anything. Still \"Submission CSV Not Found\"...</p>",
      "rawMarkdown": "FYI, I tried replacing my use of the `.SliceLocation` attributes with `.SliceThickness` and `.InstanceNumber` (after verifying that every single image file in the train set that doesn't require the GDCM library to open has these), but it didn't change anything. Still \"Submission CSV Not Found\"...",
      "votes": null
    },
    {
      "id": "948147",
      "postDate": "07/27/2020 17:32:48",
      "content": "<p>Hello everyone,\nI have the same problem and I try to solve it but unfortunately without success. I merged the test data with the training data and then used that to merge with the submission file, but I got the same error: \"Submission CSV Not Found\".\nBelow is the log file from my notebook.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2651354%2F8d58d85b0ee01271e2135efd47893c63%2Fkaggle_log.png?generation=1595871061983594&amp;alt=media\" alt=\"\"></p>",
      "rawMarkdown": "Hello everyone,\nI have the same problem and I try to solve it but unfortunately without success. I merged the test data with the training data and then used that to merge with the submission file, but I got the same error: \"Submission CSV Not Found\".\nBelow is the log file from my notebook.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2651354%2F8d58d85b0ee01271e2135efd47893c63%2Fkaggle_log.png?generation=1595871061983594&amp;alt=media)",
      "votes": null
    },
    {
      "id": "948303",
      "postDate": "07/27/2020 19:38:39",
      "content": "<p>At this point I've decided to go for a new approach: I created an ultra-simple notebook that outputs what are essentially dummy \"predictions\", and succeeded in submitting it. Now I will add chunks of my code one by one, and see at what point it stops submitting........</p>",
      "rawMarkdown": "At this point I've decided to go for a new approach: I created an ultra-simple notebook that outputs what are essentially dummy \"predictions\", and succeeded in submitting it. Now I will add chunks of my code one by one, and see at what point it stops submitting........",
      "votes": null
    },
    {
      "id": "949326",
      "postDate": "07/28/2020 15:04:26",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> You're previous submission was failing trying to fetch a file that did not exist, that's why I hinted at that. When a cell hits an exception, it doesn't continue processing that cell, which means it doesn't write out submission.csv if that's done later in that same cell, and leads to you seeing a 'no submission.csv' error.</p>\n\n<p>Your new submission has another, different code bug causing a similar issue (exception thrown in cell before submission.csv is written out).</p>",
      "rawMarkdown": "archimedus You're previous submission was failing trying to fetch a file that did not exist, that's why I hinted at that. When a cell hits an exception, it doesn't continue processing that cell, which means it doesn't write out submission.csv if that's done later in that same cell, and leads to you seeing a 'no submission.csv' error.\n\nYour new submission has another, different code bug causing a similar issue (exception thrown in cell before submission.csv is written out).",
      "votes": null
    },
    {
      "id": "949333",
      "postDate": "07/28/2020 15:07:21",
      "content": "<p><a href=\"/yeayates21\">@yeayates21</a> We sadly can't provide much feedback on submissions (like an output log) because it could be used to leak private data. I agree it's rather painful but it's one of the few ways we can mitigate the potential for abusing the system to cheat.</p>",
      "rawMarkdown": "yeayates21 We sadly can't provide much feedback on submissions (like an output log) because it could be used to leak private data. I agree it's rather painful but it's one of the few ways we can mitigate the potential for abusing the system to cheat.",
      "votes": null
    },
    {
      "id": "949421",
      "postDate": "07/28/2020 16:18:22",
      "content": "<p>Thank you, that information will be useful in narrowing down where the problem lies.</p>",
      "rawMarkdown": "Thank you, that information will be useful in narrowing down where the problem lies.",
      "votes": null
    },
    {
      "id": "949513",
      "postDate": "07/28/2020 17:20:25",
      "content": "<p><a href=\"/herbison\">@herbison</a> Perhaps a \"sanitized\" log could be quite a useful feature, something that tells the submitter basic info such as the cell number/line number that trips up the private run and/or the general type of error (file missing/unreadable, wrong shapes/lengths, arithmetic, CSV result diverging from specifications, etc.), without providing <em>all</em> the runtime info (just a few \"white-listed\" items).</p>\n\n<p>I agree that it's important to keep some opacity in the system to deter would-be cheaters, but there has to be a grey area between full logs and no info at all!</p>",
      "rawMarkdown": "herbison Perhaps a \"sanitized\" log could be quite a useful feature, something that tells the submitter basic info such as the cell number/line number that trips up the private run and/or the general type of error (file missing/unreadable, wrong shapes/lengths, arithmetic, CSV result diverging from specifications, etc.), without providing *all* the runtime info (just a few \"white-listed\" items).\n\nI agree that it's important to keep some opacity in the system to deter would-be cheaters, but there has to be a grey area between full logs and no info at all!",
      "votes": null
    },
    {
      "id": "949585",
      "postDate": "07/28/2020 18:17:42",
      "content": "<p>Another possible improvement would be to make failed submissions not count towards the daily total. Daily submission limits make sense if the goal is to prevent people from abusing their public scores as a validation score to fine-tune their models (even though this would be good strategy to get an overfitting model, since the scores are based on only 15% of the test data), but when you're trying to debug an issue that happens in the hidden run (while the run on your end seems to work perfectly) and you have to rely on imprecise trial-and-error (since you have to run the entire notebook - twice! - every time you want to test a hypothesis regarding the location of the error, and get no feedback other than a pass/fail), making the fails count toward the daily submission limit just adds one more (seemingly useless) hurdle...</p>",
      "rawMarkdown": "Another possible improvement would be to make failed submissions not count towards the daily total. Daily submission limits make sense if the goal is to prevent people from abusing their public scores as a validation score to fine-tune their models (even though this would be good strategy to get an overfitting model, since the scores are based on only 15% of the test data), but when you're trying to debug an issue that happens in the hidden run (while the run on your end seems to work perfectly) and you have to rely on imprecise trial-and-error (since you have to run the entire notebook - twice! - every time you want to test a hypothesis regarding the location of the error, and get no feedback other than a pass/fail), making the fails count toward the daily submission limit just adds one more (seemingly useless) hurdle...",
      "votes": null
    },
    {
      "id": "949616",
      "postDate": "07/28/2020 18:43:50",
      "content": "<p><a href=\"/archimedus\">@archimedus</a> Basically any information we send back from the result (different errors, content of those errors, logs etc.) could be used as bits in an information leaking system, so we really have to limit how much feedback exists.</p>\n\n<p>The failed submissions count towards totals both because they could be part of a leaking system, and because Kaggle has to control compute usage. Every run of every notebook and all the resources it requires costs something, and we want to make sure every user is given their fair chance, and so individual users don't consume all available compute.</p>\n\n<p>I do understand your frustration with limited submits though. When I was in university I submitted my homework to a very similar system, with a max submits per day. What the university advised students (and I pass on to Kagglers), is to:\n- submit as early as you can to make the most of your limited number of refreshes\n- spend more time critically looking at your code to understand where things can go wrong, and carefully handle them\nRelying on the evaluation system to verify your code works won't be very helpful in production system where you'll need to be able to think ahead of possible issues and protect against them.</p>",
      "rawMarkdown": "archimedus Basically any information we send back from the result (different errors, content of those errors, logs etc.) could be used as bits in an information leaking system, so we really have to limit how much feedback exists.\n\nThe failed submissions count towards totals both because they could be part of a leaking system, and because Kaggle has to control compute usage. Every run of every notebook and all the resources it requires costs something, and we want to make sure every user is given their fair chance, and so individual users don't consume all available compute.\n\nI do understand your frustration with limited submits though. When I was in university I submitted my homework to a very similar system, with a max submits per day. What the university advised students (and I pass on to Kagglers), is to:\n- submit as early as you can to make the most of your limited number of refreshes\n- spend more time critically looking at your code to understand where things can go wrong, and carefully handle them\nRelying on the evaluation system to verify your code works won't be very helpful in production system where you'll need to be able to think ahead of possible issues and protect against them.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 942433,
      "author_name": "archimedus",
      "author_url": "",
      "post_date": "07/23/2020 18:34:14",
      "content": "<p>The Status columns in my Submissions tab says \"Submission CSV Not Found\"... But when I look at the actual notebook, the Output section does contain the CSV...</p>",
      "votes": null,
      "replies": [
        {
          "id": 942544,
          "author_name": "juliaelliott",
          "author_url": "",
          "post_date": "07/23/2020 19:59:47",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> You might need to verify that the submission file you are producing is named <code>submission.csv</code></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 942556,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/23/2020 20:13:17",
          "content": "<p><a href=\"/juliaelliott\">@juliaelliott</a> I rechecked and that's exactly how my file is named. Do we need to specify a certain directory when we output? I was using <code>output.to_csv(\"submission.csv\", index=False)</code> but using Kaggle's \"Copy file path\" function next to <em>output &gt; kaggle/working</em> in the right panel's \"Data\" tab seems to suggest I add <code>../input/output/</code> as a prefix to the file name. Does that seem right?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 942584,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/23/2020 20:57:24",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> At the top of your notebook in the viewer mode (not editing), does it say \"Version X of Y\"?</p>\n\n<p>It's very possible that the latest version has actually failed for a different reason and you're seeing a successfully run version in the viewer (by default it shows you the latest successful one).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 942615,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/23/2020 21:45:29",
          "content": "<p><a href=\"/herbison\">@herbison</a> actually, there is a discrepancy between the number of versions that seem to exist for the same notebook when looking at the \"My Submissions\" tab (currently 27) and when looking at the viewer/edit mode (currently 32). Seems a bit odd, maybe the issue has something to do with that?</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943085,
          "author_name": "dxchen",
          "author_url": "",
          "post_date": "07/24/2020 06:13:00",
          "content": "<p>Hi <a href=\"/herbison\">@herbison</a>,  I aslo meet the unknown error \"Notebook Exceeded Allowed Compute\". I'm sure that the kernel (the committed version works fine) running time is about 900s without GPU and internet is off as well. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943805,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/24/2020 15:39:21",
          "content": "<p><a href=\"/dxchen\">@dxchen</a> That means the rerun of your notebook on private data probably ran out of memory, I wouldn't be surprised if there is more private data or something.</p>\n\n<p><a href=\"/archimedus\">@archimedus</a> It does look like your commit version was also missing submission.csv, please make sure you're looking at the correct version (goto the viewer tab, click on Versions, click on the latest one).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943926,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/24/2020 17:18:03",
          "content": "<p><a href=\"/herbison\">@herbison</a> My Notebook produces the submission.csv file, but it looks like the summary somehow doesn't realize it:</p>\n\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F366c89889034fed690f1031d2cffb748%2FScreenshot%20from%202020-07-24%2013-04-28.png?generation=1595610616477857&amp;alt=media\" alt=\"\"></p>\n\n<p>Scrolling down...\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa4111e1fb5a7d289e43aa821a4e2680d%2FScreenshot%20from%202020-07-24%2013-13-32.png?generation=1595610925014576&amp;alt=media\" alt=\"\"></p>\n\n<hr>\n\n<p>There also seems to be a discrepancy between what version numbers are available. Maybe this offers a clue to the source of the problem:\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F69fccb544beeda38ead1a18ecc9e98f1%2FScreenshot%20from%202020-07-24%2013-05-33.png?generation=1595610717867342&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": [
            {
              "id": 943929,
              "author_name": "yovinyahathugoda",
              "author_url": "",
              "post_date": "07/24/2020 17:29:04",
              "content": "<p><a href=\"/archimedus\">@archimedus</a> this has happened to me in previous competitions where i get the error <strong>Submission CSV Not Found</strong>. </p>\n\n<p>The problem is your submission kernel is crashing while predicting on the hidden test data used to calculate the public LB score and the kernel stops before the csv file is created to calculate the score. It is usually when GPU memory allocation exceeds the limit or due to a small syntax error you may have missed.</p>\n\n<p>The  <strong>Notebook Exceeded Allowed Compute</strong> error will be shown if your submission kernel exceeds available RAM, but GPU memory exceeding does not usually trigger this error.</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 943930,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/24/2020 17:30:51",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> Correction, I checked and your recent commit does have the submission.csv, but the private rerun hits an exception in your code just before writing it out, please review your code for possible assumptions about the private data files.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943986,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/24/2020 18:17:00",
          "content": "<p><a href=\"/herbison\">@herbison</a> I'm not sure what sort of assumptions might be relevant (I'm a relative novice at all this). To the best of my knowledge I'm not doing anything really special with the files <em>per se</em>.</p>\n\n<p>Are you talking about file paths? For all inputs I have the base path <code>\"../input/osic-pulmonary-fibrosis-progression/\"</code>; my output has no prefix, I just use <code>output.to_csv(\"submission.csv\", index=False)</code> directly. Apart from that:\n- I'm not using any input files other than the default package provided for the competition.\n- I use pydicom.dcmread() (in a <em>with</em> block) to read the image files\n- I use Pandas' read_csv() function to read the input CSV's</p>\n\n<p>Is there some line of code I could insert in my notebook to throw up a message about the sort of exceptions you're thinking of?</p>\n\n<p>Unless perhaps this is what you're referring to?\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F0e03793e87b97d14edc2b9f8f0e62d5f%2FScreenshot%20from%202020-07-24%2014-15-30.png?generation=1595614573212920&amp;alt=media\" alt=\"\"></p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 943995,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/24/2020 18:21:56",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> Make sure that every file you try to read actually exists.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 944184,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/24/2020 22:42:50",
          "content": "<p><a href=\"/herbison\">@herbison</a> All the files my code tries to access should exist.\n- CSV's don't seem to cause any trouble.\n- For image files my code loops through lists of files obtained using the os library. \n- The file paths are generated partly with a list of patient ID's, but I ran a version of my notebook that printed a line every time a new path was tried to say whether or not it's valid, and all results were \"True\".</p>\n\n<p>I'm currently committing a version in which I commented out a <code>del</code> statement (which follows immediately some file-reading <code>with</code> blocks - I was having trouble with memory usage so I figured I'd try doing some garbage-collection explicitly). Who knows, maybe that'll work.</p>\n\n<p>Another potential source of the problem: when you're talking about private data files, do you mean the testing files that are used to score my model but which I won't see on my end? If so, my code relies on the DCM files' <code>.pixel_array</code>, <code>.Rows</code>, <code>.Columns</code>, <code>.SliceLocation</code> attributes. If too many of a patient's image files lack these attributes, my code just skips those examples. It doesn't cause too much trouble on the training end (apart from wasting a few training examples) and none on the visible test set, but if the same issues surface in the private data files maybe my code ends up not making predictions for one or more test patient, and rather than the submission checker saying that my CSV file is incomplete (because of missing predictions) it just says that my CSV file isn't there.</p>",
          "votes": null,
          "replies": [
            {
              "id": 944189,
              "author_name": "richardepstein",
              "author_url": "",
              "post_date": "07/24/2020 22:54:05",
              "content": "<p>have you posted a link to your notebook so we can try and help you find your error? Are you looking for dicom files with hard-coded names (like dcm.1, which might not always exist in the test set?)</p>",
              "votes": null,
              "replies": []
            }
          ]
        },
        {
          "id": 946875,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/26/2020 23:12:53",
          "content": "<p>FYI, I tried replacing my use of the <code>.SliceLocation</code> attributes with <code>.SliceThickness</code> and <code>.InstanceNumber</code> (after verifying that every single image file in the train set that doesn't require the GDCM library to open has these), but it didn't change anything. Still \"Submission CSV Not Found\"...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949326,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/28/2020 15:04:26",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> You're previous submission was failing trying to fetch a file that did not exist, that's why I hinted at that. When a cell hits an exception, it doesn't continue processing that cell, which means it doesn't write out submission.csv if that's done later in that same cell, and leads to you seeing a 'no submission.csv' error.</p>\n\n<p>Your new submission has another, different code bug causing a similar issue (exception thrown in cell before submission.csv is written out).</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949421,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/28/2020 16:18:22",
          "content": "<p>Thank you, that information will be useful in narrowing down where the problem lies.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 942533,
      "author_name": "roshanr11",
      "author_url": "",
      "post_date": "07/23/2020 19:52:43",
      "content": "<p>This is an unfortunate issue. I would suggest perusing some of the suggested solutions to the same issue encountered by another user in <a href=\"https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/152425\">this thread</a>. Hope this helps -- good luck!</p>",
      "votes": null,
      "replies": [
        {
          "id": 942558,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/23/2020 20:14:43",
          "content": "<p>Thanks <a href=\"/roshanr11\">@roshanr11</a>, I will look into these suggestions!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 944260,
      "author_name": "yeayates21",
      "author_url": "",
      "post_date": "07/25/2020 01:08:56",
      "content": "<p>getting similar error :(</p>\n\n<p>I'm not sure the log you get when you run your notebook is relevant to the error when you submit your csv file.  I have notebooks that show the same error log after running the notebook but when I submit I get a leaderboard score just fine.  </p>\n\n<p>Unfortunately we don't get a log when we submit for leaderboard scoring, so the only relevant information is \"Submission CSV Not Found\" which makes no sense when the notebook ran without error.</p>\n\n<p>I might just move on, come back to it later and try to run it again (i.e. unplug and re-plug the router so to speak lol).  Might try the copy and paste the code into a new notebook idea as well, maybe clean up the code as best I can.</p>",
      "votes": null,
      "replies": [
        {
          "id": 946817,
          "author_name": "yeayates21",
          "author_url": "",
          "post_date": "07/26/2020 21:30:27",
          "content": "<p>So I found a solution to my problem, but it doesn't make much sense.  Probably something very particular to this competition, but this notebook\\ person (<a href=\"https://www.kaggle.com/titericz/tabular-simple-eda-linear-model\">https://www.kaggle.com/titericz/tabular-simple-eda-linear-model</a>) merged the test data with the training data and then used that to merge with the submission file, and that worked for me.  I kept all of the same features, but that data pipeline worked.  No idea why, but I'll take it I guess.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 948303,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/27/2020 19:38:39",
          "content": "<p>At this point I've decided to go for a new approach: I created an ultra-simple notebook that outputs what are essentially dummy \"predictions\", and succeeded in submitting it. Now I will add chunks of my code one by one, and see at what point it stops submitting........</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949333,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/28/2020 15:07:21",
          "content": "<p><a href=\"/yeayates21\">@yeayates21</a> We sadly can't provide much feedback on submissions (like an output log) because it could be used to leak private data. I agree it's rather painful but it's one of the few ways we can mitigate the potential for abusing the system to cheat.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949513,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/28/2020 17:20:25",
          "content": "<p><a href=\"/herbison\">@herbison</a> Perhaps a \"sanitized\" log could be quite a useful feature, something that tells the submitter basic info such as the cell number/line number that trips up the private run and/or the general type of error (file missing/unreadable, wrong shapes/lengths, arithmetic, CSV result diverging from specifications, etc.), without providing <em>all</em> the runtime info (just a few \"white-listed\" items).</p>\n\n<p>I agree that it's important to keep some opacity in the system to deter would-be cheaters, but there has to be a grey area between full logs and no info at all!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949585,
          "author_name": "archimedus",
          "author_url": "",
          "post_date": "07/28/2020 18:17:42",
          "content": "<p>Another possible improvement would be to make failed submissions not count towards the daily total. Daily submission limits make sense if the goal is to prevent people from abusing their public scores as a validation score to fine-tune their models (even though this would be good strategy to get an overfitting model, since the scores are based on only 15% of the test data), but when you're trying to debug an issue that happens in the hidden run (while the run on your end seems to work perfectly) and you have to rely on imprecise trial-and-error (since you have to run the entire notebook - twice! - every time you want to test a hypothesis regarding the location of the error, and get no feedback other than a pass/fail), making the fails count toward the daily submission limit just adds one more (seemingly useless) hurdle...</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 949616,
          "author_name": "herbison",
          "author_url": "",
          "post_date": "07/28/2020 18:43:50",
          "content": "<p><a href=\"/archimedus\">@archimedus</a> Basically any information we send back from the result (different errors, content of those errors, logs etc.) could be used as bits in an information leaking system, so we really have to limit how much feedback exists.</p>\n\n<p>The failed submissions count towards totals both because they could be part of a leaking system, and because Kaggle has to control compute usage. Every run of every notebook and all the resources it requires costs something, and we want to make sure every user is given their fair chance, and so individual users don't consume all available compute.</p>\n\n<p>I do understand your frustration with limited submits though. When I was in university I submitted my homework to a very similar system, with a max submits per day. What the university advised students (and I pass on to Kagglers), is to:\n- submit as early as you can to make the most of your limited number of refreshes\n- spend more time critically looking at your code to understand where things can go wrong, and carefully handle them\nRelying on the evaluation system to verify your code works won't be very helpful in production system where you'll need to be able to think ahead of possible issues and protect against them.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 948147,
      "author_name": "vukpetar",
      "author_url": "",
      "post_date": "07/27/2020 17:32:48",
      "content": "<p>Hello everyone,\nI have the same problem and I try to solve it but unfortunately without success. I merged the test data with the training data and then used that to merge with the submission file, but I got the same error: \"Submission CSV Not Found\".\nBelow is the log file from my notebook.\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2651354%2F8d58d85b0ee01271e2135efd47893c63%2Fkaggle_log.png?generation=1595871061983594&amp;alt=media\" alt=\"\"></p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "942245": "Hi! I'm trying to submit my solution but I'm getting an error. I followed the instructions in these threads ([first](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/166533), [second](https://www.kaggle.com/c/osic-pulmonary-fibrosis-progression/discussion/168276)), namely:\n- make sure you iterate over all patients of the test set\n- make sure you cover all weeks from -12 to 133 inclusively\n- turn the internet off in the notebook\n- make sure the name of the submission is \"submission.csv\"  \n\nI also made sure all my column names and formats were OK. Yet, I still can't submit...\n\nHere's the execution info:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa155866847b3d2b75fa6600577a2d717%2FScreenshot%20from%202020-07-23%2012-46-13.png?generation=1595523261224339&amp;alt=media)",
    "942433": "The Status columns in my Submissions tab says \"Submission CSV Not Found\"... But when I look at the actual notebook, the Output section does contain the CSV...",
    "942533": "This is an unfortunate issue. I would suggest perusing some of the suggested solutions to the same issue encountered by another user in [this thread](https://www.kaggle.com/c/tweet-sentiment-extraction/discussion/152425). Hope this helps -- good luck!",
    "942544": "archimedus You might need to verify that the submission file you are producing is named `submission.csv`",
    "942556": "juliaelliott I rechecked and that's exactly how my file is named. Do we need to specify a certain directory when we output? I was using `output.to_csv(\"submission.csv\", index=False)` but using Kaggle's \"Copy file path\" function next to *output &gt; kaggle/working* in the right panel's \"Data\" tab seems to suggest I add `../input/output/` as a prefix to the file name. Does that seem right?",
    "942558": "Thanks @roshanr11, I will look into these suggestions!",
    "942584": "archimedus At the top of your notebook in the viewer mode (not editing), does it say \"Version X of Y\"?\n\nIt's very possible that the latest version has actually failed for a different reason and you're seeing a successfully run version in the viewer (by default it shows you the latest successful one).",
    "942615": "herbison actually, there is a discrepancy between the number of versions that seem to exist for the same notebook when looking at the \"My Submissions\" tab (currently 27) and when looking at the viewer/edit mode (currently 32). Seems a bit odd, maybe the issue has something to do with that?",
    "943085": "Hi @herbison,  I aslo meet the unknown error \"Notebook Exceeded Allowed Compute\". I'm sure that the kernel (the committed version works fine) running time is about 900s without GPU and internet is off as well.",
    "943805": "dxchen That means the rerun of your notebook on private data probably ran out of memory, I wouldn't be surprised if there is more private data or something.\n\n@archimedus It does look like your commit version was also missing submission.csv, please make sure you're looking at the correct version (goto the viewer tab, click on Versions, click on the latest one).",
    "943926": "herbison My Notebook produces the submission.csv file, but it looks like the summary somehow doesn't realize it:\n\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F366c89889034fed690f1031d2cffb748%2FScreenshot%20from%202020-07-24%2013-04-28.png?generation=1595610616477857&amp;alt=media)\n\nScrolling down...\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2Fa4111e1fb5a7d289e43aa821a4e2680d%2FScreenshot%20from%202020-07-24%2013-13-32.png?generation=1595610925014576&amp;alt=media)\n\n\n---\nThere also seems to be a discrepancy between what version numbers are available. Maybe this offers a clue to the source of the problem:\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F69fccb544beeda38ead1a18ecc9e98f1%2FScreenshot%20from%202020-07-24%2013-05-33.png?generation=1595610717867342&amp;alt=media)",
    "943929": "archimedus this has happened to me in previous competitions where i get the error **Submission CSV Not Found**. \n\nThe problem is your submission kernel is crashing while predicting on the hidden test data used to calculate the public LB score and the kernel stops before the csv file is created to calculate the score. It is usually when GPU memory allocation exceeds the limit or due to a small syntax error you may have missed.\n\nThe  **Notebook Exceeded Allowed Compute** error will be shown if your submission kernel exceeds available RAM, but GPU memory exceeding does not usually trigger this error.",
    "943930": "archimedus Correction, I checked and your recent commit does have the submission.csv, but the private rerun hits an exception in your code just before writing it out, please review your code for possible assumptions about the private data files.",
    "943986": "herbison I'm not sure what sort of assumptions might be relevant (I'm a relative novice at all this). To the best of my knowledge I'm not doing anything really special with the files *per se*.\n\nAre you talking about file paths? For all inputs I have the base path `\"../input/osic-pulmonary-fibrosis-progression/\"`; my output has no prefix, I just use `output.to_csv(\"submission.csv\", index=False)` directly. Apart from that:\n- I'm not using any input files other than the default package provided for the competition.\n- I use pydicom.dcmread() (in a *with* block) to read the image files\n- I use Pandas' read_csv() function to read the input CSV's\n\nIs there some line of code I could insert in my notebook to throw up a message about the sort of exceptions you're thinking of?\n\nUnless perhaps this is what you're referring to?\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F3503700%2F0e03793e87b97d14edc2b9f8f0e62d5f%2FScreenshot%20from%202020-07-24%2014-15-30.png?generation=1595614573212920&amp;alt=media)",
    "943995": "archimedus Make sure that every file you try to read actually exists.",
    "944184": "herbison All the files my code tries to access should exist.\n- CSV's don't seem to cause any trouble.\n- For image files my code loops through lists of files obtained using the os library. \n- The file paths are generated partly with a list of patient ID's, but I ran a version of my notebook that printed a line every time a new path was tried to say whether or not it's valid, and all results were \"True\".\n\nI'm currently committing a version in which I commented out a `del` statement (which follows immediately some file-reading `with` blocks - I was having trouble with memory usage so I figured I'd try doing some garbage-collection explicitly). Who knows, maybe that'll work.\n\nAnother potential source of the problem: when you're talking about private data files, do you mean the testing files that are used to score my model but which I won't see on my end? If so, my code relies on the DCM files' `.pixel_array`, `.Rows`, `.Columns`, `.SliceLocation` attributes. If too many of a patient's image files lack these attributes, my code just skips those examples. It doesn't cause too much trouble on the training end (apart from wasting a few training examples) and none on the visible test set, but if the same issues surface in the private data files maybe my code ends up not making predictions for one or more test patient, and rather than the submission checker saying that my CSV file is incomplete (because of missing predictions) it just says that my CSV file isn't there.",
    "944189": "have you posted a link to your notebook so we can try and help you find your error? Are you looking for dicom files with hard-coded names (like dcm.1, which might not always exist in the test set?)",
    "944260": "getting similar error :(\n\nI'm not sure the log you get when you run your notebook is relevant to the error when you submit your csv file.  I have notebooks that show the same error log after running the notebook but when I submit I get a leaderboard score just fine.  \n\nUnfortunately we don't get a log when we submit for leaderboard scoring, so the only relevant information is \"Submission CSV Not Found\" which makes no sense when the notebook ran without error.\n\nI might just move on, come back to it later and try to run it again (i.e. unplug and re-plug the router so to speak lol).  Might try the copy and paste the code into a new notebook idea as well, maybe clean up the code as best I can.",
    "946817": "So I found a solution to my problem, but it doesn't make much sense.  Probably something very particular to this competition, but this notebook\\ person (https://www.kaggle.com/titericz/tabular-simple-eda-linear-model) merged the test data with the training data and then used that to merge with the submission file, and that worked for me.  I kept all of the same features, but that data pipeline worked.  No idea why, but I'll take it I guess.",
    "946875": "FYI, I tried replacing my use of the `.SliceLocation` attributes with `.SliceThickness` and `.InstanceNumber` (after verifying that every single image file in the train set that doesn't require the GDCM library to open has these), but it didn't change anything. Still \"Submission CSV Not Found\"...",
    "948147": "Hello everyone,\nI have the same problem and I try to solve it but unfortunately without success. I merged the test data with the training data and then used that to merge with the submission file, but I got the same error: \"Submission CSV Not Found\".\nBelow is the log file from my notebook.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-user-content/o/inbox%2F2651354%2F8d58d85b0ee01271e2135efd47893c63%2Fkaggle_log.png?generation=1595871061983594&amp;alt=media)",
    "948303": "At this point I've decided to go for a new approach: I created an ultra-simple notebook that outputs what are essentially dummy \"predictions\", and succeeded in submitting it. Now I will add chunks of my code one by one, and see at what point it stops submitting........",
    "949326": "archimedus You're previous submission was failing trying to fetch a file that did not exist, that's why I hinted at that. When a cell hits an exception, it doesn't continue processing that cell, which means it doesn't write out submission.csv if that's done later in that same cell, and leads to you seeing a 'no submission.csv' error.\n\nYour new submission has another, different code bug causing a similar issue (exception thrown in cell before submission.csv is written out).",
    "949333": "yeayates21 We sadly can't provide much feedback on submissions (like an output log) because it could be used to leak private data. I agree it's rather painful but it's one of the few ways we can mitigate the potential for abusing the system to cheat.",
    "949421": "Thank you, that information will be useful in narrowing down where the problem lies.",
    "949513": "herbison Perhaps a \"sanitized\" log could be quite a useful feature, something that tells the submitter basic info such as the cell number/line number that trips up the private run and/or the general type of error (file missing/unreadable, wrong shapes/lengths, arithmetic, CSV result diverging from specifications, etc.), without providing *all* the runtime info (just a few \"white-listed\" items).\n\nI agree that it's important to keep some opacity in the system to deter would-be cheaters, but there has to be a grey area between full logs and no info at all!",
    "949585": "Another possible improvement would be to make failed submissions not count towards the daily total. Daily submission limits make sense if the goal is to prevent people from abusing their public scores as a validation score to fine-tune their models (even though this would be good strategy to get an overfitting model, since the scores are based on only 15% of the test data), but when you're trying to debug an issue that happens in the hidden run (while the run on your end seems to work perfectly) and you have to rely on imprecise trial-and-error (since you have to run the entire notebook - twice! - every time you want to test a hypothesis regarding the location of the error, and get no feedback other than a pass/fail), making the fails count toward the daily submission limit just adds one more (seemingly useless) hurdle...",
    "949616": "archimedus Basically any information we send back from the result (different errors, content of those errors, logs etc.) could be used as bits in an information leaking system, so we really have to limit how much feedback exists.\n\nThe failed submissions count towards totals both because they could be part of a leaking system, and because Kaggle has to control compute usage. Every run of every notebook and all the resources it requires costs something, and we want to make sure every user is given their fair chance, and so individual users don't consume all available compute.\n\nI do understand your frustration with limited submits though. When I was in university I submitted my homework to a very similar system, with a max submits per day. What the university advised students (and I pass on to Kagglers), is to:\n- submit as early as you can to make the most of your limited number of refreshes\n- spend more time critically looking at your code to understand where things can go wrong, and carefully handle them\nRelying on the evaluation system to verify your code works won't be very helpful in production system where you'll need to be able to think ahead of possible issues and protect against them."
  },
  "source": "meta"
}