{
  "id": 183496,
  "title": "Submission file missing - help needed",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/183496",
  "author_name": "",
  "post_date": "2020-09-16T22:06:40.052239700Z",
  "votes": null,
  "comment_count": 11,
  "views": 0,
  "content": "<p>HI all!<br>\nI would be grateful if I could get some help with my notebook. By the numbers' I'm getting when I run it locally, it's not competitive at the leaderboard, but still, I was wondering about what could be wrong with it.</p>\n<p>Here it is: <a href=\"https://www.kaggle.com/douglaskgaraujo/bayesian-log-linear-hierarchical-model\" target=\"_blank\">https://www.kaggle.com/douglaskgaraujo/bayesian-log-linear-hierarchical-model</a></p>\n<p>When it's run both locally and on Kaggle it apparently generates a submission.csv - so I'm puzzled what the issue might be.</p>\n<p>Many many thanks in advance,<br>\nDoug</p>",
  "messages": [
    {
      "id": "1013695",
      "postDate": "09/16/2020 22:06:40",
      "content": "<p>HI all!<br>\nI would be grateful if I could get some help with my notebook. By the numbers' I'm getting when I run it locally, it's not competitive at the leaderboard, but still, I was wondering about what could be wrong with it.</p>\n<p>Here it is: <a href=\"https://www.kaggle.com/douglaskgaraujo/bayesian-log-linear-hierarchical-model\" target=\"_blank\">https://www.kaggle.com/douglaskgaraujo/bayesian-log-linear-hierarchical-model</a></p>\n<p>When it's run both locally and on Kaggle it apparently generates a submission.csv - so I'm puzzled what the issue might be.</p>\n<p>Many many thanks in advance,<br>\nDoug</p>",
      "rawMarkdown": "HI all!\nI would be grateful if I could get some help with my notebook. By the numbers' I'm getting when I run it locally, it's not competitive at the leaderboard, but still, I was wondering about what could be wrong with it.\n\nHere it is: https://www.kaggle.com/douglaskgaraujo/bayesian-log-linear-hierarchical-model\n\nWhen it's run both locally and on Kaggle it apparently generates a submission.csv - so I'm puzzled what the issue might be.\n\nMany many thanks in advance,\nDoug",
      "votes": null
    },
    {
      "id": "1014020",
      "postDate": "09/17/2020 06:38:02",
      "content": "<p>nice work</p>",
      "rawMarkdown": "nice work",
      "votes": null
    },
    {
      "id": "1014056",
      "postDate": "09/17/2020 07:16:06",
      "content": "<p>It's not immediately obvious what the issue is. The output submission.csv that your notebook creates from the publically available test.csv looks right. </p>\n<p>So, I'd assume the issue is related to something that is different when using the real leaderboard test.csv vs. the one that is available when you develop code. However, since you seem to exclude test patients from your training data (like it apparently is in the real test data), it is not immediately clear why you would run into an issue (perhaps you are not fully doing that in every way?). Perhaps worth double-checking your script handles the following correctly (I had one mess up with that, so it seems like someone else could fall for these issues): (1) number of test records is different than in the test.csv you have when you develop (so, e.g. if you assign patient numbers from 0,1,2,… and rely on that range not being exceeded that could be an issue), and (2) patient IDs may or may not overlap (I think they do not, but best if the code runs either way) between train.csv and test.csv when you submit (while they do on what you have available publically). </p>",
      "rawMarkdown": "It's not immediately obvious what the issue is. The output submission.csv that your notebook creates from the publically available test.csv looks right. \n\nSo, I'd assume the issue is related to something that is different when using the real leaderboard test.csv vs. the one that is available when you develop code. However, since you seem to exclude test patients from your training data (like it apparently is in the real test data), it is not immediately clear why you would run into an issue (perhaps you are not fully doing that in every way?). Perhaps worth double-checking your script handles the following correctly (I had one mess up with that, so it seems like someone else could fall for these issues): (1) number of test records is different than in the test.csv you have when you develop (so, e.g. if you assign patient numbers from 0,1,2,... and rely on that range not being exceeded that could be an issue), and (2) patient IDs may or may not overlap (I think they do not, but best if the code runs either way) between train.csv and test.csv when you submit (while they do on what you have available publically).",
      "votes": null
    },
    {
      "id": "1014068",
      "postDate": "09/17/2020 07:25:40",
      "content": "<p>Many thanks for the comments, <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> ! Insightful. From your reply, I gather it could be something related then to the way I am using LabelEncoder(). I'll explore that and perhaps follow <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> 's original design where I think (trying to remember by heart now) he encodes patient IDs using both sets, presumably to avoid the issue I'm facing. </p>\n<p>But thanks again, and I'll share an update here after I have some (hopefully) good news!</p>",
      "rawMarkdown": "Many thanks for the comments, @bjoernholzhauer ! Insightful. From your reply, I gather it could be something related then to the way I am using LabelEncoder(). I'll explore that and perhaps follow @carlossouza 's original design where I think (trying to remember by heart now) he encodes patient IDs using both sets, presumably to avoid the issue I'm facing. \n\nBut thanks again, and I'll share an update here after I have some (hopefully) good news!",
      "votes": null
    },
    {
      "id": "1014069",
      "postDate": "09/17/2020 07:26:09",
      "content": "<p>Thank you, <a href=\"https://www.kaggle.com/vijaysimhareddyp\" target=\"_blank\">@vijaysimhareddyp</a> !</p>",
      "rawMarkdown": "Thank you, @vijaysimhareddyp !",
      "votes": null
    },
    {
      "id": "1015625",
      "postDate": "09/18/2020 09:57:37",
      "content": "<p>I faced the same problem. Check settings before saving version of your notebook, <br>\nAfter clicking \"save version\"  &gt;&gt;&gt; Choose \"always save outputs\" before to save it finally. You can then select the output in Submission section then.<br>\nThis solved the problem for me. <br>\nI hope it is helpful.</p>",
      "rawMarkdown": "I faced the same problem. Check settings before saving version of your notebook, \nAfter clicking \"save version\"  >>> Choose \"always save outputs\" before to save it finally. You can then select the output in Submission section then.\nThis solved the problem for me. \nI hope it is helpful.",
      "votes": null
    },
    {
      "id": "1016295",
      "postDate": "09/18/2020 19:48:19",
      "content": "<p><a href=\"https://www.kaggle.com/amritpal333\" target=\"_blank\">@amritpal333</a> thank you for your suggestion and your insight! I have now completed my notebook, making sure to have the \"always save outputs\" option checked. Unfortunately for me it did not work, so I will continue investigating the cases. In any case, I appreciate your sharing of your experience and I am glad you were able to submit your solution. Best of luck in this competition!</p>",
      "rawMarkdown": "amritpal333 thank you for your suggestion and your insight! I have now completed my notebook, making sure to have the \"always save outputs\" option checked. Unfortunately for me it did not work, so I will continue investigating the cases. In any case, I appreciate your sharing of your experience and I am glad you were able to submit your solution. Best of luck in this competition!",
      "votes": null
    },
    {
      "id": "1017912",
      "postDate": "09/19/2020 10:36:15",
      "content": "<p>Would anyone in the Kaggle team have an insight into why is this happening?</p>",
      "rawMarkdown": "Would anyone in the Kaggle team have an insight into why is this happening?",
      "votes": null
    },
    {
      "id": "1018063",
      "postDate": "09/19/2020 12:27:24",
      "content": "<p>Just reaching. Could this give you a divide by zero error:</p>\n<p>df_test['FVC_inf'] = df_test['FVC_pred'] / df_test['sigma']</p>",
      "rawMarkdown": "Just reaching. Could this give you a divide by zero error:\n\n                                                                \ndf_test['FVC_inf'] = df_test['FVC_pred'] / df_test['sigma']",
      "votes": null
    },
    {
      "id": "1022472",
      "postDate": "09/22/2020 15:05:29",
      "content": "<p>Hi Doug, this may help: <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a></p>\n<p>For fairness reasons, we are unfortunately not able to provide private debugging information.</p>",
      "rawMarkdown": "Hi Doug, this may help: https://www.kaggle.com/code-competition-debugging\n\nFor fairness reasons, we are unfortunately not able to provide private debugging information.",
      "votes": null
    },
    {
      "id": "1022513",
      "postDate": "09/22/2020 15:34:38",
      "content": "<p>Many thanks for the pointer, <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> ! Will definitely check it out. And of course, if it would impact fairness then it's best not to provide private debugging information. Many thanks!</p>",
      "rawMarkdown": "Many thanks for the pointer, @wcukierski ! Will definitely check it out. And of course, if it would impact fairness then it's best not to provide private debugging information. Many thanks!",
      "votes": null
    },
    {
      "id": "1022518",
      "postDate": "09/22/2020 15:36:11",
      "content": "<p>Many thanks for the suggestion, <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> ! I have wrapped my division as follows: df_test['FVC_pred'] / max(0.01, df_test['sigma']) and still it can't submit. I'll continue searching. I appreciate your suggestion!</p>",
      "rawMarkdown": "Many thanks for the suggestion, @richardepstein ! I have wrapped my division as follows: df_test['FVC_pred'] / max(0.01, df_test['sigma']) and still it can't submit. I'll continue searching. I appreciate your suggestion!",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1014056,
      "author_name": "bjoernholzhauer",
      "author_url": "",
      "post_date": "09/17/2020 07:16:06",
      "content": "<p>It's not immediately obvious what the issue is. The output submission.csv that your notebook creates from the publically available test.csv looks right. </p>\n<p>So, I'd assume the issue is related to something that is different when using the real leaderboard test.csv vs. the one that is available when you develop code. However, since you seem to exclude test patients from your training data (like it apparently is in the real test data), it is not immediately clear why you would run into an issue (perhaps you are not fully doing that in every way?). Perhaps worth double-checking your script handles the following correctly (I had one mess up with that, so it seems like someone else could fall for these issues): (1) number of test records is different than in the test.csv you have when you develop (so, e.g. if you assign patient numbers from 0,1,2,… and rely on that range not being exceeded that could be an issue), and (2) patient IDs may or may not overlap (I think they do not, but best if the code runs either way) between train.csv and test.csv when you submit (while they do on what you have available publically). </p>",
      "votes": null,
      "replies": [
        {
          "id": 1014068,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/17/2020 07:25:40",
          "content": "<p>Many thanks for the comments, <a href=\"https://www.kaggle.com/bjoernholzhauer\" target=\"_blank\">@bjoernholzhauer</a> ! Insightful. From your reply, I gather it could be something related then to the way I am using LabelEncoder(). I'll explore that and perhaps follow <a href=\"https://www.kaggle.com/carlossouza\" target=\"_blank\">@carlossouza</a> 's original design where I think (trying to remember by heart now) he encodes patient IDs using both sets, presumably to avoid the issue I'm facing. </p>\n<p>But thanks again, and I'll share an update here after I have some (hopefully) good news!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1015625,
      "author_name": "amritpal333",
      "author_url": "",
      "post_date": "09/18/2020 09:57:37",
      "content": "<p>I faced the same problem. Check settings before saving version of your notebook, <br>\nAfter clicking \"save version\"  &gt;&gt;&gt; Choose \"always save outputs\" before to save it finally. You can then select the output in Submission section then.<br>\nThis solved the problem for me. <br>\nI hope it is helpful.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1016295,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/18/2020 19:48:19",
          "content": "<p><a href=\"https://www.kaggle.com/amritpal333\" target=\"_blank\">@amritpal333</a> thank you for your suggestion and your insight! I have now completed my notebook, making sure to have the \"always save outputs\" option checked. Unfortunately for me it did not work, so I will continue investigating the cases. In any case, I appreciate your sharing of your experience and I am glad you were able to submit your solution. Best of luck in this competition!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1017912,
      "author_name": "douglaskgaraujo",
      "author_url": "",
      "post_date": "09/19/2020 10:36:15",
      "content": "<p>Would anyone in the Kaggle team have an insight into why is this happening?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1018063,
          "author_name": "richardepstein",
          "author_url": "",
          "post_date": "09/19/2020 12:27:24",
          "content": "<p>Just reaching. Could this give you a divide by zero error:</p>\n<p>df_test['FVC_inf'] = df_test['FVC_pred'] / df_test['sigma']</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022472,
          "author_name": "wcukierski",
          "author_url": "",
          "post_date": "09/22/2020 15:05:29",
          "content": "<p>Hi Doug, this may help: <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a></p>\n<p>For fairness reasons, we are unfortunately not able to provide private debugging information.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022513,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/22/2020 15:34:38",
          "content": "<p>Many thanks for the pointer, <a href=\"https://www.kaggle.com/wcukierski\" target=\"_blank\">@wcukierski</a> ! Will definitely check it out. And of course, if it would impact fairness then it's best not to provide private debugging information. Many thanks!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1022518,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/22/2020 15:36:11",
          "content": "<p>Many thanks for the suggestion, <a href=\"https://www.kaggle.com/richardepstein\" target=\"_blank\">@richardepstein</a> ! I have wrapped my division as follows: df_test['FVC_pred'] / max(0.01, df_test['sigma']) and still it can't submit. I'll continue searching. I appreciate your suggestion!</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1014020,
      "author_name": "vijaysimhareddyp",
      "author_url": "",
      "post_date": "09/17/2020 06:38:02",
      "content": "<p>nice work</p>",
      "votes": null,
      "replies": [
        {
          "id": 1014069,
          "author_name": "douglaskgaraujo",
          "author_url": "",
          "post_date": "09/17/2020 07:26:09",
          "content": "<p>Thank you, <a href=\"https://www.kaggle.com/vijaysimhareddyp\" target=\"_blank\">@vijaysimhareddyp</a> !</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1013695": "HI all!\nI would be grateful if I could get some help with my notebook. By the numbers' I'm getting when I run it locally, it's not competitive at the leaderboard, but still, I was wondering about what could be wrong with it.\n\nHere it is: https://www.kaggle.com/douglaskgaraujo/bayesian-log-linear-hierarchical-model\n\nWhen it's run both locally and on Kaggle it apparently generates a submission.csv - so I'm puzzled what the issue might be.\n\nMany many thanks in advance,\nDoug",
    "1014020": "nice work",
    "1014056": "It's not immediately obvious what the issue is. The output submission.csv that your notebook creates from the publically available test.csv looks right. \n\nSo, I'd assume the issue is related to something that is different when using the real leaderboard test.csv vs. the one that is available when you develop code. However, since you seem to exclude test patients from your training data (like it apparently is in the real test data), it is not immediately clear why you would run into an issue (perhaps you are not fully doing that in every way?). Perhaps worth double-checking your script handles the following correctly (I had one mess up with that, so it seems like someone else could fall for these issues): (1) number of test records is different than in the test.csv you have when you develop (so, e.g. if you assign patient numbers from 0,1,2,... and rely on that range not being exceeded that could be an issue), and (2) patient IDs may or may not overlap (I think they do not, but best if the code runs either way) between train.csv and test.csv when you submit (while they do on what you have available publically).",
    "1014068": "Many thanks for the comments, @bjoernholzhauer ! Insightful. From your reply, I gather it could be something related then to the way I am using LabelEncoder(). I'll explore that and perhaps follow @carlossouza 's original design where I think (trying to remember by heart now) he encodes patient IDs using both sets, presumably to avoid the issue I'm facing. \n\nBut thanks again, and I'll share an update here after I have some (hopefully) good news!",
    "1014069": "Thank you, @vijaysimhareddyp !",
    "1015625": "I faced the same problem. Check settings before saving version of your notebook, \nAfter clicking \"save version\"  >>> Choose \"always save outputs\" before to save it finally. You can then select the output in Submission section then.\nThis solved the problem for me. \nI hope it is helpful.",
    "1016295": "amritpal333 thank you for your suggestion and your insight! I have now completed my notebook, making sure to have the \"always save outputs\" option checked. Unfortunately for me it did not work, so I will continue investigating the cases. In any case, I appreciate your sharing of your experience and I am glad you were able to submit your solution. Best of luck in this competition!",
    "1017912": "Would anyone in the Kaggle team have an insight into why is this happening?",
    "1018063": "Just reaching. Could this give you a divide by zero error:\n\n                                                                \ndf_test['FVC_inf'] = df_test['FVC_pred'] / df_test['sigma']",
    "1022472": "Hi Doug, this may help: https://www.kaggle.com/code-competition-debugging\n\nFor fairness reasons, we are unfortunately not able to provide private debugging information.",
    "1022513": "Many thanks for the pointer, @wcukierski ! Will definitely check it out. And of course, if it would impact fairness then it's best not to provide private debugging information. Many thanks!",
    "1022518": "Many thanks for the suggestion, @richardepstein ! I have wrapped my division as follows: df_test['FVC_pred'] / max(0.01, df_test['sigma']) and still it can't submit. I'll continue searching. I appreciate your suggestion!"
  },
  "source": "meta"
}