{
  "id": 188317,
  "title": "How can I debug my code that works on commit but fails submission?",
  "url": "/competitions/osic-pulmonary-fibrosis-progression/discussion/188317",
  "author_name": "",
  "post_date": "2020-10-02T17:41:54.778617Z",
  "votes": 3,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Hi all </p>\n<p>time is getting short now so I'm really hoping some kind soul can help me figure out where my error is.  I'm getting a bit desperate and have read all the submission scoring error threads - like this <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a> but to avail.</p>\n<p>My NB is:</p>\n<ul>\n<li>based on the ubiquitous <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> -higher-lb-score-by-tuning-mloss-around-6-81 </li>\n<li>with my own  Effnet trained on augmented CT scans + lung volumes + their hist features - (trained thanks to <a href=\"https://www.kaggle.com/khoongweihao\" target=\"_blank\">@khoongweihao</a> NB) and with the volume and hist calculated on the training set using the NB from <a href=\"https://www.kaggle.com/hfutybx\" target=\"_blank\">@hfutybx</a> </li>\n</ul>\n<p>So having trained the effnet I now have a NB where I run inference on the test set images -so of course I need to on the fly calculate the lung volumes and hist features soI can feed the CTscans and these extra features into my saved Effnet.</p>\n<p>That all works when I run the NB and it gives me a submission file that has the right columns of Patient_ID, FV, confidence with 730 rows (the -12 to 133 weeks for each of the 5 test.csv patients).</p>\n<p>But THEN when I submit the submission.csv it fails on a scoring error.  I think its because the Effnet model is creating NaNs in my predictions on either the public private LB test data.  I think this because if I do this line to fill NaNs in my Effnet model submission (df here) it does not get a submission error (and does get a much worse score - as you'd expect by sticking zeros in for FVC)</p>\n<p>Please help - some hints on how to troubleshoot it would be amazing please.</p>",
  "messages": [
    {
      "id": "1035415",
      "postDate": "10/02/2020 17:41:54",
      "content": "<p>Hi all </p>\n<p>time is getting short now so I'm really hoping some kind soul can help me figure out where my error is.  I'm getting a bit desperate and have read all the submission scoring error threads - like this <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a> but to avail.</p>\n<p>My NB is:</p>\n<ul>\n<li>based on the ubiquitous <a href=\"https://www.kaggle.com/reighns\" target=\"_blank\">@reighns</a> -higher-lb-score-by-tuning-mloss-around-6-81 </li>\n<li>with my own  Effnet trained on augmented CT scans + lung volumes + their hist features - (trained thanks to <a href=\"https://www.kaggle.com/khoongweihao\" target=\"_blank\">@khoongweihao</a> NB) and with the volume and hist calculated on the training set using the NB from <a href=\"https://www.kaggle.com/hfutybx\" target=\"_blank\">@hfutybx</a> </li>\n</ul>\n<p>So having trained the effnet I now have a NB where I run inference on the test set images -so of course I need to on the fly calculate the lung volumes and hist features soI can feed the CTscans and these extra features into my saved Effnet.</p>\n<p>That all works when I run the NB and it gives me a submission file that has the right columns of Patient_ID, FV, confidence with 730 rows (the -12 to 133 weeks for each of the 5 test.csv patients).</p>\n<p>But THEN when I submit the submission.csv it fails on a scoring error.  I think its because the Effnet model is creating NaNs in my predictions on either the public private LB test data.  I think this because if I do this line to fill NaNs in my Effnet model submission (df here) it does not get a submission error (and does get a much worse score - as you'd expect by sticking zeros in for FVC)</p>\n<p>Please help - some hints on how to troubleshoot it would be amazing please.</p>",
      "rawMarkdown": "Hi all \n\ntime is getting short now so I'm really hoping some kind soul can help me figure out where my error is.  I'm getting a bit desperate and have read all the submission scoring error threads - like this https://www.kaggle.com/code-competition-debugging but to avail.\n\nMy NB is:\n- based on the ubiquitous @reighns -higher-lb-score-by-tuning-mloss-around-6-81 \n- with my own  Effnet trained on augmented CT scans + lung volumes + their hist features - (trained thanks to @khoongweihao NB) and with the volume and hist calculated on the training set using the NB from @hfutybx \n\nSo having trained the effnet I now have a NB where I run inference on the test set images -so of course I need to on the fly calculate the lung volumes and hist features soI can feed the CTscans and these extra features into my saved Effnet.\n\nThat all works when I run the NB and it gives me a submission file that has the right columns of Patient_ID, FV, confidence with 730 rows (the -12 to 133 weeks for each of the 5 test.csv patients).\n\nBut THEN when I submit the submission.csv it fails on a scoring error.  I think its because the Effnet model is creating NaNs in my predictions on either the public private LB test data.  I think this because if I do this line to fill NaNs in my Effnet model submission (df here) it does not get a submission error (and does get a much worse score - as you'd expect by sticking zeros in for FVC)\n\nPlease help - some hints on how to troubleshoot it would be amazing please.",
      "votes": null
    },
    {
      "id": "1035448",
      "postDate": "10/02/2020 18:02:49",
      "content": "<p>Maybe the error comes from the actual submission file itself? I know that I had a problem when pushing my tensors into a .csv, but maybe you could share some code?</p>",
      "rawMarkdown": "Maybe the error comes from the actual submission file itself? I know that I had a problem when pushing my tensors into a .csv, but maybe you could share some code?",
      "votes": null
    },
    {
      "id": "1035548",
      "postDate": "10/02/2020 19:25:10",
      "content": "<p>Sorry - I meant to add some code.  In fact here is the NB made public - the first half is the processing for the newly trained Effnet and img_sub is the submission file that seems to get NANs from the Test data.  The second half is the MQ method unchanged from the original.</p>\n<p>(btw- I've tried to credit all the sources I've used here - please tell me if any missed)</p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/cascadenite/mytrained-model-effnet-mq-nan-filling</a></p>",
      "rawMarkdown": "Sorry - I meant to add some code.  In fact here is the NB made public - the first half is the processing for the newly trained Effnet and img_sub is the submission file that seems to get NANs from the Test data.  The second half is the MQ method unchanged from the original.\n\n(btw- I've tried to credit all the sources I've used here - please tell me if any missed)\n\n[https://www.kaggle.com/cascadenite/mytrained-model-effnet-mq-nan-filling](url)",
      "votes": null
    },
    {
      "id": "1035557",
      "postDate": "10/02/2020 19:41:32",
      "content": "<p>I had the same problem and I had to burn 5-6 submissions for only debugging it. Try to disable every suspicious part in your code and try submitting them.</p>",
      "rawMarkdown": "I had the same problem and I had to burn 5-6 submissions for only debugging it. Try to disable every suspicious part in your code and try submitting them.",
      "votes": null
    },
    {
      "id": "1035560",
      "postDate": "10/02/2020 19:52:47",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - me too, I've gone through 10 submissions to find out where the submission scoring error comes from.  Its so frustrating because:</p>\n<ul>\n<li>the effnet model runs and doesn't throw any fatal errors either in commit or on the private data</li>\n<li>the MQ model is unchanged and so runs fine and gives the same results as you'd expect.</li>\n<li>its only when I blend the two model predictions (img_sub and reg_sub) without imputing NANs in my img_sub file that the submission scoring error occurs.</li>\n</ul>\n<p>The problem I have is that I can't think of any way to see where my effnet model is producing NaNs.  I thought it might be because the private LB dataset has GDCM-needing files but the organisers say that's not the case. Without a log or a way of looking at my submission files when they run on the test data I'm a bit stuck….</p>",
      "rawMarkdown": "Hi @gunesevitan - me too, I've gone through 10 submissions to find out where the submission scoring error comes from.  Its so frustrating because:\n- the effnet model runs and doesn't throw any fatal errors either in commit or on the private data\n- the MQ model is unchanged and so runs fine and gives the same results as you'd expect.\n- its only when I blend the two model predictions (img_sub and reg_sub) without imputing NANs in my img_sub file that the submission scoring error occurs.\n\nThe problem I have is that I can't think of any way to see where my effnet model is producing NaNs.  I thought it might be because the private LB dataset has GDCM-needing files but the organisers say that's not the case. Without a log or a way of looking at my submission files when they run on the test data I'm a bit stuck....",
      "votes": null
    },
    {
      "id": "1036869",
      "postDate": "10/04/2020 11:08:38",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/abhaykatoch\" target=\"_blank\">@abhaykatoch</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - I know you aren't waiting with baited breath but since you were kind enough to reply I thought I'd update you.  I found my error - I had made my categorical Gender wrongly! <br>\nI had coded:<br>\ndata.loc[data.Sex == 'Male','Male']=1</p>\n<p>but no Female =0.  Since the sample test set is all Male it didn't catch this omission until it runs on the real test data…</p>\n<p>What a rookie error :)</p>",
      "rawMarkdown": "Hi @abhaykatoch @gunesevitan - I know you aren't waiting with baited breath but since you were kind enough to reply I thought I'd update you.  I found my error - I had made my categorical Gender wrongly! \nI had coded:\ndata.loc[data.Sex == 'Male','Male']=1\n\nbut no Female =0.  Since the sample test set is all Male it didn't catch this omission until it runs on the real test data...\n\nWhat a rookie error :)",
      "votes": null
    },
    {
      "id": "1036882",
      "postDate": "10/04/2020 11:26:53",
      "content": "<p>I'm glad that you solved it cuz I know how frustrating it is :)</p>\n<p>FYI this is how I did the encoding</p>\n<pre><code>for df in [self.train, self.test]:\n    df['Male'] = 0\n    df['Female'] = 0\n    df.loc[df['Sex'] == 0, 'Male'] = 1\n    df.loc[df['Sex'] == 1, 'Female'] = 1\n</code></pre>",
      "rawMarkdown": "I'm glad that you solved it cuz I know how frustrating it is :)\n\nFYI this is how I did the encoding\n\n```\nfor df in [self.train, self.test]:\n    df['Male'] = 0\n    df['Female'] = 0\n    df.loc[df['Sex'] == 0, 'Male'] = 1\n    df.loc[df['Sex'] == 1, 'Female'] = 1\n```",
      "votes": null
    },
    {
      "id": "1037210",
      "postDate": "10/04/2020 17:27:23",
      "content": "<p>Yeah, little errors like those always get me too. I think for mine I used <code>sklearn.preprocessing.LabelEncoder</code>. </p>",
      "rawMarkdown": "Yeah, little errors like those always get me too. I think for mine I used `sklearn.preprocessing.LabelEncoder`.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1035448,
      "author_name": "abhaykatoch",
      "author_url": "",
      "post_date": "10/02/2020 18:02:49",
      "content": "<p>Maybe the error comes from the actual submission file itself? I know that I had a problem when pushing my tensors into a .csv, but maybe you could share some code?</p>",
      "votes": null,
      "replies": [
        {
          "id": 1036869,
          "author_name": "cascadenite",
          "author_url": "",
          "post_date": "10/04/2020 11:08:38",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/abhaykatoch\" target=\"_blank\">@abhaykatoch</a> <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - I know you aren't waiting with baited breath but since you were kind enough to reply I thought I'd update you.  I found my error - I had made my categorical Gender wrongly! <br>\nI had coded:<br>\ndata.loc[data.Sex == 'Male','Male']=1</p>\n<p>but no Female =0.  Since the sample test set is all Male it didn't catch this omission until it runs on the real test data…</p>\n<p>What a rookie error :)</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1036882,
          "author_name": "gunesevitan",
          "author_url": "",
          "post_date": "10/04/2020 11:26:53",
          "content": "<p>I'm glad that you solved it cuz I know how frustrating it is :)</p>\n<p>FYI this is how I did the encoding</p>\n<pre><code>for df in [self.train, self.test]:\n    df['Male'] = 0\n    df['Female'] = 0\n    df.loc[df['Sex'] == 0, 'Male'] = 1\n    df.loc[df['Sex'] == 1, 'Female'] = 1\n</code></pre>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1037210,
          "author_name": "abhaykatoch",
          "author_url": "",
          "post_date": "10/04/2020 17:27:23",
          "content": "<p>Yeah, little errors like those always get me too. I think for mine I used <code>sklearn.preprocessing.LabelEncoder</code>. </p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1035548,
      "author_name": "cascadenite",
      "author_url": "",
      "post_date": "10/02/2020 19:25:10",
      "content": "<p>Sorry - I meant to add some code.  In fact here is the NB made public - the first half is the processing for the newly trained Effnet and img_sub is the submission file that seems to get NANs from the Test data.  The second half is the MQ method unchanged from the original.</p>\n<p>(btw- I've tried to credit all the sources I've used here - please tell me if any missed)</p>\n<p><a href=\"url\" target=\"_blank\">https://www.kaggle.com/cascadenite/mytrained-model-effnet-mq-nan-filling</a></p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1035557,
      "author_name": "gunesevitan",
      "author_url": "",
      "post_date": "10/02/2020 19:41:32",
      "content": "<p>I had the same problem and I had to burn 5-6 submissions for only debugging it. Try to disable every suspicious part in your code and try submitting them.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1035560,
          "author_name": "cascadenite",
          "author_url": "",
          "post_date": "10/02/2020 19:52:47",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/gunesevitan\" target=\"_blank\">@gunesevitan</a> - me too, I've gone through 10 submissions to find out where the submission scoring error comes from.  Its so frustrating because:</p>\n<ul>\n<li>the effnet model runs and doesn't throw any fatal errors either in commit or on the private data</li>\n<li>the MQ model is unchanged and so runs fine and gives the same results as you'd expect.</li>\n<li>its only when I blend the two model predictions (img_sub and reg_sub) without imputing NANs in my img_sub file that the submission scoring error occurs.</li>\n</ul>\n<p>The problem I have is that I can't think of any way to see where my effnet model is producing NaNs.  I thought it might be because the private LB dataset has GDCM-needing files but the organisers say that's not the case. Without a log or a way of looking at my submission files when they run on the test data I'm a bit stuck….</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1035415": "Hi all \n\ntime is getting short now so I'm really hoping some kind soul can help me figure out where my error is.  I'm getting a bit desperate and have read all the submission scoring error threads - like this https://www.kaggle.com/code-competition-debugging but to avail.\n\nMy NB is:\n- based on the ubiquitous @reighns -higher-lb-score-by-tuning-mloss-around-6-81 \n- with my own  Effnet trained on augmented CT scans + lung volumes + their hist features - (trained thanks to @khoongweihao NB) and with the volume and hist calculated on the training set using the NB from @hfutybx \n\nSo having trained the effnet I now have a NB where I run inference on the test set images -so of course I need to on the fly calculate the lung volumes and hist features soI can feed the CTscans and these extra features into my saved Effnet.\n\nThat all works when I run the NB and it gives me a submission file that has the right columns of Patient_ID, FV, confidence with 730 rows (the -12 to 133 weeks for each of the 5 test.csv patients).\n\nBut THEN when I submit the submission.csv it fails on a scoring error.  I think its because the Effnet model is creating NaNs in my predictions on either the public private LB test data.  I think this because if I do this line to fill NaNs in my Effnet model submission (df here) it does not get a submission error (and does get a much worse score - as you'd expect by sticking zeros in for FVC)\n\nPlease help - some hints on how to troubleshoot it would be amazing please.",
    "1035448": "Maybe the error comes from the actual submission file itself? I know that I had a problem when pushing my tensors into a .csv, but maybe you could share some code?",
    "1035548": "Sorry - I meant to add some code.  In fact here is the NB made public - the first half is the processing for the newly trained Effnet and img_sub is the submission file that seems to get NANs from the Test data.  The second half is the MQ method unchanged from the original.\n\n(btw- I've tried to credit all the sources I've used here - please tell me if any missed)\n\n[https://www.kaggle.com/cascadenite/mytrained-model-effnet-mq-nan-filling](url)",
    "1035557": "I had the same problem and I had to burn 5-6 submissions for only debugging it. Try to disable every suspicious part in your code and try submitting them.",
    "1035560": "Hi @gunesevitan - me too, I've gone through 10 submissions to find out where the submission scoring error comes from.  Its so frustrating because:\n- the effnet model runs and doesn't throw any fatal errors either in commit or on the private data\n- the MQ model is unchanged and so runs fine and gives the same results as you'd expect.\n- its only when I blend the two model predictions (img_sub and reg_sub) without imputing NANs in my img_sub file that the submission scoring error occurs.\n\nThe problem I have is that I can't think of any way to see where my effnet model is producing NaNs.  I thought it might be because the private LB dataset has GDCM-needing files but the organisers say that's not the case. Without a log or a way of looking at my submission files when they run on the test data I'm a bit stuck....",
    "1036869": "Hi @abhaykatoch @gunesevitan - I know you aren't waiting with baited breath but since you were kind enough to reply I thought I'd update you.  I found my error - I had made my categorical Gender wrongly! \nI had coded:\ndata.loc[data.Sex == 'Male','Male']=1\n\nbut no Female =0.  Since the sample test set is all Male it didn't catch this omission until it runs on the real test data...\n\nWhat a rookie error :)",
    "1036882": "I'm glad that you solved it cuz I know how frustrating it is :)\n\nFYI this is how I did the encoding\n\n```\nfor df in [self.train, self.test]:\n    df['Male'] = 0\n    df['Female'] = 0\n    df.loc[df['Sex'] == 0, 'Male'] = 1\n    df.loc[df['Sex'] == 1, 'Female'] = 1\n```",
    "1037210": "Yeah, little errors like those always get me too. I think for mine I used `sklearn.preprocessing.LabelEncoder`."
  },
  "source": "meta"
}