{
  "id": 196883,
  "title": "Submission errror",
  "url": "/competitions/riiid-test-answer-prediction/discussion/196883",
  "author_name": "",
  "post_date": "2020-11-13T08:20:00.257554500Z",
  "votes": null,
  "comment_count": 7,
  "views": 0,
  "content": "<p>Sorry for repetition but I am not able to submit. <br>\nI want to do a baseline submission (1st submission for this competition) but i am always getting submission error. I have tried everything mentioned in this discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124</a>. Also I took 7.5M rows from train as validation set (each batch between 1000 and 4000 rows) and it seems working fine. Also I ran on example set, even there it is working fine. And eventhough the test batches as mentioned in this discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009</a> is very small (&lt;30), my submission is giving error within 10 minutes. </p>\n<ol>\n<li>Could there be batches in the test set with only lecture row ? If yes, in that case should I pass only the empty dataframe ?</li>\n<li>Will print statements cause error ?</li>\n<li>Also during the hidden set, the prior group responses and prior group answers correct will be string right ? i.e. the 1st row… Bcoz thats how it is there in the example test. So I used json to convert to array. </li>\n</ol>\n<p>Also I would be thankful if you guys could give other possible reasons too. <br>\nMoreover instead of showing just error in the submission page, would not it be nice to show like 1st 100 characters of the exception or at least error type and line number ?  Or at least provide us a better test set which is a good representative of hidden test set given that we receive at least one topic under this context daily? </p>\n<p>Edit:<br>\nI found out the reasons. There are two, </p>\n<ol>\n<li>I was calculating auc score for each batch, and sklearn metrics.roc_auc_score throws error when the batch contains only single class</li>\n<li>I was predicting True or False (i.e I gave pred &gt;= 0.5), and it was throwing error for some reason in the hidden set metrics calculation. But I can easily calculate the score with it on kaggle notebooks. </li>\n</ol>",
  "messages": [
    {
      "id": "1077086",
      "postDate": "11/13/2020 08:20:00",
      "content": "<p>Sorry for repetition but I am not able to submit. <br>\nI want to do a baseline submission (1st submission for this competition) but i am always getting submission error. I have tried everything mentioned in this discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124</a>. Also I took 7.5M rows from train as validation set (each batch between 1000 and 4000 rows) and it seems working fine. Also I ran on example set, even there it is working fine. And eventhough the test batches as mentioned in this discussion <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009</a> is very small (&lt;30), my submission is giving error within 10 minutes. </p>\n<ol>\n<li>Could there be batches in the test set with only lecture row ? If yes, in that case should I pass only the empty dataframe ?</li>\n<li>Will print statements cause error ?</li>\n<li>Also during the hidden set, the prior group responses and prior group answers correct will be string right ? i.e. the 1st row… Bcoz thats how it is there in the example test. So I used json to convert to array. </li>\n</ol>\n<p>Also I would be thankful if you guys could give other possible reasons too. <br>\nMoreover instead of showing just error in the submission page, would not it be nice to show like 1st 100 characters of the exception or at least error type and line number ?  Or at least provide us a better test set which is a good representative of hidden test set given that we receive at least one topic under this context daily? </p>\n<p>Edit:<br>\nI found out the reasons. There are two, </p>\n<ol>\n<li>I was calculating auc score for each batch, and sklearn metrics.roc_auc_score throws error when the batch contains only single class</li>\n<li>I was predicting True or False (i.e I gave pred &gt;= 0.5), and it was throwing error for some reason in the hidden set metrics calculation. But I can easily calculate the score with it on kaggle notebooks. </li>\n</ol>",
      "rawMarkdown": "Sorry for repetition but I am not able to submit. \nI want to do a baseline submission (1st submission for this competition) but i am always getting submission error. I have tried everything mentioned in this discussion https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124. Also I took 7.5M rows from train as validation set (each batch between 1000 and 4000 rows) and it seems working fine. Also I ran on example set, even there it is working fine. And eventhough the test batches as mentioned in this discussion https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009 is very small (<30), my submission is giving error within 10 minutes. \n\n1. Could there be batches in the test set with only lecture row ? If yes, in that case should I pass only the empty dataframe ?\n2. Will print statements cause error ?\n3. Also during the hidden set, the prior group responses and prior group answers correct will be string right ? i.e. the 1st row... Bcoz thats how it is there in the example test. So I used json to convert to array. \n\nAlso I would be thankful if you guys could give other possible reasons too. \nMoreover instead of showing just error in the submission page, would not it be nice to show like 1st 100 characters of the exception or at least error type and line number ?  Or at least provide us a better test set which is a good representative of hidden test set given that we receive at least one topic under this context daily? \n\nEdit:\nI found out the reasons. There are two, \n\n1. I was calculating auc score for each batch, and sklearn metrics.roc_auc_score throws error when the batch contains only single class\n2. I was predicting True or False (i.e I gave pred >= 0.5), and it was throwing error for some reason in the hidden set metrics calculation. But I can easily calculate the score with it on kaggle notebooks.",
      "votes": null
    },
    {
      "id": "1077219",
      "postDate": "11/13/2020 10:35:49",
      "content": "<p>Hi Ajay:</p>\n<p>Your 1: based on my submissions there is no batch with only lectures.<br>\nYour 2: print statements have not caused errors for me<br>\nYour 3: json? My understanding is that submissions should (must?) be in Python or R code, not Javascript. For an example of how to handle \"prior group answers correct\" as a list in Python, see <br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815</a></p>\n<p>Suggestion: do a simple submission in which every prediction is 0.5. Does this get you an ROCAUC of 0.500. Once this has succeeded, add new features one small step at a time. Doing this revealed obscure reasons for submission errors.</p>",
      "rawMarkdown": "Hi Ajay:\n\nYour 1: based on my submissions there is no batch with only lectures.\nYour 2: print statements have not caused errors for me\nYour 3: json? My understanding is that submissions should (must?) be in Python or R code, not Javascript. For an example of how to handle \"prior group answers correct\" as a list in Python, see \nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\n\nSuggestion: do a simple submission in which every prediction is 0.5. Does this get you an ROCAUC of 0.500. Once this has succeeded, add new features one small step at a time. Doing this revealed obscure reasons for submission errors.",
      "votes": null
    },
    {
      "id": "1077227",
      "postDate": "11/13/2020 11:07:28",
      "content": "<p>Actually i meant json package in python. json.loads('[1. 2, 3]'), will produce python list [1, 2, 3]. Moreover I actually dont know any syntax in Javascript. And thanks for the help Mike. Will start from basic 0.5 prediction submission. </p>",
      "rawMarkdown": "Actually i meant json package in python. json.loads('[1. 2, 3]'), will produce python list [1, 2, 3]. Moreover I actually dont know any syntax in Javascript. And thanks for the help Mike. Will start from basic 0.5 prediction submission.",
      "votes": null
    },
    {
      "id": "1077303",
      "postDate": "11/13/2020 12:51:42",
      "content": "<p>Hi Ajay,</p>\n<p>I also experienced Submission Errors after a couple of minutes. The reason for that was, that I included a feature that depends on user' history. However, this resulted in an \"IndexerError\" for new users in one line and I instantly reveived a \"Submission Error\" after submitting.</p>\n<p>What I highly recommend is using validation loop from this notebook <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator/data\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator/data</a></p>\n<p>I also only took a subset of the observed users in the training set to make sure that calculations for new users are indeed correct.</p>",
      "rawMarkdown": "Hi Ajay,\n\nI also experienced Submission Errors after a couple of minutes. The reason for that was, that I included a feature that depends on user' history. However, this resulted in an \"IndexerError\" for new users in one line and I instantly reveived a \"Submission Error\" after submitting.\n\nWhat I highly recommend is using validation loop from this notebook https://www.kaggle.com/its7171/time-series-api-iter-test-emulator/data\n\nI also only took a subset of the observed users in the training set to make sure that calculations for new users are indeed correct.",
      "votes": null
    },
    {
      "id": "1079161",
      "postDate": "11/15/2020 17:32:21",
      "content": "<p>Off-topic question: How do you guys filter lectures out of test_set? I didnt find any reference</p>",
      "rawMarkdown": "Off-topic question: How do you guys filter lectures out of test_set? I didnt find any reference",
      "votes": null
    },
    {
      "id": "1085017",
      "postDate": "11/20/2020 15:49:50",
      "content": "<p>you can use the following code</p>\n<pre><code>test_df = test_df.loc[test_df['content_type_id'] == 0].reset_index(drop=True)\n</code></pre>\n<p>or:</p>\n<pre><code>env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>\n<p>to filter lectures out of test_set</p>",
      "rawMarkdown": "you can use the following code\n```\ntest_df = test_df.loc[test_df['content_type_id'] == 0].reset_index(drop=True)\n```\nor:\n```\nenv.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n```\nto filter lectures out of test_set",
      "votes": null
    },
    {
      "id": "1085027",
      "postDate": "11/20/2020 16:00:58",
      "content": "<p>Hello, I have figured it out afterwards</p>\n<p>but thank you still.</p>",
      "rawMarkdown": "Hello, I have figured it out afterwards\n\nbut thank you still.",
      "votes": null
    },
    {
      "id": "1105104",
      "postDate": "12/07/2020 14:22:28",
      "content": "<p>I encountered Submission Score Error after about 26mins.  Any suggestion on the possible cause?</p>",
      "rawMarkdown": "I encountered Submission Score Error after about 26mins.  Any suggestion on the possible cause?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 1077219,
      "author_name": "mikel1",
      "author_url": "",
      "post_date": "11/13/2020 10:35:49",
      "content": "<p>Hi Ajay:</p>\n<p>Your 1: based on my submissions there is no batch with only lectures.<br>\nYour 2: print statements have not caused errors for me<br>\nYour 3: json? My understanding is that submissions should (must?) be in Python or R code, not Javascript. For an example of how to handle \"prior group answers correct\" as a list in Python, see <br>\n<a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815</a></p>\n<p>Suggestion: do a simple submission in which every prediction is 0.5. Does this get you an ROCAUC of 0.500. Once this has succeeded, add new features one small step at a time. Doing this revealed obscure reasons for submission errors.</p>",
      "votes": null,
      "replies": [
        {
          "id": 1077227,
          "author_name": "ajayhayagreeve",
          "author_url": "",
          "post_date": "11/13/2020 11:07:28",
          "content": "<p>Actually i meant json package in python. json.loads('[1. 2, 3]'), will produce python list [1, 2, 3]. Moreover I actually dont know any syntax in Javascript. And thanks for the help Mike. Will start from basic 0.5 prediction submission. </p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1079161,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/15/2020 17:32:21",
          "content": "<p>Off-topic question: How do you guys filter lectures out of test_set? I didnt find any reference</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1085017,
          "author_name": "suhehs",
          "author_url": "",
          "post_date": "11/20/2020 15:49:50",
          "content": "<p>you can use the following code</p>\n<pre><code>test_df = test_df.loc[test_df['content_type_id'] == 0].reset_index(drop=True)\n</code></pre>\n<p>or:</p>\n<pre><code>env.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n</code></pre>\n<p>to filter lectures out of test_set</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 1085027,
          "author_name": "aliabdin1",
          "author_url": "",
          "post_date": "11/20/2020 16:00:58",
          "content": "<p>Hello, I have figured it out afterwards</p>\n<p>but thank you still.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 1077303,
      "author_name": "markuslill",
      "author_url": "",
      "post_date": "11/13/2020 12:51:42",
      "content": "<p>Hi Ajay,</p>\n<p>I also experienced Submission Errors after a couple of minutes. The reason for that was, that I included a feature that depends on user' history. However, this resulted in an \"IndexerError\" for new users in one line and I instantly reveived a \"Submission Error\" after submitting.</p>\n<p>What I highly recommend is using validation loop from this notebook <a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator/data\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator/data</a></p>\n<p>I also only took a subset of the observed users in the training set to make sure that calculations for new users are indeed correct.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 1105104,
      "author_name": "lsllee",
      "author_url": "",
      "post_date": "12/07/2020 14:22:28",
      "content": "<p>I encountered Submission Score Error after about 26mins.  Any suggestion on the possible cause?</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1077086": "Sorry for repetition but I am not able to submit. \nI want to do a baseline submission (1st submission for this competition) but i am always getting submission error. I have tried everything mentioned in this discussion https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124. Also I took 7.5M rows from train as validation set (each batch between 1000 and 4000 rows) and it seems working fine. Also I ran on example set, even there it is working fine. And eventhough the test batches as mentioned in this discussion https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/196009 is very small (<30), my submission is giving error within 10 minutes. \n\n1. Could there be batches in the test set with only lecture row ? If yes, in that case should I pass only the empty dataframe ?\n2. Will print statements cause error ?\n3. Also during the hidden set, the prior group responses and prior group answers correct will be string right ? i.e. the 1st row... Bcoz thats how it is there in the example test. So I used json to convert to array. \n\nAlso I would be thankful if you guys could give other possible reasons too. \nMoreover instead of showing just error in the submission page, would not it be nice to show like 1st 100 characters of the exception or at least error type and line number ?  Or at least provide us a better test set which is a good representative of hidden test set given that we receive at least one topic under this context daily? \n\nEdit:\nI found out the reasons. There are two, \n\n1. I was calculating auc score for each batch, and sklearn metrics.roc_auc_score throws error when the batch contains only single class\n2. I was predicting True or False (i.e I gave pred >= 0.5), and it was throwing error for some reason in the hidden set metrics calculation. But I can easily calculate the score with it on kaggle notebooks.",
    "1077219": "Hi Ajay:\n\nYour 1: based on my submissions there is no batch with only lectures.\nYour 2: print statements have not caused errors for me\nYour 3: json? My understanding is that submissions should (must?) be in Python or R code, not Javascript. For an example of how to handle \"prior group answers correct\" as a list in Python, see \nhttps://www.kaggle.com/c/riiid-test-answer-prediction/discussion/195815\n\nSuggestion: do a simple submission in which every prediction is 0.5. Does this get you an ROCAUC of 0.500. Once this has succeeded, add new features one small step at a time. Doing this revealed obscure reasons for submission errors.",
    "1077227": "Actually i meant json package in python. json.loads('[1. 2, 3]'), will produce python list [1, 2, 3]. Moreover I actually dont know any syntax in Javascript. And thanks for the help Mike. Will start from basic 0.5 prediction submission.",
    "1077303": "Hi Ajay,\n\nI also experienced Submission Errors after a couple of minutes. The reason for that was, that I included a feature that depends on user' history. However, this resulted in an \"IndexerError\" for new users in one line and I instantly reveived a \"Submission Error\" after submitting.\n\nWhat I highly recommend is using validation loop from this notebook https://www.kaggle.com/its7171/time-series-api-iter-test-emulator/data\n\nI also only took a subset of the observed users in the training set to make sure that calculations for new users are indeed correct.",
    "1079161": "Off-topic question: How do you guys filter lectures out of test_set? I didnt find any reference",
    "1085017": "you can use the following code\n```\ntest_df = test_df.loc[test_df['content_type_id'] == 0].reset_index(drop=True)\n```\nor:\n```\nenv.predict(test_df.loc[test_df['content_type_id'] == 0, ['row_id', 'answered_correctly']])\n```\nto filter lectures out of test_set",
    "1085027": "Hello, I have figured it out afterwards\n\nbut thank you still.",
    "1105104": "I encountered Submission Score Error after about 26mins.  Any suggestion on the possible cause?"
  },
  "source": "meta"
}