{
  "id": 191582,
  "title": "Does anyone get other errors than \"submission scoring error\"?",
  "url": "/competitions/riiid-test-answer-prediction/discussion/191582",
  "author_name": "Alex Bader",
  "post_date": "2020-10-17T09:34:10.130000",
  "votes": 12,
  "comment_count": 17,
  "views": 0,
  "content": "<p>If I submit a notebook which throws an exception after a few seconds, it reports back \"submission scoring error\". Another of my notebooks runs for ~9 hours (so probably times out) but also gives me a \"submission scoring error\". Another runs for some hours (not sure if it hit the time limit) and also gives me this error (the code being literally copied from a kernel of mine which was accepted without errors).</p>\n<p>What's going on here? Are these the only error messages we're getting in this competition? Besides being zero informative, they seem to be somewhat random as well. As if it wasn't bad enough that the public test set which we're supposed to use for debugging is somewhat useless (no lecture rows, not of a representative size, … ).</p>\n<p>Sorry for venting but I do feel like the setup of this competition is quite frustrating overall.</p>",
  "messages": [
    {
      "id": 1052086,
      "postDate": "2020-10-17T09:34:10.130Z",
      "content": "<p>If I submit a notebook which throws an exception after a few seconds, it reports back \"submission scoring error\". Another of my notebooks runs for ~9 hours (so probably times out) but also gives me a \"submission scoring error\". Another runs for some hours (not sure if it hit the time limit) and also gives me this error (the code being literally copied from a kernel of mine which was accepted without errors).</p>\n<p>What's going on here? Are these the only error messages we're getting in this competition? Besides being zero informative, they seem to be somewhat random as well. As if it wasn't bad enough that the public test set which we're supposed to use for debugging is somewhat useless (no lecture rows, not of a representative size, … ).</p>\n<p>Sorry for venting but I do feel like the setup of this competition is quite frustrating overall.</p>",
      "rawMarkdown": "If I submit a notebook which throws an exception after a few seconds, it reports back \"submission scoring error\". Another of my notebooks runs for ~9 hours (so probably times out) but also gives me a \"submission scoring error\". Another runs for some hours (not sure if it hit the time limit) and also gives me this error (the code being literally copied from a kernel of mine which was accepted without errors).\n\nWhat's going on here? Are these the only error messages we're getting in this competition? Besides being zero informative, they seem to be somewhat random as well. As if it wasn't bad enough that the public test set which we're supposed to use for debugging is somewhat useless (no lecture rows, not of a representative size, ... ).\n\nSorry for venting but I do feel like the setup of this competition is quite frustrating overall.",
      "votes": 11
    },
    {
      "id": 1052408,
      "postDate": "2020-10-17T17:25:59.133Z",
      "content": "<p>I have mixed feelings here. On the one hand, I've had the same frustrations as you -- debugging when you can't see the traceback or even disambiguate between programming and memory errors is an obvious nightmare. On the other hand, this <em>is</em> a \"code competition\" after all -- it's good to be forced into a more real-world scenario where designing your own rigorous test cases and measuring your pipeline's efficiency before deployment is critical. Certainly frustrating, but it makes for good habit building in a way that a traditional competition doesn't. I'm trying to think of it as an opportunity to practice more coding/engineering skills instead of just an annoying obstacle. </p>\n<p>I think it would have been a nice compromise to give us access to a larger and more representative practice test set to parse with the API, but at least we have tons of data to use for pipeline testing. Again, more realistic to have to design these tests from scratch than to be handed them for free.</p>\n<p>Re: the \"submission scoring error\", I believe that kaggle intentionally throws only 1 error message in order to prevent test set probing. The more granular the error types, the more possible it is to use submissions to effectively probe. A simple example: say you want to know the rough # of new users in test and there are 10 error types possible -- you can come up with 10 buckets and have your kernel throw an error type corresponding to each, so then use only 1 sub to quickly narrow down the number of new users to one of those 10 buckets. So lack of granularity is horrible for debugging, but does come with the big upside of helping to insulate the competition from probing risks.  </p>",
      "rawMarkdown": "I have mixed feelings here. On the one hand, I've had the same frustrations as you -- debugging when you can't see the traceback or even disambiguate between programming and memory errors is an obvious nightmare. On the other hand, this *is* a \"code competition\" after all -- it's good to be forced into a more real-world scenario where designing your own rigorous test cases and measuring your pipeline's efficiency before deployment is critical. Certainly frustrating, but it makes for good habit building in a way that a traditional competition doesn't. I'm trying to think of it as an opportunity to practice more coding/engineering skills instead of just an annoying obstacle. \n\nI think it would have been a nice compromise to give us access to a larger and more representative practice test set to parse with the API, but at least we have tons of data to use for pipeline testing. Again, more realistic to have to design these tests from scratch than to be handed them for free.\n\nRe: the \"submission scoring error\", I believe that kaggle intentionally throws only 1 error message in order to prevent test set probing. The more granular the error types, the more possible it is to use submissions to effectively probe. A simple example: say you want to know the rough # of new users in test and there are 10 error types possible -- you can come up with 10 buckets and have your kernel throw an error type corresponding to each, so then use only 1 sub to quickly narrow down the number of new users to one of those 10 buckets. So lack of granularity is horrible for debugging, but does come with the big upside of helping to insulate the competition from probing risks.  ",
      "votes": 5,
      "replies": [
        {
          "id": 1052754,
          "postDate": "2020-10-18T08:03:10.743Z",
          "rawMarkdown": "",
          "votes": -1,
          "isDeleted": true
        }
      ]
    },
    {
      "id": 1052219,
      "postDate": "2020-10-17T13:19:31.230Z",
      "content": "<p>I am facing the same issue. Trying to make my first valid submission since yesterday, but no luck. But did get a few good pointers from the discussion here. Hopefully will be able to figure out what's the issue and get started on LB.</p>",
      "rawMarkdown": "I am facing the same issue. Trying to make my first valid submission since yesterday, but no luck. But did get a few good pointers from the discussion here. Hopefully will be able to figure out what's the issue and get started on LB.",
      "votes": 2
    },
    {
      "id": 1052473,
      "postDate": "2020-10-17T19:07:56.230Z",
      "content": "<p>I get the same error </p>",
      "rawMarkdown": "I get the same error "
    },
    {
      "id": 1052173,
      "postDate": "2020-10-17T12:18:49.760Z",
      "content": "<p>I'm struggling with submission errors too. The \"submission scoring error\" is the only error i have seen so far, although mine never went beyond 2mins for lb submission : (  (except simple basic model that doesn't interact with user history too much, these were ok)<br>\nthe error doesnt seem like to be reflecting what <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">kaggle page</a> said about this error:</p>\n<blockquote>\n  <p>Submission Scoring Error: Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.</p>\n</blockquote>\n<p>Also I agree the given example test (csv or with api) is not that helpful, I even tried to manually modify it to hopefully cover some cases other than all questions, but it still passed my feature engineering code yet fails in 2mins during submission, maybe i need to include more test cases lol …</p>",
      "rawMarkdown": "I'm struggling with submission errors too. The \"submission scoring error\" is the only error i have seen so far, although mine never went beyond 2mins for lb submission : (  (except simple basic model that doesn't interact with user history too much, these were ok)\nthe error doesnt seem like to be reflecting what [kaggle page](https://www.kaggle.com/code-competition-debugging) said about this error:\n> Submission Scoring Error: Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.\n\nAlso I agree the given example test (csv or with api) is not that helpful, I even tried to manually modify it to hopefully cover some cases other than all questions, but it still passed my feature engineering code yet fails in 2mins during submission, maybe i need to include more test cases lol ...\n",
      "replies": [
        {
          "id": 1052176,
          "postDate": "2020-10-17T12:23:02.487Z",
          "content": "<p>Glad I'm not the only one…<br>\nA bug I fixed a while earlier is that the list of previous answers which is returned with each test_df is as long as the previous test_df including lectures (i.e., an answer choice and correctness is given back now only for question rows, but also for lecture rows). That was a bit unintuitive and gave me quite a headache until I found it…. maybe it's helpful. <br>\nOtherwise no idea why it would fail so quickly.</p>",
          "rawMarkdown": "Glad I'm not the only one...\nA bug I fixed a while earlier is that the list of previous answers which is returned with each test_df is as long as the previous test_df including lectures (i.e., an answer choice and correctness is given back now only for question rows, but also for lecture rows). That was a bit unintuitive and gave me quite a headache until I found it.... maybe it's helpful. \nOtherwise no idea why it would fail so quickly.",
          "votes": 3
        },
        {
          "id": 1052189,
          "postDate": "2020-10-17T12:37:35.107Z",
          "content": "<p>hmm… interesting, I assumed it would just simply skip it (i guess that was a bad assumption since in train they have -1), then that is one of reason why mine is failing, I'm filtering away lectures when assigning answered correctly value. Thanks, that was helpful. Will try it in 12 hours (out of submission now….)</p>",
          "rawMarkdown": "hmm... interesting, I assumed it would just simply skip it (i guess that was a bad assumption since in train they have -1), then that is one of reason why mine is failing, I'm filtering away lectures when assigning answered correctly value. Thanks, that was helpful. Will try it in 12 hours (out of submission now....)",
          "votes": 1
        },
        {
          "id": 1052199,
          "postDate": "2020-10-17T12:44:54.207Z",
          "content": "<p>Good luck!</p>",
          "rawMarkdown": "Good luck!"
        },
        {
          "id": 1052203,
          "postDate": "2020-10-17T12:51:49.910Z",
          "content": "<p>I agree 100% but maybe we are over-complicating things as well?</p>\n<p>I was trying this, aggregating the test_df's so that i don't call train on every batch. It passes the dummy test but fails fast. Not quite sure why it's happening but i have invested ~15+ hrs now just to get this working on sample test and possible private test set (few_chunks). I can share what i did so as to save other's time… Do you guys know if \"answer_correctness\" can be null at the first row at random as well? I am not handling that. </p>",
          "rawMarkdown": "I agree 100% but maybe we are over-complicating things as well?\n\nI was trying this, aggregating the test_df's so that i don't call train on every batch. It passes the dummy test but fails fast. Not quite sure why it's happening but i have invested ~15+ hrs now just to get this working on sample test and possible private test set (few_chunks). I can share what i did so as to save other's time... Do you guys know if \"answer_correctness\" can be null at the first row at random as well? I am not handling that. "
        },
        {
          "id": 1052211,
          "postDate": "2020-10-17T13:05:33.873Z",
          "content": "<p>Hm my incremental training notebook works now and there I'm not handling any such cases…<br>\n<a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme</a><br>\nJust acting dumb and in every iteration setting</p>\n<p><code>previous_df['answered_correctly'] = np.array(eval(new_df.iloc[0]['prior_group_answers_correct']), dtype=np.int)</code></p>\n<p>seems to have worked out fine.<br>\nSo I think the test data is consistent in that it will always give you a set of answers fitting to the previous test chunk.</p>",
          "rawMarkdown": "Hm my incremental training notebook works now and there I'm not handling any such cases...\nhttps://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\nJust acting dumb and in every iteration setting\n\n`previous_df['answered_correctly'] = np.array(eval(new_df.iloc[0]['prior_group_answers_correct']), dtype=np.int)`\n\nseems to have worked out fine.\nSo I think the test data is consistent in that it will always give you a set of answers fitting to the previous test chunk.",
          "votes": 1
        },
        {
          "id": 1052213,
          "postDate": "2020-10-17T13:08:18.777Z",
          "content": "<p>Cool thanks! Here's what i did,</p>\n<p>NB it's buggy as of now but i am trying to see why it's failing… 🤐🤐🤐🤐;  Let me know if you make it working or see something which is wrong ;</p>\n<pre><code>import riiideducation\nenv = riiideducation.make_env()\niter_test = env.iter_test()\n\ndef batchify_retraining():\n    # inits\n    last_k_batches = pd.DataFrame()\n    _prior_group_answers_correct = [] # starting row is NULL\n    prev_answers = []\n    mod_by = 3 # tune-able\n    start_row_id, end_row_id = collections.deque(maxlen=mod_by), collections.deque(maxlen=mod_by)\n    count = -1\n\n    for idx, (test_df, sample_prediction_df) in enumerate(iter_test, start=1):        \n        # we have a batch of data with us now\n        current_group_num = idx - 1 # 0,1,2,3; idx is 1,2,3,4        \n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # &lt;re-training_block (kinda house-keeping stuff first 😅)&gt;\n            # add the correct label to the data first\n            # split the data into training and left-over as well\n            last_k_batches = last_k_batches.reset_index() # test_df returns with group_num as index set\n            # print(prev_answers, len(prev_answers), start_row_id, end_row_id, current_group_num, current_group_num-1, last_k_batches.shape)\n\n            # only consider 0,1 group_nums if % 3 is there as we don't have labels for the group_num 2 yet.\n            training_df, last_k_batches = last_k_batches[last_k_batches.group_num &lt; current_group_num-1], last_k_batches[last_k_batches.group_num &gt;= current_group_num-1] \n            # print(training_df.shape, last_k_batches.shape, len(prev_answers), last_k_batches)\n            # last_k_batches -&gt; future yet to get the labels for data.....\n            training_df.loc[:, \"prior_group_answers_correct\"] = prev_answers\n\n            # &lt;flushing the queue's as well for the processed group_nums&gt;\n            start_row_id_cp = list(start_row_id) # a copy\n            end_row_id_cp = list(end_row_id) # a copy\n            # print(start_row_id_cp, end_row_id_cp)\n            previous_answers = [] # reset previous answers\n            start_row_id = collections.deque([start_row_id_cp[-1]], maxlen=mod_by)\n            end_row_id = collections.deque([end_row_id_cp[-1]], maxlen=mod_by)\n\n            # print(start_row_id, end_row_id) # should point to last_k_batches stats\n            # print(\"re-training-the-model-now\")\n\n            training_df = training_df[training_df['content_type_id'] == 0] # remove lecture rows in re-training data\n\n            # &lt;re-train your model here all your re-training logic, new feats creation for training data etc should come here.\n\n        prev_answers += eval(test_df.iloc[0][\"prior_group_answers_correct\"])\n        _test_df = test_df.copy()\n        last_k_batches = pd.concat([last_k_batches, _test_df])\n        start_row_id.append(_test_df[_test_df['content_type_id'] == 0].head(1).row_id.values.tolist()[0])\n        end_row_id.append(_test_df[_test_df['content_type_id'] == 0].tail(1).row_id.values.tolist()[0])\n\n        # **** preds have to be made for each-block ****\n\n        _test_df = _test_df[_test_df['content_type_id'] == 0] # remove lecture rows in test data\n        label_names, X = get_features(_test_df.copy()) # get features\n\n        # &lt;predict all models section&gt;\n\n        # submit predictions section\n        test_df['answered_correctly'] = your_preds\n        env.predict(test_df.loc[:,['row_id', 'answered_correctly']])\n</code></pre>",
          "rawMarkdown": "Cool thanks! Here's what i did,\n\nNB it's buggy as of now but i am trying to see why it's failing... 🤐🤐🤐🤐;  Let me know if you make it working or see something which is wrong ;\n\n```\nimport riiideducation\nenv = riiideducation.make_env()\niter_test = env.iter_test()\n\ndef batchify_retraining():\n    # inits\n    last_k_batches = pd.DataFrame()\n    _prior_group_answers_correct = [] # starting row is NULL\n    prev_answers = []\n    mod_by = 3 # tune-able\n    start_row_id, end_row_id = collections.deque(maxlen=mod_by), collections.deque(maxlen=mod_by)\n    count = -1\n    \n    for idx, (test_df, sample_prediction_df) in enumerate(iter_test, start=1):        \n        # we have a batch of data with us now\n        current_group_num = idx - 1 # 0,1,2,3; idx is 1,2,3,4        \n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # <re-training_block (kinda house-keeping stuff first 😅)>\n            # add the correct label to the data first\n            # split the data into training and left-over as well\n            last_k_batches = last_k_batches.reset_index() # test_df returns with group_num as index set\n            # print(prev_answers, len(prev_answers), start_row_id, end_row_id, current_group_num, current_group_num-1, last_k_batches.shape)\n            \n            # only consider 0,1 group_nums if % 3 is there as we don't have labels for the group_num 2 yet.\n            training_df, last_k_batches = last_k_batches[last_k_batches.group_num < current_group_num-1], last_k_batches[last_k_batches.group_num >= current_group_num-1] \n            # print(training_df.shape, last_k_batches.shape, len(prev_answers), last_k_batches)\n            # last_k_batches -> future yet to get the labels for data.....\n            training_df.loc[:, \"prior_group_answers_correct\"] = prev_answers\n            \n            # <flushing the queue's as well for the processed group_nums>\n            start_row_id_cp = list(start_row_id) # a copy\n            end_row_id_cp = list(end_row_id) # a copy\n            # print(start_row_id_cp, end_row_id_cp)\n            previous_answers = [] # reset previous answers\n            start_row_id = collections.deque([start_row_id_cp[-1]], maxlen=mod_by)\n            end_row_id = collections.deque([end_row_id_cp[-1]], maxlen=mod_by)\n            \n            # print(start_row_id, end_row_id) # should point to last_k_batches stats\n            # print(\"re-training-the-model-now\")\n            \n            training_df = training_df[training_df['content_type_id'] == 0] # remove lecture rows in re-training data\n\n            # <re-train your model here all your re-training logic, new feats creation for training data etc should come here.\n        \n        prev_answers += eval(test_df.iloc[0][\"prior_group_answers_correct\"])\n        _test_df = test_df.copy()\n        last_k_batches = pd.concat([last_k_batches, _test_df])\n        start_row_id.append(_test_df[_test_df['content_type_id'] == 0].head(1).row_id.values.tolist()[0])\n        end_row_id.append(_test_df[_test_df['content_type_id'] == 0].tail(1).row_id.values.tolist()[0])\n        \n        # **** preds have to be made for each-block ****\n\n        _test_df = _test_df[_test_df['content_type_id'] == 0] # remove lecture rows in test data\n        label_names, X = get_features(_test_df.copy()) # get features\n\n        # <predict all models section>\n\n        # submit predictions section\n        test_df['answered_correctly'] = your_preds\n        env.predict(test_df.loc[:,['row_id', 'answered_correctly']])\n```"
        },
        {
          "id": 1052222,
          "postDate": "2020-10-17T13:26:55.753Z",
          "content": "<p>Hm overall it looks fine to me (and it probably should be, if it's working on the dummy test set. An issue I could imagine coming up (if I didn't miss anything in the code, that is) is memory. You don't seem to discard any previous information held in <code>last_k_batches</code> and <code>prev_answers</code>.<br>\nSo if you try this as kind of a dummy submission without actually calculating any predictions or features, the iterations would go quite fast and concatenating / appending would let your memory usage explode within minutes I'd say. So it might be beneficial to empty these two variables once you're done using the last X batches for feature engineering and re-training (if that's what you plan on doing). I.e., at the end of your if loop where it says <br>\n<code># &lt;re-train your model here all your re-training logic, new feats creation for training data etc should come here.</code></p>",
          "rawMarkdown": "Hm overall it looks fine to me (and it probably should be, if it's working on the dummy test set. An issue I could imagine coming up (if I didn't miss anything in the code, that is) is memory. You don't seem to discard any previous information held in `last_k_batches` and `prev_answers`.\nSo if you try this as kind of a dummy submission without actually calculating any predictions or features, the iterations would go quite fast and concatenating / appending would let your memory usage explode within minutes I'd say. So it might be beneficial to empty these two variables once you're done using the last X batches for feature engineering and re-training (if that's what you plan on doing). I.e., at the end of your if loop where it says \n`# <re-train your model here all your re-training logic, new feats creation for training data etc should come here.`",
          "votes": 1
        },
        {
          "id": 1052227,
          "postDate": "2020-10-17T13:33:05.170Z",
          "content": "<p>Yep, we can delete a bunch of things for sure, (thanks for the tip!) but mem shouldn't be the issue as it's like i have lot of them to spare. And i was doing concatenation of only 3 chunks as a starter…</p>\n<p><code>training_df, last_k_batches = last_k_batches[last_k_batches.group_num &lt; current_group_num-1], last_k_batches[last_k_batches.group_num &gt;= current_group_num-1]</code></p>\n<p>So last_k_batches does get updated (expanded and shrinked over time) but i can delete the training data for sure! If you see anything else, let me know! Thanks</p>\n<blockquote>\n  <p>So if you try this as kind of a dummy submission without actually calculating any predictions or features</p>\n</blockquote>\n<p>Thanks! Will do;</p>",
          "rawMarkdown": "Yep, we can delete a bunch of things for sure, (thanks for the tip!) but mem shouldn't be the issue as it's like i have lot of them to spare. And i was doing concatenation of only 3 chunks as a starter...\n\n`training_df, last_k_batches = last_k_batches[last_k_batches.group_num < current_group_num-1], last_k_batches[last_k_batches.group_num >= current_group_num-1] `\n\nSo last_k_batches does get updated (expanded and shrinked over time) but i can delete the training data for sure! If you see anything else, let me know! Thanks\n\n>So if you try this as kind of a dummy submission without actually calculating any predictions or features\n\nThanks! Will do;"
        },
        {
          "id": 1052234,
          "postDate": "2020-10-17T13:43:59.253Z",
          "content": "<p>Ah that's where that happens, must have skipped over it while reading! Then memory should be fine I would think yeah… I saw lots of other kernels using the whole training set without getting into trouble…</p>",
          "rawMarkdown": "Ah that's where that happens, must have skipped over it while reading! Then memory should be fine I would think yeah... I saw lots of other kernels using the whole training set without getting into trouble..."
        },
        {
          "id": 1052410,
          "postDate": "2020-10-17T17:27:59.540Z",
          "content": "<p>Wow, nice jumps! Guess you found a golden feature! Need to mimic as to why it breaks :(</p>",
          "rawMarkdown": "Wow, nice jumps! Guess you found a golden feature! Need to mimic as to why it breaks :("
        },
        {
          "id": 1053421,
          "postDate": "2020-10-19T01:41:46.983Z",
          "content": "<p>I think i found my bug! Thanks a lot :)</p>",
          "rawMarkdown": "I think i found my bug! Thanks a lot :)"
        },
        {
          "id": 1053529,
          "postDate": "2020-10-19T05:13:22.020Z",
          "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> here you go! It's much simpler than i thought :) <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856</a></p>",
          "rawMarkdown": "@spacelx here you go! It's much simpler than i thought :) https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856",
          "votes": 1
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 1052408,
      "author_name": "Joe Eddy",
      "author_url": "",
      "post_date": "2020-10-17T17:25:59.133000",
      "content": "<p>I have mixed feelings here. On the one hand, I've had the same frustrations as you -- debugging when you can't see the traceback or even disambiguate between programming and memory errors is an obvious nightmare. On the other hand, this <em>is</em> a \"code competition\" after all -- it's good to be forced into a more real-world scenario where designing your own rigorous test cases and measuring your pipeline's efficiency before deployment is critical. Certainly frustrating, but it makes for good habit building in a way that a traditional competition doesn't. I'm trying to think of it as an opportunity to practice more coding/engineering skills instead of just an annoying obstacle. </p>\n<p>I think it would have been a nice compromise to give us access to a larger and more representative practice test set to parse with the API, but at least we have tons of data to use for pipeline testing. Again, more realistic to have to design these tests from scratch than to be handed them for free.</p>\n<p>Re: the \"submission scoring error\", I believe that kaggle intentionally throws only 1 error message in order to prevent test set probing. The more granular the error types, the more possible it is to use submissions to effectively probe. A simple example: say you want to know the rough # of new users in test and there are 10 error types possible -- you can come up with 10 buckets and have your kernel throw an error type corresponding to each, so then use only 1 sub to quickly narrow down the number of new users to one of those 10 buckets. So lack of granularity is horrible for debugging, but does come with the big upside of helping to insulate the competition from probing risks.  </p>",
      "votes": 5,
      "replies": [
        {
          "id": 1052754,
          "author_name": "",
          "author_url": "",
          "post_date": "2020-10-18T08:03:10.743000",
          "content": "",
          "votes": -1,
          "replies": []
        }
      ]
    },
    {
      "id": 1052219,
      "author_name": "sandy1112",
      "author_url": "",
      "post_date": "2020-10-17T13:19:31.230000",
      "content": "<p>I am facing the same issue. Trying to make my first valid submission since yesterday, but no luck. But did get a few good pointers from the discussion here. Hopefully will be able to figure out what's the issue and get started on LB.</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 1052473,
      "author_name": "CherubRock",
      "author_url": "",
      "post_date": "2020-10-17T19:07:56.230000",
      "content": "<p>I get the same error </p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 1052173,
      "author_name": "samshipengs",
      "author_url": "",
      "post_date": "2020-10-17T12:18:49.760000",
      "content": "<p>I'm struggling with submission errors too. The \"submission scoring error\" is the only error i have seen so far, although mine never went beyond 2mins for lb submission : (  (except simple basic model that doesn't interact with user history too much, these were ok)<br>\nthe error doesnt seem like to be reflecting what <a href=\"https://www.kaggle.com/code-competition-debugging\" target=\"_blank\">kaggle page</a> said about this error:</p>\n<blockquote>\n  <p>Submission Scoring Error: Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.</p>\n</blockquote>\n<p>Also I agree the given example test (csv or with api) is not that helpful, I even tried to manually modify it to hopefully cover some cases other than all questions, but it still passed my feature engineering code yet fails in 2mins during submission, maybe i need to include more test cases lol …</p>",
      "votes": 0,
      "replies": [
        {
          "id": 1052176,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-17T12:23:02.487000",
          "content": "<p>Glad I'm not the only one…<br>\nA bug I fixed a while earlier is that the list of previous answers which is returned with each test_df is as long as the previous test_df including lectures (i.e., an answer choice and correctness is given back now only for question rows, but also for lecture rows). That was a bit unintuitive and gave me quite a headache until I found it…. maybe it's helpful. <br>\nOtherwise no idea why it would fail so quickly.</p>",
          "votes": 3,
          "replies": []
        },
        {
          "id": 1052189,
          "author_name": "samshipengs",
          "author_url": "",
          "post_date": "2020-10-17T12:37:35.107000",
          "content": "<p>hmm… interesting, I assumed it would just simply skip it (i guess that was a bad assumption since in train they have -1), then that is one of reason why mine is failing, I'm filtering away lectures when assigning answered correctly value. Thanks, that was helpful. Will try it in 12 hours (out of submission now….)</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1052199,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-17T12:44:54.207000",
          "content": "<p>Good luck!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052203,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-17T12:51:49.910000",
          "content": "<p>I agree 100% but maybe we are over-complicating things as well?</p>\n<p>I was trying this, aggregating the test_df's so that i don't call train on every batch. It passes the dummy test but fails fast. Not quite sure why it's happening but i have invested ~15+ hrs now just to get this working on sample test and possible private test set (few_chunks). I can share what i did so as to save other's time… Do you guys know if \"answer_correctness\" can be null at the first row at random as well? I am not handling that. </p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052211,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-17T13:05:33.873000",
          "content": "<p>Hm my incremental training notebook works now and there I'm not handling any such cases…<br>\n<a href=\"https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme\" target=\"_blank\">https://www.kaggle.com/spacelx/2020-r3id-incremental-learning-pytorch-creme</a><br>\nJust acting dumb and in every iteration setting</p>\n<p><code>previous_df['answered_correctly'] = np.array(eval(new_df.iloc[0]['prior_group_answers_correct']), dtype=np.int)</code></p>\n<p>seems to have worked out fine.<br>\nSo I think the test data is consistent in that it will always give you a set of answers fitting to the previous test chunk.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1052213,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-17T13:08:18.777000",
          "content": "<p>Cool thanks! Here's what i did,</p>\n<p>NB it's buggy as of now but i am trying to see why it's failing… 🤐🤐🤐🤐;  Let me know if you make it working or see something which is wrong ;</p>\n<pre><code>import riiideducation\nenv = riiideducation.make_env()\niter_test = env.iter_test()\n\ndef batchify_retraining():\n    # inits\n    last_k_batches = pd.DataFrame()\n    _prior_group_answers_correct = [] # starting row is NULL\n    prev_answers = []\n    mod_by = 3 # tune-able\n    start_row_id, end_row_id = collections.deque(maxlen=mod_by), collections.deque(maxlen=mod_by)\n    count = -1\n\n    for idx, (test_df, sample_prediction_df) in enumerate(iter_test, start=1):        \n        # we have a batch of data with us now\n        current_group_num = idx - 1 # 0,1,2,3; idx is 1,2,3,4        \n        if current_group_num % mod_by == 0 and current_group_num != 0:\n            # &lt;re-training_block (kinda house-keeping stuff first 😅)&gt;\n            # add the correct label to the data first\n            # split the data into training and left-over as well\n            last_k_batches = last_k_batches.reset_index() # test_df returns with group_num as index set\n            # print(prev_answers, len(prev_answers), start_row_id, end_row_id, current_group_num, current_group_num-1, last_k_batches.shape)\n\n            # only consider 0,1 group_nums if % 3 is there as we don't have labels for the group_num 2 yet.\n            training_df, last_k_batches = last_k_batches[last_k_batches.group_num &lt; current_group_num-1], last_k_batches[last_k_batches.group_num &gt;= current_group_num-1] \n            # print(training_df.shape, last_k_batches.shape, len(prev_answers), last_k_batches)\n            # last_k_batches -&gt; future yet to get the labels for data.....\n            training_df.loc[:, \"prior_group_answers_correct\"] = prev_answers\n\n            # &lt;flushing the queue's as well for the processed group_nums&gt;\n            start_row_id_cp = list(start_row_id) # a copy\n            end_row_id_cp = list(end_row_id) # a copy\n            # print(start_row_id_cp, end_row_id_cp)\n            previous_answers = [] # reset previous answers\n            start_row_id = collections.deque([start_row_id_cp[-1]], maxlen=mod_by)\n            end_row_id = collections.deque([end_row_id_cp[-1]], maxlen=mod_by)\n\n            # print(start_row_id, end_row_id) # should point to last_k_batches stats\n            # print(\"re-training-the-model-now\")\n\n            training_df = training_df[training_df['content_type_id'] == 0] # remove lecture rows in re-training data\n\n            # &lt;re-train your model here all your re-training logic, new feats creation for training data etc should come here.\n\n        prev_answers += eval(test_df.iloc[0][\"prior_group_answers_correct\"])\n        _test_df = test_df.copy()\n        last_k_batches = pd.concat([last_k_batches, _test_df])\n        start_row_id.append(_test_df[_test_df['content_type_id'] == 0].head(1).row_id.values.tolist()[0])\n        end_row_id.append(_test_df[_test_df['content_type_id'] == 0].tail(1).row_id.values.tolist()[0])\n\n        # **** preds have to be made for each-block ****\n\n        _test_df = _test_df[_test_df['content_type_id'] == 0] # remove lecture rows in test data\n        label_names, X = get_features(_test_df.copy()) # get features\n\n        # &lt;predict all models section&gt;\n\n        # submit predictions section\n        test_df['answered_correctly'] = your_preds\n        env.predict(test_df.loc[:,['row_id', 'answered_correctly']])\n</code></pre>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052222,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-17T13:26:55.753000",
          "content": "<p>Hm overall it looks fine to me (and it probably should be, if it's working on the dummy test set. An issue I could imagine coming up (if I didn't miss anything in the code, that is) is memory. You don't seem to discard any previous information held in <code>last_k_batches</code> and <code>prev_answers</code>.<br>\nSo if you try this as kind of a dummy submission without actually calculating any predictions or features, the iterations would go quite fast and concatenating / appending would let your memory usage explode within minutes I'd say. So it might be beneficial to empty these two variables once you're done using the last X batches for feature engineering and re-training (if that's what you plan on doing). I.e., at the end of your if loop where it says <br>\n<code># &lt;re-train your model here all your re-training logic, new feats creation for training data etc should come here.</code></p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 1052227,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-17T13:33:05.170000",
          "content": "<p>Yep, we can delete a bunch of things for sure, (thanks for the tip!) but mem shouldn't be the issue as it's like i have lot of them to spare. And i was doing concatenation of only 3 chunks as a starter…</p>\n<p><code>training_df, last_k_batches = last_k_batches[last_k_batches.group_num &lt; current_group_num-1], last_k_batches[last_k_batches.group_num &gt;= current_group_num-1]</code></p>\n<p>So last_k_batches does get updated (expanded and shrinked over time) but i can delete the training data for sure! If you see anything else, let me know! Thanks</p>\n<blockquote>\n  <p>So if you try this as kind of a dummy submission without actually calculating any predictions or features</p>\n</blockquote>\n<p>Thanks! Will do;</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052234,
          "author_name": "Alex Bader",
          "author_url": "",
          "post_date": "2020-10-17T13:43:59.253000",
          "content": "<p>Ah that's where that happens, must have skipped over it while reading! Then memory should be fine I would think yeah… I saw lots of other kernels using the whole training set without getting into trouble…</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1052410,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-17T17:27:59.540000",
          "content": "<p>Wow, nice jumps! Guess you found a golden feature! Need to mimic as to why it breaks :(</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053421,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-19T01:41:46.983000",
          "content": "<p>I think i found my bug! Thanks a lot :)</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 1053529,
          "author_name": "Aditya Soni",
          "author_url": "",
          "post_date": "2020-10-19T05:13:22.020000",
          "content": "<p><a href=\"https://www.kaggle.com/spacelx\" target=\"_blank\">@spacelx</a> here you go! It's much simpler than i thought :) <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/191856</a></p>",
          "votes": 1,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "1052086": "If I submit a notebook which throws an exception after a few seconds, it reports back \"submission scoring error\". Another of my notebooks runs for ~9 hours (so probably times out) but also gives me a \"submission scoring error\". Another runs for some hours (not sure if it hit the time limit) and also gives me this error (the code being literally copied from a kernel of mine which was accepted without errors).\n\nWhat's going on here? Are these the only error messages we're getting in this competition? Besides being zero informative, they seem to be somewhat random as well. As if it wasn't bad enough that the public test set which we're supposed to use for debugging is somewhat useless (no lecture rows, not of a representative size, ... ).\n\nSorry for venting but I do feel like the setup of this competition is quite frustrating overall.",
    "1052408": "I have mixed feelings here. On the one hand, I've had the same frustrations as you -- debugging when you can't see the traceback or even disambiguate between programming and memory errors is an obvious nightmare. On the other hand, this *is* a \"code competition\" after all -- it's good to be forced into a more real-world scenario where designing your own rigorous test cases and measuring your pipeline's efficiency before deployment is critical. Certainly frustrating, but it makes for good habit building in a way that a traditional competition doesn't. I'm trying to think of it as an opportunity to practice more coding/engineering skills instead of just an annoying obstacle. \n\nI think it would have been a nice compromise to give us access to a larger and more representative practice test set to parse with the API, but at least we have tons of data to use for pipeline testing. Again, more realistic to have to design these tests from scratch than to be handed them for free.\n\nRe: the \"submission scoring error\", I believe that kaggle intentionally throws only 1 error message in order to prevent test set probing. The more granular the error types, the more possible it is to use submissions to effectively probe. A simple example: say you want to know the rough # of new users in test and there are 10 error types possible -- you can come up with 10 buckets and have your kernel throw an error type corresponding to each, so then use only 1 sub to quickly narrow down the number of new users to one of those 10 buckets. So lack of granularity is horrible for debugging, but does come with the big upside of helping to insulate the competition from probing risks.  ",
    "1052219": "I am facing the same issue. Trying to make my first valid submission since yesterday, but no luck. But did get a few good pointers from the discussion here. Hopefully will be able to figure out what's the issue and get started on LB.",
    "1052473": "I get the same error ",
    "1052173": "I'm struggling with submission errors too. The \"submission scoring error\" is the only error i have seen so far, although mine never went beyond 2mins for lb submission : (  (except simple basic model that doesn't interact with user history too much, these were ok)\nthe error doesnt seem like to be reflecting what [kaggle page](https://www.kaggle.com/code-competition-debugging) said about this error:\n> Submission Scoring Error: Your notebook generated a submission file with incorrect format. Some examples causing this are: wrong number of rows or columns, empty values, an incorrect data type for a value, or invalid submission values from what is expected.\n\nAlso I agree the given example test (csv or with api) is not that helpful, I even tried to manually modify it to hopefully cover some cases other than all questions, but it still passed my feature engineering code yet fails in 2mins during submission, maybe i need to include more test cases lol ...\n"
  }
}