{
  "id": 204796,
  "title": "Painful Submission Scoring Error!",
  "url": "/competitions/riiid-test-answer-prediction/discussion/204796",
  "author_name": "Moiz",
  "post_date": "2020-12-16T21:46:55.357000",
  "votes": 1,
  "comment_count": 9,
  "views": 0,
  "content": "<p>Hello everyone!<br>\nIs there a way to kindly request the organizers to include more group nums in test Api. Only 108 observations are just not enough to fix all the errors during the submission process. After spending countless days working on this project and coming up with a decent CV, it is extremely painful to see the \"Submission Scoring Error\" again and again. The 4 sample groups don't even have lectures data so it's very hard to see where and why I am getting the error. The test Api in Jane Street Competition is pretty exchaustive and has more than 15000 observations which eliminates the possibility of \"Submission Scoring Error\".</p>\n<p>Is there a way we can have same thing for this competition, for example 500-1000 group nums for the initial Api submission. Please, thank you.</p>",
  "messages": [
    {
      "id": 1116216,
      "postDate": "2020-12-17T00:52:16.517Z",
      "content": "<p>Can you try the suggestion I made in another thread (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203238</a>). It might help you see where your issue is.</p>\n<blockquote>\n  <p>So I am assuming that you can run your test for-loop (in predict_2.JPG) locally but when you submit to the LB you get an error. If so, can you try placing this snippet of code at the very beginning of your testing for-loop:</p>\n<pre><code>test['content_type_id'] = np.random.randint(0, 2, len(test))\ntest['user_id'] = 1111111111\ntest['content_id'] = 222222222\n</code></pre>\n  <p>If you now get an error running it locally, then you would be able to see where your issue is. Depending on which line gives the error, you either are getting wrong dataframe shapes when encountering a content_type_id == 1, or you are not taking care of unseen users/contents properly.</p>\n  <p>Hope it helps. Good luck</p>\n</blockquote>",
      "rawMarkdown": "Can you try the suggestion I made in another thread ([https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203238](url)). It might help you see where your issue is.\n> So I am assuming that you can run your test for-loop (in predict_2.JPG) locally but when you submit to the LB you get an error. If so, can you try placing this snippet of code at the very beginning of your testing for-loop:\n> \n> ```\n> test['content_type_id'] = np.random.randint(0, 2, len(test))\n> test['user_id'] = 1111111111\n> test['content_id'] = 222222222\n> ```\n> \n> If you now get an error running it locally, then you would be able to see where your issue is. Depending on which line gives the error, you either are getting wrong dataframe shapes when encountering a content_type_id == 1, or you are not taking care of unseen users/contents properly.\n> \n> Hope it helps. Good luck",
      "votes": 1,
      "replies": [
        {
          "id": 1131815,
          "postDate": "2020-12-30T02:02:11.730Z",
          "content": "<p>Thank you for the help. I was able to solve my issue.</p>",
          "rawMarkdown": "Thank you for the help. I was able to solve my issue.",
          "votes": 1
        }
      ]
    },
    {
      "id": 1116127,
      "postDate": "2020-12-16T21:46:55.357Z",
      "content": "<p>Hello everyone!<br>\nIs there a way to kindly request the organizers to include more group nums in test Api. Only 108 observations are just not enough to fix all the errors during the submission process. After spending countless days working on this project and coming up with a decent CV, it is extremely painful to see the \"Submission Scoring Error\" again and again. The 4 sample groups don't even have lectures data so it's very hard to see where and why I am getting the error. The test Api in Jane Street Competition is pretty exchaustive and has more than 15000 observations which eliminates the possibility of \"Submission Scoring Error\".</p>\n<p>Is there a way we can have same thing for this competition, for example 500-1000 group nums for the initial Api submission. Please, thank you.</p>",
      "rawMarkdown": "Hello everyone!\nIs there a way to kindly request the organizers to include more group nums in test Api. Only 108 observations are just not enough to fix all the errors during the submission process. After spending countless days working on this project and coming up with a decent CV, it is extremely painful to see the \"Submission Scoring Error\" again and again. The 4 sample groups don't even have lectures data so it's very hard to see where and why I am getting the error. The test Api in Jane Street Competition is pretty exchaustive and has more than 15000 observations which eliminates the possibility of \"Submission Scoring Error\".\n\nIs there a way we can have same thing for this competition, for example 500-1000 group nums for the initial Api submission. Please, thank you.",
      "votes": 1
    },
    {
      "id": 1119608,
      "postDate": "2020-12-20T08:23:52.933Z",
      "content": "<p>I think submission has been a painful process for all of us. It took me more than a week before I could do a successful one.</p>\n<p>Check this thread for a collection of possible scoring errors: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124</a></p>\n<p>Also, check the official submission example code where they show how to ignore lectures.</p>\n<p>In my experience, timeout (&gt;9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.</p>\n<p>All I can say is strip your inference to the bare minimum (ie. submit 0.5 for all rows) and keep adding more of your code incrementally until it breaks. I had to use all 5 subs/day for a while until I ironed out all the bugs. The good thing is that when you remove the actual inference, the nb runs quickly enough so that you don't get desperate waiting for feedback.</p>",
      "rawMarkdown": "I think submission has been a painful process for all of us. It took me more than a week before I could do a successful one.\n\nCheck this thread for a collection of possible scoring errors: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\n\nAlso, check the official submission example code where they show how to ignore lectures.\n\nIn my experience, timeout (>9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.\n\nAll I can say is strip your inference to the bare minimum (ie. submit 0.5 for all rows) and keep adding more of your code incrementally until it breaks. I had to use all 5 subs/day for a while until I ironed out all the bugs. The good thing is that when you remove the actual inference, the nb runs quickly enough so that you don't get desperate waiting for feedback.\n\n",
      "votes": 2,
      "replies": [
        {
          "id": 1130409,
          "postDate": "2020-12-29T02:26:51.627Z",
          "content": "<blockquote>\n  <p>In my experience, timeout (&gt;9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.</p>\n</blockquote>\n<p>How did you verify that, I got both Scoring error and timeout error after 9 hours.</p>",
          "rawMarkdown": "> In my experience, timeout (>9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.\n\nHow did you verify that, I got both Scoring error and timeout error after 9 hours.",
          "replies": [
            {
              "id": 1131747,
              "postDate": "2020-12-29T23:49:39.547Z",
              "content": "<p>A working pipeline turned into 'scoring error' after 9h by simply upscaling the model size.</p>",
              "rawMarkdown": "A working pipeline turned into 'scoring error' after 9h by simply upscaling the model size.",
              "votes": 1
            }
          ]
        },
        {
          "id": 1131812,
          "postDate": "2020-12-30T02:00:56.347Z",
          "content": "<p>Thank you. I was able to debug my issue using this emulator. Amazing work<br>\n<a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator</a></p>",
          "rawMarkdown": "Thank you. I was able to debug my issue using this emulator. Amazing work\nhttps://www.kaggle.com/its7171/time-series-api-iter-test-emulator"
        }
      ]
    },
    {
      "id": 1117529,
      "postDate": "2020-12-18T07:06:42.027Z",
      "content": "<p>Might be this: Don't use pytorch/numpy squeeze on input data from the api stream since there is a batch with the size of 1, squeeze would remove this batch dimension thus causing an error. <br>\nThis lost me three entire days to debug.</p>",
      "rawMarkdown": "Might be this: Don't use pytorch/numpy squeeze on input data from the api stream since there is a batch with the size of 1, squeeze would remove this batch dimension thus causing an error. \nThis lost me three entire days to debug.",
      "votes": 2,
      "replies": [
        {
          "id": 1131813,
          "postDate": "2020-12-30T02:01:38.640Z",
          "content": "<p>Oh that's sad. I didn't use squeeze but thanks for the heads up.</p>",
          "rawMarkdown": "Oh that's sad. I didn't use squeeze but thanks for the heads up."
        }
      ]
    },
    {
      "id": 1117528,
      "postDate": "2020-12-18T07:06:42.027Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 1116216,
      "author_name": "Moeen Bagheri",
      "author_url": "",
      "post_date": "2020-12-17T00:52:16.517000",
      "content": "<p>Can you try the suggestion I made in another thread (<a href=\"url\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203238</a>). It might help you see where your issue is.</p>\n<blockquote>\n  <p>So I am assuming that you can run your test for-loop (in predict_2.JPG) locally but when you submit to the LB you get an error. If so, can you try placing this snippet of code at the very beginning of your testing for-loop:</p>\n<pre><code>test['content_type_id'] = np.random.randint(0, 2, len(test))\ntest['user_id'] = 1111111111\ntest['content_id'] = 222222222\n</code></pre>\n  <p>If you now get an error running it locally, then you would be able to see where your issue is. Depending on which line gives the error, you either are getting wrong dataframe shapes when encountering a content_type_id == 1, or you are not taking care of unseen users/contents properly.</p>\n  <p>Hope it helps. Good luck</p>\n</blockquote>",
      "votes": 1,
      "replies": [
        {
          "id": 1131815,
          "author_name": "Moiz",
          "author_url": "",
          "post_date": "2020-12-30T02:02:11.730000",
          "content": "<p>Thank you for the help. I was able to solve my issue.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 1119608,
      "author_name": "Javier Martín",
      "author_url": "",
      "post_date": "2020-12-20T08:23:52.933000",
      "content": "<p>I think submission has been a painful process for all of us. It took me more than a week before I could do a successful one.</p>\n<p>Check this thread for a collection of possible scoring errors: <a href=\"https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\" target=\"_blank\">https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124</a></p>\n<p>Also, check the official submission example code where they show how to ignore lectures.</p>\n<p>In my experience, timeout (&gt;9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.</p>\n<p>All I can say is strip your inference to the bare minimum (ie. submit 0.5 for all rows) and keep adding more of your code incrementally until it breaks. I had to use all 5 subs/day for a while until I ironed out all the bugs. The good thing is that when you remove the actual inference, the nb runs quickly enough so that you don't get desperate waiting for feedback.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1130409,
          "author_name": "Abdessalem Boukil",
          "author_url": "",
          "post_date": "2020-12-29T02:26:51.627000",
          "content": "<blockquote>\n  <p>In my experience, timeout (&gt;9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.</p>\n</blockquote>\n<p>How did you verify that, I got both Scoring error and timeout error after 9 hours.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 1131747,
              "author_name": "Javier Martín",
              "author_url": "",
              "post_date": "2020-12-29T23:49:39.547000",
              "content": "<p>A working pipeline turned into 'scoring error' after 9h by simply upscaling the model size.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 1131812,
          "author_name": "Moiz",
          "author_url": "",
          "post_date": "2020-12-30T02:00:56.347000",
          "content": "<p>Thank you. I was able to debug my issue using this emulator. Amazing work<br>\n<a href=\"https://www.kaggle.com/its7171/time-series-api-iter-test-emulator\" target=\"_blank\">https://www.kaggle.com/its7171/time-series-api-iter-test-emulator</a></p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117529,
      "author_name": "Abdessalem Boukil",
      "author_url": "",
      "post_date": "2020-12-18T07:06:42.027000",
      "content": "<p>Might be this: Don't use pytorch/numpy squeeze on input data from the api stream since there is a batch with the size of 1, squeeze would remove this batch dimension thus causing an error. <br>\nThis lost me three entire days to debug.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 1131813,
          "author_name": "Moiz",
          "author_url": "",
          "post_date": "2020-12-30T02:01:38.640000",
          "content": "<p>Oh that's sad. I didn't use squeeze but thanks for the heads up.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 1117528,
      "author_name": "",
      "author_url": "",
      "post_date": "2020-12-18T07:06:42.027000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "1116216": "Can you try the suggestion I made in another thread ([https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/203238](url)). It might help you see where your issue is.\n> So I am assuming that you can run your test for-loop (in predict_2.JPG) locally but when you submit to the LB you get an error. If so, can you try placing this snippet of code at the very beginning of your testing for-loop:\n> \n> ```\n> test['content_type_id'] = np.random.randint(0, 2, len(test))\n> test['user_id'] = 1111111111\n> test['content_id'] = 222222222\n> ```\n> \n> If you now get an error running it locally, then you would be able to see where your issue is. Depending on which line gives the error, you either are getting wrong dataframe shapes when encountering a content_type_id == 1, or you are not taking care of unseen users/contents properly.\n> \n> Hope it helps. Good luck",
    "1116127": "Hello everyone!\nIs there a way to kindly request the organizers to include more group nums in test Api. Only 108 observations are just not enough to fix all the errors during the submission process. After spending countless days working on this project and coming up with a decent CV, it is extremely painful to see the \"Submission Scoring Error\" again and again. The 4 sample groups don't even have lectures data so it's very hard to see where and why I am getting the error. The test Api in Jane Street Competition is pretty exchaustive and has more than 15000 observations which eliminates the possibility of \"Submission Scoring Error\".\n\nIs there a way we can have same thing for this competition, for example 500-1000 group nums for the initial Api submission. Please, thank you.",
    "1119608": "I think submission has been a painful process for all of us. It took me more than a week before I could do a successful one.\n\nCheck this thread for a collection of possible scoring errors: https://www.kaggle.com/c/riiid-test-answer-prediction/discussion/192124\n\nAlso, check the official submission example code where they show how to ignore lectures.\n\nIn my experience, timeout (>9h run time) also reflects as \"scoring error\" and not \"timeout error\" as in other competitions.\n\nAll I can say is strip your inference to the bare minimum (ie. submit 0.5 for all rows) and keep adding more of your code incrementally until it breaks. I had to use all 5 subs/day for a while until I ironed out all the bugs. The good thing is that when you remove the actual inference, the nb runs quickly enough so that you don't get desperate waiting for feedback.\n\n",
    "1117529": "Might be this: Don't use pytorch/numpy squeeze on input data from the api stream since there is a batch with the size of 1, squeeze would remove this batch dimension thus causing an error. \nThis lost me three entire days to debug.",
    "1117528": ""
  }
}