{
  "id": 410544,
  "title": "Difference between test set and hidden test set",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/410544",
  "author_name": "GuirLP",
  "post_date": "2023-05-15T18:02:47.752000",
  "votes": 0,
  "comment_count": 0,
  "views": 0,
  "content": "<p>Hi, when I submit my notebook I always get an error on the assert line. How is the hidden data set different from the training set? I have a \"Submission Scoring Error\" even though my csv has the correct number of rows.<br>\nI've done 80 submissions and I still haven't found the problem. <br>\nDoes anyone have any ideas?</p>\n<p>Thanks</p>\n<p>The program crashes at the assert level when it shouldn't<br>\nHere is my submission code: </p>\n<pre><code>for (test, sample_submission) in iter_test:  \n    seq = test.level_group.values[0]\n    test = test[[\"session_id\", \"index\", \"level\", \"level_group\", \"event_name\", \"room_fqid\", \"name\", \"elapsed_time\"]]\n    test = test[test[\"level_group\"] == seq].reset_index().drop([\"level_0\"], axis=1)\n    test = test[test[\"level\"] &lt; 18]\n\n    if seq == \"13-22\":\n        assert len(test[\"level\"].unique()) &gt; 5\n    sample_submission[\"correct\"] = 1\n\n    ## submission\n\n    env.predict(sample_submission)\n</code></pre>",
  "messages": [
    {
      "id": 2260550,
      "postDate": "2023-05-15T18:02:47.753Z",
      "content": "<p>Hi, when I submit my notebook I always get an error on the assert line. How is the hidden data set different from the training set? I have a \"Submission Scoring Error\" even though my csv has the correct number of rows.<br>\nI've done 80 submissions and I still haven't found the problem. <br>\nDoes anyone have any ideas?</p>\n<p>Thanks</p>\n<p>The program crashes at the assert level when it shouldn't<br>\nHere is my submission code: </p>\n<pre><code>for (test, sample_submission) in iter_test:  \n    seq = test.level_group.values[0]\n    test = test[[\"session_id\", \"index\", \"level\", \"level_group\", \"event_name\", \"room_fqid\", \"name\", \"elapsed_time\"]]\n    test = test[test[\"level_group\"] == seq].reset_index().drop([\"level_0\"], axis=1)\n    test = test[test[\"level\"] &lt; 18]\n\n    if seq == \"13-22\":\n        assert len(test[\"level\"].unique()) &gt; 5\n    sample_submission[\"correct\"] = 1\n\n    ## submission\n\n    env.predict(sample_submission)\n</code></pre>",
      "rawMarkdown": "Hi, when I submit my notebook I always get an error on the assert line. How is the hidden data set different from the training set? I have a \"Submission Scoring Error\" even though my csv has the correct number of rows.\nI've done 80 submissions and I still haven't found the problem. \nDoes anyone have any ideas?\n\nThanks\n\nThe program crashes at the assert level when it shouldn't\nHere is my submission code: \n```\nfor (test, sample_submission) in iter_test:  \n    seq = test.level_group.values[0]\n    test = test[[\"session_id\", \"index\", \"level\", \"level_group\", \"event_name\", \"room_fqid\", \"name\", \"elapsed_time\"]]\n    test = test[test[\"level_group\"] == seq].reset_index().drop([\"level_0\"], axis=1)\n    test = test[test[\"level\"] < 18]\n\n    if seq == \"13-22\":\n        assert len(test[\"level\"].unique()) > 5\n    sample_submission[\"correct\"] = 1\n    \n    ## submission\n    \n    env.predict(sample_submission)\n```"
    }
  ],
  "comments": [],
  "raw_markdown_by_id": {
    "2260550": "Hi, when I submit my notebook I always get an error on the assert line. How is the hidden data set different from the training set? I have a \"Submission Scoring Error\" even though my csv has the correct number of rows.\nI've done 80 submissions and I still haven't found the problem. \nDoes anyone have any ideas?\n\nThanks\n\nThe program crashes at the assert level when it shouldn't\nHere is my submission code: \n```\nfor (test, sample_submission) in iter_test:  \n    seq = test.level_group.values[0]\n    test = test[[\"session_id\", \"index\", \"level\", \"level_group\", \"event_name\", \"room_fqid\", \"name\", \"elapsed_time\"]]\n    test = test[test[\"level_group\"] == seq].reset_index().drop([\"level_0\"], axis=1)\n    test = test[test[\"level\"] < 18]\n\n    if seq == \"13-22\":\n        assert len(test[\"level\"].unique()) > 5\n    sample_submission[\"correct\"] = 1\n    \n    ## submission\n    \n    env.predict(sample_submission)\n```"
  }
}