{
  "id": 405218,
  "title": "Scoring error",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/405218",
  "author_name": "",
  "post_date": "2023-04-26T14:56:29.276671100Z",
  "votes": 2,
  "comment_count": 5,
  "views": 0,
  "content": "<p>I'm a newbie to Kaggle and I don't know why I have the submission file format error.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F6258d7f991f61cb98df6f105c972d7ac%2F2023-04-26%20%2011.45.02.png?generation=1682520351897621&amp;alt=media\" alt=\"\"></p>\n<p>I followed the starter notebook instructing how to use submission API, except I made all predictions to '1' just for submission test purpose. Below is the code.</p>\n<pre><code>class my_model():\n    def predict(self, X):\n        return 1\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    test_df = feature_engineer(test)\n    grp = test_df.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        gbtm = my_model()\n        test_ds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df.loc[:, test_df.columns != 'level_group'])\n        predictions = gbtm.predict(test_ds)\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        n_predictions = np.array((predictions &gt; best_threshold)).astype(int)\n        sample_submission.loc[mask,'correct'] = n_predictions.flatten()\n\n    env.predict(sample_submission[['session_id', 'correct']])\n</code></pre>\n<p>I found one thing that the order of items in my submission file is different from that of the reference notebook. However, I don't have any clue why the orders are different.</p>\n<ol>\n<li>mine<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2Ff298e9a142419864b7468811588bbb66%2F2023-04-26%20%2011.51.16.png?generation=1682520825227694&amp;alt=media\" alt=\"\"></li>\n<li>ref<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F5fdaa82c21b8436b9420ed6160185539%2F2023-04-26%20%2011.51.56.png?generation=1682520840769585&amp;alt=media\" alt=\"\"></li>\n</ol>\n<p>Can anyone give me an advice?</p>",
  "messages": [
    {
      "id": "2236129",
      "postDate": "04/26/2023 14:56:29",
      "content": "<p>I'm a newbie to Kaggle and I don't know why I have the submission file format error.<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F6258d7f991f61cb98df6f105c972d7ac%2F2023-04-26%20%2011.45.02.png?generation=1682520351897621&amp;alt=media\" alt=\"\"></p>\n<p>I followed the starter notebook instructing how to use submission API, except I made all predictions to '1' just for submission test purpose. Below is the code.</p>\n<pre><code>class my_model():\n    def predict(self, X):\n        return 1\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    test_df = feature_engineer(test)\n    grp = test_df.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        gbtm = my_model()\n        test_ds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df.loc[:, test_df.columns != 'level_group'])\n        predictions = gbtm.predict(test_ds)\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        n_predictions = np.array((predictions &gt; best_threshold)).astype(int)\n        sample_submission.loc[mask,'correct'] = n_predictions.flatten()\n\n    env.predict(sample_submission[['session_id', 'correct']])\n</code></pre>\n<p>I found one thing that the order of items in my submission file is different from that of the reference notebook. However, I don't have any clue why the orders are different.</p>\n<ol>\n<li>mine<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2Ff298e9a142419864b7468811588bbb66%2F2023-04-26%20%2011.51.16.png?generation=1682520825227694&amp;alt=media\" alt=\"\"></li>\n<li>ref<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F5fdaa82c21b8436b9420ed6160185539%2F2023-04-26%20%2011.51.56.png?generation=1682520840769585&amp;alt=media\" alt=\"\"></li>\n</ol>\n<p>Can anyone give me an advice?</p>",
      "rawMarkdown": "I'm a newbie to Kaggle and I don't know why I have the submission file format error.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F6258d7f991f61cb98df6f105c972d7ac%2F2023-04-26%20%2011.45.02.png?generation=1682520351897621&alt=media)\n\nI followed the starter notebook instructing how to use submission API, except I made all predictions to '1' just for submission test purpose. Below is the code.\n```\nclass my_model():\n    def predict(self, X):\n        return 1\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    test_df = feature_engineer(test)\n    grp = test_df.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        gbtm = my_model()\n        test_ds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df.loc[:, test_df.columns != 'level_group'])\n        predictions = gbtm.predict(test_ds)\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        n_predictions = np.array((predictions > best_threshold)).astype(int)\n        sample_submission.loc[mask,'correct'] = n_predictions.flatten()\n    \n    env.predict(sample_submission[['session_id', 'correct']])\n```\n\nI found one thing that the order of items in my submission file is different from that of the reference notebook. However, I don't have any clue why the orders are different.\n1. mine\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2Ff298e9a142419864b7468811588bbb66%2F2023-04-26%20%2011.51.16.png?generation=1682520825227694&alt=media)\n2. ref\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F5fdaa82c21b8436b9420ed6160185539%2F2023-04-26%20%2011.51.56.png?generation=1682520840769585&alt=media)\n\nCan anyone give me an advice?",
      "votes": null
    },
    {
      "id": "2236647",
      "postDate": "04/27/2023 02:58:09",
      "content": "<p>I found that the order of items is not related this issue, as the reference notebook successfully submitted.</p>",
      "rawMarkdown": "I found that the order of items is not related this issue, as the reference notebook successfully submitted.",
      "votes": null
    },
    {
      "id": "2237038",
      "postDate": "04/27/2023 10:34:16",
      "content": "<p>I found a working example for submission success. I think there were two reasons of failure.</p>\n<p>First problem was submission with 1s for every questions. Below is the simplified version of code. The scoring time to resulting out failure was too fast (&lt; 10min) in this case.</p>\n<pre><code>import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    for i in range(len(sample_submission)):\n        sample_submission.loc[i, 'correct'] = 1\n\n    env.predict(sample_submission)\n</code></pre>\n<p>Second problem took long time to be judged (&gt; 30min), and the reason was not sure, just assumed that the indexing of the 'sample_submission' was not ordered by 'session_id'. As I changed the code like below, the submission was succeed.</p>\n<pre><code>import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = 1 if np.random.random() &gt; 0.5 else 0\n\n    env.predict(sample_submission)\n</code></pre>\n<p>If there are newbies having troubles just like mine, hope this would help!</p>",
      "rawMarkdown": "I found a working example for submission success. I think there were two reasons of failure.\n\nFirst problem was submission with 1s for every questions. Below is the simplified version of code. The scoring time to resulting out failure was too fast (< 10min) in this case.\n```\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    for i in range(len(sample_submission)):\n        sample_submission.loc[i, 'correct'] = 1\n    \n    env.predict(sample_submission)\n```\n\nSecond problem took long time to be judged (> 30min), and the reason was not sure, just assumed that the indexing of the 'sample_submission' was not ordered by 'session_id'. As I changed the code like below, the submission was succeed.\n```\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = 1 if np.random.random() > 0.5 else 0\n    \n    env.predict(sample_submission)\n```\n\nIf there are newbies having troubles just like mine, hope this would help!",
      "votes": null
    },
    {
      "id": "2241546",
      "postDate": "05/01/2023 15:37:05",
      "content": "<p>Just try to write it in another way. It is quite hard to debug. </p>",
      "rawMarkdown": "Just try to write it in another way. It is quite hard to debug.",
      "votes": null
    },
    {
      "id": "2243097",
      "postDate": "05/02/2023 17:15:35",
      "content": "<p>I'm running into a scoring issue, not sure why but are you saying that there are time limits set during the scoring phase?</p>",
      "rawMarkdown": "I'm running into a scoring issue, not sure why but are you saying that there are time limits set during the scoring phase?",
      "votes": null
    },
    {
      "id": "2244077",
      "postDate": "05/03/2023 12:23:20",
      "content": "<p>No, I just said the time I failed was a kind if hint differentiating the reason of failure.<br>\nAnd now I succeed submitting by using 'mask' as in my code in the last comment.</p>",
      "rawMarkdown": "No, I just said the time I failed was a kind if hint differentiating the reason of failure.\nAnd now I succeed submitting by using 'mask' as in my code in the last comment.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2236647,
      "author_name": "taegeunlim",
      "author_url": "",
      "post_date": "04/27/2023 02:58:09",
      "content": "<p>I found that the order of items is not related this issue, as the reference notebook successfully submitted.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2237038,
      "author_name": "taegeunlim",
      "author_url": "",
      "post_date": "04/27/2023 10:34:16",
      "content": "<p>I found a working example for submission success. I think there were two reasons of failure.</p>\n<p>First problem was submission with 1s for every questions. Below is the simplified version of code. The scoring time to resulting out failure was too fast (&lt; 10min) in this case.</p>\n<pre><code>import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    for i in range(len(sample_submission)):\n        sample_submission.loc[i, 'correct'] = 1\n\n    env.predict(sample_submission)\n</code></pre>\n<p>Second problem took long time to be judged (&gt; 30min), and the reason was not sure, just assumed that the indexing of the 'sample_submission' was not ordered by 'session_id'. As I changed the code like below, the submission was succeed.</p>\n<pre><code>import jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = 1 if np.random.random() &gt; 0.5 else 0\n\n    env.predict(sample_submission)\n</code></pre>\n<p>If there are newbies having troubles just like mine, hope this would help!</p>",
      "votes": null,
      "replies": [
        {
          "id": 2243097,
          "author_name": "joseviteri",
          "author_url": "",
          "post_date": "05/02/2023 17:15:35",
          "content": "<p>I'm running into a scoring issue, not sure why but are you saying that there are time limits set during the scoring phase?</p>",
          "votes": null,
          "replies": [
            {
              "id": 2244077,
              "author_name": "taegeunlim",
              "author_url": "",
              "post_date": "05/03/2023 12:23:20",
              "content": "<p>No, I just said the time I failed was a kind if hint differentiating the reason of failure.<br>\nAnd now I succeed submitting by using 'mask' as in my code in the last comment.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2241546,
      "author_name": "littlstar123",
      "author_url": "",
      "post_date": "05/01/2023 15:37:05",
      "content": "<p>Just try to write it in another way. It is quite hard to debug. </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2236129": "I'm a newbie to Kaggle and I don't know why I have the submission file format error.\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F6258d7f991f61cb98df6f105c972d7ac%2F2023-04-26%20%2011.45.02.png?generation=1682520351897621&alt=media)\n\nI followed the starter notebook instructing how to use submission API, except I made all predictions to '1' just for submission test purpose. Below is the code.\n```\nclass my_model():\n    def predict(self, X):\n        return 1\n\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    test_df = feature_engineer(test)\n    grp = test_df.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        gbtm = my_model()\n        test_ds = tfdf.keras.pd_dataframe_to_tf_dataset(test_df.loc[:, test_df.columns != 'level_group'])\n        predictions = gbtm.predict(test_ds)\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        n_predictions = np.array((predictions > best_threshold)).astype(int)\n        sample_submission.loc[mask,'correct'] = n_predictions.flatten()\n    \n    env.predict(sample_submission[['session_id', 'correct']])\n```\n\nI found one thing that the order of items in my submission file is different from that of the reference notebook. However, I don't have any clue why the orders are different.\n1. mine\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2Ff298e9a142419864b7468811588bbb66%2F2023-04-26%20%2011.51.16.png?generation=1682520825227694&alt=media)\n2. ref\n![](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F13518318%2F5fdaa82c21b8436b9420ed6160185539%2F2023-04-26%20%2011.51.56.png?generation=1682520840769585&alt=media)\n\nCan anyone give me an advice?",
    "2236647": "I found that the order of items is not related this issue, as the reference notebook successfully submitted.",
    "2237038": "I found a working example for submission success. I think there were two reasons of failure.\n\nFirst problem was submission with 1s for every questions. Below is the simplified version of code. The scoring time to resulting out failure was too fast (< 10min) in this case.\n```\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    for i in range(len(sample_submission)):\n        sample_submission.loc[i, 'correct'] = 1\n    \n    env.predict(sample_submission)\n```\n\nSecond problem took long time to be judged (> 30min), and the reason was not sure, just assumed that the indexing of the 'sample_submission' was not ordered by 'session_id'. As I changed the code like below, the submission was succeed.\n```\nimport jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\n\nfor (test, sample_submission) in iter_test:\n    grp = test.level_group.values[0]\n    a,b = limits[grp]\n    for t in range(a,b):\n        mask = sample_submission.session_id.str.contains(f'q{t}')\n        sample_submission.loc[mask,'correct'] = 1 if np.random.random() > 0.5 else 0\n    \n    env.predict(sample_submission)\n```\n\nIf there are newbies having troubles just like mine, hope this would help!",
    "2241546": "Just try to write it in another way. It is quite hard to debug.",
    "2243097": "I'm running into a scoring issue, not sure why but are you saying that there are time limits set during the scoring phase?",
    "2244077": "No, I just said the time I failed was a kind if hint differentiating the reason of failure.\nAnd now I succeed submitting by using 'mask' as in my code in the last comment."
  },
  "source": "meta"
}