{
  "id": 412512,
  "title": "Problem with the new API (jo_wilder_310)",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/412512",
  "author_name": "Jack (Japan)",
  "post_date": "2023-05-24T03:16:46.195000",
  "votes": 63,
  "comment_count": 17,
  "views": 0,
  "content": "<p>After several verifications, I have come to the conclusion that <strong>both data provided by the new API (not just the latter data, as I mentioned <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409898#2266542\" target=\"_blank\">here</a>) are shuffled</strong>. It was difficult to detect, however, because it had not occurred for the three sample data at the time of the commit of the notebook.</p>\n<p>Thus, <strong>if your feature engineering process depends on the original order of the data, you need to sort the session data by the index column beforehand, or else it will not be processed as intended, and will make your LB score worse</strong>. This problem does not occur with the old API.</p>\n<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> I think this fact would be very confusing to participants since training data and sample test data always seem to be sorted by index (the first dataframe from API) or question number (the second dataframe from API). Would you consider addressing this issue?</p>",
  "messages": [
    {
      "id": 2271582,
      "postDate": "2023-05-24T03:16:46.197Z",
      "content": "<p>After several verifications, I have come to the conclusion that <strong>both data provided by the new API (not just the latter data, as I mentioned <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409898#2266542\" target=\"_blank\">here</a>) are shuffled</strong>. It was difficult to detect, however, because it had not occurred for the three sample data at the time of the commit of the notebook.</p>\n<p>Thus, <strong>if your feature engineering process depends on the original order of the data, you need to sort the session data by the index column beforehand, or else it will not be processed as intended, and will make your LB score worse</strong>. This problem does not occur with the old API.</p>\n<p><a href=\"https://www.kaggle.com/philculliton\" target=\"_blank\">@philculliton</a> I think this fact would be very confusing to participants since training data and sample test data always seem to be sorted by index (the first dataframe from API) or question number (the second dataframe from API). Would you consider addressing this issue?</p>",
      "rawMarkdown": "After several verifications, I have come to the conclusion that **both data provided by the new API (not just the latter data, as I mentioned [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409898#2266542)) are shuffled**. It was difficult to detect, however, because it had not occurred for the three sample data at the time of the commit of the notebook.\n\nThus, **if your feature engineering process depends on the original order of the data, you need to sort the session data by the index column beforehand, or else it will not be processed as intended, and will make your LB score worse**. This problem does not occur with the old API.\n\n@philculliton I think this fact would be very confusing to participants since training data and sample test data always seem to be sorted by index (the first dataframe from API) or question number (the second dataframe from API). Would you consider addressing this issue?",
      "votes": 62
    },
    {
      "id": 2271603,
      "postDate": "2023-05-24T03:49:59.900Z",
      "content": "<p>That means, this is enough right?</p>\n<pre><code>import jo_wilder_310\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nfor (test, sample_submission) in iter_test:\n    test = test.sort_values(by = 'index')\n    ...\n</code></pre>",
      "rawMarkdown": "That means, this is enough right?\n```\nimport jo_wilder_310\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nfor (test, sample_submission) in iter_test:\n    test = test.sort_values(by = 'index')\n    ...\n```",
      "votes": 3,
      "replies": [
        {
          "id": 2271627,
          "postDate": "2023-05-24T04:20:36.990Z",
          "content": "<p>Yes. In addition, you have to be mindful of the correspondence with the question number when substituting predictions from the model to the sample_submission DataFrame (not sorted by question number).</p>",
          "rawMarkdown": "Yes. In addition, you have to be mindful of the correspondence with the question number when substituting predictions from the model to the sample_submission DataFrame (not sorted by question number).",
          "votes": 3,
          "replies": [
            {
              "id": 2271736,
              "postDate": "2023-05-24T05:56:29.633Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2272507,
              "postDate": "2023-05-24T14:45:28.020Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2272534,
              "postDate": "2023-05-24T14:57:57.527Z",
              "content": "<p>sorting the sample_submission improves the score but is still not at par with the old API</p>\n<pre><code>for (test, sample_submission) in (iter_test):\n    test = test.sort_values(by = 'index')\n    sample_submission['questions'] = sample_submission['session_id'].str.split('_').apply(lambda x: x[1])\n    sample_submission = sample_submission.sort_values(by = 'questions')\n    sample_submission = sample_submission[['session_id', 'correct']]\n</code></pre>",
              "rawMarkdown": "sorting the sample_submission improves the score but is still not at par with the old API\n```\nfor (test, sample_submission) in (iter_test):\n    test = test.sort_values(by = 'index')\n    sample_submission['questions'] = sample_submission['session_id'].str.split('_').apply(lambda x: x[1])\n    sample_submission = sample_submission.sort_values(by = 'questions')\n    sample_submission = sample_submission[['session_id', 'correct']]\n```",
              "votes": 4
            },
            {
              "id": 2274390,
              "postDate": "2023-05-26T00:19:58.440Z",
              "content": "<p><a href=\"https://www.kaggle.com/navinkumarmnk\" target=\"_blank\">@navinkumarmnk</a> It seems to me that there is something wrong with your feature engineering process (perhaps in the feature_engineer function) rather than being caused by the API.</p>\n<p><a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> In my case, with the proper sorting process, I was able to get decent results with the new API. However, there may be other issues that I am not aware of.</p>",
              "rawMarkdown": "@navinkumarmnk It seems to me that there is something wrong with your feature engineering process (perhaps in the feature_engineer function) rather than being caused by the API.\n\n@pjmathematician In my case, with the proper sorting process, I was able to get decent results with the new API. However, there may be other issues that I am not aware of."
            },
            {
              "id": 2274417,
              "postDate": "2023-05-26T01:51:43.913Z",
              "rawMarkdown": "",
              "isDeleted": true
            },
            {
              "id": 2274861,
              "postDate": "2023-05-26T09:54:04.153Z",
              "content": "<p>I have addressed the sort issue for index and question, but still the LB score is still destroyed.<br>\nI have created a new <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413002\" target=\"_blank\">thread</a> so if anyone is experiencing the same problem, we can share information.</p>",
              "rawMarkdown": "I have addressed the sort issue for index and question, but still the LB score is still destroyed.\nI have created a new [thread](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413002) so if anyone is experiencing the same problem, we can share information."
            }
          ]
        }
      ]
    },
    {
      "id": 2282967,
      "postDate": "2023-06-01T01:19:00.470Z",
      "content": "<p>Thanks for sharing, seems like now this is a problem for ALL of us 😔, replying just to make this post resurface.</p>",
      "rawMarkdown": "Thanks for sharing, seems like now this is a problem for ALL of us 😔, replying just to make this post resurface.",
      "votes": 1
    },
    {
      "id": 2279380,
      "postDate": "2023-05-29T10:20:46.670Z",
      "content": "<p>Thanks a lot! Its help me to solve my problem.</p>",
      "rawMarkdown": "Thanks a lot! Its help me to solve my problem.",
      "votes": 1
    },
    {
      "id": 2274299,
      "postDate": "2023-05-25T19:23:23.513Z",
      "content": "<p>Thanks a lot! Saved me a bunch of time, wasn't sure what was the issue</p>",
      "rawMarkdown": "Thanks a lot! Saved me a bunch of time, wasn't sure what was the issue",
      "votes": 1
    },
    {
      "id": 2272004,
      "postDate": "2023-05-24T08:34:49.563Z",
      "content": "<p>Thanks for the detailed info on the problem. Just to confirm, <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> is there any other problem with the old API ? </p>",
      "rawMarkdown": "Thanks for the detailed info on the problem. Just to confirm, @rsakata is there any other problem with the old API ? ",
      "votes": 1,
      "replies": [
        {
          "id": 2272603,
          "postDate": "2023-05-24T15:50:01.510Z",
          "content": "<p>Maybe, there is not a problem with the old API.</p>",
          "rawMarkdown": "Maybe, there is not a problem with the old API.",
          "votes": 1
        }
      ]
    },
    {
      "id": 2277412,
      "postDate": "2023-05-27T17:42:20.873Z",
      "content": "<p>Always impressed by the curiosity needed to find such things, thx for your contribution</p>",
      "rawMarkdown": "Always impressed by the curiosity needed to find such things, thx for your contribution",
      "votes": 2
    },
    {
      "id": 2284238,
      "postDate": "2023-06-01T20:37:39.210Z",
      "content": "<p>available run local :<br>\n<strong>jo_wilder_310.py</strong>  <br>\n                   kaggle datasets download -d liudacheldieva/competition-game-jo-wilder-py </p>\n<p>Example of run in:<br>\n   <a href=\"https://www.kaggle.com/code/liudacheldieva/40ipynb-sucsess\" target=\"_blank\">https://www.kaggle.com/code/liudacheldieva/40ipynb-sucsess</a><br>\nDescription:<br>\n    <strong>local run:</strong><br>\nimport jo_wilder_310 as jo_wilder<br>\niter_test =jo_wilder.make_env_iter_test()<br>\nenv = iter_test</p>\n<p><strong>kaggle run:</strong><br>\nimport jo_wilder_310 as jo_wilder<br>\nenv = jo_wilder.make_env()<br>\niter_test = env.iter_test()</p>",
      "rawMarkdown": "available run local :\n**jo_wilder_310.py**  \n                   kaggle datasets download -d liudacheldieva/competition-game-jo-wilder-py \n\nExample of run in:\n   https://www.kaggle.com/code/liudacheldieva/40ipynb-sucsess\nDescription:\n    **local run:**\nimport jo_wilder_310 as jo_wilder\niter_test =jo_wilder.make_env_iter_test()\nenv = iter_test\n\n   **kaggle run:**\nimport jo_wilder_310 as jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()"
    },
    {
      "id": 2274527,
      "postDate": "2023-05-26T04:31:57.173Z",
      "content": "<p>Isnt  feature engineering automatic with algorithms like py caret</p>",
      "rawMarkdown": "Isnt  feature engineering automatic with algorithms like py caret"
    },
    {
      "id": 2272870,
      "postDate": "2023-05-24T20:08:07.600Z",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> 🤘</p>",
      "rawMarkdown": "Thanks for sharing @rsakata 🤘",
      "votes": 3
    }
  ],
  "comments": [
    {
      "id": 2271603,
      "author_name": "pjmathematician",
      "author_url": "",
      "post_date": "2023-05-24T03:49:59.900000",
      "content": "<p>That means, this is enough right?</p>\n<pre><code>import jo_wilder_310\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nfor (test, sample_submission) in iter_test:\n    test = test.sort_values(by = 'index')\n    ...\n</code></pre>",
      "votes": 3,
      "replies": [
        {
          "id": 2271627,
          "author_name": "Jack (Japan)",
          "author_url": "",
          "post_date": "2023-05-24T04:20:36.990000",
          "content": "<p>Yes. In addition, you have to be mindful of the correspondence with the question number when substituting predictions from the model to the sample_submission DataFrame (not sorted by question number).</p>",
          "votes": 3,
          "replies": [
            {
              "id": 2271736,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-24T05:56:29.633000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2272507,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-24T14:45:28.020000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2272534,
              "author_name": "pjmathematician",
              "author_url": "",
              "post_date": "2023-05-24T14:57:57.527000",
              "content": "<p>sorting the sample_submission improves the score but is still not at par with the old API</p>\n<pre><code>for (test, sample_submission) in (iter_test):\n    test = test.sort_values(by = 'index')\n    sample_submission['questions'] = sample_submission['session_id'].str.split('_').apply(lambda x: x[1])\n    sample_submission = sample_submission.sort_values(by = 'questions')\n    sample_submission = sample_submission[['session_id', 'correct']]\n</code></pre>",
              "votes": 4,
              "replies": []
            },
            {
              "id": 2274390,
              "author_name": "Jack (Japan)",
              "author_url": "",
              "post_date": "2023-05-26T00:19:58.440000",
              "content": "<p><a href=\"https://www.kaggle.com/navinkumarmnk\" target=\"_blank\">@navinkumarmnk</a> It seems to me that there is something wrong with your feature engineering process (perhaps in the feature_engineer function) rather than being caused by the API.</p>\n<p><a href=\"https://www.kaggle.com/pjmathematician\" target=\"_blank\">@pjmathematician</a> In my case, with the proper sorting process, I was able to get decent results with the new API. However, there may be other issues that I am not aware of.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2274417,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-05-26T01:51:43.913000",
              "content": "",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2274861,
              "author_name": "tonic",
              "author_url": "",
              "post_date": "2023-05-26T09:54:04.153000",
              "content": "<p>I have addressed the sort issue for index and question, but still the LB score is still destroyed.<br>\nI have created a new <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/413002\" target=\"_blank\">thread</a> so if anyone is experiencing the same problem, we can share information.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2282967,
      "author_name": "Woprime",
      "author_url": "",
      "post_date": "2023-06-01T01:19:00.470000",
      "content": "<p>Thanks for sharing, seems like now this is a problem for ALL of us 😔, replying just to make this post resurface.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2279380,
      "author_name": "cruelkanade",
      "author_url": "",
      "post_date": "2023-05-29T10:20:46.670000",
      "content": "<p>Thanks a lot! Its help me to solve my problem.</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2274299,
      "author_name": "Informhunter",
      "author_url": "",
      "post_date": "2023-05-25T19:23:23.513000",
      "content": "<p>Thanks a lot! Saved me a bunch of time, wasn't sure what was the issue</p>",
      "votes": 1,
      "replies": []
    },
    {
      "id": 2272004,
      "author_name": "NikhilMishra",
      "author_url": "",
      "post_date": "2023-05-24T08:34:49.563000",
      "content": "<p>Thanks for the detailed info on the problem. Just to confirm, <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> is there any other problem with the old API ? </p>",
      "votes": 1,
      "replies": [
        {
          "id": 2272603,
          "author_name": "Jack (Japan)",
          "author_url": "",
          "post_date": "2023-05-24T15:50:01.510000",
          "content": "<p>Maybe, there is not a problem with the old API.</p>",
          "votes": 1,
          "replies": []
        }
      ]
    },
    {
      "id": 2277412,
      "author_name": "JEANMPIA",
      "author_url": "",
      "post_date": "2023-05-27T17:42:20.873000",
      "content": "<p>Always impressed by the curiosity needed to find such things, thx for your contribution</p>",
      "votes": 2,
      "replies": []
    },
    {
      "id": 2284238,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-01T20:37:39.210000",
      "content": "<p>available run local :<br>\n<strong>jo_wilder_310.py</strong>  <br>\n                   kaggle datasets download -d liudacheldieva/competition-game-jo-wilder-py </p>\n<p>Example of run in:<br>\n   <a href=\"https://www.kaggle.com/code/liudacheldieva/40ipynb-sucsess\" target=\"_blank\">https://www.kaggle.com/code/liudacheldieva/40ipynb-sucsess</a><br>\nDescription:<br>\n    <strong>local run:</strong><br>\nimport jo_wilder_310 as jo_wilder<br>\niter_test =jo_wilder.make_env_iter_test()<br>\nenv = iter_test</p>\n<p><strong>kaggle run:</strong><br>\nimport jo_wilder_310 as jo_wilder<br>\nenv = jo_wilder.make_env()<br>\niter_test = env.iter_test()</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2274527,
      "author_name": "Muhammad Ammar Jamshed",
      "author_url": "",
      "post_date": "2023-05-26T04:31:57.173000",
      "content": "<p>Isnt  feature engineering automatic with algorithms like py caret</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 2272870,
      "author_name": "Gaurav Srivastava",
      "author_url": "",
      "post_date": "2023-05-24T20:08:07.600000",
      "content": "<p>Thanks for sharing <a href=\"https://www.kaggle.com/rsakata\" target=\"_blank\">@rsakata</a> 🤘</p>",
      "votes": 3,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2271582": "After several verifications, I have come to the conclusion that **both data provided by the new API (not just the latter data, as I mentioned [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/409898#2266542)) are shuffled**. It was difficult to detect, however, because it had not occurred for the three sample data at the time of the commit of the notebook.\n\nThus, **if your feature engineering process depends on the original order of the data, you need to sort the session data by the index column beforehand, or else it will not be processed as intended, and will make your LB score worse**. This problem does not occur with the old API.\n\n@philculliton I think this fact would be very confusing to participants since training data and sample test data always seem to be sorted by index (the first dataframe from API) or question number (the second dataframe from API). Would you consider addressing this issue?",
    "2271603": "That means, this is enough right?\n```\nimport jo_wilder_310\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nfor (test, sample_submission) in iter_test:\n    test = test.sort_values(by = 'index')\n    ...\n```",
    "2282967": "Thanks for sharing, seems like now this is a problem for ALL of us 😔, replying just to make this post resurface.",
    "2279380": "Thanks a lot! Its help me to solve my problem.",
    "2274299": "Thanks a lot! Saved me a bunch of time, wasn't sure what was the issue",
    "2272004": "Thanks for the detailed info on the problem. Just to confirm, @rsakata is there any other problem with the old API ? ",
    "2277412": "Always impressed by the curiosity needed to find such things, thx for your contribution",
    "2284238": "available run local :\n**jo_wilder_310.py**  \n                   kaggle datasets download -d liudacheldieva/competition-game-jo-wilder-py \n\nExample of run in:\n   https://www.kaggle.com/code/liudacheldieva/40ipynb-sucsess\nDescription:\n    **local run:**\nimport jo_wilder_310 as jo_wilder\niter_test =jo_wilder.make_env_iter_test()\nenv = iter_test\n\n   **kaggle run:**\nimport jo_wilder_310 as jo_wilder\nenv = jo_wilder.make_env()\niter_test = env.iter_test()",
    "2274527": "Isnt  feature engineering automatic with algorithms like py caret",
    "2272870": "Thanks for sharing @rsakata 🤘"
  }
}