{
  "id": 413004,
  "title": "For those suffering from the API...",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/413004",
  "author_name": "",
  "post_date": "2023-05-26T10:28:00.464741900Z",
  "votes": 14,
  "comment_count": 1,
  "views": 0,
  "content": "<p>I've seen a lot of API-related posts on the Discussion tab lately.</p>\n<p>At least I have been submitting without problems since the beginning of April to the end of May in the following way.</p>\n<p>Code example :</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n (test, sample_submission)  iter_test:\n    sample_submission[] = sample_submission[]..split().apply( x: (x[]))\n    sample_submission = sample_submission.sort_values(by=, ascending=).reset_index()\n    sample_submission = sample_submission[[, ]] \n\n    grp = test.level_group.values[]\n    session_ids = test.session_id.unique()\n     session_id  session_ids:   \n        test_session = test[test[] == session_id]\n        test_session = test_session.sort_values(by=)  \n        df = (pl.from_pandas(test_session).drop([, , ]).with_columns(columns))\n        df = feature_engineer(df, grp)\n\n        preds_ens = []\n         fold  (, FOLD_COUNT): \n            a,b = limits[grp]\n            preds = []\n             q  (a, b):\n                FEATURES = features_dict[(q)]\n                model = models_list[q-][fold]\n                pred = model.predict_proba(df[FEATURES])[,]\n                preds.append((pred, ))\n            preds_ens.append(preds)\n\n        preds_ens = np.average(np.array(preds_ens), axis=)\n        mask = sample_submission.session_id..contains()\n        :\n            sample_submission.loc[mask,] = preds_ens\n         Exception  e:\n            (e)\n     DEBUG:\n        display(sample_submission)\n    env.predict(sample_submission)\n</code></pre>\n<p>It may not be efficient code, but at least the submission reliability is high.</p>\n<p>I hope it will be of help.</p>\n<p>Edit 1 : It is estimated that the method of providing API data has changed after June 1, 2023. Specifically, the elapsed_time or index order is not sorted. So, I added a sorting procedure to the code above.</p>\n<p>Edit 2 : Submission questions are also out of order. So I add code to sort questions.</p>",
  "messages": [
    {
      "id": "2274879",
      "postDate": "05/26/2023 10:28:00",
      "content": "<p>I've seen a lot of API-related posts on the Discussion tab lately.</p>\n<p>At least I have been submitting without problems since the beginning of April to the end of May in the following way.</p>\n<p>Code example :</p>\n<pre><code>limits = {:(,), :(,), :(,)}\n (test, sample_submission)  iter_test:\n    sample_submission[] = sample_submission[]..split().apply( x: (x[]))\n    sample_submission = sample_submission.sort_values(by=, ascending=).reset_index()\n    sample_submission = sample_submission[[, ]] \n\n    grp = test.level_group.values[]\n    session_ids = test.session_id.unique()\n     session_id  session_ids:   \n        test_session = test[test[] == session_id]\n        test_session = test_session.sort_values(by=)  \n        df = (pl.from_pandas(test_session).drop([, , ]).with_columns(columns))\n        df = feature_engineer(df, grp)\n\n        preds_ens = []\n         fold  (, FOLD_COUNT): \n            a,b = limits[grp]\n            preds = []\n             q  (a, b):\n                FEATURES = features_dict[(q)]\n                model = models_list[q-][fold]\n                pred = model.predict_proba(df[FEATURES])[,]\n                preds.append((pred, ))\n            preds_ens.append(preds)\n\n        preds_ens = np.average(np.array(preds_ens), axis=)\n        mask = sample_submission.session_id..contains()\n        :\n            sample_submission.loc[mask,] = preds_ens\n         Exception  e:\n            (e)\n     DEBUG:\n        display(sample_submission)\n    env.predict(sample_submission)\n</code></pre>\n<p>It may not be efficient code, but at least the submission reliability is high.</p>\n<p>I hope it will be of help.</p>\n<p>Edit 1 : It is estimated that the method of providing API data has changed after June 1, 2023. Specifically, the elapsed_time or index order is not sorted. So, I added a sorting procedure to the code above.</p>\n<p>Edit 2 : Submission questions are also out of order. So I add code to sort questions.</p>",
      "rawMarkdown": "I've seen a lot of API-related posts on the Discussion tab lately.\n\nAt least I have been submitting without problems since the beginning of April to the end of May in the following way.\n\n\nCode example :\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nfor (test, sample_submission) in iter_test:\n    sample_submission['questions'] = sample_submission['session_id'].str.split('_q').apply(lambda x: int(x[1]))\n    sample_submission = sample_submission.sort_values(by='questions', ascending=True).reset_index()\n    sample_submission = sample_submission[['session_id', 'correct']] # Edit2 : Sort questions\n\n    grp = test.level_group.values[0]\n    session_ids = test.session_id.unique()\n    for session_id in session_ids:   # Processing by session_id\n        test_session = test[test['session_id'] == session_id]\n        test_session = test_session.sort_values(by='elapsed_time')  # Edit1 : Sort by elapsed_time\n        df = (pl.from_pandas(test_session).drop([\"fullscreen\", \"hq\", \"music\"]).with_columns(columns))\n        df = feature_engineer(df, grp)\n\n        preds_ens = []\n        for fold in range(0, FOLD_COUNT): # If you use fold ensemble...\n            a,b = limits[grp]\n            preds = []\n            for q in range(a, b):\n                FEATURES = features_dict[str(q)]\n                model = models_list[q-1][fold]\n                pred = model.predict_proba(df[FEATURES])[0,1]\n                preds.append(round(pred, 5))\n            preds_ens.append(preds)\n\n        preds_ens = np.average(np.array(preds_ens), axis=0)\n        mask = sample_submission.session_id.str.contains(f'{session_id}_q')\n        try:\n            sample_submission.loc[mask,'correct'] = preds_ens\n        except Exception as e:\n            print(e)\n    if DEBUG:\n        display(sample_submission)\n    env.predict(sample_submission)\n```\n\nIt may not be efficient code, but at least the submission reliability is high.\n\nI hope it will be of help.\n\n\nEdit 1 : It is estimated that the method of providing API data has changed after June 1, 2023. Specifically, the elapsed_time or index order is not sorted. So, I added a sorting procedure to the code above.\n\nEdit 2 : Submission questions are also out of order. So I add code to sort questions.",
      "votes": null
    },
    {
      "id": "2274907",
      "postDate": "05/26/2023 11:18:03",
      "content": "<p>Great, I have been struggling with the order issue of the new API. Thank you for sharing.</p>",
      "rawMarkdown": "Great, I have been struggling with the order issue of the new API. Thank you for sharing.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2274907,
      "author_name": "dwchen",
      "author_url": "",
      "post_date": "05/26/2023 11:18:03",
      "content": "<p>Great, I have been struggling with the order issue of the new API. Thank you for sharing.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2274879": "I've seen a lot of API-related posts on the Discussion tab lately.\n\nAt least I have been submitting without problems since the beginning of April to the end of May in the following way.\n\n\nCode example :\n```python\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nfor (test, sample_submission) in iter_test:\n    sample_submission['questions'] = sample_submission['session_id'].str.split('_q').apply(lambda x: int(x[1]))\n    sample_submission = sample_submission.sort_values(by='questions', ascending=True).reset_index()\n    sample_submission = sample_submission[['session_id', 'correct']] # Edit2 : Sort questions\n\n    grp = test.level_group.values[0]\n    session_ids = test.session_id.unique()\n    for session_id in session_ids:   # Processing by session_id\n        test_session = test[test['session_id'] == session_id]\n        test_session = test_session.sort_values(by='elapsed_time')  # Edit1 : Sort by elapsed_time\n        df = (pl.from_pandas(test_session).drop([\"fullscreen\", \"hq\", \"music\"]).with_columns(columns))\n        df = feature_engineer(df, grp)\n\n        preds_ens = []\n        for fold in range(0, FOLD_COUNT): # If you use fold ensemble...\n            a,b = limits[grp]\n            preds = []\n            for q in range(a, b):\n                FEATURES = features_dict[str(q)]\n                model = models_list[q-1][fold]\n                pred = model.predict_proba(df[FEATURES])[0,1]\n                preds.append(round(pred, 5))\n            preds_ens.append(preds)\n\n        preds_ens = np.average(np.array(preds_ens), axis=0)\n        mask = sample_submission.session_id.str.contains(f'{session_id}_q')\n        try:\n            sample_submission.loc[mask,'correct'] = preds_ens\n        except Exception as e:\n            print(e)\n    if DEBUG:\n        display(sample_submission)\n    env.predict(sample_submission)\n```\n\nIt may not be efficient code, but at least the submission reliability is high.\n\nI hope it will be of help.\n\n\nEdit 1 : It is estimated that the method of providing API data has changed after June 1, 2023. Specifically, the elapsed_time or index order is not sorted. So, I added a sorting procedure to the code above.\n\nEdit 2 : Submission questions are also out of order. So I add code to sort questions.",
    "2274907": "Great, I have been struggling with the order issue of the new API. Thank you for sharing."
  },
  "source": "meta"
}