{
  "id": 419445,
  "title": "Predict using data from the previous level group will result in an Submission Scoring Error.",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/419445",
  "author_name": "",
  "post_date": "2023-06-26T02:18:08.321534800Z",
  "votes": -1,
  "comment_count": 5,
  "views": 0,
  "content": "<p>Hi kaggle staff and participants.</p>\n<p>I am in trouble.</p>\n<p>I predict the following.<br>\n・q1-q3 are predicted using data with level_group '0-4'.<br>\n・q4-q13 are predicted using data from level_group '0-4' and level_group '5-12'.<br>\n・q14-q18 are predicted using data from level_group '0-4', level_group '5-12', and level_group <br>\n   '13-22'</p>\n<p><strong>I am able to submit.\nHowever, about 3 hours after the first evaluation of the hidden test set, an Submission Scoring Error occurs.</strong></p>\n<p>Here is the pseudo code.</p>\n<pre><code>env = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nlimits = {:(,), :(,), :(,)}\nhistorical_meta = defaultdict()\nbest_threshold = \n\n\nsession_df_dict = {}\n\n (test, sample_submission)  iter_test:\n    grp = test.level_group.values[]\n    session_id = test.session_id.values[]\n\n    test = test.sort_values(by=).reset_index(drop=)\n\n    \n    test = (\n        pl.from_pandas(test)\n        .rename(\n            {:}\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .clip(, )\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            pl.col().cast(pl.Float32),\n            pl.col().fill_null(),\n            pl.col().fill_null()\n        )\n    )\n\n    \n     grp == :\n\n        \n        test = _make_exog_easy(test, grp)\n\n        \n        keys = +(session_id)\n\n        \n         keys   session_df_dict.keys():\n            session_df_dict[keys] = test\n\n    \n     grp == :\n\n        \n        test = _make_exog_normal(test, grp)\n\n        \n        keys = +(session_id)\n         keys  session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on=, how=)\n\n        \n        keys = +(session_id)\n\n        \n         keys   session_df_dict.keys():\n            session_df_dict[keys] = test\n\n    \n     grp == :\n        \n        test = _make_exog_hard(test, grp)\n\n        \n        keys = +(session_id)\n         keys  session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on=, how=)\n\n    test = test.fillna()\n\n    sample_submission[] = sample_submission[].apply( x: (x.split()[]))\n    sample_submission[] = sample_submission[].apply( x: (x.split()[-][:]))\n    sample_submission = sample_submission.sort_values(by=)\n\n    min_q, max_q = limits[grp]\n\n     n  (min_q, max_q):\n        feature_lists = importance_dict[(n)]\n        col_lists = ((feature_lists)-((test)))\n         (col_lists) &gt; :\n             var  col_lists:\n                test[var] = \n\n         pred = []\n          i  (, ):\n              (, mode=)  fp:\n                 model = pickle.load(fp)\n             pred.append(model.predict_proba(test[feature_lists])[:, ][])\n\n         sample_submission.loc[(sample_submission[]==session_id)&amp;(sample_submission[]==n), ] = mean(pred)\n\n\n    sample_submission[] = sample_submission[].apply( x:  x&gt;best_threshold  )\n    env.predict(sample_submission[[, ]])\n</code></pre>\n<p>Please help. Thank you in advance for your help.</p>",
  "messages": [
    {
      "id": "2317806",
      "postDate": "06/26/2023 02:18:08",
      "content": "<p>Hi kaggle staff and participants.</p>\n<p>I am in trouble.</p>\n<p>I predict the following.<br>\n・q1-q3 are predicted using data with level_group '0-4'.<br>\n・q4-q13 are predicted using data from level_group '0-4' and level_group '5-12'.<br>\n・q14-q18 are predicted using data from level_group '0-4', level_group '5-12', and level_group <br>\n   '13-22'</p>\n<p><strong>I am able to submit.\nHowever, about 3 hours after the first evaluation of the hidden test set, an Submission Scoring Error occurs.</strong></p>\n<p>Here is the pseudo code.</p>\n<pre><code>env = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nlimits = {:(,), :(,), :(,)}\nhistorical_meta = defaultdict()\nbest_threshold = \n\n\nsession_df_dict = {}\n\n (test, sample_submission)  iter_test:\n    grp = test.level_group.values[]\n    session_id = test.session_id.values[]\n\n    test = test.sort_values(by=).reset_index(drop=)\n\n    \n    test = (\n        pl.from_pandas(test)\n        .rename(\n            {:}\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .clip(, )\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            (pl.col().sub(pl.col().shift()))\n            .fill_null()\n            .()\n            .over([, ])\n            .alias()\n        )\n        .with_columns(\n            pl.col().cast(pl.Float32),\n            pl.col().fill_null(),\n            pl.col().fill_null()\n        )\n    )\n\n    \n     grp == :\n\n        \n        test = _make_exog_easy(test, grp)\n\n        \n        keys = +(session_id)\n\n        \n         keys   session_df_dict.keys():\n            session_df_dict[keys] = test\n\n    \n     grp == :\n\n        \n        test = _make_exog_normal(test, grp)\n\n        \n        keys = +(session_id)\n         keys  session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on=, how=)\n\n        \n        keys = +(session_id)\n\n        \n         keys   session_df_dict.keys():\n            session_df_dict[keys] = test\n\n    \n     grp == :\n        \n        test = _make_exog_hard(test, grp)\n\n        \n        keys = +(session_id)\n         keys  session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on=, how=)\n\n    test = test.fillna()\n\n    sample_submission[] = sample_submission[].apply( x: (x.split()[]))\n    sample_submission[] = sample_submission[].apply( x: (x.split()[-][:]))\n    sample_submission = sample_submission.sort_values(by=)\n\n    min_q, max_q = limits[grp]\n\n     n  (min_q, max_q):\n        feature_lists = importance_dict[(n)]\n        col_lists = ((feature_lists)-((test)))\n         (col_lists) &gt; :\n             var  col_lists:\n                test[var] = \n\n         pred = []\n          i  (, ):\n              (, mode=)  fp:\n                 model = pickle.load(fp)\n             pred.append(model.predict_proba(test[feature_lists])[:, ][])\n\n         sample_submission.loc[(sample_submission[]==session_id)&amp;(sample_submission[]==n), ] = mean(pred)\n\n\n    sample_submission[] = sample_submission[].apply( x:  x&gt;best_threshold  )\n    env.predict(sample_submission[[, ]])\n</code></pre>\n<p>Please help. Thank you in advance for your help.</p>",
      "rawMarkdown": "Hi kaggle staff and participants.\n\nI am in trouble.\n\nI predict the following.\n・q1-q3 are predicted using data with level_group '0-4'.\n・q4-q13 are predicted using data from level_group '0-4' and level_group '5-12'.\n・q14-q18 are predicted using data from level_group '0-4', level_group '5-12', and level_group \n   '13-22'\n\n**I am able to submit.\nHowever, about 3 hours after the first evaluation of the hidden test set, an Submission Scoring Error occurs.**\n\nHere is the pseudo code.\n\n```python\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nhistorical_meta = defaultdict(list)\nbest_threshold = 0.625\n\n#Declare a variable in the dictionary to store past level group data.\nsession_df_dict = {}\n\nfor (test, sample_submission) in iter_test:\n    grp = test.level_group.values[0]\n    session_id = test.session_id.values[0]\n\n    test = test.sort_values(by=\"index\").reset_index(drop=True)\n\n    #Processing test data.\n    test = (\n        pl.from_pandas(test)\n        .rename(\n            {'elapsed_time':'elapsed_time_total'}\n        )\n        .with_columns(\n            (pl.col('elapsed_time_total').sub(pl.col('elapsed_time_total').shift(1)))\n            .fill_null(0)\n            .clip(0, 1e9)\n            .over(['session_id', 'level_group'])\n            .alias('elapsed_time')\n        )\n        .with_columns(\n            (pl.col('room_coor_x').sub(pl.col('room_coor_x').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('room_coor_x_move')\n        )\n        .with_columns(\n            (pl.col('room_coor_y').sub(pl.col('room_coor_y').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('room_coor_y_move')\n        )\n        .with_columns(\n            (pl.col('screen_coor_x').sub(pl.col('screen_coor_x').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('screen_coor_x_move')\n        )\n        .with_columns(\n            (pl.col('screen_coor_y').sub(pl.col('screen_coor_y').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('screen_coor_y_move')\n        )\n        .with_columns(\n            pl.col('page').cast(pl.Float32),\n            pl.col('fqid').fill_null('fqid_None'),\n            pl.col('text_fqid').fill_null('text_fqid_None')\n        )\n    )\n    \n    #When the level group is '0-4'.\n    if grp == '0-4':\n\n        #Generate features from data with level groups 0-4.\n        test = _make_exog_easy(test, grp)\n        \n        #Let \"easy_str(session_id)\" be the key of the dictionary.\n        keys = 'easy_'+str(session_id)\n\n        #If \"easy_str(session_id)\" does not exist in the dictionary, add the above test data frame.\n        if keys not in session_df_dict.keys():\n            session_df_dict[keys] = test\n\n    #When the level group is '5-12'.\n    elif grp == '5-12':\n\n        #Generate features from data with level groups 0-4.\n        test = _make_exog_normal(test, grp)\n\n        #If the data of level_group'0-4' of the corresponding session ID is in the dictionary, it is combined with the data of level_group'5-12'.\n        keys = 'easy_'+str(session_id)\n        if keys in session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on='session_id', how='left')\n\n        #Let \"easy_normal_str(session_id)\" be the key of the dictionary.\n        keys = 'easy_normal_'+str(session_id)\n        \n        #If \"easy_normal_str(session_id)\" does not exist in the dictionary, add the above test data frame.\n        if keys not in session_df_dict.keys():\n            session_df_dict[keys] = test\n    \n    #When the level group is '13-22'.\n    elif grp == '13-22':\n        #Generate features from data with level groups 0-4.\n        test = _make_exog_hard(test, grp)\n    \n        #If the data of level_group'0-4' and level_group'5-12' of the corresponding session ID is in the dictionary, it is combined with the data of level_group'13-22'.\n        keys = 'easy_normal_'+str(session_id)\n        if keys in session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on='session_id', how='left')\n\n    test = test.fillna(0)\n        \n    sample_submission['session'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[0]))\n    sample_submission['q'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[-1][1:]))\n    sample_submission = sample_submission.sort_values(by='q')\n        \n    min_q, max_q = limits[grp]\n\n    for n in range(min_q, max_q):\n        feature_lists = importance_dict[str(n)]\n        col_lists = list(set(feature_lists)-set(list(test)))\n        if len(col_lists) > 0:\n            for var in col_lists:\n                test[var] = 0\n            \n         pred = []\n         for i in range(1, 6):\n             with open(f'../input/model/LGB_question_first{n}_{i}.lgb', mode='rb') as fp:\n                 model = pickle.load(fp)\n             pred.append(model.predict_proba(test[feature_lists])[:, 1][0])\n\n         sample_submission.loc[(sample_submission['session']==session_id)&(sample_submission['q']==n), 'correct'] = mean(pred)\n\n    \n    sample_submission['correct'] = sample_submission['correct'].apply(lambda x:1 if x>best_threshold else 0)\n    env.predict(sample_submission[['session_id', 'correct']])\n```\n\nPlease help. Thank you in advance for your help.",
      "votes": null
    },
    {
      "id": "2317822",
      "postDate": "06/26/2023 02:54:58",
      "content": "<p>Yesterday I also encountered <code>Submission Scoring Error</code> in my inference process. I checked my feature engineering code and finally found it was a <code>ZeroDivisionError</code>(there was a line of code like <code>df['c'] = df['a']/df['b']</code>). Maybe you could also check yours.</p>",
      "rawMarkdown": "Yesterday I also encountered `Submission Scoring Error` in my inference process. I checked my feature engineering code and finally found it was a `ZeroDivisionError`(there was a line of code like `df['c'] = df['a']/df['b']`). Maybe you could also check yours.",
      "votes": null
    },
    {
      "id": "2317825",
      "postDate": "06/26/2023 03:11:58",
      "content": "<p>Thank you!!</p>",
      "rawMarkdown": "Thank you!!",
      "votes": null
    },
    {
      "id": "2317829",
      "postDate": "06/26/2023 03:23:06",
      "content": "<p>The simplest possibility could be memory overflow. The session dictionary might get pretty big(?) and might be memory inefficient(?) and there's a very small memory cap at 8GB.</p>",
      "rawMarkdown": "The simplest possibility could be memory overflow. The session dictionary might get pretty big(?) and might be memory inefficient(?) and there's a very small memory cap at 8GB.",
      "votes": null
    },
    {
      "id": "2317858",
      "postDate": "06/26/2023 03:58:05",
      "content": "<p>Thank you!!</p>\n<p>I checked the following.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a></p>\n<p>I think that if memory overflow,  not Submission Scoring Error occurs, but Notebook Exceeded Allowed Compute occurs.</p>",
      "rawMarkdown": "Thank you!!\n\nI checked the following.\n[https://www.kaggle.com/code-competition-debugging](url)\n\nI think that if memory overflow,  not Submission Scoring Error occurs, but Notebook Exceeded Allowed Compute occurs.",
      "votes": null
    },
    {
      "id": "2318883",
      "postDate": "06/26/2023 16:35:17",
      "content": "<p>it didnt give the ZeroDivisionError ?</p>",
      "rawMarkdown": "it didnt give the ZeroDivisionError ?",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2317822,
      "author_name": "takanashihumbert",
      "author_url": "",
      "post_date": "06/26/2023 02:54:58",
      "content": "<p>Yesterday I also encountered <code>Submission Scoring Error</code> in my inference process. I checked my feature engineering code and finally found it was a <code>ZeroDivisionError</code>(there was a line of code like <code>df['c'] = df['a']/df['b']</code>). Maybe you could also check yours.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2317825,
          "author_name": "tomfuj",
          "author_url": "",
          "post_date": "06/26/2023 03:11:58",
          "content": "<p>Thank you!!</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2318883,
          "author_name": "qurious",
          "author_url": "",
          "post_date": "06/26/2023 16:35:17",
          "content": "<p>it didnt give the ZeroDivisionError ?</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2317829,
      "author_name": "roberthatch",
      "author_url": "",
      "post_date": "06/26/2023 03:23:06",
      "content": "<p>The simplest possibility could be memory overflow. The session dictionary might get pretty big(?) and might be memory inefficient(?) and there's a very small memory cap at 8GB.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2317858,
          "author_name": "tomfuj",
          "author_url": "",
          "post_date": "06/26/2023 03:58:05",
          "content": "<p>Thank you!!</p>\n<p>I checked the following.<br>\n<a href=\"url\" target=\"_blank\">https://www.kaggle.com/code-competition-debugging</a></p>\n<p>I think that if memory overflow,  not Submission Scoring Error occurs, but Notebook Exceeded Allowed Compute occurs.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2317806": "Hi kaggle staff and participants.\n\nI am in trouble.\n\nI predict the following.\n・q1-q3 are predicted using data with level_group '0-4'.\n・q4-q13 are predicted using data from level_group '0-4' and level_group '5-12'.\n・q14-q18 are predicted using data from level_group '0-4', level_group '5-12', and level_group \n   '13-22'\n\n**I am able to submit.\nHowever, about 3 hours after the first evaluation of the hidden test set, an Submission Scoring Error occurs.**\n\nHere is the pseudo code.\n\n```python\nenv = jo_wilder_310.make_env()\niter_test = env.iter_test()\n\nlimits = {'0-4':(1,4), '5-12':(4,14), '13-22':(14,19)}\nhistorical_meta = defaultdict(list)\nbest_threshold = 0.625\n\n#Declare a variable in the dictionary to store past level group data.\nsession_df_dict = {}\n\nfor (test, sample_submission) in iter_test:\n    grp = test.level_group.values[0]\n    session_id = test.session_id.values[0]\n\n    test = test.sort_values(by=\"index\").reset_index(drop=True)\n\n    #Processing test data.\n    test = (\n        pl.from_pandas(test)\n        .rename(\n            {'elapsed_time':'elapsed_time_total'}\n        )\n        .with_columns(\n            (pl.col('elapsed_time_total').sub(pl.col('elapsed_time_total').shift(1)))\n            .fill_null(0)\n            .clip(0, 1e9)\n            .over(['session_id', 'level_group'])\n            .alias('elapsed_time')\n        )\n        .with_columns(\n            (pl.col('room_coor_x').sub(pl.col('room_coor_x').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('room_coor_x_move')\n        )\n        .with_columns(\n            (pl.col('room_coor_y').sub(pl.col('room_coor_y').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('room_coor_y_move')\n        )\n        .with_columns(\n            (pl.col('screen_coor_x').sub(pl.col('screen_coor_x').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('screen_coor_x_move')\n        )\n        .with_columns(\n            (pl.col('screen_coor_y').sub(pl.col('screen_coor_y').shift(1)))\n            .fill_null(0)\n            .abs()\n            .over(['session_id', 'level_group'])\n            .alias('screen_coor_y_move')\n        )\n        .with_columns(\n            pl.col('page').cast(pl.Float32),\n            pl.col('fqid').fill_null('fqid_None'),\n            pl.col('text_fqid').fill_null('text_fqid_None')\n        )\n    )\n    \n    #When the level group is '0-4'.\n    if grp == '0-4':\n\n        #Generate features from data with level groups 0-4.\n        test = _make_exog_easy(test, grp)\n        \n        #Let \"easy_str(session_id)\" be the key of the dictionary.\n        keys = 'easy_'+str(session_id)\n\n        #If \"easy_str(session_id)\" does not exist in the dictionary, add the above test data frame.\n        if keys not in session_df_dict.keys():\n            session_df_dict[keys] = test\n\n    #When the level group is '5-12'.\n    elif grp == '5-12':\n\n        #Generate features from data with level groups 0-4.\n        test = _make_exog_normal(test, grp)\n\n        #If the data of level_group'0-4' of the corresponding session ID is in the dictionary, it is combined with the data of level_group'5-12'.\n        keys = 'easy_'+str(session_id)\n        if keys in session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on='session_id', how='left')\n\n        #Let \"easy_normal_str(session_id)\" be the key of the dictionary.\n        keys = 'easy_normal_'+str(session_id)\n        \n        #If \"easy_normal_str(session_id)\" does not exist in the dictionary, add the above test data frame.\n        if keys not in session_df_dict.keys():\n            session_df_dict[keys] = test\n    \n    #When the level group is '13-22'.\n    elif grp == '13-22':\n        #Generate features from data with level groups 0-4.\n        test = _make_exog_hard(test, grp)\n    \n        #If the data of level_group'0-4' and level_group'5-12' of the corresponding session ID is in the dictionary, it is combined with the data of level_group'13-22'.\n        keys = 'easy_normal_'+str(session_id)\n        if keys in session_df_dict.keys():\n             test = pd.merge(test, session_df_dict[keys], on='session_id', how='left')\n\n    test = test.fillna(0)\n        \n    sample_submission['session'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[0]))\n    sample_submission['q'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[-1][1:]))\n    sample_submission = sample_submission.sort_values(by='q')\n        \n    min_q, max_q = limits[grp]\n\n    for n in range(min_q, max_q):\n        feature_lists = importance_dict[str(n)]\n        col_lists = list(set(feature_lists)-set(list(test)))\n        if len(col_lists) > 0:\n            for var in col_lists:\n                test[var] = 0\n            \n         pred = []\n         for i in range(1, 6):\n             with open(f'../input/model/LGB_question_first{n}_{i}.lgb', mode='rb') as fp:\n                 model = pickle.load(fp)\n             pred.append(model.predict_proba(test[feature_lists])[:, 1][0])\n\n         sample_submission.loc[(sample_submission['session']==session_id)&(sample_submission['q']==n), 'correct'] = mean(pred)\n\n    \n    sample_submission['correct'] = sample_submission['correct'].apply(lambda x:1 if x>best_threshold else 0)\n    env.predict(sample_submission[['session_id', 'correct']])\n```\n\nPlease help. Thank you in advance for your help.",
    "2317822": "Yesterday I also encountered `Submission Scoring Error` in my inference process. I checked my feature engineering code and finally found it was a `ZeroDivisionError`(there was a line of code like `df['c'] = df['a']/df['b']`). Maybe you could also check yours.",
    "2317825": "Thank you!!",
    "2317829": "The simplest possibility could be memory overflow. The session dictionary might get pretty big(?) and might be memory inefficient(?) and there's a very small memory cap at 8GB.",
    "2317858": "Thank you!!\n\nI checked the following.\n[https://www.kaggle.com/code-competition-debugging](url)\n\nI think that if memory overflow,  not Submission Scoring Error occurs, but Notebook Exceeded Allowed Compute occurs.",
    "2318883": "it didnt give the ZeroDivisionError ?"
  },
  "source": "meta"
}