{
  "id": 388365,
  "title": "Saving features from level group '0-4' for later use",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/388365",
  "author_name": "",
  "post_date": "2023-02-17T05:02:19.035183700Z",
  "votes": 22,
  "comment_count": 21,
  "views": 0,
  "content": "<p>Hi there,</p>\n<p>My notebooks are throwing submission scoring error when I am trying to save &amp; load features between different level groups. The motivation: when predicting for level group 5-12, it makes sense to also use the features from level group 0-4.</p>\n<p>Needless to say, the commit works totally fine, it is only during the submission time that the error happens.</p>\n<p>First of all, I verified it has nothing to do with submission, but rather it is a standard notebook error. Hence, it should say <code>notebook threw exception</code> rather than <code>submission scoring error</code>. This is confusing.</p>\n<p>It happens after one minute of the submission, so it must be one of the very first sessions doing that.<br>\nIt seems to me that there is a session where level group 0-4 is presented after level group 5-12.</p>\n<p>Has anyone encountered something similar? Or am I missing something obvious here?</p>\n<p>This is the pseudo-code for the part that causes the problem:</p>\n<pre><code>session_features = {}\n (sample_submission, test)  iter_test:\n   sample_submission[] = sample_submission[].apply( x: (x.split()[-][:]) )\n   sample_submission[] = \n   sample_submission.loc[sample_submission[]&lt;=,] = \n   sample_submission.loc[sample_submission[]&gt;=,] = \n   cur_level_group = sample_submission[].values[]\n   sample_submission[] = sample_submission[].apply( x: (x.split()[]) )\n   cur_ssid = sample_submission[].values[]\n\n   \n     cur_level_group == :\n        session_features[cur_ssid] = {}\n        session_features[cur_ssid][cur_level_group] = features\n    :\n        session_features[cur_ssid][cur_level_group] = features\n   \n     cur_level_group == :\n        \n     cur_level_group == :\n        sample_submission= sample_submission.merge(session_features[cur_ssid][], how = , on = )\n    :\n        sample_submission = sample_submission.merge(session_features[cur_ssid][], how = , on = )\n        sample_submission = sample_submission.merge(session_features[cur_ssid][], how = , on = )\n</code></pre>",
  "messages": [
    {
      "id": "2148016",
      "postDate": "02/17/2023 05:02:19",
      "content": "<p>Hi there,</p>\n<p>My notebooks are throwing submission scoring error when I am trying to save &amp; load features between different level groups. The motivation: when predicting for level group 5-12, it makes sense to also use the features from level group 0-4.</p>\n<p>Needless to say, the commit works totally fine, it is only during the submission time that the error happens.</p>\n<p>First of all, I verified it has nothing to do with submission, but rather it is a standard notebook error. Hence, it should say <code>notebook threw exception</code> rather than <code>submission scoring error</code>. This is confusing.</p>\n<p>It happens after one minute of the submission, so it must be one of the very first sessions doing that.<br>\nIt seems to me that there is a session where level group 0-4 is presented after level group 5-12.</p>\n<p>Has anyone encountered something similar? Or am I missing something obvious here?</p>\n<p>This is the pseudo-code for the part that causes the problem:</p>\n<pre><code>session_features = {}\n (sample_submission, test)  iter_test:\n   sample_submission[] = sample_submission[].apply( x: (x.split()[-][:]) )\n   sample_submission[] = \n   sample_submission.loc[sample_submission[]&lt;=,] = \n   sample_submission.loc[sample_submission[]&gt;=,] = \n   cur_level_group = sample_submission[].values[]\n   sample_submission[] = sample_submission[].apply( x: (x.split()[]) )\n   cur_ssid = sample_submission[].values[]\n\n   \n     cur_level_group == :\n        session_features[cur_ssid] = {}\n        session_features[cur_ssid][cur_level_group] = features\n    :\n        session_features[cur_ssid][cur_level_group] = features\n   \n     cur_level_group == :\n        \n     cur_level_group == :\n        sample_submission= sample_submission.merge(session_features[cur_ssid][], how = , on = )\n    :\n        sample_submission = sample_submission.merge(session_features[cur_ssid][], how = , on = )\n        sample_submission = sample_submission.merge(session_features[cur_ssid][], how = , on = )\n</code></pre>",
      "rawMarkdown": "Hi there,\n\nMy notebooks are throwing submission scoring error when I am trying to save & load features between different level groups. The motivation: when predicting for level group 5-12, it makes sense to also use the features from level group 0-4.\n\nNeedless to say, the commit works totally fine, it is only during the submission time that the error happens.\n\nFirst of all, I verified it has nothing to do with submission, but rather it is a standard notebook error. Hence, it should say `notebook threw exception` rather than `submission scoring error`. This is confusing.\n\nIt happens after one minute of the submission, so it must be one of the very first sessions doing that.\nIt seems to me that there is a session where level group 0-4 is presented after level group 5-12.\n\nHas anyone encountered something similar? Or am I missing something obvious here?\n\nThis is the pseudo-code for the part that causes the problem:\n```python\nsession_features = {}\nfor (sample_submission, test) in iter_test:\n   sample_submission['q'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[-1][1:]) )\n   sample_submission['level_group'] = '5-12'\n   sample_submission.loc[sample_submission['q']<=3,'level_group'] = '0-4'\n   sample_submission.loc[sample_submission['q']>=14,'level_group'] = '13-22'\n   cur_level_group = sample_submission['level_group'].values[0]\n   sample_submission['ssid'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[0]) )\n   cur_ssid = sample_submission['ssid'].values[0]\n   \n   #save features\n    if cur_level_group == '0-4':\n        session_features[cur_ssid] = {}\n        session_features[cur_ssid][cur_level_group] = features\n    else:\n        session_features[cur_ssid][cur_level_group] = features\n   #load features\n    if cur_level_group == '0-4':\n        pass\n    elif cur_level_group == '5-12':\n        sample_submission= sample_submission.merge(session_features[cur_ssid]['0-4'], how = 'left', on = 'ssid')\n    else:\n        sample_submission = sample_submission.merge(session_features[cur_ssid]['0-4'], how = 'left', on = 'ssid')\n        sample_submission = sample_submission.merge(session_features[cur_ssid]['5-12'], how = 'left', on = 'ssid')\n```",
      "votes": null
    },
    {
      "id": "2148050",
      "postDate": "02/17/2023 05:41:22",
      "content": "<p>Maybe you are going OOM as the features are saved in a dictionary? I don't know if kaggle notebooks throw a different error for that. You could try saving and reading the features from disk.</p>",
      "rawMarkdown": "Maybe you are going OOM as the features are saved in a dictionary? I don't know if kaggle notebooks throw a different error for that. You could try saving and reading the features from disk.",
      "votes": null
    },
    {
      "id": "2148058",
      "postDate": "02/17/2023 05:47:27",
      "content": "<p>I got the same thing when i use dict to cache the previous level. Everything is ok in interactive session, but failed immediately in submitting.</p>",
      "rawMarkdown": "I got the same thing when i use dict to cache the previous level. Everything is ok in interactive session, but failed immediately in submitting.",
      "votes": null
    },
    {
      "id": "2148059",
      "postDate": "02/17/2023 05:47:30",
      "content": "<p>Thanks for your suggestion, but shouldn't then the error be <code>memory error</code>? Also, it happens within one minute since the submission start, and I doubt the instance would run out of memory so fast</p>",
      "rawMarkdown": "Thanks for your suggestion, but shouldn't then the error be `memory error`? Also, it happens within one minute since the submission start, and I doubt the instance would run out of memory so fast",
      "votes": null
    },
    {
      "id": "2148060",
      "postDate": "02/17/2023 05:48:00",
      "content": "<p>Hallelujah! I am not crazy :) </p>",
      "rawMarkdown": "Hallelujah! I am not crazy :)",
      "votes": null
    },
    {
      "id": "2148062",
      "postDate": "02/17/2023 05:48:21",
      "content": "<p>did you manage to find a workaround?</p>",
      "rawMarkdown": "did you manage to find a workaround?",
      "votes": null
    },
    {
      "id": "2148069",
      "postDate": "02/17/2023 05:56:07",
      "content": "<p>I don't know whether some bugs in kaggle api or just not allow us to use previous levels when submitting.</p>",
      "rawMarkdown": "I don't know whether some bugs in kaggle api or just not allow us to use previous levels when submitting.",
      "votes": null
    },
    {
      "id": "2148080",
      "postDate": "02/17/2023 06:03:37",
      "content": "<p>But it ok when I just concat the previous levels dataframe together when loop starts.</p>",
      "rawMarkdown": "But it ok when I just concat the previous levels dataframe together when loop starts.",
      "votes": null
    },
    {
      "id": "2148100",
      "postDate": "02/17/2023 06:20:30",
      "content": "<p>How do you store dataframe from previous levels?</p>",
      "rawMarkdown": "How do you store dataframe from previous levels?",
      "votes": null
    },
    {
      "id": "2148112",
      "postDate": "02/17/2023 06:28:07",
      "content": "<p>Τhe only error that is clear is memory error and you get a message. My take :<br>\n1..Simulate submission loop with the test data.Watch the grouping , every batch must have one USER and one LEVEL<br>\n2..Watch the masks, if you use them, in submission loops</p>",
      "rawMarkdown": "Τhe only error that is clear is memory error and you get a message. My take :\n1..Simulate submission loop with the test data.Watch the grouping , every batch must have one USER and one LEVEL\n2..Watch the masks, if you use them, in submission loops",
      "votes": null
    },
    {
      "id": "2148113",
      "postDate": "02/17/2023 06:30:12",
      "content": "<p>just use nested dict.</p>",
      "rawMarkdown": "just use nested dict.",
      "votes": null
    },
    {
      "id": "2148115",
      "postDate": "02/17/2023 06:31:27",
      "content": "<p>The notebook works perfectly with available test data.</p>\n<p>But it breaks on the hidden test data.</p>",
      "rawMarkdown": "The notebook works perfectly with available test data.\n\nBut it breaks on the hidden test data.",
      "votes": null
    },
    {
      "id": "2148275",
      "postDate": "02/17/2023 09:57:47",
      "content": "<p>so do you store all raw events in memory? does it not run out of memory?</p>",
      "rawMarkdown": "so do you store all raw events in memory? does it not run out of memory?",
      "votes": null
    },
    {
      "id": "2148367",
      "postDate": "02/17/2023 11:31:07",
      "content": "<p>just update every session's status, not load the whole rows.</p>",
      "rawMarkdown": "just update every session's status, not load the whole rows.",
      "votes": null
    },
    {
      "id": "2148385",
      "postDate": "02/17/2023 11:54:13",
      "content": "<p>Hi qyxs, could you please describe \"nested dict\" and \"update every session's status\" in detail👀? I also got the same error.</p>",
      "rawMarkdown": "Hi qyxs, could you please describe \"nested dict\" and \"update every session's status\" in detail👀? I also got the same error.",
      "votes": null
    },
    {
      "id": "2148989",
      "postDate": "02/17/2023 20:11:28",
      "content": "<p>Chris said that in Kaggle API, <code>level_group</code> is fetched the other way around: first is <code>13-22</code>, then <code>5-12</code>, and finally <code>0-4</code> (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479)</a>.</p>\n<p>I haven't found out why he figured it out, but my guess is that he tried multiple caching direction (0-4 to 5-12 to 13-22; 0-4 to 13-22 to 5-12; 13-22 to 5-12 to 0-4; etc.) and found out the one that works (i.e. submission not throwing an error).</p>\n<p>Either way, if the order is not random for each user batch and follows a strict order, we will have some kind of other-group data to work with. However I personally prefer the chronological order, which is naturally how the final product is going to be.</p>",
      "rawMarkdown": "Chris said that in Kaggle API, `level_group` is fetched the other way around: first is `13-22`, then `5-12`, and finally `0-4` (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479).\n\nI haven't found out why he figured it out, but my guess is that he tried multiple caching direction (0-4 to 5-12 to 13-22; 0-4 to 13-22 to 5-12; 13-22 to 5-12 to 0-4; etc.) and found out the one that works (i.e. submission not throwing an error).\n\nEither way, if the order is not random for each user batch and follows a strict order, we will have some kind of other-group data to work with. However I personally prefer the chronological order, which is naturally how the final product is going to be.",
      "votes": null
    },
    {
      "id": "2149084",
      "postDate": "02/17/2023 23:35:24",
      "content": "<p>Thanks, this confirms there is an unintended leak</p>",
      "rawMarkdown": "Thanks, this confirms there is an unintended leak",
      "votes": null
    },
    {
      "id": "2149177",
      "postDate": "02/18/2023 03:26:13",
      "content": "<p>If I'm correct, I think your code assumes that a batch fetched by the Kaggle API contains one level group of one user at a time. Have you checked to confirm that it's not one level group for multiple users?</p>",
      "rawMarkdown": "If I'm correct, I think your code assumes that a batch fetched by the Kaggle API contains one level group of one user at a time. Have you checked to confirm that it's not one level group for multiple users?",
      "votes": null
    },
    {
      "id": "2149317",
      "postDate": "02/18/2023 06:56:23",
      "content": "<p>I've just checked myself. Confirmed that it is one session per batch.</p>\n<p><strong>However</strong>, this may change when the API is updated due to <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479\" target=\"_blank\">leakage</a>.</p>",
      "rawMarkdown": "I've just checked myself. Confirmed that it is one session per batch.\n\n**However**, this may change when the API is updated due to [leakage](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479).",
      "votes": null
    },
    {
      "id": "2222819",
      "postDate": "04/15/2023 15:30:13",
      "content": "<p>Did you encounter any issues while trying to implement this with the latest API changes? I tried it on my own, but it didn't work. However, I'm not sure if the issue lies in the code or my side inference. </p>",
      "rawMarkdown": "Did you encounter any issues while trying to implement this with the latest API changes? I tried it on my own, but it didn't work. However, I'm not sure if the issue lies in the code or my side inference.",
      "votes": null
    },
    {
      "id": "2271797",
      "postDate": "05/24/2023 06:26:54",
      "content": "<p>hey Narsil!, Did you figured it out, why the submission scoring error is happening?<br>\nI am getting the same submission scoring error, when i submit my notebook for prediction </p>",
      "rawMarkdown": "hey Narsil!, Did you figured it out, why the submission scoring error is happening?\nI am getting the same submission scoring error, when i submit my notebook for prediction",
      "votes": null
    },
    {
      "id": "2282682",
      "postDate": "05/31/2023 18:31:57",
      "content": "<p>Hi <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> , may i know if your error is solved now ? and how is the solution ?  Thanks for your information advance ,</p>",
      "rawMarkdown": "Hi @juzqyxs , may i know if your error is solved now ? and how is the solution ?  Thanks for your information advance ,",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2148050,
      "author_name": "samfc10",
      "author_url": "",
      "post_date": "02/17/2023 05:41:22",
      "content": "<p>Maybe you are going OOM as the features are saved in a dictionary? I don't know if kaggle notebooks throw a different error for that. You could try saving and reading the features from disk.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2148059,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/17/2023 05:47:30",
          "content": "<p>Thanks for your suggestion, but shouldn't then the error be <code>memory error</code>? Also, it happens within one minute since the submission start, and I doubt the instance would run out of memory so fast</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2148058,
      "author_name": "juzqyxs",
      "author_url": "",
      "post_date": "02/17/2023 05:47:27",
      "content": "<p>I got the same thing when i use dict to cache the previous level. Everything is ok in interactive session, but failed immediately in submitting.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2148060,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/17/2023 05:48:00",
          "content": "<p>Hallelujah! I am not crazy :) </p>",
          "votes": null,
          "replies": [
            {
              "id": 2148062,
              "author_name": "narsil",
              "author_url": "",
              "post_date": "02/17/2023 05:48:21",
              "content": "<p>did you manage to find a workaround?</p>",
              "votes": null,
              "replies": [
                {
                  "id": 2148069,
                  "author_name": "juzqyxs",
                  "author_url": "",
                  "post_date": "02/17/2023 05:56:07",
                  "content": "<p>I don't know whether some bugs in kaggle api or just not allow us to use previous levels when submitting.</p>",
                  "votes": null,
                  "replies": []
                },
                {
                  "id": 2148080,
                  "author_name": "juzqyxs",
                  "author_url": "",
                  "post_date": "02/17/2023 06:03:37",
                  "content": "<p>But it ok when I just concat the previous levels dataframe together when loop starts.</p>",
                  "votes": null,
                  "replies": [
                    {
                      "id": 2148100,
                      "author_name": "narsil",
                      "author_url": "",
                      "post_date": "02/17/2023 06:20:30",
                      "content": "<p>How do you store dataframe from previous levels?</p>",
                      "votes": null,
                      "replies": [
                        {
                          "id": 2148113,
                          "author_name": "juzqyxs",
                          "author_url": "",
                          "post_date": "02/17/2023 06:30:12",
                          "content": "<p>just use nested dict.</p>",
                          "votes": null,
                          "replies": [
                            {
                              "id": 2148275,
                              "author_name": "narsil",
                              "author_url": "",
                              "post_date": "02/17/2023 09:57:47",
                              "content": "<p>so do you store all raw events in memory? does it not run out of memory?</p>",
                              "votes": null,
                              "replies": [
                                {
                                  "id": 2148367,
                                  "author_name": "juzqyxs",
                                  "author_url": "",
                                  "post_date": "02/17/2023 11:31:07",
                                  "content": "<p>just update every session's status, not load the whole rows.</p>",
                                  "votes": null,
                                  "replies": []
                                }
                              ]
                            }
                          ]
                        }
                      ]
                    }
                  ]
                }
              ]
            }
          ]
        },
        {
          "id": 2148385,
          "author_name": "carnozhao",
          "author_url": "",
          "post_date": "02/17/2023 11:54:13",
          "content": "<p>Hi qyxs, could you please describe \"nested dict\" and \"update every session's status\" in detail👀? I also got the same error.</p>",
          "votes": null,
          "replies": []
        },
        {
          "id": 2282682,
          "author_name": "johnfjliu",
          "author_url": "",
          "post_date": "05/31/2023 18:31:57",
          "content": "<p>Hi <a href=\"https://www.kaggle.com/juzqyxs\" target=\"_blank\">@juzqyxs</a> , may i know if your error is solved now ? and how is the solution ?  Thanks for your information advance ,</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2148112,
      "author_name": "georgem",
      "author_url": "",
      "post_date": "02/17/2023 06:28:07",
      "content": "<p>Τhe only error that is clear is memory error and you get a message. My take :<br>\n1..Simulate submission loop with the test data.Watch the grouping , every batch must have one USER and one LEVEL<br>\n2..Watch the masks, if you use them, in submission loops</p>",
      "votes": null,
      "replies": [
        {
          "id": 2148115,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/17/2023 06:31:27",
          "content": "<p>The notebook works perfectly with available test data.</p>\n<p>But it breaks on the hidden test data.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2148989,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "02/17/2023 20:11:28",
      "content": "<p>Chris said that in Kaggle API, <code>level_group</code> is fetched the other way around: first is <code>13-22</code>, then <code>5-12</code>, and finally <code>0-4</code> (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479)\" target=\"_blank\">https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479)</a>.</p>\n<p>I haven't found out why he figured it out, but my guess is that he tried multiple caching direction (0-4 to 5-12 to 13-22; 0-4 to 13-22 to 5-12; 13-22 to 5-12 to 0-4; etc.) and found out the one that works (i.e. submission not throwing an error).</p>\n<p>Either way, if the order is not random for each user batch and follows a strict order, we will have some kind of other-group data to work with. However I personally prefer the chronological order, which is naturally how the final product is going to be.</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149084,
          "author_name": "narsil",
          "author_url": "",
          "post_date": "02/17/2023 23:35:24",
          "content": "<p>Thanks, this confirms there is an unintended leak</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2149177,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "02/18/2023 03:26:13",
      "content": "<p>If I'm correct, I think your code assumes that a batch fetched by the Kaggle API contains one level group of one user at a time. Have you checked to confirm that it's not one level group for multiple users?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2149317,
          "author_name": "hoangnguyen719",
          "author_url": "",
          "post_date": "02/18/2023 06:56:23",
          "content": "<p>I've just checked myself. Confirmed that it is one session per batch.</p>\n<p><strong>However</strong>, this may change when the API is updated due to <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479\" target=\"_blank\">leakage</a>.</p>",
          "votes": null,
          "replies": []
        }
      ]
    },
    {
      "id": 2222819,
      "author_name": "thaweewatboy",
      "author_url": "",
      "post_date": "04/15/2023 15:30:13",
      "content": "<p>Did you encounter any issues while trying to implement this with the latest API changes? I tried it on my own, but it didn't work. However, I'm not sure if the issue lies in the code or my side inference. </p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2271797,
      "author_name": "navinkumarmnk",
      "author_url": "",
      "post_date": "05/24/2023 06:26:54",
      "content": "<p>hey Narsil!, Did you figured it out, why the submission scoring error is happening?<br>\nI am getting the same submission scoring error, when i submit my notebook for prediction </p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2148016": "Hi there,\n\nMy notebooks are throwing submission scoring error when I am trying to save & load features between different level groups. The motivation: when predicting for level group 5-12, it makes sense to also use the features from level group 0-4.\n\nNeedless to say, the commit works totally fine, it is only during the submission time that the error happens.\n\nFirst of all, I verified it has nothing to do with submission, but rather it is a standard notebook error. Hence, it should say `notebook threw exception` rather than `submission scoring error`. This is confusing.\n\nIt happens after one minute of the submission, so it must be one of the very first sessions doing that.\nIt seems to me that there is a session where level group 0-4 is presented after level group 5-12.\n\nHas anyone encountered something similar? Or am I missing something obvious here?\n\nThis is the pseudo-code for the part that causes the problem:\n```python\nsession_features = {}\nfor (sample_submission, test) in iter_test:\n   sample_submission['q'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[-1][1:]) )\n   sample_submission['level_group'] = '5-12'\n   sample_submission.loc[sample_submission['q']<=3,'level_group'] = '0-4'\n   sample_submission.loc[sample_submission['q']>=14,'level_group'] = '13-22'\n   cur_level_group = sample_submission['level_group'].values[0]\n   sample_submission['ssid'] = sample_submission['session_id'].apply(lambda x: int(x.split('_')[0]) )\n   cur_ssid = sample_submission['ssid'].values[0]\n   \n   #save features\n    if cur_level_group == '0-4':\n        session_features[cur_ssid] = {}\n        session_features[cur_ssid][cur_level_group] = features\n    else:\n        session_features[cur_ssid][cur_level_group] = features\n   #load features\n    if cur_level_group == '0-4':\n        pass\n    elif cur_level_group == '5-12':\n        sample_submission= sample_submission.merge(session_features[cur_ssid]['0-4'], how = 'left', on = 'ssid')\n    else:\n        sample_submission = sample_submission.merge(session_features[cur_ssid]['0-4'], how = 'left', on = 'ssid')\n        sample_submission = sample_submission.merge(session_features[cur_ssid]['5-12'], how = 'left', on = 'ssid')\n```",
    "2148050": "Maybe you are going OOM as the features are saved in a dictionary? I don't know if kaggle notebooks throw a different error for that. You could try saving and reading the features from disk.",
    "2148058": "I got the same thing when i use dict to cache the previous level. Everything is ok in interactive session, but failed immediately in submitting.",
    "2148059": "Thanks for your suggestion, but shouldn't then the error be `memory error`? Also, it happens within one minute since the submission start, and I doubt the instance would run out of memory so fast",
    "2148060": "Hallelujah! I am not crazy :)",
    "2148062": "did you manage to find a workaround?",
    "2148069": "I don't know whether some bugs in kaggle api or just not allow us to use previous levels when submitting.",
    "2148080": "But it ok when I just concat the previous levels dataframe together when loop starts.",
    "2148100": "How do you store dataframe from previous levels?",
    "2148112": "Τhe only error that is clear is memory error and you get a message. My take :\n1..Simulate submission loop with the test data.Watch the grouping , every batch must have one USER and one LEVEL\n2..Watch the masks, if you use them, in submission loops",
    "2148113": "just use nested dict.",
    "2148115": "The notebook works perfectly with available test data.\n\nBut it breaks on the hidden test data.",
    "2148275": "so do you store all raw events in memory? does it not run out of memory?",
    "2148367": "just update every session's status, not load the whole rows.",
    "2148385": "Hi qyxs, could you please describe \"nested dict\" and \"update every session's status\" in detail👀? I also got the same error.",
    "2148989": "Chris said that in Kaggle API, `level_group` is fetched the other way around: first is `13-22`, then `5-12`, and finally `0-4` (https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479).\n\nI haven't found out why he figured it out, but my guess is that he tried multiple caching direction (0-4 to 5-12 to 13-22; 0-4 to 13-22 to 5-12; 13-22 to 5-12 to 0-4; etc.) and found out the one that works (i.e. submission not throwing an error).\n\nEither way, if the order is not random for each user batch and follows a strict order, we will have some kind of other-group data to work with. However I personally prefer the chronological order, which is naturally how the final product is going to be.",
    "2149084": "Thanks, this confirms there is an unintended leak",
    "2149177": "If I'm correct, I think your code assumes that a batch fetched by the Kaggle API contains one level group of one user at a time. Have you checked to confirm that it's not one level group for multiple users?",
    "2149317": "I've just checked myself. Confirmed that it is one session per batch.\n\n**However**, this may change when the API is updated due to [leakage](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/388479).",
    "2222819": "Did you encounter any issues while trying to implement this with the latest API changes? I tried it on my own, but it didn't work. However, I'm not sure if the issue lies in the code or my side inference.",
    "2271797": "hey Narsil!, Did you figured it out, why the submission scoring error is happening?\nI am getting the same submission scoring error, when i submit my notebook for prediction",
    "2282682": "Hi @juzqyxs , may i know if your error is solved now ? and how is the solution ?  Thanks for your information advance ,"
  },
  "source": "meta"
}