{
  "id": 398250,
  "title": "No more data leak in the new test set?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/398250",
  "author_name": "Hoang Nguyen",
  "post_date": "2023-03-29T07:49:21.795000",
  "votes": 6,
  "comment_count": 7,
  "views": 0,
  "content": "<p>I made a few experimental submissions and noticed 3 points below about the <strong>new</strong> test API</p>\n<ol>\n<li>Each batch contains only one <code>session_id</code></li>\n<li>Each batch contains only one <code>level_group</code></li>\n<li>Each batch contains only one <code>event_name=checkpoint</code></li>\n</ol>\n<p>RE point 1 and 2: these have been confirmed in several discussions (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/397414\" target=\"_blank\">link 1</a>, <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732\" target=\"_blank\">link 2</a>).<br>\nRE point 3: This implies that the <strong>new</strong> hidden test set does NOT contain users who <strong>complete</strong> a chapter multiple times, which is a leak the train set has. However, there can still be a corner case where user plays a chapter more than once but completes only once (i.e. they restart in the middle of chapter), and I don't know how to check for this scenario. My questions are</p>\n<ul>\n<li>Do you have a suggestion on how to check the chapter-restart case?</li>\n<li>If you have looked into whether the test set contains repeated gameplay or data leak in general, what did you find?</li>\n</ul>\n<p>TIA!</p>",
  "messages": [
    {
      "id": 2201300,
      "postDate": "2023-03-29T07:49:21.797Z",
      "content": "<p>I made a few experimental submissions and noticed 3 points below about the <strong>new</strong> test API</p>\n<ol>\n<li>Each batch contains only one <code>session_id</code></li>\n<li>Each batch contains only one <code>level_group</code></li>\n<li>Each batch contains only one <code>event_name=checkpoint</code></li>\n</ol>\n<p>RE point 1 and 2: these have been confirmed in several discussions (<a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/397414\" target=\"_blank\">link 1</a>, <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732\" target=\"_blank\">link 2</a>).<br>\nRE point 3: This implies that the <strong>new</strong> hidden test set does NOT contain users who <strong>complete</strong> a chapter multiple times, which is a leak the train set has. However, there can still be a corner case where user plays a chapter more than once but completes only once (i.e. they restart in the middle of chapter), and I don't know how to check for this scenario. My questions are</p>\n<ul>\n<li>Do you have a suggestion on how to check the chapter-restart case?</li>\n<li>If you have looked into whether the test set contains repeated gameplay or data leak in general, what did you find?</li>\n</ul>\n<p>TIA!</p>",
      "rawMarkdown": "I made a few experimental submissions and noticed 3 points below about the **new** test API\n1. Each batch contains only one `session_id`\n2. Each batch contains only one `level_group`\n3. Each batch contains only one `event_name=checkpoint`\n\nRE point 1 and 2: these have been confirmed in several discussions ([link 1](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/397414), [link 2](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732)).\nRE point 3: This implies that the **new** hidden test set does NOT contain users who **complete** a chapter multiple times, which is a leak the train set has. However, there can still be a corner case where user plays a chapter more than once but completes only once (i.e. they restart in the middle of chapter), and I don't know how to check for this scenario. My questions are\n- Do you have a suggestion on how to check the chapter-restart case?\n- If you have looked into whether the test set contains repeated gameplay or data leak in general, what did you find?\n\nTIA!",
      "votes": 5
    },
    {
      "id": 2201471,
      "postDate": "2023-03-29T10:38:29.560Z",
      "content": "<p>Could this work?</p>\n<pre><code>for (test, sample_submission) in iter_test:\n    min_level_diff = min(test[\"level\"].diff().values)\n    if min_level_diff &lt; 0:\n        sample_submission.loc[:, \"correct\"] = 1\n    else:\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission[['session_id','correct']])\n</code></pre>\n<p>If there is a restart, the <code>min_level_diff</code> should be less than 0, and we make all <code>correct = 1</code>. The LB value may not be 0.216.</p>",
      "rawMarkdown": "Could this work?\n```\nfor (test, sample_submission) in iter_test:\n    min_level_diff = min(test[\"level\"].diff().values)\n    if min_level_diff < 0:\n        sample_submission.loc[:, \"correct\"] = 1\n    else:\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission[['session_id','correct']])\n```\nIf there is a restart, the `min_level_diff` should be less than 0, and we make all `correct = 1`. The LB value may not be 0.216.",
      "votes": 1,
      "replies": [
        {
          "id": 2201967,
          "postDate": "2023-03-29T16:45:28.373Z",
          "rawMarkdown": "",
          "isDeleted": true,
          "replies": [
            {
              "id": 2202051,
              "postDate": "2023-03-29T17:42:32.207Z",
              "rawMarkdown": "",
              "isDeleted": true
            }
          ]
        },
        {
          "id": 2202386,
          "postDate": "2023-03-30T02:46:31.103Z",
          "content": "<p>Good idea! I checked using <a href=\"https://www.kaggle.com/code/hoangnguyen719/pspfgp-api-check-check-if-gameplay-restarts\" target=\"_blank\">this notebook</a> and the LB is 0.216, so I believe (1) <strong>there is no restart within a batch</strong>.</p>\n<p>Additionally, I check for <em>restarts across batches</em> using the below code snippet and the LB is 0.216, which suggests (2) <strong>no restart across batches</strong>. From (1) and (2) (unless my code is incorrect anywhere) I suppose there is no leak in the new API?</p>\n<p>Did you find anything?</p>\n<pre><code>prediction_value = 0 # default prediction value\nseen_ssids = {} # cache of seen ssids\n\nfor (test, sample_submission) in iter_test:\n    cur_lvl_grp = test.level_group.values[0]\n    ssid = test.session_id.values[0]\n    cache = seen_ssids.get(ssid)\n\n    if cache is None: # if user never seen, add them\n        seen_ssids[ssid] = [cur_lvl_grp]\n    else:\n        last_lvl_grp = cache[-1]\n        if (\n            ((last_lvl_grp == '0-4') &amp; (cur_lvl_grp == '0-4')) |\n            ((last_lvl_grp == '5-12') &amp; (cur_lvl_grp in ['0-4', '5-12'])) |\n            (last_lvl_grp == '13-22')\n        ): # if level_group repeats, change prediction value to 1\n            prediction_value = 1\n        else: # otherwise append current level group\n            seen_ssids[ssid].append(cur_lvl_grp)\n\n    sample_submission['correct'] = prediction_value\n    env.predict(sample_submission)\n</code></pre>",
          "rawMarkdown": "Good idea! I checked using [this notebook](https://www.kaggle.com/code/hoangnguyen719/pspfgp-api-check-check-if-gameplay-restarts) and the LB is 0.216, so I believe (1) **there is no restart within a batch**.\n\nAdditionally, I check for *restarts across batches* using the below code snippet and the LB is 0.216, which suggests (2) **no restart across batches**. From (1) and (2) (unless my code is incorrect anywhere) I suppose there is no leak in the new API?\n\nDid you find anything?\n```\nprediction_value = 0 # default prediction value\nseen_ssids = {} # cache of seen ssids\n\nfor (test, sample_submission) in iter_test:\n    cur_lvl_grp = test.level_group.values[0]\n    ssid = test.session_id.values[0]\n    cache = seen_ssids.get(ssid)\n    \n    if cache is None: # if user never seen, add them\n        seen_ssids[ssid] = [cur_lvl_grp]\n    else:\n        last_lvl_grp = cache[-1]\n        if (\n            ((last_lvl_grp == '0-4') & (cur_lvl_grp == '0-4')) |\n            ((last_lvl_grp == '5-12') & (cur_lvl_grp in ['0-4', '5-12'])) |\n            (last_lvl_grp == '13-22')\n        ): # if level_group repeats, change prediction value to 1\n            prediction_value = 1\n        else: # otherwise append current level group\n            seen_ssids[ssid].append(cur_lvl_grp)\n\n    sample_submission['correct'] = prediction_value\n    env.predict(sample_submission)\n```",
          "replies": [
            {
              "id": 2205421,
              "postDate": "2023-04-01T14:30:43.983Z",
              "content": "<p>Your comment \"(1) there is no restart within a batch.\" is not true.  Code <code>min(test[\"level\"].diff().values)</code> become nan, so <code>min_level_diff &lt; 0</code> is always False.  I modified the code to fill nan and the LB become 0.219.</p>",
              "rawMarkdown": "Your comment \"(1) there is no restart within a batch.\" is not true.  Code `min(test[\"level\"].diff().values)` become nan, so `min_level_diff < 0` is always False.  I modified the code to fill nan and the LB become 0.219.",
              "votes": 3
            },
            {
              "id": 2207230,
              "postDate": "2023-04-03T08:31:56.123Z",
              "content": "<p><a href=\"https://www.kaggle.com/maruichi01\" target=\"_blank\">@maruichi01</a> Interesting - thanks for flagging that! Did you have a chance to check my other assumptions:</p>\n<ol>\n<li>Each batch contains only one session_id</li>\n<li>Each batch contains only one level_group</li>\n<li>Each batch contains only one event_name=checkpoint</li>\n<li>There is no restart across batches</li>\n</ol>",
              "rawMarkdown": "@maruichi01 Interesting - thanks for flagging that! Did you have a chance to check my other assumptions:\n1. Each batch contains only one session_id\n2. Each batch contains only one level_group\n3. Each batch contains only one event_name=checkpoint\n4. There is no restart across batches"
            }
          ]
        }
      ]
    },
    {
      "id": 2221679,
      "postDate": "2023-04-14T13:05:27.520Z",
      "content": "<p>My notes:</p>\n<ul>\n<li><strong>Each batch contains only one <code>event_name=checkpoint</code>.</strong> I ran this code and got <code>LB 0.216</code></li>\n</ul>\n<pre><code>leak = \n (test, sample_submission)  iter_test:\n  checkpoints = (test[]==).()\n   checkpoints != :\n      leak = \n   leak:\n      sample_submission[] = \n  :\n      sample_submission[] = \n  env.predict(sample_submission)\n</code></pre>\n<ul>\n<li><strong>Some people restarted the chapter</strong> (This code got <code>LB 0.463</code> ): </li>\n</ul>\n<pre><code>leak = \n (test, sample_submission)  iter_test:\n  level_diff = (test[].diff().unique()&lt;)\n   level_diff &gt; :\n      leak =    \n   leak:\n      sample_submission[] = \n  :\n      sample_submission[] = \n  env.predict(sample_submission)\n</code></pre>\n<p>if there were no restarts, then the result of <code>test['level'].diff().unique()</code> would be <code>[0, 1, nan]</code>.  <code>sum(test['level'].diff().unique()&lt;0)</code> means that for some sessions the difference between the current level and the level of the previous record is less than 0, which means restarting the chapter</p>\n<ul>\n<li><strong>Some users missed levels (e.g. firstly they were at level 1, and then they immediately went to level 3)</strong> Perhaps this is related to the previous point. (This code get <code>LB 0.43</code>)</li>\n</ul>\n<pre><code>leak = \n (test, sample_submission)  iter_test:\n  positive_level_diff = (test[].diff().unique()&gt;)\n   positive_level_diff &gt; :\n      leak = \n   leak:\n      sample_submission[] = \n  :\n      sample_submission[] = \n  env.predict(sample_submission)\n</code></pre>\n<p>if there were no level skips, then the result of <code>test['level'].diff().unique()</code> would be <code>[0, 1, nan]</code> 1 means moving to a next level. <code>sum(test['level'].diff().unique()&gt;1)&gt;0</code> means that for some sessions the difference between the current level and the level of the previous record is greater than 1, which means skipping the level</p>",
      "rawMarkdown": "My notes:\n-  **Each batch contains only one `event_name=checkpoint`.** I ran this code and got `LB 0.216`\n\n```python\nleak = False\nfor (test, sample_submission) in iter_test:\n  checkpoints = (test[\"event_name\"]=='checkpoint').sum()\n  if checkpoints !=1 :\n      leak = True\n  if leak:\n      sample_submission['correct'] = 1\n  else:\n      sample_submission['correct'] = 0\n  env.predict(sample_submission)\n```\n-  **Some people restarted the chapter** (This code got ` LB 0.463` ): \n  ```python\nleak = False\nfor (test, sample_submission) in iter_test:\n  level_diff = sum(test['level'].diff().unique()<0)\n  if level_diff >0 :\n      leak = True   \n  if leak:\n      sample_submission['correct'] = 1\n  else:\n      sample_submission['correct'] = 0\n  env.predict(sample_submission)\n```\n\nif there were no restarts, then the result of `test['level'].diff().unique()` would be `[0, 1, nan]`.  `sum(test['level'].diff().unique()<0)` means that for some sessions the difference between the current level and the level of the previous record is less than 0, which means restarting the chapter\n- **Some users missed levels (e.g. firstly they were at level 1, and then they immediately went to level 3)** Perhaps this is related to the previous point. (This code get `LB 0.43`)\n  ```python\nleak = False\nfor (test, sample_submission) in iter_test:\n  positive_level_diff = sum(test['level'].diff().unique()>1)\n  if positive_level_diff >0 :\n      leak = True\n  if leak:\n      sample_submission['correct'] = 1\n  else:\n      sample_submission['correct'] = 0\n  env.predict(sample_submission)\n```\nif there were no level skips, then the result of `test['level'].diff().unique()` would be `[0, 1, nan]` 1 means moving to a next level. `sum(test['level'].diff().unique()>1)>0` means that for some sessions the difference between the current level and the level of the previous record is greater than 1, which means skipping the level",
      "votes": 2
    }
  ],
  "comments": [
    {
      "id": 2201471,
      "author_name": "Joseph Zhou",
      "author_url": "",
      "post_date": "2023-03-29T10:38:29.560000",
      "content": "<p>Could this work?</p>\n<pre><code>for (test, sample_submission) in iter_test:\n    min_level_diff = min(test[\"level\"].diff().values)\n    if min_level_diff &lt; 0:\n        sample_submission.loc[:, \"correct\"] = 1\n    else:\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission[['session_id','correct']])\n</code></pre>\n<p>If there is a restart, the <code>min_level_diff</code> should be less than 0, and we make all <code>correct = 1</code>. The LB value may not be 0.216.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 2201967,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-03-29T16:45:28.373000",
          "content": "",
          "votes": 0,
          "replies": [
            {
              "id": 2202051,
              "author_name": "",
              "author_url": "",
              "post_date": "2023-03-29T17:42:32.207000",
              "content": "",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2202386,
          "author_name": "Hoang Nguyen",
          "author_url": "",
          "post_date": "2023-03-30T02:46:31.103000",
          "content": "<p>Good idea! I checked using <a href=\"https://www.kaggle.com/code/hoangnguyen719/pspfgp-api-check-check-if-gameplay-restarts\" target=\"_blank\">this notebook</a> and the LB is 0.216, so I believe (1) <strong>there is no restart within a batch</strong>.</p>\n<p>Additionally, I check for <em>restarts across batches</em> using the below code snippet and the LB is 0.216, which suggests (2) <strong>no restart across batches</strong>. From (1) and (2) (unless my code is incorrect anywhere) I suppose there is no leak in the new API?</p>\n<p>Did you find anything?</p>\n<pre><code>prediction_value = 0 # default prediction value\nseen_ssids = {} # cache of seen ssids\n\nfor (test, sample_submission) in iter_test:\n    cur_lvl_grp = test.level_group.values[0]\n    ssid = test.session_id.values[0]\n    cache = seen_ssids.get(ssid)\n\n    if cache is None: # if user never seen, add them\n        seen_ssids[ssid] = [cur_lvl_grp]\n    else:\n        last_lvl_grp = cache[-1]\n        if (\n            ((last_lvl_grp == '0-4') &amp; (cur_lvl_grp == '0-4')) |\n            ((last_lvl_grp == '5-12') &amp; (cur_lvl_grp in ['0-4', '5-12'])) |\n            (last_lvl_grp == '13-22')\n        ): # if level_group repeats, change prediction value to 1\n            prediction_value = 1\n        else: # otherwise append current level group\n            seen_ssids[ssid].append(cur_lvl_grp)\n\n    sample_submission['correct'] = prediction_value\n    env.predict(sample_submission)\n</code></pre>",
          "votes": 0,
          "replies": [
            {
              "id": 2205421,
              "author_name": "Maruichi01",
              "author_url": "",
              "post_date": "2023-04-01T14:30:43.983000",
              "content": "<p>Your comment \"(1) there is no restart within a batch.\" is not true.  Code <code>min(test[\"level\"].diff().values)</code> become nan, so <code>min_level_diff &lt; 0</code> is always False.  I modified the code to fill nan and the LB become 0.219.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 2207230,
              "author_name": "Hoang Nguyen",
              "author_url": "",
              "post_date": "2023-04-03T08:31:56.123000",
              "content": "<p><a href=\"https://www.kaggle.com/maruichi01\" target=\"_blank\">@maruichi01</a> Interesting - thanks for flagging that! Did you have a chance to check my other assumptions:</p>\n<ol>\n<li>Each batch contains only one session_id</li>\n<li>Each batch contains only one level_group</li>\n<li>Each batch contains only one event_name=checkpoint</li>\n<li>There is no restart across batches</li>\n</ol>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2221679,
      "author_name": "superslon",
      "author_url": "",
      "post_date": "2023-04-14T13:05:27.520000",
      "content": "<p>My notes:</p>\n<ul>\n<li><strong>Each batch contains only one <code>event_name=checkpoint</code>.</strong> I ran this code and got <code>LB 0.216</code></li>\n</ul>\n<pre><code>leak = \n (test, sample_submission)  iter_test:\n  checkpoints = (test[]==).()\n   checkpoints != :\n      leak = \n   leak:\n      sample_submission[] = \n  :\n      sample_submission[] = \n  env.predict(sample_submission)\n</code></pre>\n<ul>\n<li><strong>Some people restarted the chapter</strong> (This code got <code>LB 0.463</code> ): </li>\n</ul>\n<pre><code>leak = \n (test, sample_submission)  iter_test:\n  level_diff = (test[].diff().unique()&lt;)\n   level_diff &gt; :\n      leak =    \n   leak:\n      sample_submission[] = \n  :\n      sample_submission[] = \n  env.predict(sample_submission)\n</code></pre>\n<p>if there were no restarts, then the result of <code>test['level'].diff().unique()</code> would be <code>[0, 1, nan]</code>.  <code>sum(test['level'].diff().unique()&lt;0)</code> means that for some sessions the difference between the current level and the level of the previous record is less than 0, which means restarting the chapter</p>\n<ul>\n<li><strong>Some users missed levels (e.g. firstly they were at level 1, and then they immediately went to level 3)</strong> Perhaps this is related to the previous point. (This code get <code>LB 0.43</code>)</li>\n</ul>\n<pre><code>leak = \n (test, sample_submission)  iter_test:\n  positive_level_diff = (test[].diff().unique()&gt;)\n   positive_level_diff &gt; :\n      leak = \n   leak:\n      sample_submission[] = \n  :\n      sample_submission[] = \n  env.predict(sample_submission)\n</code></pre>\n<p>if there were no level skips, then the result of <code>test['level'].diff().unique()</code> would be <code>[0, 1, nan]</code> 1 means moving to a next level. <code>sum(test['level'].diff().unique()&gt;1)&gt;0</code> means that for some sessions the difference between the current level and the level of the previous record is greater than 1, which means skipping the level</p>",
      "votes": 2,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2201300": "I made a few experimental submissions and noticed 3 points below about the **new** test API\n1. Each batch contains only one `session_id`\n2. Each batch contains only one `level_group`\n3. Each batch contains only one `event_name=checkpoint`\n\nRE point 1 and 2: these have been confirmed in several discussions ([link 1](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/397414), [link 2](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/396202#2191732)).\nRE point 3: This implies that the **new** hidden test set does NOT contain users who **complete** a chapter multiple times, which is a leak the train set has. However, there can still be a corner case where user plays a chapter more than once but completes only once (i.e. they restart in the middle of chapter), and I don't know how to check for this scenario. My questions are\n- Do you have a suggestion on how to check the chapter-restart case?\n- If you have looked into whether the test set contains repeated gameplay or data leak in general, what did you find?\n\nTIA!",
    "2201471": "Could this work?\n```\nfor (test, sample_submission) in iter_test:\n    min_level_diff = min(test[\"level\"].diff().values)\n    if min_level_diff < 0:\n        sample_submission.loc[:, \"correct\"] = 1\n    else:\n        sample_submission.loc[:, \"correct\"] = 0\n\n    env.predict(sample_submission[['session_id','correct']])\n```\nIf there is a restart, the `min_level_diff` should be less than 0, and we make all `correct = 1`. The LB value may not be 0.216.",
    "2221679": "My notes:\n-  **Each batch contains only one `event_name=checkpoint`.** I ran this code and got `LB 0.216`\n\n```python\nleak = False\nfor (test, sample_submission) in iter_test:\n  checkpoints = (test[\"event_name\"]=='checkpoint').sum()\n  if checkpoints !=1 :\n      leak = True\n  if leak:\n      sample_submission['correct'] = 1\n  else:\n      sample_submission['correct'] = 0\n  env.predict(sample_submission)\n```\n-  **Some people restarted the chapter** (This code got ` LB 0.463` ): \n  ```python\nleak = False\nfor (test, sample_submission) in iter_test:\n  level_diff = sum(test['level'].diff().unique()<0)\n  if level_diff >0 :\n      leak = True   \n  if leak:\n      sample_submission['correct'] = 1\n  else:\n      sample_submission['correct'] = 0\n  env.predict(sample_submission)\n```\n\nif there were no restarts, then the result of `test['level'].diff().unique()` would be `[0, 1, nan]`.  `sum(test['level'].diff().unique()<0)` means that for some sessions the difference between the current level and the level of the previous record is less than 0, which means restarting the chapter\n- **Some users missed levels (e.g. firstly they were at level 1, and then they immediately went to level 3)** Perhaps this is related to the previous point. (This code get `LB 0.43`)\n  ```python\nleak = False\nfor (test, sample_submission) in iter_test:\n  positive_level_diff = sum(test['level'].diff().unique()>1)\n  if positive_level_diff >0 :\n      leak = True\n  if leak:\n      sample_submission['correct'] = 1\n  else:\n      sample_submission['correct'] = 0\n  env.predict(sample_submission)\n```\nif there were no level skips, then the result of `test['level'].diff().unique()` would be `[0, 1, nan]` 1 means moving to a next level. `sum(test['level'].diff().unique()>1)>0` means that for some sessions the difference between the current level and the level of the previous record is greater than 1, which means skipping the level"
  }
}