{
  "id": 420025,
  "title": "\"session_level\" column in test/sample_submission.csv",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/420025",
  "author_name": "",
  "post_date": "2023-06-29T01:02:45.537488300Z",
  "votes": 2,
  "comment_count": 2,
  "views": 0,
  "content": "<p>TLDR:</p>\n<ul>\n<li>This is not directly related to our solution, we don't gain much advantage from this finding.</li>\n<li>Create \"session_level\" column in the train.csv: <a href=\"https://www.kaggle.com/code/kingychiu/session-level-encoding-for-train-dataset\" target=\"_blank\">https://www.kaggle.com/code/kingychiu/session-level-encoding-for-train-dataset</a></li>\n</ul>\n<p>In test/sample_submission.csv, there is a unseen column <code>session_level</code>, if you have tried to download the given <code>jo_wilder</code> Python api and run it on your machine, you will see it cannot be run because of missing this column.</p>\n<p>It is not hard to show, this columns is actually the encoding of 2 columns <code>session_id</code> and <code>level_group</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fba162ce133d1ef44c84244c7e1a6d118%2FScreenshot%202023-06-29%20at%208.51.49%20AM.png?generation=1687999968380750&amp;alt=media\" alt=\"For each unique session_level, show the unique session_id and level\"></p>\n<p>Another thing we can confirm is that the test loop api only yield 3 iterations for each session id, because of this encoding method. This is already well known idea in this discussion, but by understanding this encoding method, we can see that it is not possible for a next iter having lower level given the same session_id. </p>\n<p>In otherwords even in the training data we are seeing some \"restarted game\": Lv1 Lv2 Lv3 Lv1' Lv2' …, but in the test loop api, they will be presented to use (Lv1 Lv1') (Lv2 Lv2') Lv3</p>",
  "messages": [
    {
      "id": "2321877",
      "postDate": "06/29/2023 01:02:45",
      "content": "<p>TLDR:</p>\n<ul>\n<li>This is not directly related to our solution, we don't gain much advantage from this finding.</li>\n<li>Create \"session_level\" column in the train.csv: <a href=\"https://www.kaggle.com/code/kingychiu/session-level-encoding-for-train-dataset\" target=\"_blank\">https://www.kaggle.com/code/kingychiu/session-level-encoding-for-train-dataset</a></li>\n</ul>\n<p>In test/sample_submission.csv, there is a unseen column <code>session_level</code>, if you have tried to download the given <code>jo_wilder</code> Python api and run it on your machine, you will see it cannot be run because of missing this column.</p>\n<p>It is not hard to show, this columns is actually the encoding of 2 columns <code>session_id</code> and <code>level_group</code>:<br>\n<img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fba162ce133d1ef44c84244c7e1a6d118%2FScreenshot%202023-06-29%20at%208.51.49%20AM.png?generation=1687999968380750&amp;alt=media\" alt=\"For each unique session_level, show the unique session_id and level\"></p>\n<p>Another thing we can confirm is that the test loop api only yield 3 iterations for each session id, because of this encoding method. This is already well known idea in this discussion, but by understanding this encoding method, we can see that it is not possible for a next iter having lower level given the same session_id. </p>\n<p>In otherwords even in the training data we are seeing some \"restarted game\": Lv1 Lv2 Lv3 Lv1' Lv2' …, but in the test loop api, they will be presented to use (Lv1 Lv1') (Lv2 Lv2') Lv3</p>",
      "rawMarkdown": "TLDR:\n- This is not directly related to our solution, we don't gain much advantage from this finding.\n- Create \"session_level\" column in the train.csv: https://www.kaggle.com/code/kingychiu/session-level-encoding-for-train-dataset\n\nIn test/sample_submission.csv, there is a unseen column `session_level`, if you have tried to download the given `jo_wilder` Python api and run it on your machine, you will see it cannot be run because of missing this column.\n\nIt is not hard to show, this columns is actually the encoding of 2 columns `session_id` and `level_group`:\n![For each unique session_level, show the unique session_id and level](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fba162ce133d1ef44c84244c7e1a6d118%2FScreenshot%202023-06-29%20at%208.51.49%20AM.png?generation=1687999968380750&alt=media)\n\nAnother thing we can confirm is that the test loop api only yield 3 iterations for each session id, because of this encoding method. This is already well known idea in this discussion, but by understanding this encoding method, we can see that it is not possible for a next iter having lower level given the same session_id. \n\nIn otherwords even in the training data we are seeing some \"restarted game\": Lv1 Lv2 Lv3 Lv1' Lv2' ..., but in the test loop api, they will be presented to use (Lv1 Lv1') (Lv2 Lv2') Lv3",
      "votes": null
    },
    {
      "id": "2322062",
      "postDate": "06/29/2023 04:14:55",
      "content": "<p>Interesting. I didn't pay much attention to that <code>session_level</code>. This implies that if an user goes Lv 1 &gt; 2 &gt; 3 &gt;1 &gt; 2…, you can use (Lv1 Lv1) to gain more information about subsequently groups?</p>\n<p>You said your solution didn't benefit much from this information - does it mean you did not use this idea or that it did not improve your model?</p>",
      "rawMarkdown": "Interesting. I didn't pay much attention to that `session_level`. This implies that if an user goes Lv 1 > 2 > 3 >1 > 2..., you can use (Lv1 Lv1) to gain more information about subsequently groups?\n\nYou said your solution didn't benefit much from this information - does it mean you did not use this idea or that it did not improve your model?",
      "votes": null
    },
    {
      "id": "2322262",
      "postDate": "06/29/2023 07:14:58",
      "content": "<p>I think most people have this information implicitly already.  Most public notebooks first filter the training data by a question and train on each question. If you want to contact the features for Q1 for Q2 Q3 etc… without sorting after the concat, you will end up having <code>(Lv1 Lv1') (Lv2 Lv2') Lv3</code> in your training data as well. Therefore we might implicitly use this idea, but we are not abusing this fact as a leak. </p>\n<p>In fact, in some of our models we sort after concat different qs/lvs, which means we enforce both the train and test data has no reverse time: <code>(Lv1 Lv1') (Lv2 Lv2') Lv3</code> -&gt; <code>Lv1 Lv2 Lv3 Lv1 Lv2</code>… We will have a solution post later for details.</p>",
      "rawMarkdown": "I think most people have this information implicitly already.  Most public notebooks first filter the training data by a question and train on each question. If you want to contact the features for Q1 for Q2 Q3 etc... without sorting after the concat, you will end up having ` (Lv1 Lv1') (Lv2 Lv2') Lv3` in your training data as well. Therefore we might implicitly use this idea, but we are not abusing this fact as a leak. \n\nIn fact, in some of our models we sort after concat different qs/lvs, which means we enforce both the train and test data has no reverse time: `(Lv1 Lv1') (Lv2 Lv2') Lv3` -> `Lv1 Lv2 Lv3 Lv1 Lv2`... We will have a solution post later for details.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2322062,
      "author_name": "hoangnguyen719",
      "author_url": "",
      "post_date": "06/29/2023 04:14:55",
      "content": "<p>Interesting. I didn't pay much attention to that <code>session_level</code>. This implies that if an user goes Lv 1 &gt; 2 &gt; 3 &gt;1 &gt; 2…, you can use (Lv1 Lv1) to gain more information about subsequently groups?</p>\n<p>You said your solution didn't benefit much from this information - does it mean you did not use this idea or that it did not improve your model?</p>",
      "votes": null,
      "replies": [
        {
          "id": 2322262,
          "author_name": "kingychiu",
          "author_url": "",
          "post_date": "06/29/2023 07:14:58",
          "content": "<p>I think most people have this information implicitly already.  Most public notebooks first filter the training data by a question and train on each question. If you want to contact the features for Q1 for Q2 Q3 etc… without sorting after the concat, you will end up having <code>(Lv1 Lv1') (Lv2 Lv2') Lv3</code> in your training data as well. Therefore we might implicitly use this idea, but we are not abusing this fact as a leak. </p>\n<p>In fact, in some of our models we sort after concat different qs/lvs, which means we enforce both the train and test data has no reverse time: <code>(Lv1 Lv1') (Lv2 Lv2') Lv3</code> -&gt; <code>Lv1 Lv2 Lv3 Lv1 Lv2</code>… We will have a solution post later for details.</p>",
          "votes": null,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2321877": "TLDR:\n- This is not directly related to our solution, we don't gain much advantage from this finding.\n- Create \"session_level\" column in the train.csv: https://www.kaggle.com/code/kingychiu/session-level-encoding-for-train-dataset\n\nIn test/sample_submission.csv, there is a unseen column `session_level`, if you have tried to download the given `jo_wilder` Python api and run it on your machine, you will see it cannot be run because of missing this column.\n\nIt is not hard to show, this columns is actually the encoding of 2 columns `session_id` and `level_group`:\n![For each unique session_level, show the unique session_id and level](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F690886%2Fba162ce133d1ef44c84244c7e1a6d118%2FScreenshot%202023-06-29%20at%208.51.49%20AM.png?generation=1687999968380750&alt=media)\n\nAnother thing we can confirm is that the test loop api only yield 3 iterations for each session id, because of this encoding method. This is already well known idea in this discussion, but by understanding this encoding method, we can see that it is not possible for a next iter having lower level given the same session_id. \n\nIn otherwords even in the training data we are seeing some \"restarted game\": Lv1 Lv2 Lv3 Lv1' Lv2' ..., but in the test loop api, they will be presented to use (Lv1 Lv1') (Lv2 Lv2') Lv3",
    "2322062": "Interesting. I didn't pay much attention to that `session_level`. This implies that if an user goes Lv 1 > 2 > 3 >1 > 2..., you can use (Lv1 Lv1) to gain more information about subsequently groups?\n\nYou said your solution didn't benefit much from this information - does it mean you did not use this idea or that it did not improve your model?",
    "2322262": "I think most people have this information implicitly already.  Most public notebooks first filter the training data by a question and train on each question. If you want to contact the features for Q1 for Q2 Q3 etc... without sorting after the concat, you will end up having ` (Lv1 Lv1') (Lv2 Lv2') Lv3` in your training data as well. Therefore we might implicitly use this idea, but we are not abusing this fact as a leak. \n\nIn fact, in some of our models we sort after concat different qs/lvs, which means we enforce both the train and test data has no reverse time: `(Lv1 Lv1') (Lv2 Lv2') Lv3` -> `Lv1 Lv2 Lv3 Lv1 Lv2`... We will have a solution post later for details."
  },
  "source": "meta"
}