{
  "id": 416963,
  "title": "Code to partially restore 1D Conv networks",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/416963",
  "author_name": "Daniel Phalen",
  "post_date": "2023-06-13T15:42:29.352000",
  "votes": 15,
  "comment_count": 10,
  "views": 0,
  "content": "<p>If you have some models similar to <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">[LB 0.694] Event-Aware TConv with Only 4 Features</a>, putting the following code in your loop will help to recover the original performance:</p>\n<pre><code> ():\n     item1[] == item2[]:\n         (item1[] - item2[]) &gt; :\n             item1[] &lt; item2[]:\n                 -\n             item1[] &gt; item2[]:\n                 \n            :\n                 \n        :\n             item1[] &lt; item2[]:\n                 -\n             item1[] &gt; item2[]:\n                 \n            :\n                 \n    :\n         item1[] - item2[]\n\n functools\n ():\n    v = [{: idx, :row[], : row[], : row[]}  idx, row  df.iterrows()]\n    v.sort(key=functools.cmp_to_key(compare_func))\n    mdf = pd.DataFrame(v)\n     (mdf) == (df)\n     df.loc[mdf.idx]\n</code></pre>\n<p>Then in the loop:</p>\n<pre><code> (test, sample_submission)  iter_test:\n    # test = test.sort.reset\n    test = sort\n</code></pre>\n<p>This is better since there are some sessions where the index is reset.  An example from training would be session <em>20110609231859892</em> </p>",
  "messages": [
    {
      "id": 2301063,
      "postDate": "2023-06-13T15:42:29.353Z",
      "content": "<p>If you have some models similar to <a href=\"https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features\" target=\"_blank\">[LB 0.694] Event-Aware TConv with Only 4 Features</a>, putting the following code in your loop will help to recover the original performance:</p>\n<pre><code> ():\n     item1[] == item2[]:\n         (item1[] - item2[]) &gt; :\n             item1[] &lt; item2[]:\n                 -\n             item1[] &gt; item2[]:\n                 \n            :\n                 \n        :\n             item1[] &lt; item2[]:\n                 -\n             item1[] &gt; item2[]:\n                 \n            :\n                 \n    :\n         item1[] - item2[]\n\n functools\n ():\n    v = [{: idx, :row[], : row[], : row[]}  idx, row  df.iterrows()]\n    v.sort(key=functools.cmp_to_key(compare_func))\n    mdf = pd.DataFrame(v)\n     (mdf) == (df)\n     df.loc[mdf.idx]\n</code></pre>\n<p>Then in the loop:</p>\n<pre><code> (test, sample_submission)  iter_test:\n    # test = test.sort.reset\n    test = sort\n</code></pre>\n<p>This is better since there are some sessions where the index is reset.  An example from training would be session <em>20110609231859892</em> </p>",
      "rawMarkdown": "If you have some models similar to [[LB 0.694] Event-Aware TConv with Only 4 Features](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features), putting the following code in your loop will help to recover the original performance:\n\n```\ndef compare_func(item1, item2):\n    if item1['level'] == item2['level']:\n        if abs(item1['index'] - item2['index']) > 50:\n            if item1['elapsed_time'] < item2['elapsed_time']:\n                return -1\n            elif item1['elapsed_time'] > item2['elapsed_time']:\n                return 1\n            else:\n                return 0\n        else:\n            if item1['index'] < item2['index']:\n                return -1\n            elif item1['index'] > item2['index']:\n                return 1\n            else:\n                return 0\n    else:\n        return item1['level'] - item2['level']\n\nimport functools\ndef sort_frame(df):\n    v = [{'idx': idx, 'index':row['index'], 'elapsed_time': row['elapsed_time'], 'level': row['level']} for idx, row in df.iterrows()]\n    v.sort(key=functools.cmp_to_key(compare_func))\n    mdf = pd.DataFrame(v)\n    assert len(mdf) == len(df)\n    return df.loc[mdf.idx]\n```\n\nThen in the loop:\n\n```\nfor (test, sample_submission) in iter_test:\n    # test = test.sort_values('index').reset_index(drop=True)\n    test = sort_frame(test)\n```\n\n\nThis is better since there are some sessions where the index is reset.  An example from training would be session *20110609231859892* ",
      "votes": 15
    },
    {
      "id": 2317705,
      "postDate": "2023-06-25T22:34:04.920Z",
      "content": "<p>Nice! Does it work notably better (slightly better?) than simply sorting on elapsed time, for you? Sort on elapsed time doesn't have the same issue, does it?</p>\n<p>I plan to try it on some of my models for both training and submit, see if it works better :)</p>",
      "rawMarkdown": "Nice! Does it work notably better (slightly better?) than simply sorting on elapsed time, for you? Sort on elapsed time doesn't have the same issue, does it?\n\nI plan to try it on some of my models for both training and submit, see if it works better :)",
      "replies": [
        {
          "id": 2317760,
          "postDate": "2023-06-26T01:19:58.297Z",
          "content": "<p>If you look in the training data, elapsed time goes slightly backward for some events that are after each other, no idea why.  When I tried models sorting on elapsed time they did not perform as well, but yours may do better.</p>",
          "rawMarkdown": "If you look in the training data, elapsed time goes slightly backward for some events that are after each other, no idea why.  When I tried models sorting on elapsed time they did not perform as well, but yours may do better.",
          "replies": [
            {
              "id": 2317796,
              "postDate": "2023-06-26T02:03:50.593Z",
              "content": "<p>Yeah, I also thought that by index is a bit more reliable (when not reset). And had seen the reset issue but never tried to fix it. So this code looks helpful :)</p>",
              "rawMarkdown": "Yeah, I also thought that by index is a bit more reliable (when not reset). And had seen the reset issue but never tried to fix it. So this code looks helpful :)"
            }
          ]
        }
      ]
    },
    {
      "id": 2307439,
      "postDate": "2023-06-18T07:36:42.383Z",
      "content": "<p><a href=\"https://www.kaggle.com/danielphalen\" target=\"_blank\">@danielphalen</a>, thank you for the code.  These models are very interesting but I can't make them work in the submission (the training is fine). The Lb score is awful (LB 0.56 while CV is 0.685). I tried your code but it is still not good 🤥. As we don't see what happens during the submission, it is hard to debug.</p>",
      "rawMarkdown": "@danielphalen, thank you for the code.  These models are very interesting but I can't make them work in the submission (the training is fine). The Lb score is awful (LB 0.56 while CV is 0.685). I tried your code but it is still not good 🤥. As we don't see what happens during the submission, it is hard to debug.",
      "replies": [
        {
          "id": 2307468,
          "postDate": "2023-06-18T08:07:01.610Z",
          "content": "<p>Are you sorting the questions correctly? As discussed <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">here</a>, the order of the questions in the sample_submission seems to be shuffled only in the hidden test set.</p>",
          "rawMarkdown": "Are you sorting the questions correctly? As discussed [here](https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512), the order of the questions in the sample_submission seems to be shuffled only in the hidden test set.",
          "replies": [
            {
              "id": 2307590,
              "postDate": "2023-06-18T09:47:27.420Z",
              "content": "<p>I think, that was exactly my problem! Trying to submit now. Thanks <a href=\"https://www.kaggle.com/jinmiyashita\" target=\"_blank\">@jinmiyashita</a> !</p>",
              "rawMarkdown": "I think, that was exactly my problem! Trying to submit now. Thanks @jinmiyashita !"
            }
          ]
        }
      ]
    },
    {
      "id": 2302354,
      "postDate": "2023-06-14T13:45:30.693Z",
      "content": "<p>Could you explain the maining of doing this (rather than directly use <code>test.sort_values('index').reset_index(drop=True)</code>)? I try this Conv network code and found LB score becomes 0.68+</p>",
      "rawMarkdown": "Could you explain the maining of doing this (rather than directly use `test.sort_values('index').reset_index(drop=True)`)? I try this Conv network code and found LB score becomes 0.68+",
      "replies": [
        {
          "id": 2302386,
          "postDate": "2023-06-14T14:04:48.497Z",
          "content": "<p>It will of course depend upon the training you use.  If you look in the training set, there are some sessions where the index resets within a <code>level_group</code>.  In the below image, the first column is the order in the <code>train.csv</code> file, the second the <code>session_id</code>, the third is the <code>index</code> column, and the fourth is <code>elapsed_time</code>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552115%2F0c0400e6aacfcdb97174e958fb1db829%2FIndexReset.png?generation=1686750958516616&amp;alt=media\" alt=\"Index Reset\"></p>\n<p>You can see that the index column resets from <code>901</code> to <code>0</code> for unknown reasons, and the elapsed time slightly goes backward, which is common.  If you did the <code>test.sort_values('index').reset_index(drop=True)</code>, these events, which are near the end of the <code>level_group</code>, would be put at the beginning of your data sample.</p>\n<p>For one of my own models, with the new API it scores terribly with no sort, 0.689 with the <code>index</code> sort, and 0.696 with the above code.  Of course YMMV.</p>",
          "rawMarkdown": "It will of course depend upon the training you use.  If you look in the training set, there are some sessions where the index resets within a `level_group`.  In the below image, the first column is the order in the `train.csv` file, the second the `session_id`, the third is the `index` column, and the fourth is `elapsed_time`:\n\n![Index Reset](https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552115%2F0c0400e6aacfcdb97174e958fb1db829%2FIndexReset.png?generation=1686750958516616&alt=media)\n\nYou can see that the index column resets from `901` to `0` for unknown reasons, and the elapsed time slightly goes backward, which is common.  If you did the `test.sort_values('index').reset_index(drop=True)`, these events, which are near the end of the `level_group`, would be put at the beginning of your data sample.\n\nFor one of my own models, with the new API it scores terribly with no sort, 0.689 with the `index` sort, and 0.696 with the above code.  Of course YMMV.",
          "votes": 1,
          "replies": [
            {
              "id": 2302463,
              "postDate": "2023-06-14T15:14:11.340Z",
              "content": "<p>So sad, my best conv model LB693 just become 694🤒</p>",
              "rawMarkdown": "So sad, my best conv model LB693 just become 694🤒"
            }
          ]
        }
      ]
    },
    {
      "id": 2301164,
      "postDate": "2023-06-13T16:56:34.973Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2317705,
      "author_name": "Robert Hatch",
      "author_url": "",
      "post_date": "2023-06-25T22:34:04.920000",
      "content": "<p>Nice! Does it work notably better (slightly better?) than simply sorting on elapsed time, for you? Sort on elapsed time doesn't have the same issue, does it?</p>\n<p>I plan to try it on some of my models for both training and submit, see if it works better :)</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2317760,
          "author_name": "Daniel Phalen",
          "author_url": "",
          "post_date": "2023-06-26T01:19:58.297000",
          "content": "<p>If you look in the training data, elapsed time goes slightly backward for some events that are after each other, no idea why.  When I tried models sorting on elapsed time they did not perform as well, but yours may do better.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2317796,
              "author_name": "Robert Hatch",
              "author_url": "",
              "post_date": "2023-06-26T02:03:50.593000",
              "content": "<p>Yeah, I also thought that by index is a bit more reliable (when not reset). And had seen the reset issue but never tried to fix it. So this code looks helpful :)</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2307439,
      "author_name": "Elias",
      "author_url": "",
      "post_date": "2023-06-18T07:36:42.383000",
      "content": "<p><a href=\"https://www.kaggle.com/danielphalen\" target=\"_blank\">@danielphalen</a>, thank you for the code.  These models are very interesting but I can't make them work in the submission (the training is fine). The Lb score is awful (LB 0.56 while CV is 0.685). I tried your code but it is still not good 🤥. As we don't see what happens during the submission, it is hard to debug.</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2307468,
          "author_name": "tonic",
          "author_url": "",
          "post_date": "2023-06-18T08:07:01.610000",
          "content": "<p>Are you sorting the questions correctly? As discussed <a href=\"https://www.kaggle.com/competitions/predict-student-performance-from-game-play/discussion/412512\" target=\"_blank\">here</a>, the order of the questions in the sample_submission seems to be shuffled only in the hidden test set.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2307590,
              "author_name": "Elias",
              "author_url": "",
              "post_date": "2023-06-18T09:47:27.420000",
              "content": "<p>I think, that was exactly my problem! Trying to submit now. Thanks <a href=\"https://www.kaggle.com/jinmiyashita\" target=\"_blank\">@jinmiyashita</a> !</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2302354,
      "author_name": "嘴爷",
      "author_url": "",
      "post_date": "2023-06-14T13:45:30.693000",
      "content": "<p>Could you explain the maining of doing this (rather than directly use <code>test.sort_values('index').reset_index(drop=True)</code>)? I try this Conv network code and found LB score becomes 0.68+</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2302386,
          "author_name": "Daniel Phalen",
          "author_url": "",
          "post_date": "2023-06-14T14:04:48.497000",
          "content": "<p>It will of course depend upon the training you use.  If you look in the training set, there are some sessions where the index resets within a <code>level_group</code>.  In the below image, the first column is the order in the <code>train.csv</code> file, the second the <code>session_id</code>, the third is the <code>index</code> column, and the fourth is <code>elapsed_time</code>:</p>\n<p><img src=\"https://www.googleapis.com/download/storage/v1/b/kaggle-forum-message-attachments/o/inbox%2F5552115%2F0c0400e6aacfcdb97174e958fb1db829%2FIndexReset.png?generation=1686750958516616&amp;alt=media\" alt=\"Index Reset\"></p>\n<p>You can see that the index column resets from <code>901</code> to <code>0</code> for unknown reasons, and the elapsed time slightly goes backward, which is common.  If you did the <code>test.sort_values('index').reset_index(drop=True)</code>, these events, which are near the end of the <code>level_group</code>, would be put at the beginning of your data sample.</p>\n<p>For one of my own models, with the new API it scores terribly with no sort, 0.689 with the <code>index</code> sort, and 0.696 with the above code.  Of course YMMV.</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2302463,
              "author_name": "嘴爷",
              "author_url": "",
              "post_date": "2023-06-14T15:14:11.340000",
              "content": "<p>So sad, my best conv model LB693 just become 694🤒</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2301164,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-06-13T16:56:34.973000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2301063": "If you have some models similar to [[LB 0.694] Event-Aware TConv with Only 4 Features](https://www.kaggle.com/code/abaojiang/lb-0-694-event-aware-tconv-with-only-4-features), putting the following code in your loop will help to recover the original performance:\n\n```\ndef compare_func(item1, item2):\n    if item1['level'] == item2['level']:\n        if abs(item1['index'] - item2['index']) > 50:\n            if item1['elapsed_time'] < item2['elapsed_time']:\n                return -1\n            elif item1['elapsed_time'] > item2['elapsed_time']:\n                return 1\n            else:\n                return 0\n        else:\n            if item1['index'] < item2['index']:\n                return -1\n            elif item1['index'] > item2['index']:\n                return 1\n            else:\n                return 0\n    else:\n        return item1['level'] - item2['level']\n\nimport functools\ndef sort_frame(df):\n    v = [{'idx': idx, 'index':row['index'], 'elapsed_time': row['elapsed_time'], 'level': row['level']} for idx, row in df.iterrows()]\n    v.sort(key=functools.cmp_to_key(compare_func))\n    mdf = pd.DataFrame(v)\n    assert len(mdf) == len(df)\n    return df.loc[mdf.idx]\n```\n\nThen in the loop:\n\n```\nfor (test, sample_submission) in iter_test:\n    # test = test.sort_values('index').reset_index(drop=True)\n    test = sort_frame(test)\n```\n\n\nThis is better since there are some sessions where the index is reset.  An example from training would be session *20110609231859892* ",
    "2317705": "Nice! Does it work notably better (slightly better?) than simply sorting on elapsed time, for you? Sort on elapsed time doesn't have the same issue, does it?\n\nI plan to try it on some of my models for both training and submit, see if it works better :)",
    "2307439": "@danielphalen, thank you for the code.  These models are very interesting but I can't make them work in the submission (the training is fine). The Lb score is awful (LB 0.56 while CV is 0.685). I tried your code but it is still not good 🤥. As we don't see what happens during the submission, it is hard to debug.",
    "2302354": "Could you explain the maining of doing this (rather than directly use `test.sort_values('index').reset_index(drop=True)`)? I try this Conv network code and found LB score becomes 0.68+",
    "2301164": ""
  }
}