{
  "id": 405173,
  "title": "Data Augmentation by Connecting Session",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/405173",
  "author_name": "",
  "post_date": "2023-04-26T12:20:47.905823900Z",
  "votes": 17,
  "comment_count": 3,
  "views": 0,
  "content": "<p>I recently experimented with Data Augmentation. I created pairs of sessions with the same label values for each level_group, then cut each session at random level points and combined them. With this method, I doubled the training data and trained my NN model, but the score wasn't improved.</p>\n<p>In creating pairs for Data Augmentation, I generated several features such as elapsed_time_diff aggregation feature and clustering for each session,  and creating pairs within the same cluster. However, there is still a difference in distribution between the augmented data and the original training data, so I am using Adversarial Validation to select samples close to the training data. But, It did not work.</p>\n<p>Has anyone had success with a similar method for Data Augmentation?</p>",
  "messages": [
    {
      "id": "2235911",
      "postDate": "04/26/2023 12:20:47",
      "content": "<p>I recently experimented with Data Augmentation. I created pairs of sessions with the same label values for each level_group, then cut each session at random level points and combined them. With this method, I doubled the training data and trained my NN model, but the score wasn't improved.</p>\n<p>In creating pairs for Data Augmentation, I generated several features such as elapsed_time_diff aggregation feature and clustering for each session,  and creating pairs within the same cluster. However, there is still a difference in distribution between the augmented data and the original training data, so I am using Adversarial Validation to select samples close to the training data. But, It did not work.</p>\n<p>Has anyone had success with a similar method for Data Augmentation?</p>",
      "rawMarkdown": "I recently experimented with Data Augmentation. I created pairs of sessions with the same label values for each level_group, then cut each session at random level points and combined them. With this method, I doubled the training data and trained my NN model, but the score wasn't improved.\n\nIn creating pairs for Data Augmentation, I generated several features such as elapsed_time_diff aggregation feature and clustering for each session,  and creating pairs within the same cluster. However, there is still a difference in distribution between the augmented data and the original training data, so I am using Adversarial Validation to select samples close to the training data. But, It did not work.\n\nHas anyone had success with a similar method for Data Augmentation?",
      "votes": null
    },
    {
      "id": "2238208",
      "postDate": "04/28/2023 09:59:54",
      "content": "<p>For me, I simply concatenate the data from sessions. E.g., to predict questions 1 to 3, I use level 0-4, to predict questions 4-13, I use level 0-4 and 5-12 (not only 5-12), and to predict question 14-18, I use all levels. Of course, model with the last level should take longer sequences. But, as like you, it doesn't seem to help in my case (CV: 0.6937 -&gt; 0.6938 for 1 fold with a transformer, I think it's just noise)</p>",
      "rawMarkdown": "For me, I simply concatenate the data from sessions. E.g., to predict questions 1 to 3, I use level 0-4, to predict questions 4-13, I use level 0-4 and 5-12 (not only 5-12), and to predict question 14-18, I use all levels. Of course, model with the last level should take longer sequences. But, as like you, it doesn't seem to help in my case (CV: 0.6937 -> 0.6938 for 1 fold with a transformer, I think it's just noise)",
      "votes": null
    },
    {
      "id": "2258237",
      "postDate": "05/14/2023 03:18:10",
      "content": "<p>This is a very interesting topic. I think the big problem is that the number of sessions is only about 20,000, especially when training with NNs.</p>",
      "rawMarkdown": "This is a very interesting topic. I think the big problem is that the number of sessions is only about 20,000, especially when training with NNs.",
      "votes": null
    },
    {
      "id": "2260665",
      "postDate": "05/15/2023 19:35:45",
      "content": "<p>I tried a somewhat similar approach doing linear interpolation between sessions with similar answers and consistently got a worse score.</p>",
      "rawMarkdown": "I tried a somewhat similar approach doing linear interpolation between sessions with similar answers and consistently got a worse score.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2238208,
      "author_name": "shinomoriaoshi",
      "author_url": "",
      "post_date": "04/28/2023 09:59:54",
      "content": "<p>For me, I simply concatenate the data from sessions. E.g., to predict questions 1 to 3, I use level 0-4, to predict questions 4-13, I use level 0-4 and 5-12 (not only 5-12), and to predict question 14-18, I use all levels. Of course, model with the last level should take longer sequences. But, as like you, it doesn't seem to help in my case (CV: 0.6937 -&gt; 0.6938 for 1 fold with a transformer, I think it's just noise)</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2258237,
      "author_name": "jinmiyashita",
      "author_url": "",
      "post_date": "05/14/2023 03:18:10",
      "content": "<p>This is a very interesting topic. I think the big problem is that the number of sessions is only about 20,000, especially when training with NNs.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2260665,
      "author_name": "sakvaua",
      "author_url": "",
      "post_date": "05/15/2023 19:35:45",
      "content": "<p>I tried a somewhat similar approach doing linear interpolation between sessions with similar answers and consistently got a worse score.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2235911": "I recently experimented with Data Augmentation. I created pairs of sessions with the same label values for each level_group, then cut each session at random level points and combined them. With this method, I doubled the training data and trained my NN model, but the score wasn't improved.\n\nIn creating pairs for Data Augmentation, I generated several features such as elapsed_time_diff aggregation feature and clustering for each session,  and creating pairs within the same cluster. However, there is still a difference in distribution between the augmented data and the original training data, so I am using Adversarial Validation to select samples close to the training data. But, It did not work.\n\nHas anyone had success with a similar method for Data Augmentation?",
    "2238208": "For me, I simply concatenate the data from sessions. E.g., to predict questions 1 to 3, I use level 0-4, to predict questions 4-13, I use level 0-4 and 5-12 (not only 5-12), and to predict question 14-18, I use all levels. Of course, model with the last level should take longer sequences. But, as like you, it doesn't seem to help in my case (CV: 0.6937 -> 0.6938 for 1 fold with a transformer, I think it's just noise)",
    "2258237": "This is a very interesting topic. I think the big problem is that the number of sessions is only about 20,000, especially when training with NNs.",
    "2260665": "I tried a somewhat similar approach doing linear interpolation between sessions with similar answers and consistently got a worse score."
  },
  "source": "meta"
}