{
  "id": 400858,
  "title": "BERT-like Approach for Representation Learning on Educational Process Data",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/400858",
  "author_name": "",
  "post_date": "2023-04-10T16:02:21.067944100Z",
  "votes": 24,
  "comment_count": 4,
  "views": 0,
  "content": "<p>Hi there!</p>\n<p>I supposed that there is a good idea to apply RNN (LSTM) models to process the event sequences in educational game. In particular I suppose that BERT-like models possibly can do good encoding of events in sessions and after this we can use classifier to predict the correctness of answers in sessions. </p>\n<p><strong>Our findings</strong><br>\nIndeed we found a few publications of applying LSTM and transformers to educational process click-stream data. But one of them -- <a href=\"https://arxiv.org/pdf/2204.13607.pdf\" target=\"_blank\">Process-BERT: A Framework for Representation Learning\non Educational Process Data</a> -- is especially interesting.  <br>\nIt is even more interesting because authors coded up a framework and have a good GitHub repo with the source code <a href=\"https://github.com/alexscarlatos/clickstream-assessments/\" target=\"_blank\">https://github.com/alexscarlatos/clickstream-assessments/</a><br>\nCode is quite good but not well documented. We spend about a week to read publication, obtain original data, understand and run this code. The original task authors solved is quite similar to this competition -- they predict correctness of answers during the online-testing using click-streams' and events' data.</p>\n<p><strong>Our asks</strong><br>\nCurrently we are in process of tuning this solution for this competition. It requires changes in data preprocessing, data loading, modelling and training. But I think it is very promising try. Do somebody want to join our team to make this tuning faster? </p>\n<p>Also it is interesting to know which other approaches of LSTM/transformers use for educational process data have you seen in relative publications? Did you try them for this competition? If yes then which results did you obtain?   </p>",
  "messages": [
    {
      "id": "2217117",
      "postDate": "04/10/2023 16:02:21",
      "content": "<p>Hi there!</p>\n<p>I supposed that there is a good idea to apply RNN (LSTM) models to process the event sequences in educational game. In particular I suppose that BERT-like models possibly can do good encoding of events in sessions and after this we can use classifier to predict the correctness of answers in sessions. </p>\n<p><strong>Our findings</strong><br>\nIndeed we found a few publications of applying LSTM and transformers to educational process click-stream data. But one of them -- <a href=\"https://arxiv.org/pdf/2204.13607.pdf\" target=\"_blank\">Process-BERT: A Framework for Representation Learning\non Educational Process Data</a> -- is especially interesting.  <br>\nIt is even more interesting because authors coded up a framework and have a good GitHub repo with the source code <a href=\"https://github.com/alexscarlatos/clickstream-assessments/\" target=\"_blank\">https://github.com/alexscarlatos/clickstream-assessments/</a><br>\nCode is quite good but not well documented. We spend about a week to read publication, obtain original data, understand and run this code. The original task authors solved is quite similar to this competition -- they predict correctness of answers during the online-testing using click-streams' and events' data.</p>\n<p><strong>Our asks</strong><br>\nCurrently we are in process of tuning this solution for this competition. It requires changes in data preprocessing, data loading, modelling and training. But I think it is very promising try. Do somebody want to join our team to make this tuning faster? </p>\n<p>Also it is interesting to know which other approaches of LSTM/transformers use for educational process data have you seen in relative publications? Did you try them for this competition? If yes then which results did you obtain?   </p>",
      "rawMarkdown": "Hi there!\n\nI supposed that there is a good idea to apply RNN (LSTM) models to process the event sequences in educational game. In particular I suppose that BERT-like models possibly can do good encoding of events in sessions and after this we can use classifier to predict the correctness of answers in sessions. \n\n**Our findings**\nIndeed we found a few publications of applying LSTM and transformers to educational process click-stream data. But one of them -- [Process-BERT: A Framework for Representation Learning\non Educational Process Data](https://arxiv.org/pdf/2204.13607.pdf) -- is especially interesting.  \nIt is even more interesting because authors coded up a framework and have a good GitHub repo with the source code https://github.com/alexscarlatos/clickstream-assessments/\nCode is quite good but not well documented. We spend about a week to read publication, obtain original data, understand and run this code. The original task authors solved is quite similar to this competition -- they predict correctness of answers during the online-testing using click-streams' and events' data.\n\n**Our asks**\nCurrently we are in process of tuning this solution for this competition. It requires changes in data preprocessing, data loading, modelling and training. But I think it is very promising try. Do somebody want to join our team to make this tuning faster? \n\nAlso it is interesting to know which other approaches of LSTM/transformers use for educational process data have you seen in relative publications? Did you try them for this competition? If yes then which results did you obtain?",
      "votes": null
    },
    {
      "id": "2217958",
      "postDate": "04/11/2023 09:44:14",
      "content": "<p>Thank you for finding and sharing this.</p>",
      "rawMarkdown": "Thank you for finding and sharing this.",
      "votes": null
    },
    {
      "id": "2217985",
      "postDate": "04/11/2023 10:11:52",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ivanisaev\" target=\"_blank\">@ivanisaev</a> , Its sounds amazing!<br>\nI could be implemented in GPT and other language models for verifying the answer.</p>",
      "rawMarkdown": "Hey @ivanisaev , Its sounds amazing!\nI could be implemented in GPT and other language models for verifying the answer.",
      "votes": null
    },
    {
      "id": "2218794",
      "postDate": "04/12/2023 03:26:19",
      "content": "<p>Thank you for sharing!<br>\nI am currently experimenting with the Transformer model. Currently the CV is 0.6917.</p>",
      "rawMarkdown": "Thank you for sharing!\nI am currently experimenting with the Transformer model. Currently the CV is 0.6917.",
      "votes": null
    },
    {
      "id": "2248329",
      "postDate": "05/06/2023 18:22:06",
      "content": "<p>📌 I promised to share my code of this experiment and it is <a href=\"https://www.kaggle.com/code/ivanisaev/jo-wilder-bert-like-lstm-model/notebook\" target=\"_blank\">in this notebook</a>.</p>",
      "rawMarkdown": "📌 I promised to share my code of this experiment and it is [in this notebook](https://www.kaggle.com/code/ivanisaev/jo-wilder-bert-like-lstm-model/notebook).",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 2217958,
      "author_name": "shambi",
      "author_url": "",
      "post_date": "04/11/2023 09:44:14",
      "content": "<p>Thank you for finding and sharing this.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2217985,
      "author_name": "ppb00x",
      "author_url": "",
      "post_date": "04/11/2023 10:11:52",
      "content": "<p>Hey <a href=\"https://www.kaggle.com/ivanisaev\" target=\"_blank\">@ivanisaev</a> , Its sounds amazing!<br>\nI could be implemented in GPT and other language models for verifying the answer.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2218794,
      "author_name": "ryotak12",
      "author_url": "",
      "post_date": "04/12/2023 03:26:19",
      "content": "<p>Thank you for sharing!<br>\nI am currently experimenting with the Transformer model. Currently the CV is 0.6917.</p>",
      "votes": null,
      "replies": []
    },
    {
      "id": 2248329,
      "author_name": "ivanisaev",
      "author_url": "",
      "post_date": "05/06/2023 18:22:06",
      "content": "<p>📌 I promised to share my code of this experiment and it is <a href=\"https://www.kaggle.com/code/ivanisaev/jo-wilder-bert-like-lstm-model/notebook\" target=\"_blank\">in this notebook</a>.</p>",
      "votes": null,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2217117": "Hi there!\n\nI supposed that there is a good idea to apply RNN (LSTM) models to process the event sequences in educational game. In particular I suppose that BERT-like models possibly can do good encoding of events in sessions and after this we can use classifier to predict the correctness of answers in sessions. \n\n**Our findings**\nIndeed we found a few publications of applying LSTM and transformers to educational process click-stream data. But one of them -- [Process-BERT: A Framework for Representation Learning\non Educational Process Data](https://arxiv.org/pdf/2204.13607.pdf) -- is especially interesting.  \nIt is even more interesting because authors coded up a framework and have a good GitHub repo with the source code https://github.com/alexscarlatos/clickstream-assessments/\nCode is quite good but not well documented. We spend about a week to read publication, obtain original data, understand and run this code. The original task authors solved is quite similar to this competition -- they predict correctness of answers during the online-testing using click-streams' and events' data.\n\n**Our asks**\nCurrently we are in process of tuning this solution for this competition. It requires changes in data preprocessing, data loading, modelling and training. But I think it is very promising try. Do somebody want to join our team to make this tuning faster? \n\nAlso it is interesting to know which other approaches of LSTM/transformers use for educational process data have you seen in relative publications? Did you try them for this competition? If yes then which results did you obtain?",
    "2217958": "Thank you for finding and sharing this.",
    "2217985": "Hey @ivanisaev , Its sounds amazing!\nI could be implemented in GPT and other language models for verifying the answer.",
    "2218794": "Thank you for sharing!\nI am currently experimenting with the Transformer model. Currently the CV is 0.6917.",
    "2248329": "📌 I promised to share my code of this experiment and it is [in this notebook](https://www.kaggle.com/code/ivanisaev/jo-wilder-bert-like-lstm-model/notebook)."
  },
  "source": "meta"
}