{
  "id": 390637,
  "title": "Does Transformer-based NN Work Well for this Competition?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/390637",
  "author_name": "william.wu",
  "post_date": "2023-02-26T14:12:31.866000",
  "votes": 7,
  "comment_count": 13,
  "views": 0,
  "content": "<p>This competition is very similar to the <a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/193250\" target=\"_blank\">Riid Competition</a>. Both of these 2 competitions ask you to predict whether the student is able to answer future questions correctly based on his past behaviors. The main difference is that Riid provides students' past correctness but this competition doesn't. However, we can still use Transformer-based NNs for this competition.</p>",
  "messages": [
    {
      "id": 2160246,
      "postDate": "2023-02-26T14:12:31.867Z",
      "content": "<p>This competition is very similar to the <a href=\"https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/193250\" target=\"_blank\">Riid Competition</a>. Both of these 2 competitions ask you to predict whether the student is able to answer future questions correctly based on his past behaviors. The main difference is that Riid provides students' past correctness but this competition doesn't. However, we can still use Transformer-based NNs for this competition.</p>",
      "rawMarkdown": "This competition is very similar to the [Riid Competition](https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/193250). Both of these 2 competitions ask you to predict whether the student is able to answer future questions correctly based on his past behaviors. The main difference is that Riid provides students' past correctness but this competition doesn't. However, we can still use Transformer-based NNs for this competition.",
      "votes": 7
    },
    {
      "id": 2227673,
      "postDate": "2023-04-19T23:47:35.473Z",
      "content": "<p>I experimented Transformer-based NN for several days. I reach 0.694 on LB with 1fold.</p>",
      "rawMarkdown": "I experimented Transformer-based NN for several days. I reach 0.694 on LB with 1fold.",
      "votes": 2,
      "replies": [
        {
          "id": 2234608,
          "postDate": "2023-04-25T10:57:02.943Z",
          "content": "<p>I am also having some Transformer-based models with CV of 0.6938, I haven't submitted them yet, but I expect LB is also around your LB with the Transformer. Yet, I am really struggling in improving it. Not sure it could reach 0.698 CV with the NN approach or not..</p>",
          "rawMarkdown": "I am also having some Transformer-based models with CV of 0.6938, I haven't submitted them yet, but I expect LB is also around your LB with the Transformer. Yet, I am really struggling in improving it. Not sure it could reach 0.698 CV with the NN approach or not..",
          "votes": 1,
          "replies": [
            {
              "id": 2234759,
              "postDate": "2023-04-25T13:23:32.493Z",
              "content": "<p>Small changes such as altering the architecture do not seem to contribute to improvement. I think it is necessary to input features that cannot be captured by the Transformer. For example, the features that are typically used in GBDT. But, I haven't tried it yet</p>",
              "rawMarkdown": "Small changes such as altering the architecture do not seem to contribute to improvement. I think it is necessary to input features that cannot be captured by the Transformer. For example, the features that are typically used in GBDT. But, I haven't tried it yet",
              "votes": 1
            },
            {
              "id": 2234770,
              "postDate": "2023-04-25T13:27:48.383Z",
              "content": "<p>Agree! I am now trying something new. I also tried adding the aggregated features, but it makes the NN converge really fast, and it didn't improve the model. I concatenate these aggregated features to the session embeddings, btw. Maybe there are some things I need to tweak.</p>",
              "rawMarkdown": "Agree! I am now trying something new. I also tried adding the aggregated features, but it makes the NN converge really fast, and it didn't improve the model. I concatenate these aggregated features to the session embeddings, btw. Maybe there are some things I need to tweak.",
              "votes": 1
            },
            {
              "id": 2246039,
              "postDate": "2023-05-04T19:18:02.973Z",
              "content": "<p><a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> may I ask how much the CV of your 0.694 LB model is if you don't mind sharing it? I have a model with 0.694 CV (1 fold) but score only 0.690 LB with the transformer. However, with a GBDT model, 1-fold CV is 0.693, and 1-fold LB is 0.695. I spent around a week searching for the bug but didn't see any. I am wondering if the Transformer model could be overfitting or not.</p>",
              "rawMarkdown": "@ryotak12 may I ask how much the CV of your 0.694 LB model is if you don't mind sharing it? I have a model with 0.694 CV (1 fold) but score only 0.690 LB with the transformer. However, with a GBDT model, 1-fold CV is 0.693, and 1-fold LB is 0.695. I spent around a week searching for the bug but didn't see any. I am wondering if the Transformer model could be overfitting or not."
            },
            {
              "id": 2246145,
              "postDate": "2023-05-04T23:20:34.917Z",
              "content": "<p>When 0.694 LB, CV is 0.6947. In my case, CV and LB are correlated!</p>",
              "rawMarkdown": "When 0.694 LB, CV is 0.6947. In my case, CV and LB are correlated!",
              "votes": 1
            },
            {
              "id": 2246315,
              "postDate": "2023-05-05T04:25:42.690Z",
              "content": "<p>Thank you very much for your sharing!</p>",
              "rawMarkdown": "Thank you very much for your sharing!"
            }
          ]
        }
      ]
    },
    {
      "id": 2160929,
      "postDate": "2023-02-27T05:16:12.123Z",
      "content": "<p>The suitability of transformer-based neural networks for the task of predicting student performance from game play largely depends on the specific requirements of the competition and the nature of the data being used.</p>\n<p>Transformer-based neural networks, such as the BERT (Bidirectional Encoder Representations from Transformers) model, have shown promising results in natural language processing (NLP) tasks due to their ability to capture long-term dependencies and handle sequential data effectively. However, predicting student performance from game play data may not involve the same type of sequential data that is common in NLP tasks.</p>\n<p>That being said, transformer-based models have also been successful in tasks involving non-sequential data, such as image and audio processing. Therefore, it's possible that they could be effective for predicting student performance from game play data, especially if the data contains complex relationships or patterns that require a deep understanding of the context.</p>\n<p>Ultimately, the effectiveness of transformer-based neural networks for this competition will depend on factors such as the size and complexity of the dataset, the specific features being used, and the performance metrics being evaluated. It's worth exploring the use of transformer-based models along with other machine learning algorithms to determine the most effective approach for the task at hand.</p>",
      "rawMarkdown": "The suitability of transformer-based neural networks for the task of predicting student performance from game play largely depends on the specific requirements of the competition and the nature of the data being used.\n\nTransformer-based neural networks, such as the BERT (Bidirectional Encoder Representations from Transformers) model, have shown promising results in natural language processing (NLP) tasks due to their ability to capture long-term dependencies and handle sequential data effectively. However, predicting student performance from game play data may not involve the same type of sequential data that is common in NLP tasks.\n\nThat being said, transformer-based models have also been successful in tasks involving non-sequential data, such as image and audio processing. Therefore, it's possible that they could be effective for predicting student performance from game play data, especially if the data contains complex relationships or patterns that require a deep understanding of the context.\n\nUltimately, the effectiveness of transformer-based neural networks for this competition will depend on factors such as the size and complexity of the dataset, the specific features being used, and the performance metrics being evaluated. It's worth exploring the use of transformer-based models along with other machine learning algorithms to determine the most effective approach for the task at hand.",
      "votes": -17,
      "replies": [
        {
          "id": 2161099,
          "postDate": "2023-02-27T07:51:20.250Z",
          "content": "<p>was it generated with chatgpt? 🤔</p>",
          "rawMarkdown": "was it generated with chatgpt? 🤔",
          "votes": 6,
          "replies": [
            {
              "id": 2161101,
              "postDate": "2023-02-27T07:53:04.820Z",
              "content": "<p>hhha, I have the same feeling</p>",
              "rawMarkdown": "hhha, I have the same feeling"
            },
            {
              "id": 2161126,
              "postDate": "2023-02-27T08:26:14.060Z",
              "content": "<p>So let's just ask chatgpt:</p>\n<blockquote>\n  <p>what is the probability the following setences created by you: paste his answer above</p>\n</blockquote>\n<p>This is the my chatgpt response:</p>\n<blockquote>\n  <p>As an AI language model, I do not have a personal agenda, so I did not create the sentences to express any specific purpose or opinion. However, I can tell you that the sentences you provided discuss the potential suitability of transformer-based neural networks for predicting student performance from game play data. The probability that these sentences were created by me is 100%, as they were generated by my language model.</p>\n</blockquote>\n<p>😂😂😂 Not too high, just 100%</p>",
              "rawMarkdown": "So let's just ask chatgpt:\n>what is the probability the following setences created by you: paste his answer above\n\nThis is the my chatgpt response:\n\n>As an AI language model, I do not have a personal agenda, so I did not create the sentences to express any specific purpose or opinion. However, I can tell you that the sentences you provided discuss the potential suitability of transformer-based neural networks for predicting student performance from game play data. The probability that these sentences were created by me is 100%, as they were generated by my language model.\n\n😂😂😂 Not too high, just 100%",
              "votes": 5
            },
            {
              "id": 2231667,
              "postDate": "2023-04-23T14:25:56.183Z",
              "content": "<p>😂😂😂😂 So funny</p>",
              "rawMarkdown": "😂😂😂😂 So funny"
            }
          ]
        }
      ]
    },
    {
      "id": 2246144,
      "postDate": "2023-05-04T23:20:15.837Z",
      "rawMarkdown": "",
      "isDeleted": true
    }
  ],
  "comments": [
    {
      "id": 2227673,
      "author_name": "Ryota",
      "author_url": "",
      "post_date": "2023-04-19T23:47:35.473000",
      "content": "<p>I experimented Transformer-based NN for several days. I reach 0.694 on LB with 1fold.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 2234608,
          "author_name": "Minh Tri Phan",
          "author_url": "",
          "post_date": "2023-04-25T10:57:02.943000",
          "content": "<p>I am also having some Transformer-based models with CV of 0.6938, I haven't submitted them yet, but I expect LB is also around your LB with the Transformer. Yet, I am really struggling in improving it. Not sure it could reach 0.698 CV with the NN approach or not..</p>",
          "votes": 1,
          "replies": [
            {
              "id": 2234759,
              "author_name": "Ryota",
              "author_url": "",
              "post_date": "2023-04-25T13:23:32.493000",
              "content": "<p>Small changes such as altering the architecture do not seem to contribute to improvement. I think it is necessary to input features that cannot be captured by the Transformer. For example, the features that are typically used in GBDT. But, I haven't tried it yet</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2234770,
              "author_name": "Minh Tri Phan",
              "author_url": "",
              "post_date": "2023-04-25T13:27:48.383000",
              "content": "<p>Agree! I am now trying something new. I also tried adding the aggregated features, but it makes the NN converge really fast, and it didn't improve the model. I concatenate these aggregated features to the session embeddings, btw. Maybe there are some things I need to tweak.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2246039,
              "author_name": "Minh Tri Phan",
              "author_url": "",
              "post_date": "2023-05-04T19:18:02.973000",
              "content": "<p><a href=\"https://www.kaggle.com/ryotak12\" target=\"_blank\">@ryotak12</a> may I ask how much the CV of your 0.694 LB model is if you don't mind sharing it? I have a model with 0.694 CV (1 fold) but score only 0.690 LB with the transformer. However, with a GBDT model, 1-fold CV is 0.693, and 1-fold LB is 0.695. I spent around a week searching for the bug but didn't see any. I am wondering if the Transformer model could be overfitting or not.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2246145,
              "author_name": "Ryota",
              "author_url": "",
              "post_date": "2023-05-04T23:20:34.917000",
              "content": "<p>When 0.694 LB, CV is 0.6947. In my case, CV and LB are correlated!</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2246315,
              "author_name": "Minh Tri Phan",
              "author_url": "",
              "post_date": "2023-05-05T04:25:42.690000",
              "content": "<p>Thank you very much for your sharing!</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2160929,
      "author_name": "Ashkan Forootan",
      "author_url": "",
      "post_date": "2023-02-27T05:16:12.123000",
      "content": "<p>The suitability of transformer-based neural networks for the task of predicting student performance from game play largely depends on the specific requirements of the competition and the nature of the data being used.</p>\n<p>Transformer-based neural networks, such as the BERT (Bidirectional Encoder Representations from Transformers) model, have shown promising results in natural language processing (NLP) tasks due to their ability to capture long-term dependencies and handle sequential data effectively. However, predicting student performance from game play data may not involve the same type of sequential data that is common in NLP tasks.</p>\n<p>That being said, transformer-based models have also been successful in tasks involving non-sequential data, such as image and audio processing. Therefore, it's possible that they could be effective for predicting student performance from game play data, especially if the data contains complex relationships or patterns that require a deep understanding of the context.</p>\n<p>Ultimately, the effectiveness of transformer-based neural networks for this competition will depend on factors such as the size and complexity of the dataset, the specific features being used, and the performance metrics being evaluated. It's worth exploring the use of transformer-based models along with other machine learning algorithms to determine the most effective approach for the task at hand.</p>",
      "votes": -17,
      "replies": [
        {
          "id": 2161099,
          "author_name": "empty",
          "author_url": "",
          "post_date": "2023-02-27T07:51:20.250000",
          "content": "<p>was it generated with chatgpt? 🤔</p>",
          "votes": 6,
          "replies": [
            {
              "id": 2161101,
              "author_name": "william.wu",
              "author_url": "",
              "post_date": "2023-02-27T07:53:04.820000",
              "content": "<p>hhha, I have the same feeling</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2161126,
              "author_name": "BUUMOO",
              "author_url": "",
              "post_date": "2023-02-27T08:26:14.060000",
              "content": "<p>So let's just ask chatgpt:</p>\n<blockquote>\n  <p>what is the probability the following setences created by you: paste his answer above</p>\n</blockquote>\n<p>This is the my chatgpt response:</p>\n<blockquote>\n  <p>As an AI language model, I do not have a personal agenda, so I did not create the sentences to express any specific purpose or opinion. However, I can tell you that the sentences you provided discuss the potential suitability of transformer-based neural networks for predicting student performance from game play data. The probability that these sentences were created by me is 100%, as they were generated by my language model.</p>\n</blockquote>\n<p>😂😂😂 Not too high, just 100%</p>",
              "votes": 5,
              "replies": []
            },
            {
              "id": 2231667,
              "author_name": "WSpark",
              "author_url": "",
              "post_date": "2023-04-23T14:25:56.183000",
              "content": "<p>😂😂😂😂 So funny</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2246144,
      "author_name": "",
      "author_url": "",
      "post_date": "2023-05-04T23:20:15.837000",
      "content": "",
      "votes": 0,
      "replies": []
    }
  ],
  "raw_markdown_by_id": {
    "2160246": "This competition is very similar to the [Riid Competition](https://www.kaggle.com/competitions/riiid-test-answer-prediction/discussion/193250). Both of these 2 competitions ask you to predict whether the student is able to answer future questions correctly based on his past behaviors. The main difference is that Riid provides students' past correctness but this competition doesn't. However, we can still use Transformer-based NNs for this competition.",
    "2227673": "I experimented Transformer-based NN for several days. I reach 0.694 on LB with 1fold.",
    "2160929": "The suitability of transformer-based neural networks for the task of predicting student performance from game play largely depends on the specific requirements of the competition and the nature of the data being used.\n\nTransformer-based neural networks, such as the BERT (Bidirectional Encoder Representations from Transformers) model, have shown promising results in natural language processing (NLP) tasks due to their ability to capture long-term dependencies and handle sequential data effectively. However, predicting student performance from game play data may not involve the same type of sequential data that is common in NLP tasks.\n\nThat being said, transformer-based models have also been successful in tasks involving non-sequential data, such as image and audio processing. Therefore, it's possible that they could be effective for predicting student performance from game play data, especially if the data contains complex relationships or patterns that require a deep understanding of the context.\n\nUltimately, the effectiveness of transformer-based neural networks for this competition will depend on factors such as the size and complexity of the dataset, the specific features being used, and the performance metrics being evaluated. It's worth exploring the use of transformer-based models along with other machine learning algorithms to determine the most effective approach for the task at hand.",
    "2246144": ""
  }
}