{
  "id": 417070,
  "title": "Submission scoring time - Why such a huge test data set?",
  "url": "/competitions/predict-student-performance-from-game-play/discussion/417070",
  "author_name": "Murugesan Narayanaswamy",
  "post_date": "2023-06-14T04:49:07.420000",
  "votes": 0,
  "comment_count": 23,
  "views": 0,
  "content": "<p>This problem involves time sequence dataset. We can classify an answer as to whether it is correct or not, only after studying the entire session activities . The entire session has to be classified as leading to correct question or not. Averaging the session details (as  done in starter notebook) loses all the important features / data. Hence, it is a time sequence problem where we cannot arrive at a classification on row-by-row basis. Hence, I tried a basic transformer model. But it got timed-out in submission scoring due to 9 hours limit!! :-(</p>\n<p>But why should the test dataset be so huge as to take 9 hours to do scoring? </p>\n<p>I suggest that competition host try halving the hidden test dataset - will it change the leadership board ranking? If it does change the leadership board ranking, then it might indicate a problem in the whole modelling of the problem. The ranking should not change based on the size of the test dataset. Such a huge size is required to capture all variations and patterns? If so, does the training dataset captures all such variations and patterns?  I guess increasing the size of the hidden dataset beyond a point would affect only at 3rd-4th decimal of the score, about which competition host is not much bothered anyway..</p>\n<p>When they require small &amp; lightweight models, why go for transformers? - I believe the problem demands such a solution - it takes just 3 hours to complete the transformer notebook with GPU.  Moreover, once trained, it can be used forever (till retraining is required) - after all, only the \"running seconds of submission evaluation' matters…(though I am not sure as of now about the efficiency of transformer model - I got all 95% F1 Score with half the dataset !! - mostly it could be overfitting as it's a naive basic model. Has anyone tried transformer models?)</p>\n<p>Also, I could not understand this 'feature engineering obsession' with other kagglers in this competition. Some of the notebooks talks about incorporating thousands of features!!!! I simply could not understand… Is this normal in Kaggle competitions? Is this specially so in this competition? (am new to kaggle competitions). This is about a student using a GUI based class room game. Where is the question of creating thousands of features?! If that's for fourth decimal of the submission score, it might not benefit much to the competition host!!</p>",
  "messages": [
    {
      "id": 2301759,
      "postDate": "2023-06-14T06:37:28.537Z",
      "content": "<blockquote>\n  <p>Also, I could not understand this 'feature engineering obsession' with other kagglers in this competition.</p>\n</blockquote>\n<p>At least '<em>feature engineering</em> obsession' makes more sense than '<em>train 100 deep learning models with different seeds then ensemble them all together just because my company has that many A100 GPUs</em> obsession', I assume🤷‍♂️</p>\n<p>Also, more feature never means a better score since you are certainly going to face overfitting with so many features at some point.</p>",
      "rawMarkdown": ">Also, I could not understand this 'feature engineering obsession' with other kagglers in this competition.\n\nAt least '*feature engineering* obsession' makes more sense than '*train 100 deep learning models with different seeds then ensemble them all together just because my company has that many A100 GPUs* obsession', I assume🤷‍♂️\n\nAlso, more feature never means a better score since you are certainly going to face overfitting with so many features at some point.",
      "votes": 5,
      "replies": [
        {
          "id": 2302468,
          "postDate": "2023-06-14T15:15:54.060Z",
          "content": "<p>I agree with you - training and ensembling hundreds of models for third decimal accuracy defeats the purpose of competition hosts, but when it comes to competition and when the community is thriving to compete and outsmart each other, such things cannot be avoided, but the spirit of the competition - i.e. to create a model that can serve the purpose or objective of the competition host should be served in any case.  </p>\n<p>In our case, if there is a slight change in the video game, either in design of screen or in navigation or in dialogues, will all these extra feature engineering continue to make sense? Feature engineering is done to make the model consider all factors involved in the phenomena being modelled, not for improving accuracy in general</p>",
          "rawMarkdown": "I agree with you - training and ensembling hundreds of models for third decimal accuracy defeats the purpose of competition hosts, but when it comes to competition and when the community is thriving to compete and outsmart each other, such things cannot be avoided, but the spirit of the competition - i.e. to create a model that can serve the purpose or objective of the competition host should be served in any case.  \n\nIn our case, if there is a slight change in the video game, either in design of screen or in navigation or in dialogues, will all these extra feature engineering continue to make sense? Feature engineering is done to make the model consider all factors involved in the phenomena being modelled, not for improving accuracy in general",
          "replies": [
            {
              "id": 2302973,
              "postDate": "2023-06-15T02:07:18.020Z",
              "content": "<p>One of the reasons people need so many different features for this competition is many people use 18 different model for each question and the features needed for different model varies a lot.</p>\n<p>Also if you have read other threads, other Kaggler are actually quite aware of the redundancy of features, and feature elimination has always been a very important topic of this competition. I'm curious about how other teams handle their strategy but I guess we will only learn when competition ends.</p>",
              "rawMarkdown": "One of the reasons people need so many different features for this competition is many people use 18 different model for each question and the features needed for different model varies a lot.\n\nAlso if you have read other threads, other Kaggler are actually quite aware of the redundancy of features, and feature elimination has always been a very important topic of this competition. I'm curious about how other teams handle their strategy but I guess we will only learn when competition ends."
            },
            {
              "id": 2302986,
              "postDate": "2023-06-15T02:28:19.323Z",
              "content": "<p>I am just wondering, given that thousands of features are being talked about, it should inlcude all permutations and combinations of existing features, but will this  feature engineering hold good if there are any slight modifications to the game in their next release of the learning video game? Or will the accuracies take a huge beating?</p>",
              "rawMarkdown": "I am just wondering, given that thousands of features are being talked about, it should inlcude all permutations and combinations of existing features, but will this  feature engineering hold good if there are any slight modifications to the game in their next release of the learning video game? Or will the accuracies take a huge beating?",
              "votes": 1
            },
            {
              "id": 2303159,
              "postDate": "2023-06-15T05:54:13.517Z",
              "content": "<blockquote>\n  <p>it should inlcude all permutations and combinations of existing features</p>\n</blockquote>\n<p>No, I only use features that give a post-permutation CV boost, there are tons of combinations I don't use even though I could easily come up with a reason why it should work.</p>\n<p>It's not the feature itself important, the important thing is the strategy of how you generate your feature. Even with a different game, I can generate that many features with the same strategy.</p>\n<p>Also, it's frustrating you are complaining about the 'meaningfulness' of 'feature engineering' while using a transformer model as the opposite example. At least features - even those that can't boost CV - have a specific meaning while hyperparameter-tuning just uses optuna to set an interval.</p>",
              "rawMarkdown": "> it should inlcude all permutations and combinations of existing features\n\nNo, I only use features that give a post-permutation CV boost, there are tons of combinations I don't use even though I could easily come up with a reason why it should work.\n\nIt's not the feature itself important, the important thing is the strategy of how you generate your feature. Even with a different game, I can generate that many features with the same strategy.\n\nAlso, it's frustrating you are complaining about the 'meaningfulness' of 'feature engineering' while using a transformer model as the opposite example. At least features - even those that can't boost CV - have a specific meaning while hyperparameter-tuning just uses optuna to set an interval.\n"
            },
            {
              "id": 2303551,
              "postDate": "2023-06-15T10:51:59.950Z",
              "content": "<blockquote>\n  <p>It should include all permutations and combinations of existing features, </p>\n</blockquote>\n<p>Feature combination should be safe? If you use any neural network, the fully connected layer is basically blending all features together. The features based on event name, and object location can be dangerous if we need to apply the model to a different game. (given that we don't have an automatic way to extract this information).</p>\n<blockquote>\n  <p>Will this feature engineering hold good if there are any slight modifications to the game in their next release of the learning video game?</p>\n</blockquote>\n<p>This is a good question. But we won't know the answer, because we don't have the information that which part of the game is subject to be changed. In this competition, the game setting in the test and the train are the same. </p>\n<p>I do agree that the problem statement of this competition is not ideal.</p>",
              "rawMarkdown": "> It should include all permutations and combinations of existing features, \n\nFeature combination should be safe? If you use any neural network, the fully connected layer is basically blending all features together. The features based on event name, and object location can be dangerous if we need to apply the model to a different game. (given that we don't have an automatic way to extract this information).\n\n> Will this feature engineering hold good if there are any slight modifications to the game in their next release of the learning video game?\n\nThis is a good question. But we won't know the answer, because we don't have the information that which part of the game is subject to be changed. In this competition, the game setting in the test and the train are the same. \n\nI do agree that the problem statement of this competition is not ideal.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2302204,
      "postDate": "2023-06-14T12:05:04.413Z",
      "content": "<p>Why do you think the test set is huge? At the start of the competition we had train_set and test_set of similar size (~12000 sessions each). After the test set was replaced, it was noticed that submissions are scored much faster than before, so the current test_set is going to be noticeably less than 12000 sessions. Can we call it a huge data set? I would rather say it's small.</p>",
      "rawMarkdown": "Why do you think the test set is huge? At the start of the competition we had train_set and test_set of similar size (~12000 sessions each). After the test set was replaced, it was noticed that submissions are scored much faster than before, so the current test_set is going to be noticeably less than 12000 sessions. Can we call it a huge data set? I would rather say it's small.",
      "votes": 3,
      "replies": [
        {
          "id": 2302456,
          "postDate": "2023-06-14T15:07:47.983Z",
          "content": "<p>The size of the test dataset should be sufficient enough to detect any overfitting done to the training data. While training, the general practice is to just use 20% of the training data for validating the model for any overfitting. Why should the size of the (hidden private) test data be as big as training data? Should the leadership board ranking depend on the size of the test data? The ranking should be same for any size greater than the size that is required to detect overfitting, and that size should not necessitate as big as size of training data. If the leadership board ranking keeps on changing as size is increased, then modelling may not be correct</p>\n<p>In this specific case, since time series API is used for evaluation, the scoring time matters and hence size of test data. In real time usage, they can use batches to evaluate and predict test data. So, running time of any individual prediction should matter.</p>",
          "rawMarkdown": "The size of the test dataset should be sufficient enough to detect any overfitting done to the training data. While training, the general practice is to just use 20% of the training data for validating the model for any overfitting. Why should the size of the (hidden private) test data be as big as training data? Should the leadership board ranking depend on the size of the test data? The ranking should be same for any size greater than the size that is required to detect overfitting, and that size should not necessitate as big as size of training data. If the leadership board ranking keeps on changing as size is increased, then modelling may not be correct\n\nIn this specific case, since time series API is used for evaluation, the scoring time matters and hence size of test data. In real time usage, they can use batches to evaluate and predict test data. So, running time of any individual prediction should matter.",
          "replies": [
            {
              "id": 2302632,
              "postDate": "2023-06-14T17:07:49.510Z",
              "content": "<p>The test set is not as large as the training set now. Currently, the training set contains ~24000 sessions, and the test set can be estimated at ~6000 sessions (I don't remember how much the submission time decreased, but more or less twice). So now the test set is 25% of training data - exactly what you wish.</p>",
              "rawMarkdown": "The test set is not as large as the training set now. Currently, the training set contains ~24000 sessions, and the test set can be estimated at ~6000 sessions (I don't remember how much the submission time decreased, but more or less twice). So now the test set is 25% of training data - exactly what you wish.",
              "votes": 1
            },
            {
              "id": 2302887,
              "postDate": "2023-06-15T00:02:55.957Z",
              "content": "<p>If you combine with private hidden test that will be used for final scoring after deadline, then the test data will be 50% of training data!!</p>\n<p>I hope they synthetically engineer test data with different GUI possibilities etc., to burst features that overfit the data…will these models based on feature engineering work with the same accuracy in next version of their video games?!</p>",
              "rawMarkdown": "If you combine with private hidden test that will be used for final scoring after deadline, then the test data will be 50% of training data!!\n\nI hope they synthetically engineer test data with different GUI possibilities etc., to burst features that overfit the data...will these models based on feature engineering work with the same accuracy in next version of their video games?!"
            },
            {
              "id": 2303102,
              "postDate": "2023-06-15T05:01:39.267Z",
              "content": "<p>That's not how it works. The test set on which submissions are scored already contains a Public part + a Private part. It's just that for now we only see the score from the Public part, and after the deadline we will see the score from the Private part. So the test data is not 50% of training data, it's less than 25%. </p>",
              "rawMarkdown": "That's not how it works. The test set on which submissions are scored already contains a Public part + a Private part. It's just that for now we only see the score from the Public part, and after the deadline we will see the score from the Private part. So the test data is not 50% of training data, it's less than 25%. "
            },
            {
              "id": 2303117,
              "postDate": "2023-06-15T05:10:52.540Z",
              "content": "<p>It would be even better if they engineer a private test data from a completely different game to burst all the features built by the feature-obsessed teams. This will check if the models will work with the same accuracy in all other video games.</p>",
              "rawMarkdown": "It would be even better if they engineer a private test data from a completely different game to burst all the features built by the feature-obsessed teams. This will check if the models will work with the same accuracy in all other video games.",
              "votes": 1
            }
          ]
        }
      ]
    },
    {
      "id": 2301660,
      "postDate": "2023-06-14T04:49:07.420Z",
      "content": "<p>This problem involves time sequence dataset. We can classify an answer as to whether it is correct or not, only after studying the entire session activities . The entire session has to be classified as leading to correct question or not. Averaging the session details (as  done in starter notebook) loses all the important features / data. Hence, it is a time sequence problem where we cannot arrive at a classification on row-by-row basis. Hence, I tried a basic transformer model. But it got timed-out in submission scoring due to 9 hours limit!! :-(</p>\n<p>But why should the test dataset be so huge as to take 9 hours to do scoring? </p>\n<p>I suggest that competition host try halving the hidden test dataset - will it change the leadership board ranking? If it does change the leadership board ranking, then it might indicate a problem in the whole modelling of the problem. The ranking should not change based on the size of the test dataset. Such a huge size is required to capture all variations and patterns? If so, does the training dataset captures all such variations and patterns?  I guess increasing the size of the hidden dataset beyond a point would affect only at 3rd-4th decimal of the score, about which competition host is not much bothered anyway..</p>\n<p>When they require small &amp; lightweight models, why go for transformers? - I believe the problem demands such a solution - it takes just 3 hours to complete the transformer notebook with GPU.  Moreover, once trained, it can be used forever (till retraining is required) - after all, only the \"running seconds of submission evaluation' matters…(though I am not sure as of now about the efficiency of transformer model - I got all 95% F1 Score with half the dataset !! - mostly it could be overfitting as it's a naive basic model. Has anyone tried transformer models?)</p>\n<p>Also, I could not understand this 'feature engineering obsession' with other kagglers in this competition. Some of the notebooks talks about incorporating thousands of features!!!! I simply could not understand… Is this normal in Kaggle competitions? Is this specially so in this competition? (am new to kaggle competitions). This is about a student using a GUI based class room game. Where is the question of creating thousands of features?! If that's for fourth decimal of the submission score, it might not benefit much to the competition host!!</p>",
      "rawMarkdown": "This problem involves time sequence dataset. We can classify an answer as to whether it is correct or not, only after studying the entire session activities . The entire session has to be classified as leading to correct question or not. Averaging the session details (as  done in starter notebook) loses all the important features / data. Hence, it is a time sequence problem where we cannot arrive at a classification on row-by-row basis. Hence, I tried a basic transformer model. But it got timed-out in submission scoring due to 9 hours limit!! :-(\n\nBut why should the test dataset be so huge as to take 9 hours to do scoring? \n\nI suggest that competition host try halving the hidden test dataset - will it change the leadership board ranking? If it does change the leadership board ranking, then it might indicate a problem in the whole modelling of the problem. The ranking should not change based on the size of the test dataset. Such a huge size is required to capture all variations and patterns? If so, does the training dataset captures all such variations and patterns?  I guess increasing the size of the hidden dataset beyond a point would affect only at 3rd-4th decimal of the score, about which competition host is not much bothered anyway..\n\nWhen they require small & lightweight models, why go for transformers? - I believe the problem demands such a solution - it takes just 3 hours to complete the transformer notebook with GPU.  Moreover, once trained, it can be used forever (till retraining is required) - after all, only the \"running seconds of submission evaluation' matters...(though I am not sure as of now about the efficiency of transformer model - I got all 95% F1 Score with half the dataset !! - mostly it could be overfitting as it's a naive basic model. Has anyone tried transformer models?)\n\nAlso, I could not understand this 'feature engineering obsession' with other kagglers in this competition. Some of the notebooks talks about incorporating thousands of features!!!! I simply could not understand... Is this normal in Kaggle competitions? Is this specially so in this competition? (am new to kaggle competitions). This is about a student using a GUI based class room game. Where is the question of creating thousands of features?! If that's for fourth decimal of the submission score, it might not benefit much to the competition host!!\n\n"
    },
    {
      "id": 2301878,
      "postDate": "2023-06-14T07:42:29.030Z",
      "content": "<p>I doubt you will get 0.95 just by using transformer, or any other model. Some of the questions are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?</p>",
      "rawMarkdown": "I doubt you will get 0.95 just by using transformer, or any other model. Some of the questions are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?",
      "replies": [
        {
          "id": 2301894,
          "postDate": "2023-06-14T07:53:09.630Z",
          "content": "<blockquote>\n  <p>I doubt you will get 0.95 just by using transformer, or any other model. Some of the problems are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?</p>\n</blockquote>\n<p>I think he means 95% of his score, not a 0.95 score.</p>\n<p>Still, 0.7 * (1-95%) = 0.035 is actually a lot for this competition.</p>",
          "rawMarkdown": "> I doubt you will get 0.95 just by using transformer, or any other model. Some of the problems are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?\n\nI think he means 95% of his score, not a 0.95 score.\n\nStill, 0.7 * (1-95%) = 0.035 is actually a lot for this competition.",
          "replies": [
            {
              "id": 2301896,
              "postDate": "2023-06-14T08:02:11.080Z",
              "content": "<p>Sorry, I'm not following. What do you mean by his score? I thought he meant he got 0.95 F1 in CV.</p>",
              "rawMarkdown": "Sorry, I'm not following. What do you mean by his score? I thought he meant he got 0.95 F1 in CV."
            },
            {
              "id": 2301922,
              "postDate": "2023-06-14T08:33:49.147Z",
              "content": "<p>When he said </p>\n<blockquote>\n  <p>'I got all 95% F1 Score with half the dataset !! ' </p>\n</blockquote>\n<p>I think he means the score of his transformer model is 95% of the top score(or his top score), not that he achieved a 0.95 score with his transformer model.</p>\n<p>I don't see his point though. I don't doubt he can achieve a 95% score with a deep learning model, but 5% of the score is still significant for most competitions and that's why people prefer traditional ml over dl model in most tabular/time-series data competitions because dl model is not competitive enough. </p>",
              "rawMarkdown": "When he said \n>'I got all 95% F1 Score with half the dataset !! ' \n\nI think he means the score of his transformer model is 95% of the top score(or his top score), not that he achieved a 0.95 score with his transformer model.\n\nI don't see his point though. I don't doubt he can achieve a 95% score with a deep learning model, but 5% of the score is still significant for most competitions and that's why people prefer traditional ml over dl model in most tabular/time-series data competitions because dl model is not competitive enough. \n\n"
            },
            {
              "id": 2301931,
              "postDate": "2023-06-14T08:39:03.813Z",
              "content": "<p>Riiiiight, that make more sense now.</p>",
              "rawMarkdown": "Riiiiight, that make more sense now."
            },
            {
              "id": 2302501,
              "postDate": "2023-06-14T15:30:04.120Z",
              "content": "<p>It is the characteristics of deep learning model to overfit any dataset. I just tried a basic model without any parameter tuning with basic data - so, it could be a case of overfitting as is normal with DL models. Have not tried cross validation yet, not sure how much time will that take…let me check</p>",
              "rawMarkdown": "It is the characteristics of deep learning model to overfit any dataset. I just tried a basic model without any parameter tuning with basic data - so, it could be a case of overfitting as is normal with DL models. Have not tried cross validation yet, not sure how much time will that take...let me check"
            }
          ]
        },
        {
          "id": 2302485,
          "postDate": "2023-06-14T15:24:05.110Z",
          "content": "<p>I created a naive-basic transformer model based on keras tutorial implementation for time sequence. I just wanted to test the model - so, in order to quickly create a workable prototype, I did not add any of categorical variables and considered only 200 rows for each session (first few sessions use 2000 rows!!).So, just half the dataset. It was giving F1 Score of 88% to 96% for most of the questions of validation test dataset (20%). Only for the questions [5,10,13,15], I got F1 scores of [0.65,0.58,0.038,0.48] - Not clear why F1 score for question 13 alone is 0.038!!</p>\n<p>But since I am not considering most of the dataset, I guess it should clearly be a case of overfitting - I wanted to find what exactly would be the score in LB, but it got timed out.. (I have not yet tried cross validation - it takes 3-4 hours in GPU for single run)</p>",
          "rawMarkdown": "I created a naive-basic transformer model based on keras tutorial implementation for time sequence. I just wanted to test the model - so, in order to quickly create a workable prototype, I did not add any of categorical variables and considered only 200 rows for each session (first few sessions use 2000 rows!!).So, just half the dataset. It was giving F1 Score of 88% to 96% for most of the questions of validation test dataset (20%). Only for the questions [5,10,13,15], I got F1 scores of [0.65,0.58,0.038,0.48] - Not clear why F1 score for question 13 alone is 0.038!!\n\nBut since I am not considering most of the dataset, I guess it should clearly be a case of overfitting - I wanted to find what exactly would be the score in LB, but it got timed out.. (I have not yet tried cross validation - it takes 3-4 hours in GPU for single run)",
          "replies": [
            {
              "id": 2302542,
              "postDate": "2023-06-14T15:54:13.830Z",
              "content": "<p>That's not how F1 works… You should probably read up on how F1 is computed, and use it on all questions rather than separately. </p>",
              "rawMarkdown": "That's not how F1 works... You should probably read up on how F1 is computed, and use it on all questions rather than separately. ",
              "votes": 2
            },
            {
              "id": 2304397,
              "postDate": "2023-06-16T01:06:29.617Z",
              "content": "<p>This is how F1 works for binary classification as per documentation :-) The concept of averaging is applicable only for multiclass problems at least as per sklearn documentation. </p>\n<p>I am not saying this competition's F1 metric - which does macro averaging between binary classes - is not appropriate. It is correct but it has its own advantages and disadvantages. Most importantly, it is very conservative</p>\n<p>One major disadvantage with the above  F1 metric (when calculated for all questions together) is - it is not interpretable (and I guess can be biased in some cases). For example, suppose I predict perfectly for 50% of questions (9 questions) and predict completely wrongly for remaining 50% of questions, if we calculate F1 score independently for each score and macro average it, then it will show F1 score as 50% . Thus, it is interpretable. On the other hand, the current metric might show something less than 50%.  Let us say 45%. But this 45% might happen in several other cases also…So, major disadvantage of this metric as used in this competition is we get no information on efficiency with respect to individual questions  </p>\n<p>But there is advantage from the perspective of the other dimension. Suppose I predict correctly all questions for 50% of users (sessions) and completely wrongly for remaining 50% of users, then F1 score that averages individual  F1 scores of all questions will not give 50% but could give greater than 50% !! But the competition metric is unbiased in this respect - it will give F1 score of something less than 50%</p>\n<p>Considering all the above aspects, the average of F1 metric of two binary classes as done in this competitions seems better  compared to averaging of individual f1 scores - except that we don't have information on efficiency with respect to predictions of individual quesions</p>",
              "rawMarkdown": "This is how F1 works for binary classification as per documentation :-) The concept of averaging is applicable only for multiclass problems at least as per sklearn documentation. \n\nI am not saying this competition's F1 metric - which does macro averaging between binary classes - is not appropriate. It is correct but it has its own advantages and disadvantages. Most importantly, it is very conservative\n\nOne major disadvantage with the above  F1 metric (when calculated for all questions together) is - it is not interpretable (and I guess can be biased in some cases). For example, suppose I predict perfectly for 50% of questions (9 questions) and predict completely wrongly for remaining 50% of questions, if we calculate F1 score independently for each score and macro average it, then it will show F1 score as 50% . Thus, it is interpretable. On the other hand, the current metric might show something less than 50%.  Let us say 45%. But this 45% might happen in several other cases also...So, major disadvantage of this metric as used in this competition is we get no information on efficiency with respect to individual questions  \n\nBut there is advantage from the perspective of the other dimension. Suppose I predict correctly all questions for 50% of users (sessions) and completely wrongly for remaining 50% of users, then F1 score that averages individual  F1 scores of all questions will not give 50% but could give greater than 50% !! But the competition metric is unbiased in this respect - it will give F1 score of something less than 50%\n\nConsidering all the above aspects, the average of F1 metric of two binary classes as done in this competitions seems better  compared to averaging of individual f1 scores - except that we don't have information on efficiency with respect to predictions of individual quesions"
            }
          ]
        },
        {
          "id": 2302533,
          "postDate": "2023-06-14T15:46:45.337Z",
          "rawMarkdown": "",
          "isDeleted": true
        },
        {
          "id": 2305923,
          "postDate": "2023-06-17T02:21:45.820Z",
          "content": "<p>There might be some mistakes in my model. Since transformer notebook is getting timed out during inference,  I tried first four questions alone with transformer models and rest with xgboost model, the LB score has come down from 0.674 to 0.597 :-) :-)</p>",
          "rawMarkdown": "There might be some mistakes in my model. Since transformer notebook is getting timed out during inference,  I tried first four questions alone with transformer models and rest with xgboost model, the LB score has come down from 0.674 to 0.597 :-) :-)\n\n"
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 2301759,
      "author_name": "Ya Xu",
      "author_url": "",
      "post_date": "2023-06-14T06:37:28.537000",
      "content": "<blockquote>\n  <p>Also, I could not understand this 'feature engineering obsession' with other kagglers in this competition.</p>\n</blockquote>\n<p>At least '<em>feature engineering</em> obsession' makes more sense than '<em>train 100 deep learning models with different seeds then ensemble them all together just because my company has that many A100 GPUs</em> obsession', I assume🤷‍♂️</p>\n<p>Also, more feature never means a better score since you are certainly going to face overfitting with so many features at some point.</p>",
      "votes": 5,
      "replies": [
        {
          "id": 2302468,
          "author_name": "Murugesan Narayanaswamy",
          "author_url": "",
          "post_date": "2023-06-14T15:15:54.060000",
          "content": "<p>I agree with you - training and ensembling hundreds of models for third decimal accuracy defeats the purpose of competition hosts, but when it comes to competition and when the community is thriving to compete and outsmart each other, such things cannot be avoided, but the spirit of the competition - i.e. to create a model that can serve the purpose or objective of the competition host should be served in any case.  </p>\n<p>In our case, if there is a slight change in the video game, either in design of screen or in navigation or in dialogues, will all these extra feature engineering continue to make sense? Feature engineering is done to make the model consider all factors involved in the phenomena being modelled, not for improving accuracy in general</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2302973,
              "author_name": "Ya Xu",
              "author_url": "",
              "post_date": "2023-06-15T02:07:18.020000",
              "content": "<p>One of the reasons people need so many different features for this competition is many people use 18 different model for each question and the features needed for different model varies a lot.</p>\n<p>Also if you have read other threads, other Kaggler are actually quite aware of the redundancy of features, and feature elimination has always been a very important topic of this competition. I'm curious about how other teams handle their strategy but I guess we will only learn when competition ends.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2302986,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-15T02:28:19.323000",
              "content": "<p>I am just wondering, given that thousands of features are being talked about, it should inlcude all permutations and combinations of existing features, but will this  feature engineering hold good if there are any slight modifications to the game in their next release of the learning video game? Or will the accuracies take a huge beating?</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2303159,
              "author_name": "Ya Xu",
              "author_url": "",
              "post_date": "2023-06-15T05:54:13.517000",
              "content": "<blockquote>\n  <p>it should inlcude all permutations and combinations of existing features</p>\n</blockquote>\n<p>No, I only use features that give a post-permutation CV boost, there are tons of combinations I don't use even though I could easily come up with a reason why it should work.</p>\n<p>It's not the feature itself important, the important thing is the strategy of how you generate your feature. Even with a different game, I can generate that many features with the same strategy.</p>\n<p>Also, it's frustrating you are complaining about the 'meaningfulness' of 'feature engineering' while using a transformer model as the opposite example. At least features - even those that can't boost CV - have a specific meaning while hyperparameter-tuning just uses optuna to set an interval.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2303551,
              "author_name": "Anthony Chiu",
              "author_url": "",
              "post_date": "2023-06-15T10:51:59.950000",
              "content": "<blockquote>\n  <p>It should include all permutations and combinations of existing features, </p>\n</blockquote>\n<p>Feature combination should be safe? If you use any neural network, the fully connected layer is basically blending all features together. The features based on event name, and object location can be dangerous if we need to apply the model to a different game. (given that we don't have an automatic way to extract this information).</p>\n<blockquote>\n  <p>Will this feature engineering hold good if there are any slight modifications to the game in their next release of the learning video game?</p>\n</blockquote>\n<p>This is a good question. But we won't know the answer, because we don't have the information that which part of the game is subject to be changed. In this competition, the game setting in the test and the train are the same. </p>\n<p>I do agree that the problem statement of this competition is not ideal.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2302204,
      "author_name": "Alvor",
      "author_url": "",
      "post_date": "2023-06-14T12:05:04.413000",
      "content": "<p>Why do you think the test set is huge? At the start of the competition we had train_set and test_set of similar size (~12000 sessions each). After the test set was replaced, it was noticed that submissions are scored much faster than before, so the current test_set is going to be noticeably less than 12000 sessions. Can we call it a huge data set? I would rather say it's small.</p>",
      "votes": 3,
      "replies": [
        {
          "id": 2302456,
          "author_name": "Murugesan Narayanaswamy",
          "author_url": "",
          "post_date": "2023-06-14T15:07:47.983000",
          "content": "<p>The size of the test dataset should be sufficient enough to detect any overfitting done to the training data. While training, the general practice is to just use 20% of the training data for validating the model for any overfitting. Why should the size of the (hidden private) test data be as big as training data? Should the leadership board ranking depend on the size of the test data? The ranking should be same for any size greater than the size that is required to detect overfitting, and that size should not necessitate as big as size of training data. If the leadership board ranking keeps on changing as size is increased, then modelling may not be correct</p>\n<p>In this specific case, since time series API is used for evaluation, the scoring time matters and hence size of test data. In real time usage, they can use batches to evaluate and predict test data. So, running time of any individual prediction should matter.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2302632,
              "author_name": "Alvor",
              "author_url": "",
              "post_date": "2023-06-14T17:07:49.510000",
              "content": "<p>The test set is not as large as the training set now. Currently, the training set contains ~24000 sessions, and the test set can be estimated at ~6000 sessions (I don't remember how much the submission time decreased, but more or less twice). So now the test set is 25% of training data - exactly what you wish.</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 2302887,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-15T00:02:55.957000",
              "content": "<p>If you combine with private hidden test that will be used for final scoring after deadline, then the test data will be 50% of training data!!</p>\n<p>I hope they synthetically engineer test data with different GUI possibilities etc., to burst features that overfit the data…will these models based on feature engineering work with the same accuracy in next version of their video games?!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2303102,
              "author_name": "Alvor",
              "author_url": "",
              "post_date": "2023-06-15T05:01:39.267000",
              "content": "<p>That's not how it works. The test set on which submissions are scored already contains a Public part + a Private part. It's just that for now we only see the score from the Public part, and after the deadline we will see the score from the Private part. So the test data is not 50% of training data, it's less than 25%. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2303117,
              "author_name": "Alvor",
              "author_url": "",
              "post_date": "2023-06-15T05:10:52.540000",
              "content": "<p>It would be even better if they engineer a private test data from a completely different game to burst all the features built by the feature-obsessed teams. This will check if the models will work with the same accuracy in all other video games.</p>",
              "votes": 1,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 2301878,
      "author_name": "Woprime",
      "author_url": "",
      "post_date": "2023-06-14T07:42:29.030000",
      "content": "<p>I doubt you will get 0.95 just by using transformer, or any other model. Some of the questions are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?</p>",
      "votes": 0,
      "replies": [
        {
          "id": 2301894,
          "author_name": "Ya Xu",
          "author_url": "",
          "post_date": "2023-06-14T07:53:09.630000",
          "content": "<blockquote>\n  <p>I doubt you will get 0.95 just by using transformer, or any other model. Some of the problems are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?</p>\n</blockquote>\n<p>I think he means 95% of his score, not a 0.95 score.</p>\n<p>Still, 0.7 * (1-95%) = 0.035 is actually a lot for this competition.</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2301896,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-06-14T08:02:11.080000",
              "content": "<p>Sorry, I'm not following. What do you mean by his score? I thought he meant he got 0.95 F1 in CV.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2301922,
              "author_name": "Ya Xu",
              "author_url": "",
              "post_date": "2023-06-14T08:33:49.147000",
              "content": "<p>When he said </p>\n<blockquote>\n  <p>'I got all 95% F1 Score with half the dataset !! ' </p>\n</blockquote>\n<p>I think he means the score of his transformer model is 95% of the top score(or his top score), not that he achieved a 0.95 score with his transformer model.</p>\n<p>I don't see his point though. I don't doubt he can achieve a 95% score with a deep learning model, but 5% of the score is still significant for most competitions and that's why people prefer traditional ml over dl model in most tabular/time-series data competitions because dl model is not competitive enough. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2301931,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-06-14T08:39:03.813000",
              "content": "<p>Riiiiight, that make more sense now.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 2302501,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-14T15:30:04.120000",
              "content": "<p>It is the characteristics of deep learning model to overfit any dataset. I just tried a basic model without any parameter tuning with basic data - so, it could be a case of overfitting as is normal with DL models. Have not tried cross validation yet, not sure how much time will that take…let me check</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2302485,
          "author_name": "Murugesan Narayanaswamy",
          "author_url": "",
          "post_date": "2023-06-14T15:24:05.110000",
          "content": "<p>I created a naive-basic transformer model based on keras tutorial implementation for time sequence. I just wanted to test the model - so, in order to quickly create a workable prototype, I did not add any of categorical variables and considered only 200 rows for each session (first few sessions use 2000 rows!!).So, just half the dataset. It was giving F1 Score of 88% to 96% for most of the questions of validation test dataset (20%). Only for the questions [5,10,13,15], I got F1 scores of [0.65,0.58,0.038,0.48] - Not clear why F1 score for question 13 alone is 0.038!!</p>\n<p>But since I am not considering most of the dataset, I guess it should clearly be a case of overfitting - I wanted to find what exactly would be the score in LB, but it got timed out.. (I have not yet tried cross validation - it takes 3-4 hours in GPU for single run)</p>",
          "votes": 0,
          "replies": [
            {
              "id": 2302542,
              "author_name": "Woprime",
              "author_url": "",
              "post_date": "2023-06-14T15:54:13.830000",
              "content": "<p>That's not how F1 works… You should probably read up on how F1 is computed, and use it on all questions rather than separately. </p>",
              "votes": 2,
              "replies": []
            },
            {
              "id": 2304397,
              "author_name": "Murugesan Narayanaswamy",
              "author_url": "",
              "post_date": "2023-06-16T01:06:29.617000",
              "content": "<p>This is how F1 works for binary classification as per documentation :-) The concept of averaging is applicable only for multiclass problems at least as per sklearn documentation. </p>\n<p>I am not saying this competition's F1 metric - which does macro averaging between binary classes - is not appropriate. It is correct but it has its own advantages and disadvantages. Most importantly, it is very conservative</p>\n<p>One major disadvantage with the above  F1 metric (when calculated for all questions together) is - it is not interpretable (and I guess can be biased in some cases). For example, suppose I predict perfectly for 50% of questions (9 questions) and predict completely wrongly for remaining 50% of questions, if we calculate F1 score independently for each score and macro average it, then it will show F1 score as 50% . Thus, it is interpretable. On the other hand, the current metric might show something less than 50%.  Let us say 45%. But this 45% might happen in several other cases also…So, major disadvantage of this metric as used in this competition is we get no information on efficiency with respect to individual questions  </p>\n<p>But there is advantage from the perspective of the other dimension. Suppose I predict correctly all questions for 50% of users (sessions) and completely wrongly for remaining 50% of users, then F1 score that averages individual  F1 scores of all questions will not give 50% but could give greater than 50% !! But the competition metric is unbiased in this respect - it will give F1 score of something less than 50%</p>\n<p>Considering all the above aspects, the average of F1 metric of two binary classes as done in this competitions seems better  compared to averaging of individual f1 scores - except that we don't have information on efficiency with respect to predictions of individual quesions</p>",
              "votes": 0,
              "replies": []
            }
          ]
        },
        {
          "id": 2302533,
          "author_name": "",
          "author_url": "",
          "post_date": "2023-06-14T15:46:45.337000",
          "content": "",
          "votes": 0,
          "replies": []
        },
        {
          "id": 2305923,
          "author_name": "Murugesan Narayanaswamy",
          "author_url": "",
          "post_date": "2023-06-17T02:21:45.820000",
          "content": "<p>There might be some mistakes in my model. Since transformer notebook is getting timed out during inference,  I tried first four questions alone with transformer models and rest with xgboost model, the LB score has come down from 0.674 to 0.597 :-) :-)</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "2301759": ">Also, I could not understand this 'feature engineering obsession' with other kagglers in this competition.\n\nAt least '*feature engineering* obsession' makes more sense than '*train 100 deep learning models with different seeds then ensemble them all together just because my company has that many A100 GPUs* obsession', I assume🤷‍♂️\n\nAlso, more feature never means a better score since you are certainly going to face overfitting with so many features at some point.",
    "2302204": "Why do you think the test set is huge? At the start of the competition we had train_set and test_set of similar size (~12000 sessions each). After the test set was replaced, it was noticed that submissions are scored much faster than before, so the current test_set is going to be noticeably less than 12000 sessions. Can we call it a huge data set? I would rather say it's small.",
    "2301660": "This problem involves time sequence dataset. We can classify an answer as to whether it is correct or not, only after studying the entire session activities . The entire session has to be classified as leading to correct question or not. Averaging the session details (as  done in starter notebook) loses all the important features / data. Hence, it is a time sequence problem where we cannot arrive at a classification on row-by-row basis. Hence, I tried a basic transformer model. But it got timed-out in submission scoring due to 9 hours limit!! :-(\n\nBut why should the test dataset be so huge as to take 9 hours to do scoring? \n\nI suggest that competition host try halving the hidden test dataset - will it change the leadership board ranking? If it does change the leadership board ranking, then it might indicate a problem in the whole modelling of the problem. The ranking should not change based on the size of the test dataset. Such a huge size is required to capture all variations and patterns? If so, does the training dataset captures all such variations and patterns?  I guess increasing the size of the hidden dataset beyond a point would affect only at 3rd-4th decimal of the score, about which competition host is not much bothered anyway..\n\nWhen they require small & lightweight models, why go for transformers? - I believe the problem demands such a solution - it takes just 3 hours to complete the transformer notebook with GPU.  Moreover, once trained, it can be used forever (till retraining is required) - after all, only the \"running seconds of submission evaluation' matters...(though I am not sure as of now about the efficiency of transformer model - I got all 95% F1 Score with half the dataset !! - mostly it could be overfitting as it's a naive basic model. Has anyone tried transformer models?)\n\nAlso, I could not understand this 'feature engineering obsession' with other kagglers in this competition. Some of the notebooks talks about incorporating thousands of features!!!! I simply could not understand... Is this normal in Kaggle competitions? Is this specially so in this competition? (am new to kaggle competitions). This is about a student using a GUI based class room game. Where is the question of creating thousands of features?! If that's for fourth decimal of the submission score, it might not benefit much to the competition host!!\n\n",
    "2301878": "I doubt you will get 0.95 just by using transformer, or any other model. Some of the questions are too random to obtain such a score. How exactly did you evaluate your model to reach such an insane score?"
  }
}