{
  "id": 550668,
  "title": "Importance and absence of accelerometer data, and feature engineering extra columns.",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/550668",
  "author_name": "",
  "post_date": "2024-12-08T21:03:00.239235900Z",
  "votes": 3,
  "comment_count": 3,
  "views": 0,
  "content": "<p>Hello everyone, this is my first kaggle competition I am participating in. </p>\n<p>Seeing we have 2 different datasets, I tried training models only on the train dataset, trying to get to a maximum score before including the accelerometer data. The max score i managed to reach was 0.432. This is, using KNN imputation on missing train data, doing nothing on imputing test data, and not dropping any features, not even seasons. (only dropping features that dont exist in given test data.) </p>\n<p>Now on the accelerometer data. I noticed on the train data about 1/4 of entries have accelerometer data, and on the test data only 1/10. We can assume that in the hidden test data, about 1/10 - 1/4 (10%-25%) entries will have accelerometer data. </p>\n<p>Now, I tried some techniques on the accelerometer data, using an auto-encoder on the data.describe() pandas function, giving us a table of some statistics on each entries accelerometer data. Then, merging the encoded data to my regular dataset, and training as before. The QWK score didnt change at all.</p>\n<p>The models I use are XGB, LGBM and CatBoost.</p>\n<p>Looking at other public notebooks with better scores, (which are all almost exactly same) I couldnt help but notice the addition of another ~10 extra features which are non-linearly derived from given features, which are<br>\n<code>#BMI\n    df['BMI_Internet_Hours'] = df['Physical-BMI'] * df['PreInt_EduHx-computerinternet_hoursday']\n    df['BFP_BMI'] = df['BIA-BIA_Fat'] / df['BIA-BIA_BMI']\n    df['FFMI_BFP'] = df['BIA-BIA_FFMI'] / df['BIA-BIA_Fat']\n    df['FMI_BFP'] = df['BIA-BIA_FMI'] / df['BIA-BIA_Fat']\n    df['LST_TBW'] = df['BIA-BIA_LST'] / df['BIA-BIA_TBW']\n    df['BFP_BMR'] = df['BIA-BIA_Fat'] * df['BIA-BIA_BMR']\n    df['BFP_DEE'] = df['BIA-BIA_Fat'] * df['BIA-BIA_DEE']\n    df['BMR_Weight'] = df['BIA-BIA_BMR'] / df['Physical-Weight']\n    df['DEE_Weight'] = df['BIA-BIA_DEE'] / df['Physical-Weight']\n    df['SMM_Height'] = df['BIA-BIA_SMM'] / df['Physical-Height']\n    df['Muscle_to_Fat'] = df['BIA-BIA_SMM'] / df['BIA-BIA_FMI']\n    df['Hydration_Status'] = df['BIA-BIA_TBW'] / df['Physical-Weight']\n    df['ICW_TBW'] = df['BIA-BIA_ICW'] / df['BIA-BIA_TBW']</code><br>\nWhen I added that as well, my QWK went up about 0.01, taking me to about 0.442.</p>\n<p>Now, my questions are:</p>\n<p>-Why do we merge the accelerometer data, instead of training 2 voting regressors on different datasets? For example, given a hidden test set, we can see which of the hidden entries have accelerometer data (by id) and predict based on accelerometer data, which in theory should give better predictions. (I haven't tried this yet, I am out of daily submissions.)</p>\n<p>-Why add these extra columns, and how do they increase the score? I spent a lot of time doing analysis and demographics, and these seem to be the most highly-correlated features, but, when multiplying - dividing them and then creating new features, which aren't really new, how does that increase performance? Also, since we are dividing, unless we replace np.inf values with 0 or NaN we have a problem, and that problem is very common, since in the test dataset there is a ton of data missing, so its very likely that these new features will be inf, or undefined.</p>\n<p>-How come no one talks about missing data imputation much on the \"top\" published note-books? Especially on the test set? It should be obvious that a model can't make predicitons on non existing data, you can't train a model on 60 features and then expect it to make a prediciton on 2 with the rest being NaN. The missing data in the test set seems to be the bigger problem here. In train set, missing data can be imputed, I tried KNN imputing and Mean imputing with not a big difference, but in test data there isn't really a point to that, and I believe this shoudl be discussed more. In my code, i let the models handle the missing NaN data them selves, using Default Direction, since they are all decision trees.</p>\n<p>-Model ensembling: Most people use the same models as I, but also use TabNet. However, when it comes to ensembling, they use a voting regressor, but give 3 submissions on 3 different voting regressors, and then take the majority voting of those 3 submissions table. Why not just make 1 submission instead? This doesn't make sense to me. In my case, training 3 models, say a, b, c, instead of making a regressor for all 3 of them and giving predictions, it's like im doing: sub 1 (a and b only) sub 2( b and c only) sub 3 ( a and c only ) final sub: majority vote of sub 1 ,2, 3. If anything, this just makes the outcome more hard to interpret and find the model with the actual best QWK score. </p>\n<p>-The SEED: When changing the SEED from 42 to 41, my qwk score went up 0.01, from 0.430-0.440. Why is that? Are the predictions truly that sensitive to randomness? If so, does that mean that, say i get to 0.49, i just mess around with the SEED to get a better score? I read a lot of people finding this problem, especially with the TabNet, which I haven't used. </p>\n<p>-Finally, since I am new to kaggle, by taking part in this competition, I have learned a lot and I have read a ton of notebooks with scores better than mine. As i mentioned, i noticed almost all of them have extremely similar code. I my self copied the autoencoder code and the train ML function which is really general, and would probably find it online somewhere, but the rest I researched and wrote my self. It seems, to me at least, 90% of the notebooks are forks of a \"good\" notebook and play with the seed or weights on voting regressor hoping to get a better score, therefore \"earning\" a medal. Is this normal? Is this how kaggle competitions are? I could just fork the top notebook i find and do that as well, but what's really the point here…</p>\n<p>Thank you very much for reading!</p>",
  "messages": [
    {
      "id": "3067082",
      "postDate": "12/08/2024 21:03:00",
      "content": "<p>Hello everyone, this is my first kaggle competition I am participating in. </p>\n<p>Seeing we have 2 different datasets, I tried training models only on the train dataset, trying to get to a maximum score before including the accelerometer data. The max score i managed to reach was 0.432. This is, using KNN imputation on missing train data, doing nothing on imputing test data, and not dropping any features, not even seasons. (only dropping features that dont exist in given test data.) </p>\n<p>Now on the accelerometer data. I noticed on the train data about 1/4 of entries have accelerometer data, and on the test data only 1/10. We can assume that in the hidden test data, about 1/10 - 1/4 (10%-25%) entries will have accelerometer data. </p>\n<p>Now, I tried some techniques on the accelerometer data, using an auto-encoder on the data.describe() pandas function, giving us a table of some statistics on each entries accelerometer data. Then, merging the encoded data to my regular dataset, and training as before. The QWK score didnt change at all.</p>\n<p>The models I use are XGB, LGBM and CatBoost.</p>\n<p>Looking at other public notebooks with better scores, (which are all almost exactly same) I couldnt help but notice the addition of another ~10 extra features which are non-linearly derived from given features, which are<br>\n<code>#BMI\n    df['BMI_Internet_Hours'] = df['Physical-BMI'] * df['PreInt_EduHx-computerinternet_hoursday']\n    df['BFP_BMI'] = df['BIA-BIA_Fat'] / df['BIA-BIA_BMI']\n    df['FFMI_BFP'] = df['BIA-BIA_FFMI'] / df['BIA-BIA_Fat']\n    df['FMI_BFP'] = df['BIA-BIA_FMI'] / df['BIA-BIA_Fat']\n    df['LST_TBW'] = df['BIA-BIA_LST'] / df['BIA-BIA_TBW']\n    df['BFP_BMR'] = df['BIA-BIA_Fat'] * df['BIA-BIA_BMR']\n    df['BFP_DEE'] = df['BIA-BIA_Fat'] * df['BIA-BIA_DEE']\n    df['BMR_Weight'] = df['BIA-BIA_BMR'] / df['Physical-Weight']\n    df['DEE_Weight'] = df['BIA-BIA_DEE'] / df['Physical-Weight']\n    df['SMM_Height'] = df['BIA-BIA_SMM'] / df['Physical-Height']\n    df['Muscle_to_Fat'] = df['BIA-BIA_SMM'] / df['BIA-BIA_FMI']\n    df['Hydration_Status'] = df['BIA-BIA_TBW'] / df['Physical-Weight']\n    df['ICW_TBW'] = df['BIA-BIA_ICW'] / df['BIA-BIA_TBW']</code><br>\nWhen I added that as well, my QWK went up about 0.01, taking me to about 0.442.</p>\n<p>Now, my questions are:</p>\n<p>-Why do we merge the accelerometer data, instead of training 2 voting regressors on different datasets? For example, given a hidden test set, we can see which of the hidden entries have accelerometer data (by id) and predict based on accelerometer data, which in theory should give better predictions. (I haven't tried this yet, I am out of daily submissions.)</p>\n<p>-Why add these extra columns, and how do they increase the score? I spent a lot of time doing analysis and demographics, and these seem to be the most highly-correlated features, but, when multiplying - dividing them and then creating new features, which aren't really new, how does that increase performance? Also, since we are dividing, unless we replace np.inf values with 0 or NaN we have a problem, and that problem is very common, since in the test dataset there is a ton of data missing, so its very likely that these new features will be inf, or undefined.</p>\n<p>-How come no one talks about missing data imputation much on the \"top\" published note-books? Especially on the test set? It should be obvious that a model can't make predicitons on non existing data, you can't train a model on 60 features and then expect it to make a prediciton on 2 with the rest being NaN. The missing data in the test set seems to be the bigger problem here. In train set, missing data can be imputed, I tried KNN imputing and Mean imputing with not a big difference, but in test data there isn't really a point to that, and I believe this shoudl be discussed more. In my code, i let the models handle the missing NaN data them selves, using Default Direction, since they are all decision trees.</p>\n<p>-Model ensembling: Most people use the same models as I, but also use TabNet. However, when it comes to ensembling, they use a voting regressor, but give 3 submissions on 3 different voting regressors, and then take the majority voting of those 3 submissions table. Why not just make 1 submission instead? This doesn't make sense to me. In my case, training 3 models, say a, b, c, instead of making a regressor for all 3 of them and giving predictions, it's like im doing: sub 1 (a and b only) sub 2( b and c only) sub 3 ( a and c only ) final sub: majority vote of sub 1 ,2, 3. If anything, this just makes the outcome more hard to interpret and find the model with the actual best QWK score. </p>\n<p>-The SEED: When changing the SEED from 42 to 41, my qwk score went up 0.01, from 0.430-0.440. Why is that? Are the predictions truly that sensitive to randomness? If so, does that mean that, say i get to 0.49, i just mess around with the SEED to get a better score? I read a lot of people finding this problem, especially with the TabNet, which I haven't used. </p>\n<p>-Finally, since I am new to kaggle, by taking part in this competition, I have learned a lot and I have read a ton of notebooks with scores better than mine. As i mentioned, i noticed almost all of them have extremely similar code. I my self copied the autoencoder code and the train ML function which is really general, and would probably find it online somewhere, but the rest I researched and wrote my self. It seems, to me at least, 90% of the notebooks are forks of a \"good\" notebook and play with the seed or weights on voting regressor hoping to get a better score, therefore \"earning\" a medal. Is this normal? Is this how kaggle competitions are? I could just fork the top notebook i find and do that as well, but what's really the point here…</p>\n<p>Thank you very much for reading!</p>",
      "rawMarkdown": "Hello everyone, this is my first kaggle competition I am participating in. \n\nSeeing we have 2 different datasets, I tried training models only on the train dataset, trying to get to a maximum score before including the accelerometer data. The max score i managed to reach was 0.432. This is, using KNN imputation on missing train data, doing nothing on imputing test data, and not dropping any features, not even seasons. (only dropping features that dont exist in given test data.) \n\nNow on the accelerometer data. I noticed on the train data about 1/4 of entries have accelerometer data, and on the test data only 1/10. We can assume that in the hidden test data, about 1/10 - 1/4 (10%-25%) entries will have accelerometer data. \n\nNow, I tried some techniques on the accelerometer data, using an auto-encoder on the data.describe() pandas function, giving us a table of some statistics on each entries accelerometer data. Then, merging the encoded data to my regular dataset, and training as before. The QWK score didnt change at all.\n\nThe models I use are XGB, LGBM and CatBoost.\n\nLooking at other public notebooks with better scores, (which are all almost exactly same) I couldnt help but notice the addition of another ~10 extra features which are non-linearly derived from given features, which are\n`#BMI\n    df['BMI_Internet_Hours'] = df['Physical-BMI'] * df['PreInt_EduHx-computerinternet_hoursday']\n    df['BFP_BMI'] = df['BIA-BIA_Fat'] / df['BIA-BIA_BMI']\n    df['FFMI_BFP'] = df['BIA-BIA_FFMI'] / df['BIA-BIA_Fat']\n    df['FMI_BFP'] = df['BIA-BIA_FMI'] / df['BIA-BIA_Fat']\n    df['LST_TBW'] = df['BIA-BIA_LST'] / df['BIA-BIA_TBW']\n    df['BFP_BMR'] = df['BIA-BIA_Fat'] * df['BIA-BIA_BMR']\n    df['BFP_DEE'] = df['BIA-BIA_Fat'] * df['BIA-BIA_DEE']\n    df['BMR_Weight'] = df['BIA-BIA_BMR'] / df['Physical-Weight']\n    df['DEE_Weight'] = df['BIA-BIA_DEE'] / df['Physical-Weight']\n    df['SMM_Height'] = df['BIA-BIA_SMM'] / df['Physical-Height']\n    df['Muscle_to_Fat'] = df['BIA-BIA_SMM'] / df['BIA-BIA_FMI']\n    df['Hydration_Status'] = df['BIA-BIA_TBW'] / df['Physical-Weight']\n    df['ICW_TBW'] = df['BIA-BIA_ICW'] / df['BIA-BIA_TBW']`\nWhen I added that as well, my QWK went up about 0.01, taking me to about 0.442.\n\nNow, my questions are:\n\n-Why do we merge the accelerometer data, instead of training 2 voting regressors on different datasets? For example, given a hidden test set, we can see which of the hidden entries have accelerometer data (by id) and predict based on accelerometer data, which in theory should give better predictions. (I haven't tried this yet, I am out of daily submissions.)\n\n-Why add these extra columns, and how do they increase the score? I spent a lot of time doing analysis and demographics, and these seem to be the most highly-correlated features, but, when multiplying - dividing them and then creating new features, which aren't really new, how does that increase performance? Also, since we are dividing, unless we replace np.inf values with 0 or NaN we have a problem, and that problem is very common, since in the test dataset there is a ton of data missing, so its very likely that these new features will be inf, or undefined.\n\n-How come no one talks about missing data imputation much on the \"top\" published note-books? Especially on the test set? It should be obvious that a model can't make predicitons on non existing data, you can't train a model on 60 features and then expect it to make a prediciton on 2 with the rest being NaN. The missing data in the test set seems to be the bigger problem here. In train set, missing data can be imputed, I tried KNN imputing and Mean imputing with not a big difference, but in test data there isn't really a point to that, and I believe this shoudl be discussed more. In my code, i let the models handle the missing NaN data them selves, using Default Direction, since they are all decision trees.\n\n-Model ensembling: Most people use the same models as I, but also use TabNet. However, when it comes to ensembling, they use a voting regressor, but give 3 submissions on 3 different voting regressors, and then take the majority voting of those 3 submissions table. Why not just make 1 submission instead? This doesn't make sense to me. In my case, training 3 models, say a, b, c, instead of making a regressor for all 3 of them and giving predictions, it's like im doing: sub 1 (a and b only) sub 2( b and c only) sub 3 ( a and c only ) final sub: majority vote of sub 1 ,2, 3. If anything, this just makes the outcome more hard to interpret and find the model with the actual best QWK score. \n\n-The SEED: When changing the SEED from 42 to 41, my qwk score went up 0.01, from 0.430-0.440. Why is that? Are the predictions truly that sensitive to randomness? If so, does that mean that, say i get to 0.49, i just mess around with the SEED to get a better score? I read a lot of people finding this problem, especially with the TabNet, which I haven't used. \n\n\n-Finally, since I am new to kaggle, by taking part in this competition, I have learned a lot and I have read a ton of notebooks with scores better than mine. As i mentioned, i noticed almost all of them have extremely similar code. I my self copied the autoencoder code and the train ML function which is really general, and would probably find it online somewhere, but the rest I researched and wrote my self. It seems, to me at least, 90% of the notebooks are forks of a \"good\" notebook and play with the seed or weights on voting regressor hoping to get a better score, therefore \"earning\" a medal. Is this normal? Is this how kaggle competitions are? I could just fork the top notebook i find and do that as well, but what's really the point here...\n\nThank you very much for reading!",
      "votes": null
    },
    {
      "id": "3067653",
      "postDate": "12/09/2024 14:04:54",
      "content": "<p>It's not that simple. You can copy a public notebook and even get a bronze medal. And what next? What will it give if there is no knowledge? The main thing is who has what goal here. We came here to hone our knowledge and skills. And it's not that simple. Secondly, there are overtrained models, there are errors in public notebooks, sometimes your own mistakes go unnoticed, and the LB often does not reflect the real possible place in the leaderboard. By opening notebooks, many learn from this, discussing different ideas and solution techniques. The main thing is not to open notebooks with a big result a week before the end of the competition.</p>",
      "rawMarkdown": "It's not that simple. You can copy a public notebook and even get a bronze medal. And what next? What will it give if there is no knowledge? The main thing is who has what goal here. We came here to hone our knowledge and skills. And it's not that simple. Secondly, there are overtrained models, there are errors in public notebooks, sometimes your own mistakes go unnoticed, and the LB often does not reflect the real possible place in the leaderboard. By opening notebooks, many learn from this, discussing different ideas and solution techniques. The main thing is not to open notebooks with a big result a week before the end of the competition.",
      "votes": null
    },
    {
      "id": "3067860",
      "postDate": "12/09/2024 17:37:45",
      "content": "<p>Exactly! I think a lot of people just copy and submit notebooks to get more medals on their profiles, however in that way, there should be 10-20 actual submissions in the 0.490-0.5 range, meanwhile there are literally hundreds with the same 0.492 score. I understand that the score converges at that point, but when a notebook with 0.492 score is published it is a bit weird to see 5 notebooks with 0.495-0.500, and 100 it 0.492 specifically. I think you get my point here.</p>\n<p>Never the less, you are absolutely right about having to set goals before entering a competition. I have learned a ton reading other notebooks, and don't really care about my public QWK. Thank you for your reply!</p>",
      "rawMarkdown": "Exactly! I think a lot of people just copy and submit notebooks to get more medals on their profiles, however in that way, there should be 10-20 actual submissions in the 0.490-0.5 range, meanwhile there are literally hundreds with the same 0.492 score. I understand that the score converges at that point, but when a notebook with 0.492 score is published it is a bit weird to see 5 notebooks with 0.495-0.500, and 100 it 0.492 specifically. I think you get my point here.\n\nNever the less, you are absolutely right about having to set goals before entering a competition. I have learned a ton reading other notebooks, and don't really care about my public QWK. Thank you for your reply!",
      "votes": null
    },
    {
      "id": "3067907",
      "postDate": "12/09/2024 18:31:42",
      "content": "<p>My goal is to do better than all .492 and .494 notebooks, on private leaderboard.<br>\nThese notebooks depend on a seed number: when you change the seed number they lose their high score. They are very fragile and will lose their high score when the test dataset changes on the 19th.</p>",
      "rawMarkdown": "My goal is to do better than all .492 and .494 notebooks, on private leaderboard.\nThese notebooks depend on a seed number: when you change the seed number they lose their high score. They are very fragile and will lose their high score when the test dataset changes on the 19th.",
      "votes": null
    }
  ],
  "comments": [
    {
      "id": 3067653,
      "author_name": "konstantinboyko",
      "author_url": "",
      "post_date": "12/09/2024 14:04:54",
      "content": "<p>It's not that simple. You can copy a public notebook and even get a bronze medal. And what next? What will it give if there is no knowledge? The main thing is who has what goal here. We came here to hone our knowledge and skills. And it's not that simple. Secondly, there are overtrained models, there are errors in public notebooks, sometimes your own mistakes go unnoticed, and the LB often does not reflect the real possible place in the leaderboard. By opening notebooks, many learn from this, discussing different ideas and solution techniques. The main thing is not to open notebooks with a big result a week before the end of the competition.</p>",
      "votes": null,
      "replies": [
        {
          "id": 3067860,
          "author_name": "kw5t45",
          "author_url": "",
          "post_date": "12/09/2024 17:37:45",
          "content": "<p>Exactly! I think a lot of people just copy and submit notebooks to get more medals on their profiles, however in that way, there should be 10-20 actual submissions in the 0.490-0.5 range, meanwhile there are literally hundreds with the same 0.492 score. I understand that the score converges at that point, but when a notebook with 0.492 score is published it is a bit weird to see 5 notebooks with 0.495-0.500, and 100 it 0.492 specifically. I think you get my point here.</p>\n<p>Never the less, you are absolutely right about having to set goals before entering a competition. I have learned a ton reading other notebooks, and don't really care about my public QWK. Thank you for your reply!</p>",
          "votes": null,
          "replies": [
            {
              "id": 3067907,
              "author_name": "adaubas",
              "author_url": "",
              "post_date": "12/09/2024 18:31:42",
              "content": "<p>My goal is to do better than all .492 and .494 notebooks, on private leaderboard.<br>\nThese notebooks depend on a seed number: when you change the seed number they lose their high score. They are very fragile and will lose their high score when the test dataset changes on the 19th.</p>",
              "votes": null,
              "replies": []
            }
          ]
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3067082": "Hello everyone, this is my first kaggle competition I am participating in. \n\nSeeing we have 2 different datasets, I tried training models only on the train dataset, trying to get to a maximum score before including the accelerometer data. The max score i managed to reach was 0.432. This is, using KNN imputation on missing train data, doing nothing on imputing test data, and not dropping any features, not even seasons. (only dropping features that dont exist in given test data.) \n\nNow on the accelerometer data. I noticed on the train data about 1/4 of entries have accelerometer data, and on the test data only 1/10. We can assume that in the hidden test data, about 1/10 - 1/4 (10%-25%) entries will have accelerometer data. \n\nNow, I tried some techniques on the accelerometer data, using an auto-encoder on the data.describe() pandas function, giving us a table of some statistics on each entries accelerometer data. Then, merging the encoded data to my regular dataset, and training as before. The QWK score didnt change at all.\n\nThe models I use are XGB, LGBM and CatBoost.\n\nLooking at other public notebooks with better scores, (which are all almost exactly same) I couldnt help but notice the addition of another ~10 extra features which are non-linearly derived from given features, which are\n`#BMI\n    df['BMI_Internet_Hours'] = df['Physical-BMI'] * df['PreInt_EduHx-computerinternet_hoursday']\n    df['BFP_BMI'] = df['BIA-BIA_Fat'] / df['BIA-BIA_BMI']\n    df['FFMI_BFP'] = df['BIA-BIA_FFMI'] / df['BIA-BIA_Fat']\n    df['FMI_BFP'] = df['BIA-BIA_FMI'] / df['BIA-BIA_Fat']\n    df['LST_TBW'] = df['BIA-BIA_LST'] / df['BIA-BIA_TBW']\n    df['BFP_BMR'] = df['BIA-BIA_Fat'] * df['BIA-BIA_BMR']\n    df['BFP_DEE'] = df['BIA-BIA_Fat'] * df['BIA-BIA_DEE']\n    df['BMR_Weight'] = df['BIA-BIA_BMR'] / df['Physical-Weight']\n    df['DEE_Weight'] = df['BIA-BIA_DEE'] / df['Physical-Weight']\n    df['SMM_Height'] = df['BIA-BIA_SMM'] / df['Physical-Height']\n    df['Muscle_to_Fat'] = df['BIA-BIA_SMM'] / df['BIA-BIA_FMI']\n    df['Hydration_Status'] = df['BIA-BIA_TBW'] / df['Physical-Weight']\n    df['ICW_TBW'] = df['BIA-BIA_ICW'] / df['BIA-BIA_TBW']`\nWhen I added that as well, my QWK went up about 0.01, taking me to about 0.442.\n\nNow, my questions are:\n\n-Why do we merge the accelerometer data, instead of training 2 voting regressors on different datasets? For example, given a hidden test set, we can see which of the hidden entries have accelerometer data (by id) and predict based on accelerometer data, which in theory should give better predictions. (I haven't tried this yet, I am out of daily submissions.)\n\n-Why add these extra columns, and how do they increase the score? I spent a lot of time doing analysis and demographics, and these seem to be the most highly-correlated features, but, when multiplying - dividing them and then creating new features, which aren't really new, how does that increase performance? Also, since we are dividing, unless we replace np.inf values with 0 or NaN we have a problem, and that problem is very common, since in the test dataset there is a ton of data missing, so its very likely that these new features will be inf, or undefined.\n\n-How come no one talks about missing data imputation much on the \"top\" published note-books? Especially on the test set? It should be obvious that a model can't make predicitons on non existing data, you can't train a model on 60 features and then expect it to make a prediciton on 2 with the rest being NaN. The missing data in the test set seems to be the bigger problem here. In train set, missing data can be imputed, I tried KNN imputing and Mean imputing with not a big difference, but in test data there isn't really a point to that, and I believe this shoudl be discussed more. In my code, i let the models handle the missing NaN data them selves, using Default Direction, since they are all decision trees.\n\n-Model ensembling: Most people use the same models as I, but also use TabNet. However, when it comes to ensembling, they use a voting regressor, but give 3 submissions on 3 different voting regressors, and then take the majority voting of those 3 submissions table. Why not just make 1 submission instead? This doesn't make sense to me. In my case, training 3 models, say a, b, c, instead of making a regressor for all 3 of them and giving predictions, it's like im doing: sub 1 (a and b only) sub 2( b and c only) sub 3 ( a and c only ) final sub: majority vote of sub 1 ,2, 3. If anything, this just makes the outcome more hard to interpret and find the model with the actual best QWK score. \n\n-The SEED: When changing the SEED from 42 to 41, my qwk score went up 0.01, from 0.430-0.440. Why is that? Are the predictions truly that sensitive to randomness? If so, does that mean that, say i get to 0.49, i just mess around with the SEED to get a better score? I read a lot of people finding this problem, especially with the TabNet, which I haven't used. \n\n\n-Finally, since I am new to kaggle, by taking part in this competition, I have learned a lot and I have read a ton of notebooks with scores better than mine. As i mentioned, i noticed almost all of them have extremely similar code. I my self copied the autoencoder code and the train ML function which is really general, and would probably find it online somewhere, but the rest I researched and wrote my self. It seems, to me at least, 90% of the notebooks are forks of a \"good\" notebook and play with the seed or weights on voting regressor hoping to get a better score, therefore \"earning\" a medal. Is this normal? Is this how kaggle competitions are? I could just fork the top notebook i find and do that as well, but what's really the point here...\n\nThank you very much for reading!",
    "3067653": "It's not that simple. You can copy a public notebook and even get a bronze medal. And what next? What will it give if there is no knowledge? The main thing is who has what goal here. We came here to hone our knowledge and skills. And it's not that simple. Secondly, there are overtrained models, there are errors in public notebooks, sometimes your own mistakes go unnoticed, and the LB often does not reflect the real possible place in the leaderboard. By opening notebooks, many learn from this, discussing different ideas and solution techniques. The main thing is not to open notebooks with a big result a week before the end of the competition.",
    "3067860": "Exactly! I think a lot of people just copy and submit notebooks to get more medals on their profiles, however in that way, there should be 10-20 actual submissions in the 0.490-0.5 range, meanwhile there are literally hundreds with the same 0.492 score. I understand that the score converges at that point, but when a notebook with 0.492 score is published it is a bit weird to see 5 notebooks with 0.495-0.500, and 100 it 0.492 specifically. I think you get my point here.\n\nNever the less, you are absolutely right about having to set goals before entering a competition. I have learned a ton reading other notebooks, and don't really care about my public QWK. Thank you for your reply!",
    "3067907": "My goal is to do better than all .492 and .494 notebooks, on private leaderboard.\nThese notebooks depend on a seed number: when you change the seed number they lose their high score. They are very fragile and will lose their high score when the test dataset changes on the 19th."
  },
  "source": "meta"
}