{
  "id": 542244,
  "title": "Various strategies.",
  "url": "/competitions/child-mind-institute-problematic-internet-use/discussion/542244",
  "author_name": "Catadanna",
  "post_date": "2024-10-23T18:28:47.627000",
  "votes": 15,
  "comment_count": 23,
  "views": 0,
  "content": "<p>Hallo</p>\n<p>Here are various strategies I tried.</p>\n<ol>\n<li><p>I used different models only with the train/test data. <br>\n1.1 There are missing values in the labels so I tried to train only with the non null labels first. <br>\n1.2 I also tried to predict the values for the missing labels, using various classifiers, but that did not provide a better result.</p></li>\n<li><p>I tried using both train/test data merged with the parquet data. <br>\nNow, the parquet data contains several lines for one sample in train/test. <br>\n2.1 Sometimes, the data for one sample in train/test is missing in the parquet data. So one could choose either to delete these samples for training or to make predictions for it. I tried both. <br>\n2.2 I tried aggregations by sample for the information in the parquet data, only for numeric features. For each of them I created new aggregated features for mean, median, q1 and q3. These are the new features I added to the existing train features. For the test features, the test set is hidden; in the case of missing features in parquet files for a given sample, I replaced with the mean parsed for the respective feature in the train set.</p></li>\n</ol>\n<p>The CV results for the dataset using train/test and parquet files did not give a better result. Best result was given by the basic train set with only the non null labels. I used different algorithms: neural networks and gradient boosting.</p>\n<ol>\n<li><p>I divided the data in 5 splits, recorded the spits and worked only on them, in order to train/test on the same dataset and be able to compare correctly CV results obtained with different algorithms.</p></li>\n<li><p>I used Optuna in order to get the best parameters.</p></li>\n</ol>\n<p>Feel free to comment, any new idea or remark is welcome!  </p>",
  "messages": [
    {
      "id": 3026385,
      "postDate": "2024-10-23T18:28:47.627Z",
      "content": "<p>Hallo</p>\n<p>Here are various strategies I tried.</p>\n<ol>\n<li><p>I used different models only with the train/test data. <br>\n1.1 There are missing values in the labels so I tried to train only with the non null labels first. <br>\n1.2 I also tried to predict the values for the missing labels, using various classifiers, but that did not provide a better result.</p></li>\n<li><p>I tried using both train/test data merged with the parquet data. <br>\nNow, the parquet data contains several lines for one sample in train/test. <br>\n2.1 Sometimes, the data for one sample in train/test is missing in the parquet data. So one could choose either to delete these samples for training or to make predictions for it. I tried both. <br>\n2.2 I tried aggregations by sample for the information in the parquet data, only for numeric features. For each of them I created new aggregated features for mean, median, q1 and q3. These are the new features I added to the existing train features. For the test features, the test set is hidden; in the case of missing features in parquet files for a given sample, I replaced with the mean parsed for the respective feature in the train set.</p></li>\n</ol>\n<p>The CV results for the dataset using train/test and parquet files did not give a better result. Best result was given by the basic train set with only the non null labels. I used different algorithms: neural networks and gradient boosting.</p>\n<ol>\n<li><p>I divided the data in 5 splits, recorded the spits and worked only on them, in order to train/test on the same dataset and be able to compare correctly CV results obtained with different algorithms.</p></li>\n<li><p>I used Optuna in order to get the best parameters.</p></li>\n</ol>\n<p>Feel free to comment, any new idea or remark is welcome!  </p>",
      "rawMarkdown": "Hallo\n\nHere are various strategies I tried.\n\n1. I used different models only with the train/test data. \n1.1 There are missing values in the labels so I tried to train only with the non null labels first. \n1.2 I also tried to predict the values for the missing labels, using various classifiers, but that did not provide a better result.\n\n2. I tried using both train/test data merged with the parquet data. \nNow, the parquet data contains several lines for one sample in train/test. \n2.1 Sometimes, the data for one sample in train/test is missing in the parquet data. So one could choose either to delete these samples for training or to make predictions for it. I tried both. \n2.2 I tried aggregations by sample for the information in the parquet data, only for numeric features. For each of them I created new aggregated features for mean, median, q1 and q3. These are the new features I added to the existing train features. For the test features, the test set is hidden; in the case of missing features in parquet files for a given sample, I replaced with the mean parsed for the respective feature in the train set.\n\nThe CV results for the dataset using train/test and parquet files did not give a better result. Best result was given by the basic train set with only the non null labels. I used different algorithms: neural networks and gradient boosting.\n\n3. I divided the data in 5 splits, recorded the spits and worked only on them, in order to train/test on the same dataset and be able to compare correctly CV results obtained with different algorithms.\n\n4. I used Optuna in order to get the best parameters.\n\n\n\nFeel free to comment, any new idea or remark is welcome!  ",
      "votes": 13
    },
    {
      "id": 3027396,
      "postDate": "2024-10-24T18:45:37.173Z",
      "content": "<p>I've tried most of your methods.<br>\nMy experience is that parquet data is giving some useful data so we can't drop it. However, most of them are rubbish that you know. (e.g. the 1st one is count, anything that can give to model is noise. May be X, Y, Z etc will provide some useful insight to model, but I won't impute the parquet data)</p>\n<p>I've tried adding features/ dropping features<br>\nI've tried oversampling, undersampling<br>\nI've tried custom scoring!<br>\nI've tried SHAP analysis!<br>\nI've tried KNN imputation!<br>\nI've tried binary classification!<br>\nI've tried pseudo labelling!</p>\n<p>The last thing I haven't tried is copy the top score notebook and make a submission.<br>\nThere's a fatal mistake in some of the top score notebooks, I guess, that I can't say it here, that keeps me from directly copying from them.</p>\n<p>My suggestion is to add a confusion matrix and this gives you a better sense of your model, even though it's not how it's scored.</p>\n<p>Try dropping rows without Sii + LGBM + kfold + Optuna 30 trials + optimizer + a bit of data cleaning (e.g. cap max for outlier), it shall take you to 0.43-0.45.</p>\n<p>Good luck!</p>",
      "rawMarkdown": "I've tried most of your methods.\nMy experience is that parquet data is giving some useful data so we can't drop it. However, most of them are rubbish that you know. (e.g. the 1st one is count, anything that can give to model is noise. May be X, Y, Z etc will provide some useful insight to model, but I won't impute the parquet data)\n\nI've tried adding features/ dropping features\nI've tried oversampling, undersampling\nI've tried custom scoring!\nI've tried SHAP analysis!\nI've tried KNN imputation!\nI've tried binary classification!\nI've tried pseudo labelling!\n\nThe last thing I haven't tried is copy the top score notebook and make a submission.\nThere's a fatal mistake in some of the top score notebooks, I guess, that I can't say it here, that keeps me from directly copying from them.\n\nMy suggestion is to add a confusion matrix and this gives you a better sense of your model, even though it's not how it's scored.\n\nTry dropping rows without Sii + LGBM + kfold + Optuna 30 trials + optimizer + a bit of data cleaning (e.g. cap max for outlier), it shall take you to 0.43-0.45.\n\nGood luck!",
      "votes": 5,
      "replies": [
        {
          "id": 3027412,
          "postDate": "2024-10-24T19:11:45.227Z",
          "content": "<p>What surprises me about the top scoring (shared) notebooks is that people are just copying them and don't analyze the code, because the same mistake is repeating in all of them, for example how do they use the autoencoder is just not correct and maybe by some luck it gives a good score in LB (at least in the public LB) but they don't see that the original author made a mistake in the method how it's applied and everybody just blindly use it.</p>",
          "rawMarkdown": "What surprises me about the top scoring (shared) notebooks is that people are just copying them and don't analyze the code, because the same mistake is repeating in all of them, for example how do they use the autoencoder is just not correct and maybe by some luck it gives a good score in LB (at least in the public LB) but they don't see that the original author made a mistake in the method how it's applied and everybody just blindly use it.",
          "votes": 8,
          "replies": [
            {
              "id": 3027443,
              "postDate": "2024-10-24T20:11:41.773Z",
              "content": "<p>Let me name one as well here then 🤣. Using KNN imputer to fill sii in train set is not convincing me yet.</p>",
              "rawMarkdown": "Let me name one as well here then 🤣. Using KNN imputer to fill sii in train set is not convincing me yet.",
              "votes": 3
            },
            {
              "id": 3027702,
              "postDate": "2024-10-25T06:58:41.143Z",
              "content": "<p>Absolutely agree and I have an open discussion about this <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/539536\" target=\"_blank\">here </a></p>",
              "rawMarkdown": "Absolutely agree and I have an open discussion about this [here ](https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/539536)"
            },
            {
              "id": 3031086,
              "postDate": "2024-10-29T10:10:36.900Z",
              "content": "<p>If people compete with some other's notebook and, say, they win, then they must reproduce the same high score, and that might not work -- especially if they copied without understanding. </p>",
              "rawMarkdown": "If people compete with some other's notebook and, say, they win, then they must reproduce the same high score, and that might not work -- especially if they copied without understanding. "
            },
            {
              "id": 3031576,
              "postDate": "2024-10-29T20:57:07.883Z",
              "content": "<p>Oh yes, I noticed that aswell and I'm confused about how persistent this incorrect use of the autoencoder is.😬 </p>",
              "rawMarkdown": "Oh yes, I noticed that aswell and I'm confused about how persistent this incorrect use of the autoencoder is.😬 ",
              "votes": 1
            }
          ]
        },
        {
          "id": 3027445,
          "postDate": "2024-10-24T20:13:59.587Z",
          "content": "<p>I made the confusion matrix, there are some features which are highly correlated.<br>\nAs I already mentioned, I already dropped lines without label.<br>\nI am against copying high score notebooks and compete ith them.</p>",
          "rawMarkdown": "I made the confusion matrix, there are some features which are highly correlated.\nAs I already mentioned, I already dropped lines without label.\nI am against copying high score notebooks and compete ith them.",
          "votes": 1
        },
        {
          "id": 3029295,
          "postDate": "2024-10-27T05:46:48.590Z",
          "content": "<p>Can not agree more.</p>\n<p>I tried many methods and derived more than 30000 features (without using AutoEncoder) and got 0.45-0.46.<br>\nI also tried to copy AutoEncoder function to used with my owner feature engineering and not get better LB score.</p>",
          "rawMarkdown": "Can not agree more.\n\nI tried many methods and derived more than 30000 features (without using AutoEncoder) and got 0.45-0.46.\nI also tried to copy AutoEncoder function to used with my owner feature engineering and not get better LB score."
        },
        {
          "id": 3031131,
          "postDate": "2024-10-29T11:55:05.460Z",
          "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> <a href=\"https://www.kaggle.com/ggggpeushmy\" target=\"_blank\">@ggggpeushmy</a> I'm so surprised there are still quite a few of us here in the lower rank.</p>\n<ol>\n<li>Feature engineering can be useful, in some scenarios I found removing useless feature seems to boost a little.</li>\n<li>The validation score is of course important, but I guess it's not as important as other games. <br>\nMy guess is that there're cases with high hours + high SDS + low physical score but sii =0, and <br>\nthere're cases sii=3 but 0 internet hours, normal BMI, low SDS.<br>\nIf we further push the cv score/ training score, it may overfit to these cases.<br>\nIt all depends if the final test set have these silly samples or not.</li>\n<li>Combining different models will definitely help, but I haven't done it yet.<br>\nGood luck everyone.</li>\n</ol>",
          "rawMarkdown": "@catadanna @eu1234 @ggggpeushmy I'm so surprised there are still quite a few of us here in the lower rank.\n\n1. Feature engineering can be useful, in some scenarios I found removing useless feature seems to boost a little.\n2. The validation score is of course important, but I guess it's not as important as other games. \nMy guess is that there're cases with high hours + high SDS + low physical score but sii =0, and \nthere're cases sii=3 but 0 internet hours, normal BMI, low SDS.\nIf we further push the cv score/ training score, it may overfit to these cases.\nIt all depends if the final test set have these silly samples or not.\n3. Combining different models will definitely help, but I haven't done it yet.\nGood luck everyone.",
          "votes": 1,
          "replies": [
            {
              "id": 3041769,
              "postDate": "2024-11-10T20:22:49.253Z",
              "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> : </p>\n<p>Low score does not mean anything for the moment, as these scores are parsed only on a small part  of the data.<br>\nValidation score is important, but one should clean the data correctly.<br>\nCombining predictions is good practice, because if one model gives a bad prediction for a sample, the others will give a good one, so the result will be smoothed and good in the end.</p>",
              "rawMarkdown": "@tomyuen : \n\nLow score does not mean anything for the moment, as these scores are parsed only on a small part  of the data.\nValidation score is important, but one should clean the data correctly.\nCombining predictions is good practice, because if one model gives a bad prediction for a sample, the others will give a good one, so the result will be smoothed and good in the end."
            }
          ]
        }
      ]
    },
    {
      "id": 3076800,
      "postDate": "2024-12-20T09:22:38.937Z",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a><br>\nI just come back and congratulate you both. Performance's not bad, right?! At least not falling under cliff after the shakeup.</p>",
      "rawMarkdown": "@catadanna @eu1234\nI just come back and congratulate you both. Performance's not bad, right?! At least not falling under cliff after the shakeup.",
      "votes": 1,
      "replies": [
        {
          "id": 3076817,
          "postDate": "2024-12-20T09:45:20.047Z",
          "content": "<p>Thanks! I had a small boost of +1649 positions up, but still missed to select my best scoring notebook that could bring me to a silver medal position 61. A good experience anyway. Best of luck in the future challenges!</p>",
          "rawMarkdown": "Thanks! I had a small boost of +1649 positions up, but still missed to select my best scoring notebook that could bring me to a silver medal position 61. A good experience anyway. Best of luck in the future challenges!"
        },
        {
          "id": 3078236,
          "postDate": "2024-12-22T01:34:05.757Z",
          "content": "<p>Thank you! It seems I did not overfit, or at least  I overfit less than others, as I climber about 1800 places up. I trusted my CV scores when choosing the submissions to compete with.</p>",
          "rawMarkdown": "Thank you! It seems I did not overfit, or at least  I overfit less than others, as I climber about 1800 places up. I trusted my CV scores when choosing the submissions to compete with."
        }
      ]
    },
    {
      "id": 3026738,
      "postDate": "2024-10-24T07:01:14.410Z",
      "content": "<p>In 2.2:</p>\n<ul>\n<li>try more aggregation types. </li>\n<li>you are imputing missing test values with overall train average, but you can impute the average from that person's values directly or even simpler - don't impute anything, even in the main data because the boosters are dealing just fine with missing data, or at least give it a try for testing.</li>\n</ul>\n<p>And the final - LB score is very sensitive to the random seed, as you can see the LB is overflooded with very close scores because maybe half of the people are using almost the same code that is publicly available and just re-runs it with different seeds or average the results of the same model but with different seeds.</p>\n<p>My advice - try to find your way to get a stable CV-LB score pairing, when them both grow hand by hand and to not have large oscilations between them.</p>",
      "rawMarkdown": "In 2.2:\n- try more aggregation types. \n- you are imputing missing test values with overall train average, but you can impute the average from that person's values directly or even simpler - don't impute anything, even in the main data because the boosters are dealing just fine with missing data, or at least give it a try for testing.\n\nAnd the final - LB score is very sensitive to the random seed, as you can see the LB is overflooded with very close scores because maybe half of the people are using almost the same code that is publicly available and just re-runs it with different seeds or average the results of the same model but with different seeds.\n\nMy advice - try to find your way to get a stable CV-LB score pairing, when them both grow hand by hand and to not have large oscilations between them.",
      "votes": 2,
      "replies": [
        {
          "id": 3027247,
          "postDate": "2024-10-24T15:56:44.290Z",
          "content": "<p>Thank you for your answer. I did not get a very good LB score for the moment. I shall try more aggregations (sum or other quantilles). And no, I do not wish to use someone else's code and compete with it, I do not think it is fair and does not give me any credit.</p>\n<p>Concerning the missing values : I assumed that some individuals in the test set do not have corresponding parquet information at all. That is why I replace with mean from test for each feature. I have no choice, test set is hidden. I may try to replace nan with one value, let us say -1 .</p>",
          "rawMarkdown": "Thank you for your answer. I did not get a very good LB score for the moment. I shall try more aggregations (sum or other quantilles). And no, I do not wish to use someone else's code and compete with it, I do not think it is fair and does not give me any credit.\n\nConcerning the missing values : I assumed that some individuals in the test set do not have corresponding parquet information at all. That is why I replace with mean from test for each feature. I have no choice, test set is hidden. I may try to replace nan with one value, let us say -1 .",
          "votes": 1,
          "replies": [
            {
              "id": 3027406,
              "postDate": "2024-10-24T19:01:35.293Z",
              "content": "<p>Try to not impute anything as any lgbm, xgb, cat - can train with missing values.<br>\nI supose you are using a classification model that is not returning a good score so go for a regressor as it's better suited for this task and make some ensemble of them to generalize and stabilize the predictions</p>",
              "rawMarkdown": "Try to not impute anything as any lgbm, xgb, cat - can train with missing values.\nI supose you are using a classification model that is not returning a good score so go for a regressor as it's better suited for this task and make some ensemble of them to generalize and stabilize the predictions",
              "votes": 1
            },
            {
              "id": 3028121,
              "postDate": "2024-10-25T16:23:44.593Z",
              "content": "<p>Tried regression too, used different sets of columns, parsed a round at the end. CV was less performant.</p>",
              "rawMarkdown": "Tried regression too, used different sets of columns, parsed a round at the end. CV was less performant."
            },
            {
              "id": 3031502,
              "postDate": "2024-10-29T18:49:58.577Z",
              "content": "<p>And if a stranger - a public code will prompate you for new solutions and direct you in a new direction that you have not even guessed about? Suddenly this code will give you more good ideas than this discussion? Have a good day!</p>",
              "rawMarkdown": "And if a stranger - a public code will prompate you for new solutions and direct you in a new direction that you have not even guessed about? Suddenly this code will give you more good ideas than this discussion? Have a good day!"
            },
            {
              "id": 3032853,
              "postDate": "2024-10-31T13:27:44.447Z",
              "content": "<p>Of course, some competitors might post new ideas that can be useful. There are always more ideas, I do not own all of them :-).</p>",
              "rawMarkdown": "Of course, some competitors might post new ideas that can be useful. There are always more ideas, I do not own all of them :-)."
            }
          ]
        }
      ]
    },
    {
      "id": 3069391,
      "postDate": "2024-12-11T12:46:16.527Z",
      "content": "<p>Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible</p>",
      "rawMarkdown": "Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible"
    },
    {
      "id": 3029344,
      "postDate": "2024-10-27T06:44:37.217Z",
      "content": "<p>Your approach demonstrates a well-thought-out, structured attempt to address common machine-learning challenges, especially with missing labels and data augmentation from external sources. I have done work with most of the techniques and I recent learn the optuna and the meaning of how to use parquet file</p>",
      "rawMarkdown": "Your approach demonstrates a well-thought-out, structured attempt to address common machine-learning challenges, especially with missing labels and data augmentation from external sources. I have done work with most of the techniques and I recent learn the optuna and the meaning of how to use parquet file"
    },
    {
      "id": 3031137,
      "postDate": "2024-10-29T11:58:56.240Z",
      "rawMarkdown": "",
      "isDeleted": true,
      "replies": [
        {
          "id": 3032854,
          "postDate": "2024-10-31T13:28:36.907Z",
          "content": "<p>I tried gradient boosting first of all. I tried NN, did not get a better result than GB yet.</p>",
          "rawMarkdown": "I tried gradient boosting first of all. I tried NN, did not get a better result than GB yet."
        }
      ]
    }
  ],
  "comments": [
    {
      "id": 3027396,
      "author_name": "Tom Yuen",
      "author_url": "",
      "post_date": "2024-10-24T18:45:37.173000",
      "content": "<p>I've tried most of your methods.<br>\nMy experience is that parquet data is giving some useful data so we can't drop it. However, most of them are rubbish that you know. (e.g. the 1st one is count, anything that can give to model is noise. May be X, Y, Z etc will provide some useful insight to model, but I won't impute the parquet data)</p>\n<p>I've tried adding features/ dropping features<br>\nI've tried oversampling, undersampling<br>\nI've tried custom scoring!<br>\nI've tried SHAP analysis!<br>\nI've tried KNN imputation!<br>\nI've tried binary classification!<br>\nI've tried pseudo labelling!</p>\n<p>The last thing I haven't tried is copy the top score notebook and make a submission.<br>\nThere's a fatal mistake in some of the top score notebooks, I guess, that I can't say it here, that keeps me from directly copying from them.</p>\n<p>My suggestion is to add a confusion matrix and this gives you a better sense of your model, even though it's not how it's scored.</p>\n<p>Try dropping rows without Sii + LGBM + kfold + Optuna 30 trials + optimizer + a bit of data cleaning (e.g. cap max for outlier), it shall take you to 0.43-0.45.</p>\n<p>Good luck!</p>",
      "votes": 5,
      "replies": [
        {
          "id": 3027412,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-10-24T19:11:45.227000",
          "content": "<p>What surprises me about the top scoring (shared) notebooks is that people are just copying them and don't analyze the code, because the same mistake is repeating in all of them, for example how do they use the autoencoder is just not correct and maybe by some luck it gives a good score in LB (at least in the public LB) but they don't see that the original author made a mistake in the method how it's applied and everybody just blindly use it.</p>",
          "votes": 8,
          "replies": [
            {
              "id": 3027443,
              "author_name": "Tom Yuen",
              "author_url": "",
              "post_date": "2024-10-24T20:11:41.773000",
              "content": "<p>Let me name one as well here then 🤣. Using KNN imputer to fill sii in train set is not convincing me yet.</p>",
              "votes": 3,
              "replies": []
            },
            {
              "id": 3027702,
              "author_name": "Danu A.",
              "author_url": "",
              "post_date": "2024-10-25T06:58:41.143000",
              "content": "<p>Absolutely agree and I have an open discussion about this <a href=\"https://www.kaggle.com/competitions/child-mind-institute-problematic-internet-use/discussion/539536\" target=\"_blank\">here </a></p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3031086,
              "author_name": "Catadanna",
              "author_url": "",
              "post_date": "2024-10-29T10:10:36.900000",
              "content": "<p>If people compete with some other's notebook and, say, they win, then they must reproduce the same high score, and that might not work -- especially if they copied without understanding. </p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3031576,
              "author_name": "Lennart Haupts",
              "author_url": "",
              "post_date": "2024-10-29T20:57:07.883000",
              "content": "<p>Oh yes, I noticed that aswell and I'm confused about how persistent this incorrect use of the autoencoder is.😬 </p>",
              "votes": 1,
              "replies": []
            }
          ]
        },
        {
          "id": 3027445,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2024-10-24T20:13:59.587000",
          "content": "<p>I made the confusion matrix, there are some features which are highly correlated.<br>\nAs I already mentioned, I already dropped lines without label.<br>\nI am against copying high score notebooks and compete ith them.</p>",
          "votes": 1,
          "replies": []
        },
        {
          "id": 3029295,
          "author_name": "WangJiazhen",
          "author_url": "",
          "post_date": "2024-10-27T05:46:48.590000",
          "content": "<p>Can not agree more.</p>\n<p>I tried many methods and derived more than 30000 features (without using AutoEncoder) and got 0.45-0.46.<br>\nI also tried to copy AutoEncoder function to used with my owner feature engineering and not get better LB score.</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3031131,
          "author_name": "Tom Yuen",
          "author_url": "",
          "post_date": "2024-10-29T11:55:05.460000",
          "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a> <a href=\"https://www.kaggle.com/ggggpeushmy\" target=\"_blank\">@ggggpeushmy</a> I'm so surprised there are still quite a few of us here in the lower rank.</p>\n<ol>\n<li>Feature engineering can be useful, in some scenarios I found removing useless feature seems to boost a little.</li>\n<li>The validation score is of course important, but I guess it's not as important as other games. <br>\nMy guess is that there're cases with high hours + high SDS + low physical score but sii =0, and <br>\nthere're cases sii=3 but 0 internet hours, normal BMI, low SDS.<br>\nIf we further push the cv score/ training score, it may overfit to these cases.<br>\nIt all depends if the final test set have these silly samples or not.</li>\n<li>Combining different models will definitely help, but I haven't done it yet.<br>\nGood luck everyone.</li>\n</ol>",
          "votes": 1,
          "replies": [
            {
              "id": 3041769,
              "author_name": "Catadanna",
              "author_url": "",
              "post_date": "2024-11-10T20:22:49.253000",
              "content": "<p><a href=\"https://www.kaggle.com/tomyuen\" target=\"_blank\">@tomyuen</a> : </p>\n<p>Low score does not mean anything for the moment, as these scores are parsed only on a small part  of the data.<br>\nValidation score is important, but one should clean the data correctly.<br>\nCombining predictions is good practice, because if one model gives a bad prediction for a sample, the others will give a good one, so the result will be smoothed and good in the end.</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3076800,
      "author_name": "Tom Yuen",
      "author_url": "",
      "post_date": "2024-12-20T09:22:38.937000",
      "content": "<p><a href=\"https://www.kaggle.com/catadanna\" target=\"_blank\">@catadanna</a> <a href=\"https://www.kaggle.com/eu1234\" target=\"_blank\">@eu1234</a><br>\nI just come back and congratulate you both. Performance's not bad, right?! At least not falling under cliff after the shakeup.</p>",
      "votes": 1,
      "replies": [
        {
          "id": 3076817,
          "author_name": "Danu A.",
          "author_url": "",
          "post_date": "2024-12-20T09:45:20.047000",
          "content": "<p>Thanks! I had a small boost of +1649 positions up, but still missed to select my best scoring notebook that could bring me to a silver medal position 61. A good experience anyway. Best of luck in the future challenges!</p>",
          "votes": 0,
          "replies": []
        },
        {
          "id": 3078236,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2024-12-22T01:34:05.757000",
          "content": "<p>Thank you! It seems I did not overfit, or at least  I overfit less than others, as I climber about 1800 places up. I trusted my CV scores when choosing the submissions to compete with.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    },
    {
      "id": 3026738,
      "author_name": "Danu A.",
      "author_url": "",
      "post_date": "2024-10-24T07:01:14.410000",
      "content": "<p>In 2.2:</p>\n<ul>\n<li>try more aggregation types. </li>\n<li>you are imputing missing test values with overall train average, but you can impute the average from that person's values directly or even simpler - don't impute anything, even in the main data because the boosters are dealing just fine with missing data, or at least give it a try for testing.</li>\n</ul>\n<p>And the final - LB score is very sensitive to the random seed, as you can see the LB is overflooded with very close scores because maybe half of the people are using almost the same code that is publicly available and just re-runs it with different seeds or average the results of the same model but with different seeds.</p>\n<p>My advice - try to find your way to get a stable CV-LB score pairing, when them both grow hand by hand and to not have large oscilations between them.</p>",
      "votes": 2,
      "replies": [
        {
          "id": 3027247,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2024-10-24T15:56:44.290000",
          "content": "<p>Thank you for your answer. I did not get a very good LB score for the moment. I shall try more aggregations (sum or other quantilles). And no, I do not wish to use someone else's code and compete with it, I do not think it is fair and does not give me any credit.</p>\n<p>Concerning the missing values : I assumed that some individuals in the test set do not have corresponding parquet information at all. That is why I replace with mean from test for each feature. I have no choice, test set is hidden. I may try to replace nan with one value, let us say -1 .</p>",
          "votes": 1,
          "replies": [
            {
              "id": 3027406,
              "author_name": "Danu A.",
              "author_url": "",
              "post_date": "2024-10-24T19:01:35.293000",
              "content": "<p>Try to not impute anything as any lgbm, xgb, cat - can train with missing values.<br>\nI supose you are using a classification model that is not returning a good score so go for a regressor as it's better suited for this task and make some ensemble of them to generalize and stabilize the predictions</p>",
              "votes": 1,
              "replies": []
            },
            {
              "id": 3028121,
              "author_name": "Catadanna",
              "author_url": "",
              "post_date": "2024-10-25T16:23:44.593000",
              "content": "<p>Tried regression too, used different sets of columns, parsed a round at the end. CV was less performant.</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3031502,
              "author_name": "Zaakcii Ru",
              "author_url": "",
              "post_date": "2024-10-29T18:49:58.577000",
              "content": "<p>And if a stranger - a public code will prompate you for new solutions and direct you in a new direction that you have not even guessed about? Suddenly this code will give you more good ideas than this discussion? Have a good day!</p>",
              "votes": 0,
              "replies": []
            },
            {
              "id": 3032853,
              "author_name": "Catadanna",
              "author_url": "",
              "post_date": "2024-10-31T13:27:44.447000",
              "content": "<p>Of course, some competitors might post new ideas that can be useful. There are always more ideas, I do not own all of them :-).</p>",
              "votes": 0,
              "replies": []
            }
          ]
        }
      ]
    },
    {
      "id": 3069391,
      "author_name": "MZhu",
      "author_url": "",
      "post_date": "2024-12-11T12:46:16.527000",
      "content": "<p>Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3029344,
      "author_name": "Sumit_08",
      "author_url": "",
      "post_date": "2024-10-27T06:44:37.217000",
      "content": "<p>Your approach demonstrates a well-thought-out, structured attempt to address common machine-learning challenges, especially with missing labels and data augmentation from external sources. I have done work with most of the techniques and I recent learn the optuna and the meaning of how to use parquet file</p>",
      "votes": 0,
      "replies": []
    },
    {
      "id": 3031137,
      "author_name": "",
      "author_url": "",
      "post_date": "2024-10-29T11:58:56.240000",
      "content": "",
      "votes": 0,
      "replies": [
        {
          "id": 3032854,
          "author_name": "Catadanna",
          "author_url": "",
          "post_date": "2024-10-31T13:28:36.907000",
          "content": "<p>I tried gradient boosting first of all. I tried NN, did not get a better result than GB yet.</p>",
          "votes": 0,
          "replies": []
        }
      ]
    }
  ],
  "raw_markdown_by_id": {
    "3026385": "Hallo\n\nHere are various strategies I tried.\n\n1. I used different models only with the train/test data. \n1.1 There are missing values in the labels so I tried to train only with the non null labels first. \n1.2 I also tried to predict the values for the missing labels, using various classifiers, but that did not provide a better result.\n\n2. I tried using both train/test data merged with the parquet data. \nNow, the parquet data contains several lines for one sample in train/test. \n2.1 Sometimes, the data for one sample in train/test is missing in the parquet data. So one could choose either to delete these samples for training or to make predictions for it. I tried both. \n2.2 I tried aggregations by sample for the information in the parquet data, only for numeric features. For each of them I created new aggregated features for mean, median, q1 and q3. These are the new features I added to the existing train features. For the test features, the test set is hidden; in the case of missing features in parquet files for a given sample, I replaced with the mean parsed for the respective feature in the train set.\n\nThe CV results for the dataset using train/test and parquet files did not give a better result. Best result was given by the basic train set with only the non null labels. I used different algorithms: neural networks and gradient boosting.\n\n3. I divided the data in 5 splits, recorded the spits and worked only on them, in order to train/test on the same dataset and be able to compare correctly CV results obtained with different algorithms.\n\n4. I used Optuna in order to get the best parameters.\n\n\n\nFeel free to comment, any new idea or remark is welcome!  ",
    "3027396": "I've tried most of your methods.\nMy experience is that parquet data is giving some useful data so we can't drop it. However, most of them are rubbish that you know. (e.g. the 1st one is count, anything that can give to model is noise. May be X, Y, Z etc will provide some useful insight to model, but I won't impute the parquet data)\n\nI've tried adding features/ dropping features\nI've tried oversampling, undersampling\nI've tried custom scoring!\nI've tried SHAP analysis!\nI've tried KNN imputation!\nI've tried binary classification!\nI've tried pseudo labelling!\n\nThe last thing I haven't tried is copy the top score notebook and make a submission.\nThere's a fatal mistake in some of the top score notebooks, I guess, that I can't say it here, that keeps me from directly copying from them.\n\nMy suggestion is to add a confusion matrix and this gives you a better sense of your model, even though it's not how it's scored.\n\nTry dropping rows without Sii + LGBM + kfold + Optuna 30 trials + optimizer + a bit of data cleaning (e.g. cap max for outlier), it shall take you to 0.43-0.45.\n\nGood luck!",
    "3076800": "@catadanna @eu1234\nI just come back and congratulate you both. Performance's not bad, right?! At least not falling under cliff after the shakeup.",
    "3026738": "In 2.2:\n- try more aggregation types. \n- you are imputing missing test values with overall train average, but you can impute the average from that person's values directly or even simpler - don't impute anything, even in the main data because the boosters are dealing just fine with missing data, or at least give it a try for testing.\n\nAnd the final - LB score is very sensitive to the random seed, as you can see the LB is overflooded with very close scores because maybe half of the people are using almost the same code that is publicly available and just re-runs it with different seeds or average the results of the same model but with different seeds.\n\nMy advice - try to find your way to get a stable CV-LB score pairing, when them both grow hand by hand and to not have large oscilations between them.",
    "3069391": "Thank you for your insights. I think I will try more evaluation indicators to comprehensively evaluate my model. If I find any, I will share it as soon as possible",
    "3029344": "Your approach demonstrates a well-thought-out, structured attempt to address common machine-learning challenges, especially with missing labels and data augmentation from external sources. I have done work with most of the techniques and I recent learn the optuna and the meaning of how to use parquet file",
    "3031137": ""
  }
}